Audio signal encoding system
Abstract
Apparatus and methods for including a code (68) having at least one code frequency component in an audio signal (60) are provided. The abilities of various frequency components in the audio signal to mask the code frequency component to human hearing are evaluated (64), and based on these evaluations an amplitude (76) is assigned to the code frequency component. Methods and apparatus for detecting a code in an encoded audio signal are also provided. A code frequency component in the encoded audio signal is detected based on an expected code amplitude or on a noise amplitude within a range of audio frequencies including the frequency of the code component.

Term
Term ended
Expired 27 March 2015, 11.5 years ago.
- Priority
- Filed
- Granted
- Expired
- Today
7 claims: 1 independent, 6 dependent
- 1Zastrzeżenia patentowe 1. System kodowania sygnału dźwiękowego, zaopatrzony w końcówkę wejściową sygnału dźwiękowego oraz zespół do włączania składowych częstotliwości kodu do sygnału dźwiękowego, znamienny tym, że zawiera zespół szacowania zdolności maskowania składowych częstotliwości cyfrowego sygnału dźwiękowego ze względu na słyszalność ludzkiego ucha (34) połączony z końcówką wejściową (30) i zespół przypisywania amplitud indywidualnych składowych częstotliwości kodu wynikających z szacowania zdolności maskowania (40), połączony z zespołem szacowania zdolności maskowania składowych częstotliwości cyfrowego sygnału dźwiękowego ze względu na słyszalność ludzkiego ucha (34).
- 2System kodowania sygnału dźwiękowego według zastrz. 1, znamienny tym, że zespół szacowania zdolności maskowania składowych częstotliwości cyfrowego sygnału dźwiękowego ze względu na słyszalność ludzkiego ucha (34) jest zaopatrzony w miernik poziomu sygnału składowych częstotliwości cyfrowego sygnału dźwiękowego.
- 3System kodowania sygnału dźwiękowego według zastrz. 1, znamienny tym, że jest zaopatrzony w zespół do dekodowania zakodowanego sygnału dźwiękowego dla wykrywania składowej częstotliwości kodu.
- 4System kodowania sygnału dźwiękowego według zastrz. 1, znamienny tym, że zespół przypisywania amplitud indywidualnych składowych częstotliwości kodu wynikających z szacowania zdolności maskowania (40) jest zaopatrzony w wejście danych (42) oraz generator składowych częstotliwości kodu w odpowiedzi na dane źródła i/lub identyfikacji.
- 5System kodowania sygnału dźwiękowego według zastrz. 1, znamienny tym, że zawiera sumator sumujący sygnał dźwiękowy i składowe częstotliwości kodu.
- 6System kodowania sygnału dźwiękowego według zastrz. 1, znamienny tym, że zawiera procesor cyfrowy sumujący sygnał dźwiękowy i składowe częstotliwości kodu.
- 7System kodowania sygnału dźwiękowego według zastrz. 1, znamienny tym, że zespół szacowania zdolności maskowania składowych częstotliwości cyfrowego sygnału dźwiękowego ze względu na słyszalność ludzkiego ucha (34) jest zaopatrzony w zespół wykrywania mocy sygnału składowych częstotliwości cyfrowego sygnału dźwiękowego, urządzenie do wyznaczania współczynników maskowania na podstawie częstotliwości i mocy sygnału cyfrowego, połączone z zespołem wykrywania mocy sygnału składowych częstotliwości cyfrowego sygnału dźwiękowego, selektor współczynników maskowania na podstawie zdolności maskowania, połączony z urządzeniem do wyznaczania współczynników maskowania na podstawie częstotliwości i mocy sygnału cyfrowego, przy czym zespół przypisywania amplitud indywidualnych składowych częstotliwości kodu wynikających z szacowania zdolności maskowania (40) zawiera zespół przypisywania amplitud indywidualnych składowych częstotliwości kodu na podstawie współczynników maskowania.
Independent claims7
292 paragraphs in 17 sections, as filed
The present invention relates to an audio coding system.
Over the years, many ways have been proposed to mix codes with audio signals in such a way that codes can be reliably reproduced from audio signals, while they remain inaudible when converting audio signals into sound. Achieving both of these goals is crucial for practical applications. For example, broadcasters and broadcast producers, as well as those recording music from public distribution, would not tolerate the addition of audible codes to programs or recordings.
183 307
Methods for encoding audio signals have been repeatedly proposed, starting at least from US Patent No. 3,004,104, to Hembroke, of October 10, 1961. Hembroke showed a coding method in which the energy of the audio signal in a narrow frequency band was selectively removed to encode the signal. The problem with this method arises when the noise or signal interference reintroduces energy into the narrow frequency band, thereby overriding the code.
According to another method, described in US Patent No. 3,845,391 to Crosby, it was proposed to eliminate the narrow frequency band from the audio signal and place the code there. This method has the same disadvantage as the previous solution, as indicated in US Patent No. 4,703,476 to Howard, which is associated with Crosby's patent. However, Howard's patent was only intended to improve Crosby's solution without going beyond its fundamental assumptions.
Coding of binary signals at frequencies passing through the whole sound band has also been proposed. The problem with this solution is that in the absence of sound signal components to mask the frequency of the code, it may become audible. This method assumes that the codes are similar to noise, suggesting that their presence will be ignored by listeners. However, in many cases this assumption is not true, for example in the case of classical music containing portions of relatively low content of a sound signal, or in the event of a speech break.
Another method has been proposed, according to which two-tone multi-frequency codes (DTMF) are inserted into the audio signal. The meaning of the DTMF code is detected based on its frequency and duration. However, the components of the audio signals may be confused for one or both tones of each DTMF code, whereby either the presence of the code may be unnoticed by the detector, or the signal components may be misread as elements of the DTMF code. In addition, it should be noted that the DTMF code has a common tone with other DTMF codes. As a result, the signal component corresponding to the tone of another DTMF code may be mixed with the tone of the DTMF code, which is also present in the signal, which leads to incorrect detection.
The audio coding system, provided with an audio input input terminal and an assembly for incorporating code frequency components into the audio signal, according to the invention is distinguished by that it includes a team for estimating the ability to mask the frequency components of a digital audio signal due to audibility of the human ear connected to the input terminal, and a team for assigning amplitudes of individual code frequency components resulting from the estimation of masking capacity, combined with a team for estimating the ability to mask frequency components of the digital audio signal for audibility human ear.
The team for estimating the ability to mask the frequency components of a digital audio signal due to audibility of the human ear is preferably provided with a level meter of the signal components of the digital audio signal.
The audio coding system is preferably provided with a means for decoding the coded audio signal for detecting the frequency component of the code.
The unit for assigning the amplitudes of individual code frequency components resulting from the estimation of masking capability is preferably provided with data input and a code frequency component generator in response to source data and / or identification.
The audio coding system preferably includes an adder that sums the audio signal and the frequency components of the code. The audio coding system preferably includes a digital processor that sums the audio signal and the frequency components of the code.
The team for estimating the ability to mask the frequency components of a digital audio signal due to audibility of the human ear is preferably provided with a component for detecting the signal strength of the frequency components of the digital audio signal, a device for determining masking factors based on the frequency and power of the digital signal, combined with a component detecting the signal strength
183 307 digital audio signal frequency, masking factor selector based on masking ability, connected to a device for determining masking factors based on frequency and power of a digital signal, where the set of assigning amplitudes of individual frequency components of the code resulting from the estimation of masking capacity includes the set of assigning amplitudes of individual frequency components code based on masking coefficients.
The solution according to the invention makes it possible to overcome the disadvantages of the hitherto proposed solutions, because it enables the incorporation of codes into audio signals in such a way that, as sounds, the codes are inaudible to the human ear, but can be reliably detected by the decoding device and reliably read from the audio signals.
The object of the invention, in an embodiment, is explained in the drawing, in which Fig. 1 is a block diagram of the encoder operation, Fig. 2 is a block diagram of the operation of the digital encoder, Fig. 3 is a block diagram of the coding system used to encode audio signals provided in analog form , Fig. 4 - spectral diagrams used to illustrate frequency components corresponding to different code symbols, after coding in accordance with the embodiment of Fig. 3, Fig. 5 and 6 - block diagrams used to illustrate the operation of the embodiment of Fig. 3, Figures 7A to 7C - the flowchart executed by the program procedure, used in the embodiment of Fig. 3, Figures 7D to 7E - flowchart for the illustration an alternative program procedure used in the embodiment of Fig. 3, Fig. 7F - graph showing a linear approximation of the single-tone masking relationship, Fig. 8 - block diagram of an encoder using the analog circuit, Fig. 9 - block diagram of the weighting factor determination system for the embodiment of Fig. 8, Fig. 10 - block diagram of the decoder according to some features of the present invention, Fig. 11 - block diagram of the decoder according to the embodiment of the present invention using digital signal processing, Fig. 12A and 12B - a flowchart describing the operation of the decoder of Fig. 11, Fig. 13 - a block diagram of a decoder according to some embodiments of the present invention, Fig. 14 - block diagram of an embodiment of an analog decoder, Fig. 15 - block diagram of a component detector according to the embodiment of Fig. 14, Figs. 15, 16 and 17 - block diagrams of the device incorporated in the system for producing audibility estimates of widely distributed information.
The present invention implements methods for incorporating code into audio signals to optimize the likelihood of accurate reproduction of information in code in a signal, ensuring that when the audio signal is reproduced in the form of sound, the code is not audible, even if it falls in an audible frequency range.
Referring to Fig. 1, a block diagram of the encoder operation is shown. The audio signal to be encoded is taken at the input terminal 30. The audio signal may represent, for example, a radio program, the television audio part, music or any other type of audio signals reproduced in a similar manner. In addition, the audio signal can be used for private communication, such as telephone transmission or for some type of personal recording. However, these examples of use are listed here for illustrative purposes only and do not limit the field of application of the invention.
In the function block designated 34 in Fig. 1, the ability of one or more components of the downloaded audio signal to mask sounds having frequencies corresponding to the code component or frequency components to be added to the audio signal is calculated. Multiple calculations for a single code frequency may be performed, separate calculations for each of the multiple code frequencies may be performed, multiple calculations for each of the multiple code frequencies may be performed, one or more joint calculations for multiple code frequencies may be performed, or it may be a combination of one or more of the above operations used. Each calculation is performed based on the frequency of one or more code components to be masked, and the frequency (one or more) of the component audio signal whose masking ability is currently being calculated. In addition, if the code component and the component or components of the sound masking do not fall in substantially the same signal intervals, so that they could be reproduced as sound at significantly different time intervals, the effects of differences in the signal intervals between the component or code components that are masked , and the masking component or components of the program are also considered.
Advantageously, in some embodiments, multiple calculations are made for each code component by separately considering the ability of different parts of the audio signal to mask each code component. In one embodiment, the ability of each of a plurality of substantially single tonal components of an audio signal to mask a code component is calculated based on the frequency of the audio signal component, its amplitude (as it will be defined) and the time dependence of the component code, such masking is referred to herein as tonal masking .
The term amplitude is used herein to describe any one or more signal value that can be used to estimate the masking ability so as to select the size of the code component, to detect its presence in the reproduced signal, and for any other purposes, what can be used here for energy, power, voltage, current and pressure, whether measured in absolute or relative terms, and whether instantaneous or accumulated values are considered. Accordingly, the amplitude can be measured as window mean, arithmetic mean, by integrating the square root of the value, by the accumulation of absolute or relative discrete values, or otherwise.
In other embodiments, in addition to, or alternatively to, tonal masking estimates, the ability of audio components from a relatively narrow frequency band close enough to a given code component of masking (which is herein referred to as narrow band masking) is calculated. In yet other embodiments, the ability of multiple code components in a relatively wide frequency band to mask component is calculated. Depending on the necessity or possibility, the abilities of the components of the sound program are calculated in the signal intervals preceding or following a given, one or more components to mask it in a non-simultaneous manner. This method of estimation is particularly useful when the components of the audio signal in a given signal interval have insufficient amplitudes to enable the inclusion of code components with sufficiently large amplitudes in the same interval, making them distinguishable from noise.
Preferably, the combination of two or more tonal masking capabilities, narrowband masking and broadband masking capabilities (and, if necessary or appropriate, non-simultaneous masking capabilities) are calculated for multiple code components. When the code components are close enough in the frequency domain, separate calculations need not be performed for each component.
In some other preferred embodiments, the sliding tonal analysis is performed instead of separate tonal, narrow or wideband analyzes, without having to classify the sound program as tonal, narrowband or broadband.
Advantageously, when a combination of masking capability is calculated, each calculation provides the maximum allowable amplitude for one or more code components, so that by comparing all calculations that have been performed and that relate to a given component, the maximum amplitude is selected, ensuring that each the component will be masked by the sound signal when it is played as sound, so that it will not be heard by the human ear. By maximizing the amplitude of each component, the probability of detecting its presence based on its amplitude is also increased. Of course, it is not necessary to use the greatest possible amplitude, which is only necessary for decoding, in order to be able to distinguish a sufficiently large number of code components from components of an audio signal or other noise.
The calculation effects are discharged, as indicated by 36 in Fig. 1, into the code generator 40. The code generation can be performed in many different ways. One particularly preferred method is to assign a single set of frequency components
183 307 code to each of the numerous data states or symbols, so that during a given signal interval, the respective data state is represented by the presence of its respective set of code frequency components. In this way, the overlap of the detected code with the audio components is reduced because, in a preferably high percentage of signal intervals, a sufficiently large number of code components will be detectable despite the overlapping of the program's audio components over the other components. In addition, the process of implementing masking estimates is simplified when the frequencies of code components are known prior to their generation.
Other forms of coding can also be implemented. For example, frequency shift keying (FSK), frequency modulation (FM), hopping coding, diffuse spectral coding, or a combination of these methods can be used. Other coding methods that may be used to implement the present invention will be apparent from its description.
The data to be encoded is taken at the input 42 of the code generator 40, which responds by creating a unique group of code frequency components and assigning the amplitude of each of them based on calculations taken from output 36. The code frequency components thus generated are delivered to the first input of the summing system 46, which receives a sound signal to be encoded at its second input. The circuit 46 adds the frequency components of the code to the audio signal and discharges the encoded audio signal at its output terminal 50. The circuit 46 can be an analog or digital summation system, depending on the form of the signals fed to it. Adding can also be implemented in software, and in this case, a digital processor for performing masking estimates and code generation can also be used to add code with an audio signal. In one embodiment, the code is provided as time domain data in digital form, which is then aggregated in the time domain of audio data. In another example, the audio signal is converted to the frequency domain in digital form and added to the code in a similar way represented as digital frequency domain data. In most applications, the aggregated frequency domain data is then converted to time domain data.
It follows that masking estimation, as well as code generation functions, can be performed by digital or analog processing, or by a combination thereof. In addition, although the audio signal can be received in analog form at the output terminal 30 and added to the code components in analog form by the circuit 46, as shown in Fig. 1, then, alternatively, the audio signal may be converted into digital form upon its receipt, added to the code components in digital form and output in digital or analog form. For example, when a signal is to be recorded on a compact disc or digital audio tape, it may be output in digital form, whereas if it is to be distributed by traditional radio or television, it may have an analog output form. Various other combinations of analog and digital processing can be implemented.
In some embodiments, the code components of only one code symbol at a time are included in the audio signal. However, in other embodiments, the components of the plurality of code symbols are simultaneously incorporated into the audio signal. For example, in some embodiments, components of one symbol occupy one frequency band and components of another symbol occupy another frequency band at the same time. Alternatively, the components of one symbol may be in the same band as another, or their bands may overlap as long as these components are distinguishable, for example, by assigning substantially different frequencies or frequency ranges.
An embodiment of the digital decoder is shown in Fig. 2. In this embodiment, the audio signal in analog form is taken at the input terminal 60 and converted to digital form by the A / D converter 62. The digital sound signal is brought to the masking estimation, as indicated as function block 64,
183 307 before which the digital audio signal is divided into frequency components, for example by fast Fourier transformation (FFT), wavelet transformation or other time-to-frequency transformations, or by digital filtering. Then, the masking capacities of the frequency components of the sound signal in the respective frequency packet are determined, determining the tonal masking ability, narrowband masking ability and broadband masking ability (and, if necessary and appropriate, non-simultaneous masking ability). Alternatively, the masking capabilities of the frequency components of the audio signal in a given frequency packet are calculated with sliding tonal analysis.
The encoded data is taken from the input terminal 68 and, for each data state corresponding to a given signal interval, a corresponding group of code components is generated, which is marked with a functional block of signal generation 72, and is subjected to a level determination, which is marked with a functional block 76, which also downloads appropriate masking estimates. Signal generation can be implemented, for example, by means of a dictionary table storing each of the code components as time domain data, or by interpolation of stored data. The code components can either be permanently stored or generated after the system initialization of Fig. 2 and then stored in memory, such as RAM, in order to be discharged in response to data retrieved from terminal 68. The component values can also be calculated at the time they are generated.
The level determination is performed for each of the code components based on the respective masking estimates described above, and the code components whose amplitude has been determined to ensure inaudibility are added to the digital audio signal, which is indicated by the sum symbol 80. Depending on the amount of time needed to perform the above operations, it may be beneficial to delay the digital audio signal, as indicated by 82, by temporarily storing in memory. If the audible signal is not delayed, after performing the FFT and estimating the masking for the first audible signal interval, the code components with a fixed amplitude are added to the second audible signal interval following the first interval. If the audio signal is delayed, code components with a fixed amplitude can be added to the first interval and at the same time masking estimation can be used. In addition, if a portion of the audio signal during the first compartment provides greater capability to mask the code component added during the second compartment than a portion of the audio signal during the second compartment for the same code component, then the amplitude can be assigned to the code component based on the ability to simultaneously mask the portion of the audio signal in the first compartment. In this way, simultaneous and non-simultaneous masking capabilities can be estimated, and the optimal amplitude can be assigned to each code component based on the most-favorable estimation.
In some applications, such as radio broadcasts or analog recording (such as, for example, on a classic tape cassette), the encoded audio signal in digital form is converted to analog form by a digital-to-analogue converter (DAC) 84. However, when the signal is to be transmitted or saved in digital form, the DAC 84 can be omitted.
The various functions shown in Fig. 2 can be implemented, for example, by a digital signal processor or personal computer, workstation or large computer system, or other digital computer.
Figure 3 is a block diagram of a coding system for coding audio signals provided in analog form, such as classic radio broadcasts. In the system of Fig. 3, the main processor 90, which may be, for example, a personal computer, manages the selection and generation of information to be encoded and incorporated into an analog audio signal taken from the input terminal 94. The main processor 90 is coupled to the keyboard 96 and to the monitor 100, such as the CRT monitor, so that the user can select the desired information to be encoded by selecting the available information from the menu,
183 307 displayed on monitor 100. Typical encoding information in a radio broadcast signal may include channel or station identification information, a program or information segment, and / or a time code.
When the requested information has been entered into the main processor 90, the processor discharges data representing information symbols to the digital signal processor (DSP) 104, which in turn encodes each symbol obtained from the main processor 90 in the form of a unique set of code signal components, as described below. According to one embodiment, the main processor generates a four-state data stream, i.e., a data stream in which each data unit can assume one of four different states, each of which represents a unique symbol, including synchronization symbols called E and S, and two information symbols 1 and 0 representing the respective binary state. Of course, any number of distinguishable data states can be used. For example, instead of two information symbols, three-state data can be represented by three different symbols, which allows correspondingly more information to be transferred in a data stream of a given size.
For example, when the signal is about speech, it is preferable to send the symbol for a relatively long period of time compared to a broadcast having a substantially more continuous energy content to allow for natural breaks in speech. Accordingly, in order to provide sufficiently high information throughput in this case, the number of possible information symbols can be advantageously increased. For symbols representing up to five bits, signal transmission lengths of 2, 3 and 4 seconds provide increasing probability of correct decoding. In some such embodiments, the initial symbol (E) is decoded when the energy of the FFT packet for this symbol is greatest, when the average energy minus the standard energy deviation for this symbol is greater than the average energy plus the average standard energy deviation for all other symbols, and when the shape of the energy versus time graph is generally bell-shaped, with a peak in the temporal inter-symbol border.
In the embodiment shown in Fig. 3, when DSP 104 receives the symbols of a given information to be encoded, it responds by generating a unique set of code frequency components for each symbol, which set it outputs to output 106. Also referring to Fig. 4, shown there are spectral plots for each of the four data symbols S, E, 0 and 1 of the exemplary data set described. As shown in fig. 4, in this embodiment the symbol S is represented by a unique group of ten code frequency components from f1 to f10 spaced at equal intervals in the frequency range from the frequency value slightly greater than 2 kHz to the frequency value slightly less than 3 kHz. The symbol E is represented by the second unique group of ten code frequency components fll to f20, spaced in the frequency spectrum at equal intervals from the first frequency value slightly greater than 2 kHz to the frequency value slightly less than 3 kHz, where each of the code components from f11 to f20 has a unique frequency value different from all values in the same group as well as from all frequencies from f1 to f10. The symbol 0 is represented by another unique group of ten code frequency components from f21 to f30, also distributed in the frequency spectrum at equal intervals from the first frequency value slightly greater than 2 kHz to the frequency value slightly less than 3 kHz, where each of the code components has a unique a frequency value different from all values in the same group as well as from all frequencies from f1 to f20. Finally, symbol 1 is represented by another unique group of ten code frequency components from f31 to f40, spaced in the frequency spectrum at equal intervals from the first frequency value slightly greater than 2 kHz to the frequency value slightly less than 3 kHz, where each of the code components from f31 to f40 has a unique frequency value different from all frequencies from f1 to f40. By using multiple code frequency components for each of the data states, so that the code components of each nu are 303 substantially different in frequency, the presence of noise (such as non-coded components of an audio signal or other noise) in a common detection band with any given noise component data status is less likely and overlapping is less likely.
In other embodiments, it is preferable to represent symbols by multiple frequency components, e.g., ten tones or code frequency components that are not uniformly spaced in the frequency domain and that do not have the same offset between the symbols. Avoiding the integral relationship between code frequencies for a symbol by grouping tones reduces the effects of inter-frequency rumbling or empty rooms, i.e. places where echo reflected from the walls affects correct decoding. The following sets of code tone frequency components for four symbols are provided for suppressing the effects of empty rooms, where f / to f / 0 represent the corresponding frequency components (expressed in Hertz) of the code for each of the four symbols:
<td></td><td> 0”</td><td>/ "Ł?<sup>1</sup></td><td>"S"</td><td>"E"</td>
<td>f /</td><td> /046.9</td><td> /054.7</td><td> /062.5</td><td> /070.3</td>
<td>f2</td><td> //95.3</td><td> /203./</td><td> //79.7</td><td> //87.5</td>
<td>β</td><td> /35/.6</td><td> /343.8</td><td> /335.9</td><td> /328./</td>
<td>f4</td><td> /492.2</td><td> /484.4</td><td> /507.8</td><td> /500.0</td>
<td>f5</td><td> 1656.3</td><td> /664/</td><td> /67/.9</td><td> /679.7</td>
<td>f6</td><td> 1859.4</td><td> /867.2</td><td> /843.8</td><td> /85/.6</td>
<td>f7</td><td> 2078./</td><td> 2070.3</td><td> 2062.5</td><td> 2054.7</td>
<td>f8</td><td> 2296.9</td><td> 2289/</td><td> 2304.7</td><td> 23/2.5</td>
<td>f9</td><td> 2546.9</td><td> 2554.7</td><td> 2562.5</td><td> 2570.3</td>
<td>f / 0</td><td> 2859.4</td><td> 2867.2</td><td> 2843.8</td><td> 285/.6</td>
Generally speaking, in the examples shown above, the spectral content of the code changes relatively little when DSP / 04 switches its output from any of the states S, E, 0 and / to any other of them. In accordance with one aspect of the present invention, in some preferred embodiments, each frequency component of the code of each symbol is paired with the frequency component of each of the other data states, such that the difference between them is less than their critical bandwidth. For any pair of pure tones, the critical bandwidth is the frequency range in which the frequency interval between two tones can change without significantly increasing the volume. When the distance between adjacent tones for each of the data states S, E, 0 and /, and when each tone of each of the data states is paired with the corresponding tone of each of the other of these states, so that the frequency difference between them is less than critical bandwidth for this pair, there will generally be no change in volume when switching from any of the data states S, E, / and 0 to any other data state during playback as a sound. Furthermore, by minimizing the frequency difference between the frequency components of the code of each pair, the respective probabilities of detecting each state of data upon receipt does not essentially depend on the transmission path. Another advantage of pairing the components of different data states is that the masking estimation performed for the code component of the first data state will be substantially accurate for the next data state, if switching occurs.
Alternatively, with a solution providing a heterogeneous distribution of code tones aimed at minimizing the effects of empty rooms, it is visible that the frequencies selected for each of the code frequency components from f / to f / 0 are grouped
183 307 around a certain frequency, e.g. frequency components for f1, f2 and f3 are located adjacent 1055 Hz, 1180 Hz and 1340 Hz, respectively. In this particular embodiment, the tones are spaced twice the distance of the FFT resolution, for example for 4 Hz the tones are shown at 8 Hz intervals, and are selected to be in the middle of the FFT packet's frequency range. Also, the order of the different frequencies that are assigned to the frequency components of the code from f1 to f10 to represent different symbols 0, 1, S, E changes in each group. For example, the frequencies selected for the components f1, f2 and β correspond to the symbols (0, 1, S, E), (S, E, 0, 1) and (E, S, 1, 0), from the smallest to the largest, respectively, that is (1046.9, 1054.7, 1062.5, 1070.3), (1179.7, 1187.5, 1195.3, 1203.1), (1328.1, 1335.9, 1343.8, 1351.6). The advantage of such a solution is that even if there is an empty space that interferes with the appropriate reception of the code component, generally the same tone is eliminated from each of the symbols, making it easier to decode the symbol from the other components. In contrast, if the empty space eliminates a component from one of the symbols, but not the other, it is more difficult to properly decode the symbol.
It can be said that more or less than 4 symbols can be used for coding. In addition, each data state or symbol can be represented by less or more than ten tones of code, and while it is preferable to represent each data state by the same number of tones, it is not necessary that in all applications the number of tones it uses to represent each data state is itself. Advantageously, each of the code tones differs in frequency form from each of the other code tones to increase the likelihood of each data state being distinguished during decoding. However, it is not necessary in all applications that none of the code tone frequencies are common to two or more data states.
Figure 5 is a functional block diagram that is referred to when explaining the coding operation performed by the embodiment of Figure 3. As mentioned above, DSP 104 receives data from the main processor 90 sending a series of data states to DSP 104 as the frequency components of the code. Advantageously, DSP 104 generates a time table of the time domain representation of each of the code frequency components f1 to f40, which are then stored in its RAM, represented as memory 110 in Fig. 5. In response to data received from the main processor 90, DSP 104 generates the corresponding address, which is used as an input address to memory 110, as indicated by 112 in Fig. 5, which causes the memory 110 to output time-domain data for each of the ten frequency components corresponding to the data state to be output at a given time.
Also referring to Fig. 6, which is a functional block diagram illustrating certain operations performed by DSP 104, memory 110 stores time-domain value sequences for each of the frequency components for each of the symbols S, E, 0 and 1. In this particular embodiment, if the code frequency components are in the range from 2 kHz to about 3 kHz, a sufficiently large number of time-domain samples are stored in memory 110 for each of the frequency components from f1 to f40, so that they can be removed from higher than the Nyquist frequency of the highest frequency code component. Time-domain code components are discharged at a correspondingly high frequency from memory 110, which stores time-domain components for each of the code frequency components representing assumed durations, so that time-domain components (n) are stored for each of the code frequency components from f1 to f40 for (n) ranges tl to tn as shown in Fig. 6. For example, if the symbol S is to be encoded during a given signal interval, during the first interval t1, memory 110 discharges time domain components f1 to f10 corresponding to that interval. In the next time interval, components in the time domain from f1 to f10 for interval t2 are removed from memory 110. This process is sequentially continued for intervals from t3 to tn, and back again from tl until the end of the encoded symbol S.
183 307
In some embodiments, instead of draining all ten code components, i.e., from fi to f10, over a period of time, only those components in the critical band of audio signals are discharged. This is essentially a conservative approach to ensuring the inaudibility of code components.
Referring again to Figure 5, DSP 104 also serves to determine the time-domain amplitudes removed from memory 110 in such a way that when the frequency components of the code are reproduced as sound, they will be masked by the components of the audio signal into which they were incorporated, such that remain inaudible to the human ear. In addition, the DSP 104 also receives the audio signal received from input terminal 94 after appropriate filtering and analog-to-digital conversion. More specifically, the encoder in Fig. 3 includes an analog bandpass filter 120 that is used to substantially remove the frequency components of the audio signal from outside of the band involved to calculate the masking ability of the received audio signal, which in a smaller embodiment is in the range from about 1.5 kHz to about 3.2 kHz. The filter 120 is also used to remove high-frequency components from an audio signal that may cause aliasing when the signal is subsequently digitized by an analog-to-digital (A / D) converter 124 operating at a sufficiently high sampling rate.
As shown in Figure 3, the digital audio signal is provided by the A / D converter 124 to DSP 104, where, as indicated by 130 in Figure 5, the program audio signal is subjected to frequency range division. In this particular embodiment, the division of the frequency range is performed as a Fast Fourier Transform (FFT), which is performed periodically with or without time overlap to produce successive frequency packs, each of which has a given frequency width. Other ways to segregate the frequency components of audio signals are also available, such as wavelet transformation, discrete Walsh-Hadamard transformation, discrete Hadamard transformation, discrete cosine transformation, as well as numerous filtering methods.
After DSP 104 divided the frequency components of the digital audio signal into successive frequency packets, as mentioned above, it proceeds to assess the ability of the various frequency components present in the audio signal to mask various code components discharged by memory 110, and to produce appropriate amplitude coefficient settings, which are used to determine the amplitudes of various code frequency components, so that they will be masked by the sound signal of the program when played as sound so that they remain inaudible to the human ear. These operations are represented by block 134 in Fig. 5.
For audio signal components that are substantially simultaneous with the code frequency components to be masked (but which precede the code frequency components for a short period of time), the ability to mask the components of the sound program is estimated on the basis of tonal as well as on the basis of narrow masking bandwidth, and based on broad masking band as described below. For each of the code frequency components that is discharged from memory 110 at a given time, the tonal masking capability is calculated for each of the multiple frequency components of the audio signal based on the energy level in each of the respective packets that these components fall into, as well as based on the frequency dependencies of each packet and the corresponding code frequency component. The estimation in each case (tonal, narrowband, broadband masking) may be in the form of an amplitude determination factor or other measurement enabling the assignment of the code component amplitude in such a way that the code component will be masked by an audio signal. Alternatively, the estimate may be a sliding tonal analysis.
In the case of narrowband masking, in this embodiment, for each relevant code frequency component, the energy content of frequency components below the assumed level in the assumed frequency band is calculated, including the corresponding code frequency component for obtaining a separate estimate
183 307 camouflage capabilities. In some implementations, narrowband masking capability is measured based on the energy content of the signal frequency components below the average packet energy level in the assumed frequency band. In this implementation, the energy levels of the components below the average energy of the packet (being the energy threshold) are added together to produce a narrowband energy level by which the corresponding code components are identified. Instead, a different narrowband energy level may be generated by selecting a threshold component other than the average energy level. Furthermore, in accordance with other embodiments, the average energy level of all components of the audio signal may be used as the narrowband energy level to assign the narrowband masking estimate to the corresponding code component. According to yet other embodiments, the tonal energy content of the audio signal components in the assumed frequency band is used for this purpose, while in other embodiments the level of the minimum component in the assumed frequency band is used.
Finally, in some implementations, the broadband energy content of the audio signal is determined to calculate the ability of the audio signal to mask the corresponding code frequency component by broadband masking. In this embodiment, the estimation of wideband masking is based on the minimum narrowband energy level determined during the estimation of narrowband masking described above. That is, if four separate assumed frequency bands were tested during the estimation of narrowband masking, as described above, and the wideband noise is included in the minimum narrowband energy level among all four assumed (designated) frequency bands, then this minimum narrowband energy level is multiplied by a ratio equal to the ratio of the frequency range to the width of the assumed frequency band having the minimum narrowband energy level. The result indicates the allowable total code power level. If the total allowable code power level is labeled P and the code contains ten code components, each of them is then assigned an amplitude determination factor to result in a component power level that is 10 dB smaller than P. Alternatively, broadband noise is calculated for an assumed relatively wide band comprising code components, by selecting one of the methods described above for estimating a narrow band energy level, but using audio components from the entire assumed relatively wide band. After determining the broadband noise in the selected manner, an appropriate estimate of the wideband masking is assigned to each respective code component.
The amplitude determination factor for each code frequency component is then selected on the basis that one of the tonal, wideband or narrowband masking estimates gives the highest acceptable level of amplitude for the respective component. This increases the likelihood that each corresponding frequency component of the code will be masked in such a way that it will remain inaudible to the human ear.
The amplitude determination factors are selected for each of tonal, narrowband and wideband masking based on the following factors and conditions. In the case of tonal masking, coefficients are assigned based on the frequencies of the components of the audio signal, whose masking abilities are estimated, and the frequency of one or more masked code components. In addition, a given sound signal in any selected range provides the ability to mask a given code component in the same range (i.e. simultaneous masking) at a maximum level greater than that at which the same sound signal is able to mask the same code component before or after the selected interval (i.e. non-simultaneous masking). The conditions in which the encoded audio signal will be heard by the listeners in an appropriate manner are also taken into account. For example, if the sound of a TV signal is to be encoded, distorting effects of a typical environment to
183 307 listings are usefully considered if in these circumstances some frequencies are more distorted than others. Reception and playback equipment (such as graphic equalizers) may use similar effects. Effects related to the environment or equipment can be compensated by selecting low enough amplitude coefficients to provide masking under the expected conditions.
In some embodiments, only one of the tonal, narrowband or broadband capabilities is estimated. In other embodiments, two of these different types of masking capabilities are estimated, and in yet other embodiments, all three are used.
In some embodiments, sliding tonal analysis is performed to evaluate the masking ability of the sound signal. Sliding tonal analysis generally meets the masking rules for narrowband, broadband noise and single tones without the need for sound classification. In sliding tonal analysis, the sound signal is considered as a set of discrete tones, each of which is centered in the appropriate FFT frequency pack. Generally, sliding tonal analysis first calculates the sound signal strength in each FFT packet. Then, for each code tone, the effects of masking discrete tones of the audio signal in each FFT packet divided by frequency by no more than the critical bandwidth of the audio tone are calculated based on the strength of the sound signal in each such packet, using masking relationships for masking by single tones. The masking effects of all relevant discrete tones from the audio signal are added for each code tone, then set for the number of tones in the critical tone band of the audio signal and the complexity of the audio signal. As explained below, in some embodiments, the complexity of the program material is empirically determined based on the power ratio in the respective tones of the sound signal and the root of the sum of squares of power in those tones of the sound signal. The complexity is used to state that narrowband noise and broadband noise, each, provide much better masking effects than those obtained by simple tone summation used to model narrowband and wideband noise.
In certain embodiments that use sliding tone analysis, the assumed number of audio signal samples is first subjected to high FFT, which provides high resolution but requires longer processing time. Then subsequent portions of the assumed number of samples are subjected to a relatively smaller FFT, which is faster, but provides lower resolution. The amplitude coefficients determined from the large FFT are combined with those determined from the smaller FFT, which essentially corresponds to the time weighting of the higher high FFT frequency accuracy by the greater time precision of the smaller FFT.
In the embodiment shown in Fig. 5, after selecting the appropriate amplitude determination factor for each of the frequency components of the code output from memory 110, DSP 104 sets the amplitude of each frequency component accordingly, which is indicated by the functional block determining the amplitude 114. In other embodiments, each code frequency component is initially generated in such a way that its amplitude matches its respective determination factor. Referring to Fig. 6, the amplitude determination operation performed by DSP 104 in this embodiment leads to the multiplication of the selected ten values from the code frequency in the time domain from f1 to f40 for the current time interval from tl to tn by the respective determination factors GA1 to GA10, and then DSP 104 adds time-domain components with a fixed amplitude to produce the total code signal that is output to output 106. Referring to Fig. 3 and 5, the total code signal is converted by a digital-to-analogue converter (DAC) 140 and fed to the first input of the summing circuit 142. The summing circuit 142 receives a sound signal from the input terminal 94 at the second input and adds the total analog code signal to the analog sound signal and escorts him at the tip of 146.
When used to broadcast radio broadcasts, the encoded audio signal modulates the carrier wave and is transmitted by air. On NtSC television, frequency 14
183 The encoded audio signal modulates the subcarrier and is mixed with the video component so that the combined signal is used for carrier modulation in air transport. Of course, classic TV and radio signals can also be transmitted by cable (for example, classical or optical fiber), satellite, or other means. In other applications, the encoded audio may be recorded either for distribution in a recorded form, or for later distribution, or for another wide distribution method. Coded audio can also be used in point-to-point transmissions. Various other applications, transmission methods and recording methods are of course possible.
Figures 7A to 7C show flowcharts showing the course of program procedures performed by DSP 104 to implement the tonal, narrowband and broadband estimates of the functions described above. Figure 7A illustrates the main DSP 104 loop. The program is initiated by an instruction from the main processor 90 (step 150), after which the DSP 104 initializes its hardware registers (step 152) and then proceeds to step 154 to determine the unweighted data in the time domain of the code component, as shown in Fig. 6, which it is then stored in memory to be read when needed to generate time-domain code components as mentioned above. Alternatively, this step may be omitted if the code components are stored permanently in ROM or other non-volatile memory. It is also possible to calculate code frequency component data on demand, which, however, places a higher processing load. Another way is to produce unweighted code components in analog form and then determine the amplitudes of the analog components using the weighting factors generated by the digital processor.
After the time domain data has been calculated and collected, in step 156, DSP 104 sends a request to the main processor 90 asking for the next information to be encoded. The information is in the form of a stream of characters, integers or other unique data symbols identifying code component groups that are discharged by DSP 104 in the order that is assumed by the information. In other embodiments, the master processor, knowing the speed of discharging data from the DSP, determines for itself when to provide the next information to the DSP by appropriately setting the timer and providing the information in a time-synchronized manner. In another alternative embodiment, the output from the DSP 104 is coupled to a decoder to receive the output code components for decoding and feedback information to the main processor as the DSP output, so that the main computer can determine when to provide the next information to the DSP 104. In yet other embodiments, the functions of the main processor 90 and DSP 104 are performed by a single processor.
After the next information has been received from the main processor, in step 156, DSP proceeds to the generation of code components for each information symbol in order and provides combined, balanced code frequency components to its output 106. This process is represented by a loop identified by 160 in Fig. 7A.
After entering loop 160, DSP 104 permits interrupts 1 and 2, and proceeds to the procedure for determining the weighting factors 162, which will be described in connection with the flowchart of Figures 7B and 7C. Referring first to Fig. 7B, after entering subroutine 162, DSP first determines whether sufficient amount of audio signal samples has been collected to enable high resolution FFT to perform spectral analysis of the audio signal content in the last assumed audio signal interval, which is designated as step 163. To start it is necessary that sufficient samples of audio signals are collected to perform the FFT. However, if overlapping FFT is used, during subsequent loop passes, a correspondingly smaller number of samples must be collected before the next FFT.
As will be seen in Fig. 7B, DSP remains in the small loop 163 waiting for the necessary collection of samples. Under the influence of interrupt 1, the A / D converter 124 provides a new digital sample of the program audio signal, which is stored in the DSP data buffer 104, which is designated as subprocess 164 in Fig. 7A.
183 307
Returning to Fig. 7B, after the DSP has accumulated enough data samples, the processing proceeds to step 168, in which said high resolution FFT is performed on the audio signal data samples of the last audio signal interval. Then, as indicated by 170, the corresponding weighting or amplitude coefficient is calculated for each of the frequency components of the code in the currently encoded symbol. At step 172, the frequency packet generated by the high frequency FFT (step 168) that provides the ability to mask the highest level of the corresponding code component based on a single tone (dominant tone) is determined as described above.
Referring also to Fig. 7C, at step 176, the weighting factor for the dominant tonne is determined and stopped for comparison with the respective masking capabilities provided by wideband and narrowband masking, and, if found to be the most preferred masking method, is used as a weighting factor for setting the amplitude of the current code frequency component. In the next step 180, an assessment of narrowband masking and wideband masking is performed as previously described. Then, in step 182, it is determined whether narrowband masking provides the best ability to mask the corresponding code component, and if so, in step 184 the weighting factor is changed based on narrowband masking. In the next step 186, it is determined whether broadband masking provides the best ability to mask the corresponding code frequency component, and if so, in step 190 the weighting factor is changed based on broadband masking. Then, in step 192, it is determined whether the weighting factors have been chosen for each code frequency component that is currently to be returned to the current symbol representation, and if not, the loop is reinitialized to select the weighting factor for the next code frequency component. However, if the weighting factors for all components have been selected, then the subprocedure is completed, which is marked as step 194.
After the appearance of interrupt 2, the processing proceeds to the subroutine 200, in which the functions shown in Fig. 6 are performed. That is, in the subroutine 200, the weighting factors calculated in the subroutine 162 are used to multiply the corresponding time domain values of the current symbol to be discharged and then the time-weighted values of the code components are added and derived as a weighted total code signal to DAC 140. Each code symbol is discharged within a predetermined period of time, after which processing proceeds to step 156 from step 202.
Figures 7D and 7E show flowcharts for implementing tonal slide analysis for calculating masking effects in an audio signal. At step 702, variables such as large FFT and smaller FFT sample sizes, the number of smaller FFT to large FFT, and the number of code tones per symbol are initiated, e.g. 2048, 256, 8 and 10, respectively.
In steps 704-708, a number of samples corresponding to a large FFT are analyzed. At step 704, audio samples are taken. At step 706, the program material strength in each FFT package is retrieved. At step 708, an acceptable code tone power is obtained in each respective FFT packet, due to the effects of all the corresponding beeps in that packet, for each tone. The flow chart of Fig. 7E shows step 708 in greater detail. In steps 710-712, a number of samples corresponding to a smaller FFT are analyzed. At step 714, the allowable code powers determined from the large FFT in step 708 and from the low FFT in step 712 are combined for the portion of the samples that were subjected to the smaller FFT. In step 716, the code tones are mixed with the beep to create the encoded sound, and in step 718 the encoded sound is led to the DAC 140. In step 720, it is decided whether to repeat steps 710-718, i.e., whether the remaining parts of the audio signal samples that have passed a large FFT and not passed a smaller one. Then in step 722, if there are no more sound samples, the next number of samples, corresponding to a large FFT, is analyzed.
Figure 7E shows details of steps 708 and 712, determining the allowable code strength in each FFT packet. Generally, this procedure models the signal
183 307 sound as containing a set of tones (see examples below), calculates the effect of masking by each tone of the audio signal of each code tone, sums up masking effects and sets the code tone density and signal complexity.
At step 752, the band involved is determined. For example, let the coding band used be 800 Hz to 3200 Hz and the sampling frequency 44/00 samples per second. The initial package starts at 800 Hz and the last package is at 3200 Hz.
At step 754, the masking effect of each corresponding tone of the audio signal for each code tone in this packet is determined using a masking curve for each code tone, and compensation is made for the non-zero packet width of the audio signal FFT by determining (/) the first masking value based on the assumption that all of the sound signal power is at the upper end of the packet and determining (2) the second masking value based on the assumption, that all of the sound signal power is at the bottom end of the packet, after which the one from the masking value that is smaller is chosen.
Figure 7F shows an approximation of a single-tone masking curve for a tone of an audio signal at a frequency of fPGM that is about 2200 Hz in this example, according to the work of JJ Zwislocki under the title Masking: Experimental and Theoretical Aspects of Simultaneous, Forward, Backward and Central Masking, / 978 , Zwicker et al., Psychoacoustics edition: Facts and Models, pages 283-3 / 6, Springer-Yerlag, New York. The critical bandwidth (CB) is defined by Zwislocki as follows:
critical band = 0.002 * fPGM / .5 + / 00
According to the following definitions, where the mask is the tone of the sound signal:
/ ± 0.3 critical bands / / - / 6 dB from the mask / / - 26 dB from the mask / / - 24 dB on the kyttycrne band / / - 7 dB on the critical band /
BRKPOINT = 0.3
PEAKFAC = 0.025 // 9
BEATFAC - 0.0025 / 2 mNEG = - 2.40 mPOS = - 0.70 cf = code frequency mf = mask frequency cband = critical band around fPGM masking factor, mfactor, can be calculated as follows: brkpt = cband * BRKPOINT if on a negative slope of the curve with Fig. 7F mfactor = PEAKFAC * / O ^ mNEG ^ f-brkpt-cfycband) if on the flat part of the curve in Fig. 7 mfactor = BEATRAC if on the positive slope of the curve in Fig. 7F mfactor = PEAKFAC * / 0 ** (mPOS * mf-brkpt-cf) / cband).
Specifically, the first mfactor is calculated based on the assumption that all the sound signal power is at the bottom end of its packet, the second mfactor is calculated assuming that all the sound signal power is at the top end of its packet, and the smaller of these two factors is selected as the masking value provided by the tone of the audio signal for the selected code tone. At step 754, processing is performed for each specific tone of the audio signal for each tone of code.
At step 756, each code tone is determined by each of said masking factors corresponding to the tones of the audio signal. In this embodiment, the masking factor is multiplied by the strength of the sound signal in a given packet.
At step 758, the result of multiplication of the masking coefficients by the sound signal strength is summed for each packet, providing acceptable power for each tone of code.
183 307
At step 760, the allowable powers in the critical band on each side of the code tone that is being calculated and for the complexity of the audio signal are determined for the number of code tones. The number of code tones in the critical band calculated by CTSUM is calculated. The setting factor, ADJFAC, is given by:
ADJFAC = GLOBAL * (PSUM / PRSS) 1.5 / CTSUM, where GLOBAL is the derating factor specifying the encoder inaccuracy due to time delays in FFT performance, (PSUM / PRSS) 1.5 is the experimental complexity correction factor, and 1 / CTSUM represents simple power sharing a beep for all code tones to be masked. PSUM is the sum of the levels of masking tone assigned to mask the code tones that ADJFAC is determined. The root of the sum of power squares (PRSS) is determined by
PRSS = SQRT (2 (Ę<sup>2</sup>)) where i - FFT packets in the band
For example, assuming that the total masking power in the band is evenly distributed over one, two or three tones, then:
<td>No. tone</td><td>tone power</td><td>Ptotal</td><td>PRSS</td>
<td> 1</td><td> 10</td><td> 1*10= 10</td><td> 10</td>
<td> 2</td><td> 5,5</td><td> 2*5 = 10</td><td>SQRT (2 * 5<sup>2</sup>) = 7.07</td>
<td> 3</td><td> 3.3, 3.3, 3.3</td><td> 3*3.3 = 10</td><td>SQRT (3 * 3.3<sup>2</sup>) = 5.77</td>
Hence, PRSS measures the degree of masking power focus (increasing values) or scattering (decreasing values) of program material.
In step 762 of Fig. 7E, it is determined if there are still packets in the band under consideration and if so, they are processed as described above.
Examples of masking calculations will now be presented. The audio signal symbol is assumed at 0 dB, so that the values obtained are the maximum code tone powers relative to the sound signal power. Four cases are considered: single 2500 Hz tone; three tones at 2000, 2500 and 3000 Hz, narrow band noise modeled as 75 tones in the critical band with a center at 2600, where 75 tones are evenly spaced every 5 Hz in the range from 2415 to 2785 Hz; and broadband noise modeled as 351 tones evenly spaced every 5 Hz in the range from 1750 to 3250 Hz. For each case, the result obtained from the sliding tonal analysis (STA) is compared with the calculated result of the best of three types of analysis: single tone, narrowband noise and broadband noise.
<td></td><td colspan="2">Single tone</td><td colspan="2">Many tones</td><td colspan="2">Narrow band noise</td><td colspan="2">Broadband noise</td>
<td>tons of code</td><td>STA</td><td>best</td><td>STA</td><td>best</td><td>STA</td><td>best</td><td>STA</td><td>best</td>
<td>(Hz)</td><td>(DB)</td><td>with 3 (dB)</td><td>(DB)</td><td>with 3 (dB)</td><td>(DB)</td><td>with 3 (dB)</td><td>(DB)</td><td>with 3 (dB)</td>
<td> 1</td><td> 2</td><td> 3</td><td> 4</td><td> 5</td><td> 6</td><td> 7</td><td> 8</td><td> 9</td>
<td> 1976</td><td> -50</td><td> -49</td><td> -28</td><td> -30</td><td> -19</td><td>ON</td><td> 14</td><td> 12</td>
<td> 2070</td><td> -45</td><td> -45</td><td> -22</td><td> -32</td><td> -14</td><td>ON</td><td> 13</td><td> 12</td>
<td> 2163</td><td> -40</td><td> -39</td><td> -29</td><td> -25</td><td> -9</td><td>ON</td><td> 13</td><td> 12</td>
<td> 2257</td><td> -34</td><td> -33</td><td> -28</td><td> -28</td><td> -3</td><td>ON</td><td> 12</td><td> 12</td>
<td> 2351</td><td> -28</td><td> -27</td><td> -20</td><td> -28</td><td> 1</td><td>ON</td><td> 12</td><td> 12</td>
183 307
cd table
<td> 1</td><td> 2</td><td> 3</td><td> 4</td><td> 5</td><td> 6</td><td> 7</td><td> 8</td><td> 9</td>
<td> 2444</td><td> -34</td><td> -34</td><td> -23</td><td> -33</td><td> 2</td><td> 7</td><td> 13</td><td> 12</td>
<td> 2538</td><td> -34</td><td> -34</td><td> -24</td><td> -34</td><td> 3</td><td> 7</td><td> 13</td><td> 12</td>
<td> 2632</td><td> -24</td><td> -24</td><td> -18</td><td> -24</td><td> 5</td><td> 7</td><td> 14</td><td> 12</td>
<td> 2726</td><td> -26</td><td> -26</td><td> -21</td><td> -26</td><td> 5</td><td> 7</td><td> 14</td><td> 12</td>
<td> 2819</td><td> -27</td><td> -27</td><td> -22</td><td> -27</td><td> 6</td><td>ON</td><td> 15</td><td> 12</td>
For example, in the tonal slide analysis (STA) for one tone case, the masking tone is 2500 Hz, which corresponds to a critical bandwidth of 0.02 * 25001.5 + 100 = 350 Hz. The breaking points of the curve of Fig. 7F are at 2500 ± 0.3 * 350, i.e. 2395 and 2605 Hz. The code frequency of 1976 can be seen on the negatively inclined part of the Fig. 7F curve, so the masking factor is:
mfactor = 0.025119 * 10-2.4 * (2500 - 105 - 1976) / 350 = 3.365 * 10-5 = -44.7 dB
There are three tones of codes in the 1976 Hz critical band, so masking power is divided between them:
4.364 * 10-5 / 3 = -49.5 dB
This result is rounded to -50 dB and is shown in the upper left corner of the results table.
In the best of three analysis, tonal masking is calculated according to the single-tone method explained above with reference to Fig. 7F.
In the best of three analysis, narrowband noise masking is calculated by first calculating the average power in a critical band centered around the frequency of the code tone being considered. Tonal ones with power greater than average power are not considered as part of the noise and are removed. The sum of remaining power is the power of narrowband noise. The maximum allowable code tone power is - 6 dB narrowband noise power for all tones in the critical band of the code tone under consideration.
In the best of three analysis, broadband noise masking is calculated by calculating the power of narrowband noise for critical bandwidth centers at 2000, 2280, 2600 and 2970 Hz. The minimum of the calculated narrowband noise powers are multiplied by the ratio of the total bandwidth to the corresponding critical bandwidth to find the broadband noise power. For example, if a band at 2600 Hz has a critical bandwidth of 370 Hz and the least power, its narrowband noise power is multiplied by 1322Hz / 370Hz = 3.57 to obtain broadband noise power. The permissible code tone power is - 3 dB broadband noise. When we deal with 10 tones of code, the maximum power allowed for each of the tones is 10 dB lower, i.e. -13 dB of wideband noise power.
As you can see, the calculations of the sliding tonal analysis generally correspond to the calculations of the best of three, which indicates that the sliding tonal analysis is a reliable method. In addition, the results obtained by this analysis for the case of many tones are better, that is, they enable greater code tone powers than the best of three analyzes, which indicates that sliding tonal analysis is suitable even for cases that do not fit well into one of the calculations, the best of three.
Referring to Fig. 8, a block diagram of an embodiment of an encoder that uses an analog circuit is shown. The analog encoder receives an audio signal in analog form at the input terminal 210, from which the audio signal is supplied as input to the N generator components, from 2201 to 220N, each of which generates the corresponding code component from Cl to CN. For simplicity and clarity, only one of the generator arrangements 2201 to 220N is shown in Figure 8. For the controlled generation of code components for the corresponding data symbol to be included in
183 307 of the audio signal, so that an encoded audio signal is created, input to each of the generator components is input through appropriate terminals from 2221 to 222N, which serve as permission inputs for the respective generator components. Each symbol is encoded as a subset of the code components from Cl to CN by selectively applying the enable signal to certain generator components from 2201 to 220N. The generated code components corresponding to each data symbol are delivered to the inputs of summation circuit 226, which also takes an input audio signal from input terminal 210, which circuit is used to add code components to the input audio signal to produce an encoded audio signal that is output to the output .
Each of the component generator systems is similar in design and contains a suitable system for determining the weighting factor, from 2301 to 230N, a suitable signal generator, from 2321 to 232N, and a corresponding switching system, from 2341 to 234N. Each of the signal generators, from 2321 to 232N, respectively, produces a different code frequency component and provides the generated component to the appropriate switching system, 2341 to 234N, each of which has a second input shorted to ground and the output is coupled with the input to the appropriate multiplier from 2361 to 236N . In response to receiving an enabling signal at the appropriate data input terminal, from 2201 to 220N, each of the switching circuits, from 2341 to 234N, responds by coupling the output of the respective signal generator, from 2321 to 232N, with the input of the appropriate multiplier, from 2361 to 236N . Meanwhile, in the absence of a enable signal at the data input, each switching system, from 2341 to 234N, shorts its output to ground, so that the output of the respective multiplier, from 2361 to 236N, is zero.
Each weighting factor determination system, from 2301 to 230N, is used to estimate the ability of the frequency components of the sound signal in the appropriate frequency band of this signal to mass the code component generated by the appropriate generator system, from 2321 to 232N, to produce a weighting factor, which is then given to enter the appropriate multiplier system, from 2361 to 236N, to determine the amplitude of the relevant code component, to ensure that it will be masked by the part of the sound signal that has been calculated by the weighting factor determination system. Referring also to Fig. 9, the design of each of the weighting factor determination systems, 2301 to 230N, designated as an example system 230, is shown in block form. System 230 includes a masking filter 240 that receives an audio signal at the input and is used to extract the portion of the audio signal to be used to calculate the weighting factor to be given to the respective of the multipliers from 2361 to 236N. The properties of the masking filter are also selected so as to balance the amplitudes of the frequency components of the audio signal due to their ability to mask the corresponding code component.
The portion of the audio signal selected by the masking filter 240 is provided to the absolute value determination system 242, which produces an output representing the absolute value of the signal portion in the frequency band after passing through the masking filter 240. The output of the 242 absolute value system is provided as input to a 244 scaling amplifier, having a gain selected to produce a signal that, when multiplied by the output of the appropriate switch, from 2341 to 234N, creates a code component at the output of the appropriate multiplier, ad 2361 to 236N, which will ensure that the multiplied code component is masked by the selected part of the audio signal that has passed through the masking filter 240, when playing a coded audio signal as sound. Each weighting factor determination system, from 2301 to 230N, therefore produces a signal representing an estimate of the ability of a selected portion of the sound signal to mask the corresponding code component.
In other embodiments of the analog encoders of the present invention, a plurality of weighting factor determination systems are provided to the generator of each code component, and each of the plurality of weighting factor determination systems corresponding to a given code component calculates the ability of the various parts of the sound signal to have 20
183 307 when a particular component is scanned when the encoded audio signal is played as sound. For example, multiple weighting factor determination systems may be provided, each of which calculates the ability of a portion of the audio signal in a relatively narrow frequency band (such that the energy of the audio signal in such a band will most likely consist of a single frequency component) to mask the corresponding component code when the encoded audio signal is played as sound. A further weighting factor determination system for the same code component, for calculating the sound signal energy capacity, in a critical band whose central frequency is the code component, may be provided to mask the code component when the encoded audio signal is played as sound.
In addition, although various elements of the embodiment of Figures 8 and 9 are implemented as analog circuits, it is possible to provide the same functions performed by digital circuits.
decoding
Decoders and decoding methods that are particularly suited for decoding audio signals encoded by the methods of the invention described above, as well as generally for decoding codes contained in audio signals such that the codes can be distinguishable from the rest of the signal based on amplitude, will now be described. According to certain features of the present invention, and with reference to the block diagram of Fig. 10, the presence of at least one of the code components in the encoded audio signal is detected by establishing the expected amplitude or amplitudes of the at least one code component based on the sound signal level or the noise signal level of the bear signal, or both, as indicated by function block 250. One or more signals representing such an appropriate amplitude or amplitude are provided at 252 in Fig. 10, for determining the presence of a code component by detecting a signal corresponding to the expected amplitude or amplitudes, as indicated by functional block 254. The decoders of the present invention are particularly well suited for detecting the presence of code components that are masked by other components of the audio signal if the amplitude relationship between the components of the code and other components of the sound signal are, to some extent, pre-established.
Figure 11 is a block diagram of an embodiment of a decoder according to the present invention that performs digital signal processing for extracting codes from encoded audio signals received by the decoder in an analog form. Decoder with fig. 11 has an input terminal 260 for receiving an encoded analog audio signal, which may be, for example, a signal taken from a microphone, from radio or television broadcasts, reproduced as sound by a receiver, or an encoded analog audio signal in the form of electrical signals directly from such a receiver. Such encoded analogue sound can be created by playing a sound recording, such as on a compact disc or tape. Analog processing 262 circuits are coupled to input 260 to receive coded analog sound and are used to amplify the signal, automatic control of low-pass filtering gain, preventing aliasing, before analog-to-digital conversion. In addition, analog processing circuits 262 are used to perform bandpass filtering to ensure that the output signals are limited to the frequency range at which the code may occur. Analogue processing 262 discharges processed analog signals to an analog-to-digital (A / D) converter 263, which converts the received signals to digital form and delivers them to a digital signal processor (DSP) 266, which already processes digital signals to detect the presence of code components and designates the code symbols they represent. Digital signal processor 266 is coupled to memory 270 (containing program and data memories) and input / output (I / O) 272 to receive external commands (e.g., decode initiation order or command to remove stored codes), and to remove decoded information .
183 307
The operation of the digital decoder of Fig. 11, which decodes encoded audio signals by means of the device of Fig. 3, will now be described. The analog processing system 262 serves as a bandpass filter for encoded audio signals, with a passband of approximately 1.5 kHz to 3.1 kHz and DSP 266 samples the filtered analog signals at a sufficiently high frequency. The digital audio signal is then divided by DSP 266 into frequency component ranges or FFT processing packets. More specifically, an overlapping window FFT is performed on the assumed number of most recent data points, so that a new FFT is periodically performed on a set of sufficient new samples. The data is weighted as described below, and an FFT is performed to produce the assumed number of frequency packs, each having a assumed width. The energy B (i) of each frequency packet in the range covering the component frequencies of the code is calculated by DSP 266.
Noise level estimation is performed around each packet in which the noise component may occur. Accordingly, when the decoder of Fig. 11 is used to decode signals encoded by the embodiment of Fig. 3, there are 40 frequency packs at which code components may occur. For each frequency packet, the noise level is estimated as follows. First, the average energy E (j) in frequency packets is calculated in a window extending to frequencies below and above a particular frequency packet j (that is, the packet in which the code component may occur), according to the following relationship:
ΕΟ) = ^ ΣΒ (0 where i = (jw) -> (j + w), and w represents the window size below and above the package under consideration, expressed in the number of packages. Then the noise level NS (j) in the frequency package j is calculated, according to the following formula:
<sup>ns</sup>(J) = (Σ<sup>Βη</sup>(Ϋ) / Σ<sup>δ</sup>ω) where Bn (i) equals B (i) (energy level in the package i) if B (i) <E (j), and equals 0 otherwise, and 5 (i) equals 1 if B (i) <EG), and 0 otherwise. That is, it is assumed that the noise components contain components having a level less than the average energy level in a particular window surrounding the package under consideration and thus contain components of the audio signal that fall below the average energy level.
When the noise level for the packet in question has been estimated, the signal-to-noise ratio SNR (j) for this packet is estimated by dividing the energy level B (j) in the packet by the estimated noise level NS (j). SNR (j) values are used to detect and timing synchronization symbols as well as data symbol states, as described below. Various methods may be used to eliminate components of the audio signal as potentially non-code components on a statistical basis. For example, it can be assumed that the packet with the highest signal-to-noise ratio contains an audio signal component. Another option is to exclude packages that have SNR (j) above the assumed value. Another option is to eliminate packages with the highest and / or lowest SNR (j).
When used to detect the presence of codes in the audio signals encoded with the device of Fig. 3, the device of Fig. 11 collects data indicating the presence of code components in each of the considered packages in a cyclical manner, for at least the main part of the assumed compartment in which the symbol can be found code. Accordingly, the following process is repeated many times and current component data is collected for each of the considered packages in a given time frame. Methods for determining the appropriate detection time frames based on synchronization codes will be wrapped below in larger ones
183 307 details. When DSP 266 collects data for the appropriate frame, it determines as follows which of the possible code signals were present in the signal. DSP 266 then stores the detected code symbol in memory 270 together with a time stamp for identifying the point at which the signal was detected based on the DSP internal timing signal. Then, in response to the appropriate command for DSP 266 received from the I / O 272 system, DSP causes memory 270 to discharge stored code symbols and timestamps through the I / O 272 systems.
The flowcharts of Figures 12A and 12B illustrate the sequences of operations performed by DSP 266 when decoding an encoded symbol in an analog audio signal received at an input terminal 260. Referring first to Fig. 12A, after initiating the decoding process, DSP 266 enters the main program loop in step 450, in which the SYNCH flag is set so that DSP 266 first begins the operation of detecting sync symbols E and S in the input audio signal in the assumed information order. After DSP 266 completes step 450, DSP invokes the DET subroutine, which is illustrated in the flow chart of FIG. 12B, used to search for the presence of code components representing synchronization symbols in the sound signal.
Referring to Fig. 12B, at step 454, DSP collects and stops samples of the input audio signal repeatedly until a sufficient number of the FFTs described above have been collected. After completing this operation, the collected data is subjected to a weighing function, such as the cosine square weighing function, · Kaiser-Bessel function, Gauss (Poisson function), Hanning function or other appropriate weighing function, as marked as step 456, setting the data windows. However, when the code components are clear enough, no weighing is necessary. The data windows are then subjected to an overlapping FFT, indicated by step 460.
When the FFT has been completed, the SYNCH flag is tested in step 462 to see if it is set (in this case, the sync symbol is expected) or whether it is reset (in this case, the data bit symbol is expected). Because DSP initially sets the SYNCH flag to detect the presence of code components representing synchronization symbols, the program proceeds to step 466 because the frequency domain data obtained from the FFT in step 460 is calculated to determine whether this data indicates the presence of components representing the synchronization symbol E or the synchronization symbol S .
To detect the presence and timing of synchronization symbols, the sum of the values, SNR (j) is first determined for each possible synchronization symbol and data symbol. At any given time during the process of detecting synchronization symbols, a specific symbol is expected. As the first step in detecting the expected symbol, it is determined whether the sum of its respective SNR (j) values is greater than any other. If so, the detection threshold is set based on noise levels in frequency packets that may contain code components. That is, if at any given time, only one code symbol is included in the encoded audio signal, only a quarter of the packages considered will contain the code components. The remaining three-quarters of the packages will contain noise, i.e., components of the sound program and / or other additional energies. The detection threshold is generated as the average SNR (j) for all forty packages considered, but can be determined by a multiplication factor to take into account the effects of neutral noise and / or to compensate for the observed error. Once the detection threshold has been established, the sum of the SNR0 values of the expected synchronization symbol is compared with the detection threshold to determine whether it is greater or not above this threshold. If so, the detection of the expected sync symbol is noted. Upon finding this fact, as indicated by step 470, the program returns to the main processing loop of Fig. 12A to step 472, where it is determined (as described below) whether the decoded data pattern corresponds to the assumed eligibility criterion. If not, processing returns to step 450 to re-start the sync symbol presence test in the audio signal, but if these criteria are met, it is determined whether the expected sync pattern (i.e., the expected E and S symbol sequence) has been completely received and detected, indicated as step 474.
However, after first passing through the DET subroutine, insufficient data will be retained to determine if the pattern meets the eligibility criteria, so that from step 474 processing returns to the DET subroutine to perform further FFT and calculate the presence of the synchronization symbol. When the DET subroutine has been performed the predetermined number of times, when processing returns to step 472, DSP determines whether the collected data meets the qualification criteria for the synchronization pattern.
That is, when DET was performed such an assumed number of times, the corresponding number of calculations was performed in step 466 of the DET subroutine. The number of detected occurrences of the E symbol is used in one embodiment as a measure of the amount of energy of the E symbol during an appropriate period of time. However, other E-symbol energy measurements (such as all SNRs of E packages that exceed the average pack energy) can also be used here. After the DET subroutine is called again and the calculation is performed in step 466, in step 472, the most recent calculation is added to the calculation accumulated in the assumed range, and the oldest calculation among those previously collected is rejected. This process continues during numerous passes through the DET subroutine, and in step 472 the energy peak of the E symbol is sought. If the peak was not found, it leads to the conclusion that the synchronization pattern was not found and processing returns from step 472 to step 450, where it sets the SYNCH flag again and starts searching for the synchronization pattern.
In the case where such a maximum energy of the E symbol has been found, the calculation process performed in step 472 after subroutine 452 is continued each time using the same number of calculations from step 466, but discarding the oldest calculation and adding the most recent, so that the moving data window is used for this purpose. During this process, after the assumed number of passes, in step 472 it is determined whether there has been a transition from the E symbol to the S symbol. In one embodiment, this is determined as the point where all SNRs of packages S, resulting from step 466, in the moving window will cross all SNRs of packages E during the same interval. Once such a transition point has been found, processing continues as previously described, i.e. the maximum energy of the S symbol is sought, which is designated as the largest number of S detections in the moving data window. If such a maximum is not found, or if no such maximum occurs in the assumed time frame after the maximum energy of the S symbol, processing proceeds from step 472 back to step 450 and the search for the synchronization pattern begins again.
If the above criteria are met, the sync pattern is declared in step 474 and processing proceeds to step 480 to determine the appropriate bit ranges based on the energy maxima of the E and S symbols, and the detected transition point. Instead of such a process of detecting the presence of a synchronization pattern, other strategies may be used. In another embodiment, when the synchronization pattern does not meet the criteria such as those described above, but approximates the qualification pattern (i.e., the detected pattern is not explicitly unclassifiable), determining whether the synchronization pattern has been detected can be suspended for further analysis, which is based on calculations performed (as described below) to determine the presence of data bit ranges following the potential synchronization pattern. Based on the total of the detected data, i.e. during the suspected synchronization pattern interval and during the expected bit range, a retrospective qualification of a possible synchronization pattern can be performed.
Returning to Fig. 12A, once the synchronization pattern has been positively classified, in step 480, as mentioned above, the time dependencies of the bit are determined based on two maxima and a transition point. That is, these values are averaged to determine the expected start and end points of subsequent data bit ranges. When this has been done, in step 482 the SYNCH flag is reset to indicate that DSP will be looking for the presence of possible bit states. Then the DET 452 subroutine is
183 307 called again and also referring to Fig. / 2B, this is done in the same manner as described above up to step 462, in which the SYNCH status flag indicates that a state bit should be determined and processing proceeds to step 486. In at step 486, DSP looks for the presence of code components indicating the state of the zero or one bit, as described above.
When this is completed, in step 470, processing returns to the main processing loop of Fig. / 2A in step 490, where it is determined whether enough data has been received to determine the state of the bit. To do this, multiple passes through the 452 subroutine must be made, so that after the first pass the processing returns to the DET 452 subroutine for further calculations based on the new FFT. Once subprocedure 452 has been performed the assumed number of times, the accumulated data is calculated in step 486 to determine whether the received data indicates a state of zero, a state of one or an indefinite state (which can be specified based on data parity). That is, the sum of SNR packages 0 is compared to the sum of SNr packages /. Which of them is larger determines the state of the data, and if they are equal, the state is undefined. Alternatively, if the SNR sums of packages 0 and / are not equal, but are close to each other, an undefined state may also be declared. Also, if there are more data symbols, the symbol for which the largest sum of SNr was found is declared as the received symbol.
When the processing returns to step 490 again, the bit state determination is detected and the processor proceeds to step 492, where the DSP stores data in memory 270, indicating the states of the respective bits making up the word with the predetermined number of symbols represented by the encoded symbols in the received audio signal. Then, in step 496, it is determined whether the received data applies to all bits of the coded word or information. If not, processing returns to the DET 452 subroutine to determine the bit state of the next expected information symbol. However, if it is determined in step 496 that the last information symbol has been received, processing returns to step 450 for setting the SYNCH flag to examine the presence of synchronization symbols represented by code components in the encoded audio signal.
Referring to Fig. / 3, in some embodiments, the non-coding components of the audio signal and / or other noise (generally referred to simply as noise in this context) are used to produce a comparative value, such as a threshold, as indicated by function block 276. One or more parts of the encoded audio signal are compared with the comparative value, which is designated as functional block 277, in order to detect the presence of code components. Advantageously, the encoded audio signal is first processed to isolate components in the band (or multiple bands) that may contain code components, and then they are collected for a period of time to average noise, as indicated by functional block 278.
Referring now to Fig. / 4, an embodiment of an analog decoder according to the invention is shown in block form. The decoder of Fig. / 4 includes an input terminal that is coupled to four groups of component detectors 282, 284, 286 and 288. Each group of component detectors 282 to 288 is used to detect the presence of code components in the input audio signal representing the corresponding code symbol. In the embodiment of Fig. / 4, the decoder device is designed to detect the presence of each of the 4N code components, where N is an integer, so that the code consists of four different symbols, each of which is represented by a unique group of N code components. Accordingly, four groups 282 to 288 contain 4N component detectors.
An embodiment of one of the 4N component detectors from groups 282 to 288 is shown in block form in Fig. / 5 and is referred to as component detector 290. Component detector 290 has input 292 coupled to input 280 of the decoder in Fig. / 4 for receiving encoded audio signal. Component detector 290 has an upper branch of the system containing a noise estimation filter 294, which in one embodiment is in the form of a bandpass filter to pass the energy of the audio signal in a band with a center frequency corresponding to the detected code component. In an alternative and preferred embodiment, the noise estimation filter 294 consists of two filters, one of which has a pass band extending up from the frequency of the corresponding detected component code, and the other filter has a pass band extending down from the detected component frequency code, so that both filters together transmit energy having frequencies above and below (but not containing) the frequency component, which is to be detected but occurring in its vicinity. The output of the noise estimation filter 294 is connected to the input of the absolute value determination system 296, which produces an output signal representing the absolute value of the output from the noise estimation filter 294, fed to the input of the integration circuit 300, which collects the signals supplied to it and discharges the value representing the signal energy in parts of the spectrum of the frequency adjacent but not containing it, with the frequency component to be detected, and returns this value to the non-inverting input of the 302 differential amplifier, which works like a logarithmic amplifier.
The component detector of Fig. 15 also has a lower branch comprising a signal estimation filter 306, whose input is coupled to the input 292, to receive an encoded audio signal for passing a frequency band substantially narrower than the broad band of the noise estimation filter 294, so that the estimation filter 306 passes signal components essentially only at the frequency of the code component to be detected. The signal estimation filter 306 has an output coupled to the input of the absolute value determination system 308, which is used to produce at its output a signal representing the absolute value of the signal from the signal estimation filter 306. The output from the absolute value determination system 308 is coupled to the input of the integration circuit 310. The integration circuit 310 collects the values discharged from the circuit 308 and creates an output signal representing energy in the narrow band of the signal estimation filter for the assumed period of time.
Integration circuits 300 and 310 each have reset inputs connected to each other to receive a common reset signal applied to terminal 312. The reset signal is provided by control system 314 shown in Fig. 14, which generates a reset signal periodically.
Returning to Fig. 15, the output from the integration circuit 310 is provided to the inverting input of the amplifier 302, which operates in such a way that it outputs a signal representing the difference between the output from the integration circuit 310 and the integration circuit 300. If the 302 amplifier is a logarithmic amplifier, the range of possible output values is narrowed to reduce the dynamic output range, for use with a window comparator 316 that detects the presence or absence of a code component during a given interval, as determined by control 314 by application of the zeroing signal. The window comparator 316 discharges the code presence signal when the input provided from the 302 amplifier falls between the lower threshold, applied as a fixed value to the input terminal of the lower comparator threshold 316, and the fixed upper threshold, applied to the input terminal of the upper comparator threshold 316.
Referring again to Figure 14, each of the N component detectors 290 of each detector group couples the output of the respective window comparator 316 to the input of the logic determination code 320. The system 320, controlled by the control system 314, collects various signals of the presence of code components from 4N detector systems components 290 for a large number of zeroing cycles, depending on the setting by control 314. At the end of the detection interval of a given symbol, determined as it will be described, the code determination logic 320 determines whether the code symbol has been received as the symbol for which the largest number of components was detected during the interval, and outputs the signal indicating the detected code symbol to output 322 . The output signal can be stored in memory, included in larger information or data file, transmitted or used in another way (for example as a control signal).
Symbol detection intervals for the decoders described above with reference to Figs. 11, 12A, 12B, 14 and 15 can be established based on the time dependencies of sym26
183 307 it hurts synchronization sent in each encoded information that has assumed duration and order. For example, the encoded information contained in the audio signal may consist of two data compartments of the encoded E symbol followed by two data compartments of the encoded S symbol, both of which are described with reference to Fig. 4. The decoders in Fig. 11, 12A, 12B, 14 and 15 operate to initially test for the presence of the first expected synchronization symbol, i.e. the encoded E symbol, which is transmitted in a predetermined time interval and determines its transmission interval. Then, the decoders examine the presence of the code components characterizing the S symbol and when they detect, the decoders determine its transmission interval. Based on the designated transmission intervals, a transition point from the E symbol to the S symbol is determined, and from this point detection intervals are determined for each data bit symbol. During the detection interval, the decoder collects the code components to determine the corresponding symbol transmitted during this interval in the manner described above.
Although the various elements of the embodiment of Figures 14 and 15 are implemented by analog circuits, it is obvious that it is possible for the same functions to be performed in whole or in part by a digital circuit.
Referring to Figures 16 and 17, a system for estimating the audibility of widely disseminated information such as television and radio programs is shown. Fig. 16 is a block diagram of a radio broadcast station for transmitting audio signals in ether that has been encoded to identify the station along with the broadcast time. If desired, the identification of the program or segment that is transmitted can also be enabled. The 340 program source, such as a compact disc player, digital cassette player or live sound source, is controlled by the station manager via control device 342 to control the evacuation of audio signals for broadcasting. The output 344 from the sound program source is coupled to the input of the encoder 348, according to the embodiment of Fig. 3 and comprising DSP 104, bandpass filter 120, analogue to digital converter (A / D) 124, digital to analogue converter (DAC) 140 and summing circuit 142. Control device 342 includes main processor 90, keyboard 96 and monitor 100 in the example of Fig. 3, so that the main processor included in the controller 342 is coupled to the DSP included in the encoder 348 of Fig. 16. The encoder 348 operates under the control of the controller 342 and incorporates the encoded information periodically into the sound to be transmitted, and the information contains appropriate identification data. The encoder 348 discharges the encoded sound to the input of the radio transmitter 350, which modulates the carrier wave with an encoded sound program and transmits it to the ether via a 352 antenna. The main processor included in the control device 342 is programmed with the help of the keyboard to control the encoder so as to discharge the corresponding coded information containing the station identification data. The main processor automatically generates transmission time data using the clock reference system it contains.
Also referring to Figure 17, the system's personal monitoring device 380 is enclosed in a housing 382 that is small enough to be carried by a member of an auditorium member participating in supervised auditor estimation. Each member of the auditorium is provided with a personal monitoring device, such as a 380 device, which is carried by an auditorium member during an audit period, such as an assumed period of one week. The personal monitoring device 380 includes an omni-directional microphone for collecting sounds surrounding a member of the auditorium, including radio programs being played as sound by the loudspeaker of the radio receiver, such as the radio receiver 390 shown in Fig. 17.
The personal monitoring device 380 also includes a 394 signal processing system having a 386 microphone-coupled input, used to amplify the output of this microphone and subject it to band filtering to suppress frequency from outside the audio frequency band, containing various code frequency components placed in the audio signal. program through the 348 encoder of Fig. 16, as well as to perform anti-aliasing filtering prior to analog-to-digital conversion.
183 307
The digital circuit of the personal monitoring device 380 is shown in Figure 17 in the form of a functional block diagram comprising a decoder block and a control block, both of which can be implemented as a digital signal processor. The program and data memory 404 is coupled to a decoder 400 for receiving detected codes for storage, as well as to a control block 402 for controlling write and read operations from the memory 404. The input / output (I / O) 406 is coupled to memory 404 for receiving data output by a personal monitoring device 380, as well as for storing information such as program instructions. The I / O 406 is also coupled to control block 402 to control the input and output operations of the 380 device.
The decoder 400 operates as the decoder of Fig. 1, described earlier, and drains the time code and station identification data for storage in memory 404. The 380 personal monitoring device is advantageously capable of cooperating with a docking station as described in USA Patent Application No. 08 / 101,558, filed on August 2, 1993, entitled Compliance Incentives for Audience Monitoring / Recording Devices, which is often combined with this invention and is cited herein by reference. In addition, the personal monitoring device is provided with additional features of a portable reception-oriented monitoring device such as described in the application just mentioned.
The supervising station communicates via modem via telephone lines with a central data processing device to load time and station identification data into it, which is used to create reports on the audience and / or viewership of programs. The central device may also send information to the supervising station for use by itself or for providing it to the device 380, such as an executable program. The central device may also provide information to the supervising station and / or devices 380 via an RF channel, such as existing FM broadcasts encoded with such information in a manner consistent with the present invention. The surveillance station and / or device 380 are provided with an FM receiver (not shown for greater clarity) that demodulates the encoded FM broadcasts and delivers them to the decoder according to the present invention. The encoded FM broadcast may also be provided by cable or by another transmission center.
In addition to monitoring by personal monitoring units, stationary units (such as set-top units) can be used. The set-top units may be combined to receive the encoded audio signal from the receiver or may use a microphone such as the 386 microphone in Fig. 17. Therefore, settop units can monitor selected channels, while monitoring the composition of the auditorium, with or without the present invention.
Other applications relate to the encoding and decoding methods of the present invention. In one application, commercial product soundtracks are provided with identification codes to ensure that commercial products have been transmitted (by radio or television) as agreed for a given time.
In yet other applications, the control signals are transmitted in the form of codes generated in accordance with the present invention. In this type of application, the interactive toy receives and decodes the coded control signal contained in the audio part of the television or radio program, or in the audio recording, and performs the appropriate action. In another application, the protective control codes are included in the sound part of radio or television broadcasts or sound recordings, so that the receiving or reproducing device, by decoding these codes, can perform a protective control for preventive reception and / or playback of the broadcast or recording. Also, control codes may be included in the transmission on cell phones to prevent unauthorized use of the cell phone ID. In yet another application, the codes are included in telephone transmissions to distinguish voice and data transmissions for appropriate controlled selection of the transmission path to avoid loss of transmitted data.
183 307
Various transmitter identification functions can be implemented to confirm the authenticity of military transmissions and voice communications in air transport. Monitoring devices are also considered. In one such application, market research participants wear personal monitoring devices that receive coded information added to public addresses or similar sound signals at wholesalers or sales locations that record attendee presence. According to yet another application, employees wear personal monitoring devices to receive coded information added to the audible signals in the workplace to monitor their presence in the right places.
Secure communication can also be implemented using the present invention. In one such application, underwater secured communication is performed by the coding and decoding assemblies of the present invention, either by assigning levels of code components in such a way that the codes are masked by the underwater sounds of the medium, or by a sound source arising at the place of the transmitter. In another application, secured paging transmissions are made by including masking codes in the radio signal transmission that are received and decoded by a paging device.
The encoding and decoding methods of the present invention can also be used to confirm the authenticity of a voice. For example, on a telephone, the stored voice recording can be compared to a live voice. In another application example, data such as the security number and / or time of day can be coded and combined with a voice statement and then decoded and used for automatic control of the voice processing of the speech. The control device in this arrangement may either be attached to a telephone or other voice communication device, or may be a separate unit used when the voice statement is to be stored directly, without sending by telephone lines or otherwise. A further application is to provide an authenticity code to the memory of a mobile telephone such that the voice stream includes an authenticity code, thereby enabling detection of unauthorized transmission.
It is also possible to achieve better use of the communication channel bandwidth by including data in voice or other audio transmissions. In this type of application, data readings performed by aircraft equipment are included in the aircraft-to-ground voice transmission to inform ground staff about the operation of these devices, without requiring separate voice or data channels.
Cassette piracy, unauthorized copying of copyrighted works, such as audio / video recordings and music, can also be detected by encoding a unique identification number in the sound part of each authorized copy using the encoding method of the present invention. If the encoded identification number is encoded in multiple copies, unauthorized copying becomes evident.
Further application designates programs that have been recorded using the VCR, comprising a decoder according to the present invention. Video programs (such as entertainment, commercial programs) are encoded in accordance with the present invention by an identification code identifying the program. When the VCR is set in recording mode, the recorded audio parts of the signals are delivered to the decoder for detecting identification codes. The detected codes are stored in the VCR memory for later use when generating the recording report.
Data identifying copyrighted work that has been posted by the station or otherwise transmitted by the manufacturer may be retained using the present invention to ensure the implementation of protection rights. The works are coded with appropriate identification codes that uniquely identify them. A monitoring unit that receives broadcast signals or is otherwise transmitted by one or more stations or broadcasters provides the audio part of the signal to a decoder according to the invention which detects the identification codes present therein.
183 307
The detected codes are stored in memory for use to generate a report that can be used to determine compliance with copyright.
Proposed decoders that comply with the Motion Picture Experts Group (MPEG) 2 standard already use some of the processing components with the acoustic expansion required to extract the encoded data in accordance with the present invention, so that recording prevention using means (e.g., preventing unauthorized recording of copyrighted works) using the codes according to the present invention are well suited for MPEG 2 decoders. A suitable decoder according to the present invention is placed in the recording device, or in addition to it, and detects the presence of a forbidden copy code in the sound provided from the recording. The recording device responds when such a code is detected by preventing the corresponding sound signal and any accompanying signals such as video signals from being recorded. The copyright information encoded according to the present invention is within the band and does not require separate timing or synchronization, and naturally accompanies the audio material.
In yet other applications, programs transmitted by air or other means, or programs recorded on tape, disc or other means, contain parts of the sound encoded with control signals for use by one or more devices controlled by the listener or viewer. For example, a program describing a route that a cyclist could travel may include a sound portion coded according to the invention with control signals used by a stationary exercise bike to control pedal resistance or pedal travel in accordance with the visible slope of the described route. When the user pedals on a stationary bike, he sees the program on a TV screen or any other screen, and the sound part of the program is played as sound. A microphone in a stationary bicycle captures the reproduced sound and the decoder according to the present invention detects the control signals contained in the sound, thus ensuring control of the resistance of the exercise bike pedals.
It follows from the above that the methods of the invention can be implemented in whole or in part using analog or digital circuits, and that all or part of the signal processing functions can be performed by circuits with built-in control circuits or using digital signal processors, microprocessors, microcomputers, multiprocessor systems (e.g., parallel processors), or the like.
183 307
<img file="PL183307B1_D0001.tif" />
<\ ι k
183 307
<img file="PL183307B1_D0002.tif" />
FfG.3
183 307
ΓΟ · <Ο
CT3
WHAT
WHAT
ŁO
ΓΟ
From
From ο
From σ>
οο what
LO ro ου <ο
ΙΟ σ>
from
From
From what
From
UT5
From «3 from? 3
From
From
From
2L "Ο ~ ο
AT"
N
-sc
ΓΟ s?
OO ro ro ro ro
ŁO ro <3ro ro ro
From Ro
From
F / G4
WHAT
183 307
<img file="PL183307B1_D0003.tif" />
183 307
<img file="PL183307B1_D0004.tif" />
CO t_l_l -tSL
183 307
<img file="PL183307B1_D0005.tif" />
183 307
FIG. 7B
<img file="PL183307B1_D0006.tif" />
183 307
<img file="PL183307B1_D0007.tif" />
KONIECPOPPRocąpuRy
194
183 307
<img file="PL183307B1_D0008.tif" />
183 307
<img file="PL183307B1_D0009.tif" />
no (ΖΚΟ & ΙΟΝΕ'Ϊ --— \
764
183 307
FIG. 7F
<img file="PL183307B1_D0010.tif" />
183 307
<img file="PL183307B1_D0011.tif" />
183 307
Ο τΟ «Λ
<img file="PL183307B1_D0012.tif" />
Ob
<img file="PL183307B1_D0013.tif" />
ό
183 307
<img file="PL183307B1_D0014.tif" />
<img file="PL183307B1_D0015.tif" />
183 307
<img file="PL183307B1_D0016.tif" />
CD
C \ J
183 307
<img file="PL183307B1_D0017.tif" />
183 307
FIG. 12B
<img file="PL183307B1_D0018.tif" />
183 307
<img file="PL183307B1_D0019.tif" />
<img file="PL183307B1_D0020.tif" />
183 307
FIG. 1
<img file="PL183307B1_D0021.tif" />
θ oo
F
183 307
<img file="PL183307B1_D0022.tif" />
183 307
OJ
<img file="PL183307B1_D0023.tif" />
<img file="PL183307B1_D0024.tif" />
183 307
F / G / 7
<img file="PL183307B1_D0025.tif" />
<img file="PL183307B1_D0026.tif" />
183 307
<img file="PL183307B1_D0027.tif" />
FIG. /
UP Department of Publications. Circulation of 70 copies Price PLN 6.00.
Contents17
23 sheets
Sheet 1 Sheet 2 Sheet 3 Sheet 4 Sheet 5 Sheet 6 Sheet 7 Sheet 8 Sheet 9 Sheet 10 Sheet 11 Sheet 12 Sheet 13 Sheet 14 Sheet 15 Sheet 16 Sheet 17 Sheet 18 Sheet 19 Sheet 20 Sheet 21 Sheet 22 Sheet 23
120 members in 26 offices
Priority claims12
| Document | Office | Kind | Date |
|---|---|---|---|
| 22101994 | United States of America | A | |
| 22101994 | United States of America | A | |
| 40801095 | United States of America | A | |
| 40801095 | United States of America | A | |
| 9503797 | United States of America | W | |
| 9503797 | United States of America | W | |
| 94221019 | – | – | – |
| 95408010 | – | – | – |
| 95US9503797 | – | – | – |
| US19940221019 | – | – | – |
| US19950408010 | – | – | – |
| WO1995US03797 | – | – | – |
Members120
| Document | Office | Kind | |
|---|---|---|---|
| IL113190D0 | Israel | D0 | |
| US5450490A | United States of America | A | |
| CA2185790A1 | Canada | A1 | |
| WO9527349A1 | World Intellectual Property Organization (WIPO) | A1 | |
| AU2196995A | Australia | A | |
| NO964062D0 | Norway | D0 | |
| SE9603570D0 | Sweden | D0 | |
| GB9620181D0 | United Kingdom | D0 | |
| DK105996A | Denmark | A | |
| NO964062L | Norway | L | |
| HU9602628D0 | Hungary | D0 | |
| FI963827A | Finland | A | |
| FI963827L | Finland | L | |
| SE9603570L | Sweden | L | |
| GB2302000A | United Kingdom | A | |
| GB2302000A8 | United Kingdom | A8 | |
| EP0753226A1 | European Patent Office (EPO) | A1 | |
| PL316631A1 | Poland | A1 | |
| LU88820A1 | Luxembourg | A1 | |
| DE19581594T1 | Germany | T1 | |
| CZ284096A3 | Czechia | A3 | |
| CN1149366A | China | A | |
| KR970702635A | Republic of Korea | A | |
| MX9604464A | Mexico | A | |
| BR9507230A | Brazil | A | |
| HUT76453A | Hungary | A | |
| JPH10500263A | Japan | A | |
| US5764763A | United States of America | A | |
| NZ283612A | New Zealand | A | |
| GB9818342D0 | United Kingdom | D0 | |
| GB9818347D0 | United Kingdom | D0 | |
| GB9818349D0 | United Kingdom | D0 | |
| GB9818352D0 | United Kingdom | D0 | |
| GB9818353D0 | United Kingdom | D0 | |
| GB9818354D0 | United Kingdom | D0 | |
| GB9818355D0 | United Kingdom | D0 | |
| GB2325826A | United Kingdom | A | |
| GB2325827A | United Kingdom | A | |
| GB2325828A | United Kingdom | A | |
| GB2325829A | United Kingdom | A | |
| GB2325830A | United Kingdom | A | |
| GB2325831A | United Kingdom | A | |
| GB2325832A | United Kingdom | A | |
| GB9823987D0 | United Kingdom | D0 | |
| GB2302000B | United Kingdom | B | |
| GB2302000B8 | United Kingdom | B8 | |
| GB2325826B | United Kingdom | B | |
| GB2325827B | United Kingdom | B | |
| GB2325828B | United Kingdom | B | |
| GB2325829B | United Kingdom | B | |
| GB2325830B | United Kingdom | B | |
| GB2325831B | United Kingdom | B | |
| GB2325832B | United Kingdom | B | |
| GB2327582A | United Kingdom | A | |
| GB2327582B | United Kingdom | B | |
| AU709873B2 | Australia | B2 | |
| PL177808B1 | Poland | B1 | |
| AU6442299A | Australia | A | |
| IL113190A | Israel | A | |
| NZ331166A | New Zealand | A | |
| EP0753226A4 | European Patent Office (EPO) | A4 | |
| HU0004765D0 | Hungary | D0 | |
| HU0004766D0 | Hungary | D0 | |
| HU0004767D0 | Hungary | D0 | |
| HU0004768D0 | Hungary | D0 | |
| HU0004769D0 | Hungary | D0 | |
| HU0004770D0 | Hungary | D0 | |
| PL180441B1 | Poland | B1 | |
| HU219256B | Hungary | B | |
| IL133700D0 | Israel | D0 | |
| IL133701D0 | Israel | D0 | |
| IL133702D0 | Israel | D0 | |
| IL133703D0 | Israel | D0 | |
| IL133704D0 | Israel | D0 | |
| IL133705D0 | Israel | D0 | |
| IL133706D0 | Israel | D0 | |
| IL133707D0 | Israel | D0 | |
| HU219627B | Hungary | B | |
| HU219628B | Hungary | B | |
| CZ288497B6 | Czechia | B6 | |
| HU219667B | Hungary | B | |
| HU219668B | Hungary | B | |
| NZ502630A | New Zealand | A | |
| ATA902795A | Austria | A | |
| PL183307B1This record | Poland | B1 | |
| PL183573B1 | Poland | B1 | |
| US6421445B1 | United States of America | B1 | |
| AT410047B | Austria | B | |
| SE519882C2 | Sweden | C2 | |
| US2003081781A1 | United States of America | A1 | |
| AU763243B2 | Australia | B2 | |
| IL133702A | Israel | A | |
| IL133703A | Israel | A | |
| IL133700A | Israel | A | |
| IL133704A | Israel | A | |
| IL133706A | Israel | A | |
| IL133707A | Israel | A | |
| IL133701A | Israel | A | |
| IL133705A | Israel | A | |
| PL187110B1 | Poland | B1 |
1 legal event, as the office reported them to INPADOC
Events
| Event | Code | |
|---|---|---|
| Decisions on the lapse of the protection rightsLapsedLAPS | LAPS |
Numbers
- Publication, DOCDB
- 183307
- Publication, EPODOC
- PL183307B
- Application
- 95333767
- Application, DOCDB
- 33376795
- Application, EPODOC
- PL19950333767
Titles2
- English
- AUDIO SIGNAL ENCODING SYSTEM
- Polish
- System kodowania sygnału dźwiękowego
Classification
- CPC, 14
- H04H20/31
- H04H20/14
- H04H60/13
- H04H60/17
- H04H60/37
- H04H60/40
- H04H60/44
- H04H60/45
- H04H60/58
- H04H60/63
- H04H60/66
- H04K1/02
- H04L27/10
- H04L27/30
- IPC, 16
- G10L19 00
- G10L19 018
- H04N5 38
- G11B20 10
- H04H20 31
- H04H60 13
- H04H60 17
- H04H60 37
- H04H60 40
- H04H60 44
- H04H60 45
- H04H60 58
- H04H60 63
- H04H60 66
- H04M11 06
- H04N5 60