Method of and apparatus for code detecting
Abstract
Apparatus and methods for including a code (68) having at least one code frequency component in an audio signal (60) are provided. The abilities of various frequency components in the audio signal to mask the code frequency component to human hearing are evaluated (64), and based on these evaluations an amplitude (76) is assigned to the code frequency component. Methods and apparatus for detecting a code in an encoded audio signal are also provided. A code frequency component in the encoded audio signal is detected based on an expected code amplitude or on a noise amplitude within a range of audio frequencies including the frequency of the code component.

Term
Term ended
Expired 27 March 2015, 11.5 years ago.
- Priority
- Filed
- Granted
- Expired
- Today
6 claims: 4 independent, 2 dependent
- 1Zastrzeżenia patentowe 1. Sposób detekcji kodu w zakodowanym sygnale dźwiękowym, zawierającym liczne składowe częstotliwości sygnału dźwiękowego i przynajmniej jedną składową częstotliwości kodu o dobranej amplitudzie i częstotliwości dźwięku do maskowania składowej częstotliwości kodu, ze względu na słyszalność ludzkiego ucha, przez przynajmniej jedną z licznych składowych częstotliwości sygnału dźwiękowego, znamienny tym, że w kolejnych etapach ustala się spodziewaną amplitudę kodu przynajmniej jednej składowej częstotliwości kodu na podśtawie zakodowanego sygnału dźwiękowego i wykrywa się składową częstotliwość kodu w zakodowanym sygnale dźwiękowym na podstawie spodziewanej amplitudy kodu.
- 2Sposób detekcji kodu w zakodowanym sygnale dźwiękowym, zawierającym liczne składowe częstotliwości sygnału dźwiękowego i przynajmniej jedną składową częstotliwości kodu posiadającą założoną amplitudę i założoną częstotliwość dźwięku, do wyróżniania przynajmniej jednej składowej częstotliwości kodu licznych składowych częstotliwości sygnału dźwiękowego, znamienny tym, że w kolejnych etapach wyznacza się amplitudę składowej częstotliwości zakodowanego sygnału dźwiękowego w pierwszym zakresie częstotliwości dźwięku, zawierającym założoną częstotliwość dźwięku przynajmniej jednej składowej częstotliwości kodu, ustala się amplitudę szumu dla pierwszego zakresu częstotliwości dźwięku i wykrywa się obecność przynajmniej jednej składowej częstotliwości kodu w pierwszym zakresie częstotliwości dźwięku na podstawie ustalonej amplitudy jego szumu i wyznaczonej amplitudy zawartych w nim składowych częstotliwości.
- 3Urządzenie do detekcji kodu w zakodowanym sygnale dźwiękowym, maskowanego i niesłyszalnego przez ludzkie ucho, znamienne tym, że zawiera elektroniczny analizator sygnału, ustalający spodziewaną amplitudę kodu na podstawie zakodowanego sygnału dźwiękowego oraz połączony z analizatorem elektroniczny detektor kodu, wykrywającego składowe częstotliwości kodu w zakodowanym sygnale dźwiękowym na podstawie spodziewanej amplitudy kodu ustalonej przez analizator.
- 4Urządzenie według zastrz. 3, znamienne tym, że zawiera detektor określonej częstotliwości składowej częstotliwości kodu w zakodowanym sygnale dźwiękowym dla określonej częstotliwości, przy czym elektroniczny analizator sygnału ustalający spodziewaną amplitudę kodu jest zaopatrzony w komparator amplitudy wykrytej składowej częstotliwości kodu i spodziewanej amplitudy kodu.
- 5Urządzenie według zastrz. 4, znamienne tym, że detektor określonej częstotliwości stanowi separator rozdzielający zakodowany sygnał dźwiękowy na grupy składowych częstotliwości, z których każda zawiera przynajmniej jedną składową częstotliwości z odpowiedniego zakresu częstotliwości, gdzie pierwsza z grup składowych częstotliwości posiada odpowiedni zakres częstotliwości zawierający częstotliwość dźwięku przynajmniej jednej składowej częstotliwości kodu.
- 6Urządzenie do detekcji kodu w zakodowanym sygnale dźwiękowym, znamienne tym, że zawiera elektroniczny analizator sygnału, wyznaczający spodziewaną amplitudę składowej częstotliwości kodu w zakodowanym sygnale dźwiękowym w pierwszym zakresie częstotliwości, zawierającym składową częstotliwość kodu, przelicznik szumu, wyznaczający amplitudę szumu w pierwszym zakresie częstotliwości, oraz połączony z analizatorem sygnału i przelicznikiem szumu detektor elektroniczny, wykrywający obecność składowej częstotliwości kodu na podstawie amplitudy szumu wyliczonej przez przelicznik szumu i amplitudy składowej częstotliwości kodu wyznaczonej przez analizator w pierwszym zakresie częstotliwości. 180 441
Independent claims6
287 paragraphs in 18 sections, as filed
The present invention relates to a method and apparatus for detecting a code used to encode an audio signal.
Over the years, many ways have been proposed for mixing codes with audio signals in such a way that the codes can be reliably reproduced from the audio signals while remaining inaudible when the audio signals are processed into sound. Achieving both of these goals is crucial for practical applications. For example, broadcasters and program producers, as well as recorders of music for public distribution, would not tolerate the inclusion of audible codes on programs or recordings.
Methods of encoding audio signals have been proposed many times, beginning at least with US Patent 3,004,104 to Hembroke on October 10, 1961. Hembroke has shown a method of encoding in which the energy of an audio signal in a narrow frequency band has been selectively removed to encode the signal. The problem with this method arises when noise or signal interference re-injects energy into a narrow frequency band whereby the code is obscured.
In accordance with another method, described in U.S. Patent No. 3,845,391 to Crosby, it has been proposed to eliminate a narrow frequency band from the audio signal and include a code there. This method has the same disadvantage as the previous solution as pointed out in US Patent No. 4,703,476 to Howard, which is related to the Crosby patent. However, Howard's patent was only intended to improve Crosby's solution without going beyond its fundamental tenets.
It has also been proposed to encode binary signals in frequencies passing through the entire audio band. The problem with this solution is that in the absence of audio signal components to mask the code frequency, it may become audible. This method assumes that the codes are noise-like in nature, which suggests that their presence will be ignored by listeners. However, in many cases this assumption is not true, for example in the case of classical music with relatively little audio signal content, or in the event of a speech break.
Another method has been proposed whereby two-tone multi-frequency (DTMF) codes are inserted into an audio signal. The meaning of the DTMF code is detectable by its frequency and duration. However, the components of the audio signals may be confused for one or both tones of each DTMF code, whereby either the presence of the code may be noticed by the detector or the signal components may be misread as elements of the DTMF code. Moreover, it should be stated that the DTMF code has a common tone with other DTMF codes. As a result, a signal component corresponding to a tone of another DTMF code may be mixed with a tone of a DTMF code that is simultaneously present in the signal, leading to erroneous detection.
A method of detecting a code in an encoded audio signal, comprising a plurality of frequency components of the audio signal and at least one code frequency component with the amplitude and audio frequency selected for masking the code frequency component for the audibility of the human ear through at least one of the multiple frequency components of the audio signal, according to the invention, it is distinguished in that in the following steps an expected code amplitude of at least one code frequency component is determined from the encoded audio signal and the code frequency component of the encoded audio signal is detected from the code amplitude.
A detection method in an encoded audio signal comprising a plurality of frequency components of an audio signal and at least one code frequency component having a predetermined amplitude and predetermined audio frequency, for distinguishing at least one code frequency component from a plurality of frequency components of the audio signal, according to the invention is distinguished by the fact that that in successive steps the amplitude of the frequency component of the encoded audio signal in the first audio frequency range containing the assumed audio frequency of at least one code frequency component is determined, the noise amplitude for the first audio frequency range is determined and the presence of at least one frequency component is detected
180 441 of the code in the first audio frequency range on the basis of the determined amplitude of its noise and the determined amplitude of the frequency components contained therein.
The device for detecting a code in an encoded audio signal, masked and inaudible by the human ear, according to the invention is distinguished by the fact that it comprises an electronic signal analyzer that determines the expected amplitude of a code on the basis of the encoded audio signal, and an electronic code detector connected to the analyzer detects the frequency components of the code in coded audio signal based on the expected code amplitude determined by the analyzer.
The device preferably comprises a detector of a predetermined frequency of the code frequency component in the coded audio signal for a predetermined frequency, the electronic signal analyzer for determining the expected code amplitude preferably being provided with an amplitude comparator of the detected code frequency component and the expected code amplitude.
The specific frequency detector is preferably a separator separating the encoded audio signal into groups of frequency components, each of which comprises at least one frequency component from a respective frequency range, the first of the frequency component groups having a corresponding frequency range containing the audio frequency of at least one code frequency component.
The apparatus for detecting a code in an encoded audio signal according to the invention is characterized in that it comprises an electronic signal analyzer for determining an expected amplitude of the code frequency component in an encoded audio signal in a first frequency range, including a code frequency component, a noise converter, determining the amplitude of the noise in the first frequency range , and an electronic detector connected to the signal analyzer and noise converter, detecting the presence of the code frequency component on the basis of the noise amplitude calculated by the noise converter and the amplitude of the code frequency component determined by the analyzer in the first frequency range.
The solution according to the invention makes it possible to overcome the disadvantages of the solutions proposed so far. The inventive method and device for decoding make it possible to reliably read the codes from the audio signals. The subject of the invention, in an embodiment, has been explained in the drawing, in which Fig. 1 shows a block diagram of an encoder, Fig. 2 - a digital encoder operation diagram, Fig. 3 - block diagram of an encoder used to encode audio signals provided in analog form, fig. 4 - spectral distributions for illustrating frequency component patterns corresponding to different code symbols, after coding according to the embodiment of fig. 3, fig. 5 and 6 - diagrams block for illustration of the operation of the embodiment of fig. 3, fig. 7A to 7C - the network operated by the program routine used in the embodiment of Fig. 3, Figs. 7D to 7E - flowchart to illustrate an alternative program routine used in the embodiment of Fig. 3, Fig. 7F - graph showing linear approximation of the single tone masking relationship, Fig. 8 is a block diagram of an analog encoder, Fig. 9 is a block diagram of a weighting factor determining circuit for the embodiment of Fig. Fig. 8, Fig. 10 is a block diagram of a decoder according to some features of the present invention, Fig. 11 - a block diagram of a decoder according to an embodiment of the present invention using digital signal processing, Figs. 12A and 12B - flowchart describing the operation of the decoder in Fig. 11, Fig. 13 - a block diagram of a decoder according to some embodiments of the present invention, Fig. 14 - a block diagram of an embodiment of an analog decoder according to the present invention, fig. 15 a block diagram of a component detector according to the embodiment of Fig. 14, Figs. 15, 16, and 17; block diagrams of an apparatus incorporated in a system for producing listenability estimates for widely distributed information.
The present invention implements methods for incorporating code into audio signals to optimize the likelihood of accurately reproducing the information in the code in the signal, ensuring that when the audio signal is played back as audio, the code is not heard, even if it falls within an audible frequency range.
180 441
Referring to Fig. 1, a block diagram of an encoder operation is shown. The audio signal to be encoded is taken at the input terminal 30. The audio signal may represent, for example, a radio program, an audio part of a television signal, music or any other type of audio signals reproduced in a similar manner. In addition, the audio signal may be used for private communication such as transmission by telephone or for personal recording of some kind. However, these application examples are mentioned herein for purposes of illustration only and do not limit the field of application of the invention.
In the function block designated by 34 in Fig. 1, the ability of one or more components of an acquired audio signal to mask sounds having frequencies corresponding to a code frequency component or component to be added to the audio signal is calculated. Multiple computations for a single code frequency may be performed, a separate computation may be performed for each of a plurality of code frequencies, a plurality of computations may be performed for each of a plurality of code frequencies, one or more common computations for a plurality of code frequencies may be performed, or a combination of one or more of the above operations used. Each calculation is performed on the basis of the frequency of one or more code components to be masked and the frequency of one or more components of the audio signal whose masking ability is currently being calculated. In addition, if the code component and the audio masking component or components do not fall within substantially the same signal intervals, such that they could be reproduced as sound at significantly different time intervals, the effects of differences in the signal interval between the code component or components that are masked. and masking component or program components are also taken into account.
Advantageously, in some embodiments, multiple calculations are performed for each code component by separately considering the ability of different portions of the audio signal to mask each code component. In one embodiment, the ability of each of the plurality of substantially single tonal components of an audio signal to mask a code component is calculated based on the frequency of the audio signal component, its "amplitude" (as will be defined), and the timing of the code component, such masking is referred to herein as "Tonal masking".
The term "amplitude" is used herein to denote any one or more signal determining values that can be used to estimate masking capacity so as to size the code component to detect its presence in a reproduced signal and for any other purpose. , and may include values such as energy, power, voltage, current and pressure, whether measured in absolute or relative terms, and regardless of whether instantaneous or accumulative values are considered. Accordingly, amplitude may be measured as window mean, arithmetic mean, by integrating the square root of values, by accumulating absolute or relative discrete values, or by other means.
In other embodiments, in addition to or alternatively to tonal masking estimates, the ability of the audio signal components of a relatively narrow frequency band sufficiently close to the given code component to be masked (referred to herein as "narrowband" masking) is calculated. In still other embodiments, the ability of multiple code components over a relatively wide frequency band to mask the component is calculated. Depending on the necessity or possibility, the capabilities of the components of the sound program in the signal intervals preceding or following a given one or more components are calculated for its masking in a non-simultaneous manner. This method of estimation is particularly useful when the components of the audio signal in a given signal interval are of insufficient amplitude to enable the inclusion of code components with sufficiently large amplitudes in the same interval, which would make them distinguishable from noise.
Advantageously, a combination of two or more tonal masking capabilities, narrowband masking and broadband masking capabilities (and, if
180 441 necessary or appropriate, non-simultaneous masking capabilities) are computed for multiple code components. When the code components are close enough in the frequency domain, separate computations for each component need not be performed.
In some other preferred embodiments, a sliding tonal analysis is performed in place of separate tonal, narrow, or broadband analyzes, without the need to classify the audio program as tonal, narrowband, or wideband.
Advantageously, when a combination of masking capacity is computed, each computation provides the maximum allowable amplitude for one or more code components such that by comparing all computations that have been performed that pertain to a given component, the maximum amplitude is chosen to ensure that each the component will be masked by the audio signal when it is played back as audio so that it will not be audible to the human ear. By maximizing the amplitude of each component, the probability of detecting its presence based on its amplitude is also increased. Of course, it is not necessary to use the largest possible amplitude, which is only necessary for decoding, in order to be able to distinguish a sufficiently large number of code components from the components of an audio signal or other noise.
The effects of the computation are fed, as indicated at 36 in Fig. 1, to the code generator 40. The code generation can be performed in many different ways. One particularly advantageous method is to assign a single set of code frequency components to each of a plurality of data states or symbols whereby, during a given signal interval, the corresponding data state is represented by the presence of its corresponding set of code frequency components. In this way, the overlap of the detected code with the audio signal components is reduced because, in the preferably high percentage of the signal intervals, a sufficiently large number of code components will be detectable despite the program audio overlap with the other components. Moreover, the process of implementing masking estimates is simplified when the frequencies of the code components are known prior to their generation.
Other forms of encoding may also be implemented. For example, frequency shift keying (FSK), frequency modulation (FM), hopping coding, diffuse spectral coding, or a combination of these methods may be used. Other coding methods that can be used to implement the present invention will be apparent from the description thereof.
The data to be encoded is taken at input 42 of code generator 40, which responds by generating a unique group of code frequency components and assigning an amplitude to each of them based on calculations taken from output 36. The code frequency components thus produced are supplied to the first input of summing circuit 46, which it receives an audio signal to be encoded on its other input. Circuit 46 adds the code frequency components to the audio signal and outputs the encoded audio signal at its output terminal 50. Circuit 46 may be an analog or digital adder, depending on the form of the signals fed to it. The summation may also be implemented in software, and, in this case, a digital processor to perform masking estimates and code production may also be used to sum the code with the audio signal. In one embodiment, the code is provided as time-domain data in digital form, which is then summed across the time-domain audio data. In another example, the audio signal is digitized into the frequency domain and added to a code similarly represented as frequency domain digital data. In most applications, the summed frequency domain data is then converted to time domain data.
It can be seen from the above that masking estimation as well as code production functions can be performed by digital or analog processing, or by a combination thereof. Moreover, although the audio signal may be received in analog form at the output terminal 30 and added to the code components in analog form by the circuit 46 as shown.
180 441 of Figure 1, alternatively, the audio signal may be digitized upon reception, added to digital code components, and output in digital or analog form. For example, when the signal is to be recorded on a CD or digital audio tape, it can be digitized, while if it is to be distributed by traditional radio or television, it can be output in analog form. Various other combinations of analog and digital processing can be implemented.
In some embodiments, code components of only one code symbol at a time are included in the audio signal. However, in other embodiments, components of multiple code symbols are simultaneously incorporated into the audio signal. For example, in some embodiments, components of one symbol occupy one frequency band and components of another symbol simultaneously occupy another frequency band. Alternatively, the components of one symbol may be in the same band of another, or their bands may overlap, as long as the components are distinguishable, for example by assigning substantially different frequencies or frequency ranges.
An embodiment of a digital decoder is shown in Fig. 2. In this embodiment, an analog audio signal is taken at the input terminal 60 and converted to digital form by the A / D converter 62. The digital audio signal is fed to an estimate of the masking, denoted by functional block 64, before which the digital audio signal is split into frequency components, for example by Fast Fourier Transform (FFT), wavelet transformation or other transformations in the domain time to frequency, or by digital filtering. Next, the masking capabilities of the frequency components of the audio signal in the respective frequency bundle are calculated, specifying the tonal masking ability, narrowband masking ability, and wideband masking ability (and, if necessary and appropriate, non-simultaneous masking ability). Alternatively, the masking capabilities of the frequency components of the audio signal in a given frequency bundle are computed with sliding tonal analysis.
The encoded data is taken from input terminal 68 and, for each data state corresponding to a given signal interval, a corresponding group of code components is produced as denoted by functional signal generation block 72 and leveled as indicated by function block 76 which also takes appropriate masking estimates. Signal generation may be implemented, for example, with a dictionary table storing each of the code components as time-domain data, or by interpolating the stored data. The code components may either be stored permanently or generated upon initialization of the system of FIG. 2 and then placed in a memory, such as RAM, to be discharged in response to data retrieved from terminal 68. Component values can also be computed while they are generated.
A level determination is performed for each of the code components based on the respective masking estimates described above, and the code components whose amplitude has been determined to be inaudible are added to the digital audio signal as indicated by the summation symbol 80. Depending on the amount of time required to perform the above operations, it may be advantageous to delay the digital audio signal, indicated by 82, by temporarily storing it in memory. If the audio signal is not delayed, after performing the FFT and estimating the masking for the first interval of the audio signal, the code components with a predetermined amplitude are added to the second interval of the audio signal following the first interval. If the audio signal is delayed, code components with a predetermined amplitude can be added to the first interval and a masking estimate can be used simultaneously. Moreover, if a portion of the audio signal at the time of the first interval has a greater ability to mask the code component added during the second interval than the portion. sound signal during the second interval for the same code component,
180 441 then the amplitude may be assigned to a code component based on the ability to mask a portion of the audio signal in the first interval non-simultaneously. In this way, simultaneous and non-simultaneous masking capabilities can be estimated, and an optimal amplitude can be assigned to each code component based on the most favorable of the estimates.
In certain applications such as radio broadcasts or analog recording (such as for example on a classic tape cassette), the encoded digital audio signal is converted into analog form by a digital to analog converter (DAC) 84. However, when the signal is to be transmitted or digitized, the DAC 84 may be omitted.
The various functions shown in Fig. 2 may be implemented by, for example, a digital signal processor or a personal computer, workstation or large computer system, or other digital computer.
Figure 3 shows a block diagram of an encoder system for coding audio signals provided in analog form, such as classic radio broadcasts. In the system of Fig. 3, a host processor 90, which may be a personal computer, for example, manages the selection and generation of information to be encoded and incorporated into an analog audio signal taken from input terminal 94. The main processor 90 is coupled to a keyboard 96 and a monitor 100, such as a CRT monitor, such that a user can select the desired information to be encoded by selecting from a menu of available information displayed on the monitor 100. Typical information to be encoded in a radio broadcast signal may include identifying information. channel or station, program or information segment, and / or time code.
Once the requested information has been entered into main processor 90, the processor outputs data representing the information symbols to the digital signal processor (DSP) 104, which in turn encodes each symbol obtained from main processor 90 into a unique set of code signal components, as described below. According to one embodiment, the main processor generates a four-state data stream, that is, a data stream in which each data unit can assume one of four different states, each representing a unique symbol, including synchronization symbols called "E" and "S". ", And two information symbols" 1 "and" 0 "representing the respective binary state. Of course, any number of distinguishable data states may be used. For example, instead of two information symbols, three-state data may be represented by three different symbols, which allows correspondingly more information to be carried in a data stream of a given size.
For example, when the signal relates to speech, it is preferable to transmit the symbol for a relatively long period of time compared to a broadcast having a substantially more continuous energy content, in order to allow natural breaks present in speech to occur. Accordingly, in order to provide a sufficiently large information throughput in this case, the number of possible information symbols can be advantageously increased. For symbols representing up to five bits, signal transmission lengths of 2, 3, and 4 seconds provide increasingly greater probabilities of correct decoding. In some such embodiments, a starting symbol ("E") is decoded when the FFT packet energy for that symbol is greatest when the average energy minus energy standard deviation for that symbol is greater than the average energy plus the average energy standard deviation for all other symbols. , and where the shape of the energy versus time plot is generally bell-shaped, peaking at the intersymbol temporal border.
In the embodiment shown in Fig. 3, when the DSP 104 receives symbols of a given information to be encoded, it responds by generating a unique set of code frequency components for each symbol that it outputs to 106. Referring also to Fig. 4, there are spectral plots for each of the four data symbols S, E, 0, and 1 of the exemplary data set described. As shown in Fig. 4, in this embodiment, the S symbol is represented by a unique group of ten
180 441 code frequency components fl to flO spaced uniformly over the frequency range extending from a frequency value slightly greater than 2 kHz to a frequency value slightly less than 3 kHz. The symbol E is represented by a second unique group of ten code frequency components from f1 to f20, distributed in the frequency spectrum at equal intervals from the first frequency value slightly greater than 2 kHz to the frequency value slightly less than 3 kHz, where each of the code components is from Η1 to f20 has a unique frequency value different from all the values of the same group as well as from all the frequencies f to flO. The symbol 0 is represented by another unique group of ten code frequency components f21 to f30, also spaced in the frequency spectrum at equal intervals from the first frequency value slightly greater than 2 kHz to the frequency value slightly less than 3 kHz, each of the code components having a unique a frequency value different from all the values of the same group as well as from all the frequencies f to f20. Finally, symbol 1 is represented by another unique group of ten code frequency components f31 to f40 spaced equidistantly in the frequency spectrum from a first frequency value slightly greater than 2 kHz to a frequency value slightly less than 3 kHz, each of the code components f31 to f40 has a unique frequency value different from all the frequencies f to f40. By using a plurality of code frequency components for each data state such that the code components of each state are substantially different from each other in frequency, the presence of noise (such as non-coded components of an audio signal or other noise) in a common detection band with any noise component of a given data state is less likely and less likely to overlap.
In other embodiments, it is preferable to represent the symbols by a plurality of frequency components, such as ten tones or code frequency components that are not uniformly spaced in the frequency domain and that do not have the same offset between the symbols. Avoiding an integral relationship between the code frequencies for a symbol by grouping the tones reduces the effects of inter-frequency rumble or empty rooms, that is, places where echoes reflected from walls affect proper decoding. The following sets of code tone frequency components for the four symbols are provided for void room cancellation, where f to flO represent the corresponding frequency components (in Hertz) of the code for each of the four symbols:
<td></td><td> 0”</td><td> 1” ,5 <sup>1</sup></td><td>"S"</td><td>"E"</td>
<td>fl</td><td> 1046.9</td><td> 1054.7</td><td> 1062.5</td><td> 1070.3</td>
<td>f2</td><td> 1195.3</td><td> 1203.1</td><td> 1179.7</td><td> 1187.5</td>
<td>β</td><td> 1351.6</td><td> 1343.8</td><td> 1335.9</td><td> 1328.1</td>
<td>f4</td><td> 1492.2</td><td> 1484.4</td><td> 1507.8</td><td> 1500.0</td>
<td>f5</td><td> 1656.3</td><td> 1664.1</td><td> 1671.9</td><td> 1679.7</td>
<td>f6</td><td> 1859.4</td><td> 1867.2</td><td> 1843.8</td><td> 1851.6</td>
<td>f7</td><td> 2078.1</td><td> 2070.3</td><td> 2062.5</td><td> 2054.7</td>
<td>f8</td><td> 2296.9</td><td> 2289.1</td><td> 2304.7</td><td> 2312.5</td>
<td>f9</td><td> 2546.9</td><td> 2554.7</td><td> 2562.5</td><td> 2570.3</td>
<td>flO</td><td> 2859.4</td><td> 2867.2</td><td> 2843.8</td><td> 2851.6</td>
Generally speaking, in these examples shown above, the spectral content of the code varies relatively little as the DSP 104 switches its output from any of the states S, E, 0, and 1 to any of the other states. In accordance with one aspect of the present invention
180 441, in certain preferred embodiments, each code frequency component of each symbol is paired with the frequency component of each of the other data states such that the difference therebetween is less than their critical bandwidth. For any pair of pure tones, the critical bandwidth is the frequency range over which the frequency spacing between two tones can vary without substantially increasing the loudness. When the spacing between adjacent tones for each of the data states S, E, 0, and 1, and when each tone of each data state is paired with a corresponding tone of each of the other states such that the frequency difference therebetween is less than critical bandwidth for this pair, substantially no loudness variation will occur as the transition from any of the data states S, E, 1 and 0 to any other data state occurs during playback as audio. Moreover, by minimizing the frequency difference between the code frequency components of each pair, the respective probability of detecting each data state on reception does not substantially depend on the transmission path. Another advantage of pairing components of different data states is that the masking estimate performed for the code component of the first data state will be substantially accurate for the next data state if a handover takes place.
Alternatively, in a non-uniform code tone spacing solution to minimize the void-room effects, it can be seen that the frequencies selected for each of the code frequency components f to fl 0 are grouped around a frequency, e.g. the frequency components for f, f2 and f3 are located adjacent to 1055 Hz, 1180 Hz and 1340 Hz, respectively. In this particular embodiment, the tones are spaced twice as wide as the FFT resolution, e.g., for a resolution of 4 Hz, tones are shown in steps of 8 Hz, and are selected to be in the middle of the frequency range of the FFT pack. Also, the order of the different frequencies that are assigned to the code frequency components f1 to f0 to represent the different symbols 0, 1, S, E varies with each group. For example, the frequencies selected for the components f, f2, and f3 correspond to the symbols (0, 1, S, E), (S, E, 0, 1), and (E, S, 1, 0), from the lowest to the highest, respectively. that is (1046.9, 1054.7, 1062.5, 1070.2), (1179.7, 1187.5, 1195.3, 1203.1), (1328.1, 1335.9, 1343.8, 1351.6). The advantage of this is that even if there is empty space that interferes with the corresponding reception of a code component, generally the same tone is eliminated from each symbol, making it easier to decode the symbol from the other components. In contrast, if empty space eliminates a component from one of the symbols but not the others, it is more difficult to correctly decode the symbol.
It can be stated that more or less than 4 symbols may be used for encoding. Moreover, each data state or symbol may be represented by fewer or more than ten code tones, and while it is preferable to represent each data state by the same number of tones, it is not necessary for all applications that the number of tones used to represent each data state is same. Advantageously, each of the code tones differs in frequency form from each of the other code tones to increase the probability of distinguishing each data state at the time of decoding. However, it is not necessary in all applications that no code tone frequency be common to two or more data states.
Figure 5 is a flowchart to which reference is made in explaining the coding operation performed by the embodiment of Fig. 3. As mentioned above, the DSP 104 receives data from the main processor 90 sending a series of data states to the DSP 104 as code frequency components. Advantageously, the DSP 104 generates a dictionary table of time domain representation of each of the code frequency components f1 through f40, which are then stored in its RAM, represented as memory 110 in Fig. 5. In response to data received from main processor 90, The DSP 104 generates a corresponding address that is used as an input address to memory 110 as indicated by 112 in FIG. 5, which causes memory 110 to output time-domain data for each of the ten frequency components corresponding to the data state to be discharged at a given time.
180 441
Also referring to Fig. 6, which is a functional block diagram illustrating some operations performed by the DSP 104, memory 110 stores sequences of time-domain values for each of the frequency components for each of the symbols S, E, 0, and 1. In this particular embodiment, if the code frequency components are in the range of 2kHz to about 3kHz, a sufficiently large number of time-domain samples are stored in memory 110 for each of the frequency components f1 through f40, whereby they can be withdrawn from a frequency greater than the Nyquist frequency of the code component with the highest frequency. The time-domain code components are removed at a suitably high frequency from the memory 110, which stores the time-domain components for each of the code frequency components representing assumed durations, such that the time-domain components (n) are stored for each of the code frequency components from f1 to f40 for the (n) ranges t1 to tn as shown in Fig. 6. For example, if an S symbol is to be coded during a given signal interval, during the first interval t1, memory 110 removes the time domain components f1 to f0 corresponding to that interval. In the next time slot, the time domain components f1 to f0 for the interval t2 are removed from the memory 110. This process continues sequentially for the intervals t3 to tn and back from t1 until the end of the coded symbol S.
In some embodiments, instead of draining all ten code components, i.e., f1 to f0, in a time interval, only those components lying in the critical band of the audio signal tones are drained. This is essentially a conservative approach to making code components inaudible.
Referring again to Fig. 5, the DSP 104 also serves to determine the amplitudes of the time domain components output from the memory 110 such that when the code frequency components are played back as audio, they will be masked by the audio signal components into which they were incorporated so that they remain. inaudible to the human ear. In addition, the audio signal received from the input terminal 94 after appropriate filtering and analog-to-digital conversion is also fed to the DSP 104. More specifically, the encoder of Fig. 3 includes an analog bandpass filter 120 that serves to substantially remove the frequency components of the audio signal from the out-of-range band to calculate the masking capacity of the received audio signal, which in the present embodiment is in the range of about 1.5 kHz to about 3.2 kHz. The filter 120 also serves to remove high frequency components from the audio signal that may cause aliasing when the signal is sequentially digitized by an analog-to-digital (A / D) converter 124 operating at a sufficiently high sampling rate.
As shown in Fig. 3, the digital audio signal is provided by the A / D converter 124 to the DSP 104, where, as indicated by 130 in Fig. 5, the program audio signal is subjected to a frequency range split. In this particular embodiment, the frequency range division is performed as a Fast Fourier Transform (FFT) which is performed periodically with or without temporal overlap to produce successive frequency bins each having a predetermined frequency width. Other methods of segregating the frequency components of the audio signals are also available, such as the wavelet transform, the discrete Walsh - Hadamard transform, the discrete Hadamard transform, the discrete cosine transform, as well as numerous filtering methods.
After the DSP 104 has divided the chisel frequencies of the digital audio signal into successive frequency packets, as mentioned above, it proceeds to estimate the ability of the various frequency components present in the audio signal to mask the different code components output by the memory 110, and to produce appropriate amplitude coefficient settings. which are used to determine the amplitudes of different components of the code frequency, so that they will be masked by the program's audio signal when played back as sound, so that they will not be heard by the human ear. These operations are represented by block 134 in FIG. 5.
For audio signal components that are substantially simultaneous with the code frequency components to mask (but that precede the frequency components
180 441 code for a short period of time), the ability to mask the components of an audio program is estimated on a tonal basis as well as a narrow band masking basis and a wide band masking basis as described below. For each of the code frequency components that are output at a given time from the memory 110, the tonal masking capability is computed for each of a plurality of frequency components of the audio signal based on the energy level in each of the respective parcels to which the components fall as well as the frequency dependencies of each packet and the corresponding code frequency component. The estimation in each case (tonal, narrowband, wideband masking) may be in the form of an amplitude finding factor or other measurement allowing the amplitude of the code component to be assigned such that the code component will be masked by the audio signal. Alternatively, the estimation may be a sliding tonal analysis.
In the case of narrowband masking, in this embodiment, a frequency energy content below a predetermined level in a predetermined frequency band is calculated for each respective code frequency component, including computing the corresponding code frequency component to obtain a separate masking capability estimate. In some implementations narrowband masking capability is measured from the energy content of the frequency components of the signal below the average packet energy level in a predetermined frequency band. In this implementation, the component energy levels below the pack's average energy (which is the energy threshold) are summed to produce a narrowband energy level whereby appropriate code components are identified. Instead, a different narrowband energy level may be produced by selecting a threshold component other than the average energy level. Moreover, in accordance with other embodiments, the average energy level of all components of the audio signal may be used as the narrowband energy level for assigning a narrowband masking estimate to the corresponding code component. In accordance with still other embodiments, the tonal energy content of the audio components in a predetermined frequency band is used for this purpose, while in other embodiments the level of the minimum component is used in a predetermined frequency band.
Finally, in some implementations the wideband energy content of the audio signal is determined to calculate the ability of the audio signal to mask the corresponding frequency component of the code by wideband masking. In this embodiment, the wideband masking estimation is based on the minimum narrowband energy level determined in the narrowband masking estimation time described above. That is, if four separate predetermined frequency bands have been tested in the narrowband masking estimation as described above, and broadband noise is included in the minimum narrowband energy level of all four predicted (designated) frequency bands, then this minimum narrowband energy level is multiplied by a factor equal to the ratio of the frequency range to the width of the predetermined frequency band having the minimum narrowband energy level. The result indicates the acceptable overall code strength level. If the total allowable code power level is marked with P and the code has ten code components, each of them is then assigned an amplitude fix factor to result in a component power level that is 10 dB less than P. Alternatively, the broadband noise is calculated to assume a relatively wide bandwidth covering the code components by selecting one of the methods described above for estimating the narrowband energy level but using the audio signal components from the entire assumed relatively wide band. After the wideband noise has been determined in the selected manner, an appropriate wideband masking estimate is assigned to each appropriate code component.
The amplitude finding factor for each code frequency component is then selected on the basis that one of the total, wideband, or narrowband masking estimates gives the highest acceptable amplitude level for the respective
180 441 component. This increases the probability that any relevant component of the code frequency will be masked so that it remains inaudible to the human ear.
The amplitude settling factors are selected for each of the tonal, narrowband, and wideband masking based on the following factors and conditions. In the case of tonal masking, the coefficients are assigned based on the frequencies of the audio components whose masking abilities are estimated, and the frequency, one or more, of the masked code component. Moreover, a given audible signal in any selected interval provides the ability to mask a given code component in the same interval (i.e. simultaneous masking) with a maximum level greater than that at which the same audio signal is in a state of masking the same code component occurring before or after the selected interval (i.e. non-simultaneous masking). The conditions under which the coded audio signal will be heard appropriately by listeners are also taken into account. For example, if the sound of a television signal is to be encoded, the distorting effects of a typical listening environment are advantageously taken into account if, in such circumstances, some frequencies are more distorted than others. Reception and playback equipment (such as graphic equalizers) may use similar effects. Environmental or equipment effects can be compensated by selecting sufficiently low amplitude setting factors to provide masking under the expected conditions.
In some embodiments, only one of the tonal narrowband or wideband capabilities is estimated. In other embodiments, two of these different types of masking capabilities are estimated, and in still other embodiments, all three are used.
In some embodiments, a sliding tonal analysis is performed to estimate the masking ability of an audio signal. The sliding tonal analysis essentially follows the masking rules for narrowband, broadband, and single-tone noise without having to classify the sound. In sliding tonal analysis, an audio signal is considered as a set of discrete tones, each centered in a respective FFT frequency bundle. Generally, the sliding tonal analysis first computes the strength of the audio signal in each FFT bin. Then, for each code tone, the masking effects of the discrete tones of the audio signal in each FFT frequency batch divided by no more than the critical bandwidth of the audio tone are computed from the audio signal strength in each such batch, using the masking relationship for single-tone masking. The masking effects of all corresponding discrete tones from the audio signal are summed for each code tone, then set for the number of tones in the critical tone band of the audio signal and the complexity of the audio signal. As explained below, in some embodiments, program material complexity is empirically determined from the ratio of the power in the respective tones of the audio signal and the square root of the power in those tones of the audio signal. Complexity is used to account for the fact that narrowband noise and wideband noise each provide much better masking effects than those obtained by simply summing tones used to model narrowband and broadband noise.
In some embodiments that use sliding tonal analysis, a predetermined number of audio samples is first subjected to a high FFT, which provides high resolution but requires more processing time. Then successive aliquots of the predetermined number of samples are subjected to a relatively smaller FFT, which is faster but provides lower resolution. The amplitude coefficients determined from the large FFT are combined with those determined from the smaller FFT, which essentially corresponds to time weighting of the higher "frequency precision" of the large FFT over the greater "time precision" of the smaller FFT.
In the embodiment shown in Fig. 5, after selecting a suitable amplitude fix factor for each of the code frequency components output from the memory 110, the DSP 104 determines the amplitude of each of the frequency components respectively as indicated by the amplitude determination function block 114. In other examples,
180 441 implementations, each code frequency component is initially generated such that its amplitude matches its corresponding fixation factor. Referring to fig. 6, the amplitude determining operation performed by the DSP 104 in this embodiment leads to the multiplication of the selected ten values from the code frequency in the time domain f1 to f40 for the current time interval f1 to tn by the respective fixation factors GAI to GA10, and then the DSP 104 adds time-domain components at a predetermined amplitude to produce an overall code signal that is output to output 106. Referring to Fig. 3 and 5, the total code signal is converted by the digital-to-analog converter (DAC) 140 and fed to the first input of adder 142. Adder 142 receives an audio signal from input terminal 94 on the second input and adds the total analog code signal to the analog audio signal and brings it to end 146.
In a radio broadcast application, the encoded audio signal modulates the carrier wave and is transmitted over the air. In NTSC television, the frequency of the encoded audio signal modulates the subcarrier and is mixed with a component of the video signal such that the combined signal is used for carrier modulation when transmitted over the air. Classic TV and radio signals can of course also be transmitted via cable (for example classic or optical fiber), via satellite, or otherwise. In other applications, the encoded audio can be recorded either for pre-recorded distribution or for later distribution or for some other wide distribution method. The encoded audio can also be used for point-to-point broadcasts. Various other uses, transmission methods and recording methods are of course possible.
Figures 7A through 7C show a flowchart showing the course of program procedures performed by the DSP 104 to implement the tonal, narrowband and wideband estimates of the functions described above. Figure 7A illustrates the main loop of the DSP program 104. The program is initialized by an instruction from main processor 90 (step 150), whereupon DSP 104 initializes its hardware registers (step 152) and then proceeds to step 154 for determining data of the unweighted time domain code component as shown in Fig. 6, which is it is then stored in memory so as to be read as needed for the generation of time-domain code components as mentioned above. Alternatively, this step may be omitted if the code components are stored permanently in ROM or other non-volatile memory. It is also possible to compute the code frequency component data on demand, which, however, causes a higher processing overhead. Another way is to generate unweighted code components in analog form and then determine the amplitudes of the analog components using weighting factors produced by the digital processor.
After the time-domain data has been calculated and collected, in step 156, the DSP 104 sends a request to the main processor 90 asking for the next information to be encoded. The information is in the form of a stream of characters, integers, or other unique data symbols identifying groups of code components that are outputted by the DSP 104 in the order that is assumed by the information. In other embodiments, the main processor, knowing the DSP data output rate, determines for itself when to provide the next information to the DSP by appropriately setting the timer and providing the information in a time-synchronized manner. In another alternative embodiment, the output of the DSP 104 is coupled to a decoder to receive the output code components to decode them and feed the information back to the main processor as an output from the DSP, so that the host computer can determine when to provide the next information to the DSP 104. in yet other embodiments, functions of main processor 90 and DSP 104 are performed by a single processor.
After the next information has been received from the main processor, in step 156, the DSP proceeds to generate the code components for each in the order of the information symbol and provides the combined, weighted code frequency components to its output 106. This process is represented by a loop identified by 160 in Fig. 7A.
180 441
Upon entering loop 160, the DSP 104 enables interrupts 1 and 2 and proceeds to the "weighting factor determination" procedure 162 which will be described in connection with the flowchart of Figs. 7B and 7C. Referring first to fig. 7B, upon entering procedure 162, the DSP first determines whether sufficient audio samples have been collected to allow a high resolution FFT to be performed to perform a spectral analysis of the audio signal content in the last predetermined interval of the audio signal, indicated by step 163. To start. it is necessary that a sufficient number of audio signal samples be collected to perform the FFT. However, if an overlapping FFT is used, on successive loop passes a correspondingly smaller number of samples must be collected before the next FFT execution.
As will be seen in Fig. 7B, the DSP remains in the small loop 163 waiting for the necessary sample collection. Upon interruption 1, the A / D converter 124 provides a new digital sample of the program audio signal that is stored in the data buffer DSP 104, indicated as subroutine 164 in FIG. 7A.
Returning to Fig. 7B, after the DSP has collected enough data samples, processing proceeds to step 168 where said high resolution FFT is performed on the audio signal data samples of the last audio signal interval. Then, as indicated by 170, an appropriate weighting or amplitude determination factor is computed for each of the code frequency components in the currently encoded symbol. In step 172, the one of the frequency bud produced by the high frequency FFT (step 168) that provides the ability to mask the top level of the corresponding code component based on a single tone ("tonal dominant") is determined as described above.
Referring also to Fig. 7C, in step 176, a weighting factor for tonal dominant is determined and held for comparison with the corresponding masking capacities provided by wideband and narrowband masking, and, if found to be the most preferred masking method, is used as the weighting factor for setting the amplitude of the current frequency component of the code. In a next step 180, estimation of narrowband masking capability and wideband masking is performed as previously described. Then, in step 182, it is determined whether narrowband masking provides the best masking ability of the corresponding code component, and if so, in step 184, the weighting factor is changed based on narrowband masking. In a next step 186, it is determined if the wideband masking provides the best masking ability of the corresponding code frequency component, and if so, in step 190 the weighting factor is changed based on the wideband masking. Then, in step 192, it is determined if weighting factors have been selected for each code frequency component to be currently output to the current symbol representation and if not, the loop is re-initialized to select a weighting factor for the next code frequency component. However, if weighting factors for all components have been selected, then the subroutine is terminated as indicated by step 194.
After interrupt 2 has occurred, processing proceeds to subroutine 200 in which the functions shown in FIG. 6 are performed. That is, in routine 200, the weighting factors calculated in routine 162 are used to multiply the respective time domain values of the current symbol to be output. and then the time-domain-weighted values of the code components are added and outputted as vases, the total code signal to the DAC 140. Each code symbol is withdrawn for a predetermined period of time, after which processing proceeds to step 156 in step 202.
Figures 7D and 7E show the flowcharts for implementing a sliding tonal analysis for computing masking effects in an audio signal. In step 702, variables such as sample size of large FFT and minor FFT, number of minor FFTs per major FFT, and number of code tones per symbol, e.g., 2048, 256, 8, and 10, are initialized, respectively.
In steps 704-708, the number of samples corresponding to a large FFT is analyzed. In step 704, the audio signal is sampled. In step 706, power is taken
180 441 program material in each FFT bundle. In step 708, an acceptable code tone power in each respective FFT bin is obtained, in view of the effects of all the corresponding audio signal tones in that bin, for each of the tones. The flowchart of Fig. 7E shows 708 in greater detail.
In steps 710-712, a number of samples corresponding to the smaller FFT are analyzed. In step 714, the permissible code powers determined from the large FFT in step 708 and from the small FFT in step 712 are combined for the part of the samples that underwent the lower FFT. In step 716, code tones are mixed with an audio signal to create a coded audio, and in a step 718 the encoded audio is output to the DAC 140. In step 720 it is decided to repeat steps 710-718, that is, the remainder of the audio samples having passed the high FFT and not the lower. Then, in step 722, if there are no more sound samples, the next number of samples corresponding to the large FFT is analyzed.
Figure 7E shows details of steps 708 and 712 for determining the allowed code strength in each FFT bin. In general, this procedure models an audio signal as containing a set of tones (see examples below), calculates the effect of masking by each of the tones of the audio signal of each code tone, sums the masking effects, and determines the code tone density and signal complexity.
In step 752, the band involved is determined. For example, let the encoding bandwidth used be 800 Hz to 3200 Hz and the sampling rate 44100 samples per second. The starting packet starts at 800 Hz and the last packet is at 3200 Hz.
In step 754, the masking effect of each corresponding audio signal tone for each code tone in that bundle is determined using the masking curve for each code tone, and compensation is made for the non-zero burst width FFT of the audio signal by determining (1) a first masking value based on assuming that the entire strength of the audio signal is at the upper end of the bundle, and determining (2) the second masking value based on the assumption, that the entire strength of the audio signal is at the lower end of the packet, and the masking value whichever is the smaller is selected.
Figure 7F shows an approximation of the single-tone masking curve for a tone of an audio signal at a frequency of fPGM that is about 2200 Hz in this example, according to JJ Zwislock's work on "Masking: Experimental and Theoretical Aspects of Simultaneous, Forward, Backward and Central Masking". 1978, Zwicker et al., Psychoacoustics: Facts and Models edition, pages 283-316, Springer-Verlag, New York. The critical bandwidth (CB) is defined by Zwislocki as follows: critical band 0.002 * fPGMl, 5 + 100
According to the following definitions, where "mask" is the tone of the audible signal:
BRKPOINT PEAKFAC BEATFAC mNEG mPOS cf mf cband = 0.3 (± 0.3 critical bands) = 0.025119 (- 16 dB from mask) = 0.002512 (-26 dB from mask) = -2.40 (- 24 dB per critical band) = -0.70 (- 7 dB per critical band) = code frequency = mask frequency = critical band around fPGM the masking factor, mfactor, can be calculated as follows: brkpt = cband * BRKPOINT if on the negative slope of the curve from fig. 7F mfactor = PEAKFAC * 10 ** (mNEG * mf-brkpt-cf) / cband) if on the flat part of the curve in Fig. 7 mfactor = BEATRAC if on the positive slope of the curve in Fig. 7F mfactor = PEAKFAC * 10 ** (mPOS * mf-brkpt-cf) / cband).
180 441
Specifically, the first mfactor is calculated on the assumption that all of the audio signal power is at the lower end of its packet, the second mfactor is calculated assuming that all of the audio signal power is at the upper end of its packet, and the smaller of the two factors is chosen as the mask value provided by the beep tone for the selected code tone. In step 754, processing is performed for each specific tone of the audio signal for each tone of the code.
In step 756, each code tone is set by each of said masking factors corresponding to the audio signal tones. In this embodiment, the masking factor is multiplied by the audio signal power in a given batch.
In step 758, the result of multiplying the masking factors by the audio signal power is summed for each batch to provide an acceptable power for each code tone.
In step 760, for the number of code tones, allowable powers in the critical band on each side of the code tone that is being computed and for the complexity of the audio signal are determined. The number of code tones in the critical band as indicated by CTSUM is calculated. The fix factor, ADJFAC, is given by the formula:
ADJFAC = GLOBAL * (PSUM / PRSS) 1.5 / CTSUM, where GLOBAL is the derating factor for encoder inaccuracy due to time delays in performing the FFT, (PSUM / PRSS) 1.5 is the experimental complexity correction factor and 1 / CTSUM represents a simple division of the audio signal power for all code tones to be masked. The PSUM is the sum of the tone masking power levels associated with code tone masking, the ADJFAC of which is determined. The root of the sum of the squares of the power (PRSS) is given by
PRSS = SQRT (Σί (Ρ<sup>2</sup>0) where i = FFT packets in the band
For example, assuming the total tone masking power in the band is evenly distributed over one, two, or three tones, then:
<td>tone no</td><td>tone power</td><td>PSUM</td><td>PRSS</td>
<td> 1</td><td> 10</td><td> 1*10 = 10</td><td> 10</td>
<td> 2</td><td> 5,5</td><td> 2*5 = 10</td><td>SQRT (2 * 5<sup>2</sup>) = 7.07</td>
<td> 3</td><td> 3,3,3,3,3,3</td><td> 3*3,3 = 10</td><td>SQRT (3 * 3.3<sup>2</sup>) = 5,77</td>
Hence, PRSS measures the degree of aggregation of the masking power (increasing values) or spreading (decreasing values) of the program material.
In step 762 of Fig. 7E, it is determined if there are still packets in the band under consideration, and if so, they are processed as described above.
Examples of masking calculations will now be presented. The symbol of the audio signal at 0 dB is taken so that the values obtained are the maximum code tone strengths relative to the audio signal strength. Four cases are considered: single tone 2500 Hz; three 'tones at 2000, 2500 and 3000 Hz, narrowband noise modeled as 75 tones in a critical band centered at 2600, where 75 tones are evenly spaced every 5 Hz in the range 2415 to 2785 Hz; and broadband noise modeled as 351 tones evenly spaced every 5 Hz in the range from 1750 to 3250 Hz. For each case, the result obtained from the Sliding Tonal Analysis (STA) is compared with the calculated result of the best of the three analysis types: single tone, narrowband noise, and broadband noise.
180 441
<td rowspan="2">Code tone (Hz)</td><td colspan="2">Single tone</td><td colspan="2">Many tones</td><td colspan="2">Narrowband noise</td><td colspan="2">Broadband noise</td>
<td>STA (dB)</td><td>best of 3 (dB)</td><td>STA (dB)</td><td>best of 3 (dB)</td><td>STA (dB)</td><td>best of 3 (dB)</td><td>STA (dB)</td><td>best of 3 (dB)</td>
<td> 1976</td><td> -50</td><td> -49</td><td> -28</td><td> -30</td><td> -19</td><td>ON</td><td> 14</td><td> 12</td>
<td> 2070</td><td> -45</td><td> -45</td><td> -22</td><td> -32</td><td> -14</td><td>ON</td><td> 13</td><td> 12</td>
<td> 2163</td><td> -40</td><td> -39</td><td> -29</td><td> -25</td><td> -9</td><td>ON</td><td> 13</td><td> 12</td>
<td> 2257</td><td> -34</td><td> -33</td><td> -28</td><td> -28</td><td> -3</td><td>ON</td><td> 12</td><td> 12</td>
<td> 2351</td><td> -28</td><td> -27</td><td> -20</td><td> -28</td><td> 1</td><td>ON</td><td> 12</td><td> 12</td>
<td> 2444</td><td> -34</td><td> -34</td><td> -23</td><td> -33</td><td> 2</td><td> 7</td><td> 13</td><td> 12</td>
<td> 2538</td><td> -34</td><td> -34</td><td> -24</td><td> -34</td><td> 3</td><td> 7</td><td> 13</td><td> 12</td>
<td> 2632</td><td> -24</td><td> -24</td><td> -18</td><td> -24</td><td> 5</td><td> 7</td><td> 14</td><td> 12</td>
<td> 2726</td><td> -26</td><td> -26</td><td> -21</td><td> -26</td><td> 5</td><td> 7</td><td> 14</td><td> 12</td>
<td> 2819</td><td> -27</td><td> -27</td><td> -22</td><td> -27</td><td> 6</td><td>ON</td><td> 15</td><td> 12</td>
For example, in the sliding tone analysis (STA) for a one-tone case, the masking tone is 2500 Hz, which corresponds to a critical bandwidth of 0.02 * 25001.5 + 100 = 350 Hz. The curve breakpoints of FIG. 7F are at 2500 0.3 * 350, that is, 2395 and 2605 Hz. The code frequency of 1976 as can be seen is on the negatively sloping portion of the curve of Fig. 7F, so the masking factor is:
mfactor = 0.025119 * 10 -2.4 * (2500 - 105 - 1976) / 350 = 3.365 * 10-5 = -44.7 dB.
There are three code tones in the 1976 Hz critical band, so the masking power is split between them:
4.364 * 10-5 / 3 = -49.5 dB
This result is rounded to -50 dB and is shown in the upper left corner of the result table.
In the "best of three" analysis, tonal masking is calculated according to the single tone method explained above with reference to Fig. 7F.
In the "best of three" analysis, narrowband noise masking is calculated by first counting the average power over a critical band centered around the frequency of the code tone under consideration. Tonals with a power greater than the average power are not taken into account as part of the noise and are removed. The sum of the remaining power is the power of narrowband noise. The maximum permissible code tone power is - 6 dB narrowband noise power for all tones in the critical band of the code tone under consideration.
In the "best of three" analysis, the broadband noise masking is calculated by counting the narrowband noise power for the critical bandwidth measures at 2000, 2280, 2600 and 2970 Hz. The minimum of the calculated narrowband noise power is multiplied by the ratio of the total bandwidth to the corresponding critical bandwidth to find the wideband noise power. For example, if a band centered at 2600 Hz has a critical bandwidth of 370 Hz and has the lowest power, its narrowband noise power is multiplied by 1322 Hz / 370 Hz = 3.57 to obtain the wideband noise power. The acceptable code tone power is - 3 dB of broadband noise. When there are 10 code tones, the maximum power allowed for each tone is 10 dB less, i.e. -13 dB of broadband noise power.
As can be seen, the calculation of the sliding tonal analysis generally corresponds to that of "best of three", which indicates that the sliding tonal analysis is a reliable method. Moreover, the results obtained by this analysis for the multi-tone case are better, i.e. enable
180 441 more code tone strengths than "best of three" analyzes, indicating that sliding tonal analysis is appropriate even for cases that do not fit well with one of the "best of three" calculations.
Referring to Fig. 8, a block diagram of an embodiment of an encoder which uses an analog circuit is shown. The analog encoder receives an audio signal in analog form at the input terminal 210, from which the audio signal is provided as input to N component generator circuits 2201 to 220N, each generating a corresponding code component C1 to CN. For simplicity and clarity, only one of the generator circuits 2201 to 220N is shown in Fig. 8. For the controlled generation of code components for the respective data symbol to be included in the audio signal such that an encoded audio signal is generated, each of the generator components is supplied with input data via respective terminals 2221 to 222N which serve as enable inputs for the corresponding components of the generator systems. Each symbol is encoded as a subset of the code components C1 through CN by selectively applying a enable signal to certain components of the generator 2201 through 220N. The generated code components corresponding to each data symbol are provided to the inputs of adder 226 which also takes input audio from input terminal 210, which is used to add code components to input audio to produce an encoded audio signal that is outputted to the output. .
Each of the component generator circuits is similar in structure and includes a respective weighting factor determining circuit 2301 to 230N, a corresponding signal generator 2321 to 232N, and a respective switch circuit 2341 to 234N. Each of the signal generators 2321 to 232N produces a different code frequency component, respectively, and provides the generated component to the respective switching circuit, 2341 to 234N, each having a second input shorted to ground and an output coupled to the input to the corresponding of the multipliers 2361 to 236N. . Responding to receiving an enable signal at the respective data input terminal 2201 to 220N, each of the switches 2341 to 234N responds by coupling the output of the respective signal generator 2321 to 232N to the input of the corresponding multiplier 2361 to 236N . Meanwhile, in the absence of a data input enable signal, each switch 2341 to 234N shorts its output to ground so that the output of the corresponding multiplier 2361 to 236N is zero.
Each weighting factor calculator 2301 to 230N is used to estimate the ability of the frequency components of an audio signal in the corresponding frequency band of that signal to massage the code component produced by the corresponding generator circuit 2321 to 232N to produce a weighting factor which is then reported. input to the appropriate multiplier, 2361 to 236N, to determine the amplitude of the corresponding code component to ensure, that it will be masked by the portion of the audio signal that has been computed by the weighting factor calculator. Referring also to Fig. 9, the structure of each of the weighting factor determining circuits 2301 to 230N, designated exemplary circuit 230, is shown in block form. Circuit 230 includes a masking filter 240 which receives an audio signal at its input and serves to extract a portion of the audio signal to be used to calculate a weighting factor to be applied to the corresponding multipliers 2361 to 236N. The properties of the masking filter are further chosen to balance the amplitudes of the frequency components of the audio signal with respect to. their ability to mask the relevant code component.
The portion of the audio signal selected by the masking filter 240 is provided to an absolute evaluator 242 which produces an output representing the absolute value of the portion of the signal in a frequency band after passing through the masking filter 240. The output of the absolute evaluation circuit 242 is provided as an input to a scaling amplifier 244 having a gain selected to produce a signal that, when multiplied by the output of the corresponding switch, 2341 to 234N, forms
180 441 with the code component at the output of the corresponding multiplier, from 2361 to 236N, which will ensure that the multiplied code component is masked by the selected portion of the audio signal that has passed through the masking filter 240 when reproducing the encoded audio signal as audio. Each weighting factor calculator 2301 to 230N thus produces a signal representing an estimate of the ability of a selected portion of the audio signal to mask the corresponding code component.
In other embodiments of the analog encoders of the present invention, a plurality of weighting factor determiners are provided to the code generator of each code component, and each of the plurality of weighting factor determiners corresponding to a given code component calculates the ability of various portions of an audio signal to mask that particular component when the encoded signal is sound is played as sound. For example, a plurality of weighting factor calculators may be provided, each calculating the ability of a portion of an audio signal over a relatively narrow frequency band (such that the energy of the audio signal in such a band will very likely consist of a single frequency component) to mask the corresponding component. code when encoded audio is played as audio. A further weighting factor determining circuitry for the same corresponding code component may be provided to calculate the energy capacity of an audio signal in a critical band whose center frequency is the code component to mask the code component when the encoded audio signal is played as audio.
Moreover, although various elements of the embodiment of Figs. 8 and 9 are implemented as analog circuits, it is possible to provide the same functions performed by the digital circuits.
Decoding
Decoders and decoding methods which are particularly suited for decoding audio signals encoded by the above-described methods according to the invention, as well as generally for decoding codes contained in audio signals, such that the codes can be distinguished from the rest of the signal on the basis of amplitude, will now be described. In accordance with certain features of the present invention, and with reference to the block diagram of Fig. 10, the presence of at least one code component in an encoded audio signal is detected by adjusting the expected amplitude or amplitudes of at least one code component to produce either the audio signal level or the bearish signal noise level, or both, as indicated by function block 250. One or more signals representing such a suitable amplitude or amplitude are provided at 252 in FIG. 10, for determining the presence of a code component by detecting a signal corresponding to the expected amplitude or amplitudes as indicated by function block 254. The decoders of the present invention are particularly well suited for detecting the presence of code components which are masked by other components of the audio signal if the amplitude relationship between code components and other components of the audio signal is, to some extent, presumed.
Figure 11 shows a block diagram of an embodiment of a decoder according to the present invention which performs digital signal processing to extract codes from encoded audio signals received by the decoder in analog form. The decoder in Fig. 11 has an input terminal 260 for receiving an encoded analog audio signal, which may be, for example, a signal taken from a microphone, from a radio or television broadcast, reproduced as audio by a receiver, or an encoded analog audio signal in the form of electrical signals directly from such a receiver. Such encoded analog sound may be created by playing back an audio recording such as on a compact disc or tape cassette. Analog processing circuits 262 are coupled to input 260 to receive encoded analog audio and serve to amplify the signal, automatically control the low-pass filtering gain to prevent aliasing prior to analog-to-digital conversion. In addition, the analog processing circuits 262 are used to perform bandpass filtering to ensure that the signals are at the output are limited to the frequency range in which
180 441 code may occur. The analog processing circuits 262 feed the processed analog signals to an analog-to-digital (A / D) converter 263, which digitizes the received signals and supplies them to a digital signal processor (DSP) 266, which already processes digital signals to detect the presence of code components and designates the code symbols that are represented by them. The digital signal processor 266 is coupled to a memory 270 (including program and data memories) and input / output (I / O) circuits 272 to receive external instructions (e.g., a decode initiate command or an instruction to remove accumulated codes) and to output decoded information. .
The operation of the digital decoder of Fig. 11, which decodes encoded audio signals with the apparatus of Fig. 3, will now be described. Analog processing circuit 262 serves as a bandpass filter for encoded audio signals, with a passband extending from approximately 1.5 kHz to 3.1 kHz, and DSP 266 samples the filtered analog signals at a sufficiently high frequency. The digital audio signal is then divided by the DSP 266 into ranges of frequency components or FFT processing bundles. More specifically, an overlapping window FFT is performed on a predicted number of most recent data points such that a new FFT is performed periodically on a sufficient number of new samples. The data is weighted as described below and an FFT is performed to produce a predetermined number of frequency bins each with a predetermined width. The energy B (i) of each of the frequency packets over the range of the code component frequencies is computed by DSP 266.
Noise level estimation is performed around each batch where a noise component may occur. Accordingly, when the decoder of Fig. 11 is used to decode signals encoded by the embodiment of Fig. 3, there are 40 frequency bins in which code components may occur. For each frequency bundle, the noise level is estimated as follows. First, the average energy E (j) in the frequency bundles in the window extending at frequencies below and above a particular frequency bundle j (i.e., the bundle in which the code component may occur) is calculated according to the following relationship:
Εφ = —ĄsB (i) 2w + 1 where i = (jw) -> (j + w), and w represents the size of the window below and above the package under consideration, expressed as the number of packages. Then the noise level NS (j) in the frequency bin j is calculated according to the following formula:
NS (j) = (ΣΒη (ΐ)) / Σδ (ΐ)) where Bn (i) equals B (i) (energy level in the i packet), if B (i) <E (j), i equals 0 otherwise and δ (ί) equals 1 if B (i) <E (j), and 0 otherwise. That is, the noise components are assumed to contain components having a level less than the average energy level in the particular window surrounding the bundle under consideration, and therefore contain components of the audio signal that fall below the average energy level.
Once the noise level for a bundle under consideration has been estimated, the signal-to-noise ratio SNR (j) for that bundle is estimated by dividing the energy level B (j) in the bundle under consideration by the estimated noise level NS (j). The SNR (j) values are used to detect and timing the sync symbols as well as data symbol states as described below. Various methods may be used to eliminate components of the audio signal as potentially not code components on a statistical basis. For example, it can be assumed that the bundle having the highest signal-to-noise ratio has an audio signal component. Another possibility is to exclude those packages that have SNR (j) above the assumed value. Another possibility is to eliminate packets having the highest and / or lowest SNR (j).
When used to detect the presence of codes in audio signals encoded with the device of Fig. 3, the device of Fig. 11 collects presence indicative data.
180 441 code components in each of the packages under consideration in a cyclical fashion for at least a major portion of a predetermined range within which the code symbol may be found. Accordingly, the following process is repeated many times and the present component data is collected for each of the bundles considered in a given time frame. The methods of determining the respective detection time frames from the timing codes will be described in greater detail below. Once the DSP 266 has collected the data for the corresponding frame, it determines as described below which of the possible code signals was present in the signal. DSP 266 then stores the detected code symbol in memory 270 together with a time stamp to identify the point at which the signal was detected from the internal DSP clock signal. Then, in response to the appropriate command for DSP 266 received from I / O circuit 272, the DSP causes memory 270 to output stored code symbols and time stamps through I / O circuit 272.
The flowcharts of Figs. 12A and 12B illustrate a sequence of operations performed by DSP 266 in decoding an encoded symbol in analog audio signal received at input terminal 260. Referring first to Fig. 12A, after initiating the decoding process, DSP 266 enters the main program loop in step 450 where a SYNCH flag is determined so that DSP 266 first starts an operation of detecting sync symbols E and S in the input audio signal in a predetermined order of information. After DSP 266 executes step 450, the DSP calls a subroutine DET which is illustrated in the flowchart of Fig. 12B for searching for the presence of code components representing the sync symbols in the audio signal.
Referring to Fig. 12B, in step 454, the DSP collects and retains input audio samples repeatedly until a sufficient number of FFTs described above have been accumulated. Upon completion of this operation, the collected data is subjected to a weighing function such as a cosine square weight function, Kaiser-Bessel function, Gauss (Poisson) function, Hanning function or other appropriate weighing function as indicated by step 456, to establish data windows. However, when the code components are sufficiently clear, weighing is not necessary. The data windows are then subjected to an overlapping FFT as indicated by step 460.
When the FFT has completed, in step 462 the SYNCH flag is tested to see if it is set (expect a sync symbol in this case) or that it is cleared (in this case a data bit symbol is expected). Since the DSP initially sets the SYNCH flag to detect the presence of code components representing sync symbols, the program proceeds to step 466 because frequency domain data obtained from FFT in step 460 is computed to determine whether this data indicates the presence of components representing sync symbol E or sync symbol. S.
To detect the presence and timing of sync symbols, first a sum of SNR values (j) for each possible sync symbol and data symbol is determined. At any given time during the sync symbol detection process, a specific symbol is expected. As a first step in detecting the expected symbol, it is determined whether the sum of its respective SNR (j) values is greater than any of the others. If so, then the detection threshold is established based on the noise levels in the frequency packets, which may include code components. That is, if, at any given time, only one code symbol is included in an encoded audio signal, only a quarter of the packets considered will contain the code components. The remaining three-fourths of the packets will contain noise, that is, sound program components and / or other additional energies. The detection threshold is produced as the average of the SNR (j) values for all forty packages considered, but can be set by a multiplication factor to take into account the effects of neutral noise and / or to compensate for the observed error amount. Once the detection threshold has been established, the sum of the SNR (j) values of the expected timing symbol is compared with the detection threshold to determine whether or not it is greater than the threshold. If so, detection of the expected sync symbol is recorded. In stating this fact, as denoted as
180 441, step 470, the program returns to the main processing loop of Fig. 12A to step 472, where it is determined (as described below) whether the pattern of decoded data corresponds to a predetermined eligibility criterion. If not, processing returns to step 450 to restart testing for the presence of the sync symbol in the audio signal, and if these criteria are met, it is determined whether the expected sync pattern (i.e., the expected sequence of E and S symbols) has been fully received and detected. which is marked as step 474.
However, after first going through the DET routine, insufficient data will be held to determine if the pattern meets the qualification criteria, so from step 474 processing returns to the DET routine to perform a further FFT and compute the presence of a sync symbol. Once the subroutine DET has been executed the predetermined number of times, when processing returns to step 472, the DSP determines whether the collected data meets the qualification criteria for the timing pattern.
That is, when the DET has been performed the predetermined number of times, a corresponding number of calculations have been made at step 466 of the DET procedure. The number of detected instances of the "E" symbol is used in one embodiment as a measure of the amount of "E" symbol energy during a corresponding period of time. However, other "E" symbol energy measurements (such as all "E" packet SNRs that exceed the average packet energy) may also be applied here. After calling the DET procedure again and executing the calculation further in step 466, in step 472 this most recent calculation is added to the computation stored in the predetermined range, and the oldest computation previously collected is discarded. This process continues during multiple passes through the subroutine DET, and at 'step 472, peak energy of the "E" symbol is searched for. If the peak is not found, this leads to the finding that no sync pattern has been found and processing returns from step 472 to step 450 where it re-sets the SYNCH flag and begins searching for a sync pattern.
In the event that such an "E" symbol energy maximum is found, the computation performed in step 472 after subroutine 452 continues each time using the same number of computations from step 466, but discarding the oldest computation and adding the newest, so that a data window is used for this purpose. While this process is running, after the predetermined number of passes, step 472 determines if there has been a transition from an "E" symbol to an "S" symbol. In one embodiment, this is determined as the point where all of the "S" packet SNRs resulting from step 466 in the moving window will exceed all of the "E" packet SNRs during the same interval. Once such a transition point has been found, processing continues as previously described, that is, the maximum energy of the "S" symbol is searched, which is denoted as the largest number of "S" detections in the moving data window. If no such maximum is found, or if such a maximum does not occur in the predetermined time frame after the "S" symbol maximum energy, processing proceeds from step 472 back to step 450, and the search for the synchronization pattern begins again.
If the above criteria are met, the presence of a sync pattern is declared in step 474, and processing proceeds to 'step 480 to determine the respective bit intervals based on the energy peaks of the "E" and "S" symbols and the detected transition point. Instead of such a timing pattern detection process, other strategies may be used. In another embodiment, when the timing pattern does not meet criteria such as those described earlier, but approximates a qualifying pattern (i.e., the detected pattern is not explicitly unclassified), determining whether the timing pattern has been detected may be suspended for further analysis. it is based on a calculation performed (as described below) to determine the presence of data bit intervals following a potential synchronization pattern. Based on the totality of the detected data, i.e. during the suspect timing pattern interval and during the expected bit interval, a retrospective qualification of the possible timing pattern can be performed.
Returning to Fig. 12A, once the timing pattern has been positively classified, in step 480, as mentioned above, bit timing is determined based on
180 441 of the two maxima and the transition point. That is, these values are averaged to determine the expected start and end points of successive data bit intervals. Once this has been done, in step 482 the SYNCH flag is cleared to indicate that the DSP will be looking for the presence of possible bit states. Then the DET routine 452 is called again and, referring also to Fig. 12B, this continues in the same manner as described above up to step 462, in which the SYNCH flag indicates that a status bit should be determined and processing proceeds to step 486. In step 486, the DSP looks for the presence of code components indicating the status of the bit. zero or one as described above.
When this is complete, in step 470, processing returns to the main processing loop of Fig. 12A in step 490, where it is determined whether a sufficient portion of data has been received to determine the state of the bit. To do this, multiple passes through subroutine 452 must be made, such that after the first pass, processing returns to routine DET 452 to perform further calculations based on the new FFT. After subroutine 452 is executed the predetermined number of times, the accumulated data is calculated in step 486 to determine whether the received data indicates state zero, state one, or undefined state (which can be determined based on parity of the data). That is, the sum of the SNR of the packets "0" is compared to the sum of the SNR of the packet "1". Which one is greater determines the state of the data, and if they are equal, the state is undefined. Alternatively, if the sums of the SNRs of the "0" and "1" packets are not equal, but are close to each other, an undefined state may also be declared. Also, if there are more data symbols, the symbol with the highest sum of SNRs found is declared as the received symbol.
When processing returns to step 490, determining the state of the bit is detected and the processor proceeds to step 492 where the DSP stores data in memory 270 indicating the states of the respective bits constituting a word with a predetermined number of symbols represented by coded symbols in a received audio signal. Then, in step 490, it is determined whether the received data is for all bits of the coded word or information. If not, processing returns to the DET subroutine 452 to determine the bit state of the next expected information symbol. However, if it is determined in step 496 that the last information symbol has been received, processing returns to step 450 to set a SYNCH flag to examine for the presence of sync symbols represented by code components in an encoded audio signal.
Referring to Fig. 13, in some embodiments, the non-code components of the audio signal and / or other noise (generally referred to simply as "noise" in this context) are used to produce a comparative value, such as a threshold, as indicated by function block 276. One or more portions of the coded audio signal are compared to a comparison value, indicated as functional block 277, to detect the presence of code components. Advantageously, the encoded audio signal is firstly processed to isolate the components in the band (or multiple bands) that may contain code components and then collected over a period of time for noise averaging as indicated by function block 278.
Referring now to Fig. 14, an embodiment of an analog decoder according to the invention is shown in block form. The decoder of Fig. 14 includes an input terminal that is coupled to four groups of component detectors 282, 284, 286, and 288. Each group of component detectors 282 to 288 is for detecting the presence of code components in the input audio signal representing a corresponding code symbol. In the embodiment of Fig. 14, the decoder apparatus is such as to detect the presence of each of the 4N code components, where N is an integer, such that the code consists of four different symbols each represented by a unique group of N code components. Accordingly, the four groups 282 to 288 contain 4N component detectors.
An embodiment of one of the 4N component detectors of groups 282 to 288 is shown in block form in Fig. 15 and is identified therein as a component detector 290. Component detector 290 has an input 292 coupled to an input 280 of the decoder of Fig. 14 to receive an encoded signal. sound. The component detector 290 has an upper branch
180 441 of the circuit, including a noise estimation filter 294, which, in one embodiment, is a bandpass filter for transmitting audio signal energy over a band centered at a frequency corresponding to a detected code component. In an alternative and preferred embodiment, the noise estimation filter 294 consists of two filters, one of which has a passband extending upward from the frequency of the corresponding detected code component, and the other filter having a passband extending downward from the frequency of the detected code component. so that both filters together pass energy having frequencies above and below (but not including) the frequency of the component, to be detected but present in its vicinity. The output of the noise estimator 294 is coupled to an input of the absolute evaluator 296 that produces an output representing the absolute value of the output from the noise estimation filter 294, fed to the input of the integrator 300, which collects signals fed thereto and outputs a value representing the signal energies in adjacent, but not including, parts of the spectrum with the frequency of the component to be detected, and feeds the value to the non-inverting input of differential amplifier 302 which operates as a logarithmic amplifier.
The component detector of Fig. 15 also has a lower branch including a signal estimation filter 306 whose input is coupled to input 292 to receive an encoded audio signal to pass a frequency band substantially narrower than the wide band of the noise estimation filter 294 such that the estimation filter signal 306 passes signal components essentially only with the frequency of the code component to be detected. Signal estimator 306 has an output coupled to an input of absolute evaluator 308, which is for outputting a signal representing the absolute value of the signal from signal estimation filter 306. The output of the absolute evaluator 308 is coupled to an input of the integrator 310. The integrator 310 collects the values output from the circuit 308 and produces an output signal representing the energy in the narrow bandwidth of the signal estimation filter for a predetermined period of time.
The integrators 300 and 310 each have reset inputs interconnected to receive a common reset signal applied to terminal 312. The reset signal is provided by a controller 314 shown in FIG. 14 that generates a reset signal periodically.
Returning to FIG. 15, the output from integrator 310 is fed to the inverted input of amplifier 302, which is operated to produce an output signal representing the difference between the output from integrator 310 and from integrator 300. If the amplifier 302 is a logarithmic amplifier, the range of possible outputs is narrowed to reduce the dynamic range of the output, for application to a window comparator 316 that detects the presence or absence of a code component during a given interval, as determined by the control circuit 314 by applying a reset signal. The window comparator 314 outputs a code presence signal in the event that the input supplied from the amplifier 302 falls between a low threshold applied as a predetermined value to the low threshold input terminal of comparator 316 and a predetermined high threshold input terminal of comparator high threshold 316.
Referring again to FIG. 14, each of the N component detectors 290 of each group of detectors couples the output of the corresponding window comparator 316 to the input of code determination logic 320. Circuit 320, controlled by controller 314, collects different code presence signals from 4N component detector arrays. 290 for a large number of zero cycles, depending on the setting by the controller 314. At the end of the detection interval of a given symbol as determined as described, code determination logic 320 determines whether the code symbol has been received as the symbol for which the highest number of components has been detected during the interval, and outputs 322 to an output indicating the detected code symbol. . The output signal may be stored in memory, included in a larger information or data file, transmitted or otherwise used (for example as a control signal).
180 441
Symbol detection intervals for the decoders described above with reference to Figs. 11, 12A, 12B, 14, and 15 may be established based on the timing of the sync symbols transmitted in each coded information having a predetermined duration and order. For example, the coded information contained in an audio signal may be composed of two data slices of an encoded E symbol followed by two data slices of an encoded S symbol, both described with reference to Fig. 4. Decoders of Figs. 11, 12A, 12B, 14, and 15 are operated to initially look for the presence of a first expected sync symbol, i.e., an E-coded symbol, that is transmitted at a predetermined time interval and defines its transmission interval. Then, the decoders examine the presence of the code components characterizing the S symbol and, upon detection, the decoders determine its transmission interval. Based on the determined transmission intervals, a transition point from an E symbol to an S symbol is determined, and, from that point, detection intervals are determined for each data bit symbol. During the detection interval, the decoder accumulates the code components to determine the corresponding symbol transmitted during the interval in the manner described above.
Although the various elements of the embodiment of Figs. 14 and 15 are implemented by analog circuits, it is obvious that it is possible for the same functions to be performed, in whole or in part, by a digital circuit.
Referring to Figs. 16 and 17, a system for estimating the listenership of broadly disseminated information such as television and radio broadcasts is shown. Figure 16 shows a block diagram of a radio broadcast station for broadcasting audio signals over the air which have been coded to identify the station with transmission time. If desired, identification of the program or segment that is broadcast may also be enabled. The radio program source 340, such as a CD player, digital cassette player, or live sound source, is controlled by the station manager via a controller 342 to control the output of the audio signals for broadcasting. The output 344 from the audio program source is coupled to the input of the encoder 348, according to the embodiment of Fig. 3 and including a DSP 104, a bandpass filter 120, an analog to digital (A / D) converter 124, a digital to analog converter (DAC) 140, and an adder 142. The controller 342 includes a main processor 90, a keyboard 96, and a monitor 100 in the example of FIG. 3 so that the main processor of control device 342 is coupled to the DSP included in encoder 348 of FIG. 16. The encoder 348 operates under the control of the controller 342 and incorporates the coded information periodically into the audio to be transmitted, and the information includes corresponding identifying data. Encoder 348 outputs the encoded audio to the input of radio transmitter 350, which modulates the carrier wave with the encoded audio program and transmits it through the air through antenna 352. The main processor of the controller 342 is programmed with the keyboard to control the encoder so as to output appropriate coded information including station identification data. The main processor automatically generates transmit time data using a reference clock within it.
Referring also to Fig. 17, the personal system monitoring device 380 is enclosed in a housing 382 that is small enough to be carried by a person of an audience member participating in supervised audience estimation. Each audience member is provided with a personal monitoring device, such as a device 380, that is carried by a member of the audience during the examination period, such as an assumed one-week period. Personal monitoring device 380 includes an omni-directional microphone picking up sounds surrounding an audience member, including radio broadcasts played as sound through the loudspeaker of a radio receiver, such as radio receiver 390 shown in Fig. 17.
The personal monitoring device 380 also includes a signal processing circuit 394 having an input coupled to the microphone output 386 to amplify the output of the microphone and subject it to bandpass filtering to suppress frequencies outside of the audio frequency band containing various code frequency components.
180 441 placed in the program audio signal by encoder 348 in Fig. 16, as well to perform filtering for anti-aliasing by analog-to-digital conversion.
The digital circuitry of personal monitor 380 is shown in FIG. 17 in the form of a functional block diagram including a decoder block and a control block, both of which may be implemented as a digital signal processor. Program and data memory 404 is coupled to a decoder 400 to receive detected codes for storage, as well as to a control block 402 for controlling write and read operations from memory 404. The input / output (I / O) circuitry 406 is coupled to a memory 404 for receiving data output by the personal monitoring device 380, as well as for storing information such as a program instruction. The (I / O) circuit 406 is also coupled to a control block 402 to control the input and output operations of the device 380.
The decoder 400 acts as the decoder of FIG. 1, previously described, and outputs the time code and station identification data for storage in memory 404. The Personal Monitoring Device 380 is advantageously operable with a "docking station" as described in U.S. Patent Application No. 08 / 101,558, filed, Aug 2, 1993, entitled "Compilance Incentives for Audience Monitoring / Recording Devices" which is often associated with the present invention and is cited herein by reference. Moreover, the personal monitoring device is provided with the additional features of a portable broadcast-capable monitoring device as described in the aforementioned application.
The surveillance station communicates via modem over telephone lines with a central data processing device to download time and station identification data to it for generating audience and / or viewing reports. The central device may also send information to the supervisory station for use by itself or for delivery to the device 380, such as an executable program. The central device may also provide information to the surveillance station and / or device 380 over the RF channel, such as existing FM broadcasts broadcast encoded with such information in a manner in accordance with the present invention. The surveillance station and / or device 380 are provided with an FM receiver (not shown for greater clarity) which demodulates the encoded FM broadcasts and delivers them to a decoder according to the present invention. The scrambled FM broadcast may also be delivered by cable or by another broadcast center.
In addition to monitoring by personal monitoring units, stationary units (such as set-top units) may be used. The set-top units may be coupled to receive the encoded audio signal in electrical form from the receiver, or may use a microphone, such as microphone 386 in FIG. 17. Thus, settop units may monitor selected channels, while monitoring the composition of the audience, with or without using the present invention.
Other uses relate to the encoding and decoding methods of the present invention. In one application, commercial product soundtracks are provided with identification codes to ensure that commercial products have been shipped (by radio or television) as contracted for a given period of time.
In still other applications, the control signals are transmitted as codes generated in accordance with the present invention. In an application of this type, the interactive toy receives and decodes an encoded control signal included in the audio portion of a television or radio broadcast or an audio recording, and performs the appropriate action. In another application, the protective control codes are incorporated into the audio portion of radio or television broadcasts or audio recordings, so that the receiving or reproducing apparatus, by decoding these codes, can perform protective control to preventively prevent the reception and or playback of the broadcast or recordings. Also, control codes may be incorporated into cell phone transmissions to prevent unauthorized use of the cell phone ID. In yet another embodiment, codes are incorporated into telephone transmissions to distinguish transmissions
180 441 voice and data to suitably select the transmission path to avoid loss of transmitted data.
Various transmitter identification functions for authenticating military transmissions and air transport voice communications may be implemented. Monitoring devices are also contemplated here. In one such application, market research participants wear personal monitoring devices that receive coded information added to public addresses or similar audible tones at wholesalers or points of sale that record the attendance of the participants. In yet another application, workers wear personal monitoring devices to receive coded information added to audible signals at the workplace to monitor their presence at appropriate locations.
Secure communication may also be implemented using the present invention. In one such application, the underwater secure communication is performed by the coding and decoding teams of the present invention, either by assigning levels of code components such that the codes are masked by underwater sounds of the medium or by an audio source generated at the transmitter location. In another application, secure paging transmissions are performed by incorporating masking codes in the airborne audio signal transmission that are received and decoded by the paging device.
The encoding and decoding methods of the present invention can also be used for voice authentication. For example, in a telephone, a stored voice record can be compared with a live voice. In another embodiment, data such as a security number and / or time of day may be encoded and combined with a spoken speech, then decoded and used for automatic voice processing control of the speech. The control device in such an arrangement may either be connected to a telephone or other voice communication device, or may be a separate unit used when speech is to be stored directly, without transmission over telephone lines or otherwise. A further application is to provide an authentication code to the memory of a portable telephone such that the voice stream contains the authentication code, thereby allowing detection of an unauthorized transmission.
It is also possible to achieve better use of the bandwidth of the communication channel by incorporating data into voice or other audio transmissions. In an application of this type, data readings from aircraft devices are incorporated into ground-to-ground voice communications to inform ground crews of the operation of the devices without having to allocate separate voice or data channels.
Cassette piracy, unauthorized duplication of copyrighted works such as audio / video recordings and music can also be detected by encoding a unique identification number in the audio portion of each authorized copy with the encoding method of the present invention. If the encoded identification number is encoded into multiple copies, the unauthorized copying becomes evident.
A further application identifies programs that were recorded using a VCR comprising a decoder according to the present invention. Video programs (such as entertainment, commercial, etc.) are encoded in accordance with the present invention with an identification code identifying the program. When the VCR is set to record mode, the recording of the audio portion of the signals is provided to the decoder for ID code detection. The detected codes are stored in the VCR's memory for later use when generating a recording report.
Data identifying copyrighted works that have been broadcast by a station or otherwise transmitted by the producer may be retained using the present invention to ensure that protective works are performed. The works are coded with appropriate identification codes that uniquely identify them. A monitoring unit which receives broadcast signals or otherwise broadcast by one or more stations or broadcasters, provides an audio part
180 441 of the signal to an inventive decoder which detects the identification codes present therein. The detected codes are stored in memory for use in generating a report that can be used to determine copyright compliance.
Proposed decoders compatible with the Motion Picture Experts Group (MPEG) 2 standard already use some of the acoustically expanded processing elements required to extract encoded data according to the present invention, such that recording prevention using methods (e.g., preventing unauthorized recording of copyrighted works) using the codes according to the present invention are well suited to MPEG 2 decoders. A corresponding decoder according to the present invention is placed in or in addition to the recording device and detects the presence of a forbidden copy code in the sound provided from the recording. The recording device responds when detecting such a code by preventing recording of the corresponding audio signal and any accompanying signals such as video signals. The copyright information encoded according to the present invention is in-band and does not require separate timing or timing, and naturally accompanies the audio material.
In still other applications, programs transmitted by air, cable, or other means, or programs recorded on tape, disc or otherwise, include portions of sound encoded with control signals for use by one or more listener or viewer controlled devices. For example, a program describing a route a cyclist could travel may include an audio portion encoded in accordance with the invention with control signals used by the stationary exercise bike to control pedal resistance or travel of the pedals in accordance with the apparent slope of the described route. When the user pedals on the stationary bike, he sees the program on the TV screen or any other screen, and the audio part of the program is played as sound. A microphone on a stationary bicycle picks up the sound being played, and a decoder according to the present invention detects the control signals contained in the sound, thereby controlling the resistance of the pedals of the exercise bicycle.
It follows from the above that the methods of the invention may be implemented in whole or in part using analog or digital circuits, and that all or part of the signal processing functions may be performed by circuits with built-in control circuits or by using digital signal processes, microprocessors, microcomputers. , multiprocessor systems (e.g., parallel processors), or the like.
180 441
<img file="PL180441B1_D0001.tif" />
FIG. /
180 441
<img file="PL180441B1_D0002.tif" />
<img file="PL180441B1_D0003.tif" />
180 441
<img file="PL180441B1_D0004.tif" />
<img file="PL180441B1_D0005.tif" />
180 441
<img file="PL180441B1_D0006.tif" />
FIG. 4
180 441
<img file="PL180441B1_D0007.tif" />
180 441
<img file="PL180441B1_D0008.tif" />
FIG. 6
180 441
<img file="PL180441B1_D0009.tif" />
180 441
FIG.7B
START THE WEIGHING TRAINERS
<img file="PL180441B1_D0010.tif" />
_______ I ______ Z =
CALCULATE FFT HIGH <sub>vol </sub>R0ZJ) iZieu: Z05C |
168
<img file="PL180441B1_D0011.tif" />
---------! --------
J> UA OF EACH Kodo COMPONENT IN THE CURRENT ^ / Μβοι, υ
FIND A JM7 HINNAL TONAL IN THE NEIGHBORHOOD
172
<img file="PL180441B1_D0012.tif" />
<img file="PL180441B1_D0013.tif" />
180 441
<img file="PL180441B1_D0014.tif" />
STOP THE COEFFICIENT WEIGHING FACTOR ON A DOMINANT TONAL BASIS
<img file="PL180441B1_D0015.tif" />
176
180
CALCULATE NARROW AND BROADBAND MASKING
<img file="PL180441B1_D0016.tif" />
SET WEIGHING ON NARROW BAND MASKING
190^
SET WEIGHING TO BROADBAND MASKING
<img file="PL180441B1_D0017.tif" />
<img file="PL180441B1_D0018.tif" />
180 441
720
722
<img file="PL180441B1_D0019.tif" />
180 441
<img file="PL180441B1_D0020.tif" />
FIG. 7E
180 441
FIG. 7F
POWER
<img file="PL180441B1_D0021.tif" />
<sup>f</sup>PG<sub>M.</sub>-0.3CB fpg<sub>M.</sub> fp<sub>GM</sub>+ 0.3CB
180 441
<img file="PL180441B1_D0022.tif" />
180 441
<img file="PL180441B1_D0023.tif" />
<img file="PL180441B1_D0024.tif" />
<img file="PL180441B1_D0025.tif" />
180 441
<img file="PL180441B1_D0026.tif" />
FIG. ΙΟ
180 441
<img file="PL180441B1_D0027.tif" />
<img file="PL180441B1_D0028.tif" />
<img file="PL180441B1_D0029.tif" />
180 441
START
FIG. 12A
SET «SYM
450
YES call j> Er<sup>at</sup> -»------
<img file="PL180441B1_D0030.tif" />
---------1^-482
SET <sup>at</sup>blTi> <
<img file="PL180441B1_D0031.tif" />
- • DO.PA] BIT IN WORDS
492
<img file="PL180441B1_D0032.tif" />
180 441
FIG. ! 2B
<img file="PL180441B1_D0033.tif" />
180 441
<img file="PL180441B1_D0034.tif" />
ABOUT
Ν
<img file="PL180441B1_D0035.tif" />
180 441
FIG. 14
<img file="PL180441B1_D0036.tif" />
180 441
<img file="PL180441B1_D0037.tif" />
180 441
<img file="PL180441B1_D0038.tif" />
<img file="PL180441B1_D0039.tif" />
180 441
FIG.17
<img file="PL180441B1_D0040.tif" />
<img file="PL180441B1_D0041.tif" />
<img file="PL180441B1_D0042.tif" />
Publishing Department of the UP RP. Circulation of 70 copies. Price PLN 6.00.
Contents18
65 sheets
Sheet 1 Sheet 2 Sheet 3 Sheet 4 Sheet 5 Sheet 6 Sheet 7 Sheet 8 Sheet 9 Sheet 10 Sheet 11 Sheet 12 Sheet 13 Sheet 14 Sheet 15 Sheet 16 Sheet 17 Sheet 18 Sheet 19 Sheet 20 Sheet 21 Sheet 22 Sheet 23 Sheet 24 Sheet 25 Sheet 26 Sheet 27 Sheet 28 Sheet 29 Sheet 30 Sheet 31 Sheet 32 Sheet 33 Sheet 34 Sheet 35 Sheet 36 Sheet 37 Sheet 38 Sheet 39 Sheet 40 Sheet 41 Sheet 42 Sheet 43 Sheet 44 Sheet 45 Sheet 46 Sheet 47 Sheet 48 Sheet 49 Sheet 50 Sheet 51 Sheet 52 Sheet 53 Sheet 54 Sheet 55 Sheet 56 Sheet 57 Sheet 58 Sheet 59 Sheet 60 Sheet 61 Sheet 62 Sheet 63 Sheet 64 Sheet 65
120 members in 26 offices
Priority claims12
| Document | Office | Kind | Date |
|---|---|---|---|
| 22101994 | United States of America | A | |
| 22101994 | United States of America | A | |
| 40801095 | United States of America | A | |
| 40801095 | United States of America | A | |
| 9503797 | United States of America | W | |
| 9503797 | United States of America | W | |
| 221019 | – | – | – |
| 408010 | – | – | – |
| US9503797 | – | – | – |
| US19940221019 | – | – | – |
| US19950408010 | – | – | – |
| WO1995US03797 | – | – | – |
Members120
| Document | Office | Kind | |
|---|---|---|---|
| IL113190D0 | Israel | D0 | |
| US5450490A | United States of America | A | |
| CA2185790A1 | Canada | A1 | |
| WO9527349A1 | World Intellectual Property Organization (WIPO) | A1 | |
| AU2196995A | Australia | A | |
| NO964062D0 | Norway | D0 | |
| SE9603570D0 | Sweden | D0 | |
| GB9620181D0 | United Kingdom | D0 | |
| DK105996A | Denmark | A | |
| NO964062L | Norway | L | |
| HU9602628D0 | Hungary | D0 | |
| FI963827A | Finland | A | |
| FI963827L | Finland | L | |
| SE9603570L | Sweden | L | |
| GB2302000A | United Kingdom | A | |
| GB2302000A8 | United Kingdom | A8 | |
| EP0753226A1 | European Patent Office (EPO) | A1 | |
| PL316631A1 | Poland | A1 | |
| LU88820A1 | Luxembourg | A1 | |
| DE19581594T1 | Germany | T1 | |
| CZ284096A3 | Czechia | A3 | |
| CN1149366A | China | A | |
| KR970702635A | Republic of Korea | A | |
| MX9604464A | Mexico | A | |
| BR9507230A | Brazil | A | |
| HUT76453A | Hungary | A | |
| JPH10500263A | Japan | A | |
| US5764763A | United States of America | A | |
| NZ283612A | New Zealand | A | |
| GB9818342D0 | United Kingdom | D0 | |
| GB9818347D0 | United Kingdom | D0 | |
| GB9818349D0 | United Kingdom | D0 | |
| GB9818352D0 | United Kingdom | D0 | |
| GB9818353D0 | United Kingdom | D0 | |
| GB9818354D0 | United Kingdom | D0 | |
| GB9818355D0 | United Kingdom | D0 | |
| GB2325826A | United Kingdom | A | |
| GB2325827A | United Kingdom | A | |
| GB2325828A | United Kingdom | A | |
| GB2325829A | United Kingdom | A | |
| GB2325830A | United Kingdom | A | |
| GB2325831A | United Kingdom | A | |
| GB2325832A | United Kingdom | A | |
| GB9823987D0 | United Kingdom | D0 | |
| GB2302000B | United Kingdom | B | |
| GB2302000B8 | United Kingdom | B8 | |
| GB2325826B | United Kingdom | B | |
| GB2325827B | United Kingdom | B | |
| GB2325828B | United Kingdom | B | |
| GB2325829B | United Kingdom | B | |
| GB2325830B | United Kingdom | B | |
| GB2325831B | United Kingdom | B | |
| GB2325832B | United Kingdom | B | |
| GB2327582A | United Kingdom | A | |
| GB2327582B | United Kingdom | B | |
| AU709873B2 | Australia | B2 | |
| PL177808B1 | Poland | B1 | |
| AU6442299A | Australia | A | |
| IL113190A | Israel | A | |
| NZ331166A | New Zealand | A | |
| EP0753226A4 | European Patent Office (EPO) | A4 | |
| HU0004765D0 | Hungary | D0 | |
| HU0004766D0 | Hungary | D0 | |
| HU0004767D0 | Hungary | D0 | |
| HU0004768D0 | Hungary | D0 | |
| HU0004769D0 | Hungary | D0 | |
| HU0004770D0 | Hungary | D0 | |
| PL180441B1This record | Poland | B1 | |
| HU219256B | Hungary | B | |
| IL133700D0 | Israel | D0 | |
| IL133701D0 | Israel | D0 | |
| IL133702D0 | Israel | D0 | |
| IL133703D0 | Israel | D0 | |
| IL133704D0 | Israel | D0 | |
| IL133705D0 | Israel | D0 | |
| IL133706D0 | Israel | D0 | |
| IL133707D0 | Israel | D0 | |
| HU219627B | Hungary | B | |
| HU219628B | Hungary | B | |
| CZ288497B6 | Czechia | B6 | |
| HU219667B | Hungary | B | |
| HU219668B | Hungary | B | |
| NZ502630A | New Zealand | A | |
| ATA902795A | Austria | A | |
| PL183307B1 | Poland | B1 | |
| PL183573B1 | Poland | B1 | |
| US6421445B1 | United States of America | B1 | |
| AT410047B | Austria | B | |
| SE519882C2 | Sweden | C2 | |
| US2003081781A1 | United States of America | A1 | |
| AU763243B2 | Australia | B2 | |
| IL133702A | Israel | A | |
| IL133703A | Israel | A | |
| IL133700A | Israel | A | |
| IL133704A | Israel | A | |
| IL133706A | Israel | A | |
| IL133707A | Israel | A | |
| IL133701A | Israel | A | |
| IL133705A | Israel | A | |
| PL187110B1 | Poland | B1 |
1 legal event, as the office reported them to INPADOC
Events
| Event | Code | |
|---|---|---|
| Decisions on the lapse of the protection rightsLapsedLAPS | LAPS |
Numbers
- Publication, DOCDB
- 180441
- Publication, EPODOC
- PL180441B
- Application
- 95333769
- Application, DOCDB
- 33376995
- Application, EPODOC
- PL19950333769
Titles
- English
- METHOD OF AND APPARATUS FOR CODE DETECTING
Classification
- CPC, 14
- H04H20/31
- H04H20/14
- H04H60/13
- H04H60/17
- H04H60/37
- H04H60/40
- H04H60/44
- H04H60/45
- H04H60/58
- H04H60/63
- H04H60/66
- H04K1/02
- H04L27/10
- H04L27/30
- IPC, 16
- G10L19 00
- G10L19 018
- H04N5 38
- G11B20 10
- H04H20 31
- H04H60 13
- H04H60 17
- H04H60 37
- H04H60 40
- H04H60 44
- H04H60 45
- H04H60 58
- H04H60 63
- H04H60 66
- H04M11 06
- H04N5 60