Voice detection and discrimination apparatus and method
Abstract
Voice detection and discrimination apparatus for controlling a voice-operated system, comprising a terminal element (1, 2) for hearing protection in order to protect the ear by providing acoustic attenuation, an internal electroacoustic transducer element (M2) in an inner side of the terminal element (1, 2) for the ears for the purpose of detecting a first acoustic field and providing a first electronic signal representing said first acoustic field, an outer electroacoustic transducer element (M1) on an outer side of the terminal element (1, 2) for the ears in order to detect a second acoustic field and provide a second electronic signal representing said second acoustic field, an electronic unit (11, E3) connected to said electroacoustic transducer elements (M1, M2) and comprising first means of comparison (27) to compare said first and second electronic signals in order to obtain the difference between said two electronic signals, and second means of comparison (28) to compare said difference with given criteria, and output means (13, E12) connected to said electronic unit to provide an output signal depending on said second comparison, said output signal being an input signal to the system (29) activated by voice.

Term
Term ended
Projected expiry passed 28 February 2022, 4.6 years ago.
- Priority and filed
- Published
- Projected expiry
- Today
29 claims: 2 independent, 27 dependent
- 1ES 2 295 313 T3 REIVINDICACIONES 1. Aparato para detección y discriminación de voz para controlar un sistema accionado por voz, que comprende un elemento terminal (1, 2) de protección para los oídos a efectos de proteger el oído al proporcionar atenuación acústica, un elemento transductor electroacústico interior (M2) en un lado interior del elemento terminal (1, 2) para los oídos a efectos de detectar un primer campo acústico y proporcionar una primera señal electrónica que representa dicho primer campo acústico, un elemento transductor electroacústico exterior (M1) en un lado exterior del elemento terminal (1, 2) para los oídos a efectos de detectar un segundo campo acústico y proporcionar una segunda señal electrónica que representa dicho segundo campo acústico, una unidad electrónica (11, E3) conectada con dichos elementos transductores electroacústicos (M1, M2) y que comprende primeros medios de comparación (27) para comparar dichas primera y segunda señales electrónicas a efectos de obtener la diferencia entre dichas dos señales electrónicas, y segundos medios de comparación (28) para comparar dicha diferencia con unos criterios dados, y medios de salida (13, E12) conectados a dicha unidad electrónica para proporcionar una señal de salida dependiendo de dicha segunda comparación, siendo dicha señal de salida una señal de entrada al sistema (29) accionado por voz.
- 2Aparato, según la reivindicación 1, en el que la señal de salida comprende instrucciones, órdenes, código o similares para dicho sistema accionado por voz, tal como un sistema de comunicación.
- 3Aparato, según la reivindicación 1, en el que los criterios se predeterminan y se almacenan en medios de almacenamiento (E8, E9, E10) del aparato.
- 4Aparato, según la reivindicación 1, que comprende medios de procesamiento (E3) para adaptar los criterios durante el uso del aparato.
- 5Aparato, según la reivindicación 1, en el que los criterios comprenden una diferencia predeterminada en la intensidad de la señal acústica tal como es detectada por los dos elementos transductores electroacústicos (M1, M2) en un entorno ruidoso, típicamente en el intervalo de -20 a -40 dB, y una disminución permisible predeterminada de dicha diferencia en la intensidad de la señal acústica, típicamente en el intervalo de 2 a 10 dB.
- 6Aparato, según la reivindicación 1, que comprende un primer filtro electrónico (22) para filtrar la señal electrónica desde el elemento transductor electroacústico interior (M2).
- 7Aparato, según la reivindicación 1, que comprende un segundo filtro electrónico (25) para filtrar la señal electrónica desde el elemento transductor electroacústico exterior (M1).
- 8Aparato, según la reivindicación 6 ó 7, que comprende filtros de paso de banda electrónicos (22, 25), típicamente con una banda de paso en el intervalo de 150 a 700 Hz.
- 9Aparato, según la reivindicación 8, que comprende un detector (23, 26) de la intensidad de la señal conectado con dichos filtros de paso de banda (22, 25) a efectos de proporcionar entradas para un bloque de comparación (27) de la intensidad de la señal.
- 10Aparato, según la reivindicación 9, que comprende un bloque de decisión (28) conectado con dicho bloque de comparación (27) de la intensidad de la señal para proporcionar dicha señal de salida a dicho sistema (29) accionado por voz.
- 11Aparato, según la reivindicación 1, en el que al menos una parte de la unidad electrónica (11, E3) y de los medios de salida (13) está contenida en una unidad que está separada de dicho terminal (1,2) para los oídos, pero conectada al mismo.
- 12Aparato, según la reivindicación 1, en el que dicho elemento terminal (1,2) para los oídos comprende una sección de tapón (1) para los oídos y una sección de cierre (2) que forman un protector auditivo (1,2) para colocación en un oído y para protección de la función auditiva.
- 13Aparato, según la reivindicación 1, en el que los medios de comparación (27) de la intensidad de la señal comprenden medios para estimar la atenuación acústica del elemento terminal (1, 2) para los oídos a efectos de establecer dicha diferencia predeterminada. ES 2 295 313 T3
- 14Aparato, según la reivindicación 13, que comprende un seguidor de mínimos (60) conectado con los detectores (23, 26) para obtener estimaciones temporales exponenciales de dicha atenuación acústica.
- 15Aparato, según la reivindicación 1, en el que la unidad electrónica (11, E3) comprende medios (E3) de análisis de señales para extraer una señal de voz del elemento transductor electroacústico interior (M2), para transmisión al sistema accionado por voz.
- 16Aparato, según la reivindicación 15, que comprende medios (E3) de análisis de señales para detectar la presencia de componentes específicos del habla, tales como palabras, en la señal desde el elemento transductor electroacústico interior (M2), a efectos de formar órdenes, instrucciones o código para el sistema (29) accionado por voz.
- 17Aparato, según la reivindicación 1, en el que el elemento o elementos transductores electroacústicos (M1, M2) es o son micrófonos.
- 18Método para detectar una voz y para controlar un sistema accionado por voz, que utiliza un elemento terminal (1, 2) para los oídos a efectos de proteger el oído al proporcionar atenuación acústica y que comprende las siguientes etapas;detectar (M2) la intensidad de la señal acústica en el lado interior de dicho elemento terminal para los oídos, detectar (M1) la intensidad de la señal acústica en el lado exterior de dicho elemento terminal para los oídos, obtener (23, 26, 27) un valor de diferencias que representa la diferencia en la intensidad de la señal acústica entre el lado interior y exterior de dicho elemento terminal para los oídos, decidir (28) si está presente una voz usando dicho valor que representa dicha diferencia y dichos criterios dados (28), proporcionar una señal de salida, dependiendo de dicha decisión, usando medios de salida (13, E12), siendo la señal de salida una entrada al sistema (29) accionado por voz.
- 19Método, según la reivindicación 18, que comprende proporcionar una señal de salida (13) en forma de una señal de control, órdenes, instrucciones, código o similares para dicho sistema (29) accionado por voz, tal como un sistema de comunicación.
- 20Método, según la reivindicación 18, que comprende escribir y leer criterios predeterminados hasta medios de almacenamiento (E8, E9, E10) y desde los mismos en el elemento terminal (1, 2) para los oídos.
- 21Método, según la reivindicación 18, que comprende determinar si dicha diferencia en la intensidad (27) de la señal acústica entre el lado interior y exterior de dicho elemento terminal (1, 2) para los oídos ha disminuido más de una cantidad dada, típicamente de 2 a 10 dB, desde una cantidad predeterminada, típicamente de -20 a -40 dB.
- 22Método, según la reivindicación 18, que comprende filtrar (22, 25) las señales desde los dos elementos transductores electroacústicos (M2, M1) usando filtros de paso de banda (22, 25).
- 23Método, según la reivindicación 18, que comprende;analizar (E3) la señal de voz para determinar sus características, tales como duración, frecuencia y amplitud, proporcionar dicha señal de salida (13) dependiendo de dichas características de voz determinadas.
- 24Método, según la reivindicación 23, que comprende;obtener un valor que representa el promedio temporal exponencial de la diferencia en el nivel de la señal acústica entre el lado interior y exterior del elemento terminal (1, 2) para los oídos.
- 25Método, según la reivindicación 18, que comprende ajustar dichos criterios según el promedio temporal exponencial de dicha diferencia. ES 2 295 313 T3
- 26Método, según la reivindicación 18, que comprende extraer información de las señales acústicas que son detectadas por los elementos transductores electroacústicos (M1, M2).
- 27Método, según la reivindicación 26, que comprende analizar la señal desde los elementos transductores electroacústicos (M1, M2) para la presencia de componentes del habla particulares, tales como palabras.
- 28Método, según la reivindicación 18, que comprende antes del uso normal, realizar una operación de calibración en la que se obtenga una estimación de la atenuación del protector auditivo al determinar la diferencia en los niveles sonoros, tal como es detectada por los dos elementos transductores electroacústicos en un período en el que no está presente voz, sino sólo ruido.
- 29Método, según la reivindicación 18, que comprende durante el funcionamiento normal, realizar una operación de verificación en la que se obtenga una estimación de la atenuación del protector auditivo al determinar la diferencia en los niveles sonoros tal como es detectada por los dos elementos transductores electroacústicos en un período en el que no está presente voz, sino sólo ruido, y en el que la atenuación obtenida se compare con un valor de atenuación almacenado predeterminado.
Independent claims29
104 paragraphs in 1 section, as filed
ES 2 295 313 T3
DESCRIPTION
Apparatus and method for voice detection and discrimination.
Invention sector
The present invention relates to an apparatus for voice detection and discrimination in a hearing protection device and, more particularly, to a VOX (voice actuated transmission / exchange) apparatus, for determining whether an acoustic voice signal is present or not in a hearing protection device.
The present invention also relates to a method for detecting a voice using the speech detection and discrimination apparatus of the invention in a hearing protection device.
Application sector
Voice-activated control is used extensively in communication systems, such as radio transceivers, intercom systems, recording equipment, etc., and in speech-based human-machine interfaces.
This invention is intended for use in noisy environments, for example in environments where some source of acoustic noise predominates, which makes it difficult to hear, or where there may be a risk of impairing hearing function. In such environments, there may be for example heavy machinery or heavy vehicle traffic in the vicinity. In other environments, there may be a large number of people, for example in sports stadiums, such as football stadiums or the like, where the public generates a lot of noise.
In particular, the primary application for the invention is in situations where it is desirable for people to wear a hearing protection device, although some means of communication are still required, for example to talk to other people or to give orders to others. voice-operated equipment.
Previous technique
Current devices for capturing a person's speech in a very noisy environment, as a basis for voice detection, represent a technological challenge, and take various forms. The most common types include:
- A microphone very close to the mouth, supported on an arm to support it. The microphone is made with a feature that highlights the near field from the mouth. This type is sometimes called "noise canceling."
- A vibration sensor in contact with the throat, which captures the vibrations of the vocal cords.
- A vibration sensor in contact with the wall of the auditory meatus, the external auditory canal, which captures the vibrations of the tissue in the head.
- A similar sensor in contact with the cheekbone.
- A microphone that picks up sound in a closed space inside the auditory meatus.
Voice detection is based on several techniques:
- Measurement of the intensity of a band-pass filtered microphone / vibration sensor signal,
- Advanced signal processing of a signal picked up by a microphone (a review of methods can be found in Bishnu S. Atal, Lawrence R. Rabiner: "A Pattern Recognition Approach to Voice-Unvoiced-Silence Classification with Applications to Speech Recognition", IEEE Transactions on Acoustics, Speech and Signal Processing, volume ASSP-24, number 3, June 1976, pages 209-212.
Today's devices often do not work properly in noisy environments. Typically, the following types of errors occur:
- The device is not activated by normal voice.
- Noise falsely activates the device in the event of no speech.
Furthermore, it is noted that US 5,577,511 discloses an occlusion meter in which two microphones are used to detect sound pressure levels outside and inside an ear canal in which an earphone is placed. . The difference between these sound pressure levels provides an indication of the degree of occlusion by the receiver.
ES 2 295 313 T3
Related requests
The present invention is a further development of the ear terminal described in international patent applications PCT / N001 / 00357, PCT / N001 / 00358, PCT / N001 / 00359, PCT / N001 / 00360 and PCT / N001 / 00361, all them pending publication.
Objectives of the invention
An object of the present invention is to provide an apparatus for speech detection and discrimination, which provides greatly improved performance in noisy environments.
Yet another object of the invention is to provide a VOX (voice actuated transmission / exchange) apparatus that is capable of detecting and discriminating a voice in a noisy environment.
Yet another object of the invention is to provide an apparatus for speech detection and discrimination, with improved speech detection capability and in which false triggering due to acoustic noise has been reduced.
In particular, yet another object of the invention is to provide a VOX (voice actuated transmission / exchange) apparatus suitable for use with electronic communication systems that are used in noisy environments.
Furthermore, another object of the present invention is to provide a method for detecting a voice in order to control a voice-operated system, by using a terminal element for the ears also intended to protect the hearing function by providing acoustic attenuation.
Characteristics of the invention
According to the invention, these objectives are achieved with an apparatus for voice detection and discrimination for controlling a voice-operated system, comprising a terminal ear protection element for the purpose of protecting the ear by providing acoustic attenuation. An inner electro-acoustic transducer element on an inner side of the ear terminal element detects a first acoustic field and provides a first electronic signal representing said first acoustic field. An outer electro-acoustic transducer element on an outer side of the ear terminal element detects a second acoustic field and provides a second electronic signal representing said second acoustic field. An electronic unit is connected with said electroacoustic transducer elements. The electronic unit comprises first comparison means for comparing said first and second electronic signals in order to obtain the difference between the two mentioned electronic signals. The electronic unit also comprises second comparison means for comparing said difference with given criteria. Output means connected to said electronic unit provide an output signal depending on said second comparison. The output signal is used as an input signal to the voice actuated system.
According to the invention, these objectives are also achieved with a corresponding method for detecting a person's own voice and for controlling a voice-operated system, which uses a terminal element for the ears in order to protect the ear by providing acoustic attenuation. The method comprises the following steps: the intensity of the acoustic signal is detected on the inner side of said terminal element for the ears using a first electroacoustic transducer element. The intensity of the acoustic signal on the outer side of the ear terminal element is detected using a second electroacoustic transducer element. A difference value is obtained which represents the difference in the intensity of the acoustic signal between the inner and outer side of the terminal element for the ear. Using the value representing the difference obtained and the criteria given, it is decided whether a voice is present. Depending on the decision, an output signal is provided using output means. The output signal is used as input to the voice actuated system.
Further preferred embodiments of the invention are defined in the independent claims.
Brief description of the drawings
Fig. 1 shows an embodiment of the invention in which voice detection is included in a hearing protection voice communication terminal in the form of an earplug.
Figure 2 is a block diagram showing the main functional units of the electronic circuits of the apparatus according to the invention.
Fig. 3 is a block diagram showing a possible signal processing arrangement according to the invention.
Fig. 4 is a block diagram showing a possible signal processing arrangement according to the invention, in which the input of decision reference signals is obtained from a minimum follower.
Figure 5 is a block diagram of the two stages of decision block processing.
ES 2 295 313 T3
Figure 6 shows a possible frequency response characteristic of band pass filters.
Fig. 7 is a block diagram showing a preferred embodiment of signal detection.
Figure 8 shows a typical sound attenuation characteristic of a hearing protector with a polymer foam sealing section and active sound control (ANR).
Figure 9 shows the operation of a minimum tracker.
Detailed description of the invention
Figure 1 shows an embodiment according to the invention in which voice detection and discrimination are included in a hearing protection voice communication terminal (1), (2) based on an earplug. The earplug (1), (2) comprises a main section (1) containing two electroacoustic transducer elements (M1) and (M2) and a sound generator (SG). The main section (1) is designed to provide a comfortable and secure placement in the concha (the bowl-shaped cavity at the entrance to the ear canal). This can be achieved by using individually molded ear pieces (1), which are held in place by the outer ear or by the application to at least part of the ear plug (1), (2) of a surrounding pressure. flexible against the outer ear structure. A sealing section (2) is attached to the main section. The sealing section (2) can be an integral part of the ear plug (1), (2), or it can be interchangeable. The sound input of the electroacoustic transducer element (M1) is connected to the outside of the plug (1), (2) for the ears, capturing external sounds. The electroacoustic transducer element (M2) is connected to the inside of the auditory meatus (3) by means of an acoustic transmission tube (T1). The acoustic transmission tube (T1) may contain additional and optional acoustic filtering elements. The filter elements can consist, for example, of acoustic resistance elements in the form of porous or sintered insert elements, and / or acoustic adaptation elements in the form of small cavities, separately or in combination. An outlet of the sound generator (SG) is opened into the interior part of the auditory meatus (3) by means of an acoustic transmission tube (T2) between the sound generator (SG) and the part directed inwards of the sealing section (2). The acoustic transmission tube (T2) may also contain additional and optional acoustic filtering elements.
When smaller electro-acoustic transducer elements (M2) and sound generators (SG) are available, it will be possible to mount the electro-acoustic transducer element (M2) and the sound generator (SG) in the innermost part of the sealing section ( 2). There is then no need for the transmission tubes (T1) and (T2).
The two electroacoustic transducer elements (M1), (M2) and the sound generator (SG) are connected to an electronic unit (11), which can be connected to other equipment by a connection interface (13), which can transmit signals digital or analog, or both, and optionally power.
The electronic unit (11) and the power supply (12), for example a battery, can be included in the main section (1) or in a separate section.
One or both of the electroacoustic transducer elements (M1), (M2) may be, in a preferred embodiment, microphones, such as standard miniature electret microphones, similar to those used in hearing aids. Newly developed silicon microphones can also be used.
The sound generator (SG) can be based, in a preferred embodiment, on the electromagnetic or electrodynamic principle, similar to the sound generators applied in hearing aids.
The main section of the ear plug (1), (2) can be made of standard polymer materials used for normal hearing aids. The sealing part (2) may be made of a shape-retaining, slowly re-expanding, elastic polymeric foam, similar to PVC, PUR or other suitable materials for earplugs.
For some applications (with less extreme noise levels), the earplug may be molded in one piece (1), (2), combining the main section (1) and the sealing section (2). The material for this design can be a typical material used for passive earplugs (elazin, acryl).
It is also possible to make the plug (1), (2) for the ears in a piece comprising the main section (1) and the sealing section (2), all made of the polymer foam mentioned above, but then the tubes ( T1), (T2), (T3) must be made of a wall material that prevents the channels (T1), (T2), (T3) from collapsing when the sealing section (2) is inserted into the meatus auditory (3). The wall material should be non-porous, typically a plastic or rubber material with sufficient rigidity to keep the tube open while allowing the tube to bend to conform to the geometry of the auditory meatus.
When a user wears the apparatus according to the invention in a noisy environment, there will normally be a significant difference in sound, also called the intensity of the acoustic field or acoustic signal between the interior and
ES 2 295 313 T3 outside the apparatus, detected by the two electroacoustic transducer elements (M1) and (M2). When the wearer of the hearing protector speaks, the signal from their own voice produces a signal at the inner electro-acoustic transducer element (M2). Then the difference in signal intensity, detected by the two electroacoustic transducer elements (M1), (m2) decreases. When such a decrease in signal difference is detected, it is interpreted as that a voice signal is present, and the output means (13) generates an output for controlling a voice-actuated system (29).
In many applications of an apparatus according to the invention, the surrounding pressure can vary during use. Consequently, a pressure equalization is required between the two sides of the earplug system. This is achieved by using a very thin conduit (T3), (T4) or a valve that compensates for static pressure differences, while maintaining strong low-frequency sound attenuation. A safety valve (V) can be incorporated to take care of rapid decompression in the pressure compensation system (T3), (T4). The pressure compensation means (T3), (T4), (V) are not strictly required in a basic embodiment of the apparatus according to the invention, but may be an optional feature included in another embodiment of the invention.
Figure 2 is a block diagram showing the main functional units of the electronic circuits of the apparatus according to the invention. The electroacoustic transducer element (M1) picks up ambient sound. A signal from the electroacoustic transducer element (M1) is amplified in the amplifier (E1) and, in a basic embodiment of the invention, provided directly to the signal processing unit (E3). The signal from the electroacoustic transducer element (M1) in a preferable alternative embodiment of the invention is sampled and digitized in an analog-to-digital converter (E2) and fed to a processing unit (E3), which may be a digital signal processor. (DSP), a microprocessor (pP), or a combination of both. A signal from the electroacoustic transducer element (M2), which captures the sound in the auditory meatus (3) between the isolation section (2) and the eardrum (4), is amplified in the amplifier (E4). The amplified signal can be provided directly to the processing unit (E3) or it can be sampled and digitized in the analog-to-digital converter (E5) before being fed to the processing unit (E3).
In the event that the apparatus according to the invention is being used to control a voice-operated system (29) in the form of a communication system, for example using the techniques described in the aforementioned related applications, it will be useful to include a function of lock, as explained below.
An incoming communication signal can be input to the processing unit (E3) through the digital interface (E12). This communication signal is converted into analog form in the digital to analog converter (E7) and fed to the analog output amplifier (E6) that drives the sound generator (SG). The sound signal produced by the loudspeaker (SG) is fed to the eardrum (4) through the conduit (T2), into the auditory meatus (3).
When the incoming communication signal is input to the same terminal that is used for voice activated control, it is necessary to apply a blocking function in the form of an additional decision status signal to the decision process. This additional decision status signal is typically dependent on the incoming communication signal. The additional decision status signal will prohibit or block the detection of the incoming communication signal, such as if it were the users' own voice, during the time periods when the incoming communication signal is active. This detection prohibition signal is applied to the decision block (28) and to the corresponding detail block (40), in Figures 3 and 5, respectively.
However, an incoming communication signal can be input into a further hearing protection terminal (1), (2) located on the ear opposite the one that houses the terminal used for voice activated control. In this case, the above-mentioned locking function is not strictly required.
The processing unit (E3) is connected to storage media which can be RAM (random access memory) (E8), ROM (read-only memory) (E9) or EEPROM (electrically erasable programmable read-only memory ) (E10), or combinations thereof. Memories E8, E9 and E10 are used, in a preferred embodiment of the invention, to store computer programs, filter coefficients, analysis data and other relevant data.
The storage means (E8), (E9), (E10) typically contain the criteria to be used by the processing unit (E3) during the operation of the device. The criteria may typically comprise data provided during assembly of the apparatus, data provided as part of a calibration procedure, possibly associated with particular users, data obtained from adaptive processes during operation of the device, or input data, for example provided by a user. , through the digital interface (E12).
The electronic circuits (11) can be connected to other electrical units by an interface, such as a bi-directional digital interface (E12). Communication with other electrical units can be done by cable or wirelessly through a digital radio link. The Bluetooth standard for short-range digital radio is a possible wireless communication candidate for this digital interface (E12).
ES 2 295 313 T3
In a preferred embodiment of the invention, the signals that can be transmitted through this interface are:
- A program code for the processing unit (E3).
- Analysis data from the processing unit (E3).
- Synchronization data when using two terminals (1), (2) for the ears in a binaural mode.
- Audio signals digitized in both directions to and from a terminal (1), (2) for the ears.
- Control signals to control the operation of the ear terminal.
- Digital measurement signals to diagnose the behavior of the ear terminal.
Based on the signals received by the electronic circuits (11) through communication with other electrical units, on the signals stored in the terminal element itself (1), (2) for the ears, or on the signals detected by the electroacoustic transducer elements (M1), (M2), the signal processor (E3) can generate an output signal for the sound generator (SG). In the digital version of the invention, the digital signal generated in the processing unit (E3) is converted into analog form in the digital to analog converter (E7) and fed to the analog output amplifier (E6) that drives the loudspeaker (SG ). The sound signal produced by the loudspeaker (SG) is fed to the eardrum (4) through the conduit (T2), into the auditory meatus (3), as previously described.
A manual control signal can be generated in the manual control unit (E11) and fed to the processing unit (E3). The manual control signal can be generated by operating buttons, switches, etc., and can be used to start and stop the appliance, to change the operating mode, etc. In an alternative embodiment, a speech signal may constitute the control signals for the processing unit (E3). In this case, the detected speech signal would typically be compared to a predetermined, for example, prerecorded, stored representation of a speech signal, such as a digital recording. A manual control signal can also be provided by a remote unit supplying output signals adapted to be received by the apparatus, for example via interface (13).
The electrical circuits are powered by the power source (12a), which can be a primary or rechargeable battery arranged in the earplug or in a separate unit, or they can be powered by a connection to other equipment, for example a radio. Communication.
The block diagram in figure 3 shows a possible signal processing arrangement according to the invention. Signal processing can be completely analog, or amplified signals from (M1) and (M2) can be converted from analog to digital, and all signal filtering and processing can be done in the digital domain, as shown in figure 2. The signal from the electroacoustic transducer element (M1) is amplified in the amplifier (21) and filtered in a band pass filter (22) before being fed to a signal intensity detector. Likewise, the signal from the electroacoustic transducer element (M2) is amplified in the amplifier (24) and filtered in a band pass filter (25) before being fed to a signal intensity detector (26). The outputs from the two signal intensity detectors (23), (26) are compared in the signal intensity comparison unit (27), which also provides as an output a difference signal representing the difference in the intensity of the signal from the two electroacoustic transducer elements.
This output of the difference signal is input to the decision block (28). The difference signal has a negative value in dB. When the difference signal is less negative than a certain limit, the decision is made that the user of the equipment is speaking. When the difference signal is more negative than this limit, the user is considered not to be speaking. The difference signal is typically in the range of -20 to -40 dB. Typically, a suitable value of the limit will be in the range of 2 to 10 dB less negative than a typical value of the difference signal only in the case of noise, for any single apparatus.
The limit is typically stored in the storage means E8, E9, E10 of the apparatus and can be a predetermined value for any individual apparatus. The limit is entered as a decision reference for decision block (28), as indicated in figure 3.
However, the limit can be generated during the use of the device to accommodate the slow drift in the performance of the device. An adaptive process can be carried out in the processing unit (E3) in order to obtain this adapted limit. This limit could also be generated in a calibration procedure carried out at regular intervals.
Similarly, the limit can be generated during use of the device to accommodate individual differences between users of the device.
In an alternative, the input of decision reference signals can be obtained using a minimum follower (60), as shown in Figure 4. The minimum follower takes an input signal from the output of the comparison block ( 27) of the signal strength.
ES 2 295 313 T3
A possible characteristic of band pass filters is shown in figure 6. The diagram shows the frequency response with upper and lower cutoff frequencies of 150 Hz and 700 Hz, respectively, a high-pass slope of +18 dB / octave, and a low-pass slope of -6 dB / octave. This frequency characteristic selects a frequency range in which the intensity of the user's voice signal in the closed space inside the auditory meatus is high. At the same time, it suppresses low-frequency noise that can be dominant in typical environments (vehicles, factories, etc.).
A preferred embodiment of signal intensity detection (23), (26) is illustrated in Figure 7. The band pass filtered signal from each of the band pass filters (23), (25) is rectified in a rectifier unit (31) and passed to a low pass filter (32). A suitable low-pass filter time constant is 10 ms. The output signal of the low pass filter (32) is provided at the input of a logarithmic converter (33). The logarithm of the signal from (32), in a digital version of the invention, is calculated in the comparison block (27) and is provided at the output thereof. In an analog version, the logarithm of the signal from (32) is obtained using a logarithmic analog converter (33). The outputs from the logarithmic converters (33) are fed to the signal strength comparison unit (27).
In the signal intensity comparison unit (27) a momentary exponential subtraction of the logarithmic value that is obtained from the detector (23) of the signal intensity can be performed, from the logarithmic value that is obtained from the detector ( 26) of the signal strength.
If said signal subtraction is performed when there is no voice signal produced by the user, the subtraction result provides a measure of the attenuation of the hearing protection function of the apparatus (1), (2) according to the invention.
Furthermore, if said momentary subtraction is carried out before the use of the apparatus, it is possible to obtain a calibration value, which is a typical attenuation value of the attenuation of the hearing protection function of the apparatus (1), (2) according to the invention . Preferably, said calibration value would be stored internally in storage means E8, E9, E10 in the apparatus. In some applications, such a calibration could be performed for an intended user of the apparatus.
When said calibration operation has been performed for an apparatus and for a particular user, which is a subsequent signal subtraction performed when there is no voice signal produced by the user, the result of the subtraction will be a verification of the continued correct operation of the sound attenuation function of the hearing protection function of the apparatus (1), (2) according to the invention. Correct operation is verified if the result of the subtraction is approximately equal to the value obtained in the calibration operation.
If a calibration operation is performed in a controlled environment with a controllable noise signal generator, it is possible to obtain a calibrated frequency-dependent characteristic attenuation of the hearing protection function of the apparatus (1), (2) according to the invention.
The decision block (28) makes a decision based on the fact that the difference in signal intensities (calculated at 27) between the electroacoustic transducer elements (M1) and (M2) only due to external sounds is independent of the sound character and the sound level. This difference only depends on the sound attenuation properties of the hearing protector. These properties are normally independent of the intensity of the sound, but depend on the frequency of the sound. A typical sound attenuation characteristic of a hearing protector with a polymer foam sealing section and active sound control (ANR) is shown in Figure 8.
When the user of the hearing protector speaks, who is in a noisy environment, the signal of his voice produces a strong signal in the electroacoustic transducer element (M2) (especially in the frequency range 100 Hz to 1 kHz) and is lowered the difference, as measured by block (27) in Figure 3. When the frequency-dependent attenuation of the hearing protector is known, a very accurate operation of the voice detector is achieved compared to a detector that is based solely on one level of the signal. The detector can operate over a limit that is only a few dB lower than the attenuation of the hearing protector, without risking false detection due to external noise.
The attenuation characteristic of the hearing protector can be known a priori, or it can be estimated before normal use or during use, as explained above. An exponential temporal estimation of the attenuation can be made by using a minimum tracker (60) on the signal from the signal intensity comparison unit (27) in Figure 3.
The result of the signal strength comparison (28) will typically be a constant or slowly varying signal mainly due to the characteristics of the hearing protector (1), (2), with rather short peaks added by the voice signal when the user speaks. Speech peaks will typically have a duration of 10 to 30 milliseconds. The interval between the peaks will typically be 10 to 500 milliseconds long. The decision block (28) can consequently contain two signal processing stages, the first stage (40) being a momentary decision comparing the composite input signal with a reference signal, and the second stage (41 ) a timer that can be reactivated with a fixed delay typically 500 milliseconds. The end of
ES 2 295 313 T3 delay that can be reactivated is bridging the interval between speech peaks. The output from the reactivatable delay constitutes the final signal signifying the presence of speech.
Figure 9 shows the operation of the minimum follower (60). The minimum follower (60) provides the output with the minimum value of the input signal as a function of time. When the input signal is greater than the output signal, the output value slowly increases upward until it reaches the input value.
The minimum follower (60) is preferably implemented as a digital filter with a first input (50) to a minimum (min) function block (52) having an output (51). In a feedback loop of the min function block (52), a sample delay block (54) provides a sample delay (1 / z), a multiplication block (53) with a multiplication factor (1+ δ) provides a multiplication function, the output of which provides a second input to the min function block (52). A minimum value input (55) set to a minimum value (ε) provides a third input to the min function block (52). The min function block (52) outputs the smallest value of its three inputs, as indicated in Figure 9.
The minimum value (ε) provided to the minimum value input (55) is used to prevent the minimum follower (60) from dropping to a value of zero. If the minimum follower (60) enters the null value, it will not recover from this value. The initial value in the delay should therefore be (ε). The range of the positive constants (δ) and (ε) depends on the number format (integer or floating point) and the number of bits in the digital filter implementation.
Due to the syllabic nature of speech, the output of the minimum tracker (60) represents a measure of the acoustic attenuation of the hearing protector. The limit in the decision block (28) can then be adjusted according to the exponential time estimate of the acoustic attenuation.
Part of the electronic unit (11), (E3) and of the output means (13) may be contained in a unit that is separate from said terminal (1), (2) for the ears, but connected to it. In some situations, the use of signal processing units located in a separate unit may be required, for example due to limited space in the terminal element for the ears, or due to additional processing functions in auxiliary signal processing units. The output means may require, in some applications, radio transmitters with undesirable output power sources near the users head. In this case, part of the output means could be arranged at a certain distance away from the head, but connected via a suitable communication interface.
The processing unit (E3) may comprise signal analysis means for detecting the presence of speech components, such as words, in the signal from the inner electro-acoustic transducer element (M2). After detecting certain components in the signal, particular commands, instructions or code can be searched from the storage means E8, E9, E10 for transmission to the voice-operated system using the output means 13. The signal analysis means may also comprise means for determining the duration, frequency content and amplitude of the signal from the inner electro-acoustic transducer element (M2). In particular, the signal analysis means comprise means for separating the speech signal from the total signal detected by the inner electro-acoustic transducer element. Signal analysis is typically performed as software modules that perform a combination of signal processing functions, such as digital filtering.
9 sheets
Sheet 1 Sheet 2 Sheet 3 Sheet 4 Sheet 5 Sheet 6 Sheet 7 Sheet 8 Sheet 9
14 members in 9 offices
Priority claims4
| Document | Office | Kind | Date |
|---|---|---|---|
| 0200084 | Norway | W | |
| 0200084 | Norway | W | |
| 02703994 | – | – | – |
| WO2002NO00084 | – | – | – |
Members14
| Document | Office | Kind | |
|---|---|---|---|
| CA2474792A1 | Canada | A1 | |
| US2003165246A1 | United States of America | A1 | |
| WO03073790A1 | World Intellectual Property Organization (WIPO) | A1 | |
| AU2002237590A1 | Australia | A1 | |
| US6728385B2 | United States of America | B2 | |
| EP1479265A1 | European Patent Office (EPO) | A1 | |
| EP1479265B1 | European Patent Office (EPO) | B1 | |
| AT380450T | Austria | T | |
| ATE380450T1 | Austria | T1 | |
| DE60223945D1 | Germany | D1 | |
| DK1479265T3 | Denmark | T3 | |
| ES2295313T3This record | Spain | T3 | |
| DE60223945T2 | Germany | T2 | |
| CA2474792C | Canada | C |
Numbers
- Publication
- 2295313
- Publication, DOCDB
- 2295313
- Publication, EPODOC
- ES2295313T
- Application
- 2703994
- Application, DOCDB
- 02703994
- Application, EPODOC
- ES20020703994T
Titles2
- Spanish
- APARATO Y METODO PARA DETECCION Y DISCRIMINACION DE VOZ.
- English
- DEVICE AND METHOD FOR VOICE DETECTION AND DISCRIMINATION.
Classification
- CPC, 8
- H04R1/1016
- H04R1/1083
- H04R2420/07
- G10L25/78
- H04R2460/05
- H04R3/12
- H04R5/04
- A61F11/145
- IPC, 5
- H04R25 02
- A61F11 08
- G10L25 78
- H04R1 10
- H04R5 033