Multi-sensory speech enhancement using a speech-state model
13 claims: 6 independent, 7 dependent
- 1雑音低減音声信号の一部を表す雑音低減値の推定値を決定する方法であって、 代替センサを使用して代替センサ信号を生成するステップと、 気導マイクロホン信号を生成するステップと、 前記代替センサ信号および前記気導マイクロホン信号を使用して、周波数成分のセットのそれぞれに関する音声状態の別個の尤度を推定し、かつ前記別個の尤度を結合して前記音声状態の前記尤度を形成することによって、音声状態S t の尤度L(S t )を推定するステップと、 前記音声状態の前記尤度を使用して として前記雑音低減値 を推定するステップであって、 ここでπ S は状態についての事後値であり、 によって与えられ、 ここで、 であり、 ここで、 であり、 であり、ここで、M * は、Mの複素共役であり、X t は雑音低減値であり、Y t は前記気導マイクロホン信号のフレームtに関する値であり、B t は前記代替センサ信号のフレームtに関する値であり、 は気導マイクロホンにおけるセンサ雑音の分散値であり、 は代替センサにおけるセンサ雑音の分散値であり、 は周囲雑音であり、Gは周囲雑音への前記代替センサのチャネル応答であり、Hはクリーン音声信号への前記代替センサの前記チャネル応答であり、Sは全ての音声状態の前記セットであり、 は音声状態を与えられた雑音低減値の確率をモデル化する分布に関する分散値である、 推定するステップとを備えることを特徴とする方法。
- 2音声状態の前記尤度の前記推定値を使用して、前記気導マイクロホン信号のフレームが音声を含むかどうか判断するステップをさらに備えることを特徴とする請求項1に記載の方法。
- 3音声を含まないと判断される前記気導マイクロホン信号のフレームを使用して雑音源の分散値を決定し、前記雑音源の前記分散値を使用して前記低減値を推定するステップをさらに備えることを特徴とする請求項2に記載の方法。
- 4前のフレームの雑音低減値の推定値と現在のフレームの前記気導マイクロホン信号のフィルタリング済みバージョンとの線形結合として前記分布の前記分散値を推定するステップをさらに含むことを特徴とする請求項1に記載の方法。
- 5前記気導マイクロホン信号の前記フィルタリング済みバージョンは、周波数に依存するフィルタを使用して形成されることを特徴とする請求項4に記載の方法。
- 6前記気導マイクロホン信号の前記フィルタリング済みバージョンは、信号対雑音比に依存するフィルタを使用して形成されることを特徴とする請求項4に記載の方法。
- 7前記雑音低減値の前記推定値を使用することによって反復を実施して、前記雑音低減値の新しい推定を形成するステップをさらに備えることを特徴とする請求項1に記載の方法。
- 8コンピュータ によって実行されると コンピュータ に以下のステップを実行させるコ プログラム を 記録した コンピュータ 読み取り可能な 記憶媒体であって、 コンピュータに、 代替センサを使用して生成された代替センサ信号を 生成 するステップと、 気 導マイクロホン信号を 生成 するステップと、 前記代替センサ信号および前記気導マイクロホン信号を使用して、周波数成分のセットのそれぞれに関する音声状態の別個の尤度を推定し、かつ前記別個の尤度を結合して前記音声状態の前記尤度を形成することによって、音声状態S t の尤度L(S t )を推定するステップと、 前記音声状態の前記尤度を使用して として前記雑音低減値 を推定するステップであって、 ここでπ S は状態についての事後値であり、 によって与えられ、 ここで、 であり、 ここで、 であり、 であり、ここで、M * は、Mの複素共役であり、X t は雑音低減値であり、Y t は前記気導マイクロホン信号のフレームtに関する値であり、B t は前記代替センサ信号のフレームtに関する値であり、 は気導マイクロホンにおけるセンサ雑音の分散値であり、 は代替センサにおけるセンサ雑音の分散値であり、 は周囲雑音であり、Gは周囲雑音への前記代替センサのチャネル応答であり、Hはクリーン音声信号への前記代替センサの前記チャネル応答であり、Sは全ての音声状態の前記セットであり、 は音声状態を与えられた雑音低減値の確率をモデル化する分布に関する分散値である、 推定するステップと を実行させる プログラム を 記録した ことを特徴とするコンピュータ 読み取り可能な 記憶媒体。
- 9音声状態の前記尤度の前記推定値を使用して、前記気導マイクロホン信号のフレームが音声を含むかどうか判断 するステップを含むことを特徴とする請求項8に記載のコンピュータ 読み取り可能な 記憶媒体。
- 10音声を含まないと判断される前記気導マイクロホン信号のフレームを使用して雑音源の分散値を決定し、前記雑音源の前記分散値を使用して前記低減値を推定するステップ を含むことを特徴とする請求項 9 に記載 の コンピュータ 読み取り可能な 記憶媒体。
- 11前のフレームの雑音低減値の推定値と現在のフレームの前記気導マイクロホン信号のフィルタリング済みバージョンとの線形結合として前記分布の前記分散値を推定するステップをさらに含む ことを特徴とする 請求項8に記載のコンピュータ読み取り可能な記憶媒体 。
- 12前記気導マイクロホン信号の前記フィルタリング済みバージョンは、周波数に依存するフィルタを使用して形成される ことを特徴とする請求項11に記載の コンピュータ読み取り可能な記憶媒体 。
- 13前記気導マイクロホン信号の前記フィルタリング済みバージョンは、信号対雑音比に依存するフィルタを使用して形成される ことを特徴とする請求項11に記載の コンピュータ読み取り可能な記憶媒体 。
Independent claims13
153 paragraphs, as filed
The present invention relates to high quality multi-sensor voice using a voice state model.
A common problem with speech recognition and speech transmission is the corruption of speech signals due to additional noise. Specifically, audio damage from another speaker has proven difficult to detect and / or correct.
Recently, systems have been developed that attempt to eliminate noise by using a combination of an alternative sensor such as a bone conduction microphone and an air conduction microphone. Various techniques have been developed that use alternative sensor signals and airborne microphone signals to form high quality audio signals that have less noise than airborne microphone signals.
<p> However, completely high-quality voice has not been achieved, and further improvement in the formation of high-quality voice signals is required.</p>
<p> The method and device determine the likelihood of the voice state based on the alternative sensor signal and the airborne microphone signal. The likelihood of this voice state is used to estimate the clean voice value of the clean voice signal.</p>
FIG. 1 shows an example of a suitable computing system environment 100 in which embodiments of the present invention can be implemented. The computing system environment 100 is merely an example of a suitable computing environment and does not imply any limitation on the use or scope of functionality of the present invention. The computing environment 100 should not be construed as having a dependency or requirement for any one or a combination of the components shown in the exemplary operating environment 100.
Embodiments of the present invention can operate in a plurality of other general purpose or special purpose computing system environments or configurations. Examples of well-known computing systems, environments and / or configurations that may be suitable for use in embodiments of the present invention include, but are not limited to, personal computers, server computers, portable or laptop devices, multis. Includes processor systems, microprocessor-based systems, set-top boxes, programmable home appliances, network PCs, mini-computers, mainframe computers, telephone systems, distributed computing environments including any of the above systems or devices, and the like.
Embodiments of the present invention can be described in the general context of computer-executable instructions such as program modules executed by a computer. In general, a program module includes routines, programs, objects, components, data structures, etc. that perform a particular task or implement a particular abstract data type. The present invention is intended to be performed in a distributed computing environment in which tasks are performed by remote processing devices linked via a communication network. In a distributed computing environment, program modules are placed on both local and remote computer storage media, including memory storage.
Referring to FIG. 1, an exemplary system for carrying out the present invention includes a general purpose computing device in the form of a computer 110. The components of the computer 110 are not limited to that, but may include a processing device 120, a system memory 130, and a system bus 121 that connects various system components including the system memory to the processing device 120. The system bus 121 may be any of a plurality of types of bus structures, including a memory bus or memory controller, a peripheral bus and a local bus using any of the various bus architectures. For example, these architectures are not limited, but include Industry Standard Architecture (ISA) buses, Micro Channel Architecture (MCA) buses, and Enhanced ISA (EISA) buses. , Video Electronics Standards Association (VESA) Includes Association (Association) local buses and Peripheral Component Interconnect (PCI) buses, also known as mezzanine buses.
The computer 110 generally includes various computer-readable media. Computer-readable media can include any usable medium accessible by the computer 110, including both volatile and non-volatile media, removable and non-removable media. For example, but not for limitation, a computer-readable medium may include a computer storage medium and a communication medium. Computer storage media are volatile and non-volatile, removable and non-removable, implemented in any way or technology for storing information such as computer readable instructions, data structures, program modules or other data. Includes both media. Computer storage media include, but are not limited to, RAM, ROM, EEPROM, flash memory or other memory technologies, CD-ROMs, and digital versatile discs (DVDs: digital versatile). Can be used to store disk) or other optical disk storage devices, magnetic cassettes, magnetic tapes, magnetic disk storage devices or other magnetic storage devices, or desired information and can be accessed by computer 110. Any other medium is included. Communication media generally implement computer-readable instructions, data structures, program modules or other data as modulated data signals such as carrier waves or other transport mechanisms, and also include any information delivery medium. The term "modulated data signal" means a signal in which one or more of its properties have been set or modified in such a way as to encode the information in the signal. For example, not for limitation, communication media include wired media such as wired networks and direct wired connections, as well as wireless media such as acoustic, RF, infrared, and other wireless media. Any combination of the above contents should also be included within the scope of a computer-readable medium.
System memory 130 includes computer storage media in the form of volatile and / or non-volatile memory such as read-only memory 131 and random access memory 132. The basic input / output system (BIOS) 133, which contains basic routines that help transfer information between elements in the computer 110, such as at boot time, is generally stored in ROM 131. RAM 132 generally includes data and / or program modules that have immediate access to processing equipment 120 and / or are currently operating by it. For illustration purposes, not to limit, FIG. 1 shows an operating system 134, an application program 135, other program modules 136, and program data 137.
Computer 110 may also include other removable / non-removable, volatile / non-volatile computer storage media. For illustration purposes only, FIG. 1 shows a hard disk drive 141 that reads from or writes to a non-removable non-volatile magnetic medium, a magnetic disk drive 151 that reads from or writes to a removable non-volatile magnetic disk 152, and An optical disk drive 155 is shown which reads from or writes to a removable non-volatile optical disk 156 such as a CD-ROM or other optical medium. Other removable / non-removable, volatile / non-volatile computer storage media that can be used in an exemplary operating environment include, but are not limited to, magnetic tape cassettes, flash memory cards, and digital versatile discs. , Digital videotapes, solid-state RAM, solid-state ROM, etc. The hard disk drive 141 is generally connected to system bus 121 via a non-removable memory interface such as interface 140, and the magnetic disk drive 151 and optical disk drive 155 are generally connected to system bus 121 via a removable memory interface such as interface 150. Will be done.
The drives and related computer storage media discussed above and shown in FIG. 1 provide computer-readable instructions, data structures, program modules, and other data storage for the computer 110. In FIG. 1, for example, the hard disk drive 141 is shown as storing an operating system 144, an application program 145, other program modules 146, and program data 147. Note that these components may be the same as or different from the operating system 134, the application program 135, the other program modules 136 and the program data 137. The operating system 144, the application program 145, the other program modules 146, and the program data 147 are given different numbers here, at least to indicate that they are different copies.
The user can enter commands and information into the computer 110 using an input device such as a keyboard 162, a microphone 163, and a pointing device 161 such as a mouse, trackball or touchpad. Other input devices (not shown) may include joysticks, gamepads, satellite dishes, scanners, and the like. These and other input devices are often connected to the processing device 120 via a user input interface 160 coupled to the system bus, such as a parallel port, game port or universal serial bus (USB). It may be connected by other interfaces and bus structures. A monitor 191 or other type of display device is also connected to the system bus 121 via an interface such as the video interface 190. In addition to the monitor, the computer can also include other peripheral output devices such as a speaker 197 and a printer 196 which may be connected by an output peripheral interface 195.
Computer 110 operates in a networked environment using a logical connection to one or more remote computers, such as remote computer 180. The remote computer 180 can be a personal computer, a portable device, a server, a router, a network PC, a peer device, or other common network node, and generally has many or all of the elements mentioned above with respect to the computer 110. Including. The logical connections shown in Figure 1 include a local area network (LAN) 171 and a wide area network (WAN) 173, but can also include other networks. Such networking environments are common in offices, enterprise-scale computer networks, intranets, and the Internet.
When used within a LAN networking environment, computer 110 is connected to LAN 171 via a network interface or adapter 170. When used within a WAN networking environment, the computer 110 generally includes a modem 172, or other means of establishing communication over the WAN 173, such as the Internet. The modem 172, which may be internal or external, may be connected to system bus 121 via user input interface 160 or other suitable mechanism. In a networked environment, the program modules shown for computer 110, or parts thereof, may be stored in remote memory storage. For illustration purposes, not for limitation, FIG. 1 shows a remote application program 185 residing within the remote computer 180. It will be appreciated that the network connections shown are exemplary and other means of establishing communication links between computers may be used.
FIG. 2 is a block diagram of a mobile device 200, which is an exemplary computing environment. The mobile device 200 includes a microprocessor 202, a memory 204, an input / output (I / O: input / output) component 206, and a communication interface 208 for communicating with a remote computer or other mobile device. In one embodiment, the components mentioned above are combined to communicate with each other via the appropriate bus 210.
The memory 204 is a random access memory (RAM) with a battery backup module (not shown) so that the information stored in the memory 204 is not lost when the overall power to the mobile device 200 is cut off. : random access memory) implemented as a non-volatile electronic memory. One part of memory 204 is preferably allocated as addressable memory for program execution, and another part of memory 204 is preferably used for storage such as simulating storage on a disk drive. To.
Memory 204 includes operating system 212, application program 214, and object store 216. During operation, the operating system 212 is preferably run from memory 204 by processor 202. In a preferred embodiment, the operating system 212 is a WINDOWS® CE branded operating system commercially available from Microsoft Corporation. The operating system 212 is preferably designed for mobile devices and implements database functions that can be used by application 214 through a set of published application programming interfaces and methods. Objects in object store 216 are maintained by application 214 and operating system 212, at least in part, in response to calls to exposed application programming interfaces and methods.
Communication interface 208 represents a plurality of devices and technologies that allow the mobile device 200 to send and receive information. Devices include wired and wireless modems, satellite receivers, and broadcast tuners, to name a few. The mobile device 200 can also be directly connected to the computer for data exchange with the computer. In such cases, the communication interface 208 can be an infrared transceiver, or a serial or parallel communication connection, all of which can transmit streaming information.
The input / output component 206 includes various input devices such as touch sensor screens, buttons, rollers and microphones, as well as various output devices including audio generators, vibrating devices and displays. The devices listed above are for illustration purposes only and not all need to be present on the mobile device 200. In addition, other I / O devices may be connected to or seen with the mobile device 200 within the scope of the present invention.
FIG. 3 shows a basic block diagram of an embodiment of the present invention. In FIG. 3, the speaker 300 produces an audio signal 302 (X) detected by the air conduction microphone 304 and the alternative sensor 306. Examples of alternative sensors include a throat microphone that measures the vibration of the user's throat, placed on or near the bone or skull of the user's face (such as the jawbone), or in the user's ear and generated by the user. It is a bone conduction sensor that detects the vibration of the skull and jaw corresponding to voice. The air conduction microphone 304 is a type of microphone commonly used to convert an audio air-wave into an electrical signal.
The air conduction microphone 304 receives ambient noise 308 (V) generated by one or more noise sources 310 and produces its own sensor noise 305 (U). Ambient noise 308 can also be detected by an alternative sensor 306, depending on the type of ambient noise and the level of ambient noise. However, in embodiments of the present invention, the alternative sensor 306 is generally less sensitive to ambient noise than the airborne microphone 304. Therefore, in general, the alternative sensor signal 316 (B) generated by the alternative sensor 306 contains less noise than the air conduction microphone signal 318 (Y) generated by the air conduction microphone 304. The alternative sensor 306 is less sensitive to ambient noise, but produces some sensor noise 320 (W).
The path from speaker 300 to alternative sensor signal 316 can be modeled as a channel with channel response H. The path from ambient noise 308 to alternative sensor signal 316 can be modeled as a channel with channel response G.
The alternative sensor signal 316 (B) and the air conduction microphone signal 318 (Y) are supplied to the clean signal estimator 322 that estimates the clean signal 324. The clean signal estimate 324 is supplied to the voice process 328. The clean signal estimate 324 may be a time domain signal or a Fourier transform vector. If the clean signal estimate 324 is a time domain signal, the voice process 328 may take the form of a listener, voice coding system or voice recognition system. If the clean signal estimate 324 is a Fourier transform vector, the speech process 328 is generally a speech recognition system, or includes an inverse Fourier transform to transform the Fourier transform vector into a waveform.
Within the clean signal estimator 322, the alternate sensor signal 316 and the microphone signal 318 are converted into the frequency domain used for the clean voice estimation. As shown in FIG. 4, the alternate sensor signal 316 and the air conduction microphone signal 318 are fed to the analog-to-digital converters 404 and 414, respectively, to generate a sequence of digital values, each of which has a frame configuration. Grouped into frames of values by vessels 406 and 416. In one embodiment, AD converters 404 and 414 sample a 16 kHz, 16 bit analog signal per sample, thereby creating 32 kilobytes of audio data per second, and frame constructors 406 and 416 every 10 milliseconds. Create each new frame containing 20 milliseconds worth of data.
Each frame of the data provided by the frame constructors 406 and 416 is transformed into the frequency domain using the Fast Fourier Transforms (FFTs) 408 and 418, respectively.
The frequency domain values of the alternative sensor signal and the air conduction microphone signal are supplied to the clean signal estimator 420, which uses the frequency domain values to estimate the clean audio signal 324.
In some embodiments, the clean audio signal 324 is transformed back into the time domain using the inverse Fast Fourier Transform 422. This creates a time domain version of the clean audio signal 324.
The present invention uses the model of the system of FIG. 3 including the voice state of clean voice to generate a high quality voice signal. Figure 5 shows the model graphically.
In the model of FIG. 5, the clean voice 500 depends on the voice state 502. The air conduction microphone signal 504 depends on sensor noise 506, ambient noise 508 and clean audio signal 500. The alternative sensor signal 510 depends on the sensor noise 512, the clean audio signal 500 when passing through the channel response 514, and the ambient noise 508 when passing through the channel response 516.
In the present invention, the model of FIG. 5 is a noisy observation value Y.<sub>t</sub>And B<sub>t</sub>Clean audio signal from X<sub>t</sub>Used to estimate multiple speech states S<sub>t</sub>Identify the likelihood of.
In one embodiment of the invention, the clean audio signal estimate and the state likelihood for the clean audio signal estimate are formed by first assuming a Gaussian distribution of noise components in this system model. Therefore,
<maths num="1"><img file="JP5000647B2_D0001.tif" /></maths>
And here, each noise component has its own variance value.
<maths num="2"><img file="JP5000647B2_D0002.tif" /></maths>
Modeled as a zero average Gaussian with, V is the ambient noise, U is the sensor noise in the airborne microphone, and W is the sensor noise in the alternative sensor. In Equation 1, g is an adjustment parameter that allows adjustment of the dispersion value of ambient noise.
Further, this embodiment of the present invention
<maths num="3"><img file="JP5000647B2_D0003.tif" /></maths>Variance value as
<maths num="4"><img file="JP5000647B2_D0004.tif" /></maths>
Model the probability of a clean audio signal, given the state as a zero-mean Gaussian with.
In one embodiment of the invention, the prior probabilities of a given state are assumed to be uniform so that all states have equal potential. Specifically, the prior probabilities are
<maths num="5"><img file="JP5000647B2_D0005.tif" /></maths>
Defined as, where N<sub>s</sub>Is the number of voice states available in the model.
In the description of the following equations for determining the estimated value of the clean speech signal and the likelihood of the speech state, all variables are modeled in the complex spectral region. Each frequency component (Bin) is processed independently of other frequency components. For ease of notation, this method is described below for a single frequency component. Those skilled in the art will recognize that the calculations are performed for each frequency component of the spectral version of the input signal. For variables that change over time, the subscript t is attached to the variable.
Noisy observation value Y<sub>t</sub>And B<sub>t</sub>Clean audio signal from X<sub>t</sub>To estimate the conditional probabilities p (X)<sub>t</sub>| Y<sub>t</sub>, B<sub>t</sub>) Is maximized, and this conditional probability is the probability of a clean voice signal when a noisy air-conducted microphone signal and a noisy alternative sensor signal are given. The clean voice signal estimate is the voice state S in this model<sub>t</sub>Because it depends on, this conditional probability is
<maths num="6"><img file="JP5000647B2_D0006.tif" /></maths>
Where {S} represents the set of all audio states, p (X)<sub>t</sub>| Y<sub>t,</sub>B<sub>t,</sub>S<sub>t</sub>= s) is the current noisy observation and X given the voice state s<sub>t</sub>Likelihood of, p (S)<sub>t</sub>= s | Y<sub>t,</sub>B<sub>t</sub>) Is the likelihood of the speech state s given a noisy observation. In the present invention, any number of possible voice states may be used, including voiced, fricative, nasal and back vowel voice states. In some embodiments, separate speech states are provided for each set of speech units, such as phonemes. However, in one embodiment, only two voice states are provided, one for voice and one for non-voice.
Under some embodiments, a single state is used for all of the frequency components. Therefore, each frame has a single voice state variable.
The term on the right side of Equation 6 is
<maths num="7"><img file="JP5000647B2_D0007.tif" /></maths>
It can be calculated as, the conditional probability of a clean voice signal given an observation value can be estimated by the coupling probability, observation value and state of the clean voice signal, given by the observation value. It is shown that the conditional probabilities of the states when they are given can be approximated by integrating the combined probabilities, observations and states of the clean voice signals over all possible clean voice values.
Using the Gaussian assumptions about the noise distribution discussed above in Equations 1-3, the coupling probabilities, observations and states of the clean audio signal are:
<maths num="8"><img file="JP5000647B2_D0008.tif" /></maths>
Can be calculated like, here,
<maths num="9"><img file="JP5000647B2_D0009.tif" /></maths>
Is the prior probability of the state given by the uniform probability distribution in Equation 5, G is the channel response of the alternative sensor to ambient noise, and H is the channel response of the alternative sensor signal to the clean voice signal. Complex terms between vertical bars, such as, | G |, indicate the magnitude of the complex value.
The alternative sensor's channel response G for background audio is estimated from the signals of air microphone Y and alternative sensor B over the last D frame that the user is not speaking. Specifically, G is
<maths num="10"><img file="JP5000647B2_D0010.tif" /></maths>
Where D is the number of frames that the user is not speaking, but the background audio is present. Here we assume that G is constant over all time frames D. In other embodiments, instead of using all D frames equally, a technique known as "exponential aging" is used so that the latest frame contributes more to the G estimation than the older frame. use.
The channel response H of the alternative sensor to the clean audio signal is estimated from the signals of the air microphone Y and the alternative sensor B over the last T frame the user is speaking. Specifically, H is
<maths num="11"><img file="JP5000647B2_D0011.tif" /></maths>
Where T is the number of frames the user is talking about. Here we assume that H is constant over all time frames T. In other embodiments, instead of using all T-frames equally, a technique known as "exponential aging" is used so that the latest frame contributes more to the estimation of H than the older frame.
State p (S<sub>t</sub>= s | Y<sub>t,</sub>B<sub>t</sub>) Conditional likelihood using the approximation of Equation 8 and the coupling probability calculation of Equation 9.
<maths num="12"><img file="JP5000647B2_D0012.tif" /></maths> Equation 12
Calculated as, it is
<maths num="13"><img file="JP5000647B2_D0013.tif" /></maths>
It can be simplified as follows.
Taking a closer look at Equation 13, the first term models in a sense the correlation between the alternative sensor channel and the airborne microphone channel, and the second term describes the observations in the air microphone channel. It is clear that the state model and the noise model are used for this purpose. The third term is simply the prior, which is a uniform distribution in one embodiment.
Given the observations calculated by Equation 13, there are two possible applications for the likelihood of the state. First, it can be used to build a voice state classifier, which can be used to determine the variance of the noise source from voice-free frames. , Can be used to classify into those with or without audio. It can also be used to provide "soft" weights when estimating clean audio signals, as further shown below.
As mentioned above, each of the variables in the above equations is defined for a particular frequency component in the complex spectral region. Therefore, the likelihood of Equation 13 is for a state associated with a particular frequency component. However, since there is only a single state variable for each frame, the likelihood of the frame's state is formed by summing the likelihood over the frequency components as follows.
<maths num="14"><img file="JP5000647B2_D0014.tif" /></maths>
Here, L (S<sub>t</sub>(f)) = p (S)<sub>t</sub>(f) | Y<sub>t</sub>(f), B<sub>t</sub>(f)) is the likelihood of the frequency component f defined in Equation 13. This product is determined over all frequency components except the DC and Nyquist frequencies. Note that if the likelihood calculation is performed in the log-likelihood region, the products in the above equation will be replaced by the sum.
The above likelihood is
<maths num="15"><img file="JP5000647B2_D0015.tif" /></maths>
Can be used to build a speech / non-speech classifier based on the likelihood ratio test so that, in Equation 15, the frame is considered to contain speech if the ratio r is greater than 0. Made, otherwise it is considered audio-free.
The likelihood of the voice state can be used to form an estimate of a clean voice signal. In one embodiment, this estimation is
<maths num="16"><img file="JP5000647B2_D0016.tif" /></maths>
Is formed using the minimum mean square estimate (MMSE) based on Equation 6 above, where E (X)<sub>t</sub>| Y<sub>t,</sub>B<sub>t</sub>) Is the expected value of the clean audio signal given the observed value, and E (X)<sub>t</sub>| Y<sub>t,</sub>B<sub>t,</sub>S<sub>t</sub>= s) is the expected value of the clean voice signal given the observed value and the voice state.
Expected value E (X) using equations 7 and 9<sub>t</sub>| Y<sub>t,</sub>B<sub>t,</sub>S<sub>t</sub>Conditional probability p (X) from which = s) can be calculated<sub>t</sub>| Y<sub>t,</sub>B<sub>t,</sub>S<sub>t</sub>= s) can be determined as follows.
<maths num="17"><img file="JP5000647B2_D0017.tif" /></maths>
by this,
<maths num="18"><img file="JP5000647B2_D0018.tif" /></maths> Equation 18
here,
<maths num="19"><img file="JP5000647B2_D0019.tif" /></maths>
<maths num="20"><img file="JP5000647B2_D0020.tif" /></maths>
The expected value of is brought, M<sup>*</sup>Is the complex conjugate of M.
Therefore, the clean audio signal X<sub>t</sub>MMSE estimates for
<maths num="21"><img file="JP5000647B2_D0021.tif" /></maths>
Given by, where π<sub>S</sub>Is the posterior for the state,
<maths num="22"><img file="JP5000647B2_D0022.tif" /></maths>
Given by, where L (S<sub>t</sub>= s) is given by Equation 14. Therefore, the estimation of a clean speech signal is partially based on the relative likelihood of a particular speech state, which provides a soft weight for estimating the clean speech signal.
In the above calculation, H was assumed to be known with high accuracy. However, in reality, H is simply known for its limited accuracy. In an additional embodiment of the invention, H is a Gaussian random variable.
<maths num="23"><img file="JP5000647B2_D0023.tif" /></maths>
Modeled as. In one such embodiment, all of the above calculations are marginalized over all possible values of H. However, this makes mathematics awkward. In one embodiment, an iterative process is used to overcome this awkwardness. During each iteration, H in equations 13 and 20<sub>0</sub>Replaced by
<maths num="24"><img file="JP5000647B2_D0024.tif" /></maths>
Is
<maths num="25"><img file="JP5000647B2_D0025.tif" /></maths>
Replaced by, here,
<maths num="26"><img file="JP5000647B2_D0026.tif" /></maths>
Is an estimate of the clean audio signal determined from the previous iteration. The clean audio signal is then estimated using Equation 21. Then this new estimate of the clean audio signal is
<maths num="27"><img file="JP5000647B2_D0027.tif" /></maths>
Is set as the new value for and the next iteration is performed. The iteration ends when the estimation of the clean audio signal is stable.
FIG. 6 provides a method of estimating a clean audio signal using the above equations. At step 600, the input speech frame is identified where the user is not speaking. These frames are then the variance values of the ambient noise.
<maths num="28"><img file="JP5000647B2_D0028.tif" /></maths>
, Alternative sensor noise variance
<maths num="29"><img file="JP5000647B2_D0029.tif" /></maths>
And air-conducted microphone noise variance
<maths num="30"><img file="JP5000647B2_D0030.tif" /></maths>
Used to determine.
Alternate sensor signals can be inspected to identify frames where the user is not speaking. If the energy of the alternative sensor signal is low, it can first be assumed that the speaker is not speaking, as the alternative sensor signal produces a background voice signal value that is much smaller than the noise signal value. The values of the air-conducted microphone signal and the alternative sensor signal of the frame without voice are stored in the buffer.
<maths num="31"><img file="JP5000647B2_D0031.tif" /></maths>
<maths num="32"><img file="JP5000647B2_D0032.tif" /></maths>
Is used to calculate the noise variance value, as in. Where N<sub>v</sub>Is the number of noise frames in speech used to form the variance value, and V is a set of noise frames when the user is not speaking.
<maths num="33"><img file="JP5000647B2_D0033.tif" /></maths>
Refers to an alternative sensor signal after the leak has been revealed.
<maths num="34"><img file="JP5000647B2_D0034.tif" /></maths>
It is calculated as, in some embodiments, as an alternative,
<maths num="35"><img file="JP5000647B2_D0035.tif" /></maths>
It is calculated as.
In some embodiments, the technique of identifying non-voice frames based on the low energy level of the alternative sensor signal is simply performed during the initial frame of training. After the initial value has been formed for the noise variance value, it may be used to determine which frame contains voice and which frame does not contain voice using the likelihood ratio of Equation 15. ..
In one particular embodiment, the estimated variance value
<maths num="36"><img file="JP5000647B2_D0036.tif" /></maths>
The value of g, which is an adjustment parameter that can be used to increase or decrease, is set to 1. This suggests complete reliability in the noise estimation procedure. In different embodiments of the invention, different g values may be used.
Dispersion value of air conduction microphone noise
<maths num="37"><img file="JP5000647B2_D0037.tif" /></maths>
Is estimated based on the observation that airborne microphones are less prone to sensor noise than alternative sensors. Therefore, the dispersion value of the air conduction microphone is
<maths num="38"><img file="JP5000647B2_D0038.tif" /></maths>
Can be calculated as
In step 602, the voice distribution value
<maths num="39"><img file="JP5000647B2_D0039.tif" /></maths>
Is estimated using a noise suppression filter with time smoothing. The suppression filter is a generalization of the spectral subtraction method. Specifically, the voice distribution value is
<maths num="40"><img file="JP5000647B2_D0040.tif" /></maths>
However,
<maths num="41"><img file="JP5000647B2_D0041.tif" /></maths>here,
<maths num="42"><img file="JP5000647B2_D0042.tif" /></maths>
Calculated as, here,
<maths num="42"><img file="JP5000647B2_D0043.tif" /></maths>
Is a clean speech estimate from the previous frame, in some embodiments τ is a smoothing factor set to .2 and α is the distortion of the speech when α> 1. Controlling the range of noise reduction so that more noise is reduced at the expense of the increase in noise, β provides a means of adding background noise that provides a minimal noise floor and masks the perceived residual music noise. In some embodiments, γ1 = 2 and γ2 = 1/2. In some embodiments, β is set equal to 0.01 for a 20 dB noise reduction in a pure noise frame.
Therefore, in Equation 28, the variance values are the weighted sum of the estimated clean audio signals of the previous frame, and the noise suppression filter K.<sub>s</sub>Obtained as the energy of the airborne microphone signal filtered by.
In some embodiments, α is selected according to the signal-to-noise ratio and the masking principle, which masking principle goes to recognition when the same amount of noise is in the high voice energy band, rather than in the low voice energy band. It is clarified that the influence of noise is reduced, and the recognition of noise in the adjacent frequency band is reduced when high voice energy is present at a certain frequency. In this embodiment, α is
<maths num="43"><img file="JP5000647B2_D0044.tif" /></maths>
Where SNR is the signal-to-noise ratio in decibels (dB) and B is the desired signal-to-noise ratio level at which noise reduction should not be performed beyond that, α.<sub>0</sub>Is the amount of noise to be removed with a signal-to-noise ratio of 0. In some embodiments, B is set equal to 20 dB.
The following definition of signal-to-noise ratio
<maths num="44"><img file="JP5000647B2_D0045.tif" /></maths>
Using the noise suppression filter of Equation 29,
<maths num="45"><img file="JP5000647B2_D0046.tif" /></maths>
become.
This noise suppression filter provides weak noise suppression for positive signal-to-noise ratios and stronger noise suppression for negative signal-to-noise ratios. In fact, for a sufficiently negative signal-to-noise ratio, all observed signals and noise are removed and the only signal present is the noise floor, which is the "not" of the noise suppression filter of Equation 33. It has been added back by the "case" branch.
In some embodiments, α<sub>0</sub>Is frequency dependent so that different amounts of noise are removed for different frequencies. In one embodiment, this frequency dependence is α<sub>0</sub>(k) = α<sub>0min</sub>+ (Α<sub>0max</sub>-α<sub>0min</sub>) k / 225 formula 34 30Hz α so that<sub>o o</sub>And 8KHz α<sub>0</sub>Formed using linear interpolation between and, where k is the number of frequency components, α<sub>0min</sub>Is the desired α at 30Hz<sub>0</sub>Is the value of<sub>0max</sub>Is the desired α at 8KHz<sub>0</sub>It is assumed that there are 256 frequency components.
After the voice variance value is determined in step 602, this variance value is used in step 604 to determine the likelihood of each voice state using equations 13 and 14 above. The voice state likelihood is then used in step 606 to determine a clean voice estimate for the current frame. As mentioned above, in embodiments where a Gaussian distribution is used to represent H, steps 604 and 606 use the latest estimates of the clean audio signal at each iteration and also address the Gaussian model of H. Iteratively used to make changes to the formulas discussed above.
Although the present invention has been described with reference to specific embodiments, one of ordinary skill in the art will recognize that changes in form and detail may be made without departing from the spirit and scope of the invention.
<figref num="1">It is a block diagram of a computing environment in which an embodiment of the present invention can be implemented.</figref><figref num="2">It is a block diagram of an alternative computing environment in which an embodiment of the present invention can be implemented.</figref><figref num="3">It is a block diagram of the general voice processing system of this invention.</figref><figref num="4">It is a block diagram of the system for high quality voice by one Embodiment of this invention.</figref><figref num="5">It is a figure which shows the model based on the voice quality improvement by one Embodiment of this invention.</figref><figref num="6">It is a flowchart for improving voice quality by one Embodiment of this invention.</figref>
72 sheets
Sheet 1 Sheet 2 Sheet 3 Sheet 4 Sheet 5 Sheet 6 Sheet 7 Sheet 8 Sheet 9 Sheet 10 Sheet 11 Sheet 12 Sheet 13 Sheet 14 Sheet 15 Sheet 16 Sheet 17 Sheet 18 Sheet 19 Sheet 20 Sheet 21 Sheet 22 Sheet 23 Sheet 24 Sheet 25 Sheet 26 Sheet 27 Sheet 28 Sheet 29 Sheet 30 Sheet 31 Sheet 32 Sheet 33 Sheet 34 Sheet 35 Sheet 36 Sheet 37 Sheet 38 Sheet 39 Sheet 40 Sheet 41 Sheet 42 Sheet 43 Sheet 44 Sheet 45 Sheet 46 Sheet 47 Sheet 48 Sheet 49 Sheet 50 Sheet 51 Sheet 52 Sheet 53 Sheet 54 Sheet 55 Sheet 56 Sheet 57 Sheet 58 Sheet 59 Sheet 60 Sheet 61 Sheet 62 Sheet 63 Sheet 64 Sheet 65 Sheet 66 Sheet 67 Sheet 68 Sheet 69 Sheet 70 Sheet 71 Sheet 72
Every citation, both ways
| Document | Relation | Office |
|---|---|---|
| JP2005157354A | Cites | Japan |
| JP2004102287A | Cites | Japan |
| JP2000330597A | Cites | Japan |
| JP1115491A | Cites | Japan |
| JP4505670A | Cites | Japan |
| JP2004503983A | Cites | Japan |
22 members in 11 offices
Priority claims9
| Document | Office | Kind | Date |
|---|---|---|---|
| 11168770 | United States of America | – | |
| 16877005 | United States of America | A | |
| 16877005 | United States of America | A | |
| 2006022863 | United States of America | W | |
| 2006022863 | United States of America | W | |
| 2005168770 | – | – | – |
| 2006022863 | – | – | – |
| US20050168770 | – | – | – |
| WO2006US22863 | – | – | – |
Members22
| Document | Office | Kind | |
|---|---|---|---|
| US2006293887A1 | United States of America | A1 | |
| WO2007001821A2 | World Intellectual Property Organization (WIPO) | A2 | |
| MX2007015446A | Mexico | A | |
| EP1891624A2 | European Patent Office (EPO) | A2 | |
| KR20080019222A | Republic of Korea | A | |
| JP2009501940A | Japan | A | |
| WO2007001821A3 | World Intellectual Property Organization (WIPO) | A3 | |
| RU2007149546A | Russian Federation | A | |
| EP1891624A4 | European Patent Office (EPO) | A4 | |
| CN101606191A | China | A | |
| US7680656B2 | United States of America | B2 | |
| BRPI0612668A2 | Brazil | A2 | |
| EP1891624B1 | European Patent Office (EPO) | B1 | |
| AT508454T | Austria | T | |
| ATE508454T1 | Austria | T1 | |
| RU2420813C2 | Russian Federation | C2 | |
| DE602006021741D1 | Germany | D1 | |
| CN101606191B | China | B | |
| JP5000647B2This record | Japan | B2 | |
| JP2012155339A | Japan | A | |
| KR101224755B1 | Republic of Korea | B1 | |
| JP5452655B2 | Japan | B2 |
17 legal events, as the office reported them to INPADOC
Over the term
Point at a mark for the eventEvents
| Event | Code | |
|---|---|---|
| Cancellation because of no payment of annual feesLAPS | LAPS | |
| Receipt of annual feesJAPANESE INTERMEDIATE CODE: R250R250 | R250 | |
| Receipt of annual feesJAPANESE INTERMEDIATE CODE: R250R250 | R250 | |
| Receipt of annual feesJAPANESE INTERMEDIATE CODE: R250R250 | R250 | |
| Written notification of registration of transferJAPANESE INTERMEDIATE CODE: R350R350 | R350 | |
| Request for change of ownership or part of ownershipJAPANESE INTERMEDIATE CODE: R313113S111 | S111 | |
| Renewal fee payment (event date is renewal date of database)FPAY | FPAY | |
| Certificate of patent or registration of utility modelJAPANESE INTERMEDIATE CODE: R150R150 | R150 | |
| First payment of annual fees (during grant procedure)JAPANESE INTERMEDIATE CODE: A61A61 | A61 | |
| Written decision to grant a patent or to grant a registration (utility model)JAPANESE INTERMEDIATE CODE: A01A01 | A01 | |
| Written decision to grant a patent or to grant a registration (utility model)JAPANESE INTERMEDIATE CODE: A01A01 | A01 | |
| Decision of grant or rejection writtenTRDD | TRDD | |
| Written amendmentJAPANESE INTERMEDIATE CODE: A523A521 | A521 | |
| Notification of reasons for refusalJAPANESE INTERMEDIATE CODE: A131A131 | A131 | |
| Report on retrievalJAPANESE INTERMEDIATE CODE: A971007A977 | A977 | |
| Written amendmentJAPANESE INTERMEDIATE CODE: A523A521 | A521 | |
| Written request for application examinationJAPANESE INTERMEDIATE CODE: A621A621 | A621 |
Numbers
- Publication
- 5000647
- Publication, DOCDB
- 5000647
- Publication, EPODOC
- JP5000647B
- Application
- 2008519337
- Application, DOCDB
- 2008519337
- Application, EPODOC
- JP20080519337
Titles2
- Japanese
- 音声状態モデルを使用したマルチセンサ音声高品質化
- English
- High quality multi-sensor voice using voice state model
Classification
- CPC, 3
- G10L21/0208
- G10L15/20
- G10L2021/02165
- IPC, 2
- G10L21 02
- H04R1 00
