Adapting to adverse acoustic environment in speech processing using playback training data
Claim Score by NHIP
Abstract
An arrangement is provided for an automatic speech recognition mechanism to adapt to an adverse acoustic environment. Some of the original training data, collected from an original acoustic environment, is played back in an adverse acoustic environment. The playback data is recorded in the adverse acoustic environment to generate recorded playback data. An existing speech model is then adapted with respect to the adverse acoustic environment based on the recorded playback data and/or the original training data.

Term
Term ended
Projected expiry passed 9 November 2023, 2.9 years ago.
- Priority and filed
- Published
- Projected expiry
- Today
27 claims: 9 independent, 18 dependent
- 1Broadest claimClaim Score 76, broad(NHIP)A method, comprising:playing back at least some of original training data, collected in a first acoustic environment, in an adverse acoustic environment to generate playback data;recording the playback data in the adverse acoustic environment to generate recorded playback data;and generating an adapted speech model from an existing speech model with respect to the adverse acoustic environment based on the recorded playback data and the original training data.
- 6A method for deriving an adapted speech model in an adverse acoustic environment, comprising:playing back at least some of original training data, collected in a first acoustic environment, in an adverse acoustic environment to generate playback data;recording the playback data in the adverse acoustic environment to generate recorded playback data;and re-training an existing speech model using the recorded playback data to generate an adapted speech model which adapts the existing speech model, derived based on the original training data with respect to the first acoustic environment, to the adverse acoustic environment.
- 8A method for deriving an adapted speech model in an adverse acoustic environment, comprising:playing back at least some of original training data, collected in a first acoustic environment, in an adverse acoustic environment to generate playback data;recording the playback data in the adverse acoustic environment to generate recorded playback data;estimating the discrepancy between the original training data and the recorded playback data;and adapting an existing speech model based on the discrepancy to generate an adapted speech model with respect to the adverse acoustic environment.
- 11A system, comprising:an original training data set collected in a first acoustic environment for training at least one existing speech model;a training data sampling and playback mechanism for selecting at least some of the original training data and for playing back the at least some of the original data in an adverse acoustic environment to generate playback data;and an adverse acoustic environment adaptation mechanism for deriving at least one adapted speech model, adapting to the adverse acoustic environment, based on the playback data and the original training data.
- 14A system for deriving an adapted speech model in an adverse acoustic environment, comprising:a training data sampling and playback mechanism for selecting at least some of original training data and for generating playback data in an adverse acoustic environment based on the selected original training data;and an adverse acoustic environment adaptation mechanism comprising a speech model re-training mechanism for re-training an existing speech model using the playback data to derive an adapted speech model with respect to the adverse acoustic environment.
- 16A system for deriving an adapted speech model, comprising:a training data sampling and playback mechanism for selecting at least some of original training data and for generating playback data in an adverse acoustic environment based on the selected original training data;and an adverse acoustic environment adaptation mechanism comprising a discrepancy based speech model adaptation mechanism for adapting an existing speech model, derived based on the original training data with respect to the first acoustic environment, to obtain an adapted speech model based on discrepancy between the original training data and the playback data.
- 18A machine-accessable medium encoded with data, the data, when accessed, causing:playing back at least some of original training data, collected in a first acoustic environment, in an adverse acoustic environment to generate playback data;recording the playback data in the adverse acoustic environment to generate recorded playback data;and generating an adapted speech model from an existing speech model with respect to the adverse acoustic environment based on the recorded playback data and the original training data.
- 23A machine-accessable medium encoded with data for deriving an adapted speech model in an adverse acoustic environment, the data, when accessed, causing:playing back at least some of original training data, collected in a first acoustic environment, in an adverse acoustic environment to generate playback data;recording the playback data in the adverse acoustic environment to generate recorded playback data;and re-training an existing speech model using the recorded playback data to generate an adapted speech model which adapts the existing speech model, derived based on the original training data with respect to the first acoustic environment, to the adverse acoustic environment.
- 25A machine-accessable medium encoded with data for deriving an adapted speech model in an adverse acoustic environment, the data, when accessed, causing:playing back at least some of original training data, collected in a first acoustic environment, in an adverse acoustic environment to generate playback data;recording the playback data in the adverse acoustic environment to generate recorded playback data;estimating the discrepancy between the original training data and the recorded playback data;and adapting an existing speech model based on the discrepancy to generate an adapted speech model with respect to the adverse acoustic environment.
Independent claims9
50 paragraphs in 4 sections, as filed
RESERVATION OF COPYRIGHT
[0001] This patent document contains information subject to copyright protection. The copyright owner has no objection to the facsimile reproduction by anyone of the patent document or the patent, as it appears in the U.S. Patent and Trademark Office files or records but otherwise reserves all copyright rights whatsoever.
BACKGROUND
[0002] Aspects of the present invention relate to automated speech processing. Other aspects of the present invention relate to adaptive automatic speech recognition.
[0003] In a society that is becoming increasingly “information anywhere and anytime”, voice enabled solutions are often deployed to provide voice information services. For example, a telephone service company may offer a voice enabled call center so that inquiries from customers may be automatically directed to appropriate agents. In addition, voice information services may be necessary to users who communicate using devices that do not have a platform on which information can be exchanged in conventional textual form. In these applications, an automatic speech recognition system may be deployed to enable voice-based communications between a user and a service provider.
[0004] Automatic speech recognition systems usually rely on a plurality of automatic speech recognition models trained based on a given corpus, consisting of a collection of speech from diversified speakers recorded in one or more different acoustic environment. The speech models built based on such given corpus capture both the characteristics of spoken words and that of the acoustic environment in which the spoken words are uttered. The accuracy of an automatic speech recognition system depends on the appropriateness of the speech models it relies on. In other words, if an automatic speech recognition system is deployed in an acoustic environment similar to the acoustic environment in which the training corpus is collected, the recognition accuracy tends to be higher than when it is deployed in a different acoustic environment. For example, if speech models are built based on a training corpus collected in a studio environment, using these speech models to perform speech recognition in an outdoor environment may result in very poor accuracy.
[0005] An important issue in developing an automatic speech recognition system that may potentially be deployed in an adverse acoustic environment involves how to adapt underlying speech models to an (adverse) acoustic environment. There are two main categories of existing approaches to adapt an automatic speech recognition system. One is to re-train the underlying speech models using new training data collected from the deployment site (or adverse acoustic environment). With this approach, both the original training corpus and speech models established therefrom are completely abandoned. In addition, to ensure reasonable performance, it usually requires a new corpus of a comparable size. This often means that a large amount of new training data needs to be collected from the adverse acoustic environment at the deployment site.
[0006] A different approach is to adapt, instead of re-training, speech models established based on an original corpus. To do so, a relatively smaller new corpus needs to be generated in an adverse acoustic environment. The new corpus is then used to determine how to adapt existing speech models (via, for example, changing the parameters of the existing models). Although less effort may be required to collect new training data, the original corpus is also put in no use.
[0007] Collecting training data is known to be an expensive operation. The need of acquiring new training data at every new deployment site not only increases the cost but also often frustrates users. In addition, in some situations, it may even be impossible. For example, if a speech recognition system is installed on a hand held device which is used by military personnel in battlefield scenarios, it may be simply not possible to re-collect training data at every locale. Furthermore, abandoning an original corpus, which is collected with high cost and effort, wastes resources.
BRIEF DESCRIPTION OF THE DRAWINGS
[0008] The present invention is further described in terms of exemplary embodiments, which will be described in detail with reference to the drawings. These embodiments are nonlimiting exemplary embodiments, in which like reference numerals represent similar parts throughout the several views of the drawings, and wherein:
[0009]FIG. 1 depicts a framework that facilitates an automated speech recognition mechanism to adapt to an adverse acoustic environment based on playback speech training data, according to embodiments of the present invention;
[0010]FIG. 2 illustrates exemplary types of automatic speech recognition models;
[0011]FIG. 3(<i>a</i>) is an exemplary segment of speech training data in waveform;
[0012]FIG. 3(<i>b</i>) is an exemplary segment of playback speech training data recorded in an adverse acoustic environment;
[0013]FIG. 3(<i>c</i>) is an exemplary discrepancy between a segment of speech training data and its playback recorded in an adverse acoustic environment, according to an embodiment of the present invention;
[0014]FIG. 4 is a high level functional block diagram of an exemplary embodiment of an adverse acoustic environment adaptation mechanism;
[0015]FIG. 5 is a high level functional block diagram of another exemplary embodiment of an adverse acoustic environment adaptation mechanism;
[0016]FIG. 6 is a flowchart of an exemplary process, in which an automatic speech recognition framework adapts to an adverse acoustic environment based on playback speech training data, according to an embodiment of the present invention;
[0017]FIG. 7 is a flowchart of an exemplary process, in which playback speech training data recorded in an adverse acoustic environment is used to generate new automatic speech recognition models, according to an embodiment of the present invention; and
[0018]FIG. 8 is a flowchart of an exemplary process, in which discrepancy between original speech training data and its corresponding playback recorded in an adverse acoustic environment and original speech training data is used to adapt existing automatic speech recognition models, according to an embodiment of the present invention.
DETAILED DESCRIPTION
[0019] The processing described below may be performed by a properly programmed general-purpose computer alone or in connection with a special purpose computer. Such processing may be performed by a single platform or by a distributed processing platform. In addition, such processing and functionality can be implemented in the form of special purpose hardware or in the form of software being run by a general-purpose computer. Any data handled in such processing or created as a result of such processing can be stored in any memory as is conventional in the art. By way of example, such data may be stored in a temporary memory, such as in the RAM of a given computer system or subsystem. In addition, or in the alternative, such data may be stored in longer-term storage devices, for example, magnetic disks, rewritable optical disks, and so on. For purposes of the disclosure herein, a computer-readable media may comprise any form of data storage mechanism, including such existing memory technologies as well as hardware or circuit representations of such structures and of such data.
[0020]FIG. 1 depicts a framework <b>100</b> that facilitates an automated speech recognition mechanism <b>180</b> to adapt to an adverse acoustic environment, according to embodiments of the present invention. The framework <b>100</b> comprises an automatic speech recognition mechanism <b>105</b>, a training data sampling and playback mechanism <b>140</b>, a re-recording mechanism <b>150</b>, and an adverse acoustic environment adaptation mechanism <b>160</b>. The automatic speech recognition mechanism <b>105</b> takes input speech <b>102</b> as input and generates output text <b>104</b> as output. The output text <b>104</b> represents the textual information automatically transcribed, by the automatic speech recognition mechanism <b>105</b>, from the input speech <b>102</b>.
[0021] The automatic speech recognition mechanism <b>105</b> includes a set of training data <b>110</b>, an automatic speech recognition (ASR) model generation mechanism <b>120</b>, a set of ASR models <b>130</b>, an automatic speech recognizer <b>180</b>, and an adaptation mechanism <b>170</b>. Automatic speech recognition is usually performed according to some pre-established models (e.g., ASR models <b>130</b>). Such models may characterize speech with respect to different aspects. For example, the linguistic aspect of speech may be modeled so that the grammatical structure of a language may be properly characterized. The pronunciation aspect of speech may be characterized using an acoustic model, which is specific to a particular language.
[0022]FIG. 2 illustrates exemplary types of automatic speech recognition models. The ASR models <b>130</b> may include a language model <b>210</b>, an acoustic model <b>220</b>, . . . , a background model <b>230</b>, and one or more transformation filters <b>240</b>. The acoustic model <b>220</b> may describe spoken words of a particular language in terms of phonemes. That is, each spoken word is represented as a sequence of phonemes. The acoustic model <b>220</b> may further include a plurality of specific acoustic models. For example, it may include a female speaker model <b>250</b>, a male speaker model <b>260</b>, or models for different speech channels <b>270</b>. Creating separate female and male speaker models may improve the performance of the automatic speech recognizer <b>180</b> due to the fact that speech from female speakers usually present different features than that of male speakers. Modeling different speech channels may achieve the same (improve performance). Speech transmitted over a phone line may present very different acoustic characteristics from that recorded in a studio.
[0023] The background model <b>230</b> may characterize an acoustic environment in which spoken words in the speeches in the training data <b>110</b> are uttered and recorded. An acoustic environment corresponds to the acoustic surroundings, which may consist of environmental sound or background noises. For example, near the lobby of a hotel, the acoustic surrounding may consist of background speech from the people in the lobby and some music played. At a beach site, acoustic surroundings may include sound of wave and people's voice in the background. The training data <b>110</b> is usually a combination of background sound and dominant speech, which may be modulated on top of the background sound. From the training data <b>110</b>, a background model may be derived from the portions of the training data <b>110</b> where dominant speech is not present.
[0024] The transformation filters <b>240</b> may be generated to map an ASR model to a new ASR model. For example, a transformation filter may be applied to an existing acoustic model to yield a different acoustic model. A transformation filter may be derived based on application needs. For example, when acoustic environment changes, some of the ASR models may need to be revised to adapt to the adverse acoustic environment. In this case, the changes with respect to the original acoustic environment may be used to generate one or more transformation filters that may be applied, during speech recognition, to provide the automatic speech recognition mechanism <b>105</b> adapted models. This will be discussed in more detail in referring to FIG. 5.
[0025] The ASR model generation mechanism <b>120</b> creates the ASR models <b>130</b> based on the training data <b>110</b>. The automatic speech recognizer <b>180</b> uses such generated models to recognize spoken words from the input speech <b>102</b>. Appropriate ASR models may be selected with respect to the input speech <b>102</b>. For example, when the input speech <b>102</b> contains speeches from a female speaker, the female speaker model <b>250</b> may be used to recognize the spoken words uttered by the female speaker. When the input speech <b>102</b> represents speech signals transmitted over a phone line, a channel model modeling transmitting medium corresponding to phone lines may be retrieved to assist the automatic speech recognizer <b>180</b> to take into account the aspects of transmitting medium during processing.
[0026] The adaptation mechanism <b>170</b> selects appropriate ASR models according to the nature of the input speech <b>102</b>. The automatic speech recognizer <b>180</b> may first determine whether the input speech <b>102</b> contains speech from male or female speakers. Such information is used by the adaptation mechanism <b>170</b> to select proper ASR models. The input speech <b>102</b> may also contain different segments of speech, some of which may be from female speakers and some from male speakers. In this case, the automatic speech recognizer <b>180</b> may segment the input speech <b>102</b> first, label the segments accordingly, and provide the segmentation results to the adaptation mechanism <b>170</b> to select appropriate ASR models.
[0027] The automatic speech recognition mechanism <b>105</b> may be deployed at different sites. A deployment site may have an acoustic environment that is substantially similar to the acoustic environment in which the training data <b>110</b> is collected. It may also correspond to an acoustic environment that is substantially different (adverse) from the original acoustic environment in which the training data <b>110</b> is collected. Without adaptation, the performance of the automatic speech recognizer <b>180</b> may depend on how similar an adverse acoustic environment is to the original acoustic environment. To optimize the performance in an adverse acoustic environment, the ASR models <b>130</b> may be adapted to reflect the change of the underlying acoustic environment. Such changes may include the noise level change in the underlying speech environment or the change in noise type when a different type of microphone is used.
[0028]FIG. 3(<i>a</i>) presents an exemplary segment of speech data collected in an original acoustic environment. The segment of speech data is illustrated in waveform. FIG. 3(<i>b</i>) presents an exemplary segment corresponding to the same speech data, but played back and recorded in an adverse acoustic environment. As seen in the plots in FIG. 3(<i>a</i>) and FIG. <b>3</b>(<i>b</i>), the portions of the waveforms corresponding to speech parts (the portion with substantially higher volume) are substantially similar. Yet the background noise seen in the two waveforms seems quite different. The difference represents the change in the adverse acoustic environment. FIG. 3(<i>c</i>) illustrates the discrepancy between the segment of speech data in FIG. 3(<i>a</i>) and its playback version recorded in the adverse acoustic environment. The discrepancy may be computed by subtracting one waveform from the other. There may be other means to compute such discrepancy. For example, the discrepancy may also be represented in the form of spectrum.
[0029] The discrepancy between the waveforms of the same speech recorded in different acoustic environments may be used to adapt the ASR models <b>130</b> (that are established based on the training data <b>110</b> recorded in an original acoustic environment) to the adverse acoustic environment. To do so, the training data sampling and playback mechanism <b>140</b> (see FIG. 1) may select part or the entire set of the training data <b>110</b> and play back the selected part of the training data <b>110</b> in the adverse acoustic environment. The re-recording mechanism <b>150</b> then records the playback training data in the adverse acoustic environment. An adverse acoustic environment may correspond to a particular physical locale in which the surrounding sound, the device used to playback the selected training data, as well as the equipment used to record the playback training data may all attribute to the adverse acoustic environment. To achieve optimal performance, the environment in which the training data is played back and re-recorded may be set up to match the environment in which the input speech <b>102</b> is to be collected. For example, a microphone that is the same type as that to be used in generating the input speech <b>102</b> may be used.
[0030] With the re-recorded playback training data, the adverse acoustic environment adaptation mechanism <b>160</b> updates the ASR models <b>130</b> so that they appropriately characterize the adverse acoustic environment. There may be different means to adapt. Specific means to adapt may depend on particular applications. FIG. 4 depicts a high level functional diagram of an exemplary implementation of the adverse acoustic environment adaptation mechanism <b>160</b>. In this exemplary embodiment, the adverse acoustic environment adaptation mechanism <b>160</b> utilizes playback training data recorded in an adverse acoustic environment to re-train relevant ASR models. For example, both the acoustic model <b>220</b> and the background model <b>230</b> may be re-trained based on the re-recorded playback training data.
[0031] In this embodiment, the adverse acoustic environment adaptation mechanism <b>160</b> may be realized with a speech model re-training mechanism <b>415</b>, which comprises at least some of an acoustic model re-training mechanism <b>420</b> and a background model retraining mechanism <b>430</b>. Both the acoustic model re-training mechanism <b>420</b> and the background model re-training mechanism <b>430</b> use the re-recorded playback training data <b>410</b> as input to train an underlying model. During the re-training process, both the acoustic model re-training mechanism <b>420</b> and the background model re-training mechanism <b>430</b> attempt to capture the characteristics of the adverse acoustic environment and accordingly generate appropriate models. The functionalities of the acoustic model re-training mechanism <b>420</b> and the background model re-training mechanism <b>430</b> may be substantially the same as that of the corresponding portion of the ASR model generation mechanism <b>120</b>, except the training data used is different. They may also invoke certain part of the ASR model generation mechanism <b>120</b> to directly utilize existing capabilities of the automatic speech recognition mechanism <b>105</b> to perform re-training tasks.
[0032] The new models generated during a re-training session are used to either replace the original models or stored as separate set of models (so that each set of models is used for a different acoustic environment).
[0033] A different means to adapt to an adverse acoustic environment is to update or revise existing ASR models, instead of re-generate such models. In this case, discrepancy between the original training data <b>10</b> and its playback version may be used to determine how relevant existing ASR models may be revised to reflect the changes in an adverse acoustic environment. FIG. 5 depicts a high level functional diagram of a different exemplary embodiment of the adverse acoustic environment adaptation mechanism <b>160</b> that adapts the ASR models to an adverse acoustic environment via updating existing models. The adverse acoustic environment adaptation mechanism <b>160</b>, in this case, may be implemented with a discrepancy based speech model adaptation mechanism <b>505</b>, comprising at least some of a discrepancy detection mechanism <b>510</b>, a transformation update mechanism <b>520</b>, an acoustic model update mechanism <b>530</b>, and a background model update mechanism <b>540</b>.
[0034] An ASR model, including an acoustic model or a background model, may be typically represented by a set of model parameters. For example, a Gaussian mixture may be used as a background model to characterize an acoustic environment (background). In this case, the parameters of the Gaussian mixture (normally including a set of mean vectors and covariance matrices) describe a distribution in a high dimensional feature space that models the distribution of acoustic features computed from the underlying acoustic environment. One type of such acoustic features is the cepstral feature computed from training data. Background models generated from different acoustic environments may be distinguishable from the difference in their model parameter values. Similarly, since an acoustic environment may also affect acoustic models, acoustic models derived with respect to one acoustic environment may differ from those derived from an adverse acoustic environment. The difference may be reflected in their model parameter values.
[0035] Two exemplary different means of adapting an existing model are described below. An existing model may be updated directly to produce an adapted model. This means that the parameters of an existing model may be revised according to some criteria. Another means to adapt an existing model is to devise a transformation (filter) to map existing model parameter values to transformed model parameter values that constitute an adapted model. For example, a linear transformation may be devised and applied to each existing model parameter to produce new model parameters. A transformation may be devised according to some criteria. Both the criteria used to revise model parameters and the criteria used to devise a transformation may be determined with respect to the acoustic environment changes. For example, the discrepancy between the original acoustic environment and the present adverse acoustic environment may be used to determine how an ASR model should adapt.
[0036] The discrepancy detection mechanism <b>510</b> computes the discrepancy between the original training data <b>110</b> (from an original acoustic environment) and the re-recorded training data <b>410</b> (from an adverse acoustic environment). Based on the discrepancy, the acoustic model update mechanism <b>530</b> revises the parameters of the existing acoustic model <b>220</b> to produce an updated (or adapted) acoustic model that is appropriate to the present adverse acoustic environment. Similarly, the background model update mechanism <b>540</b> revises the parameters of the existing background model <b>230</b> (if any) to produce an updated background model.
[0037] The first exemplary means of adapting an existing model in an adverse acoustic environment is to directly update model parameters based on given new training data recorded in an adverse acoustic environment. That is, based on the re-recorded training data <b>410</b>, the acoustic model update mechanism <b>530</b> or the background model update mechanism <b>540</b> may directly revise the parameter values of the underlying existing models based on detected discrepancy between the original training data <b>110</b> and the re-recorded training data <b>410</b>. Existing techniques may be used to achieve this task. For instance, if an existing ASR model is a Gaussian mixture, a known technique called maximum aposteriori estimation (MAP) may be used to adapt its model parameters, including mean vectors and covariance matrices, in accordance with the discrepancy between the original acoustic environment and the adverse acoustic environment.
[0038] Another approach to adapt existing model parameters is to integrate existing models with updated background models. For example, since the detected discrepancy may represent mainly the difference in two different acoustic environments, such discrepancy may be used to update an existing background model or to change the parameter values of the existing background model. Such updated background model may then be integrated with the existing acoustic model to generate an updated acoustic model corresponding to the adverse environment. Existing techniques may be applied to achieve the integration. For example, parallel model combination (PMC) approach combines a background noise model represented as a Gaussian mixture with existing noise-free acoustic models to generate models for noisy speech.
[0039] The second exemplary means of adapting an existing model in an adverse acoustic environment is to devise an appropriate transformation (e.g., transformation filters <b>240</b>), which transforms the parameter values of an existing model to produce an adapted model. Such transformation may be applied to each of the model parameters. For example, assume a Gaussian mixture is used for an existing acoustic model and each Gaussian in the mixture is characterized by a mean vector and a covariance matrix. The mean vectors and the covariance matrices of all Gaussians (model parameters) constitute an acoustic model. Here a mean vector of a Gaussian may correspond to the mean of certain features such as the cepstral features that are computed from the training data <b>110</b>.
[0040] To adapt a Gaussian mixture model with known model parameter values to another Gaussian mixture with different model parameter values, a linear transformation may be devised and applied to the mean vectors and covariance matrices of the original Gaussian mixture. The transformed mean vectors and covariance matrices effectively yield a new Gaussian mixture corresponding to a different distribution in a high dimensional feature space. Such a linear transformation may be derived according to both the existing model and the discrepancy (between the original training data <b>110</b> and the re-recorded training data <b>410</b>). Different techniques may be used to derive a linear transformation. For example, maximum likelihood linear regression (MLLR) is a known technique to derive an appropriate linear transformation based on the discrepancy between training data, recorded in one acoustic environment (e.g., the original training data <b>110</b>), and training data, recorded in an adverse acoustic environment (e.g., the re-recorded training data <b>410</b> as shown in FIG. 3).
[0041] Model parameters may also be transformed via a non-linear means. A non-linear transformation may be derived using existing techniques. For example, codeword dependent cepstral normalization (CDCN) is a technique that applies a linear transformation on original model parameters by taking into account of the minimum mean square of the cepstral features computed using a log transformation on spectral features.
[0042] Different means of adapting an existing model may achieve similar net outcome. For example, transforming model parameters to derive new parameters may yield substantially the same new model parameter values as what can be derived from directly changing the model parameters. In applications, the choice of the means to adapting model parameters may be determined based on specific system set up or other considerations.
[0043] Different mechanisms in the adverse acoustic environment adaptation mechanism <b>160</b> may be properly invoked according to the adaptation strategy. If existing model parameters are to be revised directly, the acoustic model update mechanism <b>530</b> and the background model update mechanism <b>540</b> (if any background model is present) may be invoked. The adapted model generated may be stored either separately from the corresponding existing models or used to replace the corresponding existing model.
[0044] If existing models are to be adapted through proper transformations, the transformation update mechanism <b>520</b> may be invoked. The transformation generated may be stored in the transformation filters <b>240</b>. In this case, whenever the ASR models are needed in speech processing, appropriate transformation filters may need to be retrieved first and then transformed prior to being applied in speech recognition processing. A different strategy in implementation may be to apply a derived transformation as soon as it is generated to an existing model to produce a new model, which is then stored either separately from the corresponding existing model or to used to replace the existing model.
[0045] Different mechanisms in the second embodiment of the adverse acoustic environment adaptation mechanism <b>160</b> may also work together to achieve the adaptation. For example, a derived transformation may be sent to appropriate mechanisms (e.g., either the acoustic model update mechanism <b>530</b> or the background model update mechanism <b>540</b>) so that the transformation is applied to an existing model to generate a new model. In a third embodiment of the present invention, the first embodiment of the adverse acoustic environment adaptation mechanism <b>160</b> described earlier (adaptation via re-training ASR models) may be integrated with the second embodiment so that different strategies of adaptation can be all made available. The determination of an adaptation strategy in a particular application may be dynamically made according to considerations such as the deployment setting and the specific goals of the application.
[0046]FIG. 6 is a flowchart of an exemplary process, in which the automatic speech recognition framework <b>100</b> adapts to an adverse acoustic environment based on playback training data recorded in an adverse acoustic environment, according to an embodiment of the present invention. Part or all of the training data <b>110</b> is first selected, at act <b>610</b>, for playback purposes. Such selected training data is then played back, at act <b>620</b>, and re-recorded, at act <b>630</b>, in the adverse acoustic environment to generate the re-recorded training data <b>410</b>. Based on the re-recorded training data <b>410</b>, adapted ASR models are generated, at act <b>640</b>, based on both the original training data <b>110</b> and the re-recorded training data <b>410</b>. With the adapted ASR models, when the automatic speech recognition mechanism <b>105</b> receives, at act <b>650</b>, the input speech <b>102</b> produced in the adverse acoustic environment, it performs, at act <b>660</b>, speech recognition using the adapted ASR models.
[0047]FIG. 7 is a flowchart of an exemplary process, in which one exemplary embodiment of adverse acoustic environment adaptation mechanism <b>160</b> uses playback training data recorded in an adverse acoustic environment to generate adapted ASR models, according to an embodiment of the present invention. The re-recorded training data <b>410</b> is first retrieved at act <b>710</b>. Such data is used to re-train, at act <b>720</b>, relevant ASR models. New ASR models that adapts to the adverse acoustic environment are generated from the retraining at act <b>730</b> and are used to update, at act <b>740</b>, the existing ASR models.
[0048]FIG. 8 is a flowchart of an exemplary process, in which discrepancy between playback speech training data recorded in an adverse acoustic environment and the original speech training data <b>110</b> is used to adapt the ASR models <b>130</b>, according to an embodiment of the present invention. The original training data <b>110</b> and the re-recorded training data <b>410</b> are first retrieved at act <b>810</b>. Discrepancy between the original training data <b>110</b> and the re-recorded playback training data <b>410</b> is estimated at act <b>820</b>. If appropriate transformations are to be generated, determined at act <b>830</b>, the transformation update mechanism <b>520</b> is invoked to generate, at act <b>840</b>, transformations that are capable of adapting corresponding existing ASR models to the adverse acoustic environment.
[0049] When an existing acoustic model is to be adapted, determined at act <b>850</b>, the acoustic model update mechanism <b>530</b> is invoked to update, at act <b>860</b>, the existing acoustic model. The update may include either directly revising the parameters of the acoustic model or using a transformation to map the existing acoustic model to a new adapted acoustic model. When an existing background model is to be adapted, determined at act <b>870</b>, the background model update mechanism <b>540</b> is invoked to update, at act <b>880</b>, the existing background model. Similarly, such update may include either directly revising the parameters of the existing background model or applying a transformation to the existing background model to generate an adapted background model.
[0050] While the invention has been described with reference to the certain illustrated embodiments, the words that have been used herein are words of description, rather than words of limitation. Changes may be made, within the purview of the appended claims, without departing from the scope and spirit of the invention in its aspects. Although the invention has been described herein with reference to particular structures, acts, and materials, the invention is not to be limited to the particulars disclosed, but rather can be embodied in a wide variety of forms, some of which may be quite different from those of the disclosed embodiments and extends to all equivalent structures, acts, and, materials, such as are within the scope of the appended claims.
Contents4
9 sheets
Sheet 1 Sheet 2 Sheet 3 Sheet 4 Sheet 5 Sheet 6 Sheet 7 Sheet 8 Sheet 9
Every citation, both ways
| Document | Relation | Office | Cited during |
|---|---|---|---|
| US9552825B2 | Cited by | United States of America | Search report |
| US2013096915A1 | Cited by | United States of America | Pre-grant |
| US2014278415A1 | Cited by | United States of America | Pre-grant |
| US2007078652A1 | Cited by | United States of America | Pre-grant |
| US9990936B2 | Cited by | United States of America | Search report |
| US2014316778A1 | Cited by | United States of America | Pre-grant |
| US7933771B2 | Cited by | United States of America | Search report |
| US8938388B2 | Cited by | United States of America | Applicant |
| US2008147411A1 | Cited by | United States of America | Pre-grant |
| US2004181409A1 | Cited by | United States of America | Pre-grant |
| US7533023B2 | Cited by | United States of America | Search report |
| US7337113B2 | Cited by | United States of America | Search report |
| US10154342B2 | Cited by | United States of America | Search report |
| US2010318354A1 | Cited by | United States of America | Pre-grant |
| US2017078791A1 | Cited by | United States of America | Pre-grant |
| GB2493413B | Cited by | United Kingdom | Search report |
| US2017309291A1 | Cited by | United States of America | Pre-grant |
| US2004158457A1 | Cited by | United States of America | Pre-grant |
| US2018061398A1 | Cited by | United States of America | Search report |
| GB2493413A | Cited by | United Kingdom | Search report |
| US9009039B2 | Cited by | United States of America | Search report |
| US9741341B2 | Cited by | United States of America | Applicant |
| US8131544B2 | Cited by | United States of America | Search report |
| US10283115B2 | Cited by | United States of America | Search report |
| US8972256B2 | Cited by | United States of America | Search report |
| US2009228272A1 | Cited by | United States of America | Pre-grant |
| US2004002867A1 | Cited by | United States of America | Pre-grant |
| US2012109646A1 | Cited by | United States of America | Pre-grant |
| US2003061037A1 | Cites | United States of America | Pre-grant |
| US5475792A | Cites | United States of America | Pre-grant |
| US5737485A | Cites | United States of America | Pre-grant |
| US5960397A | Cites | United States of America | Pre-grant |
| US6912417B1 | Cites | United States of America | Pre-grant |
2 members in 1 office
Priority claims2
| Document | Office | Kind | Date |
|---|---|---|---|
| 11593402 | United States of America | A | |
| US20020115934 | – | – | – |
Members2
| Document | Office | Kind | |
|---|---|---|---|
| US2003191636A1 | United States of America | A1 | |
| US7072834B2 | United States of America | B2 |
6 legal events, as the office reported them to INPADOC
Over the term
Point at a mark for the eventEvents
| Event | Code | |
|---|---|---|
| Lapse for failure to pay maintenance feesLapsedPATENT EXPIRED FOR FAILURE TO PAY MAINTENANCE FEES (ORIGINAL EVENT CODE: EXP.)LAPS | LAPS | |
| Information on status: patent discontinuationPATENT EXPIRED DUE TO NONPAYMENT OF MAINTENANCE FEES UNDER 37 CFR 1.362STCH | STCH | |
| Fee payment procedureMAINTENANCE FEE REMINDER MAILED (ORIGINAL EVENT CODE: REM.)FEPP | FEPP | |
| Fee paymentFPAY | FPAY | |
| Fee paymentFPAY | FPAY | |
| AssignmentAS | AS |
Numbers
- Publication, DOCDB
- 2003191636
- Publication, EPODOC
- US2003191636
- Application
- 10115934
- Application, DOCDB
- 11593402
- Application, EPODOC
- US20020115934
Titles
- English
- Adapting to adverse acoustic environment in speech processing using playback training data
Classification
- CPC, 3
- G10L15/065
- G10L15/20
- G10L21/0216
- IPC, 2
- G10L15 06
- G10L15 20
- USPC, 3
- 704226000
- 704E15010
- 704E15039