Method and system for dynamic adaptation of speech synthesizer for increasing legibility of speech synthesized by it
Abstract
FIELD: method and system for adaptation of speech synthesizer using data received in real time scale. ^ SUBSTANCE: during realization of method and system for dynamically modifying synthesized speech on basis of inputted text and a set of values of dynamic control parameters, synthesized speech is generated. Then on basis of input signal, characterizing legibility of speech by listener perceiving it, data received in real time scale are generated, on basis of which one or several values of dynamic control parameters are modified. ^ EFFECT: increased legibility of synthesized speech. ^ 3 cl, 6 dwg
Term
Term ended
Expired 7 March 2022, 4.6 years ago.
- Priority
- Filed
- Granted
- Expired
- Today
30 claims: 3 independent, 27 dependent
- 1Способ модификации синтезированной речи, заключающийся в том, что на основе вводимого текста и множества значений параметров динамического управления генерируют синтезированную речь, на основе входного сигнала, характеризующего разборчивость речи воспринимающим ее слушателем, формируют поступающие в реальном масштабе времени данные и на основе этих поступающих в реальном масштабе времени данных модифицируют одно или несколько значений параметров динамического управления, в результате чего повышается разборчивость синтезированной речи, причем, по меньшей мере, один из параметров динамического управления определяют как просодический параметр, используемый для синтеза вводимого текста.
- 2Способ по п.1, в котором поступающие в реальном масштабе времени данные формируют на основе фонового шума, присутствующего в окружающем пространстве, в котором воспроизводится синтезированная речь.
- 3Способ по п.2, в котором фоновый шум преобразуют в электрический сигнал, из базы данных, в которой хранятся модели шумовых помех, выбирают одну или несколько моделей шумовых помех и на основе электрического сигнала и моделей шумовых помех определяют характеристики фонового шума, представляя их в виде поступающих в реальном масштабе времени данных.
- 4Способ по п.3, в котором электрический сигнал для определения его временных характеристик подвергают анализу во временной области.
- 5Способ по п.3, в котором электрический сигнал для определения его частотных характеристик подвергают анализу в частотной области.
- 6Способ по п.3, в котором стадия определения характеристик фонового шума предусматривает выполнение операций, выбранных из группы, преимущественно включающей выявление в фоновом шуме помех высокого уровня, выявление в фоновом шуме помех низкого уровня, выявление в фоновом шуме кратковременных помех, выявление в фоновом шуме длительных помех, выявление в фоновом шуме изменяющихся помех, выявление в фоновом шуме постоянных помех, определение пространственного местонахождения источников фонового шума, выявление потенциальных источников фонового шума и выявление речи в фоновом шуме.
- 7Способ по п.1, в котором получают поступающие в реальном масштабе времени данные, на основе поступающих в реальном масштабе времени данных определяют релевантные характеристики синтезированной речи, имеющие соответствующие относящиеся к ним параметры динамического управления, и значения параметров динамического управления изменяют в соответствии с регулировочными значениями, внося таким путем необходимые изменения в релевантные характеристики синтезированной речи.
- 8Способ по п.7, в котором в качестве релевантных характеристик синтезированной речи изменяют релевантные характеристики, описывающие особенности говорящего.
- 9Способ по п.8, в котором в качестве релевантных характеристик синтезированной речи изменяют релевантные характеристики, описывающие особенности голоса.
- 10Способ по п.9, в котором изменяемыми характеристиками являются параметры, выбранные из группы, преимущественно включающей темп речи, тембр, громкость, параметрическую ассимиляцию звуков, частоту формант и ширину полосы частот формант, образование звуков в голосовой щели, смещение энергетического спектра речи, пол, возраст и индивидуальность.
- 11Способ по п.8, в котором в качестве релевантных характеристик синтезированной речи изменяют релевантные характеристики, описывающие стиль речи.
- 12Способ по п.11, в котором изменяемыми характеристиками являются параметры, выбранные из группы, преимущественно включающей динамическую просодию и артикуляцию.
- 13Способ по п.7, в котором в качестве релевантных характеристик синтезированной речи изменяют релевантные характеристики, описывающие эмоциональность.
- 14Способ по п.13, в котором изменяемой характеристикой является актуальность воспроизводимого в виде синтезированной речи сообщения.
- 15Способ по п.7, в котором в качестве релевантных характеристик синтезированной речи изменяют релевантные характеристики, описывающие особенности выговора.
- 16Способ по п.15, в котором изменяемыми характеристиками являются параметры, выбранные из группы, преимущественно включающей произношение и артикуляцию.
- 17Способ по п.7, в котором в качестве релевантных характеристик синтезированной речи изменяют релевантные характеристики, описывающие особенности содержащейся в синтезированной речи информации.
- 18Способ по п.17, в котором изменяемыми характеристиками являются параметры, выбранные из группы, преимущественно включающей повтор, плеоназм и лексику.
- 19Способ по п.1, в котором для создания эффекта изменения пространственного местоположения источника синтезированной речи используют полифоническую обработку звука на основе поступающих в реальном масштабе времени данных.
- 20Способ по п.1, в котором поступающие в реальном масштабе времени данные формируют на основе информации, вводимой слушателем.
- 21Способ по п.1, в котором синтезированную речь используют для воспроизведения голосовых сообщений в автомобиле.
- 22Способ модификации одного или нескольких параметров динамического управления синтезатором речи, заключающийся в том, что получают поступающие в реальном масштабе времени данные, на основе этих поступающих в реальном масштабе времени данных определяют релевантные характеристики синтезированной речи, имеющие соответствующие относящиеся к ним параметры динамического управления, и значения параметров динамического управления изменяют в соответствии с регулировочными значениями, внося таким путем необходимые изменения в релевантные характеристики синтезированной речи.
- 23Способ по п.22, в котором в качестве релевантных характеристик синтезированной речи изменяют релевантные характеристики, описывающие особенности говорящего.
- 24Способ по п.23, в котором в качестве релевантных характеристик синтезированной речи изменяют релевантные характеристики, описывающие особенности голоса.
- 25Способ по п.23, в котором в качестве релевантных характеристик синтезированной речи изменяют релевантные характеристики, описывающие стиль речи.
- 26Способ по п.22, в котором в качестве релевантных характеристик синтезированной речи изменяют релевантные характеристики, описывающие эмоциональность.
- 27Способ по п.22, в котором в качестве релевантных характеристик синтезированной речи изменяют релевантные характеристики, описывающие особенности выговора.
- 28Способ по п.22, в котором в качестве релевантных характеристик синтезированной речи изменяют релевантные характеристики, описывающие особенности содержащейся в синтезированной речи информации.
- 29Система адаптации синтезатора речи, имеющая преобразующий текст в речь синтезатор, который на основе вводимого текста и множества значений параметров динамического управления генерирует синтезированную речь, систему аудиоввода, которая на основе фонового шума, присутствующего в окружающем пространстве, в котором воспроизводится синтезированная речь, формирует поступающие в реальном масштабе времени данные, и функционально связанное с этими синтезатором и системой аудиоввода устройство управления адаптацией, которое на основе поступающих в реальном масштабе времени данных модифицирует одно или несколько значений параметров динамического управления, что обеспечивает уменьшение взаимных помех между фоновым шумом и синтезированной речью.
- 30Система адаптации по п.29, в которой система аудиоввода имеет преобразователь акустического сигнала в электрический.
Independent claims30
28 paragraphs in 5 sections, as filed
BACKGROUND OF THE INVENTION
TECHNICAL FIELD OF THE INVENTION
The present invention relates to speech synthesis. The invention relates in particular to a method and a system that allow, based on the incoming real-time data increase intelligibility of synthesized speech in the dynamic mode.
SUMMARY OF THE INVENTION
Recently, systems have been developed, the purpose of which is to improve the intelligibility of the reproduced into synthesized speech sound and to improve the perception of a listener in a variety of environments, such as inside a car in the cockpit as well as in residential and office rooms. Thus, for example, as a result of recent developments aimed at improving the performance, respectively, sound quality car audio systems have been created equalizers, which allow either manually or automatically adjust the spectral content of audio reproduced sound. Unlike conventional systems in which such an adjustment is carried out manually listener through a variety of audio controls, in more recent developments is provided for selective control of the playback of sound in the environment in which the listener. An approach based on the use of equalizers in the audio systems typically require knowledge of a significant amount of information on the conditions that are expected to prevail in the surrounding area, which will be operated by the audio system. Thus, this type of adaptation to the conditions of its audio playback is limited to regulation and audio output parameters in relation to the car is usually tied to a specific brand and model.
In addition, for many years due to air traffic control and military communications used the phonetic alphabet based in the pronunciation of words in their letters to replace the words beginning with the same letter (ie, for example, in the English letter " A "corresponds to the word" alpha ", the letter" b "corresponds to the word" bravo "," C "corresponds to the word" Charlie ", etc.), and to avoid the possibility of ambiguous interpretation of individual spoken letters in difficult conditions due. The basis of this approach, therefore, as is the theoretical assumption that the presence of noise in the communication channel and / or background noise, some sounds are inherently possess greater intelligibility compared to the other.
As another example, speech enhancement signal processing may be called a mobile or cellular telephone for reducing the aurally perceptible distortions arising from signal transmission on uplink / downlinks or via the base station. It should be noted that this approach is aimed at the elimination of distortions caused by noise in the communication channel (or noise arising in the convolutional coding of the signal), and does not allow to take into account the background (or additive) noise present in the environment in which the listener. Another example of a speech enhancement system is the conventional echo cancellation, which is generally used in a conference call.
It should also be noted that none of the above methods for improving sound reproduction does not allow to modify the synthesized speech in the dynamic mode. However, currently there is an urgent need to develop such methods dynamic modification of the synthesized speech as speech synthesis is rapidly gaining popularity, given the progress made in recent years in improving the output characteristics of speech synthesizers. However, despite all the progress made in recent years in this area progress still remains unresolved a number of problems related to speech synthesis. In particular, one such problem is that even in the development of all conventional speech synthesizers to set their control parameters to certain values must beforehand have information about conditions that are expected to prevail in the surrounding space, which will be used by the speech synthesizer . It is clear that this approach is totally inflexible and allows for the possibility of a particular speech synthesizer in a relatively limited set of environmental conditions, which can be optimal operation of a speech synthesizer. Based on the foregoing, it is desirable to develop a method and a system that would allow to modify the synthesized speech based on the received real-time data and thus improve its clarity.
This and other objects are achieved by the inventive method for modifying synthesized speech. This method consists in the fact that, based on the input text and a set of parameter values for dynamic management generate synthesized speech. Further, based on an input signal indicative of its intelligibility receptive listener, formed in the incoming real-time data. Then, in accordance with the method of the invention on the basis of the received real-time data is modified one or more parameter values of dynamic control, thereby improving the intelligibility of synthesized speech. Modifying these parameters control the speech synthesizer in a dynamic mode rather than at the stage of its development, provides a high level of adaptation, which can not be achieved with conventional approaches.
The present invention also provides a method of modifying one or more parameters dynamically control the speech synthesizer. This method consists in the fact that they receive incoming real-time data and on the basis of the received real-time data define the relevant characteristics of synthesized speech. Such relevant characteristics of the synthesized speech are appropriate, their related parameters of dynamic control. Then, in accordance with the method of the invention the dynamic control parameters is changed in accordance with set values, so by making the necessary changes in the relevant characteristics of synthesized speech.
Another object of the present invention is a speech synthesizer adaptation system having converts text to speech (TTS) synthesizer, an audio input system and a control device adaptation. This synthesizer generates synthesized speech based on the input text and a set of parameter values for dynamic management. Incoming audio input system generates real-time data based on background noise present in the environment in which the synthesized speech is played back. The control device is operatively associated with the adaptation of these synthesizer and audio in the system. Such adaptation of the control device on the basis of the incoming real-time data modifies one or more parameters of the dynamic control that provides a reduction in interference between the synthesized background noise and speech.
It is noted that the foregoing general description and the following detailed description of the invention are illustrative only and are primarily intended to illustrate the general principles and concepts underlying the invention. The accompanying drawings, the description further serve to illustrate the inventive solutions, and in accordance with it are part of the present description. These drawings, which show various features of the invention and the embodiments thereof, together with the description serve to explain the underlying principles of the invention and functional features of the proposed system in it.
BRIEF DESCRIPTION OF DRAWINGS
Various features and advantages of the present invention in more detail in the following description and claims with reference to the accompanying drawings, in which:
Figure 1 - Diagram of the inventive adaptation of a speech synthesizer,
Figure 2 - a block diagram illustrating a process of modifying synthesized speech in accordance with the present invention,
Figure 3 - a block diagram illustrating a process of forming the incoming real-time data based on the input signal according to one embodiment of the present invention,
4 - block diagram illustrating a process for determining the characteristics of the background noise and render it as the incoming real-time data according to one embodiment of the present invention,
Figure 5 - a block diagram illustrating a process of modifying one or more parameters of the dynamic control according to one embodiment of the present invention, and
6 - a diagram which shows the relevant characteristics and corresponding dynamic control parameters according to one embodiment of the present invention.
DETAILED DESCRIPTION OF PREFERRED EMBODIMENTS
1 shows a preferred embodiment made by the system 10 to adapt the speech synthesizer. Typically, such adaptation system 10 has a converting text to speech (TTS) synthesizer 12, which on the basis of the input text 16 and 42 set the values of parameters of dynamic control generates synthesized speech 14. On the basis of the background noise 22 present in a surrounding space 24 which reproduces 14 synthesized speech, audio input system 18 formed in the incoming real-time data (PRMVD) 20. With these synthesizer 12 and the system 18 is functionally connected audio input device 26 is an adaptation of the control. Such adaptation control unit 26 based on the received real-time data 20 modifies one or more values of the dynamic control parameters 42 that provides a reduction in interference between the background noise 22 and the synthesized speech 14. To transform sound vibrations into electrical audio input system 18 in the preferred embodiment It has an acoustic transducer into an electrical signal, such as a microphone.
The background noise 22 can be created by a number of different sources, some of which are shown by way of example in the drawing. Such sources of background noise that interferes with the intelligibility of speech reproduced by the synthesizer, classified by type and characteristics. For example, some sources of noise, in particular police car siren 28 and flying aircraft (not shown) create short-term high level of interference noise, typically with rapidly changing characteristics. Other sources of noise, such as operating mechanisms 30 installed on the manufacturing, and air (not shown) typically produce long constant background noise is low. The third source of noise, such as radio 32 and various household apparatus (not shown) often creates continuous noise interference, particularly in the form of music and singing with characteristics similar to those of the synthesized speech 14. The source of noise interference can be, in addition, the present 24 in the surrounding people talk to each other 34, the characteristics of speech which are practically identical to those of the synthesized speech 14. In addition, in the prevailing conditions of the surrounding area 24 can also affect the playback of the synthesized speech 14. The conditions in the surrounding space 24, and thereby exerted their influence can dynamically change over time.
It should be noted that the present invention is not limited to that shown in the figure as an example of adaptation of the system 10, which receives the real-time data 20 generated based on background noise 22 present in the surrounding space 24 where the synthesized speech is played back 14. For example, entering a real-time data 20 may also be formed on the basis of information entered by the listener 36 via an appropriate input device 19, as described in more detail below.
2 shows a block diagram 38 illustrating a process of modifying the synthesized speech. In accordance with this flow chart, at step 40 based on the input text 16 and a set of values 42 of parameters dynamically controlling generated synthesized speech. In step 44, based on the input signal 46 characterizes intelligibility listener perceives it, formed in the incoming real-time data 20. As mentioned above, the input signal source 46 may itself serve as the background noise in the environment either the listener (or other user) . However, in any case, the input signal 46 contains data relating to speech, and in accordance with this is an important source of information used to adapt the speech in the dynamic mode. At step 48 based on the received real-time data 20 modifies one or more values of the dynamic control parameters 42, thereby improving the intelligibility of synthesized speech.
As mentioned above, in one embodiment of the present invention, the incoming real-time data 20 generated based on the background noise present in the environment in which the synthesized speech is played back. Accordingly, Figure 3 illustrates a preferred process of forming the incoming real-time data 20 at step 44. According to that shown in this figure, the flowchart in step 52, the background noise 22 is converted into an electric signal 50. Then in step 54 from the corresponding base data, wherein the stored noise interference pattern (not shown) selects one or more models 56 noise. Thereafter, at step 58 based on the electric signal 50, and noise interference patterns 56 can determine characteristics of the background noise and to present them as moving in real-time data 20.
4 is a block diagram illustrating a preferred process for determining the characteristics of the background noise at step 58. According to that shown in this figure, the flowchart, first in step 60 an electric signal 50 to determine the temporal characteristics is analyzed in the time domain. The resulting analysis data 62 of change in time of the electrical signal comprise much of the information that is used when performing discussed herein operations. Similarly, at step 64 an electrical signal 50 is analyzed in the frequency domain to obtain information 66 about its frequency characteristics. It should be noted that the order of operations in steps 60 and 64 is not significant and does not affect the final result.
It should also be noted that in step 58, which identifies the characteristics of the background noise detection type provides various kinds of noise that are present in the background noise. As an example of such noise interference present in the background noise, it can be mentioned, but not limited to, high level interference, low level interference, momentary interference, long interference, varying interference and noise constant. In step 58, which identifies the characteristics of the background noise can also be provided for operations to identify potential sources of background noise, for identifying speech in the background noise and to determine the location of all such sources of background noise.
5 shows a block diagram an example of which is illustrated in more detail the preferred process of modifying the values of the dynamic control parameters 42. According to that shown in this figure, the flowchart upon obtaining step 68 incoming real-time data 20 is then based on them at the next step 70 determines relevant characteristics of synthesized speech 72. Such relevant characteristics 72 synthesized speech are appropriate, their related parameters of dynamic control. Next, at step 74 dynamic control parameter values are changed in accordance with set values, resulting in 72 relevant characteristics of synthesized speech and make the necessary changes.
Figure 6 shows in greater detail the possible relevant characteristics of synthesized speech 72 described above. Usually such relevant characteristics can be divided into 72 76 characteristics that describe the characteristics of the speaker, in the 77 characteristics that describe the emotion, the performance of 78, describing the features of a reprimand, and 79 characteristics that describe the features contained in the synthesized voice information. Specifications 76 describing the characteristics of the speaker, in turn, can be divided into 80 specifications describing the features of the voices, and the characteristics of 82, describing the particular style of speech. The parameters that affect the characteristics of 80, describing the features of voice, include, but are limited to the rate of speech, tone (fundamental frequency), loudness, parametric assimilation sounds Formant (formant frequencies and bandwidth of formants), the formation of sounds glottis, the shift of the energy spectrum of speech, gender, age and personality. The parameters that affect the 82 characteristics that describe particular style of speech, include, but are limited to, dynamic prosody (the rhythm, stress and intonation) and articulation. Thus, in particular, the intelligibility of the speech can be improved by accurate pronunciations final consonants, etc., potentially allowing better intelligibility of synthesized speech.
To attract the attention of the listener can also use the parameters related to the characteristics of 77, describing the emotions, such as the relevance of playing in a synthesized voice message. Among 78 characteristics that describe the features of a reprimand, include pronunciation and articulation (formant, etc.). It is obvious that the characteristics of 79, describing the features contained in the synthesized voice information includes parameters such as pleonasm, repetition and vocabulary. For example, the presence or absence of a pleonasm in speech is determined using phrases and words-synonyms (for example, in English, to play a voice message indicating the current time of day at 5 pm can be used the phrase "five pm" or the phrase "five o ' clock in the afternoon "(" five o'clock in the afternoon ")). Replay involves the selective repetition of certain parts of the message, reproduced using synthesized speech, in order to make a clearer focus on it contains important information. In addition, the use of a limited vocabulary and limited syntax providing simplification of the language, can also help improve intelligibility.
In respect of Figure 1, it should also be noted that to create the effect of changing the spatial location of the source of the synthesized speech 14 in conjunction with an audio output system 84 may be used polyphonic audio processing based on the received real-time data 20.
From the above description to those skilled in the art will recognize that the solution according to the invention allows the possibility of its practical implementation in various ways. Accordingly, the present invention is not limited to the specific embodiments thereof, an example of which is discussed above, and suggests the possibility of making them in various, obvious to those skilled changes and modifications on the basis of the description, claims and drawings appended hereto.
Contents5
Every citation, both ways
| Document | Relation | Office | Cited during |
|---|---|---|---|
| US9230558B2 | Cited by | United States of America | Applicant |
| US9711157B2 | Cited by | United States of America | Applicant |
| US9449606B2 | Cited by | United States of America | Applicant |
| US9236062B2 | Cited by | United States of America | Applicant |
| RU2598326C2 | Cited by | Russian Federation | Search report |
| US8983851B2 | Cited by | United States of America | Applicant |
| US9805735B2 | Cited by | United States of America | Applicant |
| US10629215B2 | Cited by | United States of America | Applicant |
| US11024323B2 | Cited by | United States of America | Applicant |
| US11869521B2 | Cited by | United States of America | Applicant |
| US9275652B2 | Cited by | United States of America | Applicant |
| US9043203B2 | Cited by | United States of America | Applicant |
| US12080305B2 | Cited by | United States of America | Applicant |
| RU2487429C2 | Cited by | Russian Federation | Search report |
| US12080306B2 | Cited by | United States of America | Applicant |
10 members in 6 offices
Priority claims3
| Document | Office | Kind | Date |
|---|---|---|---|
| 80092501 | United States of America | A | |
| 09800925 | – | – | – |
| US20010800925 | – | – | – |
Members10
| Document | Office | Kind | |
|---|---|---|---|
| US2002128838A1 | United States of America | A1 | |
| WO02073596A1 | World Intellectual Property Organization (WIPO) | A1 | |
| EP1374221A1 | European Patent Office (EPO) | A1 | |
| JP2004525412A | Japan | A | |
| CN1549999A | China | A | |
| EP1374221A4 | European Patent Office (EPO) | A4 | |
| US6876968B2 | United States of America | B2 | |
| RU2003129075A | Russian Federation | A | |
| RU2294565C2This record | Russian Federation | C2 | |
| CN1316448C | China | C |
1 legal event, as the office reported them to INPADOC
Events
| Event | Code | |
|---|---|---|
| The patent is invalid due to non-payment of feesMM4A | MM4A |
Numbers
- Publication, DOCDB
- 2294565
- Publication, EPODOC
- RU2294565
- Application
- 200312907509
- Application, DOCDB
- 2003129075
- Application, EPODOC
- RU20030129075
Titles2
- English
- METHOD AND SYSTEM FOR DYNAMIC ADAPTATION OF SPEECH SYNTHESIZER FOR INCREASING LEGIBILITY OF SPEECH SYNTHESIZED BY IT
- Russian
- СПОСОБ И СИСТЕМА ДИНАМИЧЕСКОЙ АДАПТАЦИИ СИНТЕЗАТОРА РЕЧИ ДЛЯ ПОВЫШЕНИЯ РАЗБОРЧИВОСТИ СИНТЕЗИРУЕМОЙ ИМ РЕЧИ
Classification
- CPC, 2
- G10L13/033
- G10L21/0364
- IPC, 7
- G10L13 08
- G10L15 10
- G10L13 00
- G10L13 02
- G10L13 06
- G10L19 14
- G10L21 02