Speech synthesis using concatenation of speech waveforms
5 claims: 3 independent, 2 dependent
- 1A speech synthesizer comprising:a speech database (141) referencing speech waveforms;a speech waveform selector (131), in communication with the speech database, that selects waveforms referenced by the database using designators that correspond to a phonetic transcription input;and a speech waveform concatenator (151), in communication with the speech database, that concatenates waveforms selected by the speech waveform selector to produce a speech signal output, wherein, for at least one ordered sequence of a first waveform and a second waveform, the concatenator selects a location of a trailing edge of the first waveform and a location of a leading edge of the second waveform, each location being selected so as to produce an optimization of a phase match between the first and second waveforms based on similarity in shape in regions near the locations using a cross-correlation technique.
Independent claims3
69 paragraphs in 2 sections, as filed
Technical Field
0001The present invention relates to a speech synthesizer based on concatenation of digitally sampled speech units from large database of such samples and associated phonetic, symbolic, and numeric descriptors.
Background Art
0002A concatenation-based speech synthesizer uses pieces of natural speech as building blocks to reconstitute an arbitrary utterance. A database of speech units may hold speech samples taken from an inventory of pre-recorded natural speech data. Using recordings of real speech preserves some of the inherent characteristics of a real person's voice. Given a correct pronunciation, speech units can then be concatenated to form arbitrary words and sentences. An advantage of speech unit concatenation is that it is easy to produce realistic coarticulation effects, if suitable speech units are chosen. It is also appealing in terms of its simplicity, in that all knowledge concerning the synthetic message is inherent to the speech units to be concatenated. Thus, little attention needs to be paid to the modeling of articulatory movements. However speech unit concatenation has previously been limited in usefulness to the relatively restricted task of neutral spoken text with little, if any, variations in inflection.
0003A tailored corpus is a well-known approach to the design of a speech unit database in which a speech unit inventory is carefully designed before making the database recordings. The raw speech database then consists of carriers for the needed speech units. This approach is well-suited for a relatively small footprint speech synthesis system. The main goal is phonetic coverage of a target language, including a reasonable amount of coarticulation effects. No prosodic variation is provided by the database, and the system instead uses prosody manipulation techniques to fit the database speech units into a desired utterance.
0004For the construction of a tailored corpus, various different speech units have been used (see, for example, <nplcit id="ncit0001" npl-type="s"><text>Klatt, D.H., "Review of text-to-speech conversion for English," J. Acoust. Soc. Am. 82(3), September 1987</text></nplcit>). Initially, researchers preferred to use phonemes because only a small number of units was required - approximately forty for American English - keeping storage requirements to a minimum. However, this approach requires a great deal of attention to coarticulation effects at the boundaries between phonemes. Consequently, synthesis using phonemes requires the formulation.of complex coarticulation rules.
0005Coarticulation problems can be minimized by choosing an alternative unit. One popular unit is the diphone, which consists of the transition from the center of one phoneme to the center of the following one. This model helps to capture transitional information between phonemes. A complete set of diphones would number approximately 1600, since there are approximately (40)<sup>2</sup> possible combinations of phoneme pairs. Diphone speech synthesis thus requires only a moderate amount of storage. One disadvantage of diphones is that they lead to a large number of concatenation points (one per phoneme), so that heavy reliance is placed upon an efficient smoothing algorithm, preferably in combination with a diphone boundary optimization. Traditional diphone synthesizers, such as the ITS-3000 of Lernout & Hauspie Speech And Language Products N.V., use only one candidate speech unit per diphone. Due to the limited prosodic variability, pitch and duration manipulation techniques are needed to synthesize speech messages. In addition, diphones synthesis does not always result in good output speech quality.
0006Syllables have the advantage that most coarticulation occurs within syllable boundaries. Thus, concatenation of syllables generally results in good quality speech. One disadvantage is the high number of syllables in a given language, requiring significant storage space. In order to minimize storage requirements while accounting for syllables, demi-syllables were introduced. These half-syllables, are obtained by splitting syllables at their vocalic nucleus. However the syllable or demi-syllable method does not guarantee easy concatenation at unit boundaries because concatenation in a voiced speech unit is always more difficult that concatenation in unvoiced speech units such as fricatives.
0007The demi-syllable paradigm claims that coarticulation is minimized at syllable boundaries and only simple concatenation rules are necessary. However this is not always true. The problem of coarticulation can be greatly reduced by using word-sized units, recorded in isolation with a neutral intonation. The words are then concatenated to form sentences. With this technique, it is important that the pitch and stress patterns of each word can be altered in order to give a natural sounding sentence. Word concatenation has been successfully employed in a linear predictive coding system.
0008Some researchers have used a mixed inventory of speech units in order to increase speech quality, <i>e</i>.<i>g</i>., using syllables, demi-syllables, diphones and suffixes (see, <nplcit id="ncit0002" npl-type="b"><text>Hess, W.J., "Speech Synthesis - A Solved Problem, Signal processing VI: Theories and Applications," J. Vandewalle, R. Boite, M. Moonen, A. Oosterlinck (eds.), Elsevier Science Publishers B.V., 1992</text></nplcit>).
0009To speed up the development of speech unit databases for concatenation synthesis, automatic synthesis unit generation systems have been developed (see, <nplcit id="ncit0003" npl-type="b"><text>Nakajima, S., "Automatic synthesis unit generation for English speech synthesis based on multi-layered context oriented clustering," Speech Communication 14 pp. 313-324, Elsevier Science Publishers B.V., 1994</text></nplcit>). Here the speech unit inventory is automatically derived from an analysis of an annotated database of speech - i.e. the system 'learns' a unit set by analyzing the database. One aspect of the implementation of such systems involves the definition of phonetic and prosodic matching functions.
0010A new approach to concatenation-based speech synthesis was triggered by the increase in memory and processing power of computing devices. Instead of limiting the speech unit databases to a carefully chosen set of units, it became possible to use large databases of continuous speech, use non-uniform speech units, and perform the unit selection at run-time. This type of synthesis is now generally known as corpus-based concatenative speech synthesis.
0011The first speech synthesizer of this kind was presented in <nplcit id="ncit0004" npl-type="b"><text>Sagisaka, Y., "Speech synthesis by rule using an optimal selection of non-uniform synthesis units," ICASSP-88 New York vol.1 pp. 679-682, IEEE, April 1988</text></nplcit>. It uses a speech database and a dictionary of candidate unit templates, i.e. an inventory of all phoneme sub-strings that exist in the database. This concatenation-based synthesizer operates as follows <ol id="ol0001" compact="compact"><li>(1) For an arbitrary input phoneme string, all phoneme sub-strings in a breath group are listed,</li><li>(2) All candidate phoneme sub-strings found in the synthesis unit entry dictionary are collected,</li><li>(3) Candidate phoneme sub-strings that show a high contextual similarity with the corresponding portion in the input string are retained,</li><li>(4) The most preferable synthesis unit sequence is selected mainly by evaluating the continuities (based only on the phoneme string) between unit templates,</li><li>(5) The selected synthesis units are extracted from linear predictive coding (LPC) speech samples in the database,</li><li>(6) After being lengthened or shortened according to the segmental duration calculated by the prosody control module, they are concatenated together.</li></ol>
0012Step (3) is based on an appropriateness measure - taking into account four factors: conservation of consonant-vowel transitions, conservation of vocalic sound succession, long unit preference, overlap between selected units. The system was developed for Japanese, the speech database consisted of 5240 commonly used words.
0013A synthesizer that builds further on this principle is described in <nplcit id="ncit0005" npl-type="s"><text>Hauptmann, A.G., "SpeakEZ: A first experiment in concatenation synthesis from a large corpus," Proc. Eurospeech '93, Berlin, pp.1701-1704,1993</text></nplcit>. The premise of this system is that if enough speech is recorded and catalogued in a database, then the synthesis consists merely of selecting the appropriate elements of the recorded speech and pasting them together. It uses a database of 115,000 phonemes in a phonetically balanced corpus of over 3200 sentences. The annotation of the database is more refined than was the case in the Sagisaka system: apart from phoneme identity there is an annotation of phoneme class, source utterance, stress markers, phoneme boundary, identity of left and right context phonemes, position of the phoneme within the syllable, position of the phoneme within the word, position of the phoneme within the utterance, pitch peak locations.
0014Speech unit selection in the SpeakEZ is performed by searching the database for phonemes that appear in the same context as the target phoneme string. A penalty for the context match is computed as the difference between the immediately adjacent phonemes surrounding the target phoneme with the corresponding phonemes adjacent to the database phoneme candidate. The context match is also influenced by the distance of the phoneme to its left and right syllable boundary, left and right word boundary, and to the left and right utterance boundary.
0015Speech unit waveforms in the SpeakEZ are concatenated in the time domain, using pitch synchronous overlap-add (PSOLA) smoothing between adjacent phonemes. Rather than modify existing prosody according to ideal target values, the system uses the exact duration, intonation and articulation of the database phoneme without modifications. The lack of proper prosodic target information is considered to be the most glaring shortcoming of this system.
0016Another approach to corpus-based concatenation speech synthesis is described in <nplcit id="ncit0006" npl-type="s"><text>Black, A.W., Campbell, N., "Optimizing selection of units from speech databases for concatenative synthesis," Proc. Eurospeech '95, Madrid, pp. 581-584, 1995</text></nplcit>, and in<nplcit id="ncit0007" npl-type="s"><text> Hunt, A.J., Black, A.W., "Unit selection in a concatenative speech synthesis system using a large speech database," ICASSP-96, pp. 373-376, 1996</text></nplcit>. The annotation of the speech database is taken a step further to incorporate acoustic features: pitch (F<sub>0</sub>), power and spectral parameters are included. The speech database is segmented in phone-sized units. The unit selection algorithm operates as follows: <ol id="ol0002" compact="compact"><li>(1) A unit distortion measure D<sub>u</sub>(u<sub>i</sub>, t<sub>i</sub>) is defined as the distance between a selected unit u<sub>i</sub> and a target speech unit t<sub>i</sub>, i.e. the difference between the selected unit feature vector {uf<sub>1</sub>, uf<sub>2</sub>..., uf<sub>n</sub>} and the target speech unit vector {tf<sub>1</sub>, tf<sub>2</sub>,..., tf<sub>n</sub>} multiplied by a weights vector W<sub>u</sub> {w<sub>1</sub>, w<sub>2</sub>,..., w<sub>n</sub>}.</li><li>(2) A continuity distortion measure D<sub>c</sub>(u<sub>i</sub>, u<sub>i-1</sub>) is defined as the distance between a selected unit and its immediately adjoining previous selected unit, defined as the difference between a selected units unit's feature vector and its previous one multiplied by a weight vector W<sub>c</sub>.</li><li>(3) The best unit sequence is defined as the path of units from the database which minimizes: <maths id="math0001"><math display="block"><mstyle displaystyle="true"><munderover><mo>∑</mo><mrow><mi>i</mi><mo>=</mo><mn>1</mn></mrow><mi>n</mi></munderover></mstyle><mfenced><msub><mi>D</mi><mi>c</mi></msub><mo></mo><mfenced><msub><mi>u</mi><mi>i</mi></msub><msub><mi>u</mi><mrow><mi>i</mi><mo>-</mo><mn>1</mn></mrow></msub></mfenced><mo>*</mo><msub><mi>W</mi><mi>c</mi></msub><mo>+</mo><msub><mi>D</mi><mi>u</mi></msub><mfenced><msub><mi>u</mi><mi>i</mi></msub><msub><mi>t</mi><mi>i</mi></msub></mfenced><mo>*</mo><msub><mi>W</mi><mi>u</mi></msub></mfenced></math><img file="EP1501075B1_D0001.tif" /></maths></li></ol> where n is the number of speech units in the target utterance.
0017In continuity distortion, three features are used: phonetic context, prosodic context, and acoustic join cost. Phonetic and prosodic context distances are calculated between selected units and the context (database) units of other selected units. The acoustic join cost is calculated between two successive selected units. The acoustic join cost is based on a quantization of the mel-cepstrum, calculated at the best joining point around the labeled boundary.
0018A Viterbi search is used to find the path with the minimum cost as expressed in (3). An exhaustive search is avoided by pruning the candidate lists at several stages in the selection process. Units are concatenated without doing any signal processing (i.e., raw concatenation).
0019A clustering technique is presented in<nplcit id="ncit0008" npl-type="s"><text> Black, A.W., Taylor, P.," Automatically clustering similar units for unit selection in speech synthesis," Proc. Eurospeech '97, Rhodes, pp. 601-604, 1997</text></nplcit>, that creates a CART (classification and regression tree) for the units in the database. The CART is used to limit the search domain of candidate units, and the unit distortion cost is the distance between the candidate unit and its cluster center.
0020As an alternative to the mel-cepstrum, <nplcit id="ncit0009" npl-type="s"><text>Ding, W., Campbell, N., "Optimising unit selection with voice source and formants in the CHATR speech synthesis system," Proc. Eurospeech '97, Rhodes, pp. 537-540, 1997</text></nplcit>, presents the use of voice source parameters and formant information as acoustic features for unit selection.
0021Each of the references mentioned above is hereby incorporated herein by reference.
0022<nplcit id="ncit0010" npl-type="s"><text>Hunt and Black, in "Unit selection in a concatenative speech synthesis system using a large speech database", IEEE 1996</text></nplcit>, describes a speech synthesizer with a large speech database. The database uses phonemes.
0023<nplcit id="ncit0011" npl-type="s"><text>Banga and Garcia Mateo, "Shape invariant pitch-synchronous text-to-speech conversion" in ICASSP 90, the International conference on acoustics, speech and signal processing 1990</text></nplcit>, describes a text to speech system that uses, in an example, diphones.
0024The invention as claimed in claim 1 provides a speech synthesizer, and the embodiment includes: <ul id="ul0001" list-style="none" compact="compact"><li>a speech database referencing speech waveforms;</li><li>a speech waveform selector, in communication with the speech database, that selects waveforms referenced by the database using designators that correspond to a phonetic transcription input; and</li><li>a speech waveform concatenator, in communication with the speech database that concatenates waveforms selected by the speech waveform selector to produce a speech signal output,</li></ul> wherein, for at least one ordered sequence of a first waveform and a second waveform, the concatenator selects (i) a location of a trailing edge of the first waveform and (ii) a location of a leading edge of the second waveform, each location being selected so as to produce an optimization of a phase match between the first and second waveforms in regions near the locations.
0025In related embodiments, the phase match is achieved by changing the location only of the leading edge and by changing the location only of the trailing edge. Optionally, or in addition, the optimization is determined on the basis of similarity in shape of the first and second waveforms in the regions near the locations. In further embodiments, similarity is determined using a cross-correlation technique, which optionally is normalized cross correlation. Optionally or in addition, the optimization is determined using at least one non-rectangular window. Also optionally or in addition, the optimization is determined in a plurality of successive stages in which time resolution associated with the first and second waveforms is made successively finer. Optionally, or in addition, the change in resolution is achieved by downsampling.
Brief Description of the Drawings
0026The present invention will be more readily understood by reference to the following detailed description taken with the accompanying drawings, in which: <ul id="ul0002" list-style="none" compact="compact"><li><figref idref="f0001">Fig. 1</figref> illustrates speech synthesizer according to a representative embodiment.</li><li><figref idref="f0002">Fig. 2</figref> illustrates the structure of the speech unit database in a representative embodiment.</li></ul>
Detailed Description of Specific Embodiments Overview
0027A representative embodiment of the present invention, known as the RealSpeak™ Text-to-Speech (TTS) engine, produces high quality speech from a phonetic specification, that can be the output of a text processor, known as a target, by concatenating parts of real recorded speech held in a large database. The main process objects that make up the engine, as shown in <figref idref="f0001">Fig. 1</figref>, include a text processor 101, a target generator 111, a speech unit database 141, a waveform selector 131, and a speech waveform concatenator <b>151</b>.
0028The speech unit database 141 contains recordings, for example in a digital format such as PCM, of a large corpus of actual speech that are indexed in individual speech units by their phonetic descriptors, together with associated speech unit descriptors of various speech unit features. In one embodiment, speech units in the speech unit database 141 are in the form of a diphone, which starts and ends in two neighboring phonemes. Other embodiments may use differently sized and structured speech units. Speech unit descriptors include, for example, symbolic descriptors <i>e.g</i>., lexical stress, word position, etc.-and prosodic descriptors <i>e.g</i>. duration, amplitude, pitch, etc.
0029The text processor 101 receives a text input, <i>e.g</i>., the text phrase "Hello, goodbye!" The text phrase is then converted by the text processor 101 into an input phonetic data sequence. In <figref idref="f0001">Fig. 1</figref>, this is a simple phonetic transcription-#'hE-10#'Gud-bY#. In various alternative embodiments, the input phonetic data sequence may be in one of various different forms. The input phonetic data sequence is converted by the target generator 111 into a multi-layer internal data sequence to be synthesized. This internal data sequence representation, known as extended phonetic transcription (XPT), includes phonetic descriptors, symbolic descriptors, and prosodic descriptors such as those in the speech unit database 141.
0030The waveform selector 131 retrieves from the speech unit database 141 descriptors of candidate speech units that can be concatenated into the target utterance specified by the XPT transcription. The waveform selector 131 creates an ordered list of candidate speech units by comparing the XPTs of the candidate speech units with the XPT of the target XPT, assigning a node cost to each candidate. Candidate-to-target matching is based on symbolic descriptors,such as phonetic context and prosodic context, and numeric descriptors and determines how well each candidate fits the target specification. Poorly matching candidates may be excluded at this point.
0031The waveform selector 131 determines which candidate speech units can be concatenated without causing disturbing quality degradations such as clicks, pitch discontinuities, etc. Successive candidate speech units are evaluated by the waveform selector 131 according to a quality degradation cost function. Candidate-to-candidate matching uses frame-based information such as energy, pitch and spectral information to determine how well the candidates can be joined together. Using dynamic programming, the best sequence of candidate speech units is selected for output to the speech waveform concatenator 151.
0032The speech waveform concatenator 151 requests the output speech units (diphones and/or polyphones) from the speech unit database 141 for the speech waveform concatenator 151. The speech waveform concatenator 151 concatenates the speech units selected forming the output speech that represents the target input text.
0033Operation of various aspects of the system will now be described in greater detail.
Speech Unit Database
0034As shown in <figref idref="f0002">Fig. 2</figref>, the speech unit database 141 contains three types of files: <ol id="ol0003" compact="compact"><li>(1) a speech signal file 61</li><li>(2) a time-aligned extended phonetic transcription (XPT) file 62, and</li><li>(3) a diphone lookup table 63.</li></ol>
Database Indexing
0035Each diphone is identified by two phoneme symbols - these two symbols are the key to the diphone lookup table 63. A diphone index table 631 contains an entry for each possible diphone in the language, describing where the references of these diphones can be found in the diphone reference table 632. The diphone reference table 632 contains references to all the diphones in the speech unit database 141. These references are alphabetically ordered by diphone identifier. In order to reference all diphones by identity it is sufficient to specify where a list starts in the diphone lookup table 63, and how many diphones it contains. Each diphone reference contains the number of the message (utterance) where it is found in the speech unit database 141, which phoneme the diphone starts at, where the diphone starts in the speech signal, and the duration of the diphone.
XPT
0036A significant factor for the quality of the system is the transcription that is used to represent the speech signals in the speech unit database 141. Representative embodiments set out to use a transcription that will allow the system to use the intrinsic prosody in the speech unit database 141 without requiring precise pitch and duration targets. This means that the system can select speech units that are matched phonetically and prosodically to an input transcription. The concatenation of the selected speech units by the speech waveform concatenator 151 effectively leads to an utterance with the desired prosody.
0037The XPT contains two types of data: symbolic features (i.e., features that can be derived from text) and acoustic features (i.e., features that can only be derived from the recorded speech waveform). To effectively extract speech units from the speech unit database 141, the XPT typically contains a time aligned phonetic description of the utterance. The start of each phoneme in the signal is included in the transcription; The XPT also contains a number of prosody related cues, e.g., accentuation and position information. Apart from symbolic information, the transcription also contains acoustic information related to prosody, e.g. the phoneme duration. A typical embodiment concatenates speech units from the speech unit database 141 without modification of their prosodic or spectral realization. Therefore, the boundaries of the speech units should have matching spectral and prosodic realizations. The necessary information required to verify this match is typically incorporated into the XPT by a boundary pitch value and spectral data. The boundary pitch value and the spectrum are calculated at the polyphone edges.
Database Storage
0038Different types of data in the speech unit database 141 may be stored on different physical media, <i>e</i>.<i>g</i>., hard disk, CD-ROM, DVD, random-access memory (RAM), etc. Data access speed may be increased by efficiently choosing how to distribute the data between these various media. The slowest accessing component of a computer system is typically the hard disk. If part of the speech unit information needed to select candidates for concatenation were stored on such a relatively slow mass storage device, valuable processing time would be wasted by accessing this slow device. A much faster implementation could be obtained if selection-related data were stored in RAM. Thus in a representative embodiment, the speech unit database 141 is partitioned into frequently needed selection-related data 21-stored in RAM, and less frequently needed concatenation-related data 22-stored, for example, on CD-ROM or DVD. As a result, RAM requirements of the system remain modest, even if the amount of speech data in the database becomes extremely large (~Gbytes). The relatively small number of CD-ROM retrievals may accommodate multi-channel applications using one CD-ROM for multiple threads, and the speech database may reside alongside other application data on the CD (<i>e.g</i>., navigation systems for an auto-PC).
0039Optionally, speech waveforms may be coded and/or compressed using techniques well-known in the art.
Waveform Selection
0040Initially, each candidate list in the waveform selector 131 contains many available matching diphones in the speech unit database 141. Matching here means merely that the diphone identities match. Thus in an example of a diphone'#1' in which the initial '1' has primary stress in the target, the candidate list in the waveform selector 131 contains every '#1' found in the speech unit database 141, including the ones with unstressed or secondary stressed '1'. The waveform selector 131 uses Dynamic Programming (DP) to find the best sequence of diphones so that: <ol id="ol0004" compact="compact"><li>(1) the database diphones in the best sequence are similar to the target diphones in terms of stress, position, context, etc., and</li><li>(2) the database diphones in the best sequence can be joined together with low concatenation artifacts.</li></ol> In order to achieve these goals, two types of costs are used - a <i>NodeCost</i> which scores the suitability of each candidate diphone to be used to synthesize a particular target, and a <i>TransitionCost</i> which scores the 'joinability' of the diphones. These costs are combined by the DP algorithm, which finds the optimal path.
Cost Functions
0041The cost functions used in the unit selection may be of two types depending on whether the features involved are symbolic (<i>i</i>.<i>e</i>., non numeric <i>e.g.</i>, stress, prominence, phoneme context) or numeric (<i>e</i>.<i>g</i>., spectrum, pitch, duration).
Cost Functions for Symbolic Features
0042For scoring candidates based on the similarity of their symbolic features (<i>i</i>.<i>e</i>., non numeric features) to specified target units, there are 'grey' areas between what is a good match and what is a bad match. The simplest cost weight function would be a binary 0/1. If the candidate has the same value as the target, then the cost is 0; if the candidate is something different, then the cost is 1. For example, when scoring a candidate for its stress (sentence accent (strongest), primary, secondary, unstressed (weakest)) for a target with the strongest stress, this simple system would score primary, secondary or unstressed candidates with a cost of 1. This is counter-intuitive, since if the target is the strongest stress, a candidate of primary stress is preferable to a candidate with no stress.
0043To accommodate this, the user can set up tables which describe the cost between any 2 values of a particular symbolic feature. Some examples are shown in Table 1 and Table 2 in the Tables Appendix which are called 'fuzzy tables' because they resemble concepts from fuzzy logic. Similar tables can be set up for any or all of the symbolic features used in the NodeCost calculation.
0044Fuzzy tables in the waveform selector 131 may also use special symbols, as defined by the developer linguist, which mean 'BAD' and 'VERY BAD'. In practice, the linguist puts a special symbol /1 for BAD, or /2 for VERY BAD in the fuzzy table, as shown in Table 1 in the Tables Appendix, for a target prominence of 3 and a candidate prominence of 0. It was previously mentioned that the normal minimum contribution from any feature is 0 and the maximum is 1. By using /1 or /2 the cost of feature mismatch can be made much higher than 1, such that the candidate is guaranteed to get a high cost. Thus, if for a particular feature the appropriate entry in the table is /1, then the candidate will rarely be used, and if the appropriate entry in the table is /2, then the candidate will almost never be used. In the example of Table 1, if the target prominence is 3, using a /1 makes it unlikely that a candidate with prominence 0 will ever be selected.
Context Dependent Cost Functions
0045The input specification is used to symbolically choose the best combination of speech units from the database which match the input specification. However, using fixed cost functions for symbolic features, to decide which speech units are best, ignores well-known linguistic phenomena such as the fact that some symbolic features are more important in certain contexts than others.
0046For example, it is well-known that in some languages phonemes at the end of an utterance, i.e.,the last syllable, tend to be longer than those elsewhere in an utterance. Therefore, when the dynamic programming algorithm searches for candidate speech units to synthesize the last syllable of an utterance, the candidate speech units should also be from utterance-final syllables, and so it is desirable that in utterance-final position, more importance is placed on the feature of "syllable position". This sort of phenomena varies from language to language, and therefore it is useful to have a way of introducing context-dependent speech unit selection in a rule-based framework, so that the rules can be specified by linguistic experts rather than having to manipulate the actual parameters of the waveform selector 131 cost functions directly. Thus the weights specified for the cost functions may also be manipulated according to a number of rules related to features, e.g. phoneme identities. Additionally, the cost functions themselves may also be manipulated according to rules related to features, e.g. phoneme identities. If the conditions in the rule are met, then several possible actions can occur, such as <ol id="ol0005" compact="compact"><li>(1) For symbolic or numeric features, the weight associated with the feature may be changed — increased if the feature is more important in this context, decreased if the feature is less important. For example, because 'r' often colors vowels before and after it, an expert rule fires when an 'r' in vowel-context is encountered which increases the importance that the candidate items match the target specification for phonetic context.</li><li>(2) For symbolic features, the fuzzy table which a feature normally uses may be changed to a different one.</li><li>(3) For numeric features, the shape of the cost functions can be changed. Some examples are shown in Table 3 in the Tables Appendix, in which * is used to denote 'any phone', and [] is used to surround the current focus diphone. Thus r[at]# denotes a diphone 'at' in context r_#.</li></ol>
Scalability
0047System scalability is also a significant concern in implementing representative embodiments. The speech unit selection strategy offers several scaling possibilities. The waveform selector 131 retrieves speech unit candidates from the speech unit database 141 by means of lookup tables that speed up data retrieval. The input key used to access the lookup tables represents one scalability factor. This input key to the lookup table can vary from minimal-<i>e.g</i>., a pair of phonemes describing the speech unit core-to more complex-<i>e.g</i>., a pair of phonemes + speech unit features (accentuation, context,...). A more complex the input key results in fewer candidate speech units being found through the lookup table. Thus, smaller (although not necessarily better) candidate lists are produced at the cost of more complex lookup tables.
0048The size of the speech unit database 141 is also a significant scaling factor, affecting both required memory and processing speed. The more data that is available, the longer it will take to find an optimal speech unit. The minimal database needed consists of isolated speech units that cover the phonetics of the input (comparable to the speech data bases that are used in linear predictive coding-based phonetics-to-speech systems). Adding well chosen speech signals to the database, improves the quality of the output speech at the cost of increasing system requirements.
0049The pruning techniques described above also represents a scalability factor which can speed up unit selection. A further scalability factor relates to the use of a speech coding and/or speech compression techniques to reduce the size of the speech database.
Signal Processing/Concatenation
0050The speech waveform concatenator 151 performs concatenation-related signal processing. The synthesizer generates speech signals by joining high-quality speech segments together. Concatenating unmodified PCM speech waveforms in the time domain has the advantage that the intrinsic segmental information is preserved. This implies also that the natural prosodic information, including the micro-prosody, is transferred to the synthesized speech. Although the intra-segmental acoustic quality is optimal, attention should be paid to the waveform joining process that may cause inter-segmental distortions. The major concern of waveform concatenation is in avoiding waveform irregularities such as discontinuities and fast transients that may occur in the neighborhood of the join. These waveform irregularities are generally referred to as concatenation artifacts.
0051It is thus important to minimize signal discontinuities at each junction. The concatenation of two segments can be performed by using the well-known weighted overlap-and-add (OLA) method. The overlap and-add procedure for segment concatenation is in fact nothing else than a (non-linear) short time fade-in/fade-out of speech segments. To get high-quality concatenation, we locate a region in the trailing part of the first segment and we locate a region in the leading part of the second segment, such that a phase mismatch measure between the two regions is minimized. This process is performed as follows: <ul id="ul0003" list-style="bullet"><li>We search for the maximum normalized cross-correlation between two sliding windows, one in the trailing part of the first speech segment and one in the leading part of the second speech segment.</li><li>The trailing part of the first speech segment and the leading part of the second speech segment are centered around the diphone boundaries as stored in the lookup tables of the database.</li><li>In the preferred embodiment the length of the trailing and leading regions are of the order of one to two pitch periods and the sliding window is bell-shaped.</li></ul> In order to reduce the computational load of the exhaustive search, the search can be performed in multiple stages. The first stage performs a global search as described in the procedure above on a lower time resolution. The lower time resolution is based on cascaded downsampling of the speech segments. Successive stages perform local searches at successively higher time resolutions around the optimal region determined in the previous stage.
Conclusion
0052Representative embodiments can be implemented as a computer program product for use with a computer system. Such implementation may include a series of computer instructions fixed either on a tangible medium, such as a computer readable medium (<i>e</i>.<i>g</i>., a diskette, CD-ROM, ROM, or fixed disk) or transmittable to a computer system, via a modem or other interface device, such as a communications adapter connected to a network over a medium. The medium may be either a tangible medium (<i>e.g</i>., optical or analog communications lines) or a medium implemented with wireless techniques (<i>e.g</i>., microwave, infrared or other transmission techniques). The series of computer instructions embodies all or part of the functionality previously described herein with respect to the system. Those skilled in the art should appreciate that such computer instructions can be written in a number of programming languages for use with many computer architectures or operating systems. Furthermore, such instructions may be stored in any memory device, such as semiconductor, magnetic, optical or other memory devices, and may be transmitted using any communications technology, such as optical, infrared, microwave, or other transmission technologies. It is expected that such a computer program product may be distributed as a removable medium with accompanying printed or electronic documentation (<i>e</i>.<i>g</i>., shrink wrapped software), preloaded with a computer system (<i>e</i>.<i>g</i>., on system ROM or fixed disk), or distributed from a server or electronic bulletin board over the network (<i>e</i>.<i>g</i>., the Internet or World Wide Web). Of course, some embodiments of the invention may be implemented as a combination of both software (<i>e</i>.<i>g</i>., a computer program product) and hardware. Still other embodiments of the invention are implemented as entirely hardware, or entirely software (<i>e</i>.<i>g</i>., a computer program product).
Glossary
0053The definitions below are pertinent to both the present description and the claims following this description. "Diphone" is a fundamental speech unit composed of two adjacent half-phones. Thus the left and right boundaries of a diphone are in-between phone boundaries. The center of the diphone contains the phone-transition region. The motivation for using diphones rather than phones is that the edges of diphones are relatively steady-state, and so it is easier to join two diphones together with no audible degradation, than it is to join two phones together. "High level" linguistic features of a polyphone or other phonetic unit include, with respect to such unit, accentuation, phonetic context, and position in the applicable sentence, phrase, word, and syllable. "Large speech database" refers to a speech database that references speech waveforms. The database may directly contain digitally sampled waveforms, or it may include pointers to such waveforms, or it may include pointers to parameter sets that govern the actions of a waveform synthesizer. The database is considered "large" when, in the course of waveform reference for the purpose of speech synthesis, the database commonly references many waveform candidates, occurring under varying linguistic conditions. In this manner, most of the time in speech synthesis, the database will likely offer many waveform candidates from which to select. The availability of many such waveform candidates can permit prosodic and other linguistic variation in the speech output, as described throughout herein, and particularly in the Overview. "Low level" linguistic features of a polyphone or other phonetic unit includes, with respect to such unit, pitch contour and duration. "Non-binary numeric" function assumes any of at least three values, depending upon arguments of the function. "Polyphone" is more than one diphone joined together. A triphone is a polyphone made of 2 diphones. "SPT (simple phonetic transcription)" describes the phonemes. This transcription is optionally annotated with symbols for lexical stress, sentence accent, etc... Example (for the word 'worthwhile') : #'werT-'wYl# "Triphone" has two diphones joined together. It thus contains three components - a half phone at its left border, a complete phone, and a half phone at its right border. "Weighted overlap and addition of first and second adjacent waveforms" refers to techniques in which adjacent edges of the waveforms are subjected to fade-in and fade-out.
TABLES APPENDIX
0054<tables id="tabl0001" num="0001"><table frame="bottom"><title>Table 1a XPT Transcription Example</title><tgroup cols="12" colsep="0"><colspec colnum="1" colname="col1" colwidth="17mm" /><colspec colnum="2" colname="col2" colwidth="17mm" /><colspec colnum="3" colname="col3" colwidth="15mm" /><colspec colnum="4" colname="col4" colwidth="15mm" /><colspec colnum="5" colname="col5" colwidth="15mm" /><colspec colnum="6" colname="col6" colwidth="15mm" /><colspec colnum="7" colname="col7" colwidth="15mm" /><colspec colnum="8" colname="col8" colwidth="15mm" /><colspec colnum="9" colname="col9" colwidth="15mm" /><colspec colnum="10" colname="col10" colwidth="15mm" /><colspec colnum="11" colname="col11" colwidth="15mm" /><colspec colnum="12" colname="col12" colwidth="16mm" /><thead><row><entry namest="col1" nameend="col12" align="left" valign="top">XPT: 26 phonemes - 2029.400024 ms - CLASS: S</entry></row></thead><tbody><row rowsep="0"><entry namest="col1" nameend="col2" align="left">PHONEME :</entry><entry>#</entry><entry>y</entry><entry>k</entry><entry>U</entry><entry>d</entry><entry>n</entry><entry>b</entry><entry>i</entry><entry>s</entry><entry>U</entry></row><row rowsep="0"><entry namest="col1" nameend="col2" align="left">DIFF :</entry><entry>0</entry><entry>0</entry><entry>0</entry><entry>0</entry><entry>0</entry><entry>0</entry><entry>0</entry><entry>0</entry><entry>0</entry><entry>0</entry></row><row><entry rowsep="0" /><entry rowsep="0" /><entry rowsep="0" /><entry rowsep="0" /><entry rowsep="0" /><entry rowsep="0" /><entry rowsep="0" /><entry rowsep="0" /><entry rowsep="0" /><entry rowsep="0" /><entry rowsep="0" /><entry rowsep="0" /></row><row rowsep="0"><entry namest="col1" nameend="col2" align="left">SYLL_BND :</entry><entry>S</entry><entry>S</entry><entry>A</entry><entry>B</entry><entry>A</entry><entry>B</entry><entry>A</entry><entry>B</entry><entry>A</entry><entry>N</entry></row><row rowsep="0"><entry namest="col1" nameend="col2" align="left">BND_TYPE-> :</entry><entry>N</entry><entry>W</entry><entry>N</entry><entry>S</entry><entry>N</entry><entry>W</entry><entry>N</entry><entry>W</entry><entry>N</entry><entry>N</entry></row><row><entry rowsep="0" /><entry rowsep="0" /><entry rowsep="0" /><entry rowsep="0" /><entry rowsep="0" /><entry rowsep="0" /><entry rowsep="0" /><entry rowsep="0" /><entry rowsep="0" /><entry rowsep="0" /><entry rowsep="0" /><entry rowsep="0" /></row><row rowsep="0"><entry namest="col1" nameend="col2" align="left">sent_acc :</entry><entry>U</entry><entry>U</entry><entry>S</entry><entry>S</entry><entry>U</entry><entry>U</entry><entry>U</entry><entry>U</entry><entry>S</entry><entry>S</entry></row><row rowsep="0"><entry namest="col1" nameend="col2" align="left">PROMINENCE :</entry><entry>0</entry><entry>0</entry><entry>3</entry><entry>3</entry><entry>0</entry><entry>0</entry><entry>0</entry><entry>0</entry><entry>3</entry><entry>3</entry></row><row><entry rowsep="0" /><entry rowsep="0" /><entry rowsep="0" /><entry rowsep="0" /><entry rowsep="0" /><entry rowsep="0" /><entry rowsep="0" /><entry rowsep="0" /><entry rowsep="0" /><entry rowsep="0" /><entry rowsep="0" /><entry rowsep="0" /></row><row rowsep="0"><entry namest="col1" nameend="col2" align="left">TONE :</entry><entry>X</entry><entry>X</entry><entry>X</entry><entry>X</entry><entry>X</entry><entry>X</entry><entry>X</entry><entry>X</entry><entry>X</entry><entry>X</entry></row><row><entry rowsep="0" /><entry rowsep="0" /><entry rowsep="0" /><entry rowsep="0" /><entry rowsep="0" /><entry rowsep="0" /><entry rowsep="0" /><entry rowsep="0" /><entry rowsep="0" /><entry rowsep="0" /><entry rowsep="0" /><entry rowsep="0" /></row><row rowsep="0"><entry namest="col1" nameend="col2" align="left">SYLL_IN_WRD :</entry><entry>F</entry><entry>F</entry><entry>I</entry><entry>I</entry><entry>F</entry><entry>F</entry><entry>F</entry><entry>F</entry><entry>F</entry><entry>F</entry></row><row rowsep="0"><entry namest="col1" nameend="col2" align="left">SYLL_IN_PHRS :</entry><entry>L</entry><entry>1</entry><entry>2</entry><entry>2</entry><entry>M</entry><entry>M</entry><entry>P</entry><entry>P</entry><entry>L</entry><entry>L</entry></row><row rowsep="0"><entry namest="col1" nameend="col2" align="left">syll_count-> :</entry><entry>0</entry><entry>0</entry><entry>1</entry><entry>1</entry><entry>2</entry><entry>2</entry><entry>3</entry><entry>3</entry><entry>4</entry><entry>4</entry></row><row rowsep="0"><entry namest="col1" nameend="col2" align="left">syll_count<- :</entry><entry>0</entry><entry>4</entry><entry>3</entry><entry>3</entry><entry>2</entry><entry>2</entry><entry>1</entry><entry>1</entry><entry>0</entry><entry>0</entry></row><row rowsep="0"><entry namest="col1" nameend="col2" align="left">SYLL_IN_SENT :</entry><entry>I</entry><entry>I</entry><entry>M</entry><entry>M</entry><entry>M</entry><entry>M</entry><entry>M</entry><entry>M</entry><entry>M</entry><entry>M</entry></row><row rowsep="0"><entry namest="col1" nameend="col2" align="left">NR_SYLL_PHRS :</entry><entry>1</entry><entry>5</entry><entry>5</entry><entry>5</entry><entry>5</entry><entry>5</entry><entry>5</entry><entry>5</entry><entry>5</entry><entry>5</entry></row><row><entry rowsep="0" /><entry rowsep="0" /><entry rowsep="0" /><entry rowsep="0" /><entry rowsep="0" /><entry rowsep="0" /><entry rowsep="0" /><entry rowsep="0" /><entry rowsep="0" /><entry rowsep="0" /><entry rowsep="0" /><entry rowsep="0" /></row><row rowsep="0"><entry namest="col1" nameend="col2" align="left">WRD_IN_SENT :</entry><entry>I</entry><entry /><entry>M</entry><entry>M</entry><entry>M</entry><entry>M</entry><entry>M</entry><entry>M</entry><entry>f</entry><entry>f</entry></row><row rowsep="0"><entry namest="col1" nameend="col2" align="left">PHRS_IN_SENT :</entry><entry>n</entry><entry>n</entry><entry>n</entry><entry>n</entry><entry>n</entry><entry>n</entry><entry>n</entry><entry>n</entry><entry>n</entry><entry>n</entry></row><row><entry rowsep="0" /><entry rowsep="0" /><entry rowsep="0" /><entry rowsep="0" /><entry rowsep="0" /><entry rowsep="0" /><entry rowsep="0" /><entry rowsep="0" /><entry rowsep="0" /><entry rowsep="0" /><entry rowsep="0" /><entry rowsep="0" /></row><row rowsep="0"><entry namest="col1" nameend="col2" align="left">Phon_Start :</entry><entry>0.0</entry><entry>50.0</entry><entry>120.7</entry><entry>250.7</entry><entry>302.5</entry><entry>325.6</entry><entry>433.1</entry><entry>500.7</entry><entry>582.7</entry><entry>734.7</entry></row><row rowsep="0"><entry namest="col1" nameend="col2" align="left">Mid_F0 :</entry><entry>-48.0</entry><entry>23.7</entry><entry>-48.0</entry><entry>27.4</entry><entry>27.0</entry><entry>25.8</entry><entry>24.0</entry><entry>22.7</entry><entry>-48.0</entry><entry>23.3</entry></row><row rowsep="0"><entry namest="col1" nameend="col2" align="left">Avg_F0 :</entry><entry>-48.0</entry><entry>23.2</entry><entry>-48.0</entry><entry>27.4</entry><entry>26.3</entry><entry>25.7</entry><entry>23.8</entry><entry>22.4</entry><entry>-4.8.0</entry><entry>23.2</entry></row><row rowsep="0"><entry namest="col1" nameend="col2" align="left">Slope_F0 :</entry><entry>0.0</entry><entry>-28.6.</entry><entry>0.0</entry><entry>0.0</entry><entry>-165.8</entry><entry>-2.2</entry><entry>84.2</entry><entry>-34.6</entry><entry>0.0</entry><entry>-29.1</entry></row><row><entry namest="col1" nameend="col2" align="left">CepVecInd :</entry><entry>37</entry><entry>0</entry><entry>2</entry><entry>1</entry><entry>16</entry><entry>21</entry><entry>8</entry><entry>20</entry><entry>1</entry><entry>0</entry></row><row><entry /><entry /><entry /><entry /><entry /><entry /><entry /><entry /><entry /><entry /><entry /><entry /></row><row rowsep="0"><entry>r</entry><entry>h</entry><entry>i</entry><entry>w</entry><entry>S</entry><entry>z</entry><entry>s</entry><entry>t</entry><entry>I</entry><entry>l</entry><entry>S</entry><entry>s</entry></row><row><entry rowsep="0">0</entry><entry rowsep="0">0</entry><entry rowsep="0">0</entry><entry rowsep="0">0</entry><entry rowsep="0">0</entry><entry rowsep="0">0</entry><entry rowsep="0">0</entry><entry rowsep="0">0</entry><entry rowsep="0">0</entry><entry rowsep="0">0</entry><entry rowsep="0">0</entry><entry rowsep="0">0</entry></row><row><entry rowsep="0" /><entry rowsep="0" /><entry rowsep="0" /><entry rowsep="0" /><entry rowsep="0" /><entry rowsep="0" /><entry rowsep="0" /><entry rowsep="0" /><entry rowsep="0" /><entry rowsep="0" /><entry rowsep="0" /><entry rowsep="0" /></row><row rowsep="0"><entry>B</entry><entry>A</entry><entry>B</entry><entry>A</entry><entry>N</entry><entry>B</entry><entry>A</entry><entry>N</entry><entry>N</entry><entry>B</entry><entry>S</entry><entry>A</entry></row><row rowsep="0"><entry>P</entry><entry>N</entry><entry>W</entry><entry>N</entry><entry>N</entry><entry>W</entry><entry>N</entry><entry>N</entry><entry>N</entry><entry>W</entry><entry>S</entry><entry>N</entry></row><row><entry rowsep="0" /><entry rowsep="0" /><entry rowsep="0" /><entry rowsep="0" /><entry rowsep="0" /><entry rowsep="0" /><entry rowsep="0" /><entry rowsep="0" /><entry rowsep="0" /><entry rowsep="0" /><entry rowsep="0" /><entry rowsep="0" /></row><row rowsep="0"><entry>X</entry><entry>X</entry><entry>X</entry><entry>X</entry><entry>X</entry><entry>X</entry><entry>X</entry><entry>X</entry><entry>X</entry><entry>X</entry><entry>X</entry><entry>X</entry></row><row rowsep="0"><entry>S</entry><entry>U</entry><entry>U</entry><entry>U</entry><entry>U</entry><entry>U</entry><entry>S</entry><entry>S</entry><entry>S</entry><entry>S</entry><entry>U</entry><entry>S</entry></row><row rowsep="0"><entry>3</entry><entry>0</entry><entry>0</entry><entry>0</entry><entry>0</entry><entry>0</entry><entry>3</entry><entry>3</entry><entry>3</entry><entry>3</entry><entry>0</entry><entry>3</entry></row><row><entry rowsep="0" /><entry rowsep="0" /><entry rowsep="0" /><entry rowsep="0" /><entry rowsep="0" /><entry rowsep="0" /><entry rowsep="0" /><entry rowsep="0" /><entry rowsep="0" /><entry rowsep="0" /><entry rowsep="0" /><entry rowsep="0" /></row><row rowsep="0"><entry>F</entry><entry>F</entry><entry>F</entry><entry>F</entry><entry>F</entry><entry>F</entry><entry>F</entry><entry>F</entry><entry>F</entry><entry>F</entry><entry>I</entry><entry>F</entry></row><row><entry rowsep="0">L</entry><entry rowsep="0">l</entry><entry rowsep="0" /><entry rowsep="0">2</entry><entry rowsep="0">2</entry><entry rowsep="0">2</entry><entry rowsep="0">M</entry><entry rowsep="0">M</entry><entry rowsep="0">M</entry><entry rowsep="0">M</entry><entry rowsep="0">P</entry><entry rowsep="0">L</entry></row><row rowsep="0"><entry>4</entry><entry>0</entry><entry>0</entry><entry>1</entry><entry>1</entry><entry>1</entry><entry>2</entry><entry>2</entry><entry>2</entry><entry>2</entry><entry>3</entry><entry>4</entry></row><row rowsep="0"><entry>0</entry><entry>4</entry><entry>4</entry><entry>3</entry><entry>3</entry><entry>3</entry><entry>2</entry><entry>2</entry><entry>2</entry><entry>2</entry><entry>1</entry><entry>0</entry></row><row rowsep="0"><entry>M</entry><entry>M</entry><entry>M</entry><entry>M</entry><entry>M</entry><entry>M</entry><entry>M</entry><entry>M</entry><entry>M</entry><entry>M</entry><entry>M</entry><entry>F</entry></row><row rowsep="0"><entry>5</entry><entry>5</entry><entry>5</entry><entry>5</entry><entry>5</entry><entry>5</entry><entry>5</entry><entry>5</entry><entry>5</entry><entry>5</entry><entry>5</entry><entry>5</entry></row><row><entry rowsep="0" /><entry rowsep="0" /><entry rowsep="0" /><entry rowsep="0" /><entry rowsep="0" /><entry rowsep="0" /><entry rowsep="0" /><entry rowsep="0" /><entry rowsep="0" /><entry rowsep="0" /><entry rowsep="0" /><entry rowsep="0" /></row><row rowsep="0"><entry>f</entry><entry>i</entry><entry>i</entry><entry>M</entry><entry>M</entry><entry>M</entry><entry>M</entry><entry>M</entry><entry>M</entry><entry>M</entry><entry>F</entry><entry>F</entry></row><row rowsep="0"><entry>n</entry><entry>f</entry><entry>f</entry><entry>f</entry><entry>f</entry><entry>f</entry><entry>f</entry><entry>f</entry><entry>f</entry><entry>f</entry><entry>f</entry><entry>f</entry></row><row><entry rowsep="0" /><entry rowsep="0" /><entry rowsep="0" /><entry rowsep="0" /><entry rowsep="0" /><entry rowsep="0" /><entry rowsep="0" /><entry rowsep="0" /><entry rowsep="0" /><entry rowsep="0" /><entry rowsep="0" /><entry rowsep="0" /></row><row rowsep="0"><entry>826.6</entry><entry>894.7</entry><entry>952.7</entry><entry>1023.2</entry><entry>1053.6</entry><entry>1112.7</entry><entry>1188.7</entry><entry>1216.7</entry><entry>1288.7</entry><entry>1368.7</entry><entry>1429.9</entry><entry>1481-.8</entry></row><row rowsep="0"><entry>22.1</entry><entry>20.0</entry><entry>21.4</entry><entry>18.9</entry><entry>20.0</entry><entry>19.5</entry><entry>-48.0</entry><entry>-48.0</entry><entry>21.4</entry><entry>20.0</entry><entry>19.5</entry><entry>-48.0</entry></row><row><entry rowsep="0" /><entry rowsep="0" /><entry rowsep="0" /><entry rowsep="0" /><entry rowsep="0" /><entry rowsep="0" /><entry rowsep="0" /><entry rowsep="0" /><entry rowsep="0" /><entry rowsep="0" /><entry rowsep="0" /><entry rowsep="0" /></row><row rowsep="0"><entry>22.0</entry><entry>20.2</entry><entry>21.3</entry><entry>19.1</entry><entry>19.9</entry><entry>-48.0</entry><entry>-48.0</entry><entry>-48.0</entry><entry>21.2</entry><entry>20.0</entry><entry>19.6</entry><entry>-48.0</entry></row><row rowsep="0"><entry>-6.9</entry><entry>2.2</entry><entry>-23.1</entry><entry>-5.9</entry><entry>5.5</entry><entry>0.0</entry><entry>0.0</entry><entry>0.0</entry><entry>-27.0</entry><entry>0.0</entry><entry>-9.2</entry><entry>0.0</entry></row><row><entry>21</entry><entry>1</entry><entry>22</entry><entry>2</entry><entry>33</entry><entry>11</entry><entry>38</entry><entry>30</entry><entry>25</entry><entry>28</entry><entry>58</entry><entry>35</entry></row><row><entry /><entry /><entry /><entry /><entry /><entry /><entry /><entry /><entry /><entry /><entry /><entry /></row><row rowsep="0"><entry>1</entry><entry>i</entry><entry>p</entry><entry>#</entry><entry /><entry /><entry /><entry /><entry /><entry /><entry /><entry /></row><row rowsep="0"><entry>0</entry><entry>0</entry><entry>0</entry><entry>0</entry><entry /><entry /><entry /><entry /><entry /><entry /><entry /><entry /></row><row><entry rowsep="0" /><entry rowsep="0" /><entry rowsep="0" /><entry rowsep="0" /><entry rowsep="0" /><entry rowsep="0" /><entry rowsep="0" /><entry rowsep="0" /><entry rowsep="0" /><entry rowsep="0" /><entry rowsep="0" /><entry rowsep="0" /></row><row rowsep="0"><entry>N</entry><entry>N</entry><entry>B</entry><entry>S</entry><entry /><entry /><entry /><entry /><entry /><entry /><entry /><entry /></row><row rowsep="0"><entry>N</entry><entry>N</entry><entry>P</entry><entry>N</entry><entry /><entry /><entry /><entry /><entry /><entry /><entry /><entry /></row><row><entry rowsep="0" /><entry rowsep="0" /><entry rowsep="0" /><entry rowsep="0" /><entry rowsep="0" /><entry rowsep="0" /><entry rowsep="0" /><entry rowsep="0" /><entry rowsep="0" /><entry rowsep="0" /><entry rowsep="0" /><entry rowsep="0" /></row><row rowsep="0"><entry>X</entry><entry>X</entry><entry>X</entry><entry>X</entry><entry /><entry /><entry /><entry /><entry /><entry /><entry /><entry /></row><row><entry rowsep="0" /><entry rowsep="0" /><entry rowsep="0" /><entry rowsep="0" /><entry rowsep="0" /><entry rowsep="0" /><entry rowsep="0" /><entry rowsep="0" /><entry rowsep="0" /><entry rowsep="0" /><entry rowsep="0" /><entry rowsep="0" /></row><row rowsep="0"><entry>S</entry><entry>S</entry><entry>S</entry><entry>U</entry><entry /><entry /><entry /><entry /><entry /><entry /><entry /><entry /></row><row rowsep="0"><entry>3</entry><entry>3</entry><entry>3</entry><entry>0</entry><entry /><entry /><entry /><entry /><entry /><entry /><entry /><entry /></row><row rowsep="0"><entry>F</entry><entry>F</entry><entry>F</entry><entry>F</entry><entry /><entry /><entry /><entry /><entry /><entry /><entry /><entry /></row><row rowsep="0"><entry>L</entry><entry>L</entry><entry>L</entry><entry>L</entry><entry /><entry /><entry /><entry /><entry /><entry /><entry /><entry /></row><row rowsep="0"><entry>4</entry><entry>4</entry><entry>4</entry><entry>0</entry><entry /><entry /><entry /><entry /><entry /><entry /><entry /><entry /></row><row rowsep="0"><entry>0</entry><entry>0</entry><entry>0</entry><entry>0</entry><entry /><entry /><entry /><entry /><entry /><entry /><entry /><entry /></row><row><entry rowsep="0">F</entry><entry rowsep="0">F</entry><entry rowsep="0">F</entry><entry rowsep="0">F</entry><entry rowsep="0" /><entry rowsep="0" /><entry rowsep="0" /><entry rowsep="0" /><entry rowsep="0" /><entry rowsep="0" /><entry rowsep="0" /><entry rowsep="0" /></row><row><entry rowsep="0">5</entry><entry rowsep="0">5</entry><entry rowsep="0">5</entry><entry rowsep="0">1</entry><entry rowsep="0" /><entry rowsep="0" /><entry rowsep="0" /><entry rowsep="0" /><entry rowsep="0" /><entry rowsep="0" /><entry rowsep="0" /><entry rowsep="0" /></row><row><entry rowsep="0" /><entry rowsep="0" /><entry rowsep="0" /><entry rowsep="0" /><entry rowsep="0" /><entry rowsep="0" /><entry rowsep="0" /><entry rowsep="0" /><entry rowsep="0" /><entry rowsep="0" /><entry rowsep="0" /><entry rowsep="0" /></row><row><entry rowsep="0">F</entry><entry rowsep="0">F</entry><entry rowsep="0">F</entry><entry rowsep="0">F</entry><entry rowsep="0" /><entry rowsep="0" /><entry rowsep="0" /><entry rowsep="0" /><entry rowsep="0" /><entry rowsep="0" /><entry rowsep="0" /><entry rowsep="0" /></row><row><entry rowsep="0">f</entry><entry rowsep="0">f</entry><entry rowsep="0">f</entry><entry rowsep="0">f</entry><entry rowsep="0" /><entry rowsep="0" /><entry rowsep="0" /><entry rowsep="0" /><entry rowsep="0" /><entry rowsep="0" /><entry rowsep="0" /><entry rowsep="0" /></row><row><entry rowsep="0" /><entry rowsep="0" /><entry rowsep="0" /><entry rowsep="0" /><entry rowsep="0" /><entry rowsep="0" /><entry rowsep="0" /><entry rowsep="0" /><entry rowsep="0" /><entry rowsep="0" /><entry rowsep="0" /><entry rowsep="0" /></row><row><entry rowsep="0">1619.0</entry><entry rowsep="0">1677.6</entry><entry rowsep="0">1840.7</entry><entry rowsep="0">1979.4</entry><entry rowsep="0" /><entry rowsep="0" /><entry rowsep="0" /><entry rowsep="0" /><entry rowsep="0" /><entry rowsep="0" /><entry rowsep="0" /><entry rowsep="0" /></row><row><entry rowsep="0">20.0</entry><entry rowsep="0">17.2</entry><entry rowsep="0">13.3</entry><entry rowsep="0">9.4</entry><entry rowsep="0" /><entry rowsep="0" /><entry rowsep="0" /><entry rowsep="0" /><entry rowsep="0" /><entry rowsep="0" /><entry rowsep="0" /><entry rowsep="0" /></row><row><entry rowsep="0">19.8</entry><entry rowsep="0">17.2</entry><entry rowsep="0">-48.0</entry><entry rowsep="0">-48.0</entry><entry rowsep="0" /><entry rowsep="0" /><entry rowsep="0" /><entry rowsep="0" /><entry rowsep="0" /><entry rowsep="0" /><entry rowsep="0" /><entry rowsep="0" /></row><row><entry rowsep="0">-30.8</entry><entry rowsep="0">-29.8</entry><entry rowsep="0">0.0</entry><entry rowsep="0">0.0</entry><entry rowsep="0" /><entry rowsep="0" /><entry rowsep="0" /><entry rowsep="0" /><entry rowsep="0" /><entry rowsep="0" /><entry rowsep="0" /><entry rowsep="0" /></row><row><entry>21</entry><entry>14</entry><entry>26</entry><entry>1</entry><entry /><entry /><entry /><entry /><entry /><entry /><entry /><entry /></row></tbody></tgroup></table></tables><tables id="tabl0002" num="0002"><table frame="all"><tgroup cols="4"><colspec colnum="1" colname="col1" colwidth="48mm" /><colspec colnum="2" colname="col2" colwidth="26mm" /><colspec colnum="3" colname="col3" colwidth="47mm" /><colspec colnum="4" colname="col4" colwidth="47mm" /><thead><row><entry namest="col1" nameend="col4" align="center" valign="top">SYMBOLIC FEATURES (XPT)</entry></row><row><entry valign="top">name & acronym</entry><entry valign="top">applies to</entry><entry valign="top">possible values</entry><entry valign="top">When?</entry></row></thead><tbody><row rowsep="0"><entry><i>phonetic differentiator</i></entry><entry>phoneme</entry><entry>0 (not annotated)</entry><entry>no annotation symbol present after phoneme</entry></row><row rowsep="0"><entry><i>DIFF</i></entry><entry /><entry>1 (annotated with first symbol)</entry><entry>first annotation symbol present after phoneme</entry></row><row rowsep="0"><entry /><entry /><entry>2 (annotated with second symbol)</entry><entry>second annotation symbol</entry></row><row><entry /><entry /><entry>etc</entry><entry>etc</entry></row><row rowsep="0"><entry>phoneme position in syllable</entry><entry>phoneme</entry><entry>A(fter syllable boundary)</entry><entry>phoneme after syllable boundary</entry></row><row rowsep="0"><entry>SYLL_BND</entry><entry /><entry>B(efore syllable boundary)</entry><entry>phoneme before, but not after, syllable boundary</entry></row><row rowsep="0"><entry /><entry /><entry>S(urrounded by syllable boundaries)</entry><entry>phoneme surrounded by syllable boundaries, or phoneme is silence</entry></row><row><entry /><entry /><entry>N(ot near syllable boundary)</entry><entry>phoneme not before or after syllable boundary</entry></row><row rowsep="0"><entry>type of boundary following phoneme</entry><entry>phoneme</entry><entry>N(o)</entry><entry>no boundary following phoneme</entry></row><row rowsep="0"><entry>BND_TYPE-></entry><entry /><entry>S(yllable)</entry><entry>Syllable boundary following phoneme</entry></row><row rowsep="0"><entry /><entry /><entry>W(ord)</entry><entry>Word boundary following phoneme</entry></row><row><entry /><entry /><entry>P(hrase)</entry><entry>Phrase boundary following phoneme</entry></row><row rowsep="0"><entry>lexical stress</entry><entry>syllable</entry><entry>(P)rimary</entry><entry>phoneme in syllable with primary stress</entry></row><row rowsep="0"><entry>lex_str</entry><entry /><entry>(S)econdary</entry><entry>phoneme in syllable with secondary stress</entry></row><row><entry /><entry /><entry>(U)nstressed</entry><entry>phoneme in syllable without lexical stress, or phoneme is silence</entry></row><row rowsep="0"><entry>sentence accent</entry><entry>syllable</entry><entry>(S)tressed</entry><entry>phoneme in syllable with sentence accent</entry></row><row><entry>sent_acc</entry><entry /><entry>(U)nstressed</entry><entry>phoneme in syllable without sentence accent, or phoneme is silence</entry></row><row rowsep="0"><entry>prominence</entry><entry>syllable</entry><entry>0</entry><entry>lex_str = U and sent_acc = U</entry></row><row rowsep="0"><entry>PROMINENCE</entry><entry /><entry>1</entry><entry>lex_str = S and sent_acc =U</entry></row><row rowsep="0"><entry /><entry /><entry>2</entry><entry>lex_str = P and sent_acc= U</entry></row><row><entry /><entry /><entry>3</entry><entry>sent_acc = S.</entry></row><row rowsep="0"><entry>t<i>one value</i> TONE</entry><entry>syllable (mora)</entry><entry>X (missing value)</entry><entry>phoneme in syllable (mora) without tone marker, or phoneme = #, or optional feature is not</entry></row><row rowsep="0"><entry /><entry /><entry>L(ow tone)</entry><entry>supported phoneme in mora with tone = L</entry></row><row rowsep="0"><entry /><entry /><entry>R(ising tone)</entry><entry>phoneme in mora with tone = R</entry></row><row rowsep="0"><entry /><entry /><entry>H(igh tone)</entry><entry>phoneme in mora with tone = H</entry></row><row><entry /><entry /><entry>F(alling tone)</entry><entry>phoneme in mora with tone = F</entry></row><row rowsep="0"><entry>syllable position in word</entry><entry>syllable</entry><entry>I(nitial)</entry><entry>phoneme in first syllable of multi-syllabic word</entry></row><row rowsep="0"><entry>SYTLL_IN_WRD</entry><entry /><entry>M(edial)</entry><entry>phoneme neither in first nor last syllable of word</entry></row><row><entry /><entry /><entry>F(inal)</entry><entry>phoneme in last syllable of word (including mono-syllabic words), or phoneme is silence</entry></row><row><entry>syllable count in phrase (from first) syll_count-></entry><entry>syllable</entry><entry>0:.N-1(N=nr syll in phrase)</entry><entry /></row><row><entry>syllable count in phrase (from last) syll_count<-</entry><entry>syllable</entry><entry>N-1..0(N=nr syll in phrase)</entry><entry /></row><row rowsep="0"><entry>syllable position in phrase</entry><entry>syllable</entry><entry>1 (first)</entry><entry>syll_count-> = 0</entry></row><row rowsep="0"><entry>SYLL_IN_PHRS</entry><entry /><entry>2 (second)</entry><entry>syll_count-> = 1</entry></row><row rowsep="0"><entry /><entry /><entry>I (nitial)</entry><entry>syll_count-> < 0.3 *N</entry></row><row rowsep="0"><entry /><entry /><entry>M(edial)</entry><entry>all other cases</entry></row><row rowsep="0"><entry /><entry /><entry>F(inal)</entry><entry>syll_count<- < 0.3*N</entry></row><row rowsep="0"><entry /><entry /><entry>P(enultimate)</entry><entry>syll_count<- = N-2 ,</entry></row><row><entry /><entry /><entry>L(ast)</entry><entry>syll_count<- = N-1</entry></row><row rowsep="0"><entry>syllable position in sentence</entry><entry>syllablle</entry><entry>I(nitial)</entry><entry>first syllable in sentence following initial silence, and</entry></row><row rowsep="0"><entry>SYLL_IN_SENT</entry><entry /><entry>M(edial)</entry><entry>initial silence all other cases</entry></row><row><entry /><entry /><entry>F(inal)</entry><entry>last syllable in sentence receding final silence, mono-syllable, and final silence</entry></row><row><entry>number of syllables in phrase NR_SYLL_PHRS</entry><entry>phrase</entry><entry>N (number of syll)</entry><entry /></row><row rowsep="0"><entry>word position in sentence</entry><entry>word</entry><entry>1(nitial)</entry><entry>first word in sentence</entry></row><row rowsep="0"><entry>WRD_IN_SENT</entry><entry /><entry>M(edial)</entry><entry>not first or last word in sentence or phrase or phrase</entry></row><row rowsep="0"><entry /><entry /><entry>f(inal in phrase, but sentence medial)</entry><entry>last word in phrase, but not last word in sentence</entry></row><row><entry rowsep="0" /><entry rowsep="0" /><entry rowsep="0">i(initial in phrase, but sentence medial)</entry><entry rowsep="0">first word in phrase, but not first word in sentence</entry></row><row><entry /><entry /><entry>F(inal)</entry><entry>last word in sentence</entry></row><row rowsep="0"><entry>phrase position in sentence</entry><entry>phrase</entry><entry morerows="1" rowsep="1">n(ot final) f(inal)</entry><entry>not last phrase in sentence</entry></row><row><entry>PHRS_IN_SENT</entry><entry /><entry>last phrase in sentence</entry></row></tbody></tgroup></table></tables><tables id="tabl0003" num="0003"><table frame="all"><title>Table 1b - XPT Descriptors</title><tgroup cols="3"><colspec colnum="1" colname="col1" colwidth="75mm" /><colspec colnum="2" colname="col2" colwidth="33mm" /><colspec colnum="3" colname="col3" colwidth="59mm" /><thead><row><entry namest="col1" nameend="col3" align="left" valign="top">ACOUSTIC FEATURES (XPT)</entry></row><row><entry valign="top">name & acronym</entry><entry valign="top">applies to</entry><entry valign="top">possible values</entry></row></thead><tbody><row><entry>start of phoneme in signal Phon_Start</entry><entry>phoneme</entry><entry>0..length_of_signal</entry></row><row><entry>pitch at diphone boundary in phoneme Mid_F0</entry><entry>diphone e boundary</entry><entry>expressed in semitones</entry></row><row><entry>average pitch value within the phoneme Avg_F0</entry><entry>phoneme</entry><entry>expressed in semitones</entry></row><row><entry>pitch slope within phoneme Slope_F0</entry><entry>phoneme</entry><entry>expressed in semitones per second</entry></row><row><entry>cepstral vector index at diphone boundary in phoneme CepVecInd</entry><entry>diphone boundary</entry><entry>unsigned integer value (usually 0..128)</entry></row></tbody></tgroup></table></tables><tables id="tabl0004" num="0004"><table frame="all"><title>Table 2 Example of a fuzzy table for prominence matching</title><tgroup cols="6"><colspec colnum="1" colname="col1" colwidth="24mm" /><colspec colnum="2" colname="col2" colwidth="12mm" /><colspec colnum="3" colname="col3" colwidth="12mm" /><colspec colnum="4" colname="col4" colwidth="12mm" /><colspec colnum="5" colname="col5" colwidth="12mm" /><colspec colnum="6" colname="col6" colwidth="12mm" /><thead><row><entry valign="top" /><entry valign="top" /><entry namest="col3" nameend="col6" align="left" valign="top">Candidate Prominence</entry></row><row><entry valign="top" /><entry valign="top" /><entry valign="top">0</entry><entry valign="top">1</entry><entry valign="top">2</entry><entry valign="top">3</entry></row></thead><tbody><row><entry rowsep="0">Target</entry><entry>0</entry><entry>0</entry><entry>0.1</entry><entry>0.5</entry><entry>1.0</entry></row><row><entry rowsep="0">Prominence</entry><entry>1</entry><entry>0.2</entry><entry>0</entry><entry>0.1</entry><entry>0.8</entry></row><row><entry rowsep="0" /><entry>2</entry><entry>0.8</entry><entry>0.3</entry><entry>0</entry><entry>0.2</entry></row><row><entry /><entry>3</entry><entry>1.0</entry><entry>1.0</entry><entry>0.3</entry><entry>0</entry></row></tbody></tgroup></table></tables><tables id="tabl0005" num="0005"><table frame="all"><title>Table 3 Example of a fuzzy table for the left context phone</title><tgroup cols="8"><colspec colnum="1" colname="col1" colwidth="16mm" /><colspec colnum="2" colname="col2" colwidth="10mm" /><colspec colnum="3" colname="col3" colwidth="10mm" /><colspec colnum="4" colname="col4" colwidth="10mm" /><colspec colnum="5" colname="col5" colwidth="10mm" /><colspec colnum="6" colname="col6" colwidth="10mm" /><colspec colnum="7" colname="col7" colwidth="10mm" /><colspec colnum="8" colname="col8" colwidth="10mm" /><thead><row><entry valign="top" /><entry valign="top" /><entry namest="col3" nameend="col8" align="left" valign="top">Candidate left context phone</entry></row><row><entry valign="top" /><entry valign="top" /><entry valign="top">a</entry><entry valign="top">e</entry><entry valign="top">I</entry><entry valign="top">p</entry><entry valign="top">...</entry><entry valign="top">$</entry></row></thead><tbody><row><entry rowsep="0">Target</entry><entry>a</entry><entry>0</entry><entry>0.2</entry><entry>0.4</entry><entry>1.0</entry><entry>...</entry><entry>0.8</entry></row><row><entry rowsep="0">Left</entry><entry>e</entry><entry>0.1</entry><entry>0</entry><entry>0.8</entry><entry>1.0</entry><entry>...</entry><entry>0.8</entry></row><row><entry rowsep="0">Context</entry><entry>i</entry><entry>0.9</entry><entry>0.8</entry><entry>0</entry><entry>1.0</entry><entry>...</entry><entry>0.2</entry></row><row><entry rowsep="0">Phone</entry><entry>P</entry><entry>1.0</entry><entry>1.0</entry><entry>1.0</entry><entry>0</entry><entry>...</entry><entry>1.0</entry></row><row><entry rowsep="0" /><entry>...</entry><entry>...</entry><entry>...</entry><entry>...</entry><entry>...</entry><entry>...</entry><entry>...</entry></row><row><entry /><entry>$</entry><entry>0.2</entry><entry>0.8</entry><entry>0.8</entry><entry>1.0</entry><entry>...</entry><entry>0</entry></row></tbody></tgroup></table></tables><tables id="tabl0006" num="0006"><table frame="all"><title>Table 4 Example of a fuzzy table for prominence matching</title><tgroup cols="6"><colspec colnum="1" colname="col1" colwidth="24mm" /><colspec colnum="2" colname="col2" colwidth="12mm" /><colspec colnum="3" colname="col3" colwidth="12mm" /><colspec colnum="4" colname="col4" colwidth="12mm" /><colspec colnum="5" colname="col5" colwidth="12mm" /><colspec colnum="6" colname="col6" colwidth="12mm" /><thead><row><entry valign="top" /><entry valign="top" /><entry namest="col3" nameend="col6" align="left" valign="top">Candidate Prominence</entry></row><row><entry valign="top" /><entry valign="top" /><entry valign="top">0</entry><entry valign="top">1</entry><entry valign="top">2</entry><entry valign="top">3</entry></row></thead><tbody><row><entry rowsep="0">Target</entry><entry>0</entry><entry>0</entry><entry>0.1</entry><entry>0.5</entry><entry>1.0</entry></row><row><entry rowsep="0">Prominence</entry><entry>1</entry><entry>0.2</entry><entry>0</entry><entry>0.1</entry><entry>0.8</entry></row><row><entry rowsep="0" /><entry>2</entry><entry>0.8</entry><entry>0.3</entry><entry>0</entry><entry>0.2</entry></row><row><entry /><entry>3</entry><entry>/1</entry><entry>1.0</entry><entry>0.3</entry><entry>0</entry></row></tbody></tgroup></table></tables><tables id="tabl0007" num="0007"><table frame="all"><title>Table 5 Examples of context-dependent weight modifications</title><tgroup cols="3"><colspec colnum="1" colname="col1" colwidth="37mm" /><colspec colnum="2" colname="col2" colwidth="65mm" /><colspec colnum="3" colname="col3" colwidth="65mm" /><thead><row><entry>Rule</entry><entry valign="top">Action</entry><entry valign="top">Justification</entry></row></thead><tbody><row><entry>*[r*]*</entry><entry>Make the left context more important</entry><entry>r can be colored by the preceding vowel</entry></row><row><entry>r[V*]*, V=any vowel</entry><entry>Make the left context more important</entry><entry>The vowel can be colored by the r.</entry></row><row><entry>*[X]*,X=unvoiced stop</entry><entry>Make the left context more important</entry><entry>If left context is s then X is not aspirated. This encourages exact matching for s[X*]* , but also includes some side effects.</entry></row><row><entry>*[*V]r</entry><entry>Make the right context more important</entry><entry>Vowel coloring</entry></row><row><entry>*[X*]*X=non-sonorant</entry><entry>Make syllable position weights and prominence weights zero.</entry><entry>Sonorants are more sensitive to position and prominence than non-sonorants</entry></row></tbody></tgroup></table></tables><tables id="tabl0008" num="0008"><table frame="all"><title>Table 6 Transition Cost Calculation Features (Features marked * only 'fire' on accented vowels)</title><tgroup cols="5"><colspec colnum="1" colname="col1" colwidth="28mm" /><colspec colnum="2" colname="col2" colwidth="34mm" /><colspec colnum="3" colname="col3" colwidth="35mm" /><colspec colnum="4" colname="col4" colwidth="35mm" /><colspec colnum="5" colname="col5" colwidth="35mm" /><thead><row><entry valign="top">Feature number</entry><entry valign="top">Feature</entry><entry valign="top">Lowest cost if....</entry><entry valign="top">Highest cost if..</entry><entry valign="top">Type of scoring</entry></row></thead><tbody><row><entry>1</entry><entry>Adjacent in database (i.e., adjacent in donor recorded item)</entry><entry>The two speech units are in adjacent position in same donor word</entry><entry>They are not adjacent</entry><entry>0/1</entry></row><row><entry>2</entry><entry>Pitch difference</entry><entry>There is no pitch difference</entry><entry>There is a big pitch difference</entry><entry>Bigger mismatch = bigger cost (also depends on cost function)</entry></row><row><entry>3</entry><entry>Cepstral distance</entry><entry>There is cepstral continuity</entry><entry>There is no cepstral continuity</entry><entry>Bigger mismatch = bigger cost (also depends on cost function)</entry></row><row><entry>4</entry><entry>Duration pdf</entry><entry>The duration of the phone (the 2 demiphones joined together) is within expected limits for the target phone ID, accent and position</entry><entry>The duration of the phone is outside that expected for the target phone ID, accent and position</entry><entry>Bigger mismatch = bigger cost</entry></row><row><entry>5</entry><entry>Vowel pitch continuity Acc-acc or unacc-unacc (for declination)</entry><entry>Pitch of this accented(unacc) syl is same or slightly lower than the previous accented (unacc) syl in this phrase</entry><entry>Pitch is higher than previous acc (unacc)syl, or pitch is much lower than previous acc (unacc) syl</entry><entry>Flat-bottomed function cost</entry></row><row><entry>6</entry><entry>Vowel pitch continuity Unacc-Acc* (for rising pitch from unacc-acc)</entry><entry>Pitch is same or slightly higher than the previous unaccented syllable in this phrase</entry><entry>Pitch is lower than previous unacc syl, or pitch is much higher than previous acc syl.</entry><entry>Flat bottomed asymmetric cost function.</entry></row></tbody></tgroup></table></tables><tables id="tabl0009" num="0009"><table frame="all"><title>Table 7 - Weight function shapes used in Transistion Cost calculation</title><tgroup cols="10"><colspec colnum="1" colname="col1" colwidth="21mm" /><colspec colnum="2" colname="col2" colwidth="16mm" /><colspec colnum="3" colname="col3" colwidth="16mm" /><colspec colnum="4" colname="col4" colwidth="16mm" /><colspec colnum="5" colname="col5" colwidth="16mm" /><colspec colnum="6" colname="col6" colwidth="16mm" /><colspec colnum="7" colname="col7" colwidth="17mm" /><colspec colnum="8" colname="col8" colwidth="17mm" /><colspec colnum="9" colname="col9" colwidth="17mm" /><colspec colnum="10" colname="col10" colwidth="17mm" /><thead><row><entry valign="top"><b>Transition Cost Feature</b></entry><entry namest="col2" nameend="col10" align="left" valign="top"><b>Shape of cost function</b></entry></row></thead><tbody><row><entry>1 Adjacent in database</entry><entry namest="col2" nameend="col10" align="left">If items are adjacent cost =0. Otherwise cost=1)</entry></row><row><entry>2 Pitch Difference</entry><entry namest="col2" nameend="col10" align="left"><img file="EP1501075B1_D0002.tif" /> Pitch(right demiphone)-pitch(letf demiphone) R = range</entry></row><row><entry>3 Cepstral Distance</entry><entry namest="col2" nameend="col10" align="left"><img file="EP1501075B1_D0003.tif" /> Cepstral distance between left demiphone and right demiphone</entry></row><row><entry>4 Duration PDF</entry><entry namest="col2" nameend="col10" align="left"><img file="EP1501075B1_D0004.tif" /> Duration of phone (=dur of left demiphone+dur of right demiphone)</entry></row><row><entry>5 Vowel pitch continuity (I)</entry><entry namest="col2" nameend="col10" align="left"><img file="EP1501075B1_D0005.tif" /> Pitch(now)-pitch(prev syl with same accentuation)</entry></row><row><entry>6 Vowel pitch continuity(II)</entry><entry namest="col2" nameend="col10" align="left"><img file="EP1501075B1_D0006.tif" /> Pitch(now)pitch(prev unacc syl)</entry></row></tbody></tgroup></table></tables><tables id="tabl0010" num="0010"><table frame="all"><title>Table 8 Example of a cost function table for categorical variable</title><tgroup cols="6"><colspec colnum="1" colname="col1" colwidth="15mm" /><colspec colnum="2" colname="col2" colwidth="15mm" /><colspec colnum="3" colname="col3" colwidth="15mm" /><colspec colnum="4" colname="col4" colwidth="15mm" /><colspec colnum="5" colname="col5" colwidth="15mm" /><colspec colnum="6" colname="col6" colwidth="15mm" /><thead><row><entry valign="top" /><entry valign="top" /><entry namest="col3" nameend="col6" align="left" valign="top">x2</entry></row><row><entry valign="top" /><entry valign="top" /><entry valign="top">a</entry><entry valign="top">e</entry><entry valign="top">...</entry><entry valign="top">Z</entry></row></thead><tbody><row><entry morerows="3">x1</entry><entry>a</entry><entry>0.0</entry><entry>0.4</entry><entry>...</entry><entry>0.1</entry></row><row><entry>e</entry><entry>0.1</entry><entry>0.0</entry><entry>...</entry><entry>0.2</entry></row><row><entry>...</entry><entry>...</entry><entry>...</entry><entry>...</entry><entry>...</entry></row><row><entry>z</entry><entry>0.9</entry><entry>1.0</entry><entry>...</entry><entry>0</entry></row></tbody></tgroup></table></tables><tables id="tabl0011" num="0011"><table frame="all"><title>Table 9 - Duration PDF Table</title><tgroup cols="3" colsep="0" rowsep="0"><colspec colnum="1" colname="col1" colwidth="27mm" /><colspec colnum="2" colname="col2" colwidth="22mm" /><colspec colnum="3" colname="col3" colwidth="24mm" colsep="1" /><thead><row><entry namest="col1" nameend="col3" align="left" valign="top">[FEATURES]</entry></row></thead><tbody><row><entry>CLASS</entry><entry namest="col2" nameend="col3" align="left">#$?DFLNPRSV</entry></row><row><entry>ACCENT</entry><entry namest="col2" nameend="col3" align="left">YN</entry></row><row><entry>PHRASEFINAL</entry><entry namest="col2" nameend="col3" align="left">YN</entry></row><row><entry /><entry /><entry /></row><row><entry>[DATA]</entry><entry /><entry /></row><row><entry># N N</entry><entry align="char" char=".">48.300000</entry><entry align="char" char="." charoff="31">114.800000</entry></row><row><entry># N Y</entry><entry align="char" char=".">0.000000</entry><entry align="char" char="." charoff="31">1000.000000</entry></row><row><entry># Y N</entry><entry align="char" char=".">0.000000</entry><entry align="char" char="." charoff="31">1000.000000</entry></row><row><entry># Y Y</entry><entry align="char" char=".">0.00000</entry><entry align="char" char="." charoff="31">1000.000000</entry></row><row><entry>$ N N</entry><entry align="char" char=".">35.300000</entry><entry align="char" char="." charoff="31">60.700000</entry></row><row><entry>$ N Y</entry><entry align="char" char=".">56.300000</entry><entry align="char" char="." charoff="31">93.900000</entry></row><row><entry>$ Y N</entry><entry align="char" char=".">0.000000</entry><entry align="char" char="." charoff="31">1000.000000</entry></row><row><entry>$ Y Y</entry><entry align="char" char=".">0.000000</entry><entry align="char" char="." charoff="31">1000.000000</entry></row><row><entry>? N N</entry><entry align="char" char=".">50.900000</entry><entry align="char" char="." charoff="31">84.000000</entry></row><row><entry>? N Y</entry><entry align="char" char=".">59.200000</entry><entry align="char" char="." charoff="31">89.400000</entry></row><row><entry>? Y N</entry><entry align="char" char=".">51.400000</entry><entry align="char" char="." charoff="31">83.500000</entry></row><row><entry>? Y Y</entry><entry align="char" char=".">51.500000</entry><entry align="char" char="." charoff="31">88.400000</entry></row><row><entry>D N N</entry><entry align="char" char=".">96.400000</entry><entry align="char" char="." charoff="31">148.700000</entry></row><row><entry>D N Y</entry><entry align="char" char=".">154.000000</entry><entry align="char" char="." charoff="31">249.500000</entry></row><row><entry>D Y N</entry><entry align="char" char=".">117.400000</entry><entry align="char" char="." charoff="31">174.400000</entry></row><row><entry>D Y Y</entry><entry align="char" char=".">176.800000</entry><entry align="char" char="." charoff="31">275.500000</entry></row><row><entry>F N N</entry><entry align="char" char=".">39.000000</entry><entry align="char" char="." charoff="31">90.100000</entry></row><row rowsep="1"><entry>F Y N</entry><entry align="char" char=".">56.200000</entry><entry align="char" char="." charoff="31">122.90000</entry></row></tbody></tgroup></table></tables>
Contents2
8 sheets
Sheet 1 Sheet 2 Sheet 3 Sheet 4 Sheet 5 Sheet 6 Sheet 7 Sheet 8
Every citation, both ways
| Document | Relation | Office |
|---|---|---|
| EP0608833A | Cites | European Patent Office (EPO) |
| EP0848372A | Cites | European Patent Office (EPO) |
| US2002178006A1 | Cites | United States of America |
| BANGA E R ET AL: "SHAPE-INVARIANT PITCH-SYNCHRONOUS TEXT-TO-SPEECH CONVERSION" PROCEEDINGS OF THE INTERNATIONAL CONFERENCE ON ACOUSTICS, SPEECH, AND SIGNAL PROCESSING (ICASSP),US,NEW YORK, IEEE, 1995, pages 656-659, XP000658079 ISBN: 0-7803-2432-3 | Non-patent | – |
19 members in 8 offices
Priority claims3
| Document | Office | Kind | Date |
|---|---|---|---|
| 108201P | United States of America | – | |
| 10820198 | United States of America | P | |
| 99972346 | European Patent Office (EPO) | A |
Members19
| Document | Office | Kind | |
|---|---|---|---|
| CA2354871A1 | Canada | A1 | |
| WO0030069A2 | World Intellectual Property Organization (WIPO) | A2 | |
| AU1403100A | Australia | A | |
| WO0030069A3 | World Intellectual Property Organization (WIPO) | A3 | |
| EP1138038A2 | European Patent Office (EPO) | A2 | |
| JP2002530703A | Japan | A | |
| US6665641B1 | United States of America | B1 | |
| AU772874B2 | Australia | B2 | |
| US2004111266A1 | United States of America | A1 | |
| EP1501075A2 | European Patent Office (EPO) | A2 | |
| EP1138038B1 | European Patent Office (EPO) | B1 | |
| AT298453T | Austria | T | |
| ATE298453T1 | Austria | T1 | |
| DE69925932D1 | Germany | D1 | |
| DE69925932T2 | Germany | T2 | |
| US7219060B2 | United States of America | B2 | |
| EP1501075A3 | European Patent Office (EPO) | A3 | |
| EP1501075B1This record | European Patent Office (EPO) | B1 | |
| DE69940747D1 | Germany | D1 |
31 legal events, as 4 offices reported them to INPADOC
Over the term
Point at a mark for the eventEvents
| Event | Code | Office | |
|---|---|---|---|
| Lapsed in a contracting state [announced via postgrant information from national office to epo]LapsedPG25 | PG25 | EP | |
| Patent expired after termination of 20 yearsExpiredPE20 | PE20 | GB | |
| Expiry of rightR071 | R071 | DE | |
| Annual fee paid to national office [announced via postgrant information from national office to epo]GrantedPGFP | PGFP | EP | |
| Annual fee paid to national office [announced via postgrant information from national office to epo]GrantedPGFP | PGFP | EP | |
| Annual fee paid to national office [announced via postgrant information from national office to epo]GrantedPGFP | PGFP | EP | |
| Fee paymentPLFP | PLFP | FR | |
| Fee paymentPLFP | PLFP | FR | |
| Fee paymentPLFP | PLFP | FR | |
| No opposition filedOpposition26N | 26N | EP | |
| No opposition filed within time limitOppositionORIGINAL CODE: 0009261PLBE | PLBE | EP | |
| Information on the status of an ep patent application or granted ep patentGrantedSTATUS: NO OPPOSITION FILED WITHIN TIME LIMITSTAA | STAA | EP | |
| Corresponds to:REF | REF | EP | |
| Divisional application: reference to earlier applicationAC | AC | EP | |
| Designated contracting statesAK | AK | EP | |
| European patent grantedGrantedFG4D | FG4D | GB | |
| (expected) grantORIGINAL CODE: 0009210GRAA | GRAA | EP | |
| Grant fee paidORIGINAL CODE: EPIDOSNIGR3GRAS | GRAS | EP | |
| Despatch of communication of intention to grant a patentORIGINAL CODE: EPIDOSNIGR1GRAP | GRAP | EP | |
| First examination report despatched17Q | 17Q | EP | |
| Designated contracting statesAK | AK | EP | |
| Deferred search report published (deleted)D17D | D17D | EP | |
| Search report despatchedORIGINAL CODE: 0009013PUAL | PUAL | EP | |
| Information related to the publication of a search report (a3 document) modified or deletedORIGINAL CODE: 0009199SEPUPUAF | PUAF | EP | |
| Designation fees paidAKX | AKX | EP | |
| Request for examination filed17P | 17P | EP | |
| Designated contracting statesAK | AK | EP | |
| Search report despatchedORIGINAL CODE: 0009013PUAL | PUAL | EP | |
| Divisional application: reference to earlier applicationAC | AC | EP | |
| Designated contracting statesAK | AK | EP | |
| Public reference made under article 153(3) epc to a published international application that has entered the european phaseORIGINAL CODE: 0009012PUAI | PUAI | EP |
Numbers
- Publication
- 1501075
- Application
- 40777237
Titles3
- German
- Sprachsynthese mittels Verknüpfung von Sprachwellenformen
- English
- Speech synthesis using concatenation of speech waveforms
- French
- Synthèse de la parole par concaténation de formes d'ondes de parole
Classification
- IPC, 2
- G10L13 07
- G10L13 06
Designated states3
- Contracting states, 3
- Germany
- France
- United Kingdom
