Speech synthesis
Claim Score by NHIP
Abstract
A method of controlling production of an aural advertisement including: enabling user adaptation of audio features that are associated with advertising copy text in a display; and creating a speech synthesizer command dependent upon a textual content of the advertising copy text and the adapted audio features that controls a voice synthesizer to produce an aural advertisement.

Term
Projected expiry 15 February 2027.
- Priority
- Filed
- Published
- Today
- Projected expiry
18 claims: 3 independent, 15 dependent
- 1A method of controlling production of an aural advertisement comprising:enabling user adaptation of audio features that are associated with advertising copy text in a display;and creating a speech synthesizer command dependent upon a textual content of the advertising copy text and the adapted audio features that controls a voice synthesizer to produce an aural advertisement.
- 14An apparatus comprising:a display for displaying advertising copy text;a user input device for adapting audio features that are visually associated with the advertising copy text;a processor for creating voice synthesizer commands that depend upon the advertising copy text and the adapted audio features;and an audio output device for outputting speech synthesized by a voice synthesizer.
- 17Broadest claimClaim Score 89, very broad(NHIP)A method of manufacturing an aural advertisement comprising:receiving advertising copy;and performing speech synthesis of the advertising copy to produce and record a data structure that is operable to control an audio device to render the advertising copy as synthesized speech.
Independent claims3
132 paragraphs in 5 sections, as filed
FIELD OF THE INVENTION
0001Embodiments of the present invention relate to speech synthesis. In particular embodiments of the invention relate to enabling users who are not experts at speech synthesis to control the production of synthesized speech.
BACKGROUND TO THE INVENTION
0002Recordings of speech have many uses. For example an answer phone may require a short section of speech to welcome a caller, another perhaps to say a caller is away at a meeting and when they will be back. The amount of speech sections required for a device varies from a few, in a device such as an answer phone, to hundreds in a call centre application to possibly thousands in an interactive computer game.
0003Recording good quality speech is expensive and often impossible for small enterprises or individuals who cannot afford a recording studio and pay an experienced voiceover artist.
0004A speech synthesis system can be used to create speech. However, the way a sentence is produced by the speech synthesis system may not match a user's requirements. In a recording studio environment the voiceover artist can be instructed to produce the speech section differently. For synthetic speech, techniques do exist to modify the way the speech is produced but they are complex and require an experienced engineer or phonetician to set the input parameters to the algorithms.
0005It would therefore be desirable to provide a technical solution that enables unskilled users to prepare and modify synthetic speech.
0006It would be desirable to provide new applications for speech synthesis by enabling unskilled users to prepare and modify synthetic speech.
BRIEF DESCRIPTION OF THE INVENTION
0007According to some embodiments of the invention there is provided a method of controlling production of an aural advertisement comprising: enabling user adaptation of audio features that are associated with advertising copy text in a display; and creating a speech synthesizer command dependent upon a textual content of the advertising copy and the adapted audio features that controls a voice synthesizer to produce an aural advertisement.
0008According to some embodiments of the invention there is provided an apparatus comprising: a display for displaying advertising copy text; a user input device for adapting audio features that are visually associated with the advertising copy text; a processor for creating voice synthesizer commands that depend upon the advertising copy text and the adapted audio features; and an audio output device for outputting speech synthesized by a voice synthesizer.
0009According to some embodiments of the invention there is provided a control interface for controlling production of an aural advertisement comprising: a display for displaying a advertising copy text; and a user input device for adapting audio features that are associated with the advertising copy text; and means for creating voice synthesizer commands that depend upon the advertising copy text and also the adapted audio features.
0010According to some embodiments of the invention there is provided a method of manufacturing an aural advertisement comprising: receiving advertising copy; and performing speech synthesis of the advertising copy to produce a data structure that is operable to control an audio device to render the advertising copy as synthesized speech.
0011Embodiments of the invention allow novices to produce high quality audio using text to speech synthesis.
0012According to some embodiments of the invention there is provided a method of controlling text to speech synthesis comprising: enabling user adaptation of a presentation of text in a display; creating a speech synthesizer command that depends upon a text content and also upon the presentation of the text and that controls a voice synthesizer to produce synthesized speech that renders the text as speech with characteristics that are dependent upon the presentation of the text.
0013The presentation of text may be dependent upon the visual appearance of the text and independent of any characters used to form the text. The text may comprise phrases and the visual appearance of the text may be dependent upon the visual appearance of each phrase. The visual appearance of a phrase may be dependent upon one or more of typeface, size of typeface, positioning of text, and orientation of text.
0014A presentation of text may be dependent upon the presence of flags in association with the text.
0015The characteristics of synthesized speech may include amplitude, speech rate, pitch and intonation.
0016The speech synthesizer command may control audio features of the synthesized speech.
0017The speech synthesizer command may constrain operation of the speech synthesizer.
0018The text may be comprised of phrases, the speech synthesizer command forcing an alternative rendition of at least one phrase within the text.
0019The text may be comprised of phrases, a command forcing a particular rendition of at least one phrase within the text.
0020The speech synthesizer command may mandate a particular audio feature in the synthesized speech.
0021The speech synthesizer command may comprise mark up text.
0022The method may further comprise: receiving from the voice synthesizer mark up text representing synthesized speech; converting the mark up text into a particular presentation of the text that represents the synthesized speech and its characteristics; and displaying the particular presentation of the text for user adaptation.
0023The method may further comprise: producing synthesized speech that renders the text as speech with characteristics that are dependent upon the presentation of the text.
0024The text may comprise phrases, the method further comprising adapting a presentation of the text in the display by independently adapting the presentation of individual phrases.
0025The method may enable a user to adapt automatically and simultaneously the presentation of multiple independent phrases within the text.
0026According to some embodiments of the invention there is provided a method of controlling the rendering of text as audio using speech synthesis comprising: displaying text with a first presentation; creating a first command that depends upon the first presentation and controls a voice synthesizer to produce first synthesized speech that renders the text and that has first characteristics; enabling user adaptation of the presentation of text to form a second, different, presentation of the same text; creating a second, different, command that depends upon the second presentation of the text and controls the voice synthesizer to produce second synthesized speech that renders the text and that has second characteristics at least some of which are different to the first characteristics.
0027According to embodiments of the invention there is provided an apparatus for controlling text to speech synthesis comprising: a display for displaying a presentation of text; a user input device for adapting the presentation of text; a processor for creating voice synthesizer commands that depend upon the text and also the presentation of the text; and an audio output device for outputting speech synthesized by the voice synthesizer in response to voice synthesizer commands.
0028The user input may be arranged to adapt a presentation of text by varying, for phrases within the text, one or more of a typeface type used for a phrase, a size of typeface used for a phrase, a position of a phrase relative to other phrases and an orientation of a phrase.
0029The text may comprise phrases, the user input enabling a user to adapt a presentation of text by independently adapting the presentation of individual phrases.
0030The user input may be arranged to adapt a presentation of a phrase by applying one or more flags to the phrase
0031A voice synthesizer command may control characteristics of the synthesized speech A voice synthesizer command may control operation of the speech synthesizer.
0032A voice synthesizer command may forces one or more alternative renditions of at least one phrase within the text, a particular rendition of at least one phrase within the text and digital signal processing at the speech synthesizer.
0033The voice synthesizer command may comprise mark up text.
0034The apparatus may be arranged to control the display to display the text with a particular presentation, for user adaptation, that represents synthesized speech output by the output device.
0035According to embodiments of the invention there is provided a control interface for text to speech synthesis comprising: a display for displaying a presentation of text; and a user input device for adapting the presentation of text; and means for creating voice synthesizer commands that depend upon the text and also the presentation of the text.
0036The user input device may be arranged to adapt a presentation of text by varying, for whole phrases within the text, one or more of a typeface type used for a phrase, a size of typeface used for a phrase, a position of a phrase relative to other phrases and an orientation of a phrase.
0037The text may comprise phrases, the user input enabling user adaptation of the presentation of the text on a phrase by phrase basis only.
0038According to embodiments of the invention there is provided a computer program product comprising computer program instructions which when loaded into a processor control the processor to enable user adaptation of a presentation of displayed text; and automatic creation of voice synthesizer commands that depend upon a content of displayed text and also the presentation of the displayed text.
0039According to embodiments of the invention there is provided a speech synthesizer comprising: an input for receiving speech synthesis commands; a synthesizer for performing speech synthesis to produce synthesized speech based upon the received speech synthesize commands; and an output for providing a text based description of the produced synthesized speech.
0040The speech synthesizer may further comprise an interpreter for interpreting received speech synthesis commands to identify specified audio features for the synthesized speech and identify constraints on the synthesis process.
0041The synthesizer process may include unit selection, the synthesizer constraining a result of the synthesis process by mandating the selection of units specified in a speech synthesis command.
0042The synthesizer process may include unit selection, the synthesizer constraining a result of the synthesis process by preventing the selection of specific units in the result.
0043The synthesizer may constrain a result of the synthesis process by mandating that audio features specified in a speech synthesis command are included in a result of a synthesis process. The synthesizer process may include unit selection, the synthesizer performing digital signal processing to achieve the specified audio features if they cannot be obtained by unit selection.
0044According to embodiments of the invention there is provided a speech synthesizer comprising: an input for receiving speech synthesis commands; a synthesizer for performing speech synthesis to produce synthesized speech based upon the received speech synthesize commands; and an interpreter for interpreting received speech synthesis commands to identify specified audio features for the synthesized speech and constraints on the synthesis process.
BRIEF DESCRIPTION OF THE DRAWINGS
0045For a better understanding of the present invention reference will now be made by way of example only to the accompanying drawings in which:
0046<figref idref="DRAWINGS">FIG. 1</figref> schematically illustrates an apparatus for controlling text to speech synthesis for which the speech synthesizer is local;
0047<figref idref="DRAWINGS">FIG. 2</figref> schematically illustrates an apparatus for controlling text to speech synthesis for which the speech synthesizer is remote;
0048<figref idref="DRAWINGS">FIG. 3A</figref> illustrates a presentation of text before user adaptation;
0049<figref idref="DRAWINGS">FIG. 3B</figref> illustrates a new presentation of text, after the presentation illustrated in <figref idref="DRAWINGS">FIG. 3A</figref> has been adapted by a user;
0050<figref idref="DRAWINGS">FIG. 4A</figref> illustrates a presentation of text before user adaptation
0051<figref idref="DRAWINGS">FIG. 4B</figref> illustrates a new presentation of text, after the presentation illustrated in <figref idref="DRAWINGS">FIG. 4A</figref> has been adapted by a user;
0052<figref idref="DRAWINGS">FIG. 5</figref> illustrates a presentation of text after user adaptation;
0053<figref idref="DRAWINGS">FIG. 6</figref> illustrates a mapping between presentation parameters and respective acoustic effects of the presentation parameters on speech synthesis;
0054<figref idref="DRAWINGS">FIG. 7</figref> schematically illustrates a method of controlling text to speech synthesis;
0055<figref idref="DRAWINGS">FIG. 8</figref> schematically illustrates a text-to-speech synthesizer;
0056<figref idref="DRAWINGS">FIG. 9</figref> illustrates a main editing window of a different embodiment in which speech characteristics of synthesized text are controlled by user selection of selectable options positioned adjacent the text; and
0057<figref idref="DRAWINGS">FIG. 10</figref> illustrates a pronunciation window for accessed via the main editing window of <figref idref="DRAWINGS">FIG. 9</figref>.
DETAILED DESCRIPTION OF EMBODIMENTS OF THE INVENTION
0058Some embodiments of the present invention provide a graphical user interface (GUI) that enables intuitive user control of text to speech synthesis. The GUI allows a user to vary how text is synthesized by adapting the presentation of that text. By adapting the presentation of text a user is thus able to manually intervene in automatic speech synthesis of that text. <figref idref="DRAWINGS">FIGS. 3A</figref>, <b>3</b>B, <b>4</b>B and <b>5</b> illustrate different presentations of the same text <b>70</b> and each different presentation controls a different synthesis of the text. The presentation of text is independent of the content of the text i.e. the particular arrangement of characters used. It relates to how the text looks and not what it says.
0059Different presentation features (defined by presentation parameters) result in different speech synthesis control and <figref idref="DRAWINGS">FIG. 6</figref> illustrates an example of how different presentation parameters may be mapped to different speech synthesis characteristics.
0060The GUI enables the complex control of synthesis required to produce desired synthetic speech to be carried out by a user who has no technical background in speech synthesis and speech technology.
0061<figref idref="DRAWINGS">FIG. 1</figref> schematically illustrates an apparatus <b>2</b> for controlling text to speech synthesis and providing the GUI. The apparatus <b>2</b>, in this example, comprises a display <b>4</b> for displaying a presentation of text <b>70</b>; a user input device <b>6</b> for adapting the presentation of text <b>70</b>; a processor <b>8</b> for creating voice synthesizer commands that depend upon the text and also the presentation of the text; and an audio output device <b>10</b> for outputting speech synthesized by the voice synthesizer.
0062The processor <b>8</b> is arranged to receive input control signals from the user input device <b>6</b>, which may be, for example, a computer mouse, a keyboard etc and the processor <b>8</b> is arranged to provide output control signals to the display <b>4</b> and also to the audio output device <b>10</b>. The processor <b>8</b> is also arranged to read from and write to a memory <b>12</b>.
0063The memory <b>12</b> stores a computer program <b>14</b> which when loaded into the processor <b>8</b> enables the processor <b>8</b> to perform steps in the method illustrated in <figref idref="DRAWINGS">FIG. 7</figref>. The programmed processor <b>8</b> provides command means for creating voice synthesizer commands <b>73</b> that depend upon the text <b>70</b> in the display <b>4</b> and also upon the presentation of the text <b>70</b> in the display. This command means in combination with the display <b>4</b> and user input device <b>6</b> provides a control interface via the GUI for text to speech synthesis.
0064The computer program <b>14</b> may arrive at the apparatus <b>2</b> via an electromagnetic carrier signal that may be temporarily stored in a memory buffer or be copied from a physical entity <b>3</b> such as a computer program product, a memory device or a record medium such as a CD-ROM or DVD.
0065The apparatus <b>2</b> in this example has software <b>16</b> which when loaded in the processor <b>2</b> provides a text-to-speech synthesizer. In other examples, a dedicated text-to-speech synthesizer <b>22</b> may be provided in the apparatus. In other examples, a text-to-speech synthesizer <b>22</b> may be located remotely from the apparatus <b>2</b> as illustrated in <figref idref="DRAWINGS">FIG. 2</figref>.
0066In <figref idref="DRAWINGS">FIG. 2</figref>, the apparatus <b>2</b> is substantially as described in relation to <figref idref="DRAWINGS">FIG. 1</figref> except the memory <b>12</b> does not store text-to-speech synthesis software <b>16</b> and the apparatus <b>2</b> further comprises a network adapter <b>18</b>, connected to processor <b>8</b>, that enables the apparatus <b>2</b> to communicate with a remote speech synthesizer <b>22</b> via a network <b>20</b>. The network <b>20</b> is preferably a packet-switched network such as the Internet.
0067The text content ‘What is your date of birth?’ is illustrated with different presentations <b>30</b> in <figref idref="DRAWINGS">FIGS. 3A</figref>, <b>3</b>B, <b>4</b>A, <b>4</b>B and <b>5</b>. The text <b>70</b> is comprised of phrases (e.g. words) <b>32</b>. The presentation <b>34</b> of each phase <b>32</b> may be defined by a set of presentation parameters that are separately and independently adjusted by a user using the user input device <b>6</b>.
0068The user adjustable presentation parameters for a phrase <b>32</b> include its typeface, its font size, its position in a transverse direction relative to other phrases, the orientation of the phrase, the presence of different flags.
0069There is an intuitive mapping between presentation parameters for a phrase <b>32</b> and the speech characteristics of that phrase when it is rendered as a synthetic speech utterance. An example of such a mapping is illustrated in <figref idref="DRAWINGS">FIG. 6</figref>.
0070The mapping of presentation parameters for a phrase to speech characteristics for that phrase enables a user without knowledge of the complex language required to control a speech synthesizer directly to control the text-to-speech synthesis.
0071<figref idref="DRAWINGS">FIG. 6</figref> illustrates a mapping <b>40</b> presented as a table. The first column of the table comprises presentation parameters <b>42</b>. The second column describes the respective effects of the presentation parameters, when applied to a phrase, on the speech characteristics of that phrase when synthesized.
0072The presentation parameters <b>42</b> include a first set of presentation parameters <b>42</b>A that relate to the visual appearance of a phrase and, in particular, the visual appearance of the characters of the phrase. They include, for example, typeface size, transverse position, orientation and character spacing. This first set of presentation parameters <b>42</b>A, when applied to a phrase, specifies audio features for the synthesized speech when that phrase is synthesized. The audio features include amplitude, pitch, intonation and speech rate. In this example, font height is related to amplitude, font length to speech rate, transverse position (height on the page) to pitch, orientation to intonation- upwards to indicate rising pitch, downwards to indicate falling pitch.
0073The presentation parameters <b>42</b> include a second set of presentation parameters <b>42</b>B that identify flags <b>50</b> that can be visually applied to a phrase. They include, for example, an ‘alternative’ flag <b>50</b>A and a ‘fixed’ flag <b>50</b>B. This second set of presentation parameters <b>42</b>B, when applied to a phrase, specifies constraints on the operation of the speech synthesizer <b>22</b>. The alternative flag <b>50</b>A, which is this example is a small graphic symbol such as a cross or skull and crossbones, when associated with a phrase requires a different rendition for that phrase by the speech synthesizer than a previous rendition of that phrase by the speech synthesizer. The fixed flag <b>50</b>B, which is this example is a small graphic symbol such as an anchor or padlock, when associated with a phrase requires that the same rendition for that phrase by the speech synthesizer as the previous rendition.
0074Font colour may be used to mark how strongly a user has a preference for a specific change. The color red (grey in the figures) indicates the strongest preference.
0075The presentation parameters for a phrase are encoded as an XML code portion containing that phrase. The various XML code portions are concatenated to form an XML document which functions as a voice synthesizer command. Each XML code portion defines speech characteristics for the contained phrase.
0076The encoding of presentation parameters for a phrase as an XML code portion containing that phrase may use a new XML tag—the USEL tag. The attributes of the USEL tag are as follows:
0000<tables id="TABLE-US-00001" num="00001"><table frame="none" colsep="0" rowsep="0" pgwide="1"><tgroup align="left" colsep="0" rowsep="0" cols="3"><colspec colname="1" colwidth="35pt" align="left" /><colspec colname="2" colwidth="49pt" align="left" /><colspec colname="3" colwidth="175pt" align="left" /><thead><row><entry namest="1" nameend="3" align="center" rowsep="1" /></row><row><entry /><entry>Permitted</entry><entry /></row><row><entry>Attribute</entry><entry>Values</entry><entry>Function</entry></row><row><entry namest="1" nameend="3" align="center" rowsep="1" /></row></thead><tbody valign="top"><row><entry>variant</entry><entry>0–9</entry><entry>This forces use of a specified synthesis version (i.e. 0</entry></row><row><entry /><entry /><entry>means default, 1 means first alternative from the</entry></row><row><entry /><entry /><entry>default, 2 means second alternative from the default).</entry></row><row><entry /><entry /><entry>In the example of Figs this function is specified using one</entry></row><row><entry /><entry /><entry>or more alternative flag 50A.</entry></row><row><entry>force</entry><entry>TRUE/FALSE</entry><entry>If true the specified audio features for the synthesized</entry></row><row><entry /><entry /><entry>speech are forced to occur using digital signal</entry></row><row><entry /><entry /><entry>processing if unit selection cannot find the correct units.</entry></row><row><entry /><entry /><entry>In the example of Figs this function is specified using a</entry></row><row><entry /><entry /><entry>different color.</entry></row><row><entry>unit_ids</entry><entry>List of ids</entry><entry>Use these items in the database for synthesis rather</entry></row><row><entry /><entry>(e.g. ‘pl p23</entry><entry>than searching the database. This XML can only be</entry></row><row><entry /><entry>p45’)</entry><entry>constructed automatically based on previous synthesis.</entry></row><row><entry /><entry /><entry>In the example of Figs this function is specified using a</entry></row><row><entry /><entry /><entry>fixed flag 50B.</entry></row><row><entry namest="1" nameend="3" align="center" rowsep="1" /></row></tbody></tgroup></table></tables>
0077Commands for changes in amplitude, pitch and duration are translated into industry standard SSML XML using the SSML prosody tag.
0078Although XML has been described any mark-up which connects commands to text could be used.
0079<figref idref="DRAWINGS">FIG. 7</figref> schematically illustrates method <b>50</b> for controlling text to speech synthesis. The figure illustrates the process when the speech synthesizer <b>22</b> and apparatus <b>2</b> are remote as illustrated, for example, in <figref idref="DRAWINGS">FIG. 2</figref>. However, the method <b>50</b> is also applicable if the speech synthesizer <b>22</b> is located within the apparatus <b>2</b> as illustrated, for example, in <figref idref="DRAWINGS">FIG. 1</figref>.
0080At step <b>51</b> text <b>70</b> is input to the apparatus <b>2</b>. The text <b>70</b> may be entered either interactively using the user input device <b>6</b> or from a file.
0081Next at step <b>52</b>, the text <b>70</b> is sent to the speech synthesizer <b>22</b>. At step <b>53</b>, the speech synthesizer performs text-to-speech synthesis producing results <b>70</b> which are sent to the apparatus <b>2</b> at step <b>54</b>. The results <b>72</b> include a data structure <b>72</b>A for rendering synthesized speech as an audio output and an XML document <b>72</b>B that describes the rendered speech.
0082At step <b>55</b>, the processor <b>8</b> of the apparatus <b>2</b> temporarily stores the received data structure <b>72</b>A and the received XML document in the memory <b>12</b>. The processor <b>8</b> uses the received data structure <b>72</b>A to render synthesized speech via the audio output device <b>10</b>. The processor <b>8</b> uses the received XML document <b>72</b>B to display the text <b>70</b> on the display <b>4</b> with a presentation that corresponds to the rendered synthesized speech. The processor <b>8</b> identifies from the received XML document <b>72</b>B the presentation parameters associated with each phrase in the text and adapts the presentation of a phrase on the display <b>4</b> to conform to the associated presentation parameters. The user can therefore listen to the synthesized speech and also simultaneously view a presentation of the text of the synthesized speech that represents the speech characteristics of the synthesized speech.
0083The user at step <b>56</b> decides whether the rendered synthesized speech is acceptable. If the user selects an option indicating that it is acceptable, the data structure <b>72</b>A is saved in the memory <b>12</b> and the method <b>50</b> ends. If the user selects an option indicating that it is not acceptable, a cue editing process begins in which a user is able to modify the speech characteristics of phrases within the text.
0084A step <b>58</b>, user adaptation of the presentation of the text <b>70</b> is enabled. The presentation parameters associated with a phrase <b>32</b> may be adjusted by selecting the phrase <b>32</b> and selecting from a list of presentation parameter options. Selecting a presentation option changes the presentation of the phrase in the display.
0085An expressive speech macro may be selected by a user to define a set of presentation parameters for a group of phrases. For example, an expressive speech macro may be used to convey emotions in the speech such as happiness and sadness. The happiness macro will define speech characteristics (and corresponding presentation parameters) that may increase pitch range, slightly increase the rate of speech and potentially add non standard XML tags used by the speech synthesis system to select more cheerful speech material over the entire series of phrases. The sadness macro will define speech characteristics (and corresponding presentation parameters) that may reduce pitch range, slow rate of speech and add non standard XML tags used by the speech synthesis system to select less cheerful speech material over the entire phrase. Neutral may be another macro.
0086When the user adaptation of the presentation of the text <b>70</b> is finished, at step <b>60</b>, the processor <b>8</b> creates a speech synthesizer command <b>73</b> from the presentation parameters associated with the phrases of the text as described previously. The synthesizer command <b>73</b> is then sent to the speech synthesizer at step <b>61</b>.
0087At step <b>62</b>, the speech synthesizer performs text-to-speech synthesis producing results <b>74</b> which are sent to the apparatus <b>2</b> at step <b>63</b>. The results <b>74</b> include a data structure <b>74</b>A for rendering synthesized speech and an XML document <b>74</b>B that describes the rendered speech.
0088At step <b>64</b>, the processor <b>8</b> of the apparatus <b>2</b> temporarily stores the received data structure <b>74</b>A and the received XML document <b>74</b>B in the memory <b>12</b>. The processor <b>8</b> uses the received data structure <b>74</b>A to render synthesized speech via the audio output device <b>10</b>. The processor <b>8</b> uses the received XML document <b>74</b>B to display the text <b>70</b> on the display <b>4</b> with a presentation that corresponds to the rendered synthesized speech. The processor <b>8</b> identifies from the received XML document <b>74</b>B the presentation parameters associated with each phrase <b>32</b> in the text <b>70</b> and controls the presentation <b>34</b> of a phrase <b>32</b> on the display <b>4</b> to conform to the associated presentation parameters. The user can therefore listen to the synthesized speech and also simultaneously view a presentation of the text <b>70</b> of the synthesized speech that represents the speech characteristics of the synthesized speech.
0089The method then returns to step <b>56</b>. It will therefore be appreciated that a user may make many iterative changes to the synthesized speech. This has the advantages that a user can rectify mistakes or misjudgements easily and a user can enforce their preference for a speech characteristic if it is not provided by the speech synthesizer despite being requested.
0090<figref idref="DRAWINGS">FIGS. 3A and 3B</figref> shows an example of how a user may alter the font position and size to request a change to a phrase. <figref idref="DRAWINGS">FIG. 3A</figref> shows the text <b>70</b> having a presentation that corresponds to the synthesized speech rendered using the data structure <b>72</b>A. The stress is on ‘is’ and ‘birth’ (large font size) with a dull falling intonation pattern. <figref idref="DRAWINGS">FIG. 3B</figref> shows a user modified version of the presentation of the text <b>70</b>. The stress is now on ‘What’, ‘your’ and ‘birth’ with a cheerful rising intonation pattern.
0091<figref idref="DRAWINGS">FIG. 4A</figref> shows the text <b>70</b> having a presentation that corresponds to the synthesized speech rendered using the data structure <b>72</b>A. The stress is on ‘is’ and ‘birth’ (large font size) with a dull falling intonation pattern. <figref idref="DRAWINGS">FIG. 4B</figref> shows the user using ‘preference functionality’ to modify the presentation of the text and control the operation of the speech synthesizer. ‘Birth’ is demanded to have a rising intonation. It is marked in red font <b>36</b> (shown as grey in the figure) indicating that the user is insisting that this rising intonation is applied even if poor synthesis results. The user is happy with the synthesis of ‘your date’ so fixes it with two fixed flags <b>50</b>B. The user dislikes the synthesis for ‘what is’ and requests an unspecified alternative using a alternative flag <b>50</b>A.
0092These changes would produce XML code in the speech synthesis command <b>73</b> in the following format:
0000<tables id="TABLE-US-00002" num="00002"><table frame="none" colsep="0" rowsep="0"><tgroup align="left" colsep="0" rowsep="0" cols="2"><colspec colname="offset" colwidth="28pt" align="left" /><colspec colname="1" colwidth="189pt" align="left" /><thead><row><entry /><entry namest="offset" nameend="1" align="center" rowsep="1" /></row></thead><tbody valign="top"><row><entry /><entry><usel variant=‘l’></entry></row><row><entry /><entry> what is</entry></row><row><entry /><entry></usel></entry></row><row><entry /><entry><usel unit_ids=‘p98 p789 p457 p9 p67 p1234’></entry></row><row><entry /><entry> your date</entry></row><row><entry /><entry></usel></entry></row><row><entry /><entry><prosody contour=’ (O%,+5Hz) (IOO%,+20Hz)‘></entry></row><row><entry /><entry> <usel forcel=‘l’></entry></row><row><entry /><entry> birth</entry></row><row><entry /><entry> </usel></entry></row><row><entry /><entry></prosody></entry></row><row><entry /><entry namest="offset" nameend="1" align="center" rowsep="1" /></row></tbody></tgroup></table></tables>
0093<figref idref="DRAWINGS">FIG. 4</figref> shows an example of a user rejecting default synthesis. In this example the user has listened to 3 alternatives to the rendition of the proper name ‘Cereproc’ before finding a version they like. The three cross symbols are translated into XML code in the speech synthesis command of the following format:
0000<br />welcome to<use1 variant=‘3’>cereproc</use1>
0094The speech synthesizer <b>22</b> may, as schematically illustrated in <figref idref="DRAWINGS">FIG. 8</figref>, comprise a number of functional blocks including a speech synthesis command interpreter <b>80</b>, a descriptor unit <b>86</b> and a speech synthesis system having a front end <b>82</b> and a back end <b>84</b>.
0095The speech synthesis command interpreter <b>80</b> receives and interprets speech synthesis commands <b>73</b>. It extracts the text <b>70</b> from the speech command and provides it to the front-end <b>82</b>. It also interprets the XML within a speech synthesis command to produce commands that control speech characteristics of the synthesized speech. These commands include user specified target audio features such as, for each phrase, one or more of the pitch, duration, amplitude, and units used (if this output is supported by the speech synthesis system) and, possibly, speech synthesis constraints that constrain the synthesis process.
0096At the front end <b>82</b> the text <b>70</b> is normalised and then split into phrases and then phonemes. A phrase is defined as a sequence of speech sounds surrounded by silence although the length of the silence could be very short, for example 5 milliseconds.
0097The back-end <b>84</b> receives the target audio features and the speech synthesis constraints from the speech synthesis command interpreter <b>80</b> and the phrases from the front-end <b>82</b>. The back-end <b>84</b> uses an indexed database of speech sounds (units) from a recorded speaker or speakers. The units for a particular phoneme are indexed by their different speech characteristics such as pitch, intonation, amplitude etc.
0098The synthesis process uses unit selection from the indexed database. Unit selection synthesis works by taking the target audio features and a join or concatenation function which measures how well two sections of speech connect together. A database search is then carried out using the Viterbi algorithm to find the optimal sequence of speech chunks (units) that fulfill the target requirements AND join together well.
0099The back-end unit <b>84</b> provides as its output the data structure <b>74</b>A representing the synthesized speech.
0100The descriptor unit <b>86</b> is connected to the back-end <b>84</b> of the speech synthesizer and it produces XML output <b>74</b>B which describes how the synthesized speech has been realised in terms of pitch, amplitude, duration and units selected. The XML output <b>74</b>B has the same format as speech synthesis commands <b>73</b>.
0101Speech synthesis constraints are used to cope with the possibility that some target audio features may not be realised because either the required units do not exist in the database or units that satisfy the target audio features cannot be joined together well.
0102Certain XML code in a speech synthesis command <b>73</b> indicates that it is preferable for specific speech units to be used to synthesize a phrase typically because such units have been used before and the user has specified using the fixed flag <b>50</b>B that the speech synthesis for that phrase should not change. An example of the XML code is:
0000<tables id="TABLE-US-00003" num="00003"><table frame="none" colsep="0" rowsep="0"><tgroup align="left" colsep="0" rowsep="0" cols="2"><colspec colname="offset" colwidth="35pt" align="left" /><colspec colname="1" colwidth="182pt" align="left" /><thead><row><entry /><entry namest="offset" nameend="1" align="center" rowsep="1" /></row></thead><tbody valign="top"><row><entry /><entry><usel unit_ids=‘p98 p789 p457 p9 p67 p1234’></entry></row><row><entry /><entry> your date</entry></row><row><entry /><entry></usel></entry></row><row><entry /><entry namest="offset" nameend="1" align="center" rowsep="1" /></row></tbody></tgroup></table></tables>
0103In this example, the use<b>1</b> unit_ids attribute specifies a set of units to use for ‘your date’. The back-end <b>84</b> of the speech synthesizer <b>22</b> may respond to receiving this command by pruning out all possible alternative units to the ones specified before unit selection for the text is performed.
0104Certain XML code in a speech synthesis command <b>73</b> indicates a strength of user preference concerning a particular target speech requirements. Typically, a specified target requirement is treated as desirable unless it is flagged <b>36</b> as mandatory. In the absence of suitable units digital signal processing should be used to interpolate and generate a suitable unit. The digital signal processing may be used, for example, to alter amplitude, pitch duration for a unit to enable the target speech requirements to be met. An example of the XML code is:
0000<tables id="TABLE-US-00004" num="00004"><table frame="none" colsep="0" rowsep="0"><tgroup align="left" colsep="0" rowsep="0" cols="2"><colspec colname="offset" colwidth="28pt" align="left" /><colspec colname="1" colwidth="189pt" align="left" /><thead><row><entry /><entry namest="offset" nameend="1" align="center" rowsep="1" /></row></thead><tbody valign="top"><row><entry /><entry><prosody contour=‘ (O%,+5Hz) (IOO%,+20Hz)’></entry></row><row><entry /><entry> <usel force=‘l’></entry></row><row><entry /><entry> birth</entry></row><row><entry /><entry> </usel></entry></row><row><entry /><entry></prosody></entry></row><row><entry /><entry namest="offset" nameend="1" align="center" rowsep="1" /></row></tbody></tgroup></table></tables>
0105The SSML prosody tag requests that the pitch in ‘birth’ is raised and is higher still towards the end. The usel force attribute instructs the back-end <b>84</b> of the speech synthesis system to use a corresponding unit if available or if a corresponding unit is not available to generate units using digital signal processing techniques that force this pitch rise.
0106Certain XML code in a speech synthesis command indicates that it is preferable for specific renditions of a phrase not to be used. This code corresponds to the alternate flag <b>50</b>A. An example of the XML code used to flag a phrase in this way is:
0000<tables id="TABLE-US-00005" num="00005"><table frame="none" colsep="0" rowsep="0"><tgroup align="left" colsep="0" rowsep="0" cols="2"><colspec colname="offset" colwidth="84pt" align="left" /><colspec colname="1" colwidth="133pt" align="left" /><thead><row><entry /><entry namest="offset" nameend="1" align="center" rowsep="1" /></row></thead><tbody valign="top"><row><entry /><entry><usel variant=‘l’></entry></row><row><entry /><entry> what is</entry></row><row><entry /><entry></usel></entry></row><row><entry /><entry namest="offset" nameend="1" align="center" rowsep="1" /></row></tbody></tgroup></table></tables>
0107This flag instructs the back-end speech synthesis system to prune out, after an initial unit selection search, the first unit selection for ‘what is’ and carry out the search again. Re-synthesis occurs N times where N is the value of the variant attribute. It occurs once for ‘what is’ of <figref idref="DRAWINGS">FIG. 4B</figref> and three times for ‘Cereproc’ of <figref idref="DRAWINGS">FIG. 5</figref>). Each time the selected units are pruned out and the synthesis repeated.
0108One application for the apparatus <b>2</b> is as an advertisement creation apparatus.
0109This is a useful practical application in that the apparatus <b>2</b> enables an unskilled user to produce an aural advertisement without having to hire and direct a voice artist. The results of the invention are an advertisement that has been produced at lower cost. The invention therefore represents a technical improvement to a manufacturing process and is a tangible invention.
0110The apparatus <b>2</b> illustrated in <figref idref="DRAWINGS">FIG. 1</figref> is capable of controlling the manufacture or production of an aural advertisement and also capable of manufacturing the aural advertisement. The apparatus <b>2</b> illustrated in <figref idref="DRAWINGS">FIG. 1</figref> is capable of controlling the manufacture or production of an aural advertisement and the remote speech synthesizer <b>22</b> is capable of manufacturing the aural advertisement.
0111At the apparatus <b>2</b>, the display <b>4</b> displays the advertising copy text <b>70</b>. A user input device is used to adapt audio features that are visually associated with the advertising copy text <b>70</b>.
0112Audio features may be visually associated with the advertising copy text <b>70</b> as previously described with reference to <figref idref="DRAWINGS">FIGS. 3A</figref>, <b>3</b>B, <b>4</b>A, <b>4</b>B and <b>5</b> namely via the presentation of the text. In this case, audio features are adapted by adapting the presentation of the text as described above.
0113Audio features may alternatively be visually associated with the advertising copy text <b>70</b> as described below with reference to <figref idref="DRAWINGS">FIGS. 9 and 10</figref> namely using explicit user selectable options that are positioned adjacent the text. In this case, audio features are adapted by changing the selected options as described below.
0114The processor <b>8</b> creates voice synthesizer commands <b>73</b> that depend upon the advertising copy text and the adapted audio features (step <b>60</b> of <figref idref="DRAWINGS">FIG. 7</figref>).
0115The voice synthesizer <b>22</b>, in response to the speech synthesizer command <b>73</b> including advertising copy text, produces synthesized speech encoded in a data structure <b>74</b>A (step <b>62</b> of <figref idref="DRAWINGS">FIG. 7</figref>). The speech synthesizer thus manufacturers an aural advertisement.
0116The data structure <b>74</b>A is stored in memory <b>12</b> by processor <b>8</b> and is then processed by the processor <b>2</b> which then controls the audio output device <b>10</b> to render the advertising copy as synthesized speech (step <b>64</b> of <figref idref="DRAWINGS">FIG. 7</figref>).
0117The display <b>4</b>, user input device <b>6</b> and processor <b>8</b> in combination provide a control interface for controlling the production of an aural advertisement that enables a user to easily produce speech synthesizer commands that depend upon the advertising copy text and also user adapted audio features.
0118As an alternative to adapting audio features by adapting a presentation of displayed text, it is possible to adapt audio features associated with the text of the advertising copy as illustrated in <figref idref="DRAWINGS">FIGS. 9 and 10</figref>.
0119<figref idref="DRAWINGS">FIG. 9</figref> illustrates a main editing window of the control interface <b>90</b>. The text <b>70</b> is displayed (unformatted) in window <b>91</b>. It includes phrases (e.g. words) <b>32</b> and each phrase <b>32</b> has an associated adjacent emphasise option <b>92</b>, a version number option <b>93</b> and an edit option <b>94</b>.
0120The audio features for a phrase <b>32</b>N are controlled by selecting one or more of the options <b>92</b>N, <b>93</b>N or <b>94</b>N associated with that phrase.
0121Selecting the emphasise option <b>92</b> for a phrase has the equivalent effect of increasing the font size in the embodiments illustrated in <figref idref="DRAWINGS">FIGS. 3-5</figref>. It indicates that when this phrase is synthesized it should be synthesized with increased amplitude.
0122Selecting the Mth version number <b>93</b> for a phrase has the equivalent effect of marking the phrase with M ‘alternative’ flags <b>50</b>A in the embodiments illustrated in <figref idref="DRAWINGS">FIGS. 3-5</figref>. It indicates that this phrase should be synthesized M times, with the combination of units selected during synthesis of the phrase being prevented from being selected in subsequent re-synthesis.
0123Selecting the edit option <b>94</b> associated with a phrase opens a pronunciation window <b>95</b> for the phrase as illustrated in <figref idref="DRAWINGS">FIG. 10</figref>. This window identifies the syllables <b>96</b> of the phrase and it allows a user to scroll and select, for individual syllables, alternative phonemes <b>97</b>. It also allows a user to select options <b>98</b> for emphasising individual syllables. A play option is provided for each syllable so that a user can understand the effect of any adaptation to the audio features of the syllable.
0124A text input window <b>99</b> is also provided where an expert user can explicitly enter phonemes.
0125Although embodiments of the present invention have been described in the preceding paragraphs with reference to various examples, it should be appreciated that modifications to the examples given can be made without departing from the scope of the invention as claimed.
0126Whilst endeavoring in the foregoing specification to draw attention to those features of the invention believed to be of particular importance it should be understood that the Applicant claims protection in respect of any patentable feature or combination of features hereinbefore referred to and/or shown in the drawings whether or not particular emphasis has been placed thereon.
Contents5
5 sheets
Sheet 1 Sheet 2 Sheet 3 Sheet 4 Sheet 5
Every citation, both ways
| Document | Relation | Office | Cited during |
|---|---|---|---|
| US2010088097A1 | Cited by | United States of America | Pre-grant |
| US2008235025A1 | Cited by | United States of America | Pre-grant |
| US2010312563A1 | Cited by | United States of America | Pre-grant |
| US2019019497A1 | Cited by | United States of America | Search report |
| US2010312565A1 | Cited by | United States of America | Pre-grant |
| US10636412B2 | Cited by | United States of America | Applicant |
| US8332225B2 | Cited by | United States of America | Applicant |
| US10079011B2 | Cited by | United States of America | Search report |
| US2014257818A1 | Cited by | United States of America | Pre-grant |
| US8433573B2 | Cited by | United States of America | Search report |
| US5799279A | Cites | United States of America | Pre-grant |
| US6085161A | Cites | United States of America | Pre-grant |
| US6564186B1 | Cites | United States of America | Pre-grant |
| US6810378B2 | Cites | United States of America | Pre-grant |
| US6856958B2 | Cites | United States of America | Pre-grant |
2 members in 2 offices
Priority claims4
| Document | Office | Kind | Date |
|---|---|---|---|
| 0624474 | United Kingdom | A | |
| 06244743 | United Kingdom | – | |
| 06244743 | – | – | – |
| GB20060024474 | – | – | – |
Members2
| Document | Office | Kind | |
|---|---|---|---|
| GB2444539A | United Kingdom | A | |
| US2008140407A1 | United States of America | A1 |
20 transactions on the USPTO file
Abandoned after 1 non-final rejection.
- Non-final rejections
- 1
- Final rejections
- 0
- RCEs
- 0
- Appeals
- 0
Over time
Point at a mark for the transactionTransactions
| Event | Code | |
|---|---|---|
| Mail Abandonment for Failure to Respond to Office ActionAbandonedMABN2 | MABN2 | |
| Aband. for Failure to Respond to O. A.AbandonedABN2 | ABN2 | |
| Mail Non-Final RejectionNon-final rejectionMCTNF | MCTNF | |
| Non-Final RejectionNon-final rejectionCTNF | CTNF | |
| Case Docketed to Examiner in GAUDOCK | DOCK | |
| Case Docketed to Examiner in GAUDOCK | DOCK | |
| PG-Pub Issue NotificationPG-ISSUE | PG-ISSUE | |
| Request for Foreign Priority (Priority Papers May Be Included)RQPR | RQPR | |
| Case Docketed to Examiner in GAUDOCK | DOCK | |
| IFW TSS Processing by Tech Center CompleteTSSCOMP | TSSCOMP | |
| Application Dispatched from OIPEOIPE | OIPE | |
| Sent to Classification ContractorPGPC | PGPC | |
| Application Is Now CompleteCOMP | COMP | |
| Additional Application Filing FeesADDFLFEE | ADDFLFEE | |
| A statement by one or more inventors satisfying the requirement under 35 USC 115, Oath of the ApplicOATHDECL | OATHDECL | |
| Applicant has submitted new drawings to correct Corrected Papers problemsCORRDRW | CORRDRW | |
| Notice Mailed--Application Incomplete--Filing Date AssignedINCD | INCD | |
| Cleared by OIPE CSRL194 | L194 | |
| IFW Scan & PACR Auto Security ReviewSCAN | SCAN | |
| Initial Exam Team nnIEXX | IEXX |
2 legal events, as the office reported them to INPADOC
Over the term
Point at a mark for the eventEvents
| Event | Code | |
|---|---|---|
| Information on status: application discontinuationABANDONED -- FAILURE TO RESPOND TO AN OFFICE ACTIONSTCB | STCB | |
| AssignmentAS | AS |
Numbers
- Publication
- 20080140407
- Publication, DOCDB
- 2008140407
- Publication, EPODOC
- US2008140407
- Application
- 11706770
- Application, DOCDB
- 70677007
- Application, EPODOC
- US20070706770
Titles
- English
- Speech synthesis
Classification
- CPC, 2
- G10L21/06
- G10L13/00
- IPC, 3
- G10L13 00
- G10L13 04
- G10L21 06
- USPC, 5
- 704260000
- 704E13004
- 704E13008
- 704E13013
- 704E21019