Automated extraction of semantic content and generation of a structured document from speech
6 claims: 4 independent, 2 dependent
- 1Zastrzeżenia patentowe 1. Sposób obejmujący etapy:(A) identyfikowania probabilistycznego modelu języka obejmującego wiele probabilistycznych modeli języka powiązanych z wieloma podstrukturami dokumentu;oraz (B) wykorzystywania dekodera rozpoznawania mowy aby zastosować probabilistyczny model języka w mówionym strumieniu audio, aby utworzyć dokument obejmujący zawartość zorganizowaną w wiele podstruktur, przy czym zawartość w każdej z wielu podstruktur jest utworzona przez rozpoznawanie mowy wykorzystujące probabilistyczny model języka powiązany z podstrukturą, przy czym wiele probabilistycznych modeli języka jest zorganizowanych w hierarchii, oraz przy czym etap (B) zawiera etapy: (B) (1) identyfikowania ścieżki przez hierarchię, zawierającego etapy: (B) (1) (a) identyfikowania wielu ścieżek przez hierarchię;(B) (1) (b) dla każdej z wielu ścieżek P, tworzenia potencjalnego strukturalnego dokumentu dla mówionego strumienia audio przez wykorzystywanie dekodera rozpoznawania mowy do rozpoznawania mówionego strumienia audio wykorzystującego modele języka na ścieżce P;(B)(1)(c) zastosowania metryki do wielu potencjalnych strukturalnych dokumentów utworzonych w etapie (B) (1) (b) , aby utworzyć wiele ocen przydatności dla wielu potencjalnych strukturalnych dokumentów;oraz (B) (1) (d) wybierania ścieżki, która utworzy potencjalny strukturalny dokument posiadający najwyższą ocenę przydatności (B)(2) generowania dokumentu posiadającego strukturę odpowiadającą ścieżce zidentyfikowanej w etapie (B) (1) .
- 2Sposób według zastrz. 1, w którym wiele probabilistycznych modeli języka obejmuje co najmniej jeden n-gramowy model j ęzyka.
- 3Sposób według zastrz. 1, w którym wiele probabilistycznych modeli języka obejmuje co najmniej jeden model języka skończonego stanu.
- 4Sposób według zastrz. 1, w którym wiele podstruktur obejmuje podstrukturę stanowiącą semantyczną koncepcję.
- 5Sposób według zastrz. 4, w którym semantyczna koncepcja zawiera kurację.
- 6Sposób według zastrz. 1, ponadto obejmujący etap:(C) interpretowania dokumentu, aby utworzyć interpretację wskazującą strukturę dokumentu. Mówiony strumień audio 100 Urządzenie do rozpoznawania mowy 104 FIG. 1A (STAN TECHNIKI) FIG. 1B 200 202 204 FIG. 2 302 φ S) Φ Φ φ 4* Λ (Λ 3 Φ Λ Ω. Ο 2 Φ Ł 4* η Ξ * 3 Ε ** ί ο σι φ φ Έ Φ φ 4* Φ 3 Φ φ. Ε π 4* C Φ Φ α α. ο u Ά ο^ W 2 φ 'α. σ ο _ φ α. ο c Φ ο π w Ε φ ο. ο χθ § Ο 3 3) .2 Ń ϋ Ł. Φ Λ Ο Έ π φ C φ 3 -V α. ο φ ϋ Φ Π § ϋ 01 Q ω α. (Μ Ο Ο Ν Φ ϋ Φ Ε ω w Ν * ϋ Φ Ν s_ α Φ c ο Έ φ Ε Ν Φ S ϋ φ Ν Ο. φ ο FIG. 4 502 310 Ό CM οι co CS CN ΤΤ Λ --Ο Ε β 3 β Ν Λ β β ν CN es φ Ε ε co 91 Φ 3 β β β ε φ ζ 0 CM co CS I ΙΟ Λ 9 β VI •Ν Φ V ΟΙ Ol σι θ' s 91 Φ Ε σι π α σι η β 3 β Ν V CS V Ν ΟΙ Ο ΟΙ σ s cs cs •Ν Φ β Ν Φ 3 £ α ίο Ν Λ ' β 3 β Ν V 1/J Ε β 3 β Ν V β 91 β β s 'φ 91 Φ 3 β ν cs Λ Λ ο φ β Ν Φ u Φ β ε φ 0 β ε φ ζ β Ν ε ε u φ Ε β 3 β Ν Λ φ (Λ Φ VI Ν β ε φ ζ ν s V» 9ΐ φ Ν CS β 3 β Ν Λ β δ ν β σι β β β ε Ν Φ ο 2 φ α Φ Ν υ Ο 9» Φ Ν ω φ Ε β e m ri Ε β 3 ν (Λ CS V 318 FIG. 6 FIG. 7 320 312a 306a 808a FIG. 9 304 304 1030 FIG. 11A FIG. 11B 1156 FIG.11C FIG. 11D _ Selektor ścieżki 1202 dokumentu 308 FIG. 12A 1300 FIG. 13 1400 Rozpoznaj mówiony strumień audio, aby utworzyć transkrypcję 1402 Złóż transkrypcję używając hierarchiczny model języka, aby tworzyć—1404 strukturalny dokument tekstowy FIG. 14 1500 FIG. 15 DOKUMENTY PRZYTOCZONE W OPISIE Lista przytoczonych przez Zgłaszającego dokumentów została zamieszczona wyłącznie do informacji czytelnika i nie stanowi części składowej europejskiego dokumentu patentowego. Została ona zestawiona z największą starannością;EUP nie ponosi jednakże żadnej odpowiedzialności za ewentualne błędy lub braki. Literatura patentowa przytoczona • US 20020123891 A [0019]
Independent claims6
159 paragraphs in 4 sections, as filed
Description
REFERENCE TO RELATED NOTIFICATIONS
[0001] This application is related to the co-filed US patent application entitled "Document Transcription System Training. BACKGROUND"
Field of the Invention
[0002] The invention relates to automatic speech recognition and, more particularly, to automatic speech transcription technology. Condition tecńniki
[0003] It is desirable, in many contexts, to generate a written document from human speech. In the legal profession, e.g. transcriptionists write down statements made in court and out-of-court proceedings in order to present a written record of the testimony. Similarly, in the medical profession, transcripts are made of diagnoses, prognoses, prescriptions, and other information dictated by doctors and other medical professionals. Transcripts in these and other fields usually need to be very accurate (measured in terms of the degree of correspondence between the semantic content (signification) of the original word and the semantic content of the resulting transcript) due to the fact that one relies on the transcription received and the damage that could be caused by due to inaccuracies (e.g., delivery of an incorrect prescription drug to a patient). However, a high degree of reliability may be difficult to achieve for many different reasons, such as differences in: (1) the characteristics of the people whose speech is being processed (e.g., accent, voice volume, dialect, speed of speech), (2) external conditions (e.g. background noise), (3) transcriptionist or transcription system (e.g. imperfect hearing or sound capture capabilities, imperfect language comprehension), or (4) recording / transmission medium (e.g. paper, analog audio cassette, analog telephone network, compression algorithms used in digital telephone networks, and sounds / artifacts resulting from cell phone channels).
[0004] Initially, transcription was only performed by the human transcriptionist who listened to the speech, either in real time (i.e. personally "taking notes) or listening to the recording. One advantage is that the human transcriptionist may possess specific knowledge, e.g., knowledge of medicine and medical terminology, which enables the interpretation of ambiguities in speech, thereby improving the accuracy of the transcript. Such transcriptionists, however, have many disadvantages. For example, they write transcripts relatively slowly, their accuracy decreases over time due to fatigue.
[0005] There are various automated speech recognition systems for recognizing human speech in general and for speech transcription in particular. Speech recognition systems that create transcriptions are referred to herein as "automatic transcription systems or" automatic dictation systems. Out-of-the-box dictation software, for example, can be used by personal computer users to dictate documents in a word processor, as an alternative to typing such documents using the keyboard.
[0006] Automated dictation systems typically attempt to record speech-in-word. Such a record in which there is a one-to-one mapping between the words spoken in the audio stream and the words in the transcription is referred to herein as "verbatim record." Automated dictation systems are not perfect and may not produce perfectly accurate transcripts.
[0007] However, in some cases a literal notation is not desirable. In fact, maybe transcriptionists may deliberately make various changes to written transcription. The transcriptist can, for example, filter out spontaneous speech effects (e.g. pauses, fluctuations, and false starts), dismiss irrelevant remarks and comments, convert the data to a standard format, insert headings and other support material, change the order of speech to suit the structure of a written report .
[0008] In the medical field, for example, oral reports made by physicians are often transcribed in the form of written reports having standard formats. For example, referring to FIG. IB, an example of a structured and formatted medical report 111 is shown. Report 111 includes a plurality of sections 112-138 that appear in a specific order when report 111 is displayed. In the particular example shown in FIG. IB, the report includes a header section 112, a private section 122, a target section 134, an evaluation section 136, and a scheduling section 138. The sections may include text as well as subsections. For example, the header section 112 includes a hospital name section 120 (containing the text "General Hospital), a patient name section 114 (containing the text" Jane Doe), a table number section 116 (containing the text "851D), and a report date section 118 (containing the text" Jane Doe) " "10/1/1993).
[0009] Similarly, private section 122 includes various private patient information, including both textual and in medical history sections 124, medications section 126, allergy sections 128, family history sections 130, and social history sections 132. Objective section 134 includes various objective sections. patient information such as weight and blood pressure. Although not shown in FIG. IB, the information in the target section may include sub-sections containing the information presented. The rating section 136 includes a textual assessment of the patient's condition, and the subsection plan 138 includes a textual description of the treatment plan.
[0010] It should be noted that the information may appear in a different form in the report 111 from that information that was spoken by the dictating physician. For example, the date in the date section 118 of a report might be pronounced "October, 1st 1993," 1st October 93, or some other form. The transcriptist, however, rewrites such speech as "1/10/1993 in the date section 118 of the report because the hospital specified in hospital section 120 requires such a date in the written report expressed in such format.
[0011] Similarly, the information in the medical report 111 may not be in the same order as the original audio recording, due to the need to match report format requirements or other reasons. For example, the dictating physician may first dictate the target section 134, then the private section 122, then the header 120. The written report 111, however, includes header 120 first, then private section 122, then the target section.
134. Such a report structure may, for example, be required for medical reports in a hospital identified in hospital section 120.
[0012] Report start 111 could be generated based on a spoken audio stream such as: "this is doctor smitń u October 1st, 1993, patient ID is 85 1 d um then the patient's family history has been reviewed ... It is clear that a literal transcription of this speech would be difficult to understand and not useful.
[0013] It should be noted, for example, that specific words such as "next is a" do not appear in the written report 111. Likewise, statements with pauses such as "uh do not appear in the written report 111. Additionally, the written report 111 composes the original text. speech into specific sections 112-140 by pre-arranging speech. These examples demonstrate that Written Report 111 is not an exact transcription of speech dictated by a physician.
[0014] In summary, a report such as report 111 may be more desirable than a verbatim transcription for a number of reasons (e.g., because it organizes the information so that it can be understood). It would therefore be desirable that the automated transcription system could generate a structured report (not a verbatim transcription) based on unstructured speech.
[0015] Referring to FIG. 1A, a data flow diagram of a prior art system 100 for generating a structured document 110 based on spoken audio stream 102 is shown. Such a system creates a structured text document 110 from spoken audio stream 102 using a two-step process: (1) automatic speech recognition device 104 generates verbatim transcription 106 based on spoken audio stream 102; and (2) natural language processor 108 identifies a structure in transcription 106 and thereby creates a structured document 110 that has the same content as transcription 106, but which is structured in a structure (e.g., report format) identified by natural language processor 108.
[0016] For example, some existing systems attempt to generate a structured text document by: (1) analyzing the spoken audio stream 102 to recognize and highlight the spoken content in the audio stream 102 with conditional or unconditional structured cues in the audio stream 102; (2) converting a portion of the content of the spoken audio stream 102 into raw text; and (3) using the identified structural clues to convert raw text into a structured report 110. Examples of conditional structural clues include formatting commands (e.g., "new paragraph," newline, "next item) and paragraph identifiers (e.g.," discoveries "). , "Conclusions," summary). Examples of unconditional structural clues include long pauses that can signify paragraph boundaries, prosodic clues that indicate ending numbers, and the spoken content itself.
[0017] For a variety of reasons, described in more detail below, the structured document 110 formed by system 100 may be suboptimal. For example, structured document 110 may contain incorrect rewrites (i.e., unrecognized) words, and the structure of structured document 110 may not reflect the desired structure of the document, and the content of the spoken audio stream 102 may be inserted in the wrong substructure (e.g., section, paragraph, or sentence) in a structured document.
[0018] Furthermore, in addition to or instead of generating a structured document 110 based on the spoken audio stream 102, it may be desirable to extract semantic content (such as information about medications, allergies, or a patient's previous disease described in audio stream 102) from the spoken audio stream 102. Although such semantic content may be useful for generating a structured document 110, such content may also be useful for other purposes, such as supplementing a database with patient information that may be analyzed in a stand-alone document 110. Prior art systems such as the system 100 shown in FIG. 1, however, are typically designed to generate a structured document 110 based primarily or solely on syntactic information in the spoken audio stream 102. Such systems, therefore, are not useful for extracting semantic content.
[0019] Therefore, improved techniques for generating structured documents based on spoken audio streams are required.
[0019a] US 2002/0123891 discloses a method for converting speech to text using a hierarchy of contextual models. In the disclosed method, the contextual model that best reflects one or more statements is used to convert speech to text. The preamble of the independent claims is based on this document.
SUMMARY
[0020] Techniques for automatically generating structured documents based on speech are disclosed, including the identification of essential concepts and their interpretation. In one embodiment, the structured document generator uses an integrated process to generate a structured text document (such as a structured text medical report) based on a spoken audio stream. Spoken audio stream can be identified using a language model that has a plurality of sub-models arranged in a hierarchical structure. Each of the sub-models may correspond to a concept that may appear in the spoken audio stream. For example, the sub-models may correspond to sections of the document. The sub-models may, for example, be n-gram models of a language or a context-free grammar. Different parts of the spoken audio stream can be recognized using different sub-models. The resulting structured text document may have a hierarchical structure that corresponds to the hierarchical structure of the language submodels that were used to generate the structured text document.
[0021] For example, in one aspect of the invention, there is provided a method according to independent claim 1.
[0022] Other features and advantages of the various aspects and embodiments of the invention will become apparent from the following description and claims.
BRIEF DESCRIPTION OF THE DRAWINGS
[0023] FIG. 1A is a system data flow diagram from state of the art to generating a structured document based on a spoken audio stream;
[0024] FIG. IB shows a text medical report generated from a spoken report;
[0025] FIG. 2 is a diagram of a method that is performed in one embodiment of the invention to generate a structured text document based on a spoken document;
[0026] FIG. 3 is a data flow diagram of a system that performs the method of FIG. 2 in one embodiment of the invention;
[0027] FIG. 4 shows an example of a spoken audio stream in one embodiment of the invention;
[0028] FIG. 5 shows a structured text document according to one embodiment of the invention;
[0029] FIG. 6 is an example of a composite document that is assembled based on the structured text document of FIG. In accordance with one embodiment of the invention;
[0030] FIG. 7 is a diagram of a method that is performed by the structured document generator of FIG. 3 in one embodiment according to the invention to generate a structured text document;
[0031] FIG. 8 is a data flow diagram illustrating part of the system of FIG. 3 with details relevant to the method of FIG. 7 in accordance with one embodiment of the invention;
[0032] FIG. 9 is a diagram illustrating mappings between language models, document substructures corresponding to language models, and content candidates created with language models in accordance with one embodiment of the invention;
[0033] FIG. 10A is a diagram illustrating a hierarchical language model in accordance with one embodiment of the invention;
[0034] FIG. 10B is a diagram illustrating paths through the hierarchical language model of FIG. 10A in accordance with one embodiment of the invention;
[0035] FIG. 10C is a diagram illustrating a hierarchical language model according to another embodiment of the invention;
[0036] FIG. 11A is a diagram of a method that is performed by the structured document generator of FIG. 3 to generate a structured text document in accordance with one embodiment of the invention;
[0037] FIG. 11B is a flowchart of a method that uses an integrated process to select paths through a hierarchical language model and to generate a speech-based structured text document in accordance with one embodiment of the invention;
[0038] FIG. 11C-11D are flowcharts of methods that are performed in one embodiment of the invention to calculate a suitability score for a candidate document;
[0039] FIG. 12A is a data flow diagram illustrating part of the system of FIG. 3 with details relevant to the method of FIG. 11A in accordance with one embodiment of the invention;
[0040] FIG. 12B is a data flow diagram illustrating an embodiment of the document structure generator of FIG. 3 which executes the method of FIG. 11B in one embodiment of the present invention;
[0041] FIG. 13 is a diagram of a method that is used in one embodiment of the invention to generate hierarchical language models for use in generating a structured text document;
[0042] FIG. 14 is a diagram of a method that is used in one embodiment of the invention to generate a structured text document using separate speech recognition and structured analyzing steps; and
[0043] FIG. 15 is a data flow diagram of a system that performs the method of FIG. 14 in accordance with one embodiment of the invention.
DETAILED DESCRIPTION
[0044] Referring to FIG. 2, there is shown a flow chart of a method 200 that is performed in one embodiment of the invention to generate a structured text document based on a spoken document. With reference to FIG. 3, a data flow diagram of a system 300 for performing the method 200 of FIG. 2 according to one embodiment of the invention.
[0045] System 300 includes a spoken audio stream 302 which may, for example, be a live or recorded spoken audio stream of a medical report dictated by a physician. With reference to FIG. 4, a textual representation of an example of spoken audio stream 302 is shown. In FIG. 4, the text between the percent signs represents spoken punctuation (e.g., "% comma%,"% break%, and "% colon%) and clear structured cues (e.g.,"% speech-paragraph%) in the audio stream 302. As seen in the audio stream 302 shown in FIG. 4, a verbatim transcription of the audio stream 302 would not be useful for the purpose of understanding a diagnosis, prognosis, or other information contained in a medical report represented by the audio stream 302.
[0046] System 300 also includes a probabilistic language model 304. The term "probabilistic language model" as used herein refers to any language model that assigns probabilities to the order of a spoken word.
Context-free (probabilistic) grammar and ngram models of language 306a-e are examples of "probabilistic language models, for the term as used herein.
[0047] Generally, a context-free grammar identifies multiple spoken forms for concepts and assigns probabilities to each of the spoken forms. Finite-state grammar is an example of context-free grammar. For example, a finite state grammar for the date October 1, 1993, may include the spoken form "October 1st 1993 with a probability of 0.7, the spoken form" ten one 93 with a probability of 0.2, and the spoken form "October 1st 93 with a probability of 0, 1. The probabilities associated with each spoken form are the estimated probabilities that the concept will be spoken in spoken form on the specified audio stream. A finite state grammar, therefore, is one type of probabilistic language model.
[0048] Generally, the n-gram model of a language determines the probabilities by which n words will appear in a spoken audio stream in a certain order. Considering, for example, the "unigram language model for which n = l. For each word in a language, unigram determines the probability that the word will occur in the spoken document. The "bigram" language model (for which n = 2) determines the probability that the word pairs will occur in the spoken document. For example, the bigram model may determine the conditional probabilities that the word "cat will occur in the spoken document, assuming the previous word in the document was" this. Similarly, the trigram language model determines the probabilities of word third order, and so on. The probabilities determined by the n-gram language models and the finite state grammar can be obtained by teaching such documents to use learned speech and practice text as described in more detail in the above-indicated patent application entitled, "Document Transcription System Training. [0049] The probabilistic tongue model 304 includes multiple submodels 306a-e, each of which is a probabilistic tongue model. The sub-models 306a-e may include n-gram language models and / or finite-state grammar in any combination. Furthermore, as described in more detail below, each of the sub-models 306a-e may further include sub-models, and so on. Although five sub-models are shown in FIG. 3, the probabilistic language model 304 may include any number of sub-models.
[0050] The purpose of the system 300 shown in FIG. 3 is to create a structured text document 310 that includes content from spoken audio stream 302, the content being organized into a specific structure, and where the concepts are identified and interpreted in a device-readable form. The structured text document 310 includes multiple substructures 312a-f, such as sections, paragraphs, and / or sentences. Each of substructures 312a-f may further include substructures, and so on. Although six substructures are shown in FIG. 3, the structured text document 310 may include any number of sub-structures.
[0051] For example, referring to FIG. 5, an example of a structured text document 310 is shown. In the example shown in FIG. 5, structured text document 310 is an XML document. The structured text document 310 may, however, be implemented in any format. As shown in FIG. 5, structured document 310 includes six substructures 312a-f, each of which may represent a document section 310.
[0052] For example, structured document 310 includes a header section 312a that includes document metadata 310, such as the title 314 of the document 310 ("no contrast chest CT) and the date 316 on which the document 310 was dictated (" <data> 22-APR-2003 </data> "). Note that the content in the header section 312a was obtained from the start of the spoken audio stream 302 (FIG. 4). Further, it should be noted that the header section 312a includes both a plain text (i.e., title 314) and a substructure (e.g., date 316) representing the concept that has been interpreted in device-readable form as a triplet of values (day-month-year). .
[0053] Displaying the date in machine-readable form allows the date to be stored in a database with ease and processed more easily than when the date is stored as text. For example, if multiple dates in the audio stream 302 have been recognized and stored in device-readable form, such dates can easily be compared to each other by a computer. As another example, statistical information on the content of the audio stream 302, such as the mean time between visits to a physician, can be easily generated if the dates are stored in computer-readable form. This advantage of the inventive embodiments is generally applicable not only to the date, but to recognizing any semantic content and saving such content in a device-readable form.
[0054] Structured document 310 further includes a comparison section 312b that includes content describing previous studies performed on the same patient as the subject of the document 310. Note that the content in the compare section 312b was obtained from the portion of the audio stream 302 beginning with "compare to and ending" on April 6, 2001, but the comparison section 312b does not include the text "compare to", which exemplifies the guidance section. The use of such cues to recognize a section start or other document substructure will be described in more detail below.
[0055] In summary, the structured document 310 also includes a technique section 312c that describes techniques that were performed in procedures performed on a patient; discovery section 312d, which describes the physician's findings; and an application section 312e that describes the physician's conclusions related to the patient.
[0056] XML documents, such as, for example, the structured document 310 shown in FIG. 5, are not normally intended for viewing by the end user. On the contrary, such documents are conventionally folded in a format which is more readable to the end user. System 300, for example, includes a folding engine 314 that assembles a structured text document 310 based on the style sheet 316 to create a composite document 318. Techniques for generating style sheets and composing documents according to style sheets are well known to those skilled in the art.
[0057] Referring to FIG. 6, an example of a composite document 318 is shown. The composite document 318 includes five sections 602a-e, each of which may correspond to one or more of the six substructures 312a-f in the structured text document 310. More specifically, the composite document 318 includes a header section 602a section. comparisons 602b, technique section 602c, discovery section 602d, and conclusion section 602e. It should be noted that there may or may not be a one-to-one mapping between the sections in the composite document 318 and the substructures in the structured text document 310. For example, each of the substructures 312af need not represent a distinct type of document section. If, for example, two or more substructures 312af represent the same type of section (such as a header section), the folding engine 314 may place both substructures in the same section of composite document 318.
[0058] System 300 includes a structured document generator 308 that identifies a probabilistic language model 304 (step 202), and uses the language model 304 to recognize spoken audio stream 302 and thereby to form a structured text document 310 (step 204). Structured document generator 308 may, for example, include an automatic speech recognition decoder 320 that creates each of the substructures 312a-f in structured text document 310 using the corresponding of the probabilistic language model 304. As is known to those skilled in the art, the decoder is part of the component of a speech recognition device that converts audio into text. The decoder 320 may, for example, form a sub-model 312a by using the sub-model 306a to recognize the first part of the spoken audio stream 302. Likewise, the decoder 320 may form the sub-structure 312b by using the sub-model 306b to recognize the second part of the spoken audio stream 302.
It should be noted that there is no need for a one-to-one mapping between the sub-models 306a-e of the language model 304 and the sub-structures 312a-f in structured document 310. For example, a speech recognition decoder may use the sub-model 306a to recognize the first part of spoken audio stream 302 and thus form a substructure 312a, and use the same sub-model 306a to recognize the second part of spoken audio stream 302 and thereby form a substructure 312b. In such a case, multiple substructures in structured text document 310 may have content for a single semantic structure (e.g., a section or paragraph).
[0060] Sub-model 306a may, for example, be a "header language model" that is used to recognize portions of spoken audio stream 302 including content in header section 312a; sub-model 306b may, for example, be a comparison language model, which is used to recognize portions of spoken audio stream 302 including content in comparison section 312b; and so on. Each such language model can be taught to use learned text from the relevant sections of the training documents. For example, a header sub-model 306a can be trained to use text from the header section of a plurality of training documents, and a comparison sub-model can be trained to use text from a comparison section of a plurality of training documents.
Having generally described the features of various embodiments according to the invention, embodiments according to the invention will now be described in more detail. With reference to FIG. 7, a method diagram is shown that is performed by a structured document generator 308 in one embodiment of the invention to generate a structured text document 310 (FIG. 2, step 204). With reference to FIG. 8, a data flow diagram is shown illustrating a portion of system 300 in detail relevant to the method of FIG. 7.
[0062] In the example shown in FIG. 8, the structured document generator 308 includes a segment identifier 814 that identifies a plurality of S segments 802ac in spoken audio stream 302 (step 701). 802a-c segments may, for example, represent concepts such as sections, paragraphs, sentences, words, dates, times, or codes. Although only three segments 802a-c are shown in FIG. 8, spoken audio stream 302 may include any number of parts. Although for simplicity, all of the segments 802a-c are identified at step 701 of FIG. 7, prior to executing the remainder of method 700, segment identification 802ac may be performed concurrently with recognizing the audio stream 302 and generating the structured document 310, as will be described in more detail below with reference to FIG. 11B and 12B.
[0063] Structured document generator 308 loops each segment S in spoken audio stream 302 (step 702). As described above, structured document generator 308 includes a speech recognition decoder 320, which may, for example, include one or more conventional speech recognition decoders for speech recognition using different kinds of language models. As further described above, each of the 306ae sub-models can be an Ngram language model, a context-free grammar, or a combination thereof.
For example, assume that the structured document generator 308 is currently processing spoken audio stream segment 802a 302. The structured document generator 308 selects a plurality of 804 of the sub-models 306a-e that recognizes the current segment S. Sub-models 804 may, for example, be all of the 306a-e sub-models or a subset of the 306a-e sub-models. Speech recognition decoder 320 recognizes the current S segment (e.g., segment 802a) with each selectable submodel 804, thereby creating multiple content candidates 808 corresponding to segment S (step 704). In other words, each of the content candidates 808 is formed by using a speech recognition decoder 320 to recognize the current S segment by using one of the sub-models 804 separately. Note that each of the content candidates 808 may include not only recognized text, but also other types of content such as concepts (e.g., dates, times, codes, medications, allergies, signs, etc.) encoded in device-readable form. .
[0065] Structured document generator 308 includes a final content selector 810 that selects one of the content candidates 808 as final content 812 for segment S (step 706). Final content selector 810 may use any techniques that are well known to those skilled in the art to select the speech recognition output that best matches the speech from which the outputs are derived.
Structured document generator 308 keeps track of the sub-model that is used to create each of the content candidates 808. Suppose, for example, that the sub-models 304 include all of the sub-models 306a-e, and that the content candidates 808 thus include five content candidates per segment. 802a-c (one created using each of the 306a-e sub-models). For example, referring to FIG. 9, a diagram is provided illustrating the mapping between document substructure 312a-f, sub-models 306a-e, and content candidates 808a-e. As described above, each of the sub-models 306a-e may be associated with one or more corresponding sub-structures 312a-f in structured text document 310. These relationships are indicated in FIG. 9 by mappings 902a-e between substructures 312a-ea and submodels 306a-e. Structured document generator 308 may store such mappings 902a-in the table or use other means.
[0067] When the speech recognition decoder 320 recognizes a segment S (e.g., segment 802a) from each of the sub-models 306a-e, a corresponding content candidate 808a-e is created. For example, content candidate 808a is text that is formed when speech recognition decoder 320 recognizes segment 802a from sub-model 306a, content candidate 808b is text that is formed when speech recognition decoder 320 recognizes segment 802a from sub-model 306b. , and so on. Structured document generator 308 may record mapping between content candidates 808a-e and corresponding sub-models 306a-in the model-content candidate mapping set 816. [0068] Thus, when the structured document generator 308 selects one of the content candidates 808ae as final content 812 for segment S (step 706), the final mapping identifier 818 may use mapping 816 and the selected final content 812 to recognize the language submodel that created the candidate. content that has been selected as the final content 812 (step 708). For example, if content candidate 808c is selected as final content 812, as seen in FIG. 9, the final mapping identifier 818 may identify the sub-model 306c as the sub-model that made the content candidate 808c. Final mapping identifier 818 may collect each identified sub-model in mapping set 820 such that, at any time, mappings 820 identify the order of language sub-models that were used to generate the final contents that were selected for inclusion in structured text document 310.
Once a sub-model corresponding to the final content 812 has been identified, the structured document generator 308 can identify the document sub-structure associated with the identified sub-model (step 710). For example, if the sub-model 306c was identified in step 708 as seen in FIG. 9, the sub-structure document 312c is associated with the sub-model 306c.
[0070] Structured content inserting device 822 inserts final content 812 into the identified substructure of structured text document 310 (step 712). For example, if substructure 312c is identified in step 710, text insertion machine 514 inserts final content 812 into substructure 312c.
Structured document generator repeats steps 704712 for the remaining segments 802b-c of the spoken audio stream 302 (step 714), thereby generating final content 812 for each of the remaining segments 802b-c and inserting final content 812 into the corresponding document substructures 312a-f. text 310. Upon completion of method 700, structured text document 310 includes text corresponding to spoken audio stream 302, and final model content mappings 820 identifying the orders of language sub-models that have been used by speech recognition decoder 320 to generate text in structured text document 310.
[0072] It should be noted that in the process of recognizing spoken audio stream 302, method 700 may not only generate text corresponding to spoken audio, but may also identify semantic information represented by audio and record such semantic information in device-readable form. For example, again referring to FIG. 5, comparison section 312b includes a date element wherein the specified date is represented as a triplet containing individual values for day ("06), month (" APR), and year ("2001). Other examples of semantic concepts in the medical field include vital signs, medications and their dosages, allergies, medical codes, etc. The retrieval and display of semantic information enables the automatic processing of such information to be performed. It should be noted that the specific form in which the semantic information is represented in FIG. 5 is only an example and does not limit the invention. [0073] Referring again to step 701 in which the method 700 shown in FIG. 7A identifies the set of segments 802a-c before identifying sub-models to be used to recognize segments 802a-c. It should be noted, however, that the structured document generator 308 may integrate the 802a-c segment identification process with the sub-model identification process to use them to recognize 802a-c segments and with a process for performing speech recognition on the 802a-c segments. Examples of techniques that may be used to perform such integrated segmentation and recognition will be described in more detail below with reference to FIG. 11B and 12B.
[0074] Having generally described the operation of the method shown in FIG. 7, the use of the method of FIG. 7 in the example of audio stream 302 shown in FIG. 4. Assuming the first part of the spoken audio stream 302 is the spoken stream of speech: "Chest CT without contrast April 22nd 2003. This portion may be selected in step 702 and recognized using all language sub-models 306a-e in step 704 to create multiple content candidates 808a-e. As described above, we assume that the sub-model 306a is the "header language model, and the sub-model 306b is the" comparison language model, and the sub-model 306c is the "technique" language model, the sub-model 306d is the "discovery, and" language model. the 306e sub-model is the inference.
[0075] Since sub-model 306a is a language model that has been taught speech recognition in the document header section 310 (e.g., substructure 312a), it is likely that a content candidate 808a created using sub-model 306a will match the words in the above-indicated audio portion more closely. exactly than other 808b-e content candidates. Assuming that content candidate 808a is selected as final content 812 for this audio portion, content inserting device 822 will insert the final content 812 formed by sub-model 306a into header section 312a of structured text document 310.
[0076] We assume that the second part of the spoken audio stream is the spoken speech stream: "compared to the earlier survey of March 26, 2002 and April 6, 2001. This part can be selected in step 702 and recognized using all sub-models of the language 306a-ew. step 704 to create multiple content candidates 808a-e. Since sub-model 306b is a language model that was taught to recognize speech in the document section "comparison 310 (e.g., sub-structure 312b), it is likely that a content candidate 808b created using sub-model 306b will match the words in the above-indicated audio portion. more specifically than the other content candidates 808a and 808c-e. Assuming that content candidate 808b is selected as final content 812 for this audio portion, text insertion device 514 will insert the final content 812 formed by sub-model 306b into the comparison section 312b of structured text document 310.
[0077] The remainder of the audio stream 302 shown in FIG. 4 may similarly be recognized and located in the corresponding substructures 312a-f in structured text document 310. It should be noted that although the content in spoken audio stream 302 shown in FIG. 4 occurs in the same order as sections 312a-f in structured textual document 310, it is not required according to the invention. On the contrary, the content may be present in the audio stream 302 in a different order. Each of segments 802a-c of the audio stream 302 is recognized by the speech recognition decoder 320, and the resulting final content 812 is placed in the corresponding of substructures 312a-f. As a result, the order of the textual content in substructures 312a-f may not be the same as the order of the content in the spoken audio stream. It should be noted, however, that even if the order of the textual content is the same in both the audio streams 302 and the structured text document 310, the interpreting engine 314 (FIG. 3) can orient the textual content of the document 310 in any desired order.
[0078] In another embodiment of the invention, the probabilistic language model 304 is a hierarchical language model. Specifically, in this embodiment, multiple sub-models 306a-e are organized in a hierarchy. As described above, the sub-models 306a-e may further include additional sub-models, and so on such that the hierarchy of the language model 304 may span multiple levels.
[0079] Referring to FIG. 10A, a diagram is shown illustrating an example of a language model 304 in a hierarchical form. The language model 304 includes multiple nodes 1002, 306a-e, 1006a-e, and 1010 and 1012. The square nodes 1002, 306b-e, and 1006e and 1012 use a probabilistic finite state grammar for a model highly constrained by concepts (such as report section order, Hints, Dates, and Times section). The elliptical nodes 306a, 1006a-d, and 1010 use statistical (n-gram) language models to model the less constrained language.
[0080] As used herein, the term "concept includes, for example, sections for date, times, numbers, codes, medications, medical history, diagnoses, prescriptions, expressions, lists, and directions. The concept can be said in many ways. Any way of saying a certain concept is referred to as "spoken form of the concept." There is sometimes a distinction between "semantic concepts and" syntactic concepts. As used herein, the term "concept includes, but is not limited to, and is not limited to a specific definition of" semantic concept or "syntactic concept or to distinguish between them.
[0081] Considering, for example, October 1, 1993, which is an example of a concept in which the term is used. Spoken forms of this concept include spoken expressions, "October 1st 1993," October 1st
1993, and "ten dash one dash 93. Texts such as" October 1, 1993 and "10/01/1993 are examples of" written forms of this concept.
[0082] Referring to the sentence "John Jones has pneumonia. This sentence, which is the concept in which the term is used, may be spoken in a number of ways, such as spoken phrases, "John Jones has pneumonia," Jones patient is diagnosed with pneumonia, and "Patient Jones is diagnosed with pneumonia. The written sentence "John Jones has pneumonia is an example of" a written form of the same concept.
[0083] Although language models for low level concepts such as dates and times are not shown in FIG. 10A (except for sub-model 1012), the hierarchical language model 304 may include sub-models for such low-level concepts. For example, the n-gram sub-models 306a, 1006a-d, and 1010 can assign probabilities to a word order representing dates, times, and other low-level concepts.
[0084] The language model 304 includes a node base 1002 that includes a finite state grammar representing the probabilities of occurrence of the sub-nodes 306a-e of the node 1002. The node base 1002 may, for example, indicate the probabilities of the header sections, comparisons, tecnas, discoveries, and conclusions of the document. 310 occurring in a specific order in the spoken audio stream 302.
Crossing down one level in the first language model 304, node 306a is a "header node, which is an N gram model of the language representing the probabilities of a word in a part of the spoken audio stream 302 to be included in the header section 312a of the structured document." text 310.
[0086] Node 306b includes a "comparison finite state grammar" representing the probabilities of multiple spoken alternatives hint for comparison section 312b of a text document. A finite state grammar with comparison node 306b may, for example, include clues such as "compare to," compare for, "previously is, and" prior research are. " A finite state grammar can include the probabilities of each of these clues. Such probabilities may, for example, be based on observed cue usage frequencies in a speech practice set for the same speaker or in the same domain as the spoken audio stream 302. Such frequencies may be obtained, for example, using the techniques disclosed in the above referenced patent application. entitled "Document Transcription System Training.
[0087] The comparison node 306b includes a "comparison content sub-node 1006a, which is an n-gram model of a language representing the probabilities of a word in a part of the spoken audio stream 302 to be included in the comparison section 312b of the text document 310. The comparison content node 1006a has date node 1012 as a child. As will be described in more detail below, date node 1012 is a finite state grammar representing the probabilities of a date speaking in various ways.
[0088] Nodes 306c and 306d may be understood similarly. Node 306c includes a finite state grammar "tecńniki" representing the probabilities of multiple alternative spoken forms of clues for technique section 312c of text document 310. Engineering node 306c includes a "content" sub-node 1006b, which is a ngram language model representing the probabilities of a word in a part of spoken audio stream 302 to be included in technique section 312c of text document 310. Likewise, node 306d includes a discovery finite state grammar. representing the probabilities of multiple oral alternate forms of clues for discovery section 312d of text document 310. Discovery node 306d includes a "discovery content sub-node 1006c, which is an n-gram language model representing the probabilities of a word occurring in a part of spoken audio stream 302 to be included in discovery section 312d of text document 310.
Inference node 306e is similar to nodes 306bd in that it includes a finite state grammar for section hand recognition and a sub-node 1006d that includes an ngram model of the language for section recognition. Additionally, however, the inference node 306e includes an additional sub-node 1006e, which in turn includes the sub-node 1010. This indicates that the content of the conclusions section can be recognized using either the language model in the conclusion content node 1006d or the node "enum 1006e" governed by the finite state grammar based language model with the corresponding conclusion node 306e. Node "enum 1006e includes a finite state grammar indicating the probabilities associated with different ways of speaking the numbering indicia (such as" number one, "number two," first, "second," third, and so on).
The conclusion content node 1010 may include the same language model as the conclusion content node 1006d.
[0090] Having described the hierarchical structure of the tongue model 304 in one embodiment of the invention, examples of techniques that may be used to generate a structured document 310 using the tongue model 304 will now be described. Referring to FIG. 11A, there is shown a flowchart of a method that is performed by a structured document generator 308 in one embodiment of the present invention to generate a structured text document 310 (FIG. 2, step 204). With reference to FIG . 12A, a data flow diagram is shown illustrating a portion of system 300 in detail relevant to the method of FIG. 11A.
[0091] The structured document generator 308 includes a path selector 1202 that identifies the paths 1204 through the hierarchical language model 304 (step 1102). Path 1204 is sequentially ordered with nodes in a hierarchical language model 304. Nodes may be moved multiple times in path 1204. Examples of path generation techniques 1204 will be described in more detail below with reference to FIG. 11B and 12B.
[0092] Referring to FIG. 10B, an example of path 1204 is shown. Path 1204 includes points 1020aj which define the order in which the nodes are moved in the language model 304. Points 1020a-j are referred to as "points, not" nodes, to distinguish them from nodes 1002, 306a. -e, 1006a-e, and 1010 in the 304 language model.
[0093] In the example shown in FIG. 10B, path 1204 moves the following language model 304 nodes in order:
(1) the base of the node 1002 (item 1020a); (2) header compactness node 306a (point 1020b); (3) comparison node 306b (point 1020c); (4) content comparison node 1006a (point 1020d); (5) technique node 306c (point 1020e); (6) engineering content node 1006b (item 1020f); (7) discovery node 306d (point 1020g); (8) discovery content node 1006c (point 1020h); (9) conclusions node 306e (point 10201); and (10) conclusion content node 1006d (point 1020j).
[0094] As shown with reference to FIG. 4, recognizing spoken audio stream 302 using language sub-models along the track 1204 shown in FIG. 10B results in optimal speech recognition because speech in the audio stream 302 occurs in the same order as the language sub-models on the path 1204 shown in FIG. 10B. For example, spoken audio stream 302 starts with the speech that is best recognized by the header content of the language model 306a ("Chest CT without contrast April 22nd 2003), then the speech that is best recognized by comparing the language model 306b (" comparison to ), then the speech that is best recognized by comparing the content of the language model 1006a ("earlier studies of March 26, 2002 and April 6, 2001), and so on.
[0095] After identifying the track 1204, the structured document generator 308 recognizes the spoken audio stream 302 using the language models shifted through the tracks 1204 to form the structured text document 310 (step 1104). As described in more detail below with reference to FIG. 11B and 12B, the speech recognition step and the generation of the structured text document 1104 may be integrated into the track identification step 1102, rather than performed alone.
[0096] More specifically, the structured document generator 308 may include a node numerator 1206 that iterates each node of the N model 1208 shifted through the selected paths 1204 (step 1106). For each such node N, speech recognition decoder 320 may recognize the portion of the audio stream 302 corresponding to the language model at node N to form the corresponding structured text T (step 1108). Structured document generator 308 may insert text T 1210 into a substructure of structured text document 310 corresponding to node N 1208 of the language model 304 (step 1110).
[0097] For example, when the node N is a comparison node 306b (FIG. 10A), the comparison node 306b may be used to recognize the "compare to" text in spoken audio stream 302 (FIG. 4). Since the comparison node 306b corresponds to the document substructure (e.g., comparison section 312b) more than the content, the result of the speech recognition performed in step 1108 in this case may be a document substructure, i.e., an empty "comparison section." Such a section may be inserted into the structured document 310 in step 1110, for example, in a format matching the labels "<compare>" and "</ compare>".
[0098] When node N is the content compare node 1006a (FIG. 10A), the content compare node 1006a may be used to recognize the text "earlier surveys of March twenty-six, 2002 and April 6, 2001 in spoken audio stream 302 (FIG. 4). thereby creating the structured text of "prior studies of <data> 26MAR-2002 </data> and <data> 06-APR-2001 </data>" as shown in FIG. 5. This structured text may then be inserted into the compare section 312b in step 1110 (e.g., between the labels "<compare>" and "</ compare>" as shown in FIG. 5).
[0099] Structured document generator 308 repeats steps 1108-1110 for the remaining N nodes moved by paths 1204 (step 1112), thereby inserting a plurality of structured text documents 1210 into structured text document 310. The end result of the method shown in FIG. 11A is to create a structured text document 310 that includes text having a structure that corresponds to the path structure'1204 in the language model 304. For example, as shown in FIG. 10B, the structure of the depicted path shifts the language model nodes corresponding to the header, comparison, technique, discovery, and conclusion sections, respectively. The resulting structured text document 310 (as shown, for example, in FIG. 5) similarly includes header, comparison, technique, discovery, and conclusion sections sequentially. The structured text document 310 thus has the same structure as the language model paths 1204 that was used to create the structured text document 310.
[0100] Above, it is shown that the structured document generator 308 places the recognized structured text 1210 into the corresponding substructures of the structured text document 310 (FIG. 11A, step 1110). As shown in FIG. 5, structured text document 310 may be implemented as an XML document or other document that supports placed structures. In this case, it is necessary to insert each of the recognized structural texts
1210 within the corresponding substructures so that the final structured text document 310 has a structure that corresponds to that of the path 1204. It is clear to those skilled in the art how to use the final mapping model content 820 (FIG. 8) to use path 1204 to shift the structure of the language model 304 and thus create such a structured document.
[0101] The system shown in FIG. 12A includes a track selector 1202 that selects tracks 1204 by a language model 304. The method shown in FIG. 11A then uses the selected tracks 1204 to generate a structured text document 310. In other words, in FIG. 11A and 12A, the selection of the path steps and the creation of the structured document are performed separately. However, this does not limit the invention.
[0102] In other words, referring to FIG. 11B, a flow chart of a method 1150 is shown that integrates the steps of selecting a path and generating a structured document. With reference to FIG. 12B, an embodiment of a document structure generator 308 that performs method 1150 in FIG. 11B in one embodiment of the present invention. Generally, method 1150 of FIG. 11B searches for a possible path through the hierarchy of the language model 304 (FIG. 10A) starting at the base of node 1002 and working outward. Any of a number of techniques, including those known to those of skill in the art, may be used to search through the language model hierarchy. Since method 1150 identifies partial paths through the language model hierarchy, method 1150 uses a speech recognition decoder 320 to recognize an ascendingly larger portion of a spoken audio stream 302 using language models along the partial paths, thereby creating partial potential structured documents. Method 1150 assigns suitability ratings to each of the partial potential structural documents. The suitability rating for each potential structured document is a measure of how well the paths that make up the potential structured document are executed. Method 1150 extends the partial paths, thereby continuing to search through the language model hierarchy until the entire spoken audio stream 302 has been recognized. Structured document generator 308 selects the potential structured document with the highest suitability rating as the final structured text document 310.
[0103] More specifically, method 1150 initializes one or more candidate paths 1224 through the language model 304 (step 1152). For example, candidate paths 1224 may be initialized to include single paths consisting of the base of node 1002. The term "frame refers to a short period of time, such as 10 milliseconds. Method 1150 initializes the pointer of the audio stream to point to the first frame in the audio stream 302 (step 1153). For example, in the embodiment shown in FIG. 12B, the structured document generator 308 includes an audio stream numbering device 1240 that provides a portion 1242 of the audio stream 302 to a speech recognition decoder 320. After starting method 1150, portion 1242 may itself include the first frame of the audio stream 302.
[0104] Speech recognition decoder 320 recognizes the present portion 1242 of the audio stream 302 using language sub-models in candidate paths 1224 to generate one or more candidate structural partial documents 1232 (step 1154). It should be noted that the documents 1232 are only partial documents 1232 as they were generated based on part of the audio stream 302. When step 1154 is performed first, speech recognition decoder 320 may simply recognize the first frame of the audio stream 302 using the language model at the base of node 1002 of language model 304.
[0105] It should be noted that the techniques disclosed above with reference to FIG. 11A and FIG. 12A may be used by the speech recognition decoder 320 to generate candidate structural partial documents 1232 using candidate paths 1224. More specifically, speech recognition decoder 320 may use the methods shown in FIG. 11A for a portion of the audio stream 1242 using each of the candidate paths 1224 as paths identified in step 1102 (FIG. 11A).
[0106] Referring to FIG. 11B and 12B, fit evaluator 1234 generates suitability ratings 1236 for each candidate structural partial documents 1232 (step 1156). Suitability ratings 1236 are measured based on how correctly the structural partial document candidates 1232 represent the corresponding portions of the audio stream 302. Generally, a suitability score for a single candidate document can be generated by: (1) generating suitability scores for each of the nodes in the respective candidate paths 1224; and (2) using a synthesis function to synthesize the individual node suitability scores generated in step (1) in an overall suitability evaluation for a potential structural document. Examples of techniques that may be used to generate the fitness score candidate 1236 will be described in more detail below with reference to FIG. 11C.
[0107] If the structured document generator 308 were to try to search for all possible paths through the language model hierarchy 304, the computational resources requiring the evaluation of each possible path could become too expensive and / or time consuming due to the exponential increase in the number of possible paths. Thus, in the embodiment shown in FIG. 12B, path trimmer 1230 uses candidate suitability estimates 1236 to remove mismatched paths from candidate path 1224, thereby creating a set of trimmed paths 1222 (step 1158).
[0108] If the entire audio stream 302 has been recognized (step 1160), the final document selector 1238 selects, from among the structural partial document candidates 1232, the candidate structured document having the highest suitability rating, and provides the selected document as the final structured text document 310 (step 1164) . If the entire audio stream 302 is not recognized, track extender 1220 extends the trimmed tracks 1222 in the language model 304 to create a new candidate track set 1224 (step 1162). For example, if the trimmed paths 1222 consist of a single path including the base of node 1002, path extender 1220 may extend that path through one node down the hierarchy shown in FIG. 10A to form multiple candidate paths extending from the base of node 1002, such as the base paths of node 1002 to the contents of the node header 306a, the base path paths
1002 for comparing node 306b, path with the base of node 1002 to technical node 306c, and so on. Various techniques for extending path 1224 to perform depth-first, width-first, or other types of hierarchical searches are well known to those skilled in the art.
[0109] Audio stream numbering 1240 extends portion 1242 of audio stream 302 to include the next frame in audio stream 302 (step 1163). Steps 1154-1160 are then repeated using the new candidate paths 1224 to recognize a portion 1242 of the audio stream 302. Thereby, the entire audio stream 302 can be recognized using the corresponding sub-models in the language model 304.
[0110] As described above with reference to FIG. 11B and 12B, suitability scores 1236 may be generated for each partial document candidate 1232 created by the structured document generator 308 while evaluating candidate 1224 paths by the language model 304. Examples of techniques will be described below for generating suitability scores, or for partial structural candidate partial document 1232 shown in FIG. 12B or more generally for structured documents. [0111] For example, referring to FIG. 10A, note that the compared contents of node 1006a have date node 1012 as a child. Assuming the text "Chest CT scan without contrast April 22nd 2003 was recognized as the text corresponding to the comparison of the contents of node 1006a. Note that the content comparison of node 1006a was used to recognize the text "CT Chest without contrast and the date of node 1012, which is a child of the content comparison of node 1006a, was used to generate the text" April twenty-second 2003. The suitability score of this text can, therefore, be computed by using node content comparison 1006a to compute a first suitability score for the text "Chest CT without contrast with the following arbitrary date, computing a second suitability score for April 22, 2003 based on the date node. 1012, and multiplying the first and second suitability ratings. [0112] Referring to FIG. 11C, a method diagram is shown which is performed in one embodiment of the present invention to calculate a suitability score for a candidate document, and which can therefore be used to implement step 1156 of method 1150 shown in FIG. 11B. The suitability of S is initialized to the value of one of the estimated potential structural documents (step 1172). The method assigns the current node index N to point to a node base in the candidate path corresponding to the candidate document (step 1174).
[0113] The method calls a function called Match () with the values of N and S (step 1176) and returns the result as a suitability score for the candidate document (step 1178). As will be described in more detail below, the Match () function generates a suitability score of S using hierarchical factorization by the moving candidate paths corresponding to the candidate document.
[0114] Referring to FIG. 11D, a diagram of an Alignment function () 1180 according to one embodiment of the present invention is shown. Function 1180 identifies the probabilities P (W (N)) that the text W corresponding to the current node N has been recognized by the language model associated with that node, and multiplies the probabilities by the current value of S to generate a new value for S (step 1184).
[0115] If Node N has no children (step 1186), the value of S is returned (step 1194). If node N has children, then Match () 1180 is called multiple times on each of the child nodes, and the result is multiplied by the value of S to form new S values (steps 1188-1192). The obtained S value is returned (step 1194).
[0116] After the method shown in FIG. 11C, the S value represents the suitability rating for the entire potential structural document, and the S value is returned, e.g., for use in method 1150 shown in FIG. 11B (step 1194).
[0117] For example, recall the text "Chest CT without contrast April 22, 2003. A suitability rating (probability) of such a text can be obtained by identifying the probability of the text" Chest CT without contrast <DATA> ", where <DATE> means any date multiplied by the conditional probability of the text "April 22, 2003 occurring as long as the text represents a date.
[0118] More generally, the result of the method shown in FIG. 11C is a hierarchical word order probability derivation according to the language model 304, enabling individual probability estimates associated with each language model node to be seamlessly combined with probability estimates associated with other nodes. This probability structure enables the system to model and use statistical language models with built-in probabilistic finite-state grammar and finite-state grammar with built-in statistical language models.
[0119] As described above, the nodes in the language model 304 represent language sub-models that determine the probabilities of a word order occurring in the spoken audio stream 302. In the description above, it is assumed that probabilities have already been assigned in the language modeling methods. Tecńnik examples will now be disclosed for assigning probability to language sub-models (such as n-gram language models and context-free grammar) in the language model 304.
[0120] Referring to FIG. 13, a method diagram 1300 is shown that is used in one embodiment of the invention to generate a language model 304. A plurality of nodes are selected for use in the language model (step 1302). Nodes may, for example, be selected by a transcriptionist or other person skilled in the art. The nodes may be selected in an attempt to capture all kinds of concepts that may exist in spoken audio stream 302. For example, in the medical field, nodes (such as those shown in FIG. 10A) may be selected to represent sections of a medical report and concepts (such as dates, times, medications, allergies, symptoms, and medical codes) that are expected to occur. that they will appear in the medical report.
[0121] A concept model and language model type may be assigned to each of the nodes selected in step 1302 (steps 1304-1306). For example, node 306b (FIG. 10A) may be assigned the concept of "section pointers comparison and assigned to a type of language model" finite state grammar. Similarly, node 1006a may be assigned the concept of "content comparison and language model type" n-gram language model.
[0122] The nodes selected in step 1302 may be organized into a hierarchical structure (step 1308). For example, nodes 1002, 306a-e, 1006a-e, and 1010 may be organized into the hierarchical structure shown in FIG. 10A to identify and enforce structural relationships between nodes.
[0123] Each of the nodes selected in step 1302 can then be trained using text corresponding to a specific concept (step 1310). For example, a set of training documents may be identified. The training document set may, for example, be a set of existing medical reports or other documents in the same field as spoken audio stream 302. Training documents can be manually tagged to indicate the existence and placement of structure within the document, such as sections, sub-sections, dates, times, codes, and other concepts. Such marking may, for example, be performed automatically on formatted documents, or manually by a transcriptionist or other person skilled in the art. Examples of node learning techniques selected in step 1302 are described in the above-referenced patent application entitled "Document Transcryption System Training.
[0124] Conventional language model learning techniques may be used in step 1310 to teach language models with a specific concept for each concept that is tagged in the training documents. For example, text with all header sections marked in the training documents may be used to train the language model node 306a constituting the header section. Thereby, the language models for each of the nodes 1002, 306a-e, 1006a-e, and 1010 in the tongue model 304 shown in FIG. 10A can be learned. The result of the method 1300 shown in FIG. 13 is a hierarchical language model having learned capabilities that can be used to generate a structured text document 310 as described above. This hierarchical language model can then be used, for example, to segment a training text multiple times, for example using the techniques disclosed above in connection with FIG. 11B and 12B. The segmented training text can be used to maintain the hierarchical model of the language. The segmentation and retraining process can be repeated many times to improve the quality of the language model many times over.
[0125] In the above-described examples, the structured document generator 308 both recognizes the spoken audio stream 302 and generates the structured text document 310 using an integrated process to generate intermediate non-structured transcription. Such tans, however, are only disclosed by way of example and do not limit the invention.
[0126] Referring to FIG. 14, a schematic of a method 1400 is shown which is used in a further embodiment of the invention to generate a structured text document 310 using separate speech recognition and structural analysis steps. With reference to FIG. 15, a data flow diagram in system 1500 that performs method 1400 in FIG. 14 according to one embodiment of the invention.
[0127] The speech recognition decoder 320 recognizes the spoken audio stream 302 using the language model 1506 to transcribe 1502 the spoken audio stream 302.
It should be noted that the tongue model 1506 may be a conventional language model that is different from the tongue model 304. More specifically, the tongue model 1506 may be a conventional monolithic language model. The language model 1506 may, for example, be generated using the same training corpus as is used to train the language model 304. While parts of the training body may be used to train the nodes of the language model 304, the entire body may be used to teach the language model 1506. The speech recognition decoder 320 may, therefore, use conventional speech recognition techniques to recognize spoken audio stream 302 using the model. 1506 and thereby to create the 1502 transcription.
[0128] It should be noted that transcription 1502 may be "uniform transcription 1502 of spoken audio stream 302, and not as in previous examples of structured document. Transcription 1502 may, for example, include the order of a unitary text resembling the text shown in FIG. 4 (which represents the spoken audio stream 302 as text).
[0129] System 1500 also includes a structured parser 1504 that uses a hierarchical language model 304 to analyze transcription 1502 and thereby to create a structured text document 310 (step 1404).
Structural parser 1504 can use the techniques disclosed above with respect to FIG. 11C and 12B to: (1) create multiple potential structured documents having the same content as transcription 1502, but having structures corresponding to different paths through the language model 304; (2) generate suitability assessments for each of the potential structural documents; and (3) select the potential structured text document with the highest suitability rating as the final structured text document. In contrast to the technique disclosed above with reference to FIG. 11C and 12B, however, step 1404 may be performed without performing speech recognition to generate each of the potential structural documents. On the contrary, once transcription 1502 is created using a speech recognition decoder 320, potential structured documents can be generated based on transcription 1502 without performing additional speech recognition.
[0130] Furthermore, the structured parser 1504 does not need to use the full language model 304 to create a structured text document 310. On the contrary, structured parser 1504 may use a scaled "language skeleton model, such as the language model 1030 shown in FIG. 10C. Note that the exemplary language model 1030 shown in FIG. 10C is the same as the tongue model 304 shown in FIG. 10A, except that in the language skeleton model 1030 the contents of the language model nodes 306a, 1006a-d, and 1010 have been replaced with universally accepted language models 1032a-f, also referred to as don't care language models. The 1032a-f language models accept any text that is supplied to them as input. The guideline path of the 306b-e language models of the 1030 framework language model enables the structured parser 1504 to analyze transcription 1502 as valid substructures in structured document 310. The use of universally accepted language models 1032a-f, however, enables structured parser 1504 to perform such structured analysis without incurring (which is usually of importance) for learning the content of language models such as models 306a, 1006a-d, and 1010 shown in FIG. 10A.
[0131] It should be noted that the language framework 1030 may still include language models, such as the date language model 1012, conforming to lower-level concepts. As a result, language framework 1030 can be used to generate structured document 310 transcribed 1502 without invoking higher content of language models, while the ability to parse concepts of structured document 310 is maintained.
[0132] Among the advantages of the invention are one or more of the following. The techniques disclosed herein replace the traditional global language model with a combination of specialized local language models that are more fit within a section of the document than a single generic language model. There are many advantages to this language model.
[0133] For example, the use of a language model that includes sub-models each corresponding to a specific concept is advantageous as it allows the most appropriate language model to be used for the speech recognition corresponding to each concept. In other words, if each of the sub-models corresponds to different concepts, then each of the sub-models can be used to perform speech recognition in speech representing the respective concept. Because speech properties can vary from concept to concept, using such a concept language model may result in better recognition than that which would be created using a monolithic language model for all concepts.
[0134] Although sub-models of the language model may correspond to sections of the document, this is not a limitation of the invention. Conversely, each sub-model in the language model can correspond to any concept such as section, paragraph, sentence, date, time, or an ICD9 code. As a result, sub-models in the language model can be fitted to specific concepts with greater precision than would be possible if only sections of the specific language models were used. The use of such concept language models for different concepts can further improve the accuracy of speech recognition.
[0135] Moreover, hierarchical language models designed according to embodiments in accordance with the invention may have multi-level hierarchical structures with built-in sub-models within themselves. As a result of; sub-models in the language model may be applied to the part of the spoken audio stream 302 at different levels, with the most appropriate language model applied at each level. For example, the "header section language model" may be applied generally to speech within the header section of a document, while the "date" language model may be applied specifically to speech relating to dates in the header section. The ability to place language models and apply the placed language models to different parts of speech can further improve recognition accuracy by selecting the most appropriate language model to be applied to each part of the spoken audio stream.
[0136] Another benefit of using a language model that includes multiple sub-models is that the techniques disclosed herein may use such a language model to generate a structured text document from a spoken audio stream using a single integrated process, rather than as in the prior art dual the step-by-step process 100 shown in FIG. 1A, wherein the speech recognition step is followed by the natural language processing step. In the two-step process 100 shown in FIG. 1A, the steps performed by the speech recognition device 104 and the natural language processor 108 are completely disconnected. Since the automatic speech recognition device 104 and the natural language processor 108 operate independently of each other, the output 106 of the automatic speech recognition device 104 is a verbatim transcription of the spoken content in the audio stream 102. Verbatim transcription 106 thus includes text corresponding to all spoken statements in the audio stream 102, whether or not such statements are relevant to the final desired structured text document. Such utterances may include, for example, hesitations, irrelevant words or repetitions, as well as structured prompts or task-related words. In addition, the natural language processor 108 relies on correctly detecting and transcribing a specific keyword and / or key phrases, such as structured prompts. If these keywords / phrases are incorrectly recognized by the automatic speech recognition device 104, the identification of the structural entities by the natural language processor 108 may be incorrect. Conversely, the method 200 shown in FIG. 2, speech recognition and natural language processing are integrated, thereby allowing the language model to affect the recognition of the word in the audio stream 302 and the generation of structure in the structured text document 310, thereby improving the quality of the structured document 310.
[0137] In addition to generating the structured document 310, the techniques disclosed herein may also be used to retrieve and interpret semantic content from the audio stream 302. For example, the date language model 1012 (FIGS. 10A-10B) may be used to identify portions of the stream. audio 302 that represent a date and to record representations of such dates in a computer-readable form. For example, the techniques disclosed herein can be used to recognize a spoken phrase "October 1, 1998 as a date, and to record the date in computer-readable form such as" month = 10, day = 1, year = 1998. Saving such concepts in a computer-readable form enables the content of such concepts to be easily processed by a computer, such as by dated document sections or by identifying stored medications after a given date. In addition, the techniques disclosed herein enable the user to define different parts (e.g., sections) of the document, and select which concepts to download in each section. The techniques disclosed herein, therefore, enable the recognition and processing of semantic content in spoken audio streams. Such techniques may be used in place of or in addition to storing the retrieved information in a structured document.
[0138] Fields such as the medical and legal fields, in which there are a large number of already recorded audio streams that can be used as exercise text, may have certain benefits from the tactics disclosed herein. Such practice text may be used to teach a language model 304 using the techniques disclosed with reference to FIG. 13. Since documents in such domains may require well-defined structures, and since such structures can be easily identified in existing documents, it can be easy (albeit time-consuming) to correctly identify parts of such existing documents so as to use this in learning each language model node with a specific concept in the 304 language model. As a result, each node of the language model can be appropriately trained to recognize relevant concepts, thereby increasing recognition accuracy and increasing the ability of the system to generate documents having the required structure.
[0139] Furthermore, the techniques disclosed herein can be applied in such fields and do not require any alteration to the existing process in which the audio is recorded and transcribed. In the medical field, for example, doctors can continue to dictate medical reports as they have done so far. The techniques disclosed herein can be used to generate documents having a desired structure no matter how the spoken audio stream is dictated. Alternative techniques that require changes in the work, i.e. techniques that require speakers to be written down (by reading the practice text), requiring the speakers to modify their way of speaking (e.g. always saying certain concepts using specific spoken forms), or that require the generation of transcripts in specific format can be very expensive to implement in areas such as medical and legal. Such changes may, in fact, be inconsistent with institutional or legal requirements related to the reporting structure (for example, imposed by reporting requirements for an insurer). The techniques disclosed herein, on the contrary, are capable of generating the audio stream 302 in any way and in any form.
[0140] Additionally, individual sub-models 306a-e of the language model 304 can be easily updated without affecting the rest of the language model. For example, the header content sub-model 306a may be replaced with another header content sub-model that corresponds differently to how the document header is dictated. The modular structure of the language model 304 'allows the sub-models to be modified / replaced so that they can be made without the need to modify the rest of the language model 304. As a result, portions of the language model 304 can be easily updated to reflect different document dictation conventions.
[0141] Moreover, the structured text document 310 that is formed by the various embodiments of the invention may be used to train a language model. For example, teaching the techniques described in the above-referenced patent application entitled "Document Transcription
The Training system can use the structured text document 310 to re-learn and thus improve the language model 304. The retrained language model 304 can then be used to create another structured text document, which in turn can be used to re-learn the language model 304. This retrained language model 304. a multiple process can be used to improve the quality of structured documents that are created over time.
[0142] It is clear that although the invention has been described above with reference to specific embodiments, the above embodiments are exemplary only, and do not limit or define the scope of the invention. Other embodiments, including but not limited to the following, are also within the scope of the claims. For example, the elements and components described herein may be further divided into additional components or combined together to obtain fewer components fulfilling the same functions.
[0143] Spoken audio stream 302 may be any audio stream such as live audio stream received directly or indirectly (for example over a telephone or IP connection), or an audio stream recorded in any medium and in any format. In distributed speech recognition, DSR (DSR) distributed speach recognition), the client performs processing to the audio stream to create a processed audio stream that is broadcast to a server that performs speech recognition on the processed audio stream. The audio stream 302 may, for example, be a processed audio stream produced by a DSR client.
[0144] Although in the above examples each node in the language model 304 is described as containing a language model that conforms to a specific concept, this is not a requirement of the invention. For example, a node may include a language model that results from the interpolation of a specific concept language model associated with a node with one or more: (1) global background language models, or (2) language models with a specific concept associated with other nodes.
[0145] In the examples above, a distinction can be made between "grammar and" text. It is clear that a text can be represented as a grammar in which there is a single spoken form having a probability of one. Thus, documents that are described herein as including both text and grammar may be implemented using only grammar if desired. Moreover, finite-state grammar is just one type of context-free grammar, which is a kind of linguistic model that enables the representation of many alternative spoken forms of a concept. Thus, any description of the techniques that are used in a finite state grammar may be applied more generally to any type of grammar. Moreover, while the above description may refer to finite state grammar and ngram language models, these are only examples of the kinds of language models that may be used in conjunction with embodiments of the invention. The embodiments of the invention are not limited to use in connection with any particular type of language model.
[0146] The invention is not limited to any of the fields described (such as medical and legal reporting), but is generally applicable to any structured documents.
[0147] The techniques described above may be implemented in, for example, hardware, software, firmware, or any combination thereof. The techniques described above may be implemented in one or more computer programs executed on a programmable computer including a processor, a processor readable storage medium (including, e.g., volatile and non-volatile memory and / or memory elements), at least one input device, and at least one input device. one output device. Program code may be applied to data input using an input device to perform the described functions and to generate output data. The output may be delivered to one or more output devices.
[0148] Any computer program as defined in the following claims may be implemented in any programming language, such as assembly language, machine language, high-level language, or object-oriented programming language. The programming language may, for example, be a compiled or interpreted programming language.
[0149] Any such computer program may be implemented in a computer program product physically resided in a computer readable memory device to be executed by a computer processor. The steps of the inventive method may be performed by a computer processor executing a program physically resident on a computer readable medium to perform the functions of the invention by handling the input data and generating the output data. Suitable processors include, for example, both general and special purpose microprocessors. Generally, the processor receives instructions and data from read-only memory and / or RAM. Memory devices suitable for physically containing computer program instructions thereon include, for example, all forms of non-volatile memory, such as semiconductor memory devices including EPROM, EEPROM, and flash memory devices; magnetic disks such as internal hard disks and removable disks; magnetic optical discs; and CD-ROM. Each of the above-mentioned can be replaced with or incorporated into specially designed electronic ASICs or directly programmable FPGA gate arrays. The computer may also generally receive programs and data from a storage medium such as an internal drive (not shown) or a removable drive. These elements are also available from conventional PCs or workstations as well as other computers suitable for executing computer programs implementing the described methods that may be used in conjunction with any digital printing device or marking device, monitor, or other raster output device capable of producing color. or grayscale pixels on paper, film, monitor screen, or other output medium.
22559 / ΡΕ / 12
ΕΡ1787288
Contents4
33 sheets
Sheet 1 Sheet 2 Sheet 3 Sheet 4 Sheet 5 Sheet 6 Sheet 7 Sheet 8 Sheet 9 Sheet 10 Sheet 11 Sheet 12 Sheet 13 Sheet 14 Sheet 15 Sheet 16 Sheet 17 Sheet 18 Sheet 19 Sheet 20 Sheet 21 Sheet 22 Sheet 23 Sheet 24 Sheet 25 Sheet 26 Sheet 27 Sheet 28 Sheet 29 Sheet 30 Sheet 31 Sheet 32 Sheet 33
59 members in 9 offices
Priority claims6
| Document | Office | Kind | Date |
|---|---|---|---|
| 92351704 | United States of America | A | |
| 05789851 | European Patent Office (EPO) | A | |
| 2005029354 | United States of America | W | |
| EP20050789851 | – | – | – |
| US20040923517 | – | – | – |
| WO2005US29354 | – | – | – |
Members59
| Document | Office | Kind | |
|---|---|---|---|
| US2006041427A1 | United States of America | A1 | |
| US2006041428A1 | United States of America | A1 | |
| CA2577721A1 | Canada | A1 | |
| CA2577726A1 | Canada | A1 | |
| WO2006023622A2 | World Intellectual Property Organization (WIPO) | A2 | |
| WO2006023631A2 | World Intellectual Property Organization (WIPO) | A2 | |
| WO2006034152A2 | World Intellectual Property Organization (WIPO) | A2 | |
| US2006074656A1 | United States of America | A1 | |
| US2007033032A1 | United States of America | A1 | |
| WO2006023631A3 | World Intellectual Property Organization (WIPO) | A3 | |
| WO2007018842A2 | World Intellectual Property Organization (WIPO) | A2 | |
| WO2006034152A3 | World Intellectual Property Organization (WIPO) | A3 | |
| WO2006023622A3 | World Intellectual Property Organization (WIPO) | A3 | |
| EP1787287A2 | European Patent Office (EPO) | A2 | |
| EP1787288A2 | European Patent Office (EPO) | A2 | |
| WO2007018842A3 | World Intellectual Property Organization (WIPO) | A3 | |
| EP1908055A2 | European Patent Office (EPO) | A2 | |
| JP2008511024A | Japan | A | |
| EP1787288A4 | European Patent Office (EPO) | A4 | |
| EP1908055A4 | European Patent Office (EPO) | A4 | |
| JP2009503560A | Japan | A | |
| US2009048833A1 | United States of America | A1 | |
| EP1787287A4 | European Patent Office (EPO) | A4 | |
| US7584103B2 | United States of America | B2 | |
| EP1908055B1 | European Patent Office (EPO) | B1 | |
| AT454691T | Austria | T | |
| ATE454691T1 | Austria | T1 | |
| DE602006011622D1 | Germany | D1 | |
| US2010299135A1 | United States of America | A1 | |
| US7844464B2 | United States of America | B2 | |
| US2010318347A1 | United States of America | A1 | |
| JP4940139B2 | Japan | B2 | |
| EP1787288B1 | European Patent Office (EPO) | B1 | |
| US8335688B2 | United States of America | B2 | |
| PL1787288T3This record | Poland | T3 | |
| ES2394726T3 | Spain | T3 | |
| US8412521B2 | United States of America | B2 | |
| US2013103400A1 | United States of America | A1 | |
| US2013166297A1 | United States of America | A1 | |
| JP5284785B2 | Japan | B2 | |
| US2013304453A9 | United States of America | A9 | |
| US8694312B2 | United States of America | B2 | |
| US8731920B2 | United States of America | B2 | |
| US8768706B2 | United States of America | B2 | |
| US2014249818A1 | United States of America | A1 | |
| US2014309995A1 | United States of America | A1 | |
| US2014343939A1 | United States of America | A1 | |
| CA2577721C | Canada | C | |
| CA2577726C | Canada | C | |
| US9135917B2 | United States of America | B2 | |
| US9190050B2 | United States of America | B2 | |
| US2016005402A1 | United States of America | A1 | |
| US9286896B2 | United States of America | B2 | |
| US2016078861A1 | United States of America | A1 | |
| US2016196821A1 | United States of America | A1 | |
| US9454965B2 | United States of America | B2 | |
| EP1787287B1 | European Patent Office (EPO) | B1 | |
| US9520124B2 | United States of America | B2 | |
| US9552809B2 | United States of America | B2 |
Numbers
- Publication, DOCDB
- 1787288
- Publication, EPODOC
- PL1787288T
- Application
- 789851
- Application, DOCDB
- 05789851
- Application, EPODOC
- PL20050789851T
Titles2
- English
- AUTOMATED EXTRACTION OF SEMANTIC CONTENT AND GENERATION OF A STRUCTURED DOCUMENT FROM SPEECH
- Polish
- Automatyczne wyodrębnianie semantycznej zawartości oraz generowanie strukturalnego dokumentu z mowy
Classification
- CPC, 2
- G10L15/1815
- G16H15/00
- IPC, 2
- G10L15 18
- G06F19 00
