Automated extraction of semantic content and generation of a structured document from speech
Abstract
A procedure comprising the steps: (A) identify a probabilistic linguistic model that includes a plurality of probabilistic linguistic models associated with a plurality of substructures of a document; and (B) use a speech recognition decoder to apply the probabilistic linguistic model 5 to a spoken audio stream to produce a document that includes organized content in the plurality of substructures, in which the content in each of the plurality of substructures is produced by recognizing speech using the substructure, in which the plurality of probabilistic linguistic models are organized in a hierarchy, and in which stage (B) comprises the stages of: (B) (1) identifying a path through the hierarchy, comprising the stages of: (B) (1) (a) identifying a plurality of path a through hierarchy (B) (1) (b) for each of the plurality of paths P, produce a structured document candidate for the spoken audio stream using the speech recognition decoder to recognize the spoken audio stream using the linguistic models in the P path; B (1) © apply a measurement to the plurality of candidate structured documents produced in the stage (B) (1) (b) to produce a plurality of relevance scores for the plurality of structured candidate documents; and (B) (1) (d) select the trajectory that produces the candidate structured documents that have the highest relevant score; (B) (2) generate the document that has a structure that corresponds to the trajectory identified in step (B) (1).

Term
Term ended
Projected expiry passed 18 August 2025, 1.1 years ago.
- Priority
- Filed
- Published
- Projected expiry
- Today
6 claims: 1 independent, 5 dependent
- 1ES 2 394 726 T3 REIVINDICACIONES 1. - Un procedimiento que comprende las etapas:(A) identificar un modelo lingüístico probabilista que incluye una pluralidad de modelos lingüísticos probabilistas asociada a una pluralidad de subestructuras de un documentos;y (B) utilizar un descodificador de reconocimiento de habla para aplicar el modelo lingüístico probabilista a un flujo de audio hablado para producir un documento que incluye contenido organizado en la pluralidad de subestructuras, en el cual el contenido en cada una de la pluralidad de subestructuras es producido reconociendo el habla usando la subestructura, en el cual la pluralidad de modelos lingüísticos probabilistas están organizados en una jerarquía, y en el cual la etapa (B) comprende las etapas de: (B)(1) identificar una trayectoria a través de la jerarquía, que comprende las etapas de: (B)(1)(a) identificar una pluralidad de trayectoria a través de la jerarquía (B)(1)(b) para cada una de la pluralidad de trayectorias P, producir un documento estructurado candidato para el flujo de audio hablado usando el descodificador de reconocimiento de habla para reconocer el flujo de audio hablado usando los modelos lingüísticos en la trayectoria P;B(1)© aplicar una medición a la pluralidad de documentos estructurados candidatos producidos en la etapa (B)(1)(b) para producir una pluralidad de puntuaciones de pertinencia para la pluralidad de documentos estructurados candidatos;y (B)(1)(d) seleccionar la trayectoria que produce los documentos estructurados candidatos que tienen la mayor puntuación pertinente;(B)(2) generar el documento que tiene una estructura que corresponde a la trayectoria identificada en la etapa (B) (1).
- 2- El procedimiento de la reivindicación 1, en el cual la pluralidad de modelos lingüísticos probabilistas incluye al menos un modelo lingüístico de n-gramas.
- 3- El procedimiento de la reivindicación 1, en el cual la pluralidad de modelos lingüísticos probabilistas incluye al menos un modelo lingüístico de estado finito.
- 4- El procedimiento de la reivindicación 1, en el cual la pluralidad de subestructura incluye una subestructura que representa un concepto semántico.
- 5- El procedimiento de la reivindicación 4, en el cual el concepto semántico comprende una medicación.
- 6- El procedimiento de la reivindicación 1, que comprende, además, una etapa de:(C) presentar el documento para producir una representación que indica la estructura del documento.
Independent claims6
169 paragraphs in 10 sections, as filed
ES 2 394 726 T3
DESCRIPTION
Automatic extraction of semantic content and generation of a structured document from speech
Cross reference to related requests
This application is related to the United States patent application entitled "Document Transcription System Training".
Background
Field of the invention
The present invention relates to automatic speech recognition, and more particularly to techniques for automatically transcribing speech.
Related art
It is desirable in many contexts to generate a written document based on human speech. In the legal profession, for example, transcriptionists transcribe testimonies given in court proceedings and in depositions to produce a written transcript of the testimony. Also, in the medical profession, transcripts of diagnoses, prognoses, prescriptions, and other information dictated by physicians and other medical professionals are produced. Transcripts in these and other fields typically need to be very precise (measured in terms of the degree of correspondence between the semantic content (meaning) of the original speech and the semantic content of the resulting transcript) due to the reliance placed on the resulting transcripts and the detriment that could cause inaccuracy (such as giving a wrong drug prescription to a patient). High degrees of reliability can, however, be difficult to obtain consistently for several reasons, such as variations in: (1) the characteristics of the speakers whose speech is transcribed (e.g., accent, volume, dialect, speed) ; (2) external conditions (for example, background noise); (3) the transcriber or transcription system (eg imperfect audio listening or capture capabilities, imperfect understanding of language); or (4) recording / transmission medium (eg, paper, analog audio tape, analog telephone network, compression algorithms applied in digital telephone networks, and noise / artifacts due to cell phone channels).
At first, transcription was only carried out by human transcriptionists who listened to speech, speech, either in real time (for example, in person "taking dictation") or listening to a recording. An advantage of human transcriptionists is that they may have field-specific knowledge, such as knowledge of medicine and medical terminology, which allows them to interpret ambiguities in speech and thus improve the accuracy of transcription. Human transcriptionists, however, have several drawbacks. For example, human transcriptionists produce transcripts at a relatively slow rate, and their accuracy decreases over time as a result of fatigue.
Various automatic speech recognition systems exist to recognize human speech generally and to transcribe speech in particular. Speech recognition systems that create transcripts are called "automated transcription systems" or "automated dictation systems." Ready-to-use disk software, for example, can be used by personal computer users to dictate documents into a word processor as an alternative to typing such documents using a keyboard.
Automated dictation systems typically attempt to produce a word-for-word transcript of speech. Such a transcription, in which there is a one-to-one correspondence between the words in the spoken audio stream and the words in the transcript, is referred to herein as "verbatim transcription". Automated dictation systems are not perfect and therefore can fail to produce literally perfect transcripts.
In some circumstances, however, a verbatim transcription is not desirable. In fact, transcriptionists can intentionally introduce various changes to the written transcript. A transcript can, for example, filter out spontaneous speech effects (e.g. pause expressions, hesitations, false starts), discard irrelevant remarks and comments, convert data into a standard format, insert headings or other explanatory material, or change the sequence of the speech. speech to adjust the structure of a written report.
In the medical field, for example, spoken reports produced by physicians are often transcribed into written reports in standard formats. For example with reference to FIG. 1B, an example of a structured and formatted medical report 111 is shown. The report 111 includes a variety of sections 112-138 that appear in a predetermined sequence when the report 111 is displayed. In the particular example shown in Figure 1B, the report includes a heading section 112, a subjective section 122, an objective section 134, an evaluation section 136, and a plan section 138. Sections can include text as well as subsections. . For example, heading section 112 includes a hospital name section 120 (containing the text "General Hospital"), a patient name section 114 (containing the text "Jane DOE", a card number section 116 (containing the text “851D”), and a report date section 118 (containing the text (1/10/1993 ”).
ES 2 394 726 T3
Likewise, the subjective section 122 includes various subjective information about the patient, included both in the text and in a medical history section 124, a medications section 126, an allergy section 128, a family history section 130, and a section of social history 132. Objective section 134 includes various objective information. Although not illustrated in Figure 1B, the information in the objective section may include subsections that contain the illustrated information. Evaluation section 136 includes a textual evaluation of the patient's condition, and plan subsection 138 includes a textual description of a treatment plan.
It should be noted that the information may appear in a different form in the 111 report than the way such information was dictated by the physician. For example, the date in the report date section 118 may have been stated as “October, 1993,” “October 1, 1993,” or another way. The transcriber, however, transcribed such speech using the text “10/1/1993) in the date section of report 118, perhaps because the hospital specified in the hospital section 120 requires that the dates be expressed in the reports written with such format.
Also, the information in the medical report 111 may not appear in the same sequence as in the original audio recording, due to the need to conform to a required report format or for some other reason. For example, the prescribing physician may have dictated objective section 134 first, followed by subjective section 122, and then heading 120. The written report 111, however, contains the heading 120 first, followed by the subjective section 122 and then the objective section 134. Such a report structure may, for example, be necessary for medical reports at the hospital specified in the hospital section 120.
The beginning of report 111 may have been generated based on a spoken audio stream such as the following: "Dr. Smith uh October 1 um 1993, patient identity eighty five one d um below is the patient's family history which I have reviewed .... ”It should be apparent that a verbatim transcript of this speech would be difficult to understand and not particularly helpful.
It should be noted, for example, that some words, such as "next is a" do not appear in the written report 111. Likewise, the expression that pauses as "uh" does not appear in the written report 111. Furthermore, the written report 111 Organize the original speech into the predefined sections 112-140 by rearranging the speech. As these examples illustrate, Written Report 111 is not a verbatim transcript of the doctor's speech that he delivers.
In summary, a report such as Report 111 may be more desirable than a verbatim transcript for a number of reasons (for example, because it organizes the information in a way that makes it easier to understand). Therefore it would be desirable for an automated transcription system to be able to generate a structured report (rather than a verbatim transcript) based on unstructured speech.
Referring to FIG. 1A, there is shown a data flow diagram of a prior art system 100 for generating a structured document 110 based on a spoken audio stream 102. Such a system produces the structured text document 110 from the spoken audio stream 102 using a two-step procedure: (1) an automated speech recognizer 104 generates a verbatim transcription 106 based on the spoken audio stream 102; and (2) a natural language processor 108 identifies the structure in the transcript 106 and thus creates the structured document 110, which has the same content as the transcript106, but is organized within the structure (e.g. report format ) identified by the natural language processor 108.
For example, some existing systems try to generate structured text documents; (1) analyzing spoken audio stream 102 to identify and distinguish the spoken content in the audio stream 102 from explicit or implicit structural tracks in the audio stream 102; (2) converting the "content" portions of the spoken audio stream 102 into raw text; and (3) using the identified structural clues to convert the raw text into the 110 structured report. Examples of explicit structural clues include formatting instructions (eg "new paragraph", "new line", "next point") and paragraph identifiers (eg "findings", impressions "conclusions"). Examples of implicit structural clues include long pauses that can indicate paragraph boundaries, prosodic cues that indicate the end of an enumeration, and the spoken content itself.
For various reasons described in more detail hereinafter, the structured document 110 produced by the system 100 may not be optimal. For example, structured document 110 may contain incorrectly transcribed (e.g., misrecognized) words, structure of structured document 110 may fail to reflect the desired structure of the document, and content of spoken audio stream 102 may be inserted into substructures. (for example, sections, paragraphs, or sentences) in the structured document.
Also, in addition to or instead of generating the structured document 110 based on the spoken audio stream 102, it may be desirable to extract the semantic content (such as information about the patient's previous medications, allergies, or diseases described in the audio stream 102 ) of the spoken audio stream 102. While such semantic content may be useful for generating structured document 110, such content may also be useful for other purposes, such as popularizing a database of patient information that can be analyzed independently of document 110. Prior art systems such as the system 100 shown in Figure 1, however, are typically intended to generate the structured document 110 based primarily or only on syntactic information in the spoken audio stream 102. Such systems are not, for therefore useful for extracting semantic content.
ES 2 394 726 T3
What is needed, however, is improved techniques for generating structured documents based on spoken audio streams.
US 2002/0123891 discloses a speech-to-text conversion procedure using a hierarchy of contextual models. In the disclosed method, the contextual model that more accurately reflects one or more spoken user expressions is used to convert speech to text. The preamble to the independent claims appended hereto is based on this document.
Summary
Techniques are disclosed to automatically generate structured documents based on speech, including identification of relevant concepts and their interpretation. In one embodiment, a structured document generator uses an integrated method to generate a structured text document (such as a structured text medical report) based on a spoken audio stream. The spoken audio stream can be recognized using a linguistic model that includes a plurality of sub-models arranged in a hierarchical structure. Each of the sub-models can correspond to a concept that is intended to appear in the spoken audio stream. For example, submodels can correspond to document sections. Submodels can, for example, be n-gram language models or contextless grammars. Different portions of the spoken audio stream can be recognized using different sub-models. The resulting structured text document can have a hierarchical structure that corresponds to the hierarchical structure of the linguistic sub-models that were used to generate the structured text document.
For example, in one aspect of the present invention, there is provided a method according to independent claim 1.
Other features and advantages of various aspects and embodiments of the present invention will become apparent from the following description and from the claims.
Brief description of the drawings
Figure 1A is a data flow diagram of a prior art system for generating a structured document based on a spoken audio stream;
Figure 1B illustrates a generated text medical report based on a spoken report;
Figure 2 is a flow chart of a method that is executed in an embodiment of the present invention to generate a structured text document based on a spoken document;
Figure 3 is a data flow diagram of a system carrying out the method of Figure 2 in one embodiment of the present invention;
Figure 4 illustrates an example of a spoken audio stream in one embodiment of the present invention;
Figure 5 illustrates a structured text document in accordance with one embodiment of the present invention;
Figure 6 is an example of a displayed document that is displayed based on the structured text document of Figure 5 in accordance with one embodiment of the present invention.
Figure 7 is a flow chart of a procedure that is executed by the structured document generator of Figure 3 in an embodiment of the present invention to generate a structured text document.
Figure 8 is a data flow diagram illustrating a portion of the system of Figure 3 in detail related to the method of Figure 7 in accordance with one embodiment of the present invention.
Figure 9 is a diagram illustrating correspondences between language models, document substructures corresponding to the language models, and candidate content produced using the language models in accordance with one embodiment of the present invention;
Figure 10A is a diagram illustrating a hierarchical language model in accordance with one embodiment of the present invention;
Figure 10B is a diagram illustrating a path through the hierarchical language model of Figure 10A in accordance with one embodiment of the present invention;
Figure 10C is a diagram illustrating a hierarchical language model in accordance with one embodiment of the present invention;
FIG. 11A is a flow chart of a method that is executed by the structured document generator of FIG. 3 to generate a structured text document in accordance with one embodiment of the present invention;
Figure 11B is a flowchart of a method using an integrated process to select a path through a hierarchical linguistic model and to generate a structured speech-based text document in accordance with one embodiment of the present invention;
Figures 11C-11D are flowcharts of procedures that are performed in one embodiment of the present invention to calculate a relevance score for a candidate document;
Figure 12A is a data flow diagram illustrating a portion of the system of Figure 3 in detail related to the method of Figure 11A in accordance with one embodiment of the present invention.
ES 2 394 726 T3
Figure 12B is a data flow diagram illustrating an embodiment of the structured document generator of Figure 3 that performs the method of Figure 11B in one embodiment of the present invention.
Figure 13 is a flow chart of a method that is used in one embodiment of the present invention to generate a hierarchical linguistic model for use in generating structured text documents.
Figure 14 is a flow diagram of a method that is used in an embodiment of the present invention to generate a structured text document using different stages of speech recognition and structural analysis; Y
Figure 15 is a flow chart of a system applying the method of Figure 14 in accordance with one embodiment of the present invention.
Detailed description
Referring to FIG. 2, a flow chart of a method 200 that is applied in an embodiment of the present invention is shown to generate a structured text document based on a spoken document. Referring to Figure 3, there is shown a data flow diagram of a system 300 for applying the method 200 of Figure 2 in accordance with one embodiment of the present invention.
System 300 includes a spoken audio stream 302, which may, for example, be a live or recorded spoken audio stream of a medical report dictated by a physician. Referring to FIG. 4, a text representation of an example of the spoken audio stream 302 is shown. In figure 4, the text between the percentage signs represents the spoken punctuation (for example "% comma%", "% point%", and "% colon%") explicit structural cues (for example "% new paragraph%" ) in audio stream 302. It can be seen from the audio stream 302 illustrated in Figure 4 that a verbatim transcription of the audio stream 302 would not be particularly helpful in order to understand the diagnosis, prognosis or other information contained in the medical report represented by the stream of audio 302.
System 300 also includes a probabilistic linguistic model 304. The term "probabilistic linguistic model" used herein refers to any linguistic model that assigns probabilities to spoken word sequences. Context-free (probabilistic) grammars and 306a-e n-gram linguistic models are both examples of "probabilistic linguistic models" as this term is used herein.
In general, a context-free grammar specifies a plurality of spoken forms for a concept and associates probabilities to each of the spoken forms. A finite state grammar is an example of a contextless grammar. For example, a finite state grammar for the date October 1, 1993, might include the spoken form October 1, 1993 "with a probability of 0.7, the spoken form ten ninety-three" with a probability of 0, 2 and the spoken form "first of October ninety-three" with a probability of 0.1. The probability associated with each spoken form is an estimated probability that the concept will be spoken in that spoken form in a particular audio stream. A finite state grammar is therefore a type of probabilistic linguistic model.
In general, an n-gram linguistic model specifies the probability that a particular sequence of n words will occur in a spoken audio stream. Consider, for example, a “unigram” linguistic model, for which n = 1. For each word in a language, a unigram specifies the probability that the word will occur in a spoken document. A “bigrama” linguistic model (for which n = 2) specifies probabilities that pairs of words will occur in a spoken document. For example, a bigrama model can specify the conditional probability that the word “cat” will occur in a spoken document since the previous word in the document was “the”. Likewise, a "trigram" linguistic model specifies the probabilities of three words, and so on. The probabilities specified by n-gram linguistic models and finite state grammars can be obtained by forming such documents using a training speech and a training text, as described in more detail in the above referenced patent application entitled "Document Transcription System Training ”.
The probabilistic linguistic model 304 includes a plurality of submodels 306a-e, each of which is a probabilistic linguistic model. The 306a-e submodels can include n-gram linguistic models. The 306a-e submodels can include n-gram linguistic models and / or finite state grammars in a combination. Also, as described in more detail hereinafter, each of the sub-models 306a-e may contain additional sub-models, and so on. Although five submodels are shown in Figure 3, the probabilistic linguistic model 304 can include any number of submodels.
The objective of the system 300 shown in Figure 3 is to produce a structured text document 310 that includes the content of the spoken audio stream 302, in which the content is organized in a particular structure and where the concepts are identified and interpreted in a machine-readable form. Structured text document 310 includes a plurality of substructures 312a-f such as sections, paragraphs, and / or sentences. Each of the 312a-f substructures can include additional substructures, and so on. Although six substructures are shown in Figure 3, the structured text document 310 can include any number of substructures.
ES 2 394 726 T3
For example, referring to FIG. 5, an example of the structured text document 310 is shown. In the example illustrated in FIG. 5, the structured text document 310 is an XML document. The structural text document
310 it can, however, be applied in any way. As shown in figure 5, the structured document
310 includes six substructures 312a-f, each of which may represent a section of document 310.
For example, structured document 310 includes heading section 312a that includes metadata about document 310, such as a title 314 of document 310 ("Noncontrast Chest CT Scan") and the date 316 on which the document was issued. 310 (“<date> Apr 22, 2003 </date>”). Note that the content in the header section 312a was obtained from the beginning of the spoken audio stream 302 (Figure 4). Also, it should be noted that heading section 312a includes both plain text (for example, title 314) and a substructure (for example, date 316) that represents a concept that has been interpreted in a machine-readable way as a triplet of values (day-month-year).
Representing the date in a computer-readable form allows the date to be stored in a database and processed more easily than if the date were stored in text form. For example, if multiple dates in the audio stream 302 have been recognized and stored in machine-readable form, such dates can easily be compared with each other by a computer. In another example, statistical information about the content of the audio stream 302, such as the mean time between medical visits, can be easily generated if the dates are stored in computer-readable form. This advantage of embodiments of the present invention generally applies not only to dates but to the recognition of any type of semantic content and the storage of such content in machine-readable form.
The structured document 310 further includes a comparison section 312b, which includes content that describes previous studies carried out on the same patient as the patient who is the subject of the document (report) 310. It should be noted that the content in the comparison section 312b was obtained from the portion of the audio stream 302 that begins with “comparison with” and ends with “April six, two thousand and one”, but that the comparison section 312b it does not include the "compare to" text that is in an example of a section hint. The use of such indicia to identify the beginning of a section or other document substructure will be described in more detail hereinafter.
In summary, structured document 310 includes a technical section 312c, which describes techniques that have been performed in procedures performed on the patient; a findings section 312d, which describes the physician's findings; and a printing section 312e, describing the physician's impressions of the patient.
XML documents, such as the exemplary structural document 310 illustrated in FIG. 5, are not typically intended to be directly viewed by an end user. Instead, such documents are typically represented in a way that is easier to read before being presented to the end user. The system 300, for example, includes a presentation engine 314 that presents the structured text document 310 based on a sheet. of style sheets 316 for producing a displayed document 318. Techniques for generating style sheets and for presenting documents according to style sheets are well known to those skilled in the art.
Referring to Figure 6, an example of the filed document 318 is shown. The filed document 318 includes five sections 602a-e, each of which may correspond to one or more of the six substructures 312a-f in the text document. structured 310. More specifically, the presented document 318 includes a header section 602a, a comparison section 602b, a technical section 602c, a findings section 602d, and a print section 602e. It should be noted that there may or may not be a one-to-one correspondence between sections in the displayed document 318 in the structured text document 310. For example, each of the substructures 312a-f need not represent a different type of document section. If, for example, two or more substructures 312a-f represent the same type of section (such as a header section), the presentation engine 314 can present both substructures in the same section of the presented document 318.
The system 300 includes a structured document generator 308, which identifies the probabilistic linguistic model 304 (step 202), and uses the linguistic model 304 to recognize the spoken audio stream 302 and thus produce the structured text document 310 (step 204). The structured document generator 308 may, for example, include an automatic speech recognition decoder 320 that produces each of the substructures 312a-f in the structured text document 310 that uses a corresponding submodel of the submodels 306a-e in the Probabilistic Linguistic Model 304. As is well known to one of ordinary skill in the art, a decoder is a component of a speech recognizer that converts audio to text. Decoder 320 may, for example, produce substructure 312a using submodel 306a to recognize a first portion of the spoken audio stream 302. Likewise, decoder 320 may produce substructure 312b using submodel 306b to recognize a second portion of the audio stream. spoken audio 302.
It should be noted that there is no need for a one-to-one correspondence between submodels 306a-e in linguistic model 304 and substructures 312a-f in structured document 310. For example, the speech recognition decoder may use submodel 306a to recognize a first portion of the spoken audio stream 302 and thereby produce the substructures 312a, and use the same submodel 306a to recognize a second
ES 2 394 726 T3 portion of the spoken audio stream 302 and thereby produce the substructure 312b. In such a case, multiple substructures in the structured text document 310 may contain the content of a single semantic structure (eg, section or paragraph).
Submodel 306a may, for example, be a "header" linguistic model that is used to recognize portions of the spoken audio stream 302 that contain content in header section 312a; submodel 306b may, for example, be a "comparison" linguistic model that is used to recognize portions of the spoken audio stream 302 that contains content in comparison section 312b; and so on. Each such linguistic model can be trained using training text from the corresponding section of the training documents. For example, the header submodel 306a can be trained using text from the header sections of a plurality of training documents, and the comparison submodel can be trained using text from the comparison sections of the plurality of training documents.
Having generally described features of various embodiments of the present invention, embodiments of the present invention will now be described in more detail. Referring to Figure 7, there is shown a flow diagram of a procedure that is applied by structured document generator 308 in one embodiment of the present invention to generate structured text document 310 (Figure 2, step 204). Referring to Figure 8, a data flow diagram is shown illustrating a portion of the system 300 in detail relevant to the method of Figure 7.
In the example illustrated in FIG. 8, the structured document generator 308 includes a segment identifier 814 that identifies a plurality of S segments 802a-c in the spoken audio stream 302 (step 701). Segments 802a-c can, for example, represent concepts such as sections, paragraphs, sentences, words, dates, times, or codes. Although only three segments 802a-c are shown in FIG. 8, the spoken audio stream 302 can include any number of portions. Although for ease of explanation, all 802a-c segments are identified in step 701 of Figure 7 before performing the remainder of procedure 700, identification of 802a-c segments can be performed concurrently with recognition of audio stream 302 and generating structured document 310, as will be described in more detail hereinafter with respect to Figures 11B and 12B.
The structured document generator 308 loops each S segment in the spoken audio stream 302 (step 702). As described above, structured document generator 308 includes speech recognition decoder 320, which may, for example, include one or more conventional speech recognition decoders that include different language models. Also as described above, each of the submodels 306a-e can be an n-gram linguistic model, a contextless grammar, or a combination thereof.
It is assumed by way of example that structured document generator 308 is currently processing segment 802a of spoken audio stream 302. Structured document generator 308 selects a plurality 804 of sub-models 306a-e with which to recognize the current S segment. The sub-models 804 may, for example, be all linguistic sub-models 306a-e or a subset of the sub-models 306a-e. The speech recognition decoder 320 recognizes the current segment S (eg, segment 802a) with each of the selected sub-models 804, thereby producing a plurality of candidate content 808 corresponding to segment S (step 704). In other words, each of the candidate content 808 is produced using the speech recognition decoder 320 to recognize the current segment S using a different sub-model of the sub-models 804. Note that each of the candidate content 808 may include not only recognized text but also other types of content such as concepts (eg, dates, times, codes, medications, allergies, vital signs, etc.). Encoded in machine readable form.
The structured document generator 308 includes a final content selector 810 that selects one of the candidate content 808 as final content 812 for segment S (706). The final content selector 810 may use any of a variety of techniques that are well known to one of ordinary skill in the art to select the speech recognition result that most closely matches the speech from which it is derived.
The structured document generator 308 keeps track of the submodel that is used to produce each of the candidate content 808. As an example, the submodels 304 include all of the submodels 306a-e, and that the candidate content 808 includes at least both five candidate contents per 802a-c segment (one produced using each of the 306a-e submodels). For example, referring to FIG. 9, a diagram is shown illustrating mappings between document substructures 312a-f, submodels 306a-e, and candidate contents 808a-e. As described above, each of the submodels 306a-e can be associated with one or more corresponding substructures 312a-f in structured text document 310. These mappings are indicated in Figure 9 by mappings 902a-e between substructures 312a-e and submodels 306a-e. The structured document generator 308 may maintain such mappings 902a-e in a table or use other means.
When the speech recognition decoder 320 recognizes the S segment (eg, segment 802a) with each of the sub-models 306a-e, it produces the corresponding candidate content 808a-e, For example, the candidate content 808a is the text that is produced when the candidate content decoder 320 recognizes the segment
ES 2 394 726 T3
802a with submodel 306a, candidate content 808b is the text that is produced when speech recognition decoder 320 recognizes segment 802a with submodel 306b, and so on. Structured document generator 308 may record the mapping between candidate content 808a-e and corresponding submodels 306a-e in a candidate model-content mapping set 816.
Therefore, when the structured document generator 308 selects one of the candidate contents 808a-e as the final content 812 for the segment S (step 706), a final match identifier 818 can use the matches 816 and the selected final content 812 to identify the linguistic submodel that produced the candidate content that has been selected as the final content 812 (step 708). For example, if candidate content 808c is selected as final content 812, it can be seen in FIG. 9 that final match identifier 818 can identify submodel 306C as the submodel that produced candidate content 808c. The final correspondence identifier 818 can accumulate each identified submodel in the correspondence set 820, such that at any given time the matches 820 identify the sequence of linguistic submodels that were used to generate the final contents that have been selected for inclusion in structured text document 310.
Once the sub-model corresponding to the final content 812 has been identified, the structured document generator 308 can identify the document substructure associated with the identified sub-model (step 710). For example, if submodel 306c has been identified in step 708, it can be seen from FIG. 9 that document substructure 312c is associated with submodel 306c.
A structured content inserter 822 inserts the final content 812 into the identified substructure of the structured text document 310 (step 712). For example, if substructure 312c is identified in step 710, text inserter 514 inserts final content 812 into substructure 312c.
The structured document generator repeats steps 704-712 for the remainder of the 802b-c segments of the spoken audio stream 302 (step 714), thereby generating the final content 812 for each of the remaining 802b-ce segments by inserting the final content 812 in the appropriate substructures of the substructures 312a-f of the text document 310. At the conclusion of procedure 700, the structured text document 310 includes text that corresponds to the spoken audio stream 302, and the final model-content matches 820 identify the sequence of linguistic sub-models that were used by the speech recognition decoder 320 to generate the text in structured text document 310.
It should be noted that in the process of recognizing the spoken audio stream 302, the procedure 700 can not only generate text that corresponds to the spoken audio, but can also identify semantic information represented by the audio and store such semantic information in a readable form. machine. For example, referring back to Figure 5, comparison section 312b includes a date element in which a particular date is represented as a triplet containing individual values for the day ("06"), month ("APR "), And the year (" 2001 "). Other examples for semantic concepts in the medical field include vital signs, medications and their dosages, allergies, medical codes, etc. The extraction and representation of semantic information in this way facilitates the process of applying automated processing on such information. It should be noted that the particular way of representing semantic information in Figure 5 is merely an example and does not constitute a limitation of the present invention.
As will be recalled from step 701, the procedure 700 shown in FIG. 7A identifies the set of 802a-c segments before identifying the sub-models to use to recognize the 802a-c segments. It should be noted, however, that the structured document generator 308 can integrate the process of identifying the 802a-c segments with the process of identifying the sub-models to be used to recognize the 802a-c segments, and with the process of applying speech recognition over 802a-c segments. Examples of techniques that can be used to apply such integrated targeting and recognition will be described in more detail hereinafter with respect to Figures 11B and 12B.
Having generally described the operation of the procedure illustrated in FIG. 7, we now consider applying the procedure of FIG. 7 to the exemplary audio stream 302 shown in FIG. 4. The first portion of the spoken audio stream 302 is assumed to be the spoken flow of the expressions "CT scan of the thorax without contrast April 22, two thousand and three". This portion can be selected at step 702 and recognized using all language sub-models 306a-e at step 704 to produce a plurality of candidate content 808a-e. As described above, assuming that submodel 306a-e is a "headline" linguistic model, that submodel 306b is a "comparison" linguistic model, that submodel 306c is a technique linguistic model, that submodel 306d is a linguistic model of "findings", and the one that submodel 306e is a linguistic model of "impressions"
Due to the fact that the 306a submodel is a linguistic model that has been trained to recognize speech in the "header" section of document 310 (eg, the 312a substructure), it is likely that the 808a candidate content produced using the 306a submodel matches the words of the aforementioned audio portion to a greater extent than the other content candidate 808b-e. Assuming that the candidate content 808a is selected as the final content 812 for this audio portion, the content inserter 822 will insert the
ES 2 394 726 T3 final content 812 produced by submodel 306a in header section 312a of structured text document 310.
The second portion of the spoken audio stream is supposed to be the spoken stream of expressions "compared to previous studies of March 6, 2000 and April 6, 2001." This portion can be selected at step 702 and recognized using all language sub-models 306a-e at step 704 to produce a plurality of candidate content 808a-e. Due to the fact that submodel 306b is a linguistic model that has been trained to recognize speech in the "comparison" section of document 310 (eg, substructure 312b), it is likely that the candidate content 808b produced using submodel 306b will match with the words of the candidate content portion to a greater extent than the other content referenced above 808a and 808c-e. Assuming that the candidate content 808b is selected as the final content 812 for this audio portion, the text inserter 514 will insert the final content 812 produced by the submodel 306 into the comparison section 312b of the structured text document 310.
The remainder of the audio stream 302 illustrated in FIG. 4 can be recognized and inserted into the appropriate substructures of substructures 312a-f in the similarly structured text document 310. Note that although the content of the spoken audio stream 302 illustrated in Figure 4 appears in the same sequence as sections 312a-f in structured text document 310, it is not a condition of the present invention. Instead, the content can appear on the audio stream 302 in any order. Each of the segments 802a-c of the audio stream 302 is recognized by the speech recognition decoder 320, and the resulting final content 812 is inserted into the appropriate substructure of the substructures 312a-f. Consequently, the order of the text content in the substructures 312a-f may not be the same as the order of the content in the spoken audio stream. It should be noted, however, that even if the order of the text content is the same in both the audio stream 302 as it is in the structured text document 310, the presentation engine 314 (Figure 3) can render the text content of the document 310 in any desired order.
In another embodiment of the present invention, the probabilistic language model 304 is a hierarchical language model. In particular, in this embodiment the plurality of sub-models 306a-e are organized in a hierarchy. As described above, submodels 306a-e may further include additional submodels, and so on, such that the hierarchy of linguistic model 304 may include multiple levels.
Referring to FIG. 10A, a diagram is shown illustrating an example of the linguistic model 304 in a hierarchical manner. Linguistic model 304 includes a plurality of nodes 1002, 306a-e, 1006a-e, and 1010 and 1012. Square nodes 1002, 306b-e, and 1006e and 1012 use probabilistic finite state grammars to model very limited concepts (such as such as the order of report sections, section hints, dates, and times). Elliptical nodes 306a, 1006a-d, and 1010 use statistical linguistic models (of n-grams) to model a less limiting language.
The term "concept" as used herein includes, for example, dates, times, numbers, codes, medications, medical history, diagnoses, prescriptions, expressions, enumerations, and section indicia. A concept can be expressed verbally in many ways. Each way of verbally expressing a particular concept is referred to herein as the "spoken form" of the concept. Sometimes a distinction is made between "semantic" concepts and "syntactic" concepts. The term "concept" as used herein includes both semantic concepts and syntactic concepts, but is not limited to either of them and no particular definition of "semantic concept" or "syntactic concept" is based or on any distinction. Between both.
Consider, for example, the date of October 1, 1993, which is an example of a concept as this term is used herein. The spoken forms of this concept include the spoken expressions "October 1, nineteen hundred and ninety-three", October one, ninety-three ", one indent ten indent ninety-three". Text such as “October 1, 1993) and“ 10/01/1993 ”are examples of“ spoken forms ”of this concept.
The phrase "John Jones has pneumonia" is now considered. This phrase, which is a concept as this term is used herein, can be expressed verbally in various ways, such as the spoken expressions, "John Jones has pneumonia" and "Jones patient diagnosed with pneumonia." The written phrase "John Jones has pneumonia" is an example of a "written form" of the same concept.
Although the linguistic models for low-level concepts such as dates and times are not shown in Figure 10A (except for submodel 1012), the hierarchical linguistic model 304 may include submodels for such low-level concepts. For example, the n-gram submodels 306a, 1006a-d, and 1010 can assign probabilities to sequences of words that represent dates, times, and other low-level concepts.
The linguistic model 304 includes the root node 1002, which contains a finite state grammar that represents the probabilities of occurrence of the subnodes 306a-e of the node 1002. The root node 1002 can, for example, indicate probabilities of the heading sections, comparisons, findings and impressions of document 310 that appear in particular orders in the spoken audio stream 302.
Moving down one level in the hierarchy of the 304 linguistic model, node 306a is a "header" node, which is an n-gram linguistic model that represents word occurrence probabilities in portions of the spoken audio stream 302 intended for inclusion. in heading section 312a of structured text document 310.
ES 2 394 726 T3
Node 306b contains a "comparison" finite state grammar representing occurrence probabilities of various alternative spoken forms of clues for the comparison section 312b of the text document. The finite state grammar in comparison node 306 may, for example, include clues such as "compare to", "compare to", "before is", and prior studies are ". The finite state grammar can include a probability for each of these clues. Such probabilities can, for example, be based on observed usage frequencies of the cues in a training speech set for the same speaker or in the same field as the spoken audio stream 302. Such frequencies can be obtained, for example, using the techniques disclosed in the aforementioned patent application entitled "Document Transcription System Training".
The comparison node 306e includes a "comparison content" subnode 1006a, which is a linguistic n-gram model representing word occurrence probabilities in portions of the spoken audio stream 302 intended for inclusion in the body of the speech section. comparison 312b of text document 310. The comparison content node 1006a has a date node 1012 as a child. As described in more detail hereinafter, the date node 1012 is a finite state grammar that represents probabilities of the date being verbally expressed in various ways.
Nodes 306c and 306d can be understood similarly. Node 306c contains a "state of the art" finite grammar representing occurrence probabilities of various alternative spoken forms of clues for technical section 312c of text document 310. Technical node 306c includes a "technical content subnode" 1006b, which is an n-gram language model representing word occurrence probabilities in portions of the spoken audio stream 302 intended for inclusion in the body of technical section 312c. of text document 310. In addition, node 306d contains a "findings" finite state grammar representing occurrence probabilities of various alternative spoken forms of clues for finding section 312d of text document 310. The findings node 306d includes a “findings content” submode 1006c, which is a linguistic model of n-grams representing word occurrence probabilities in portions of the spoken audio stream 302 intended for inclusion in the body of the section. of finding 312d of text document 310.
Print node 306 is similar to nodes 306b-d- in that it includes a finite state grammar 1006 that includes an n-gram linguistic model for recognizing section hints and a submode 1006d that includes an n-gram linguistic model to recognize the content of sections. In addition, however, the print node 306e includes an additional submode 1006e which in turn includes a submode 1010. This indicates that the content of the impressions section can be recognized using either the linguistic model at the impressions content node 1006d or the "enum" node 1006e, governed by the linguistic model based on the finite state grammar that corresponds to the node of impressions 306e. The "enum" node 1006e contains a finite state grammar indicating probability associated with different ways of verbally expressing enumeration cues (such as "number one", "number two", "first", second "," third ", and so on). The impressions content node 1010 may include the same linguistic model as the impressions content node 1006d.
Having described the hierarchical structure of language model 304 in one embodiment of the present invention, examples of techniques that can be used to generate structured document 310 using language model 304 will now be described. Referring to FIG. 11A, there is shown a flow diagram of a procedure that is applied by the structured document generator 308 in an embodiment of the present invention 308 in an embodiment of the present invention to generate the structured text document 310 ( Figure 2, step 204). Referring to FIG. 12A, a data flow diagram is shown illustrating a portion of the system 300 in detail relevant to the procedure of FIG. 11A.
The structured document generator 308 includes a path selector 1202 that identifies a path 1204 through the hierarchical language model 304 (step 1102). Path 1204 is an ordered sequence of nodes in hierarchical linguistic model 304. The nodes may be multiple times traversed in path 1204. Examples of techniques for generating path 1204 will be described in more detail hereinafter with respect to Figures 11B and 12B.
Referring to Figure 10B, an example of path 1204 is illustrated. Path 1204 includes points 1020a-j, which specify a sequence in which to traverse nodes in linguistic model 304. Points 1020a-j are called " dots ”instead of“ nodes ”to distinguish them from nodes 1002, 306a-e, 1006a-e, and 1010 in the 304 language model.
In the example illustrated in Figure 10B, path 1204 traverses the following nodes of linguistic model 304 in sequence: (1) root node 1002 (point 1020a); (2) header content node 306a (dot 1020b): (3) comparison node 306b (dot 1020c); (4) comparison content node 1006a (item 1020d); (5) technical node 306c (item 1020e); (6) technical content node 1006b (item 1020f); (7) findings node 306d (point 1020g); (8) finding content node 1006c (item 1020h); (9) print node 306e (point 1020i); and (10) print content node 1006d (item 1020j).
As can be seen with reference to Figure 4, recognizing the spoken audio stream 302 using the linguistic sub-models found along the path 1204 illustrated in Figure 10B will result in optical speech recognition, as speech audio stream 302 occurs in the same sequence as the audio streams.
ES 2 394 726 T3 language sub-models in the path 1204 illustrated in FIG. 10B. For example, the spoken audio stream 302 begins with speech that is best recognized by the header content linguistic model 306a (Noncontrast Chest CT Scan "April 22 two thousand and three"), followed by speech that is best recognized by the linguistic comparison model 306b (“comparison with”), followed by speech that is best recognized by the 1006a comparison content language model (prior to the March 6, 2002 and April 6, 2001 studies (, and so on.
Having identified the path 1204, the structured document generator 308 recognizes the spoken audio stream 302 using the linguistic models traversed by the path 1204 to produce the structured text document 310 (step 1104). As described in more detail hereinafter with respect to Figures 11B and 12B, the speech recognition and generation of the structured text document from step 1104 can be integrated with the path identification from step 1102, rather than performed. separately..
More specifically, structured document generator 308 may include a node enumerator 1206 that repeats each of the N language model nodes 1208 traversed by the selected path 1204 (step 1106). For each such node N, the speech recognition decoder 320 can recognize the portion of the audio stream 302 that corresponds to the linguistic model at node N to produce the corresponding T structured text (step 1108). The structured document generator 308 can insert the text T 1210 into the structure of the structured text document 310 that corresponds to node N 1208 of the linguistic model 304 (step 1110).
For example, when node N is compare node 306b (FIG. 10A), compare node 306b can be used to recognize "compare with" text in spoken audio stream 302 (FIG. 4). Since the comparison node 306b corresponds to a document substructure (for example, the comparison section 312b) instead of the content, the result of the speech recognition performed in step 1108 in this case may be a document substructure, namely an empty "comparison" section. Such a section may be inserted into the structured document 310 at step 1110, for example, in the form of match tags "<comparison>" and "/ comparison>".
When Node N is Comparison Content Node 1006a (Figure 10A), Comparison Content Node 1006a can be used to recognize the text "before studies of March 26, two thousand two and April 6, two thousand one "In the spoken audio stream 302 (FIG. 4). Producing in this way the structured text “previous studies of <date> 26-MAR-2002 </date> and <date> 26-APR-2001 </date>, as shown in figure 5. This structured text can then be inserted into comparison section 312b at step 1110 (eg, between the "<comparison>" and "> / comparison>" tags, as shown in FIG. 5).
Structured document generator 308 repeats steps 1108-1110 for the remaining nodes N traversed by path 1204 (step 1112), thereby inserting a plurality of structured texts 1210 into structured text document 310. The end result of the procedure illustrated in FIG. 11A is the creation of the structured text document 310, which contains text that has a structure that corresponds to the structure of the '1204 path through the linguistic model 304. For example, it can be seen in FIG. 10B that the illustrated path structure traverses the linguistic model nodes that correspond to the heading, comparison, technique, findings, and prints sections in sequence. The resulting structured text document 310 (as illustrated, for example, in FIG. 5) similarly includes the header, comparison, technique, findings, and print sequences in sequence. The structured text document 310 therefore has the same structure as the linguistic model path 1204 that was used to create the structured text document 310.
It has previously been established that structured document generator 308 inserts recognized structured text 1210 into the appropriate substructures of structured text document 310 (FIG. 11A, step 1110). As shown in Figure 5, the structured text document 310 can be applied as an XML document or another document that supports nested structures. In such a case, it is necessary to insert each of the recognized structured texts 1210 into the appropriate substructure so that the final structured text document 310 has a structure that corresponds to the structure of the path 1204. One of ordinary skill in the art will understand how to use the final content mappings of model 820 (FIG. 8) to use path 1204 to traverse the structure of linguistic model 304 and therefore to create such a structured document.
The system illustrated in Figure 12A includes a path selector 1202, which selects a path 1204 through the linguistic model 304. The procedure illustrated in Figure 11A then uses the selected path 1204 to generate the structured text document 310. Said of otherwise, in FIGS. 11A and 12A, the path selection and structured document creation steps are performed separately. This is not, however, a limitation of the present invention.
Instead, referring to FIG. 11B, a flow chart of a procedure 1150 is shown that integrates the steps of path selection and generation of structured documents. Referring to FIG. 12B, there is shown an embodiment of the structured document generator 308 that applies the method 1150 of FIG. 11B in one embodiment of the present invention. In general, the procedure 1150 of FIG. 11B searches for possible paths through the hierarchy of the linguistic model 304 (FIG. 10A), starting at root node 1002 and expanding outward. Any of a number of techniques, including techniques well known to those skilled in the art
ES 2 394 726 T3 technique can be used to search through the hierarchy of linguistic models. Since procedure 1150 identifies partial trajectories through the linguistic model hierarchy, procedure 1150 uses speech recognition decoder 320 to recognize increasing portions of the spoken audio stream 302 using the linguistic patterns found throughout the partial paths thereby creating candidate structured partial documents. Procedure 1150 assigns scores to each of the partial candidate structured documents. The relevance score for each candidate structured document is a measure of how well the trajectory that produced the candidate structured document performed. The method 1150 expands the partial paths thus continuing to search through the language model hierarchy, until the entire spoken audio stream 302 has been recognized. The structured document generator 308 selects the candidate structured document that has the highest relevance score as the final structured text document 310.
More specifically, procedure 1150 initiates one or more candidate paths 1224 through language model 304 (step 1152). For example, candidate paths 1224 may be started to contain a single path consisting of root node 1002. The term "frame" refers herein to a short period of time, such as 10 milliseconds. Procedure 1150 initiates an audio stream pointer to point to the first frame in audio stream 302 (step 1153). For example, in the embodiment illustrated in FIG. 12B, the structured document generator 308 contains an audio stream enumerator 1240 that provides a portion 1242 of the audio stream 302 to the speech recognition decoder 320. Upon initiation of procedure 1150, portion 1242 may contain only the first frame of audio stream 302.
The speech recognition decoder 320 recognizes the current portion 1242 of the audio stream 302 using the linguistic submodels in the candidate path (s) 1224 to generate one or more candidate structured partitions 1232 (step 1154). It should be noted that documents 1232 are only partial documents 1232 because they have been generated based on only a portion of the audio stream 302. When step 1154 is applied first, the speech recognition decoder 320 can simply recognize the first frame of the audio stream 302 using the language model at the root node 1002 of the language model 304.
It should be noted that the techniques disclosed above with respect to FIG. 11A and FIG. 12A can be used by the speech recognition decoder 320 to generate the candidate structured part documents 1232 using the candidate paths 1224. More specifically, speech decoder 320 may apply the procedure illustrated in FIG. 11A to audio stream portion 1242 using each of the candidate paths 1224 as the path identified in step 1102 (FIG. 11A).
Returning to Figures 11B and 12B, a relevance evaluator 1234 generates relevance scores 1236 for each of the candidate structured sub-documents 1232 (step 1156). Relevance scores 1236 are measurements of how well the candidate structured subdocuments 1232 represent the corresponding portion of the audio stream 302. In general, the relevance score for an individual candidate document can be generated by: (1) generating relevance scores for each of the nodes in the corresponding path of candidate paths 1224; and (2) using a synthesis function to synthesize the individual node relevance score generated in step (1) into an overall relevance score for the candidate structured document. Examples of techniques that can be used to generate candidate relevance scores 1236 will be described in more detail hereinafter with respect to FIG. 11C.
If the structured document generator 308 were to attempt to search all possible trajectories through the hierarchy of the linguistic model 304, the computing resources necessary to evaluate each possible trajectory could be cost and time prohibitive. Due to the exponential growth in the number of possible trajectories. Therefore, in the embodiment illustrated in Figure 12B, a path pruner 1230 uses candidate relevance scores 1236 to remove misfit paths from candidate paths 1224, thereby producing a set of pruned paths 1222 (step 1158). ).
If the entire audio stream 302 has been recognized (step 1160), a final document selector 1238 selects, among candidate structured partial documents 1232, the candidate structured document that has the highest relevance score, and provides the selected document as the final structured text document 310 (step 1164). If the entire audio stream 302 has not been rearranged, a path extender extends the pruned paths 1222 within the linguistic model 304 to produce a new set of candidate paths 1224 (step 1162). If, for example, candidate paths 1222 consist of a single path that contains root node 1002, path extender 1220 may extend this path one node down the hierarchy illustrated in Figure 10A to produce a plurality of candidate paths that extend from root node 1002, such as a path from root node 1002 to header content node 306a, a path from root node 1002 to comparison node 306b, a path from root node 1002 to technical node 306c, and so on. Various techniques for extending trajectories 1224 to perform depth, breadth, or other types of hierarchical searches are well known to those of skill in the art.
Audio stream enumerator 1240 extends portion 1242 of audio stream 302 to include the next frame in audio stream 302 (step 1163). Steps 1154-1160 are then repeated using the new candidate paths.
ES 2 394 726 T3
1224 to recognize portion 1242 of audio stream 302. Thus the entire audio stream 302 can be recognized using appropriate sub-models in language model 304.
As described above with respect to Figures 11B and 12B, relevance scores 1236 can be generated for each of the candidate structured sub-documents 1232 produced by the structured document generator 308 while the candidate paths 1224 are evaluated through the model. linguistic 304. Examples of techniques for generating relevance scores will now be described, either for candidate structured sub-documents 1232 illustrated in FIG. 12B or for structured documents more generally.
For example, referring to FIG. 10A, it should be noted that the comparison content node 1006a has a date node 1012 as a descendant node. The text "Chest CT scan without contrast April 22, two thousand and three" is assumed to have been recognized as text corresponding to comparison content node 1006a. It should be noted that the comparison content node 1006a was used to recognize the text "CT scan of the thorax without contrast" and that the date node 1012, which is a descendant node of the comparison content node 1006a, was used to generate the text "April 22, two thousand and three." The relevance score for this text can therefore be calculated using the comparison content node 1006a to calculate a first relevance score for the text "Noncontrast Chest CT Scan Followed by any date, by calculating a second relevance score for the text "April 22nd two thousand three" based on the node dated 1012, and multiplying the first and second relevance scores.
Referring to Figure 11C, there is shown a flow chart of a procedure that is applied in one embodiment of the present invention to calculate a relevance score for a candidate document, and that can therefore be used to apply step 1156 of the Procedure 1150 illustrated in FIG. 11B. A relevance score S is started at a value of one for the candidate structured document being evaluated (step 1172). The procedure assigns a current node pointing N to point to the root node in the candidate path that corresponds to the candidate document (step 1174).
The procedure requires a function called Relevance () with the values N and S (step 1176) and returns the result as the relevance score for the candidate document (step 1178). As will now be described in more detail, the Relevance () function generates the relevance score S using hierarchical factorization traversing the candidate path that corresponds to the candidate document.
Referring to FIG. 11D, a flow chart of the Relevance () function 1180 according to one embodiment of the present invention is shown. The function 1180 identifies the probability P (W (N)) that the text W corresponding to the current node N has been recognized by the linguistic model associated with that node, and multiplies the probability by the current value of S to produce a new value for S (step 1184).
If node N has no downstream node (step 1186), the value of S is restored (step 1194). If node N has a descending node, then the Relevance () function 1180 is recursively required on each of the descending nodes, with the results being multiplied by the value of S to produce new values of S (steps 1188-1192 ). The resulting value of S is restored (step 1194).
Upon completion of the procedure illustrated in Figure 11C, the value of S represents a relevance score for the entire candidate structured document, and the value of S is restored, for example, for use in the procedure 1150 illustrated in Figure 11B (step 1194).
For example, going back to the text “CT scan of the thorax without contrast April 22, two thousand and three”. The relevance score (probability) of this text can be obtained by identifying the probability of the text “Chest CT scan without contrast <DATE>”, where <DATE> indicates any date, multiplied by the conditional probability of the text “twenty-two of April two one thousand three ”that occurs since the text represents a date.
More generally, the effect of the procedure illustrated in Figure 11C is to hierarchically incorporate the probabilities of word sequences according to the hierarchy of the 304 linguistic model, allowing the individual probability estimates associated with each linguistic model node to be easily combined with the estimates probability associated with other nodes, This probabilistic framework allows the system to model and use statistical linguistic models with built-in probabilistic finite-state grammars and built-in statistical linguistic models.
As described above, the nodes in the linguistic model 304 represent linguistic submodels that specify the probabilities of occurrence of sequences of words in the spoken audio stream 302. In the above discussion, it has been assumed that the probabilities have already been assigned in such linguistic models. Examples of techniques for assigning probabilities to linguistic submodels (such as linguistic models of ngrams and contextless grammars) in the 304 linguistic model will now be disclosed.
Referring to Figure 13, there is shown a flow diagram of a method 1300 that is used in one embodiment of the present invention to generate the linguistic model 304. A plurality of nodes are selected for use in the linguistic model (step 1302 ). The nodes can, for example, be selected by a transcriber or other person skilled in the relevant field. Nodes can be selected in an attempt to capture all types of
ES 2 394 726 T3 concepts that can occur in the spoken audio stream 302. For example, in the medical field, you can select nodes (such as those shown in Figure 10A) that represent the sections of a medical report and the conceptor (such as dates, times, medications, allergies, vital signs, and medical codes) expected to be on a medical report.
A concept and any type of linguistic model can be assigned to each of the selected nodes in step 1302 (steps 1304-1306). For example, node 306b (FIG. 10A) may be assigned to the concept "comparison section hint" and assigned to the linguistic model type "finite state grammar". Similarly, node 1006a can be assigned to the concept "comparison content" and the language model type "n-gram language model".
The nodes selected in step 1302 can be arranged in a hierarchical structure (step 1308). For example, nodes 1002, 306a-e, 1006a-e, and 1010 can be arranged in the hierarchical structure illustrated in Figure 10A to represent and enforce structural dependencies between nodes.
Each of the nodes selected in step 1302 can then be trained using text that represents a corresponding concept (step 1310). For example, a set of training documents can be identified. The set of training documents can, for example, be a set of existing medical reports or other documents in the same domain as the spoken audio stream 302. Training documents can be manually marked to indicate the existence and location of structures in the document, such as sections, subsections, dates, times, codes, and other concepts. Such marking can, for example, be performed automatically on formatted documents, or manually by transcriptionists or other qualified persons in the relevant field. Examples of techniques for training the selected nodes in step 1302 are described in the above referenced patent application entitled "Document Transcription System Training"
Conventional language model training techniques can be used in step 1310 to train the specific concept language models for each of the concepts that are marked in the training documents. For example, the text of all marked "heading" sections in the training documents can be used to train the linguistic model node 306a that represents the heading section. In this way, the linguistic models for each of the nodes 1002, 306a-e, 1006a-e, and 1010 can be trained in the linguistic model 304 illustrated in FIG. 10A. The result of procedure 1300 illustrated in FIG. 13 is a hierarchical linguistic model that has training capabilities, which can be used to generate structured text document 310 in the manner described above. This hierarchical linguistic model can then be used, for example, to repeatedly re-segment the training text, such as using the techniques disclosed above e in conjunction with Figures 11B and 12B. The resegmented training text can be used to retain the hierarchical linguistic model. This process of retargeting and retraining can be applied repeatedly to repeatedly improve the quality of the linguistic model.
In the examples described above, the structured document generator 308 recognizes the spoken audio stream 302 and generates the structured text document 310 using an integrated process, generating an unstructured intermediate transcript. Such techniques, however, are disclosed merely by way of example and do not constitute limitations of the present invention.
Referring to FIG. 14 there is shown a flow diagram of a method 1400 that is used in another embodiment of the present invention to generate the structured text document 310 using various stages of speech recognition and structural analysis. Referring to FIG. 15, there is shown a data flow diagram of a system 1500 that performs the method 1400 of FIG. 14 in accordance with one embodiment of the present invention.
The speech recognition decoder 320 recognizes the spoken audio stream 302 using a linguistic model 1506 to produce a transcript 1502 of the spoken audio stream 302. It should be noted that the linguistic model 1506 may be a conventional linguistic model that is distinct from the linguistic model. 304. More specifically, linguistic model 1506 may be a conventional monolithic linguistic model. The 1506 language model can, for example, be generated using the same training body as that used to train the 304 language model. While portions of the training body can be used to train the 304 language model, the entire body can be used to train the 1506 linguistic model. The speech recognition decoder 320 can thus use conventional speech recognition techniques to recognize the spoken audio stream 302 using the linguistic model 1506 and thereby produce the transcript 1502.
It should be noted that transcript 1502 may be a "flat" transcript 1502 of the spoken audio stream 302, rather than a structured document as in the previous examples disclosed above. The transcript 1502 may, for example, include a plain text sequence that resembles the text illustrated in Figure 4 (illustrating the spoken audio stream 302 in text form).
System 1500 also includes a structural analyzer 1504, which uses hierarchical linguistic model 304 to analyze transcription 1502 and thereby produce structured text document 310 (step 1404). The structural analyzer 1504 can use the techniques previously disclosed with respect to Figures 11C and 12B to: (1)
ES 2 394 726 T3 producing multiple candidate structured documents that have the same content as transcript 1502 but with structures corresponding to different trajectories through linguistic model 304; (2) generate a relevance score for each of the candidate structured documents; and (3) selecting the candidate structured document with the highest relevance score as the final structured text document. Contrary to the techniques disclosed above with respect to Figures 11C and 12B, however, step 1404 may be applied without performing speech recognition to generate each of the candidate structured documents. Instead, once transcript 1502 is produced using speech recognition decoder 320, candidate structured documents can be generated based on transcript 1502 without performing additional speech recognition.
Also, the structural analyzer 1504 does not need to use the entire linguistic model 304 to produce the structured text document 310. Instead, the structural analyzer 1504 may use a scaled down "skeletal" linguistic model, such as the illustrated linguistic model 1030. in Figure 10C. It should be noted that the exemplary language model 1030 shown in Figure 10C is the same as the language model 304 shown in Figure 10A, except that in the skeletal language model 1030 the content language model nodes 306a, 1006a-d, and 1010 have been replaced by universally accepted linguistic models 1032a-f, also referred to as the "Nevermind" linguistic model. The 1032a-f language models will accept whatever text is provided as input. The heading hint language model 306b-e in the skeletal linguistic model 1030 allows the structural parser 1504 to analyze the transcript 1502 into the correct substructures in the structured document 310. The use of the universally accepted linguistic models 1032a-f, however, allows the structural analyzer 1504 to perform such structural analysis without incurring the expense (typically considerable) of training content linguistic models, such as the 306a models. 1006a-d and 1010 shown in Figure 10A.
It should be noted that the skeletal linguistic model 1030 may also continue to include linguistic models, such as the linguistic dating model 1012, which corresponds to low-level concepts. As a consequence, the skeletal linguistic model 1030 can be used to generate structured document 310 from transcript 1502 without incurring the overhead of training content linguistic models, while retaining the ability to analyze lower-level concepts in structured document 310.
Among the advantages of the invention are one or more of the following. The techniques disclosed herein replace the traditional global linguistic model with a combination of specialized local linguistic models that are better suited to the section of a document than a single generic linguistic model. Such a linguistic model has several advantages.
For example, the use of a linguistic model that contains sub-models, each of which corresponds to a particular concept, is advantageous because it allows the most appropriate linguistic model to be used to recognize the speech that corresponds to each concept. In other words, if each of the sub-models corresponds to a different concept, then each of the sub-models can be used to apply speech recognition to the speech that represents the corresponding concept. Since the characteristics of speech can vary from concept to concept, the use of such concept-specific language models can produce better recognition results than would be produced using a monolithic language model for all concepts.
Although the submodels of a linguistic model may correspond to sections of a document, this is not a limitation of the present invention. Instead, each submodel in the linguistic model can correspond to any concept, such as a section, paragraph, sentence, date, time, or ICD9 code. Consequently, submodels in the linguistic model can match particular concepts with a degree of precision greater than would be possible if only section-specific linguistic models were used. The use of such concept-specific linguistic models for a wide variety of concepts can further improve speech recognition accuracy.
Also, hierarchical language models designated in accordance with embodiments of the present invention may have multi-level hierarchical structures, with the effect of nesting sub-models within each other. Due; The sub-models in the linguistic model can be applied to portions of the spoken audio stream 302 with various levels of granularity, with the most appropriate linguistic model being applied at each level of granularity. For example, a “heading section” linguistic model can be applied generally to speech within the heading section of a document, while a “date” linguistic model can be applied specifically to speech representing dates in the header section. This ability to nest linguistic models and apply nested linguistic models to different portions of speech can further improve recognition accuracy by allowing the most appropriate linguistic model to be applied to each portion of a spoken audio stream.
Another advantage of using a linguistic model that includes a plurality of sub-models is that the techniques disclosed herein can use such a linguistic model to generate a structured text document from a spoken audio stream using a single integrated process, in location of the prior art two-stage process 100 illustrated in FIG. 1A wherein the speech recognition stage is followed by a natural language processing stage. In the two-stage process 100 illustrated in FIG. 1A the steps carried out by speech recognizer 104 and natural language processor 108 are completely decoupled. Due to the automatic speech recognizer 104 and the natural language processor 108 they operate independently of each other.
ES 2 394 726 T3 other, the result 106 of the automatic speech recognizer 104 is a verbatim transcription of the spoken content in the audio stream 102. The verbatim transcription 106 thus contains the text corresponding to all the words spoken in the audio stream. audio 102, whether these words are relevant or not relevant to the final desired structured text document. Such words can include, for example, doubts, strange words or repetitions, as well as structural clues or words related to the task. In addition, the natural language processor 108 relies on the successful detection and transcription of certain keywords and / or key expressions, such as structural clues. If these key words / expressions are misrecognized by the automatic speech recognizer 104, the identification of structural entities by the natural language processor 108 may be adversely affected. In contrast, in the procedure 200 illustrated in Figure 2, speech recognition and natural language processing are integrated, thus allowing the linguistic model to influence both word recognition in the 4302 audio stream and generating structure in structured text document 310, thereby improving the overall quality of structured document 310.
In addition to generating structured document 310, the techniques disclosed herein can also be used to extract and interpret semantic content from audio stream 302. For example, the linguistic date model 1012 (Figures 10A-10B) can be used to identifying portions of the audio stream 302 that represent dates, and storing representations of such dates in computer-readable form. For example, the techniques disclosed herein can be used to identify the spoken expression "October 1, nineteen hundred and ninety-three" as a date and store the date in a computer-readable form, such as "month = 10, day = 1, year = 1998). Storing such concepts in a computer-readable form allows the content of such concepts to be easily processed by a computer, for example by selecting document sections by date or by identifying medications prescribed before a given date. In addition, the techniques disclosed herein allow the user to define different portions (eg, sections) of the document and to choose which concepts to extract from each section. The techniques disclosed herein thus facilitate the recognition and processing of semantic content in spoken audio streams. Such techniques can be applied instead of or in addition to storing extracted information in a structured document.
Areas such as the medical and legal fields, in which there are large bodies of pre-existing recorded audio streams for use as training text, can be particularly advantageous in the techniques disclosed herein. Such training text can be used to train linguistic model 304 using techniques disclosed above with respect to Figure 13. Since documents in such domains may be necessary to have well-defined structures, and since such structures can be easily identified in existing documents, it can be relatively easy (albeit time-consuming) to correctly identify portions of such concept-specific documents to its use in training each of the specific concept linguistic model nodes in the 304 linguistic model. As a consequence, each of the linguistic model nodes can be well trained to recognize the corresponding concept, thereby increasing the recognition accuracy and increasing the ability of the system to generate documents with the required structure.
Also, the techniques disclosed herein can be applied in such settings without requiring any changes to the existing process by which audio is recorded and transcribed. In the medical field, for example, doctors can continue to dictate medical reports in their usual way. The techniques disclosed herein can be used to generate documents with the desired structure regardless of the manner in which the spoken audio stream is dictated. Alternative techniques that require workflow changes, such as techniques that require speakers to register (reading training text), that require speakers to modify their way of speaking (for example, always saying concepts using predetermined spoken forms) , or require the transcripts to be generated in a particular format, can be cost prohibitive for application in fields such as the medical and legal fields. Such changes may, in fact, be inconsistent with institutional or legal needs related to the structure of the report (such as insurance reporting requirements). The techniques disclosed herein, in contrast, allow the audio stream 302 to be generated in any way and have any shape.
Also, the individual submodels 306a-e in the language model 304 can be easily updated without affecting the rest of the language model. For example, the header content submodel 306a-e can be replaced with a different header content submodel that is represented differently by the way the document header was rendered. 'The modular structure of the 304 language model' allows such sub-model modification / substitution to be carried out without the need to modify any part of the 304 language model. As a consequence, the parts of the 304 language model can be easily updated to reflect different agreements. dictation of documents.
Also, the structured text document 310 that is produced by various embodiments of the present invention can be used to train a linguistic model. For example, the training techniques described in the above referenced patent application entitled "Document Transcription System Training" can use structured text document 310 to retrain and thereby improve linguistic model 304. The retrained language model 304 can then be used to produce subsequent structured text documents, which can in turn be used to retrain the language model 304. This iterative process can be used to improve the quality of the structured documents. that occur over time.
ES 2 394 726 T3
It is to be understood that although the invention has been described above in terms of particular embodiments, the above embodiments are provided for illustrative purposes only, and do not limit or define the scope of the invention. Various other embodiments, including but not limited to the following, are also within the scope of the claims. For example, the elements and components described herein can be further divided into additional components or joined together to form fewer components to perform the same functions.
The spoken audio stream 302 can be any audio stream, such as a directly or indirectly received live audio stream (such as over a telephone or IP connection) or an audio stream recorded on any medium and in any format. In distributed speech recognition (DSR), a client performs preprocessing on an audio stream to produce a processed audio stream that is transmitted to a server, which performs speech recognition on the processed audio stream. . Audio stream 302 may, for example, be a processed audio stream produced by a DSR client.
Although each node in the linguistic model 304 is described in the above examples as containing a linguistic model that corresponds to a particular concept, it is not a requirement of the present invention. For example, a node may include a linguistic model that results from the interpolation of a specific concept linguistic model associated with the node with one or more of: (1) background global linguistic models with other nodes, or (2) specific linguistic models of concept associated with other nodes.
In the examples above, a distinction can be made between "grammars" and "text". It should be appreciated that the text can be represented as a grammar, in which it is a single spoken form that has only one probability. Therefore, the documents that are described herein as included in both the text and the grammars can be applied only using grammars if desired. Furthermore, a single-state grammar is simply a type of contextless grammar, which is a type of linguistic model that allows multiple alternative spoken forms of a concept to be applied more generally to any other type of grammar. Also, while the above description may refer to finite state grammars and n-gram language models, there are simply examples of types of language models that can be used in conjunction with embodiments of the present invention. Embodiments of the present invention are not limited to use in conjunction with any particular type or types of language model (s).
The invention is not limited to any of the fields described (such as medical and legal reports), but is generally applied to any type of structured documents.
The techniques described above can be applied, for example, in hardware, software, firmware, or any combination thereof. The techniques described above can be applied in one or more computer programs running on a programmable computer including a processor, a processor-readable storage medium (including, for example, volatile and non-volatile memory and / or storage elements) , at least one input device, and at least one output device. The program code can be applied to the input entered using the input device to carry out the described functions and generate the output. The output can be provided to one or more output devices.
Each computer program within the following claims may be applied in any programming language, such as assembly language, machine language, a high-level procedural programming language, or an object-oriented programming language. The programming language can, for example, be a compiled or interpreted programming language.
Each such computer program can be implemented in a computer program product tangibly embodied on a machine-readable storage device for execution by a computer processor. The steps of the method of the invention can be carried out by a computer processor that executes a program materialized in a tangible way on a computer-readable medium to apply the functions of the invention that work on the input and generate the output. Suitable processors include, by way of example, both general-purpose and special-purpose microprocessors. Generally, the processor receives instructions and data from read-only memory and / or random access memory. Storage devices suitable for tangibly embodying computer program instructions include, for example, all forms of non-volatile memory, such as semiconductor memory devices, including EPROM, EEPROM, and flash memory devices; magnetic drives such as internal hard drives and removable drives; magneto-optical discs; and CDROM. Any of the above can be supplemented with, or incorporated into, specially designed ASICs (Application Specific Integrated Circuits or FPGAs (Field Programmable Gate Arrays). A computer can generally also receive programs and data from a storage medium such as a internal disc (not shown) or removable disc). These elements will also be found in a conventional desktop or workstation computer as well as other computers suitable for running computer programs that apply the procedures described in this document, which can be used in conjunction with any digital printing engine or marking engine. , display monitor, or other raster output device capable of producing color or grayscale pixels on paper, film, display screen, or other means of egress.
Contents10
21 sheets
Sheet 1 Sheet 2 Sheet 3 Sheet 4 Sheet 5 Sheet 6 Sheet 7 Sheet 8 Sheet 9 Sheet 10 Sheet 11 Sheet 12 Sheet 13 Sheet 14 Sheet 15 Sheet 16 Sheet 17 Sheet 18 Sheet 19 Sheet 20 Sheet 21
60 members in 10 offices
Priority claims9
| Document | Office | Kind | Date |
|---|---|---|---|
| 923517 | United States of America | – | |
| 92351704 | United States of America | A | |
| 92351704 | United States of America | A | |
| 2005029354 | United States of America | W | |
| 2005029354 | United States of America | W | |
| 923517 | – | – | – |
| PCTUS2005029354 | – | – | – |
| US20040923517 | – | – | – |
| WO2005US29354 | – | – | – |
Members60
| Document | Office | Kind | |
|---|---|---|---|
| US2006041427A1 | United States of America | A1 | |
| US2006041428A1 | United States of America | A1 | |
| CA2577721A1 | Canada | A1 | |
| CA2577726A1 | Canada | A1 | |
| WO2006023622A2 | World Intellectual Property Organization (WIPO) | A2 | |
| WO2006023631A2 | World Intellectual Property Organization (WIPO) | A2 | |
| WO2006034152A2 | World Intellectual Property Organization (WIPO) | A2 | |
| US2006074656A1 | United States of America | A1 | |
| US2007033032A1 | United States of America | A1 | |
| WO2006023631A3 | World Intellectual Property Organization (WIPO) | A3 | |
| WO2007018842A2 | World Intellectual Property Organization (WIPO) | A2 | |
| WO2006034152A3 | World Intellectual Property Organization (WIPO) | A3 | |
| WO2006023622A3 | World Intellectual Property Organization (WIPO) | A3 | |
| EP1787287A2 | European Patent Office (EPO) | A2 | |
| EP1787288A2 | European Patent Office (EPO) | A2 | |
| WO2007018842A3 | World Intellectual Property Organization (WIPO) | A3 | |
| EP1908055A2 | European Patent Office (EPO) | A2 | |
| JP2008511024A | Japan | A | |
| EP1787288A4 | European Patent Office (EPO) | A4 | |
| EP1908055A4 | European Patent Office (EPO) | A4 | |
| JP2009503560A | Japan | A | |
| US2009048833A1 | United States of America | A1 | |
| EP1787287A4 | European Patent Office (EPO) | A4 | |
| US7584103B2 | United States of America | B2 | |
| EP1908055B1 | European Patent Office (EPO) | B1 | |
| AT454691T | Austria | T | |
| ATE454691T1 | Austria | T1 | |
| DE602006011622D1 | Germany | D1 | |
| US2010299135A1 | United States of America | A1 | |
| US7844464B2 | United States of America | B2 | |
| US2010318347A1 | United States of America | A1 | |
| JP4940139B2 | Japan | B2 | |
| EP1787288B1 | European Patent Office (EPO) | B1 | |
| DK1787288T3 | Denmark | T3 | |
| US8335688B2 | United States of America | B2 | |
| PL1787288T3 | Poland | T3 | |
| ES2394726T3This record | Spain | T3 | |
| US8412521B2 | United States of America | B2 | |
| US2013103400A1 | United States of America | A1 | |
| US2013166297A1 | United States of America | A1 | |
| JP5284785B2 | Japan | B2 | |
| US2013304453A9 | United States of America | A9 | |
| US8694312B2 | United States of America | B2 | |
| US8731920B2 | United States of America | B2 | |
| US8768706B2 | United States of America | B2 | |
| US2014249818A1 | United States of America | A1 | |
| US2014309995A1 | United States of America | A1 | |
| US2014343939A1 | United States of America | A1 | |
| CA2577721C | Canada | C | |
| CA2577726C | Canada | C | |
| US9135917B2 | United States of America | B2 | |
| US9190050B2 | United States of America | B2 | |
| US2016005402A1 | United States of America | A1 | |
| US9286896B2 | United States of America | B2 | |
| US2016078861A1 | United States of America | A1 | |
| US2016196821A1 | United States of America | A1 | |
| US9454965B2 | United States of America | B2 | |
| EP1787287B1 | European Patent Office (EPO) | B1 | |
| US9520124B2 | United States of America | B2 | |
| US9552809B2 | United States of America | B2 |
Numbers
- Publication
- 2394726
- Publication, DOCDB
- 2394726
- Publication, EPODOC
- ES2394726T
- Application
- 5789851
- Application, DOCDB
- 05789851
- Application, EPODOC
- ES20050789851T
Titles2
- Spanish
- Extracción automática de contenido semántico y generación de un documento estructurado a partir del habla
- English
- Automatic extraction of semantic content and generation of a structured document from speech
Classification
- CPC, 2
- G10L15/1815
- G16H15/00
- IPC, 2
- G06F19 00
- G10L15 18