Untitled record
Abstract
This application discloses a method implemented of recognizing a keyword in a speech that includes a sequence of audio frames further including a current frame and a subsequent frame. A candidate keyword is determined for the current frame using a decoding network that includes keywords and filler words of multiple languages, and used to determine a confidence score for the audio frame sequence. A word option is also determined for the subsequent frame based on the decoding network, and when the candidate keyword and the word option are associated with two distinct types of languages, the confidence score of the audio frame sequence is updated at least based on a penalty factor associated with the two distinct types of languages. The audio frame sequence is then determined to include both the candidate keyword and the word option by evaluating the updated confidence score according to a keyword determination criterion.

Term
No projected expiry on record.
- Priority
- Filed
- Published
- Today
20 claims: 20 independent, 0 dependent
- 1protection items عناصر الحماية 1- A method for recognizing a keyword in a speech, comprising:receiving a sequence of audio frames comprising a current frame and a subsequent frame;Define the candidate keyword for the current frame using a predefined decoding network that includes 1- طريقة لمتعرف عمى كممة أساسية في خطاب keyword in a speech ، تشتمل عمى: استقبال متوالية من إطا ارت صوت receiving a sequence of audio frames تشتمل عمى إطار حالي current frame واطار الحق subsequent frame ؛ تحديد الكممة األساسية المرشحة لإلطار الحالي باستخدام شبكة فك شفرة decoding network محددة مسبقا تتضمن 5 Base quantifiers and filler words from multiple languages, audio frame sequence binding 5 كممات أساسية وكممات حشو filler words من لغات متعددة ، ربط متوالية إطار الصوت which is a confidence score associating the audio frame sequence التي يتم confidence score بدرجة الثقة associating the audio frame sequence Partially defined according to the candidate keyword;specifying the quantum option for the subsequent frame using the quantum candidate and the predefined decoding network;When a candidate keyword and an option word are associated with two distinct types of language;Confidence Score Update تحديدىا جزئيا وفقا لمكممة األساسية المرشحة؛ تحديد خيار الكممة لإلطار الالحق باستخدام الكممة األساسية المرشحة وشبكة فك التشفير decoding network المحددة مسبقاً؛ وعندما يتم ربط الكممة األساسية المرشحة وخيار الكممة باثنين من األنواع المتميزة لمغات؛ تحديث درجة الثقة 10 confidence score for the sequence of the sound frame based on the part parameter which is predetermined according to the two distinct types of language;mute choice and vocal model for the subsequent frame;Determining the sequence of the sound frame includes both the candidate keyword and the word choice by evaluating the confidence score that has been updated according to the criterion for determining a keyword. 10 confidence score الخاصة بمتوالية إطار الصوت بناءً عمى معامل الج ازء الذي يتم تحديده مسبقا وفقا الثنين من األنواع المتميزة لمغات؛ خيار الكممة والنموذج الصوتي لإلطار الالحق؛ ويشتمل تحديد متوالية إطار الصوت عمى كل من الكممة األساسية المرشحة وخيار الكممة بواسطة تقييم درجة الثقة confidence score التي تم تحديثيا وفقا لمعيار تحديد كممة أساسية.
- 215 2- الطريقة وفقاً لعنصر الحماية رقم 1، حيث يتم تحديد مجموعة من الكممات األساسية المرشحة، 15th 2- The method according to claim #1, where a set of candidate key quantifiers is identified, comprising the candidate keyword, for the current frame of the audio frame sequence, and each candidate keyword is associated with at least one keyword option, where a subset of the candidate keyword is selected to be included in the audio frame sequence with At least one word choice of its own التي تشتمل عمى الكممة األساسية المرشحة ، لإلطار الحالي لمتوالية اإلطار الصوتي audio frame sequence ، ويتم ربط كل من الكممة األساسية المرشحة مع خيار الكممة واحد عمى األقل، وحيث يتم تحديد مجموعة فرعية من الكممة األساسية المرشحة ليتم إد ارجيا في متوالية اإلطار الصوتي audio frame sequence مع خيار الكممة الواحد عمى األقل الخاص بيا ذي 20 Relevance based on the criterion of identifying a keyword. 20 الصمة بناءً عمى معيار تحديد كممة أساسية.
- 33- The method according to element of protection No. 2, where the subsequent frame is the last frame of the audio frame sequence, and according to the criterion for specifying a key word, the candidate key word associated with the confidence score is chosen from a set of candidate key words such as the associated key word In the current frame of the audio frame sequence 3- الطريقة وفقا لعنصر الحماية رقم 2، حيث يتمثل اإلطار الالحق في اإلطار األخير لمتوالية اإلطار الصوتي audio frame sequence، ووفقا لمعيار تحديد كممة أساسية ، يتم اختيار الكممة األساسية المرشحة المرتبطة بدرجة الثقة confidence score المفضمة من مجموعة من الكممات األساسية المرشحة مثل كممة أساسية مرتبطة باإلطار الحالي لمتوالية اإلطار الصوتي . audio frame sequence 25 . audio frame sequence 25 ٤٥٤٣ ٤٥٤٣ -٣٤- -٣٤-
- 44- The method according to claim number 2, where, according to the criterion for determining a key word, each set of candidate key words is linked with a confidence score that is relevant to it. 4- الطريقة وفقا لعنصر الحماية رقم 2، حيث أنو وفقا لمعيار تحديد كممة أساسية ، يتم ربط كل من مجموعة من الكممات األساسية المرشحة مع درجة الثقة confidence score ذات الصمة بيا For the audio frame sequence, the confidence score of 5 is greater than the threshold value of the base quantity. لمتوالية اإلطار الصوتي audio frame sequence ، وتكون درجة الثقة confidence score 5 ذات الصمة أكبر من القيمة الحدية threshold value لمكممة األساسية.
- 55- The method according to claim number 2, where after selecting the candidate keyword subset to be included in the audio frame sequence with at least one word choice of its own, the corresponding confidence score is updated 5- الطريقة وفقاً لعنصر الحماية رقم 2، حيث أنو بعد تحديد المجموعة الفرعية لمكممات األساسية المرشحة ليتم إد ارجيا في متوالية اإلطار الصوتي audio frame sequence مع خيار الكممة الواحد عمى األقل الخاص بيا ذي الصمة، يتم تحديث درجة الثقة confidence score المقابمة 10 It is determined to exceed a threshold value for a basic word according to the criterion for determining a basic word. 10 ويتم تحديدىا لتتجاوز قيمة حدية لمكممة األساسية وفقاً لمعيار تحديد كممة أساسية.
- 66- The method according to protection element No. 1, where according to the criterion for determining a key quantity, the confidence score of the audio frame sequence is greater than the threshold value of the base quantity. 6- الطريقة وفقا لعنصر الحماية رقم 1، حيث أنو وفقا لمعيار تحديد كممة أساسية ، تكون درجة الثقة confidence score لمتوالية اإلطار الصوتي audio frame sequence أكبر من القيمة الحدية threshold value لمكممة األساسية. 15 15
- 77- The method according to protection element No. 1, where the pre-defined decoding network is linked to two or more English, Chinese, Japanese, Russian, French, German and the like, and includes a subset of basic words and a subset of words filler words for each of two or more languages. 7- الطريقة وفقا لعنصر الحماية رقم 1، حيث يتم ربط شبكة فك التشفير decoding network المحددة مسبقاً باثنين أو أكثر من المغات اإلنجميزية، الصينية، اليابانية، الروسية، الفرنسية، األلمانية وما شابو ذلك، وتشتمل عمى مجموعة فرعية من الكممات األساسية ومجموعة فرعية من كممات الحشو filler words لكل من اثنين أو أكثر من المغات. 20 20
- 88- The method according to claim No. 1, where each keyword from a pre-defined decoding network includes one or more triplex speakers. 8- الطريقة وفقا لعنصر الحماية رقم 1، حيث تشتمل كل كممة أساسية من شبكة فك التشفير decoding network المحددة مسبقاً عمى واحدة أو أكثر من سماعات ثالثية.
- 99- The method according to claim #1, which is according to the decoding structure 9- الطريقة وفقا لعنصر الحماية رقم 1، حيث أنو وفقا لييكل فك التشفير decoding 25 structure of a predefined decoding network, each keyword in a predefined decoding network is associated with at least one keyword used with 25 structure لشبكة فك التشفير decoding network المحددة مسبقا، يتم ربط كل كممة أساسية في شبكة فك التشفير decoding network المحددة مسبقا بكممة واحدة عمى األقل تستخدم مع ٤٥٤٣ ٤٥٤٣ -٣٥- -٣٥- The imprint base quantum in a real speech is argy in the decoding . network الكممة األساسية ذات الصمة في خطاب حقيقي واد ارجيا في في شبكة فك التشفير decoding . network . network
- 1010- The method according to claim No. 9, which is according to the decoding code 10- الطريقة وفقا لعنصر الحماية رقم 9، حيث أنو وفقا لييكل فك التشفير decoding 5 structure of a predefined decoding network, producing each keyword in a subset of keywords and at least one related word that is used with the related keyword from two different languages. 5 structure لشبكة فك التشفير decoding network المحددة مسبقا، تنتج كل كممة أساسية في مجموعة فرعية من الكممات األساسية والكممة الواحدة عمى األقل ذات الصمة التي تستخدم مع الكممة األساسية ذات الصمة من اثنين من المغات المختمفة.
- 1111- The method in accordance with claim No. 1 also includes:11- الطريقة وفقا لعنصر الحماية رقم 1، تشتمل أيضا عمى: 10 Create a part coefficient table to include a set of part factors, each of which is linked to two different languages, where the part coefficient used to update the confidence score for the audio frame sequence is determined by extracting the part coefficient table based on two distinct types From the different languages of the candidate basic quantifier and the option of the quantum. 10 إنشاء جدول معامل الج ازء ليشتمل عمى مجموعة من عوامل الج ازء يتم ربط كل منيا باثنين من المغات المختمفة، حيث يتم تحديد معامل الج ازء المستخدم لتحديث درجة الثقة confidence score لمتوالية اإلطار الصوتي audio frame sequence بواسطة استخ ارج جدول معامل الج ازء بناءً عمى نوعين متميزين من المغات المختمفة لمكممة األساسية المرشحة وخيار الكممة.
- 1215 12- الطريقة وفقا لعنصر الحماية رقم 1، تشتمل أيضا عمى:15th 12- The method in accordance with claim No. 1 also includes: Creating a predefined decoder network, in which base words and filler words from multiple languages are grouped according to their language types, also includes: إنشاء شبكة فك تشفير محددة مسبقاً، حيث يتم تجميع الكممات األساسية وكممات الحشو filler words من لغات متعددة وفقاً ألنواع المغات الخاصة بيا، يشتمل أيضاً عمى: Create a start node and an end node;Create a set of language nodes, each representing a type of language, binding each language node to a start node ;خمق عقدة بداية start node وعقدة نياية end node ؛ خمق مجموعة من عقد المغة language nodes يمثل كل منيا نوع من المغة، ربط كل عقدة لغة بعقدة بداية start node ؛ 20 Associate each language node with a subset of related keywords and a subset of related filler words as both from the corresponding language;For each base word, converting the keyword of the triphone sequences, creating a three-phoneme node for each triphone of the triphone sequences of the triphone sequence, connecting the speaker node 20 ربط كل عقدة لغة بمجموعة فرعية من الكممات األساسية ذات الصمة ومجموعة فرعية من كممات الحشو filler words ذات الصمة كمتاىما من المغة المقابمة؛ لكل كممة أساسية، تحويل الكممة األساسية converting the keyword ذات الصمة لمتوالية من السماعات الثالثية triphone sequences ، إنشاء عقدة سماعة ثالثية ذات صمة لكل سماعة ثالثية من متوالية السماعات الثالثية triphone sequences من الكممة األساسية ذات الصمة، ربط عقد السماعة الثالثية 25 of triphone sequences together to form a triphone node sequence including a master treble node and an edematous triphone node, linking the treble node 25 من متوالية السماعات الثالثية triphone sequences معا لتشكيل متوالية عقد سماعة ثالثية بما في ذلك عقدة سماعة ثالثية رئيسية وعقدة سماعة ثالثية ذيمية، ربط عقدة السماعة الثالثية ٤٥٤٣ ٤٥٤٣ -٣٦- -٣٦- The main sympathetic node has the parasympathetic node and the edematous tertiary speaker node has an end node. For each stuffing word, an emollient padding node is created, and the 'immunopositive' padding node is created between the corresponding idiom node and the end node;And connect a start node and end node الرئيسية ذات الصمة بعقدة المغة المقابمة وعقدة السماعة الثالثية الذيمية ذات الصمة بعقدة نياية end node ؛ لكل كممة حشو، يتم إنشاء عقدة حشو ذات صمة واق ارن عقدة الحشو ذات الصمة بين عقدة المغة المقابمة وعقدة نياية end node ؛ وربط عقدة بداية start node وعقدة نياية . end node . end node 5 5
- 1313- The method according to protection element No. 12, where the candidate keyword and the word option are selected to be associated with two distinct types of languages, where one of a set of language nodes is linked between the candidate keyword and the word option on the pre-defined decoding network . 13- الطريقة وفقا لعنصر الحماية رقم 12، حيث يتم تحديد الكممة األساسية المرشحة وخيار الكممة ليتم ربطيا باثنين من األنواع المميزة لمغات، حيث يتم ربط واحدة من مجموعة من عقد المغة language nodes بين الكممة األساسية المرشحة وخيار الكممة عمى شبكة فك التشفير decoding network المحددة مسبقا. 10 10
- 1414- The method according to protection element No. 12, whereby according to the decoding structure of the predefined decoding network, each keyword in the decoding network is connected to the predefined decoding network with at least one word that is used With the relevant main word in verbal speech. 14- الطريقة وفقا لعنصر الحماية رقم 12، حيث أنو وفقا ل ىيكل فك التشفير decoding structure لشبكة فك التشفير decoding network المحددة مسبقا، يتم ربط كل كممة أساسية في شبكة فك التشفير decoding network عمى شبكة فك التشفير decoding network المحددة مسبقاً بكممة واحدة عمى األقل تستخدم مع الكممة األساسية ذات الصمة في خطاب فعمي. 15 15
- 1515- An electronic badge that contains:15- وسيمة إلكترونية ، تحتوي عمى: one or more processors;A memory has generalizations stored in it, which when executed by one or more processors cause the processors to perform operations that include: واحد أو أكثر من المعالجات؛ وذاكرة بيا تعميمات مخزنة عمييا، والتي عند تنفيذىا بواسطة واحد أو أكثر من المعالجات تجعل المعالجات تقوم بإج ارء عمميات تشتمل عمى: Receiving a sequence of sound frames comprising a current frame and a subsequent frame استقبال متوالية من إطا ارت صوت تشتمل عمى إطار حالي current frame واطار الحق 20 after frame follows the current frame;Define a keyword candidate for the current framework using a predefined decoding network that includes keywords and filler words from multiple languages;associating the audio frame sequence with a confidence degree partially determined by the candidate keyword;Selecting a quantum option for the subsequent frame using the main quantum filter and the decoding network 20 subsequent frame يتبع اإلطار الحالي؛ تحديد كممة أساسية مرشحة لإلطار الحالي باستخدام شبكة فك شفةر decoding network محددة مسبقاً تشمل عمى الكممات األساسية وكممات حشو filler words من لغات متعددة ؛ ربط متوالية إطار الصوت associating the audio frame sequence بدرجة ثقة يتم تحديدىا جزئيا وفقا لمكممة األساسية المرشحة؛ تحديد خيار كممة لإلطار الالحق باستخدام الكممة األساسية المرشحة وشبكة فك شفرة decoding network 25 Predefined;When a candidate base word and an option word are associated with two distinct types of language, . is done 25 محددة مسبقاً؛ عند ربط الكممة األساسية المرشحة وخيار الكممة بنوعين متميزين من المغات، يتم Update the confidence score of the sound frame sequence based on a section coefficient تحديث درجة الثقة confidence score الخاصة بمتوالية إطار الصوت بناءً عمى معامل جازء ٤٥٤٣ ٤٥٤٣ -٣٧- -٣٧- a pre-set penalty factor according to the two distinct types of language, the choice of punch and audio model for the later frame;Determine that the audio frame sequence includes both the candidate keyword and the word choice by evaluating the confidence score that was updated according to a key word criterion. penalty factor محدد مسبقا وفقا الثنتين من األنواع المتميزة لمغات، خيار الكممة ونموذج سمعي لإلطار الالحق ؛ وتحديد أن متوالية اإلطار الصوتي audio frame sequence تشتمل عمى كل من الكممة األساسية المرشحة وخيار الكممة من خالل تقييم درجة الثقة confidence score التي تم تحديثيا وفقا لمعيار تحديد كممة أساسية.
- 1616- The electronic tag in accordance with protection element No. 15, where, according to the criterion for determining a base quantity, the confidence score of the audio frame sequence is greater than the threshold value of the key quantity. 16- الوسيمة اإللكترونية وفقا لعنصر الحماية رقم 15، حيث أنو وفقا لمعيار تحديد كممة أساسية ، تكون درجة الثقة confidence score لمتوالية اإلطار الصوتي audio frame sequence أكبر من القيمة الحدية threshold value لمكممة األساسية.
- 1710 17- The electronic merchandise in accordance with the element of protection No. 15, which includes the operations that are performed 10 17- الوسيمة اإللكترونية وفقا لعنصر الحماية رقم 15، حيث تشتمل العمميات التي يتم إج ارؤىا By wizards also blindness:بواسطة المعالجات أيضا عمى: Create a part coefficient table to include a set of part factors, each of which is linked to two different languages, where the part coefficient used to update the confidence score for the audio frame sequence is determined by extracting the 15 part coefficient table based on two types Distinguished from the different languages of the candidate basic quantity and the option of the word. إنشاء جدول معامل الج ازء ليشتمل عمى مجموعة من عوامل الج ازء يتم ربط كل منيا باثنين من المغات المختمفة، حيث يتم تحديد معامل الج ازء المستخدم لتحديث درجة الثقة confidence score لمتوالية اإلطار الصوتي audio frame sequence بواسطة استخ ارج جدول معامل 15 الج ازء بناءً عمى نوعين متميزين من المغات المختمفة لمكممة األساسية المرشحة وخيار الكممة.
- 1818- The electronic medium according to the element of protection No. 15, where the pre-defined decoding network is linked to two or more English, Chinese, Japanese, Russian, French, German and the like, and includes a subset of basic words 18- الوسيمة اإللكترونية وفقا لعنصر الحماية رقم 15، حيث يتم ربط شبكة فك التشفير decoding network المحددة مسبقاً باثنين أو أكثر من المغات اإلنجميزية، الصينية، اليابانية، الروسية، الفرنسية، األلمانية وما شابو ذلك، وتشتمل عمى مجموعة فرعية من الكممات األساسية 20 A subset of filler words for each of two or more languages. 20 ومجموعة فرعية من كممات الحشو filler words لكل من اثنين أو أكثر من المغات.
- 1919- A storage medium that can be read by a non-transitory computer readable medium, with stored circulars on which, when executed by one or more processors, the processors perform operations that include:19- وسط تخزين يمكن ق ارءتو بواسطة حاسوب غير انتقالي -non-transitory computer readable medium، بو تعميمات مخزنة عميو والتي عند تنفيذىا بواسطة واحد أو أكثر من المعالجات تجعل المعالجات تقوم بإج ارء عمميات تشتمل عمى: 25 receiving a sequence of audio frames comprising a current frame and a subsequent frame that follows the current frame;25 استقبال متوالية من إطا ارت صوت receiving a sequence of audio frames تشتمل عمى إطار حالي current frame واطار الحق subsequent frame يتبع اإلطار الحالي؛ ٤٥٤٣ ٤٥٤٣ -٣٨- -٣٨- Define a keyword candidate for the current framework using a predefined decoding network that includes keywords and filler words from multiple languages;associating the audio frame sequence with a confidence degree partially determined by the candidate keyword;Selecting a muzzle option for the subsequent frame using the main word تحديد كممة أساسية مرشحة لإلطار الحالي باستخدام شبكة فك شفرة decoding network محددة مسبقا تشمل عمى الكممات األساسية وكممات حشو filler words من لغات متعددة ؛ ربط متوالية إطار الصوت associating the audio frame sequence بدرجة ثقة يتم تحديدىا جزئيا وفقا لمكممة األساسية المرشحة؛ تحديد خيار كممة لإلطار الالحق باستخدام الكممة األساسية 5 The filter and decoding network is predefined;When a candidate keyword and word choice are associated with two distinct types of language, the confidence score of the audio frame sequence is updated based on a pre-determined penalty factor according to the two distinct language types, the word choice and an audio model of the subsequent frame;Determine that the audio frame sequence includes each of the candidate keywords 5 المرشحة وشبكة فك شفرة decoding network محددة مسبقاً؛ عند ربط الكممة األساسية المرشحة وخيار الكممة بنوعين متميزين من المغات، يتم تحديث درجة الثقة confidence score الخاصة بمتوالية إطار الصوت بناءً عمى معامل جازء penalty factor محدد مسبقا وفقا الثنتين من األنواع المتميزة لمغات، خيار الكممة ونموذج سمعي لإلطار الالحق ؛ وتحديد أن متوالية اإلطار الصوتي audio frame sequence تشتمل عمى كل من الكممة األساسية المرشحة 10 And the choice of the word by evaluating the degree of confidence that has been updated according to the criterion for determining a basic word. 10 وخيار الكممة من خالل تقييم درجة الثقة التي تم تحديثيا وفقا لمعيار تحديد كممة أساسية.
- 2020- A storage medium that can be read by a non-transitory computer readable medium in accordance with the element of protection No. 19, where the operations that are performed by processors also include:20- وسط تخزين يمكن ق ارءتو بواسطة حاسوب غير انتقالي -non-transitory computer readable medium وفقا لعنصر الحماية رقم 19، حيث تشتمل العمميات التي يتم إج ارؤىا بواسطة المعالجات processors أيضا عمى: 15 إنشاء جدول معامل الج ازء ليشتمل عمى مجموعة من عوامل الج ازء يتم ربط كل منيا باثنين من المغات المختمفة، حيث يتم تحديد معامل الج ازء المستخدم لتحديث درجة الثقة confidence score لمتوالية اإلطار الصوتي audio frame sequence بواسطة استخ ارج جدول معامل الج ازء بناءً عمى نوعين متميزين من المغات المختمفة لمكممة األساسية المرشحة وخيار الكممة. 15th Create a part coefficient table to include a set of part factors, each of which is linked to two different languages, where the part coefficient used to update the confidence score for the audio frame sequence is determined by extracting the part coefficient table based on two distinct types From the different languages of the candidate basic quantifier and the option of the quantum. ٤٥٤٣ ٤٥٤٣ -٣٩- -٣٩- Update the degree of stress تحديثدرجةأتثدة later audio frame اطار صوتي لاحق F6F2F3F;F5F6F7F8F9F10F11F12 Fn F٦F٢F٣F؛F٥F٦F٧F٨F٩F١٠F١١F١٢ Fn audio frame اطار صوتي حاتي barracks ثكن ا ٤٥٤٣ ٤٥٤٣
Independent claims20
355 paragraphs, as filed
full description
hidden invention
The applications that they disclose generally relate to speech recognition, and in particular, to the detection of key words in speech data in more than one language.
In automatic speech recognition (ASR), the word . is
<p>5 The basic word in a word is associated with a specific objective meaning, and is represented in a stereotypical way by a noun or phrase. On the contrary, the word filling follows the basic words in a natural way and does not assume a major and ambiguous role..</p>
The key word is revealed when the start and end time points of the base word are specified in the data of a speech received by an electronic tag. As a result of the discovery of the main quantities, the speech data is determined by the system of the disclosure of the basic words to include the usual basic quantities
<p>10 And filler words. The current key word detection systems are mainly implemented based on two models, a non-vocabulary data model and a syllable/voice recognition model.</p>
In the key word detection system based on the unacceptable data model, a decoding network is used to identify the key words in the received speech data.
<p>15th To frame the pre-selected network. According to the decoding network, the key word detection system identifies each part (for example, a frame) of speech data as being associated with the base word or filler word. The basic quantifiers are the relevant culture temperature to determine if the detection of the quantitative is done correctly.</p>
<p>20 With information about their status within the speech data.</p>
٤٥٤٣
-٣-
On the other hand, the key word detection system based on the voice/syllable recognition model detects the key words in the received speech data on the basis of the entire context of the speech data. In response to the challenges, it is necessary to take out an audio network or a verbal segment of the received speech data, and the basic words of the speech data are revealed from the network.
<p>5 The sound or the syllable using context search technology.</p>
When more than one language and one share the knowledge of speech, the discovery systems of the current terminology usually require two separate phases, the phase of acquaintance with the languages and the phase of discovery of the basic words. During the language recognition phase, a specific language is selected for the speech data to be received, and during the subsequent key word detection phase, this is again determined by the key words engine.
<p>10 The explorer reveals the basic words associated with this particular purpose. Then it is combined between the basic quantities that were detected and externally as a result of identification from the detection system of basic words.</p>
Nevertheless, the performance of the detection system is impeded in the current basic word, which includes two or more languages, often by means of the phase of acquaintance, which is suppressed. The impact of the accuracy of acquaintance on the meanings during the acquaintance phase on the ambiguity in a direct manner on the results of the scout for the basic term
<p>15th The scout phase includes the basic equipment. In terms of specifics, accurate identification of omissions generally requires speech data that lasts for an extended length (for example, from 3 to 5 seconds), and this condition inevitably encounters some obstacles to the flow of the key word to reveal the key word later. The main vocabularies present are an active part, and in particular, when nouns from multiple languages are mixed together in a single sentence (for example, in linked speech statements).</p>
<p>20 baa “昨晚的演唱会high不high”), thus making an inaccurate identification of languages and basic idioms. Moreover, there is a need for precise definitions of basic idioms, that is, a discourse that contains two or more words.</p>
٤٥٤٣
-٤-
General description of the invention
The above-mentioned and other deficiencies are not implemented or the problems associated with traditional methods of network connection are eliminated or eliminated by the application that is disclosed below. In some embodiments,
<p>5 The application is in an electronic medium that contains one or more processors, memory, and one or more modules, flashes, or groups of instructions stored in memory to perform multiple functions. It is possible to include instructions for performing these functions in any computer program whose functions are complete for implementation by one or more processors.</p>
One aspect of the requisition is a method to be implemented on an electronic basis for the knowledge of my work as a basic concept. The method includes a number of successive receptions, according to the current framework.
frame and subsequent frame subsequent frame Follows the current frame, specifying a candidate keyword for the current frame using a pre-defined decoder grid including the base and filler quantities in multiple languages. The method also includes interlinking the sequence of the sound frame with a degree of confidence that is partially determined according to the candidate keyword, and selecting a word choice for the subsequent frame using the basic word.
<p>15th The filter and the FAC network has a previously specified code. When the filtered basic word association and the choice of language with two distinct types of language, the degree of confidence of the sound frame sequence is updated based on a pre-determined fraction coefficient according to the two distinct types of language, the word choice and audio models for the append frame. The method also includes a challenge that the audio frame sequence includes both the candidate basic word and the word choice by evaluating the degree of confidence that has been updated according to the criterion for determining a key word.</p>
<p>20 Another aspect of the requirement is an electronic medium comprising one or more processors and their memory with at least one program (including instructions) stored on them, which when executed by one or more processors that make the processors resist by executing the basic word-challenging operations. The single program of the above-mentioned stores includes no more than two circulars that make the electronic media carry out operations in the manner described above.</p>
٤٥٤٣
-٥-
Another aspect of the request is a storage medium that can be read by a non-transitional computer that stores at least one program designed to be executed by at least one processor from an electronic medium. The single program shall include at least circulars that make the e-branding perform the operations mentioned in the method described above.
<p>5 Models and examples can be clarified. These skills are in no way in the light of the descriptions and drawings contained in this specification.</p>
Brief explanation of drawings
You are clearly aware of the aforementioned applications, as well as other applications. As a result of the following detailed description of several aspects of the selection, the most important objective is taken into account in relation to communications. 10 Raise similar reference numbers into corresponding parts through multiple projections of the drawings.
Figure 1 illustrates an exemplary speech data that includes both sets of audio frames according to some demand models.
Figures 2 illustrates an ideal network of code that includes both basic components and applica- tions in multiple languages according to some application models.
<p>15th Figure 3 illustrates a flowchart representing a method for defining my work as a basic word in a letter i according to</p>
for some models.
Figure 4 shows another ideal decoder network. According to some requisition models.
Figure 5 illustrates a flowchart representing a method for the detection of a basic concept according to some application models.
<p>20 Figure 6 shows an exemplary decoder network comprising base words and filler words from multiple languages according to some application models.</p>
٤٥٤٣
-٦-
Figure 7a shows a box diagram of an electronic tag that detects basic quantities in a letter according to some request models.
Figure 7b illustrates a box chart for a typical unit of knowledge of any letter and electronic message shown in Figure 7a according to some application forms.
5 Similar reference numbers refer to adjacent sections from multiple plot points.
Description:
Reference is made now to the forms in detail, the eights which are clarified in the attached drawings. In the following detailed description, several specific details are explained to provide a complete understanding of the subject matter provided in this document. However, it is clear to the savvy people that the issues can be implemented without these details.
<p>10 Al-Muhaaddah. So, there are other cases. Why do you describe methods, procedures, components, and drugs that are well known in detail so as not to cause ambiguity in the aspects of the models.</p>
To clarify the objectives, the solution and the technical advantages of the current requirement are much clearer, this request is also described below in detail with reference to the attached drawings.
Figure 1 shows the data of an exemplary speech 10 comprising a set of sound frames (F1,
<p>15th Fn, ..., F2) according to some student models. In some applications, when an electronic medium receives a DC from an audio signal, it converts the audio signal into digital to become digital audio data (for example, digital speech data of a human speech being received). These 10 speech data are then divided into groups of 10 sound frames. In one instance, each sound frame lasts for 10-20 milliseconds, and speech data 10 has 50-100 frames per second.</p>
<p>20 The speech data 10 that is originally received by the electronic medium contain basic words from one or more languages, and each pronunciation of the base word optional includes a sequence of frames. These split sound frames are then analyzed from the speech data 10 to identify these key words in the speech data 10.</p>
٤٥٤٣
-٧-
In different models of the request, the set of audio frames is processed successively according to a predefined decoder network, so that one or more basic words of speech data are recognized 10. In the previously specified decoder network, each base word must be defined in a competition with a limited number Not basic or gaskets. Therefore, when the current voice framework is determined, it is not related to a fundamental quantity.
<p>5 Specific, the context of the suffix and one of the suffix sounds should be equal to the basic nouns or the noun phrases associated with the identifier. A subset of this limited number may be chosen from the basic words or the filler words as complementary options for the audio frame to be attached based on a similar operation to the subsequent audio frame.</p>
In a specific example, the current sound frame associated with the base word, 'love', is pre-selected to follow
<p>10 "Ego", "food", "running", "sports", "nothing", and other groups of choices are other qualifiers.. In some ideal exemplary decoding networks, the number of words following the word 'love' is greater. However, when the affixed sound frame is received, the sound models are derived from this received voice frame, and helps to narrow the word choices to a subset of the selected word choices in a contest to follow the word “hab” according to the decoder network. fi</p>
15th In some models, the number of option words is reduced from more than 100 to less than 5.
During the recognition of the basic words of speech data 10, the degree of confidence is updated as a group of sound frames are linked in the following sequence to the basic words or the words of the corresponding texts. And the representation of the degree of this culture is equal. Similar to the basic complements and the data of the speech 10 that are fully acquainted. Uncle Sabil, for example, is the stubbornness of the current sound frame (example F5) with the basic sleeve and is not complete.
20 Process the subsequent audio frame (example F6) with a dimension, and the culture temperature is updated to indicate the same frames.
The voices before the subsequent voice frame (for example, F5-F1) and the corresponding words in the acquaintance are based on an option based on a network that avoided the corresponding symbol previously, as well as the other, as well as the other, F6. The relevant calculation is based on the silent sound model.
25 At a preferred cultural level, and therefore it is linked to the accompanying sound framework F6. What are the types of models, you are a challenge
٤٥٤٣
-٨-
Subgroups There are more quantitative options with the appended sound framework F6, when the relevant cultures meet the criteria for determining the main quantitative quantitative previously selected. It is also necessary to evaluate this sub-group of speech options, as more are known about the following sound frameworks, based on the construction of the transparent network. More details are explained. Try to treat the cultures below with reference to 5 Figures 2-6.
In some student models, the audio frames of speech data that are fully received include 10 words of two or more languages, and the basic words are determined through the method of recognition of the specific main words that accommodate the need for acquaintance with efficiency, or of two aspects. . In contrast to many previous technical solutions, various models of this Request 10 do not involve a lengthy run to identify an individual language type of discourse before identifying the key words in the discourse. Instead, the keyword recognition method in this application uses a previously defined decoder network that pairs words (for example, keywords and filler words) from different languages together, taking into account changes in the language types of fully-defined words On the other hand, while calculating the degree of culture and its conversations in turn is concerned with the recognition of the basic words associated with the speech data 10. Building 15 Accordingly, the method of recognizing the basic word improves the efficiency and accuracy of word recognition.
From multiple languages in speech.
Figure 2 illustrates an exemplary code-facing network 20 that includes both basic and appendices quantities in multiple languages (for example, languages 1 to K) according to some student models. The network also includes code-breaking 20 total words (K) 18 in total. 20 KW11 and a set of filler sleeves (for example, from Sabil, FW14-FW11) for each language (for example from Sabil
Objective 1). The number of basic nouns and the number of fill-in nouns are different for each language. In some models, the main word is a word that is linked to a specific objective meaning, while the filler qualifier is represented in a transitive sound or as a value quantifier in a verbal conversation that is only spoken. Examples of filler words include, for example, “um” and “ah”.
٤٥٤٣
-٩-
In some applications, the current phoneme frame (for example, frame F5) is defined to be associated with a base word (for example, KW13) of a particular language (for example, null 1). According to the decoder network 20, the word KW13 is linked from Optionally with a language 1 KW16, a language 1 FW12 filling mask, or a 2 language KW22 sleeve, a later muzzle is selected for the frame
5 F5 (for example, as linked to the frame F6) The basic word KW16, the filler FW12 and the KW22 key.
What are the specifications of these students, the basic terms associated with the current scriptural framework are sometimes shortened to a “candidate word”.
10 Both options for the subsequent opaque frame are linked to the degree of sound. In some models, the pitch of the sound is calculated based on the comparison between the phonetic models of the subsequent word frame and the corresponding word choice. And the bonding of the great voices is done by Msato. Blind Man-Chapping Between the phonetic models of the affixed muzzle frame and the corresponding muffler option. In some models, the degree of sound is used to talk about the degree of culture associated with the process of identifying these basic words. In one example, the degree of confidence is a degree
15th What are the sound levels of the sound frames that are complete for familiarization with them?
(Example, tire F6-F1).
Fi culicosis Alttbaiqaat, Anadma gets up mooring Khiaar Alkmmah) Amay SABEL Almthaal, KW22 (Bataar Alsaaot annexation Bnaua to Gaah Mtmiazh Aaan Ataar Alsaaot Alhaala, Iaatm Hjaam Drjah Althagaah Boisath officer Amaamaal Gaa Aze Kabbaal Thdithaao Bdrjaah Alsaaaot Almkabmaah to Khiaaar Alkmmaah. To Aazlk, Aaaatm Tmthiaal Drjaah Althagaah I love you
20 The word associated with the next sound frame is as follows:
(1) CS(Fs) CS(Fc) AS(WO)
(2) CS(Fs) CS(Fc)(Lc,Ls) AS(WO) or
where FC(CS) and CS(Fs) are the confidence scores associated with the present and subsequent audio frames, respectively; and AS(WO) is the pitch associated with the corresponding word choice. and δ(Lc, Ls)
٤٥٤٣
-١٠-
It is a term that deals with the part of the part associated with two distinct types of language for the basic components associated with the current market framework and the corresponding value option. Equation (1) is applied to the stubbornness of the basic quantum bond with the current sound frame and the corresponding word choice for the subsequent sound frame in the same language, while equation (2) is applied when they are linked to two different languages.
5 In some embodiments, the component coefficient is predetermined according to the predetermined part coefficient table. Table 1 represents a typical penalty factor table according to some models. The part coefficient table includes a group of part factors, each of which is linked to the two different variables, and the part coefficient used to update the degree of confidence of the audio frame sequence is determined by searching for the part coefficient table based on the two types of the completed language. Paired with a frame
10 Current audio and corresponding muzzle option. In some embodiments, the magnitude of the part coefficients indicates the relative probability of using two distinct language words together. For example, the odds are linked as a basic Chinese term with another Chinese term. Therefore, the operands in the first part of it 1 (not found in a part) are linked to situations in which it follows an option as the native language associated with a special frame. Man Nahia Akhar.
15th An English word, Japanese or Russian, and therefore, the part coefficients of 0.9, 0.7 or 0.5 are linked with a meaning whereby the English, Japanese or Russian word option is followed by the ordinal suffix suffix.
<tr><td><p>Russian</p></td><td><p>Japanese</p></td><td><p>English</p></td><td><p>tray</p></td><td></td></tr><tr><td><p>5.0</p></td><td><p>0.7</p></td><td><p>0.9</p></td><td><p>1</p></td><td><p>tray</p></td></tr><tr><td><p>0.7</p></td><td><p>0.8</p></td><td><p>1</p></td><td><p>0.9</p></td><td><p>English</p></td></tr><tr><td><p>0.5</p></td><td><p>1</p></td><td><p>0.7</p></td><td><p>0.7</p></td><td><p>Japanese</p></td></tr><tr><td><p>1</p></td><td><p>0.6</p></td><td><p>0.7</p></td><td><p>0.6</p></td><td><p>Russian</p></td></tr>
٤٥٤٣
-١١-
Table No. 1. Part coefficient table
The above description only shows several examples of the current disclosure to provide the principle and implementation of the current demand, and does not mean in any way whatsoever is the scope of the current demand. Any adjustments, equivalents, improvements, and the like that fall into the spirit and principle of the current demand should be included in the scope of the current demand.
5 After updating the degree of confidence with the part coefficients and sound degrees for quantitative options, one or more quantitative options for the auxiliary sound frame are chosen in accordance with the criterion for determining the basic quantifier specified earlier. In some models, according to the criterion for determining a basic word, the grammatical option associated with the highest degree of weight is selected as a basic word in conjunction with the suffix sound framework. After that, the process of acquaintance continues on the basic aspects of processing the following voice framework, for example, the F7 framework, following the voice framework
10 Later, until the good sound frame is processed in the speech data, for example, Fn, and is linked to a key word.
In some applications, according to the criterion for specifying a base quantifier, the confidence score should be greater than the threshold value . Moreover, in some cases, most of the word and monosyllabic options meet this criterion for defining a basic word, and therefore, they are linked to the voice framework F6. The acquaintance process continues
15th On the basic word of processing the following audio frame, for example, frame F7, based on most of the same and single options for frame F6. Removing some of the most one-of-a-kind quantum options for frame F6 can make the quantum options corresponding to frame F7 decrease the maximum value to the minimum value.
Therefore, in some applications, the distance between the current sound frame and the previous sound frame determines the recognition certainty of this previous sound frame. A significantly larger distance is associated with the selection of a key word for the corresponding previous audio frame, while on the other hand, a smaller distance can be linked to several options for sampling, which is capable of stifling a previous audio frame that races against the current audio frame.
٤٥٤٣
-١٢-
Figure 3 shows a process flow chart representing method 30 to identify a key word in a letter according to some demand models. Method 30 is optionally controlled by generalizations that are stored in a storage medium that can be read on a non-traffic computer device and executed by one or more host system processors. All the processes shown in Figure 3 can correspond to the instructions stored 5 in a computer memory or a storage medium that can be read by a portable computer. The storage medium that can be read by a blind computer may include a storage medium in the form of a magnetic disk or an optical disk, and storage means in the solid state such as a flash memory or a non-volatile memory stick or other means. . Instructions stored on a storage medium that can be read by a computer on one or more can include: source code, assembly 10 language code, object code, or another instruction format interpreted by one or more processors . Some of the processes in Method 30 can be combined and/or the order of some of the processes can be changed.
The 30 key word recognition method is implemented on an electronic medium that receives (32) sound frame sequences. The sound frame sequence includes a current frame and a subsequent frame that follows the current frame.
15th A basic filter (34) for the current framework is defined using a previously defined network, containing keywords and fillers from multiple languages.
In some embodiments, the previously specified decipher network is linked to two or more English, Chinese, Japanese, Russian, French, German and other languages, and includes a subset of the basic words and a subset of the acronyms of all two or more languages. In addition to that, in some models, both as a basic term for the FAC code network specified in a single competition, include one or more
More than treble headphones.
In some models, according to the structure of the FAC network a specified cipher, and each basic word is linked in the FAC network to a specified cipher with a single word, and the least is used with a basic expression with a stigma in a real speech and the writing is done externally. In some cases, according to the code of decoding the code, 25 both as a basic quantitative categorization are not included in a subgroup of the basic quantifiers, and as a categorization and one of the relevant categories.
٤٥٤٣
-١٣-
They are not used in any of the basic words of the same character, as there are two different languages. For example, the first word is represented by “說” in the Chinese language, and the second word is “no” in the English language according to a specific decoding structure.
After selecting the candidate word, the audio frame sequence (36) is linked to the confidence that was selected.
5 Partially according to the filter quantum. Then the electronic tag (38) resists by selecting an option quantum for the subsequent frame using the filter word and a decoder network specified in the competition. The stubbornness of the filter muzzle link and the chord option with two distinct types of objectives, the update of the punctuation grades is done according to the previous languages (40) The two are distinct types of language, muffler options, and phonetic models for the next frame.
10 As it is clarified above with reference to the figure 2, so in some forms, a table of the foreign operators is created in order to include a group of two foreign operators, as well as an accompanying statement. The acquaintance of the part coefficients used to update the culture degree of the phonetic frame sequence is done by searching for the part coefficient table based on two distinct language types for the candidate base word and the word choice. In some applications, the confidence level is adjusted for the sound frame sequence
<p>15th By means of the corresponding part coefficients, which are determined in a competition according to the two distinct types of objects. In some models, the degree of weight is a cumulative degree of weight, based on the sound degrees of the current framework and all frameworks before the current framework, and the confidence degree that was updated also includes the degree of soundness that was specified in the framework.</p>
After updating the confidence score of the audio frame sequence, the electronic tag (42) determines that the sequence
<p>20 The audio frame includes both the candidate keyword and the word choice by evaluating the degree of confidence that is complete in conversation according to the criteria for determining keywords. For any specific example, according to the criteria for determining basic quantities, the culture temperature of the acoustic frame sequence is greater than the minimum value of basic quantities.</p>
٤٥٤٣
-١٤-
In some models, a set of candidate keywords is defined for the current framework of the audio-frame sequence, and the above-mentioned candidate keyword is one of a set of candidate keywords. Each filtered base quantum is linked to one at least one quantum option according to a predefined decoded network. A first subset of the MAN filter principal quantifiers is excluded
<p>5 Key words recognized from the phonemic frame sequence, in particular because the relevant at least one word choices fail to meet the key word recognition criterion.</p>
On the other hand, a second sub-group of the candidate basic words is selected to be included in the phonetic frame sequence with both word choices and one of the words that are unique based on the criteria for determining the basic words. Amy and Jaw Al-Hadeed, the degree of the corresponding confidence level of the sound frame sequence
<p>10 On the criteria of defining the basic quantities of each of them as a candidate basic product in the first sub-group, what is the status of the nominated basic quantities and the choice of the word that is relevant to each of the above sections. I love you</p>
In the example, the second sub-group of the candidate keywords is selected to be included in the sequence of the audio frame, with the options of the word for each one of the words at least, the corresponding culture level is updated and the corresponding values are selected to be challenged according to the quantity values.
<p>15th basic. Someone skilled in the art recognizes that this determination of the second subset of candidate keywords is rather tentative, for when other word choices are made. In the decoder network for the subsequent frames of the next frame, the degree of confidence needs to be updated and can fail in the standard definition of key words.</p>
In addition to this, in some models, the attached framework represents the way to define the whole
<p>20 The above-mentioned key 30 In the good frame of the audio frame sequence, according to the key word definition criterion, the candidate keyword associated with the preferred degree of confidence, for example, the greatest confidence, is chosen from a set of candidate keywords as a key word associated with the current frame of the voice frame sequence .</p>
٤٥٤٣
-١٥-
In some models, according to the criterion for determining the basic quantities, all the groups of the filtered basic quantities are linked to the value of the significant culture so that the sequence of the sound frame has a larger value than the basic values.
In some applications of the current demand, the specific decoder network is designed according to the architecture
<p>5 Given decoding, base quantifiers and filler quantifiers from multiple languages are grouped according to their language types. The decoding structure includes a start node, an end node, language nodes, filler nodes, and a three-way speaker nodes for each nodes as a basic executor. Therefore, to create a previously defined network, which resists electronic media marketing by creating a primary node, an affiliate node, and a group, as long as the canceled contracts represent a type of non-compliance, and each link is linked to a network.</p>
<p>10 start start node . Likewise, each language node is linked to a subset of related basic words and a subset of related filler words, as well as from the corresponding language.</p>
For each basic quantifier, the corresponding primary quantum is converted to the three-speaker security sequence, and a three-speaker node is created for each triple-speaker from the relevant three-speaker sequence. Then, the treble nodes of the treble cascade are joined together to form a
<p>15th The sequence of a tertiary stethoscope node, including a main treble node and an edematous treble node, and the sympathetic treble node of the main treble speaker is attached to the opposite node and the sympathetic treble node to a terminal node. On the other hand, for each filler word, a node is created as a wadding muffler with a punctuation mark also between the opposite node and the end node. And both sleeves are tied between a beginning knot and an end knot.</p>
<p>20 In some embodiments, a keyword candidate and option keyword are specified to be associated with two distinct types of language, when one of a set of language nodes is bound between the keyword candidate and option keyword on a predefined decoder network.</p>
٤٥٤٣
-١٦-
In some models, according to the structure of the decoder network defined above, the connection of both basic words in any decoder network on the pre-defined decoder network is concerned with at least one word to be used together with the relevant keyword in a real speech.
It should be realized that the special arrangement in which the operations are described in Figure 3 are just perfect and not
<p>5 The purpose of it is to indicate that the intimidation is complete, and it is pure intimidation that cannot be carried out with the execution of operations. He gets acquainted with a person of normal skill, so he knows different ways to create automatic inbox messages, as it is explained in this document. In addition to that, refrain from mentioning the details of the other operations. Described in this document in relation to Method 200 (for example, Figure 5) is also applicable in a manner similar to Method 30 above in relation to Figure 3.</p>
10 For brevity, these details will not be repeated in this document.
Figure 4 illustrates a typical FAC chafer network. 100 according to some student models. As well as the structures in Figure 4, the difficulty of using the network of ciphers of 100 by the system of acquaintance with the basic words, based on the network of the rejected words, in the case of the previous fan, the need for syllable segments of the basic words also requires expansion in the context.
15th Do you have the final coding of the current state of HMM as a belief for the illustrator of the statement? Basic utterances are described using precise phonemic models, and HMM modeling is generally used for the context-dependent treble. Reference is made to the forms in the front forms. On the other hand, the filler syllable is a part that does not represent basic words in the speech syllable, and in general, the most coarse phonemic models are used, for example, the phonetic occlusion models after aggregation. Refer to these models
20 hidden models.
However, in existing multilingual word recognition technology, language recognition generally requires an audio signal of at least a certain length (eg, 3-5 seconds), which can result in some obstacles trying to prevent the recognition flow of the basic word In addition, the key word recognition technology in the case of the previous two thousand cannot handle the actual applications.
25 With a situation where multiple languages are scrambled.
٤٥٤٣
-١٧-
In some models of the current demand, a new solution is proposed to recognize the key word based on a framework that is also created based on the network of rejected words. When a decoding space is created for the proposed base word recognition solution, it is directly necessary to embed the decoded information in the decoder space, so as to avoid effectively affecting the detection flow in the language recognition stage; Phi crypto decoder man
<p>5 Current student models. It is possible to modify the strategic bases of passing codes using the null information, and the task of identifying the multilingual basic word can be completed using a single detection engine.</p>
Compared to the existing base word recognition system based on the non-acceptable word network, this technical solution proposed in the current application mainly involves two improvements: (1) the design of a multilingual decoding network that includes information from languages; and (2) the application of recognition algorithms
<p>10 algorithm on the multilingual basic words of a multilingual decoding network. In the decoding process, the token score is adjusted by the judges on the language information, and the part parameter is entered to convert the language.</p>
As I use the term “degree of students” and the attached protection elements, the term “symbolic degree” equals the term “degree of confidence”, and these two terms are used interchangeably.
<p>15th Figure 5 shows a flowchart representing the method of identifying the 200 basic words according to some student models. Controllers are executed in Method 200, optionally, by means of circulars that are stored in any storage media that can be read on a portable spare computer and executed by one or more processors of the host system. It is possible for both of the operations shown in Figure 5 to correspond to the instructions stored in the memory of a computer or storage media that can be read by a spare computer.</p>
<p>20 It may include storage media that can be read by a computer, whether it is storing magnetic or optical discs, and storage devices in the wrong cases, for example, memory, audio or other means. Instructions stored on a storage media that can be read by a computer on one or more can include: a source cipher, an assembly language cipher, an object cipher, or other formats that are interpreted by one or more processors. Some operations can be compared</p>
25 In method 200 and/or the order of some operations can be changed.
٤٥٤٣
-١٨-
What is the error 201, please create a network, do not provide information and include all information for the purpose. The basic quantifiers in the decoder network are grouped according to the disallowed information. In the process of producing a decoder network, it is possible to start a beginning contract and an end contract in this document, and the following steps are carried out for each information for the purpose of Li, where i is the language number:
<p>5 • Create a language node NLI, and create a side of a node to NLi.</p>
Load a list of the basic quantities and a list of the filler segment that matches the information of the Li language.
<p>• Executing the following procedures for each basic quantitative kj in the list of basic quantifiers, where j is the number of the basic quantitative:</p>
<p>• Convert the base word Kj to tertiary speaker sequences, and create a node for each treble speaker to form 10 node sequences. Then design aspects between node sequences; The design of the side of the closed node NLi to the first node in the node sequence and the side of the final node in the sequence of the node to the end node;</p>
<p>• Perform the following procedures for each filler section Fk in the list of filling section, where k is the expression of the number of the filler section:</p>
15th • Create an NFK node corresponding to the Fk padding segment.
<p>• Create a side of a language node NLi to an NFK and a side of an NFK to a Fian node;</p>
create an aspect of a start node to a start node; And
<p>• Create a decoding network.</p>
What is the error 202?
20 encryption. In some models, when the suppressed information of the fully disclosed base complements is inconsistent, the section parameter for the fully disclosed base complements is set. In some models that include a code propagation to reveal the keywords, when the language state node is checked, it is
٤٥٤٣
-١٩-
Determining whether the canceled information on any rule in the case of the canceled is in agreement with the information on the canceled symbols by comparison. If two pieces of information do not match, a break parameter is set for the symbol insertion.
Most preferred, the previous officer to schedule the work schedule of the section corresponding to the differences in the exemption category. What are some models, when the missing information about the basic quantities that the disclosures contain are other than that?
5 Consistently, the section parameter table is retrieved from the section factor table to determine the section parameter that has been adjusted for the detected basic quantifiers.
In step 203, the detected key words are evaluated based on the fraction coefficient.
In some models, the boundary value of the basic word can be preset in this document, and the weight of the basic words to be explored is calculated according to the culture algorithms and the part factor. The base quantifiers 10 are removed when the computed confidence is less than the boundary value of the base quantifier.
In some models, another parameter is added in relation to each of the sections of the file, so that it is possible to identify the basic words with greater ease, and thus improve the overall ratio. In addition, if some of the keywords are more important to the task of disclosure, a larger weighting factor can be given to the keywords, while a smaller weighting factor is given to the other keywords. The value officer is able to
15th Limiting the degree of code in the code passing process, thus increasing the decoding speed.
It should be realized that the special arrangement completed with the description of the operations in Figure No. 5 is only ideal, and the purpose is not to indicate that the above-mentioned arrangement is complete and clear, or the only one that can be implemented in the statements. He gets acquainted with a person of normal skill, so he knows different ways to create automatic inbox messages, as it is explained in this document. In addition to that, he left to mention the details of 20 other operations. described in this document in relation to Method 30 (for example, Figure 3)
It also applies in a manner similar to the aforementioned Paradigm 200 in relation to Figure 5. For brevity, because these details are repeated in this document.
٤٥٤٣
-٢٠-
Figure 6 shows another typical decoder network. 300 includes basic and filler quantifiers from multiple languages according to some application forms. It can be seen from Figure 6 that the base quantifiers and padding segments are grouped according to the language information in the decoding network.
The language case contract corresponding to the predicate vocabularies and filler syllables is added before both basic quantifiers 5 and each filler syllable. For example, the wording of the cover 1 corresponds to the basic words 11 to the 1n of the cover 1 , and the sections of the filler 11 to the section 1 m of the cover 1. For example, the word of the cover k corresponds to the number of the basic words k1 k to the sections of the filler k1 kn .k
In some of the processes of publishing ideal symbols, once the abolition status doctrine is achieved, it must be determined whether the information canceled from the doctrine agrees with the information canceled from the code by means of comparison, and if the size of the symbol does not apply to the extent of 10, two pieces of the code must be set.
In some applications, a 300 multi-language decoder network is created from the following steps:
In step one, the NStart node and the NEnd node are started;
In the second step, the multilingual list is checked according to the following sub-steps, 1-2 [-24] These sub-steps are performed sequentially for each Li language, sub-steps 2-3 [and 215-4] also include sub-steps 2-3 [-1 - 2 - 3-4] and Sub Errors 2-4-1] -
2-4-2], Amma Al-Tawali. Specifically, they include sub-errors 1-2 [-2-4] of the following:
<p>• 1-2] Establishment of the NStart NLi contract, and the establishment of the NStart NLi contract.</p>
<p>• 2-2] Download the list of main words and a list of fillers corresponding to the Li.</p>
<p>• 2-3] Perform the following procedures for each basic quantifier Kj in the list of basic quantifiers, which includes</p>
<p>20 Optionally uncle:</p>
<p>○ 2- 3- 1] Convert Kj into tertiary speaker sequences T1, T2, T2, TP,....</p>
<p>○ 2-3-2] Create a node for each speaker, where the node sequences are recorded in the N1 image,</p>
<p>NP,...,N2;</p>
٤٥٤٣
-٢١-
<p>○ 2- 3- 3] sequentially create sides from N1 to N2, from N2 to N3,..., and from 1-Np to</p>
.Np
<p>○ 2-3-4] Establishment of two sides of the Np contract to the N1 and the side of the Np to the contract</p>
Nend;
5 • 2-4] Perform the following actions on each filler segment Fj in the list of filler segment list, which includes
Optionally uncle:
<p>○ 2- 4-1] Create an NFj node corresponding to the padding segment Fj.</p>
<p>○ 2-2-4] Establishment of any part of the NLi contract to the NFj and the part of the NFj to any other contract</p>
Nend;
<p>10 In the third step, by returning the polynomial list checker, the designs of Janaab Man Nyaya NEnd are completed to the NStart node.</p>
In the fourth step, a tailored multilingual decoder network is created.
In some models, during the discovery of the base polymorphic word, the following steps can be performed sequentially, for example 1] The first step; [3] Statue of Sin 3; 4] Statue of Sin 4; 5 [figurines]
<p>15th step 5; and 6 [represents step 6. Furthermore, it can include [2] on Sub-Stream 1-2]; Sub-step 2.1 includes [2-1-1], 2-1-2, [2-1-3, 2-1-4]. is wanted</p>
Steps 1 [-6] are as follows:
<p>• 1[ Give the node an initial active token , where the degree to be initialized</p>
<p>1.</p>
<p>20 • 2] Read the data of the speech of the following frame, and implement the following steps, until all the data are processed</p>
the speech:
<p>• 1-2] Perform the following steps for each active Tk token, until all active tokens are processed:</p>
٤٥٤٣
-٢٢-
<p>○ 2.1.1] Move Tk from the current state node Si forward to the side of the muzzle grille, where</p>
A new node is set as Sj, and the new symbol is represented by Tp;
<p>○ 1-2-2] If Sj is a player in a language, update the score (Tp) of the TP code according to (Score(Tp)=δ(Lang(Tp),Lang(Si))×Score(Tk, where The equivalence of degree (Tk) is a cumulative statement of the inclusion of the sound model on all the node paths through which the symbol passes in a transport.</p>
From the start node to the Si node, then the continuity to 1-2-1] to keep the token moving forward; otherwise, procedure 2.1.3]; Where (·)Lang represents a function to obtain a node or information for a language for a symbol, and (·) is a statement about a part function, used to determine the punishment of a degree when converting from one language to another language. The value 1 is taken when the language information is consistent.
<p>10 ○ 2.1.3] Updating the TP code degree using voice models according to the error data</p>
current frame.
<p>○ 2.1.4] Judging whether the new token TP is active according to the audit strategy.</p>
<p>• 3] Active code registration for planar blind registration. In all active symbols connected to an end node such as</p>
.Tfinal
<p>15th • 4] Returning all the information obtained from the Tfinal track, and returning all the main words to me</p>
path.
<p>• 5] evaluate each key word detected using a confidence algorithm; And</p>
<p>• 6] Produce a list of the final key words that have been detected.</p>
In some embodiments, the part function (•) is referred to using a two-dimensional table (for example, 20 Table No. 1). In this case, the part function defines the part operators related to Chinese and English in four languages, Japanese and Russian.
٤٥٤٣
-٢٣-
In some forms, there are 3 activity codes (T3, T2, T1), so what time do you need specific codes, and the corresponding information is as follows:
<tr><td><p>Class</p></td><td><p>the language</p></td><td></td></tr><tr><td><p>50</p></td><td><p>Chinese</p></td><td><p>T1</p></td></tr><tr><td><p>30</p></td><td><p>Japanese</p></td><td><p>T2</p></td></tr><tr><td><p>20</p></td><td><p>Chinese</p></td><td><p>T3</p></td></tr>
Table No. 2
While searching the network, there are 5 nodes (S1, S2, S3, S4, S5) to which the code can be moved forward.
5 It worked on the network path, and the corresponding information was in the following manner (the degree of the final columns represented in the degree of the phonetic model of the speech data of the next frame on the node).
<tr><td><p>Class</p></td><td><p>the language</p></td><td><p>Type</p></td><td></td></tr><tr><td><p>0</p></td><td><p>English</p></td><td><p>language knot</p></td><td><p>S1</p></td></tr><tr><td><p>6</p></td><td><p>Chinese</p></td><td><p>normal knot</p></td><td><p>S2</p></td></tr><tr><td><p>0</p></td><td><p>Russian</p></td><td><p>language knot</p></td><td><p>S3</p></td></tr><tr><td><p>9</p></td><td><p>Japanese</p></td><td><p>normal knot</p></td><td><p>S4</p></td></tr>
٤٥٤٣
-٢٤-
<tr><td><p>5</p></td><td><p>Chinese</p></td><td><p>normal knot</p></td><td><p>S5</p></td></tr>
Table No. 3
A corresponding relationship is represented by the two symbols and cases in the following: T1 transports can be forwarded to S1 and S3, and T2 can be transferred to S4, and T3 can be transferred to S2 and S5, and when the mobile is complete, the corresponding nodes are as follows To a language node, it is necessary to transfer the symbol continuously, to complete the processing on the speech data.
<tr><td></td><td><p>S1</p></td><td><p>S2</p></td><td><p>S3</p></td><td><p>S4</p></td><td><p>S5</p></td></tr><tr><td><p>T1</p></td><td><p>T4</p></td><td></td><td><p>T5</p></td><td></td><td></td></tr><tr><td><p>T2</p></td><td></td><td></td><td></td><td><p>T6</p></td><td></td></tr><tr><td><p>T3</p></td><td></td><td><p>T7</p></td><td></td><td></td><td><p>T8</p></td></tr>
Table No. 4
The information for each symbol can be updated in the following way:
<tr><td><p>Class</p></td><td><p>the language</p></td><td></td></tr><tr><td><p>δ(Chinese, English) x score(T1) = 45</p></td><td><p>Chinese=< English</p></td><td><p>T4</p></td></tr><tr><td><p>δ(Chinese, Russian) x Score(T1) = 25</p></td><td><p>Chinese =< Russian</p></td><td><p>T5</p></td></tr><tr><td><p>Score(T2) + Score(S4) = 39</p></td><td><p>Japanese</p></td><td><p>T6</p></td></tr>
٤٥٤٣
-٢٥-
<tr><td><p>Score(T3) + Score(S2) = 26</p></td><td><p>Chinese</p></td><td><p>T7</p></td></tr><tr><td><p>Score(T3) + Score(S5) = 25</p></td><td><p>Chinese</p></td><td><p>T8</p></td></tr>
Table No. 5
The definition of the part function (•)δ can be formulated with reference to the historical paths of the encoder, for example, the number of Mart type language in the changes of the paths.
Although the foregoing describes a specific example of the part function in detail, people who are skilled in art can realize that this description is only exemplary, but it is not intended to limit the models of the current demand.
Based on the above detailed analysis, the current demand model also provides a key word detection system.
Figure 7a shows a box diagram of an electronic tag 70 that detects basic quantities in a discourse according to some request models. In some applications, the server system includes at least 14 clouds
10 Single or more than 710 processors (for example, semi-centralized) and memory 720 for storing data, parchments, and instructions to be executed by one or more 710 processors. In some applications, server system 14 also includes one or more 730 interfaces, communication, In/Out (740) O/I, one or more 750 communication buses communicate with these components.
In some embodiments, the 740 O /I interface has a 742 input unit and a 744 display unit.
15th Examples of the 742 input unit include a keyboard, mouse, touchpad, a wireless controller, a function key, a camera, a joystick, a microphone, a camera, and the like. In addition to that, presenter 744 contradicts the information that has been entered before the user or submitted to the user for reference. It includes an eight-and-a-half and a sharp-edged viewer 744, an example of a Siege-of-Amy, a liquid crystal display (LCD) and a dual-mode LED display.
20 It emits an organic light-emitting diode (OLED). In some applications,
٤٥٤٣
-٢٦-
The input point 742 and the display unit 744 on a touch-sensitive screen display a graphical user interface (GUI).
In some models, 530 communication buses include circuits (sometimes called chips) that connect and control communications between system components. In some models, these include
5 The 730 also interfaces with the 732 receiver and the 734 transmitter.
In some models, the memory includes 720 high-speed random access memory, for example DDR RAM, SRAM, DRAM, or other random access memory. in the solid state; Optionally including non-volatile memory, such as one or more magnetic disk storage devices, optical disk storage devices, flash memory devices, or other storage devices. non-volatile
10 in the solid state. In some embodiments, the memory 720 includes one or more storage devices that can be completed in one or more capacities of processors 710. In some models, the memory includes 720, or alternatively a 720 device(s) within non-volatile memory , is a storage media that can be read on a non-transferable computer.
In some forms, it resists memory 720, or alternatively, storage media that can be read by me, 15 spare computers, such as memory, 720, which stores the following columns, modules, generalizations, or subsets:
<p>• Operational System 701 includes procedures for handling various basic system servers and performing tasks that depend on devices.</p>
<p>• The O/I interface module module 702 which includes procedures for dealing with 20 different basic entry/exit permits through one or more entry/exit permits, where</p>
It also includes a sharp and straightforward interface 702 O / I, as well as a sharp and advanced display that controls the display and the graphical user interface.
<p>• A 703 communication module used to connect the server system 14 to other computing devices. (Example, service units and customer services), single or multiple jobs</p>
٤٥٤٣
-٢٧-
Communication network 750 (wireless or wireless), one or more communication networks, eg the Internet, other networks. Wide area, protected networks, metropolitan area networks, etc.
<p>• A 400 speech recognition module that recognizes a structured speech</p>
<p>5 On a pre-defined decoder network 704, including main and filler words, multiple languages and stored in memory 720; And</p>
<p>• Network decoding 704.</p>
Further details of the attempt and the sharpness of acquaintance with the above 400 letter are also explained by referring to Figures 1-6.
<p>10 Figure 7B illustrates a box chart for the standardization of knowledge of the letters of 400 people and electronic media shown in Figure 7A, according to some student models. As shown in Figure 5, the device includes a decoder network generation unit 501, a base word detection unit 502 and a word evaluation unit 503. The decoder network generation unit 501 is used to generate the decoding network that includes language information, and collect the basic words in a network. Decoding the encryption according to the information provided. And use and sharp scouts</p>
<p>15th About the 502 key words to reveal the key words in a speech entered using a decoder network, where, when the information that is excluded from the basic complements is complete and the disclosure about them is consistent, it is necessary to adjust the parameters of all the functions. And the assessment unit is used to evaluate the basic quantities 503, based on part of the transactions, and the basic quantities that contain the information about it.</p>
<p>20 In some embodiments, the 503 base quantity evaluation unit is used to preset the boundary value of the base quantity, and calculate the confidence of the base quantities fully disclosed according to the confidence algorithms and the weighting factor, where the base terms are removed when the calculated confidence is less than the base quantity's limit value.</p>
٤٥٤٣
-٢٨-
In some embodiments, a 501 decoder network generation unit is used to launch a start node and an end node, and perform the following steps for all Li language information, where i represents the language number:
• Create a language node NLi, and create a side of a start node to NLi.
Load a list of the basic quantities and a list of the filler segment that matches the information of the Li language.
<p>5 • Executing the following procedures for each basic quantitative quantitative Kj in the list of basic quantifications, where j is the expression</p>
About the basic quantum number:
<p>• convert the base word Kj into tertiary speaker sequences, and create a node for each treble speaker to form a node sequence; Next design aspects between node sequences; Design one side of the NLi node to the first node in the node sequence and one side of the final node in the node sequence to a final node;</p>
10 • Perform the following procedures for each tamping segment Fk in the existing tamping segment, where k is an expression
Filling clip number:
<p>• Create an NFK node corresponding to the Fk padding segment;</p>
<p>• Create a side of a language node NLi to an NFK and a side of an NFK to a Fian node;</p>
create an aspect of a start node to a start node; And
15th • Output the decoding network.
Fi Avaad Alnmaaazj, Tsaatkhaddm and Haadh Ckshaav Alkmmah Aolsasaaah 502, Vaaa Amamiaah Anchaaar Ahaaarh for Ckshaav Alkmmah Aolsasaaah, stubbornness Alossaol Alai Akaadh Halh to Gah, Thaadd, Palmqarnah, Maaa if Kanat Mamomaaat to Gaah Amay Akadh Halh Amoah Ttaaabak Mamomaat to Gaah Amay Aahaaarh, if Kanat Almuammaahumican Mttaaabaktin Parts , you set a partial coefficient for a symbol degree.
20 In some models, they are used and sharp as the basic quantity index 502 used in the previous controller for a partial coefficient table corresponding to the differences in the final classes; And when you accept information that is not complete
٤٥٤٣
-٢٩-
The complete base parameter detects matching variables, and retrieves the partial coefficient table to determine the set of the basic word partial coefficients that have been detected.
It is possible for a user to conduct a treatment procedure to reveal the basic equipment, usually peripherals, and it is possible for those who wish to cooperate, for example, by siege: earphones, smart phones, smart phones,
5 Hand computers, personal computers (PCs, tablets)
personal digital assistants (PDAs) or tablet PCs
And what a shabu.
Although the above mentioned a specific example of a party in detail, people skilled in the field can realize that, the list is only for illustration, and is not intended to limit the scope of protection for the present claim. It includes
10 The handshake program specifically Microsoft Internet Explorer, Mozilla's Firefox, Apple's Safari browser, Google, Opera, Apple's Safari GreenBrowser, Chrome and many others.
Forced blindness from the remembrance of what was previously done to each other by not blinking the handshake in detail, it is possible for the passer-by in the field to realize that, the models of the current students are not limited to my blind handshake, but it would have been possible to apply them
15th Applied to any application (App) A program can be used for a long time to view a web page of a server or a file in a file system and allow users to interact with things, and an application can be a combination of the currently common handshakes, or it can be a I blink again.
My web page browsing function
It is noted that the method and device for detecting the basic quantity proposed in the current demand models are not limited to
20
On the current demand forms, it can also be used in any other form..
For example, it is possible to follow a certain standard program interface, the key word detection method is written as an auxiliary program that is installed on PCs, portable peripherals and the like, or it can be downloaded in a picture of a program that the user downloads and uses. When written as a plugin, it can be implemented in several forms with auxiliary scripts such as dll, .ocx and cab. The proposed basic quantum detector method can also be implemented
٤٥٤٣
-٣٠-
What are the models of current students, through special technologies, for example, the assistant flash program, the program?
RealPlayer plugin, MMS plugin, MIDI plugin, sata plugin, and programme.
ActiveX plugin.
5 The proposed key word detection method can be stored in the current students’ models on multiple storage media in the way of storing instructions or storing a set of information. The storage media includes, but is not limited to: floppy disk, DVD, CD, hard disk, flash memory, U disk, CF card, SD card, MMC card, SM card, memory card, xD card, etc.
In addition, the proposed basic quantum detection method can also be applied in the current demand models for storage media based on NAND flash such as U disk, CF card, SD card, 10 SDHC card, MMC card, SM card, memory card, xD card and so on .
In short, in the current student models, a coding network is generated that integrates linguistic information, and the basic words are placed in groups according to the decoded network information. The search for basic quantifications in a substance, such as fed water, is done using a grille. In some models, when the information of the language of the basic words that was detected is not consistent, a partial coefficient of the basic quantifiers is set.
15th Which complete the disclosure about, and the evaluation of the basic quantities that complete the disclosure is based on the partial transactions. There are multiple models for the current student models, the creation of information for the purposes of the coding network is directly carried out, and the basic words in different languages are placed in groups according to the language information. The aforementioned order works effectively in order to avoid the effect of acquaintance on eliminating the basic word detector, and makes the detection of the basic word in a word that differs in different languages more efficient and accurate.
20 In addition, in the coding process of models for the current requirement, the degree of a code is adjusted by judging the information from languages, and a subfunction is entered to transform into languages, so that a single detection engine can be completed as a multilingual basic quality only by a single detection engine.
٤٥٤٣
-٣١-
The previous description shall be mentioned as examples of the current requirements, which are not intended to limit the current requirements. A modification, equivalent replacement, and development that is carried out without deviating from the spirit and principle of the current demand should fall within the scope of the protection of the current demand.
My blindness is forced from the models described above, what is the concept, I am not, what is meant is to restrict students to their protection
5 Parallel models. On the other hand, the students shall include substitutes, equations and rewards to ensure the spirit and scope of the applied elements of protection. Completely, I usually want certain details here, but it is time to give a detailed explanation of the topics mentioned here. Other than that, it will be the total security of a person of normal ability in the field to be able to practice the subject matter of the patent without the specific details mentioned. What are the other cases. I did not elaborate and describe Tariq known, procedures, components, and circles, as we do not participate in making some aspects of the invention.
10 Unnecessarily vague uncle.
The terms used to describe the student contained herein have been used for the purpose of describing specific models only and are not intended to restrict students. By calculating what is used in the description of students and the added protection elements, expressions indicating the singular include plural forms as well, unless the context clearly indicates otherwise. It shall also be understood that the term “and/or” used herein refers to and includes all possible combinations.
<p>15th For one or more of the aforementioned terms. It is also understood that the terms "jointly", "includes" and/or "includes", when used in the present description, define and exist for the features, operations, elements, and/or components mentioned, but this does not preclude the existence or addition of one or more. From features, processes, elements, other components. and/or Mina groups.</p>
As used here, the term “if” can be interpreted to mean “when” or “when.”
<p>20 upon 'or 'in response to determining' or 'in response to a determination'</p>
in response to detecting or "in response to detection" according with a determination
<p>And that the aforementioned case is true, depending on the context. Similarly, the phrase if the aforementioned case is proven true [or “when the aforementioned case is valid]” can be interpreted as meaning “intransigence upon determining” or “in response to a determination response” determining" or "according to"</p>
or " upon detecting" according with a determination 25 to determine
٤٥٤٣
-٣٢-
In response to detecting, the above-mentioned case is true, depending on the context.
Despite the clarification of some of the various drawings, many of the logical stages in a specific order, it is possible to rearrange the stages that do not depend on the order, and it is also possible in the case of other stages. Or else. wow
5 The time in which the rearrangement or merging of some of the other stages was mentioned. Specifically, some of the other stages. It would be obvious to people of average skill in the field, so we didn't bother mentioning a glorious list of alternatives.
In addition to that, it is necessary to realize whether it is possible to implement the stages, whether it is a picture of a hard component, a tool, software, or a combination thereof.
10 Complete the previous description, for the purpose of clarification, with reference to specific forms. However, the illustrative discussions above are not intended to be long or to restrict the application to specifically disclosed images.
Many modifications and variations can be implemented in light of the above generalizations. The models have been selected and described.
In order to clarify the principles of the application and its practical applications, to enable other people skilled in the field
Using students in the best way possible and usually using models with usually equations
By the account of what is appropriate
<p>15th For specific use.</p>
٤٥٤٣
-٣٣-
8 sheets
Sheet 1 Sheet 2 Sheet 3 Sheet 4 Sheet 5 Sheet 6 Sheet 7 Sheet 8
2 priority claims, no other members on record
Priority claims2
| Document | Office | Kind | Date |
|---|---|---|---|
| 201310355905 | China | A | |
| 2013103559056 | China | – |
Numbers
- Publication
- 4543
- Publication, DOCDB
- 4543
- Application
- 114350692
- Application, DOCDB
- 114350692
Titles2
- English
- Disclosure of the key word to recognize the letter
- Arabic
- الكشف عن الكلمة الأساسية للتعرف على خطاب
Classification
- CPC, 3
- G10L15/08
- G10L15/083
- G10L2015/088