Nova Patents
SA4543B1

Untitled record

Abstract

This application discloses a method implemented of recognizing a keyword in a speech that includes a sequence of audio frames further including a current frame and a subsequent frame. A candidate keyword is determined for the current frame using a decoding network that includes keywords and filler words of multiple languages, and used to determine a confidence score for the audio frame sequence. A word option is also determined for the subsequent frame based on the decoding network, and when the candidate keyword and the word option are associated with two distinct types of languages, the confidence score of the audio frame sequence is updated at least based on a penalty factor associated with the two distinct types of languages. The audio frame sequence is then determined to include both the candidate keyword and the word option by evaluating the updated confidence score according to a keyword determination criterion.

SA4543B1, drawing sheet 1
Sheet 1 of 8

Term

No projected expiry on record.

  1. Priority
  2. Filed
  3. Published
  4. Today

20 claims: 20 independent, 0 dependent

  1. 1
    protection items عناصر الحماية 1- A method for recognizing a keyword in a speech, comprising:receiving a sequence of audio frames comprising a current frame and a subsequent frame;Define the candidate keyword for the current frame using a predefined decoding network that includes 1- طريقة لمتعرف عمى كممة أساسية في خطاب keyword in a speech ، تشتمل عمى: استقبال متوالية من إطا ارت صوت receiving a sequence of audio frames تشتمل عمى إطار حالي current frame واطار الحق subsequent frame ؛ تحديد الكممة األساسية المرشحة لإلطار الحالي باستخدام شبكة فك شفرة decoding network محددة مسبقا تتضمن 5 Base quantifiers and filler words from multiple languages, audio frame sequence binding 5 كممات أساسية وكممات حشو filler words من لغات متعددة ، ربط متوالية إطار الصوت which is a confidence score associating the audio frame sequence التي يتم confidence score بدرجة الثقة associating the audio frame sequence Partially defined according to the candidate keyword;specifying the quantum option for the subsequent frame using the quantum candidate and the predefined decoding network;When a candidate keyword and an option word are associated with two distinct types of language;Confidence Score Update تحديدىا جزئيا وفقا لمكممة األساسية المرشحة؛ تحديد خيار الكممة لإلطار الالحق باستخدام الكممة األساسية المرشحة وشبكة فك التشفير decoding network المحددة مسبقاً؛ وعندما يتم ربط الكممة األساسية المرشحة وخيار الكممة باثنين من األنواع المتميزة لمغات؛ تحديث درجة الثقة 10 confidence score for the sequence of the sound frame based on the part parameter which is predetermined according to the two distinct types of language;mute choice and vocal model for the subsequent frame;Determining the sequence of the sound frame includes both the candidate keyword and the word choice by evaluating the confidence score that has been updated according to the criterion for determining a keyword. 10 confidence score الخاصة بمتوالية إطار الصوت بناءً عمى معامل الج ازء الذي يتم تحديده مسبقا وفقا الثنين من األنواع المتميزة لمغات؛ خيار الكممة والنموذج الصوتي لإلطار الالحق؛ ويشتمل تحديد متوالية إطار الصوت عمى كل من الكممة األساسية المرشحة وخيار الكممة بواسطة تقييم درجة الثقة confidence score التي تم تحديثيا وفقا لمعيار تحديد كممة أساسية.
  2. 2
    15 2- الطريقة وفقاً لعنصر الحماية رقم 1، حيث يتم تحديد مجموعة من الكممات األساسية المرشحة، 15th 2- The method according to claim #1, where a set of candidate key quantifiers is identified, comprising the candidate keyword, for the current frame of the audio frame sequence, and each candidate keyword is associated with at least one keyword option, where a subset of the candidate keyword is selected to be included in the audio frame sequence with At least one word choice of its own التي تشتمل عمى الكممة األساسية المرشحة ، لإلطار الحالي لمتوالية اإلطار الصوتي audio frame sequence ، ويتم ربط كل من الكممة األساسية المرشحة مع خيار الكممة واحد عمى األقل، وحيث يتم تحديد مجموعة فرعية من الكممة األساسية المرشحة ليتم إد ارجيا في متوالية اإلطار الصوتي audio frame sequence مع خيار الكممة الواحد عمى األقل الخاص بيا ذي 20 Relevance based on the criterion of identifying a keyword. 20 الصمة بناءً عمى معيار تحديد كممة أساسية.
  3. 3
    3- The method according to element of protection No. 2, where the subsequent frame is the last frame of the audio frame sequence, and according to the criterion for specifying a key word, the candidate key word associated with the confidence score is chosen from a set of candidate key words such as the associated key word In the current frame of the audio frame sequence 3- الطريقة وفقا لعنصر الحماية رقم 2، حيث يتمثل اإلطار الالحق في اإلطار األخير لمتوالية اإلطار الصوتي audio frame sequence، ووفقا لمعيار تحديد كممة أساسية ، يتم اختيار الكممة األساسية المرشحة المرتبطة بدرجة الثقة confidence score المفضمة من مجموعة من الكممات األساسية المرشحة مثل كممة أساسية مرتبطة باإلطار الحالي لمتوالية اإلطار الصوتي . audio frame sequence 25 . audio frame sequence 25 ٤٥٤٣ ٤٥٤٣ -٣٤- -٣٤-
  4. 4
    4- The method according to claim number 2, where, according to the criterion for determining a key word, each set of candidate key words is linked with a confidence score that is relevant to it. 4- الطريقة وفقا لعنصر الحماية رقم 2، حيث أنو وفقا لمعيار تحديد كممة أساسية ، يتم ربط كل من مجموعة من الكممات األساسية المرشحة مع درجة الثقة confidence score ذات الصمة بيا For the audio frame sequence, the confidence score of 5 is greater than the threshold value of the base quantity. لمتوالية اإلطار الصوتي audio frame sequence ، وتكون درجة الثقة confidence score 5 ذات الصمة أكبر من القيمة الحدية threshold value لمكممة األساسية.
  5. 5
    5- The method according to claim number 2, where after selecting the candidate keyword subset to be included in the audio frame sequence with at least one word choice of its own, the corresponding confidence score is updated 5- الطريقة وفقاً لعنصر الحماية رقم 2، حيث أنو بعد تحديد المجموعة الفرعية لمكممات األساسية المرشحة ليتم إد ارجيا في متوالية اإلطار الصوتي audio frame sequence مع خيار الكممة الواحد عمى األقل الخاص بيا ذي الصمة، يتم تحديث درجة الثقة confidence score المقابمة 10 It is determined to exceed a threshold value for a basic word according to the criterion for determining a basic word. 10 ويتم تحديدىا لتتجاوز قيمة حدية لمكممة األساسية وفقاً لمعيار تحديد كممة أساسية.
  6. 6
    6- The method according to protection element No. 1, where according to the criterion for determining a key quantity, the confidence score of the audio frame sequence is greater than the threshold value of the base quantity. 6- الطريقة وفقا لعنصر الحماية رقم 1، حيث أنو وفقا لمعيار تحديد كممة أساسية ، تكون درجة الثقة confidence score لمتوالية اإلطار الصوتي audio frame sequence أكبر من القيمة الحدية threshold value لمكممة األساسية. 15 15
  7. 7
    7- The method according to protection element No. 1, where the pre-defined decoding network is linked to two or more English, Chinese, Japanese, Russian, French, German and the like, and includes a subset of basic words and a subset of words filler words for each of two or more languages. 7- الطريقة وفقا لعنصر الحماية رقم 1، حيث يتم ربط شبكة فك التشفير decoding network المحددة مسبقاً باثنين أو أكثر من المغات اإلنجميزية، الصينية، اليابانية، الروسية، الفرنسية، األلمانية وما شابو ذلك، وتشتمل عمى مجموعة فرعية من الكممات األساسية ومجموعة فرعية من كممات الحشو filler words لكل من اثنين أو أكثر من المغات. 20 20
  8. 8
    8- The method according to claim No. 1, where each keyword from a pre-defined decoding network includes one or more triplex speakers. 8- الطريقة وفقا لعنصر الحماية رقم 1، حيث تشتمل كل كممة أساسية من شبكة فك التشفير decoding network المحددة مسبقاً عمى واحدة أو أكثر من سماعات ثالثية.
  9. 9
    9- The method according to claim #1, which is according to the decoding structure 9- الطريقة وفقا لعنصر الحماية رقم 1، حيث أنو وفقا لييكل فك التشفير decoding 25 structure of a predefined decoding network, each keyword in a predefined decoding network is associated with at least one keyword used with 25 structure لشبكة فك التشفير decoding network المحددة مسبقا، يتم ربط كل كممة أساسية في شبكة فك التشفير decoding network المحددة مسبقا بكممة واحدة عمى األقل تستخدم مع ٤٥٤٣ ٤٥٤٣ -٣٥- -٣٥- The imprint base quantum in a real speech is argy in the decoding . network الكممة األساسية ذات الصمة في خطاب حقيقي واد ارجيا في في شبكة فك التشفير decoding . network . network
  10. 10
    10- The method according to claim No. 9, which is according to the decoding code 10- الطريقة وفقا لعنصر الحماية رقم 9، حيث أنو وفقا لييكل فك التشفير decoding 5 structure of a predefined decoding network, producing each keyword in a subset of keywords and at least one related word that is used with the related keyword from two different languages. 5 structure لشبكة فك التشفير decoding network المحددة مسبقا، تنتج كل كممة أساسية في مجموعة فرعية من الكممات األساسية والكممة الواحدة عمى األقل ذات الصمة التي تستخدم مع الكممة األساسية ذات الصمة من اثنين من المغات المختمفة.
  11. 11
    11- The method in accordance with claim No. 1 also includes:11- الطريقة وفقا لعنصر الحماية رقم 1، تشتمل أيضا عمى: 10 Create a part coefficient table to include a set of part factors, each of which is linked to two different languages, where the part coefficient used to update the confidence score for the audio frame sequence is determined by extracting the part coefficient table based on two distinct types From the different languages ​​of the candidate basic quantifier and the option of the quantum. 10 إنشاء جدول معامل الج ازء ليشتمل عمى مجموعة من عوامل الج ازء يتم ربط كل منيا باثنين من المغات المختمفة، حيث يتم تحديد معامل الج ازء المستخدم لتحديث درجة الثقة confidence score لمتوالية اإلطار الصوتي audio frame sequence بواسطة استخ ارج جدول معامل الج ازء بناءً عمى نوعين متميزين من المغات المختمفة لمكممة األساسية المرشحة وخيار الكممة.
  12. 12
    15 12- الطريقة وفقا لعنصر الحماية رقم 1، تشتمل أيضا عمى:15th 12- The method in accordance with claim No. 1 also includes: Creating a predefined decoder network, in which base words and filler words from multiple languages ​​are grouped according to their language types, also includes: إنشاء شبكة فك تشفير محددة مسبقاً، حيث يتم تجميع الكممات األساسية وكممات الحشو filler words من لغات متعددة وفقاً ألنواع المغات الخاصة بيا، يشتمل أيضاً عمى: Create a start node and an end node;Create a set of language nodes, each representing a type of language, binding each language node to a start node ;خمق عقدة بداية start node وعقدة نياية end node ؛ خمق مجموعة من عقد المغة language nodes يمثل كل منيا نوع من المغة، ربط كل عقدة لغة بعقدة بداية start node ؛ 20 Associate each language node with a subset of related keywords and a subset of related filler words as both from the corresponding language;For each base word, converting the keyword of the triphone sequences, creating a three-phoneme node for each triphone of the triphone sequences of the triphone sequence, connecting the speaker node 20 ربط كل عقدة لغة بمجموعة فرعية من الكممات األساسية ذات الصمة ومجموعة فرعية من كممات الحشو filler words ذات الصمة كمتاىما من المغة المقابمة؛ لكل كممة أساسية، تحويل الكممة األساسية converting the keyword ذات الصمة لمتوالية من السماعات الثالثية triphone sequences ، إنشاء عقدة سماعة ثالثية ذات صمة لكل سماعة ثالثية من متوالية السماعات الثالثية triphone sequences من الكممة األساسية ذات الصمة، ربط عقد السماعة الثالثية 25 of triphone sequences together to form a triphone node sequence including a master treble node and an edematous triphone node, linking the treble node 25 من متوالية السماعات الثالثية triphone sequences معا لتشكيل متوالية عقد سماعة ثالثية بما في ذلك عقدة سماعة ثالثية رئيسية وعقدة سماعة ثالثية ذيمية، ربط عقدة السماعة الثالثية ٤٥٤٣ ٤٥٤٣ -٣٦- -٣٦- The main sympathetic node has the parasympathetic node and the edematous tertiary speaker node has an end node. For each stuffing word, an emollient padding node is created, and the 'immunopositive' padding node is created between the corresponding idiom node and the end node;And connect a start node and end node الرئيسية ذات الصمة بعقدة المغة المقابمة وعقدة السماعة الثالثية الذيمية ذات الصمة بعقدة نياية end node ؛ لكل كممة حشو، يتم إنشاء عقدة حشو ذات صمة واق ارن عقدة الحشو ذات الصمة بين عقدة المغة المقابمة وعقدة نياية end node ؛ وربط عقدة بداية start node وعقدة نياية . end node . end node 5 5
  13. 13
    13- The method according to protection element No. 12, where the candidate keyword and the word option are selected to be associated with two distinct types of languages, where one of a set of language nodes is linked between the candidate keyword and the word option on the pre-defined decoding network . 13- الطريقة وفقا لعنصر الحماية رقم 12، حيث يتم تحديد الكممة األساسية المرشحة وخيار الكممة ليتم ربطيا باثنين من األنواع المميزة لمغات، حيث يتم ربط واحدة من مجموعة من عقد المغة language nodes بين الكممة األساسية المرشحة وخيار الكممة عمى شبكة فك التشفير decoding network المحددة مسبقا. 10 10
  14. 14
    14- The method according to protection element No. 12, whereby according to the decoding structure of the predefined decoding network, each keyword in the decoding network is connected to the predefined decoding network with at least one word that is used With the relevant main word in verbal speech. 14- الطريقة وفقا لعنصر الحماية رقم 12، حيث أنو وفقا ل ىيكل فك التشفير decoding structure لشبكة فك التشفير decoding network المحددة مسبقا، يتم ربط كل كممة أساسية في شبكة فك التشفير decoding network عمى شبكة فك التشفير decoding network المحددة مسبقاً بكممة واحدة عمى األقل تستخدم مع الكممة األساسية ذات الصمة في خطاب فعمي. 15 15
  15. 15
    15- An electronic badge that contains:15- وسيمة إلكترونية ، تحتوي عمى: one or more processors;A memory has generalizations stored in it, which when executed by one or more processors cause the processors to perform operations that include: واحد أو أكثر من المعالجات؛ وذاكرة بيا تعميمات مخزنة عمييا، والتي عند تنفيذىا بواسطة واحد أو أكثر من المعالجات تجعل المعالجات تقوم بإج ارء عمميات تشتمل عمى: Receiving a sequence of sound frames comprising a current frame and a subsequent frame استقبال متوالية من إطا ارت صوت تشتمل عمى إطار حالي current frame واطار الحق 20 after frame follows the current frame;Define a keyword candidate for the current framework using a predefined decoding network that includes keywords and filler words from multiple languages;associating the audio frame sequence with a confidence degree partially determined by the candidate keyword;Selecting a quantum option for the subsequent frame using the main quantum filter and the decoding network 20 subsequent frame يتبع اإلطار الحالي؛ تحديد كممة أساسية مرشحة لإلطار الحالي باستخدام شبكة فك شفةر decoding network محددة مسبقاً تشمل عمى الكممات األساسية وكممات حشو filler words من لغات متعددة ؛ ربط متوالية إطار الصوت associating the audio frame sequence بدرجة ثقة يتم تحديدىا جزئيا وفقا لمكممة األساسية المرشحة؛ تحديد خيار كممة لإلطار الالحق باستخدام الكممة األساسية المرشحة وشبكة فك شفرة decoding network 25 Predefined;When a candidate base word and an option word are associated with two distinct types of language, . is done 25 محددة مسبقاً؛ عند ربط الكممة األساسية المرشحة وخيار الكممة بنوعين متميزين من المغات، يتم Update the confidence score of the sound frame sequence based on a section coefficient تحديث درجة الثقة confidence score الخاصة بمتوالية إطار الصوت بناءً عمى معامل جازء ٤٥٤٣ ٤٥٤٣ -٣٧- -٣٧- a pre-set penalty factor according to the two distinct types of language, the choice of punch and audio model for the later frame;Determine that the audio frame sequence includes both the candidate keyword and the word choice by evaluating the confidence score that was updated according to a key word criterion. penalty factor محدد مسبقا وفقا الثنتين من األنواع المتميزة لمغات، خيار الكممة ونموذج سمعي لإلطار الالحق ؛ وتحديد أن متوالية اإلطار الصوتي audio frame sequence تشتمل عمى كل من الكممة األساسية المرشحة وخيار الكممة من خالل تقييم درجة الثقة confidence score التي تم تحديثيا وفقا لمعيار تحديد كممة أساسية.
  16. 16
    16- The electronic tag in accordance with protection element No. 15, where, according to the criterion for determining a base quantity, the confidence score of the audio frame sequence is greater than the threshold value of the key quantity. 16- الوسيمة اإللكترونية وفقا لعنصر الحماية رقم 15، حيث أنو وفقا لمعيار تحديد كممة أساسية ، تكون درجة الثقة confidence score لمتوالية اإلطار الصوتي audio frame sequence أكبر من القيمة الحدية threshold value لمكممة األساسية.
  17. 17
    10 17- The electronic merchandise in accordance with the element of protection No. 15, which includes the operations that are performed 10 17- الوسيمة اإللكترونية وفقا لعنصر الحماية رقم 15، حيث تشتمل العمميات التي يتم إج ارؤىا By wizards also blindness:بواسطة المعالجات أيضا عمى: Create a part coefficient table to include a set of part factors, each of which is linked to two different languages, where the part coefficient used to update the confidence score for the audio frame sequence is determined by extracting the 15 part coefficient table based on two types Distinguished from the different languages ​​of the candidate basic quantity and the option of the word. إنشاء جدول معامل الج ازء ليشتمل عمى مجموعة من عوامل الج ازء يتم ربط كل منيا باثنين من المغات المختمفة، حيث يتم تحديد معامل الج ازء المستخدم لتحديث درجة الثقة confidence score لمتوالية اإلطار الصوتي audio frame sequence بواسطة استخ ارج جدول معامل 15 الج ازء بناءً عمى نوعين متميزين من المغات المختمفة لمكممة األساسية المرشحة وخيار الكممة.
  18. 18
    18- The electronic medium according to the element of protection No. 15, where the pre-defined decoding network is linked to two or more English, Chinese, Japanese, Russian, French, German and the like, and includes a subset of basic words 18- الوسيمة اإللكترونية وفقا لعنصر الحماية رقم 15، حيث يتم ربط شبكة فك التشفير decoding network المحددة مسبقاً باثنين أو أكثر من المغات اإلنجميزية، الصينية، اليابانية، الروسية، الفرنسية، األلمانية وما شابو ذلك، وتشتمل عمى مجموعة فرعية من الكممات األساسية 20 A subset of filler words for each of two or more languages. 20 ومجموعة فرعية من كممات الحشو filler words لكل من اثنين أو أكثر من المغات.
  19. 19
    19- A storage medium that can be read by a non-transitory computer readable medium, with stored circulars on which, when executed by one or more processors, the processors perform operations that include:19- وسط تخزين يمكن ق ارءتو بواسطة حاسوب غير انتقالي -non-transitory computer readable medium، بو تعميمات مخزنة عميو والتي عند تنفيذىا بواسطة واحد أو أكثر من المعالجات تجعل المعالجات تقوم بإج ارء عمميات تشتمل عمى: 25 receiving a sequence of audio frames comprising a current frame and a subsequent frame that follows the current frame;25 استقبال متوالية من إطا ارت صوت receiving a sequence of audio frames تشتمل عمى إطار حالي current frame واطار الحق subsequent frame يتبع اإلطار الحالي؛ ٤٥٤٣ ٤٥٤٣ -٣٨- -٣٨- Define a keyword candidate for the current framework using a predefined decoding network that includes keywords and filler words from multiple languages;associating the audio frame sequence with a confidence degree partially determined by the candidate keyword;Selecting a muzzle option for the subsequent frame using the main word تحديد كممة أساسية مرشحة لإلطار الحالي باستخدام شبكة فك شفرة decoding network محددة مسبقا تشمل عمى الكممات األساسية وكممات حشو filler words من لغات متعددة ؛ ربط متوالية إطار الصوت associating the audio frame sequence بدرجة ثقة يتم تحديدىا جزئيا وفقا لمكممة األساسية المرشحة؛ تحديد خيار كممة لإلطار الالحق باستخدام الكممة األساسية 5 The filter and decoding network is predefined;When a candidate keyword and word choice are associated with two distinct types of language, the confidence score of the audio frame sequence is updated based on a pre-determined penalty factor according to the two distinct language types, the word choice and an audio model of the subsequent frame;Determine that the audio frame sequence includes each of the candidate keywords 5 المرشحة وشبكة فك شفرة decoding network محددة مسبقاً؛ عند ربط الكممة األساسية المرشحة وخيار الكممة بنوعين متميزين من المغات، يتم تحديث درجة الثقة confidence score الخاصة بمتوالية إطار الصوت بناءً عمى معامل جازء penalty factor محدد مسبقا وفقا الثنتين من األنواع المتميزة لمغات، خيار الكممة ونموذج سمعي لإلطار الالحق ؛ وتحديد أن متوالية اإلطار الصوتي audio frame sequence تشتمل عمى كل من الكممة األساسية المرشحة 10 And the choice of the word by evaluating the degree of confidence that has been updated according to the criterion for determining a basic word. 10 وخيار الكممة من خالل تقييم درجة الثقة التي تم تحديثيا وفقا لمعيار تحديد كممة أساسية.
  20. 20
    20- A storage medium that can be read by a non-transitory computer readable medium in accordance with the element of protection No. 19, where the operations that are performed by processors also include:20- وسط تخزين يمكن ق ارءتو بواسطة حاسوب غير انتقالي -non-transitory computer readable medium وفقا لعنصر الحماية رقم 19، حيث تشتمل العمميات التي يتم إج ارؤىا بواسطة المعالجات processors أيضا عمى: 15 إنشاء جدول معامل الج ازء ليشتمل عمى مجموعة من عوامل الج ازء يتم ربط كل منيا باثنين من المغات المختمفة، حيث يتم تحديد معامل الج ازء المستخدم لتحديث درجة الثقة confidence score لمتوالية اإلطار الصوتي audio frame sequence بواسطة استخ ارج جدول معامل الج ازء بناءً عمى نوعين متميزين من المغات المختمفة لمكممة األساسية المرشحة وخيار الكممة. 15th Create a part coefficient table to include a set of part factors, each of which is linked to two different languages, where the part coefficient used to update the confidence score for the audio frame sequence is determined by extracting the part coefficient table based on two distinct types From the different languages ​​of the candidate basic quantifier and the option of the quantum. ٤٥٤٣ ٤٥٤٣ -٣٩- -٣٩- Update the degree of stress تحديثدرجةأتثدة later audio frame اطار صوتي لاحق F6F2F3F;F5F6F7F8F9F10F11F12 Fn F٦F٢F٣F؛F٥F٦F٧F٨F٩F١٠F١١F١٢ Fn audio frame اطار صوتي حاتي barracks ثكن ا ٤٥٤٣ ٤٥٤٣
Independent claims20