Method and system for non-intrusive speaker verification using behavior model
Abstract
According to the present invention, a system and method for authenticating a user's identity includes a conversation system (114) for receiving input from a user (110) and converting the input into a formal command. The behavior checker (118) is coupled to the conversation system (114) to extract features from the input. The characteristics include user behavior patterns. The behavior checker (118) is adapted to compare the input behavior with a behavior model (214) to determine whether to approve the user to interact with the system.

Term
Term ended
Projected expiry passed 12 December 2021, 4.8 years ago.
- Priority and filed
- Published
- Projected expiry
- Today
40 claims: 3 independent, 37 dependent
- 1一种用于验证用户身份的系统,包括:一个会话系统,用于接收来自用户的输入并将所述输入转换成形式命令;以及一个与会话系统相耦合的行为检验器,用于从输入中提取特征,这些特征包括用户的行为模式,行为检验器适于对输入行为以及一个行为模型进行比较,从而确定是否批准该用户与系统进行交互。
- 2如权利要求1所述的系统,其中会话系统包括一个自然语言理解单元,用于解释作为输入而被接收的语音。
- 3如权利要求1所述的系统,其中所述输入包括语音、笔迹、文本和手势中的至少一个。
- 4如权利要求1所述的系统,其中行为检验器包括一个特征提取器,用于从输入中提取特征向量。
- 5如权利要求4所述的系统,其中特征向量包括语言模型得分、声学模型得分、自然语言和理解得分中的至少一个。
- 6如权利要求4所述的系统,其中特征向量包括命令预测得分和发音得分中的至少一个。
- 7如权利要求4所述的系统,其中特征向量包括关于系统响应用户的信息。
- 8如权利要求4所述的系统,其中特征向量包括用户命令之间的持续时间以及用户和系统之间的对话状态中的至少一个。
- 9如权利要求4所述的系统,其中特征向量包括用户使用的输入模态类型。
- 10如权利要求1所述的系统,其中行为模型包括多个模型。
- 11如权利要求1所述的系统,其中行为检验器包括一个概率计算器,所述概率计算器适于根据用户行为来计算批准该用户与系统进行交互的第一概率。
- 12如权利要求11所述的系统,其中行为检验器包括一个模型构造器,用于为用户构造一个行为模型,所述行为模型由概率计算器使用,以便对其中的行为与用户的当前行为进行比较。
- 13如权利要求11所述的系统,还包括:一个声学和生物识别检验器,用于确定来自用户的声学和生物识别信息,并且根据用户的声学或生物识别信息来确定一个批准该用户与系统进行交互的第二概率;以及所述行为检验器包括一个概率混合器,所述概率混合器适于将第一概率与第二概率相结合,从而证实批准该用户与系统进行交互。
- 14如权利要求11所述的系统,其中将第一概率与一个阈值概率进行比较,以便确定是否批准该用户使用所述系统。
- 15一种基于行为来验证用户的方法,包括以下步骤:向一个用于接收来自用户的输入并将所述输入转换成形式命令的会话系统提供输入;从所述输入中提取包含了用户行为模式的特征;以及将输入行为与一个行为模型进行比较,从而确定是否批准该用户与系统进行交互。
- 16如权利要求15所述的方法,其中会话系统包括一个自然语言理解单元,所述方法还包括对作为输入而被接收的语音进行解释的步骤。
- 17如权利要求15所述的方法,其中所述输入包括语音、笔迹、文本和手势中的至少一个。
- 18如权利要求15所述的方法,其中行为检验器包括一个特征提取器,所述方法还包括从输入中提取特征向量的步骤。
- 19如权利要求18所述的方法,其中特征向量包括语言模型得分、声学模型得分、自然语言理解得分中的至少一个。
- 20如权利要求18所述的方法,其中特征向量包括命令预测得分和发音得分中的至少一个。
- 21如权利要求18所述的方法,其中特征向量包括关于系统响应用户的信息。
- 22如权利要求18所述的方法,其中特征向量包括用户命令之间的持续时间以及用户和系统之间的对话状态中的至少一个。
- 23如权利要求18所述的方法,其中特征向量包括用户使用的输入模态类型。
- 24如权利要求15所述的方法,其中行为检验器包括一个概率计算器,所述方法还包括根据在概率计算器上计算第一概率的步骤,所述第一概率指示的是:是否基于用户行为而批准该用户与系统进行交互。
- 25如权利要求24所述的方法,其中行为检验器包括一个模型构造器,所述方法还包括为用户构造一个行为模型的步骤,所述行为模型由概率计算器使用,以便对其中的行为与用户的当前行为进行比较。
- 26如权利要求24所述的方法,还包括:一个声学和生物识别检验器,用于确定来自用户的声学和生物识别信息,所述方法还包括步骤:基于用户的声学或生物识别信息来确定一个批准该用户与系统进行交互的第二概率;以及通过使用一个概率混合器将第一概率与第二概率相结合,以便证实批准该用户与系统进行交互。
- 27如权利要求24所述的方法,其中将第一概率与一个阈值概率进行比较,以便确定是否批准该用户使用所述系统。
- 28一种可以由机器读取的程序存储设备,其中实际包含了一个可以由机器执行的指令程序,以便执行根据行为来验证用户的方法步骤,所述方法步骤包括:向一个用于接收来自用户的输入并将所述输入转换成形式命令的会话系统提供输入;从所述输入中提取包含了用户行为模式的特征;以及将输入行为与一个行为模型进行比较,从而确定是否批准该用户与系统进行交互。
- 29如权利要求28所述的程序存储设备,其中会话系统包括一个自然语言理解单元,所述方法还包括对作为输入而被接收的语音进行解释的步骤。
- 30如权利要求28所述的程序存储设备,其中所述输入包括语音、笔迹、文本和手势中的至少一个。
- 31如权利要求28所述的程序存储设备,其中行为检验器包括一个特征提取器,所述方法还包括从输入中提取特征向量的步骤。
- 32如权利要求31所述的程序存储设备,其中特征向量包括语言模型得分、声学模型得分、自然语言理解得分中的至少一个。
- 33如权利要求31所述的程序存储设备,其中特征向量包括命令预测得分和发音得分中的至少一个。
- 34如权利要求31所述的程序存储设备,其中特征向量包括关于系统响应用户的信息。
- 35如权利要求31所述的程序存储设备,其中特征向量包括用户命令之间的持续时间以及用户和系统之间的对话状态中的至少一个。
- 36如权利要求31所述的程序存储设备,其中特征向量包括用户使用的输入模态类型。
- 37如权利要求28所述的程序存储设备,其中行为检验器包括一个概率计算器,所述方法还包括在概率计算器上计算第一概率的步骤,所述第一概率指示的是:是否基于用户行为而批准该用户与系统进行交互。
- 38如权利要求37所述的程序存储设备,其中行为检验器包括一个模型构造器,所述方法包括为用户构造一个行为模型的步骤,所述行为模型由概率计算器使用,以便对其中的行为与用户的当前行为进行比较。
- 39如权利要求37所述的程序存储设备,还包括:一个声学和生物识别检验器,用于确定来自用户的声学和生物识别信息,所述方法还包括步骤:基于用户的声学或生物识别信息来确定一个批准该用户与系统进行交互的第二概率;以及通过使用一个概率混合器将第一概率与第二概率相结合,以便证实批准该用户与系统进行交互。
- 40如权利要求37所述的程序存储设备,其中将第一概率与一个阈值概率进行比较,以便确定是否批准该用户使用所述系统。
Independent claims40
54 paragraphs, as filed
Method and system for non-interference speaker verification using behavior model
Technical field
The present invention relates to a natural language understanding system, in particular to a method and system for non-interference verification of users based on user behaviors.
Background technique
Traditional methods for speaker verification (or identification) rely on specific input from the user and used only for this purpose. These methods include providing voice samples and answering biometric questions. Once verified, the speaker is allowed to access the target system, and no further verification is usually performed. Even if additional verification is performed, the additional verification requires more specific input from the user for verification purposes. And this will disturb the user.
The prior art speaker verification system (for those systems that do not have a spoken input modality, it can also be a user verification system) based on one or more of the following standards to confirm the identity of a designated user: 1. User Who it is, this can be determined by the user's voice, fingerprints, handwriting, etc.
2. What the user knows, which can be determined by passwords or answers to certain biometric questions (for example, what is the mother's maiden name?) and other information.
3. What the user has, such as identifying documents, keys, cell phones with specific numbers, etc.
If the impersonator knows or possesses information such as the key or maiden name, all the above methods for verification will be invalid.
Therefore, there is a need for a method and system for determining user identity based on user behavior. And there is a need for a non-interfering user authentication system.
Summary of the invention
According to the present invention, a system for verifying user identity includes a conversation system for receiving input from the user and converting the input into formal commands. A behavior checker is coupled to the conversation system to extract features from the input. These characteristics include user behavior patterns. The behavior checker is adapted to compare the input behavior with a behavior model to determine whether to approve the user to interact with the system.
In an alternative embodiment, the conversation system may include a natural language understanding unit for interpreting speech received as input. The input may include at least one of voice, handwriting, text, and gesture. The behavior checker may include a feature extractor for extracting feature vectors from the input. The feature vector may include at least one of a language model score, an acoustic model score, a natural language and comprehension score, a command prediction score, and/or a pronunciation score. And the feature vector may include at least one of the information of the system responding to the user, the duration between user commands, the dialogue state between the user and the system, and/or the input modality type used by the user. The behavior model can include multiple models. The behavior checker can include a probability calculator. The probability calculator is adapted to calculate the first probability of approving the user to interact with the system according to the user's behavior. The behavior checker may include a model builder that constructs a behavior model for the user, and the behavior model is used by the probability calculator to compare the behavior therein with the current behavior of the user. The system can also include an acoustic and biometric checker to determine those acoustic and biometric information from the user and based on the users acoustic or biometric information to determine a second probability of approving the users interaction with the system, and to check the behavior The device may include a probability mixer adapted to combine the first frequency with the second frequency in order to verify that the user is authorized to interact with the system. If the user is approved to use the system, the first probability and a threshold probability can be compared.
According to the present invention, a method for authenticating a user based on behavior includes the steps of: providing input to a conversation system for receiving input from a user and converting the input into a formal command, and extracting the input from the input. The characteristics of user behavior patterns and the comparison of input behavior with a behavior model to determine whether to approve the user to interact with the system.
In other methods, the conversation system may include a natural language understanding unit, and the method may also include the step of interpreting the speech received as input. These inputs may include at least one of voice, handwriting, text, and gestures. The feature vector may include at least one of a language model score, an acoustic model score, a natural language and comprehension score, a command prediction score, and/or a pronunciation score. The feature vector may include at least one of the information that the system responds to the user, the duration between user commands, the dialogue state between users, and/or the input modality type used by the user. The behavior checker may include a probability calculator, and the method may include a step of calculating a first probability for indicating whether to approve the user to interact with the system on the probability calculator based on the user behavior.
In other methods, the behavior checker may include a model builder, and the method may include the step of constructing a behavior model for the user, wherein the probability calculator uses the behavior model to compare the behavior of the model with the current behavior of the user. . And it may include an acoustic and biometric checker for determining the acoustic and biometric information from the user, and the method further includes the following steps: based on the users acoustic or biometric information, determining whether to approve the users The second probability of the system interacting, and the first probability and the second probability are combined by using a probability mixer to confirm that the user is approved to interact with the system. And the first probability and a threshold probability can be compared to determine whether to approve the user to use the system. The methods and method steps of the present invention can be implemented by a program storage device that can be read by a machine, which actually contains program instructions that can be executed by the machine, so as to perform the method steps of verifying the user based on behavior.
These and other objectives, features, and advantages of the present invention will become clear from the following detailed description of exemplary embodiments of the present invention as understood in conjunction with the accompanying drawings.
Description of the drawings
In the following description of the preferred embodiments of the present invention, the present invention will be described in detail with reference to the accompanying drawings, in which: FIG. 1 is a block diagram/flow chart of an exemplary system/method using behavior verification according to the present invention; and FIG. 2 It is a block diagram of an exemplary behavior checker according to the present invention.
detailed description
The present invention provides a system and method for continuously verifying user identity based on how the user interacts with the target system. The verification can be implemented by comparing the user's current behavior and past behavior. No additional dedicated input from the user is required (except for initial verification), so the system is interference-free. In addition, verification can be performed continuously, and if sufficient evidence can be obtained to reject a user who is in a session, the user can be disconnected before it causes more damage.
In an alternative embodiment, even if initial verification is not necessary and basic level access (for example, to those non-confidential information) can be provided to all users, full access can be provided once additional verification is performed through non-interference processing.
In the present invention, a new dimension is provided to the speaker (or user) verification paradigm by introducing a new standard, which is how the user acts. For example, it is possible to distinguish between a user who usually uses "Howdy" to greet and an imposter who uses "Hello" or "How are you" to greet, and it is also possible to distinguish between the user and a person who starts a conversation without greeting. The pretenders are distinguished. Likewise, it is possible to distinguish between imposters who try to search for several confidential files and legitimate users who do not normally perform such searches. Although any single interaction with the system is not sufficient to make a decision, the information collected after several user-system interactions is sufficient to make a correct verification decision.
An advantage of the present invention is that, in the sense that it does not require additional dedicated input from the user for verification purposes, speaker verification is non-interference, the user can interact with the system as usual, and the information needed for verification is determined by Background processing is automatically collected. The comparison between the user's current behavior and the known past behavior is also done automatically by the system without causing any interference or inconvenience to the user.
It should be understood that various hardware, software, or a combination of the two can be used to implement the elements shown in FIGS. 1 to 2. Preferably, these elements are executed in one or more software of a general-purpose digital computer that is properly programmed and has a processor, memory, and input/output interface. Referring now to the drawings, in which the same numbers indicate the same or similar components, first referring to FIG. 1, which shows an exemplary system/method employing behavior verification according to the present invention. For the target system 100 that requires speaker verification, the system preferably can provide certain parameters about how the user 110 interacts with the system. For example, the system 100 may allow the user 110 to interact with the system using several different input modalities, such as typing text, spoken pronunciation, handwriting input, gestures, and so on. The system 100 can use technologies such as speech recognition, handwriting recognition, and image recognition together with natural language understanding and dialogue management to interpret user input and convert it into a form suitable for execution by one or more computers of the system 100. The system 100 can be connected to a number of different applications 116, such as e-mail, electronic calendar, banking, stock or mutual fund transactions, travel services, electronic data sheets, editing programs, etc., and the system 100 allows users to interact with these applications To interact. The system 100 can also provide parameters that describe how the user interacts with the system, such as parameters related to speech recognition or natural language understanding.
As shown in FIG. 1, an example of the system 100 shown in the figure includes a behavior checker 118. The input from the user 110 is expected to be spoken pronunciation, but may also be in other modalities, such as handwriting input, typed text, or gestures. When using spoken input, the conversation system 114 first uses a speech recognition engine 113 known in the prior art to convert the spoken pronunciation into text. For example, if the application 116 is an email application, the user can say "do I have any new messages", and the spoken pronunciation will be converted into a corresponding text string by the speech recognition engine. In addition, appropriate technologies known in the art, such as the handwriting recognition engine 117, are used to convert input that is not in a spoken form, such as handwriting input, into a corresponding text string. This is also true for interpreting gestures or other modalities. An appropriate recognition engine is used in all of them. In this way, all input will be converted into a recognizable form understood by the system 100.
A natural language understanding (NLU) engine 115 is then used to analyze text strings or other formatted signals to convert them into commands suitable for execution by the system 100 within the application 116. For example, sentences such as "do I have any new messages" or "can you check my mailbox" have the same meaning, and they can all be converted into a formal command in the form of CheckNewMail(). The formal command is then submitted to the application 116 for executing the command. In addition, the dialog engine 120 or the dialog manager can also be used to manage the dialog with the user and perform certain other functions, such as ambiguity resolution.
Therefore, the conversation system may include a speech and other input recognition engine, a natural language understanding (NLU) engine 115, and a dialogue engine 120. In the art, methods for constructing a conversation system are known.
The system 100 includes an acoustic and biometric checker 112. The acoustic and biometric verification device 112 is responsible for identifying and verifying the identity of the user 110. Nominally, the verification is performed before the user 110 is allowed to access the system 100. The verification process may include matching the acoustic signature of a certain person who claims to be the designated user and the known acoustic signature of the claimed user. This process is an acoustic verification process. The verification process may also include biometric verification, thereby prompting a person who claims to be a user to answer specific questions, such as a password, mother's maiden name, social security number, and so on. In the art, methods for acoustic and biometric verification are well known.
According to the present invention, during use, the behavior checker 118 is responsible for continuously performing additional user identity verification. The following describes the details of the behavior checker with reference to FIG. 2. The behavior checker 118 receives input from the conversation system 114 and the acoustic and biometric checker 112 and provides its output to the acoustic and biometric checker 112.
Referring to FIG. 2, the feature extractor 204 is responsible for extracting a feature set from the data provided by the conversation system 114, and constructing a feature vector v containing n features.
v=[ν1, ......, νn] (1) The value of n should be selected by the system designer, and the value of n can depend on the type of accuracy and/or recognition required by the system. The features ν1,..., Νn extracted by the feature extractor 204 may include one or more of the following features, and may also include other similar features. The following list of features is illustrative and is not considered as limiting the invention. In addition, the features described here can be used alone or in combination with other features, so as to determine one or more appropriate feature vectors according to the present invention. The feature may include one or more of the following features: 1) Language model score: The speech recognition engine uses a language model or a group of language models to perform the recognition. When more than one language model is used, some of these models can be personalized to a designated user (sometimes called a personal buffer, which is constructed using words and phrases frequently spoken by the designated user). of). The language model score is generated and used internally, and will be lost after the end of recognition. However, these scores contain information that can characterize the user, especially relative to the choice of frequently used words and phrases. For example, if the user usually says "begin dictation", you can detect a person who says "let us create the text for this The impersonator of "message can also be distinguished between a user who usually uses short and concise phrases to issue commands and a person who uses long sentences. Therefore, the language model score can be saved and used as a feature in the feature vector. It should be noted that there is no need to reject imposters based on a single phrase or multiple phrases. Instead, a cumulative behavior score can be maintained for a specified user session, and the score can be cycled relative to a threshold Check to determine whether the user is an impostor or whether the user has not been authenticated for using the system.
2) Acoustic model score: In the speech recognition engine, the acoustic model score (sometimes called the fast matching score and the detailed matching score) and other intermediate outputs are used internally, and they are discarded after recognition. Similar to the language model score, the acoustic model score also includes information related to characterizing a user, and can detect any deviation from the normal score range of a specified task and use it to identify an impostor. Therefore, it will be very useful to add an acoustic model to the feature vector.
3) Natural language understanding (NLU) score: The NLU engine also generates an internal score, which will be discarded after the conversion of the text to the formal command is completed. These scores also include information that can be used in characterizing users. The NLU engine usually includes two or more stages (for example, a marking stage and a conversion stage), and all these scores can be added to the feature vector, so that any deviation from the normal score range of a specified task can be detected.
In addition to these scores, other inputs can also be coded to use them as features. The other inputs can be the second choice of formal commands or the second choice of marked sentences from the intermediate marking stage. For example, the user can say "OpenSteve", which may result in the highest form command OpenMessage (name=Steve) corresponding to opening a message from Steve, and a first corresponding to opening a folder named Steve 2. Select the form command OpenFolder (folder=Steve). However, the impostor may understand better and may say something similar to "Open the message from Steve". In this case, the first choice form command is likely to be the same, but the second choice The commands can be different.
4) Command prediction score: The user often displays a pattern in the sequence of commands issued by him and the combination of commands frequently used to complete a task. Therefore, a system that predicts the user's next command based on past behavior can be used to improve the accuracy of the conversational system and make the system take the initiative to suggest the next command to the user. The system can be G.Ramaswamy and J. . Kleindienst filed October 30, 1999, named "Adaptive Command Predictor for a Natural Language Dialog System", jointly assigned US Patent Application 09/431,034, which application is incorporated herein by reference. However, in addition to these applications, the scores generated by the command prediction system can also be used to detect imposters. If someone issues a command that the actual user has never used (and therefore will get a very low command prediction score), or if someone issues a series of commands that are not the highest-level prediction commands (again, the command prediction score will be very low). Low), then this usual command or sequence of commands can indicate the presence of an imposter. Therefore, the command prediction score is a very good feature added to the feature vector.
5) Pronunciation model: In most languages, certain words have more than one pronunciation. For example, in English, the word "the" has the following universal pronunciation.
the |DH AHthe |DH AXthe |DH IY Most users tend to use a pronunciation of these words. An impostor who does not understand the pronunciation of certain words may use an alternative pronunciation. In this case, in order to detect imposters, the feature vector may include a set of features used to encode the pronunciation of these words.
6) Other input scores: If the system supports other input modalities, such as handwriting recognition or image recognition, then, similar to the language model and acoustic model scores from speech recognition, the scores from these recognition engines can also be added to the feature vector. in.
7) System response: Conversational systems not only accept spoken input from the user, but they also maintain a conversation with the user and generate responses that are presented to the user. The system of the present invention can check what kind of response the system usually generates for the user, and can use this information to detect an imposter. And a response such as "I could not find that message", "there is no such meeting" or "you do not own anyshares in that mutual fund" means that the user does not understand his previous interaction with the system, and the The user is likely to be an impostor. Similarly, some users are very strict and may issue commands such as "send this to Steve" that may not require additional explanation, but other users may not be very clear and will send the same commands as "send this to Steve", so The command may require additional dialogue to eliminate ambiguity. The system may use a question in the form of "do you mean Steve Jonesor Steve Brown?" to prompt the user. In this case, an imposter who is more rigorous or unclear than the actual user can be detected.
In order to use the system response as a feature in the feature vector, the standard system response can be put into different categories (negative response, positive response, confirmation, description, etc.). When generating a response, the identification of the category can be used as a feature To input.
8) Multi-modal interaction model: For those systems that support multi-modal input (voice, keyboard, mouse, handwriting, gestures, etc.), the input modalities commonly used by users to complete a task can be analyzed according to the present invention Combine and detect an impostor who uses different sets of input modalities for the same task. For example, some users may like to click the "Save" button to save a file, while others may prefer to use a voice command for this task. Therefore, it is very useful to add the input mode used to complete a certain task as an additional feature in the feature vector.
9) Counter-party status: Some systems may allow the user to have multiple open transactions at any given time (the user does not need to complete a task before moving to the next task). In this case, you can add features that represent the number of currently open transactions and the time elapsed since the earliest transaction started. And you can use this information again to construct feature vectors that represent a specific user. The conversation state can also include the type or duration of active actions performed on the system. For example, when logging into a system and then checking stock prices, a particular user can always access email.
10) Duration between commands: Different users may interact with the system at different rates. However, in the duration between commands, such as the time when a user pauses between commands, a designated user often shows regularity. Therefore, the duration between the end of the last command and the beginning of the current command can be explicitly entered as a feature.
All of the above characteristics describe how the user interacts with the system. Additional features that characterize how a given user behaves can also be used. These additional features that can be used can also be characteristics of how a given user conducts an activity. For example, when the system is initialized, these additional features can be trimmed by the user system and added to the feature vector v. The conversation system 114 provides all the data needed to calculate v.
The feature extractor 204 extracts a feature vector v for each input from the user, and sends it to the behavior data storage 206 and the probability calculator 210. The behavior data storage 206 is used to store all the feature vectors collected for a certain user, and the model builder 208 uses the behavior data storage 206 to construct a behavior model 214 for each approved user. In an embodiment of the present invention, a simple behavior model is constructed, which contains only the mean vector m and the covariance matrix Σ of the feature vector set (v's). In this case, when a sufficient number of samples of the feature vector v are collected, the model builder 208 calculates the mean vector m and the covariance matrix Σ for the designated user. When sufficient additional feature vectors are collected, the process will be repeated periodically. The mean vector m and the covariance matrix Σ are stored in the behavior model 214. In the art, the calculation of the mean vector and the covariance matrix are known. The feature vector is continuously collected, and the behavior model 214 is updated at periodic intervals to adapt it to any gradual changes in user behavior.
Then, for example, based on the behavior model 214, the probability calculator 210 calculates the probability P given by the following equation: P=e-12(vm)TΣ-1(vm)(2π)n2|Σ|12 ---(2)]]> The probability describes the likelihood that the specified input from the correct user can have. A higher value of P will correspond to a greater likelihood of input from a correct or approved user.
The probability mixer 212 obtains the probability score P and performs two steps. First, it calculates the weighted average of the probability score P from equation (2) for the current input and the selected number of previous inputs. If the probability score of the current input is expressed as P(t), and for i=1,...,m, the score of the i-th previous input is expressed as P(ti), where m is the consideration Enter the total number first, then the probability mixer 212 can calculate the cumulative behavior score Pb(t) at the current moment, the score is given by the following equation: Pb(t)=αtP(t)+αt-1P(t-1 )+...+αt-m(tm) (3) where the non-negative weight α satisfies αt+αt-1+......+αt-m=1 and αtαt-1αt-m0. The value of m is a system parameter that determines the number of previous probability scores considered and can be selected by the system designer. The intent of calculating the average over several scores is to ensure that no single false score will result in a false judgment.
The second step performed by the probability mixer 212 is to further mix the behavior score Pb(t) and the acoustic score Pα(t) for the current input provided by the acoustic and biometric checker 112 (FIG. 1). The acoustic score Pα(t) can be the standard acoustic score used in speaker verification, and if the current user input is in spoken form, it can be calculated using voice samples from the current user input (if the current input is not in spoken language Form, other approximations can be used, such as setting Pα(t)=Pα(t-1), or the approximation of the acoustic score starting from the most recent past input). The probability mixer 212 uses the following equation to calculate Ptotal(t)Ptotal(t)=βaPα(t)+βbPb(t) (4) where the non-negative weighting satisfies βi satisfies βa+βb=1, and the weighting can be determined by the system The designer chooses and can also be modified later by the user according to his or her preferences.
The probability mixer 212 compares the value of Ptotal(t) with a predetermined threshold Pth, and if Ptotal(t)<Pth, it sends a message to the acoustic and biometric checker 112 that the user may be an imposter. In one embodiment, the acoustic and biometric checker 112 will interrupt the user and ask the user to perform a more comprehensive verification process. If the additional verification fails, the user is no longer allowed to use the system. If the additional verification is successful, the user is allowed to interact with the system until the probability mixer 212 generates a future warning message.
In another embodiment, the user is allowed to continue to interact with the system, but the user is denied access to the system's sensitive materials. Material sensitivity may include a rating and the access rating for sensitive materials may be based on a score involving a threshold. For example, a group of employees may be allowed to access a system, however, certain employees must be excluded from sensitive materials. Employee behavior can be used to exclude unauthorized employees from sensitive materials.
The threshold Pth is a system parameter that can be selected by the system designer. However, depending on the expected performance level, the threshold can also be modified by the user.
Another embodiment of the present invention will now be described. The model builder 208 constructs two or more models and saves the collection of models in the behavior model 214. In order to construct each of these models, first use any standard clustering algorithm to divide the set of feature vectors v into multiple clusters, and the clustering algorithm may be, for example, the well-known K-means clustering algorithm. For each cluster i, the mean vector mi and the covariance matrix i will be calculated, and equation (2) will be modified to P=maxi[e-12(v-mi)TΣi-1(v-mi )(2π)n2|Σi|12]---(5)]]>Equations (3) and (4) remain the same, but they will be calculated from the above equation (5) The value of P. For example, the purpose of constructing feature vector clusters is to accommodate different behaviors displayed by the same user in different periods corresponding to different devices or different tasks used. Therefore, clusters can be explicitly constructed based on interaction-related factors, such as the accessed applications (email, calendar, stock trading, etc.), access devices (phones, cellular phones, notebook computers, desktop computers, personal digital assistants, etc.) ) Or other factors instead of using a clustering algorithm.
The preferred embodiments of the system and method for non-interference speaker verification using behavioral models (the models are illustrative and not restrictive) have been described here, but it should be pointed out that, according to the above teachings, the technology in the field Personnel can make various modifications and changes. Therefore, it should be understood that various changes can be made in the disclosed specific embodiments, and the changes are included in the spirit and scope of the present invention as summarized in the appended claims. Therefore, the present invention is described in combination with the detailed information and characteristics required by the patent law, in which the patent certificate statement and the content to be protected are stated in the appended claims.
2 sheets
Sheet 1 Sheet 2
Every citation, both ways
| Document | Relation | Office | Cited during |
|---|---|---|---|
| CN100437577C | Cited by | China | Search report |
| CN111462733A | Cited by | China | Search report |
| CN104954343A | Cited by | China | Search report |
| CN103738295A | Cited by | China | Search report |
| CN103019378A | Cited by | China | Search report |
| CN105489218A | Cited by | China | Search report |
15 members in 7 offices
Priority claims3
| Document | Office | Kind | Date |
|---|---|---|---|
| 0147910 | United States of America | W | |
| 0147910 | United States of America | W | |
| WO2001US47910 | – | – | – |
Members15
| Document | Office | Kind | |
|---|---|---|---|
| US6490560B1 | United States of America | B1 | |
| US2003046072A1 | United States of America | A1 | |
| WO03050799A1 | World Intellectual Property Organization (WIPO) | A1 | |
| AU2002230762A1 | Australia | A1 | |
| KR20040068548A | Republic of Korea | A | |
| CN1522431AThis record | China | A | |
| EP1470549A1 | European Patent Office (EPO) | A1 | |
| JP2005512246A | Japan | A | |
| CN1213398C | China | C | |
| EP1470549A4 | European Patent Office (EPO) | A4 | |
| JP4143541B2 | Japan | B2 | |
| WO03050799A9 | World Intellectual Property Organization (WIPO) | A9 | |
| AU2002230762A8 | Australia | A8 | |
| US7689418B2 | United States of America | B2 | |
| EP1470549B1 | European Patent Office (EPO) | B1 |
6 legal events, as the office reported them to INPADOC
Over the term
Point at a mark for the eventEvents
| Event | Code | |
|---|---|---|
| Expiry of patent termCX01 | CX01 | |
| Succession or assignment of patent rightASS | ASS | |
| Transfer of patent application or patent right or utility modelC41 | C41 | |
| Grant of patent or utility modelGrantedC14 | C14 | |
| Entry into substantive examinationC10 | C10 | |
| PublicationC06 | C06 |
Numbers
- Publication
- 1522431
- Publication, DOCDB
- 1522431
- Publication, EPODOC
- CN1522431
- Application
- 18234100
- Application, DOCDB
- 01823410
- Application, EPODOC
- CN20018003410
Titles3
- Chinese
- 使用行为模型来进行无干扰的说话者验证的方法和系统
- English
- Method and system for non-interference speaker verification using behavior model
- Chinese
- 使用行为模型来进行无干扰 的说话者验证的方法和系统
Classification
- CPC, 6
- G06F21/32
- G10L17/22
- G06F21/316
- G10L17/24
- G10L15/18
- G10L25/51
- IPC, 5
- G06F15 00
- G06F21 31
- G06F21 32
- G06F40 00
- G10L17 00