Method for disambiguating features in unstructured text
20 claims: 2 independent, 18 dependent
- 1サーチ質問をエンドユーザ装置から受信することに応答して、 システムのノードにより、1つ以上の抽出された特徴に一致する1つ以上の候補レコードを識別するために共起特徴を含む候補レコードのセットをサーチし、候補レコードに一致する抽出された特徴は、一次特徴であり、前記ノードはインメモリデータベースをホストするメインメモリを含み、前記インメモリデータベースはクラスターの知識ベースを格納し、各クラスターは独特の識別子(独特のID)を伴う曖昧性除去された一次特徴及び関連付けられた二次特徴を含み、 ノードにより、抽出された特徴の各々を1つ以上のマシン発生トピック識別子(トピックID)と関連付け、 ノードにより、トピックIDの関連度に基づき一次特徴の各々を互いに曖昧性除去し、 ノードにより、トピックIDの関連度に基づき各一次特徴に関連した二次特徴のセットを識別し、 ノードにより、トピックIDの関連度に基づき二次特徴の関連セットにおける二次特徴の各々から一次特徴の各々を曖昧性除去し、 前記インメモリデータベースから前記知識ベースからデータが検索されるときに、リアルタイムで、ノードにより、各一次特徴を二次特徴の関連セットにリンクして、新たなクラスターを形成し、 ノードのインメモリデータベースの曖昧性除去モジュールにより、曖昧性除去された一次特徴を伴う既存の知識クラスターへの比較的一致するスコアの指定により前記新たなクラスターの各々が既存の知識ベースクラスターに一致するかどうか決定し、 一致があるときには、既存の知識ベースクラスターにおける各一致する一次特徴に対応する既存の独特のIDを決定しそして前記新たなクラスターを含むように既存の知識ベースクラスターを更新し、及び 一致がないときには、新たな知識ベースクラスターを生成し、そしてその新たな知識ベースクラスターの一次特徴に新たな独特のIDを指定し、及び 既存の独特のID及び新たな独特のIDの一方を一次特徴として前記エンドユーザ装置へ送出する、ことを含む方法。
- 2ノードにより、抽出された特徴に一致する候補レコードの各々を比較し、及びノードにより、その比較に基づいて前記抽出された特徴の各々に重み付けされた一致スコア結果を指定する、ことを更に含む、請求項1に記載の方法。
- 3ノードにより、抽出された特徴の各々を、重み付けされた特徴属性のセットに関連付けることを更に含む、請求項2に記載の方法。
- 4ノードにより、1つ以上の重み付けされた特徴属性に基づいて抽出された特徴の各々の関連度を決定することを更に含む、請求項3に記載の方法。
- 51つ以上の抽出された特徴をノードの抽出モジュールにより確認及び抽出し、1つ以上の抽出された特徴において1つ以上の一次特徴を識別し、及び ノードの抽出モジュールにより、抽出された特徴の各々をデータベースに記憶する、ことを更に含む、請求項1に記載の方法。
- 6ノードの抽出モジュールにより、各特徴に抽出確度スコアを指定することを更に含む、請求項5に記載の方法。
- 7各々の一次特徴は、1つ以上の特徴属性のセットに関連付けられる、請求項1に記載の方法。
- 8特徴属性は、トピックID、ドキュメント識別子(ドキュメントID)、特徴タイプ、特徴名、信頼性スコア、及び特徴位置より成るグループから選択される、請求項7に記載の方法。
- 9各関連特徴は、予め定義されたクラスターハイアラーキーに従って下位順序特徴のセットに関連付けられる、請求項1に記載の方法。
- 10ノードにより、候補レコードのセットの曖昧キーサーチを遂行することを更に含む、請求項1に記載の方法。
- 11ノードのリンクオンザフライモジュールにより、関連トピックIDの共起及び1つ以上の特徴属性に基づいて2つ以上のデータソースをリンクすることを更に含む、請求項7に記載の方法。
- 12ノードにより、データソースにおける抽出された特徴が第2データソースにおいて共起するかどうかを、その抽出された特徴を第2データソースにおける特徴と比較することで決定し、及び ノードにより、前期比較に基づいてデータソースの各々をリンクする、ことを更に含む、請求項1に記載の方法。
- 13ノードにより、異なるデータソースからの抽出された特徴の共起を分析して、抽出された特徴の曖昧性除去の精度を改善することを更に含む、請求項1に記載の方法。
- 14ノードにより、1つ以上の新たなデータソースを連続的に受け取り、 ノードにより、1つ以上の抽出される特徴を連続的に抽出し、 ノードにより、1つ以上の抽出された特徴において候補サーチを連続的に遂行し、 ノードにより、抽出された特徴を連続的に曖昧性除去し、及び ノードにより、抽出された特徴を1つ以上の新たなクラスターへ連続的にリンクする、ことを更に含む、請求項1に記載の方法。
- 15コンピュータ実行可能なインストラクションが記憶された非一時的コンピュータ読み取り可能な媒体であって、プロセッサによって実行されると、 サーチ質問をエンドユーザ装置から受信することに応答して、 システムのノードにより、1つ以上の抽出された特徴に一致する1つ以上の候補レコードを識別するために共起特徴を含む候補レコードのセットをサーチし、前記ノードはインメモリデータベースをホストするメインメモリを含み、前記インメモリデータベースはクラスターの知識ベースを格納し、各クラスターは独特の識別子(独特のID)を伴う曖昧性除去された一次特徴及び関連付けられた二次特徴を含み、 ノードにより、抽出された特徴の各々を1つ以上のマシン発生トピック識別子(トピックID)と関連付け、 ノードにより、トピックIDの関連度に基づき一次特徴の各々を互いに曖昧性除去し、 ノードにより、トピックIDの関連度に基づき各一次特徴に関連した二次特徴のセットを識別し、 ノードにより、トピックIDの関連度に基づき二次特徴の関連セットにおける二次特徴の各々から一次特徴の各々を曖昧性除去し、 前記インメモリデータベースから前記知識ベースからデータが検索されるときに、リアルタイムで、ノードにより、各一次特徴を二次特徴の関連セットにリンクして、新たなクラスターを形成し、 ノードのインメモリデータベースの曖昧性除去モジュールにより、曖昧性除去された一次特徴を伴う既存の知識クラスターへの比較的一致するスコアの指定により前記新たなクラスターの各々が既存の知識ベースクラスターに一致するかどうか決定し、 一致があるときには、既存の知識ベースクラスターにおける各一致する一次特徴に対応する既存の独特のIDを決定しそして前記新たなクラスターを含むように既存の知識ベースクラスターを更新し、及び 一致がないときには、新たな知識ベースクラスターを生成し、そしてその新たな知識ベースクラスターの一次特徴に新たな独特のIDを指定し、及び 既存の独特のID及び新たな独特のIDの一方を一次特徴として前記エンドユーザ装置へ送出する、ことを含む機能が実行される、コンピュータ実行可能なインストラクションが記憶された非一時的コンピュータ読み取り可能な媒体。
- 16前記インストラクションは、更に、ノードにより、抽出された特徴に一致する候補レコードの各々を比較し、及びノードにより、その比較に基づいて前記抽出された特徴の各々に重み付けされた一致スコア結果を指定する、ことを含む、請求項15に記載の非一時的コンピュータ読み取り可能な媒体。
- 17前記インストラクションは、更に、ノードにより、抽出された特徴の各々を、重み付けされた特徴属性のセットに関連付けることを含む、請求項16に記載の非一時的コンピュータ読み取り可能な媒体。
- 18前記インストラクションは、更に、ノードにより、1つ以上の重み付けされた特徴属性に基づいて抽出された特徴の各々の関連度を決定することを含む、請求項17に記載の非一時的コンピュータ読み取り可能な媒体。
- 19前記インストラクションは、更に、 ノードの抽出モジュールにより、1つ以上の抽出された特徴を確認し及び抽出し、その1つ以上の抽出された特徴において1つ以上の一次特徴を識別し、及び ノードの抽出モジュールにより、抽出された特徴の各々をデータベースに記憶する、ことを含む、請求項15に記載の非一時的コンピュータ読み取り可能な媒体。
- 20前記インストラクションは、更に、ノードの抽出モジュールにより、各特徴に抽出確度スコアを指定することを含む、請求項19に記載の非一時的コンピュータ読み取り可能な媒体。
Independent claims20
83 paragraphs, as filed
The present invention generally relates to data management, and more particularly to data management systems and methods for extracting and storing material from source items received over a network.
Searching for information about an entity (ie, people, location, organization) in a large collection of documents, including sources such as networks, is often ambiguous, with inaccurate text processing capabilities, inaccurate features during knowledge extraction. It can lead to inaccurate data analysis.
The latest system, a large number of Argo, such as PageRank and hyperlink-induced topic search (HITS) algorithm are using clustering and ranking of linkage base in rhythm. The basic idea behind this and related solutions is that existing links typically exist between related pages or concepts. The limitation of clustering-based technology is that sometimes the context does not have the context information needed to deambiguate an entity, resulting in inaccurate decontamination results. Similarly, documents about different entities in the same or superficially similar context may be inaccurately clustered.
Other systems attempt to disambiguate an entity by referencing one or more external dictionaries (or knowledge bases) of the entity. In such a system, the entity's context is compared to a possible match entity in the dictionary, and the closest match is returned. The constraints associated with current dictionary-based technologies arise from the ever-increasing number of entities and therefore the lack of a dictionary containing representations of all entities in the world. Therefore, if the document context matches a dictionary entity, this technique will only identify the most similar entities in the diskionary, not necessarily the correct entity outside the dictionary.
<p num="0005"> Most methods use only entities and key phrases in the decontamination process. Therefore, there is still a need for an accurate entity deambiguation technique capable of performing accurate data analysis.</p>
<p num="0006"> In one embodiment, a method of disambiguating a feature is described. The method includes multiple modules such as one or more feature extraction modules, one or more deambiguation modules, one or more scoring modules, and one or more linking modules.</p><p num="0007"> Feature disambiguation is partially supported by extracting topics from the feature's surrounding documents and using the multi-component extension of the Latent Dirichlet Allocation (MC-LDA) topic model. Here, each component is modeled for each secondary feature stored in an existing knowledge base or extracted in an incoming document. In addition, the linking or disambiguation process is modeled as topic inference from MC-LDA, which gives automatic weight estimation during MC-LDA training and easily applies it during inference.</p><p num="0008"> This normative method improves the accuracy of entity deambiguation beyond what can be achieved without document linking considerations. Good deambiguation can be achieved by considering the document linkage and the document and entity relationships implied by the link.</p><p num="0009"> In one embodiment, the method searches a set of candidate records by a node of the system hosting the in-memory database to identify one or more candidates that match one or more extracted features and makes a candidate. Matching extracted features are primary features; by node, each of the extracted features is associated with one or more machine-generated topic identifiers (topic IDs); by node, primary features based on topic ID relevance. Disambiguate each other of each other; the node identifies the set of secondary features associated with each primary feature based on the relevance of the topic ID; the node identifies the set of secondary features associated with the topic ID relevance. Disambiguate each of the primary features from each of the secondary features in; nodes link each primary feature to a related set of secondary features to form a new cluster; nodes create new clusters Determines if it matches the knowledge-based cluster, and if there is a match, the in-memory database server computer's decontamination module provides an existing unique identifier (unique ID) for each matching primary feature in the knowledge-based cluster. And update the knowledge-based cluster to include the new cluster, and when there is no match, the node creates a new knowledge-based cluster and new to the primary features of the new knowledge-based cluster. Includes specifying a unique ID; and sending the node either an existing unique ID or a new unique ID as a primary feature;</p><p num="0010"> In another embodiment, a computer-executable instruction stored on a non-temporary computer-readable medium is one or more that matches one or more extracted features by the node of the system hosting the in-memory database. A set of candidate records is searched to identify candidates for, and the extracted features that match the candidates are primary features; each of the extracted features by a node is one or more machine-generated topic identifiers (topics). Associate with ID); Node disambiguates each of the primary features based on topic ID relevance; Node identifies a set of secondary features associated with each primary feature based on topic ID relevance. The node deambiguates each of the primary features from each of the secondary features in the secondary feature association set based on the topic ID relevance; the node links each primary feature to the secondary feature association set. The node determines if the new cluster matches the existing knowledge-based cluster, and if there is a match, the node corresponds to each matching primary feature in the knowledge-based cluster. Determine a unique identifier (unique ID) and update the knowledge-based cluster to include the new cluster, and if there is no match, generate a new knowledge-based cluster, and of the new knowledge-based cluster Includes specifying a new unique ID for the primary feature; and sending one of the existing unique ID and the new unique ID as the primary feature by the node;</p><p num="0011"> Additional features and effects of one embodiment will be described in the description below, and some will be apparent from that description. The object and other effects of the present invention are realized and achieved by the normative embodiments described below, the scope of claims and the structures specifically pointed out in the accompanying drawings.</p><p num="0012"> It should be understood that both the above general description and the following detailed description are merely normative explanations and provide further explanation of the invention as set forth in the claims.</p><p num="0013"> The present disclosure can be better understood by reference to the accompanying drawings. The accompanying drawings constitute a portion of the present specification, show embodiments of the present invention, and describe the present invention together with the specification. The components in the drawings are not necessarily on the correct scale, but rather are emphasized in exemplifying the principles of the present disclosure. In the drawings, reference numbers indicate corresponding parts throughout the different drawings.</p>
<figref num="1">It is a flowchart of a method of deambiguating a feature in an unstructured text by a normative embodiment.</figref><figref num="2">FIG. 6 is a flow chart of the steps performed by the deambiguation module used in the method of deambiguating a feature, according to a normative embodiment.</figref><figref num="3">FIG. 5 is a flow chart of steps performed by a link-on-the-fly module used in a method of disambiguating a feature, according to a normative embodiment.</figref><figref num="4">FIG. 5 illustrates a system used to implement a method of disambiguating a feature by a normative embodiment.</figref><figref num="5">A graphical representation of a multi-component conditional independent Latent Dirichlet Allocation (MC-LDA) topic model according to a normative embodiment.</figref><figref num="6">A multi-component conditional independent latency allocation (MC-LDA) topic model with a normative embodiment shows an embodiment of the Gibbs sampling equation.</figref><figref num="7">A multi-component conditional independence rate allocation (MC-LDA) topic model with normative embodiments shows an embodiment of a probability change inference algorithm for training and inference.</figref><figref num="8">A table showing sample topics for a multi-component conditional independent Latent Dirichlet Allocation (MC-LDA) topic model according to normative embodiments.</figref>
<u style="single">Definition</u> The following terms used herein have the following definitions.
"Document" refers to a separate electronic representation of information that has a starting point and an ending point.
"Multi-document" refers to a document with tokens, different forms of named entities, and key phrases organized into separate "bag-of-surface-forms" components.
"Database" refers to a system that contains a combination of clusters and modules that are suitable for storing one or more aggregates and for processing one or more questions.
"Corpus" refers to a collection of one or more documents.
A "raw corpus" or "document stream" refers to a corpus that is constantly supplied when a new document is uploaded to the network.
"Features" are information that is at least partially derived from a document.
A "feature attribute" refers to metadata associated with a feature, such as, among other things, the location of the feature in a document, a confidence score.
A "cluster" refers to a collection of features.
"Something knowledge base" refers to a base that contains features / entities.
"Link on the fly module" or "link OTF" refers to a linking module that updates data as the raw corpus is updated.
"Memory" refers to a hardware component suitable for storing and retrieving information at a sufficiently high speed.
"Module" refers to a computer software component suitable for performing one or more defined tasks.
"Sentiment" refers to an objective assessment associated with a document, part of a document, or feature.
A "topic" refers to a set of semantic information that is at least partially derived from a corpus.
The "topic identifier" or "topic ID" is an identifier that points to a specific instance of a topic.
A "topic collection" refers to a particular set of topics derived from a corpus, each topic having a unique identifier (unique ID).
"Topic classification" refers to specifying a particular topic identifier as a feature of a document.
A "question" refers to a request to retrieve information from one or more suitable databases.
<u style="single">Detailed explanation</u> The preferred embodiments shown in the accompanying drawings will be described in detail below. The above-described embodiment is merely an example. Those skilled in the art will recognize that the particular embodiments described herein can be replaced by a number of other components and embodiments within the scope of the present invention.
The present disclosure describes a method of deambiguating features in unstructured text. Although normative embodiments describe the practice of disambiguating features in accordance with the present disclosure, it is intended that the systems and methods described herein can be configured to be used appropriately within the scope of the present disclosure.
Existing knowledge bases include unambiguous features and related features, which leads to unreliable text analysis. The viewpoint of the present disclosure enhances the accuracy of disambiguation of features and entities, and therefore enhances the accuracy of text analysis.
According to one embodiment, the method disclosed here for disambiguating features is used in the initial data corpus to perform document capture and feature extraction and topic classification for each document contained in the initial corpus. And other text analysis. Each feature is identified and recorded, among other things, as a document name, type, location, and confidence score.
FIG. 1 is a flowchart of Method 100 showing a plurality of steps to deambiguate a feature in unstructured text. According to one embodiment, feature deambiguation method 100 begins when a new document entry step 102 is performed in the existing knowledge base. Feature extraction step 104 is performed on the document. According to one embodiment, features relate, among other things, to different feature attributes such as topic identifier (topic ID), document identifier (document ID), feature type, feature name, reliability score and feature location. ing.
According to various embodiments, the document input in step 102 is sourced from a mass corpus or a raw corpus (such as a corpus of internet or network connections), which is then fed every second.
According to different embodiments, during feature extraction step 104, one or more feature verification and extraction algorithms are used to analyze the unstructured text of document input step 102. A score is specified for each extracted feature. The score indicates the accuracy level of the features that are correctly extracted with the correct attributes.
In addition, during feature extraction step 104, one or more primary features are identified from the document input in step 102. Each primary feature is associated with a set of feature attributes and one or more secondary features. Each secondary feature is associated with a set of feature attributes. In certain embodiments, the one or more secondary features have one or more tertiary features, each with its own set of feature attributes.
The relative weight or relevance of each feature in the document input of step 102 is determined taking into account the feature attributes. In addition, a weighted scoring model is used to determine the relevance of associations between features.
Following feature extraction step 104, the features extracted from the document input in step 102 and all related information are as part of feature disambiguation request step 108 while including the features in MemDB in step 106. Loaded into an in-memory database (MemDB).
In one embodiment, MemDB forms part of a disambiguation computer server environment with one or more processors performing the steps described in relation to FIGS. 1-8. In one embodiment, the MemDB is a computer module that includes one or more search controllers, multiple search nodes, a collection of compressed data, and an ambiguous submodule. A search controller is selectively associated with one or more search nodes. Each search node can independently perform an ambiguous key search through a collection of compressed data and return a set of scored results to its associated search controller.
Feature disambiguation step 108 is performed by the deambiguation submodule in MemDB. The feature disambiguation 108 process includes a machine-generated topic ID, which is used to classify features, documents, or corpora. The relevance of individual features and specific topic IDs is determined using a disambiguation algorithm. In a document, the same feature is associated with one or more topic IDs based on the context of different occurrences of the feature within the document.
A set of features (same topic, approaching terms and entities, key phrases, events and facts) extracted from a document is when two or more features across different documents are a single feature, or they are separate features. If, then it is compared to a set of features from other documents using a decontamination algorithm defined at a certain precision level. In one example, the co-occurrence of two or more features across a collection of documents in a database is analyzed to improve the accuracy of the feature deambiguation process 108. In one embodiment, an overall scoring algorithm is used to determine the probability that the features are the same.
In one embodiment, a knowledge base is generated within MemDB as part of the feature deambiguation process 108. This knowledge base is used to temporarily store a cluster of associated deambiguous primary features and associated secondary features. When a new document is loaded into MemDB, the new deambiguous feature set is compared to the existing knowledge base to determine the relationship between features and features, and the new features and already extracted features. Determine if there is a match with.
If the compared features match, the knowledge base is updated, the feature IDs of the matching features are returned to the user and / or the requesting application or process, and further, prominent means based on the frequency of matching. It can be attached with a feature ID, which captures its popularity index in a given corpus. If the compared features do not match any of the features already extracted, then the deambiguous entity or feature is given a unique feature ID, and that unique feature ID is associated with the cluster that defines the feature. And stored in the MemDB knowledge base. Then, in step 110, the feature ID of the deambiguous feature is returned to the source through the system interface. In certain embodiments, the feature ID of the disambiguated feature comprises a secondary feature, a cluster of features, an associated feature attribute, or other required data. Features The disambiguation submodule used for deambiguation step 108 is described in detail below with reference to FIG.
<u style="single">Ambiguity removal submodule</u> FIG. 2 is a flow chart of process 200 performed by the deambiguation submodule used in the unstructured text of feature deambiguation step 108 of method 100 (FIG. 1), according to one embodiment. The deambiguation process 200 begins after the MemDB has been featured in step 106 of FIG. The extracted features given in step 202 are used to perform a candidate search in step 204, and a search for the extracted features is performed through all candidate records, including co-occurrence features.
According to various embodiments, the candidate is a primary feature with a set of related secondary features used in the feature disambiguation process 108.
The disambiguation result is improved by the co-occurrence of the topic ID and the relevance within the topic ID. The relevance of topic IDs, even across different topic models, can be found in a large corpus with topic IDs. The related topic ID can be used during record linkage step 206 to link to a document that does not contain the exact topic ID but contains one or more related topic IDs. This solution improves the recall of related features that should be included in record linkage step 206, and in some cases improves the disambiguation results.
Once a set of potentially related documents has been identified and the relevant primary and secondary features within those documents have been extracted, the attributes of the feature, between the features and features of the same document (meaningful context). Relationships, relative weights of features, and other variables are used during the record linkage process 206 to deambiguate primary and secondary features across those documents. Each record is then linked to another record to determine a cluster of deambiguous primary features and their associated secondary features. The algorithm used for record linkage 206 can overcome spelling errors or transliteration and other challenges in mining unstructured datasets.
Cluster comparison step 208 includes designating relatively matching scores for clusters of deambiguated features, defining different acceptance thresholds for different applications. The defined accuracy level determines which scores are considered positive match searches and which scores are considered negative match searches (step 210). Each new cluster is given a unique ID and is temporarily stored in the knowledge base. Each new cluster contains a new set of deambiguous primary and secondary features. If the new cluster matches a cluster already stored in the knowledge base, the system updates the knowledge base in step 212 and returns a matching feature ID to the user and / or requesting application or process. Is carried out in step 214. Knowledge base update 212 means associating an additional secondary feature with one primary feature, or adding a feature attribute that was not previously associated with a primary or secondary feature.
If the evaluated cluster is given a score lower than the threshold of the positive match search 210, the system performs an ID designation unique to the primary feature of the cluster in step 216, and in step 212. , Update the knowledge base. The system then performs match ID return process 214. Record linkage step 206 will be described in more detail with reference to FIG.
<u style="single">Link on the fly submodule</u> FIG. 3 is a flow chart of process 300 performed by a link-on-the-fly (link OTF) submodule used in method 100 for disambiguating features, according to one embodiment. The Link OTF Process 300 can constantly evaluate, score, link, and cluster information feeds. Link OTF submodules use multiple algorithms to perform record linkage 206. The candidate search results of step 204 are constantly fed to the link OTF module 300. Following data entry, a match scoring algorithm is applied (step 302), where one or more match scoring algorithms are applied simultaneously to multiple search nodes in MemDB, while, among other things, the string edit distance. Perform ambiguous key searches to evaluate and score relevant results, taking into account multiple feature attributes such as phonemes and meanings.
A linking algorithm application step 304 is then added to compare all candidate records identified during the match scoring algorithm application step 302 with each other. Application of Linking Algorithm 304 includes the use of one or more analytical linking algorithms that can filter and evaluate the scored results of ambiguous key searches performed within MemDB's multiple search nodes. In one example, the co-occurrence of two or more features across a collection of identified candidate records in MemDB is analyzed to improve the accuracy of the process. Application 304 of the linking algorithm considers different weighted models and reliability scores associated with different feature attributes.
After applying the linking algorithm step 304, the linked result is placed in a cluster of related features, and in step 306 it is returned as part of the return of the cluster of linked records.
FIG. 4 is a diagram illustrating an embodiment of the system 400 that disambiguates features in the unstructured text described above with reference to FIG. This system 400 hosts an in-memory database and contains one or more nodes.
According to one embodiment, the system 400 performs computer instructions for multiple special purpose computer modules 401, 402, 411, 412 and 414 (described below) to disambiguate features in one or more documents. Has one or more processors. As shown in FIG. 4, document input modules 401, 402 receive a document from an internet-based source and / or the raw corpus of the document. A large number of new documents are uploaded to the document input module 402 over the network connection 404. Therefore, the source is constantly updated with new knowledge by user workstation 406, and such new knowledge is not pre-linked in a static way. Therefore, the number of documents to be evaluated increases indefinitely.
This evaluation is achieved via MemDB computer 408. The MemDB408 facilitates a fast decontamination process and facilitates the decontamination process on the fly, which facilitates the reception of up-to-date information that is intended to contribute to the MemDB408. Various methods are used to link features, which essentially use a weighted model to determine which entity types are most important and which have the greater weights. , And, based on the confidence score, determine how reliable the correct feature extraction and deambigation was performed, and determine that the correct features are directed towards the resulting feature cluster. As shown in Figure 4, the more system nodes function in parallel, the more efficient the process becomes.
According to various embodiments, when a new document arrives at system 400 through document input modules 401, 402 and network connection 404, feature extraction is performed via extraction module 411, followed by feature disambiguation. Performed in a new document via the MemDB 408 feature decontamination submodule 414. In one embodiment, after the feature deambiguation of the new document has been performed, the extracted new feature 410 is included in the MemDB to pass through the link OTF submodule 412, where the features are compared. The feature ID of the feature 110 that has been, linked, and deambiguated is returned to the user as a result of the question. In addition to the feature ID, the resulting feature cluster that defines the deambiguous feature may be optionally returned.
The MemDB computer 408 is a database that stores data in a record controlled by a database management system (DBMS) (not shown) that is configured to store the data record in the device main memory. This is in contrast to traditional databases and DBMS modules that are stored in "disk" memory. Traditional disk storage requires the processor (CPU) to execute read and write commands to the device's hard disk, so instructions for the CPU to locate (ie seek) and retrieve memory locations for data. Is required to perform some form of operation with the data at that memory location after performing. An in-memory database system accesses data stored in main memory and appropriately addressed, thus reducing the number of instructions performed by the CPU, and seeks associated with the CPU seeking data on the hard disk. Eliminate time.
An in-memory database is implemented in a distributed computing architecture, which is a computing system that includes one or more nodes configured to aggregate each node's resources (eg, memory, disk, processor). As disclosed herein, an embodiment of a computing system that hosts an in-memory database distributes and stores data records in the database among one or more nodes. In some embodiments, these nodes are formed into "clusters" of nodes. In certain embodiments, these clusters of nodes store parts or "aggregates" of database information.
Various embodiments use an evolving and efficiently linkable feature knowledge base configured to memorize secondary features such as co-occurrence topics, key phrases, proximity terms, events, facts and trend popularity indices. Features of computer execution to provide disambiguation technology. The embodiments disclosed herein are elaborate graph clustered solutions from a simple conceptual distance scale based on the dimensions of related secondary features that help analyze a given extracted feature against a feature stored in the knowledge base. It is carried out through various linking algorithms that can change up to the decision. In addition, those embodiments extend the existing feature knowledge base with the ability to not only update the secondary features of the existing feature entry, but also extend it by discovering new features that can be added to the knowledge base. Evolving solutions can be introduced.
An embodiment of the disambiguation solution uses a topic modeling solution to provide an auto-weighting (over all secondary features) linking process that is modeled as topic inference. To support this auto-weighted linking process, those embodiments build a new topic modeling solution called Multi-Component LDA (MC-LDA) that can support a large number of components (secondary features) conditionally independently. Extend traditional LDA topic modeling to do so. Also, embodiments of the modeling solution can automatically learn component weights during training and use them for inference (linking) on disambiguation. The MC-LDA solution introduced for decontamination can be scaled for an additional number of secondary features that can be introduced to improve decontamination accuracy.
FIG. 5 is a graphic representation of an embodiment of the Multi-Component Conditional Independent Latent Dirichlet Allocation (MC-LDA) Topic Computer Modeling Solution used by System 400 in FIG. In the embodiments shown herein, each component block represents, for example, modeling of each secondary feature over a knowledge base, performed via MemDB408 in FIG. 4 initialized with the parameters shown in FIG.
FIG. 6 shows an embodiment of the Gibbs sampling equation of the MC-LDA topic model used in FIG. 5 described above. This embodiment of the sampling solution aids System 400 in FIG. 4 in automatically and efficiently training the weights of individual components (secondary features).
Figure 7 shows, for example, a stochastic change for training and reasoning in the MC-LDA topic model of Figure 5-6, which is performed via MemDB408 of System 400 of Figure 4, which is initialized with the parameters shown in Figure 7. An embodiment of computer execution of an inference algorithm is shown. An embodiment of this inference method models the linking / disambiguation process as topic inference by taking all secondary features (extracted from the document) as input and giving a weighted topic as output. Easily applied to. These weighted topics can then be used to calculate the similarity score for the stored feature knowledge base entries.
Figure 8 is a table showing sample topics for the MC-LDA topic model. FIG. 8 shows a top scoring surface form for each component of the model, executed, for example, through MemDB408 in system 400 of FIG. 4, according to one embodiment.
Example # 1 provides method 100 for disambiguating features in unstructured text when the feature (primary feature) is football player John Doe and the user wants to monitor news mentioning John Doe. It applies. According to one embodiment, document input 102 describing John Doe is uploaded to the network. Features of document input 102 are extracted, included in MemDB408, deambiguous, linked to a cluster of secondary features related to the primary feature (John Doe), and compared to an existing cluster of similar features. To. Method 100 outputs different feature IDs and associated clusters of feature IDs, which output all relevant secondary features to John Doe, such as engineer John Doe; teacher John Doe; and football player John Doe ;. Including. Other primary features with similar secondary features, such as nicknames or abbreviations, are possible. Football player John From the same team as Doe, "JD" football players of the same age and experience are considered to have the same primary characteristics. Therefore, all documents related to football player John Doe are easily accessible.
Example # 2 applies method 100 for deambiguating a feature in an unstructured document when the primary feature is an image. According to certain embodiments, method 100 comprises extracting features 104, where features are, among other things, general attributes such as edges and shapes, or above all, such as tanks, individuals and watches. Specific attribute. For example, a new image is entered, where the image has secondary features such as a particular shape (eg, square, personal or car shape), and the secondary features are extracted and included in MemDB408. Here, a match is found among all other images with similar secondary features. According to the embodiments shown herein, the feature includes only the image, i.e. the text is not included as the feature.
Example # 3 applies method 100 for deambiguating a feature in unstructured text when the primary feature is an event. According to one embodiment, when a question is asked, Method 100 allows the user to receive, among other things, results related to an event such as an earthquake, fire, or outbreak of an infectious disease. Method 100 performs feature extraction 104 and feature deambiguation 108 to find the event-related feature and give the feature ID of the deambiguous feature 110.
Example # 4 is an embodiment of Method 100 when one or more events are expected to occur. According to certain embodiments, the user pre-directs the feature and event prior to the operation and therefore links between different features associated with the event are pre-established. When the relevant features appear in the network with a high incidence, Method 100 expects the event to occur based on the increase in the number of related features. When an imminent event is detected, the user is alerted. For example, a user working for the Ministry of Health from Thailand chooses to receive an alarm about the outbreak of dengue infectious disease. For example, when another user 406 from a social network uploads a comment containing signs or inclusion of dengue fever to the hospital, Method 100 disambiguates all relevant comments from the social network and includes relevant information. Considering the number of users 406, anticipate the outbreak of dengue infectious disease and alert the Ministry of Health staff. Therefore, health ministry officials will gain additional evidence and take further steps to the affected communities to prevent the spread of infectious diseases.
Example # 5 is the application of Method 100 when the primary feature is the name of a geographical location. According to one embodiment, method 100 is used to deambiguate the name of a city, with different scoring weights associated with secondary features in the deambigation submodule. For example, Method 100 is used to deambiguate Paris, Texas from Paris, France.
Example # 6 is a social network where the primary feature is, among other things, emotions related to an individual, event, or company, and the emotions are, among other things, positive or negative comments about an individual, event, or company. It applies Method 100 for deambiguating features in unstructured text when sourced from a suitable source, including. According to one embodiment, Method 100 is used to ascertain the acceptability that the company has in the general public.
Example # 7 is an embodiment of Method 100 that includes human confirmation to increase the reliability score of the feature. According to one embodiment, the link OTF process 300 (Figure 4) is assisted by the user, who indicates whether the deambiguous features were correctly deambiguous, and has two different clusters. Indicates whether it must be, and this is what Method 100 dictates (taking into account all features and topic co-occurrence information) when two different primary features the user knows are the same. Means. Therefore, the reliability score associated with the cluster is high, and therefore the probability that the feature is correctly deambiguous is high.
Example # 8 is an embodiment of Method 100 using the deambiguation process 200 and the link OTF process 300. In this example, the linking algorithm used in Applying the Linking Algorithm 304 is configured to give a reliability score higher than 0.85 within a period of 1000 ms.
Example # 9 is an embodiment of Method 100 using the deambiguation process 200 and the link OTF process 300. In this example, the linking algorithm used in Applying the Linking Algorithm 304 is configured to give a reliability score higher than 0.80 within a period of 300 ms or less. The algorithm used in this example responds within a shorter period of time than the algorithm used in Example # 8, but generally returns a lower confidence score.
Example # 10 is an embodiment of Method 100 using the deambiguation process 200 and the link OTF process 300. In this example, the linking algorithm used in Applying the Linking Algorithm 304 is typically configured to give a reliability score higher than 0.90 within a period greater than 3000 ms. The algorithm used in this example generally gives a response with a higher confidence score than that returned by the algorithm used in Example # 8, but generally requires a significantly longer period of time.
Example # 11 is an example of 100 methods for deambiguating features in unstructured text to perform eDiscovery in a large corpus of documents from multiple sources. Given a large corpus of documents from multiple resources, the application of Method 100 to disambiguate all features in those documents allows all features to be found in the corpus. The collection of discovered features can be further used to discover all documents related to the feature and to discover related features.
The description of the above method and the process flow diagram are provided by way of example only and are not intended to require or imply that the steps of the various embodiments must be performed in the order presented. As will be apparent to those skilled in the art, the steps in the above embodiment may be performed in any order. Words such as "then", "next", etc. do not limit the order of the steps, and these words are simply used to guide the reader through a description of the method. Only. The process flow diagram shows operations as a series of processes, but many operations can be performed in parallel or simultaneously. In addition, the order of operations may be rearranged. Processes correspond to methods, functions, procedures, subroutines, subprograms, and so on. When a process corresponds to a function, its termination corresponds to the return of the function to the calling function or the main function.
The various exemplary logic blocks, modules, circuits and algorithm steps described in connection with the embodiments disclosed herein may be embodied as electronic hardware, computer software, or a combination thereof. To articulate this compatibility of hardware and software, various exemplary components, blocks, modules, circuits, and steps have been generally described with respect to their functionality. Whether such a function is embodied as hardware or software depends on the specific application and design constraints imposed on the entire system. Those skilled in the art can embody the functions described herein in various ways for each specific application, but the judgment of such embodying should not be construed as deviating from the scope of the present invention.
The embodiments embodied in computer software are embodied in software, firmware, middleware, microcode, a hardware description language, or a combination thereof. A code segment or machine-executable instruction represents a procedure, function, subprogram, program, routine, subroutine, module, software package, class, or combination of instructions, data structures, or program statements. A code segment is coupled to another code segment or hardware circuit by passing and / or receiving information, data, arguments, parameters or memory content. Information, arguments, parameters, data, etc. are passed, transferred or transmitted via appropriate means, including memory sharing, message passage, token passage, network transmission, etc.
The actual software code or special control hardware used to implement these systems and methods is not limiting the invention. Therefore, the operation and behavior of systems and methods has been described without reference to specific software code, with the understanding that software and control hardware can be designed to implement the systems and methods based on the description herein. ..
When implemented in software, functions are stored as one or more instructions or codes on a non-temporary computer-readable or processor-readable storage medium. The steps of the methods or algorithms disclosed herein are performed in processor executable software modules present in computer readable or processor readable storage media. Non-temporary computer-readable or processor-readable media include both computer storage media and tangible storage media that facilitate the transfer of computer programs from one location to another. A non-temporary processor-readable storage medium is an available medium accessed by a computer. By way of example, but not limited to, such non-temporary processor readable media are RAM, ROM, EEPROM, CD-ROM or other optical disk storage, magnetic disk storage or other magnetic storage devices, or Includes other tangible storage media used to store desired program code in the form of instructions or data structures and accessed by a computer or processor. Disk used here (disk & Discs include compact discs (CDs), laser discs®, optical discs, digital diversity discs (DVDs), floppy discs, and Blu-ray discs, where discs are typically data. On the other hand, a disc is for optically reproducing data with a laser. The above combinations are also included within the range of computer-readable media. In addition, the operation of the method or algorithm exists as one or a combination or set of codes and / or instructions on non-temporary processor readable media and / or computer readable media that are integrated into computer program products.
It is clear that the various components of the technology can be located in remote parts of distributed networks and / or the Internet, or within dedicated secure, unsecured and / or cryptographic systems. Therefore, it is clear that the components of the system can be combined into one or more devices or co-located on a particular node of a distributed network such as a telecommunications network. As is clear from the above description, for computational efficiency reasons, system components can be placed anywhere in the distributed network without affecting the operation of the system. In addition, those components can be embedded in a dedicated machine.
In addition, the various links connecting the elements are wired or wireless links or combinations thereof, or other known or upcoming elements capable of supplying and / or communicating data to and from the connected elements. It is clear that there is. The term module as used herein refers to known or upcoming hardware, software, firmware, or a combination thereof capable of performing a function associated with an element. Also, the terms decision, calculation and computing, and variants thereof used herein are used interchangeably and include any type of method, process, mathematical operation or technique.
The above description of the embodiments disclosed herein is made to enable those skilled in the art to practice or utilize the present invention. Various changes to these embodiments will be readily apparent to those of skill in the art, and the general principles defined herein apply to other embodiments without departing from the spirit or scope of the invention. Therefore, the present invention is not limited to the embodiments shown herein, but should be harmonized with the scope of claims and the broadest scope consistent with the principles and novel features disclosed herein.
The embodiments described above are merely examples. One of ordinary skill in the art will recognize a number of alternative components and embodiments that have been replaced with respect to the particular examples described herein and are still within the scope of the present invention.
400: System 401, 402: Document input module 404: Network connection 406: User workstation 408: MemDB computer 410: New features extracted 411: Extraction node 412: Link OTF submodule
8 sheets
Sheet 1 Sheet 2 Sheet 3 Sheet 4 Sheet 5 Sheet 6 Sheet 7 Sheet 8
Every citation, both ways
| Document | Relation | Office |
|---|---|---|
| US20020165847A1 | Cites | United States of America |
| US20130290665A1 | Cites | United States of America |
| JP2013239162A | Cites | Japan |
| JP2003150442A | Cites | Japan |
76 members in 10 offices
Priority claims9
| Document | Office | Kind | Date |
|---|---|---|---|
| 201361910739 | United States of America | P | |
| 201361910739 | United States of America | P | |
| 61910739 | United States of America | – | |
| 2014067918 | United States of America | W | |
| 2014067918 | United States of America | W | |
| 61910739 | – | – | – |
| US201361910739P | – | – | – |
| US2014067918 | – | – | – |
| WO2014US67918 | – | – | – |
Members76
| Document | Office | Kind | |
|---|---|---|---|
| WO2011080491A1 | World Intellectual Property Organization (WIPO) | A1 | |
| FR2955009A1 | France | A1 | |
| US2012282910A1 | United States of America | A1 | |
| CN102783219A | China | A | |
| EP2522183A1 | European Patent Office (EPO) | A1 | |
| EP2522183B1 | European Patent Office (EPO) | B1 | |
| ES2425097T3 | Spain | T3 | |
| PL2522183T3 | Poland | T3 | |
| US8948736B2 | United States of America | B2 | |
| US2015087286A1 | United States of America | A1 | |
| US9025892B1 | United States of America | B1 | |
| US2015154079A1 | United States of America | A1 | |
| US2015154194A1 | United States of America | A1 | |
| US2015154200A1 | United States of America | A1 | |
| US2015154233A1 | United States of America | A1 | |
| US2015154263A1 | United States of America | A1 | |
| US2015154264A1 | United States of America | A1 | |
| US2015154283A1 | United States of America | A1 | |
| US2015154286A1 | United States of America | A1 | |
| US2015154297A1 | United States of America | A1 | |
| CA2932399A1 | Canada | A1 | |
| CA2932402A1 | Canada | A1 | |
| WO2015084724A1 | World Intellectual Property Organization (WIPO) | A1 | |
| WO2015084726A1 | World Intellectual Property Organization (WIPO) | A1 | |
| WO2015084760A1 | World Intellectual Property Organization (WIPO) | A1 | |
| CA2932403A1 | Canada | A1 | |
| WO2015099961A1 | World Intellectual Property Organization (WIPO) | A1 | |
| US2015234899A1 | United States of America | A1 | |
| US9177254B2 | United States of America | B2 | |
| US9201744B2 | United States of America | B2 | |
| US9223875B2 | United States of America | B2 | |
| US9239875B2 | United States of America | B2 | |
| US2016019466A1 | United States of America | A1 | |
| US2016019470A1 | United States of America | A1 | |
| US2016098433A1 | United States of America | A1 | |
| US2016110446A1 | United States of America | A1 | |
| US9344879B2 | United States of America | B2 | |
| US2016140235A1 | United States of America | A1 | |
| US9348573B2 | United States of America | B2 | |
| US9355152B2 | United States of America | B2 | |
| US2016154714A1 | United States of America | A1 | |
| US2016162283A1 | United States of America | A1 | |
| US2016171057A1 | United States of America | A1 | |
| US2016196277A1 | United States of America | A1 | |
| US2016234679A1 | United States of America | A1 | |
| US9424294B2 | United States of America | B2 | |
| US9430547B2 | United States of America | B2 | |
| EP3077919A1 | European Patent Office (EPO) | A1 | |
| EP3077927A1 | European Patent Office (EPO) | A1 | |
| EP3077930A1 | European Patent Office (EPO) | A1 | |
| KR20160124742A | Republic of Korea | A | |
| KR20160124743A | Republic of Korea | A | |
| KR20160124744A | Republic of Korea | A | |
| CN106164890A | China | A | |
| CN106164897A | China | A | |
| US2016364471A1 | United States of America | A1 | |
| JP2016541069A | Japan | A | |
| US2017031788A9 | United States of America | A9 | |
| JP2017504874A | Japan | A | |
| CN106462575A | China | A | |
| JP2017505936A | Japan | A | |
| EP3077919A4 | European Patent Office (EPO) | A4 | |
| US9659108B2 | United States of America | B2 | |
| CN102783219B | China | B | |
| US9693221B2 | United States of America | B2 | |
| EP3077927A4 | European Patent Office (EPO) | A4 | |
| US9710517B2 | United States of America | B2 | |
| US9720944B2 | United States of America | B2 | |
| US2017220647A1 | United States of America | A1 | |
| CN107182049A | China | A | |
| EP3077930A4 | European Patent Office (EPO) | A4 | |
| US9785521B2 | United States of America | B2 | |
| CN107257543A | China | A | |
| JP6284643B2This record | Japan | B2 | |
| US9916368B2 | United States of America | B2 | |
| CA2932402C | Canada | C |
16 legal events, as the office reported them to INPADOC
Over the term
Point at a mark for the eventEvents
| Event | Code | |
|---|---|---|
| Receipt of annual feesJAPANESE INTERMEDIATE CODE: R250R250 | R250 | |
| Receipt of annual feesJAPANESE INTERMEDIATE CODE: R250R250 | R250 | |
| Receipt of annual feesJAPANESE INTERMEDIATE CODE: R250R250 | R250 | |
| Receipt of annual feesJAPANESE INTERMEDIATE CODE: R250R250 | R250 | |
| Written notification of registration of transferJAPANESE INTERMEDIATE CODE: R350R350 | R350 | |
| Request for change of ownership or part of ownershipJAPANESE INTERMEDIATE CODE: R313113S111 | S111 | |
| Written request for registration of change of domicileJAPANESE INTERMEDIATE CODE: R313531S531 | S531 | |
| Receipt of annual feesJAPANESE INTERMEDIATE CODE: R250R250 | R250 | |
| Certificate of patent or registration of utility modelJAPANESE INTERMEDIATE CODE: R150R150 | R150 | |
| First payment of annual fees (during grant procedure)JAPANESE INTERMEDIATE CODE: A61A61 | A61 | |
| Written decision to grant a patent or to grant a registration (utility model)JAPANESE INTERMEDIATE CODE: A01A01 | A01 | |
| Decision of grant or rejection writtenTRDD | TRDD | |
| Report on accelerated examinationJAPANESE INTERMEDIATE CODE: A971005A975 | A975 | |
| Request for written amendment filedJAPANESE INTERMEDIATE CODE: A523A521 | A521 | |
| Written request for application examinationJAPANESE INTERMEDIATE CODE: A621A621 | A621 | |
| Explanation of circumstances concerning accelerated examinationJAPANESE INTERMEDIATE CODE: A871A871 | A871 |
Numbers
- Publication
- 6284643
- Publication, DOCDB
- 6284643
- Publication, EPODOC
- JP6284643B
- Application
- 2016536850
- Application, DOCDB
- 2016536850
- Application, EPODOC
- JP20160536850
Titles2
- Japanese
- 非構造化テキストにおける特徴の曖昧性除去方法
- English
- How to disambiguate features in unstructured text
Classification
- CPC, 5
- G06F16/35
- G06F16/3344
- G16B40/00
- G16C20/70
- G06F40/279
- IPC, 2
- G06F17 30
- G06F40 00
