Techniques for similarity analysis and data enrichment using knowledge sources
Abstract
The present disclosure relates to performing similarity metric analysis and data enhancement using knowledge sources. The data enhancement service can identify similar related data by comparing the input dataset with the reference dataset stored in the knowledge source. The similarity metric may be calculated to correspond to the semantic similarity of two or more datasets. By using similarity metrics, datasets can be identified based on the metadata attributes and data values of these datasets, which can facilitate indexing and high-performance retrieval of data values. .. Input datasets can be labeled with categories based on the dataset that show the best match with the input dataset. Additional information about the dataset can be obtained by querying the knowledge source using the similarity between the input dataset and the dataset provided by the knowledge source. Recommendations may be provided to the user with additional information.

Term
9 yearsto projected expiry
Projected expiry 25 September 2035, counted from filing; an application has no term until it is granted.
- Priority
- Filed
- Published
- Today
- Projected expiry
24 claims: 11 independent, 13 dependent
- 1方法であって、 入力データセットを1つ以上の入力データソースから受けるステップと、 データ強化サービスのコンピューティングシステムによって、前記入力データセットを、参照ソースから取得した1つ以上の参照データセットと比較するステップと、 前記コンピューティングシステムによって、前記1つ以上の参照データセット各々について類似性メトリックを計算するステップとを含み、前記類似性メトリックは、前記入力データセットとの比較における前記1つ以上の参照データセット各々の類似性の程度を示し、 前記コンピューティングシステムによって、前記類似性メトリックに基づいて前記入力データセットと前記1つ以上の参照データセットとの間の一致を識別するステップと、 前記コンピューティングシステムによって、前記1つ以上の参照データセット各々について計算した前記類似性メトリックを示しかつ前記入力データセットと前記1つ以上の参照データセットとの間の前記識別した一致を示すグラフィカルインターフェイスを生成するステップと、 前記グラフィカルインターフェイスを用いて、前記1つ以上の参照データセット各々について計算した前記類似性メトリックを示しかつ前記入力データセットと前記1つ以上の参照データセットとの間の前記識別した一致を示すグラフィカルなビジュアライゼーションをレンダリングするステップとを含む、方法。
- 2前記1つ以上の参照データセットは、ドメインに対応付けられた用語を含み、前記類似性メトリックは、前記1つ以上の参照データセット各々について計算されたマッチングスコアであり、前記マッチングスコアは、前記参照データセットに関するメトリックを示す第1の値と前記入力データセットと前記参照データセットとの比較に基づくメトリックを示す第2の値とを含む1つ以上の値を用いて計算される、請求項1に記載の方法。
- 3前記グラフィカルなビジュアライゼーションはレンダリングされることによって前記マッチングスコアの計算に用いられる1つ以上の値を示す、請求項2または3に記載の方法。
- 4前記1つ以上の値は、前記入力データセットと前記データセットとの間で一致する用語の度数値と、前記データセットの母集団値と、前記入力データセットと前記データセットとの間で一致する異なる用語の数を示す固有マッチング値と、前記データセット内の用語の数を示すドメイン値と、前記データセットのキュレーションの程度を示すキュレーションレベルとを含む、請求項1から4のいずれか一項に記載の方法。
- 5前記コンピューティングシステムによって、アグリゲーションサービスから取得した増補データに基づいて増補リストを生成するステップと、 前記増補リストに基づいて前記入力データセットを増補するステップとをさらに含み、 前記1つ以上の参照データセットと比較される前記入力データは、前記増補リストに基づいて増補される、請求項1に記載の方法。
- 6前記方法はさらに、 前記コンピューティングシステムによって、前記1つ以上の参照データセットに基づいてインデックス付トライグラム表を生成するステップを含み、 増補後の前記入力データセットにおけるワードごとに、 前記ワードのトライグラムを作成するステップと、 前記トライグラム各々を前記インデックス付トライグラム表と比較するステップと、 前記トライグラムのうちの第1のトライグラムと一致する、トライグラムに対応付けられた前記インデックス付トライグラム表におけるワードを識別するステップと、 前記ワードをトライグラム増補データセットに格納するステップとを含み、 前記トライグラム増補データセットを前記1つ以上の参照データセットと比較するステップと、 前記比較に基づいて前記トライグラム増補データセットと前記1つ以上の参照データセットとの間の一致を判断するステップとを含み、 前記入力データセットと前記1つ以上の参照データセットとの間の一致を識別するステップは、前記比較に基づく前記トライグラム増補データセットと前記1つ以上の参照データセットとの間の一致を用いて実行される、請求項5に記載の方法。
- 7前記1つ以上の参照データセットの少なくとも一部を表わすデータ構造を生成するステップをさらに含み、前記データ構造における各ノードは、前記1つ以上の参照データセットから抽出された1つ以上のストリングの中の異なる文字を表わし、 前記入力データセットは、前記データ構造をトラバースすることによって前記1つ以上の参照データセットと比較される、請求項1から6のいずれか一項に記載の方法。
- 8前記類似性メトリックは、前記入力データセットとの比較における前記1つ以上の参照データセットの共通部分のカーディナリティに基づく値として計算され、 前記値は前記カーディナリティによって正規化され、 前記値は、前記1つ以上の参照データセットのサイズに基づく第1のファクタだけ減じられ、前記値は、前記1つ以上の参照データセットのタイプに基づく第2のファクタだけ減じられる、請求項7に記載の方法。
- 9前記類似性メトリックは、前記1つ以上の参照データセットのうちの各参照データセットについて、前記入力データセットと前記参照データセットとの間のコサイン類似度を求めることによって計算される、請求項1から8のいずれか一項に記載の方法。
- 10前記一致を識別するステップは、前記1つ以上の参照データセットのうち、前記1つ以上の参照データセット各々について計算した前記類似性メトリックに基づく類似性の程度が最大である参照データを求めるステップを含む、請求項1から9のいずれか一項に記載の方法。
- 11前記入力データセットは1つ以上のデータ列にフォーマットされる、請求項1から10のいずれか一項に記載の方法。
- 12データ強化システムであって、 複数の入力データソースと、 クラウドコンピューティングインフラストラクチャシステムとを備え、前記クラウドコンピューティングインフラストラクチャシステムは、 少なくとも1つの通信ネットワークを通して前記複数の入力データソースに通信可能に結合されかつ複数のデータターゲットに通信可能に結合された1つ以上のプロセッサと、 前記1つ以上のプロセッサに結合されたメモリとを含み、前記メモリは、データ強化サービスを提供することを指示する命令を格納し、前記命令は、前記1つ以上のプロセッサによって実行されたときに、前記1つ以上のプロセッサに、 入力データセットを前記複数の入力データソースのうちの1つ以上の入力データソースから受けることと、 前記入力データセットを、参照ソースから取得した1つ以上の参照データセットと比較することと、 前記1つ以上の参照データセット各々について類似性メトリックを計算することとを実行させ、前記類似性メトリックは、前記入力データセットとの比較における前記1つ以上の参照データセット各々の類似性の程度を示し、 前記類似性メトリックに基づいて前記入力データセットと前記1つ以上の参照データセットとの間の一致を識別することと、 前記1つ以上の参照データセット各々について計算した前記類似性メトリックを示しかつ前記入力データセットと前記1つ以上の参照データセットとの間の前記識別した一致を示すグラフィカルインターフェイスを生成することと、 前記1つ以上の参照データセット各々について計算した前記類似性メトリックを示しかつ前記入力データセットと前記1つ以上の参照データセットとの間の前記識別した一致を示すグラフィカルなビジュアライゼーションをレンダリングすることとを実行させる、データ強化システム。
- 13前記1つ以上の参照データセットは、ドメインに対応付けられた用語を含み、前記類似性メトリックは、前記1つ以上の参照データセット各々について計算されたマッチングスコアであり、前記マッチングスコアは、前記参照データセットに関するメトリックを示す第1の値と前記入力データセットと前記参照データセットとの比較に基づくメトリックを示す第2の値とを含む1つ以上の値を用いて計算され、前記グラフィカルなビジュアライゼーションはレンダリングされることによって前記マッチングスコアの計算に用いられる1つ以上の値を示す、請求項12に記載のデータ強化システム。
- 14前記1つ以上の値は、前記入力データセットと前記データセットとの間で一致する用語の度数値と、前記データセットの母集団値と、前記入力データセットと前記データセットとの間で一致する異なる用語の数を示す固有マッチング値と、前記データセット内の用語の数を示すドメイン値と、前記データセットのキュレーションの程度を示すキュレーションレベルとを含む、請求項13に記載のデータ強化システム。
- 15前記命令はさらに、前記1つ以上のプロセッサによって実行されたときに、前記1つ以上のプロセッサに、 アグリゲーションサービスから取得した増補データに基づいて増補リストを生成することと、 前記増補リストに基づいて前記入力データセットを増補することと、 前記1つ以上の参照データセットに基づいてインデックス付トライグラム表を生成することとを実行させ、 増補後の前記入力データセットにおけるワードごとに、 前記ワードのトライグラムを作成することと、 前記トライグラム各々を前記インデックス付トライグラム表と比較することと、 前記トライグラムのうちの第1のトライグラムと一致する、トライグラムに対応付けられた前記インデックス付トライグラム表におけるワードを識別することと、 前記ワードをトライグラム増補データセットに格納することとを実行させ、 前記トライグラム増補データセットを前記1つ以上の参照データセットと比較することと、 前記比較に基づいて前記トライグラム増補データセットと前記1つ以上の参照データセットとの間の一致を判断することとを実行させ、 前記1つ以上の参照データセットと比較される前記入力データは、前記増補リストに基づいて増補され、 前記入力データセットと前記1つ以上の参照データセットとの間の一致を識別することは、前記比較に基づく前記トライグラム増補データセットと前記1つ以上の参照データセットとの間の一致を用いて実行される、請求項12から14のいずれか一項に記載のデータ強化システム。
- 16非一時的なコンピュータ可読記憶媒体であって、前記非一時的なコンピュータ可読記憶媒体に格納された命令を含み、前記命令は、1つ以上のプロセッサによって実行されたときに、前記1つ以上のプロセッサに、 入力データセットを1つ以上の入力データソースから受けることと、 データ強化サービスのコンピューティングシステムによって、前記入力データセットを、参照ソースから取得した1つ以上の参照データセットと比較することと、 前記コンピューティングシステムによって、前記1つ以上の参照データセット各々について類似性メトリックを計算することとを実行させ、前記類似性メトリックは、前記入力データセットとの比較における前記1つ以上の参照データセット各々の類似性の程度を示し、 前記コンピューティングシステムによって、前記類似性メトリックに基づいて前記入力データセットと前記1つ以上の参照データセットとの間の一致を識別することと、 前記コンピューティングシステムによって、前記1つ以上の参照データセット各々について計算した前記類似性メトリックを示しかつ前記入力データセットと前記1つ以上の参照データセットとの間の前記識別した一致を示すグラフィカルインターフェイスを生成することと、 前記グラフィカルインターフェイスを用いて、前記1つ以上の参照データセット各々について計算した前記類似性メトリックを示しかつ前記入力データセットと前記1つ以上の参照データセットとの間の前記識別した一致を示すグラフィカルなビジュアライゼーションをレンダリングすることとを実行させる、非一時的なコンピュータ可読記憶媒体。
- 17方法であって、 入力データセットを1つ以上の入力データソースから受けるステップと、 データ強化サービスのコンピューティングシステムによって、前記入力データセットを、参照ソースから取得した1つ以上の参照データセットと比較するステップと、 前記コンピューティングシステムによって、前記1つ以上の参照データセット各々について類似性メトリックを計算するステップとを含み、前記類似性メトリックは、前記入力データセットとの比較における前記1つ以上の参照データセット各々の類似性の程度を示し、 前記コンピューティングシステムによって、前記類似性メトリックに基づいて前記入力データセットと前記1つ以上の参照データセットとの間の一致を識別するステップと、 前記入力データセットをマッチング情報とともに格納するステップとを含み、前記マッチング情報は、前記1つ以上の参照データセット各々について計算した類似性メトリックを示しかつ前記入力データセットと前記1つ以上の参照データセットとの間の前記識別した一致を示す、方法。
- 18前記入力データセットと前記1つ以上の参照データセットとの間の一致の識別に基づいて、前記入力データセットのカテゴリラベルを識別するステップと、 前記カテゴリラベルに対応付けて前記入力データセットを格納するステップとをさらに含む、請求項17に記載の方法。
- 19前記類似性メトリックは、Jaccard係数、Tversky係数、またはDice-Sorensen係数のうちの1つ以上を用いて計算される、請求項17または18に記載の方法。
- 20前記入力データセットは、グラフマッチングまたは意味類似性マッチングのうちの1つ以上を用いて、前記1つ以上の参照データセットと比較される、請求項17から19のいずれか一項に記載の方法。
- 21データ強化システムであって、 複数の入力データソースと、 クラウドコンピューティングインフラストラクチャシステムとを備え、前記クラウドコンピューティングインフラストラクチャシステムは、 少なくとも1つの通信ネットワークを通して前記複数の入力データソースに通信可能に結合されかつ複数のデータターゲットに通信可能に結合された1つ以上のプロセッサと、 前記1つ以上のプロセッサに結合されたメモリとを含み、前記メモリは、データ強化サービスを提供することを指示する命令を格納し、前記命令は、前記1つ以上のプロセッサによって実行されたときに、前記1つ以上のプロセッサに、 入力データセットを1つ以上の入力データソースから受けることと、 前記入力データセットを、参照ソースから取得した1つ以上の参照データセットと比較することと、 前記1つ以上の参照データセット各々について類似性メトリックを計算することとを実行させ、前記類似性メトリックは、前記入力データセットとの比較における前記1つ以上の参照データセット各々の類似性の程度を示し、 前記類似性メトリックに基づいて前記入力データセットと前記1つ以上の参照データセットとの間の一致を識別することと、 前記入力データセットをマッチング情報とともに格納することとを実行させ、前記マッチング情報は、前記1つ以上の参照データセット各々について計算した類似性メトリックを示しかつ前記入力データセットと前記1つ以上の参照データセットとの間の前記識別した一致を示す、データ強化システム。
- 22前記命令はさらに、前記1つ以上のプロセッサによって実行されたときに、前記1つ以上のプロセッサに、 前記入力データセットと前記1つ以上の参照データセットとの間の一致の識別に基づいて、前記入力データセットのカテゴリラベルを識別することと、 前記カテゴリラベルに対応付けて前記入力データセットを格納することとを実行させる、請求項21に記載のデータ強化システム。
- 23前記類似性メトリックは、Jaccard係数、Tversky係数、またはDice-Sorensen係数のうちの1つ以上を用いて計算される、請求項21または22に記載のデータ強化システム。
- 24前記入力データセットは、グラフマッチングまたは意味類似性マッチングのうちの1つ以上を用いて、前記1つ以上の参照データセットと比較される、請求項21または23に記載のデータ強化システム。
Independent claims24
247 paragraphs, as filed
Cross-reference of related applications This application is filed on September 24, 2015 and is entitled "TECHNIQUES FOR SIMILARITY ANALYSIS AND DATA ENRICHMENT USING KNOWLEDGE SOURCES" in US Non-Provisional Patent Application No. 14 / 864,485 claiming interests and priority based on the following applications: Claim underlying interests and priorities.
1) US provisional application entitled "METHOD FOR SEMANTIC ENTITY EXTRACTION BASED ON GRAPH MATCHING WITH AN EXTERNAL KNOWLEDGEBASE AND SIMILARITY RANKING OF DATASET METADATA FOR SEMANTIC INDEXING, SEARCH, AND RETRIEVAL" filed on September 26, 2014. No. 056,468 2) US Provisional Application No. 62 / 163,296, filed May 18, 2015 and entitled "CATEGORY LABELING" 3) US Provisional Application No. 62 / 203,806, filed on August 11, 2015 and entitled "SIMILARITY METRIC ANALYSIS AND KNOWLEDGE SCORING SYSTEM" This application relates to the following applications.
1) US Provisional Application No. 62 / 056,471 filed on September 26, 2014 and entitled "DECLARATIVE LANGUAGE AND VISUALIZATION SYSTEM FOR RECOMMENDED DATA TRANSFORMATIONS AND REPAIRS" 2) US provisional application No. 62 / 056,474, filed on September 26, 2014 and entitled "DYNAMIC VISUAL PROFILING AND VISUALIZATION OF HIGH VOLUME DATASETS AND REAL-TIME SMART SAMPLING AND STATISTICAL PROFILING OF EXTREMELY LARGE DATASETS" 3) US Provisional Application No. 62 / 056,475, filed on September 26, 2014 and entitled "AUTOMATED ENTITY CORRELATION AND CLASSIFICATION ACROSS HETEROGENEOUS DATASETS" 4) Filed on September 26, 2014, "DECLARATIVE EXTERNAL DATA" US Provisional Application No. 62 / 056,476 entitled "SOURCE IMPORTATION, EXPORTATION, AND METADATA REFLECTION UTILIZING HTTP AND HDFS PROTOCOLS" The entire contents of the above patent application are incorporated herein by reference for all purposes.
background This disclosure generally relates to the preparation and analysis of data. More specifically, techniques for performing similarity metric analysis and data enhancement using knowledge sources are disclosed.
Before "big data" systems were able to analyze data and provide useful results, it was necessary to add data to big data systems and format them for analysis. This data onboarding presents a challenge to today's cloud and "big data" systems. Typically, the data added to a big data system is noisy (eg, the data is incorrectly formatted, incorrect, expired, contains duplicates, etc.). When analyzing data (for example, for reporting, predictive modeling, etc.), inadequate signal-to-noise ratio of the data means that the results are not useful. As a result, current solutions require a virtually manual process to clean and curate data and / or analysis results. However, these manual processes cannot be scaled. As the amount of data added and analyzed increases, manual processes become impossible to implement.
Other similar related data may be identified by implementing a big data system and analyzing the data. The amount of data processed becomes an issue. Moreover, depending on the structure of the data to be analyzed or the lack of this structure, the data to be analyzed may pose even greater challenges in determining the content and relationships of the data.
You may want to implement machine learning to analyze your data. For example, unsupervised machine learning may be implemented using a data analysis tool (eg Word2Vec) to determine similarities between data. However, unsupervised machine learning may not be able to provide information indicating groups or categories that correspond to relevant data. For this reason, unsupervised learning may not be able to determine the genus or category of a set of relevant species (eg, terms). Meanwhile, curated knowledge sources (eg Max Planck Institute for) Supervised machine learning based on Informatics YAGO) can provide better results in determining groups or categories of data. Supervised learning can have inconsistent and / or incomplete results. The data provided by the curated knowledge source can be sparse and the quality can vary from curator to curator. Categories identified based on the use of supervised learning may not provide the correct categorization of similar related data. Multiple knowledge sources implement different categorizations, which can make it difficult to integrate multiple knowledge sources. Analyzing data to determine similarities and relationships can be burdensome due to misspelling of terms in the data being analyzed. If the data contains misspellings, it may not be easy to identify similar data.
<p num="0008"> Certain embodiments of the invention address the above and other issues. A brief overview This disclosure generally relates to the preparation and analysis of data. More specifically, techniques for performing similarity metric analysis and data enrichment using knowledge sources are disclosed.</p><p num="0009"> The disclosure generally relates to a data enhancement service, which extracts, repairs, and enhances datasets to obtain more precise entity resolution and correlation for later indexing and clustering. .. Data enhancement services may include a visual recommendation engine and language for performing large-scale data preparation, repair, and enhancement of heterogeneous datasets. This allows the user to select and see how the recommended enhancements (eg transformations and repairs) affect the user's data and make adjustments as needed. Data enhancement services can receive user feedback through the user interface and can filter recommendations based on user feedback. In some embodiments, the data enhancement service can identify patterns in the data by analyzing the dataset.</p><p num="0010"> In some embodiments, the data enhancement service can identify similar related data by comparing the input dataset with a reference dataset stored in the knowledge source. Matching of input data to reference datasets can be performed without supervised training (eg machine learning), and extraction accuracy can be gradually improved through adaptive feedback from end users. In some embodiments, similarity metrics can be calculated that correspond to the semantic similarity of two or more datasets. By using the similarity metric, a dataset can be identified based on its metadata attributes and data values. This can make indexing and high-performance retrieval of data values easier.</p><p num="0011"> As mentioned above, the amount of data processed becomes an issue, especially depending on the structure of the data to be analyzed or the lack of this structure. Spelling mistakes and differences in the curation of reference data can lead to categorical misclassification, which makes it difficult to identify similar or related data. The techniques described herein provide more sophisticated similarity metrics, which can improve the automatic identification of highly relevant datasets that have semantic similarity to the input dataset. The input dataset may be enhanced with data from the relevant dataset by identifying more similar related datasets. Enhanced input datasets allow users to understand and manage large amounts of data that would otherwise be difficult to manage. For example, the user may determine if a dataset is relevant to a particular topic, and if so, whether there is relevant data for this topic. In some embodiments, the reference dataset may be updated to reflect the relationship with the input data based on the similarity metric. In this way, the reference dataset can be enhanced for later use in determining similarity with other input datasets.</p><p num="0012"> In some embodiments, the data enhancement service can render a graphical interface that displays similarity metrics for each of the plurality of reference datasets that are compared to the input dataset. The graphical interface allows the user to choose a transformation based on one of the reference datasets for which the similarity metric is to be shown. In this way, the similarity metric allows the user to selectively select reference data to enhance the dataset from the data source.</p><p num="0013"> In some embodiments, the techniques disclosed herein provide a method of presenting to the user a classification of data received from a data source. This technique offers advantages over unsupervised machine learning, where it may not be possible to determine the genus or category of a relevant set of species (eg, terms). The technique also provides a more stable and perfect data classification by combining unsupervised machine learning techniques with merging multiple sources for supervised machine learning. Such techniques can take into account different levels of term curation and misspelling or categorization errors.</p><p num="0014"> In some embodiments, the data enhancement service can determine similarity metrics by comparing terms in the input dataset with terms in the dataset from the knowledge source. Similarity metrics can be calculated using the various techniques disclosed herein. The similarity metric may be expressed as a score. The input dataset may be compared to a plurality of datasets, each of which may be associated with a category (eg, domain). Similarity metrics may be calculated to compare each dataset with the input dataset. Thus, the similarity metric may indicate the degree of agreement (eg, the highest degree of similarity is indicated by the maximum value) so that higher degree of agreement can be identified based on the value of the similarity metric. Similarity metrics obtained using one or more of the techniques disclosed herein may provide greater accuracy in matching input datasets with datasets provided by knowledge sources.</p><p num="0015"> The data enhancement service may determine the similarity between an input dataset and one or more datasets by implementing several different techniques. The input dataset can be associated with or labeled with the category (eg domain) that corresponds to the dataset that has the greatest match (eg, maximum similarity) with this input dataset. Therefore, the input dataset can be modified or enhanced using category names. The category name allows the user to better identify the input dataset. In at least one embodiment, the data enhancement service can more precisely label the categories of input data by combining unsupervised learning techniques with supervised learning techniques. By using the similarity of the input data to the dataset given by the knowledge source, the knowledge source can be queried to obtain additional information about the dataset. Recommendations can be provided to users with additional information.</p><p num="0016"> In some embodiments, a computing system may be implemented to perform a data similarity metric analysis in comparison to a dataset given by a knowledge source. Computing systems can implement data enhancement services. The computing system may be configured to implement the methods and operations described herein. The system may include multiple input data sources and multiple data targets. The system may include a cloud computing infrastructure system with one or more processors communicably coupled to multiple input data sources and communicably coupled to multiple data targets through at least one communication network. A cloud computing infrastructure system may include memory coupled to one or more of the above processors. The memory includes an instruction instructing it to provide a data enhancement service, and when this instruction is executed by one or more of the above processors, one or more of the methods or operations described herein will perform one or more of the above. Executed by the processor of. Yet another embodiment relates to a system and a machine-readable tangible storage medium, which uses or stores instructions for the methods and operations described herein.</p><p num="0017"> In at least one embodiment, the method comprises receiving an input data set from one or more input data sources. The input dataset may be formatted into one or more strings of data. This method may include comparing the input dataset with one or more reference datasets obtained from the reference source. The reference source may be a knowledge source provided by a knowledge service. The input dataset may be compared to one or more reference datasets using one or more of graph matching or semantic similarity matching. The method may include calculating similarity metrics for each of the above one or more reference datasets. The similarity metric indicates the degree of similarity of each of one or more reference datasets in comparison to the input dataset. This method may include identifying a match between the input dataset and one or more reference datasets based on the similarity metric. In some embodiments, the method shows a similarity metric calculated for each of the one or more reference datasets and identifies an identified match between the input dataset and the one or more reference datasets. To show the steps to generate the graphical interface shown and the similarity metrics calculated for each of the one or more reference datasets and to show the identified match between the input dataset and the one or more reference datasets. Graphical visualizations can include rendering using a graphical interface. In some embodiments, the method shows a similarity metric calculated for each of the one or more reference datasets and identifies an identified match between the input dataset and the one or more reference datasets. To store the input dataset with the matching information shown and to identify the category label of the input dataset based on the identification of a match between the input dataset and one or more reference datasets. , Above data</p><p num="0018"> In some embodiments, the one or more reference datasets include terms associated with a domain. The similarity metric may be a matching score calculated for each of the above one or more reference datasets. The matching score may be calculated using one or more values that include a first value that indicates the metric for the reference dataset and a second value that indicates the metric based on the comparison between the input dataset and the reference dataset. Good. The graphical visualization may show one or more of the above values that are rendered and used to calculate the matching score. One or more of the above values are the degrees and values of the terms that match between the input dataset and the dataset, the population value of the dataset, and the number of different terms that match between the input dataset and the dataset. Can include a unique matching value that indicates, a domain value that indicates the number of terms in the dataset, and a curation level that indicates the degree of curation of the dataset.</p><p num="0019"> In some embodiments, the method may further include generating an augmentation list based on augmentation data obtained from the aggregation service and augmenting the input dataset based on the augmentation list. .. Input data that is compared to one or more reference datasets may be augmented based on augmented lists. The method may further include generating an indexed trigram table based on one or more of the reference datasets described above. This method involves creating a trigram of this word for each word in the augmented input dataset, comparing each trigram with an indexed trigram table, and the first of the above trigrams. It may include identifying a word in the indexed trigram table associated with the trigram that matches the trigram, and storing this word in the trigram augmented dataset. This method determines the match between the trigram augmented data set and one or more reference datasets based on the step of comparing the trigram augmented dataset with one or more reference datasets. Can include steps and. The step of identifying a match between the input dataset and one or more reference datasets is performed using the match between the trigram augmented dataset and one or more reference datasets based on the comparison. You may.</p><p num="0020"> In some embodiments, the method may include generating a data structure that represents at least a portion of the one or more reference datasets described above. Each node in this data structure represents a different character in one or more strings extracted from the one or more reference datasets above. The input dataset may be compared to one or more reference datasets by traversing the data structure. The similarity metric may be calculated as a value based on the cardinality of the intersection of the above one or more reference datasets in comparison to the input dataset. This value may be normalized by cardinality. This value may be decremented by a first factor based on the size of the one or more reference datasets above, and this value may be decremented by a second factor based on the type of the one or more reference datasets above. You may.</p><p num="0021"> In some embodiments, the similarity metric for each reference dataset of the one or more reference datasets is calculated by finding the cosine similarity between the input dataset and the reference dataset. May be good. The similarity metric may be calculated using one or more of the Jaccard coefficient, Tversky coefficient, or Dice-Sorensen coefficient. The step of identifying the match is a step of finding the reference data having the maximum degree of similarity based on the similarity metric calculated for each of the one or more reference data sets among the one or more reference data sets. Can include.</p><p num="0022"> What has been said so far, along with other features and embodiments, will become clearer with reference to the following specification, claims, and accompanying drawings.</p>
<figref num="1">A simplified high-level diagram of a data enhancement system according to an embodiment of the present invention is shown.</figref><figref num="2">A simplified block diagram of a technology stack according to an embodiment of the present invention is shown.</figref><figref num="3">A simplified block diagram of a data enhancement system according to an embodiment of the present invention is shown.</figref><figref num="4A">An example of a user interface that provides interactive data enhancement according to an embodiment of the present invention is shown.</figref><figref num="4B">An example of a user interface that provides interactive data enhancement according to an embodiment of the present invention is shown.</figref><figref num="4C">An example of a user interface that provides interactive data enhancement according to an embodiment of the present invention is shown.</figref><figref num="4D">An example of a user interface that provides interactive data enhancement according to an embodiment of the present invention is shown.</figref><figref num="5A">An example of various user interfaces that provide visualization of a dataset according to an embodiment of the present invention is shown.</figref><figref num="5B">An example of various user interfaces that provide visualization of a dataset according to an embodiment of the present invention is shown.</figref><figref num="5C">An example of various user interfaces that provide visualization of a dataset according to an embodiment of the present invention is shown.</figref><figref num="5D">An example of various user interfaces that provide visualization of a dataset according to an embodiment of the present invention is shown.</figref><figref num="6">A representative graph according to the embodiment of the present invention is shown.</figref><figref num="7">A typical state table according to the embodiment of the present invention is shown.</figref><figref num="8">An example of a case-insensitive graph according to an embodiment of the present invention is shown.</figref><figref num="9">The figure which shows the similarity of the data set according to the embodiment of this invention is shown.</figref><figref num="10">An example of a graphical interface for displaying knowledge scoring of different knowledge domains according to an embodiment of the present invention is shown.</figref><figref num="11">An example of automated data analysis according to an embodiment of the present invention is shown.</figref><figref num="12">An example of trigram modeling according to the embodiment of the present invention is shown.</figref><figref num="13">An example of category labeling according to an embodiment of the present invention is shown.</figref><figref num="14">A similarity analysis is shown to determine the ranked categories according to embodiments of the present invention.</figref><figref num="15">A similarity analysis is shown to determine the ranked categories according to embodiments of the present invention.</figref><figref num="16">A similarity analysis is shown to determine the ranked categories according to embodiments of the present invention.</figref><figref num="17">The flow chart of the process of similarity analysis according to the embodiment of this invention is shown.</figref><figref num="18">The flow chart of the process of similarity analysis according to the embodiment of this invention is shown.</figref><figref num="19">A simplified diagram of a distributed system for realizing the embodiment is shown.</figref><figref num="20">FIG. 3 is a simplified block diagram of one or more components of a system environment that can be serviced as a cloud service according to an embodiment of the present disclosure.</figref><figref num="21">A typical computer system that can be used to realize an embodiment of the present invention is shown.</figref>
Detailed explanation In the following description, the embodiments of the present invention will be fully understood by stating specific details for the sake of explanation. However, it will be clear that various embodiments can be implemented without these specific details. The drawings and descriptions are not intended to be limiting.
The disclosure generally relates to a data enhancement service, which extracts, repairs, and enhances datasets to obtain more precise entity resolution and correlation for later indexing and clustering. In some embodiments, the data enhancement service includes an extensible semantic pipeline that exposes data to a data target by processing the data at multiple stages, from collecting the data to analyzing the data.
In one embodiment of the invention, data is processed through a pipeline (also referred to herein as a semantic pipeline) that includes various processing stages before loading the data into a data warehouse (or other data target). In some embodiments, the pipeline may include a collection stage, a preparation stage, a profile stage, a conversion stage, and an open stage. During processing, the data can be analyzed, prepared and enhanced. The resulting data can then be exposed to one or more data targets (eg local storage systems, cloud-based storage services, web services, data warehouses, etc.) (eg, given to downstream processes). .. Various data analyzes can be performed on the data at this target. This data has been repaired and enhanced, and analysis of it will give useful results. Therefore, because the data onboarding process is automated, scaling can handle very large datasets that cannot be manually processed due to their volume.
In some embodiments, the data can be analyzed to extract entities from this data and the data can be repaired based on the extracted entities. For example, misspelling, misaddressing, and other common mistakes present complex problems for big data systems. If the amount of data is small, such errors can be manually identified and corrected. However, for very large datasets (eg billions of nodes or records), such manual processing is not possible. In certain embodiments of the invention, the data enhancement service can use the knowledge service to analyze the data. Entity in the data can be identified based on the content of the knowledge service. For example, the entity may be an address, business name, location, personal name, ID number, or the like.
FIG. 1 shows a simplified high-level diagram 100 of a data enhancement service according to an embodiment of the present invention. As shown in FIG. 1, the cloud-based data enhancement service 102 can receive data from various data sources 104. In some embodiments, the client can issue a data enhancement request to the data enhancement service 102, which is one or more of the data source 104 (or a portion thereof, eg, a particular table,). Identify the dataset, etc.). The data enhancement service 102 may then request processing of data from the identified data source 104. In some embodiments, the data source may be sampled and the sampled data is analyzed for enhancement, thereby making large datasets more manageable. The identified data can be received and added to a distributed storage system (such as a Hadoop distributed storage (HDFS) system) accessible from data enhancement services. Data may be processed semantically by a number of processing steps (described herein as pipelines or semantic pipelines). These processing stages may include a preparation stage 108, a strengthening stage 110, and an open stage 112. In some embodiments, the data can be processed in one or more batches by the data enhancement service. In some embodiments, it is possible to provide a streaming pipeline that processes while receiving data.
In some embodiments, the preparation stage 108 may include various processing substages. This may include automatically detecting the data source format and performing content extraction and / or repair. When a data source format is detected, the data source can be automatically normalized to a format that the data enhancement service can handle. In some embodiments, once the data source is prepared, it can be processed by the enhancement stage 110. In some embodiments, the inbound data source can be loaded into a distributed storage system 105 accessible from the data enhancement service, such as an HDFS system communicatively coupled to the data enhancement service. The distributed storage system 105 provides a temporary storage space for the collected data files, which also provides storage for intermediate processing files and as temporary storage for pre-publication results. Can be done. In some embodiments, the augmented or enhanced results can also be stored in the distributed storage system. In some embodiments, the metadata captured during the enhancement associated with the collected data source can be stored in the distributed storage system 105. System-level metadata (for example, showing data source location, results, processing history, user sessions, execution history, and configuration) is stored in a distributed storage system or in a separate repository accessible to data enhancement services. can do.
In certain embodiments, the enhancement process 110 uses a semantic bus (also referred to herein as a pipeline or semantic pipeline) and one or more natural language (NL) processors connected to this bus. The data can be analyzed. The NL processor automatically identifies the data source column, determines the type of data in a particular column, names this column if the input does not have a schema, and / or metadata that describes the column and / or the data source. Can be provided. In some embodiments, the NL processor can identify and extract entities (eg, people, places, objects, etc.) from the text in a column. The NL processor can also identify and / or build relationships within and between data sources. Data can be repaired (eg, correct typos or format errors) and / or enhanced (eg, include additional relevant information in the extracted entities) based on the extracted entities, as described further below.
In some embodiments, publish stage 112 may provide the metadata of the data source captured during the enhancement and any enhancement or repair of the data source to one or more visualization systems for analysis. Yes (for example, recommended data transformations, enhancements, and / or other modifications can be displayed to the user). The public subsystem can send the processed data to one or more data targets. The data target can correspond to a place where processed data can be sent. This location may be, for example, a location in memory, a computing system, a database, or a system that provides services. For example, data targets include Oracle Storage Cloud Service (OSCS), URLs, third-party storage services, web services, and Oracle Business Intelligence (BI), Database as a service. as a Service) and other cloud services such as Database Schema as a Service may be included. In some embodiments, the syndication engine provides the customer with a set of APIs that are subject to browsing, selection, and subscription to results. When subscribed and new results are generated, the result data can be provided as a direct feed to the endpoint of an external web service or as a bulk file download.
FIG. 2 shows a simplified block diagram 200 of a technology stack according to an embodiment of the present invention. In some embodiments, the data enhancement service can be implemented using the logic technology stack shown in Figure 2. This technology stack provides access to data enhancement services through one or more client devices (eg, using thin clients, thick clients, web browsers, or other applications running on client devices). It may include experience (UX) layer 202. The scheduler service 204 can manage the results / responses received through the UX layer, and can manage the underlying infrastructure, on which the data enhancement service runs.
In some embodiments, the processing stage described above with reference to FIG. 1 may include a large number of processing engines. For example, the preparation processing stage 108 may include a collection / preparation engine, a profiling engine, and a recommendation engine. If data is collected during the preparatory process, this data (or a sample thereof) can be stored in a distributed data storage system 210 (such as a "big data" cluster). The enhancement processing stage 110 may include a meaning / statistics engine, an entity extraction engine, and a repair / transformation engine. As further described below, the strengthening process stage 110 can utilize the information obtained from the knowledge service 206 during the strengthening process. Enhanced actions (for example, adding and / or transforming data) can be performed on the data stored in the distributed storage system 210. Data transformations may include missing data or modifications to enhance the data by adding data. Data transformation can include correcting errors in the data or repairing the data. The publishing process stage 112 may include a publishing engine, a syndication engine, and a metadata result manager. In some embodiments, different open source techniques can be used to implement several features within different processing stages and / or processing engines. For example, file format detection may use Apache Tika.
In some embodiments, management service 208 can monitor changes made to the data during enhancement process 110. Change monitoring can include tracking which users accessed the data, which data conversions were performed, and other data. This allows the data enhancement service to roll back the enhancement action.
The technology stack 200 can be implemented in an environment such as cluster 210 ("big data cluster") for big data operations. Cluster 210 can be implemented using Apache Spark, which provides a set of libraries for implementing a distributed computing framework compatible with distributed file systems (DFS) such as HDFS. Apache Spark can send map, mitigation, filter, sort, or sample cluster processing job requests to a valid resource manager like YARN. In some embodiments, the cluster 210 can be implemented using, for example, a distributed file system product provided by Cloudera®. For example, DFS provided by Cloudera® may include HDFS and YARN.
FIG. 3 shows a simplified block diagram of an interactive visualization system according to an embodiment of the present invention. As shown in FIG. 3, the data enhancement service 302 can receive data enhancement requests from one or more clients 304. The data enhancement system 300 may implement the data enhancement service 302. The data enhancement service 302 can receive data enhancement requests from one or more clients 304. Data enhancement service 302 may include one or more computers and / or servers. The data enhancement service 302 may be a module composed of several subsystems and / or modules, some of which may not be shown. The number of subsystems and / or modules in Data Enhancement Service 302 may be greater or less than the number shown, and may be a combination of two or more subsystems and / or modules, or different. It may be a configuration or deployment subsystem and / or module. In some embodiments, the data enhancement service 302 includes a user interface 306, a collection engine 328, a recommendation engine 308, a knowledge service 310, a profile engine 326, a conversion engine 322, a preparation engine 312, and a public engine. Can include 324 and. Elements that implement the data enhancement service 302 can function to implement the semantic processing pipeline as described above.
The data enhancement system 300 may include a semantic processing pipeline according to an embodiment of the present invention. All or part of the semantic processing pipeline may be implemented by the data enhancement service 102. When adding a data source, this data source and / or the data stored therein can be processed through the pipeline before loading the data source. The pipeline may include one or more processing engines that are configured to process the data and / or the data source before exposing the processed data to one or more data targets. The processing engine may include a collection engine that extracts raw data from the new data source and provides this raw data to the preparation engine. The preparation engine can identify the format associated with this raw data and can convert this raw data to a format that the data enhancement service 302 can process (eg, normalize this raw data). The profile engine can extract and / or generate metadata associated with the normalized data, and the transformation engine transforms the normalized data based on the metadata (eg repair and / or). Can be strengthened). The resulting enhanced data may be given to the public engine and sent to one or more data targets. Each processing engine will be further described below.
In some embodiments, the data enhancement service 302 may be provided by a computing infrastructure system (eg, a cloud computing infrastructure system). A computing infrastructure system can be implemented in a cloud computing environment with one or more computing systems. Computing infrastructure systems may be communicatively coupled to one or more data sources, such as those described herein, or to one or more data targets through one or more communication networks.
Client 304 may include a variety of client devices (desktop computers, laptop computers, tablet computers, mobile devices, etc.). Each client device may contain one or more client applications 304. Data enhancement service 302 can be accessed through this application. For example, browser applications, thin clients (eg mobile applications), and / or thick clients can run on client devices, allowing users to interact with the data enhancement service 302. The embodiments shown in FIG. 3 are merely examples and are not intended to unreasonably limit the claimed embodiments of the present invention. Those skilled in the art will recognize numerous modifications, alternatives, and modifications. For example, the number of client devices may be greater or less than the number of devices shown.
The types of client devices 304 can be diverse. This includes, but is not limited to, mobile or handheld devices such as personal computers, desktops, laptops, mobile phones, tablets, and other types of devices. The communication network facilitates communication between the client device 304 and the data enhancement service 302. The types of communication networks can be different. This communication network may include one or more communication networks. Examples of communication networks 106 include the Internet, wide area networks (WANs), local area networks (LANs), Ethernet (registered trademarks) networks, public or private networks, wired networks, wireless networks, and combinations thereof. Not limited to these. IEEE Communication may be facilitated using different communication protocols, including wired and wireless protocols, such as the 802.XX protocol suit, TCP / IP, IPX, SAN, AppleTalk, Bluetooth, and other protocols. In general, the communication network may include any kind of communication network or infrastructure that facilitates communication between the client and the data enhancement service 302.
The user can interact with the data enhancement service 302 through user interface 306. Client 304 displays user data and recommendations for transforming user data by rendering a graphical user interface, and sends instructions ("conversion instructions") to data enhancement service 302 through user interface 306 and / Or you can receive it. The user interfaces disclosed herein, such as those shown in FIGS. 4A-4D, 5A-5D, and 10, may be rendered by the data enhancement service 302 or through the client 304. For example, the user interface may be generated by user interface 306 or rendered by data enhancement service 302 on any one of the clients 304. The user interface may be provided from the data enhancement system 302 over the network as part of a service (eg, a cloud service) or network accessible application. In at least one example, the operator of Data Enhancement Service 302 may operate one of Client 304 to access and interact with any of the user interfaces disclosed herein. Good. The user may add data sources by sending instructions to user interface 306 (eg, providing data source access and / or location information, etc.).
The data enhancement service 302 may collect data using the collection engine 328. The collection engine 328 can act as an initial processing engine when a data source is added. The collection engine 328 can facilitate the secure, reliable, and reliable upload of user data from one or more data sources 309 to the data enhancement service 302. In some embodiments, the collection engine 328 can extract data from one or more data sources 309 and store it in a distributed storage system 305 within the data enhancement service 302. Data collected from one or more data sources 309 and / or one or more clients 304 can be processed and stored in the distributed storage system 305 as described above with reference to FIGS. 1 and 2. .. Data enhancement service 302 can receive data from client data store 307 and / or from one or more data sources 309. The distributed storage system 305 can serve as a temporary storage of uploaded data during the remaining processing stages of the pipeline prior to data disclosure to one or more data targets 330. Once the upload is complete, the preparation engine 312 can be called to normalize the uploaded dataset.
The received data may include structured data, unstructured data, or a combination thereof. Structural data can be based on data structures, including, but not limited to, arrays, records, relational database tables, hash tables, linked lists, or other types of data structures. As mentioned above, data sources include public cloud storage service 311, private cloud storage service 313, various other cloud services 315, URL or web-based data source 317, or any other accessible data source. obtain. Data enhancement requests from client 304 include data sources and / or specific data (tables, columns, files, or any other structured or unstructured data available through data source 309 or client data store 307). Can be identified. Then, the data enhancement request service 302 may access the specified data source and acquire the specific data specified in the data enhancement request. The data source can be identified by address (eg URL), by storage provider name, or by other identifier. In some embodiments, access to the data source may be controlled by an access control service. The client 304, user with identification (e.g. user name and password) input request and / or data enrichment server may indicate a request for giving permission to access the data source for bis 302.
In some embodiments, the data uploaded from one or more data sources 309 can be transformed into a wide variety of formats. The preparation engine 312 can convert the uploaded data into a common normalized format for processing by the data enhancement service 302. Normalization may be performed by routines and / or techniques implemented using instructions or code such as Apache Tika supplied by Apache®. The normalized format allows you to see the normalized data retrieved from the data source. In some embodiments, the preparation engine 312 is capable of reading a number of different file types. The readiness engine 312 normalizes the data into character separated forms (for example, tab separated values, comma separated values, etc.), or , JavaScript (registered trademark) object notation for hierarchical data (JavaScript) It can be an Object Notation (JSON) document. In some embodiments, various file formats can be recognized and normalized. For example, Microsoft Excel® format (eg XLS or XLSX), Microsoft Word® format (eg DOC or DOX), Portable Document Format (PDF), Hierarchical Formats like JSON, and Extended Markup Languages (for example) Can support standard file formats such as XML). In some embodiments, various binary coded file formats and serialized object data can also be read and decrypted. In some embodiments, the data can be fed to the pipeline in Unicode format (UTF-8) encoding. The preparation engine 312 can perform context extraction and conversion to the file types predicted by the data enhancement service 302, as well as extract document-level metadata from the data source.
Data set normalization can include converting the raw data in the dataset into a format that can be processed by the data enhancement service 302, especially the profile engine 326. In one example, normalizing a dataset to create a normalized dataset involves modifying a dataset with a certain format to a format that has been tuned as a normalized dataset. It is a format different from the above format. The dataset may be normalized by identifying one or more columns of data in this dataset and modifying the format of the data corresponding to this column to the same format. For example, data in a dataset with dates of different formats may be normalized by changing the format of this date to a common format that the profile engine 326 can handle. Data may also be normalized by modifying or converting from a non-tabular format to a tabular format with one or more columns of data.
After normalizing the data, the normalized data can be sent to the profile engine 326. The profile engine 326 identifies the types of data stored in these columns by analyzing the normalized data column by column and provides information about how the data is stored in these columns. Can be identified. Although this disclosure describes the profile engine 326 as performing operations on data in many cases, the data processed by the profile engine 326 has already been normalized by the preparation engine 312. In some embodiments, the data processed by the profile engine 326 may include unnormalized data because it is in a format that the profile engine 326 can process (eg, a normalized format). The output or result of the profile engine 326 may be metadata (eg, source profile) that indicates profile information about the data from the source. Metadata can indicate one or more patterns and / or classifications of data with respect to the data. As further described below, the metadata may include statistical information based on the analysis of the data. For example, the profile engine 326 can output a large amount of metric and pattern information for each identified column, and can identify and match schema information in the form of column names and types.
The metadata generated by the profile engine 326 may be used by other elements of the data enhancement service, such as the recommendation engine 308 and the transformation engine 322, to perform the operations described herein for the data enhancement service 302. In some embodiments, the profile engine 326 can provide metadata to the recommendation engine 308.
The recommendation engine 308 can identify repair, transformation, and data enhancement recommendations for the data processed by the profile engine 326. The metadata generated by the profile engine 326 can be used to determine recommendations for data based on the statistical analysis and / or classification that this metadata presents. In some embodiments, recommendations can be provided to the user through a user interface or other web service. Recommendations are available for data repair or enhancement, how to compare these recommendations with past user activity, and / or how to classify unknown items based on existing knowledge or patterns. It can be tailored to the business user so that the recommendation describes at a high level. Knowledge service 310 has access to one or more knowledge graphs or other knowledge sources 340. This knowledge source may include publicly available information published by websites, web services, curated knowledge stores, and other sources. The recommendation engine 308 can request (eg, query) the knowledge service 310 for data that can be recommended to the user for the data obtained from the source.
In some embodiments, the transformation engine 322 can present to the user, columnwise, sampled data or sample rows of the input dataset through user interface 306. Data enhancement service 302 may indicate to the user the recommended transformations through user interface 306. This conversion may be associated with a conversion instruction. The conversion instruction may include code and / or a function call to perform the conversion action. The conversion instruction may be invoked by the user based on a selection in user interface 306, for example, by selecting a recommendation for conversion or by receiving an input indicating an operation (eg, an operator command). May be called. In one example, the conversion instruction may include an instruction to rename at least one column of data based on entity information. It may also receive other conversion instructions to rename at least one column of data to the default name. The default name may include a predetermined name. The default name may be any specified name if the name of the column of data cannot be determined or the name of this column is not defined. The conversion instruction may include a conversion instruction for reformatting at least one column based on the entity information and an instruction for obfuscating at least one column of data based on the entity information. In some embodiments, the transformation instruction may include an enhancement instruction for adding one or more columns of data obtained from the knowledge service based on entity information.
The user can perform transformation actions through user interface 306, and the transformation engine 322 can apply the data retrieved from the data source to these actions and display the results. It gives immediate feedback to the user and can be used to visualize and verify the effect of the configuration of the conversion engine 322. In some embodiments, the transformation engine 322 can receive pattern and / or metadata information (eg, column names and types) from the profile engine 326 and the recommendation engine 308 that provides the nomination action. .. In some embodiments, the transformation engine 322 can provide a user event model that facilitates undo, redo, delete, and edit events by coordinating and tracking changes to the data. This model can capture the dependencies between actions to keep the current configuration consistent. For example, if a column is deleted, the recommendation conversion action provided by the recommendation engine 308 for this column may also be deleted. Similarly, if a conversion action results in a new column being inserted and this action being deleted, then any action performed on this new column will be deleted.
As mentioned above, the received data can be analyzed during processing, and the recommendation engine 308 can perform one or more recommended transformations on this data, including enhancements, repairs, and other transformations. Can be shown. The transformations recommended for data enhancement may consist of a set of transformations, each transformation being a transformation action or atomic transformation performed on the data. The transformation may be performed on the data previously transformed by another transformation in the above set. The set of transformations may be performed in parallel or in a particular order so that the data obtained after performing the set of conversions is enhanced. A set of conversions may be performed according to the conversion specifications. The conversion specification may include a conversion instruction indicating when and how to perform each set of conversions to the data generated by the profile engine 326, and a recommendation to enhance the data determined by the recommendation engine 308. Examples of atomic conversions can include, but are not limited to, conversions to headers, conversions, deletions, splits, joins, and repairs. A series of changes may be made to the data transformed according to a set of transformations. Each of these changes results in enhanced intermediate data. The data generated in the intermediate steps for a set of transformations is a Resilient Distributed Dataset (RDD), text, data recording format, file format, or any other format, or a combination thereof. It may be stored in a format such as.
In some embodiments, the data generated as a result of an operation performed by any element of Data Enhancement Service 302 may be RDD, text, document format, or any other format, or any of these. It may be stored in an intermediate data format, including combinations. Further operations for data enhancement service 302 may be performed using the data stored in the intermediate format.
The table below shows an example of the conversion. Table 1 outlines the types of conversion actions.
<tables num="1"><img id="000003" he="160" wi="154" file="JP2017536601A_D0001.tif" img-format="tif" img-content="drawing" /></tables>
Table 2 shows the conversion actions that do not belong to the category types shown in Table 1.
<tables num="2"><img id="000004" he="58" wi="155" file="JP2017536601A_D0001.tif" img-format="tif" img-content="drawing" /></tables>
Table 3 below shows examples of conversion example types. Specifically, Table 3 shows examples of transformation actions and describes the types of transformations that correspond to these actions. For example, the transformation action may include filtering the data based on the detection of the presence of words from the white list in the data. If the user wants to track communications (eg tweets) that include "Android" or "iPhone®", the conversion action can be added with the above two words containing the given whitelist. This is just one example of how data can be enhanced for the user.
<tables num="3"><img id="000005" he="188" wi="155" file="JP2017536601A_D0001.tif" img-format="tif" img-content="drawing" /></tables>
The recommendation engine 308 can generate a recommendation for the conversion engine 322 by using the information from the knowledge service 310 and the knowledge source 340, and instructs the conversion engine 322 to generate a conversion script that converts the data. can do. The conversion script can include programs, code, or instructions. This conversion script can be executed by one or more processing units so that the received data can be converted. In this way, the recommendation engine 308 can serve as an intermediary between the user interface 306 and the knowledge service 310.
As described above, the profile engine 326 can determine if there is any pattern by analyzing the data from the data source, and if there is any pattern, determine if the pattern can be classified. be able to. Once the data retrieved from the data source is normalized, parsing this data may identify one or more attributes or fields within the structure of the data. Patterns can be identified using a collection of regular expressions, each with a label ("tag") and defined by a category. The pattern may be identified by comparing the data with different types of patterns. Examples of identifiable pattern types are, but are not limited to, integers, fractions, dates or date / time strings, URLs, domain addresses, IP addresses, email addresses, version numbers, locale identifiers, UUIDs and other hexadecimal digits. Includes legal identifiers, social security numbers, US private box numbers, typical US street address patterns, postal codes, US phone numbers, room numbers, credit card numbers, unique nomenclature, personal information, and credit card issuers obtain.
In some embodiments, the profile engine 326 may identify patterns in the data based on a set of regular expressions defined by semantic or syntactic constraints. By using regular expressions, the shape and / or structure of the data can be determined. The profile engine 326 may determine patterns in data based on one or more regular expressions by implementing operations or routines (eg, calling the API of a routine that performs processing on a regular expression). For example, a regular expression for a pattern based on syntactic constraints may be applied to the data to determine if this pattern in the data is identifiable.
The profile engine 326 can identify patterns in the data processed by the profile engine 326 by performing parsing tasks using one or more regular expressions. Regular expressions may be arranged according to the hierarchy. Patterns may be identified based on the order of regular expression complexity. Multiple patterns may match the data to be analyzed, and the more complex pattern is selected. As further described below, the profile engine 326 may perform statistical analysis to distinguish between patterns and patterns based on the application of regular expressions used to determine these patterns.
In some embodiments, the metadata descriptive attributes within this data may be analyzed by processing the unstructured data. The metadata itself can provide information about the data. By comparing this metadata, it is possible to identify similarities and / or determine the type of information. By comparing the information identified based on the data, the type of data (eg, business information, personal identification information, or address information) can be recognized and the data corresponding to the pattern can be identified.
According to embodiments, the profile engine 326 may distinguish patterns and / or text in the data by performing statistical analysis. The profile engine 326 may generate metadata that includes statistical information based on statistical analysis. Once the patterns have been identified, the profile engine 326 may distinguish between the plurality of patterns by seeking statistical information (eg, pattern metrics) for each of the different patterns. Statistical information may include standard deviations for different patterns of recognition. Metadata, including statistics, may be provided to other components of Data Enhancement Service 302, such as Recommendation Engine 308. For example, providing metadata to the recommendation engine 308 may allow the recommendation engine 308 to make recommendations for data enhancement based on the identified pattern. The recommendation engine 308 can obtain additional information about the pattern by querying the knowledge service 310 using the pattern. Knowledge service 310 may include or have access to one or more knowledge sources 340. Knowledge sources may include publicly available information published by websites, web services, curated knowledge stores, and other sources.
The profile engine 326 may distinguish identified patterns in the data by performing statistical analysis. For example, a pattern metric (eg, the statistical frequency of different patterns in the data) may be calculated for each of the different identified patterns in the data by evaluating the data analyzed by the profile engine 326. Each set of pattern metrics is calculated for different patterns within the identified patterns. The profile engine 326 may determine the differences between the pattern metrics calculated for different patterns. Based on this difference, one pattern may be selected from the identified patterns. For example, one pattern may be distinguished from another based on the frequency of the pattern in the data. In another example, if the data consists of dates with multiple different formats and each of these formats corresponds to a different pattern, the profile engine 326 will convert the date to a standard format in addition to normalization. Then, the standard deviation of each format may be obtained from different patterns. In this example, the profile engine 326 can statistically distinguish between multiple formats if there is a format with the lowest standard deviation. The pattern corresponding to the format of the data with the lowest standard deviation may be selected as the best pattern of data.
The profile engine 326 may determine the classification of the pattern to be identified. The profile engine 326 may determine whether the identified pattern can be classified within the knowledge domain by communicating with the knowledge service 310. Knowledge service 310 may determine one or more possible domains associated with the data based on techniques described herein, such as matching techniques and similarity analysis. Knowledge service 310 may provide profile engine 326 with a classification of one or more domains that may resemble the data identified in the pattern. The knowledge service 310 may provide a similarity metric indicating the degree of similarity to the domain for each domain identified by the knowledge service 310. The techniques disclosed herein for similarity metric analysis and scoring may be applied by recommendation engine 308 to determine the classification of data processed by profile engine 326. The metadata generated by the profile engine 326 may include information about the knowledge domain, if applicable, and a metric showing the similarity to the data analyzed by the profile engine 326.
The profile engine 326 may distinguish the identified text in the data by performing a statistical analysis, whether or not the patterns in the data are identified. The text may be part of a pattern, and text analysis may be used to further identify the pattern, if any. The profile engine 326 may determine whether the text can be classified into one or more domains by requesting the knowledge service 310 to perform a domain analysis on the text. Knowledge service 310 may serve to provide information about one or more domains applicable to the text being analyzed. The analysis performed by Knowledge Service 310 to determine the domain may be performed using techniques described herein, such as similarity analysis used to determine the domain of data.
In some embodiments, the profile engine 326 may identify textual data in the dataset. The text data may correspond to each identified entity in a set of entities. The classification may be determined for each identified entity. The profile engine 326 may require the knowledge service to identify the entity classification. When determining a set of classifications for a set of entities (for example, entities in a column), the profile engine 326 can distinguish between sets of classifications by calculating a set of metrics ("classification metrics"). Good. Each set of metrics may be calculated for each classification in a set of classifications. The profile engine 326 distinguishes a set of metrics by comparing them to each other and this set of entities. The closest classification may be determined as the classification of. The classification of a set of entities may be selected based on the classification representing this set of entities.
The knowledge service 310 can use the knowledge source 340 to match the context of the pattern identified by the profile engine 326. The knowledge service 310 may compare the identified patterns in the data, or the data if in the text, with the entity information of various entities stored in the knowledge source. Entity information may be obtained from one or more knowledge sources 340 using knowledge service 310. Examples of well-known entities may include social security numbers, telephone numbers, addresses, proper nouns, or other personal information. By comparing the data with the entity information of various entities, it may be determined whether or not it matches one or more entities based on the identified pattern. For example, Knowledge Service 310 can match the pattern "XXX-XX-XXXX" with the format of a US Social Security number. In addition, Knowledge Service 310 may determine that the social security number is protected or confidential and that its disclosure leads to various punishments.
In some embodiments, the profile engine 326 can distinguish between the plurality of classifications identified for the data processed by the profile engine 326 by performing statistical analysis. For example, if the text is categorized by multiple domains, the profile engine 326 can process the data to statistically determine the appropriate classification determined by the knowledge service 310. Statistical analysis of the classification may be included in the metadata generated by the profile engine 326.
In addition to pattern identification, the profile engine 326 can analyze the data statistically. The profile engine 326 can characterize the content of large amounts of data, and column by column the overall statistics about this data and the content of this data, such as its value, pattern, type, syntax, meaning and its statistical characteristics. Analysis can be provided. For example, numerical data can be analyzed statistically, for example, N, mean, maximum, minimum, standard deviation, skewness, kurtosis, and / or a 20-bin histogram (N is greater than 100). Includes large eigenvalues greater than K). The content may be categorized for the next analysis.
In one example, the overall statistics are, but not limited to, the number of rows, the number of columns, the number of unfilled and filled columns and how they change, rows that overlap with different rows, headers. It may include the number of columns classified by information, type or subtype, as well as the number of columns with security or other warnings. Column-specific statistics include the row (for example, K maximum frequency, K minimum frequency eigenvalue, eigenpattern, and type (if applicable)), frequency distribution, and text metric (for example, text length, token count,). Punctuation points, pattern-based tokens, and various useful text characteristics derived (minimum, maximum, average), token metrics, data types and subtypes, statistical analysis of numeric columns, mostly structured L maximum / minimum probabilities found in columns of undata, simple or compound terms or n-grams, as well as reference knowledge categories, date / time pattern discovery and formatting, reference data matching, and reference data matching, matched by this unique vocabulary. , May include the causative column heading label.
Using the resulting profile, by classifying the content for the next analysis, directly or indirectly, suggesting the transformation of the data, identifying the relationships between the data sources and retrieving before. You can validate the newly acquired data before applying a set of transformations designed based on the profile of the data you have created.
One or more conversion recommendations can be generated by feeding the metadata created by the profile engine 326 to the recommendation engine 308. Data can be enhanced with entities that match the identified pattern of data. This data is enhanced with the entities identified by the classification determined using Knowledge Service 310. In some embodiments, the knowledge service 310 may be provided with data related to the identified pattern (eg, city and state) to obtain entities matching the identified pattern from the knowledge source 340. .. For example, the entity information may be received by calling the knowledge service 310 and calling routines corresponding to the identified patterns (eg getCities () and getStates ()). The information received from this knowledge service 310 may include a list of entities (eg, a canonical list) that has the appropriate spelling information about the entity (eg, the city and state of the appropriate spelling). The entity information corresponding to the matching entity obtained from Knowledge Service 310 can be used to enhance the data, such as normalize the data, repair the data, and / or augment the data.
In some embodiments, the recommendation engine 308 can generate conversion recommendations based on the matching pattern received from the knowledge service 310. For example, for data containing Social Security numbers, the recommendation engine can recommend transformations that obfuscate the entry (for example, truncation, randomization, or deletion of all or part of the entry). Other examples of conversions include reforming data (for example, reforming dates in data), renaming data, enhancing data (for example, inserting values or mapping data to categories), and finding and replacing data (for example,). It can include modifying the spelling of data), changing character cases (eg changing cases from upper to lower case), and filtering based on blacklist or whitelist terms. In some embodiments, the recommendation may be tailored to a particular user to explain at a high level which data repairs or enhancements are available. For example, an obfuscation recommendation may indicate that the first five digits of the entry should be removed. In some embodiments, recommendations may be generated based on past user activity (eg, providing recommendation conversions used when previously identifying sensitive data).
The conversion engine 322 can generate a conversion script (for example, a script for obfuscating the social security number) based on the recommendation provided by the recommendation engine 308. Conversion scripts can transform data by performing operations. In some embodiments, the transformation script may implement a linear transformation of the data. The linear transformation may be achieved through an API (eg Spark API). The conversion action may be performed by an operation called using the API. The conversion script may be configured based on the conversion operation defined using the API. The operation may be performed on the basis of recommendations.
In some embodiments, the conversion engine 322 can automatically generate a conversion script to repair the data at the data source. Repairs can include automatically renaming columns, replacing strings or patterns within columns, correcting cases of text, reformatting data, and so on. For example, the conversion engine 322 can convert a date column based on a modification of the date format in the column or a conversion recommendation from the recommendation engine 308 by generating a conversion script. Recommendations may be selected from multiple recommendations to enhance or modify data from data sources processed by the profile engine 326. The recommendation engine 308 may make recommendations based on the metadata or profile provided by the profile engine 326. The metadata can indicate a column of dates identified for different formats (eg MM / DD / YYYY, DD-MM-YY, etc.). The conversion script generated by the conversion engine 322 can, for example, split and / or combine columns based on suggestions from the recommendation engine 308. The transformation engine 322 may also delete columns based on the data source profile received from the profile engine 326 (for example, an empty column or a column that contains information that the user does not want).
Conversion scripts can be defined using syntax that describes operations on one or more algorithms (eg Spark operator tree). Thus, the syntax can describe operator-tree transformation / simplification. The conversion script may be generated based on recommendations selected by the user or requested by the user through interaction through a graphical user interface. Examples of recommended transformations are described with reference to FIGS. 4A, 4B, 4C, and 4D. Based on the conversion operation specified by the user through the graphical user interface, the conversion engine 322 performs the conversion operation according to this operation. The dataset may be enhanced by recommending conversion operations to the user.
As further described below, client 304 may display recommendations that describe or otherwise indicate each recommended transformation. If the user chooses to run a conversion script, the selected conversion script may run on all or more data from the data source in addition to the data analyzed to determine the recommended conversion. it can. The resulting transformed data can then be published by the publishing engine 324 to one or more data targets 330. In some embodiments, the data target is a data store that is different from the data source. In some embodiments, the data target may be the same data store as the data source. The data target 330 may include a public cloud storage service 332, a private cloud storage service 334, various other cloud services 336, a URL or web-based data target 338, or any other accessible data target.
In some embodiments, the recommendation engine 308 can query the knowledge service 310 for other data associated with the identified platform. For example, if the data contains a column of city names, the relevant data (eg location, state, population, country, etc.) can be identified and recommendations to enhance the dataset with the relevant data can be displayed. Examples of recommendation display and data conversion through the user interface are shown below with reference to FIGS. 4-4D.
Knowledge service 310 may include matching module 312, similarity metric module 314, knowledge scoring module 316, and categorization module 318. As further described below, the matching module 312 can compare the data with the reference data available through the knowledge service 310 by implementing a matching method. Knowledge service 310 may include one or more knowledge sources 340 or have access to one or more knowledge sources 340. Knowledge sources may include publicly available information published by websites, web services, curated knowledge stores, and other sources. The matching module 312 may implement one or more matching methods as described in this disclosure. The matching module 312 may implement a data structure for storing the state associated with the matching method applied.
The similarity metric module 314 can implement a method for determining semantic similarity between two or more datasets. It can also be used to match user data against reference data available through Knowledge Service 330. The similarity metric module 314 may perform the similarity metric analysis described in the present disclosure, including description with reference to FIGS. 6-15.
The categorization module 318 can perform operations to implement automated data analysis. In some embodiments, the categorization module 318 can analyze the input dataset using an unsupervised machine learning tool such as Word2Vec. Word2Vec can take text input (eg a text corpus from a large data source) and generate a vector representation of each input word. The resulting model may then be used to identify how relevant the arbitrarily entered set of words is. For example, a Word2Vec model built with a large text corpus (eg, a news aggregator or other data source) can be used to find the corresponding numeric vector for each input word. When these vectors are analyzed, they may be determined to be "close" (in the Euclidean sense) in vector space. Although this can identify input words as being related (for example, identifying input words that are clustered in close proximity to each other in a vector space), Word2Vec has a label that describes the word (for example, "maker"). ) May not be useful for identifying. The categorization module 318 may implement an operation for categorizing related words using a curated knowledge source 340 (eg, YAGO of the Max Planck Institute for Informatics). The categorization module 318 can add other relevant data to the input dataset using the information from the knowledge source 340.
In some embodiments, the categorization module 318 may implement operations to further refine the categorization of related terms by performing trigram modeling. Trigram modeling can be used to compare word pairs for category identification. The input dataset can be augmented with related terms.
The matching module 312 compares words from the augmented dataset with categories of data from knowledge source 340 by implementing a matching method (eg, graph matching) with an input dataset that can contain additional data. be able to. The similarity metric module 314 can implement a method for determining the semantic similarity between the augmented dataset and each category in the knowledge source 340 to identify the name of that category. Category names may be chosen based on the maximum similarity metric. The similarity metric may be calculated based on the number of terms in the dataset that match the category name. The category may be selected based on the maximum number of matching terms based on the similarity metric. The techniques and operations performed for similarity analysis and categorization are further described in this disclosure, including description with reference to FIGS. 6-15.
In some embodiments, the categorization module 318 can augment the input dataset and use information from the knowledge source 340 to add other relevant data to the input dataset. For example, a data analysis tool such as Word2Vec can be used to identify words that are semantically similar to the words contained in the input dataset from a knowledge source such as a text corpus from a news collection service. In some embodiments, the categorization module 318 processes data retrieved from knowledge source 340 (YAGO, etc.) by implementing trigram modeling to generate a table of words indexed by category. can do. The categorization module 318 can then create a trigram for each word in the augmented dataset and match that word with the word from the indexed knowledge source 340.
The categorization module 318 uses the augmented dataset (or trigram-matched augmented dataset) to require the matching module 312 to compare words from the augmented dataset with categories of data from knowledge source 340. be able to. For example, each category of data in Knowledge Source 340 can be represented in a tree structure. The root node of the tree structure represents a category, and each leaf node represents a word that belongs to that category. The similarity metric module 314 can implement a method for determining the semantic similarity between the augmented dataset and each category in the knowledge source 510 (eg, Jaccard index or other similarity metric). The name of the category that matches the augmented dataset (eg, with the largest similarity metric) can then be applied to the input dataset as a label.
In some embodiments, the similarity metric module 314 measures the similarity between two datasets A and B to the size of the intersection of datasets A and B to the size of the union of these datasets. It can be judged by finding the ratio. For example, the similarity metric may be calculated based on the ratio of 1) the intersection of a dataset (eg, an augmented dataset) to a category, and 2) the size of a combination of these. The similarity metric may be calculated for comparison between the dataset and the category, as described above. Therefore, the "best match" may be determined based on the comparison of similarity metrics. The dataset used for this comparison may be augmented with labels corresponding to the categories for which the best match was determined using the similarity metric.
As mentioned above, other similarity metrics may be used in addition to or in place of the Jaccard coefficient. Those skilled in the art will appreciate that any similarity metric can be used for the above techniques. Some examples of alternative similarity metrics include, but are not limited to, the Dice-Sorensen coefficient, the Tversky coefficient, the Tanimoto metric, and the cosine similarity metric.
In some embodiments, the categorization module 318 utilizes a data analysis tool such as Word2Vec to provide a degree of agreement between the data from the knowledge source 340 and the input data that can be augmented with the data from the knowledge source. The exact metric (eg, score) shown may be calculated. Scores ("knowledge scores") may provide more knowledge about the similarity between the input dataset and the category being compared. The knowledge score may allow the data enhancement service 302 to select the category name that best represents the input data.
In the above technique, the category classification module 318 may count the number of term matches in the input dataset for the names of candidate categories (eg, genus) in the knowledge source 340. From the result of this comparison, a value representing a whole integer can be obtained. Thus, this value indicates the degree of agreement between terms, but may not indicate the degree of agreement between the input dataset and the various terms in the knowledge source.
The categorization module 318 may use Word2Vec to determine the comparative similarity between each term in the knowledge source (eg, a term representing a species) and a term in the input data (eg, a species). The categorization module 318 can use Word2Vec to calculate similarity metrics (eg, cosine similarity or distance) between an input dataset and one or more terms obtained from a knowledge source. Cosine similarity may be calculated as the cosine angle between a data set of terms (eg, a domain or genus) obtained from a knowledge source and an input dataset of terms. The cosine similarity metric may be calculated in the same way as the Tanimoto metric. The following equation shows an example of the cosine similarity metric.
<maths num="1"><img id="000006" he="20" wi="151" file="JP2017536601A_D0001.tif" img-format="tif" img-content="drawing" /></maths>
By calculating the similarity metric based on cosine similarity, each term in the input dataset is a whole-valued integer, such as a value that indicates the percentage of similarity between that term and the candidate category. value integer) may be regarded as 1 /. For example, as a result of calculating the similarity metric between the tire manufacturer and the surname, the similarity metric may be 0.3. On the other hand, as a result of calculating the similarity metric between the tire manufacturer and the company name, the similarity metric may be 0.5. By finely comparing the incomplete integer values that represent the similarity metric, category names with a high degree of matching can be made more accurate. The category name with the highest degree of matching may be selected as the most appropriate category name based on the similarity metric closest to the value 1. In the above example, based on the similarity metric, the company name is likely to be in the correct category. Therefore, the category classification module 318 may associate the "company" instead of the "last name" with the data string provided by the user, including the tire manufacturer.
Knowledge scoring module 316 can determine information about knowledge groups (eg domains or categories). Information about knowledge groups can be displayed in a graphical user interface, such as the example shown in Figure 10. Information about a knowledge domain can include a metric (eg, a knowledge score) that indicates the similarity between the knowledge domain and the input dataset of the term. The input data may be compared to the data from knowledge source 340. The input dataset may correspond to a column of data from the dataset specified by the user. The knowledge score can indicate the degree of similarity between the input dataset and one or more terms provided by the knowledge source. Each term corresponds to a knowledge domain. The columns of data may contain terms that belong to the knowledge domain in some cases.
In at least one embodiment, the knowledge scoring module 316 can determine a more accurate matching score. This score may correspond to a value calculated using a scoring formula. The scoring formula may be used to determine the semantic similarity between the terms of the two datasets, eg, the input dataset and the domain (eg, candidate category) obtained from the knowledge source. The domain whose matching score shows the best match (eg, maximum matching score) may be selected as the domain that has the greatest similarity to the input dataset. Therefore, the terms in the input dataset may be associated with the domain name as a category.
A scoring formula may be applied to an input dataset and a domain (eg, a category of terms taken from a knowledge source) to obtain a score that indicates the degree of agreement between this input data and the domain. A domain can have one or more terms that collectively define a domain. Scores may be used to determine the domains with the most similar input datasets. The input dataset may be associated with terms that describe the domains to which the input dataset is most similar.
The scoring formula may be based on one or more factors related to the domain compared to the input dataset. Factors in the scoring formula are, but are not limited to, degree values (eg, term frequencies where the input dataset matches terms in the domain), population values (eg, number of terms in the input dataset), and unique matching values. A range of values (for example, the number of different terms that match the input dataset and the domain), domain values (for example, the number of terms in the domain), and values that indicate how curated the domain is (for example, 0.0-100.0). It may include a curation level that indicates a constant value within. In at least one embodiment, the scoring formula may be defined as a function score (f, p, u, n, c). In this case, the scoring equation is calculated by the equation (1 + c / 100) * (f / p) * (log (u + 1) / log (n + 1)), where "f" is the logarithm. "C" represents the curation level, "p" represents the population value, "u" represents the unique matching value, and "n" represents the domain value.
The calculation of the scoring equation can be further explained with reference to FIG. In the reduced example, the input dataset (eg, a column of data in a table) may be defined as having a short text value of 100, and the knowledge source is a city domain, each with 1000 terms corresponding to the city. It may be defined by a domain containing (eg, "city") and a surname domain (eg, "last_name"), each having 800 terms corresponding to the surname. The input dataset may have 80 rows (each row corresponds to one term) that matches 60 terms (eg cities) in the city domain, and the input dataset may have 55 in the surname domain. It may have 65 lines that match an individual term (eg, last name). City domains may be defined at curation level 10 and surname domains may be defined at curation level 0 (eg, uncurated). Applying the scoring formula based on the values in this example and calculating the knowledge score for the city domain according to the score (80,100,60,1000,10), it is 0.5236209875770231 (for example, 52 or 52% of 100). Become. The knowledge score for the surname domain is calculated according to the score (65,100,55,800,0) to be 0.39134505184782975 (for example, 39 out of 100 or 39%). In this reduced example, based on knowledge scoring, the input dataset has a higher degree of match and a higher degree of similarity to the urban domain than the surname domain. Following this example, FIG. 10 shows an example of various domains with terms that are compared to the input dataset and a score ("matching") calculated using a scoring formula.
In some embodiments, the scoring formula may be based on more or less factors than above. By adjusting or modifying this equation, a score that better represents the match may be generated.
The profile engine 326 can statistically analyze the data in addition to pattern identification and matching. The profile engine 326 can characterize the content of large amounts of data, and column by column the overall statistics about the data and the content of the data, such as its value, pattern, type, syntax, meaning, and its statistical characteristics. Analysis and can be provided. For example, numerical data can be analyzed statistically, for example, N, mean, maximum, minimum, standard deviation, skewness, kurtosis, and / or a 20-bin histogram (N is greater than 100). Includes large eigenvalues greater than K). The content may be categorized for the next analysis. In some embodiments, the profile engine 326 can analyze data by one or more NL processors. It automatically identifies the columns in the data source, determines the type of data in a particular column, names the columns if there is no schema in the input, and / or the metadata that describes the columns and / or the data source. Can be provided. In some embodiments, the NL processor can identify and extract text entities in a column (eg, people, places, objects, etc.). The NL processor can also identify and / or build relationships within and / or between data sources.
In one example, the overall statistics are, but not limited to, the number of rows, the number of columns, the number of unfilled and filled columns and how they change, rows that overlap with different rows, headers. It may include the number of columns classified by information, type or subtype, as well as the number of columns with security or other warnings. Column-specific statistics include the row (for example, K maximum frequency, K minimum frequency eigenvalue, eigenpattern, and type (if applicable)), frequency distribution, and text metric (for example, text length, token count,). Punctuation points, pattern-based tokens, and various useful text characteristics derived (minimum, maximum, average), token metrics, data types and subtypes, statistical analysis of numeric columns, mostly structured L maximum / minimum probability found in columns of undata, simple or compound terms or ngrams, and reference knowledge categories matched by this unique vocabulary, date / time pattern discovery and formatting, reference data matching, and The causative column heading label may be included.
Using the resulting profile, by classifying the content for the next analysis, directly or indirectly, suggesting the transformation of the data, identifying the relationships between the data sources and retrieving before. You can validate the newly acquired data before applying a set of transformations designed based on the profile of the data you have created.
In some embodiments, the user interface 306 can generate one or more graphical visualizations based on the metadata provided by the profile engine 326. As mentioned above, the data provided by the profile engine 326 may include statistical information showing metrics for the data processed by the profile engine 326. Examples of graphical visualizations of the metric of profiled data are shown in Figures 5A-5D. Graphical visualizations can include graphical dashboards (eg, visualization dashboards). Graphical dashboards can show multiple metrics. Each of these metrics represents a real-time metric of the data relative to the time the data was profiled. Graphical visualizations may be displayed in the user interface. For example, by sending a graphical visualization to a client device, the client device can view the graphical visualization in the client device's user interface. In some embodiments, graphical visualizations may provide profiling results.
In addition, structural analysis by the profile engine 326 allows the recommendation engine to better direct its queries to knowledge services, resulting in faster processing and less load on system resources. For example, this information can be used to limit the scope of knowledge to be queried to prevent the knowledge service 310 from matching columns of numeric data to location names.
4A-4D show examples of user interfaces that provide interactive data enhancement according to embodiments of the present invention. As shown in FIG. 4A, a typical interactive user interface 400 can display at least a portion of the conversion script 402, the recommended conversion 404, and the data 406 to be analyzed / transformed. The conversion script 402 listed on the panel may include a conversion 406 that has already been applied to the data and is visible on the panel. Each conversion script 402 can be written in a simple declarative language that is easy for business users to understand. The conversion script 402 listed on the panel may be automatically applied to the data and reflected in a portion of the data 406 displayed in the interactive user interface 400. For example, the conversion script 402 listed on the panel contains a rename string to describe its contents. Column 408 shown in interactive user interface 400 has already been renamed according to conversion script 402 (for example, column 0003 has been renamed to date_time_02 and column 0007 has been renamed to "url"). However, the recommended conversion 404 is not automatically applied to the user's data.
As shown in FIG. 4B, the user can see the recommendation 404 in the recommendation panel and can identify the data to be changed based on this recommendation. For example, Recommendation 410 recommends renaming to "Col_0008 to city". Recommendations are written for business users to understand (rather than, for example, code or pseudocode), so users can easily identify the corresponding data 412. As shown in Figure 4B, data 412 contains a column of strings (represented as a row in user interface 400). The profile engine 326 can analyze the data to determine that it contains a string of two or less words (or tokens). This pattern can be given to the recommendation engine 308, which can query the knowledge service 310. In this case, Knowledge Service 310 matched this data pattern against the city name, and Recommendation 408 was generated to rename the column accordingly.
In some embodiments, the recommendation 404 listed on the panel may be applied towards the user (eg, in response to an instruction to apply the transformation) or may be applied automatically. Good. For example, in some embodiments, the knowledge service 310 can give a confidence score for a given pattern match. A threshold can be set on the recommendation engine 308 so that matches with a confidence score higher than this threshold are automatically applied.
When accepting a recommendation, the user may select the acceptance icon 414 (upward arrow icon in this example) associated with this recommendation. As shown in Figure 4C, the accepted recommendation 414 then goes to the panel of conversion script 402 and automatically applies the conversion to the corresponding data 416. For example, in the embodiment shown in FIG. 4C, Col_0008 has been renamed to "city" according to the transformation selected.
In some embodiments, the data enhancement service 302 may propose to add yet another string of data to the data source. Continuing with the example of "city", as shown in Figure 4D, transformation 418 is accepted to enhance the data with a new column containing the city's population and details of the city's location, including longitude and latitude. ing. Once selected, the user's dataset is enhanced to include this additional information 420. This dataset will then contain information that was previously not comprehensively and automatically available to the user. At this point, the user's dataset can be used to create a national map of location and population zones associated with other data in the dataset (for example, this corresponds to a corporate website transaction). May be attached).
5A-5D show examples of various user interfaces that provide visualization of datasets according to embodiments of the present invention.
FIG. 5A shows an example of a user interface that provides visualization of a dataset according to an embodiment of the present invention. As shown in Figure 5A, a typical interactive user interface 500 is a profile overview 502 ("profile result"), a conversion script 504, a recommended conversion 506, and at least one of the data to be analyzed / converted. The part 508 and can be displayed. The transformation 504 listed on the panel may include a transformation 508 that has already been applied to the data and can be seen on the panel.
Profile overview 502 may include overall statistics (eg, total rows and total columns) and column-specific statistics. Column-specific statistics can be generated by analyzing the data processed by the data enhancement service 302. In some embodiments, column-specific statistics can be generated based on the column information obtained by analysis of the data processed by the data enhancement service 302.
Profile overview 502 may include a map of the United States (eg, a "heat map"). The map shows different regions of the United States in different colors, based on statistics identified from the data 508 analyzed. This statistic may indicate how often these locations were identified as being associated with the data. In an example for illustration, the data may represent purchase transactions at an online retail store, where each transaction is based on, for example, a shipping / billing address or a recorded IP address). Can be associated with a location. Profile overview 502 may indicate the location of the transaction based on the processing of data representing the purchase transaction. In some embodiments, the visualization can be modified based on user input to help the user search the data to find useful correlations. These features will be further described below.
Figures 5B, 5C, and 5D show examples of the results of interactive data enhancement of datasets. Figure 5B shows a user interface 540 that may include a profile metric panel 542. Panel 542 can provide a summary of the metrics associated with the selected data source. As shown in FIG. 5C, in some embodiments, the profile metric panel 560 may include a metric 562 for a particular column rather than the entire data set. For example, the user may select a particular column on the user's client device and then view the profile 564 for the corresponding column. In this example, the profiler shows that there is a 92% match between column_0008 and a known city in the knowledge source. In some embodiments, the high probability allows the conversion engine to automatically label col_0008 as "city".
Figure 5D shows a profile metric panel 580 that can contain an overall metric 582 (for example, a metric related to the entire dataset) and a column-by-column visualization 584. The column-by-column visualization 584 can be selected and / or used by the user to navigate the data (eg, by clicking, dragging, swiping, etc.). The above example shows a simple conversion to a small dataset. Similar or more complex processing can be automatically applied to large datasets containing billions of records.
FIG. 6 shows a representative graph according to an embodiment of the present invention. In some embodiments, it may be useful to identify character strings in the text data. A character string may be a string that can be treated "in the literal sense" because there is no embedded syntax (like a regular expression). Performs full string matching when searching for character strings in the dataset. Character string matching can be performed by representing one or more character strings in a single data structure. This data structure may be used in combination with the graph matching method. In the graph matching method, all matching character strings of the input string are found at the same time by executing one pass for the input string. This improves matching efficiency because you only have to run the path once for the text to find all the strings.
In some embodiments, the graph matching method may be implemented as a variant of the Aho-Corasik algorithm. Graph matching works by storing character strings in a tree-like data structure and repeatedly traversing the tree looking for all possible matches with previously known characters in the input text. The data structure may be a tree whose nodes are the characters of the character string. The first character of any character string may be a child of the root node. The second character of any character string may be a child of the node corresponding to the first character. Figure 6 shows a tree of the following words: can (1), car (2), cart (3), cat (4), catch (5), cup (6), cut (7), and ten (8). Is shown. The last node of each word shows the number corresponding to that word.
In some embodiments, the graph matching method can track a particular list of matches. The partial match can include a pointer to a node in the tree and a character offset in the input string that corresponds to where the partial match is introduced. The graph matching method can be initialized with one partial match corresponding to the root node of the character at offset 1. Checks if the root node has a child with a given character when the first character is read. If such a child exists, the partial match node is advanced to the child node. Add the new partial match and offset 2 characters corresponding to the root node to the partial match list before reading the next character. Repeat this for all characters. For each character, evaluate all partial matches in the list with the current character. If the partial match node has a child corresponding to the current character, the partial match is maintained. Otherwise it must be deleted. If the partial match is maintained, the node proceeds to the child node corresponding to the current character. In some embodiments, if the node of the partial match corresponds to the end of the word (indicated by a number in FIG. 6), an exact match is generated and this can be added to the list of returned values. Before moving on to the next character, create a new partial match that corresponds to the root node and the character offset of the next character.
FIG. 7 shows a typical state table according to the embodiment of the present invention. For illustration, Table 700 shows the state of the graph matching method when examining the input string "cacatch". The following table shows the internal state stored by this method for each character in the input string after partial match evaluation. As shown in FIG. 7, the graph matching method can distinguish between the letters 2 to 4 "cat" (word 4) and the letters 2 to 6 "catch" (word 5). A partial match is represented as a pair that contains the character offset and the nodes in the tree given by the character. In some embodiments, the character offset can act as a partial match identifier (eg, the character offset can remain constant even if the node is updated), and the portion introduced at each character offset. There is only one match.
In some embodiments, new partial matches can be introduced for each character. Since the root node of the tree 600 has two children ("c" and "t"), the newly generated partial match will have the letters "c" and "t" as shown in lines 1, 3, 5, and 6. If "t" is found, it will proceed. Partial match "(1, a)" cannot be advanced on line 3. This is because "a" has no child with the letter "c". On the other hand, the partial match "(3, a)" is advanced to "(3, t)" on line 5. This is because the node "a" has a child with the letter "t". At this point, the partial match "(3, t)" is not an exact match for the word "cat" indicated by the number 4 in the "t" node. Therefore, 3 to 5 matches are found for the word "cat".
FIG. 8 shows an example of a case-insensitive graph according to an embodiment of the present invention. In some embodiments, case-sensitive matching can be performed on a second case-sensitive tree. A tree like the tree shown in Figure 7 is transformed into a grid that contains both uppercase (uppercase) and lowercase (lowercase) characters for each character, and any "case" of one character is both of the following characters. It can be pointed to a "case". Tree 800 represents a case-sensitive tree / grid for the word "can". As shown in Figure 8, for the word "can", there are routes for all combinations of cases (eg, "Can", "CAn", "CaN", "caN", etc.) in the tree.
However, case-sensitive and case-sensitive entries cannot exist in the same tree. This is because case-sensitive entries have a detrimental effect on the case sensitivity of case-sensitive entries. In addition, not all letters have a lowercase (lowercase) and an uppercase (uppercase) when performing case-sensitive matching. Therefore, the tree does not necessarily contain the corresponding character pairs. As an example, Tree 802 shows a match structure when the word "b2b" is added to the tree in a case-sensitive manner.
In some embodiments, case-sensitive matching can be supported by adding a second case-sensitive tree containing a character string added for case-sensitive matching. Graph matching may then be performed as described above, except that two partial matches for each character are added to the list of partial matches corresponding to the root nodes of each of the two trees.
FIG. 9 shows a diagram showing the similarity of datasets according to embodiments of the present invention. Embodiments of the present invention can semantically analyze datasets to determine semantic similarity between these datasets. Semantic similarity between datasets can be expressed as a semantic metric. For example, given a customer list C and a reference list R for an item, the "semantic similarity" between C and R is the Jaccard coefficient, the Sorensen-Dice coefficient (also known as the Sorensen coefficient of Dice's coefficient), and Tversky. It can be calculated using many well-known functions such as coefficients. However, existing methods do not adequately match close datasets. For example, as shown in Figure 9, all provincial capitals are cities, but not all cities are provincial capitals. Therefore, given a dataset C containing a list of 50 cities, 49 of them are state capitals and one is a non-state capital, and C matches the "list of cities" instead of the "list of state capitals". There is a need. Conventional methods such as the Jaccard coefficient and the Dice coefficient treat the dataset symmetrically, i.e., these methods do not distinguish between customer data and reference data. Then, as a result, there may be a situation in which some customer data is not matched with the reference data.
In some embodiments, the method of determining the similarity metric takes into account variations in the size of the reference dataset by using the natural logarithm. As a result, if the customer list has 100 items, the reference list for 1000 items that matches all the items in the customer list is twice (and) the reference list for 10,000 items that match all the customer items. It is an item (less than 10 times). The formula that describes how to find the similarity metric is shown below.
<maths num="2"><img id="000007" he="19" wi="151" file="JP2017536601A_D0001.tif" img-format="tif" img-content="drawing" /></maths>
In the equation, R is the reference dataset, C is the customer dataset, and α and β are adjustable coefficients. In some embodiments, the defaults are α = 0.1 and β = 0.1.
This method improves similarity matching by negatively weighting the characteristics of certain unwanted datasets. For example, as the reference set size increases, the α term increases and the similarity metric decreases. In addition, for curated reference datasets (generally assumed to be high value datasets), the β term is 0. However, for uncurated datasets, the β term is 1, which significantly reduces the similarity metric.
In some embodiments, this method can incorporate a vertex rank. This method does not result in a normalized similarity metric. Therefore, the similarity metric is multiplied by a negative weight as follows.
<maths num="3"><img id="000008" he="18" wi="163" file="JP2017536601A_D0001.tif" img-format="tif" img-content="drawing" /></maths>
FIG. 10 shows an example of a graphical interface 1000 that displays knowledge scoring for different knowledge domains according to embodiments of the present invention. As mentioned earlier, the graphical interface 1000 can display a graphical visualization of the domain matching data displayed by the data enhancement service 302. Graphical interface 1000 presents data that provides users with scoring formula-based statistics for different knowledge domains. A knowledge domain may include multiple terms related to a particular category (eg, domain), such as those identified in column 1002 ("domain"). Each domain 1002 can contain multiple terms and may be defined by a knowledge source. By curating the knowledge source, the terms associated with each domain 1002 may be maintained. The graphical interface 1000 presents various values that define the domain 1002 and a matching score 1016 ("score") obtained for each domain using a scoring formula. Each domain 1002 has a degree number 1004 (for example, the frequency of matching terms between the input dataset and the terms in the domain), a population value of 1006 (for example, the number of terms in the input dataset), and a matching value of 1008 (for example). For example, the percentage of terms that match the domain calculated based on dividing the degree number 1004 by the population), the unique matching value 1010 (for example, the number of different terms that match between the input dataset and the domain), It can have values such as size 1012 (eg, domain count indicating the number of terms in the domain) and selection values (eg, indicating the percentage of selected terms for the domain). Although not shown, the graphical interface 1000 has a range of values that indicate the degree to which the domain has been curated (eg 0.0-100. A curation level indicating a constant value in 00) may be indicated. Score 1016 may be calculated using a function for finding similarities such as the scores (f, p, u, n, c) above, based on one or more values for the domain. In addition to domain-related values, scores can provide precise measurements. This allows the user to make a better evaluation of the domain that best matches the input dataset. The closest matching domain may be used to name the data about the domain or map the data about the domain to the input dataset.
Some embodiments, such as those described with reference to FIGS. 11-18, may be described as processes shown in flowcharts, flow diagrams, data flow diagrams, structural diagrams, or block diagrams. Flowcharts may describe operations as sequential processes, but many of these operations can be performed in parallel or simultaneously. In addition, the order of operations may be reconfigured. The process ends when the operation is complete, but may have additional steps not included in the drawing. Processes can correspond to methods, functions, procedures, subroutines, subprograms, and so on. If the process corresponds to a function, its end may correspond to the return of that function to the calling function or main function.
The processes described herein, such as the processes described with reference to FIGS. 11-18, are software (eg, code, instructions, programs), hardware, executed by one or more processing units (eg, processor cores). , Or a combination of these, which can be implemented. The software may be stored in memory (eg, a memory device, a non-temporary computer-readable storage medium). In some embodiments, the processes shown in the flowcharts herein can be implemented by a data enhancement service, such as the computing system of the data enhancement service 302. The particular sequence of processing steps in this disclosure is not intended to be limiting. Other sequence of steps may also be performed according to alternative embodiments. For example, alternative embodiments of the present invention may perform the steps outlined above in other order. In addition, the individual steps shown in the drawings may include multiple substeps that can be performed in different orders suitable for the individual steps. In addition, other steps may be added or removed depending on the particular application. Those skilled in the art will recognize numerous modifications, modifications and alternatives.
In some aspects of some embodiments, each process of FIGS. 11-18 can be performed by one or more processing units. A processing unit may include one or more processors, including single-core or multi-core processors, one or more cores of processors, or a combination thereof. In some embodiments, a processing unit may include one or more dedicated coprocessors such as graphics processors, digital signal processors (DSPs), and the like. In some embodiments, some or all of the processing units can be implemented using customized circuits such as application specific integrated circuits (ASICs) or field programmable gate arrays (FPGAs).
Figure 11 shows an example of automated data analysis. As shown in 1100, unsupervised machine learning techniques, such as Word2Vec, can be used to analyze input datasets. Word2Vec can take text input (eg a text corpus from a large data source) and generate a vector representation of each input word. The resulting model can then be used to identify how closely the words of any input word set are related. For example, use K-means clustering (or other vector analysis) to analyze the vectors corresponding to a set of input words and see how similar the input words are and how "close" the corresponding vectors in the vector space are. Can be judged based on.
As shown in 1100, the set of words entered can include "Bridgestone", "Firestone", and "Michelin". A Word2Vec model built with a large text corpus (eg news aggregator or other data source) can be leveraged to identify the corresponding numeric vector for each input word. When analyzing these vectors, the vectors may be determined to be "close" (in the Euclidean sense) in vector space. As shown in FIG. 11, three input words are clustered in close proximity to each other in the vector space. This allows you to identify that the input word is related, but you cannot use Word2Vec to identify the label that describes the word (eg, "tire maker").
1102 shows how to categorize using a curated data source. Curated data source (Max Planck Institute for Informatics YAGO, etc.) can provide an ontology (eg, the type, characteristics, and formal names and definitions of relationships that exist for a particular domain). Curated data sources may be used to identify entities in the input dataset through graph matching. This allows genus labels (eg categories) to be identified for input datasets containing various species (eg words). However, as shown in 1102, genus labels can be incomplete or inaccurate (for example, different curators may categorize species differently). In the example shown in 302, such inaccuracies can result in matching of the same set of input words for different genera (eg for the genus Bridgestone and Michelin are tire makers). Matched and Firestone matched to the genus People in the Tire Industry). This provides a way to identify the genus for the input of species data, but the accuracy of the identified genus is limited by the accuracy and integrity of the curated data source.
FIG. 12 shows an example of the Trigram Modeling 1200. Traditionally, trigrams have been used to perform automatic spelling correction. As shown in Figure 12, each input word can be decomposed into trigrams, which can be indexed to create a table containing the words containing that trigram (eg, trigram "ANT". Is associated with "antique", "giant", etc.). When used for automatic spelling correction, dictionaries are used as the data source and trigrams can be used to identify the words that most closely resemble the misspellings entered. For example, a similarity metric can indicate similarity by the number of trigrams shared by a misspelled input word and a word from a data source (eg, a dictionary).
According to certain embodiments, trigram modeling can be used to compare word pairs to identify categories. As will be further described below with reference to FIG. 13, a trigram can be used to create an indexed table (similar to a database index) of a data source (such as a curated YAGO data source). In an indexed table, each trigram may be a primary key associated with a plurality of words containing the trigram. Each column in this table may correspond to the respective category associated with the word containing the trigram. Upon receiving the input dataset, matching words can be identified by dividing each word in the dataset into trigrams and comparing them to an indexed table. You can then identify the best match for the category by comparing the matching words to the data source. The judgment of statistical matching in trigram modeling is "Scalable string matching as a component for" It may be done using the technique described in US Patent Application No. 13 / 596,844 (Philip Ogren et al.) entitled "unsupervised learning in semantic meta-model development".
FIG. 13 shows an example of category labeling 1300 according to an embodiment of the present invention. It can receive input dataset 1302 as shown in Figure 13. Input dataset 1302 may include, for example, a column of text strings. In this example, the input dataset 1302 may include, for example, a column of text strings. In this example, the input dataset 1302 contains the strings "Bridgestone", "Firestone", and "Michelin". At 1304, a data analysis tool (eg Word2Vec) can be used to identify data similar to the input dataset. In some embodiments, the data analysis tool can create a word augmentation list by pre-processing data obtained from a data source, such as a news aggregation service. Similar words may then be identified by comparing the input dataset 1302 with the word augmentation list. For example, Word2Vec can be used to identify a vector for each string contained in the input dataset 1302. Vector analysis (eg, K-means clustering) can be used to identify other words in the word augmentation list that are "close" to the words in the input dataset. You can create augmented dataset 1306 containing similar words from the word augmentation list. As shown in FIG. 13, this involves creating augmented dataset 1306 by adding the word "Goodyear" to input dataset 1302.
In some embodiments, the augmented dataset 1306 can then be compared to the knowledge source 1308 to identify categories that match this augmented dataset. As shown in FIG. 13, knowledge source 1308 may include data organized by category. In some embodiments, each category can be represented by a root node, and each root node can have one or more leaf nodes representing the data belonging to that category. For example, Knowledge Source 1308 includes "tire manufacturers" and "people in the tire industry" as at least two categories. Each category contains data that belongs to that category (Michelin and Bridgestone belong to the tire manufacturer, Harvey Samuel Firestone belongs to the people of the tire industry).
According to certain embodiments, the augmented dataset can be compared to the knowledge source category using similarity metrics such as the Jaccard index. Similarity metrics allow you to compare one list with another and assign a value that indicates the similarity between the two datasets.
<maths num="4"><img id="000009" he="18" wi="139" file="JP2017536601A_D0001.tif" img-format="tif" img-content="drawing" /></maths>
Even if the knowledge source 1308 is incomplete or inaccurate, the similarity metric can identify the "best fit" category.
In some embodiments, the similarity metric may be calculated based on the Tanimoto metric. It calculates the generalization of the numeric vector of the Jaccard index for the Boolean vector. The following formula represents the Tanimoto metric.
<maths num="5"><img id="000010" he="20" wi="139" file="JP2017536601A_D0001.tif" img-format="tif" img-content="drawing" /></maths>
The dots represent the vector dot product. As shown above, the Jaccard index measures the similarity between two datasets A and B by the ratio of the size of the intersection of these datasets A and B to the size of the merged set of these datasets. Can be judged by asking for. As shown in 1314, the intersection of the augmented dataset 1306 with the category "tiremaker" is 2 (Michelin and Bridgestone), and the union size is 4, so the similarity metric is 0.5. The intersection metric is 0.25 because the intersection of the augmented dataset 1306 with the category People in the tire industry is 1 (Firestone) and the union size is 4. Thus, the "best match" is the "tire manufacturer" and the data enhancement service can enhance the input dataset by labeling the "tire manufacturer" column.
As mentioned above, other similarity metrics may be used in addition to or in place of the Jaccard coefficient. Those skilled in the art will recognize that any similarity metric can be used for the above techniques. Some examples of alternative similarity metrics include, but are not limited to, the Dice-Sorensen coefficient, the Tversky coefficient, the Tanimoto metric, and the cosine similarity metric.
Embodiments of the present invention are generally described with reference to "big data" systems and services. This is for the purpose of clarifying the explanation and is not limited. Those skilled in the art will appreciate that embodiments of the present invention may also be implemented in other systems and services outside the context of "big data."
In some embodiments, Trigram Statistical Analysis 1312 may be applied to augmented dataset 1306. The trigram modeling module preprocesses knowledge source 1308 into an indexed table with a primary key of trigram and each column containing at least one word of knowledge source. Each column with the same trigram can correspond to a different category within Knowledge Source 1308. A trigram can be created for each word in the augmented dataset 1306 and compared to the indexed trigram table. The result of this comparison is a list of words that are related to each other in Knowledge Source 1308, and the one with the highest degree of matching among the related words can be added to the trigram matching dataset. The trigram match dataset is then compared to the categories of knowledge source 1308 as described above to identify the category with the highest degree of match. The input dataset 1302 can then be labeled with a match category when the input dataset 1302 is exposed to one or more downstream data targets downstream.
In some embodiments, the indexed trigram table generated from knowledge source 1308 has a primary index column with trigrams (sorted alphabetically) and each category and subcategory with the same trigram on its leaf node. Can include a second column with a list of.
14 to 16 show similarity analysis for determining ranked categories according to embodiments of the present invention. Figure 14 shows the system 1400 for automated data analysis. System 1400 allows you to find categories and rankings associated with these categories for the term input dataset 1402. Data 1402 may be categorized using curated data obtained from knowledge sources by implementing automated data analysis.
As shown, the input dataset (eg, data 1402) can be obtained from a user-provided input source. Data 1402 may be formatted into two or more columns depending on the source. The data enhancement service 1408 can be implemented using a virtual computing environment such as Java® Virtual Machine (JVM). The data enhancement service 1408 may accept data 1402 as input. Data enhancement service 1402 may obtain curated data 1406 (eg, a curated list) from a curated data source (such as YAGO from the Max Planck Institute for Informatics). The curated data 1406 can provide an ontology (eg, the type, characteristics, and formal names and definitions of relationships that exist for a particular domain). Examples of curated data can include geonames that indicate zip codes with different geographic locations.
The data enhancement service 1408 can use a curated data source to perform a semantic analysis of data 1402 to determine similarity or proximity to curated data 1406. Semantic similarity between datasets can be represented by a similarity metric (eg, a value). For example, given an input dataset and a curated list, the similarity between this input dataset and the curated list can be calculated using a number of comparison functions, such as the Tversky metric. Based on the similarity metric obtained by comparison, the rank of proximity can be determined from the comparison between data 1402 and curated data 1406. Based on the comparison, category 1404 in the curated data 1406 can be judged and ranked by the similarity metric. By evaluating the ranked category 1404, the highest ranked category that can be associated with the data 1402 can be identified.
In another embodiment, FIG. 15 shows another example 1500 of a system for automated data analysis. System 1500 allows you to find categories and rankings associated with these categories for the term input dataset 1502. System 1500 may be implemented when it is determined that the categories identified by the similarity analysis described with reference to FIG. 14 are not in close proximity and these do not appear to be good matches.
With reference to FIG. 15, the data 1502 received from the user may be formatted into two or more columns depending on the source. The data enhancement service 1520 can be implemented using a virtual computing environment such as the Java® Virtual Machine (JVM). Data 1502 may be categorized using curated data obtained from knowledge sources by implementing automated data analysis. System 1500 may be implemented by aggregating data 1502 with aggregation services before performing similarity analysis. The data 1502 may be augmented with augmented data obtained from a source other than the data source from which the input data is supplied (eg, an aggregation service). For example, the input dataset may be augmented with data from a knowledge source, such as a text corpus from a news aggregation service (eg, Google® News Corpus) that is different from the source of the reference dataset. For example, a data analysis tool such as Word2Vec can be used to identify words (eg, synonyms) that are semantically similar to those contained in a dataset from a knowledge source. A word augmentation list can be generated by pre-processing the data obtained from the knowledge source. Similar words may then be identified by comparing the input dataset with the word augmentation list. For example, Word2Vec can be used to identify a vector for each string contained in data 1502. Vector analysis (eg, K-means clustering) can be used to identify other words in the word augmentation list that are "close" to the words in the input dataset. A augmented dataset containing similar words in a word augmentation list can be generated. The input dataset can be augmented with the augmented dataset. Input datasets with augmented data may be used for the rest of the process shown in Figure 15.
After augmentation of data 1502, the data enhancement service 1502 can determine similarity or proximity to the data obtained from the knowledge source 1516 by semantically analyzing the data 1502. In some embodiments, the knowledge source 1516 may provide data from a curated data source. Curated data can include curated categories and types within one or more files. Types may include taxonomy of terms to better identify categories for data 1502. In some embodiments, a curated list may be generated from the knowledge source 1516 by implementing the intermediate system 1514. System 15 may generate a list curated by a one-time operation under development in offline mode only, or retrieve the table from the beginning (from knowledge source 1516) each time the system is initialized. You may devise (from the data).
The curated data may be stored in a distributed storage system (eg HDFS). In some embodiments, the curated data may be stored in an indexed RDD1512.
The semantic similarity between the data may be determined by comparing the curated data with the already augmented data 1502. Semantic similarity between datasets can be represented by a similarity metric (eg, a value). For example, given an input dataset and a curated list, the similarity between this input dataset and the curated list can be calculated using a number of comparison functions, such as the Tversky metric. Based on the similarity metric obtained by comparison, the rank of proximity can be determined by comparing the data 1502 with the curated data. Categories can be identified by using the rank of proximity for data 1502. Category ranking 1510 may identify the highest ranked category that can be associated with data 1502.
Figure 16 shows process 1600 comparing an input dataset with a classification of a set of curated data obtained from a knowledge source such as YAGO. This set of classifications may include alternative spelling and alternative classification 1630. This set of classifications has subclasses in the higher category 1628 (for example, living). It may be arranged in a hierarchy included in thing)). For example, organism category 1628 may have an organism 1626 as a subclass, and the subclasses of organisms are person 1622, 1624. The subclass of a person is the type of person (eg intellectual), the subclass of the type of person is the training / profession 1618 (eg scholar), and the subclass of the scholar is the philosopher 1616. is there. This set of classifications may be used for similarity analysis and compared to the input dataset. Alternative spelling may be selectively weighted during matching with the input dataset as described, for example, with reference to FIGS. 14 and 15. Weights may be determined using a broader category or classification for comparison with the input dataset.
In the example shown in FIG. 16, the major relationship 1632 may be identified using one of the classifications 1630, 1614. Classification 1614 may include preferred spelling terms 1604 such as Aristotle 1606 and misspelling 1602 such as Aristotel 1608. A misspelling, eg Aristotel 1608, may be mapped to a label 1612 that corresponds to a term, eg, Aristotle 1606, with the correct spelling 1610 (eg Aristotle).
17 and 18 show a flow chart of the similarity analysis process according to some embodiments of the present invention. In some embodiments, the processes shown in the flowcharts of FIGS. 17 and 18 herein can be implemented by the computing system of the data enhancement service 302. Flowchart 1700 illustrates the process of similarity analysis, which determines their similarity by comparing the input data with one or more reference datasets. Similarity may be expressed as a degree of similarity that allows the user of the data enhancement service to identify the relevant dataset and enhance the input dataset.
Flowchart 1700 begins with step 1702, which receives an input dataset from one or more input data sources (eg, data source 309 in FIG. 3). In some embodiments, the input dataset is formatted into one or more strings of data.
At step 1704, the input dataset may be compared to one or more reference datasets obtained from the reference source. For example, a resource source is a knowledge source such as knowledge source 340. Comparing an input dataset with a reference dataset can include comparing each term individually or collectively in a comparison between two datasets. An input dataset can contain one or more terms. A reference data can contain one or more terms. For example, a reference dataset contains terms associated with a category (eg, domain or genus). The reference dataset may be curated by a knowledge service.
In some embodiments, the input data may be augmented with augmented data obtained from a source different from the data source from which the input data is supplied. For example, the input dataset may be augmented with data from a knowledge source that is different from the source of the reference dataset. For example, data analysis tools such as Word2Vec can be used to identify words that are semantically similar to those contained in input datasets from knowledge sources such as text corpora from news aggregation services. The data obtained from the knowledge source can be preprocessed to generate a word augmentation list. Similar words may then be identified by comparing the input dataset with the word augmentation list. For example, Word2Vec can be used to identify a vector for each string contained in the input dataset 602. Vector analysis (eg, K-means clustering) can be used to identify other words in the word augmentation list that are "close" to the word in the input dataset. From the word augmentation list, augmentation datasets containing similar words can be generated. The input dataset can be augmented with the augmented dataset. Input datasets with augmented data may be used for the rest of the process shown in Figure 17.
In some embodiments, data structures may be generated that are used to compare the input dataset with one or more reference datasets. This process may include generating data structures that represent at least a portion of one or more reference datasets to be compared. Each node in the data structure may represent a different character in one or more strings extracted from one or more reference datasets. The input dataset may be compared to one or more reference datasets that generated the data structure.
At step 1706, similarity metrics may be calculated for each of one or more reference datasets. The similarity metric may indicate the degree of similarity of each of the one or more reference datasets in comparison to the input dataset.
In some embodiments, the similarity metric is a matching score calculated for each of one or more reference datasets. For example, a matching score for one reference dataset, one or more, where the first value indicates the metric for this reference dataset and the second value indicates the metric based on the comparison between the input dataset and the reference dataset. It may be calculated using the value of. One or more of the above values are the degree value of the term that matches between the input dataset and the dataset, the population value of the dataset, the unique matching value of the dataset, and the input dataset and the dataset. It can include a unique matching value that indicates the number of different terms that match between, a domain value that indicates the number of terms in the dataset, and a curation level that indicates the degree of curation of the dataset. Matching score is calculated by implementing the scoring function (1 + c / 100) * (f / p) * (log (u + 1) / log (n + 1)) with one or more values. You may. The variables of the scoring function are "f" for degree value, "c" for curation level, "p" for population value, "u" for unique matching value, and domain value. Can include "n".
In some embodiments, the similarity metric may be calculated as a value based on the intersection cardinality of one or more reference datasets in comparison to the input dataset. This value may be normalized by cardinality. This value may be decremented by a first factor based on the size of the one or more reference datasets above, and this value may be decremented by a second factor based on the type of the one or more reference datasets above. You may.
In some embodiments, the similarity metric for each reference dataset in one or more reference datasets may also be calculated by determining the cosine similarity between the input dataset and this reference dataset. Good. As mentioned above, the cosine metric (eg cosine similarity or cosine distance) between the input dataset and one or more terms in the reference dataset is the reference dataset (eg domain or genus) obtained from a knowledge source. It may be calculated as the cosine angle between and the input dataset of the term. By calculating the similarity metric based on cosine similarity, each term in the input dataset is a fraction of a full integer, such as a value that indicates the percentage of similarity between that term and the candidate category. Can be regarded as.
In step 1708, the match between the input dataset and one or more reference datasets is identified based on the similarity metric. In some embodiments, identifying this match is the greatest degree of similarity based on the similarity metric calculated for each of the one or more reference datasets above. Includes determining the reference data in the set. By comparing the similarity metrics calculated for each of the above one or more reference datasets with each other, the reference datasets showing the closest match for the similarity metrics may be identified. The closest match may be identified as corresponding to the similarity metric with the highest value. The input dataset may be modified to include data contained in the reference dataset with the greatest degree of similarity.
The input dataset may be associated with other data, such as terms that describe or label this input dataset (eg, domain or category). Other data may be determined based on a reference dataset that may be curated. Other data may be obtained from the source that provided the reference dataset.
In step 1710, generate a graphical interface that is calculated for each of one or more reference datasets and shows a similarity metric that represents an identified match between the input dataset and this one or more reference datasets. May be good. In some embodiments where the similarity metric is a matching score, the graphical interface indicates the value used to calculate the matching score.
At step 1712, a graphical visualization may be rendered using a graphical interface. For example, you may want to display a graphical interface that results in the rendering of a graphical visualization. A graphical interface may contain data used to determine how to render a graphical visualization. In some embodiments, the graphical interface may be sent to another device (eg, a client device) for rendering. The graphical visualization may show the similarity metric calculated for each of one or more reference datasets, or may show an identified match between the input dataset and one or more reference datasets. .. Examples of graphical visualizations are illustrated with reference to FIGS. 5 and 10.
In some embodiments, the input dataset is calculated for each one or more reference datasets and presents a similarity metric that represents an identified match between the input dataset and the one or more reference datasets. It may be stored together with matching information.
Further, in some embodiments, the process shown in Flowchart 1700 may include one or more other steps after augmentation of the input data. The augmented input dataset may be used to identify matches with the reference dataset. In such an embodiment, the process may include generating an indexed trigram table based on one or more reference datasets. For each word in the augmented input dataset, create a trigram for that word, compare each trigram to the indexed trigram table, and the first of the trigrams in the indexed trigram table. Identifies the word associated with the trigram that matches the trigram of, and stores that word in the trigram augmentation dataset. The trigram augmented dataset may be compared to one or more reference datasets. Based on this comparison, a match between the trigram augmented dataset and one or more reference datasets may be determined. Identifying a match between the input dataset and one or more reference datasets in step 1708 uses a match between the trigram augmented dataset and one or more reference datasets based on a comparison. May include.
The flowchart may start at step 1802, which receives an input dataset from one or more data sources. At step 1804, the input dataset may be compared to one or more datasets stored by the knowledge source. The input dataset can contain one or more terms. Each of the above one or more datasets may contain one or more terms.
At step 1806, similarity metrics may be calculated for each of the one or more datasets compared to the input dataset. In some embodiments, the similarity metric is calculated by determining the cosine similarity between the input dataset and this dataset for each dataset of one or more datasets. The cosine similarity may be calculated as the cosine angle between the input dataset and the dataset being compared to this input dataset.
At step 1808, you may determine a match between one or more datasets and the input dataset. This match may be determined based on similarity metrics calculated for each of one or more datasets. Determining a match can include identifying the similarity metric that has the highest value in a set of similarity metrics. This set of similarity metrics may include similarity metrics calculated for each of the above one or more datasets.
At step 1810, a graphical user interface may be generated. The graphical user interface may show similarity metrics calculated for each of one or more datasets. The graphical user interface may show a match between one or more datasets and the input dataset. This match is determined based on the similarity metric calculated for each of one or more datasets. At step 1812, the graphical user interface may be rendered to display the similarity metrics calculated for each of one or more datasets. The graphical user interface may show the similarity metric with the highest value of the set of similarity metrics. This set of similarity metrics may include similarity metrics calculated for each of one or more datasets.
FIG. 19 shows a simplified diagram of the distributed system 1900 for implementing the embodiment. In the embodiments shown, the distributed system 1900 includes one or more client computing devices 1902, 1904, 1906, and 1908, which are web browsers, dedicated clients (through one or more networks 1910). For example, it is configured to run and operate client applications such as Oracle Forms). Server 1912 may be communicatively coupled to remote client computing devices 1902, 1904, 1906, and 1908 over network 1910.
In various embodiments, Server 1912 may be adapted to run one or more services or software applications, such as services and applications that provide processing related to the analysis and modification of documents (eg, web pages). .. In certain embodiments, Server 1912 may also provide other services or soft applications. This can include non-virtual and virtual environments. In some embodiments, these services are provided to users of client computing devices 1902, 1904, 1906, and / or 1908, either as web-based or cloud services, or under a software as a service (SaaS) model. Can be provided. Users operating client computing devices 1902, 1904 and 1906, and / or 1908 may utilize the services provided by these components by interacting with Server 1912 using one or more client applications. ..
In the configuration shown in FIG. 19, the software components 1918, 1920 and 1922 of system 1900 are shown to be implemented on server 1912. In other embodiments, one or more of the components of System 1900 and / or the services provided by these components are also by one or more of the client computing devices 1902, 1904, 1906, and / or 1908. It may be implemented. The user operating the client computing device can then utilize one or more client applications to use the services provided by these components. These components may be implemented in hardware, firmware, software, or a combination thereof. It should be recognized that a variety of different system configurations are possible that may differ from the distributed system 1900. The embodiment shown in FIG. 19 is therefore an example of a distributed system for implementing the system of the embodiment and is not intended to be limiting.
Client computing devices 1902, 1904, 1906, and / or 1908 may include various types of computing systems. For example, the client device is a portable handheld device (eg iPhone) that runs software such as Microsoft Windows Mobile® and / or various mobile operating systems such as iOS, Windows Phone, Android, BlackBerry 10, Palm OS, etc. (Registered Trademarks), mobile phones, iPad®, computing tablets, personal digital assistants (PDAs)) or wearable devices (eg Google) Glass® head-mounted display) may be included. These devices may support a variety of Internet-related applications, email, short message service (SMS) applications, and may use a variety of other communication protocols. Client computing devices also include, for example, personal computers and / or laptop computers running different versions of Microsoft Windows®, Apple Macintosh®, and / or Linux® operating systems. It may include a general purpose personal computer. Client computing devices are not limited, but for example Google Chrome A workstation computer running either of the various UNIX® or UNIX® operating systems available on the market, including various GNU / Linux® operating systems such as the OS. You may. Client computing devices can also communicate over network 1910, thin client computers, internet-enabled gaming systems (eg Microsoft Xbox game consoles with or without Kinect® gesture input devices), and / or personal messaging. It may include electronic devices such as devices.
Although Figure 19 shows a distributed system 1900 with four client computing devices, it can support any number of client computing devices. Other devices, such as devices with sensors, may interact with server 1912.
The network 1910 of the distributed system 1900 has a variety of available networks, including but not limited to TCP / IP (Transmission control protocol / Internet protocol), SNA (systems network architecture), IPX (Internet packet exchange), AppleTalk®, etc. It may be any type of network well known to those skilled in the art that can support data communications using any of the protocols. As an example, Network 1910 is a local area network (LAN), Ethernet®, token ring based network, wide area network, internet, virtual network, virtual private network (VPN), intranet, extranet, public. Exchange network (PSTN), infrared network, wireless network (eg Institute of Electrical and) A network operating under any of the Electronics (IEEE) 802.11 protocol suite, Bluetooth®, and / or any other wireless protocol), and / or any set of the above and / or other networks. It may be a combination.
Server 1912 is one or more general-purpose computers, dedicated server computers (including, for example, PC (personal computer) servers, UNIX (registered trademark) servers, midrange servers, mainframe computers, rack mount servers, etc.), server farms, etc. It may be configured with a server cluster, or any other suitable configuration and / or combination. Server 1912 may include one or more virtual machines running a virtual operating system, or other computing architecture with virtualization. You may maintain a virtual storage device for your server by virtualizing one or more flexible pools of logical storage devices. The virtual network can be controlled by Server 1912 using software-defined networking. In various embodiments, the server 1912 may be adapted to run one or more services or software applications described in previous disclosures. For example, server 1912 may correspond to a server for performing the above processing according to the embodiment of the present disclosure.
Server 1912 may run any of the above operating systems and operating systems including commercially available server operating systems. Server 1912 also includes a variety of additional server applications, including HTTP (hypertext transport protocol) servers, FTP (file transfer protocol) servers, CGI (common gateway interface) servers, JAVA® servers, database servers, and more. And / or can run any of the mid-tier applications. Typical database servers include, but are not limited to, those commercially available from Oracle, Microsoft, Sybase, IBM (International Business Machines), and the like.
In some implementation examples, the server 1912 may include one or more applications for analyzing and integrating data feeds and / or event updates received from users of client computing devices 1902, 1904, 1906, and 1908. .. As an example, data feeds and / or event updates include, but are not limited to, Twitter® feeds, Facebook® updates, or real-time updates received from one or more third party sources and continuous data streams. Can include. These may include real-time events related to sensor data applications, stock quote display devices, network performance measurement tools (eg network monitoring and traffic management applications), clickstream analysis tools, automotive traffic monitoring, etc. Server 1912 may also include one or more applications for displaying data feeds and / or real-time events through one or more display devices of client computing devices 1902, 1904, 1906, and 1908.
The distributed system 1900 may also include one or more databases 1914 and 1916. These databases may provide a mechanism for storing information such as user dialogue information, usage pattern information, adaptive rule information, and other information used in embodiments of the present invention. Databases 1914 and 1916 can reside in different locations. As an example, one or more of databases 1914 and 1916 may be on a non-temporary storage medium that is local to (and / or is in) the server 1912. Alternatively, databases 1914 and 1916 may be located remote from server 1912 and communicate with server 1912 via a network-based or dedicated connection. In one set of embodiments, databases 1914 and 1916 may be in a storage area network (SAN). Similarly, any files needed to perform functions attributed to Server 1912 may be stored in a location local to Server 1912 and / or remote from Server 1912, as appropriate. In one set of embodiments, databases 1914 and 1916 may include relational databases, such as databases provided by Oracle, that are adapted to store, update, and retrieve data in response to SQL-formatted instructions. ..
In some embodiments, the document analysis and modification service may be provided as a service via a cloud environment. FIG. 20 is a simplified block diagram of one or more components of a system environment 2000 that may provide a service as a cloud service according to an embodiment of the present disclosure. In the embodiment shown in FIG. 20, System Environment 2000 interacts with Cloud Infrastructure System 2002, which provides cloud services including services for dynamically modifying documents (eg, web pages) according to usage patterns. Includes one or more client computing devices 2004, 2006, and 2008 that users may use to. The cloud infrastructure system 2002 may include one or more computers and / or servers that may include those previously mentioned for Server 2012.
It should be recognized that the cloud infrastructure system 2002 shown in Figure 20 may have components other than those shown. Further, the embodiment shown in FIG. 20 is only an example of a cloud infrastructure system in which the embodiment of the present invention can be incorporated. In some other embodiments, the cloud infrastructure system 2002 may have more or fewer components than shown, may combine two or more components, or have different configurations or It may have an arrangement component.
Client computing devices 2004, 2006, and 2008 may be devices similar to those previously described for 1902, 1904, 1906, and 1908. Client computing devices 2004, 2006, and 2008 may be configured to operate client applications such as the following, which client applications include, for example, a user of a client computing device with a cloud infrastructure system 2002. A web browser, a dedicated client (eg Oracle Forms), or some other application that can be used to interact and use the services provided by the cloud infrastructure system 2002. A typical system environment 2000 is shown with three client computing devices, but can support any number of client computing devices. Other devices, such as devices with sensors and the like, may interact with the cloud infrastructure system 2002.
Network 2010 may facilitate the communication and exchange of data between clients 2004, 2006, 2008 and the cloud infrastructure system 2002. Each network can support data communication using any of the protocols available on various markets, including those mentioned above for Network 2010, any type of network well known to those of skill in the art. It may be.
In certain embodiments, the services provided by the cloud infrastructure system 2002 may include a number of services made available to users of the cloud infrastructure system on demand. In addition to services related to dynamically modifying documents according to usage patterns, various other services may also be provided. These services include, but are not limited to, online data storage and backup solutions, web-based e-mail services, hosted office package and document collaboration services, database processing, managed technical support services, and more. The services provided by the cloud infrastructure system can be dynamically scaled to the needs of its users.
In certain embodiments, the specific instantiation of the service provided by the cloud infrastructure system 2002 may be referred to herein as a "service instance." Generally, any service made available to a user from a cloud service provider's system via a communication network such as the Internet is called a "cloud service". Typically, in a public cloud environment, the servers and systems that make up the cloud service provider's system are different from the customer's own on-premises servers and systems. For example, the cloud service provider's system may host the application, and the user may order and use the application on demand via a communication network such as the Internet.
In some examples, services in computer network cloud infrastructure are storage devices, hosted databases, hosted web servers, software provided to users by cloud vendors or in other ways well known in the technology. It may include protected computer network access to applications or other services. For example, a service can include password-protected access to storage devices in the cloud over the Internet. As another example, the service can include a web service-based hosted relational database and scripting language middleware engine for private use by networked developers. As another example, the service can include access to an email software application hosted on a cloud vendor's website.
In certain embodiments, the Cloud Infrastructure System 2002 is a self-service, application-based, elastically scalable, reliable, highly effective, and secure way to give a customer an application, middleware, etc. And may include a set of database service offerings. An example of such a cloud infrastructure system is the Oracle Public Cloud provided by the assignee of the present application.
Cloud infrastructure system 2002 may also provide computational and analytical services related to "big data". The term "big data" is generally very much that analysts and researchers can store and manipulate to visualize large amounts of data, discover trends, and / or otherwise interact with the data. Used when referring to large datasets. This big data and related applications can be hosted and / or operated by infrastructure systems at many levels and at various scales. Dozens, hundreds, or thousands of processors linked in parallel, by acting on such data, demonstrate it or simulate external forces on the data or what it represents. Can be done. These datasets are structured and / or unstructured data such as data otherwise organized according to a structured model in the database (eg emails, images, data blobs ((binary large)). object) can include binary large objects), web pages, complex event handling). From companies, government agencies, research organizations, private individuals, individuals or organizations with the same purpose, or other entities by increasing the ability of embodiments to direct more (or less) computational resources to targets relatively quickly. Cloud infrastructure systems can be made more effective in performing tasks on large datasets based on the demands of.
In various embodiments, the cloud infrastructure system 2002 can be adapted to automatically provision, manage, and track customer applications for services provided by the cloud infrastructure system 2002. Cloud infrastructure system 2002 may provide cloud services through different deployment models. For example, the service may be provided under a public cloud model in which the cloud infrastructure system 2002 is owned by an organization that sells cloud services (eg owned by Oracle), and the service is for the general public or different industrial enterprises. It will be available. As another example, the service is provided under a personal cloud model in which Cloud Infrastructure System 2002 operates only for a single organization and may provide services for one or more entities within the organization. obtain. Cloud services may also be provided under a community cloud model in which the services provided by Cloud Infrastructure System 2002 and Cloud Infrastructure System 2002 are shared by some organizations within the relevant community. Cloud services can also be offered under a hybrid cloud model, which is a combination of two or more different models.
In some embodiments, the services provided by Cloud Infrastructure Systems 2002 are Software as a Service (SaaS) category, Platform as a Service (PaaS) category, Infrastructure as a Service (IaaS) category, or hybrid services. May include one or more services offered under other service categories, including. Customers may order one or more services provided by Cloud Infrastructure System 2002 by application order. Then, the cloud infrastructure system 2002 performs a process for providing a service in the customer's application order.
In some embodiments, the services provided by the cloud infrastructure system 2002 may include, but are not limited to, application services, platform services and infrastructure services. In some examples, application services may be provided via a SaaS platform by a cloud infrastructure system. The SaaS platform can be configured to provide cloud services that fall into the SaaS category. For example, a SaaS platform can provide the ability to build and communicate a complete set of on-demand applications on an integrated development and deployment platform. The SaaS platform can manage and control the underlying software and infrastructure for providing SaaS services. By utilizing the services provided by the SaaS platform, customers can take advantage of applications running on cloud infrastructure systems. Customers can get application services without the need for customers to purchase separate licenses and support. A variety of different SaaS services can be provided. Examples include, but are not limited to, services that provide solutions for sales performance management, corporate integration and business flexibility for large organizations.
In some embodiments, the platform service may be provided via the PaaS platform by the cloud infrastructure system 2002. The PaaS platform can be configured to provide cloud services that fall into the PaaS category. Examples of platform services are, but are not limited to, services that allow organizations (such as Oracle) to integrate existing applications on a shared common architecture and new applications that leverage the shared services provided by the platform. Can include the ability to build. The PaaS platform can manage and control the underlying software and infrastructure for providing PaaS services. Customers can get the PaaS services provided by Cloud Infrastructure Systems 2002 without the need for customers to purchase separate licenses and support. Examples of platform services are, but are not limited to, Oracle Java® Cloud Service (JCS), Oracle Database Cloud. Includes Service (DBCS) and others.
By leveraging the services provided by the PaaS platform, customers can adopt programming languages and tools supported by cloud infrastructure systems and also control the deployed services. In some embodiments, the platform services provided by the cloud infrastructure system are database cloud services, middleware cloud services (eg Oracle Fusion Middleware). services), and Java® cloud services may be included. In one embodiment, the database cloud service may support a shared service deployment model that allows an organization to pool database resources and present the database to customers as a service in the form of a database cloud. Middleware cloud services may provide a platform for customers to deploy and deploy various business applications, and Java® cloud services allow customers to deploy Java® applications in cloud infrastructure systems. You may provide a platform for doing so.
A variety of different infrastructure services can be provided by the IaaS platform in cloud infrastructure systems. Infrastructure services facilitate the management and control of basic computing resources such as storage, networks, and other basic computing resources for customers using the services provided by the SaaS and PaaS platforms. ..
In certain embodiments, the cloud infrastructure system 2002 may also include infrastructure resources 2030 to provide resources used to provide various services to customers of the cloud infrastructure system. In one embodiment, Infrastructure Resource 2030 is a pre-integrated and optimized combination of hardware such as servers, storage, and network resources to perform services provided by the PaaS and SaaS platforms, as well as May include other resources.
In some embodiments, resources in the cloud infrastructure system 2002 may be shared by multiple users and dynamically reassigned on a per-request basis. In addition, resources may be allocated to users in different time zones. For example, Cloud Infrastructure System 2002 makes the resources of a cloud infrastructure system available to a first set of users in a first time zone for a certain number of times, and then another set of users located in different time zones. Resource utilization may be maximized by reallocating the same resource to.
In certain embodiments, it may provide a number of internally shared services 2032 shared by different components or modules of cloud infrastructure system 2002. These internally shared services are for enabling, but not limited to, security and identity services, integration services, corporate repository services, corporate manager services, virus scanning and whitelisting services, high availability, post-save repair services, and cloud support. It may include services, email services, notification services, file transfer services, etc.
In certain embodiments, the cloud infrastructure system 2002 may provide comprehensive management of cloud services (eg, SaaS, PaaS and IaaS services) in the cloud infrastructure system. In one embodiment, the cloud management function may include the ability to provision, manage, and track customer applications received by the cloud infrastructure system 2002.
In one embodiment, as shown in FIG. 20, there is one cloud management function, such as order management module 2020, order orchestration module 2022, order provisioning module 2024, order management and monitoring module 2026, and identity management module 2028. It can be provided by the above modules. These modules may include general purpose computers, dedicated server computers, server farms, server clusters, or any other suitable configuration and / or combination of one or more computers and / or servers. , Can be provided using them.
In a typical operation, at 2034, a customer using a client device such as Client Device 2004, 2006, or 2008 requests one or more services provided by Cloud Infrastructure System 2002, by Cloud Infrastructure System 2002. You may interact with Cloud Infrastructure System 2002 by placing an order for the application for one or more of the services offered. In certain embodiments, the customer may access a cloud user interface (UI) such as Cloud UI 2012, Cloud UI 2014, and / or Cloud UI 2016 and place an application order through these UIs. The order information received by the cloud infrastructure system 2002 in response to the customer placing an order includes information that identifies the customer and one or more services provided by the cloud infrastructure system 2002 that the customer intends to apply for. May include.
In 2036, the order information received from the customer may be stored in the order database 2018. If this is a new order, a new record may be created for this order. In one embodiment, the order database 2018 may be one of several databases operated by the cloud infrastructure system 2018 and operated with other system elements.
At 2038, the order information may be transferred to the order management module 2020. The order management module 2020 may be configured to perform order-related billing and accounting functions, such as confirming an order and then filling out the order.
At 2040, information about the order may be communicated to the order orchestration module 2022. The order orchestration module 2022 is configured to coordinate the provisioning of services and resources for orders placed by customers. In some instances, the order orchestration module 2022 may use the services of the order provisioning module 2024 for provisioning. In certain embodiments, the order orchestration module 2022 allows management of the business process associated with each order and applies business logic to determine if the order should proceed to provisioning.
Upon receiving an order for a new application in 2042, as shown in the embodiment shown in FIG. 20, the order orchestration module 2022 sends a request to the order provisioning module 2024 to allocate resources and place the application order. Configure the resources needed to carry out. The order provisioning module 2024 allows the allocation of resources for services ordered by customers. The order provisioning module 2024 provides a level of abstraction between the cloud services provided by Cloud Infrastructure System 2000 and the physical implementation layer used to provision resources to provide resource services. This allows the order orchestration module 2022 to be separated from implementation details such as whether services and resources are actually provisioned on the fly or pre-provisioned and assigned / assigned after a request.
In 2044, once services and resources have been provisioned, a notification may be sent to the applying customer indicating that the requested service is currently available. In some instances, you may send the customer information (eg, a link) that will allow them to start using the service requested by the customer.
At 2046, customer application orders may be managed and tracked by the Order Management and Monitoring Module 2026. In some instances, the Order Management and Monitoring Module 2026 may be configured to collect usage statistics regarding customer usage of the requested service. For example, statistics may be collected for storage usage, data transfer, number of users, and system uptime and system downtime.
In certain embodiments, the cloud infrastructure system 2000 may include an identity management module 2028. Identity Management Module 2028 is configured to provide identity services such as access management and authorization services in Cloud Infrastructure System 2000. In some embodiments, the identity management module 2028 may control information about customers who want to use the services provided by the cloud infrastructure system 2002. Such information includes information that authenticates the identity of such customers and actions that those customers are authorized to perform on various system resources (eg files, directories, applications, communication ports, memory segments, etc.). It can include information to describe. Identity management module 2028 may also include the management of descriptive information about each customer and how and by whom the descriptive information can be accessed and modified.
FIG. 21 shows a typical computer system 2100 that can be used to implement embodiments of the present invention. In some embodiments, the computer system 2100 may be used to implement any of the various servers and computer systems described above. As shown in FIG. 21, the computer system 2100 includes various subsystems including a processor 2104 that communicates with a number of peripheral subsystems via the bus subsystem 2102. These peripheral subsystems may include a processing accelerator 2106, an I / O subsystem 2108, a storage subsystem 2118, and a communication subsystem 2124. The storage subsystem 2118 may include a tangible computer-readable storage medium 2122 and system memory 2110.
The bus subsystem 2102 provides a mechanism for the various components and subsystems of the computer system 2100 to communicate with each other in a purposeful manner. Bus subsystem 2102 is schematically shown as a single bus, but alternative embodiments of the bus subsystem may utilize multiple buses. Bus subsystem 2102 may be any of several types of bus structures, including memory buses or memory controllers, peripheral buses, and local buses, using any of a variety of bus architectures. For example, such an architecture can be implemented as a Mezzanine bus manufactured according to the IEEE P1386.1 standard, etc., Industry Standard Architecture (ISA) bus, Micro Channel Architecture (MCA) bus, Enhanced ISA (EISA) bus, Video Electronics Standards. Association (VESA) Local Bus and Peripheral Component It may include an Interconnect (PCI) bus.
The processing subsystem 2104 controls the operation of the computer system 2100 and may include one or more processing units 2132, 2134, and the like. The processor may include one or more processors, including single-core or multi-core processors, one or more cores of processors, or a combination thereof. In some embodiments, the processing subsystem 2104 may include one or more dedicated coprocessors such as graphics processors, digital signal processors (DSPs), and the like. In some embodiments, some or all of the processing units of processing subsystem 2104 may be implemented using customized circuits such as application specific integrated circuits (ASICs) or field programmable gate arrays (FPGAs). Good.
In some embodiments, the processing unit of processing subsystem 2104 can execute instructions stored in system memory 2110 or computer-readable storage medium 2122. In various embodiments, the processing unit can execute various programs or code instructions, and can maintain a plurality of programs or processes that are executed at the same time. At any given time, some or all of the program code to be executed may reside on system memory 2110 and / or computer-readable storage medium 2122, which may optionally include one or more storage devices. With proper programming, the processing subsystem 2104 can provide the various functions described above for dynamically modifying a document (eg, a web page) according to usage patterns.
In certain embodiments, the processing accelerator 2106 accelerates the entire processing performed by the computer system 2100 by performing customized processing or offloading some of the processing performed by the processing subsystem 2104. Can be provided for.
The I / O subsystem 2108 may include devices and mechanisms for inputting information to and / or outputting information from or through the computer system 2100. In general, when the term "input device" is used, it is intended to include all possible types of devices and mechanisms for inputting information into the computer system 2100. User interface Input devices include, for example, pointing devices such as keyboards, mice or trackballs, touchpads or touchscreens built into the display, scroll wheels, clickwheels, dials, buttons, switches, keypads, voice command recognition. It may include voice input devices with systems, microphones, and other types of input devices. User Interface Input Devices are also Microsoft Kinect® motion sensors, Microsoft that allow users to control and interact with input devices. It may include motion detection and / or gesture recognition devices such as an Xbox® 360 game controller, a device that provides an interface for receiving input using gestures and voice commands. User interface input devices may also include eye gesture recognition devices such as the Google Glass® blink detector. It detects the user's eye activity (eg "blinking" while shooting and / or menu selection) and transforms the eye gesture as input to an input device (eg Google Glass®). In addition, the user interface input device may include a voice recognition detector that allows the user to interact with a voice recognition system (eg, a Siri® navigator) by voice command.
Other examples of user interface input devices are, but are not limited to, audio / visual devices such as 3D (3D) mice, joysticks or pointing sticks, gamepads and graphic tablets, and speakers, digital cameras, digital video cameras, portable media. Includes players, webcams, image scanners, fingerprint scanners, bar code readers 3D scanners, 3D printers, laser rangers, and line-of-sight tracking devices. In addition, the user interface input device may include, for example, a medical imaging input device such as a computed tomography apparatus, a magnetic resonance imaging apparatus, a positron emission tomography apparatus, a medical ultrasonic examination apparatus, and the like. User interface input devices may also include audio input devices such as MIDI keyboards, digital musical instruments, and the like.
User interface output devices may include non-visual displays such as display subsystems, indicator lights, or audio output devices. The display subsystem may be a flat panel device such as a cathode ray tube (CRT), a liquid crystal display (LCD) or one using a plasma display, a projection device, a touch screen or the like. In general, when the term "output device" is used, it is intended to include all possible types of devices and mechanisms for outputting information from computer system 2100 to a user or other computer. For example, user interface output devices include, but are not limited to, monitors, printers, speakers, headphones, car navigation systems, plotters, audio output devices, and various modems that visually convey text, graphics, and audio / video information. Display devices may be included.
Storage subsystem 2118 provides a repository or data store for storing information used by computer system 2100. Storage subsystem 2118 provides a tangible, non-transitory computer-readable storage medium for storing basic programming and data structures that provide the functionality of some embodiments. Software (programs, code modules, instructions) that provide the above functionality when executed by processing subsystem 2104 may be stored in storage subsystem 2118. This software may be executed by one or more processing units of processing subsystem 2104. Storage subsystem 2118 may also provide a repository for storing data used in accordance with the present invention.
Storage subsystem 2118 may include one or more non-temporary memory devices, including volatile and non-volatile memory devices. As shown in FIG. 21, the storage subsystem 2118 includes system memory 2110 and a computer-readable storage medium 2122. System memory 2110 includes a large number of volatile main random access memory (RAM) for storing instructions and data during program execution, and non-volatile read-only memory (ROM) or flash memory for storing fixed instructions. Memory can be included. In some implementations, a basic input / output system (BIOS) containing basic routines that assist in the transfer of information between elements within the computer system 2100, such as during boot, will typically be stored in ROM. .. RAM typically contains the data and / or program modules that the processing subsystem 2104 is currently processing and executing. In some implementation examples, system memory 2110 may include multiple different types of memory, such as static random access memory (SRAM) or dynamic random access memory (DRAM).
As an example, but not a limitation, the system memory 2110 may include a client application, a web browser, a mid-tier application, a relational database management system (RDBMS), etc., as shown in FIG. May include operating system 2116. As an example, the operating system 2116 includes various versions of Microsoft Windows®, Apple Macintosh®, and / or Linux® operating systems, various UNIX®s available on the market, or UNIX® operating systems (including but not limited to various GNU / Linux® operating systems, Google Chrome® OS, etc.) and / or iOS, Windows® Phone, Android ( Registered trademark) OS, BlackBerry (registered trademark) 10 OS, Palm (registered trademark) It may include mobile operating systems such as OS operating systems.
The computer-readable storage medium 2122 may store programming and data structures that provide the functionality of some embodiments. Software (programs, codes, modules, instructions) that provide the above functions when executed by the processor of processing subsystem 2104 may be stored in storage subsystem 2118. As an example, the computer-readable storage medium 2122 is a non-volatile memory such as a hard disk drive or a magnetic disk drive, or a CD. It may include optical disk drives such as ROMs, DVDs, Blu-Ray® discs, or other optical media. Computer-readable storage medium 2122 may include, but is not limited to, Zip® drives, flash memory cards, universal serial bus (USB) flash drives, secure digital (SD) cards, DVD discs, digital video tapes, and the like. Computer-readable storage media 2122 also includes non-volatile memory-based solid state drives (SSDs) such as flash memory-based SSDs, corporate flash drives, solid-state ROMs, etc., solid-state RAMs, dynamic RAMs, static RAMs, DRAM-based SSDs. , Magnetic resistance RAM (MRAM) SSDs and other volatile memory based SSDs, as well as hybrid SSDs that use a combination of DRAM and flash memory based SSDs. The computer-readable medium 2122 may provide storage for computer-readable instructions, data structures, program modules, and other data for the computer system 2100.
In certain embodiments, the storage subsystem 2100 may also include a computer-readable storage medium reader 2120 that can be further connected to the computer-readable storage medium 2122. Along with system memory 2110 and optionally in combination with system memory 2110, computer readable storage medium 2122 is a remote, local, fixed, and / or removable storage device plus for comprehensively storing computer readable information. It may include a storage medium.
In certain embodiments, the computer system 2100 may provide support for running one or more virtual machines. The computer system 2100 can run hypervisor-like programs to facilitate the configuration and management of virtual machines. Memory, computation (eg processors, cores), I / O, and networking resources may be allocated to each virtual machine. Typically, each virtual machine runs its own operating system, which may be the same as or different from the operating system run by other virtual machines run by computer system 2100. Therefore, multiple operating systems can be run simultaneously by computer system 2100. Each virtual machine generally runs independently of the other virtual machines.
Communication subsystem 2124 provides an interface to other computer systems and networks. The communication subsystem 2124 functions as an interface for receiving data from a system other than the computer system 2100 and transmitting the data to the system other than the computer system 2100. For example, communication subsystem 2124 allows the establishment of communication channels to one or more client devices over the Internet for sending and receiving information to and from client devices.
Communication subsystem 2124 may support both wired and / or wireless communication protocols. For example, in certain embodiments, the communications subsystem 2124 (eg, mobile phone technology, 3G, 4G or advanced data network technologies such as EDGE (enhanced data rates for global evolution), WiFi (IEEE 802.11 series standards, or other). Radio frequency (RF) transceiver components, Global Positioning System (GPS) receiver components, and / or other for accessing wireless voice and / or data networks (using mobile communication technology, or any combination thereof). May include components. In some embodiments, the communication subsystem 2124 can provide wired network connectivity (eg, Ethernet®) in addition to or instead of a wireless interface.
Communication subsystem 2124 can receive and transmit various forms of data. For example, in some embodiments, the communications subsystem 2124 may receive input communications in the form of structured and / or unstructured data feeds 2126, event streams 2128, event updates 2130, and the like. For example, communications subsystem 2124 provides web feeds such as Twitter (registered trademark) feeds, Facebook (registered trademark) updates, Rich Site Summary (RSS) feeds, and / or real-time updates from one or more third-party sources. Such as, may be configured to receive (or send) real-time data feeds 2126 from users of social media networks and / or other communications services.
In certain embodiments, the communication subsystem 2124 may include an event stream 2128 and / or an event update 2130 of a real-time event that may be essentially continuous or infinite without an explicit end, of a continuous data stream. It can be configured to receive form data. Examples of applications that generate continuous data may include, for example, sensor data applications, stock quotes, network performance measurement tools (eg, network monitoring and traffic management applications), clickstream analysis tools, automotive traffic monitoring, and the like.
Communication subsystem 2124 may also communicate structured and / or unstructured data feeds 2126, event streams 2128, event updates 2130, etc. with one or more streaming data source computers coupled to computer system 21001 It can be configured to output to more than one database.
Computer system 2100 includes handheld portable devices (eg iPhone® mobile phones, iPad® computing tablets, PDA), wearable devices (eg Google Glass® headmount displays), personal computers, workstations. , Mainframes, kiosks, server racks, or any other type of data processing system.
Since the nature of computers and networks is constantly changing, the description of computer system 2100 shown in FIG. 21 is only intended as a concrete example. Many other configurations are possible with more or fewer components than the system shown in Figure 21. Based on the disclosures and teachings provided herein, one of ordinary skill in the art will recognize other ways and / or methods for implementing various embodiments.
In certain embodiments of the present invention, a data enhancement system is provided. The data enhancement system can be run in a cloud computing environment that includes a computing system, and the data enhancement system can communicate with multiple input data sources (eg, data source 104 shown in Figure 1) through at least one communication network. Combined with.
The data enhancement system includes a matching unit, a similarity metric unit, and a categorization unit. The matching unit, the similarity metric unit, and the category classification unit may be, for example, the matching module 312, the similarity metric module 314, and the category classification unit 318 shown in FIG. 3, respectively.
The matching unit is configured to compare input datasets received from multiple input data sources with one or more reference datasets obtained from a reference source (eg, knowledge source 340 shown in FIG. 3). The similarity metric section is configured to calculate the similarity metric for each one or more reference datasets, where the similarity metric is the degree of similarity for each one or more reference datasets in comparison to the input dataset. The similarity metric section is configured to identify a match between an input dataset and one or more reference datasets based on the similarity metric. The categorization section produces a graphical user interface that shows the similarity metrics calculated for each of one or more reference datasets and shows the identified match between the input dataset and one or more reference datasets. It is configured to do. In addition, a graphical visualization that shows the similarity metrics calculated for each of the one or more reference datasets and shows the identified match between the input dataset and the one or more reference datasets. Rendered using the user interface.
In certain embodiments of the invention, the data enhancement system further includes a knowledge scoring unit that may correspond to, for example, the knowledge scoring module 316 shown in FIG.
In certain embodiments of the invention, the one or more reference data sets include terms associated with a domain, and the similarity metric is a matching score calculated for each of the one or more reference data sets. The score uses one or more values by the Knowledge Scoring Department, including a first value that indicates the metric for the reference data set and a second value that indicates the metric based on the comparison between the input data set and the reference data set. Is calculated and rendered a graphical visualization to display one or more values used to calculate the matching score.
In certain embodiments of the invention, the one or more values are the degree values of terms that match between the input dataset and the dataset, the population value of the dataset, and the input dataset and the dataset. It includes a unique matching value that indicates the number of different terms that match between, a domain value that indicates the number of terms in the dataset, and a curation level that indicates the degree of curation of the dataset.
In certain embodiments of the invention, the categorization unit further generates an augmented list based on the augmented data obtained from the aggregation service, augments the input dataset based on the augmented list, into one or more reference datasets. Generate an indexed trigram table based on it, create a trigram for each word in the augmented input dataset, compare each trigram to the indexed trigram table, and out of the trigrams. Identify a word in the indexed trigram table associated with a trigram that matches the first trigram, store this word in the trigram augmented dataset, and include one or more trigram augmented datasets. It is configured to compare to a reference dataset and, based on this comparison, determine a match between the trigram augmented dataset and one or more reference datasets. Input data that is compared to one or more reference datasets above is augmented based on a augmented list, and identifying a match between an input dataset and one or more reference datasets is a comparison-based trial. Performed with a match between the gram augmentation dataset and one or more reference datasets.
In one embodiment of the invention, another data enhancement system is provided. The data enhancement system can be run in a cloud computing environment that includes a computing system, and the data enhancement system can communicate with multiple input data sources (eg, data source 104 shown in Figure 1) through at least one communication network. Combined with.
The data enhancement system includes a matching unit and a similarity metric unit. The matching unit and the similarity metric unit may be, for example, the matching module 312 and the similarity metric module 314 shown in FIG. 3, respectively.
The matching unit is configured to compare input datasets received from multiple input data sources with one or more reference datasets obtained from a reference source (eg, knowledge source 340 shown in FIG. 3). The similarity metric section is configured to calculate the similarity metric for each one or more reference datasets, where the similarity metric is the degree of similarity for each one or more reference datasets in comparison to the input dataset. The similarity metric section is configured to identify a match between an input dataset and one or more reference datasets based on the similarity metric. The input dataset is stored with matching information that shows the similarity metrics calculated for each of the one or more reference datasets and shows the identified match between the input dataset and the one or more reference datasets. To.
In one embodiment of the invention, the data enhancement system further includes a categorization unit, which may correspond to, for example, the categorization module 318 shown in FIG. The categorization section is configured to identify the category label of an input dataset based on the identification of a match between the input dataset and one or more reference datasets, and the input dataset is associated with this category label. Is stored.
In certain embodiments of the invention, the similarity metric is calculated using one or more of the Jaccard coefficient, Tversky coefficient, or Dice-Sorensen coefficient.
In certain embodiments of the invention, the input dataset is compared to one or more reference datasets using one or more of graph matching or semantic similarity matching.
Instead of a particular operating process for a unit / module (eg, engine) above, a corresponding step / component of a related method / system embodiment sharing the same concept may be referenced, which reference is to the relevant unit. It will be apparent to those skilled in the art that it will be considered as a disclosure of the module. Therefore, for the sake of brevity, some specific operational processes may not be repeated or detailed.
It will also be apparent to those skilled in the art that the units / modules can be implemented in electronic devices as software, hardware, and / or a combination of software and hardware. Components described as separate components may or may not be physically separated. In particular, the components according to each embodiment of the present invention may be integrated into one physical component or may be present in various separate physical components. All various implementations of the unit in electronic devices are within the scope of protection of the present invention.
It should be understood that units, devices, and devices may be implemented in the form of well-known or upcoming software, hardware, and / or combinations of such software and hardware.
It will be apparent to those skilled in the art that the operations shown in Figure 3 can be implemented in the form of software, hardware, and / or such software and hardware combinations, depending on the particular application environment. .. It will be apparent to those skilled in the art that at least some of the steps can be implemented by executing the instructions on a general purpose processor that stores the instructions in memory. It will also be apparent to those skilled in the art that at least some of the steps can be implemented by a variety of hardware, including but not limited to DSPs, FPGAs and ASICs. For example, the "operation" in some embodiments may be implemented by executing an instruction in a CPU that implements the function of the "operation" or a dedicated processor such as a DSP, FPGA, or ASIC.
Although specific embodiments of the invention have been described, various modifications, modifications, alternative configurations, and equivalents are also within the scope of the invention. The embodiments of the present invention are not limited to operations in a specific specific data processing environment, but function freely in a plurality of data processing environments. In addition, although embodiments of the present invention have been described using a particular set of transactions and steps, it will be apparent to those skilled in the art that the scope of the invention is not limited to the series of transactions and steps described above. The various features and aspects of the above embodiments may be used individually or jointly.
Further, although the embodiments of the present invention have been described using specific combinations of hardware and software, it must be recognized that other combinations of hardware and software are also included in the scope of the invention. Embodiments of the present invention can be realized with hardware only, software alone, or a combination thereof. The various processes described herein can be implemented on the same processor or on different processors in any combination. Therefore, if a component or module is described as being configured to perform a particular operation, such a configuration may be, for example, by designing an electronic circuit to perform that operation, or. This can be achieved by programming programmable electronic circuits (such as microprocessors) to perform their operations, or by arbitrarily combining them. Processes can be exchanged using a variety of techniques, including but not limited to traditional interprocess communication techniques, where different process pairs may use different techniques, or the same process pair may use different techniques from time to time. May be used.
Therefore, the specification and drawings should be considered in an exemplary sense rather than a limiting sense. However, it will be clear that additions, reductions, deletions, and other modifications and changes can be made without departing from the broad spirit and scope set forth in the claims. Therefore, although specific embodiments of the present invention have been described, these embodiments are not intended to be limited. Various amendments and equivalents are included in the claims below. Modifications include any relevant combination of disclosed features.
The data enhancement services described herein are sometimes referred to as IMI, ODECS, and / or Big Data Prep.
36 sheets
Sheet 1 Sheet 2 Sheet 3 Sheet 4 Sheet 5 Sheet 6 Sheet 7 Sheet 8 Sheet 9 Sheet 10 Sheet 11 Sheet 12 Sheet 13 Sheet 14 Sheet 15 Sheet 16 Sheet 17 Sheet 18 Sheet 19 Sheet 20 Sheet 21 Sheet 22 Sheet 23 Sheet 24 Sheet 25 Sheet 26 Sheet 27 Sheet 28 Sheet 29 Sheet 30 Sheet 31 Sheet 32 Sheet 33 Sheet 34 Sheet 35 Sheet 36
Every citation, both ways
| Document | Relation | Office | Category | Cited during | Relevant claims |
|---|---|---|---|---|---|
| JP2024507797A | Cited by | Japan | – | Search report | – |
| JP2024501893A | Cited by | Japan | – | Search report | – |
| US11379506B2 | Cited by | United States of America | – | Applicant | – |
| US10936599B2 | Cited by | United States of America | – | Applicant | – |
| US10810472B2 | Cited by | United States of America | – | Applicant | – |
| US11417131B2 | Cited by | United States of America | – | Applicant | – |
| JP2020057365A | Cited by | Japan | – | Search report | – |
| JP2022538591A | Cited by | Japan | – | Search report | – |
| US11500880B2 | Cited by | United States of America | – | Applicant | – |
| KR20230142109A | Cited by | Republic of Korea | – | Search report | – |
| US10885056B2 | Cited by | United States of America | – | Applicant | – |
| US10915820B2 | Cited by | United States of America | – | Applicant | – |
| JP2023514964A | Cited by | Japan | – | Search report | – |
| US2012101975A1 | Cites | United States of America | Y | Search report | 17 |
| US2013110792A1 | Cites | United States of America | XY | Search report | 1-3,5,7,9-16,18,17 |
28 members in 5 offices
Priority claims24
| Document | Office | Kind | Date |
|---|---|---|---|
| 201462056468 | United States of America | P | |
| 201462056468 | United States of America | P | |
| 62056468 | United States of America | – | |
| 201562163296 | United States of America | P | |
| 201562163296 | United States of America | P | |
| 62163296 | United States of America | – | |
| 201562203806 | United States of America | P | |
| 201562203806 | United States of America | P | |
| 62203806 | United States of America | – | |
| 14864485 | United States of America | – | |
| 201514864485 | United States of America | A | |
| 201514864485 | United States of America | A | |
| 2015052190 | United States of America | W | |
| 2015052190 | United States of America | W | |
| 14864485 | – | – | – |
| 62056468 | – | – | – |
| 62163296 | – | – | – |
| 62203806 | – | – | – |
| US201462056468P | – | – | – |
| US2015052190 | – | – | – |
| US201514864485 | – | – | – |
| US201562163296P | – | – | – |
| US201562203806P | – | – | – |
| WO2015US52190 | – | – | – |
Members28
| Document | Office | Kind | |
|---|---|---|---|
| US2016092090A1 | United States of America | A1 | |
| US2016092474A1 | United States of America | A1 | |
| US2016092475A1 | United States of America | A1 | |
| US2016092476A1 | United States of America | A1 | |
| US2016092557A1 | United States of America | A1 | |
| WO2016049437A1 | World Intellectual Property Organization (WIPO) | A1 | |
| WO2016049460A1 | World Intellectual Property Organization (WIPO) | A1 | |
| WO2016049437A9 | World Intellectual Property Organization (WIPO) | A9 | |
| CN106687952A | China | A | |
| CN106796595A | China | A | |
| EP3198482A1 | European Patent Office (EPO) | A1 | |
| EP3198484A1 | European Patent Office (EPO) | A1 | |
| JP2017534108A | Japan | A | |
| JP2017536601AThis record | Japan | A | |
| US10210246B2 | United States of America | B2 | |
| US2019138538A1 | United States of America | A1 | |
| US10296192B2 | United States of America | B2 | |
| JP6568935B2 | Japan | B2 | |
| US10891272B2 | United States of America | B2 | |
| US10915233B2 | United States of America | B2 | |
| US10976907B2 | United States of America | B2 | |
| JP2021061063A | Japan | A | |
| CN106796595B | China | B | |
| US2021223947A1 | United States of America | A1 | |
| CN106687952B | China | B | |
| US11379506B2 | United States of America | B2 | |
| JP7148654B2 | Japan | B2 | |
| US11693549B2 | United States of America | B2 |
13 legal events, as the office reported them to INPADOC
Over the term
Point at a mark for the eventEvents
| Event | Code | |
|---|---|---|
| Receipt of annual feesJAPANESE INTERMEDIATE CODE: R250R250 | R250 | |
| Receipt of annual feesJAPANESE INTERMEDIATE CODE: R250R250 | R250 | |
| Receipt of annual feesJAPANESE INTERMEDIATE CODE: R250R250 | R250 | |
| Receipt of annual feesJAPANESE INTERMEDIATE CODE: R250R250 | R250 | |
| Certificate of patent or registration of utility modelJAPANESE INTERMEDIATE CODE: R150R150 | R150 | |
| First payment of annual fees (during grant procedure)JAPANESE INTERMEDIATE CODE: A61A61 | A61 | |
| Written decision to grant a patent or to grant a registration (utility model)JAPANESE INTERMEDIATE CODE: A01A01 | A01 | |
| Decision of grant or rejection writtenTRDD | TRDD | |
| Request for written amendment filedJAPANESE INTERMEDIATE CODE: A523A521 | A521 | |
| Notification of reasons for refusalJAPANESE INTERMEDIATE CODE: A131A131 | A131 | |
| Report on retrievalJAPANESE INTERMEDIATE CODE: A971007A977 | A977 | |
| Request for written amendment filedJAPANESE INTERMEDIATE CODE: A523A521 | A521 | |
| Written request for application examinationJAPANESE INTERMEDIATE CODE: A621A621 | A621 |
Numbers
- Publication
- 2017536601
- Publication, DOCDB
- 2017536601
- Publication, EPODOC
- JP2017536601
- Application
- 2017516310
- Application, DOCDB
- 2017516310
- Application, EPODOC
- JP20170516310
Titles2
- Japanese
- 知識ソースを用いた類似性分析およびデータ強化の技術
- English
- Similarity analysis and data enhancement techniques using knowledge sources
Classification
- CPC, 10
- G06F16/334
- G06F16/35
- G06Q30/02
- G06F16/215
- G06F16/248
- G06F16/254
- G06F16/9024
- G06Q30/0631
- G06F16/285
- G06F16/355
- IPC, 1
- G06F17 30
Designated states5
- Regional, 4
- Zimbabwe
- Turkmenistan
- Türkiye
- Togo
- National, 1
- United States of America