Nova Patents
CA2845995A1

Data clustering

Abstract

In various embodiments, systems, methods, and techniques are disclosed for generating a collection of clusters of related data from a seed. Seeds may be generated based on seed generation strategies or rules. Clusters may be generated by, for example, retrieving a seed, adding the seed to a first cluster, retrieving a clustering strategy or rules, and adding related data and/or data entities to the cluster based on the clustering strategy. Various cluster scores may be generated based on attributes of data in a given cluster. Further, cluster metascores may be generated based on various cluster scores associated with a cluster. Clusters may be ranked based on cluster metascores. Various embodiments may enable an analyst to discover various insights related to data clusters, and may be applicable to various tasks including, for example, financial fraud detection.

CA2845995A1, drawing sheet 1
Sheet 1 of 11

Term

Projected expiry 13 March 2034.

  1. Priority
  2. Filed
  3. Published
  4. Today
  5. Projected expiry

45 claims: 6 independent, 39 dependent

  1. 1
    CA 02845995 2014-03-13 CLAIMS:1. A computer-implemented method for generating a cluster of related data entities, the method comprising: communicating with one or more electronic data stores storing a plurality of data entities and respective data entity attributes, the one or more electronic data stores in communication with one or more hardware computer processors, the one or more hardware computer processors configured with specific computer executable instructions, and the plurality of data entities related to financial data and including at least one of: an account data entity, a transaction data entity, a customer data entity, or a phone number data entity;retrieving, from the one or more electronic data stores and by the one or more hardware computer processors, a seed data entity selected from the plurality of data entities;adding, by the one or more hardware computer processors, the seed data entity to a data entity cluster;identifying, by the one or more hardware computer processors, at least one data entity attribute associated with the seed data entity;determining, by the one or more hardware computer processors and based on a cluster strategy and the at least one data entity attribute, one or more additional data entities related to the seed data entity, wherein the cluster strategy is configured to identify related data entities for detection of possible fraudulent financial activity;and adding, by the one or more hardware computer processors and based on the cluster strategy, the one or more additional data entities to the data entity cluster, -32CA 02845995 2014-03-13 wherein the data entity cluster is useable by an analyst to determine various data entities related to each other and to possible fraudulent financial activity.
  2. 9
    A non-transitory computer-readable storage medium storing computerexecutable instructions that, when executed by a computer system, configure the computer system to perform operations comprising:communicating with one or more electronic data stores storing a plurality of data entities and respective data entity attributes, the plurality of data entities related to financial data and including at least one of: an account data entity, a transaction data entity, a customer data entity, or a phone number data entity;retrieving, from the one or more electronic data stores, a seed data entity selected from the plurality of data entities;-34CA 02845995 2014-03-13 adding the seed data entity to a data entity cluster;identifying at least one data entity attribute associated with the seed data entity;determining, based on a cluster strategy and the at least one data entity attribute, one or more additional data entities related to the seed data entity, wherein the cluster strategy is configured to identify related data entities for detection of possible fraudulent financial activity;and adding, based on the cluster strategy, the one or more additional data entities to the data entity cluster, wherein the data entity cluster is useable by an analyst to determine various data entities related to each other and to possible fraudulent financial activity.
  3. 17
    A computer system comprising:one or more non-transitory computer readable storage devices configured to store one or more software programs;and one or more hardware computer processors in communication with the one or more non-transitory computer readable storage devices and configured to execute the one or more software programs in order to cause the computer system to: communicate with one or more electronic data stores storing a plurality of data entities and respective data entity attributes, the plurality of data entities related to financial data and including at least one of: an account data entity, a transaction data entity, a customer data entity, or a phone number data entity;retrieve, from the one or more electronic data stores and by the one or more hardware computer processors, a seed data entity selected from the plurality of data entities;add, by the one or more hardware computer processors, the seed data entity to a data entity cluster;identify, by the one or more hardware computer processors, at least one data entity attribute associated with the seed data entity;determine, by the one or more hardware computer processors and based on a cluster strategy and the at least one data entity attribute, one or more additional data entities related to the seed data entity, wherein the cluster -37CA 02845995 2014-03-13 strategy is configured to identify related data entities for detection of possible fraudulent financial activity;and add, by the one or more hardware computer processors and based on the cluster strategy, the one or more additional data entities to the data entity cluster, wherein the data entity cluster is useable by an analyst to determine various data entities related to each other and to possible fraudulent financial activity.
  4. 25
    A computer-implemented method for prioritizing a plurality of data entity clusters, the method comprising:communicating with one or more electronic data stores storing a plurality of data entities and respective data entity attributes, the one or more electronic data stores in communication with one or more hardware computer processors, the one or more hardware computer processors configured with specific computer executable instructions, and the plurality of data entities related to financial data and including at least one of: an account data entity, a transaction data entity, a customer data entity, or a phone number data entity;generating, by the one or more hardware computer processors, a plurality of data entity clusters that each store a collection of related data entities, each of the data entity clusters generated by: identifying and adding a seed data entity to the data entity cluster;and determining and adding to the data entity cluster, based on a cluster strategy configured to identify related data entities for detection of possible fraudulent financial activity, one or more additional data entities related to the seed data entity;identifying, by the one or more hardware computer processors, a scoring strategy for prioritizing the plurality of data entity clusters;for each data entity cluster: evaluating, by the one or more hardware computer processors and based on the scoring strategy, the data entity cluster;and assigning, by the one or more hardware computer processors and based on the evaluation, a score to the data entity cluster;and ranking the plurality of data entity clusters according to the respective assigned scores, -40CA 02845995 2014-03-13 wherein the plurality of ranked data entity clusters are useable by an analyst to determine various data entities related to each other and to possible fraudulent financial activity.
  5. 32
    A non-transitory computer-readable storage medium storing computerexecutable instructions that, when executed by a computer system, configure the computer system to perform operations comprising:communicating with one or more electronic data stores storing a plurality of data entities and respective data entity attributes, the plurality of data entities related to financial data and including at least one of: an account data entity, a transaction data entity, a customer data entity, or a phone number data entity;generating a plurality of data entity clusters that each store a collection of related data entities, each of the data entity clusters generated by: identifying and adding a seed data entity to the data entity cluster;and determining and adding to the data entity cluster, based on a cluster strategy configured to identify related data entities for detection of possible fraudulent financial activity, one or more additional data entities related to the seed data entity;identifying a scoring strategy for prioritizing the plurality of data entity clusters;for each data entity cluster: evaluating, based on the scoring strategy, the data entity cluster;and assigning, based on the evaluation, a score to the data entity cluster;and -42CA 02845995 2014-03-13 ranking the plurality of data entity clusters according to the respective assigned scores, wherein the plurality of ranked data entity clusters are useable by an analyst to determine various data entities related to each other and to possible fraudulent financial activity.
  6. 39
    A computer system comprising:one or more non-transitory computer readable storage devices configured to store one or more software programs;and one or more hardware computer processors in communication with the one or more non-transitory computer readable storage devices and configured to execute the one or more software programs in order to cause the computer system to: communicate with one or more electronic data stores storing a plurality of data entities and respective data entity attributes, the plurality of data entities related to financial data and including at least one of: an account data entity, a transaction data entity, -44 CA 02845995 2014-03-13 a customer data entity, or a phone number data entity;generate, by the one or more hardware computer processors, a plurality of data entity clusters that each store a collection of related data entities, each of the data entity clusters generated by: identifying and adding a seed data entity to the data entity cluster;and determining and adding to the data entity cluster, based on a cluster strategy configured to identify related data entities for detection of possible fraudulent financial activity, one or more additional data entities related to the seed data entity;identify, by the one or more hardware computer processors, a scoring strategy for prioritizing the plurality of data entity clusters;for each data entity cluster: evaluate, by the one or more hardware computer processors and based on the scoring strategy, the data entity cluster;and assign, by the one or more hardware computer processors and based on the evaluation, a score to the data entity cluster;and rank the plurality of data entity clusters according to the respective assigned scores, wherein the plurality of ranked data entity clusters are useable by an analyst to determine various data entities related to each other and to possible fraudulent financial activity.