Nova Patents
EP3627345B1

Data clustering

Abstract

This record has no abstract on file.

EP3627345B1, drawing sheet 1
Sheet 1 of 10

Term

7.5 yearsleft in the term

Expires 13 March 2034.

  1. Priority
  2. Filed
  3. Granted
  4. Today
  5. Expires

14 claims: 10 independent, 4 dependent

  1. 1
    A computer-implemented method for prioritizing a plurality of data entity clusters, the method comprising:communicating with one or more electronic data stores (820, 830) storing a plurality of data entities and respective data entity attributes, the one or more electronic data stores in communication with one or more hardware computer processors (860), the one or more hardware computer processors configured with specific computer executable instructions;generating, by the one or more hardware computer processors, a plurality of data entity clusters that each store a collection of related data entities, wherein generating each of the data entity clusters comprises: designating a seed data entity (302) as an initial data entity of the data entity cluster (252, 252-1);accessing, based on a cluster strategy (232), two or more search protocols (237-1, 237-2, ... , 237-M);executing a first search protocol of the two or more search protocols on the one or more electronic data stores to identify one or more data entities (305-1) related to the seed data entity;adding the one or more data entities to the data entity cluster;iteratively: executing each of the other search protocols of the two or more search protocols on the one or more electronic data stores to identify one or more additional data entities (305-2, ... , 305-6) related to the one or more data entities previously added to the data entity cluster, wherein the search protocol uses at least one data entity attribute (303-2, ... , 303-6) associated with at least one of the one or more data entities previously added to the data entity cluster as a data input parameter to the search protocol, and adding the one or more additional data entities to the data entity cluster (252-1);identifying, by the one or more hardware computer processors, a scoring strategy for prioritizing the plurality of data entity clusters;for each data entity cluster: evaluating, by the one or more hardware computer processors and based on the scoring strategy, the data entity cluster;and assigning, by the one or more hardware computer processors and based on the evaluation, a score to the data entity cluster;ranking the plurality of data entity clusters according to the respective assigned scores;and storing the plurality of data entity clusters.
  2. 5
    The method of any preceding claim, wherein storing the plurality of data entity clusters comprises storing the ranked plurality of data entity clusters.
  3. 6
    The method of any preceding claim, wherein assigning a score to the data entity cluster comprises:determining a plurality of base scores for the data entity cluster;determining, based on the plurality of base scores, an overall score for the data entity cluster;and assigning the overall score to the data entity cluster.
  4. 7
    The method of any preceding claim, further comprising:presenting, on an electronic display viewable by a user, a listing and/or summary of one or more of the plurality of data entity clusters according to the ranking, and optionally wherein the listing and/or summary presents the seed data entity associated with each of the one or more or the plurality of clusters.
  5. 8
    The method of any preceding claim, further comprising:identifying a second scoring strategy for prioritizing the plurality of data entity clusters;for each data entity cluster: evaluating one or more attributes of the collection of related data entities stored by the data entity cluster according to the second scoring strategy;and assigning a second score to the data entity cluster;re-ranking the plurality of data entity clusters according to the respective assigned second scores;and presenting, on an electronic display viewable by a user, a listing and/or summary of one or more of the plurality of data entity clusters according to the re-ranking.
  6. 9
    The method of any preceding claim, wherein executing the first of the one two or more search protocols comprises:identifying at least one data entity attribute (303-1) associated with the seed data entity;evaluating the plurality of data entities to determine the one or more data entities (305-1) sharing the at least one data entity attribute with the seed data entity.
  7. 10
    The method of any preceding claim, wherein executing each of the other search protocols of the two or more search protocols comprises:identifying the at least one data entity attribute (303-2, ... , 303-6) associated with the at least one of the one or more data entities previously added to the data entity cluster;and evaluating, based on performing the search protocol, the plurality of data entities to determine one or more additional data entities sharing the at least one data entity attribute with the at least one of the one or more data entities previously added to the data entity cluster.
  8. 11
    The method of any preceding claim, further comprising:comparing, by the one or more hardware computer processors, data entities associated with a first data entity cluster (252-1) to data entities associated with a second data entity cluster (252-C);and in response to determining that at least one data entity associated with the first data entity cluster shares an attribute with and/or is related to at least one data entity associated with the second data entity cluster, merging the first data entity cluster (252-1) and the second data entity cluster (252-C).
  9. 12
    The method of any preceding claim, wherein the first search protocol of the two or more search protocols searches for particular data entities in a first electronic data store, and the other search protocols of the two or more search protocols searches for particular data entities in a second electronic data store.
  10. 13
    A computer program comprising machine-readable instructions that, when executed by computer apparatus, cause the computer apparatus to perform the method of any preceding claim.
  11. 14
    Apparatus configured to perform the method of any of claims 1 to 12.