US10169610B2

Data privacy employing a k-anonymity model with probabalistic match self-scoring

Summary by NHIP

Data protection via k-anonymity

The system determines a desired duplication rate and generates a self-score threshold using Receiver Operating Characteristic curves. It produces self-scores by comparing quasi-identifiers, then modifies attributes of failing records to be less specific while enabling access to those satisfying the threshold.

Claim Score by NHIP

Read claim 1, the broadest

Abstract

According to one embodiment of the present invention, a system for protecting data determines a desired duplication rate based on a level of desired anonymity for the data and generates a threshold for data records within the data based on the desired duplication rate. The system produces a data record score for each data record based on comparisons of attributes for that data record, compares the data record scores to the threshold, and controls access to the data records based on the comparison. Embodiments of the present invention further include a method and computer program product for protecting data in substantially the same manners described above.

US10169610B2, drawing sheet 1
Sheet 1 of 3

Term

7.9 yearsleft in the term

Expires 4 August 2034.

  1. Priority
  2. Filed
  3. Granted
  4. Today
  5. Expires

14 claims: 3 independent, 11 dependent

  1. 1
    Broadest claimClaim Score 41, average(NHIP)A computer-implemented method of protecting data comprising:determining a desired duplication rate for a data set based on a level of desired anonymity for the data in the data set, wherein one or more attributes of data records within the data set that individually identify an identity for a corresponding data record are de-identified, and remaining attributes of the data records include quasi-identifiers;generating a self-score threshold for the data records within the data set based on the desired duplication rate and Receiver Operating Characteristic (ROC) curves;producing a data record self-score for each data record in the data set by comparing quasi-identifiers for that data record to quasi-identifiers of a corresponding original data record;comparing the data record self-scores to the self-score threshold;andcontrolling access to the data records of the data set based on the data record self-scores satisfying the self-score threshold, wherein controlling access further comprises: modifying attributes of data records failing to satisfy the self-score threshold to be less specific;andenabling access to data records in response to the data record self-scores satisfying the self-score threshold indicating a presence of the desired level of anonymity for the data records.
  2. 7
    A system for protecting data comprising:at least one processor configured to: determine a desired duplication rate for a data set based on a level of desired anonymity for the data in the data set, wherein one or more attributes of data records within the data set that individually identify an identity for a corresponding data record are de-identified, and remaining attributes of the data records include quasi-identifiers;generate a self-score threshold for the data records within the data set based on the desired duplication rate and Receiver Operating Characteristic (ROC) curves;produce a data record self-score for each data record in the data set by comparing quasi-identifiers for that data record to quasi-identifiers of a corresponding original data record;compare the data record self-scores to the self-score threshold;andcontrol access to the data records of the data set based on the data record self-scores satisfying the self-score threshold, wherein controlling access further comprises: modifying attributes of data records failing to satisfy the self-score threshold to be less specific;andenabling access to data records in response to the data record self-scores satisfying the self-score threshold indicating a presence of the desired level of anonymity for the data records.
  3. 11
    A computer program product for protecting data comprising:a computer readable storage medium having computer readable program code embodied therewith for execution on a processing system, the computer readable program code comprising computer readable program code configured to: determine a desired duplication rate for a data set based on a level of desired anonymity for the data in the data set, wherein one or more attributes of data records within the data set that individually identify an identity for a corresponding data record are de-identified, and remaining attributes of the data records include quasi-identifiers;generate a self-score threshold for the data records within the data set based on the desired duplication rate and Receiver Operating Characteristic (ROC) curves;produce a data record self-score for each data record in the data set by comparing quasi-identifiers for that data record to quasi-identifiers of a corresponding original data record;compare the data record self-scores to the self-score threshold;andcontrol access to the data records of the data set based on the data record self-scores satisfying the self-score threshold, wherein controlling access further comprises: modifying attributes of data records failing to satisfy the self-score threshold to be less specific;andenabling access to data records in response to the data record self-scores satisfying the self-score threshold indicating a presence of the desired level of anonymity for the data records.