US11537748B2

Self-contained system for de-identifying unstructured data in healthcare records

Summary by NHIP

Healthcare Data De-identification

The method initializes computing devices with blacklist and whitelist dictionaries to identify and remove personally identifiable information from unstructured healthcare records. It modifies the blacklist by excluding whitelist terms, augments it with record-specific entries, and replaces removed data with case-type tags while repeating the process for each record.

Claim Score by NHIP

Read claim 1, the broadest

Abstract

A method and apparatus for identifying personally identifiable information (PII) and protected health information (PHI) within unstructured data, removing the PII and PHI from the unstructured data, and replacing the removed information with case-type tags that allows the user to understand what information was removed and to tune the level of information removal in future data sets.

US11537748B2, drawing sheet 1
Sheet 1 of 8

Term

15.1 yearsleft in the term

Expires 29 October 2041, including 1,010 days of term adjustment.

  1. Priority and filed
  2. Granted
  3. Today
  4. Expires

18 claims: 2 independent, 16 dependent

  1. 1
    Broadest claimClaim Score 33, narrow(NHIP)A method for de-identifying unstructured data within data sets, the method comprising:initializing, using a computing device comprising a processor and memory, a blacklist dictionary having a structure that facilitates identification of personally identifiable information (PII) and protected health information (PHI) and that facilitates removal and replacement of PII and PHI with a case-type tag, and a whitelist dictionary comprising terms selected to remain in the data sets despite being included in the blacklist dictionary;modifying, using the processor and a register module, the blacklist dictionary by removing terms included within the whitelist dictionary to create an adjusted blacklist dictionary;augmenting, using the processor, the adjusted blacklist dictionary with a record-specific blacklist for each individual record within the data sets to create a record-specific adjusted blacklist dictionary;scrubbing, using the processor and a de-identification engine, PII and PHI from each individual record utilizing the record-specific adjusted blacklist dictionary, the scrubbing comprising: removing all elements within each individual record determined to be PII or PHI according to terms in the record-specific adjusted blacklist dictionary;replacing removed elements with a case-type tag identifying a type of information being removed according to the record-specific adjusted blacklist dictionary;and repeating, using the de-identification engine, the scrubbing of PII and PHI comprising removing steps and replacing steps for each individual record within the data sets.
  2. 10
    A system for de-identifying unstructured data within data sets, comprising:memory and a processor configured for accessing the data sets from data sources and parsing data within the data sets and identifying elements of unstructured data within the data sets and de-identifying unstructured data using a de-identification engine initializing a blacklist dictionary having a structure that facilitates identification of personally identifiable information (PII) and protected health information (PHI) and that facilitates removal and replacement of PII and PHI with a case-type tag, and a whitelist dictionary comprising terms selected to remain in the data sets despite being included in the blacklist dictionary, and comprising: a register module configured to modify the blacklist dictionary by removing terms included within the whitelist dictionary to generate an adjusted blacklist dictionary, and augmenting the adjusted blacklist dictionary with a record-specific blacklist for each individual record within the data sets to create a record-specific adjusted blacklist dictionary;a de-identification module configured to determine which unstructured elements are to be removed from each individual record of a data set by scrubbing PII and PHI from each individual record utilizing the record-specific adjusted blacklist dictionary, the scrubbing comprising: removing all elements within each individual record determined to be PII or PHI according to terms in the record-specific adjusted blacklist dictionary;and replacing removed elements with a case-type tag identifying a type of information being removed according to the record-specific adjusted blacklist dictionary;wherein the de-identification module repeats the scrubbing PII and PHI for each individual record within the data sets;and one or more user devices configured to receive input from one or more users by an input-output interface and communicate with the processor and the de-identification engine over a telecommunication network, providing functionality for the register module, de-identification module, and merging module that share a secured network connection.