EP4414902A2

Computer-based systems, computing components and computing objects configured to implement dynamic outlier bias reduction in machine learning models

Abstract

Systems and methods include processors for receiving training data for a user activity; receiving bias criteria; determining a set of model parameters for a machine learning model including: (1) applying the machine learning model to the training data; (2) generating model prediction errors; (3) generating a data selection vector to identify non-outlier target variables based on the model prediction errors; (4) utilizing the data selection vector to generate a non-outlier data set; (5) determining updated model parameters based on the non-outlier data set; and (6) repeating steps (1)-(5) until a censoring performance termination criterion is satisfied; training classifier model parameters for an outlier classifier machine learning model; applying the outlier classifier machine learning model to activity-related data to determine non-outlier activity-related data; and applying the machine learning model to the non-outlier activity-related data to predict future activity-related attributes for the user activity.

EP4414902A2, drawing sheet 1
Sheet 1 of 85

Term

14.5 yearsto projected expiry

Projected expiry 18 March 2041, counted from filing; an application has no term until it is granted.

  1. Priority
  2. Filed
  3. Published
  4. Today
  5. Projected expiry

15 claims: 2 independent, 13 dependent

  1. 1
    A method comprising:receiving, by at least one processor, a training data set of target variables representing at least one activity-related attribute for at least one user activity;receiving, by the at least one processor, at least one bias criteria used to determine one or more outliers;determining, by the at least one processor, a set of model parameters for a machine learning model comprising: (1) applying, by the at least one processor, the machine learning model having a set of initial model parameters to the training data set to determine a set of model predicted values;(2) generating, by the at least one processor, an error set of data element errors by comparing the set of model predicted values to corresponding actual values of the training data set;(3) generating, by the at least one processor, a data selection vector to identify non-outlier target variables based at least in part on the error set of data element errors and the at least one bias criteria;(4) utilizing, by the at least one processor, the data selection vector on the training data set to generate a non-outlier data set;(5) determining, by the at least one processor, a set of updated model parameters for the machine learning model based on the non-outlier data set;and (6) repeating, by the at least one processor, steps (1)-(5) as an iteration until at least one censoring performance termination criterion is satisfied so as to obtain the set of model parameters for the machine learning model as the updated model parameters, whereby each iteration re-generates the set of predicted values, the error set, the data selection vector, and the non-outlier data set using the set of updated model parameters as the set of initial model parameters;training, by the at least one processor, based at least in part on the training data set and the data selection vector, a set of classifier model parameters of an outlier classifier machine learning model to obtain a trained outlier classifier machine learning model that is configured to identify at least one outlier data element;applying, by the at least one processor, the trained outlier classifier machine learning model to a data set of activity-related data for the at least one user activity to determine: i) a set of outlier activity-related data in the data set of activity-related data, and ii) a set of non-outlier activity-related data in the data set of activity-related data;and applying, by the at least one processor, the machine learning model to the set of non-outlier activity-related data elements to predict future activity-related attribute related to the at least one user activity.
  2. 9
    A system comprising:at least one processor in communication with a non-transitory computer-readable storage medium having software instructions stored thereon, wherein the software instructions, when executed, cause the at least one processor to perform steps to: receive a training data set of target variables representing at least one activity-related attribute for at least one user activity;receive at least one bias criteria used to determine one or more outliers;determine a set of model parameters for a machine learning model comprising: (1) apply the machine learning model having a set of initial model parameters to the training data set to determine a set of model predicted values;(2) generate an error set of data element errors by comparing the set of model predicted values to corresponding actual values of the training data set;(3) generate a data selection vector to identify non-outlier target variables based at least in part on the error set of data element errors and the at least one bias criteria;(4) utilize the data selection vector on the training data set to generate a non-outlier data set;(5) determine a set of updated model parameters for the machine learning model based on the non-outlier data set;and (6) repeat steps (1)-(5) as an iteration until at least one censoring performance termination criterion is satisfied so as to obtain the set of model parameters for the machine learning model as the updated model parameters, whereby each iteration re-generates the set of predicted values, the error set, the data selection vector, and the non-outlier data set using the set of updated model parameters as the set of initial model parameters;train, based at least in part on the training data set and the data selection vector, a set of classifier model parameters of an outlier classifier machine learning model to obtain a trained outlier classifier machine learning model that is configured to identify at least one outlier data element;apply the trained outlier classifier machine learning model to a data set of activity-related data for the at least one user activity to determine: i) a set of outlier activity-related data in the data set of activity-related data, and ii) a set of non-outlier activity-related data in the data set of activity-related data;and apply the machine learning model to the set of non-outlier activity-related data elements to predict future activity-related attribute related to the at least one user activity.
Independent claims2