MY201302A

Data type recognition, model training and risk recognition methods, apparatuses and devices

Abstract

The present application provides data type recognition and model training methods and apparatuses, and computer devices. The model training method includes acquiring (102) a first sample data set, and using the first sample data set to train an anomaly detection model; and detecting (104) an abnormal sample data set from a second sample data set by means of the anomaly detection model, and using the abnormal sample data set to train a classification model. By means of this embodiment, an amount of scoring events of the classification model can be reduced, and relatively balanced sample data sets can also be provided for training, to obtain the classification model with a higher accuracy. In a particular application, data to be recognized is firstly input to the anomaly detection model, and whether the data to be recognized is first-type data can be quickly distinguished; and other data than the first-type data recognized by the anomaly detection model, is input to the classification model for recognition. The speed of online data recognition is relatively fast.(Fig. 2)

MY201302A, drawing sheet 1
Sheet 1 of 6

Term

No projected expiry on record.

  1. Priority
  2. Filed
  3. Published
  4. Today

12 claims: 6 independent, 6 dependent

  1. 1
    CLAIMS 1. A data type recognition method for recognizing data as first-type data or second-type data, wherein the method comprises:acquiring (202) data to be recognized, and using a preset anomaly detection model to detect whether the data to be recognized is first-type data;and inputting (204) other data than the first-type data recognized by the anomaly detection model, into a classification model for recognition, wherein the classification model classifies the other data as first-type data and second-type data, wherein the anomaly detection model is a first machine learning model and obtained by pretraining based on a first sample data set, and the classification model is a second machine learning model and obtained by pre-training based on a second sample data set different from the first sample data set;and the pre-training of the classification model comprises: detecting, by the anomaly detection model, an abnormal sample data set from the second sample data set;optimizing the abnormal sample data set based on a feature optimization algorithm;and using the optimized abnormal sample data set to train the classification model.
  2. 3
    A risk recognition method for recognizing data as secure data or risky data, wherein the method comprises:acquiring (302) data to be recognized, and using a preset anomaly detection model to detect whether the data to be recognized is abnormal;if the data to be recognized is detected not to be abnormal, determining (304) that the data to be recognized is secure data;and if the data to be recognized is detected to be abnormal, using (306) a preset classification model to recognize that the data to be recognized is secure data or risky data, wherein the anomaly detection model is a first machine learning model and obtained by pretraining based on a first sample data set, and the classification model is a second machine learning model and obtained by pre-training based on a second sample data set different from the first sample data set;and the pre-training of the classification model comprises: detecting, by the anomaly detection model, an abnormal sample data set from the second sample data set;optimizing the abnormal sample data set based on a feature optimization algorithm;and using the optimized abnormal sample data set to train the classification model.
  3. 5
    A computer device, comprising:a processor;and a memory for storing instructions executable by the processor, wherein the processor is configured to: acquire (202) data to be recognized, and use a preset anomaly detection model to detect whether the data to be recognized is first-type data;and input (204) other data than the first-type data recognized by the anomaly detection model, into a classification model for recognition, wherein the classification model classifies the other data as first-type data and second-type data, wherein the anomaly detection model is a first machine learning model and obtained by pretraining based on a first sample data set, and the classification model is a second machine learning model and obtained by pre-training based on a second sample data set different from the first sample data set;and the pre-training of the classification model comprises: detecting, by the anomaly detection model, an abnormal sample data set from the second sample data set;optimizing the abnormal sample data set based on a feature optimization algorithm;and using the optimized abnormal sample data set to train the classification model.
  4. 7
    A computer device, comprising:a processor;and a memory for storing instructions executable by the processor, wherein the processor is configured to: acquire (302) data to be recognized, and use a preset anomaly detection model to detect whether the data to be recognized is abnormal data;if the data to be recognized is detected not to be abnormal, determine (304) that the data to be recognized is secure data;and if the data to be recognized is detected to be abnormal, use (306) a preset classification model to recognize that the data to be recognized is secure data or risky data, wherein the anomaly detection model is a first machine learning model and obtained by pretraining based on a first sample data set, and the classification model is a second machine learning model and obtained by pre-training based on a second sample data set different from the first sample data set;and the pre-training of the classification model comprises: detecting, by the anomaly detection model, an abnormal sample data set from the second sample data set;optimizing the abnormal sample data set based on a feature optimization algorithm;and using the optimized abnormal sample data set to train the classification model.
  5. 9
    A non-transitory computer-readable storage medium having stored therein instructions that, when executed by a processor of a computer device, cause the computer device to perform a data type recognition method for recognizing data as first-type data or second-type data, wherein the method comprises:acquiring (202) data to be recognized, and using a preset anomaly detection model to detect whether the data to be recognized is first-type data;and inputting (204) other data than the first-type data recognized by the anomaly detection model, into a classification model for recognition, wherein the classification model classifies the other data as first-type data and second-type data, wherein the anomaly detection model is a first machine learning model and obtained by pretraining based on a first sample data set, and the classification model is a second machine learning model and obtained by pre-training based on a second sample data set different from the first sample data set;and the pre-training of the classification model comprises: detecting, by the anomaly detection model, an abnormal sample data set from the second sample data set;optimizing the abnormal sample data set based on a feature optimization algorithm;and using the optimized abnormal sample data set to train the classification model.
  6. 11
    A non-transitory computer-readable storage medium having stored therein instructions that, when executed by a processor of a computer device, cause the computer device to perform a risk recognition method for recognizing data as secure data or risky data, wherein the method comprises:acquiring (302) data to be recognized, and using a preset anomaly detection model to detect whether the data to be recognized is abnormal;if the data to be recognized is detected not to be abnormal, determining (302) that the data to be recognized is secure data;and if the data to be recognized is detected to be abnormal, using (304) a preset classification model to recognize that the data to be recognized is secure data or risky data, wherein the anomaly detection model is a first machine learning model and obtained by pretraining based on a first sample data set, and the classification model is a second machine learning model and obtained by pre-training based on a second sample data set different from the first sample data set;and the pre-training of the classification model comprises: detecting, by the anomaly detection model, an abnormal sample data set from the second sample data set;optimizing the abnormal sample data set based on a feature optimization algorithm;and using the optimized abnormal sample data set to train the classification model.