CA2842218C

Method and system for adaptive rule-based content scanners

Abstract

A method for scanning content, including identifying tokens within an incoming byte stream, the tokens being lexical constructs for a specific language, identifying patterns of tokens, generating a parse tree from the identified patterns of tokens, and identifying the presence of potential exploits within the parse tree, wherein said identifying tokens, identifying patterns of tokens, and identifying the presence of potential exploits are based upon a set of rules for the specific language. A system and a computer readable storage medium are also described and claimed.

CA2842218C, drawing sheet 1
Sheet 1 of 7

Term

Term ended

Expired 24 August 2025, 1.1 years ago.

  1. Priority
  2. Filed
  3. Granted
  4. Expired
  5. Today

17 claims: 3 independent, 14 dependent

  1. 1
    CLAIMS:1. A method for scanning content, comprising: expressing an exploit in terms of text strings and hexadecimal strings, and in terms of at least one Boolean expression using the strings, wherein the exploit is a portion of program code that is malicious;parsing a stream of incoming program code in a first script language, to determine if the exploit is potentially embedded therewithin in a second script language, based on said expressing, wherein the second script language is different than the first script language, said parsing including scanning the stream of incoming program code by a first scanner for scanning incoming program code in the first script language, the first scanner including a rule for invoking a second scanner for scanning incoming program code in the second script language;invoking by the first scanner, the second scanner in accordance with the rule when the first scanner encounters a script element of the exploit in the second script language embedded within the incoming program code in the first script language;wherein the first scanner accesses one or more first rule files including the rule for invoking a second scanner and first parser rules and first analyzer rules for the first script language, wherein the first parser rules define patterns in terms of tokens, tokens being lexical constructs for the first script language, and wherein the first analyzer rules identify certain combinations of tokens and patterns as being indicators of potential exploits;and further wherein the second scanner accesses one or more second rule files including second parser rules and second analyzer rules for the second script language, wherein the second parser rules define patterns in terms of tokens, tokens being lexical constructs for the second script language, and wherein the second analyzer rules identify certain combinations of tokens and patterns as being indicators of potential exploits.
  2. 9
    A system for scanning content, comprising a parser for parsing a stream of incoming program code in a first script language, to determine if an exploit is potentially embedded therewithin in a second script language, based on a formal description of the exploit expressed, in terms of text strings and hexadecimal strings, and in terms of at least one Boolean expression using the strings, wherein the exploit is a portion of program code that is malicious, and wherein the second script language is different than the first script language;CA 2842218 2018-06-26 the parser including a first scanner for scanning incoming program code in the first script language and a second scanner for scanning incoming program code in the second script language, the first scanner including a rule for invoking the second scanner when the first scanner encounters a script element of the exploit in the second script language embedded within the incoming program code in the first script language;wherein the first scanner accesses one or more first rule files including the rule for invoking a second scanner and first parser rules and first analyzer rules for the first script language, wherein the first parser rules define patterns in terms of tokens, tokens being lexical constructs for the first script language, and wherein the first analyzer rules identify certain combinations of tokens and patterns as being indicators of potential exploits;and further wherein the second scanner accesses one or more second rule files including second parser rules and second analyzer rules for the second script language, wherein the second parser rules define patterns in terms of tokens, tokens being lexical constructs for the second script language, and wherein the second analyzer rules identify certain combinations of tokens and patterns as being indicators of potential exploits.
  3. 17
    A computer-readable medium having recorded thereon computer-executable instructions that when executed by a computer perform the steps of:expressing an exploit in terms of text strings and hexadecimal strings, and in terms of at least one Boolean expression using the strings, wherein the exploit is a portion of program code that is malicious;parsing a stream of incoming program code in a first script language to determine if the exploit is potentially embedded therewith™ in a second script language, based on said expressing, wherein the second script language is different than the first script language, said parsing including scanning the stream of incoming program code by a first scanner for scanning incoming program code in the first script language, the first scanner including a rule for invoking a second scanner for scanning incoming program code in the second script language;CA 2842218 2018-06-26 invoking by the first scanner, the second scanner in accordance with the rule when the first scanner encounters a script element of the exploit in the second script language embedded within the incoming program code in the first script language;wherein the first scanner accesses one or more first rule files including the rule for invoking a second scanner and first parser rules and first analyzer rules for the first script language, wherein the first parser rules define patterns in terms of tokens, tokens being lexical constructs for the first script language, and wherein the first analyzer rules identify certain combinations of tokens and patterns as being indicators of potential exploits;and further wherein the second scanner accesses one or more second rule files including second parser rules and second analyzer rules for the second script language, wherein the second parser rules define patterns in terms of tokens, tokens being lexical constructs for the second script language, and wherein the second analyzer rules identify certain combinations of tokens and patterns as being indicators of potential exploits.