Detecting unwanted electronic mail messages based on probabilistic analysis of referenced resources
18 claims: 4 independent, 14 dependent
- 1ネットワークインタ-フェース、 該ネットワークインタ-フェースに接続された一つ以上の処理装置、および 該一つ以上の処理装置に接続され、該一つ以上の処理装置によって実行されたときに、該一つ以上の処理装置を、 それぞれが メッセージ と関連付けられた脅威を示す電子メッセージに対する1以上の基準 を規定す る複 数のルール であって、各 ルールが優先順位の値を持ち 、メ ッセージ要素型と関連付けられ、 かつルールで規定される基準の数に応じた重み値を有する複数のルール を受信しおよび記憶する手 段;受信者アカウントの宛先アドレスを有する、複数のメッセージ要素を含む電子メールメッセージを受信する手段;第一のメッセージ要素を抽出する手段;前記第一のメッセージ要素のみを 、前 記第一のメッセージ要素に対応するメッセージ要素型を有する 選択された ルールのみと、 該選択されたルールの優先順位の順序に従って 照合し、前記メッセージに対する脅威スコア値を、 前記選択されたルールのそれぞれを照合することにより得られたウイルススコア値の重み付け加算値として 決定する手段 であって、ウイルススコア値の重み付け加算値が、前記選択されたルールのそれぞれを照合することにより得られたウイルススコア値と照合されたルールの重み値との積の総和である第1の和を計算し、該第1の和を、前記照合に用いられたルールに関連する重み値の総和である第2の和で割ることで計算される手段と、 前記脅威スコア値が規定の閾値よりも大きい場合に、前記脅威スコア値を出力する手段;として機能させるロジックを備えたことを特徴とする装置。
- 2それぞれが メッセージ と関連付けられた脅威を示す電子メッセージに対する1以上の基準 を規定す る複 数のルール であって、各 ルールが優先順位の値を持ち 、メ ッセージ要素型と関連付けられ、 かつルールで規定される基準の数に応じた重み値を有する複数のルール を受信しおよび記憶する手 段;受信者アカウントの宛先アドレスを有する、複数のメッセージ要素を含む電子メールメッセージを受信する手段、 第一のメッセージ要素を抽出する手段、 前記第一のメッセージ要素のみを 、前 記第一のメッセージ要素に対応するメッセージ要素型を有する 選択された ルールのみと、 該選択されたルールの優先順位の順序に従って 照合し、前記メッセージに対する脅威スコア値を、 前記選択されたルールのそれぞれを照合することにより得られたウイルススコア値の重み付け加算値として 決定する手段 であって、ウイルススコア値の重み付け加算値が、前記選択されたルールのそれぞれを照合することにより得られたウイルススコア値と照合されたルールの重み値との積の総和である第1の和を計算し、該第1の和を、前記照合に用いられたルールに関連する重み値の総和である第2の和で割ることで計算される手段と、 前記脅威スコア値が規定の閾値よりも大きい場合に、前記脅威スコア値を出力する手段を備えたことを特徴とする装置。
- 3前記脅威スコア値が規定の閾値未満である場合にのみ、他のメッセージ要素を他のルールと照合することによって、前記メッセージに対する更新された脅威スコア値を決定する手段をさらに含むことを特徴とする請求項1または2に記載の装置。
- 4前記メッセージの本文を解読することによって前記メッセージに対する更新された脅威スコア値を決定する手段、および、前記脅威スコア値が規定の閾値未満である場合にのみ本文を一つ以上のメール本文ルールと照合する手段をさらに含むことを特徴とする請求項1または2に記載の装置。
- 5前記脅威スコア値が規定の脅威閾値と同じかそれを上回る場合に、前記メッセージを受信者アカウントに即座に配信せずに、前記メッセージを隔離キューに記憶する手段、 先入れ先出しの以外の順序で複数の隔離終了基準のうち任意の基準に基づいて前記隔離キューから前記メッセージを解放する手段であって、それぞれの隔離終了基準は一つ以上の終了アクションと関連付けられている手段、および 特定の終了基準に基づいて、前記関連付けられた一つ以上の終了アクションを選択しおよび実行する手段を更に備えたことを特徴とする請求項1または2に記載の装置。
- 6前記隔離終了基準は、メッセージ隔離時間制限の満了、前記隔離キューのオーバーフロー、前記隔離キューからの手動解放、および前記脅威スコア値を決定するために一つ以上のルールの更新を受信することを含むことを特徴とする請求項5に記載の装置。
- 7(a)前記メッセージ隔離から前記メッセージを手動解放するようにとのユーザー要求に対応して、前記メッセージを修正なしに前記受信者アカウントに配信する手段、および、 (b)隔離が満杯に成りつつあるというメッセージに対応して、前記メッセージから添付ファイルを取り除きおよび前記添付ファイルのないメッセージを前記受信者アカウントに配信する手段をさらに含むことを特徴とする請求項5に記載の装置。
- 8満了時間値をメッセージに割り当てる手段をさらに含み、 前記満了時間値は、経験則によるテストをメッセージコンテンツに適用する結果に基づいて異なることを特徴とする請求項5に記載の装置。
- 9前記隔離終了基準は、前記脅威スコア値を決定するために一つ以上のルールの更新を受信し、前記終了アクションは更新されたルールに基づいてメッセージに対する脅威スコア値を再び決定するものであることを特徴とする請求項5に記載の装置。
- 10前記一つ以上の異なる終了アクション群が、異なる隔離終了基準と関連付けられていることを特徴とする請求項5に記載の装置。
- 11コンピュータが、それぞれが メッセージ と関連付けられた脅威を示す電子メッセージに対する1以上の基準 を規定す る複 数のルール であって、各ルールが優先順位の値を持ち、メッセージ要素型と関連付けられ、かつルールで規定される基準の数に応じた重み値を有する複数のルール を受信しおよび記憶するステップ、 コンピュータが、 受信者アカウントの宛先アドレスを有する、複数のメッセージ要素を含む電子メールメッセージを受信するステップ、 コンピュータが、 第一のメッセージ要素を抽出するステップ、 コンピュータが、 前記第一のメッセージ要素のみを 、前 記第一のメッセージ要素に対応するメッセージ要素型を有する 選択された ルールのみと、 該選択されたルールの優先順位の順序に従って 照合し、前記メッセージに対する脅威スコア値を、 前記選択されたルールのそれぞれを照合することにより得られたウイルススコア値の重み付け加算値として 決定するステップ であって、ウイルススコア値の重み付け加算値が、前記選択されたルールのそれぞれを照合することにより得られたウイルススコア値と照合されたルールの重み値との積の総和である第1の和を計算し、該第1の和を、前記照合に用いられたルールに関連する重み値の総和である第2の和で割ることで計算されるステップ、 および、 コンピュータが、 前記脅威スコア値が規定の閾値よりも大きい場合に、前記脅威スコア値を出力するステップを有してなる方法。
- 12コンピュータが、 前記脅威スコア値が規定の閾値未満である場合にのみ、他のメッセージ要素を他のルールと照合することによって、前記メッセージに対する更新された脅威スコア値を決定するステップをさらに含むことを特徴とする請求項11に記載の方法。
- 13コンピュータが、 前記メッセージの本文を解読することによって前記メッセージに対する更新された脅威スコア値を決定するステップ、および、 コンピュータが、 前記脅威スコア値が規定の閾値未満である場合にのみ本文を一つ以上のメール本文ルールと照合するステップをさらに含むことを特徴とする請求項11に記載の方法。
- 14コンピュータが、 前記脅威スコア値が規定の脅威閾値と同じかそれを上回る場合に、前記メッセージを受信者アカウントに即座に配信せずに、前記メッセージを隔離キューに記憶するステップ、 コンピュータが、 先入れ先出しの以外の順序で複数の隔離終了基準のうち任意の基準に基づいて前記隔離キューから前記メッセージを解放するステップであって、それぞれの隔離終了基準は一つ以上の終了アクションと関連付けられているステップ、および コンピュータが、 特定の終了基準に基づいて、前記関連付けられた一つ以上の終了アクションを選択しおよび実行するステップを更に含むことを特徴とする請求項11に記載の方法。
- 15(a) コンピュータが、 前記メッセージ隔離から前記メッセージを手動解放するようにとのユーザー要求に対応して、前記メッセージを修正なしに前記受信者アカウントに配信するステップ、および(b) コンピュータが、 隔離が満杯に成りつつあるというメッセージに対応して、前記メッセージから添付ファイルを取り除きおよび前記添付ファイルのないメッセージを前記受信者アカウントに配信するステップをさらに含むことを特徴とする請求項14に記載の方法。
- 16コンピュータが、 満了時間値をメッセージに割り当てるステップをさらに含み、 前記満了時間値が、経験則によるテストをメッセージコンテンツに適用する結果に基づいて異なることを特徴とする請求項14に記載の方法。
- 17前記隔離終了基準は、前記脅威スコア値を決定するために一つ以上のルールの更新を受信し、前記終了アクションは更新されたルールに基づいて前記メッセージに対する脅威スコア値を再び決定するものであることを特徴とする請求項14に記載の方法。
- 18前記一つ以上の異なる終了アクション群が、異なる隔離終了基準と関連付けられていることを特徴とする請求項14に記載の方法。
Independent claims18
245 paragraphs, as filed
The present invention generally detects threats such as computer viruses, spam, and phishing attacks in electronic messages. More specifically, the present invention relates to a technique for responding to a new occurrence of a threat in an electronic message, managing a quarantine queue for the message holding the threat, and searching for the threat and scanning the message.
The approaches described in this section can be pursued, but not necessarily the ones that have been conceived or pursued. Therefore, unless otherwise indicated herein, the approach described in this chapter is not prior art to the claims in this application and is in addition to prior art by inclusion in this chapter. It will not be done.
Repeated outbreaks of message-borne viruses in computers connected to public networks have become a serious problem, especially for businesses with large private networks. Thousands of dollars in direct and indirect costs because wasted employee productivity, capital investment to buy additional hardware and software, and many viruses destroy files on shared directories It can result from lost information, and many viruses attached to files, and privacy and confidentiality breaches due to sending random files from the user's computer.
In addition, virus damage occurs in a very short time. A very high percentage of machines in the corporate network are infected between the time when a virus suddenly occurs and when the virus definition is published and deployed at a corporate email gateway that can detect and block virus-infected messages. there's a possibility that. The time frame between "sudden occurrence" and "rule deployment" is often 5 hours or more. Reducing reaction time would be extremely valuable.
In most virus outbreaks, executable attachments now act as carriers of virus code. For example, of the 17 major virus outbreaks in the last three years, 13 were sent via email attachments. Twelve of the 13 viruses sent via email attachments were sent via dangerous attachment types. As such, mail gateways on some corporate networks now block all types of executable attachments.
Obviously, current virus writers are hiding their executables. Increasingly, virus writers are hiding known dangerous file types in seemingly harmless files. For example, one virus writer may embed an executable file inside a .zip file of the type generated by WinZIP and other archiving utilities. Such .zip files are very commonly used by businesses to compress and share larger files, so are most businesses reluctant to block .zip files? Or they cannot be stopped. It is also possible to embed the executable in Microsoft Word® and several versions of Adobe Acrobat®.
<p> Based on the above, there is a clear need for an improved approach to managing the outbreak of the virus. Current technologies for preventing the delivery of messages containing large amounts of commercial junk mail (spam) and other forms of threats such as phishing attacks are also considered inadequate. Current techniques for searching for threats and scanning messages are also considered inadequate and need improvement.</p>
<p> Describes methods and devices for managing sudden outbreaks of computer viruses. In the following description, a number of specific details are provided for explanatory purposes to provide a complete understanding of the present invention. However, it will be apparent to those skilled in the art that the present invention can be carried out without these specific details. In another example, well-known structures and equipment are shown in the form of block diagrams to avoid unnecessarily obscuring the invention.</p><p> Embodiments are described herein according to the outline below: 1.0 Summary 2.0 Virus Sudden Outbreak Suppression Approach-First Embodiment-Structural and Functional Overview 2.1 Network system and virus information source 2.2 Counting suspicious messages 2.3 Message processing based on virus sudden occurrence information 2.4 Generation of virus sudden occurrence information 2.5 Use of virus sudden occurrence information 2.6 Additional features 2.7 Example use case 3.0 Approach to block spam messages 3.1 Early termination from spam scans 3.2 Spam scan judgment cache 4.0 How to detect viruses based on message heuristics, sender information, dynamic quarantine operations, and fine-grained rules 4.1 Detection using message heuristics 4.2 Virus detection for each sender 4.3 Dynamic quarantine operations including rescan 4.4 Fine-grained rules 4.5 Communication with the messaging gateway service provider 4.6 Caller White List Module 5.0 Implementation Mechanism-Hardware Overview 6.0 Extensions and alternatives 1.0 Summary The needs identified in the background above, and other needs and objectives that will become apparent in the description below, are achieved in the present invention and, in one embodiment, have the destination address of the recipient account. Receive e-mail messages; determine the virus score value for a message based on one or more rules that define the attributes of the message that are known to contain computer viruses, which attributes of the attachment to the message. Includes one or more empirical rules based on type, attachment size, and message sender, subject or body, and non-attachment signatures; if the virus score value is equal to or greater than the specified threshold. Includes methods that include storing the message in a quarantine queue without immediately delivering the message to the recipient account.</p><p> In another embodiment, the invention receives an email message with the destination address of the recipient account; determines the threat score value for the message; if the threat score value is equal to or greater than the defined threat threshold. To store the message in the quarantine queue without immediately delivering the message to the recipient account; release the message from the quarantine queue based on any of multiple quarantine termination criteria in an order other than first-in, first-out. Each quarantine termination criterion is associated with one or more termination actions; and provides methods that include selecting and performing one or more associated termination actions based on a particular termination criterion. ..</p><p> In another embodiment, the invention receives and stores a plurality of rules defining the characteristics of an electronic message indicating a threat associated with a message, where each rule has a priority value and each rule has a priority value. Associated with message element type; receive an email message with the destination address of the recipient account, the message contains multiple message elements; extract the first message element; only the first message element Determine the threat score value for a message by matching only selected rules that have the message element type corresponding to the first message element, and according to the order of priority of the selected rules; Provides a method including outputting a threat score value when is greater than a specified threshold.</p><p> In these approaches, early detection of computer viruses and other message-mediated threats is provided by applying heuristic tests to the message content and examining sender credit information when virus signature information is not available. As a result, the messaging gateway can stop delivering messages early in the event of a sudden virus outbreak, providing sufficient time to update the antivirus checker, which can remove the virus code from the message. Dynamic and flexible threat quarantine queues have various termination criteria and actions that allow early release of messages in non-first-in, first-out order. Allows early termination from parsing and scanning by matching threat rules only to selected message elements and stopping rule matching as soon as matching for one message element exceeds the threat threshold. The message scanning method to be performed will be explained.</p><p> In another embodiment, the invention includes a computer device configured to perform the steps described above and a computer-readable medium.</p>
The present invention is shown in the illustrations of the accompanying drawings as an example and not as a limitation, and similar reference numbers in the accompanying drawings indicate similar elements.
2.0 Virus Sudden Outbreak Suppression System-First Embodiment-Structural and Functional Overview 2.1 Network system and virus information source FIG. 1 is a block diagram of a system that manages the sudden outbreak of a computer virus according to one embodiment. A virus sender, whose identity and location are typically unknown, sends a virus-infected message, typically in the form of an electronic message, or email, to the Internet along with an attached executable file that carries the virus. To the public network 102 such as. The message is addressed to multiple destinations such as virus information source 104 and spam trap 106, or propagated by the action of a virus. Spam traps are email addresses or email mailboxes used to collect information about unsolicited email messages. The operation and implementation of virus information source 104 and spam trap 106 will be described in detail later. Figure 1 is for the purpose of showing a simple example. Only two destinations in the form of virus information source 104 and spam trap 106 are shown. However, in practice embodiments, such sources of virus information can be any number.
The virus sender 100 may obtain the network addresses of virus information sources 104 and spam trap 106 from public sources or by transmitting the virus to a small number of known addresses and propagating the virus.
The virus information processor 108 is communicably connected to the public network 102 and can receive information from the virus information source 104 and the spam trap 106. The virus information processing device 108 collects virus information from the virus information source 104 and the spam trap 106, generates virus sudden occurrence information, and stores the virus sudden occurrence information in the database 112. Implement certain functions detailed in the specification.
The messaging gateway 107 is connected directly from the public network 102 or directly through the firewall 111 or other network elements to the private network 110, which includes multiple termination stations 120A, 120B, 120C. The messaging gateway 107 may be integrated with a mail delivery agent 109 that processes email on the private network 110, or the mail delivery agent may be deployed separately. For example, IronPort Messaging Gateway Appliances (MGAs) such as Models C60, C30, or C10 marketed by Iron Port Systems Inc. (San Bruno, Calif.) Are for Mail Delivery Agent 109, Firewall 111, and Messaging Gateway 107. Implement the functions described herein.
In one embodiment, the messaging gateway 107 is a virus that obtains virus outbreak information from the virus information processor 108 and processes messages directed to terminal stations 120A, 120B, 120C according to the policy set by the messaging gateway. Includes information logic 114. As further described herein, virus outbreak information can include any of many types of information, which associates a virus score value, and a virus score value with a virus. Contains, but is not limited to, one or more rules associated with a given message characteristic. Such virus information logic may be integrated into the content filtering capabilities of Messaging Gateway 107, as further described herein with respect to FIG.
In one embodiment, the virus information logic 114 is implemented as an independent logic module within the messaging gateway 107. The messaging gateway 107 calls the virus information logic 114 by the message data and receives the corresponding determination. The judgment may be based on the message empirical rule. Message heuristics score messages and determine if a message is likely to be a virus.
The virus information logic 114 detects a virus based partially on the parameters of the message. In one embodiment, virus detection is performed based on one or more of the following: heuristics for emails containing executable code; heuristics for mismatched message headers; emails from known open relays. Rule of thumb; Rule of thumb for emails with mismatched content types and extensions; Rule of thumb for emails from dynamic user lists, blacklisted hosts, or senders known to be untrustworthy; and sending Reliability test results. Sender reliability test results may be generated by logic that receives sender ID values from public networks.
The messaging gateway 107 may also include an antivirus checker 116, a content filter 118, and an antispam logic 119. Antivirus checker 116 may include, for example, Sophos antivirus software. Content filter 118 provides logic that limits the delivery or reception of messages that contain content in the message subject or message body that is unacceptable by the policy associated with private network 110.
Anti-spam logic 119 scans incoming messages to determine, for example, whether the incoming message is commercial junk e-mail and according to the email receiving policy, and whether the incoming message is not desired, and anti-spam. Logic 119 applies policies that limit the delivery, destination change, or rejection of any unwanted message. In one embodiment, anti-spam logic 119 scans messages and returns a score for each message, indicating the probability that the message is spam or another type of unwanted email. Score ranges are associated with potential and likely spam thresholds that can be defined by administrators and for which users can apply the prescribed actions detailed below. In one embodiment, messages with a score of 90 or greater are spam, and messages with a score of 75 to 89 are suspected to be spanned.
In one embodiment, the anti-spam logic 119 is based in part on credit information obtained from database 112 or an external credit service such as Iron Port Systems Inc.'s Sender Base, where the sender of the message is spam, virus, or others. Determine the spam score that indicates that it is associated with the threat of. The scan may include recording an X header in the scanned message confirming that the message was successfully scanned, and include an obfuscated string that identifies the rule that matched the message. Obfuscation may include creating a hash of the rule identifier based on a private key and a one-way hash algorithm. Obfuscation ensures that only the specified parties, such as Service Provider 700 in Figure 7, break the matching rules, improving the security of the system.
The private network 110 may be a corporate network associated with a business, or any other form of network for which increased security or protection is desired. Public networks 102 and private networks 110 may use open standard protocols such as TCP / IP for communication.
The virus information source 104 includes an example of another messaging gateway 107 intervened between a public network 102 and another private network (not shown for clarity) for the purpose of protecting the other private network. It may be. In one embodiment, the virus information source 104 is Iron Port MGA. Spam Trap 106 is associated with one or more email addresses or email mailboxes associated with one or more domains. Spam trap 106 is established for the purpose of receiving unsolicited email messages or "spam" for analysis or reporting, and is typically not used in traditional email communications. For example, a spam trap can be an email address such as "dummyaccountforspam@mycompany.com", or a spam trap can be an email exchange (MX) Domain Name System (DNS) record that provides the email information received. Can be a group of email addresses grouped into. Mail delivery agent 109, or another Iron Port MGA's mail delivery agent may host spam trap 106.
In one embodiment, the virus information source 104 generates information used to manage the outbreak of a computer virus and provides it to the virus information processor 108, and the virus information processor 108 spams for the same purpose. Information from trap 106 can be obtained. For example, the virus information source 104 generates a count of messages with suspicious attachments received and provides the count to the virus information processor 108, or an external process reads the count and puts the count in a dedicated database. Allows you to remember. The messaging gateway 107 detects messages that are associated with or otherwise suspicious of a virus, creates a count of suspicious messages received within a specific time cycle, and counts the virus information processor. It also serves as a virus information source by providing it to 108 on a regular basis.
As a specific example, the functions described herein may be implemented as part of a comprehensive message data collection and reporting facility, such as Iron Port Systems Inc.'s SenderBase service. In this embodiment, the virus information processor 108 reads or receives information from the virus information source 104 and the spam trap 106, generates a count of messages with suspicious attachments or other virus indicators, and a database. Update 112 with the count and generate virus outbreak information for subsequent reading and use by the virus information logic 114 of the messaging gateway 107. Methods and devices related to the SenderBase service are described in pending application No. 10 / 857,641 of Robert Brahams et al., Filed May 28, 2004, entitled "Techniques for Determining the Credit of Message Senders". All of this content is incorporated herein by reference as if fully described herein.
Additionally or additionally, the virus information source 104 may include a SpamCop information service, or a user of the SpamCop service, accessible in the domain "spamCop.net" on the World Wide Web. The virus information source 104 may include one or more Internet service providers or other mass mail recipients.
The SenderBase and SpamCop services provide a powerful data source for detecting viruses. The service tracks information about millions of messages per day via spam trap addresses, end-user complaint reporters, DNS logs, and third-party data sources. This data can be used to quickly detect viruses using the approaches herein. Among other things, the number of messages that have a particular attachment type attached to a normal level, sent to a legitimate or spam trap address, and are not identified as a virus by an antivirus scanner is not yet known and antivirus. It provides an early warning indicator that a sudden outbreak of a virus has occurred, based on a new virus that can be detected by the scanner.
In another alternative embodiment, as a supplement to the automated approach herein, virus information source 104 manually reviews data obtained by an information service consultant or analyst, or an external source. It may be included. For example, administrators (humans) who monitor warnings from antivirus development vendors, third-party vendors, security mailing lists, spam trap data and other sources are almost always more likely than when virus definitions are published. Viruses can be detected well in advance.
Once a virus outbreak is identified based on virus outbreak information, network elements such as Messaging Gateway 107 can offer a variety of options for handling messages based on their probability of being a virus. Once the messaging gateway 107 is integrated with the mail delivery agent or mail gateway, the gateway can act on this data immediately. For example, the mail delivery agent 109 can delay the delivery of messages to the private network 110 until virus updates are received from the antivirus vendor and installed on the messaging gateway 107, resulting in delayed messages. After the virus update is received, it can be scanned by the antivirus checker 116.
The delayed message may be stored in quarantine queue 316. The messages in quarantine queue 316 may be released and delivered, deleted, or modified prior to delivery according to policies as detailed. In one embodiment, multiple quarantines 316 are established in the messaging gateway 107, and one quarantine is associated with each recipient account for computers 120A, 120B, etc. in the managed private network 110.
Although not shown in Figure 1, the virus information processor 108 can include, or can communicate with, a virus outbreak business center (VOOC), a received virus score (RVS) processor, or both. It is connected to the. Although it is separate from the VOOC and RVS processing device virus information processing device 108, it can be communicatively connected to the database 112 and the public network 102. The VOOC can be implemented as a staffing center with available personnel 24 hours a day, 7 days a week to monitor the information collected by the virus information processor 108 and stored in the database 112. .. Personnel assigned to VOCC issue virus outbreak warnings, update information stored in database 112, publish virus outbreak information so that messaging gateway 107 can access virus outbreak information, and Manual actions can be taken, such as manually initiating transmission of virus outbreak information to the messaging gateway 107 and other messaging gateways 107.
Additionally, personnel assigned to the VOCC may configure the mail delivery agent 109 to perform certain actions, such as delivering a "soft bounce." Soft bounce is performed when the mail delivery agent 109 returns a received message based on a set of rules accessible to the mail delivery agent 109. More specifically, when the mail delivery agent 109 completes an SMTP transaction by receiving an email message from the sender, the mail delivery agent 109 can access the mail delivery agent 109 as a set of stored software rules. Determines that the received message is undesired or undeliverable based on. In response to the determination that the received message is undesired or undeliverable, the mail delivery agent 109 returns the message to the bounce email address specified by the sender. When the mail delivery agent 109 returns the message to the sender, the mail delivery agent 109 may remove any attachments from the message.
In some implementations, virus outbreak information becomes available or made public in response to manual actions taken by personnel, such as personnel assigned to VOCC. In other practices, depending on the configuration of the virus information processor, VOOC, or RVS, virus outbreak information is automatically available, and then virus outbreak information and automatic actions taken are required or required. It will continue to be reviewed by personnel at VOOC who can make corrections when deemed desirable.
In one embodiment, the personnel assigned to the VOCC or the components of the system according to one embodiment have (a) patterns in receiving messages with attachments, (b) risks of attachments to received messages. Some characteristics, (c) published vendor virus alerts, (d) increased mailing list activity, (e) risky characteristics per source of messages, (f) dynamic networks associated with the sources of received messages. Does the message contain a virus based on various factors, such as the address percentage, (g) the percentage of computerized hosts associated with the source of the received message, and (h) the percentage of suspicious volume patterns? You may decide whether or not.
Each of the above factors may include various criteria. For example, the risky characteristics of an attachment to a received message are how suspicious the attachment's file name is, whether the file is associated with multiple file extensions, and similar attachments to the received message. It may be based on consideration of the amount of file size, the amount of similar file size attached to the received message, and the name of the attachment of a known virus. The patterns for receiving messages with attachments are the current percentage of messages with attachments, trends in the number of messages received with risky attachments, and messages with attachments. It may be based on considerations regarding the number of customer data sources, virus information sources 104, and spam traps 106 that are reporting an increase in.
In addition, the determination of whether a message contains a virus may be based on information sent by the customer. For example, information may be reported by the user to the system using an e-mail message received by the system in a secure environment. As a result, if the message receiver is infected with a virus, the system message receiver is configured to prevent the spread of the computer virus to other parts of the system as much as possible.
The RVS processor suddenly makes the messaging gateway 107 and other messaging gateways 107 available, for example in the form of virus score values for various attachment types or in the form of rules that associate virus score values with message characteristics. It can be implemented as an automatic system that generates generation information.
In one embodiment, the messaging gateway 107 provides a verdict cache 115 that provides local storage of verdicts from antivirus checker 116 and / or antispam logic 119 for reuse when a copied message is received. Including. The structure and function of the determination cache 115 will be further described below. In one embodiment, the messaging gateway 107 includes a log file 113 capable of storing statistical information or status messages about the functionality of the messaging gateway. Examples of information that can be recorded are message verdicts and actions taken as a result of verdicts; rules that match the message in obfuscated format; indication that a scan engine update has occurred; rule updates have occurred Indication; Includes scan engine version number, etc.
2.2 Counting suspicious messages FIG. 2 is a flow chart of a process for generating a count of suspicious messages according to one embodiment. In one implementation, the step of FIG. 2 may be performed by a virus information source such as the virus information source 104 of FIG.
At step 202, the message is received. For example, virus information source 104 or messaging gateway 107 receives a message sent by virus sender 100.
At step 204, a decision is made as to whether the message is at risk. In one embodiment, the virus checker scans the message at the virus information source 104 or messaging gateway 107 without identifying the virus, but the message is an attachment with a file type or extension known to be at risk. If it also contains, the message is determined to be at risk. For example, MS Windows® (XP Pro) file types or extensions COM, EXE, SCR, BAT, PIF, or ZIP allow virus writers to generally maliciously execute such files. Since it is used for various codes, it may be considered that there is a risk. The above are just examples of file types or extensions that can be considered risky. There are over 50 known different file types.
The decision that the message is suspicious extracts the source network address, such as the source IP value, from the message and queries the SenderBase service to determine if the source is known to be associated with spam or virus. May be done by issuing. For example, the credit score value provided by the SenderBase service may be taken into account when deciding whether a message is suspicious. The message may also be sent from an IP address associated with a host that is known to have been compromised, has a history of sending viruses, or has recently begun sending email to the Internet. , You may decide to be suspicious. The decision may be based on one or more of the following factors: (a) the type or extension of the attachment attached directly to the message, (b) the compressed file, archive, .zip file, or message. The type or extension of the file contained within other files attached directly to, and (c) the data fingerprint obtained from the attachment.
In addition, the determination of a suspicious message can be based on the size of the attachment to the suspicious message, the content of the subject of the suspicious message, the content of the body of the suspicious message, or any other characteristic of the suspicious message. You can embed other file types in some file types. For example, ".doc" and ".pdf" files may be embedded with other image file types such as ".gif" or ".bmp". Any file type embedded within the host file type may be considered when determining whether a message is suspicious. Suspicious message traits can be provided to or made available to messaging gateway 107 and used to develop rules containing virus score values associated with one or more such traits. it can.
In step 206, if the message is suspicious, then the number of suspicious messages for the current time period is incremented. For example, if the message has an EXE attachment, the count of messages with the EXE attachment is incremented.
At step 208, a count of suspicious messages is reported. For example, step 208 may involve sending a reporting message to the virus information processor 108.
In one embodiment, the virus information processor 108 continuously receives a large number of reports, such as the report of step 208, in real time. As the report is received, the virus information processor 108 updates the update database 112 with the report data, and determines and stores the virus sudden outbreak information. In one embodiment, the virus outbreak information includes a virus score value determined according to a dependent process further described below with reference to FIG.
2.3 Message processing based on virus sudden occurrence information FIG. 3 is a data flow diagram showing message processing based on virus sudden occurrence information according to one embodiment. In one implementation, the step of FIG. 3 may be performed by an MGA such as the messaging gateway 107 of FIG. Advantageously, by performing the steps shown in FIG. 3, the message may act before a positive decision is made that the message contains a virus.
At block 302, a content filter is applied to the message. Applying a content filter, in one embodiment, examines the message subject, the header values of other messages, and the message body to determine if one or more rules for content filtering are met by the content values. However, if the rules are met, it involves taking one or more actions as specified in the content policy. Execution of block 302 is optional. In this way, some embodiments may execute block 302, while other embodiments may not execute block 302.
Further, in block 302, virus outbreak information is read for use in subsequent processing steps. In one embodiment, in block 302, the messaging gateway 107 that implements FIG. 3 can periodically request the virus information processing apparatus 108 for the latest virus outbreak information at that time. In one embodiment, the messaging gateway 107 uses a secure communication protocol that prevents unauthorized persons from accessing the virus outbreak information about every 5 minutes from the virus outbreak information device 108. Read to. If the messaging gateway 107 is unable to read the virus outbreak information, the gateway can use the latest available virus outbreak information stored in the gateway.
At block 304, anti-spam processing is applied to the message, and the seemingly unsolicited message is marked or processed according to the spam policy. For example, spam messages may be quietly truncated and moved to a default mailbox or folder, or the subject line of the message may be modified to include a statement such as "potentially spam." Execution of block 304 is optional. In this way, some embodiments may execute block 304, while other embodiments may not execute block 304.
At block 306, antivirus processing is applied to the message and the message or attachment that appears to contain a virus is marked. In one embodiment, antivirus software from Sophos implements block 306. If the message is determined to be virus positive, then in block 308, the message is deleted, quarantined in quarantine queue 316, or otherwise processed according to the appropriate virus handling policy.
Alternatively, if block 306 determines that the message is not virus positive, then in block 310, a test is performed to determine if the message has been previously scanned for viruses. As further described herein, block 306 can be reached again from subsequent blocks after the message has been scanned for viruses in the past.
If the message has previously been scanned for viruses in block 306, then the process in Figure 3 is all the patterns needed to successfully identify the virus if a sudden outbreak of the virus is identified. Assume that antivirus processing 306 has been updated with, rules, or other information. Therefore, control proceeds to block 314, where previously scanned messages are delivered. If at block 310 it is determined that the message has never been scanned before, processing continues to block 312.
In block 312, a test is performed to determine whether the virus sudden outbreak information acquired in block 302 meets the specified threshold. For example, when the virus sudden occurrence information includes a virus score value (VSV), the virus score value is checked whether the virus score value is equal to or higher than the virus score threshold value.
Thresholds are specified by administrator commands in the configuration file or are received from other machines, processes or sources in a separate process. In one practice, the threshold corresponds to the probability that the message contains a virus or is associated with a sudden outbreak of a new virus. Virus that receives a score above a threshold, such executes quarantine messages in the quarantine queue 316, operating as a target of action defined by the coater. In some implementations, a single defined value is used for all messages, while in other implementations, multiple thresholds are used based on different characteristics, so that the administrator can use the messaging gateway. It can be treated more carefully than some messages, based on the type of message it receives and what it considers normal or less risky to the associated message recipient. In one embodiment, a default threshold of 3 is used based on a virus score scale from 0 to 5, where 5 is the highest (threat) risk level.
For example, virus outbreak information can include virus score values, and the network administrator determines the allowed virus score threshold, and performs all message delivery agents or the processing in Figure 3 with the virus score threshold. It can be transmitted to other processing devices. As another example, virus outbreak information can include a set of rules associated with a virus score value indicating a message characteristic indicating one or more viruses, and to the approach described herein with respect to FIG. Based on that, the virus score value can be determined based on the matching rules for the message.
The virus score threshold value set by the administrator suggests when to start delayed delivery of the message. For example, if the virus score threshold is 1, then the messaging gateway performing FIG. 3 delays the delivery of the message when the virus score value determined by the virus information processor 108 is low. If the virus score threshold is 4, then the messaging gateway performing FIG. 3 delays the delivery of the message when the virus score value determined by the virus information processor 108 is high.
If the specified threshold score value is not exceeded, then in block 314, a message is delivered.
If the message is determined to be above the virus score threshold in block 312, and the message has never been previously scanned as determined in block 310, then the message is placed in the sudden outbreak quarantine queue 316. Each message is tagged with a default hold time value, or expiration date-time value, which represents how long the message is held in the sudden occurrence quarantine queue 316. The purpose of the Sudden Quarantine Queue 316 is to deliver the message only enough time to allow the antivirus processing 306 to be updated to account for the new virus associated with the Sudden Outbreak of the detected virus. To delay.
The retention time may have any desired period. Examples of retention time values can be between 1 and 24 hours. In one embodiment, a default retention time value of 12 hours is provided. The administrator may change the retention time between any suitable retention time values at any time by issuing a command to the messaging gateway performing the processing herein. In this way, the retention time value is user configurable.
It provides one or more tools, characteristics, or user interfaces that allow an operator to monitor the status of sudden quarantine queues and quarantined messages. For example, an operator can get a list of currently quarantined messages, and the list is one of the applicable virus score values for messages that meet a given threshold, or a set of rules that match a message. You can identify why each message in the queue was quarantined, such as one or more rules. Summary information can be provided by message characteristics such as attachment type, or by rules applicable when a set of rules is used. Tools can be provided to allow the operator to review each individual message in the queue. Another characteristic can be provided to allow an operator to search for quarantined messages that meet one or more criteria. To ensure that the messaging gateway was configured correctly and that incoming messages were properly processed by the virus outbreak filter, we provided yet another tool called "tracking" to process the messages. You can simulate the message inside.
In addition, virus information processors, tools from VOOC, RVS can be provided that provide general warning information about special or serious virus risks or threats identified so far. The MGA can also include tools for contacting one or more personnel associated with the MGA when a warning is issued. For example, if a message is quarantined, a certain number of messages are quarantined, or the quarantine queue is full or reaches a specified level, an automated telephone or calling system will contact the specified individual. Can be done.
The message may terminate the abrupt quarantine queue 316 on the three circuits indicated by the designated paths 316A, 316B, 316C in FIG. The message may expire normally when the prescribed retention time for the message has expired, as shown in Route 316A. As a result, upon normal expiration, in one implementation, the sudden occurrence quarantine queue 316 operates as a FIFO (first in, first out) queue. The message is then rescanned, assuming that after the retention time has expired, the antivirus processing has been updated with any pattern file or other information needed to detect any viruses that may be present in the message. Transferred back to antivirus processing 306.
The message may be manually released from the abrupt quarantine queue 316, as indicated by route 316B. For example, one or more messages can be released from the abrupt quarantine queue 316 in response to a command issued by an administrator, operator, or other machine or process. Based on manual release, in block 318, the operator rescans or deletes a message, for example, if the operator may have received offline information indicating that a particular type of message is absolutely virus-infected. Decision is executed. In that case, the operator could have chosen to delete the message in block 320. Alternatively, the operator may have received offline information indicating that antivirus processing 306 was updated with a new pattern or other information in response to a sudden outbreak of the virus prior to the expiration of the retention time value. is there. In that case, by sending the message back to antivirus processing 306 for scanning, as indicated by route 319, the operator may choose to rescan the message without waiting for the retention time to expire. Good.
As yet another example, the operator can identify one or more messages by performing a search for messages currently held in the sudden occurrence quarantine queue 316. For example, messages identified in this way are scanned by antivirus 306 to test whether antivirus 306 has been updated with sufficient information to detect the virus involved in the outbreak of the virus. Can be selected by the operator for. If the rescan of the selected message successfully identifies the virus, the operator can manually release some or all of the messages in the sudden outbreak quarantine queue, resulting in antivirus released messages. It can be rescanned by virus processing 306. However, if a virus is detected in the test message selected by antivirus processing, then the operator waits for a later time and decides whether antivirus processing 306 has been updated to detect the virus. You can test a test message or another message to do so, or the operator can wait until the expiration time of the message expires and then release the message.
The message may expire early, for example, because the sudden outbreak quarantine queue 316 is full, as shown in route 316C. Overflow policy 322 applies to messages that expire early. For example, overflow policy 322 may require the message to be deleted, as shown in block 320. As another example, overflow policy 322 may require that the subject of the message be accompanied by an appropriate warning of the risk that the message is likely to contain a virus, as shown in block 324. For example, a message such as "may be infected" or "suspicious virus" can be added to the subject, such as at the end or beginning of the subject line of the message. Since the message with the subject was delivered via antivirus processing 306, and the message has been scanned before, processing continues from antivirus processing 306 via block 310, and the message. Is then delivered as shown in block 314.
Although not shown in Figure 3 for clarity, additional overflow policies can be applied. For example, overflow policy 322 may request removal of attachments attached to a message, followed by delivery of the message with attachments removed. Optionally, overflow policy 322 may require that attachments larger than a certain size be removed. As another example, overflow policy 322 allows the MTA to receive new messages when the sudden quarantine queue 316 is full, but before the message is received during an SMTP transaction. May be required to be rejected with a 4xx temporary error.
In one embodiment, the handling of messages according to routes 316A, 316B, 316C can be configured by the user for the entire content of the quarantine queue. Alternatively, such a policy can be configured by the user for each message.
In one embodiment, block 312 provides a warning message when the virus sudden occurrence information acquired from the virus information processing apparatus 108 meets or exceeds a specified virus score threshold, for example, when the virus score value meets or exceeds a specified virus score threshold. May involve generating and transmitting to one or more administrators. For example, a warning message sent in block 312 might have a virus score changed attachment type, current virus score, past virus score, current virus score threshold, and when virus for that type of attachment. It may include an email specifying whether the latest update of the score was received from the virus information processor 108.
In yet another embodiment, the processing of FIG. 3 is performed whenever the total number of messages in the quarantine queue exceeds a threshold set by the administrator, or when a certain amount or ratio of quarantine queue storage capacity is exceeded. May involve generating and sending a warning message to one or more administrators. Such warning messages may specify the size of the quarantine queue, the percentage of capacity used, and the like.
Sudden quarantine queue 316 may have any desired size. In one embodiment, the quarantine queue can store about 3GB of messages.
2.4 Generation of virus sudden occurrence information In one embodiment, virus outbreak information indicating the possibility of outbreak of a virus is generated based on one or more message characteristics. In one embodiment, the virus sudden outbreak information includes a numerical value such as a virus score value. Virus outbreak information includes the type of attachment to the message, the size of the attachment, the content of the message (for example, the subject line of the message or the content of the body of the message), the sender of the message, the IP address of the sender of the message, or It can be associated with one or more message characteristics, such as domain, message recipient, message sender's SenderBase credit score, or any other suitable message characteristic. As a specific example, the virus sudden occurrence information has one message characteristic, such as "EXE = 4" indicating that the virus score value is "4" for a message to which an attachment file of the type EXE is attached. Can be associated with.
In another embodiment, the virus outbreak information includes one or more rules, each of which associates a virus outbreak potential with one or more message characteristics. As a specific example, a rule of the form "if EXE and size <50, then 4" has a virus score value of "4" for a message with an attachment of type EXE and size less than 50k. Show that. A set of rules that can be applied to the messaging gateway to determine if an incoming message matches the message characteristics of the rule can be provided to the messaging gateway so that the rule is applicable to the incoming message and is therefore associated. Indicates that it should be treated based on the virus score value. The use of rule groups will be further described below with reference to FIG.
FIG. 4 is a flow chart of a method for determining a virus score value according to one embodiment. In one implementation, the step of FIG. 4 may be performed by the virus information processor 108 based on the information in the database 112 received from the virus information source 104 and the spam trap 106.
In step 401 of FIG. 4, steps 402 and 404 by a predetermined computer are executed for each source of different virus information such as virus information source 104 or spam trap 106 that can access the virus information processing device 108. Is shown.
Step 402 is a specific email attachment by combining one or more past virus score values for past time with a weighting approach that matches a larger weighting for more recent past virus score values. Accompanied by generating a weighted current average virus score value for the type of. A virus score for a particular time cycle is a score based on the number of messages with suspicious attachments received by a particular source. If the attachment meets one or more metrics, such as a particular file size, file type, etc., or if the sender's network address is known to be associated with a past virus outbreak. The message is believed to have a suspicious attachment. The decision may be based on the file size or file type or extension of the attachment.
The virus score value determination is for the SenderBase service to extract the source network address, such as the source IP address value, from the message and to determine if the source is known to be associated with spam or virus. It may also be done by issuing a query. The decision is (a) the type or extension of the attachment attached directly to the message, (b) the compressed file contained within the file, archive, .zip file, or other file attached directly to the message. It may be based on the type or extension of, and (c) the data fingerprint obtained from the attachment. Separate virus score values may be generated and stored for each attachment type found in any of the above. In addition, virus score values may be generated and stored based on the type of attachment with the highest risk found in the message.
In one embodiment, step 402 involves calculating a combination of virus score values for the last three 15-minute time widths for a given attachment type. Further, in one embodiment, the weighted values are applied to three values for a 15 minute time width, with the most recent 15 minute time frame being weighted more heavily than the earlier 15 minute time width. .. For example, in one weighting approach, a multiplier of 0.10 is applied to the virus score value for the oldest 15 minute time width (30 to 45 minutes ago) and a multiplier of 0.25 is applied to the second oldest value (15 to 30 minutes ago). And a multiplier of 0.65 is applied to the most recent virus score for the period 0-15 minutes ago.
In step 404, comparing the current average virus score value determined in step 402 with the long-term average virus score value produces a normal percentage virus score value for a particular attachment type. The current percentage of normal levels may be calculated by referring to the 30-day mean for the attachment type for all 15-minute time cycles of the 30-day period.
In step 405, all averages of the normal percentage virus score values for all sources, such as virus information source 104 and spam trap 106, are calculated, resulting in the creation of the overall ratio of normal values to a particular attachment type.
In step 406, the overall normal percentage value is mapped to the virus score value for a particular attachment type. In one embodiment, the virus score value is an integer from 0 to 5, and the normal percentage value is mapped to the virus score value. Table 1 shows an example of a virus score scale.<tables num="1"><img file="JP5118020B2_D0001.tif" /></tables>
In other embodiments, mappings to a range of score values 0 to 100, 0 to 10, 1 to 5, or any other desired value may be used. Non-integer values can be used in addition to integer score values. Instead of using the defined range of values, the probability value can be determined. For example, a higher probability is a probability in the range 0% to 100% that indicates a higher probability of a sudden outbreak of the virus, or a probability from 0 to 1 expressed as a fraction or a decimal such as 0.543.
As an optimization, and to avoid the zero-issue division that can occur with a very low 30-day count, the process of FIG. 4 can add 1 to the baseline mean calculated in step 402. In essence, adding 1 attenuates some of the data, thereby slightly increasing the noise level of the value in a preferred way.
Table 2 shows an example of data for the file type EXE in a hypothetical embodiment.<tables num="2"><img file="JP5118020B2_D0002.tif" /></tables>
In an alternative embodiment, the processing of FIGS. 2, 3 and 4 may also include logic that recognizes trends in reported data and identifies anomalies in virus scoring operations.
Most executables propagate through one or another type of email attachment, so the strategy of the approach herein is to make policy decisions based on the attachment type. focus on. In an alternative embodiment, the virus score value can be developed by considering other message data and metadata such as URL, attachment name, source network address, etc. in the message. Further, in an alternative embodiment, the virus score value may be assigned to individual messages rather than attachment type.
In yet another embodiment, another metric may be considered to determine the virus score value. For example, a virus may be identified when a large number of messages are suddenly received by the virus information processor 108 or its information source from a new host that has never sent a message. In this way, the fact that a particular message was first seen on a recent date, and the sharp rise in the amount of messages detected by the virus information processor 108, suggest an early outbreak of the virus. May be good.
2.5 Use of virus sudden occurrence information As mentioned above, virus outbreak information can simply associate a virus score value with a message characteristic such as attachment type, or virus outbreak information can each have a virus score value in a message suggesting a virus. It can contain a set of rules associated with one or more characteristics. The MGA can apply a set of rules to an incoming message to determine which rule matches the message. The MGA determines the possibility that a message contains a virus, based on a rule that matches an incoming message, for example, by determining a virus score value based on one or more virus score values derived from the matching rule. can do.
For example, the rule can be "4 for'exe'", which indicates that the message with the EXE attachment has a virus score of 4. As another example, the rule is "3 if'exe'and size <50k" to indicate that a message with an EXE attachment attached and less than 50k in size has a virus score of 3. There can be. As yet another example, the rule indicates that the virus score is 4 if the SenderBase Credit Score (SBRS) is less than "-5". It can be "4 if SBRS <-5". As another example, the rule is'PIF'and the subject is'If the message has the attachment type PIF and the subject of the message contains the string "FOOL" the virus score is 5. If it contains FOOL', it may be 5. " In general, a rule associates any number of message characteristics or other data that can be used to determine virus outbreak with an indicator that a message that matches the message characteristics or other data may contain the virus. Can be done.
In addition, the messaging gateway applies exceptions, such as in the form of one or more quarantine policies, based on virus score values determined based on matching rules, as determined in block 312 of Figure 3. Determining whether a message that meets the specified threshold in another way should be placed in the abrupt quarantine queue, or whether the message should be processed without being placed in the abrupt quarantine queue. Can be done. MGA allows messages to be delivered to email addresses or email addresses at all times regardless of virus score, or attaches messages of the specified attachment type, such as ZIP files containing PDF files. It can be configured to apply one or more policies to apply rules such as the always-delivered policy.
In general, by having a virus information processor supply rules instead of virus score values, each MGA can apply some or all of the rules in a manner determined by the MGA administrator. It provides additional flexibility to meet the needs of a particular MGA. As a result, the ability of two messaging gateways 107 to configure rule enforcement by their respective MGA administrators is determined by each MGA processing the same message, even if they use the same set of rules. Different results can be obtained in terms of possible viral attacks, and each MGA processes the same message and takes different actions depending on the configuration established by the MGA administrator. It means that you can do it.
FIG. 5 is a flow chart showing the application of a group of rules for managing the sudden outbreak of a virus according to one embodiment. The function shown in Figure 5 can be performed by the messaging gateway as part of block 312 or at any other suitable location during the processing of incoming messages.
At block 502, the messaging gateway identifies the message characteristics of the incoming message. For example, the messaging gateway 107 can determine whether a message has an attachment, and if so, the type of attachment, the size of the attachment, and the name of the attachment. As another example, the messaging gateway 107 can query the SenderBase service based on the sender's IP address in order to obtain a SenderBase credit score. For the purposes of illustrating Figure 5, assume that the message has an attachment type EXE attached, is 35k in size, and the sending post of the message has a SenderBase credit score of -2.
At block 504, the messaging gateway determines which of the rules it matches based on the message characteristics of the message. For example, for the purposes of illustrating Figure 5, assume that the rules group consists of the following five rules that associate example characteristics with the provided hypothetical virus score values: Rule 1: "3 for EXE" Rule 2: "4 for ZIP" Rule 3: "5 for EXE and size> 50k" Rule 4: "4 for EXE and sizes <50k and> 20k" Rule 5: "4 if SBRS <-5" In these rule examples, Rule 1 indicates that ZIP attachments are more likely to contain viruses than EXE attachments. The reason is that the virus score is 4 in Rule 2, but only 3 in Rule 1. Furthermore, in the example rule above, EXE attachments larger than 50k are most likely to carry the virus, but EXE attachments smaller than 50k but larger than 20k are likely to contain the virus. Indicates that it is slightly lower. The reason is probably that most suspicious messages with EXE attachments are larger than 50k in size.
In this example, where the message has an attachment of type EXE and is 35k in size, and the associated SenderBase has a credit score of -2, rules 1 and 4 match, while Rules 2, 3, and 5 do not match.
At block 506, the messaging gateway determines the virus score value used for the message based on the virus score value derived from the matching rule. The determination of the virus score value used for a message can be performed based on one of a number of approaches. The specific approach used is specified by the messaging gateway administrator and can be modified as desired.
For example, the first matching rule can be used when applying a list of rules in a listed order, and any other matching rules are ignored. Therefore, in this example, the first matching rule is rule 1, and therefore the virus score value for the message is 3.
As another example, the matching rule with the highest virus score is used. Therefore, in this example, rule 3 has the highest virus score value among the matching rules, and therefore the virus score value for the message is 5.
As yet another example, a matching rule with the most specific message characteristic group is used. Therefore, in this example, Rule 4 is the most specific matching rule because Rule 4 contains three different criteria, and therefore the virus score value for the message is 4.
As another example, virus score values from matching rules can be combined to determine the virus score value applied to a message. As a specific example, the virus score values derived from rules 1, 3, and 4 can be averaged to determine the virus score value 4 (eg, (3 + 4 + 5) ÷ 3 = 4). As another example, a weighted average of the virus score values of matching rules can be used to give greater weight to more specific rules. As a specific example, the weight for each virus score value can be the same as the number of criteria in the rule (for example, rule 1 with one criterion has a weight of 1 while three criteria. There is a rule 4 that has a weight of 3), and therefore the weighted average of rules 1, 3, and 4 results in a virus score of 4.2 (for example, (1x3 + 2x5 + 3x4). ) ÷ (1 + 2 + 3) = 4.2).
At block 508, the messaging gateway uses the virus score value determined at block 506 to determine if a defined virus score threshold is met. For example, in this example, the threshold is assumed to have a virus score value of 4. As a result, the virus score values determined in block 506 by all exemplary approaches are thresholded except when using the first matching rule and when block 506 determines that the virus score value is 3. Will meet.
If the specified threshold is determined to be met by the virus score value determined in block 508, then in block 510 one or more quarantine policies are applied to add the message to the sudden outbreak quarantine queue. To judge. For example, a messaging gateway administrator may decide that one or more users or a group of one or more users should never quarantine a message, even if a sudden outbreak of a virus is detected. .. As another example, an administrator can use a message with certain characteristics (for example, a message with an XLS attachment and a size of at least 75k) if the virus outbreak information indicates a virus attack based on a specified threshold. However, instead of being quarantined, a policy can be established to ensure that it is always delivered.
As a specific example, the legal department personnel of an organization should not be delayed by being placed in a sudden outbreak quarantine, even if the messaging gateway determines that a virus outbreak is occurring. You may receive ZIP files containing legal documents frequently. Therefore, the messaging gateway email administrator has a policy of always delivering messages with ZIP attachments to the Legal Department, even if the virus score value for the ZIP attachment meets or exceeds a specified threshold. Can be formulated.
As another embodiment, the email administrator may want to keep messages addressed to the email administrator's email address delivered. The reason is that such messages can provide information to deal with the outbreak of the virus. If the email administrator is a well-educated user, the risk of delivering virus-infected messages is low. The reason is that email administrators are likely to be able to identify and deal with infected messages before the virus can act.
In contrast to the example used to illustrate Figure 5, EXE attachments addressed to a senior technical manager at a company may have virus score values for such messages that meet or exceed the virus score threshold. Suppose the email administrator has established a policy of always delivering, if any. Therefore, if the message is addressed to any senior tech manager, the message will nevertheless be delivered instead of being put into a sudden outbreak quarantine. However, messages addressed to non-senior technical managers are quarantined (unless otherwise excluded by other applicable policies).
In one embodiment, the messaging gateway can be configured in one of two states, "calm" and "tensioned". If no message is quarantined, a calm state applies. However, if the virus outbreak information is updated and indicates that the specified threshold has been exceeded, the state is from "calm" to "tense" regardless of whether any message being received by the messaging gateway is being quarantined. It changes to "done". The tense state persists until the virus outbreak information is updated and indicates that it no longer exceeds the prescribed threshold.
In some practices, a warning message is sent to the operator or administrator whenever a change in system state (eg, "calm" to "tensioned" or "tensioned" to "peaceful") occurs. In addition, if a low virus score that once did not meet the threshold now meets or exceeds the threshold, the overall state of the system does not change (for example, the system was "calm" to "tensed" in the past. A warning can be issued even if another virus score that meets or exceeds the threshold is received from the virus information processor) while changing to, and on the other hand, in a "tensioned" state. Similarly, a warning can be issued if a high virus score that has met the threshold in the past has declined and is now below the specified threshold.
Warning messages can include one or more types of information, including but not limited to: virus outbreak information changed attachment type, current virus score, past virus score, current threshold, and. When did the latest virus outbreak information update occur?
2.6 Additional features In addition to the features described above, one or more of the following additional features can be used in a particular practice.
One additional feature is the acquisition of per-sender data specifically designed to assist in identifying viral threats. For example, if an MGA asks a service such as SenderBase to get a SenderBase credit score to connect an IP address, it can provide specific virus threat data to connect the SenderBase IP address. .. Virus threat data is based on the data collected by SenderBase for IP addresses and how often viruses are detected in IP addresses or messages originating from companies associated with IP addresses. In, and reflects the history of IP addresses. This allows the MGA to obtain a virus score from SenderBase based solely on the sender of the message, without any information or knowledge from the sender's IP address about the content of the particular message. Data on virus threats by sender can be used in place of or in addition to the virus scores determined above, or data on virus threats by sender can be used in the calculation of virus scores. Can be incorporated. For example, MGA can increase or decrease a particular virus score value based on virus threat data by sender.
Another feature is the use of dynamic or dial-up blacklists to identify messages that are likely to be infected with a virus when the dynamic or dial-up host is directly connected to an external SMTAP server. is there. Dynamic and dial-up hosts that connect to the Internet are typically expected to send outbound messages through the host's local SMTAP server. However, if the host is infected with a virus, the virus can connect the host directly to an external SMTAP server such as MGA. In such situations, the host is likely to be infected by a virus that causes the host to establish a connection with an external SMTAP server. Examples include spam and Open Relay Blocking Systeme (SORBS) dynamic hosts, and Not Just Another Bogus List (MJABL) dynamic hosts.
However, in some cases, direct connections are not initiated by viruses, such as when a novice user makes a direct connection, or when the connection is from a non-dynamic broadband host such as a DSL or cable modem. Nevertheless, as a result of such dial-ups or direct connections to external SMTAP servers from dynamic hosts, either determine a high virus score or increase the already determined virus score directly. It can reflect the increased likelihood that the connection will be caused by a virus.
Another feature is the use of a blacklist of misused hosts as a source of virus information, which tracks hosts that have been misused by viruses in the past. A host can be exploited if the server is an open relay, an open proxy, or otherwise vulnerable to anyone being able to deliver email anywhere. The blacklist of misused hosts uses one of two techniques to track the misused hosts: the content sent by the infected host and the infected host by the connection time scan. To search. An example is the Exploits Block List (XBL), which uses data from the Composite Blocking List (CBL) and the Open Proxy Monitor (OPM) and Distributed Server Boycott List (DSBL).
Another feature is that virus information processors develop a blacklist of senders and networks with a history of transmitting viruses. For example, the highest virus score can be assigned to the IP address of an individual who is known to send only the virus. A moderate virus score can be associated with the IP address of an individual who is known to send both a virus and a legitimate message that is not infected with the virus. Moderate to low virus scores can be assigned to networks containing one or more individual infected hosts.
Another feature, in addition to the discussion above, is the inclusion of an extensive set of tests that identify suspicious messages, such as identifying attachment characteristics. For example, you can use the general header test to look up a defined string or regular expression for any general message header, as in the example below: head X_MIME_FOO X-Mime = ~ / foo / head SUBJECT_YOUR Subject = ~ / your document / As another example, the general body test can be used to test the message body by searching for a defined string or regular expression, such as the example below: body HEY_PAL / hey pal | long time, no see / body ZIP_PASSWORD /\.zip password is / i As yet another example, functional tests can be used to create custom tests that test the very specific aspects of test messages, such as the example below: eval EXTENSION_EXE message_attachment_ext (.exe) eval MIME_BOUND_FOO mime_boundary (-/ d / d / d / d [af]) eval XBL_IP connecting_ip (exploited host) As another example, you can use a metatest based on multiple features like the one above to create a metarule for a rule like the example below: meta VIRUS_FOO ((SUBJECT_FOO1 || SUBJECT_FOO2) && BODY_FOO) meta VIRUS_BAR (SIZE_BAR + SUBJECT_BAR + BODY_BAR> 2) Another feature that can be used is to extend the above virus scoring approach to one or more machine learning techniques so that not all rules need to be activated, and to minimize false positives and missed detections. Is to provide an accurate classification by. For example, one or more of the following methods can be adopted: a decision tree that provides discrete answers; a cognition that provides additional scores; and a Bayesian analysis that maps probabilities to scores.
Another feature is the incorporation of the severity of the threat of a sudden outbreak of a virus based on the effects of the virus into the virus scoring. For example, if a virus causes the infected computer's hard drive to remove all content, its virus score can be increased, while the virus that simply displays the message remains unchanged or further reduced. Can have a virus score.
Another additional feature is to expand the options for handling suspicious messages. For example, a suspicious message indicates that the message is suspicious, for example, by adding a virus score to the message (for example, in the subject or body) so that the user is warned of the determined level of virus risk for the message. Can be tagged as. As another example, generate a new message to warn the recipient of an attempt to send a virus-infected message, or a new, uninfected message that contains a non-virus-infected part of the message. Can be created.
2.7 Example use case The hypothetical explanations below provide how the approaches described herein may be used to control outbreaks of the virus.
As a first use case, suppose a new virus entitled "Sprosts.ky" propagated through a Visual Basic macro embedded in Microsoft Excel®. Shortly after the virus hit, the virus score moved from 1 to 3 for attachment .xls, and Big Company, the user of the approach herein, begins delaying the delivery of Excel files. .. The Big Company network administrator receives an email stating that the .xls file is currently in quarantine. Sophos then sends a warning after an hour stating that a new update file is available to stop the virus. The network administrator then verifies that his IronPort C60 has the latest updates installed. Although network administrators have set a delay time of 5 hours for quarantine queues, Excel files are extremely important to the enterprise, so administrators cannot afford to wait another 4 hours. Therefore, the administrator is IronPort Access the C60 and manually clear the queue and send all messages with Excel files attached via Sophos antivirus checking. Administrators discover that 249 of these messages were virus-positive, and one was uninfected and was not captured by Sophos. The message is delivered with a total delay of 1.5 hours.
As a second use case, suppose the "Clegg.P" virus propagated through an encrypted zip file. The Big Company network administrator receives an email warning that the virus score has skyrocketed, but the administrator ignores the warning and relies on the automatic processing provided herein. To do. Six hours later, at dawn, the administrator receives a second page warning that the quarantine queue has reached 75% of its capacity. By the time the admin arrived at work, Clegg.P had filled Big Company's quarantine queue. Fortunately, the network administrator is IronPort The C60 has a policy of delivering messages normally if the quarantine queue overflows, and Sophos issued new updates all night before the quarantine queue overflowed. Only two users were infected before the virus score value triggered the quarantine queue, so the administrator faces only the overflowing quarantine queue. Based on the assumption that all the messages were viruses, the administrator would remove the messages from the queue at once, automatically remove them, and save them without using the load on IronPort C60. As a precautionary approach, network administrators block all encrypted .zip files during a defined future time cycle.
3.0 Approach to block "spam" messages FIG. 7 is a block diagram that may be used to block "spam" messages and for other types of email scanning approaches. In this context, the word "spam" refers to any junk e-mail, and the word "ham" refers to a legitimate mass of e-mail. The term "TI" refers to threat identification, that is, determining that a virus has suddenly occurred or spam communication has occurred.
During service provider 700, one or more TI development computers 702 are connected to corpus server cluster 706, which acts as a corpus or master repository for threat identification rules, and for evaluating threat identification rules. Apply to the message and generate a score value as a result. The mail server 704 of the service provider 700 feeds the ham email to the corpus server cluster 706. One or more spam traps 716 send spam emails to the corpus. Spam Trap 716 is an established and seeded email address for spammers, so that the address receives only spam emails. Messages received by Spam Trap 716 are converted to message signatures or checksums, which are stored in corpus server cluster 706. One or more avatars 714 give the corpus a non-confidential email for evaluation.
During service provider 700, one or more TI development computers 702 are connected to corpus server cluster 706, which acts as a corpus or master repository host for threat identification rules, and for evaluating threat identification rules. Apply to the message and generate a resulting score value. The mail server 704 of the service provider 700 feeds the ham email to the corpus server cluster 706. One or more spam traps 716 send spam emails to the corpus. Spam Trap 716 is an email address that has been established and seeded by spammers, so that the address receives only spam emails. Messages received by Spam Trap 716 are converted to message signatures or checksums, which are stored in corpus server cluster 706. One or more avatars 714 give the corpus a non-confidential email for evaluation.
Scores created by Corpus Server Cluster 706 are connected to Rule / URL Server 707, which provides rules and URLs associated with viruses, spam, and other email threats to customers as well. Issue to one or more messaging gateways 107 of service providers 700 located at. The messaging gateway 107 periodically reads out new rules via HTTPS forwarding. The Threat Operations Center (TOC) 708 may generate an interim rule for testing purposes and send it to the corpus server cluster 706. Threat Operations Center 708 refers to the personnel, tools, data and facilities involved in detecting and responding to virus threats. TOC708 also publishes rules approved for production use to Rule / URL Server 707, and whitelisted URLs that are known not to be associated with spam, viruses or other threats. To the rule-URL server. TI Team 710 may manually create other rules and provide them to the rules / URL server.
To give a clear example, FIG. 7 shows one messaging gateway 107. However, in various embodiments and commercial practices, the service provider 700 is connected to a number of field-deployed messaging gateways 107 under various customers or at their sites. The messaging gateway 107, avatar 714, and spam trap 716 connect to service provider 700 over a public network such as the Internet.
According to one embodiment, each customer's messaging gateway 107 has a local DNS URL blacklist module 718 that contains executable logic and a DNS blacklist. The structure of the DNS blacklist may include multiple DNS-type records A that map network addresses, such as IP addresses, to credit score values associated with IP addresses. The IP address is associated with the IP address of the sender of the spam message, or the root domain of a URL that has been found in the spam message or is known to be associated with a threat such as a phishing attack or virus. It may represent the server address.
Therefore, each messaging gateway 107 maintains its own DNS blacklist of IP addresses. In contrast, past approaches have been kept in a global location, where all queries must be received over network communication. This approach improves performance because DNS queries generated by MGA do not have to cross the network to reach a centralized DNS server. This approach is also easy to update; the central server can send incremental updates to the messaging gateway 107 on a regular basis. For filtered spam messages, other logic in Messaging Gateway 107 can extract one or more URLs from the message being tested, Provides input to blacklist module 718 as a list of pairs (URL, bitmask), and receives output as a list of blacklist IP address hits. If a hit is shown, the messaging gateway 107 can then block the delivery of the email and quarantine the email, or apply other policies such as removing the URL from the message before delivery.
In one embodiment, the blacklist module 718 also tests URL poisoning in email. URL poisoning is a technique used by spammers to include malicious or destructive URLs in electronic junk mail messages that also include non-malicious URLs, resulting in suspicion. A user who clicks on a URL without having it may unknowingly trigger a malicious local action, display of an advertisement, or the like. The presence of "good" URLs is intended to prevent spam detection software from marking messages as spam. In one embodiment, the blacklist module 718 can determine when a particular combination of malicious and good URLs provided as input represents a spam message.
One embodiment provides a system that retrieves DNS data and moves it to a hashed local database that can receive several database queries and then DNS responses.
The above approach may be implemented in a computer program configured as a plug-in to the SpamAssassin open source project. SpamAssassin consists of a set of Perl modules that can be used with core programs. The core program can provide a network protocol that performs message checking such as "spamd" shipped with SpamAssassin. SpamAssassin's plug-in architecture is application programming. It is extensible via the interface; programmers can add new rules of thumb and other features without changing the core code. Plugins are identified in the configuration file, loaded at runtime, and become a functional part of SpamAssassin. The API defines a heuristic form (rules for detecting words or phrases commonly used in spam) and message checking rules. In one embodiment, the heuristic is based on a dictionary of words, and the messaging gateway 107 allows the administrator to edit the contents of the dictionary to add or remove suspicious or known good words. Supports user interface. In one embodiment, the administrator can configure anti-spam logic 119 to scan messages against a company-specific content dictionary before performing other anti-spam scans. This approach allows the message to receive the lowest score first if the message contains company-specific or industry-standard words without having to perform other expensive spam scans with the computer. Become.
Moreover, in a broad sense, the approaches described above allow spam check engines to receive and use information that has formed the basis for credit determination but has not found direct use in spam checks. The information can be used to modify weighted values and other rules of thumb for spam checkers. Therefore, the spam checker can determine with higher accuracy whether or not the newly received message is spam. In addition, spam checkers are informed by the large amount of information in the corpus, which also improves accuracy.
3.1 Early termination from spam scans Anti-spam logic 119 usually works for each message in its complete form, which means that every element of each message is fully parsed, and then all registered tests It means to be executed. This gives a very accurate overall rating as to whether an email is hum or spam. However, once the message is sufficiently "spam", the message can be signaled and treated as spam. There is no additional information that needs to contribute to the binary nature of the email. If one embodiment implements spam and hum thresholds, then by terminating the message scanning function once the logic determines that the message is sufficiently "spam" to be sure that it is spam. , The performance of Anti-Spam Logic 119 is improved. As used herein, such an approach is referred to as Early Exit from anti-spam parsing or scanning.
Early termination can save a considerable amount of time by not evaluating hundreds of rules that simply further confirm that the message is spam. There are typically very few negative scoring rules, so once a given threshold is hit, Logic 119 can make a positive decision that the message is spam. Two other performance gains are also implemented using a mechanism called rule ordering and execution and on-demand parsing.
Rule ordering and execution is a mechanism that uses indicators that are reliably and quickly available. The rules are organized and grouped into test groups. After each group is run, the current score is checked and a decision is made as to whether the message is sufficiently "spam". If it is spammy, Logic 119 will stop processing the rule and announce that the message is spam.
On-demand parsing performs message parsing as part of anti-spam logic 119 only when necessary. For example, if parsing only the message headers results in the determination that the message is spam, no other parsing operation is performed. Among other things, the rules applicable to message headers can be a very good indicator of spam; if anti-spam logic 119 determines that a message is spam based on the header rules, the body will be parsed. Not done. As a result, the performance of Anti-Spam Logic 119 is improved. This is because parsing headers are more expensive to use a computer than parsing the message body.
As another example, the message body is parsed, but the HTML element is excluded if the rule applied to the non-HTML body element results in a spam verdict. HTML parsing or URI blacklisting tests (as detailed below) are only performed when needed.
FIG. 11 is a flow chart of the process of executing the message threat scan by the early termination approach. Multiple rules are received in step 1102. The rules specify the characteristics of electronic messages that indicate the threat associated with the message. Therefore, if a rule matches a message element, the message is probably threatening or spamming. Each rule has a priority value, and each rule is associated with a message element type.
At step 1104, an email message with the destination address of the recipient account is received. The message contains multiple message elements. Elements typically include headers, email body, and HTML body elements.
In step 1106, the following message elements are extracted. As shown in block 1106A, step 1106 can involve extracting a header, email body, or HTML body element. As an example, assume that only the message headers are extracted in step 1106. Extraction typically involves making a temporary copy in the data structure.
In step 1108, the next rule is selected from the group of rules for the same element type, based on the order of priority of the rules. Therefore, step 1108 reflects that for the current message element extracted in step 1106, only the rules for that element type are considered, and the rules are collated according to their priority order. For example, if the message header was extracted in step 1106, only the header rules will be collated. Unlike past approaches, the entire message is not considered at the same time, and not all rules are considered at the same time.
In step 1109, the threat score value for a message is determined by matching only the current message element to the current rule only. Alternatively, steps 1108 and 1109 can involve selecting all the rules that correspond to the current message element type and matching all such rules against the current message element. Therefore, Figure 11 determines whether early termination is possible by testing after each rule, or matching all rules for a particular message element type and then early termination is possible. Including that.
If the threat score value is greater than the specified threshold as tested in step 1110, exit from scan, parsing and collation is performed in step 1112, and the threat score value is output in step 1114. To. As a result, if the threshold is exceeded early in the scan, extraction, and rule matching processes, early termination from the scan process may be achieved and the threat score value may be output much more rapidly. In particular, if the result of the header rule is a threat score value that exceeds the threshold, the costly process of displaying HTML message elements and collating the rules with a computer can be omitted.
However, if the threat score does not exceed the threshold in step 1110, then in step 1111 a test is run to determine if all the rules for the current message element have been matched. Step 1111 is not required in the alternative form described above, where all rules for message elements are collated in step 1110 prior to testing. If there are other rules for the same message element type, control returns to step 1108 and collates those rules. If all the rules for the same message element type have already been collated, control returns to step 1106 and considers the next message element.
The processing of FIG. 11 may be performed by an anti-spam scanning engine, an anti-virus scanner, or a general threat scanning engine capable of identifying multiple different types of threats. Threats can include any one of viruses, spam, or phishing attacks.
Thus, in one embodiment, once certainty about the message nature is reached, a logic engine that performs anti-spam, anti-virus, or other message scanning actions does not test or act on the message. The engine classifies the rules into a group of priorities, so that the most effective and least costly tests are performed first. The engines are logically ordered to avoid parsing until a particular rule or set of rules requires parsing.
In one embodiment, rule priority values are assigned to the rule, and the rules can be ordered at run time. For example, a rule with a priority of -4 runs before a rule with a priority of 0, and a rule with a priority of 0 runs before a rule with a priority of 1000. In one embodiment, rule priority values are assigned by the administrator when a set of rules is created. Examples of rule priorities include -4, -3, -2, -1, BOTH, VOF, and are assigned based on rule validity, rule type, and rule overhead. For example, a header rule that is very effective and is a simple regular expression comparison may have a priority of -4 (first run). BOTH shows that rules are effective in detecting both spam and viruses. VOF shows the rules that are executed to detect the sudden outbreak of a virus.
In one embodiment, the threat identification team 710 (Figure 7) determines the classification and ordering of rules and assigns priorities. The TI Team 710 can continually evaluate the statistical effectiveness of the rules and decide how to order the rules for execution, including assigning different priorities.
In one embodiment, the message header is first parsed and the header rule is executed. Next, the message body is decrypted and the mail body rule is executed. Finally, the HTML element is displayed and the body and URI rules are executed. After each parsing step, a test is run to determine if the current spam score is above the spam request threshold. If it is large, then the parser exits and no subsequent steps are performed. Additional or alternative, the test is run after each rule has been run.
Table 3 is a matrix showing an example of the order of events in Anti-Spam Logic 119 when an early termination is performed. The HEAD line indicates that the message HEAD is parsed and header tests are performed, and that such tests support early termination and allow all priority ranges (-4 to VOF).<tables num="3"><img file="JP5118020B2_D0003.tif" /></tables>
3.2 Spam scan judgment cache With a given spam message, anti-spam logic 119 may require an enormous amount of time to determine if the message is spam. Therefore, spammers may use "poison message" attacks that repeatedly send such esoteric messages in an attempt to force system administrators to disable anti-spam logic 119. In one embodiment, to address this issue and improve performance, the anti-spam verdict of the message generated by anti-spam logic 119 is stored in the verdict cache 115 in the messaging gateway 107, and the anti-spam logic 119 , Reuse cached verdicts to process messages with the same body.
In an effective practice, a determination is called a "true determination" if the determination read from the cache is the same as the determination that would be returned by the actual scan. A decision from the cache that does not match the decision from the scan is called a "false decision". In effective practice, some performance gains are trade-off evaluated to ensure reliability. For example, in one embodiment, a digest of the "subject" column message is included as part of the key to the cache, which reduces the cache hit rate, but also reduces the chances of a false decision.
Spammers may attempt to disable the use of the verdict cache by including non-printing, invalid URL tags that have different forms of contiguous messages in the body and otherwise identical content. The use of such tags in the body of a message causes the message digest of the body to differ within such a series of messages. In one embodiment, an ambiguous digest that produces the algorithm can be used, in which the HTML element is parsed and hidden bytes are removed from the input and placed in the digest algorithm.
In one embodiment, the determination cache 115 is implemented as a determination of the Python dictionary from anti-spam logic 119. The key to the cache is the message digest. In one embodiment, the anti-spam logic 119 includes bright mail software, and the cache key includes a DCC "fuz2" message digest. Fuz2 is an MD5 hash, or a digest of the semantically unique part of the message body. Fuz2 parses the HTML and omits bytes in the message that do not affect what the user sees when viewing the message. Fuz2 also attempts to omit parts of the message that are frequently modified by spammers. For example, a subject field that starts with "Dear" is excluded from the input and included in the digest.
In one embodiment, anti-spam logic 119 initiates processing of messages eligible for spam or virus scanning, and a message digest is created and stored. If the message digest cannot be created, or if the use of decision cache 115 is disabled, the digest is set to None. The digest is used as a key to execute a reference in the determination cache 115, and determines whether or not a previously calculated determination is stored for a message having the same message body. The word "identical" means that some of the messages are identical, which the reader considers meaningful in deciding whether or not the message is spam. If a hit occurs in the cache, the cached verdict is read and no further message scans are performed. If the digest does not exist in the cache, anti-spam logic 119 is used to scan the message.
In one embodiment, the determination cache 115 has a size limit. When the size limit is reached, the longest unused input is removed from the cache. In one embodiment, each cached input expires at the end of the configurable input lifetime. The default lifetime is 600 seconds. The size limit is set to 100 times the input lifetime. Therefore, the cache requires a relatively small amount of memory, about 6MB. In one embodiment, each value in the cache is a tuple that includes the time entered, the verdict, and the time it took the anti-spam logic 119 to complete the initial scan.
In one embodiment, if the requested cache key is in the cache, then the value of the time entered is compared to the current time. If the input is still up-to-date, then the value of the cached item is returned as a verdict. When the input expires, the input is removed from the cache.
In one embodiment, some attempts may be made to compute the message digest before the determination is cached. For example, if fuz2 is available, it will be used, if not, fuz1 will be used if it is available, and if not, it will be used as a digest if "all cache parts" are available. Otherwise no cache input will be created. A digest of "all mime parts" includes, in one embodiment, a concatenation of digests of the MIME parts of a message. If there is no MIME part, the digest of the entire message body is used. In one embodiment, the "all mimes" digest is only calculated if the anti-spam logic 119 performs a message body scan for some other reason. The MIME part is extracted by scanning the text, and the marginal cost of the arithmetic digest is negligible; therefore, the operations can be combined efficiently.
In one embodiment, the decision cache is cleared at once each time the messaging gateway 107 receives a rule update from rule-URL server 707 (Figure 7). In one embodiment, whenever a configuration change in Anti-Spam Logic 119 occurs, the decision cache is cleared at once, for example, by administrative action or by loading a new configuration file.
In one embodiment, anti-spam logic 119 can scan multiple messages in parallel. Therefore, two or more identical messages can be scanned at the same time. This causes a cache miss because the cache has not yet been updated based on one of the messages. In one embodiment, the determination is cached only after one copy of the message has been completely scanned. Another copy of the same message currently being scanned is a cache miss.
In one embodiment, the anti-spam logic 119 periodically scans the entire decision cache and deletes expired decision cache inputs. In that case, anti-spam logic 119 writes log inputs to log file 113 that reports cache hits, misses, expirations, and additional counts. The anti-spam logic 119 or decision cache 115 may hold counter variables for logging or performance reporting purposes.
In another embodiment, the cached digest may be used for message filtering or antivirus determination. In one embodiment, multiple checksums are used to create a richer key that provides both a higher hit rate and a lower false positive rate. In addition, other information, such as the amount of time required to scan a long message for spam, may be stored in the verdict cache.
Optimizations can be implemented to meet the specific requirements of specific anti-spam software or logic. For example, Brightmail creates a tracker string and returns the tracker string with the message verdict; the tracker string can be attached to the message as an X-Brightmail-tracker header. The tracker string is used by the Brightmail plug-in in Microsoft Outlook® to perform language identification. The tracker string can also be sent back to Brightmail when the plugin reports a false positive.
Both the verdict and the tracker string can be different for messages with the same body. In some cases, the body is non-spam, but spam is encoded in the subject line. In one approach, the subject field of the message is included as input to the message digest algorithm along with the message body. However, the subject fields can be different if the body of both messages is clearly spam, apparently a virus, or both. For example, two messages can contain the same virus and are considered spam by Brightmail, but the subject headers may be different. Each message may have a short text attachment that differs from the other messages, and may have a different name. The names of the files in the attached files may be different. However, if both messages are scanned, the same decision will result.
In one embodiment, virus positive rules are used to improve cache hit rates. If the digest of the attachment matches the virus positive and spam positive verdicts, the past spam verdicts are reused, even if the subject and preface are different.
Some similar messages produce different tracker strings as a result of different From values and different message ID fields. Spam verdicts are the same, but apparently fake "From" values and apparently fake message-IDs allow earlier detection of verdicts and reporting of other rules to the tracker string. .. In one embodiment, the "From" header and the message-ID header are removed from the second message, the message is rescanned, and the tracker header is the same as that of the first message.
4.0 How to detect viruses based on message heuristics, sender information, dynamic quarantine operations, and fine-grained rules 4.1 Detection using message heuristics One approach provides detection of viruses that use a heuristic approach. The basic approach for detecting sudden outbreaks of viruses is Michael Olivier et al., Filed December 6, 2004, Pending Application Nos. 11 / 006,209, "Methods and Devices for Managing Computer Virus Outbreaks." It is described in.
In this context, a message heuristic refers to a group of factors used to determine the likelihood that a message is a virus if signature information about the message is not available. The heuristic may include rules for detecting commonly used words or phrases during spam. The rules of thumb may vary depending on the language used for the message text. In one embodiment, the administrative user can choose which language the rule of thumb uses in anti-spam scanning. The VSV value may be determined using message heuristics. The rules of thumb for messages may be determined by a scanning engine that performs basic anti-spam scanning and anti-virus scanning.
Since the message may contain a virus, it can be placed in quarantine storage based on the result of empirical behavior rather than the definition of a virus outbreak. Such a definition is described in the application of Oliver et al. Referenced above. Therefore, if corpus server cluster 706 contains a history of past viruses, and the message matches a pattern in the past history as a result of heuristics, then the message is in the definition of virus outbreak. It may be isolated whether they match or not. Such early quarantine provides a beneficial delay during message processing, while the TOC prepares a definition of a virus outbreak.
FIG. 8 is a graph of the number of machines infected over time in a virtual example of a sudden virus outbreak. In FIG. 8, the horizontal axis 814 represents time and the vertical axis 812 represents the number of infected machines. Point 806 represents the time it takes for an antivirus software vendor, such as Sophos, to publish an updated virus definition. The virus definition detects virus-infected messages and prevents further infection of machines in the network protected by the messaging gateway 107 using its antivirus software. Point 808 represents the time when TOC708 publishes a rule that identifies a virus outbreak for the same virus. Curve 804 changes as the number of infected machines increases over time, as shown in Figure 8, but the rate of increase decreases after point 808, and then of infected machines. The total number changes so that it eventually drops significantly past point 806. Early isolation based on the empirical rules described herein is applied at point 810 to help reduce the number of machines contained within region 816 of curve 804.
In one embodiment a variable isolation time is used. The quarantine time may be increased if heuristics indicate that it is more likely to contain the message virus. This gives the TOC or antivirus vendor maximum time to prepare a rule or definition, while applying a minimum quarantine delay to messages that are less likely to contain a virus. Therefore, quarantine time is associated with the probability that a message will contain a virus, resulting in optimal use of the quarantine buffer area, as well as minimizing quarantine time for non-essential messages.
4.2 Virus detection for each sender According to one approach, the virus score is determined in association with the IP address value of the sender of the message and stored in the database. The score therefore indicates that the message derived from the associated address may contain the virus. The premise is that a machine that sends one virus is likely to be infected by another virus or re-infected by the same or updated virus. The reason is that those machines are not well protected. Moreover, if the machine sends spam, it is more likely to send a virus next.
The IP address may identify a remote machine or may specify a machine in the corporate network protected by the messaging gateway 107. For example, an IP address may specify a machine in a corporate network that has been inadvertently infected with a virus. Such an infected machine is likely to send another message containing the virus.
In a related approach, the virus outbreak detection check can be performed simultaneously during the entire message processing as a spam check within the messaging gateway 107. Therefore, virus outbreak detection can be performed at the same time the message is parsed and spam is detected. In one embodiment, one thread performs the above actions in an ordered and continuous manner. Further, both the anti-spam detection operation and the anti-virus detection operation can be transmitted by using the result of the predetermined empirical operation.
In one embodiment, the VSV value is determined based on one or more of the following: filename extension; local and global message volume identified by sender and content. Soaring; based on the content of attachments such as "Microsoft" executables; and threat identification information by sender. In various embodiments, different sender-specific data identification information is used. Examples include dynamic or dial-up host blacklists, abused host blacklists, and virus hot zones.
Dynamic and dial-up hosts that connect to the Internet typically send outbound mail through a local SMTAP server. If the host connects directly to an external SMTAP server, such as Messaging Gateway 107, the host is probably already damaged and is sending either a spam message or an email virus. In one embodiment, the messaging gateway 107 contains logic that keeps a blacklist of dynamic hosts that have operated in the past in the manner described above, or externals such as the NJABL dynamic host list and the SORBS dynamic host list. Connect to a blacklist of retrieved dynamic hosts that can be retrieved at the source.
In this embodiment, identifying the message characteristics of an incoming message in step 502 of FIG. 5 further comprises determining whether the sender of the message is on the blacklist of dynamic hosts. If present, a higher VSV value is determined or assigned.
Step 502 includes connecting to or managing the misused host blacklist and determining whether the sender of the message is on the misused host blacklist. But it may be. An exploited host blacklist tracks hosts that are known to be infected with a virus or send spam based on the content that the infected host is sending, and Search for infected hosts by scanning the connection time. Examples include XBL (CBL and OPM) and DSBL.
In another embodiment, the service provider 700 creates and stores an internal blacklist of senders and networks with a history of transmitting viruses based on sender information received from the customer's messaging gateway 107. In one embodiment, the customer's messaging gateway 107 periodically initiates network communication to the corpus server cluster 706, and the messaging gateway 107 internal logic is spam or associated with a virus or other threat. Report the network address (for example, IP address) of the sender of the determined message. Service provider 700 logic can scan the internal blacklist on a regular basis and determine if any network address is known to send only viruses or spam. If affirmative, the logic can remember high threat level values or VSVs associated with those addresses. Moderate threat level values can be stored in association with network addresses known to send both viruses and legitimate emails. Moderate or low threat level values can be associated with networks containing one or more individual infected hosts.
Testing against a blacklist can be initiated using the above types of rules. For example, the following rule can start a blacklist test: eval DYNAMIC_IP connecting_ip (dynamic) eval HOTZONE_NETWORK connecting_ip (hotzone) eval XBL_IP connecting_ip (exploited host) 4.3 Dynamic quarantine operations including rescan In the past approach, messages are released from quarantine in first-in, first-out order. Alternatively, a first-end algorithm may be used in another embodiment. In this approach, if the quarantine buffer is full, the ordering mechanism determines which messages should be freed first. In one embodiment, the message that is considered to be the least dangerous is released first. For example, the quarantined message is released first as a result of heuristics, and the quarantined message is released second as a result of a matching virus outbreak test. To assist in this mechanism, each quarantined message is stored in the quarantine of messaging gateway 107, associated with information indicating the reason for the quarantine. Therefore, processing in the messaging gateway 107 can release the message for this reason.
The ordering may be configured in a data-driven manner by defining the order in the configuration file processed by the messaging gateway 107. Therefore, by publishing a new configuration file containing the ordering from the service provider to the customer's messaging gateway 107, those messaging gateways 107 automatically adopt the new ordering.
Similarly, when a message leaves quarantine, different actions can be taken on the quarantined message based on the threat level associated with the message as it leaves quarantine. For example, messages that are seemingly extremely threatening, but may leave quarantine as a result of an overflow, can be subject to a remove-delivery operation. In the removal-delivery operation, the attachment is removed and the message is delivered to the recipient without the attachment. Alternatively, lower threat level messages are delivered normally.
In yet another embodiment, the X header may be added to the lower threat level message. This alternative is a rule that the customer's email program (eg Eudora , "Microsoft Outlook") recognizes the X header and puts the message with the X header in a special folder (eg "latent". It is appropriate when it is composed of "dangerous messages"). In yet another alternative, the attachment for a particular threat level of the message is renamed (the message is "neutralized"), and the recipient user is proactively renamed and attached the attachment. Request that the file be made available to the application. This approach is intended to force the user to carefully examine the file before renaming it and opening it. The message may be forwarded to the administrator for evaluation. In one embodiment, any of these alternatives can be combined.
Figure 9 is a flow diagram of the approach of rescanning messages that may contain viruses. According to one embodiment, when the TOC710 releases a new threat rule to the messaging gateway 107, each messaging gateway rescans the quarantined message against the new rule. This approach offers the advantage that the message may be released from quarantine earlier. The reason is that a new rule will be used in later processing to detect that the message contains a virus. In this context, "releasing" refers to removing a message from quarantine and sending the message to an antivirus scanning process.
Alternatively, rescanning can reduce or increase message quarantine time. This minimizes the number of messages in quarantine and reduces the likelihood of releasing infected messages. For example, if the quarantine has a fixed release time and the fixed release timer expires before the antivirus vendor or other source releases the virus definition that catches the released message: Unexpected release could occur. In that scenario, the malicious message will be released automatically, and downstream processing will not catch the message.
In one embodiment, any of several events may trigger a message rescan during message quarantine. In addition, the approach in Figure 9 applies to processing messages that are in quarantine as a result of viruses, spam, or other threats or unwanted properties of the message. At step 902, the rescan timer is started and runs until it expires, and upon expiration, a rescan of all messages in the quarantine queue is triggered at step 906.
Additionally or additionally, in step 904, messaging gateway 107 receives one or more new virus threat rules, anti-spam rules, URLs, scores, or other message classification information from Rule-URL Server 707. To do. Receiving such information can also trigger a rescan at step 906. Generate a new VSV for each quarantined message using the new rules, scores and other information in the rescan step. For example, TOC Server 708 is a set of rules for virus outbreaks that are broad at first, and may later narrow the scope of the rules as more information about outbreaks becomes known, Rule-URL Server 707. It may be published via. As a result, a message that matches an earlier set of rules may not match the revised rule, and is a known misjudgment. The approach herein attempts to automatically respond to rule updates and release known false positives without intervention by the administrator of messaging gateway 107.
In one embodiment, each message in quarantine queue 316 remembers a time value indicating when the message entered quarantine, and the rescan in step 906 is in the order of quarantine intrusion time. Execute the oldest message first.
At step 908, as in step 312 of Figure 3, a test is run to determine if the new VSV for a message is equal to or above a particular threshold. The VSV threshold is set by the administrator of Messaging Gateway 107 to determine the tolerance of quarantined messages. If VSV is below the threshold, then the message can probably be released from quarantine. Therefore, control proceeds to step 910, where the normal quarantine end delivery policy is applied.
Optionally, in one embodiment, the messaging gateway 107 may implement separate reporting thresholds. If the message has a VSV above the reporting threshold, as tested in step 907, the messaging gateway 107 notifies the service provider 700 in step 909 and continues processing the message. Such notifications may provide important input to determine that a sudden outbreak of a new virus has occurred. In certain embodiments, such reporting is an aspect of "SenderBase Network Participation" (SBNP), and can be selectively enabled by administrators using configuration settings.
Applying the delivery policy in step 910 immediately queries the message delivered to the recipient in an unmodified form, removes attachments, performs content filtering, or performs other checks on the message. May include performing. Applying a delivery policy may include adding an X header to the message indicating that a virus scan will result. All applicable X headers may be added to the message in the order in which the actions occurred. Applying a delivery policy may include modifying the subject line of the message to indicate the possible presence of a virus, spam or other threat. Applying a delivery policy means changing the destination of the message to alternating recipients and remembering an archived copy of the message for subsequent analysis by other logic, system or personnel. It may be included.
In one embodiment, if a message is present in one of several quarantines, and one quarantine determines that removing the attachment is the correct action, then apply the delivery policy in step 910. Doing involves removing all attachments from the message before delivering it. For example, the messaging gateway 107 may support a virus sudden outbreak quarantine queue 316 and another quarantine queue that holds messages that appear to violate gateway policy, such as the presence of unacceptable words. Virus Sudden Occurrence Quarantine Queue 316 is configured to remove attachments immediately after overflow before delivery. Suppose the message exists on both the virus outbreak quarantine queue 316 and another policy quarantine queue, and happens to overflow the virus outbreak quarantine queue 316. The next time the administrator manually releases the same message from the policy quarantine queue, the attachment is then removed again before delivery.
In step 912, the message is delivered.
If the test in step 909 is true, the message is problematic and probably needs to be kept in quarantine.
Optionally, each message may be assigned to an expiration time value, and the expiration time value is stored in the database of messaging gateway 107 associated with quarantine queue 316. In one embodiment, the expiration time is equal to the time the message entered the quarantine queue 316 and the specified retention time. Expiration time values may vary based on message content or heuristics for messages.
At step 914, a test is run to determine if the message expiration time has expired. If it expires, the message is then removed from quarantine, but the removal of the message at that time is considered abnormal or premature termination, and therefore the abend delivery policy is applied in step 918. The message is then delivered in step 912 and is subject to the delivery policy of step 918. The delivery policy applied in step 918 may differ from the policy applied in step 910. For example, the policy in step 910 can provide unrestricted delivery, while step 918 removes attachments (for delivery of suspicious but quarantined messages longer than the expiration time). Can be required.
If the message time has not expired in step 914, then the message is kept in quarantine, as shown in step 916. If the rule that causes VSV to exceed the threshold changes, then the name and details of the rule are updated in the message database.
In various embodiments, different steps in FIG. 9 may cause the messaging gateway 107 to send one or more warning messages to an administrator or an identified user account or group. For example, a warning can be generated in steps 904, 912 or 916. An example of a warning event is reaching a specified quarantine full level or space limit; quarantine overflow; a new occurrence rule, for example, a rule that sets VSV above the quarantine threshold configured in the messaging gateway if matched. Receiving information that removes suddenly occurring rules; attempts to update new rules in the messaging gateway may be unsuccessful. Information that removes sudden rules may include receiving new rules that reduce the threat level of certain types of messages that fall below the quarantine threshold configured in the messaging gateway.
Further, a different step in FIG. 9 may cause the message gateway 107 to write one or more log inputs in the log file 113 that describe the actions taken. For example, log file input can be written if the message is released abnormally at early termination. Warnings or log entries can be sent or written when the quarantine meets the specified level. For example, a warning or log entry is sent or written when the quarantine reaches 5% full, 50% full, 75% full, and so on. The log input may include quarantine receive time, quarantine end time, quarantine end criteria, quarantine end action, number of messages in quarantine, and so on.
In another embodiment, the warning message fails scan engine update; fails rule update; fails to read rule update during a specified time cycle; rejects a specified percentage of messages; rejects a specified number of messages. ; Etc. can be shown.
FIG. 10 is a block diagram of a message flow model in a messaging gateway that implements the logic described above. Message heuristics 1002 and virus outbreak rules 1004 are provided to scan engines, such as the antivirus checker 116, which generate VSV or virus threat level (VTL) values 1005. If the VSV value exceeds the specified threshold, the message breaks into quarantine 316.
Multiple termination criteria 1006 allow a message to leave quarantine 316. Examples of termination criteria 1006 include expiration of time limit 1008, overflow 1010, manual release 1012, or rule update 1014. If the termination criterion 1006 is met, one or more termination actions 1018 occur next. Examples of termination actions 1018 include removal and delivery 1020, deletion 1022, successful delivery 1024, tagging with keywords in the message subject (eg [SPAM]) 1026, and adding an X header 1028. In another embodiment, the termination action can include warning the prescribed recipient of the message.
In one embodiment, the messaging gateway 107 holds, for each sending host associated with a message, a data structure that defines a policy that affects the message received from the host. For example, a host access table contains a Boolean attribute value that indicates whether to perform a virus outbreak scan on that host, as described herein with respect to FIGS. 3 and 9.
In addition, each message processed in the messaging gateway 107 may be stored in a data structure that holds metadata indicating which message processing to perform in the messaging gateway. Examples of metadata: VSV value of message; Rule name and corresponding rule details resulting in VSV value; Message quarantine time and overflow priority; Perform anti-spam and anti-virus scans and virus outbreak scans Flags that specify whether or not to do so; and flags that allow the content filter to be bypassed.
In one embodiment, the set of configuration information stored in the messaging gateway 107 defines additional program behavior for each potential recipient of a message from the gateway to a virus outbreak scan. Messaging gateway 107 typically manages message traffic into a finite set of users, such as employees, contractors, or other users in a private corporate network, so all such configuration information is available. May be managed for potential recipients of. For example, the per-recipient configuration values specify a list of message attachment extension types (".doc", ".ppt", etc.) that are excluded from the scan review described herein. It may be, and the value is a value indicating that the message should not be quarantined. In one embodiment, the configuration information can include a specific threshold for each recipient. Therefore, the tests in steps 312 and 908 may have different results for different recipients, depending on the associated threshold.
The messaging gateway 107 may also manage messages filtered using the techniques of FIGS. 3 and 9, VSVs of such messages, and a database table that counts the number of messages sent to message quarantine 316.
In one embodiment, each message quarantine 316 has a plurality of associated programmatic actions that control how the message terminates quarantine. Seeing FIG. 3 again, the termination action may include the manual release of the message from message quarantine 316, based on operator decision 318. The termination action may include the automatic release of the message from message quarantine 316 when the expiration timer expires, as shown in FIG. The termination action may include early termination from message quarantine 316 when the quarantine is full, as an implementation of overflow policy 322. "Early termination" means releasing a message prematurely, based on resource limits such as queue overflow, prior to the end of the expiration time value associated with the message.
Successful message termination actions and early termination actions may be organized as primary and secondary actions of the type described above for delivery policy step 910. Primary actions may include bounce, delete, attachment removal and delivery, and delivery. Secondary actions may include subject tags, X-headers, destination changes, or saves. The secondary action is not associated with the primary action, delete. In one embodiment, the secondary action, destination change, is not on the messaging gateway 107, but on the corpus server cluster 706, or on an "out-of-box" secondary quarantine queue hosted on another element within the service provider 700. , Allows you to send a message. This approach allows TI Team 710 to examine quarantined messages.
In one embodiment, the premature termination action from quarantine due to quarantine queue overflow may include any primary action, including removal and delivery of attachments. Any secondary action may be used for such early termination. The administrator of the messaging gateway 107 may select the primary action and the secondary action immediately after the early termination by issuing a configuration command to the messaging gateway using the command interface or GUI. Additional or alternative, message heuristics determined as a result of performing an antivirus scan or other message scan may differ in the corresponding early termination action.
In one embodiment, the local database in the messaging gateway 107 stores the name of the attachment of the received message in message quarantine 316, and the size of the attachment.
The rescan at step 906 may occur for a particular message, in response to other actions in Messaging Gateway 107. In one embodiment, the messaging gateway 107 implements a content filter that can change the content of a received message according to one or more rules. If the content filter changes the content of a received message that has been scanned for viruses in the past, then the VSV value of that message can change immediately after rescanning. For example, if a content filter removes an attachment from a message and a virus is in the attachment, the removed message may no longer carry a virus threat. Therefore, in one embodiment, if the content filter modifies the content of the received message, a rescan is performed in step 906.
In one embodiment, the administrator of Messaging Gateway 107 can use console commands or other user interface commands to retrieve the content of Quarantine 316. In one embodiment, the search can be performed based on the attachment name, attachment type, attachment size, and other message attributes. In one embodiment, the file type search can only be performed on messages that are in quarantine 316 and not in policy quarantine or other quarantine. This is because such a search requires a scan of the message body, which can negatively affect performance. In one embodiment, the administrator can display the contents of the virus outbreak quarantine 316 in an order sorted according to any of the above attributes.
In one embodiment, when a message is placed in quarantine 316 via the process of FIG. 3 or 9, messaging gateway 107 automatically displays a list of virus outbreak quarantines. In one embodiment, the list contains the following attribute values for each message in quarantine: the name of the sudden occurrence identifier or rule; the name of the sender; the domain of the sender; the name of the recipient; the recipient Domain; Subject; Attachment name; Attachment type; Attachment size; VSV; Isolation intrusion time; Isolation retention time.
In one embodiment, the messaging gateway 107 stores a reinsert key containing an optional unique text string that can be associated with a message manually released from quarantine 316. If the released messages have a reinsert key associated with them, the released messages cannot be quarantined during subsequent pre-delivery processing in the messaging gateway 107.
4.4 Fine-grained rules The message rule is an abstract statement, and if it matches the message in Anti-Spam Logic 119, the result is a message with a higher spam score. The rule may have a rule type. Examples of rule types include damaged hosts, suspected spam sources, header characteristics, body characteristics, URIs, and learning. In one embodiment, certain sudden occurrence rules can be applied. For example, a virus outbreak detection mechanism can determine that a given type of message with a ZIP attachment and a size of 20 kb represents a virus. The mechanism can create a rule that the customer's messaging gateway 107 quarantines messages with 20 kb ZIP attachments, but not messages with 1 MB ZIP attachments. As a result, fewer false isolation actions occur.
In one embodiment, the virus information logic 114, message for identifying a string or a regular expression defined includes logic to assist in establishing the rule or test for sage header and message body: head X_MIME_FOO X-Mime = ~ / foo / head SUBJECT_YOUR Subject = ~ / your document / body HEY_PAL / hey pal | long time, no see / body ZIP_PASSWORD /\.zip password is / i In one embodiment, the functional test can test a particular aspect of the message. Each function executes custom code to look up a message, information already captured about the message, etc. No test can be formed using a simple logical combination of common header or body tests. For example, an effective test for matching viruses without examining the contents of a file is to compare the "filename" or "name" extension of a MIME field with the claimed MIME content type. Content is suspicious if it has a "doc" extension and the content type is neither application / octet-stream nor application /. * Word. Similar comparisons can be performed on PowerPoint, Excel, image files, text files, and executables.
Another example of testing: Tests if the first line of BASE64 content matches the regular expression / ^ TV [nopqr] / that indicates a "Microsoft" executable; email high priority Tested for the absence of X-mailer or user agent headers; the message is multipart / alternative, but tested for very different content in the alternative part; It is multipart, but tests whether it contains only HTML text; searches for a specific MIME boundary format for new outbreaks;
In one embodiment, virus information logic 114 assists in establishing a meta-rule containing a plurality of associated rules. Examples include: meta VIRUS_FOO ((SUBJECT_FOO1 || SUBJECT_FOO2) && BODY_FOO) meta VIRUS_BAR (SIZE_BAR + SUBJECT_BAR + BODY_BAR>; 2) In one embodiment, virus information logic 114 establishes and tests messages against rules based on attachment size, file name keywords, encrypted files, message URLs, and antivirus logic version values. Includes logic to help you do. In one embodiment, rules regarding attachment size are established based on discrete values rather than all possible size values; for example, rules are incremented by 1K for files with file sizes 0-5K; For files sized from 5K to 1MB, you can specify the file size with a 5K increment; and a 1MB increment.
If the attachment to the message has a name that contains one or more keywords in the rule, the file name keyword rule matches the message. Encrypted file rules test whether attachments are encrypted. Such rules may be useful for isolating messages that have an encrypted container, such as encrypted ZIP files, as attachments to the message. If the message body contains one or more URLs specified in the rule, the message URL rule matches the message. In one embodiment, messages are not scanned to identify URLs unless at least one message URL is installed on the system.
If the messaging gateway 107 runs anti-virus logic with a matching version, the rules based on the anti-virus logic version value match the message. For example, a rule may specify version "7.3.1" of the AV signature, and if the messaging gateway is running AV software on a signature file with a version number, the rule matches the message. Will do.
In one embodiment, the messaging gateway 107 automatically lowers the stored VSV for a message as soon as it receives a new rule for a more specific set of messages than previously received. For example, suppose TOC708 first delivers the rule that any message with a .ZIP attachment will be assigned to VSV "3". TOC708 then delivers the rule that messages with .ZIP attachments from 30KB to 35KB have VSV "3". Correspondingly, Messaging Gateway 107 reduces the VSV of all messages with .ZIP attachments of different file sizes to the default VSV, for example "1".
In one embodiment, anti-spam logic 119 identifies legitimate emails specific to an organization based on the characteristics of the outgoing message, such as the recipient's address, recipient's domain, and frequently used words or phrases. Can be learned. In this context, outgoing messages consist of user accounts associated with computers 120A, 120B, 120C on private network 110, and through the messaging gateway 107, recipients who are logically outside the messaging gateway. A message directed to your account. Such recipient accounts are typically on a computer connected to public network 102. Since all outgoing messages pass through the messaging gateway 107 before delivery to network 102, and such outgoing messages are rarely spam, the messaging gateway can scan for such messages, and , Can automatically generate rules of thumb or rules associated with non-spam messages. In one embodiment, learning is accomplished by training the Bayesian filter in the anti-spam logic 119 against the text of the outgoing message, and then testing the incoming message with the Bayesian filter. If the trained Bayesian filter returns a high probability, the incoming message is probably not spam, according to the probability that the originating message is not spam.
In one embodiment, the messaging gateway 107 periodically polls the rules-URL server 707 to request updates for all available rules. Rule updates may be delivered using HTTPS. In one embodiment, the administrator of Messaging Gateway 107 accesses the rule update by entering the URL of the rule update and connecting to the rule-URL server 707 using a browser and proxy server or a fixed address. You can look it up. The administrator can then deliver the update to the selected messaging gateway 107 in the managed network. Receiving a rule update either displays a user notification within the interface of the messaging gateway 107, stating that the rule update was received, or that the messaging gateway was successfully connected to the rule-URL server 707. Alternatively, it may include writing the input in the log file 113.
4.5 Communication with service providers The customer messaging gateway 107 in Figure 1 can perform an "automatic phone call" or "Sender Base Network Participation" service. In these services, the messaging gateway 107 can open a connection to the service provider 700 and provide information about the messages processed by the messaging gateway 107, thus adding such information from the fields to the corpus. Or, otherwise, it can be used by service providers to improve scoring, outbreak detection, and heuristics.
In one embodiment, tree data structures and processing algorithms are used to provide efficient data communication from the messaging gateway 107 to the service provider.
Data from the service provider generated as part of anti-spam and anti-virus checking is sent to the messaging gateway 107 in the field. As a result, the service provider creates metadata that describes what data the service provider wants the messaging gateway 107 to return to the service provider. The messaging gateway 107 collates the data matching of the metadata over a time period, eg, 5 minutes. The messaging gateway 107 then returns the connection to the service provider and provides the field data according to the metadata specifications.
In this approach, the service provider can instruct the messaging gateway 107 in the field to send the different data back to the service provider by defining different metadata and delivering it to the messaging gateway 107 at different times. Therefore, the "automatic phone call" service is extensible when instructed by the service provider. MGA does not require software updates.
In one implementation, the tree is implemented as one of a plurality of hashes. There was a standard mapping of nested hashes (or dictionaries in Python) to a tree. Predetermined nodes are named in a way that allows the MGA to return data about what is what. By naming things in the tree rather than listing things based solely on location, MGA does not need to know what the service provider will do with this data. MGA only needs to find the correct data by name and send a copy of the data back to the service provider. All the MGA needs to know is the type of data, that is, whether the data is a number or a string. The MGA does not need to perform any data operations or transformations to suit the service provider.
Constants are placed on the structure of the data. The rule is that the end of the tree is always one of two. If the target data is a number, the leaf node is a counter. When the MGA sees the next incoming message, the MGA increments or decrements the counter for that clause. If the target data is a string, then the leaf node is overwritten with that string value.
The counter approach can be used to communicate any form of data. For example, if the MGA needs to return the average score value to the service provider by communication, instead of letting the service provider notify the MGA that the service provider wants the service provider to return a specific value as the average score, one is the upper value and Two counters are used, one with the lower value. MGA doesn't need to know what it is. MGA simply counts the specified value and returns it. The service provider's logic knows that the value received from the MGA is a counter and does not need to be averaged and stored.
In this way, this approach provides a method of collating and transferring data that is invisible to the user. In this method, the device that transfers the data is unaware of the particular use of the data, but can collate and provide the data. In addition, the service provider can update its software to request additional values from the messaging gateway 107, but no MGA software update is required. This allows service providers to collect data without having to modify hundreds or thousands of messaging gateways 107 in the field.
An example of data that can be communicated from the messaging gateway 107 to the service provider 700 includes an obfuscated X-header value that contains rules that match a particular message and lead to spam verdicts.
4.7 Caller White List Module In the configuration of Figure 3, the customer's messaging gateway 107 can be deployed in the customer network to receive and process both incoming and outgoing message traffic. Therefore, the messaging gateway 107 can be configured with a calling message white list. In this approach, the destination network address of the specified message leaving the messaging gateway 107 is weighted and placed on the originating message whitelist. When an incoming message is received, the outgoing message whitelist is referenced, and if the weighted value is appropriate, the incoming message with the source network address in the outgoing whitelist is delivered. That is, weighted values are taken into account when deciding whether a message should be delivered; the presence of an address in the calling whitelist does not necessarily dictate delivery. The rationale is that sending a message to an organization implies credibility, so a message received from an organization on the calling whitelist cannot be spam or threatening. That's what it means. The calling white list may be maintained by the service provider for delivery to another customer's messaging gateway 107.
Determining the weighted value may be performed using several approaches. For example, the destination address can be processed using a credit scoring system, and the weighted value can be selected based on the resulting credit score. Message identifiers are tracked and can be compared to determine if an incoming message really responds to a past message sent. A cache of message identifiers may be used. Therefore, if the Reply-To header contains the message identifier of a message previously sent by the same messaging gateway 107, the reply is likely not spam or threat.
5.0 Implementation Mechanism-Hardware Overview The approaches for controlling the outbreak of computer viruses described herein may be implemented in various ways, and the invention is not limited to any particular practice. The approach may be integrated into an email system or mail gateway device or other suitable device, or may be implemented as a stand-alone mechanism. In addition, the approach may be implemented in computer software, hardware, or a combination thereof.
FIG. 6 is a block diagram showing a computer system 600 capable of implementing one embodiment of the present invention. The computer system 600 includes a bus 602 or other communication mechanism for communicating information, and a processing device 604 connected to the bus 602 for processing information. Computer system 600 also includes main memory 606, such as random access memory (RAM) or other dynamic storage connected to bus 602 to store information and instructions performed by processing device 604. .. The main memory 606 may be used to store temporary variables or other intermediate information during the execution of instructions executed by processing device 604. The computer system 600 further includes a read-only memory (ROM) 608 or other static storage device connected to bus 602 to store static information and instructions to the processing device 604. A storage device 610, such as a magnetic disk or an optical disk, is provided and is connected to bus 602 to store information and instructions.
The computer system 600 may be connected via a bus 602 to a display 612, such as a cathode ray tube (CRT), which displays information to the computer user. Input device 614, including alphanumeric characters and other keys, is connected to bus 602 to communicate information and command selection to processing device 604. Another type of user input device is a cursor control 616 that communicates direction information and command selection, such as a mouse, trackball, stylus, or cursor direction key, to processing device 604 and controls the movement of the cursor on the display 612. is there. This input device typically has two degrees of freedom in two axes, the first axis (eg x) and the second axis (eg y), allowing the device to locate in a plane. To.
The present invention relates to the use of a computer system 600 that applies heuristic tests to message content, manages dynamic threat quarantine queues, and terminates early from parsing and scanning to scan messages. According to one embodiment of the invention, applying heuristic tests to message content, managing dynamic threat quarantine queues, and terminating early from parsing and scanning to scan messages is in main memory 606. Provided by the computer system 600, corresponding to a processor 604 that executes one or more sequences of one or more instructions contained. Such instructions may be read into main memory 606 from another computer-readable medium, such as storage device 610. Execution of the sequence of instructions contained in the main memory 606 causes processing apparatus 604 to perform the processing steps described herein. In an alternative embodiment, the invention may be practiced using wiring circuits instead of or in combination with software instructions. Therefore, embodiments of the present invention are not limited to any particular hardware circuit and software combination.
As used herein, the term "computer-readable medium" refers to any medium that participates in providing instructions to the processor 604 for execution. Such media can take many forms, including, but not limited to, non-volatile media, volatile media, and transmission media. Non-volatile media include, for example, optical or magnetic disks such as storage device 610. Volatile media include dynamic memory such as main memory 606. Transmission media are coaxial cables, copper wires and optical fibers, including wiring including bus 602. The transmission medium can also take the form of sound waves or light waves, such as those generated during the communication of radio and infrared data.
Common forms of computer-readable media are, for example, floppy (registered trademark) disks, flexible disks, hard disks, magnetic tape, or any other magnetic medium, CD-ROM, any other optical medium, punch. Cards, paper tapes, physical media with any other hole pattern, RAM, PROM, and EPROM, FLASH-EPROM, any other memory chip or cartridge, carriers described later herein, or computers. Includes any other medium that can be read.
Various computer-readable forms may be involved in delivering one or more sequences of one or more instructions to the treatment device 604 to the storage device 610 for execution. For example, the instructions may be delivered first on the magnetic disk of the remote computer. The remote computer can load the instructions into its dynamic memory and send the instructions over the telephone line using a modem. A modem from the local to the computer system 600 can receive the data over the telephone line and use an infrared transmitter to convert the data into an infrared signal. The infrared detector can receive the data carried in the infrared signal and the appropriate circuitry can put the data into bus 602. The bus 602 transports data to the main memory 606, from which the processor 604 reads and executes instructions. The instructions received by the main memory 606 may be optionally stored in the storage device 610 either before or after execution by the processing device 604.
The computer system 600 also includes a communication interface 618 connected to bus 602. Communication interface 618 provides two-way data communication coupled to network link 620 connected to local network 622. For example, the communication interface 618 may be a card or modem of an integrated services digital network (ISDN) that provides a data communication connection to the corresponding type of telephone line. As another example, the communication interface 618 may be a local area network (LAN) card that provides a data communication connection to a compatible LAN. Wireless links may also be implemented. In any such practice, the communication interface 618 transmits and receives electrical, electromagnetic or optical signals carrying digital data streams representing various types of information.
Network link 620 typically provides data communication to other data devices over one or more networks. For example, network link 620 may provide a connection to host computer 624 or data equipment operated by Internet Service Provider (ISP) 626 over local network 622. The ISP 626 then provides data communication services over a global packet data communication network, now commonly referred to as the "Internet" 628. Both local networks 622 and Internet 628 use electrical, electromagnetic or optical signals to carry digital data streams. Signals through various networks, and signals via communication interface 618 on network link 620, carry digital data to and from computer system 600, which are exemplary forms of carrier waves that carry information.
Computer system 600 sends messages and receives data, including program code, over networks, network links 620 and communication interfaces 618. In the Internet example, the server 630 may send the requested code for the application program over the Internet 628, ISP626, local network 622 and communication interface 618. According to the present invention, one such downloaded application applies heuristic tests to message content, manages dynamic threat quarantine queues, and parses and scans as described herein. Provides scanning of messages with early termination from.
The processing device 604 executes the received code as it is stored in the received and / or storage device 610, or other non-volatile storage device for later execution. In this way, the computer system 600 may acquire the application code in the form of a carrier wave.
6.0 Extensions and alternatives In the specification described above, the present invention has been described with reference to specific embodiments thereof. However, it will be clear that various modifications and modifications can be made to the invention without departing from the broader spirit and scope of the invention. The specification and drawings should therefore be construed in an exemplary sense rather than a restrictive one. The invention also includes other contexts and applications, wherein the mechanisms and processes described herein are also available for other mechanisms, methods, programs, and processes.
In addition, the predetermined processing steps are described in this specification in a particular order, and alphanumeric labels are used to identify the predetermined steps. Unless specifically stated in the present disclosure, embodiments of the present invention are not limited to any particular order in which such steps are performed. In particular, labels are used only for the convenient identification of steps and are not intended to suggest, prescribe or require a particular order in which such steps are performed. In addition, another embodiment may use more or less steps than the steps described herein.
<figref num="1">Block diagram of a system that manages the sudden outbreak of a computer virus according to one embodiment</figref><figref num="2">A flow diagram of a process that generates a count of suspicious messages executed by a virus information source, according to an embodiment.</figref><figref num="3">A data flow diagram showing message processing based on virus sudden occurrence information according to one embodiment.</figref><figref num="4">Flow diagram of a method for determining a virus score value according to one embodiment</figref><figref num="5">A flow diagram showing the application of a group of rules for managing the sudden outbreak of a virus according to one embodiment.</figref><figref num="6">A block diagram showing a computer system in which an embodiment can be implemented.</figref><figref num="7">Block diagram of a system that can be used as an approach to block "spam" messages and for other types of email scanning processing</figref><figref num="8">Graph of time vs. number of infected machines in a virtual case of sudden virus outbreak</figref><figref num="9">Flow diagram of the approach to rescan messages that may contain viruses</figref><figref num="10">Block diagram of a message flow model in a messaging gateway that implements the logic described above</figref><figref num="11">Flow diagram of processing to execute message threat scan by early termination approach</figref>
4 sheets
Sheet 1 Sheet 2 Sheet 3 Sheet 4
Every citation, both waysCites: the store holds 6 of 7
| Document | Relation | Office |
|---|---|---|
| US20040117648A1 | Cites | United States of America |
| JP2001222480A | Cites | Japan |
| JP2002123469A | Cites | Japan |
| JP2005011369A | Cites | Japan |
| JP2005005918A | Cites | Japan |
| WO2003071753A1 | Cites | World Intellectual Property Organization (WIPO) |
| 松田陽一,Windowsメール環境からの脱出 [Part2]POP3サーバーからのメール取得とスパムフィルタの設置,UNIX USER,日本,ソフトバンクパブリッシング株式会社,2004年 6月 1日,第13巻,第6号,p.56-65 | Non-patent | – |
| 宮紀雄,ネットワークソリューション講座 メール・フィルタリング活用法 情報漏えいやスパムを遮断,日経コミュニケーション,日本,日経BP社,2000年 7月 3日,第321号,p.148-153 | Non-patent | – |
| 西村卓也,試してわかった!try&review 光ファイバで自宅にWebサイトを!プロジェクト[第46回],Linux WORLD 第4巻 第5号,日本,(株)IDGジャパン,2005年 5月 1日,第4巻,第5号,p.224-231 | Non-patent | – |
37 members in 6 offices
Priority claims9
| Document | Office | Kind | Date |
|---|---|---|---|
| 60678391 | United States of America | – | |
| 67839105 | United States of America | P | |
| 67839105 | United States of America | P | |
| 2006017783 | United States of America | W | |
| 2006017783 | United States of America | W | |
| 2005678391 | – | – | – |
| 2006017783 | – | – | – |
| US20050678391P | – | – | – |
| WO2006US17783 | – | – | – |
Members37
| Document | Office | Kind | |
|---|---|---|---|
| CA2606998A1 | Canada | A1 | |
| CA2607005A1 | Canada | A1 | |
| WO2006119506A2 | World Intellectual Property Organization (WIPO) | A2 | |
| WO2006119508A2 | World Intellectual Property Organization (WIPO) | A2 | |
| WO2006119509A2 | World Intellectual Property Organization (WIPO) | A2 | |
| WO2006122055A2 | World Intellectual Property Organization (WIPO) | A2 | |
| US2007070921A1 | United States of America | A1 | |
| US2007073660A1 | United States of America | A1 | |
| US2007078936A1 | United States of America | A1 | |
| US2007079379A1 | United States of America | A1 | |
| US2007083929A1 | United States of America | A1 | |
| US2007220607A1 | United States of America | A1 | |
| EP1877904A2 | European Patent Office (EPO) | A2 | |
| EP1877905A2 | European Patent Office (EPO) | A2 | |
| JP2008545177A | Japan | A | |
| JP2008547067A | Japan | A | |
| WO2006119506A3 | World Intellectual Property Organization (WIPO) | A3 | |
| WO2006119508A3 | World Intellectual Property Organization (WIPO) | A3 | |
| WO2006119509A3 | World Intellectual Property Organization (WIPO) | A3 | |
| WO2006122055A3 | World Intellectual Property Organization (WIPO) | A3 | |
| US7548544B2 | United States of America | B2 | |
| CN101495969A | China | A | |
| CN101558398A | China | A | |
| US7712136B2 | United States of America | B2 | |
| US7836133B2 | United States of America | B2 | |
| US7854007B2 | United States of America | B2 | |
| US7877493B2 | United States of America | B2 | |
| CA2607005C | Canada | C | |
| JP4880675B2 | Japan | B2 | |
| CN101495969B | China | B | |
| CN101558398B | China | B | |
| JP5118020B2This record | Japan | B2 | |
| EP1877904A4 | European Patent Office (EPO) | A4 | |
| EP1877905A4 | European Patent Office (EPO) | A4 | |
| CA2606998C | Canada | C | |
| EP1877905B1 | European Patent Office (EPO) | B1 | |
| EP1877904B1 | European Patent Office (EPO) | B1 |
26 legal events, as the office reported them to INPADOC
Over the term
Point at a mark for the eventEvents
| Event | Code | |
|---|---|---|
| Cancellation because of no payment of annual feesLAPS | LAPS | |
| Receipt of annual feesJAPANESE INTERMEDIATE CODE: R250R250 | R250 | |
| Receipt of annual feesJAPANESE INTERMEDIATE CODE: R250R250 | R250 | |
| Receipt of annual feesJAPANESE INTERMEDIATE CODE: R250R250 | R250 | |
| Receipt of annual feesJAPANESE INTERMEDIATE CODE: R250R250 | R250 | |
| Receipt of annual feesJAPANESE INTERMEDIATE CODE: R250R250 | R250 | |
| Renewal fee payment (event date is renewal date of database)FPAY | FPAY | |
| Certificate of patent or registration of utility modelJAPANESE INTERMEDIATE CODE: R150R150 | R150 | |
| Certificate of patent or registration of utility modelJAPANESE INTERMEDIATE CODE: R150R150 | R150 | |
| First payment of annual fees (during grant procedure)JAPANESE INTERMEDIATE CODE: A61A61 | A61 | |
| Written decision to grant a patent or to grant a registration (utility model)JAPANESE INTERMEDIATE CODE: A01A01 | A01 | |
| Written decision to grant a patent or to grant a registration (utility model)JAPANESE INTERMEDIATE CODE: A01A01 | A01 | |
| Decision of grant or rejection writtenTRDD | TRDD | |
| Written amendmentJAPANESE INTERMEDIATE CODE: A523A521 | A521 | |
| Written permission of extension of timeJAPANESE INTERMEDIATE CODE: A602A602 | A602 | |
| Written request for extension of timeJAPANESE INTERMEDIATE CODE: A601A601 | A601 | |
| Written permission of extension of timeJAPANESE INTERMEDIATE CODE: A602A602 | A602 | |
| Written request for extension of timeJAPANESE INTERMEDIATE CODE: A601A601 | A601 | |
| Notification of reasons for refusalJAPANESE INTERMEDIATE CODE: A131A131 | A131 | |
| Written amendmentJAPANESE INTERMEDIATE CODE: A523A521 | A521 | |
| Written permission of extension of timeJAPANESE INTERMEDIATE CODE: A602A602 | A602 | |
| Written request for extension of timeJAPANESE INTERMEDIATE CODE: A601A601 | A601 | |
| Notification of reasons for refusalJAPANESE INTERMEDIATE CODE: A131A131 | A131 | |
| Report on retrievalJAPANESE INTERMEDIATE CODE: A971007A977 | A977 | |
| Written amendmentJAPANESE INTERMEDIATE CODE: A523A521 | A521 | |
| Written request for application examinationJAPANESE INTERMEDIATE CODE: A621A621 | A621 |
Numbers
- Publication
- 5118020
- Publication, DOCDB
- 5118020
- Publication, EPODOC
- JP5118020B
- Application
- 2008510321
- Application, DOCDB
- 2008510321
- Application, EPODOC
- JP20080510321
Titles2
- Japanese
- 電子メッセージ中での脅威の識別
- English
- Identifying threats in electronic messages
Classification
- CPC, 7
- G06Q10/107
- H04L51/212
- H04L63/123
- H04L63/126
- H04L63/145
- H04L61/4511
- H04L51/234
- IPC, 2
- G06F13 00
- G06F21 56
