System and method of indexing unique electronic mail messages and uses for same
Abstract
A system and method for identifying unique e-mail messages in a large-scale enterprise environment using external servers and database systems. Message uniqueness is determined by assigning a message tag to each message based on the attributes (500) of the email message. The message tag (506) can be calculated using a hash algorithm (504) to speed up indexing and comparison. The message tag (506) is compared with an index file of message tags associated with pre-existing email messages. If a matching message tag is found in the index file, the email message is not unique. Otherwise, the email message is unique, and the message tag is added to the index file (406). The system may include a relational database for storing the index file. A filing system and method using the uniqueness check feature of the present invention are also disclosed.
Term
Term ended
Expired 12 February 2022, 4.6 years ago.
- Priority
- Filed
- Granted
- Expired
- Today
38 claims: 5 independent, 33 dependent
- 1第 1. 一种在从电子邮件消息传送系统中所抽取的多个电子邮件消息 中标识惟一电子邮件消息的方法,所述方法包括: 从所述电子邮件消息传送系统上的邮箱中检索消息,所述消息包括 多个消息属性; 根据所述多个消息属性的至少一部分计算消息标记; 复查在与多个用户关联的单个共享的索引文件中存储的消息标记 的列表; 根据在该单个共享的索引文件中是否查找到所述消息标记来确定 所述消息是否不是已经存储在消息档案中的重复的消息;以及 如果该消息不是重复的消息,则在该单个共享的索引文件中存储该 消息标记以及在该消息档案中存储该消息。
- 2权利要求1的方法,其中所述消息标记通过连接选自所述多个消 息属性中的至少两个属性而加以计算。
- 3权利要求2的方法,其中所述消息标记还通过将散列算法施加到 所述消息标记以便构成统一的串而加以计算,其中所述统一的串具有预 定的长度。
- 4权利要求3的方法,其中所述散列算法是MD5散列算法。
- 5权利要求1的方法,其中所述多个消息属性包括发送者的名字以 及发送者的提交时间,并且其中所述消息标记通过将所述发送者的名字 连接到所述发送者的提交时间而加以计算。
- 6权利要求1的方法,其中所述多个消息属性包括发送者的名字、 发送者的提交时间以及主题,并且其中所述消息标记通过将所述发送者 的名字和所述主题连接到所述发送者的提交时间而加以计算。
- 7权利要求1的方法,其中所述索引文件被存储在关系数据库系统 中。
- 8一种在电子邮件消息传送系统以外的系统中归档多个电子邮件 消息的方法,所述方法包括: 阅读所述电子邮件消息传送系统上的第一邮箱中的第一消息,该第 一消息包括至少第一发送者的名字和至少第一发送者的提交时间; 根据第一发送者的名字和第一发送者的提交时间计算第一消息标记; 将第一消息存储到消息档案中并将第一消息标记存储到于与多个 02804805.9 第 用户相关联且与所述消息档案相关联的单个共享的索引文件中; 阅读所述电子邮件消息传送系统上的第二邮箱中的第二消息,该第 二消息包括至少第二发送者的名字和至少第二发送者的提交时间; 根据第二发送者的名字和第二发送者的提交时间来计算第二消息标 记; 比较第二消息标记和第一消息标记;以及 如果该第二消息不是已存储在该消息档案中的第一消息的重复,则 将第二消息存储到所述消息档案中,并且将第二消息标记存储到所述单 个共享的索引文件中。
- 9权利要求8的方法,其中第一消息标记被通过将第一发送者的名 字和第一发送者的提交时间连接起来以便构成第一消息串而加以计 算,并且其中第二消息标记被通过将第二发送者的名字和第二发送者的 提交时间连接起来以便构成第二消息串而加以计算。
- 10权利要求9的方法,其中第一消息标记还被通过将散列算法施 加到第一消息串以便构成第一统一的串而加以计算,其中该第一统一的 串具有预定的长度,并且其中第二消息标记还被通过将所述散列算法施 加到第二消息串以便构成第二统一的串而加以计算,其中第二统一的串 具有所述预定的长度。
- 11权利要求10的方法,其中所述散列算法是MD5散列算法。
- 12权利要求8的方法,其中该第一邮箱和第二邮箱是在所述电子 邮件消息传送系统上的不同邮箱。
- 13权利要求8的方法,其中所述索引文件被存储在关系数据库系 统中。
- 14权利要求8的方法,其中所述消息档案是关系数据库系统。
- 15一种用于标识惟一电子邮件消息的系统,其中所述系统位于电 子邮件消息传送系统的外部,所述系统包括: 与该电子邮件消息传送系统通信的重复性检查器;以及 与多个用户关联的一单个共享的索引文件且该索引文件包括多个 预定的消息标记, 其中所述重复性检查器被配置成阅读来自所述电子邮件消息传送 系统的消息,其中所述消息包括与所述消息相关联的多个属性, 其中所述重复性检查器使用其中至少两个属性来计算所述消息的 02804805.9 第 消息标记,并比较所述计算的消息标记与所述单个共享的索引文件, 其中如果所述计算的消息标记匹配于所述单个共享的索引文件中 的条目,则所述重复性检查器确定所述消息是已存储在一消息档案中的 重复消息,否则,如果所述计算的消息标记不匹配于所述单个共享的索 引文件中的条目,则所述计算的消息标记就被添加到所述单个共享的索 引文件中且该消息被存储在该消息档案中。
- 16权利要求15的系统,其中所述消息标记被通过将所述至少两个 属性连接起来以便构成消息串而加以计算。
- 17权利要求16的系统,其中所述消息标记还通过将散列算法施加 到所述消息串以便构成统一的串而加以计算,其中所述统一的串具有预 定的长度。
- 18权利要求17的系统,其中所述散列算法是MD5散列算法。
- 19权利要求15的系统,其中所述重复性检查器阅读来自在所述电 子邮件消息传送系统上的邮箱的所述消息。
- 20权利要求15的系统,其中所述多个属性包括发送者的名字和发 送者的提交时间。
- 21权利要求20的系统,其中所述多个属性还包括主题串,并且其 中所述消息标记被通过将发送者的名字、所述发送者的提交时间、以及 所述主题串连接起来以便构成消息串而加以计算。
- 22权利要求21的系统,其中该消息标记进一步通过将一散列算 法施加到所述消息串以构成统一的串而计算,其中所述统一的串具有预 定的长度。
- 23权利要求15的系统,其中所述索引文件被存储在关系数据库系 统中。
- 24-种用于归档多个电子邮件消息的系统,其中所述系统位于电 子邮件消息传送系统的外部,所述系统包括: 用于阅读来自所述电子邮件消息传送系统上的第一邮箱中的第一 消息的装置,该第一消息包括至少第一发送者的名字和至少第一发送者 的提交时间; 用于根据第一发送者的名字和第一发送者的提交时间计算第一消 息标记的装置; 用于将该第一消息存储到消息档案中并且将第一消息标记存储到 02804805.9 第 与多个用户关联且与所述消息档案相关联的一单个共享的索引文件中 的装置; 用于阅读来自所述电子邮件消息传送系统上的第二邮箱中的第二 消息的装置,该第二消息包括至少第二发送者的名字和至少第二发送者 的提交时间; 用于根据第二发送者的名字和第二发送者的提交时间计算第二消 息标记的装置; 用于比较第二消息标记与第一消息标记的装置;以及 用于在该第二消息不是该消息档案中已经存储的第一消息的重复 的情况下,将第二消息存储到所述消息档案中并且将第二消息标记存储 到所述单个共享的索引文件中的装置.
- 25权利要求24的系统,其中第一消息标记被通过将第一发送者的 名字和第一发送者的提交时间连接起来以便构成第一消息串而加以计 算,并且其中第二消息标记被通过将所述第二发送者的名字和所述第二 发送者的提交时间连接起来以便构成第二消息串而加以计算。
- 26权利要求25的系统,其中该第一消息标记进一步通过将一散列 算法施加到第一消息串以构成第一统一的串而计算,其中该第一统一的 串具有预定的长度,并且其中第二消息标记还被通过将所述散列算法施 加到第二消息串以便构成第二统一的串而加以计算,其中第二统一的串 具有所述预定的长度。
- 27权利要求26的系统,其中所述散列算法是MD5散列算法。
- 28权利要求24的系统,其中第一消息还包括第一主题串,并且第 二消息还包括第二主题串,并且其中第一消息标记被通过将所述第一发 送者的名字、所述第一发送者的提交时间以及所述第一主题串连接起来 以便构成第一消息串而加以计算,其中第二消息标记被通过将所述第二 发送者的名字、所述第二发送者的提交时间和所述第二主题串连接起来 以便构成第二消息串而加以计算。
- 29权利要求24的系统,其中所述索引文件被存储在关系数据库系 统中。
- 30权利要求24的系统,其中所述消息档案是关系数据库系统。
- 31一种从外部归档选自电子邮件消息传送系统中的多个电子邮件 消息的系统,所述系统包括: 02804805.9 第 与所述电子邮件消息传送系统通信的档案服务器; 与所述档案服务器通信的重复性检查器;以及 与所述档案服务器通信的档案消息库, 其中当所述档案服务器阅读来自所述电子邮件消息传送系统中的 消息时,与所述消息相关联的多个属性被从所述档案服务器发送到所述 重复性检查器, 其中所述重复性检查器使用其中至少两个属性计算所述消息的消 息标记,并比较所述计算的消息标记与多个用户所关联的一单个共享的 索引文件, 其中如果计算的消息标记与该单个共享的索引文件中的条目匹 配,则重复性检查器向所述档案服务器指示所述消息是已经存储在该档 案消息库中的重复消息,否则,如果所述计算的消息标记不匹配于所述 单个共享的索引文件中的条目,则所述计算的消息标记被添加到所述单 个共享的索引文件, 其中如果所述消息不是重复消息,则所述档案服务器就将所述消息 存储到所述档案消息库中。
- 32权利要求31的系统,其中所述消息标记被通过将至少两个属性 连接起来以便构成消息串而加以计算。
- 33权利要求32的系统,其中所述消息标记还被通过将散列算法施 加到所述消息串以便构成统一的串而加以计算,其中所述统一的串具有 预定的长度。
- 34权利要求33的系统,其中所述散列算法是MD5散列算法。
- 35权利要求31的系统,其中所述档案服务器阅读来自所述电子邮 件消息传送系统上的邮箱中的所述消息。
- 36权利要求35的系统,其中所述多个属性包括发送者的名字和发 送者的提交时间。
- 37权利要求36的系统,其中所述多个属性还包括主题串,并且其 中所述消息标记被通过将所述发送者的名字、所述发送者的提交时间、 以及所述主题串连接起来以便构成消息串而加以计算。
- 38权利要求37的系统,其中所述消息标记还被通过将散列算法施 加到所述消息串以便构成统一的串而加以计算,其中所述统一的串具有 预定的长度。 02804805.9
Independent claims38
50 paragraphs, as filed
The first index unique e-mail message and its use system and method This application requires No. 60/268, 092 filed on February 12, 2001 and No. 60/347, 278 filed on January 14, 2002 United States The benefits of provisional applications, all of them are introduced here for reference.
BACKGROUND OF THE INVENTION Field of the Invention The present invention generally relates to systems for managing email messages and messaging. More specifically, the present invention relates to manipulating messages extracted from email messaging systems.
BACKGROUND OF THE INVENTION Electronic mail ("email") messaging system has become a core application in many enterprises. In some organizations, a person only sends and receives a few e-mail messages a day, while in other organizations, an ordinary user can Send and receive many messages. Depending on the size of the unit, the email messaging system can handle hundreds or even thousands of messages every day. As the number and size of messages and attachments grow at a huge rate, and the key to the message library With the ever-increasing volume of business information, it is increasingly difficult to manage email servers. Excessive capacity of the email server will affect backup and recovery performance, and may lead to mission-critical information due to unintentional deletion or mail server failure The loss.
In some conventional email systems, the size of the message library can be controlled by certain thresholds, such as, for example, the limit on the number of messages that can be stored in a personal mailbox, and the cumulative size of messages that can be stored in the message library. and many more. These net values can be controlled by the system administrator, or in some cases, they can be "hard-coded" into the email messaging application. The problem with such values is that they are used to keep the message library within certain predetermined limits, and in fact do not provide any management capabilities to allow users to keep important messages for as long as they are needed.
Another method that has been used in the art to contain the size of the message library is to "archive" messages. Conventional message filing systems have been embedded in e-mail messaging applications. However, because such systems are typically dedicated software applications, email administrators may not have many options for how to archive and retrieve messages. Some systems may require system administrators to intervene when users need to retrieve archived messages. In other systems, "archive" is only to download the message to the user's local hard disk, and the user's local hard disk may not
02804805.9 Articles are easily accessed or searched to retrieve archived messages.
In those e-mail systems that do not include integrated archiving functionality, system administrators can implement manual archiving operations through the e-mail process. The backup process is typically designed to allow complete restoration of the message database (also known as the "post office") in the event of a catastrophic failure. However, this backup process typically does not provide many of the functionality desired for an archive system. For example, in some backup processes, the email administrator may have to restore the entire post office just to retrieve one or more messages from the individual users mailbox. An additional problem with the typical backup process is that email messages are based on the content of the message. Without full-text search capabilities, it is more difficult to determine whether a particular email message has been archived.
For more complex e-mail management, different units can have different e-mail archiving requirements. For example, a comprehensive "archiving plan" may be required, where the archiving process must be able to capture all messages in "real time" before the user has a chance to delete any messages. One way to perform a full archiving is to perform a full archive after the message is sent or They are intercepted when they are received and a copy of the message is placed in the archive. In this way, the message can be captured and archived before it is distributed to all recipients. Therefore, the archive file as a whole is only stored One copy of each archived message. This helps reduce the size of the archive file.
In other units, the companys strategy may not require full archiving. Instead, it may run the archiving process every week or in other cycles. This archiving process does not capture every message processed by the e-mail system, but only Capture those messages in the system that have not been deleted by the time the process is running. Unlike real-time filing systems, in periodic filing systems, messages are only captured after they have been distributed to various recipients. Third-party or external, periodic message filing systems basically read in the system All the messages stored in each mailbox are operated on. Each message read is then copied into the archive file. Because each mailbox is read independently of other mailboxes, the archive file created by this conventional filing system becomes unnecessarily large. Therefore, messages sent to multiple mailboxes will appear to be undesirable In the archive file. Although if the archiving system has accessed the internal structure of the message library, it is possible for the archiving system to archive only a single copy of each message. However, due to the proprietary nature of the e-mail system, this type of access is typically Land is not allowed.
Therefore, there is a need for a system and method for indexing unique email messages extracted from an email messaging system.
02804805.9 Summary of the invention The present invention provides a system and method for indexing unique e-mail messages extracted from an e-mail messaging system. This method includes the following steps: reading a message from a mailbox on the e-mail messaging system, wherein The message includes multiple message attributes. Examples of message attributes include the name of the sender, the submission time of the sender, the subject, etc. If the originating email messaging system is an external messaging system, the sender's name can be, for example, an email address, or if the email messaging system is a destination messaging system, it can be a canonical name. The submission time is preferably based on the submission time set by the originating mail messaging system, and may be expressed in microseconds, for example.
The invention then uses the message attributes to calculate a unique identifier or message tag, which preferably includes a string of data. For example, the sender's name and the sender's submission time can be used to calculate the message tag. If the message is unique, the message tag is stored in the index file associated with the message archive, that is, if the message tag is not unique, the message is not unique. In order to speed up the process of determining whether the message is unique, you can A hash algorithm is applied to the message tag in order to obtain a "signature" of a predetermined length of the message. Therefore, since the index record has a uniform length, the difference between the newly calculated message tag and the message tag that has been stored in the index file The comparison will be faster.
The present invention also includes an archiving system and method, in which only the only message is stored in the message archive.
Brief Description of the Drawings Fig. 1 is a schematic diagram illustrating a method for calculating a message label in a first embodiment of the present invention.
Fig. 2 is a schematic diagram illustrating a method for calculating a message tag in the second embodiment of the present invention.
Fig. 3 is a schematic diagram of an exemplary architecture of an embodiment of the present invention.
Fig. 4 is a flowchart of steps for archiving email messages according to an embodiment of the present invention.
Fig. 5 is a schematic diagram illustrating components of a uniqueness checking system according to an embodiment of the present invention. DETAILED DESCRIPTION OF THE INVENTION The present invention provides a system and method for indexing unique e-mail messages extracted from one or more e-mail messaging systems. The present invention also provides for archiving only the same e-mail
02804805. 9 System and method for unique multiple copies of the first message·The present invention uses an index file to store information about the message that has been previously extracted from the e-mail messaging system·The index file can be used to allow easy search and comparison of the The entries in the file are stored in any suitable format. For example, the index file can be a text file, an extension page, or a relational database table or table group. Whenever an e-mail message is added to the archive, the "message tag" is Generated and stored in the index file. The message tag is based on sufficient characteristics or attributes of the e-mail message to establish a unique identifier for each e-mail message.
The system and method of the present invention can be used in any application in an email messaging system where it is desired to identify duplicate messages. For example, an email archiving application can be advantageously incorporated into the system and method of the present invention to reduce or minimize the size of the archive message library. If the present invention is used in an archiving system, a temporary message tag is generated for the email message before the message is added to the archive. This temporary message tag is then compared with each message tag that has been stored in the index file. If the temporary message tag matches an existing entry in the index file, the e-mail message has already been archived. If this is the case, it is not necessary to add the message to the archive.
The following sections describe two embodiments of the invention. Each embodiment uses a different method to generate (or calculate) the message tag of the e-mail message.
The first embodiment of the present invention will be described with reference to FIG. 1. In this embodiment, the message tag can be calculated by concatenating selected message attributes to form a single text string. For example, if the email messaging system is a Microsoft Exchange system, the message may include attributes such as PR-Client-Submit-Time in box 10. PR-Sent-Representing-Emai 1-Address in box 12 and PR-Subject. Boxes 16, 18, and 20 in box 14 show the corresponding data type associated with each of these attributes. Boxes 22, 24, and 26 show examples of actual values that these attributes can have for a particular message. For example, the value of PR-Client-Submit_Time in box 10 is displayed as 0x01cl9el38106580 in box 22. The submission time in this example represents the time when the message was submitted by the sender of the message. The format of the time is as generated by the system clock on the senders email messaging server. The format of the submission time is not correct, as long as the access format is standardized for each server. That is, for all messages received from a specific server, the same time format should be used to calculate the message tag. Box 24 contains /o = sqa/ou = dogwood/cn=Recipients/cn = Crowen,
02804805.9 The first is the value of the exchange attribute PR-Sent-Email-Address in box 12. This attribute is commonly referred to in the art as the "fully qualified name" of the sender. The message tag generated based on the senders submission time and the senders fully qualified bird will be sufficient to uniquely identify most email messages. These values are connected (as illustrated by link 30) to generate a message tag 40.
As described above, using the time of submission and the name of the sender is usually sufficient to uniquely identify an email message. However, in order to increase the likelihood that the message tag represents a unique message, other attributes can be added to the string. For example, as shown in Figure 1, the PR-Subject attribute in box 14 can be included. In this example, the value of this attribute is "This is a test message", as shown in box 26. In link 32, all three attributes are connected to form a message tag 42.
The above method for generating message tags can be modified in many ways without departing from the spirit of the present invention. For example, the connection order can be changed so that the resulting message tag is constructed by connecting the submission time string to the sender's name string. Optionally, the subject may be before the sender's name, or the submission time, etc. In another variation, the sender's name may include other attributes that identify the sender of the email message. For example, the sender's name may be represented as an Internet email name, such as JDoebacme.com. This value is then used as described above. Moreover, the message tag can be generated based on other message attributes (such as message size, header information, etc.) without using any sender's information.
The message tag generated according to this embodiment will have varying lengths. That is, the length of the message tag of the first message extracted from the email messaging system may be different from the length of the message tag of the second message extracted from the email messaging system. Specifically, the reason for this is because the sender's name and the email message subject field can have different lengths. Moreover, different email messaging systems can use different implementations to calculate the submission time. Due to the variable length of the message tag, if the index file is large, searching the index file carefully may be an excessively long operation. The second embodiment is described below, which provides enhanced message marking for optimizing this search.
Second Embodiment In the second embodiment, a variable-length message tag is converted into a message tag having a predetermined length by applying a hash algorithm. In the field of cryptography, hashing algorithms are often used to generate keys for encrypting messages. They are also used to generate electronic signatures of messages. Electronic signatures can be used to verify the integrity of messages. Such signatures are also known as fingerprints of messages.
02804805.9 "Digest Information Digest". One principle that supports this hashing algorithm is: applying the algorithm to two different messages and getting the same result is "computationally infeasible". Another principle of the hashing algorithm is : As a result, the length of the message digest will be unified. It is this second principle that is useful in the environment of the present invention. That is, if the different message tags generated as described above are run through the hash algorithm, then The resulting message tag will have a uniform length and also represent the only e-mail message.
Fig. 2 is a schematic diagram illustrating the operation of the second embodiment of the present invention. The items numbered 10-42 are the same as those described above in relation to Fig. 1. The message tag 42 is generated by concatenating selected attributes to form a variable-length string (such as the string described with reference to FIG. 2). This string is then used as an input to the hash algorithm 50. In this example, the output of the hash algorithm 50 is a 64-bit number, which is represented as a hexadecimal string: "0x4764e0ccl21642b5", shown in box 60. As is known in the art, such a string It finally represents a group of 64 bits (multiple "1"s and "0"s), which can be converted into many different representations.
By generating message tags with a uniform length, the performance of search and comparison operations on index files can be greatly improved. In the preferred embodiment, the well-known "MD5" hashing algorithm is used. The MD5 hash algorithm is defined in RFC1321 at www. faqs. org/rfcl321. html, which is hereby introduced for reference. The message token generated using the MD5 hash algorithm will have a uniform length of 128 bits (that is, (if converted to ASCII characters) 16 characters or 32 hexadecimal numbers).
Architecture Figure 3 shows an architecture that can be used to implement embodiments of the present invention. The enterprise e-mail messaging system 30D includes an e-mail server 301 that provides customers 302 and 304 with e-mail services. The email messaging system 300 may be a Microsoft Exchange server, and the communication between the archive server 330 and the email messaging server 300 may be processed through a well-known messaging application programming interface (MAPI) protocol. As is well known in the art, MAPI is a messaging architecture and a client interface component. As a messaging architecture, MAPI enables multiple applications to interact with multiple messaging systems across various hardware platforms. The customer interface component, MAPI is a complete set of functions and object-oriented interfaces, which form the basis of the customer application and service provider interface of the MAPI subsystem. Compared with simple MAPI, Common Messaging Call (CMC) and CD0 library, MAPI provides the highest performance and the greatest control for messaging-based applications and service providers.
02804805.9 No. system.
Alternatively, the email messaging system 300 may be a Lotus Notes mail server and the communication may be handled through the Lotus Notes application programming interface (API) protocol. Similarly, if the email messaging system is a Simple Mail Transfer Protocol (SMTP) mail server , The communication can be processed via SMTP.
In the example shown in FIG. 3, the communication links 306 and 308 can use MAPI, SMTP, or some other protocol, depending on the capabilities of the client systems 302 and 304. E-mails can be received from the external system 320 via the Internet 322 on the communication link 321 via SMTP. In an embodiment of the present invention, the archive server 330 initiates an archive session between the communication link 332 and the email server 301 on a periodic basis. This periodic basis can be, for example, daily, weekly, monthly or Some other suitable time interval depends on the archiving needs of the enterprise. The communication link 332 may use any suitable network protocol, for example, the well-known Transmission Control/Internet Protocol (TCP/IP). In another embodiment of the present invention, the archive server 330 retrieves emails in real time or near real time.
As known in the art, the email messaging server 301 may include multiple mailboxes, directories, folders, or other "storage boxes" for associating messages with individual users. As used herein, the term "Mailbox" means a message group associated with a specific user, and where applicable, it includes any subfolder or directory created by the user to organize his email messages. In some embodiments, the mailbox may include an "inbox" for storing newly arrived email messages and an "outbox" for storing messages sent by the user.
In an embodiment where the archive server 330 extracts messages on a periodic basis, the archive server 330 reads each message in each mailbox on the email server 301. In another embodiment, the archive server 330 may be configured to only Read new messages that have been established and submitted since the last cycle of the session was completed (or started). In another embodiment, the file server 330 may be configured to only read in the inbox and outbox of the mailbox News. Regardless of the message reading scheme implemented, the archive server checks the index file to determine the uniqueness of the message.
The "uniqueness check" function can be integrated into the archive server 330 or executed on a different server. In either case, the uniqueness check function includes the calculation of the message tag, as described above. The message tag of the newly read message is compared with the index file on the database 334. The index file includes a list of message tags corresponding to all messages stored in the message archive on the database 334. If the calculated message tag matches the
02804805. 9 entry in the cited file, the message is not unique. That is, the message has been stored in the message archive and does not need to be stored a second time. Otherwise, if the fogged message is marked with If any record in the quotation file does not match, the message is unique and should be stored in the message archive. If so, the message tag is also added to the index file.
Once the message has been archived in the archive server 330, the data can be moved to other storage media without affecting the performance of the email server 301. For example, the data can be moved to a tape library system 335, an optical disc drive 336, a CD/DVD optical device 337, and so on. By moving the archived data to such storage media, the unit may reduce its long-term storage costs because these media are not as expensive as other magnetic storage media.
Fig. 4 is a flowchart illustrating the steps for archiving email messages in an embodiment of the present invention. Steps 400-406 are initialization steps and are shown for clarity. That is, once a message file and index file is provided, the process executes steps 408-420. In step 400, the first message is read from the mailbox of the email messaging server. In step 402, A message tag is calculated for the first message, and in step 404, the first message is stored in the message archive. In step 406, the message tag calculated for the first message is stored in the index file. In step 408, the second (or next) message is read from the mailbox on the email messaging server. The mailbox may be the same mailbox where the first message was read or may be a different mailbox. In step 410, the message tag of the second message is calculated, and in step 412, the second message tag is compared with the first message tag (ie, the second message tag is compared with any information that has been stored in the index file). Message tag to compare).
In step 414, the process branches, depending on the result of step 412. If the second message tag matches the first message tag (ie, if the second message tag is already in the index file), then the second message Is not unique, and the process moves to step 420. If the message is unique (ie, the message tag does not match any entry in the index file), then in step 416, the second message is stored Into the message archive, and in step 418, store the second message tag in the index file. In step 420, the process checks to see if there are more to be removed from the email messaging server. Read the message. If there are more messages, the process returns to step 408 to read the next message. Otherwise, if there are no more messages, the process ends. Figure 5 is a schematic diagram showing how the message tag is calculated in the second embodiment of the present invention
02804805.9 Figure. In FIG. 5, the email message attribute 500 is selected from the email message. As described herein, the combination of the sender's name and submission time may be sufficient to uniquely identify an email message in most applications. The selected attributes are combined to form a single string. The string may or may not include spaces. At block 502, the string is converted into a suitable bit representation. At block 504, a hash algorithm is applied to the bit string to determine the message flag at block 506.
As described herein, the present system and method of archiving and retrieving e-mail messages can be used in large-scale enterprise environments using dedicated archiving servers and database systems such as SQL or ORACLE. Optionally, the archive server can run on the same platform as the e-mail messaging server. As described above, the e-mail messaging server can be based on any suitable e-mail messaging protocol, for example, Microsoft OUTLOOK TM. Lotus NOTESTM. or proprietary or non-proprietary e-mail messaging system.
Embodiments Including Application Programs Embodiments of the present invention also include application programs that are themselves recorded on any magnetic or electrical media, and computer systems programmed using the programs. In this embodiment, the computer system programmed in this way is It is configured to traverse the mailboxes on the email messaging server to identify the messages to be added to the archive. Before the program of the present invention is executed, such a program can operate to process messages delivered to the e-mail messaging system. In this way, the program identifies and extracts existing e-mail messages for archiving. The program can also be configured to archive messages in real time, that is, when a message is processed by the email messaging system, a copy is retrieved by the archive server for archiving processing.
Embodiments of the present invention may include an embedded relational database to support high-speed searching of message metadata. In this embodiment, the key words or full text of the message are added to the message index file to quickly search for the message. In addition, some attachment content can be added to the message index. For example, attachments based on public word processing applications can be read by the archive server to enable full text search of these attachments.
The present invention provides a comprehensive solution for externally archiving e-mail messages from an e-mail messaging system. The present invention can be used by units responsible for maintaining e-mail messages for an extended period of time. For example, in some financial institutions, the Federal Secrecy and Exchange Commission (SEC) has required that all records, including e-mail messages, must be archived for a period of 5 years. These records must be stored in such a way that each record can be retrieved upon request. By storing email messages together with full-text search capability messages in an external archive
02804805.9 First, the implementation of the present invention can solve these and other needs. Moreover, by checking for duplicate messages, the size of the archive message library can be kept at a manageable level.
The foregoing disclosure of the preferred embodiment of the present invention is presented for the purpose of illustration and description. Its purpose is not to exhaust the invention nor to limit the invention to the exact form disclosed. Based on the above disclosure, many changes and modifications of the embodiments described herein are obvious to those skilled in the art The scope of the present invention is only limited by the appended claims and their equivalents.
Moreover, in describing the representative embodiments of the present invention, the method and/or process of the present invention may have been shown as steps in a specific order. However, to a certain extent, the method or process does not depend on the steps in the specific order set forth herein, and the method or process should not be limited to the steps in the specific order described. As those skilled in the art should understand, other sequences of steps are also possible. Therefore, the specific sequence of steps set forth in should not be construed as a limitation on the claims. In addition, the claims for the method and/or process of the present invention should not be limited to their steps being executed in the order described And those skilled in the art will easily understand that the order can be changed and still remain within the spirit and scope of the present invention.
02804805.9
Every citation, both ways
| Document | Relation | Office | Cited during |
|---|---|---|---|
| EP1035690A2 | Cites | European Patent Office (EPO) | Search report |
| CN1235305A | Cites | China | Search report |
| CN1250288A | Cites | China | Search report |
| US5742807A | Cites | United States of America | Search report |
| US5999932A | Cites | United States of America | Search report |
| US5999967A | Cites | United States of America | Search report |
| US6009442A | Cites | United States of America | Search report |
12 members in 7 offices
Priority claims10
| Document | Office | Kind | Date |
|---|---|---|---|
| 26809201 | United States of America | P | |
| 26809201 | United States of America | P | |
| 60268092 | United States of America | – | |
| 34723802 | United States of America | P | |
| 34723802 | United States of America | P | |
| 60347238 | United States of America | – | |
| 60268092 | – | – | – |
| 60347238 | – | – | – |
| US20010268092P | – | – | – |
| US20020347238P | – | – | – |
Members12
| Document | Office | Kind | |
|---|---|---|---|
| CA2433525A1 | Canada | A1 | |
| WO02065316A1 | World Intellectual Property Organization (WIPO) | A1 | |
| US2002122543A1 | United States of America | A1 | |
| WO02065316A9 | World Intellectual Property Organization (WIPO) | A9 | |
| EP1368739A1 | European Patent Office (EPO) | A1 | |
| KR20040007435A | Republic of Korea | A | |
| CN1531688A | China | A | |
| JP2005501308A | Japan | A | |
| CN1316397CThis record | China | C | |
| EP1368739A4 | European Patent Office (EPO) | A4 | |
| CN101030275A | China | A | |
| CN101030275B | China | B |
5 legal events, as the office reported them to INPADOC
Over the term
Point at a mark for the eventEvents
| Event | Code | |
|---|---|---|
| Grant of patent or utility modelGrantedC14 | C14 | |
| Entry into substantive examinationC10 | C10 | |
| Succession or assignment of patent rightASS | ASS | |
| Transfer of patent application or patent right or utility modelC41 | C41 | |
| PublicationC06 | C06 |
Numbers
- Publication
- 1316397
- Publication, DOCDB
- 1316397
- Publication, EPODOC
- CN1316397C
- Application
- 28048059
- Application, DOCDB
- 02804805
- Application, EPODOC
- CN2002804805
Titles2
- Chinese
- 索引惟一电子邮件消息及其使用的系统和方法
- English
- System and method for indexing unique e-mail message and its use
Classification
- CPC, 4
- G06Q10/107
- H04L51/42
- G06F15/16
- G06F16/2272
- IPC, 6
- G06F15 16
- G06F17 30
- G06Q10 00
- G06Q10 107
- G06Q10 109
- H04L12 58