Maintenance of link level consistency between database and file system
Abstract
This record has no abstract on file.
Term
Term ended
Expired 9 March 2026, 0.5 years ago.
- Priority
- Filed
- Granted
- Expired
- Today
10 claims: 4 independent, 6 dependent
- 1A method of maintaining link-level integrity of transactions between a database and a file system, which is performed by the computer associated with the database and file system and receives a file system change request for a binary large object. A step and a database log of the database, wherein the reference to the binary large object is stored in a cell of the database, and the binary large object is stored as a file of the file system outside the database. ofMaximum log sequence numberAnd the file system log of the file systemMaximum log sequence numberAnd the aboveWhen the maximum log sequence number of the database log is equal to or greater than the maximum log sequence number of the file system log.The computer associated with the database logs the file system changes associated with the file system change request to the database log, and the computer associated with the file system logs the file system changes to the file of the file system log file. A step of logging to a name, wherein the file system log file is stored as a file of the file system, and information representing the file system change is encoded in the file name of the file system log file. , A method characterized in that the computer associated with the file system comprises performing the file system modification on the binary large object. データベースとファイルシステムとの間でのトランザクションのリンクレベル整合性を維持する方法であって、前記方法は、データベースおよびファイルシステムに関連付けられるコンピュータによって実行され、 バイナリラージオブジェクトについてのファイルシステム変更要求を受信するステップであって、前記バイナリラージオブジェクトへの参照は、データベースのセルに格納され、前記バイナリラージオブジェクトは、前記データベースの外部にファイルシステムのファイルとして格納される、ステップと、 前記データベースのデータベースログの最大ログシーケンス番号と前記ファイルシステムのファイルシステムログの最大ログシーケンス番号を比較するステップと、 前記データベースログの最大ログシーケンス番号が前記ファイルシステムログの最大ログシーケンス番号以上である場合、前記データベースに関連付けられるコンピュータが、前記ファイルシステム変更要求に関連付けられるファイルシステム変更を前記データベースログにロギングするステップと、 前記ファイルシステムに関連付けられるコンピュータが、前記ファイルシステム変更をファイルシステムログファイルのファイル名にロギングするステップであって、前記ファイルシステムログファイルは、前記ファイルシステムのファイルとして格納され、前記ファイルシステム変更を表す情報は、前記ファイルシステムログファイルのファイル名に符号化される、ステップと、 前記ファイルシステムに関連付けられるコンピュータが、前記バイナリラージオブジェクトについての前記ファイルシステム変更を実行するステップと を含むことを特徴とする方法。
- 4The claim is characterized in that the step of logging the file system change to the file system log file further includes a step of encoding the information identifying the log sequence number of the database log into the file name of the file system log file. The method described in 1. 前記ファイルシステム変更をファイルシステムログファイルにロギングするステップは、前記データベースログのログシーケンス番号を識別する情報を前記ファイルシステムログファイルのファイル名に符号化するステップをさらに含むことを特徴とする請求項1に記載の方法。
- 7A computer system configured to maintain link-level consistency of transactions between a database and a file system, said computer system comprising a processor coupled to a computer-readable storage medium. A readable storage medium stores computer-executable instructions, and when the computer-executable instructions are executed by the processor, the computer system stores a reference to a binary large object in a cell of the database. And, the step of storing the binary large object as a file of the file system outside the database and the failure of the change to the binary large object are stored in the database log.maximumLog sequence number and stored in file system logmaximumIt is a step of determining by comparing with the log sequence number, and is described above.Maximum stored in the file system logThe log sequence number is encoded by the file name of the file system log file stored in the file system, the database log is stored in the database, and the file system log is stored as a file in the file system. , And a computer system characterized in that the file system logs changes in the file system and the file system executes changes to the binary large object to perform the steps of redoing the changes. .. データベースとファイルシステムとの間でのトランザクションのリンクレベル整合性を維持するように構成されたコンピュータシステムであって、 前記コンピュータシステムは、コンピュータ読み取り可能な記憶媒体に結合されたプロセッサを備え、 前記コンピュータ読み取り可能な記憶媒体は、コンピュータ実行可能命令を格納し、該コンピュータ実行可能命令は、前記プロセッサによって実行されるときに、前記コンピュータシステムに、 バイナリラージオブジェクトへの参照をデータベースのセルに格納するステップと、 前記バイナリラージオブジェクトを前記データベースの外部にファイルシステムのファイルとして格納するステップと、 前記バイナリラージオブジェクトに対する変更が失敗したことを、データベースログに格納される最大ログシーケンス番号とファイルシステムログに格納される最大ログシーケンス番号とを比較することによって判定するステップであって、前記ファイルシステムログに格納される最大ログシーケンス番号は、前記ファイルシステムに格納されるファイルシステムログファイルのファイル名に符号化され、前記データベースログは、前記データベースに格納され、前記ファイルシステムログは、前記ファイルシステムにファイルとして格納される、ステップと、 前記ファイルシステムによって前記ファイルシステムにおける変更をロギングし、前記ファイルシステムによって前記バイナリラージオブジェクトについての変更を実行することによって、前記変更をやり直すステップと を実行させることを特徴とするコンピュータシステム。
- 10A computer-readable recording medium having computer-executable instructions that perform a method of maintaining link-level integrity of transactions between a database and a file system, said computer-executable instructions being executed by a computer. When, on the computer, the database log of the databaseMaximum log sequence numberAnd the file system log of the file systemMaximum log sequence numberAnd the aboveWhen the maximum log sequence number of the database log is equal to or greater than the maximum log sequence number of the file system log., A step of logging changes to a binary large object into a record in a database log, where a reference to the binary large object is stored in a cell of the database and the binary large object is stored outside the database in a file system. Stored as a file, the database log is stored in the database, in steps, by logging file system changes to the binary large object to the file system log file by the file system, and in the file system changes. A computer comprising performing a method comprising encoding a corresponding file system object name, log sequence number, transaction identifier, and information identifying an operation descriptor into the file name of the file system log file. A readable recording medium. データベースとファイルシステムとの間でのトランザクションのリンクレベル整合性を維持する方法を実行するコンピュータ実行可能命令を有するコンピュータ読み取り可能な記録媒体であって、 前記コンピュータ実行可能命令は、コンピュータによって実行されるとき、該コンピュータに、 前記データベースのデータベースログの最大ログシーケンス番号と前記ファイルシステムのファイルシステムログの最大ログシーケンス番号を比較するステップと、 前記データベースログの最大ログシーケンス番号が前記ファイルシステムログの最大ログシーケンス番号以上である場合、バイナリラージオブジェクトに対する変更をデータベースログのレコードにロギングするステップであって、前記バイナリラージオブジェクトへの参照は、前記データベースのセルに格納され、前記バイナリラージオブジェクトは、前記データベースの外部にファイルシステムのファイルとして格納され、前記データベースログは、前記データベースに格納される、ステップと、 前記ファイルシステムによって、前記バイナリラージオブジェクトへのファイルシステム変更をファイルシステムログファイルにロギングするステップと、 前記ファイルシステム変更に対応する、ファイルシステムオブジェクト名、ログシーケンス番号、トランザクション識別子、およびオペレーション記述子を識別する情報を前記ファイルシステムログファイルのファイル名に符号化するステップと を含む方法を実行させることを特徴とするコンピュータ読み取り可能な記録媒体。
Independent claims4
61 paragraphs, as filed
The present invention generally relates to the field of database management. More specifically, the present invention relates to maintaining link-level integrity between a database and a corresponding file system.
In recent years, the use of large-scale unstructured data types such as photographs, videos, and movies has increased significantly, and along with this, there is a need to efficiently store these large-scale data streams. Traditionally, data has been stored in structures such as file systems or databases.
A file system is a hierarchical structure of data in which files are stored in folders on a storage medium such as a hard disk. Some operating systems maintain the file system and control access to files within the file system. File systems are great for streaming large amounts of unstructured data in and out of files. One of the problems with the currently known file system is that you have to manually group files (folders and subfolders), and if a user forgets where they stored a particular file, the file It will be difficult to find it again. This problem is exacerbated by technological advances, such as the breakthroughs in disk technology seen in the development of increasingly large hard disks. The ability to store huge amounts of data on a single disk can make tracking files within a file system an extremely difficult task.
Another widely used data organization method is the database. The database system stores data as one or more tables, where each row of the table has a group of related data elements about the entity and columns representing useful pieces of information about the entity that is the subject of the row. For example, maintain a database of talent information, where each row in the talent database represents an employee and each column in the talent database represents data elements such as employee name, employee social security number, and employee wage rate. can do.
The database provides some useful advantages through the file system organization of data. Database management systems are good at storing, discovering, and retrieving small pieces of structured data. More typically, there are highly flexible means of retrieving and accessing specified parts of the data stored in the database. However, databases do not specifically address the storage and access of large pieces of unstructured data called BLOBs (Binary Large Objects).
Specifically, if the database contains BLOB columns, the BLOBs are typically broken into small pieces that are distributed across the disk. The entry in the database column does not contain the BLOB itself, but a pointer to the first piece of the BLOB. This situation leads to inefficiencies in retrieving the data in the BLOB, as different pieces of the BLOB must be found and reassembled. Typically, to reduce the effects of these inefficiencies, a pointer to the first piece of the blob will return instead of immediately retrieving the blob itself.
For example, suppose your database of employee information contains a BLOB column for employee photos. Suppose a user requests a photo of a particular employee and returns a pointer to that photo. This pointer represents the physical position, for example, a 16-byte hexadecimal value that represents the actual disk address of the sector of the disk that contains the first piece of the photo. In this situation, some problems can arise. In addition to the disk address being confusing to the user, if the operating system reorganized the data on the disk, this photo may no longer be there, in this case "not found". Message will be returned to the user.
In recent years, another method of storing BLOBs has been developed, in which the BLOBs are stored in the file system as a continuous file or "FILESTREAM". Provides the FILESTREAM data storage attribute that can be used when tagging columns in the relation table. The FILESTREAM attribute indicates that the data for that column will be stored as a file in the operating system (OS) file system. The database management system manages the creation and deletion of files in the file system. There is a 1: 1 reference between files and cells (at the intersection of rows and columns) in the file system. The data in the FILESTREAM column can be manipulated in the same way as the data in other columns that use SQL or a programming language such as MICROSOFT® T-SQL.
<patcit num="1"><text>U.S. Patent Application Publication No. 2004/148308</text></patcit><patcit num="2"><text>U.S. Patent Application Publication No. 2004/148272</text></patcit>
<p num="0010"> Therefore, the FILESTREAM column is used in databases for large unstructured data. The use of the FILESTREAM data storage attribute allows large unstructured data to be stored in the file system as contiguous files while still allowing access to the database. These database management systems ensure link integrity (ie, "link level") between a row in a database that has the FILESTREAM attribute and its corresponding FILESTREAM data to ensure data integrity and to avoid database corruption. Consistency ") must be maintained. For example, if a failure such as a power failure or system crash occurs (ie, flash) before the time the changes are committed to the disk, some problems can result. For example, the database may not reflect the existence of files or directories that exist in the file system, or the database may reflect the existence of files or directories that do not exist in the file system. In this way, if the link between the database and the file is broken, the database user guarantees that the database accurately reflects the state of the current data represented by the FILESTREAM cell in the database column. The database integrity is compromised because it is not.</p><p num="0011"> Maintaining link-level integrity has typically been achieved by two different techniques: integrity checking and repair, and logging and recovery. In integrity checking and repair, the crawling task searches the database and file system for inconsistencies and potentially repairs them. These techniques are time consuming, untargeted, and consume excessive system resources.</p><p num="0012"> Traditional logging and recovery methods use logging in database logs or Transacted File. It can be adjusted using System). In the former method, file system operations are logged in the database log along with database data updates. With these techniques, if the database management system recovers the database, redo and undo (REDO) and undo (REDO) and undo (REDO) and undo (REDO) and undo (REDO) and undo (REDO) for logged file system operations and to align the file system data with the database data in the same database recovery framework. UNDO) Can trigger operations. The disadvantage of these techniques is that the database management system is usually not tightly integrated with the file system, so there is no knowledge of data flash on the file system's disks. Due to the lack of ability to coordinate data flash, the database management system uses files to achieve proper write-ahead-logging, which helps maintain link-level integrity of transactions during crash recovery. Log to log for each system operation record) must be flushed. This flushing of the logs results in one disk I / O operation per file system operation, which is usually an unacceptable amount of performance overhead.</p><p num="0013"> The latter method involves coordination using a transacted file system, where the file system itself is a transaction and can be recovered. The database management system is involved in distributed transactions coordinated by a good transaction manager (TM). During crash recovery, a good TM resolves suspicious transactions and ensures integrity between the database and the file system resource manager. The disadvantage of this approach is that the transactiond file system cannot be used with off-the-shelf operating systems. Therefore, this technique is not available for many database management systems on many OS platforms. In addition, this method has the disadvantage of adding complexity and performance costs when implemented.</p><p num="0014"> Therefore, there is a need for a mechanism to maintain link-level consistency between database columns and their corresponding FILESTREAM data in the file system while addressing the shortcomings mentioned above.</p>
<p num="0015"> In view of the disadvantages and disadvantages mentioned above, this specification discloses methods and computer-readable media for maintaining link-level integrity of transactions between a database and a file system. One method is to log the file system changes to the database log record and create a file in the file system folder that corresponds to the file system changes. When the recovery process restarts, analysis and conditional redo operations are performed based on the database logs, and conditional redo and undo operations are performed based on the files in the file system folder. The undo operation is then performed based on the database log.</p>
The subject matter of the present invention will be described with specificity for meeting legal requirements. However, this description itself is not intended to limit the scope of the claims. Rather, the inventors, etc., such that the subject matter described in the claims, along with other current or future techniques, include various steps or elements similar to those described herein. I have intended that it can be embodied by the method. Further herein, the term "step" can be used to imply various aspects of the method adopted, unless the order of the individual steps is explicitly stated. , And, except where noted, should not be construed as suggesting any particular order during or between the various steps disclosed herein.
Computing environment example FIG. 1 shows an example of a suitable computing system environment 100 in which the present invention can be implemented. The computing system environment 100 is merely an example of a suitable computing environment and is not intended to imply any limitation with respect to the scope of use or functionality of the present invention. Furthermore, the computing environment 100 should not be construed as having any dependency or requirement for any one or a combination of the components shown in the example of the operating environment 100.
The present invention is operational in the environment or configuration of many other general purpose or application computing systems. Examples of well-known computing systems, environments, and / or configurations suitable for use in the present invention include personal computers, server computers, handheld or laptop devices, multiprocessor systems, microprocessor-based systems, and more. Includes, but is not limited to, set-top boxes, programmable home appliances, networked PCs, minicomputers, mainframe computers, distributed computing environments including any of the systems or devices mentioned above, and more. ..
The present invention can be described in the general context of computer executable instructions such as program modules executed by a computer. In general, a program module includes routines, programs, objects, components, data structures, etc. that perform a particular task or perform a particular abstract data type. Typically, the functionality of the program module can be combined or distributed as desired in various embodiments. The present invention can also be implemented in a distributed computing environment where tasks are performed by remote processing devices linked over a communication network. In a distributed computing environment, program modules can be located on both local and remote computer storage media, including memory storage devices.
Referring to FIG. 1, examples of systems for carrying out the present invention include general purpose computing devices in the form of computer 110. The components of the computer 110 can include, but are not limited to, a processing unit 120, a system memory 130, and a system bus 121 that connects various system components including the system memory to the processing unit 120. Absent. The system bus 121 can be any of several types of bus structures, including memory buses or memory controllers, peripheral buses, and local buses that use any of the various bus architectures. For example, these architectures include Industry Standard Architecture (ISA) bus, Micro Channel Architecture (MCA) bus, Extended ISA (EISA) bus, and Video Electronics Standards Association (VESA). It includes, but is not limited to, the Association) local bus and the PCI (Peripheral Component Interconnect) bus, also known as the mezzanine bus.
The computer 110 typically includes various computer readable media. The computer-readable medium can be any usable medium accessible to the computer 110, including both volatile and non-volatile media, removable and non-removable media. For example, computer readable media can include, but are not limited to, computer storage media and communication media. Computer storage media are volatile and non-volatile media, removable and removable, performed by any method or technique for storing information such as computer-readable instructions, data structures, program modules, or other data. Includes both non-volatile media. Computer storage media can be RAM, ROM, EEPROM, flash memory or other memory technology, CD-ROM, digital versatile disk (DVD) or other optical disk storage, magnetic cassette, magnetic tape, magnetic disk storage, or other magnetic. It includes, but is not limited to, a storage device, or any other medium that can be used to store desired information and is accessible to the computer 110. Communication media typically embody computer-readable instructions, data structures, program modules, or other data within a modulated data signal such as a carrier wave or other transport mechanism, and include any information delivery medium. The term "modulated data signal" means a signal having one or more of its features set or modified in such a way as to encode the information in the signal. For example, communication media include, but are not limited to, wired media such as wired networks or direct wired connections, as well as wireless media such as acoustic, RF, infrared, and other wireless media. Any combination of the above shall also be included in the range of computer readable media.
System memory 130 includes computer storage media in the form of volatile and / or non-volatile memory, such as read-only memory (ROM) 131 and random access memory (RAM) 132. The basic input / output system 133 (BIOS), which contains basic routines that help transfer information between elements in the computer 110, such as at startup, is typically stored in ROM 131. RAM 132 typically contains data and / or program modules that are immediately accessible and / or currently in operation by the processing unit 120. For example, FIG. 1 shows an operating system 134, an application program 135, other program modules 136, and program data 137.
Computer 110 can also include other removable / non-removable, volatile / non-volatile computer storage media. To give a simple example, FIG. 1 shows a hard disk drive 141 that reads or writes to a non-removable non-volatile magnetic medium and a magnetic disk drive 151 that reads or writes to a removable non-volatile magnetic disk 152. And an optical disk drive 155 that reads or writes to a removable non-volatile optical disk 156 such as a CD-ROM or other optical medium. Other removable / non-removable, volatile / non-volatile computer storage media available in the operating environment example are magnetic tape cassettes, flash memory cards, digital versatile discs, digital videotapes, solid state RAM, solid state ROM, Etc., but are not limited to these. The hard disk drive 141 is typically connected to system bus 121 via a removable memory interface such as interface 140, and the magnetic disk drive 151 and optical disk drive 155 are typically connected to system bus 121 via a removable memory interface such as interface 150. Connected to.
The drives and their associated computer storage media, as described above and shown in FIG. 1, provide storage for computer-readable instructions, data structures, program modules, and other data about the computer 110. For example, in FIG. 1, hard disk drive 141 is shown as storing operating system 144, application program 145, other program modules 146, and program data 147. Note that these components may be the same as or different from the operating system 134, application program 135, other program modules 136, and program data 137. The operating system 144, the application program 145, the other program modules 146, and the program data 147 are given different numbers here to indicate that they are at least different copies. The user can enter commands and information into the computer 110 via an input device such as a keyboard 162 and a pointing device 161 commonly referred to as a mouse, trackball, or touchpad. Other input devices (not shown) may include microphones, joysticks, gamepads, satellite dishes, scanners, and the like. These and other input devices are often connected to the processing unit 120 via a user input interface 160 coupled to the system bus, but other, such as a parallel port, game port, or universal serial bus (USB). It is also possible to connect by interface and bus structure. Monitor 191 or other types of display devices are also connected to system bus 121 via an interface such as video interface 190. In addition to the monitor, the computer can be connected via the output peripheral interface 195,
Computer 110 can operate in a networked environment that uses a logical connection to one or more remote computers, such as remote computer 180. The remote computer 180 can be a personal computer, server, router, network PC, peer device, or other common network node, and although only memory storage device 181 is shown in Figure 1, usually Includes many or all of the elements mentioned above for computer 110. The logical connections shown in Figure 1 include local area networks (LAN) 171 and wide area networks (WAN) 173, but other networks can also be included. These networking environments are common in offices, enterprise-scale computer networks, intranets, and the Internet.
When used in a LAN networking environment, computer 110 is connected to LAN 171 via a network interface or adapter 170. When used in a WAN networking environment, computer 110 typically includes a modem 172 or other means for establishing communication over WAN 173, such as the Internet. The modem 172, which can be internal or external, can be connected to the system bus 121 via user input interface 160 or other suitable mechanism. In a networked environment, the program modules shown for computer 110 or parts thereof can be stored in remote memory storage devices. For example, FIG. 1 shows, but is not limited to, the remote application program 185 as resident on memory device 181. It will be appreciated that the network connections illustrated are exemplary and other means of establishing communication links between computers are available.
Examples of distributed computing frameworks or architectures In light of the concentration of personal computing and the Internet, various distributed computing frameworks have been developed and are currently under development. Individual and business users will also be provided with a seamlessly interoperable web-executable interface for applications and computing devices that will make computing activities increasingly web browser or network oriented.
For example, Microsoft's .NET platform includes servers, building block services such as web-based data storage, and downloadable device software. Generally speaking, the .NET platform has (1) the ability to collaborate the entire range of computing devices, and the ability to automatically update and synchronize user information in them, and (2) XML instead of HTML. Increased interaction with websites that can be performed with more use, (3) various applications such as email or Office Online services featuring customized access and delivery of products and services from a central origin to users for the management of software such as .NET, (4) Improve the efficiency and ease of access to information It will be centralized data storage, as well as the ability to synchronize information between users and devices, (5) the ability to integrate various communication media such as email, fax, and telephone, (6) thereby productivity. Provides the ability for developers to create reusable modules, as well as (7) many other cross-platform integration features, which improve and reduce the number of programming errors.
Although examples of embodiments are described herein in the context of software residing on a computing device, one or more parts of the invention are serviced by all .NET languages and services. Through the operating system, API, or middleware software between the coprocessor and the request object to be executed, supported by, or accessed through them, and even in other distributed computing frameworks. Similarly, it can be carried out.
Database environment In the discussion below, one of ordinary skill in the art is assumed to be familiar with the implementation of database-related FILESTREAM data storage attributes. These situations are discussed in US patent applications (Patent Document 1 and Patent Document 2), which are incorporated herein by reference in their entirety. Therefore, the details of these situations are omitted herein for the sake of clarity.
However, in order to provide additional background information and to include consideration of the following embodiments in the present specification, FIG. 2 shows an example of a database configuration in which aspects of the present invention can be implemented. Then, referring to Figure 2, (1) the client machine 280, which is the source of the query against the database 210, and (2) the database server 200, which hosts the database 210, to locate and manipulate the relevant information. (3) Three of the FILESTREAM server 240, which manages file system volume 225 capable of storing FILESTREAM values (or files) associated with a given database 210 and supports out-of-band updates to database 210. Three machines are feasible to play different roles. It will be understood that the FILESTREAM file can accommodate BLOBs stored in database columns. Each database can have one or more FILESTREAM groups 215, and each FILESTREAM group 215 can contain one or more volumes 225. This volume can reside on the FILESTREAM server 240. In the database installation example, there can be thousands of client machines, dozens of FILESTREAM servers, and one database server. A special configuration occurs when a volume containing FILESTREAM data is placed with the database.
The identifier associated with all file system volumes 225, including the FILESTREAM value (eg GUID), is registered by registering with the locator service, which allows the client machine 280 to establish a network connection with the FILESTREAM server 240. It is possible to provide a mechanism. In addition, the database server 200 registers an identifier (eg GUID) for the database 210 hosted by them. One of the possible implementations of such a locator service is a DNS (Domain Name System) service that can map volume GUIDs and database GUIDs to IP addresses.
A FILESTREAM server 240 and a client redirector can be built on the protocol to maintain coherency of SQL metadata and FILESTREAM value data between the SQL database server 200 and the FILESTREAM server 240. On the client side, the protocol extension allows client machine 280 to handle requests redirected from database server 200 to FILESTREAM server 240. To facilitate optimal offloading of activity from database server 200 to various FILESTREAM servers 240, protocol extensions include, for example, adding constraints, triggers, or columns, and deleting triggers, constraints, or columns. It can include the ability to exchange notifications about changes in the metadata of various tables. The FILESTREAM access module, or something similar implemented as part of the SQL stack, initiates propagation / invalidation of cached metadata on the FILESTREAM server 240, and they meta on the database server 200. It can include callouts needed to ensure that it is synchronized with data changes.
There are several cases where the three machine configurations shown in Figure 2 can be reduced to fewer machines. For example, in one scenario, a file system volume containing FILESTREAM data can be hosted on the same machine as the database (for example, SQL) server. In another example scenario, the table can be accessed from the machine hosting the SQL server. From the point of view of the underlying architecture, this can be either access to resources across the network or access to local resources.
Example of embodiment One embodiment provides a mechanism for maintaining transaction link-level integrity between database columns in a file system (eg structured data) and corresponding FILESTREAM data (eg unstructured data). To do. As used herein, transactional link-level integrity corresponds to a condition in which a database, such as a SQL database, has the exact status of files and directories with respect to filenames and directory names, locations, and so on. For example, if a database considers a file or directory to exist with a given name, then the file or directory must exist in the file system with that given name. Similarly, if the database considers a file or directory to not exist, then the file or directory must not exist in the file system. In certain embodiments, the types of file and / or directory operations that can be transactionally consistent include, for example, create, delete, and / or rename. As mentioned earlier, loss of transaction link-level integrity can result in errors, data corruption, or even system crashes due to the data inconsistencies that result from these conditions. In the event of a system crash, link level integrity can be lost. Therefore, the ability to reestablish link-level integrity is very important if database operations are to be resumed after a crash or other system failure.
To make it feasible to maintain link-level integrity of transactions, in some embodiments, both undo and redo of information can be logged before a file system operation is performed. In these embodiments, the logical operation in the file system is not idempotent, so recovery redo operations can be conditioned when the log sequence number (LSN) is updated. In a further embodiment, if the database transaction needs to be rolled back during recovery, the log records are flushed before the file system metadata is saved to disk to facilitate rollback of file system effects. Can act to make you. Eventually, at some point you should be able to flush the file system effect to disk so that the logs are truncated, which saves storage space.
Therefore, certain embodiments provide a mechanism for logging file system operations (eg, operations performed on a file or directory) in both the database log and the file system directory. The mechanism of one embodiment is the "ordering characteristic" that exists in common file systems, such as NTFS. Achieve write-ahead-logging by leveraging "property)" to efficiently use the logging techniques described above (eg, without the need for forced flushing to database logs). be able to. "Ordering characteristics" means that if a file system performs change A on a file or directory and then changes B on a file or directory, the file system will change B without change A on crash restart. Indicates that you will never have. In other words, at the time the system crashes (the "crash point"), regardless of the reason for the crash, either no changes, only changes A, or both changes A and B, as long as the file system itself is not corrupted. As a result, but change B without change A does not occur. In addition, in some embodiments, existing components of the database recovery framework can be used without the need for an external coordinator or transactional file system. As a result, its relatively simple and efficient operation allows a wide range of embodiments.
In the database environment described above, FILESTREAM data files can be organized into specific directories for a given database. In certain embodiments, any FILESTREAM file can have rows (which can be represented by filenames), columns, tables, and pathnames that encode the database in which the files can occur. For example, an example of a FILESTREAM file pathname for the root of a file stream group could be \ FILESTREAM-data-container \ table \ column \ rowguid. The directory structure can be changed during operations that change the table structure, such as deleting or adding columns, renaming tables, and so on. In each situation, during any of these table structure operations, appropriate locks can be acquired to exclude operations that manipulate rows in the table. In other words, the database management system can serialize concurrent file system operations performed on file system objects (files and directories) along the same higher chain in the file system hierarchy.
It will be appreciated that certain embodiments utilize this knowledge of directory rules to maintain link-level integrity of transactions. In addition, one embodiment combines a file system log-based recovery and isolation exploiting Semantics (ARIES) recovery method with a database management system ARIES recovery method. In this manner, the recovery of each system is independent of each other, and even when combined by certain embodiments, the recovery puts the database and file system data into a transactional link-level integrity state. The ARIES recovery system should be known to those of skill in the art and, for the sake of clarity, details of its implementation are not included herein.
Such a combination of ARIES recovery of a database management system and ARIES recovery of a file system can be made feasible by certain embodiments, for example, by using filenames for logging purposes within the file system. In these embodiments, the LSN of a particular operation can be encoded into the name of the file. The same name for a file will also be encoded for the name of the file system object (file or directory) that undergoes file system changes. File IDs can be reassigned by the file system during database restore and are not stable enough for logging, so such file name-based logging (ie, concatenation of LSNs and file system object names) It will be understood that the problems associated with file identity based logging are avoided.
Therefore, one embodiment can be implemented with a database that uses the FILESTREAM data attribute to store BLOBs. In these databases, file system operations (for example, creating, renaming, or deleting blobs) are logged in both the database log and the file system log. When logging an operation to a file system log, the operation can be recorded as a zero-byte file (such a zero-byte file is referred to herein as a "file system log entry"). An LSN can be assigned to each database log. For file system log entries, the corresponding database log record LSN can be used to encode the file name. The filename can also contain other coding information that describes the logged operation. Such ciphers can simply place the LSN and other descriptor information in the filename, or the LSN and any of them, without any additional process, as discussed below. It will be appreciated that other information can be stored in any type of format, for example algorithmic encryption.
In an operation, in one embodiment, if a file system operation is to be performed, the operation first logs the database log, then logs the file system log, and then the file system operation is performed. If a log entry is entered prior to the actual file system operation, then the file system log entry for the operation will exist whenever the operation itself occurs and needs to be rolled back. The recovery method should allow you to decide how to redo any operation.
File system folders can be used to log file system operations to achieve pre-reading logging. For example, a file called LOG \ X ~ A.LSN.Xact-ID.Create is created in a file system folder called LOG before file A is created under folder X. This file name has enough information encoded inside it to represent the action being performed (LSN, Xact-ID, A, and Create are LSN, transaction ID, file name, and Create, respectively). A descriptor that represents a create operation). In some embodiments, this filename is stored prior to performing the create operation. Therefore, the above-mentioned ordering characteristic indicates that if A is present, then the LOG \ X ~ A.LSN.Xact-ID.Create log entry must also be present. The embodiment can use this knowledge if the operation is to be rolled back. In some embodiments, the file system log folder must be co-located on the same volume as the file system data if the aforementioned "ordering characteristics" are only applicable within the same volume, as in the case of NTFS. Please note.
In the case of deletion of any of the directories and / or files, one embodiment provides renaming of the item to be deleted so that the item can be restored in case of rollback. As you can see, if the item to be deleted is actually deleted instead of being renamed, the delete operation cannot be rolled back because there is no data to restore the item from. In some embodiments, a file system folder named, for example, DELETED can be used for this purpose.
Now that we have discussed examples of file logging operations, we will discuss embodiments related to crash recovery. In the case of a system crash, it will be understood that at the crash point, one or more of the log entries and / or file system operations are infeasible. If the database is to be restarted, the logs can have different configurations as a result of the crash. Figures 3A-C show examples of three log scenarios that can occur as a result of a system failure such as a crash.
Figure 3A shows database log 302 and file system log 304 with log entries recorded internally. Crash point 310 represents the time when the system stopped operating due to a system crash, error, or the like. In Figure 3A, you can see that logs 302 and 304 are represented as timelines, with the left side of logs 302 and 304 being the earlier time and the right side being the more recent time. Operations A and B represent log entries with the appropriate LSN. Therefore, in FIG. 3A, there is a scenario in which changes A and B are recorded in database log 302 and file system log 304 has only change A. Operation A, as recorded in file system log 304, can be the file name described above. As discussed below with respect to FIG. 4, one embodiment performs ARIES recovery entirely based on database log 302 because database log 302 captures more information than file system log 304. be able to.
Figure 3B represents a scenario in which both database log 302 and file system log 304 flushed changes A and B prior to crash point 310. In such a scenario, in one embodiment, the database log captures at least as much information as the file system log 304 captures, so ARIES recovery can also be performed entirely based on the database log 302.
Figure 3C represents a scenario that generally takes into account some real-world problems with database logs. As mentioned above, one embodiment provides that the database log 302 can be updated prior to the file system log 304, but logs before the file system log entry is committed to the file system log 304. Cannot flush database log records frequently enough to commit to database log 302. It will be understood that flushing logs requires costly system I / O operations. Therefore, some situations can occur if the database is configured to flush its logging / entries less frequently as a trade-off for faster processing speeds. Therefore, in Figure 3C, change B was logged in both database log 302 and file system log 304, but was committed only in file system log 304, so change B remains only in file system log 304 after a system crash. Represents that. Therefore, change B is only logged in file system log 304 when the crash restarts, even if logging was originally created for database log 302. As discussed below with respect to FIG. 4, in such a scenario, in one embodiment, the file system log 304 can be treated as a logical extension to the database log 302, relating to file system operations that were not captured by the database log 302. Additional information is used to roll back the file system to a state that matches the crash point in database log 302.
Next, with reference to FIG. 4, a flow diagram is provided showing Example 400 of how to recover the database after a system crash or the like. It will be understood that Method Example 400 incorporates the ARIES algorithm. In addition, it should be understood that Method Example 400 can review the database and file system logs in relation to Figures 3A-C, as described above.
Therefore, as discussed above in relation to Figures 3A-C, in step 401, database log to find all log records accumulated since checkpoint, including log records with the latest LSN. Database logs such as 302 are analyzed. In one embodiment, step 401 can be performed in the same way as a typical ARIES analysis step in which the embodiment collects active transactions so that at rollback, the transactions that need to be rolled back are identified. I want to be understood.
In step 403, database logs such as database log 302 as discussed above in connection with Figures 3A-C, by conditionally reapplying the file system changes according to the information stored in the log. Roll forward Can be forward). For any logging, a file between the two LSNs, that is, the LSN of the logging itself (for example, "LSN1") and the file or directory name that corresponds to the redo information, or any higher directory. You can condition the redo operation by comparing it with the largest LSN that has been logged in the system log folder (for example, "LSN2" for example). If LSN1 is larger than LSN2, this indicates that the filesystem log entry for the filesystem change has not been committed to disk, based on the "ordering characteristics" described above. In addition, this also indicates that, as mentioned above, the actual filesystem changes are not committed to disk because the actual filesystem changes are always made after the filesystem log entry is created. In this case, the file system change described in the database log is re-executed.
If LSN1 is smaller than LSN2, then the actual file system changes have been committed to disk. This is because, as mentioned above, the changes made to the same file / directory name and above it are serialized within the database that holds the FILESTREAM data. In other words, no file system log entry with LSN2 is created until the first file system change corresponding to LSN1 is completed. The "ordering characteristic" ensures that if a file system log entry with LSN2 is committed to disk, the file system changes corresponding to LSN1 must be committed to disk as well. Therefore, in this case, redo is skipped.
If LSN1 is equal to LSN2, the actual file system change may or may not have been committed to disk. In this case, one embodiment can determine if the changes need to be reapplied based on the actual state of the file system. For example, if the logging / entry says "Create File A", check the file system to see if file "A" already exists. If file "A" exists, you do not need to reapply your changes. If the file "A" does not exist, the file "A" will be recreated. It should be understood that in some embodiments, step 403 can be performed in the same way as a typical ARIES redo phase.
At step 405, the type of recovery required depends on the state of the database and file system logs. In addition, a recovery preparation step can be performed in connection with step 405. For example, the log can be one of the situations described above in connection with step 403 and FIGS. 3A-C. In one embodiment, each file in the file system log folder that exceeds the end LSN of the corresponding database log can be scanned and its filename recorded in memory. The filenames can then be stored, for example, in LSN order to allow scan order, such as sequential or reverse, according to LSN for both rollforward and rollback. Since some embodiments require the file system to be recovered to the state reflected by the end of the database log, it may be necessary to roll back log entries with LSNs that exceed the end LSN of the database log. In one embodiment, if in step 405 it is found that there are no log entries in the file system log that have an LSN greater than the end LSN of the database log, then steps 407 and 409 discussed below can be skipped. This situation is described above in relation to Figures 3A-B.
At step 407, the rollforward process conditionally performs a redo operation based on any file system log entries recorded in memory by step 405 in ascending order based on the LSN. The actual conditional redo algorithm, in one embodiment, can be identical to that described above in connection with step 403. In step 409, in descending order based on LSN, step 405 allows the rollback process to be performed based on the file system log entries recorded in memory. As one of ordinary skill in the art would know, by performing an undo operation on the file system's log folder, the Compensation Log ("CLR: Compensation Log") Record) ") can be generated. These CLRs may need to be applied in the file system log folder rather than in the database log. Once the rollback process is complete within the file system, all file system effects can be flushed so that file system log entries with LSNs greater than the end LSN of the database log can be deleted.
Understand that these operations can be important because once any subsequent database log rollback is initiated, the rollback process can initiate CLR logging on both the database log and the file system log. Will be done. As a result, database log rollback CLR logging activity can cause LSN conflicts if file system logs are not cleared beyond the database log termination LSN. To flush the file system, one embodiment can track all affected files by the end of step 409 and flush those files one by one. In some embodiments, the "ordering characteristics" of the file system can also be utilized to simply create a new temporary file, flush this temporary file, and then delete it. With this new file creation and flushing, all previous file system changes (eg create, delete, rename) are flushed as well, or otherwise the "ordering characteristics" cannot be preserved. Upon completion of step 409, the file system is rolled back to the state reflected by the end of the database log.
In some embodiments, the undo operation does not occupy any special disk space. For example, canceling file creation only frees disk resources. Undoing file deletion may simply rename an existing file without increasing disk space usage. It will be appreciated that these characteristics can be important in situations where disk space is almost completely consumed.
At step 411, rollback can be performed from the end of the database log to the oldest active LSN. It should be understood that in certain embodiments, step 411 can be performed in the same manner as a typical ARIES cancellation step. Upon completion of step 411, the database and file system have been returned to a mutually matching state.
Therefore, at step 413, a checkpoint operation can be performed prior to accepting any new file system operation. In such a case, even if a system error such as a crash occurs in the future, it is not necessary to repeat the recovery process executed in steps 401 to 411 for the same operation.
It will be appreciated that Method 400 in Figure 4 is implemented by incorporating the ARIES methodology into a database and file system recovery method of an embodiment. However, it should also be understood that certain embodiments may incorporate elements that behave differently than ARIES. For example, next looking at Figure 5, shows example 500 how to roll back a file system log with a CLR. Method 500 can be performed in connection with Method 400 discussed above in connection with FIG. In a typical ARIES recovery, the CLR is not revoked. As can be seen in the discussion below, some embodiments can perform such a undo operation on the CLR.
Therefore, in step 501 it is determined whether the logging is a non-CLR logging. If it is a non-CLR log, then step 505 determines if it is a CLR-corrected log. If not corrected, the undo operation is performed in step 509. If so, the undo operation is skipped in step 507. If, as a result of the determination in step 501, the logging is not a non-CLR logging (ie, the logging is a CLR recording), then in step 503, the logging is from the end of the database log (determined by LSN, etc.). Determines if it is a CLR that corrects the previous logging. If it is such a CLR, step 509 performs an undo operation on that record. If not such a CLR, the logging is the CLR that corrects the logging after the end of the database log. In such a case, step 507 skips the undo operation. As you can see, canceling a CLR can log other CLRs (indicating the cancellation of a previous CLR). However, in one embodiment, a recovery operation for a particular file system operation can result in up to two CLRs in the file system log, so by reserving the file system log space, for example, two separate CLRs are represented2. It should be guaranteed that enough space is reserved to create one file.
It should be understood that the above procedure for handling CLRs does not store any persistence across file system log rollbacks and database log rollbacks so that each rollback process is independent of each other. Other embodiments consistently relate to which logging in the database log was corrected by the CLR in the file system log, while sticking to the ARIES algorithm to skip undo operations on the CLR. Can be remembered.
It will also be appreciated that one embodiment may choose to perform the logging and recovery described above only according to the file system log folder without duplication of logging in the database logs. However, one embodiment contemplates that the implementation of the database logging function is typically significantly better than the implementation of file system folder-based logging. Therefore, some embodiments perform recovery operations based on the database logs and use the file system logs only as an extension to those operations that were not captured by the database logs. As a result, overall recovery performance is improved compared to using only file system logs.
Although the present invention has been described above in relation to various embodiments of the drawings, the same functions of the present invention can be used without deviating from the present invention or other similar embodiments can be used. It will be appreciated that modifications and additions to the described embodiments are feasible to carry out. Therefore, it should be understood that the present invention is not limited to any single embodiment, but is within the scope and scope of the appended claims.
<figref num="1">It is a figure which shows the example of the computing environment which can carry out various aspects of this invention.</figref><figref num="2">It is a figure which shows the example of the database structure which can carry out various aspects of this invention.</figref><figref num="3A">It is a figure which shows the example of the log which can carry out various aspects of this invention.</figref><figref num="3B">It is a figure which shows the example of the log which can carry out various aspects of this invention.</figref><figref num="3C">It is a figure which shows the example of the log which can carry out various aspects of this invention.</figref><figref num="4">It is a flow chart which shows the example of the method according to various embodiments of this invention.</figref><figref num="5">It is a flow chart which shows the example of the method according to various embodiments of this invention.</figref>
Every citation, both waysCites: the store holds 7 of 8
| Document | Relation | Office |
|---|---|---|
| JP5143424A | Cites | Japan |
| JP2004252686A | Cites | Japan |
| JP2005100373A | Cites | Japan |
| US6088694A | Cites | United States of America |
| US6397351B1 | Cites | United States of America |
| US6453325B1 | Cites | United States of America |
| US2004148308A1 | Cites | United States of America |
13 members in 6 offices
Priority claims9
| Document | Office | Kind | Date |
|---|---|---|---|
| 11123563 | United States of America | – | |
| 12356305 | United States of America | A | |
| 12356305 | United States of America | A | |
| 2006008278 | United States of America | W | |
| 2006008278 | United States of America | W | |
| 2005123563 | – | – | – |
| 2006008278 | – | – | – |
| US20050123563 | – | – | – |
| WO2006US08278 | – | – | – |
Members13
| Document | Office | Kind | |
|---|---|---|---|
| US2006253502A1 | United States of America | A1 | |
| WO2006121500A2 | World Intellectual Property Organization (WIPO) | A2 | |
| WO2006121500A3 | World Intellectual Property Organization (WIPO) | A3 | |
| KR20080005501A | Republic of Korea | A | |
| EP1877906A2 | European Patent Office (EPO) | A2 | |
| JP2008541225A | Japan | A | |
| CN101460930A | China | A | |
| EP1877906A4 | European Patent Office (EPO) | A4 | |
| CN101460930B | China | B | |
| US8145686B2 | United States of America | B2 | |
| KR101255392B1 | Republic of Korea | B1 | |
| JP5259388B2This record | Japan | B2 | |
| EP1877906B1 | European Patent Office (EPO) | B1 |
20 legal events, as the office reported them to INPADOC
Over the term
Point at a mark for the eventEvents
| Event | Code | |
|---|---|---|
| Cancellation because of no payment of annual feesLAPS | LAPS | |
| Receipt of annual feesJAPANESE INTERMEDIATE CODE: R250R250 | R250 | |
| Receipt of annual feesJAPANESE INTERMEDIATE CODE: R250R250 | R250 | |
| Receipt of annual feesJAPANESE INTERMEDIATE CODE: R250R250 | R250 | |
| Receipt of annual feesJAPANESE INTERMEDIATE CODE: R250R250 | R250 | |
| Receipt of annual feesJAPANESE INTERMEDIATE CODE: R250R250 | R250 | |
| Written notification of registration of transferJAPANESE INTERMEDIATE CODE: R350R350 | R350 | |
| Request for change of ownership or part of ownershipJAPANESE INTERMEDIATE CODE: R313113S111 | S111 | |
| First payment of annual fees (during grant procedure)JAPANESE INTERMEDIATE CODE: A61A61 | A61 | |
| Renewal fee payment (event date is renewal date of database)FPAY | FPAY | |
| Certificate of patent or registration of utility modelJAPANESE INTERMEDIATE CODE: R150R150 | R150 | |
| Certificate of patent or registration of utility modelJAPANESE INTERMEDIATE CODE: R150R150 | R150 | |
| Written decision to grant a patent or to grant a registration (utility model)JAPANESE INTERMEDIATE CODE: A01A01 | A01 | |
| Decision of grant or rejection writtenTRDD | TRDD | |
| Written amendmentJAPANESE INTERMEDIATE CODE: A523A521 | A521 | |
| Notification of reasons for refusalJAPANESE INTERMEDIATE CODE: A131A131 | A131 | |
| Written amendmentJAPANESE INTERMEDIATE CODE: A523A521 | A521 | |
| Notification of reasons for refusalJAPANESE INTERMEDIATE CODE: A131A131 | A131 | |
| Written amendmentJAPANESE INTERMEDIATE CODE: A523A521 | A521 | |
| Written request for application examinationJAPANESE INTERMEDIATE CODE: A621A621 | A621 |
Numbers
- Publication
- 5259388
- Publication, DOCDB
- 5259388
- Publication, EPODOC
- JP5259388B
- Application
- 2008509997
- Application, DOCDB
- 2008509997
- Application, EPODOC
- JP20080509997
Titles2
- English
- Maintaining link-level integrity between the database and the file system
- Japanese
- データベースとファイルシステムとの間でのリンクレベル整合性の維持
Classification
- CPC, 9
- G06F17/30008
- G06F16/2308
- G06F16/2358
- G06F8/44
- G06F16/23
- G06F9/466
- G06F16/1865
- G06F16/48
- G06F40/126
- IPC, 1
- G06F12 00