AU2004264635A2

Fast application notification in a clustered computing system

Abstract

With fast notification of changes to a clustered computing system, through which a number of events are published for system state changes, applications can quickly recover and sessions can quickly be rebalanced. When a resource associated with a service experiences a change in status, such as a termination or a start/restart, a notification event is immediately published. Notification events contain information to enable subscribers to identify, based on matching a session signature, the particular sessions that are affected by the change in status, and to respond accordingly. This allows sessions to be quickly aborted and ongoing processing to be quickly terminated when a resource fails, and allows fast rebalancing of work when a resource is restarted.

AU2004264635A2, drawing sheet 1
Sheet 1 of 3

Term

Term ended

Projected expiry passed 13 August 2024, 2.1 years ago.

  1. Priority
  2. Filed
  3. Published
  4. Projected expiry
  5. Today

10 claims: 7 independent, 3 dependent

  1. 1
    CLAIMS What is claimed is:1. A method for communicating changes about a clustered computing environment that comprises a plurality of interconnected nodes that host server instances, comprising the computer-implemented steps of: receiving an indication of a status change to a resource associated with a particular service that performs work in a cluster;in response to said status change to said service resource, immediately generating first data that identifies said particular service and second data that indicates a status of said resource;publishing said first and second data to a set of one or more subscribers;and wherein said first data is used by a subscriber to identify, based on identification information that is generated in response to establishing a session with said cluster and that identifies a service associated with said session, one or more sessions with said cluster that are affected by said status change to said service resource.
  2. 2
    The method of Claim 1, wherein said status of said resource that is identified in said event is one from a group consisting of (a) termination of said resource, (b) starting of said resource, and (c) unable to restart said resource.
  3. 3
    The method of Claim 1, wherein said cluster is a database cluster, and wherein said resource is identified in said first data by identifying said database cluster that is affected by said status change. -242004264635 21 Apr 2006 ι
  4. 4
  5. 5
  6. 6
  7. 7
  8. 8
  9. 9
  10. 10
    The method of Claim 3, wherein said work is associated with a service, and wherein said resource is further identified in said first data by identifying said service that is affected by said status change. The method of Claim 4, wherein said location of said resource is further identified in said first data by identifying an instance and a node that are affected by said status change. The method of Claim 3, further comprising the computer-implemented step of:identifying one or more sessions with said database cluster that are affected by said status change, based on matching said identification information that is associated with said session with said first data that identifies said resource. The method of Claim 6, further comprising the computer-implemented step of: interrupting said one or more sessions. The method of Claim 1, wherein said resource is identified in said first data by identifying a node that is affected by said status change. The method of Claim 1, wherein said work is associated with a service, and wherein said resource is identified in said first data by identifying that a service has terminated at a particular instance and by identifying said particular instance at which said service has terminated. The method of Claim 1, wherein said work is associated with a service, and wherein said resource is identified in said first data by identifying that an entire service has terminated and by identifying said service that has terminated. -252004264635 21 Apr 2006 1 11. The method of Claim 1, wherein said resource is identified in said first data by 2 identifying that a particular instance has terminated and by identifying said 3 particular instance that has terminated. 1 12. The method of Claim 1, wherein said resource is identified in said first data by 2 identifying that all said instances have terminated and that identifies said cluster 3 with which said instances are associated. 1 13. The method of Claim 1, wherein said work is associated with a service, and 2 wherein said resource is identified in said first data by identifying that a service 3 has started at a particular instance and by identifying said particular instance on 4 which said service has started. 1 14. The method of Claim 13, wherein said resource is identified in said first data by 2 identifying the number of instances that support said service that has started. 1 15. The method of Claim 1, wherein said work is associated with a service, and 2 wherein said resource is identified in said first data by identifying that a service 3 has started on any instance and by identifying said service that has started. 1 16. The method of Claim 15, wherein said resource is identified in said first data by 2 identifying a number of instances that support said service that has started. 1 17. The method of Claim 1, wherein said resource is identified in said first data by 2 identifying that a particular instance has started and by identifying said instance 3 that has started. -262004264635 21 Apr 2006 1 18. The method of Claim 1, wherein said resource is identified in said first data by 2 identifying that an instance has started and by identifying said cluster with which 3 said instance is associated. 1 19. The method of Claim 1, wherein said resource is identified in said first data by 2 identifying that a node has terminated and by identifying said node that has 3 terminated. 1 20. The method of Claim 1, wherein the step of publishing comprises publishing said 2 first and second data through processes that are not part of clusterware that 3 manages said cluster. 1 21. The method of Claim 1, wherein a subscriber to said first and second data is a 2 connection pool manager that responds to said status change by redistributing, 3 based on said first and second data, connections to said cluster. 1 22. The method of Claim 1, wherein a subscriber to said first and second data is a 2 client application that responds to said status change by requesting, based on said 3 first and second data, redistribution within said cluster of said work that is affected 4 by said status change. 1 23. The method of Claim 1, wherein a subscriber to said first and second data is a 2 batch job that responds to said first and second data by calling for execution of a 3 routine within said cluster based on said status change. 1 24. The method of Claim 1, wherein said work is associated with a service, and 2 wherein said resource is identified in said first data by identifying that said service -272004264635 21 Apr 2006 1 25 1 26 1 27. is not restarting so that subscribing applications are interrupted from retrying to use said service. A computer-readable medium carrying one or more sequences of instructions which, when executed by one or more processors, causes the one or more processors to perform the method recited in any one of Claims 1 to 24. A system comprising: a database cluster comprising a set of server instances hosted by a set of interconnected nodes communicatively coupled to a database;cluster management software that manages resources in said cluster and distribution and performance of work in said cluster, wherein said resources are associated with respective particular services;a notification system that publishes information about status changes of said resources, for use in identifying, based on identification information that is generated when a session is established with said cluster and that identifies resources associated with said session, one or more sessions with said cluster that are affected by a respective status change;and wherein each said information about a status change to a resource comprises identification of one or more particular services associated with said resource and a status of said resource. The system of Claim 26, wherein said status of said resource that is identified in said information about a status change consists of one from a group consisting of (a) termination of said resource, (b) starting of said resource, and (c) unable to restart said resource. -282004264635 21 Apr 2006 1 28. The system of Claim 26, wherein said resource that is identified in said 2 information about a status change is associated with at least one from a group 3 consisting of (a) a service, (b) a service member that is executing on a particular 4 instance of said instances, (c) said database cluster, (d) one of said instances, and 5 (e) one of said nodes. 2 29. A method for communicating changes about a clustered computing environment, 3 substantially as hereinbefore described with reference to the accompanying 4 drawings. 2 30. A notification system for a clustered computing environment, substantially as 3 hereinbefore described with reference to the accompanying drawings.