AU2004264635B2

Fast application notification in a clustered computing system

Abstract

With fast notification of changes to a clustered computing system, through which a number of events are published for system state changes, applications can quickly recover and sessions can quickly be rebalanced. When a resource associated with a service experiences a change in status, such as a termination or a start/restart, a notification event is immediately published. Notification events contain information to enable subscribers to identify, based on matching a session signature, the particular sessions that are affected by the change in status, and to respond accordingly. This allows sessions to be quickly aborted and ongoing processing to be quickly terminated when a resource fails, and allows fast rebalancing of work when a resource is restarted.

AU2004264635B2, drawing sheet 1
Sheet 1 of 3

Term

Term ended

Expired 13 August 2024, 2.1 years ago.

  1. Priority
  2. Filed
  3. Granted
  4. Expired
  5. Today

19 claims: 5 independent, 14 dependent

  1. 1
    2004264635 18 Feb 2010 CLAIMS The claims defining the invention are as follows:, 1. A method for communicating changes about a clustered computing environment that comprises a plurality of interconnected nodes that host server instances, comprising the computer-implemented steps of: receiving an indication of a status change to a resource associated with a particular service that performs work in a cluster;in response to said status change to said resource, immediately generating first data that identifies said particular service and second data that indicates a status of said resource;publishing said first and second data to a set of one or more subscribers;and wherein said first data is used by a subscriber to identify, based on identification information that is generated in response to establishing a session with said cluster and that identifies a service associated with said session, one or more sessions with said cluster that are affected by said status change to said service resource.
  2. 2
    The method of Claim 1, wherein said status of said resource that is identified in said second data is one from a group of statuses including:(a) termination of said resource;(b) starting of said resource;and (c) unable to restart said resource.
  3. 3
    The method of Claim 1, further comprising the computer-implemented step of:interrupting said one or more sessions. -242004264635 18 Feb 2010
  4. 4
    The method of Claim 1, wherein said resource is identified in said first data by identifying a number of instances that support said service that has started.
  5. 5
    The method of Claim 1, wherein publishing comprises publishing said first and second data through processes that are not part of clusterware that manages said cluster.
  6. 6
    The method of Claim 1, wherein a subscriber to said first and second data is a connection pool manager that responds to said status change by redistributing, based on said first and second data, connections to said cluster.
  7. 7
    The method of Claim 1, wherein a subscriber to said first and second data is a client application that responds to said status change by requesting, based on said first and second data, redistribution within said cluster of said work that is affected by said status change.
  8. 8
    The method of Claim 1, wherein a subscriber to said first and second data is a batch job that responds to said first and second data by calling for execution of a routine within said cluster based on said status change. -252004264635 18 Feb 2010
  9. 9
    The method of Claim 1, wherein said work is associated with a service, and wherein said resource is identified in said first data by identifying that said service is not restarting so that subscribing applications are interrupted from retrying to use said service.
  10. 10
    A system comprising:a database cluster comprising a set of server instances hosted by a set of interconnected nodes communicatively coupled to a database;cluster management software that manages resources in said cluster and distribution and performance of work in said cluster, wherein said resources are associated with respective particular services;a notification system that publishes information about status changes of said resources, for use in identifying, based on identification information that is generated when a session is established with said cluster and that identifies resources associated with said session, one or more sessions with said cluster that are affected by a respective status change;and wherein each said information about a status change to a resource comprises identification of one or more particular services associated with said resource and a status of said resource.
  11. 11
    The system of Claim 10, wherein said status of said resource that is identified in said information about a status change consists of one from a group consisting of -26(a) termination of said resource, (b) starting of said resource, and (c) unable to restart said resource.
  12. 12
    The system of Claim 10, wherein said resource that is identified in said information about a status change is associated with at least one from a group of resources including:(a) a service;(b) a service member that is executing on a particular instance of said instances;(c) said database cluster;(d) one of said instances;and (e) one of said nodes.
  13. 13
    The method of Claim 1, wherein said resource identified in said first data by identifying at least one of the following that is affected by said status change:a) a database cluster, wherein said cluster is said database cluster;b) a node;c) an instance and a node, wherein the instance and the node further identify said location of said resource;
  14. 14
    The method of Claim 13, further comprising identifying at least one of:a) a service that is affected by said status change, wherein said work is associated with said service;or b) one or more sessions with said database cluster, wherein said one or more sessions are identified by matching said identification information that is associated with said session with said first data that identifies said resource. -272004264635 18 Feb 2010
  15. 15
    The method of Claim 13, wherein said resource is identified in said first data by identifying at least one of the following:a) that a server instance has terminated, and by identifying the particular instance at which said service has terminated;b) that an entire service has terminated, and by identifying said service that has terminated;c) that a particular instance has terminated, and by identifying said particular instance that has terminated;d) that said all instances have terminated, and by identifying said cluster with which said instances are associated;or e) that a node has terminated, and by identifying said node that has terminated.
  16. 16
    The method of Claim 13, wherein said resource is identified by said first data by identifying at least one of the following:a) that a service has started at a particular instance, and by identifying said particular instance on which said service started;b) that a service has started on any instance, and by identifying said service that has started;c) that a particular instance has started, and by identifying said instance that has started;d) that an instance has started, and by identifying said cluster with which said instance is associated;or e) a number of instances that support said service that has started. -282004264635 18 Feb 2010
  17. 17
    A computer-readable medium carrying one or more sequences of instructions which, when executed by one or more processors, causes the one or more processors to perform the method recited in any one of Claims 1 to 9 and 13 to 16.
  18. 18
    A method for communicating changes about a clustered computing environment, substantially as hereinbefore described with reference to the accompanying drawings.
  19. 19
    A notification system for a clustered computing environment, substantially as hereinbefore described with reference to the accompanying drawings. I