Video on demand digital server load balancing
Summary by NHIP
Director Server Load Balancing
The system designates a director server to receive and forward asset requests to a host processor for evaluation. This server uses an adaptable cache to detect request pairs, identify the shortest interval between start times, and stream the asset while storing it locally.
Claim Score by NHIP
Abstract
A system and method for load balancing a plurality of servers is disclosed. In a preferred embodiment, a plurality of servers in a video-on-demand or other multi-server system are divided into one or more load-balancing groups. Each server preferably maintains state information concerning other servers in its load-balancing group including information concerning content maintained and served by each server in the group. Changes in a server's content status or other state information are preferably proactively delivered to other servers in the group. When a content request is received by any server in a load-balancing group, it evaluates the request in accordance with a specified algorithm to determine whether it should deliver the requested content itself or redirect the request to another server in its group. In a preferred embodiment, this determination is a function of information in the server's state table.

Term
Term ended
Expired 27 June 2023, 3.2 years ago.
- Priority
- Filed
- Granted
- Expired
- Today
25 claims: 3 independent, 22 dependent
- 1Broadest claimClaim Score 21, narrow(NHIP)A method comprising:designating a director server from among a plurality of servers configured to service requests for assets and configured to function as a director server if designated, each of the plurality of servers further configured with a storage system and an adaptable cache configured to proactively cache assets;receiving a first request and a second request for an asset at a network interface of the designated director server;forwarding the first request and the second request to a host processor of the designated director server over a bus;determining whether or not the asset is present on the designated director server;if the asset is present on the designated director server, servicing the first request and the second request from the designated director server by: monitoring the bus by a designated director server's adaptable cache, detecting the first request and the second request on the bus at the designated director server's adaptable cache, generating, at the designated director server's adaptable cache, a list of pairs of requests for the asset comprising a first pair of requests comprising the first request and the second request, determining that the first pair of requests has the shortest interval between start times from among pairs of requests identified in the list of pairs of requests, responsive to the first request, streaming the asset from a designated director server's storage system and storing the asset on the designated director server's adaptable cache as it is being streamed from the designated director server's storage system, and responsive to the second request, streaming the asset from the designated director server's adaptable cache and transmitting an instruction from the designated director server's adaptable cache to the designated director server's storage system directing the designated director server's storage system not to respond to the second request;and if the asset is not present on the designated director server, selecting, by the designated director server based on a designated director server's state table, a first server in said plurality of servers to service the first request and the second request, said designated director server's state table comprising state information for said first server indicating that the asset is present on the first server.
- 11A method comprising:responsive to an asset being copied onto a first server in a load balancing group, modifying, by said first server, said first server's state information table by placing an identifier for the asset in said first server's state information table, wherein said first server's state information table comprises state information for servers in said load balancing group, wherein said state information comprises a plurality of parameters concerning said servers in said load balancing group, and wherein the plurality of parameters comprise asset identifiers;pushing, by said first server, said identifier for the asset to all servers in said load balancing group;modifying, by each server in said load balancing group, its respective state information table, by placing said identifier for the asset in its respective state information table;receiving by said load balancing group a designation by a management system of a director server from among servers in said load balancing group, wherein each server in said load balancing group is configured to function as a director server if designated, and wherein each server in said load balancing group is further configured with a storage system and an adaptable cache configured to proactively cache assets;receiving at a director server's network interface a first request and a second request for an asset, forwarding the first request and the second request to a director server's host processor over a bus;determining whether or not the asset is present on the director server;if the asset is present on the director server, servicing the request for the video asset by: monitoring the bus by a director server's adaptable cache, detecting the first request and the second request on the bus at the director server's adaptable cache, generating, at the director server's adaptable cache, a list of pairs of requests for the asset comprising a first pair of requests comprising the first request and the second request, determining that the first pair of requests has the shortest interval between start times from among pairs of requests associated with the list of pairs of requests, responsive to the first request, streaming the asset from a director server's storage system and storing the asset on the director server's adaptable cache as it is being streamed from the designated director server's storage system, and responsive to the second request, streaming the asset from the director server's adaptable cache and transmitting an instruction from the director server's adaptable cache to the director server's storage system directing the director server's storage system not to respond to the second request;and if the asset is not present on the director server, servicing the first request and the second request by: determining by said director server that the identifier for the asset in a director server's state information table is associated with a selected server in said load balancing group;transmitting an instruction from the director server to the selected server to stream the asset to a requesting client;and streaming by said selected server the asset to the requesting client.
- 16A method comprising:designating, by a management system, a director server from among a plurality of servers configured to service requests for assets and configured to function as the director server if designated, each of the plurality of servers further configured with a storage system and an adaptable cache configured to proactively cache video assets;receiving a first request and a second request for a first asset at a network interface of the designated director server;forwarding the first request and the second request to a host processor of the designated director server over a bus;determining that the first asset is present on the designated director server;servicing the first request and the second request from the designated director server by: monitoring the bus by a designated director server's adaptable cache, detecting the first request and the second request on the bus at the designated director server's adaptable cache, generating, at the designated director server's adaptable cache, a list of pairs of requests for the first asset comprising a first pair of requests comprising the first request and the second request, determining that the first pair of requests has the shortest interval between start times from among pairs of requests associated with the list of pairs of requests, responsive to the first request, streaming the first asset from a designated director server's storage system and storing the first asset on the designated director server's adaptable cache as it is being streamed from the designated director server's storage system, and responsive to the second request, streaming the first asset from the designated director server's adaptable cache and transmitting an instruction from the designated director server's adaptable cache to the designated director's storage system directing the designated director's storage system not to respond to the second request;receiving a third request for a second asset at the designated director server;determining that the second asset is not present on the designated director server;and selecting, by the designated director server based on a designated director server's state table, a first server in said plurality of servers to service the third request, said designated director server's state table comprising state information for said first server indicating that the second asset is present on the first server.
Independent claims3
76 paragraphs in 6 sections, as filed
CROSS REFERENCE TO RELATED APPLICATIONS
This applications is a continuation of U.S. applicaton Ser. No. 10/609,426, filed Jun. 27, 2003, and entitled “System and Method for Digital Media Server Load Balancing.”
FIELD OF THE INVENTION
This invention relates to the field of load balancing.
BACKGROUND OF THE INVENTION
Load balancing techniques exist to ensure that individual servers in multi-server systems do not become overloaded and that services retain high availability. Load balancing is especially important where it is difficult to predict the number and timing of requests that will require processing.
Most current load-balancing schemes employ simple parameters to distribute network traffic across a group of servers. These parameters are usually limited to load amount (measured by the number of received requests), server “health” or hardware status (measured by processor temperature or functioning random access memory), and server availability.
One common load-balancing architecture employs a supervisor/subordinate approach. In this architecture, a control hierarchy of devices is established in a load-balancing domain. Each server in the system is assigned to a load-balancing group that includes a central device for monitoring the status of servers in its group. The supervisor acts as the gatekeeper for requests entering the group and delegates each request to an appropriate server based on the server's relative status to that of other servers in the group.
One negative aspect of this approach is that it introduces a single point of failure into the load-balancing process. If the supervisor goes offline for any reason, incoming requests cannot be serviced. To ameliorate this problem, some load-balancing schemes employ a secondary supervisor to handle requests when the primary supervisor is unavailable. A secondary supervisor, however, introduces extra cost in terms of physical equipment and administration.
One of the earliest forms of load balancing, popular in the early 1990's, is commonly referred to as domain name service (DNS) round robin. This load-balancing scheme, described in connection with <figref idref="DRAWINGS">FIG. 1</figref>, represents an extension of the standard domain name resolution technique primarily used by Internet Web servers experiencing extremely high usage.
As shown in <figref idref="DRAWINGS">FIG. 1</figref>, in step <b>110</b>, a client requests data from a DNS server. In step <b>120</b>, the domain name server resolves the requested server name into a series of server addresses. Each address in the series corresponds to a server belonging to a single load-balancing group. Each server in the group is provided with a copy of all data to be served, so that each server replicates data stored by every other server in the group.
In step <b>130</b>, the domain name server assigns new requests by stepping through the list of server addresses, resulting in a crude and unpredictable load distribution for servers in the load-balancing group. Moreover, if the number of requests overloads the domain name server or if the server selected to service the request is at capacity, the service is ungracefully denied. In addition, if the selected server is at capacity, the new request routed by the domain name server may bring the server down.
Another major problem with DNS round robin is that the domain name server has no knowledge of server availability within the load-balancing group. If a server in the group is down, DNS round robin will nevertheless direct traffic to it.
In the mid 1990's, second generation load-balancing solutions were released. These solutions employed a dedicated load balance director (LBD), such as Cisco Systems's Local Director. The director improves the DNS round robin load-balancing scheme by periodically testing the network port connections of each server in its group and direction responses to responsive servers. One such second generation solution is discussed in “Load Balancing: A Multifaceted Solution for Improving Server Availability” (1998 Cisco Systems, Inc.,) submitted to the United States Patent and Trademark Office for their files pertaining to this patent.
A third generation of load-balancing solutions included robust, dedicated load balancing and network management devices, such as the BIG-IP™ from F5 NETWORKS.™ These devices improve server availability by monitoring server health via management protocols such as Simple Network Management Protocol (SNMP). Perhaps the biggest improvement of this generation is the ability to direct traffic based on requested content type instead of just load. For example, requests ending in “.http” are directed to Web servers, “.ftp” to file download servers, and “.ram” to REALNETWORKS'™ streaming servers. This feature enables network managers to create multiple load-balancing groups dedicated to specific content types.
Although the aforementioned load-balancing techniques are often adequate for managing multi-server systems that serve Web pages, file downloads, databases, and email, they still leave room for significant improvement. Moreover, such load-balancing schemes do not perform well in systems that serve broadcast-quality digital content, which is both time sensitive and bandwidth intensive.
SUMMARY OF THE INVENTION
A system and method for load balancing a plurality of servers is disclosed. In a preferred embodiment, a plurality of servers in a video-on-demand or other multi-server system are divided into one or more load-balancing groups. Each server preferably maintains state information concerning other servers in its load-balancing group including information concerning content maintained and served by each server in the group. Changes in a server's content status or other state information are preferably proactively delivered to other servers in the group. Thus, for example, to maintain a current inventory of assets within a load-balancing group, each server provides notification to other servers in its group when an asset that it maintains is added, removed, or modified.
When a content request is received by any server in a load-balancing group, it evaluates the request in accordance with a specified algorithm to determine whether it should deliver the requested content itself or redirect the request to another server in its group. In a preferred embodiment, this determination is a function of information in the server's state table.
The present system and method provide several benefits. First, because they employ a peer-based balancing methodology in which each server can respond to or redirect client requests, the present system and method do not present a single point of failure, as do those schemes that utilize a single load-balancing director. Second, because the present system and method proactively distribute state information within each group, each server is made aware of the current status of every server in its group prior to a client request. Consequently, when a request for content is received, it may be rapidly directed to the appropriate server without waiting for polled status results from other servers in the group. Moreover, in some preferred embodiments, the present system and method defines parameters concerning the capability of each server such as extended memory, inline adaptable cache, or other unique storage attributes, thus permitting sophisticated load-balancing algorithms that take account of multiple factors that may affect the ultimate ability of the system to most efficiently respond to client requests. Furthermore, in some preferred embodiments, the present system and method considers other media asset parameters such as whether an asset is a “new release” to help anticipate demand for the asset.
In one aspect, the present invention is directed to a method for selecting a server from a plurality of servers to service a request for content, comprising: designating a director from the plurality of servers to receive the request, wherein the designation is made on a request-by-request basis; and allocating to the director the task of selecting a server to service the request from the plurality of servers, said server having stored thereon the content, the director using a state table comprising parametric information for servers in the plurality of servers, wherein said parametric information comprises information identifying assets maintained on each server in the plurality of servers.
In another aspect of the present invention, the step of designating comprises designating the director in a round-robin fashion.
In another aspect of the present invention, the step of designating comprises designating the director on the basis of lowest load.
In another aspect of the present invention, the step of selecting further comprises selecting the director if the content is present on the director.
In another aspect of the present invention, said parametric information further comprises functional state and current load of each server.
In another aspect of the present invention, said parametric information further comprises whether each server comprises extended memory.
In another aspect of the present invention, said parametric information further comprises whether each server comprises an inline adaptable cache.
In another aspect of the present invention, said parametric information further comprises whether each asset is a new release.
In another aspect of the present invention, the method further comprises rejecting the request if the content is not present on any of the plurality of servers.
In another aspect of the present invention, the method further comprises forwarding the request to the selected server.
In another aspect of the present invention, The method further comprises redirecting the request to the selected server.
In another aspect of the present invention, the step of selecting further comprises: calculating a load factor for each server in the plurality of servers having the content; identifying as available servers one or more servers whose parameters are below threshold limits; selecting a server from the available servers having the lowest load factor; and otherwise selecting a server having the lowest load factor from the plurality of servers having the content.
In another aspect, the present invention is directed to a server for directing a request for content among a plurality of servers comprising: a state table comprising parametric information for each server in the plurality of servers, said parametric information comprising information identifying assets maintained on the plurality of servers; and a communication component for sending changes to the state table to the plurality of servers.
In another aspect of the present invention, the server is a member of a load-balancing group, and the communication component sends changes to servers in the load-balancing group.
In another aspect of the present invention, the server further comprises a redirection means for acknowledging the client request and identifying one of the plurality of servers where the requested asset is stored.
In another aspect of the present invention, the server further comprises a forwarding means for sending the client request to one of the plurality of servers where the requested asset is stored.
In another aspect of the present invention, said parametric information further comprises functional state and current load of each server.
In another aspect of the present invention, said parametric information further comprises whether each server comprises extended memory.
In another aspect of the present invention, said parametric information further comprises whether each server comprises an inline adaptable cache.
In another aspect of the present invention, said parametric information further comprises whether each asset is a new release.
BRIEF DESCRIPTION OF THE DRAWINGS
<figref idref="DRAWINGS">FIG. 1</figref> is a flowchart illustrating a DNS round robin load-balancing scheme in the prior art;
<figref idref="DRAWINGS">FIG. 2</figref> is a block diagram illustrating an exemplary designation of servers into load-balancing groups;
<figref idref="DRAWINGS">FIG. 3</figref> is a diagram illustrating exemplary state tables in a preferred embodiment of the present system and method;
<figref idref="DRAWINGS">FIG. 4</figref> is a flowchart illustrating a preferred embodiment of a process for updating state tables;
<figref idref="DRAWINGS">FIG. 5A and 5B</figref> are flowcharts illustrating a preferred embodiment of the present system and method for load-balancing in a multi-server system;
<figref idref="DRAWINGS">FIGS. 5C and 5D</figref> are block diagrams illustrating the communication paths for forwarding and redirecting client content requests in a preferred embodiment of the present system and method;
<figref idref="DRAWINGS">FIG. 6</figref> is a flowchart illustrating a preferred embodiment of the present system and method for load-balancing in a multi-server system with replicated content;
<figref idref="DRAWINGS">FIG. 7</figref> is a flowchart illustrating a preferred load-balancing algorithm; and
<figref idref="DRAWINGS">FIG. 8</figref> is a diagram illustrating an exemplary state table in a preferred embodiment of the present system and method.
DESCRIPTION OF THE PREFERRED EMBODIMENTS
<figref idref="DRAWINGS">FIG. 2</figref> shows an exemplary digital media delivery system that includes six digital media servers A-F which provide content to a plurality of clients via a network. In the exemplary embodiment of <figref idref="DRAWINGS">FIG. 2</figref>, the six servers are divided into two load-balancing groups (LBG) <b>210</b>, <b>220</b>. Servers A-C are designated as belonging to load-balancing group <b>1</b> (<b>210</b>) and servers D-F are designated as belonging to load-balancing group <b>2</b> (<b>220</b>). Load-balancing groups <b>210</b>, <b>220</b> may or may not be geographically diverse, but are logical groupings that designate which servers will share state information in a preferred embodiment of the present system and method, as described below.
Each server A-F preferably maintains state information concerning one or more parameters associated with each server in its group. Accordingly, each of servers A-C preferably maintains such state information for servers A-C and each of servers D-F preferably maintains such state information for servers D-F.
One preferred embodiment for maintaining state information concerning servers in a load-balancing group is shown in <figref idref="DRAWINGS">FIG. 3</figref>. As shown in <figref idref="DRAWINGS">FIG. 3</figref>, a first table <b>310</b> preferably comprises a row for each server in load-balancing group <b>210</b> and a second table <b>320</b> preferably comprises a row for each server in load-balancing group <b>220</b>. Each table <b>310</b>, <b>320</b> also preferably comprises a plurality of columns for storing state information concerning a plurality of parameters that may be considered in load-balancing determinations, as described below. Such parameters may include, without limitation, Moving Picture Experts Group (MPEG) bandwidth limits, total external storage capacity, addressable service groups (i.e., customer set-top boxes), current number of streams, system central processing unit (CPU) utilization, user CPU utilization, CPU idle, outgoing stream bandwidth, incoming bandwidth for network interfaces, streaming server state, system uptime, time-averaged load, server temperature, memory usage, total available cache, cache used, total external storage remaining, total external storage used, MPEG stream count limit, status of network connections, and status of power supplies, fans, and storage devices.
In a preferred embodiment, one or more of the stored parameters relate to the asset inventory of each server. The state table may also store other media asset parameters such as whether the asset is a “new release” to help anticipate demand for the asset. The state table additionally may contain parameters concerning the capability of each server such as whether it comprises extended memory or an inline adaptable cache (such as that described in U.S. patent application Ser. No. 10/609,433, now U.S. Pat. No. 7,500,055, entitled “ADAPTABLE CACHE FOR DYNAMIC DIGITAL MEDIA”, filed Jun. 27, 2003, which is hereby incorporated by reference in its entirety for each of its teachings and embodiments), or other unique storage attributes.
In a preferred embodiment, an adaptable cache is a compact storage device that can persist data and deliver it at an accelerated rate, as well as act as an intelligent controller and director of that data. Incorporating such an adaptable cache between existing storage devices and an external network interface of a media server, or at the network interface itself, significantly overcomes the transactional limitations of storage devices, increasing performance and throughput for an overall digital media system.
In one embodiment, a user request may be received at a network interface and forwarded to a host processor via an I/O bus. The host processor may send a request for the asset to the storage system the I/O bus. The adaptable cache may monitor asset requests that traverse the I/O bus and determine if the requested asset is available on the adaptable cache. If the asset is available on the adaptable cache, it may be returned to the host processor. Otherwise, if the requested resource is unavailable from the adaptable cache, the request may be forwarded to the storage system for delivery to the appropriate storage device where the resource persists. The storage device may then return the resource to the requesting application. The adaptable cache may be adapted to monitor received requests, proactively cache some or all of an asset in accordance with caching rules, and notify one or more applications or processes of content that it is currently storing. Alternatively or in addition, the adaptable cache may be adapted to direct the storage system not to respond to requests for particular assets when the assets are cached in the adaptable cache.
Alternatively or in addition, an adaptable cache may be adapted to perform interval caching wherein a sorted list of pairs of overlapping requests for the same asset is maintained that identifies pairs of requests with the shortest intervals between their start times. For these pairs, as the first request in the pair is streamed, the streamed content is also cached and then read from cache to serve the second request.
In a preferred embodiment, threshold limits may be specified for one or more of the stored parameters that represent an unacceptable condition. Use of these threshold limits in selecting a server to deliver requested content is described in more detail below.
A preferred embodiment for updating state tables <b>310</b>, <b>320</b> at each server is illustrated in <figref idref="DRAWINGS">FIG. 4</figref>. As shown in <figref idref="DRAWINGS">FIG. 4</figref>, in step <b>410</b>, all servers in a load-balancing group initially have identical state tables. In step <b>420</b>, a parameter of server A is modified (e.g., an asset is copied onto server A), thus changing server A's state. In step <b>430</b>, server A updates its own state table, and pushes the state change information to all other servers in its load-balancing group. In a preferred embodiment, this state change information is transmitted concurrently to all other servers in the group via, for example, a multicast, broadcast, or other one-to-many communications mechanism. In step <b>440</b>, the other servers update their state tables with the state change information, and the state tables of all servers are again synchronized. In addition, each server is preferably adapted to add or remove parameter columns from its state tables, so that load-balancing algorithms applied by the servers may change over time and take account of different combinations of parameters.
<figref idref="DRAWINGS">FIG. 5A</figref> illustrates a preferred embodiment for responding to a request for content from, for example, a client set-top box. In step <b>510</b>, the content request is received by a business management system (BMS) of the digital media delivery system. The business management system is preferably adapted to provide billing, authentication, conditional access, programming guide information, and other functions to client set-top boxes.
In step <b>512</b>, the business management system authenticates the client and bills the client for the request. In step <b>514</b>, the business management system designates one of the media servers of the digital media delivery system to act as a director for this request. The role of the director is to select an appropriate server to deliver the requested content to the client, as described below. In a preferred embodiment, the director may be selected by the business management system on a rotating basis. In an alternative preferred embodiment, the director may be selected on the basis of server load (i.e., the server with lowest current load is designated to act as director for the request).
In step <b>516</b>, the server designated to act as director for this request selects a server from its load-balancing group to deliver the requested content to the client. As described below, this server may be the director itself or another server in its group. Preferred embodiments for making this selection are described below in connection with <figref idref="DRAWINGS">FIGS. 5B and 6</figref>.
In step <b>518</b>, the server selected to deliver the content sets up a streaming session and notifies the business management system that it is ready to stream the requested content to the client. In step <b>520</b>, the business management system directs the client to the IP address of the selected server and delivery of the requested content is commenced.
In an alternative preferred embodiment, after selecting a server to act as director for a request, the business management system provides the director's IP address directly to the client. In this embodiment, the client contacts the director which selects a server to provide the requested content and then provides that server's IP address to the client when the streaming session is set up.
One preferred embodiment that may be utilized by a director for selecting a server to deliver requested content is now described in connection with <figref idref="DRAWINGS">FIG. 5B</figref>. As shown in <figref idref="DRAWINGS">FIG. 5B</figref>, in step <b>530</b>, a content request is received by the server designated by the business management system to act as director for the request. In step <b>540</b>, the director identifies the requested content and determines whether or not the director has this content available. If the content is available from the director itself (step <b>550</b>), it designates itself to deliver the requested content (step <b>555</b>). Otherwise, in step <b>560</b>, the server examines its state table <b>310</b> to see if the content is available from another server in its load-balancing group. If the content is not available in the group, the director rejects the request (step <b>565</b>). Otherwise, as described above, the director instructs the selected server to set up a streaming session for the client and redirects the client to submit the request to that server when the session is set up (step <b>570</b>). Two alternatives to step <b>570</b> are described respectively in connection with <figref idref="DRAWINGS">FIGS. 5C and 5D</figref>. In the first alternative, the director forwards the request to the selected server which directly respond to the client when the streaming session is set up. In the second alternative, the director redirects the client to the selected server and the client directly requests establishment of a streaming session from the selected server.
This first alternative is illustrated in an exemplary communication block diagram shown in <figref idref="DRAWINGS">FIG. 5C</figref>. More specifically, as shown in <figref idref="DRAWINGS">FIG. 5C</figref>, in step <b>530</b>, server A receives a request from the client. In step <b>570</b>, server A concludes that a different server in the group can service the request (server C in this example), and forwards the request to that server. Server C processes the forwarded request as if it were received directly from the client, and performs steps <b>540</b>-<b>555</b> shown in <figref idref="DRAWINGS">FIG. 5B</figref> to deliver the requested content to the client.
Turning to <figref idref="DRAWINGS">FIG. 5D</figref> which illustrates the second alternative, in step <b>530</b>, server A receives a request from the client. In step <b>570</b>, server A concludes that a different server in the group can service the request (server C in this example), and sends an acknowledgment to the client, redirecting the client to that server. The client then retransmits its request to server C (step <b>530</b>). Server C processes the request, and performs steps <b>540</b>-<b>555</b> shown in <figref idref="DRAWINGS">FIG. 5B</figref> to deliver the requested content to the client.
In a preferred embodiment, content may be replicated on multiple servers in a load-balancing group to satisfy request volumes that may exceed a single server's capacity. Moving or copying content from one server to another in a load-balancing group may also be used as a strategy to further distribute load within the group. <figref idref="DRAWINGS">FIG. 6</figref> illustrates a preferred system and method for load-balancing content requests in a multi-server system with replicated content.
As shown in <figref idref="DRAWINGS">FIG. 6</figref>, in step <b>610</b>, a client makes a request for content. The request is forwarded to a business management system which designates a server as director for this request and forwards the request to the director, as described above. In step <b>620</b>, the director analyzes the request to identify the requested content. In step <b>630</b>, the director determines if the requested content is present in the load-balancing group by consulting its state table. If the content is not present in the load-balancing group, the director rejects the request (step <b>640</b>).
In an alternative embodiment, the director may forward the request to a server in another load-balancing group. This alternative, however, suffers from significant drawbacks, because the director in the present embodiment has no knowledge whether the content is present in the other load-balancing groups, and a poor level of service may result depending upon the ability of a second load-balancing group to provide the content. To overcome this drawback, each server may be provided with additional state tables with information concerning servers in other load-balancing groups. Alternatively, all servers in the system may be designated as belonging to a single load-balancing group. These alternatives, however, present their own disadvantages including increased overhead to update and maintain state tables.
Returning to <figref idref="DRAWINGS">FIG. 6</figref>, if the content is available in the load-balancing group, the server examines its state table to identify those servers in the group that have the requested content (step <b>650</b>). In step <b>660</b>, the server applies a load-balancing algorithm to choose a server in its group to supply the requested content. One preferred embodiment of such an algorithm is described in more detail below. In step <b>670</b>, the client request is redirected or forwarded to the selected server as described above in connection with <figref idref="DRAWINGS">FIGS. 5B-5D</figref>. In step <b>680</b>, the selected server delivers the requested content to the client. In step <b>690</b>, the selected server updates its state table to reflect corresponding changes in its load and other parameters and communicates these state-table changes to the other servers in its load-balancing group.
A preferred embodiment of a load-balancing algorithm for selecting a server to deliver requested content is illustrated in <figref idref="DRAWINGS">FIG. 7</figref>. As shown in <figref idref="DRAWINGS">FIG. 7</figref>, in step <b>710</b>, the director examines its state table to identify all servers in its load-balancing group that have the requested content and are operational (referred to hereafter as target servers).
In step <b>720</b>, the server calculates a load factor for each of the target servers from a weighted sum of parameters indicative of load. In a preferred embodiment, the parameters used to calculate the load factor for each server are: incoming streaming bandwidth, outgoing streaming bandwidth, total storage usage, memory usage, and CPU utilization.
In step <b>730</b>, the server determines whether any target servers have exceeded a parameter threshold limit. For example, a target server may have an abundance of outgoing streaming bandwidth available, but the server's CPU utilization parameter may be very high and exceed the threshold limit established for that parameter. This target server would therefore not be a preferred choice to serve the requested content. As used herein, the term available servers refers to target servers that have not exceeded any threshold limits.
In step <b>740</b>, the server determines if there are any available servers. If so, in step <b>750</b>, the server chooses the available server having the lowest load factor to deliver the requested content. If not, then in step <b>760</b>, the server chooses the target server having the lowest load factor from all target servers.
<figref idref="DRAWINGS">FIG. 8</figref> shows an exemplary state table suitable for illustrating the above-described process of <figref idref="DRAWINGS">FIG. 7</figref>. For purposes of the present example, it is assumed that the three listed servers A-C are members of load-balancing group <b>1</b> and that each maintains the exemplary state table of <figref idref="DRAWINGS">FIG. 8</figref>. It is further assumed that a client places a request to view the asset “Dare Devil,” a feature-length film.
The director, assume server B, examines its state table and determines that the content for “Dare Devil” is stored on servers A and C. Since servers A and C are up, they are the target servers.
Server B then calculates the load factor for each of the target servers. The load factor is preferably defined to be a weighted average of parameters. For the purpose of this example, it is assumed that the bandwidth capacity, both incoming and outgoing, is 500, and the load factor is expressed as an average of each parameter, measured in percent capacity. Thus, server B would determine the load factor of server A as (4700/500+27500/500+34+37+40)/5%=35%, and the load factor of server C as (1300/500+39600/500+56+60+64)/5%=52.4%.
Next, server B determines whether both servers are available. For the purpose of this example, it is assumed that the threshold limit set for each parameter on each server is 75%. Since no threshold limits are exceeded by any target server, servers A and C are both available servers. Since there is at least one available server, server B chooses the server with the lowest load factor, namely server A.
As server A starts supplying the “Dare Devil” content, it updates its state-table parameters to reflect this fact (e.g., overall load, bandwidth, etc.). Server A preferably broadcasts these changes to all other servers in its load-balancing group, as described above.
While the invention has been described in conjunction with specific embodiments, it is evident that numerous alternatives, modifications, and variations will be apparent to persons skilled in the art in light of the foregoing description.
Contents6
11 sheets
Sheet 1 Sheet 2 Sheet 3 Sheet 4 Sheet 5 Sheet 6 Sheet 7 Sheet 8 Sheet 9 Sheet 10 Sheet 11
Every citation, both waysCites: the store holds 78 of 79
| Document | Relation | Office | Cited during |
|---|---|---|---|
| US2009327618A1 | Cited by | United States of America | Pre-grant |
| US8499096B2 | Cited by | United States of America | Applicant |
| US8102874B2 | Cited by | United States of America | Search report |
| US8886830B2 | Cited by | United States of America | Search report |
| US2022272011A1 | Cited by | United States of America | Search report |
| US2012044532A1 | Cited by | United States of America | Pre-grant |
| US2009144400A1 | Cited by | United States of America | Pre-grant |
| US8301776B2 | Cited by | United States of America | Search report |
| US8516081B2 | Cited by | United States of America | Search report |
| US2011066729A1 | Cited by | United States of America | Pre-grant |
| US11695670B2 | Cited by | United States of America | Search report |
| US2011047252A1 | Cited by | United States of America | Pre-grant |
| US2014052807A1 | Cited by | United States of America | Pre-grant |
| US8825807B2 | Cited by | United States of America | Search report |
| US2009138601A1 | Cited by | United States of America | Pre-grant |
| US2011202596A1 | Cited by | United States of America | Pre-grant |
| US2013318195A1 | Cited by | United States of America | Pre-grant |
| US2010257257A1 | Cited by | United States of America | Pre-grant |
| US11362915B2 | Cited by | United States of America | Search report |
| US8103824B2 | Cited by | United States of America | Search report |
| CN103629132A | Cited by | China | Search report |
| US8296458B2 | Cited by | United States of America | Search report |
| US2002002622A1 | Cites | United States of America | Search report |
| US2002010783A1 | Cites | United States of America | Applicant |
| US2002040402A1 | Cites | United States of America | Applicant |
| US2002059371A1 | Cites | United States of America | Applicant |
| US2002116481A1 | Cites | United States of America | Search report |
| US2002120743A1 | Cites | United States of America | Search report |
| US2002161890A1 | Cites | United States of America | Applicant |
| US2002169827A1 | Cites | United States of America | Applicant |
| US2003055910A1 | Cites | United States of America | Applicant |
| US2003088811A1 | Cites | United States of America | Search report |
| US2003115346A1 | Cites | United States of America | Search report |
| US2003158908A1 | Cites | United States of America | Applicant |
| US2003195984A1 | Cites | United States of America | Search report |
| US2004010588A1 | Cites | United States of America | Applicant |
| US2004024941A1 | Cites | United States of America | Applicant |
| US2004088412A1 | Cites | United States of America | Search report |
| US2004093288A1 | Cites | United States of America | Applicant |
| US2004125133A1 | Cites | United States of America | Search report |
| GB2342263A | Cites | United Kingdom | Search report |
| GB2343348A | Cites | United Kingdom | Applicant |
| US5353430A | Cites | United States of America | Applicant |
| US5561823A | Cites | United States of America | Applicant |
| US5586291A | Cites | United States of America | Applicant |
| US5592612A | Cites | United States of America | Applicant |
| US5761458A | Cites | United States of America | Applicant |
| US5774668A | Cites | United States of America | Applicant |
| US5809239A | Cites | United States of America | Applicant |
| US5996025A | Cites | United States of America | Applicant |
| US6038641A | Cites | United States of America | Search report |
| US6070191A | Cites | United States of America | Applicant |
| US6092178A | Cites | United States of America | Applicant |
| US6148368A | Cites | United States of America | Applicant |
| US6182138B1 | Cites | United States of America | Applicant |
| US6185598B1 | Cites | United States of America | Search report |
| US6185619B1 | Cites | United States of America | Applicant |
| US6189080B1 | Cites | United States of America | Applicant |
| US6223206B1 | Cites | United States of America | Applicant |
| US6327614B1 | Cites | United States of America | Applicant |
| US6370584B1 | Cites | United States of America | Search report |
| US6377996B1 | Cites | United States of America | Search report |
| US6466978B1 | Cites | United States of America | Applicant |
| US6535518B1 | Cites | United States of America | Applicant |
| US6587921B2 | Cites | United States of America | Applicant |
| US6665704B1 | Cites | United States of America | Applicant |
| US6718361B1 | Cites | United States of America | Applicant |
| US6728850B2 | Cites | United States of America | Applicant |
| US6748447B1 | Cites | United States of America | Search report |
| US6760763B2 | Cites | United States of America | Applicant |
| US6799214B1 | Cites | United States of America | Search report |
| US6862624B2 | Cites | United States of America | Search report |
| US6986018B2 | Cites | United States of America | Applicant |
| US7020745B1 | Cites | United States of America | Search report |
| US7043558B2 | Cites | United States of America | Applicant |
| US7080158B1 | Cites | United States of America | Search report |
| US7099915B1 | Cites | United States of America | Applicant |
| US7213062B1 | Cites | United States of America | Search report |
| US7233978B2 | Cites | United States of America | Applicant |
| US7398312B1 | Cites | United States of America | Applicant |
| US7500055B1 | Cites | United States of America | Applicant |
| US20020002622A1 | Cites | United States of America | Search report |
| US20020010783A1 | Cites | United States of America | Third party observation |
| US20020040402A1 | Cites | United States of America | Third party observation |
| US20020059371A1 | Cites | United States of America | Third party observation |
| US20020116481A1 | Cites | United States of America | Search report |
| US20020120743A1 | Cites | United States of America | Search report |
| US20020161890A1 | Cites | United States of America | Third party observation |
| US20020169827A1 | Cites | United States of America | Third party observation |
| US20030055910A1 | Cites | United States of America | Third party observation |
| US20030088811A1 | Cites | United States of America | Search report |
| US20030115346A1 | Cites | United States of America | Search report |
| US20030158908A1 | Cites | United States of America | Third party observation |
| US20030195984A1 | Cites | United States of America | Search report |
| US20040010588A1 | Cites | United States of America | Third party observation |
| US20040024941A1 | Cites | United States of America | Third party observation |
| US20040088412A1 | Cites | United States of America | Search report |
| US20040093288A1 | Cites | United States of America | Third party observation |
| US20040125133A1 | Cites | United States of America | Search report |
| GB2343348 | Cites | United Kingdom | Third party observation |
3 members in 1 office
Priority claims6
| Document | Office | Kind | Date |
|---|---|---|---|
| 60942603 | United States of America | A | |
| 60942603 | United States of America | A | |
| 46861306 | United States of America | A | |
| 10609426 | – | – | – |
| US20030609426 | – | – | – |
| US20060468613 | – | – | – |
Members3
| Document | Office | Kind | |
|---|---|---|---|
| US2007124476A1 | United States of America | A1 | |
| US7680938B2This record | United States of America | B2 | |
| US7912954B1 | United States of America | B1 |
98 transactions on the USPTO file
Allowed after 3 non-final rejections, 3 final rejections and 3 RCEs.
- Non-final rejections
- 3
- Final rejections
- 3
- RCEs
- 3
- Appeals
- 0
Over time
Point at a mark for the transactionTransactions
| Event | Code | |
|---|---|---|
| Payment of Maintenance Fee, 12th Year, Large EntityM1553 | M1553 | |
| Payment of Maintenance Fee, 8th Year, Large EntityM1552 | M1552 | |
| Change in Power of Attorney (May Include Associate POA)PA.. | PA.. | |
| Correspondence Address ChangeC.AD | C.AD | |
| Post Issue Communication - Certificate of CorrectionN423 | N423 | |
| Post Issue Communication - Certificate of CorrectionN423 | N423 | |
| Recordation of Patent Grant MailedPGM/ | PGM/ | |
| Patent Issue Date Used in PTA CalculationAllowedPTAC | PTAC | |
| Issue Notification MailedAllowedWPIR | WPIR | |
| Dispatch to FDCD1935 | D1935 | |
| Application Is Considered Ready for IssuePILS | PILS | |
| Issue Fee Payment VerifiedN084 | N084 | |
| Issue Fee Payment ReceivedIFEE | IFEE | |
| Mail Examiner's AmendmentMEX.A | MEX.A | |
| Mail Notice of AllowanceAllowedMN/=. | MN/=. | |
| Examiner's Amendment CommunicationEX.A | EX.A | |
| Notice of Allowance Data Verification CompletedAllowedN/=. | N/=. | |
| Paralegal or electronic terminal disclaimer approvedP574 | P574 | |
| Reference capture on IDSRCAP | RCAP | |
| Information Disclosure Statement (IDS) FiledM844 | M844 | |
| Information Disclosure Statement consideredIDSC | IDSC | |
| Information Disclosure Statement (IDS) FiledWIDS | WIDS | |
| Terminal Disclaimer FiledDIST | DIST | |
| Date Forwarded to ExaminerFWDX | FWDX | |
| Date Forwarded to ExaminerFWDX | FWDX | |
| Disposal for a RCE / CPA / R129AbandonedABN9 | ABN9 | |
| Request for Continued Examination (RCE)RCEX | RCEX | |
| Workflow - Request for RCE - BeginBRCE | BRCE | |
| Mail Examiner Interview Summary (PTOL - 413)MEXIN | MEXIN | |
| Examiner Interview Summary Record (PTOL - 413)EXIN | EXIN | |
| Mail Final Rejection (PTOL - 326)Final rejectionMCTFR | MCTFR | |
| Final RejectionFinal rejectionCTFR | CTFR | |
| Electronic Information Disclosure StatementEIDS. | EIDS. | |
| Information Disclosure Statement consideredIDSC | IDSC | |
| Information Disclosure Statement (IDS) FiledWIDS | WIDS | |
| Date Forwarded to ExaminerFWDX | FWDX | |
| Response after Non-Final ActionA... | A... | |
| Mail Examiner Interview Summary (PTOL - 413)MEXIN | MEXIN | |
| Examiner Interview Summary Record (PTOL - 413)EXIN | EXIN | |
| Mail Non-Final RejectionNon-final rejectionMCTNF | MCTNF | |
| Non-Final RejectionNon-final rejectionCTNF | CTNF | |
| Case Docketed to Examiner in GAUDOCK | DOCK | |
| Date Forwarded to ExaminerFWDX | FWDX | |
| Date Forwarded to ExaminerFWDX | FWDX | |
| Disposal for a RCE / CPA / R129AbandonedABN9 | ABN9 | |
| Request for Continued Examination (RCE)RCEX | RCEX | |
| Workflow - Request for RCE - BeginBRCE | BRCE | |
| Information Disclosure Statement consideredIDSC | IDSC | |
| Reference capture on IDSRCAP | RCAP | |
| Information Disclosure Statement (IDS) FiledM844 | M844 | |
| Information Disclosure Statement (IDS) FiledWIDS | WIDS | |
| Case Docketed to Examiner in GAUDOCK | DOCK | |
| Mail Final Rejection (PTOL - 326)Final rejectionMCTFR | MCTFR | |
| Final RejectionFinal rejectionCTFR | CTFR | |
| Date Forwarded to ExaminerFWDX | FWDX | |
| Correspondence Address ChangeC.AD | C.AD | |
| Response after Non-Final ActionA... | A... | |
| Mail Examiner Interview Summary (PTOL - 413)MEXIN | MEXIN | |
| Examiner Interview Summary Record (PTOL - 413)EXIN | EXIN | |
| Correspondence Address ChangeC.AD | C.AD | |
| Mail Non-Final RejectionNon-final rejectionMCTNF | MCTNF | |
| Non-Final RejectionNon-final rejectionCTNF | CTNF | |
| Date Forwarded to ExaminerFWDX | FWDX | |
| Date Forwarded to ExaminerFWDX | FWDX | |
| Date Forwarded to ExaminerFWDX | FWDX | |
| Disposal for a RCE / CPA / R129AbandonedABN9 | ABN9 | |
| Case Docketed to Examiner in GAUDOCK | DOCK | |
| Request for Continued Examination (RCE)RCEX | RCEX | |
| Request for Extension of Time - GrantedXT/G | XT/G | |
| Workflow - Request for RCE - BeginBRCE | BRCE | |
| Mail Advisory Action (PTOL - 303)MCTAV | MCTAV | |
| Advisory Action (PTOL-303)CTAV | CTAV | |
| Date Forwarded to ExaminerFWDX | FWDX | |
| Response after Final ActionA.NE | A.NE | |
| Mail Final Rejection (PTOL - 326)Final rejectionMCTFR | MCTFR | |
| Final RejectionFinal rejectionCTFR | CTFR | |
| Date Forwarded to ExaminerFWDX | FWDX | |
| Response after Non-Final ActionA... | A... | |
| Correspondence Address ChangeC.AD | C.AD | |
| Mail Non-Final RejectionNon-final rejectionMCTNF | MCTNF | |
| Non-Final RejectionNon-final rejectionCTNF | CTNF | |
| Case Docketed to Examiner in GAUDOCK | DOCK | |
| Case Docketed to Examiner in GAUDOCK | DOCK | |
| PG-Pub Issue NotificationPG-ISSUE | PG-ISSUE | |
| IFW TSS Processing by Tech Center CompleteTSSCOMP | TSSCOMP | |
| Application Dispatched from OIPEOIPE | OIPE | |
| Application Is Now CompleteCOMP | COMP | |
| Payment of additional filing fee/PreexamFLFEE | FLFEE | |
| A statement by one or more inventors satisfying the requirement under 35 USC 115, Oath of the ApplicOATHDECL | OATHDECL | |
| Information Disclosure Statement consideredIDSC | IDSC | |
| Reference capture on IDSRCAP | RCAP | |
| Information Disclosure Statement (IDS) FiledM844 | M844 | |
| Information Disclosure Statement (IDS) FiledWIDS | WIDS | |
| Preliminary AmendmentA.PE | A.PE | |
| Cleared by L&R (LARS)L128 | L128 | |
| Referred to Level 2 (LARS) by OIPE CSRL198 | L198 | |
| IFW Scan & PACR Auto Security ReviewSCAN | SCAN | |
| Initial Exam Team nnIEXX | IEXX |
11 legal events, as the office reported them to INPADOC
Over the term
Point at a mark for the eventEvents
| Event | Code | |
|---|---|---|
| Maintenance fee paymentMAFP | MAFP | |
| Maintenance fee paymentMAFP | MAFP | |
| AssignmentAS | AS | |
| Fee paymentFPAY | FPAY | |
| Certificate of correctionCC | CC | |
| Certificate of correctionCC | CC | |
| Information on status: patent grantGrantedPATENTED CASESTCF | STCF | |
| AssignmentAS | AS | |
| AssignmentAS | AS | |
| AssignmentAS | AS | |
| AssignmentAS | AS |
Numbers
- Publication
- 07680938
- Publication, DOCDB
- 7680938
- Publication, EPODOC
- US7680938
- Application
- 11468613
- Application, DOCDB
- 46861306
- Application, EPODOC
- US20060468613
Titles
- English
- Video on demand digital server load balancing
Patent term adjustment
- Applicant delay
- −112 days
- Net adjustment
- 0 days
Classification
- CPC, 3
- H04L67/1008
- H04L67/1012
- H04L67/1001
- IPC, 1
- G06F15 173
- USPC, 1
- 709226000