Data service system
Abstract
[Task] When the content server in the data service system becomes overloaded, the access request to the server is automatically controlled to maintain the safe load state.
Solution.The data service system includes an adaptive load control system in addition to the content server. The content server stores multiple content files in full content format or adaptive content format for access by external access requests. Adaptive is a content format that requires less resources than the complete form because it is a service. The adaptive load control system is to access the adaptive content format file so that the content server remains in a safe load state when the content server is overloaded, and if the load is not overloaded. Modify each access request to access the full form.

Term
Term ended
Projected expiry passed 22 March 2020, 6.5 years ago.
- Priority
- Filed
- Published
- Projected expiry
- Today
1 claim: 1 independent, 0 dependent
- 1【特許請求の範囲】 【請求項1】外部アクセス要求によるアクセスに備えて完全コンテンツ形式またはサービスのため完全形式ほど資源を必要としない適応型コンテンツ形式で各々記憶される複数のコンテンツ・ファイルを記憶するコンテンツ・サーバと、 上記コンテンツ・サーバにアクセス要求を渡すため上記コンテンツ・サーバに接続された適応型負荷制御システムと、 を備え、 上記適応型負荷制御システムが、上記コンテンツ・サーバが過負荷状態にある時、コンテンツ・サーバが安全負荷状態に維持されるように、上記適応型コンテンツ形式の対応するコンテンツ・ファイルにアクセスするようにアクセス要求を修正する、 データ・サービス・ネットワーク・システムにおけるデータ・サービス・システム。
151 paragraphs in 1 section, as filed
Description: TECHNICAL FIELD [Detailed description of the invention]
【0001】
[Technical field to which the invention belongs]
The present invention generally relates to Internet / Intranet systems and processes user access requests specifically for less resource intensive content when the web server is overloaded. It is about an adaptive web server.
【0002】
[Conventional technology]
One example of a data access network system is an internet or intranet network system. Internet / intranet network systems typically include a number of data service systems that are interconnected through interconnect networks. A data service system typically includes a web or content server that serves content for a variety of customers or applications. The customer is the owner of the content serviced by the data service system so that the subscriber or user can access the content through the computer terminal. Content servers typically utilize Internet applications such as email, electronic bulletin boards, newsgroups and worldwide web (ie WWW) access. The content served is in the form of content sites within the scope of the content server. Each site contains a large number of pages (eg WWW pages). Typically, one content site is for one customer on that server, while one particular customer can own many content sites. Content servers are also referred to as web servers.
【0003】
Often, a web page is formed from a large amount of data or executable program files such as text files, image files, audio files, video files and application program files. Each of the files is called an "object". Web pages can be accessed by users from user terminals (eg, personal computer systems or web access devices) connected to any data service system.
【0004】
As is well known, access to web pages via the Internet is typically built around a protocol called the Hyper Text Transfer Protocol (HTTP). The HTTP protocol is a request-response protocol. When a user on a client device specifies a particular web page, at least one request is generated. The number of requests depends on the complexity of the specified web page. As mentioned above, a web page can contain one or more "objects". A multi-object page is aesthetically pleasing, but each object requires its own request and its own response. Therefore, the time required for each round trip of the request and response plays a role in determining the total time the user has to wait to see the complete web page.
【0005】
[Problems to be Solved by the Invention]
The web server can be accessed by multiple users at the same time. Web servers typically handle user access requests in first-in, first-out (FIFO) order. One drawback of this type of prior art web server configuration is that the web server does not have a protection mechanism against overloaded conditions. As is well known, web servers typically have restrictions on the number of requests they can handle per second. In addition, web servers also have constraints on the response time and bandwidth they can service. Under high load, the web server can be overwhelmed by demands, resulting in longer response times to users and poorer performance. Thus, when the number of user access requests to the web server significantly exceeds the processing capacity of the web server, that is, when the load becomes excessive, the web server takes an unbearably long time. Either the request is processed or it simply becomes unresponsive. A sudden increase in server response time causes a timeout in the user connection, and the server is considered to be stopped or destroyed. In a constantly criticized application such as e-commerce, perceived server outages can lead to sales and other financial losses.
【0006】
As described above, it is necessary to improve the performance of the web server so that the overload does not continue or stop functioning when the web server becomes overloaded.
【0007】
[Means for solving problems]
In order to solve the problems of the present invention, the data service system in the data service network system provided by the present invention requires as much resources as a complete content format or a complete format for a service in preparation for access by an external access request. It includes a content server that stores multiple content files, each stored in an adaptive content format, and an adaptive load control system that is connected to the content server to pass access requests to the content server. In such a configuration, the adaptive load control system of the present invention provides the corresponding content of the adaptive content format so that the content server is maintained in a safe load state when the content server is overloaded. Modify the access request to access the file.
【0008】
One function of the present invention is to improve the performance of the web server. In addition, another function of the present invention improves the performance of the web server by processing the user access request with the content having a relatively low resource intensity when the web server becomes overloaded. Yet another feature of the present invention is an adaptive web that directs user access requests to less resource-intensive content to prevent overloading or outages when the web server is overloaded. Provide a server.
【0009】
The data service system in the data access network system of the present invention includes a content server that stores content files for access by an external access request. Each of the content files is stored in a complete content format and one or more adaptive or downgraded content formats that are less resource intensive than the complete content format. The data service system also includes an adaptive load control system connected to the content server to pass access requests to the content server. The adaptive load control system provides the corresponding content in one of the adaptive content formats so that the content server can be maintained in a safe load condition when the content server is overloaded. Modify the access request to access the file.
【0010】
Further provided is a method of maintaining a content server under safe load in a data service system of a data access network system that has a content server that stores content files for access by external access requests. .. The method includes determining the load state of the content server when the data service system receives an access request to access one of the content files stored on the content server. This step can be performed inside or outside the content server. One of the less resource-intensive adaptive content formats than the full content format so that the content server can be kept under load when it is determined to be overloaded. The access request is modified to access the corresponding content file in.
【0011】
BEST MODE FOR CARRYING OUT THE INVENTION
Figure 1 shows the data access network system 10. In one embodiment, the network system 10 is an internet system. In another embodiment, the network system 10 is an intranet system. Alternatively, the data access network system 10 may be any other known network system that utilizes a known communication protocol. As shown in FIG. 1, in the data service system 20, a large number of user terminals 11a to 11n are connected through the interconnection network 12. The user on each of the user terminals 11a to 11n can access the data service system 20 to receive the services provided by the data service system 20. The user in each of the user terminals 11a to 11n also accesses the data service system 15 via the data service system 20 and the Internet 13 to receive the services provided by the data service system 15. Can be done. User terminals 11a to 11n are also called client devices.
【0012】
Each of the user terminals 11a to 11n is located in the user's house, school or office. Each of the user terminals 11a to 11n is a network access application program (eg, Netscape® Navigator) that allows the user to access the data services provided by the data service systems 15 and 20. Web browser application programs such as).
【0013】
Each of the user terminals 11a to 11n can also be a computer system or other electronic device with data processing capabilities (eg, a web TV). The interconnect network 12 can be any known network such as Ethernet®, ISDN, T-1 or T-2 link, FDDI (Fiber Distributed Data Interface), cable or wireless LMDS network or telephone line network.
【0014】
Figure 1 shows only two data service systems 15 and 20 for the data access network system 10 for illustration purposes only. The data access network system 10 can actually include more data service systems. In addition, the Internet 13 in FIG. 1 is formed by a large number of data service systems, all connected by one network. Data communication between all data service systems is carried out using pre-determined communication protocols for Internet / intranet communication. In one embodiment, the communication protocol is the hypertext transport protocol (ie HTTP). Alternatively, other known communication protocols can be used for Internet / intranet communication.
【0015】
Each of the data service systems 15 and 20 has essentially the same functional structure. Each of the data service systems 15 and 20 is to a user or subscriber connecting to the data access system 10 via one of the data service systems 15 and 20 (web, news or advertising). It can be used by Internet / Intranet service providers (ie, ISPs) to provide data services (such as) and other services (such as e-commerce, e-mail).
【0016】
Each of the data service systems 15 and 20 has a number of servers (eg, web servers, email servers, news servers, electronic commerce servers, domain name servers, address assignment servers, process servers, advertising servers and Includes session manager server). Web servers, email servers, news servers, electronic commerce servers and advertising servers are collectively referred to as local service servers or content servers. Content servers typically store a large number of content files, including hypertext markup language (HTML) web pages, GIF images, video clips, and so on. Data transfer to and from the content server is enabled by transmission protocols such as Transmission Control Protocol (TCP) and User Datagram Protocol (UDP). Content servers support a variety of Internet applications to provide services such as access to the WWW, email, electronic bulletin boards, chat rooms, newsgroups and e-commerce. Using currently commonly available web browsers and other client applications, users can use content files stored on content servers (eg, web pages, news, etc.) via their respective user terminals. You can access the image (email).
【0017】
As mentioned above, access to the content server is centered around the HTTP protocol, which is a request-response protocol. When a user on a user terminal wants to access a content file stored on a content server in one of the data service systems 15 and 20, at least one request is generated and sent to the content server. Is done. The content server can handle multiple requests at the same time. However, the content server is limited in the number of requests it can handle per second. When the number of requests received by the content server significantly exceeds the content server's limits, the content server becomes overloaded, resulting in at least unbearably long response times and poor performance for the user. .. This means that if the request rate increases beyond the server capacity, server performance will drop dramatically, potentially leading to outages. In addition, if the client makes a request but does not receive any response, the client is likely to make more subsequent requests. This rapidly increases the number of requests received by the content server.
【0018】
In order to solve such a problem, each of the data service systems 15 and 20 according to one embodiment of the present invention, sets the load state of each content server so that each content server does not become overloaded. Includes an adaptive load control mechanism to control. Alternatively, the load control mechanism of the present invention can also be applied to each of the other servers in the data service systems 15 and 20. The details of the adaptive load control mechanism are described below with reference to FIGS. 2 to 6.
【0019】
As shown in FIG. 2, a data service system 30 according to one embodiment of the present invention includes a content server 31. The data service system 30 also includes an adaptive load control system 40 that controls the load state of the content server 31 according to one embodiment of the invention. Therefore, the data service system 30 can be referred to as an adaptive data service system or an adaptive web server. The data service system 30 may be a shift of the data service system of FIG. 1 (eg, data service system 15 or 20).
【0020】
Although details will be described later, the content server 31 stores a content file or a dynamically executable code / program for access by an access request. Thus, in the following description, content file means (1) static content file, (2) dynamic content file and (3) executable program / code. Each of the content files is stored in full content format and in adaptive or downgraded content format. Content files in adaptive content format are smaller in size and less resource intensive than the same file in full content format.
【0021】
Upon receiving the request, the adaptive load control system 40 passes the access request to the content server 31. At that time, if the content server 31 is overloaded or close to overloaded, the adaptive load control system 40 modifies the access request to access the corresponding content file in the adaptive content format. To do. If the adaptive load control system 40 determines that the content server 31 is not overloaded, the adaptive load control system 40 modifies the access request to access the corresponding content file in full content format. As a result, the content server 31 is maintained in a safe load state. Alternatively, the adaptive load control system 40 is not based on the load state of the content server 31, but on the client's capabilities (such as the client's network connectivity, latency, or client's display capability). Server 31 can also be adapted.
【0022】
Details of the data service system 30 are described below with reference to FIGS. 2 to 6. The data service system 30 can be implemented in a computer system or other data processing system. The computer system that implements the data service system 30 is a server computer system, a workstation computer system, a personal computer system, a mainframe computer system, a notebook computer system or other computer. It is a system. The data service system 30 includes a network interface 35 that interfaces with the adaptive load control system 40. Through it, the adaptive load control system 40 interfaces with the content server 31. The network interface 35 serves as an interface between the external network (not shown) and the data service system 30. Interface 35 can be implemented by any known network interface technology and will not be described in detail. The network interface 35 receives an external request to the content server 31 and passes the request to the content server 31 via the adaptive load control system 40.
【0023】
The content server 31 can be any type of content server that stores a large number of content files. Each of the content files can be accessed by an access request. The content server 31 can also include a large number of content sites, each storing a large number of content files in case of access by multiple access requests. Multiple content sites may belong to different content providers or customers whose interests may potentially conflict. In this case, the adaptive load control system 40 allows the content server 31 to be a large number of performance-isolated virtual content servers (more on this later). Each of the servers has the ability to provide an independent performance guarantee regardless of the load on the other virtual server.
【0024】
The content server 31 can be either a static server or a dynamic server. In one embodiment, the content server 31 is a static content server that stores only static content files. In another embodiment, the content server 31 can store both static and dynamic content files. As is well known, web content is generally categorized as static content, such as files, or dynamic content, such as cgi scripts. Dynamic content can be generated at runtime by a backend engine (such as a database engine) that is independent of the server itself. No adaptation in the back-end engine is handled by the adaptive load control system 40.
【0025】
According to one embodiment of the invention, each of the content files stored on the content server 31 is stored in a plurality of formats or versions. For example, a content file can be stored in its original full content format or version. Also, the same content file is stored on content server 31 in a downgraded or adaptive format or version. Figure 4 shows one example of such a form. As shown in Figure 4, the content file in adaptive or downgraded format 72 is smaller in size and inferior in visibility quality compared to the same file in full content format or version 71. .. This means that downgraded content files require less resources (eg, bandwidth) to provide services than the corresponding content files in full content format.
【0026】
For example, one content file can be displayed in 64x64 pixel resolution, 128x128 resolution, 256x256 resolution, or any other resolution, allowing content files in different formats. The 256x256 pixel resolution version of the content file clearly has more image data and requires more transmission resources than the 128x28 pixel resolution version of the same content file. If the full content version of the content file has a 256x256 pixel resolution, then the 128x128 pixel resolution or 64x64 pixel resolution version is a downgraded version of the same content file. In short, a downgraded content file contains lower resolution images, less embedded images, simple pages without backgrounds or fewer hyperlinks than the same content file in full content format. Alternatively, each of the content files stored on the content server 31 has several downgraded versions of the same content file. In the following description, the content server 31 has only two versions, for example, a full content version and one downgraded version, for each of the stored content files of the content server 31 for illustrative purposes. I will do it.
【0027】
Seeing FIG. 2 again, each access request uses a universal resource locator (ie, URL) to specify the content file stored on the content server 31. However, the URL path to the content file stored in the content server 31 does not specify which version or format of the content file to access. This means that multiple content trees contain different versions of the same URL. The path to a particular URL in a given content tree is the concatenation of the content tree name and the URL name, preceded by the name of the content server 31's root service directory. For example, two content trees are created in the root service directory "/ root": "/ full_content" (meaning full content) and "/ degraded_content" (meaning downgraded content). The "/ full_content" content tree name directs a request to access the corresponding content file in the full content version or format. The "/ degraded_content" content tree name directs a request to access the corresponding content file for the downgraded content version or format. In this case, the URL "http://www.hpl.hp.com/my_picture.jpg" is either the directory "/root/full_content/my_picture.jpg" or the directory "root/degraded_content/my_picture.jpg". Receive service from.
【0028】
This method applies even if the content file is dynamic content, such as content generated by a cgi script. Multiple content trees can contain different versions or formats of named cgi scripts (such as my_script.cgi "." / Degraded_content / "versions or formats require less resources for search scripts. This means, for example, that the "/ full_content" version or format looks for 100 matches, while looking for only the first 5 matches. Which version of the script under a given load condition The correct tree name is prepended to the script URL to determine if it should be run. For example, the URL "http://www.hpl.hp.com/my_script.cgi" would be "http: // www". Either .hpl.hp.com/root/full_content/cgi-bin/my_script.cgi "or" http://www.hpl.hp.com/root/degraded_content/cgi-bin/my_script.cgi " It will be modified as.
【0029】
Alternatively, the resource-poor "/ degraded_content / search script can be replaced with a static version or format. This is done by switching to a different content tree, for example using content server 31. An online vendor can use a dynamically generated version of the product catalog to communicate with the inventory database to display items currently in stock. When Content Server 31 is overloaded. A pre-stored static catalog version is accessed and displayed. The reason for doing this is that dynamic scripts, even downgraded or simplified ones, consume even more resources than static content. Because it does.
【0030】
The adaptive load control system 40 is connected to the content server 31. According to one embodiment of the present invention, the adaptive load control system 40 detects the overload state of the content server 31 and provides one of several alternative content qualities to provide according to the load state. Make the access request access to the file to be used. When the adaptive load control system 40 detects that the content server 31 is overloaded, the adaptive load control system 40 is downgraded in that the access request arriving at the content server 31 requires less resources to provide the service. Encourage access to the corresponding content file of. The downgrade format reduces the size of the response. The downgrade format also reduces the total number of requests because there are fewer images and other objects embedded in the simplified content. The downgrade format also removes or reduces the links embedded in the requested content file. In this way, the downgraded format requires less resources to provide the service than the complete content format.
【0031】
The adaptive load control system 40 can also deny access requests as needed. In addition, the adaptive load control system 40 supports the coexistence of multiple performance isolated virtual servers in a form in which the performance level of the virtual server and the load application decision do not affect another virtual server. Such support is needed when Content Server 31 handles multiple content sites that may belong to customers or content providers with potentially competing interests. The total load of server 31 may not be evenly distributed to such sites. In this case, the adaptive load control system 40 implements a performance isolation mechanism that is unaffected by the overload caused by another site. Using this mechanism, Content Server 31 is considered as a collection of performance-isolated virtual servers. Each performance-isolated virtual server has the ability to provide independent performance guarantees independent of overload on another virtual server. Guarantees are declared in the form of service level agreements associated with each virtual server and are used to configure the adaptive load control system 40 to allocate appropriate resource capacity. The adaptive load control system 40 makes adaptive decision-making regarding the load and resource allocation within the range of each virtual server of the content server 31.
【0032】
In addition, the adaptive load control system 40 gives priority to guaranteed service classes, while allowing non-guaranteed service classes to consume surplus capacity if available. The guaranteed class of service is defined by a service level agreement. Non-guaranteed classes are typically provided with available resources based on best efforts. Under overload conditions, the adaptive load control system 40 must first direct the non-warranty request to the downgraded content file. Arbitrary policies can be used to manage relative downgrades of client classes at different levels of importance. For example, under overloaded conditions, Class A requests must be sent to the downgraded content file before Class B requests. However, if resources remain scarce, the Class B request will be downgraded before rejecting the Class A request.
【0033】
The adaptive load control system 40 can be implemented anywhere in the data service system 30 where server requests and responses can be accessed. It implements the adaptive load control system 40 in the web server software, in the UNIX® socket library, or in the operating system of the computer system that implements the data service system 30. Means that it is possible. In addition, the adaptive load control system 40 can also be implemented at a gateway that is transparent to the server and external to it.
【0034】
Basically, the adaptive load control system 40 requires three entry points. They are (1) initialization points, (2) pre-request processing points, and (3) post-request processing points. The initialization point initializes the internally used data structure and creates an independent process that implements the various components of the adaptive load control system 40 as needed. The request preprocessing point determines the content level quality of a given request. The content adapter 41 of the adaptive load control system 40 is called from this point. The content adapter 41 can communicate with other components through shared memory. Request post-processing points are optional points for monitoring response size.
【0035】
If the server code is not available, the adaptive load control system 40 is implemented in the form of middleware (eg, dynamically linked socket libraries) used by the server 31. In a UNIX environment, a server socket is created as part of server initialization, and a listen () call is created in the socket library for incoming requests. This library routine can be modified to call the middleware initialization point. The modification can be transparent to the content server 31. The functional structure of the adaptive load control system 40 will be described in more detail below with reference to FIGS. 2 to 7.
【0036】
As shown in FIG. 2, the adaptive load control system 40 includes a requirements classification mechanism 33 connected to network interface 35. The adaptive load control system 40 also includes a load monitor 32 connected to the request classification mechanism 33. The content adapter 41 is then connected to the load monitor 32 and the content server 31. The adaptive controller 50 is connected to the content adapter 41 and the load monitor 32. Components 32, 33, 41 and 50 combine to form the adaptive load control system 40. Alternatively, the adaptive load control system 40 may include more or fewer components than described above. For example, the adaptive load control system 40 can also function without the requirements classification mechanism 33.
【0037】
The request classification mechanism 33 receives the access request from the network interface 35. The request classification mechanism 33 is used to classify received requests into a large number of classes based on different classification criteria. The request classification mechanism 33 enables priority processing for a subset of content sites or clients. The claim classification mechanism 33 requires service quality differentiation between different client classes (for example, between guaranteed and non-guaranteed clients, or between clients accessing different virtual servers). Used in client installations. In this situation, the subsequent action taken by the content adapter 41 depends on the identification of the client or the requested content. The requirements classification mechanism 33 allows multiple user classes to share the same content file or site but receive different treatments or performances. Class-based services are a mechanism that differentiates the services given to individual classes. In this way, service performance is priced based on performance or service agreement. Higher classes with greater guarantees will be priced higher than lower classes with less guarantees and more "best effort" services. The requirements classification mechanism 33 can be implemented using any known technique.
【0038】
The request classification mechanism 33 then sends the classified access request to the load monitor 32. The load monitor 32 is used to monitor the load status of the content server 31 in order to detect the overload status. The load monitor 32 can utilize a number of monitoring mechanisms to monitor the load status of the content server 31. One method of monitoring the load status of the content server 31 is to monitor the response time of the content server 31. This is because the response time is proportional to the length of the input request queue of the content server 31 (eg, the server socket listen queue in UNIX embodiments). If the server load is light, the queue tends to be short (or empty) and therefore the response time is short. When the content server 31 is overloaded, the request queue within the range of the content server 31 overflows, resulting in many times more response times. This serves as an obvious overload indicator. The advantage of this technique (which estimates the queue length by monitoring response time) is that this mechanism can be implemented outside the server process and requires code modification on content server 31. Do not do.
【0039】
Another method of monitoring whether the content server 31 is overloaded is to monitor the utilization of the server machine as well as the response time. This allows the load monitor 32 to determine how much the server is underutilized. If the utilization falls below a predetermined level, the load monitor 32 determines that the content server 31 is not in an overloaded state. In this case, to provide a faithful server load estimate, the utilization measure is not a total utilization, but a load that belongs to a valid job (such as processing a client request). There must be. For example, the presence of a low-priority process or thread that runs a busy wait loop for some events exhibits equal utilization equal to 100% even when the server is unloaded. Another concern arises when all server threads or processes are left to handle the current request and are hampered by I / O. In such cases, the derived idle time is not available because all resources are already in use and therefore should not be measured as available capacity.
【0040】
To address the concerns mentioned above, two different mechanisms have been developed in relation to the load monitor 32. The first mechanism is called the linear approximation method. In this method, the load monitor 32 utilizes a linear function of the measured request rate R and effective bandwidth BW to determine the system utilization U consumed by processing the client request. The linear function is shown as follows.
【0041】
U = aR + bBW However, the constants a and b are calculated by pre-profiling either online or offline (more on this later).
【0042】
This function can effectively estimate the load as long as the load does not exceed the server capacity. When measured under overload conditions, the linear approximation method is ineffective. Therefore, the linear approximation method can be used in combination with the response time monitoring method to determine the actual load state of the server 31. This combined load monitoring function has U = 100% if the measured response time exceeds a predetermined overload threshold. If the measured response time does not exceed a predetermined overload threshold, then U = aR + bBW.
【0043】
One advantage of this linear approximation method is that it is easy to implement. Using this method, each server process or thread Tj records its observed request rate Rj, effective bandwidth BWj and utilization Uj independently. Another monitoring process then reads the recorded values and sums them together to calculate the overall demand rate R, bandwidth BW, and utilization U. No synchronization of access to shared data structures is required as there is only one writer for any piece of data. This method enables server capacity planning for service quality assurance and provides a means of converting desired request rates and bandwidth guarantees into corresponding resource capacity allocations (eg, utilization allocations). However, this method requires the calculation of constants a and b using pre-profiling.
【0044】
An easy way to calculate the constants a and b using pre-profiling is to get some measurements of U and the corresponding R and BW. The linear regression method is then used to determine the constants a and b that best fit the equation U = aR + bB. It should be noted that R and BW can be measured online by counting the number of requests sent and the bytes returned in a given time period. Utilization U is measured using the gap estimation method described below. Given the measured R, BW, and U, the estimation theory provides a way to find a linear optimum that minimizes the error at consecutive time intervals. Estimating theory provides the equations needed to calculate and update parameters a and b online in terms of continuous new measurements.
【0045】
Alternatively, the parameters a and b can be determined by testing the content server 31 for a pre-specified workload. This test is advanced by requesting URLs of a given size at an increasing rate. This test can be continued by requesting a URL of a given size, gradually increasing the rate, until the client connection times out and marks a server overload. The maximum request rate and bandwidth reached are recorded. For the purposes of this embodiment, it can be assumed that the server utilization is 100% under this load condition. Repeated experience with requested URLs of different sizes, given different request rate and bandwidth combinations that overload the server 31. Using the set of R and BW and the point U = 100%, the line segment aR + bBW = 100 is constructed on the R and BW planes. This line segment intersects the R and BW axes at 100 / a and 100 / b, respectively, from which a and b are found.
【0046】
The second mechanism for the load monitor 32 is called the gap estimation method. This method estimates the portion of time that the content server 31 spends processing the request. Using this method, the load monitor 32 uses a global counter to track the request. The load monitor 32 increments the counter when a request is received and decrements the counter when the response exits the content server 31. The counter returns to its initial value only when there is a gap, that is, when all current requests have been processed and no additional requests have arrived yet. By summing the gaps during time T, the monitoring process or thread can calculate the total "idle time" G during period T. Idle time is the time when there are no requests in progress. Next, the utilization is estimated to be U = (TG) / T.
【0047】
The gap estimation method is shown in Figure 7. As shown in FIG. 7, the simultaneous processing of consecutive requests is interrupted by the idle time. This method does not require pre-profiling and does not require response time monitoring. Global counters require some form of locking or access synchronization as they can be updated by multiple writers.
【0048】
Referring again to FIG. 2, the load monitor 32 sends the load status information of the content server 31 to the content adapter 41 and the adaptive controller 50. In addition, the adaptive controller 50 receives a predetermined desired load value that indicates the threshold value of the overload state. Based on the comparison between the desired load value and the monitored load value, the adaptive controller 50 determines whether the content server 31 is overloaded. In case of overload, the adaptive controller 50 instructs the content adapter to modify the URL of all incoming access requests to access the corresponding content file of the downgraded content format or version. If the adaptive controller 50 determines that the content server 31 is not overloaded, the adaptive controller 50 should modify the URL of the incoming access request to access the corresponding content file in full content format or version. Instruct the content adapter.
【0049】
The desired load value for the adaptive controller 50 is, for example, a predetermined threshold T for server response time. This threshold T is set, for example, equal to (or slightly less than) the maximum server response time specified in the service level agreement, which allows adaptation when the agreement is likely to be violated. .. Without such a specification, the threshold can be derived from the preset maximum input request queue length Q and average service time S. For example, if the queue is 90% full, it is considered to indicate an overload, then the threshold T is set to T = 0.9Q × S.
【0050】
The content adapter 41 implements the adaptation by changing the directory link. Under the control of the adaptive controller 50, the content adapter 41 modifies the URL of the incoming access request to access the corresponding content file, either in full content format or in downgraded content format. For example, Content Adapter 41 uses the URL "http://www.hpl.hp.com/my_picture.jpg" to access the content file "my_picture.jpg" in full content format "http: / www". To access the .hpl.hp.com/full_content/my_picture.jpg "or the content file" my_picture.jpg "in the downgraded content format" http://www.hpl.hp.com//degraded_content/my_picture Can be converted to .jpg ". As another example If the content is dynamic content (eg http://www.hpl.hp.com/my_script.cgi), the content adapter 41 will use that URL to access the full content file "http: // To access www.hpl.hp.com/full_content/cgi-bin/my_script.cgi "or to access less resource-intensive scripts" http://www.hpl.hp.com/degraded_content/cgi-bin Can be converted to /my_script.cgi ".
【0051】
When determining that the content server 31 is overloaded, the adaptive controller 50 is downgraded, that is, the content adapter 41 to modify the URL of the incoming access request to access the corresponding content file in the adaptive content format. Instruct. As mentioned above, downgraded or adaptive content formats do not require as much resources as full content formats. If the content server 31 determines that it is not overloaded, the adaptive controller 50 instructs the content adapter 41 to modify the URL of the incoming access request to access the corresponding content file in full content format. ..
【0052】
Once the adaptation is triggered from the adaptation controller 50, the content adapter 41 will base the content adapter 41 on the root service directory 70 (based on the load state of the content server 31), as shown in FIGS. 2 and 4. In Figure 4), transparently switch between the high quality service tree 73 (Figure 4) and the downgraded quality service tree 74 (Figure 4). The switching operation is completely transparent to the content server 31.
【0053】
Figure 5 shows the results of the adaptive process by the adaptive load control system 40. Curve 81 shows the percentage of connection failures by prior art servers without the adaptive load control system 40. Curve 80 shows the percentage of connection failures by the content server 31 with the adaptive load control system 40. As can be seen in FIG. 5, as the demand rate increases, the curve 80 shows significantly less connection failures compared to the prior art curve 81. Traditional servers suffer from an increase in error rate (eg connection failure) when the load exceeds the server capacity.
【0054】
As described above, when the adaptive controller 50 determines that the content server 31 is overloaded, the adaptive controller 50 instructs the content adapter 41 to adapt the incoming access request to the downgraded content format. This puts the content server 31 in a potentially unloaded state, so that proper resource utilization will not be achieved when the adaptation is performed. This situation is shown in Figure 6. In FIG. 6, the solid line curve 85 represents the effective bandwidth of a prior art content server that does not have an adaptation mechanism according to one embodiment of the invention, and the dotted line curve 86 represents the adaptive load control system 40. Represents the effective bandwidth of the provided content server 31. As can be seen in Figure 6, the switch to a less resource-intensive content format causes the effective bandwidth of curve 86 to drop sharply for the content server 31. This means that the bandwidth utilization of the server will be inadequate.
【0055】
When the content server 31 is overloaded, this problem can be overcome by the content adapter 41 downgrading only part of the incoming access request. Part is between all and no incoming access requests. This mechanism is implemented in the content adapter 41. This mechanism uses an input control parameter G that can be considered as an external control knob that adjusts the extent of the partial downgrade. If G = Max, there is no request to be downgraded, and if G = 0, all incoming requests are downgraded. If G is reduced, the server load will be reduced.
【0056】
Suppose the total number of available content trees is Max, and the content trees are numbered from 1 to Max in ascending order of quality (tree 0 is a special tree that represents request rejection). Let I be the integer part of G and F be its decimal point. The following rules are used.
【0057】
If G is an integer (eg G = I, F = 0), content adapter 41 will serve any request from tree I. If G is not an integer, the content adapter 41 calculates a pseudo-random number N (in the range [0,1]) each time it receives each request. If N <F, the request is served from tree I + 1. Otherwise, the request is served from Tree I.
【0058】
This algorithm provides a continuous partial downgrade distribution between servicing all requests and rejecting all requests in the highest quality content tree. Next, see Figure 3 for a self-control technique that uses monitoring / feedback to automatically select the best value G to avoid overload and maintain the specified target utilization. And describe it.
【0059】
FIG. 3 shows a control / feedback loop of the adaptive load control system 40 of FIG. 2 that controls partial downgrades according to another embodiment of the invention. In this embodiment shown in FIG. 3, the partial downgrade module 60 is used to control the partial downgrade. Module 60 includes an integrated controller 62 and an adder 61. The achieved server utilization is sent from the load monitor 32 to the adder 61 of module 60. The desired load utilization is also sent to adder 61. The final result E (the difference between the two uses) is fed back to the integrated controller 62. The integrated controller 62 then adjusts the control parameter G to control the content adapter 41 to adjust the number of access requests for adaptation. Using feedback, Content Adapter 41 can quickly and automatically adjust the number of downgrade requests to be closer to the target or desired utilization. Although the present invention has been described above with reference to a specific embodiment, it is possible to make various modifications and changes to the above embodiment without departing from the idea of the present invention, as will be apparent to those skilled in the art. ..
【0060】
The present invention includes, for example, the following embodiments. (1) A content server that stores multiple content files, each stored in a complete content format or an adaptive content format that does not require as much resources as the complete format for services in preparation for access by external access requests, and the above content. -The content server includes an adaptive load control system connected to the content server to pass an access request to the server, and the content server is overloaded when the content server is overloaded. A data service system in a data service network system that modifies access requests to access the corresponding content files in the adaptive content format so that it remains in a safe load state. (2) The adaptive load control system according to (1) above, wherein the adaptive load control system modifies an access request to access a corresponding content file in full content format when the content server is not overloaded. Data service system. (3) The load monitor that monitors the load status of the content server, the load monitor, and the content server are connected to the adaptive load control system, and the load that the content server is in the overload state. The data service system according to (1) above, comprising a content adapter that modifies the access request to access the corresponding content file in the adaptive content format when the monitor indicates.
【0061】
(4) When the adaptive load control system is connected to the load monitor and the content adapter and the load monitor indicates that the content server is overloaded, the corresponding content in the adaptive content format. It also includes an adaptive controller that directs the content adapter to modify the access request to access the file, and the adaptive controller performs the load information received by the load monitor in a predetermined desire of the content server. The data service system according to (3) above, which determines whether or not the content server is overloaded by comparing with the load value of. (5) The above content adapter modifies the access request by modifying the access request URL to access the corresponding content file in either full content format or adaptive content format for each content file. The data service system according to (1) above, wherein the content server includes a service directory that directs the modified request as described above.
【0062】
(6) A method of maintaining a content server under a safe load in a data service system of a data service network system that has a content server that stores multiple content files in preparation for access by an external access request. The step of determining the load status of the content server when the data service system receives an access request to access one of the content files stored in the content server, and the above content server. If it is determined to be overloaded, the steps to modify the access request to access the corresponding content files in the adaptive content format so that the content server remains in a safe load state, and How to include.
【0063】
(7) The method according to (6) above, which includes a step of modifying the access request to access the corresponding content file in full content format when it is determined that the content server is not overloaded. .. (8) The above step of determining the load state compares the step of acquiring the actual load state of the content server using the load monitor with the above actual load state and a predetermined desired load state. The method according to (6) above, comprising the step of determining whether or not the content server is overloaded. (9) The method according to (6) above, wherein the above step of modifying the access request is performed by modifying the URL of the access request.
【0064】
[Effect of the invention]
According to the present invention, when the web server becomes overloaded, the access request to the web server is automatically controlled, and as a result, the safe load state is maintained.
[Simple explanation of drawings]
[Figure 1]
It is a block diagram which shows the data access network system including a large number of data service systems.
[Figure 2]
FIG. 5 is a block diagram showing a configuration according to an embodiment of the present invention in the adaptive data service system of FIG.
[Fig. 3]
FIG. 2 is a block diagram showing a control and feedback loop for a partial downgrade of an adaptive data service system in Figure 2.
[Fig. 4]
It is a block diagram which shows the structure of the content server of the adaptive data service system of FIG.
[Fig. 5]
It is a graph which shows the comparison of the connection failure rate between the prior art data service system and the adaptive data service system of FIG.
[Fig. 6]
It is a graph which shows the comparison of the bandwidth between the prior art data service system and the adaptive data service system of FIG.
[Fig. 7]
It is a block diagram which shows the gap estimation method used by the load monitor of FIG. 2 and FIG.
[Explanation of symbols]
30 Data service system 31 Content server 40 Adaptive load control system
8 sheets
Sheet 1 Sheet 2 Sheet 3 Sheet 4 Sheet 5 Sheet 6 Sheet 7 Sheet 8
Every citation, both ways
| Document | Relation | Office | Cited during |
|---|---|---|---|
| JP2010123015A | Cited by | Japan | Examiner |
| JP6062511B1 | Cited by | Japan | Examiner |
| JPH10133973A | Cites | Japan | Examiner |
| JPH10242997A | Cites | Japan | Search report |
| JPH1027165A | Cites | Japan | Examiner |
7 members in 4 offices
Priority claims5
| Document | Office | Kind | Date |
|---|---|---|---|
| 09299684 | United States of America | – | |
| 29968499 | United States of America | A | |
| 29968499 | United States of America | A | |
| 299684 | – | – | – |
| US19990299684 | – | – | – |
Members7
| Document | Office | Kind | |
|---|---|---|---|
| EP1049031A2 | European Patent Office (EPO) | A2 | |
| JP2000322366AThis record | Japan | A | |
| EP1049031A3 | European Patent Office (EPO) | A3 | |
| EP1049031B1 | European Patent Office (EPO) | B1 | |
| DE69939460D1 | Germany | D1 | |
| US8200837B1 | United States of America | B1 | |
| US2012233294A1 | United States of America | A1 |
10 legal events, as the office reported them to INPADOC
Over the term
Point at a mark for the eventEvents
| Event | Code | |
|---|---|---|
| Re-examination (zenchi) completed and case transferred to appeal boardAppealJAPANESE INTERMEDIATE CODE: A912A912 | A912 | |
| Transfer to examiner for re-examination before appeal (zenchi)AppealJAPANESE INTERMEDIATE CODE: A911A911 | A911 | |
| Written amendmentJAPANESE INTERMEDIATE CODE: A523A521 | A521 | |
| Decision of refusalJAPANESE INTERMEDIATE CODE: A02A02 | A02 | |
| Written amendmentJAPANESE INTERMEDIATE CODE: A523A521 | A521 | |
| Notification of acceptance of power of attorneyJAPANESE INTERMEDIATE CODE: A7422RD02 | RD02 | |
| Notification of reasons for refusalJAPANESE INTERMEDIATE CODE: A131A131 | A131 | |
| Report on retrievalJAPANESE INTERMEDIATE CODE: A971007A977 | A977 | |
| Written amendmentJAPANESE INTERMEDIATE CODE: A523A521 | A521 | |
| Written request for application examinationJAPANESE INTERMEDIATE CODE: A621A621 | A621 |
Numbers
- Publication
- 2000-322366
- Publication, DOCDB
- 2000322366
- Publication, EPODOC
- JP2000322366
- Application
- 79687
- Application, DOCDB
- 2000079687
- Application, EPODOC
- JP20000079687
Titles2
- Japanese
- データ・サービス・システム
- English
- [Title of Invention] Data Service System
Classification
- CPC, 6
- H04L67/02
- H04L65/80
- H04L67/1008
- G06F2209/5013
- G06F9/5011
- H04L65/612
- IPC, 4
- G06F13 00
- G06F17 30
- H04L29 08
- G06F15 177