A proximity-based redirection system for robust and scalable service-node location in an internetwork
Abstract
A proximity-oriented redirection system for a service-to-client accessory in a virtual overlay distribution network (300) is disclosed. The virtual overlay distribution network 300 includes addressable routers R1-R6 for routing packet traffic. A data packet is routed from the source node C1 to the destination node C2 based on the address field of the packet. The invention comprises a redirector (S1, S2) coupled to at least one of the addressing routers (R3, R4). The redirector (S1, S2) includes logic for receiving a service request from a client (312); logic for determining a selected server that is one of a plurality of servers capable of processing the service request to process the service request; and logic for generating a redirect message directed to the client to redirect the service request to the selected server.

Term
Term ended
Expired 4 March 2022, 4.6 years ago.
- Priority
- Filed
- Granted
- Expired
- Today
23 claims: 4 independent, 19 dependent
- 1테이터 패킷이 상기 패킷의 주소 필드를 기초로 하여 소스 노드로부터 종착 노드까지 라우팅되는, 패킷 트래픽을 라우팅하기 위한 주소 가능 라우터들을 포함하는 패킷-교환형 네트워크에서, 상기 패킷-교환형 네트워크는 상기 주소지정 가능 라우터들의 적어도 하나에 결합된 반향 전환기(redirector)를 포함하고, 상기 방향 전환기는 A) 클라이언트로부터 서비스 요구-상기 서비스 요구는 에니케스트 종착 주소로의 에니케스트 메시지임-를 획득하기 위한 로직;B) 상기 서비스 요구를 처리하기 위해 선태된 서버-상기 선택된 서버는 상기 서비스 요구를 조절할 수 있는 다수의 서버들 중 하나임-를 결정하는 로직;및 C) 상기 서비스 요구를 상기 선택된 서버로 방향 전환하기 위하여 상기 클라이언트로 지향된 방향 전환 메시지를 발생시키기 위한 로직을 포함하는 것을 특징으로 하는 패킷-교환형 네트워크.
- 2제1 항에 있어서, 도달능력을 상기 에니캐스트 종착 주소에 광고하기 위한 로직을 더 포함하는 것을 특징으로 하는 패킷-교환형 네트워크.
- 3제1 항에 있어서, 상기 결정 로직은 상기 다수의 서버들의 네트워크 트래픽 상태를 모니터링하는 로직;및 상기 네크워크 트래픽 상태를 기초로 하여 상기 선택된 서버를 상기 다수의 서버들로부터 선택하기 위한 로직을 포함하는 것을 특징으로 하는 패킷-교환형 네트워크.
- 4제1 항에 있어서, 상기 결정 로직은 상기 다수의 서버들의 서버 상태를 모니터링하는 논리;및 상기 서버 상태에 기초로 하여 상기 선택된 서버를 상기 다수의 서버들로부터 선택하기 위한 로직을 포함하는 것을 특징으로 하는 패킷-교환형 네트워크.
- 5제1 항에 있어서, 상기 선택된 서버는 멀티캐스팅(multicasting) 서버인 것을 특징으로 하는 패킷-교환형 네트워크.
- 6제1 항에 있어서, 상기 방향 전환기는 상기 선택된 서버인 것을 특징으로 하는 패킷-교환형 네트워크.
- 7테이터 패킷이 상기 패킷의 주소 필드에 기초하여 소스 노드로부터 종착 노드까지 라우팅되며, 패킷 트래픽을 라우팅하기 위한 주소지정 가능 라우터들을 포함하는 패킷-교환형 네트워크에서 방향 전환기를 작동시키는 방법으로서, 도달능력을 상기 방향 전환기로부터 상기 에니캐스트 종착 주소로 광고하는 단계;클라이언트로부터 서비스 요구-상기 서비스 요구는 상기 에니캐스트 종착 주소로의 에니캐스트 메시지임-를 획득하는 단계;상기 서비스 요구를 처리하기 위하여 선택된 서버-상기 선택된 서버는 상기 서비스 요구를 처리할 수 있는 다수의 서버들 중에서 하나임-를 결정하는 단계;상기 서비스 요구를 상기 선택된 서버로 방향 전환하기 위하여 상기 클라이언트로 지향된 방향 전환 메시지를 발생시키는 단계를 포함하는 것을 특징으로 하는 방향 전환기 작동 방법.
- 8제7 항에 있어서, 상기 다수의 서버들의 트래픽 상태를 모니터링하는 단계를 더 포함하는 것을 특징으로 하는 방향 전환기 작동 방법.
- 9제8 항에 있어서, 상기 결정 단계는 상기 트래픽 상태에 기초하여 상기 선택된 서버를 상기 다수의 서버들로부터로 선택하는 단계를 포함하는 것을 특징으로 하는 방향 전환기 작동 방법.
- 10제7 항에 있어서, 상기 다수의 서버들의 서버 상태를 모니터링하는 단계를 더 포함하는 것을 특징으로 하는 방향 전환기 작동 방법.
- 11제10 항에 있어서, 상기 결정 단계는 상기 서버 상태에 기초하여 상기 선택된 서버를 상기 다수의 서버들로부터 선택하는 단계를 포함하는 것을 특징으로 하는 방향 전환기 작동 방법.
- 12제7 항에 있어서, 상기 방향 전환기에서 상기 서비스 요구를 처리하는 단계를 더 포함하는 것을 특징으로 하는 방향 전환기 작동 방법.
- 13제7 항에 있어서, 상기 발생 단계는 상기 클라이언트가 상기 선택된 서버에서 멀티캐스트 그룹에 가입하도록 방향을 전환하기 위하여 상기 클라이언트로 지향된 방향 전환 메시지를 발생하는 단계를 포함하는 것을 특징으로 하는 방향 전환기 작동 방법.
- 14데이터 패킷이 상기 패킷의 주소 필드에 기초하여 소스 노드로부터 종착 노드까지 라우팅되는, 패킷 트래픽을 라우팅하기 위한 주소지정 가능 라우터들을 포함하는 패킷-교환형 네트워크에서, 상기 패킷-교환형 네트워크는 다수의 주소 가능 라우터들중 적어도 제1 라우터에 결합되고, 클라이언트와 에니캐스트 그룹의 다수 노드들 사이에 데이터 패킷들을 전파하기 위한 논리를 포함하는 적어도 하나의 서비스 노드;및 상기 다수의 주소지정 가능 라우터들중 적어도 제2 라우터에 결합된 적어도 하나의 방향 전환기를 포함하고, 상기 적어도 하나의 방향 전환기는 A) 상기 에니캐스트 그룹의 다수 노드들에 관련된 에니캐스트 종착 주소에 도달 능력을 광고하는 로직;B) 클라이언트로부터 서비스 요구-상기 서비스 요구는 상기 에니캐스트 종착 주소로의 에니케스트 메시지임-를 획득하는 로직;및 C) 상기 서비스 요구를 상기 적어도 하나의 서비스 노드로 방향 전환하기 위하여 상기 클라이언트로 지향된 방향 전환 메시지를 발생시키는 로직을 포함하는 패킷-교환형 네트워크.
- 15제14 항에 있어서, 상기 적어도 하나의 서비스 노드는 다수의 서비스 노드들을 포함하고, 상기 적어도 하나의 방향 전환기는 상기 서비스 요구를 처리하기 위해 상기 다수의 서비스 노드들로 선택된 서비스 노드를 결정하는 로직;및 상기 서비스 요구를 상기 선택된 서비스 노드로 방향 전환하기 위하여 상기 클라이언트로 지향된 방향 전환 메시지를 발생하는 로직을 포함하는 것을 특징으로 하는 패킷-교환형 네크워크.
- 16제15 항에 있어서, 상기 다수의 서비스 노드들로 선택된 서비스 노드를 결정하는 로직은 상기 다수의 서비스 노드들에서 네트워크 트래픽 상태를 모니터링하는 로직;및 상기 네크워크 트래픽 상태에 기초하여 상기 다수의 서비스 노드들로부터 상기 선택된 서비스 노드를 선택하는 논리를 포함하는 것을 특징으로 하는 패킷-교환형 네크워크.
- 17제15 항에 있어서, 상기 다수의 서비스 노드들로 선택된 서비스 노드를 결정하는 로직은 상기 다수의 서비스 노드들의 서버 상태를 모니터링하는 로직;및 상기 서버 상태에 기초하여 상기 선택된 서비스 노드를 상기 다수의 서비스 노드들로부터 선택하는 로직을 포함하는 것을 특징으로 하는 패킷-교환형 네크워크.
- 18제14 항에 있어서, 에니캐스트 그룹에서 상기 다수의 노드들의 제1 부분은 제1 지리적 위치에 배치되고 상기 에니캐스트 그룹에서 상기 다수의 노드들의 제2 부분은 제2 지리적 위치에 배치되며, 상기 방향 전환기는 상기 에니캐스트 서비스 요구를 보내는 클라이언트가 상기 에니캐스트 그룹에서 상기 다수의 노드들의 제1 부분 또는 제2 부분에 더 가까이에 있는지를 결정하는 로직;및 상기 클라이언트가 상기 에니캐스트 그룹에서 상기 다수의 노드들의 제1 부분에 더 가까이에 있는 경우 상기 서비스 요구를 제1 서비스 노드로 방향 전환하도록하며 상기 클라이언트가 상기 에니캐스트 그룹에서 상기 다수의 노드들의 제2 부분에 더 가까이에 있는 경우 상기 서비스 요구를 제2 서비스 노드로 방향 전환하도록 하는 상기 클라이언트에 지향된 상기 방향 전환 메시지를 발생하기 위한 로직을 더 포함하는 것을 특징으로 하는 패킷-교환형 네크워크.
- 19테이터 패킷이 상기 패킷의 주소 필드에 기초하여 소스 노드로부터 종착 노드까지 라우팅되며, 패킷 트래픽을 라우팅하기 위한 주소지정 가능 라우터들, 상기 주소지정 가능 라우터들의 적어도 하나에 결합된 방향 전환기 및 적어도 하나의 서비스 노드를 포함하는 패킷-교환형 네트워크를 작동시키는 방법에 있어서, 도달 능력을 상기 방향 전환기로부터 에니캐스트 종착주소에 광고하는 단계;상기 방향 전환기에서 클라이언트로부터 서비스 요구-상기 서비스 요구는 상기 에니캐스트 종착 주소로의 에니캐스트 메시지임-을 획득하는 단계;상기 서비스 요구를 상기 적어도 하나의 서비스 노드로 방향 전환하기 위해 상기 클라이언트로 지향된 방향 전환 메시지를 발생시키는 단계를 포함하는 것을 특징으로 하는 패킷-교환형 네크워크 작동 방법.
- 20제19 항에 있어서, 상기 적어도 하나의 서비스 노드는 다수의 서비스 노드들을 포함하고, 상기 발생 단계는 상기 서비스 요구를 처리하기 위해 상기 다수의 서비스 노드들로부터 선택된 서비스 노드를 결정하는 단계;및 상기 서비스 요구를 상기 선택된 서비스 노드로 방향 전환하기 위해 상기 클라이언트로 지향된 방향 전환 메시지를 발생시키는 단계를 포함하는 것을 특징으로 하는 패킷-교환형 네트워크 작동 방법.
- 21제20 항에 있어서, 상기 결정 단계는 상기 다수의 서비스 노드에서 네트워크 트래픽 상태를 모니터링하는 단계;및 상기 네크워크 트래픽 상태에 기초하여 상기 다수의 서비스 노드들로부터 상기 선택된 서비스 노드를 선택하는 단계를 포함하는 것을 특징으로 하는 패킷-교환형 네트워크 작동 방법.
- 22제20 항에 있어서, 상기 결정 단계는 상기 다수의 서비스 노드에서 서버 상태를 모니터링하는 단계;및 상기 서버 상태에 기초하여 상기 다수의 서비스 노드들로부터 상기 선택된 서비스 노드를 선택하는 단계를 포함하는 것을 특징으로 하는 패킷-교환형 네크워크 작동 방법.
- 23제19 항에 있어서, 에니캐스트 그룹에서 상기 다수의 노드들의 제1 부분은 제1 지리적 위치에 배치되고 상기 에니캐스트 그룹에서 상기 다수의 노드들의 제2 부분은 제2 지리적 위치에 배치되며, 상기 결정 단계는 상기 에니캐스트 서비스 요구를 보내는 클라이언트가 상기 에니캐스트 그룹에서 상기 다수의 노드들의 제1 부분 또는 제2 부분에 더 가까이에 있는지를 결정하는 단계;및 상기 클라이언트가 상기 에니캐스트 그룹에서 상기 다수의 노드들의 제1 부분에 더 가까이에 있는 경우 상기 서비스 요구를 제1 서비스 노드로 방향 전환하도록하며 상기 클라이언트가 상기 에니캐스트 그룹에서 상기 다수의 노드들의 제2 부분에 더 가까이에 있는 경우 상기 서비스 요구를 제2 서비스 노드로 방향 전환하도록 하는 상기 클라이언트에 지향된 상기 방향 전환 메시지를 발생하는 단계를 더 포함하는 것을 특징으로 하는 패킷-교환형 네트워크 작동 방법.
Independent claims23
169 paragraphs in 1 section, as filed
A PROXIMITY-BASED REDIRECTION SYSTEM FOR ROBUST AND SCALABLE SERVICE-NODE LOCATION IN AN INTERNETWORK
CROSS-REFERENCE TO RELATED APPLICATIONS
This application claims priority to U.S. Provisional Patent Application No. 60/152,257, filed on September 3, 1999 and pending. This application relates to U.S. Patent Applications 09/323,869 and 1999, entitled "PERFORMING MULTICAST COMMUNICATION IN COMPUTER NETWORKS BY USING OVERLAY ROUNTING," filed June 1, 1999. U.S. Patent Application 60/115,454, entitled "SYSTEM FOR PROVIDING APPLICATION-LEVEL FEATURES TO MULTICAST ROUTING IN COMPUTER NETWORKS," filed January 1; related The disclosure of each of the above-specified applications is hereby incorporated by reference for various purposes by reference.
FIELD OF THE INVENTION The present invention relates generally to the field of data networks, and more particularly to the distribution of information over data networks.
One of the central challenges in adjusting Internet infrastructure for mass adoption is the problem of distributing arbitrary content from a source site to a large number of users of the Internet in an effective, feasible and economical form. Distribution of popular news articles, video broadcasts, stock quotes, new releases of popular software, etc., where a large number of users spread out on a network all attempt to detect the same content from the same server at about the same time are all so-called flash. It can cause a flash effect. A traffic flash not only brings servers down, but also wastes bandwidth on the network as many redundant copies of the same content flow through the wide area network. For example, an emergency news event on CNN's Web site caused millions of users to fetch the text of the article and turn off their servers. Likewise, the initial launch of high-resolution movie broadcasts over the Internet may allow millions of users to access media content servers.
Two key mechanisms have been proposed for the Web to overcome the problems caused by the flash effect, ie, caching and server replication. In caching, a cache is conveniently located within the network infrastructure to intercept content requests from clients. When the cache receives a content request, it refers to the storage of the content, and if the requested data is present, the cache serves the request locally. If not, the request is provided to the originating server, and the response is returned to the client. During this process, the cache stores the response in local storage. A number of strategies have been proposed for managing local storage, such as determining when to remove objects from cache, when to recharge servers and other objects, and so on. A cache may be non-transparent if the client is explicitly configured by the cache's network address, and if the client ignores the cache and if the cache is configured by, for example, a layer-4 switch. ) to transparently intercept the content request, the cache may be transparent.
In server iteration, servers are deployed over a wide area, and clients are assigned to these distributed servers to balance the load and secure network bandwidth. These repeated servers may have some or all of the content housed on the original server, how the particular arrangement of servers is laid out, how content is distributed from the master server to them, and how clients are assigned to the appropriate servers. There are various variables for
The technology developed to support this type of server iteration and caching techniques is ad hoc and does not match the underlying Internet structure. For example, common techniques for transparent caching break TCP semantics and are therefore incompatible with some modes of underlying IP packet service, such as multipath routing. This causes many difficult administrative problems and, in particular, does not provide a responsive network structure that can be managed in a sensible fashion from a network operation center.
A similar content distribution problem involves the distribution of live streaming media to many users over the Internet. Here, the server provides live broadcast feedstock, and clients connect to the server using streaming media transport protocols to receive the broadcast. However, as more and more clients connect to the broadcast, the server and network are brought down by the task of distributing a large number of packet streams to a large number of clients.
One solution to this live broadcast problem is to borrow the efficiency of IP multicast as defined in network layer multicast or Internet architecture. In this approach, the server sends a single stream of packets to a "multicasting group" rather than sending individual copies of the stream to individual clients. In turn, the recipients involved in this stream "tune in" to the broadcast by subscribing to a mulkeycasting group (eg, by signaling subscription information using Internet Group Management Protocol (IGMP) to a nearby router). The network distributes the broadcast to each recipient by simply copying packets at fan out points in the distribution line from the source to all recipients. Thus, only one copy of each packet appears on the physical link.
Unfortunately, the wide variety of deployment and scaling issues has confused the adoption and proliferation of IP multicast in the global Internet. Most of these problems essentially require that the operation of a multicast distribution tree has a consistent view of all routers in the network as the tree looks like. For multicast, each router must have an accurate local view of a single global coherent multicast routing tree. If routers have different views of a given multicast tree in different parts of the network, routing loops and black holes will not be possible. Many other problems -- eg multicast address assignment, multicast congestion control, reliable distribution of multicast, etc. -- also plague the deployment and acceptance of IP multicast. Despite the general progress in the last two years towards commercial deployment of multicast, the resulting infrastructure is still relatively fragile, and its scope is extremely limited.
There are also practical technical barriers to the deployment of ubiquitous Internet multicast services, as well as business and economic barriers. Internet service providers have not been very successful in broadband multicast services because management, monitoring, and preparation for multicast traffic are very difficult. Moreover, it is difficult to control who generates the traffic in a multicast session and which part of the network provides that traffic. These obstacles make it difficult for multicast services to reach a good part of the Internet. Even when that happens, the process will undoubtedly take a lot of time to unfold.
To avoid the risk, others have been proposed that enable streaming-media broadcasting by means of an application-level solution called a splitter network. In this approach, a set of servers distributed in the network are located at strategic locations within networks of service providers. These servers are provided with a "splitting" capability that allows them to copy a given stream to multiple downstream servers. With this capability, servers can be arranged in a tree-like system. In this tree structure, the root server becomes the stream source of multiple downstream servers, and in turn distributes the stream to multiple copies to be provided to different tiers of the downstream servers.
Unfortunately, a distributed network of servers is plagued by a number of problems. First, the tree of distributors is statically arranged, which means that if one distributor fails, the entire sub-tree below the failed point loses service. Second, since it requires individual distributor networks composed of each physical server for maintenance of each broadcasting network, the distributor network should be directed to a single broadcasting center. Third, since the distributor abstraction is based on the extension of the media server, it is necessarily platform dependent. For example, RealNetworks-based distributor networks cannot distribute Microsoft Netshow traffic. Fourth, by not tracking recipient interest and not branching traffic from a sub-tree of a distributor network that does not have downstream recipients, distributor networks are highly bandwidth wasting. Finally, a splitter network provides weak policy control--the total bit rate consumed along the path between two splitter nodes is uncontrollable, and in a stream-aware fashion the It cannot be assigned to other classes.
In order to address many of the various problems described above, an embodiment of the present invention provides a comprehensive redirection system for content distribution in a virtual overlay broadcast network (OBN). In the present system, service nodes are arranged at strategic locations in the network infrastructure, unlike conventional systems, and these service nodes cooperate over a wide area with a cohesive, cooperative, and administrative virtual overlay network. Service node clusters are peer-to-peer across IP tunnels, exchange routing information, client subscription data, configuration control, and bandwidth provisioning efficiencies. At the same time, the server nodes can process application-on-demand requests for content, eg appearing as a web server or streaming-media server depending on the supporting service nature. In short, the service node plays a hybrid role: acting as a server as well as an application-level content router.
In embodiments of the present invention, improvements are provided for packet-switched networks. The packet-switch network in which a data packet is routed from a source node to a destination node based on address fields of the packet includes addressable routers for routing packet traffic. The enhancement includes a redirector associated with at least one of the addressable routers, logic for accepting a service request from a client, logic for determining a selected server to handle the service request, and and logic for generating a redirect message commanded to the client to echo the service request to a server, wherein the selected server is one of a plurality of servers capable of handling the service request.
1 is a diagram illustrating an example of typical components and interconnections comprising a portion of Internet connectivity;
2 is a diagram illustrating a typical overlay ISP model;
Figure 3 is a diagram illustrating a network portion 300 used for implementing anycast routing in accordance with the present invention;
4 is a diagram illustrating master and affiliate networks configured to configure interdomain anycast routing;
5 is a diagram showing how the control and service functions included in the present invention are classified within a specific ISP;
6 is a diagram illustrating a part of a data network constructed according to the present invention to implement active session failover;
Fig. 7 is a diagram showing a part of a data network constructed in accordance with the present invention to implement a wide area overflow;
8 is a diagram illustrating the use of IP multicast according to the present invention; and
9 is a diagram illustrating an embodiment of the present invention adopted for registering and connecting service installations to a service broadcasting network infrastructure.
The comprehensive redirection system of the present invention operates in tandem with a cohesive, cooperative, and administrative virtual overlay network and service nodes deployed at strategic locations across a wide cooperating network infrastructure. The overlay network structure is based on a design spirit similar to the underlying Internet structure. For example, it uses scalable addressing, adaptive routing, systematic naming, distributed management, and the like. For this reason, overlay networks enjoy the same high degree of robustness, scalability, and handling that is evident within the Internet itself. Unlike a physical Internetwork in which routers are added directly to each other through physical links, service nodes in a virtual overlay network communicate with each other using packet services provided by the underlying IP network. As such, a large area network (eg, the backbone of an entire ISP) composed of a large number of individual components (such as routers, switches and links) with a small number of service nodes to provide good content-distribution practices. Since it can only require
Another important aspect of the present invention is the "glue" interface between a client that wants to receive information content and the service nodes that actually distribute it. That is, the "Grew" interface is a mechanism by which a client can attach to a service node, request specific content and have that content effectively distributed. This is sometimes referred to as the service rendezvous problem.
Essentially, a service rendezvous is: (1) issuing a single name for a service; (2) replicate the service across the network; and (3) a system that allows each client that wants the service to receive it from the most appropriate server. To scale millions of clients, a service rendezvous mechanism must effectively distribute and load-balance client requests to service nodes spread over a large area. Moreover, to effectively utilize network bandwidth, content must flow over a minimum number of network links to reach the requesting clients. Both of these assert that clients must instruct a nearby service node that can service the request. If there is no nearby service node that can service the request, the system must be able to redirect the client to another service node in the wide-area network to service the request. It should also be possible to cluster service nodes in a specific location and have clients connect with individual nodes in the cluster based on traffic load phases. In short, the service rendezvous system should be able to provide a mechanism for service selection and utilize redirection to achieve load balancing to achieve desired results. When a local cluster is overloaded, server selection must be compensated for load balancing over a large area.
Unfortunately, server rendezvous is a difficult problem because the Internet architecture deliberately hides the underlying network infrastructure to provide redundancy and robustness to higher layer protocols, so discovering and using the selected server for a particular network transaction. difficult to do To overcome these problems, the rendezvous service described herein utilizes "anycast" routing, a network-level mechanism that can be used to route user requests to nearby service nodes based on topological locality. .
Anycasting packet transmission concept is well described in the network research literature; Yet the concept remains narrowly applied in practice due to its compatibility with existing packet transport networks. In general, the behavior of the Internet ratifies consensus agreements on standards. The draft standard is presented in a document called requests for comments (RFC). RFCs applicable to Internet operation include RFC-1546 and others
At the highest level, there are two basic approaches to implementing anycast packet transmission. The first approach is to introduce a special type of anycast address and create service interfaces for which new routing protocols and "anycast aware" exist. This will obviously lead to a tedious process of standardization and adoption by router vendors, etc. A second approach is to reuse an existing unicast address space. However, this second approach has two corresponding technical challenges that cannot be solved in this respect. The challenges include (1) support of stateful transport protocols; and (2) support of inter-domain anycast routing and route aggregation. Essentially, the embodiments of the present invention provide novel solutions to these technical challenges. For example, a solution to the problems of supporting stateful transport protocols is provided in a section of this specification entitled "Stateful Anycasting". A solution to the problem of supporting inter-domain anycast routing and route aggregation is provided in a section of this specification entitled "Interdomain Anycast Routing".
The rendezvous service described here assumes that the underlying packet transmission is not "anycast aware". However, a system based on "anycast aware" is also feasible. Anycas packet transport is used to send packets from the client to the nearest instance of the rendezvous service.
Rather, one embodiment of the present invention is a statically provided framework (where only specially configured hosts within the network infrastructure are members of a cast group) to simplify the anycast service model. In this statically provided framework, the assignment, assignment, and propagation of anycasts is addressed to a central authority, and associates a large block of anycast addresses with a single, well-connected backbone network. The backbone network may be referred to as a content backbone (CBB).
Another advantage included in an embodiment of the present invention is that the client is attached to the content distribution network at explicit, per-client service access points. This allows the infrastructure to implement user-specific authentication, monitoring, ordering, and advertising. In contrast, a purely multicast-based approach, although scalable, provides none of these features as the multicast receiver subscription process is completely anonymous.
In summary, a virtual overlay network built using anycast-based service rendezvous enjoys the following fascinating characteristics:
· The service access mechanism is highly scalable, as it discovers the nearest service node using anycast, which can be achieved by standard routing protocols deployed in new configurations;
The system provides significant bandwidth savings as requests can be routed to the nearest service node to minimize the number of network links through which the content must flow;
Service infrastructure provides particulate control, monitoring and ordering of client connections;
Facilitate large-scale deployment across heterogeneous environments managed by a diverse range of operational elements, so that the operation and formation of the infrastructure is highly distributed;
The system provides very high availability and robustness when anycast is built on standard adaptive routing protocols and service elements are clustered for redundancy, so that requests act properly and servers advertise their availability. Make sure to route only to the fields; and
· Anycast-based echo diversion service can be formed first as a content broadcasting backbone, and then at affiliate ISPs on an individual basis for track increasing user demand, so the content broadcasting network can be deployed incrementally.
In the following sections of this specification, details of the architectural model and embodiments of various system components for implementing an anycast-based redirection system for virtual overlay networks encompassed by the present invention are described.
<b>network structure</b>
1 is a diagram illustrating an example of typical components and interconnections comprising a portion of the Internet 100 . Internet service providers (ISPs) 101, 102 and 103 provide Internet access. A typical ISP operates an IP-based network across a wide area to connect individual customer networks 104 and/or individual users to the network via access devices 106 (eg, DSL, telephone modem, cable modem, etc.). do. A typical ISP peers with other ISPs through switching points 108 so that data traffic can flow from one ISP user to another ISP user. The collection of internal IP routers 110 interconnected by communication links provides connectivity between users within the ISP. Specialist border routers 112 located at the switching points send non-local traffic to and from the ISP. Often, an individual ISP network, such as ISP 103, is referred to as an autonomous system (AS) because it represents an independent and aggregated unit in terms of network routing protocols. Within an ISP, interdomain routing protocols work (eg RIP or OSPF), and in ISPs, interdomain routing protocols work (eg BGP). The term "intradomain protocol" is often used interchangeably with the term interior gateway protocol (IGP).
As the Internet and the World Wide Web grew, ISPs realized that they could achieve better end-to-end network service implementations by synthesizing two innovative architectural concepts in competition: (1) multiple at each exchange point. actively peering adjacent ISPs of and (2) co-locating data centers accommodating application services (eg, web servers) near these exchange points. The "co-locate bias" (colo) at each peering point allows application services to be copied at each peering point, allowing virtually any user anywhere in the network to enjoy high-speed connections to nearby services. .
2 is a diagram illustrating a typical overlay ISP 200 . This is because the service network so established forms an overlay structure over a number of existing ISPs, for example, ISPs 202, 204 and 206. The overlay ISP 200 is coupled to existing ISPs via routers 208 and is also coupled to a data center 210 (DC). Overlay ISPs rent machine space and network bandwidth to content providers who place their servers in colo's located in data centers.
In summary, the natural building block for CBB is the ISP colo. In embodiments of the present invention, service nodes are housed in a colos and arranged in an overlay structure over a wide area using available network connectivity. However, service nodes need not be located at specialized callo sites, and may in fact exist as part of the network. Colos are a convenient and viable deployment channel for serving nodes.
Interdomain Anycast Routing
3 is a diagram illustrating a network 300 configured to implement anycast routing in accordance with the present invention. The network 300 includes routers R1-R6 and two server devices S<sb>1 </sb>and S<sb>2</sb>), and two clients (C<sb>1 </sb>and C<sb>2 </sb>) is included. In one embodiment of network 300, both server devices advertise reachability via IGP to address block "A/A24" (ie, A is a 24-bit prefix for a 32-bit IPv4 address). As such, both server devices utilize routing advertisements to reflect server utilization in the infrastructure of network 300 . routers (R<sb>4 </sb>and R<sb>3</sb>) are configured to listen for these reachability advertisements on their attached rands 302 and 304, respectively. As a result of the IGP operation, the routers R1-R6 learn the shortest path from each client to the servers through addresses falling within the "A" potential. Therefore, if the client (C<sb>1</sb>) is the address (A<sb>1</sb>) (where A<sb>1</sb>When sending a packet to the router (R is the potential of<sb>2</sb>) is a router (R<sb>4</sb>) and, in turn, as shown in path 310, the server S<sb>2</sb>) is sent to Similarly, from the client A<sb>1</sb>The packet sent to the server S, as shown in path 312,<sb>1</sb>) is routed to If the server (S<sb>2</sb>) fails, then S for A/24<sb>2</sb>Advertisements will be generated from , and the network will recompute the shortest-path routes corresponding to A/24. Thus, as shown by path 314, there are no other nodes touting such a route, so C<sb>2</sb>from A<sb>1</sb>Packets sent to the server (S<sb>1</sb>) is routed to
One of the problems posed by the anycast routing mechanism described above is how to propagate anycast routes across a wide area to any sites that are not configured with anycast-based service nodes. Rather than requiring a new infrastructure for anycast routing, embodiments of the present invention provide a framework in which a single AS accepts a given block of anycast address and advertises using its generic interdomain protocol, namely BGP. ) and use the existing interdomain routing system. Other independent ASs can then be incrementally configured with anycast-aware service nodes, such that the IGP for those ASs can service packets sent to the anycast address block in question into service within that single AS. route to nodes.
To do this, the content backbone (CBB) resides in the "master" AS that owns the anycast blocks and advertises them to the Internet using BGP. That is, the ISP moves out of a block of its previously existing but unused address space (or requests new addresses from the Internet Assigned Numbers Authority), and allocates this block to the CBB to address these addresses. Declare that this block should only be used for anycast routing. The master AS advertises the ENICAS block--this block is called "A"---this block is referred to as "A"--over the wide area again using BGP as if it were a normal IP network. Accordingly, in the configuration described so far, packets transmitted from any place on the Internet to the address in block A are routed to the master AS.
In order to provide the services that lie under the equicast routing infrastructure, the CBB deploys service nodes in the master AS and arranges such nodes to advertise reachability for A using the master AS's IGP. Once these fragments are placed, when a packet enters the master AS (from anywhere on the Internet), it is routed to the CBB service node closest to the border router traversed by the packet as soon as it enters the master AS. Assuming that the master AS is densely peered, most users on the Internet will enjoy a short-latency, high-speed path to the service node in the master AS (CBB).
Although the architecture described so far provides a feasible mechanism for proximity-based service-node location for nodes located within the master AS, the system is limited by the fact that all service nodes belong to that master AS. A more scalable approach will allow service nodes to be installed in different ISP's networks. To do so, an affiliate AS - that is, an ISP that is not a master AS but supports rendezvous services, simply installs service nodes in exactly the same fashion as the master AS. However, the family advertises anycast blocks only within the family's domain using the family's IGP; Do not advertise out-of-domain anycasts to peers. In another embodiment, an extension to this mechanism is provided, where multiple ASs advertise anycast in BGP (ie, their external routing protocol). Such extensions are described elsewhere in this specification.
4 is a diagram illustrating a master AS 400 and affiliate networks 402, 404 and 406 configured to achieve interdomain anycast routing. Master AS 400 is anycast-based service nodes (A)<sb>1, </sb>A<sb>2 </sb>and A<sb>3</sb>), and combines with three affiliate networks through routers 408 . 4 clients (C<sb>1, </sb>C<sb>2, </sb>C<sb>3 </sb>and C<sb>4</sb>) is attached to the series, as shown. A single service node (A) in which the family 404 is arranged in its infrastructure<sb>4</sb>), but series 402 and 406 have no service nodes deployed. Thus, given the general behavior of unicast inter- and intra-domain routing protocols, C<sb>2</sb>Packets sent from block A to block A are<sb>2</sb>whereas routed to C<sb>1</sb>Packets sent to block A from<sb>1</sb>is sent to The paths 410 and 412 represent the shortest interdomain paths from the family 402 to the master AS 400 . In contrast, C<sb>3</sb>Packets sent to block A from service node A, as shown by path 414<sb>4</sb>is routed to This means that the IGP in family 404 is server node A.<sb>4</sb>will advertise its ability to reach block A and thus "hijack" packets sent to that address. Similarly, the path from series 406 to master AS 400 traverses series 404, so C<sb>4</sb>Packets sent to block A from<sb>4</sb>will also be robbed by This is a prudent and desirable feature of the architecture according to the present invention, as it scales and distributes the load of the system without the need for anycast intelligence that must be deployed everywhere for correct operation.
Although the anycast addressing and routing architecture provides a framework for scalable service rendezvous, ownership of the anycast address space is preferably centralized in the CBB and/or the master AS. Although this somewhat limits the overall redundancy of the solution, it has the advantage of centralizing the management of the address space. In this model, when content providers contract with the CBB, they are allocated an address space exclusive of the CBB block of available addresses. Providers then use this anycast address space in connection with their services, eg the host portion of a uniform resouce locator (URL). Accordingly, users who click on such web links are connected to the nearest service node in CBB or its affiliates.
Naming and Service Discovery
When a service node receives an anycast request for a service, that service must be demonstrated for the requesting client. That is, the service request must be satisfied locally (if redundancy of the master service is available locally) or must be fulfilled from the master service site. One way to direct service to the master site is to iteratively apply an anycast routing scheme from above. However, an attempt to send an anycast packet to that anycast address will fail because the packet will be rerouted to the host from which it originated. In other words, anycast is trapped in the domain that received it. Thus, the system must rely on some other mechanism for communication between the remote service node and the master service site.
In one embodiment, the service node requests some database to remap the anycast address to the master service site or set of sub-services related to the service being provided. Fortunately, distributed databases already exist to implement this type of mapping into highly scalable and robust types. The Domain Name System (DNS), which handles IP hostname-to-address mappings widely in the Internet, can be easily reused and arranged for this purpose. In particular, RFC-2052 specifies a mechanism for defining certain service entries that use DNS service (SRV) resource records. By translating a numeric anycast address into a DNS domain name according to a rather well-defined and deterministic algorithm, a service node can determine the location of services using the DNS queries resolved by the present anycast name. The requested DNS configuration may be implemented by the CBB, or the CBB may delegate authority to the original content provider to configure a DNS subdomain for a particular anycast block, so that the provider may not be able to see the pits. Likewise, it is possible to configure and manage the provided services.
An alternative method is to assign a single anycast address to the CBB and include additional information about the content delivery site in the client URL. That is, anycast routing is used to capture client requests for content published via CBB, while additional information about the URL is used to specify a specific location or other attributes for the content in question. In the remainder of this disclosure, the former method (where multiple anycast addresses are assigned to the CBB) is assumed for illustrative purposes, but how the system will be simplified so that only one anycast address can be assigned to each CBB. It will be apparent to those skilled in the art whether this is possible.
To summarize, the service rendezvous problem can be solved in a scalable fashion with two interdependence mechanisms: (1) a client associates with a service infrastructure using anycast addresses and routing. (2) The service node binds to the master service site, either explicitly through client URLs or implicitly using auxiliary information transferred through a distributed directory such as DNS. Excellent scaling performance is attributed to proximity-based anycast routing and caching and a hierarchy built into DNS.
Stateful Anycasting
One of the difficulties in performing anycast service on top of IP packet service is the dynamic nature of the routing infrastructure. Because IP allows packets to be duplicated and routed along different paths, packets sent using an anycast service may be forwarded to multiple anycast nodes simultaneously and consecutive packets to one service node and to another intermittently can be transmitted to
This is particularly problematic for transport-layer protocols such as TCP, where the endpoints of a communication channel are assumed to be fixed. As an example, consider a TCP connection to a service node via an anycast address. Assuming a halfway through the connection, anycast changes so that the client's packet abruptly connects to another service node. However, since the new service node does not know about the existing TCP connection, it sends back information of "connection reset" to the client. This causes the connection to be lost and may cause fragmentation of the service requested by the client. The most important point of this problem is that while IP cannot be stated, TCP connection can be stated.
A significant amount of research has addressed this problem, and none has provided a sufficient solution for use with the present invention. It may be possible to change the TCP protocol in a way that circumvents this problem. However, it is near impossible to change the entire installed base of the millions of deployed TCP stacks on the Internet. Other approaches have asserted methods, in which routers elicit statements within the network to ensure that a TCP connection is on its original path. This involves upgrading all routers in the Internet infrastructure and is impractical as the work is still in the research phase.
In an embodiment of the present invention, a novel scheme called stateful anycasting is employed. In this way, the client uses anycast as part of the redirect service, which is by definition a short-lived approach. That is, the client contacts the anycast mention node through the anycast service, which directs the client to the service node which is normally pointed out and has a routine. Therefore, the similarity that the redirection process fails because the anycast route is not deterministic is low. If this does not occur, the redirection process is restarted and, depending on the content, a new uncontacted service is initiated. If the redirection process is designed to be a single request and a single response, the client resolves the contradiction arising from the anycasting pathology.
If the service processing is short-lived, the request for redirection is limited. That is, a short web connection can be treated entirely as a TCP anycast connection. On the other hand, long-term connections, such as streaming media, can be sensitive to routing changes, and their statable anycasting can minimize the likelihood that route changes can cause problems. If anycast-based infrastructure is widely deployed, application vendors will be motivated to provide support for anycast services. If the routing migration causes a disruption of the service, the client will transparently modify the anycasting service to cause it again.
In other embodiments, adverse effects prior to routing may be minimized by carefully treating the operating policies of the infrastructure. Thus, a large-scale anycasting infrastructure can be created where dynamic routing changes are quite frequent and in fact IP inability to anycast is minimized. In short, the declarable anycasting method described below can provide a very useful, robust and reliable service system.
A Proximity-based Redirection System
Given the structural elements described above, this section describes an embodiment of the present invention for an anycast based redirection service that combines these elements. Some of the elements of this design are generalizations of various useful deployment and deployment scenarios and are not limited to any particular technique. Other mechanisms are geared towards special services such as streaming media broadcasting or web content delivery.
An access-based redirection system: (1) allows any application-specific redirection protocol to be used between the client and the service, and (2) a glue between the redirect service, the client, the master service site, and the CBB. ), providing a service node attachment facility for any content distribution network.
CBB has a special anycast address space attributed to the master ace. Each content provider is given one or more anycast addresses from the anycast address space. Since any service can be connected to anycast using DNS, only one address is required for each particular content provider. For the purpose of explanation, the standard content provider's DNS domain is referred to as "acme.com", and CBB's domain is referred to as "cbb.net". The anycast address assigned to CCB is 10.1.18/24, and the address assigned to acme.com is 10.1.118.27. In addition, a content provider (acme.com) generates web content and live broadcast content.
The following parts include elements comprising a local structure (defined within the colo) and elements comprising a broad local structure (defined between and across ISPs) and specific redirection algorithms based on these structures and their rationale in accordance with the present invention. Describe the general principles of
The Local Architecture
This section describes an arrangement of devices for providing proximity-based redirection service within a particular ISP, such as inside Colo, for example, and how these devices are configured and how they should be interfaced with external voices.
The content distribution architecture is decomposed into two interdependent and separable components: (1) control and direction facilities; and (2) actual service functions. That is, service typically occurs by access control triggering service provision over a data connection. In addition, access control is typically designed to be redirected to an alternate IP host. Therefore, the high-level model for the system is as follows.
The client initiates access control to the anycast address to request a service.
For anycast dialog, the agent located at the endpoint directs the client to a fixed service node area (eg addressed by a standard, non-anycast IP address).
· The client connects to the service through access control to this fixed area and initiates the service movement.
The request and data handling components based on the control are very different. For example, a control element requires handling large one-time requests and rapidly redirecting those requests, and a service element requires handling a persistent connection load such as streaming media. Also, the management requirements for these two device classes are very different, just as the sensitivity of the system to failure modes is different. For example, the control element should manage server resources so that considerations such as load balancing are included in the server selection. In this regard, the control element may monitor "Server Health" to determine which servers can redirect clients. For example, server health is based on various parameters such as server capacity, loading, expected server latency, etc., and may be monitored and indirectly received by a control element and used to select a server.
5 is a diagram illustrating an embodiment of the present invention illustrating how control and service functions are separated within a particular ISP to satisfy the above requirements. In this embodiment, one or more service nodes (SNs) and one or more anycast referral nodes (ARNs) are located on LAN segment 504 in Colo 500 . The network segment 504 is connected to the rest of the ISPs and/or to the colo router 506 which is connected to the Internet 508 .
Under this configuration, a client request 510 from any host 512 on the Internet 508 is routed to the nearest ARN 514 using proximity oriented anycast routing. The ARN 514 directs the client (path 516) to the candidate service node 518 (path 520) using the techniques described herein. Since the service node is clustered, this service model determines an arbitrary client load, and allows the system to incrementally prepare for an increase in the cluster size. In addition, the ARNs themselves may be evaluated with a local load balancing device such as a layer 4 switch.
At any given point in time, if one of the ARNs is designated as the master, eg the other ARNs are designated as backup 522 . This designation can be converted. These agents may run as separate physical components or may run within a single physical device. In this example, the SNs and ARNs are connected to a single network segment 504 via a single network interface, although the system can be easily generalized so that these agents and physical devices operate multiple local network segments. Not only can each ARN advertise its routing reachability to the anycast address space occupied by the service note structure, but the master ARN actually issues the notification. Similarly, an ISP's callo router 506 coupled to the network segment 504 is configured to receive and forward this notification. The exchange of routing information is performed by IGPs such as RIP and OSPF used in ISPs. In the above example, the single ARN is the master chosen for the anycast address, and the other ARN is designated as the backup. In another example, a master is elected for each anycast address. This allows the load to be spread across multiple active ARNs while supporting a separate set of anycast addresses. Failure of any ARN causes the election process for anycast address to begin.
There are two important steps to bootstrap the system: (1) the ARN must discover the existence and address of a service node in the SN cluster; and (2) the ARN should determine which service nodes are available and which service nodes are overloaded. One approach is to provide the ARN with a list of IP addresses of the service nodes of the service cluster. Optionally, the system may use a simple resource discovery protocol based on LAN multicast, each service node on the LAN advertises its existence to a well-known multicast group, and each ARN to guess the existence of all service nodes Focus on the group. The latter approach minimizes the configuration overhead and thus avoids the possibility of human configuration errors.
According to the multicast-based resource discovery model, a new device simply connects to the network and the system automatically starts using the new device. The technique is as follows.
The ARN(s) is a well-known multicast group G<sb>S</sb>Subscribe to
SN(s) in the service cluster send messages to group G<sb>S</sb>Announcing their presence and arbitrary information like system loads by sending them to
The ARN(s) monitors messages, forms a database of useful service nodes, stores and updates any characteristics for the load balancing, etc.
· Each database entry must be "refreshed" by its corresponding SN, otherwise it will be "timed out" and cleared by the ARN.
When receiving a new service request, the ARN selects a service node from the list of available nodes in the database and causes the client to redirect to the node.
Since all devices in the service cluster are simultaneously deployed on a single network segment or LAN, the use of IP multicast does not require special routing components outside the LAN or connected to the LAN.
<b>fail recovery</b>
So the features of the described protocol are designed to implement automatic failover recovery, resulting in very high effectiveness for the service. The system is robust against failure of ARN and failure of SN.
Since the ARN stops the SN database entries in the middle, the failure of the SN is not used for service requests. So, when a client reconnects to the service (either explicitly to the user or by user interaction), the service is restarted on another service node. If the ARN maintains a persistent state for the client, the system potentially restarts the old service embodiment rather than starting a new one from scratch (e.g., failure detection and handoff of SN If this occurs, the user will not be billed twice).
If the ARN fails, another problem may occur. By maintaining redundant ARN(s) in a single colo, this problem can be addressed using the following technique.
Each ARN is a well-known multicast group G<sb>S</sb>Subscribe to
Each ARN sends messages to group G<sb>S</sb>announce their existence by sending them to
Each ARN forms an active inter-entity database and pause entries that are not replayed according to any configurable period greater than the interactive presentation period.
The ARN with the lowest numbered network address (e.g., the ARN with a network address smaller than the addresses of all other ARNs in the database) selects itself as the master ARN and increases its reachability via the IGP. to start advertising to the anycast address.
So, if the master ARN fails, the reserve ARNs hear the condition very quickly (after a single announcement interval) and a new master ARN is selected. In this regard, as a side effect of the new IGP routing advertisements, anycast packets are routed by the call router to the new master ARN.
<b>wide area structure</b>
When describing the local architecture of devices within a single colo installation, a description of the global architecture will be provided, including how the individual service-node clusters are integrated and managed across the wide area. There are two main wide area components for realizing the embodiment of the content overlay network included in the present invention, namely:
· data services that require routing and management of large amounts of data from originating content sites to service nodes in the calls; and
There is a solution to the anycast-based service rendezvous problem that involves binding the services requested by clients to the originating content sites.
How the above problem is solved, ie how data is reliably and effectively provided to service nodes across the wide area, is beyond the scope of the present invention. For example, as described in U.S. Patent Application Serial No. 60/115,454, pending, filed January 22, 1999, entitled "System for Providing Application-Level Features for Multicast Routing in Computer Networks" It can run on the same streaming broadcast network. The content may also be implemented by a file dissemination protocol based on a flood algorithm, such as the network news transfer protocol.
This publication describes how the anycast-based redirection system interfaces with useful content delivery systems. The new framework is used for service-specific interactions to be implemented between the ARN, the SN, the client, and the originating service or content site. For example, the client may initiate a web request at anycast address A, nearest ARN ad reachability is routed to address A, and then the client redirects to a selected SN with a single HTTP redirect message. Alternatively, the web may be provided directly from the ARN.
According to an embodiment of the anycast-based redirection system, the ARN may prime the SN with application-specific information that cannot be transferred to the SN from an unaltered existing client. Here, the ARN contacts the SN and establishes a certain state Q boundary in a certain port P. The port P is allocated by the SN and returned to the ARN. The client is then redirected via port P so that the unchanged client absolutely passes the state Q via the new connection. For example, Q may indicate the wide area broadcast channel for which the service node must reserve a special streaming media transport. Since the unchanged client is not directly protocol-compliant with the CBB infrastructure, the appropriate channel subscription is carried in the state transfer Q without including the client in the conversation.
In order to avoid having to change the existing installed large base client (such as web browsers, streaming media players and servers), the client-SN interaction can be implemented using an existing service-specific protocol, e.g. HTTP for Web; It is based on RTSP for streaming media protocols or other vendor-proprietary protocols.
Once a client request is initiated and blocked by the ARN, some wide area service must be called to send the content from the CBB to the local service node (if the content does not already exist). As noted above, repeated use of anycast will fail. So, the DNS system can be standardized and used to map anycast addresses to the services in a decentralized manner.
For example, it is desirable to support the storage of web objects for the content provider "acme.com", and if the CBB assigns acme.com to the anycast address 10.1.1.18.27, the pointer at the master server is
anycast-10-1-18-24.http.tcp.cbb.net SRV www.acme.com
It can be arranged in DNS with SRV resource records such as
When the service node receives a client connection request on TCP port 80 on anycast address 10.1.118.27, the service node learns that the master host for the service is www.acme.com (anycast-10-1-18- You can query the DNS SRV record for 24.http.tcp.cbb.net). If the knowledge is cached locally, and the requested content is brought to www.acme.com, it may also be cached locally. The next request at the same anycast address for the same content can be satisfied locally. The above content stored at www.acme.com may have links explicitly referencing the reference Enicast-10-1-18-27.http.tcp.cbb.net or the site may have a user friendly name, eg , www-cbb.acme.com, i.e. the name of anycast above:
www-cbb.acme.com CNAME Note that a CNAME for anycast-10-1-18-24.http.tcp.cbb.net is available.
Also consider the desirability of supporting very large-scale streaming media broadcasts from acme.com. In addition to the master web server, we have a channel allocation service (CAS) that maps streaming media URLs to broadcast channel addresses if the channel address is similar to an application-level multicast group as described in 60/115454. need to know the location of In this case, we query the DNS for SRV resource recording indicated by the CAS, and have the following form;
Anycast-10-1-18-24.cas.tcp.cbb.net Get a record that can have SRV www.acme.com.
When the ARN receives a client connection request for a streaming media URL, the ARN queries cas.acme.com to map the URL to the broadcast channel (locally caching results for outstanding client requests) Subscribe to the channel through the CBB where the channel address is known.
By storing the service binding in the DNS in this way, any anycast service node can dynamically and automatically discover the special services following a special anycast address. This knowledge eliminates the need to arrange and update service nodes in the infrastructure. This freely simplifies the anycast-based service rendezvous mechanism and the content broadcast network.
As mentioned above, the DNS SRV records store a mapping from service names to corresponding service addresses. However, there may be cases where the ARN requires more information than a simple list of servers for a named service. The additional information may specify a service node selection algorithm or a service node setup procedure. In this case, the information for the named service may be stored in a directory server (such as LDAP or X.500) or a network of web servers. Compared with DNS, the servers provide greater flexibility and scalability in data representation.
<b>redirection algorithm</b>
Given the above components and system architecture, an embodiment of the present invention is provided to describe how an end host includes a service flow or transaction from the service-node infrastructure using stateful anycasting. do.
· The user initiates a content request by clicking on a web link, eg, expressed as a URL.
The client determines the DNS name of the resource to which the URL refers. The DNS name of the resource is determined from the anycast address managed by the authority (eg, www.acme.com is the CNAME for any-10-1-18.27.cbb.net).
The client initiates a regular application connection using the anycast address, for example, a web page request using HTTP over TCP on port 80 or a streaming media request using RTSP over TCP on port 554.
· As a side effect of the anycast routing infrastructure described above, the client's packets are routed to the nearest ARN advertisement reachability at the address, initiating a connection to the ARN. The ARN is prepared to accept requests for each deployed service, for example web requests on port 80.
In this regard, if the data is useful and of transactional nature, the ARN responds directly to the content as follows or the requesting client redirects to a service node.
The ARN selects a candidate service node from the associated service cluster. The selection decision may be based on load and validity information maintained from the local monitoring protocol as described above.
The ARN performs an application-specific conversation with S as needed by the client C to prepare to connect to S. For example, in the case of live streaming media, the ARN may indicate the broadcast channel that S tunes to the CBB overlay network via a request. As part of the conversation, S returns information to the ARN where C is required to be properly redirected to S. It is clear whether the information exists and is specific to the particular service required by the nature of the information.
The ARN responds to the original client request by a redirect message directing the client C to the selected service node S.
The client C contacts S in a client-specific manner to initiate the flow or content transaction related to the desired service. For example, the client connects to S using the streaming media control protocol and initiates live transmission of streaming media via RTS.
<b>Active Session Failover</b>
A disadvantage of the above Statepearl anycast redirection scheme is that if the selected node fails for some reason, all clients provided by the node experience an interrupted service. If the client requires a maintained service, such as streaming-media provision, the video is stopped and the client tries again. According to another embodiment, the client is modified to detect the failure of the service node and request the redirection process again before the user notifies the performance degradation in the service. Here, the redirection process is called "active session failover".
6 shows a data network 800 according to the present invention. The data network 800 represents a network transaction that describes how active session failover operates to deliver content to clients without interruption.
First, the client 802 sends a service request 820 to the anycast address A and is routed to the ARN 804 . The service request 820 requests content starting from the CBB 803 . The ARN 804 decodes the service request 820 to determine an application specific redirect message 822 to be sent to the client 802 . The application specific redirect message 822 sent by the ARN 804 causes the client 802 to be redirected to a service node 806 . The client 802 then sends a request 824 to the channel description information in the client request (eg, see 60/115454) via an application-specific protocol (eg, RTSP). send a channel description message 826 to the service node 808 using let it be The result is that content flows from the service node 808 to the client as indicated by path 828 .
Hereinafter, assume that the service node 806 fails. The client 802 notices a service stop and requests the state pearl anycast procedure again: a service request 830 is sent to the anycast address A and received by the ARN 804, and a redirect message ( 832 ), and direct the client to the new node 810 . The client may request the provision of the new service from the service node 810 , as shown at 834 . The service node 810 sends a reservation message to the node 808 as indicated at 836 . If the user uses proper buffering before presenting the streaming media signal to the user (as is the common practice to counteract network delay variations), the entire process can proceed without interruption in service. When the client attaches, the client sends a packet retransmission request to the service node 810 to properly locate the stream and retransmits only packets lost during the session failover process.
For example, due to a momentary network outage, the client may inaccurately assume the failure of the service node 806 . In this case, the client may simply ignore the redirect message 832 and continue receiving service from the service node 806 .
Overview of the wide area
A potential problem with the service rendezvous mechanism described above is that the service node installation may run out of capacity because too many clients are routed to a given service node installation. This problem may be addressed by an embodiment where the redirection system is capable of redirecting a client service request across the wide area in case of overload. For example, if all local service nodes are running at a given capacity, the redirector may select a non-local-service node and redirect the client accordingly. The redirection decision is influenced by network and server health measures. According to the above scheme, the redirector sends a periodic probe message to the candidate servers to measure the network path call. As the redirector is in the vicinity of the requesting client, a redirector-to-server measurement means an accurate estimate of the corresponding network path between the client and the candidate server.
According to an embodiment of the present invention, there are three steps for performing a global redirection:
ARNs discover candidate service nodes.
· ARNs measure network path characteristics between each service node and itself.
· ARNs query their health service nodes.
Given the information obtained by the above steps, ARNs can select the service node that is most likely to provide to the requesting client the service with the best quality. To do so, each ARN maintains an information database with load information for a predetermined number of desired service nodes. The ARN consults the information database to determine the most useful service node for each client request. To maintain the load information, the ARN may actively test network paths and service nodes. Alternatively, service nodes may monitor network load and internal load and report load information to respective ARNs.
To implement local load balancing, each ARN is arranged by the IP address of a predetermined number of nearby service nodes. ARN maintains load information for the service node. However, short-range access is done when the load is geographically concentrated, causing the ARN to fully load all nearby service nodes to deny additional service requests from clients. This may occur even if a certain number of service nodes beyond the local area are not fully utilized.
The wide-area load balancing according to the present invention overcomes the above problems. According to global load balancing, each ARN maintains an information database arranged by IP addresses of all service nodes in the network and having load information for all service nodes. Alternatively, ARNs may exchange load information using a flood algorithm.
Another embodiment of the present invention utilizes a scheme called variable-region load balancing. According to the above scheme, each ARN maintains an information database for a predetermined number of preferred service nodes; The number of preferred service nodes increases with local load. That is, as nearby service nodes approach their capacity, the ARN adds information database load information for a predetermined number of service nodes beyond the current scope of the ARN. The following provides two different methods used to discover growing long distance service nodes.
According to the first method, the ARN is prepared by the IP addresses of a predetermined number of adjacent service nodes. To identify the growing distant service nodes, it is assumed that the service nodes form an actual overlay network, the ARN simply querying the list of adjacent service nodes. This approach is referred to as "overlay network crawling".
According to the second technique, each ARN and each service node is assigned a multi-part name from a hierarchical namespace. Given the names of two service nodes, the ARN may determine the closest using the longest pattern matching. For example,
The ARN specified as arn.sanjose.california.pacificcoast.usa.northamerica uses the longest right-to-left pattern matching.
than to the above service node named sn.orlando.florida.atlanticcoast.usa.northamerica
It can be determined that it is closer to the service node named sn.seattle.washington.pacificcoast.usa.northamerica. Each ARN can reproduce a directory of all service node names and their corresponding IP addresses. The directory may be implemented using DNS or analog distributed directory technology. The variable-area load balancing scheme handles geographically concentrated load by redirecting clients to growing long distance service nodes. The scheme addresses scalable conservators by minimizing the number of ARN-to-service-node relationships. That is, the ARN monitors the number of service nodes required to supply the near-term client load. Also, as the number of nodes increases to N at a distance that is generally N hops from a given node, the rate at which ARNs test candidate service nodes is adjusted in inverse proportion to the distance.
7 shows a data network 900 according to the present invention. The data network 900 includes three local networks 902 , 904 , and 906 . As described in an embodiment of the present invention, the data network 900 is arranged to provide wide area overflow.
The local network 902 includes an ARN 902 (redirector) having an associated information database 910 . The local network 902 also includes service nodes 912 and 914 . The service nodes provide information content 928 to clients 916 , 918 , 920 , 922 , 924 , and 926 .
The local networks 904 and 906 include ARNs 930 and 932, information databases 934 and 936, and service nodes 938, 940, 942, and 944, respectively. The service nodes provide the information content 928 to a number of other clients.
The ARN 908 monitors the network load characteristics of the service nodes 912 and 914 . The load information is stored in the database 910 . The ARN 908 may monitor the network load characteristics of other service nodes. According to an embodiment of the present invention, the ARNs exchange load information with each other. For example, the load characteristics of the service nodes 938 and 940 are monitored by the ARN 930 and stored in the database 934 . The ARN 930 exchanges load information with the ARN 908 as shown at 954 . According to another embodiment of the present invention, the ARNs actively test other service nodes to determine the load characteristics. These characteristics may be stored for later use. For example, the ARN 908 tests the service node 944 as indicated at 956 and tests the service node 942 as indicated at 958 . Accordingly, there are several ways in which the ARN can determine the load characteristics of service nodes located in the local network and wide area.
At some point, client 950 attempts to receive information content 928 . When an anycast request 952 is received by the ARN 908 , the client 950 sends the anycast request 952 to the network 902 . The ARN 908 causes the client 950 to redirect to one of the local service nodes 912,914. However, the database 910 associated with the ARN 908 indicates that the local service nodes are unable to provide the requested services to the client 950 . The ARN 908 may use the information database 910 to determine which service node is best suited to handle the request from the client 950 . The selected service node is not limited to the local network with the ARN 908 . Any service node over the wide area may be selected.
The ARN 908 determines whether the service node 942 should service the request from the client 950 . The ARN 908 sends a redirect message 960 to the client 950 , causing the client to redirect to the service node 942 . The client 950 sends the request to the service node 942 using a transport layer protocol such as TCP as indicated at 962 . The service node 942 responds by providing the requested information content to the client as indicated at 964 .
Thus, using the information database and the ability to test service nodes for obtaining load characteristics, the nodes can perform global load balancing according to the present invention.
<b>technical magnification</b>
In this section, additional embodiments of the present invention are described with technical extensions to the above embodiments.
<b>Last-hop multicast</b>
The use of IP multicast can be developed locally, such as to facilitate optimization in the last-hop delivery of broadcast content. So, a client can issue an anycast request and, as a result, redirect to join a multicast group.
8 illustrates an embodiment of the present invention using IP multicast. The content provider 600 provides three service nodes SN0-SN3 for providing the information content 602 via the application level multicast tree 604 . Client 606 requests the provision of services as described above received by the ARN as indicated by path 608 . The ARN retransmits a request to the service node ARN to begin transmitting data as shown in path 610 . However, rather than starting a separate data channel for each client, the service node instructs the client to subscribe to a special multicast group 612 (via the control connection) to receive the information content. Then, the client joins the multicast group and the service node SN0 transmits the information content to the group in the local environment.
As shown, the multicast traffic is repeated at a fan-out point in the distribution path from the service node SN0 to all clients receiving the flow. At the same time, the service node SN0 contacts the upstream service node SN1 and receives the information content through a unicast connection. In this way, the content is broadcast to all relevant receivers without the need to multicast over the entire network infrastructure.
<b>transmitter attachment</b>
The system described thus relies on anycast routing to route client requests to the nearest service nodes. Similarly, anycast can be used to bridge the server at the originating site of the content to the nearest service entry point. If the content server is explicitly arranged with the broadcast infrastructure, the system can be adapted to register and link service installations to the broadcast infrastructure.
Fig. 9 shows an embodiment 700 of the present invention applied to register and link service installation to a broadcast infrastructure. The service node SN0 wants to inject a new broadcast channel from the nearby server S into the content broadcast network 704 . The service node SN0 sends a service inquiry 706 using the anycast address. The service inquiry requires the identity of the service node in the broadcast network which is most useful to serve as an endpoint for a new IP tunnel from the service node SN0. The service inquiry carries anycast and is routed to the nearest ARN, in this case A1. A1 will select the most available node and also update the channel database in broadcast network 704 indicating that new channels are available through SN0. In this case, A1 selects SN1 and sends a response 708 to SN0. The response instructs SN0 to establish a new IP tunneling circuit 702 in the service node SN1.
The originator subsystem allows the overlay broadcast network to be dynamically extended to additional servers. When using the client-subsystem described above as the originator subsystem, the originator subsystem dynamically directs client-server traffic to one or more series of tunneling circuits having tunneling endpoints closest to the client and the server, respectively. Provides a comprehensive structure for mapping. The mapping is performed by a method that is unambiguous for the client and server applications, and each access router.
<b>multiple masters</b>
According to the various embodiments described above, the master AS advertises the anycast address block via a cross-domain routing protocol. There are two extensions to the above schema. First, the system is extended to allow multiple master AS's to coexist by delimiting the anycast address space between them. That is, as long as multiple examples of the system described above use a unique address space for their anycast blocks, they will be functional and non-coherent. Second, the system is extended to allow multiple master AS's to advertise the anycast address blocks or duplicate address blocks. In this case, minimum-distance anycast routing operates at intradomain and interdomain routing levels. For example, there may be a master AS in Europe and a master AS in North America that advertise the anycast blocks externally via BGP. A packet sent from any client is sent to the nearest master AS, inside the AS the packet is routed to the nearest service node or redirected across the wide area if necessary.
The present invention provides a comprehensive redirection system for content distribution based on a virtual overlay broadcasting network. Although the present invention has been described in detail only with respect to the embodiments described above, it is apparent to those skilled in the art that changes or modifications can be made within the spirit and scope of the present invention, and such changes or modifications are included in the attached patent should be limited by the claims.
10 sheets
Sheet 1 Sheet 2 Sheet 3 Sheet 4 Sheet 5 Sheet 6 Sheet 7 Sheet 8 Sheet 9 Sheet 10
29 members in 7 offices
Priority claims8
| Document | Office | Kind | Date |
|---|---|---|---|
| 15225799 | United States of America | P | |
| 15225799 | United States of America | P | |
| 60152257 | United States of America | – | |
| 09458216 | United States of America | – | |
| 45821699 | United States of America | A | |
| 45821699 | United States of America | A | |
| US19990152257P | – | – | – |
| US19990458216 | – | – | – |
Members29
| Document | Office | Kind | |
|---|---|---|---|
| WO0118641A1 | World Intellectual Property Organization (WIPO) | A1 | |
| AU7341500A | Australia | A | |
| WO0152497A2 | World Intellectual Property Organization (WIPO) | A2 | |
| AU2282701A | Australia | A | |
| WO0152497A3 | World Intellectual Property Organization (WIPO) | A3 | |
| WO0152497A9 | World Intellectual Property Organization (WIPO) | A9 | |
| KR20020048399A | Republic of Korea | A | |
| US6415323B1 | United States of America | B1 | |
| EP1242870A1 | European Patent Office (EPO) | A1 | |
| EP1250785A2 | European Patent Office (EPO) | A2 | |
| JP2003508996A | Japan | A | |
| US2003105865A1 | United States of America | A1 | |
| AU771353B2 | Australia | B2 | |
| US6785704B1 | United States of America | B1 | |
| US2005010653A1 | United States of America | A1 | |
| EP1242870A4 | European Patent Office (EPO) | A4 | |
| US6901445B2 | United States of America | B2 | |
| KR100524258B1This record | Republic of Korea | B1 | |
| JP3807981B2 | Japan | B2 | |
| EP1250785B1 | European Patent Office (EPO) | B1 | |
| DE60036021D1 | Germany | D1 | |
| EP1865684A1 | European Patent Office (EPO) | A1 | |
| DE60036021T2 | Germany | T2 | |
| US7734730B2 | United States of America | B2 | |
| EP2320619A1 | European Patent Office (EPO) | A1 | |
| EP1865684B1 | European Patent Office (EPO) | B1 | |
| EP2320619B1 | European Patent Office (EPO) | B1 | |
| EP2838240A1 | European Patent Office (EPO) | A1 | |
| EP2838240B1 | European Patent Office (EPO) | B1 |
11 legal events, as the office reported them to INPADOC
Over the term
Point at a mark for the eventEvents
| Event | Code | |
|---|---|---|
| Annual fee paymentFPAY | FPAY | |
| Annual fee paymentFPAY | FPAY | |
| Annual fee paymentFPAY | FPAY | |
| Annual fee paymentFPAY | FPAY | |
| Annual fee paymentFPAY | FPAY | |
| Annual fee paymentFPAY | FPAY | |
| Annual fee paymentFPAY | FPAY | |
| Annual fee paymentFPAY | FPAY | |
| Written decision to grantGRNT | GRNT | |
| Decision to grant or registration of patent rightE701 | E701 | |
| Request for examinationA201 | A201 |
Numbers
- Publication
- 10-0524258
- Publication, DOCDB
- 100524258
- Publication, EPODOC
- KR100524258B
- Application
- 107002895
- Application, DOCDB
- 20027002895
- Application, EPODOC
- KR20027002895
Titles2
- Korean
- 인터네트워크에서 로버스트하고 스케일어블한 서비스-노드 위치를 위한 근접-기반 방향 전환 시스템
- English
- Proximity-based redirection system for robust and scalable service-node locations in internetworks
Classification
- CPC, 10
- H04L12/18
- G06F15/173
- H04L45/306
- H04L67/1008
- H04L67/1029
- H04L67/101
- H04L69/329
- H04L67/1001
- H04L45/22
- H04L9/40
- IPC, 4
- G06F15 173
- H04L12 56
- H04L29 06
- H04L29 08