A proximity-based redirection system for robust and scalable service-node location in an internetwork
Abstract
(57) [Summary] Services within the Virtual Overlay Delivery Network (300)-Proximity-based redirection system for client attachments. The virtual overlay delivery network (300) includes addressable routers (R1-R6) that route packet traffic. The present invention includes redirectors (S1, S2) connected to at least one of the addressable routers (R3, R4), with the logic of accepting service requests from clients (312); and processing service requests. The logic that determines the selected server, which is one of several servers that can handle the service request; to the client, which redirects the service request to the selected server. Includes logic to generate directed redirection messages.
Term
Term ended
Projected expiry passed 1 September 2020, 6.1 years ago.
- Priority
- Filed
- Published
- Projected expiry
- Today
1 claim: 1 independent, 0 dependent
- 1【特許請求の範囲】 【請求項1】 パケットトラフィックをルーティングするアドレシング可能ルータを含み、データのパケットがパケットのアドレスフィールドに基づいてソースノードから目的地ノードまでルーティングされるパケットスイッチ型ネットワークにおいて、改良が、 少なくとも1つのアドレシング可能ルータに接続されたリディレクタを含み、該リディレクタが、 A)クライアントからのサービス要求を受け入れる論理であって、該サービス要求がエニーキャスト目的地アドレスへのエニーキャストメッセージである、論理と、 B)該サービス要求を処理する選択されたサーバを決定する論理であって、該選択されたサーバが該サービス要求を処理し得る複数のサーバのうちの1つである、論理と、 C)該サービス要求を該選択されたサーバにリディレクトする、該クライアントに向けられたリディレクションメッセージを生成する論理と、 を備えた、パケットスイッチ型ネットワーク。 【請求項2】 前記エニーキャスト目的地アドレスへの到達可能性をアドバタイズする論理をさらに備えた、請求項1に記載のリディレクタ。 【請求項3】 前記決定する論理が、 前記複数のサーバのネットワークトラフィック状態を監視する論理と、 該ネットワークトラフィック状態に基づいて該複数のサーバから前記選択されたサーバを選択する論理と、 を備えた、請求項1に記載のリディレクタ。 【請求項4】 前記決定する論理が、 前記複数のサーバのサーバ状態を監視する論理と、 該サーバ状態に基づいて該複数のサーバから前記選択されたサーバを選択する論理と、 を備えた、請求項1に記載のリディレクタ。 【請求項5】 前記選択されたサーバがマルチキャスティングサーバである、請求項1に記載のパケットスイッチ型ネットワーク。 【請求項6】 前記リディレクタが前記選択されたサーバである、請求項1に記載のパケットスイッチ型ネットワーク。 【請求項7】 パケットトラフィックをルーティングするアドレシング可能ルータを含み、データのパケットがパケットのアドレスフィールドに基づいてソースノードから目的地ノードまでルーティングされるパケットスイッチ型ネットワーク内のリディレクタを動作させる方法であって、 該リディレクタからエニーキャスト目的地アドレスへの到達可能性をアドバタイズするステップと、 クライアントからのサービス要求を受け入れるステップであって、該サービス要求が前記エニーキャスト目的地アドレスへのエニーキャストメッセージである、ステップと、 該サービス要求を処理する選択されたサーバを決定するステップであって、該選択されたサーバが該サービス要求を処理し得る複数のサーバのうちの1つである、ステップと、 該サービス要求を該選択されたサーバにリディレクトする、該クライアントに向けられたリディレクションメッセージを生成するステップと、 を包含する、方法。 【請求項8】 前記複数のサーバのトラフィック状態を監視するステップをさらに包含する、請求項7に記載の方法。 【請求項9】 前記決定するステップは、前記トラフィック状態に基づいて前記複数のサーバから前記選択されたサーバを決定するステップを包含する、請求項8に記載の方法。 【請求項10】 前記複数のサーバのサーバ状態を監視するステップをさらに包含する、請求項7に記載の方法。 【請求項11】 前記決定するステップが、前記サーバ状態に基づいて前記複数のサーバから前記選択されたサーバを決定するステップを包含する、請求項10に記載の方法。 【請求項12】 前記リディレクタにおいて前記サービス要求を処理するステップをさらに包含する、請求項7に記載の方法。 【請求項13】 前記生成するステップが、前記選択されたサーバにおいて、マルチキャストグループをサブスクライブするように前記クライアントをリディレクトする、該クライアントに向けられたリディレクションメッセージを生成するステップを包含する、請求項7に記載の方法。 【請求項14】 パケットトラフィックをルーティングするアドレシング可能ルータを含み、データのパケットがパケットのアドレスフィールドに基づいてソースノードから目的地ノードまでルーティングされるパケットスイッチ型ネットワークであって、 該アドレシング可能ルータのうちの少なくとも第1のルータに接続され、エニーキャストグループ内のクライアントと複数のノードとの間でデータパケットを伝搬させる論理を有する、少なくとも1つのサービスノードと、 該アドレシング可能ルータのうちの少なくとも第2のルータに接続された少なくとも1つのリディレクタであって、該少なくとも1つのリディレクタが、 A)該エニーキャストグループ内の該複数のノードに関連するエニーキャスト目的地アドレスへの到達可能性をアドバタイズする論理と、 B)該クライアントからのサービス要求を受け入れる論理であって、該サービス要求が該エニーキャスト目的地アドレスへのエニーキャストメッセージである、論理と、 C)該サービス要求を該少なくとも1つのサービスノードにリディレクトする、該クライアントに向けられたリディレクションメッセージを生成する論理と、 を備えた、パケットスイッチ型ネットワーク。 【請求項15】 前記少なくとも1つのサービスノードが、複数のサービスノードを含み、該少なくとも1つのリディレクタが、 該複数のサービスノードから、該サービス要求を処理する選択されたサービスノードを決定する論理と、 該サービス要求を該選択されたサービスノードにリディレクトする、前記クライアントに向けられたリディレクションメッセージを生成する論理と、 を備えた、請求項14に記載のパケットスイッチ型ネットワーク。 【請求項16】 前記複数のサービスノードから前記選択されたサービスノードを決定する論理が、 該複数のサービスノードにおいてネットワークトラフィック状態を監視する論理と、 該ネットワークトラフィック状態に基づいて該複数のサービスノードから該選択されたサービスノードを選択する論理と、 を備えた、請求項15に記載のパケットスイッチ型ネットワーク。 【請求項17】 前記複数のサービスノードから前記選択されたサービスノードを決定する論理が、 該複数のサービスノードにおけるサーバ状態を監視する論理と、 該サーバ状態に基づいて該複数のサービスノードから該選択されたサービスノードを選択する論理と、 を備えた、請求項15に記載のパケットスイッチ型ネットワーク。 【請求項18】 前記エニーキャストグループ内の前記複数のノードのうちの第1の部分が、第1の地理的位置に位置づけられ、該エニーキャストグループ内の該複数のノードのうちの第2の部分が、第2の地理的位置に位置づけられ、前記リディレクタが、 該エニーキャストサービス要求を送信する前記クライアントが、該エニーキャストグループ内の該ノードのうちの該第1の部分と、該エニーキャストグループ内の該ノードのうちの該第2の部分のいずれにより近いかを決定する論理と、 該クライアントが該エニーキャストグループ内の該ノードの該第1の部分により近い場合、該サービス要求を第1のサービスノードにリディレクトし、該クライアントが該エニーキャストグループ内の該ノードの該第2の部分により近い場合、該サービス要求を第2のサービスノードにリディレクトする、該クライアントに向けられた前記リディレクションメッセージを生成する論理と、 を備えた、請求項14に記載のパケットスイッチ型ネットワーク。 【請求項19】 パケットトラフィックをルーティングするアドレシング可能ルータを含み、データのパケットがパケットのアドレスフィールドに基づいてソースノードから目的地ノードまでルーティングされるパケットスイッチ型ネットワークであって、該アドレシング可能ルータのうちの少なくとも1つと少なくとも1つのサービスノードとに接続されたリディレクタを備えたパケットスイッチ型ネットワークを動作させる方法であって、 該リディレクタからエニーキャスト目的地アドレスへの到達可能性をアドバタイズするステップと、 該リディレクタにおいてクライアントからのサービス要求を受け入れるステップであって、該サービス要求が該エニーキャスト目的地アドレスへのエニーキャストメッセージである、ステップと、 該サービス要求を該少なくとも1つのサービスノードにリディレクトする、該クライアントに向けられたリディレクションメッセージを生成するステップと、 を包含する、方法。 【請求項20】 前記少なくとも1つのサービスノードが複数のサービスノードを含み、前記生成するステップが、 該複数のノードから、前記サービス要求を処理する選択されたサービスノードを決定するステップと、 該サービス要求を該選択されたサービスノードにリディレクトする、前記クライアントに向けられたリディレクションメッセージを生成するステップと、 を包含する、請求項19に記載の方法。 【請求項21】 前記決定するステップが、 前記複数のサービスノードにおいてネットワークトラフィック状態を監視するステップと、 該ネットワークトラフィック状態に基づいて該複数のサービスノードから前記選択されたサービスノードを選択するステップと、 を包含する、請求項20に記載の方法。 【請求項22】 前記決定するステップが、 前記複数のサービスノードにおいてサーバ状態を監視するステップと、 該サーバ状態に基づいて該複数のサービスノードから前記選択されたサービスノードを選択するステップと、 を包含する、請求項20に記載の方法。 【請求項23】 前記エニーキャストグループの前記複数のノードの第1の部分が第1の地理的位置に位置づけられ、該エニーキャストグループの該複数のノードの第2の部分が第2の地理的位置に位置づけられ、前記生成するステップが、 該エニーキャストサービス要求を送信する前記クライアントが、該エニーキャストグループの該ノードの該第1の部分と、該エニーキャストグループの該ノードの該第2の部分のいずれにより近いかを決定するステップと、 該クライアントが該エニーキャストグループの該ノードの該第1の部分により近い場合、該サービス要求を第1のサービスノードにリディレクトし、該クライアントが該エニーキャストグループの該ノードの該第2の部分により近い場合、該サービス要求を第2のサービスノードにリディレクトする、該クライアントに向けられた前記リディレクションメッセージを生成する、請求項19に記載の方法。
138 paragraphs, as filed
Description: TECHNICAL FIELD [Detailed description of the invention]
【0001】
(Cross-reference of related applications) This application claims priority from the simultaneously pending US Provisional Patent Application No. 60 / 152,257 filed on September 3, 1999. This application is the US patent application No. 09 / 323,869 filed on June 1, 1999 under the name of "PERFORMING MULTICAST COMMUNICATION IN COMPUTER NETWORKS BY USING OVERLAY ROUTING" and "SYSTEM FOR PROVIDING APPLICATION-LEVEL FEATURES TO MULTICAST ROUTING". Related to US Provisional Patent Application No. 60 / 115,454 filed on January 1, 1999 under the name "IN COMPUTER NETWORKS". The entire disclosure of each of the above applications is incorporated herein by reference in its entirety for all purposes. [0002]
(Field of invention) The present invention relates generally to the field of data networks, and more particularly to the distribution of information in data networks. [0003]
(Background of invention) One of the central challenges to scaling the Internet infrastructure for mass adoption is from sourcing sites by efficient, viable, and cost-effective ways. The problem is delivering arbitrary content to many users of the Internet. Dissemination of general news articles, video broadcasts, stock quotes, new releases of popular software, etc., all try to find the same content from the same server at about the same time, with a large number of users across the network. , So-called "flash effect" can result. Traffic flash not only gives in to the server, but also wastes network bandwidth. This is because many duplicate copies of the same content flow across a wide area network. For example, breaking news on CNN's website event) can cause millions of users to fetch the text of an article from its server. Similarly, a premier show on the Internet of well-known movie broadcasts can also encourage millions of users to attempt to access media content servers. [0004]
To overcome the problems caused by the flash effect, two main mechanisms for the web, caching and server replication, have been proposed. In caching, the cache is placed in a strategic location within the network infrastructure to thwart content requests from clients. When the cache receives a content request, it stores the content. check content). Then, if the requested data exists, the cache fulfills the request locally. Otherwise, the request is relayed to the original server and the response is relayed back to the client. During this process, the cache stores the response in the cache's local storage. Many strategies for managing local enclosures have been proposed, such as deciding when to dispose of objects from the cache and when to refresh objects that may differ from the server. The cache can be non-transparent, in which case the client is explicitly configured by the network address of the cache. Alternatively, the cache can be transparent, in which case the client is ignorant of the cache, and the cache transparently interferes with content requests, for example using layer 4 switches. [0005]
In server replication, servers are deployed across a wide area and clients are assigned to these delivered servers to balance loads and save network bandwidth. These replicated servers can have some or all of the content contained on the original server, how the particular configuration of the server is deployed, how the content is from the master server to the server. There are many variations on how it is delivered and how clients are assigned to the appropriate servers. [0006]
Many technologies deployed to support these types of server replication, as well as caching technologies, are exceptional and do not fit into the underlying Internet architecture. For example, common techniques for transparent caching make TCP meaningless and therefore incompatible with certain modes of IP packet services, such as the underlying multipath routing. This leads to a number of difficult management problems and, in detail, does not provide a close network architecture that can be managed from the network operations center in a reasonable way. [0007]
A similar content delivery problem involves delivering livestreaming media to many users throughout the Internet. Here, the server produces a live broadcast feed, and the client connects to the server using the streaming media transport protocol to receive the broadcast. However, as more and more clients tune in to broadcast, servers and networks are overwhelmed by the task of delivering many packet streams to many clients. [0008]
One solution to this live broadcast problem is to affect the efficiency of network layer multicast, or IP multicast as defined in the Internet architecture. In this approach, the server sends a single packet stream to a "multicast group" rather than sending individual copies of the stream to individual clients. Receivers interested in the stream in question then broadcast by "subscribing" to the multicast group (eg, by signaling subscription information to the nearest router using the Internet Group Management Protocol, IGMP). To "tune". The network efficiently delivers the broadcast to each receiver by copying the packet from the source to all receivers only at the fan out point in the delivery path. Therefore, only one copy of each packet will appear on any physical link. [0009]
Unfortunately, a wide variety of deployment and scalability issues have disrupted the acceptance and proliferation of IP multicast on the global Internet. A uniformly consistent diagram of what the tree looks like for all routers in the network Many of these issues are fundamentally related to the fact that having a view) requires the calculation of the multicast distribution tree. In multicast, each router needs to have an accurate local diagram of a single, globally consistent multicast routing tree. Routing loops and black holes are unavoidable if routers have different views of a given multicast tree in different parts of the network. For example, many other issues such as multicast address allocation, multicast congestion control, and reliable multicast delivery also made it difficult to deploy and approve IP multicast. In the last few years, despite substantial advances in the commercial deployment of multicast, the resulting infrastructure remains relatively relatively. It is brittle and its reach is extremely limited. [0010]
In addition to the substantial technical barriers to the deployment of ubiquitous Internet Multicast services, there are also business and economic barriers. Internet service providers have been less successful in providing wide area multicast services. This is because managing, monitoring, and delivering multicast traffic is quite difficult. Moreover, it is difficult to control who in a multicast session can generate traffic and what part of the network the traffic can reach. Due to these barriers, multicast services that reach the better parts of the Internet are unlikely to emerge at any time. Even if it does appear, it will certainly take many years for the process to unfold. [0011]
To avoid the pitfalls of multicast, others have proposed enabling streaming media broadcasts with an application-level solution called a splitter network. In this approach, a set of servers distributed throughout the network is placed at strategic locations within the service provider's network. These servers are provided with "splitting" capabilities, which allow a given stream to be replicated to a large number of downstream servers. With this capability, the servers can be configured in a tree-like hierarchy, where the root server supplies streams to a large number of downstream servers, and then streams the streams to yet another tier of downstream servers. Divide into a large number of copies transferred to. [0012]
Unfortunately, the server splitter network is plagued by a number of issues. First, the splitter tree is statically configured, which means that if a single splitter breaks, the entire subtree below the break point loses service. Second, the splitter networks need to be directed to a single broadcast center, and individual splitter networks consisting of well-defined physical servers need to be maintained for each broadcast network. Third, because splitter extraction is based on an extension of the media server, it is necessarily platform dependent, for example, RealNeworks-based splitter networks cannot deliver Microsoft Netshow traffic. Fourth, splitter networks are extremely inefficient in bandwidth. This is because the splitter network does not track the interest of the receiver and does not truncate traffic from the subtree of the splitter network that does not have a downstream receiver. Finally, the splitter network is a policy control. There is a lack of control), and it is not possible to control the total bit rate consumed along the path between the two splitter nodes and assign the flow to different classes in a stream-aware manner. [0013]
(Gist of the invention) To address the wide variety of issues described above, one embodiment of the present invention provides a comprehensive redirection system for content distribution in a virtual overlay broadcast network (OBN). In this system, service nodes are placed in strategic locations throughout the network's infrastructure, but unlike previous systems, these service nodes are closely coordinated, and managed throughout a wide area. Adjusted to a virtual overlay network. Service nodes are peered to each other throughout the IP tunnel, exchanging routing information, client subscription data, configuration controls, bandwidth readiness, and more. At the same time, the service node can handle application-specific requests for content, for example, it looks like a web server or streaming media server, depending on the nature of the support service. In short, service nodes have multiple roles. That is, it acts as both a server and an application-level content router. [0014]
In one embodiment of the invention, improvements to the packet switch network are provided. The packet switch network includes addressable routers for routing packet traffic, where packet data is routed from a source node to a destination node based on the packet's address field. The improvements include a redirector coupled to at least one of the addressable routers, the logic for receiving service requests from clients, the logic for determining the selected server for processing service requests (selected server). Is one of several servers capable of processing service requests), which includes logic to generate a redirect message directed to the client to redirect the service request to the selected server. [0015]
(Description of a specific embodiment) The comprehensive redirection system of the present invention works in cooperation with service nodes installed in strategic locations throughout the network infrastructure. By linking these service nodes over a wide area, the virtual overlay network will have connectivity, coordination, and manageability. Overlay network architectures are based on design philosophies similar to the design philosophies of the underlying Internet architecture (for example, this design philosophies utilize extensible addressing, adaptive routing, hierarchical naming, decentralization, etc.). This allows the overlay architecture to also enjoy the high lossability, scalability and manageability that is obvious to the Internet itself. Unlike physical internetwork, where routers are directly attached to each other over a physical link, service nodes in a virtual overlay network communicate with each other using packet services provided by the underlying IP network. .. Therefore, virtual overlays are highly extensible. Because, in order to provide excellent content delivery performance, a large area of the network (for example, the entire backbone of the ISP) made up of many individual components (routers, switches and links, etc.) is a small number of services. This is because only nodes may be needed. [0016]
Another important aspect of the invention is the "glue" interface between the client wishing to receive the information content and the service node actually transmitting the information content. That is, a "glue" interface is a mechanism that allows a client to attach to a service node, request a particular piece of content, and have this content effectively transmitted. This is sometimes referred to as the "service rendezvous-" problem. [0017]
Basically, the service rendezvous will (1) announce one name for the service, (2) repeat this service across the network, and (3) provide this service from the best server to each client who wants the service. Accompanied by a system that can be received. To measure millions of clients, the service rendezvous mechanism needs to effectively deliver client requests to widespread service nodes and balance the load. In addition, effective use of network bandwidth requires minimizing the number of network links through which the content reaches the requesting client. What is argued from both of these points is that the client must be directed to a neighboring service node that can execute it on request. If there is no nearby service node capable of executing it on request, the system redirects the client to service nodes throughout the wide area network to execute it on request. Must be possible. In addition, it must be possible to cluster service nodes at a particular location and connect clients to individual nodes in the cluster based on traffic load conditions. That is, in order to obtain the desired result, the service rendezvous system needs to provide a mechanism for server selection and redirect to perform load balancing. If the local cluster becomes overloaded, the server selection needs to be compensated to load the balance over a wide area. [0018]
Unfortunately, service rendezvous is a difficult problem. This is because the Internet architecture deliberately hides the underlying structure of the network that imposes flexibility and lossability on higher layer protocols, making it difficult to detect and use the server selected for a particular network transaction. That's why. To overcome these problems, the rendezvous services described herein are used to route user requests to neighboring service nodes based on "anycast" routing, or phase locality. Take advantage of possible network-level mechanisms. [0019]
The concept of anycast packet forwarding is well known in the literature of network research, but due to compatibility issues with existing packet forwarding networks, the practical application of this concept is limited to a narrow range. In general, the operation of the Internet complies with the standards that are ultimately agreed upon. This standard described herein is referred to as a "request for comment" (RFC). RFCs applicable to the operation of the Internet include RFC-1546 and others. [0020]
At the highest level, there are two main approaches to performing anycast packet forwarding. The first approach is to introduce a particular kind of anycast address, create a new routing protocol, and activate an interface that is "anycast aware". This obviously involves a long process of standardization, adoption of router vendors, and so on. The second approach is to reuse part of the existing unicast address space. However, the second approach has two corresponding technical challenges that have not yet been resolved. These conundrums are: (1) Stateful transport A conundrum that supports protocol) and (2) a conundrum that supports anycast routing and collection of routes between domains. Fortunately, embodiments of the present invention provide novel solutions to these technical challenges. For example, solutions to problems supporting stateful distribution protocols are presented in the section entitled "Stateful Anycasting" herein. Solutions to problems that support interdomain anycast routing and set of routes are presented in the section entitled "interdomain anycast routing" herein. [0021] [0021]
The rendezvous service described herein assumes that base packet forwarding is not "anycast along." However, systems based on "anycast ware" networks are also feasible. Anycast packet forwarding is used to forward packets from the client to the closest instance of the rendezvous service. [0022]
Instead, one embodiment of the invention is static and temporary from a complete, dynamic framework, where the host can dynamically join and leave anycast groups. Simplify the anycast service model to the framework (where only hosts specifically configured within the network infrastructure are members of the anycast group). In this static, temporary framework, anycast address assignment, allocation and promotion is for the central authority, associating a large block of anycast addresses with one well-connected backbone network. The backbone network can be called the content backbone (CBB). [0023]
Another advantage included in the embodiments of the present invention is that the client attaches to the content distribution network at an explicit, per-client service access point. This allows infrastructure to perform user-specific certification, monitoring, customization, promotion, and more. In contrast, the pure multicast-based approach does not provide any of these features, even though it is extensible because the multicast receiver subscription process is completely anonymous. [0024]
In summary, virtual overlay networks built with anycast-based service rendezvous enjoy the following attractive characteristics: -The service access mechanism is highly extensible. This is because the closest service node is detected using anycast, which can be achieved with standard routing protocols deployed in the new configuration. -This system saves substantial bandwidth. This is because the request can be routed to the nearest service node, which minimizes the number of network links that the content must go through. The Service Infrastructure provides fine-grained control, monitoring and customization of client connections. The integration and construction of this infrastructure is highly decentralized, which facilitates large-scale deployment across heterogeneous environments managed by a diverse range of controlling entities. The system is very available and lost when anycast is built on standard applicable routing protocols and service elements are clustered for redundancy. This ensures that requests are routed only on servers that are functioning properly and that advertise their availability. -This content broadcast network can be arranged incrementally. It is possible to first build an anycast-based redirect service within the content broadcast backbone, and then build a personal-based redirect service within the relevant ISP to capture increasing user demands. Because it is possible. [0025]
The following sections of the specification describe in detail the architectural model and embodiments of various system components used to implement anycast-based redirection systems for virtual overlay broadcast networks included within the present invention. [0026]
(Network architecture) Figure 1 is an example of typical components and interconnects, including parts of the Internet 100. Internet Service Providers (ISPs) 101, 102, and 103 provide Internet access. A typical ISP is an IP-based network over a wide area to connect individual customer networks 104 and / or individual users to the network via an access device 106 (eg, DSL, telephone modem, cable modem, etc.). To operate. A typical ISP can also peer through the exchange point 108 with another ISP to allow data traffic to flow from one ISP user to another. The set of internal IP routers 110 interconnected with the communication link 111 provides connectivity between users within the ISP. The special border router 112 located at the exchange point forwards non-local traffic into and from the ISP. Often, individual ISP networks, such as ISP103, are referred to as autonomous systems (ASs) because they represent independent and aggregateable units in terms of network routing protocols. Within an ISP, intradomain routing protocols run (eg, RIP or OSPF), and across ISPs, interdomain routing protocols run (eg, BGP). The term "intra-domain protocol" is often used interchangeably with the term Interior Gateway Protocol (IGP). [0027]
As the Internet and the World Wide Web (Web) develop, ISPs combine two innovative architectural concepts together, namely: (1) actively peering with a large number of adjacent ISPs at each exchange point. By doing so and (2) installing a data center (for example, a web server) containing application services near these exchange points, it was possible to obtain better inter-terminal network service performance. Thus, the "colo" at each peer connection point replicates the application service at each peer connection point so that the user enjoys a high-speed connection to nearby services wherever they are in the network. Make it possible. [0028]
Figure 2 shows a typical architecture for the overlay ISP200, where the service network is constructed in this way, forming an overlay structure across a number of existing ISPs, for example ISP202, 204, and 206. Will be done. The overlay ISP200 combines with the existing ISP and further with the data center (DC) 210 via router 208. Overlay ISPs lend machine space and network bandwidth to content providers who have servers in ancillary facilities (colo) located in DC. [0029]
In summary, the natural construction block for CBB is the ISP facility (colo). In an embodiment of the invention, service nodes are housed in annex equipment (colo) and placed in a wide area overlay structure using available network connections. However, the service node does not have to be located at a special colo site and can actually be on any part of the network. Ancillary equipment (colo) is a convenient and effective placement channel for service nodes. [0030]
(Annycast routing between domains) FIG. 3 shows a network 300 configured to implement anycast routing according to the present invention. Network 300 is a router (R1 to R6), two server devices S<sub>1</sub>And S<sub>2</sub>, As well as two clients C<sub>1</sub>And C<sub>2</sub>To be equipped. In one embodiment of Network 300, both server devices go through an IGP to address block "A / 24" (ie, A is a 24-bit prefix for a 32-bit IPv4 address). Advertise reachability. Thus, the two server devices utilize routing advertisements to reflect server availability in the infrastructure of network 300. Router R<sub>4</sub>And R<sub>3</sub>Are configured to listen for reachability advertisements on their attached LAN 302 and 304, respectively. As a result of the IGP calculation, routers R1 to R6 in the network remember the shortest path from each client to the server via the address corresponding to the "A" prefix. Therefore, client C<sub>2</sub>Is the address A<sub>1</sub>When you send a packet to (where A<sub>1</sub>Prefix is A), Router R, as shown in Route 310<sub>2</sub>Will it router R<sub>4</sub>Transfer to server S, then transfer it to server S<sub>2</sub>Transfer to. Similarly, client C<sub>1</sub>From A<sub>1</sub>Packets sent to Server S, as shown in Route 312.<sub>1</sub>Routed to. Server S<sub>2</sub>If fails, S for A / 24<sub>2</sub>The advertisement from is over and the network recalculates the corresponding shortest path to A / 24. As a result, C because no other node advertises such a route, as shown by route 314.<sub>2</sub>From A<sub>1</sub>Packets sent to Server S<sub>1</sub>Routed to. [0031]
One of the problems caused by the anycast routing scheme described above is how the anycast route propagates over a wide area to any site that may not consist of anycast-based service nodes. Is it to do? Rather than requiring new infrastructure for anycast routing, embodiments of the present invention "own" a given anycast address block by a single AS, a conventional interdomain protocol, ie. , Advertise with BGP, simply use the framework to influence existing interdomain routing systems. The other independent ASs then have an anycast recognition service so that the IGP for those ASs routes the packets sent to that anycast address block to the service node within that single AS. It can be incrementally configured with anycast-aware service nodes. [0032]
To do this, the Content Backbone (CBB) is located in the "master" AS, which owns the anycast address block and advertises it to the Internet using BGP. That is, the ISP creates a block of pre-existing but unused address space (or requests a new address from the Internet Assigned Numbers Authority) and puts this block in this block. Assign to a CBB that declares that the block will not be used for anything other than anycast routing. The master AS uses BGP again to advertise anycast blocks (this block is called an "A") over a wide area, as if it were a normal IP network. Therefore, in the configuration described so far, an arbitrary packet transmitted from an arbitrary location in the Internet to an address in the block A is routed to the master AS. [0033]
To provide the underlying services of the anycast routing infrastructure, the CBB places service nodes within the master AS and uses the master AS's IGP to advertise their reachability to A. Configure the node. Once this part is in the right place, once the packet enters the master AS (from somewhere on the internet), it goes to the CBB service node closest to the border router that the packet passed through when it entered the master AS. Is routed. Assuming that the master ASs are dense and peer connected, most users on the Internet will enjoy a fast route with less delay to the service nodes in the master AS (CBB). [0034]
The architecture described so far provides a service node location based on proximity to nodes located within the master AS to an executable mechanism, but the system is limited by the fact that all nodes are in that master AS. Will be done. A more scalable approach allows service nodes to be installed within the network of other ISPs. To do so, affiliate ASs (ie, ISPs that support collective services but not master ASs) simply install service nodes in exactly the same way as master ASs. However, affiliates use IGPs to advertise anycast blocks only within that domain and to advertise anycast blocks outside that domain to their peers. Another embodiment provides an extension to this scheme in which multiple ASs advertise anycast blocks within BGP (ie, an external routing protocol). This extension is described in another section of this document. [0035]
Figure 4 shows the master AS400 and affiliate networks 402, 404, and 406 configured to achieve interdomain anycast routing. Master AS400 is a service node A based on anycast<sub>1</sub>, A<sub>2</sub>, And A<sub>3</sub>And connect with three affiliate networks via router 408. 4 clients C<sub>1</sub>, C<sub>2</sub>, C<sub>3</sub>, And C<sub>4</sub>However, as shown, it comes with the affiliate. Affiliates 402 and 406 do not have an internally located service node, while affiliate 404 is a single service node A that is part of its infrastructure.<sub>4</sub>Have. Therefore, assuming that unicast interdomain and intradomain routing protocols behave normally, C<sub>1</sub>Packets sent from to block A are shown by path 410, as indicated by A.<sub>1</sub>Sent to, but C<sub>2</sub>Packets sent from to block A are indicated by route 412, A.<sub>2</sub>Routed to. Routes 410 and 412 represent the shortest interdomain route from affiliate 402 to master AS400. In contrast, client C<sub>3</sub>Packets sent from to block A to service node A, as indicated by route 414.<sub>4</sub>Routed to. This is because the IGP in Affiliate 404 is Service Node A<sub>4</sub>In addition, block A makes it possible to advertise the reachability of "hijack" packets sent to that address. Similarly, C<sub>4</sub>Packets sent from to block A also have a route from affiliate 406 to master AS400 that crosses affiliate 404, so as indicated by route 416, A<sub>4</sub>Is "hijacked" by. This is a well-thought-out desired feature of the architecture according to the invention, as it measures and distributes the load of the system without the need for anycast intelligence, which is placed everywhere for precise operation. [0036]
Anycast addressing and routing architectures provide a framework for scalable service rendezvous, but ownership of anycast address spaces is preferably concentrated in the CBB and / or master AS. This limits the overall flexibility of the solution to some extent, but has the benefit of centralized address space management. In this model, when the content provider signs up with CBB, a dedicated anycast address space is allocated from the block of CBB with available addresses. Content providers then refer to those services and use this anycast address space, for example, as the host portion of a uniform resource locator (URL). Therefore, the user who clicks on such a web link is directed to the nearest service node within the CBB or its affiliate. [0037]
(Naming and Service Discovery) When a service node receives an anycast request for a service, the service must be materialized for the requesting client. That is, the service request must be met locally (if an extension of the master service is available locally) or initiated from the master service site. One way to locate services at the master site is to repeatedly apply the anycast routing architecture from above. However, since the packet is routed back to the host from which the packet came out, the attempt to send the anycast packet to the corresponding anycast address fails. In other words, the anycast packet is confined within the domain in which it was received. Therefore, the system must rely on some other mechanism for communication between the remote service node and the master service site. [0038]
In one embodiment, the service node must query some database to map the anycast address to the master service site, or even a collection of subservices for the subscribed service. Fortunately, there are already distributed databases that perform this kind of mapping in a very scalable and robust way. Overall, the Domain Name System (DNS), which handles IP hostname-to-address mapping within the Internet, can be easily reused and configured for this purpose. More specifically, RFC-2052 uses DNS Service (SRV) resource records to define a method for defining arbitrary service entries. By translating the numeric anycast address into a DNS domain name according to some clear deterministic algorithm, the service node can use this anycast name keyed DNS query to locate the service. Can be done. The required DNS configuration can be performed by the CBB, or the CBB can delegate authority to configure a DNS subdomain for a particular anycast block to the original content provider, which provider will apply for. Services can be configured and managed as appropriate. [0039]
An alternative is to assign only a single anycast address to the CBB and embed more information about the content-generating site in the client URL. That is, anycast routing is used to capture client requests for any content published via the CBB, and further information about the URL is used to determine the specific location or other attributes for the content in question. Be identified. In the rest of this disclosure, the former method (multiple anycast addresses are assigned to the CBB) is assumed for illustration purposes, but those skilled in the art will only have a single anycast address. It is clear how the system will be simplified to be assigned to each CBB. [0040]
In short, it is either explicitly carried through two independent mechanisms: (1) a client coupled to a service infrastructure that uses anycast addresses and routing, and (2) a client URL, or like DNS. The service rendezvous problem is solved in a way that it can be scaled by using a service node that is coupled to a master service site that uses auxiliary information that is implicitly carried through a distributed directory. Outstanding scaling performance is provided by proximity-based anycast routing and caching, as well as hierarchies built within DNS. [0041]
(Stateful Anycasting) One of the problems in implementing anycast services on IP packet services is the dynamic nature of the underlying routing infrastructure. Packets sent using the anycast service may be delivered to multiple anycast service nodes at the same time, as the IP allows packets to be replicated and routed (especially) along different routes. However, continuous packets may be delivered intermittently from one service node to the next. [0042]
This is particularly problematic for transport layer protocols such as TCP, where the endpoints of the communication channel are considered to be fixed. For example, consider a TCP connection to a service node via an anycast address. In the middle of the connection, the anycast route changes so that the client's packets are suddenly routed to a different service node. However, the new service node has no knowledge of the current TCP connection and therefore sends a "reset connection" signal back to the client. This can disrupt the connection and disrupt the service the client was trying to call. The difficulty with this problem is that the TCP connection is stateful and the IP is stateless. [0043]
Although considerable research has been discussed on this issue, no sufficient solution has been produced for use with the present invention. It may be possible to modify the TCP protocol to work around this issue. However, it is nearly impossible to modify the entire installed base of millions of TCP stacks distributed over the Internet. Other approaches have argued how routers can stay in the network to ensure that anycast TCP connections maintain their original path. This method is also impractical because it requires updating all routers in the Internet infrastructure and the work is still largely in the research stage. [0044]
In one embodiment of the present invention, a novel method called stateful anycasting is adopted. In this approach, the client uses anycast only as part of a redirect service, which is clearly a short-lived transaction. That is, the client contacts the anycast reference node through the anycast service, and the reference node redirects the client to the normal address and routing (unicast) service node. Therefore, it is unlikely that the redirect process will fail due to uncertainties in the underlying anycast route. If this happens, the redirection process can be restarted by the client or, depending on the context, by the new service node contacted. If the redirection process is designed for a single request and a single response, the client can easily resolve any discrepancies that arise from any casting medical conditions. [0045]
If the service transaction is short-lived (for example, data can be transferred after a few round trips), the need for redirection is limited. That is, the entire short web connection can be treated as a TCP anycast connection. On the other hand, long-lived connections such as streaming media are susceptible to route changes, while stateful anycasting can cause problems with route changes (ie, the need for changes to occur during the redirection process). Minimize. However, if infrastructure-based anycast is widely distributed, the application supplier has an incentive to provide support for the anycast service, in which case the client causes the routing transition to cause some sort of service disruption. At times, it can be modified to transparently recall the anycast service. [0046]
In other embodiments, the adverse effects of routing transitions are minimized by carefully designing the behavior of the infrastructure. Therefore, as described herein, dynamic routing changes are so frequent that, in practice, large-scale anycasting minimizes the problems caused by the statelessness of IP for anycast. Infrastructure can be built. That is, the stateful anycasting methods described herein can provide a widely available, robust, and reliable service rendezvous system. [0047]
(Proximity-based redirection system) Given the above architectural components, this section describes certain embodiments of the invention relating to anycast-based redirect services that combine these components. Obviously, some of the components of this design can be generalized to a variety of useful configuration and distribution scenarios and are not limited to the specific description herein. Other mechanisms are particularly suitable for certain services such as streaming media broadcasting or web content distribution. [0048]
Proximity-based redirection systems allow (1) any application-specific redirection protocol to be used between clients and services, and (2) redirection services, clients, master service sites, and CBBs. By providing a glue between and, it provides a service node installation facility for any content distribution network. [0049]
The CBB owns a specific anycast address space routed within the master AS. Each content provider is assigned one or more anycast addresses from the anycast address space. Only one address is required for each individual content provider, as any service can be coupled to anycast addresses using DNS. For illustration purposes, assume that the DNS domain of the Canonical Content Provider is "acme.com" and the domain of CBB is "cbb.net". The anycast address block assigned to CCB is 10.1.18/24 and the address assigned to acme.com is 10.11.8.27. It is further assumed that the content provider (acme.com) produces web content, on-demand streaming media content, and live content. [0050]
The following sections are the components according to the invention that include local architectures (specified by colo), components that include a wide range of architectures (specified between ISPs or across ISPs), and these architectures and theirs. A particular redirect algorithm based on the underlying general principles will be described. [0051]
(Local architecture) This section describes the configuration of devices that support proximity-based redirection services within a particular ISP, such as inside a roller, and how these devices are configured and interface with external components. [0052]
The content delivery architecture is essentially decomposed into two interdependent but separate components: (1) control and redirection equipment, and (2) actual service functionality. That is, the service is typically invoked by a control connection to trigger delivery of the service across the data connection. In addition, the control connection can typically be modified to be redirected to the IP host. Therefore, the high-level model for the system is as follows. [0053]
-The client initiates a control connection to the anycast address to request the service. [0054]
The agent at the endpoint for that anycast dialogue redirects the client to a fixed service-node location (ie, addressed by a standard, non-anycast IP address). [0055]
-The client attaches the service and initiates the service transfer via the control connection to this fixed location. [0056]
The requirements placed on control and data processing components vary widely. For example, the control element needs to handle a large number of short-lived requests and redirect the request quickly, and the service element needs to handle a persistent load of persistent connections such as streaming media. Also, the management requirements for these two device hierarchies are very different, as are the sensitivity of the system to failure modes. For example, the control element needs to manage server resources so that load balancing is involved in server selection. In this regard, the control element may monitor the "server health" to determine which server is redirected to the client. For example, server health is used to make server selection decisions based on various parameters such as server capacity, loading, expected server delay, etc. that can be indirectly monitored and received by control elements. .. [0057]
FIG. 5 shows an embodiment of the present invention showing how control and service functions are separated within a particular ISP to meet the above requirements. In this embodiment, the service cluster 502 of one or more service nodes (SN) and one or more anycast reference nodes (ARN) is located in the local area network segment 504 within the roller 500. The network segment 504 is coupled to the color router 506 and then to the rest of the ISP and / or Internet 508. [0058] [0058]
Under this configuration, client request 510 from any host 512 on Internet 508 is routed to the nearest ARN514 using proximity-based anycast routing. ARN514 redirects the client (route 516) to candidate service node 518 (route 520) using techniques within the scope described herein. In this service model, service nodes are clustered, so the system can be provided in an increasingly large size by scaling to any client load and increasing the cluster size. In addition, the ARN itself can be scaled with local load balancing devices such as Layer-4 switches. [0059]
At any given time, one of the ARNs (eg ARN514) is designated as the master and the remaining ARNs are designated as backup 522. This designation can change over time. These agents can be implemented as individual physical components or can all operate within one physical device. In this example, it is assumed that the SN and ARN are attached to a single network segment 504 via a single network interface, but the system has multiple local-network segments for these agents and physical devices. It can be easily generalized to work across networks. Each ARN has reachability to route ads to anycast address spaces owned by the service-node infrastructure, but only the master ARN actively generates ads. is there. Similarly, the ISP's colo router (s) 506 attached to network segment 504 are configured to listen for and propagate these advertisements. This exchange of routing information is performed regardless of the IGP used within the ISP (for example, RIP, OSPF, etc.). In the above embodiment, a single ARN is elected as the master for all anycast addresses and the remaining ARNs serve as backups. In another embodiment, a master is elected for each anycast address. This allows the load to be distributed across multiple active ARNs, with each ARN functioning as a set of disjoint anycast addresses. If any ARN fails, the election process begins for that anycast address. [0060]
There are two main steps involved in bootstrapping the system: (1) the ARN (s) must discover the existence and address of service nodes in the SN cluster; And (2) The process by which the ARN (s) must determine which service nodes are available and not overloaded. One approach is to configure the ARN to include a list of IP addresses of service nodes in the service cluster. Alternatively, the system can use a simple resource discovery protocol based on local-area network multicast, in which case each service node advertises its existence on a well-known multicast group and each ARN listens to the group and infers the existence of all service nodes. The latter approach avoids the possibility of human configuration errors by minimizing the configuration overhead. [0061]
Using this multicast-based resource discovery model, simply plug a new device into the network and the system will automatically use the new device. This technology works as follows. [0062]
· ARN (singular or plural) is a well-known multicast group G<sub>s</sub>Join. [0063]
· SNs (s) in the service cluster group G messages<sub></sub><sub>s</sub>Advertises its existence and optional information (eg, system load) by sending to. [0064]
ARN (s) monitor these messages, build a database of available service nodes, store and update optional attributes used for load balancing, etc. [0065]
-Each database entry requires an "update" by the corresponding SN. Without an "update", each database entry will "time out" and be deleted by the ARN (s). [0066]
After receiving a new service request, ARN selects a service node from the list of available nodes in the database and forwards the client to that node. [0067]
Because all devices in a service cluster are co-located on a single network segment or LAN, they are special to routing elements outside the LAN or attached to the LAN when using IP multicast. Note that no configuration is required. [0068]
(Recovery of defects) The essence of the protocol described above is that it is designed for automatic failure recovery, thus creating a very high level of service availability. This system is robust against both ARN and SN defects. [0069]
ARN "times out" entries in the SN database, so SN glitches are not used in service requests. So when the client reconnects to the service (transparently to the user or by interacting with the user), the service starts again on another service node. If the ARN maintains a persistent state for that client, the system will not charge the user twice if a SN failure is detected and a switch occurs (for example). ) Instead of starting a new service from scratch, you can resume the incarnation of an old service. [0070]
If the ARN fails, another problem can occur. By keeping redundant ARNs in a single colo, this problem can be solved using the following techniques: [0071]
-Each ARN is a well-known multicast group G<sub>s</sub>Join; Each ARN sends an advertisement message to Group G<sub>s</sub>Advertise your existence by sending to. [0072]
Each ARN builds a database of active ARN peers and times out entries that have not been updated according to a specific configurable period that is longer than the intermediate advertisement period. [0073]
An ARN with the least numbered network addresses (ie, an ARN with fewer network addresses than all remaining ARNs in the database) elects itself as the master ARN and reaches the anycast address block. Initiate the action of advertising the possibility via IGP. [0074]
Therefore, if the master ARN fails, the backup ARN will immediately detect this situation (after a single advertisement interval has elapsed) and a new master ARN will be elected. At this point, as a side effect of the new IGP route advertisement, the colo router routes the anycast packet to the new master ARN. [0075]
(Wide area architecture) Having described the local-area architecture of devices within a single colo installation range, we will discuss how to coordinate and manage individual service-node clusters within a wide area architecture. There are two main wide area components that implement the content overlay network embodiments included in the present invention. That is, -A data service that routes and manages a large amount of data from the content site from which it was created to the service node in colo. [0076]
-A solution to the problem of waiting for services using anycast. In this solution, the service requested by the client is combined with the master site from which it was created. [0077]
A solution to the former problem-that is, a method of distributing data over a wide area to service nodes without loss of reliability and efficiency is outside the scope of this disclosure. For example, it is described in the pending US Patent Application No. 60 / 115,454 entitled "System for Providing Application-level Features to Multicast Routing in Computer Networks" filed on January 22, 1999. Content can be transported using a streaming broadcast network). It is also possible to transport the content by a file distribution protocol based on a flood algorithm such as Network News Transfer Protocol (NNTP). [0078]
The present application discloses a manner in which an anycast transfer system interfaces with an available content delivery system. A new framework is used for service-specific interactions between ARNs, SNs, clients, and preferably the creator's service or content sites. For example, a client can initiate a web request to anycast address A, which is routed to the most recent ARN advertising reachability to address A, and then its recent. The leading ARN may forward the client to the selected SN with a simple HTTP forwarding message, or its web request may be served directly from the ARN. [0079]
In one embodiment of an anycast transfer system, the ARN "primes" the SN with application-specific information that cannot be transported from an existing unmodified client to the SN. be able to. Here, the ARN contacts the SN and attaches a particular state Q coupled to a particular port P. Port P can be assigned by the SN and returned to the ARN. It is then possible to forward the client to the SN over port P, which allows the unmodified client to implicitly carry state Q over this new connection. For example, Q can represent a wide range of broadcast channel addresses that a service node should subscribe to for a particular streaming media feed. Since the unmodified client does not have direct protocol conformance to the CBB infrastructure, the appropriate channel subscription is carried by state transfer Q without the need to use the client in the dialog. [0080] [0080]
The interaction between the client and the SN is replaced by existing service-specific protocols (eg, for example) so that large clients already installed (eg, web browsers and streaming media players and servers) do not need to be modified. Use HTTP for the web, RTSP for streaming media protocols, or other vendor-owned protocols. [0081]
After the client request is initiated and intercepted by the ARN, a particular wide area service must be called to pull the content from the CBB and put it on the local service node (provided that the content no longer exists). As mentioned above, repeated use of anycast causes problems. Therefore, the DNS system is used to remap the anycast address to the service in a scalable and distributed fashion. [0082]
For example, assuming you want to support caching of web objects for the content provider "acme.com" and the CBB assigns acme.com an anycast address of 10.11.8.27, a pointer to the master server. Can be configured as DNS with SRV resource records such as: [0083]
anycast-10-1-18-24.http.tcp.cbb.net SRV www.acme.com If a service node receives a client connection request for its anycast address 10.11.8.27 on TCP port 80 (ie, a standard HTTP web port), the service node will receive (anycast-10-1-18-27). Query the DNSSRV record for .http.tcp.cbb.net) to find out that the master host for this service is www.acme.com. This knowledge can be cached locally, and if the requested content is fetched from www.acme.com, that content can also be cached locally. Then, the next time there is a request for the same anycast address for the same content, the request can be fulfilled locally. Content stored on www.acme.com may have a link that explicitly references anycast-10-1-18-27.http.tcp.cbb.net, or the site may look like this: Note that you can also use a more user-friendly name (for example, www-cbb.acme.com, which is just the CNAME of the anycast name): www-cbb.acme.com CNAME anycast 10-1-18-27.http.tcp.cbb.net Consider another example where it is desirable to support very large streaming media broadcasts, also coming from acme.com. In addition to the master web server, you will also need to know the location of the "Channel Allocation Service" (CAS), which maps streaming media URLs to broadcast channel addresses. In this "channel allocation service", the channel address is similar to an application level multicast group as described in 60/115454. In this case, query DNS for CAS-oriented SRV resource records to obtain records that may have the following forms: anycast-10-1-18-24.cas.tcp.cbb.net SRV cas.acme.com When ARN receives a client connection request for a streaming media URL, ARN queries cas.acme.com to map that URL to a broadcast channel (and locally caches the results for future client requests). Now that the channel address is known, join the channel via the CBB. [0084]
By storing the service binding in DNS in this way, any anycast service node can dynamically and automatically discover a particular service bound to a particular anycast address. This greatly simplifies the configuration and management of general anycast-based service rendezvous mechanisms and content broadcast networks. [0085]
In this way, the DNS SRV record stores the mapping from the service name to the corresponding server address. However, ARN may require more information than a simple list of named services. Further information may identify the service node selection algorithm or the service node setup procedure. In these cases, information about the named service can be stored in a directory server (such as LDAP or X.500) or on the web server's network. Compared to DNS, these servers offer greater flexibility and extensibility in data representation. [0086]
(Redirect algorithm) Given the above components and system architecture, one embodiment of the invention provides how an end host calls a service flow or transaction from a service-node infrastructure using stateful anycasting. Is specified. [0087]
-The user initiates a content request, for example, by clicking on a web link represented as a URL. [0088]
-The client determines the DNS name of the resource referenced by the URL. This name ultimately determines the anycast address managed by the authorities (for example, www.acme.com is the CNAME for any-10-1-18.27.cbb.net). [0089]
The client initiates a successful application connection using that anycast address (for example, a web page request using HTTP over TCP on port 80, or streaming media using RTSP over TCP on port 554. request). [0090]
As a side effect of the anycast routing infrastructure described above, client packets are routed to the nearest ARN that advertises reachability to that address, which initiates a connection to that ARN. The ARN is prepared to accept requests for each configured service, such as web requests on port 80. [0091]
· At this point, if the data is available and should be transactional, the ARN will either respond directly with the content or redirect the requesting client to the service node as follows: ARN selects candidate service node S from its associated service cluster. Selection decisions can be made based on load and availability information retained from the local monitoring protocol as described above. [0092]
-ARN uses S to perform application-specific dialogs as necessary in preparation for client C to attach to S. For example, in the case of live broadcast streaming media, the ARN can indicate a broadcast channel. S needs to tune to the CBB overlay network over the request based on its broadcast channel. As part of this dialog, S can return the information needed to properly redirect C to S back to ARN. The existence of this information and the nature of that information is specified for the particular service requested. [0093]
-ARN responds to the original client request with a redirect message pointing client C to the selected service node S above. [0094]
Client C contacts S in a client-specific manner and initiates a flow or content transaction related to the desired service. For example, the client uses the streaming media control protocol RTSP to connect to S and initiate a live transfer of streaming media over RTP. [0095]
(Active session failover) One drawback of the Stateful Anycasting Redirection scheme above is that if the selected service node fails for any reason, all clients given that node will suffer a service interruption. If the client is calling a persistent service such as a streaming media feed, the video will stop otherwise and the client will be forced to retry. In another embodiment, the client may be modified to detect the failure of the service node and recall the redirection process before the user notices the service degradation. This process is referred to herein as "active session failover." [0096]
FIG. 6 shows a portion of a data network 800 constructed according to the present invention. Data Network 800 shows a network transaction that demonstrates how active session failover works to deliver content to clients uninterrupted. [0097]
First, client 802 sends a service request 820 to anycast address A. Service request 820 is routed to ARN804. Service request 820 requests content originating from CBB803. ARN804 decrypts the request to determine the application-specific redirect message to be sent to client 802. The redirect message 822 sent by ARN804 redirects client 802 to service node 806. The client 802 then sends a request 824 to obtain content (eg, streaming media feed, etc.) via an application-specific protocol (eg, RTSP). With this application-specific protocol, node 806 requests a streaming media channel over a wide area by sending a channel join message 826 to service node 808 using the channel description information in the client request (see, eg, 60/115454). ). As a result, the content flows from service node 808 to the client as indicated by path 828. [0098]
Now suppose service node 806 fails. Client 802 notices a disruption in service and responds by recalling the stateful anycast procedure described in the previous section: Service request 830 is sent to anycast address A and received by ARN804. ARN804 responds with redirect message 832, directing the client to the new service node 810. Here, the client may request a new service feed from service node 810, as shown in 834. Service node 810 sends a join message to node 808 as shown in 836, and the content flows back to the client as shown in 838. Assuming that the client utilizes sufficient buffering before presenting the streaming media signal to the user (a commonly used method to handle network delay variability), this entire process proceeds uninterrupted. Can be done. When the client attaches, the client sends a packet resend request to service node 810 to properly position the stream and retransmit only packets lost during the session failover process. [0099]
For example, a client could mistakenly infer service node 806's failure due to a temporary network outage. In this case, the client can simply ignore the redirect message 832 and continue to receive service from service node 806. [0100]
(Wide area overflow) One possible problem with the service rendezvous mechanism above is that a configuration on a service node can run out of capacity because too many clients are routed to that configuration. This problem can be solved in embodiments where the redirect system can redirect client service requests over a wide area in the event of an overload. For example, if all of the local service nodes are running at full capacity, the redirector may select a non-local service node and redirect clients to it. This redirect decision can then be influenced by network and server health measurements. In this approach, the redirector sends a period "probe" message to the candidate server to measure network route delay. Since redirectors are usually closer to the requesting client, these redirector-server measurements represent an accurate estimate of the corresponding network path between the client and the candidate server. [0101]
In this embodiment, there are three steps for performing a wide area redirection: -ARN discovers a candidate service node. [0102]
ARN measures network path characteristics between each service and itself. [0103]
ARN queries service nodes for their health. [0104]
Once informed by the above steps, ARN can select a service node that may provide the highest quality service for any requesting client. To do so, each ARN maintains an information database containing load information about some eligible service nodes. ARN examines its information database to determine the most accessible service node for each client request. The ARN can proactively probe network routes and service nodes to retain its load information. Alternatively, the service node can monitor network and internal loads and report load information to their respective ARNs. [0105]
To enable local area load balancing, each ARN is configured with the IP addresses of several neighboring service nodes. The ARN holds load information about those service nodes. However, this local area approach is exacerbated when the roads are geographically concentrated. This is because the ARN can fully load all of its neighboring service nodes, thereby rejecting further service requests from its clients. This can happen even if some service nodes just beyond the local area are underutilized. [0106]
Wide area load balancing according to the present invention overcomes the above problems. In wide area load balancing, each ARN is configured with the IP addresses of all service nodes in the network and maintains an information database containing load information for all service nodes. Alternatively, ARN uses a flooding algorithm to exchange load information. [0107]
Another embodiment of the invention uses a scheme called variable region load balancing. Using this scheme, each ARN maintains an information database for several eligible service nodes; and the number of eligible service nodes increases with local load. That is, as neighboring service nodes approach their respective capacities, ARN adds load information to its information database for some service nodes that are just beyond the current range of ARN. Below are two different methods that can be used to discover progressively distal service nodes. [0108]
In the first method, the ARN is given the IP addresses of several adjacent service nodes. Since it is assumed that the service nodes form a virtual overlay network, the ARN simply queries these service nodes for a list of nearby service nodes in order to identify the service nodes that are moving apart in the increasing direction. This approach may be called "overlay network crawling". [0109]
In the second technique, each ARN and each service node is assigned a multipart name from the hierarchical namespace. Given the names of two service nodes, ARN can use the longest pattern matching to determine which is closest. For example, using the longest pattern matching from right to left, arn.sanjose.california.pacificcoast.usa.northamerica The ARN name given the name is sn.orlando.florida.atlanticcoast.usa.northamerica Than the service node named sn.seattle.washington.pacificcoast.usa.northamerica It is determined that it is close to the service node named. Each ARN can search directories for all service node names and their corresponding IP addresses. This directory can be run using DNS or similar distributed directory technology. This variable domain load balancing scheme handles geographically concentrated loads by redirecting clients to service nodes that are distant in the increasing direction. This scheme solves the scalability problem by minimizing the number of ARN-to-service node relationships. That is, the ARN only monitors the number of service nodes required to handle nearby client loads. In general, the number of nodes at a distance of N hops from a given node is multiplied by N, so the speed at which ARN scrutinizes candidate service nodes is adjusted to be inversely proportional to the distance. [0110]
FIG. 7 shows a part of the data network 900 configured based on the present invention. The data network 900 includes three connected local networks 902, 904, and 906. As described in one embodiment of the invention, the network 900 is configured to provide wide area overflow. [0111]
Local network 902 includes an ARN908 (redirector) with an associated information database (DB) 910. Also, service nodes 912 and 914 are included in the local network 902. Indicates a service node that provides information content 928 to clients (C) 916, 918, 920, 922, 924, and 926. [0112]
Networks 904 and 906 include ARN930, 932, information databases 934, 936, and service nodes 938, 940, 942, and 944, respectively. These service nodes provide information content 928 to a large number of other clients. [0113]
The ARN908 monitors the network loading characteristics of its local service nodes 912 and 914. This loading information is stored in DB910. The ARN can also monitor the loading characteristics of other service nodes. In certain embodiments, the ARNs exchange loading information with each other. For example, the loading characteristics of service nodes 938 and 940 are monitored by ARN930 and stored in DB934. The ARN930 can exchange this loading information with the ARN908, as shown by arrow 954. In another embodiment, the ARN may actively scrutinize other service nodes to determine their loading characteristics. These properties are then used for future use. For example, ARN908 scrutinizes service node 944 as shown by arrow 956 and scrutinizes service node 942 as shown by arrow 958. Therefore, there are several ways in which the ARN can determine the loading characteristics of service nodes located both within the local network and over a wide area. [0114]
At some point, client 950 attempts to receive information content 928. Client 950 sends an anycast request 952 to network 902, where the anycast request 952 is received by ARN908. The ARN908 may redirect the client 950 to one of the local service nodes (912,914). However, DB910 associated with ARN908 indicates that the local service node may not be able to provide the requested service to client 950. The ARN908 uses the information DB910 to determine which service node is most appropriate for processing the request from the client 950. The selected service node is not limited to the service node in the local network having ARN908. Any service node on the wide area can be selected. [0115]
ARN908 determines that service node 942 needs to service the request from client 950. ARN908 sends a redirect message 960 to client 950, thereby redirecting the client to service node 942. Client 950 sends a request to service node 942 using a transfer layer protocol such as TCP, as shown by arrow 962. Service node 942 responds by providing the requested information content to the client, as indicated by arrow 964. [0116]
Thus, using the ability to scrutinize information databases and service nodes to obtain loading characteristics, reference nodes can achieve wide area load balancing under the present invention. [0117]
(Technical expansion) This section describes further embodiments of the invention, including technical extensions to the above embodiments. [0118]
(Last Hop Multicast) The use of IP multicast can be used as a transfer optimization in the "last hop" delivery of wide area delivery content. Therefore, the client can issue an anycast request and, as a result, be redirected to be included in the multicast group. [0119]
FIG. 8 shows an embodiment of the present invention using IP multicast. The content provider 600 provides three service nodes SN0 to SN3 to provide information content 602 via the application level multicast tree 604. Client 606 requests a service feed as described above and is received by the ARN as shown in passage 608. The ARN redirects the request to service node SN0 and initiates a data transfer as shown in passage 610. However, instead of initiating an independent data channel for each client, the service node joins a particular multicast group 612 (called Group G) and tells the client (via a control connection) to receive informational content. Command. The client then joins the multicast group and service node SN0 sends the information content to the group in the local environment. [0120]
As shown, multicast traffic is replicated only at fanout points in the delivery path from service node SN0 to all clients receiving its flow. At the same time, service node SN0 contacts upstream service node SN1 to receive information content on the unicast connection. In this way, the content delivers the content to all interested recipients without the need to enable multicast over the entire network infrastructure. [0121]
(Sender attachment) The system described above relied on anycast routing to route client requests to the nearest service node. Similarly, anycast can be used to bridge the server at the content-generated location to the nearest service entry point. If the content server is clearly configured into a delivery infrastructure, the systems described herein may be adapted to register and connect service installations to the wide area delivery infrastructure. [0122]
FIG. 9 shows Embodiment 700 of the invention adapted to register and connect a service installation to a wide area distribution infrastructure. Service node SN0 wants to plug a new wide area distribution channel from nearby server S into the content wide area distribution network 704. The service node SN0 sends a service query 706 using the anycast address. The service query requests the identity of the service node within the wide area distribution network 704, which is most available to be used as an endpoint for a new IP tunnel from SN0. The service query carries the anycast address and is routed to the nearest ARN (A1 in this case). A1 may select the most available service node and update the channel database in the wide area distribution network 704, indicating that new channels are available via SN0. In this case, A1 selects SN1 and sends response 708 to SN0. This response tells SN0 to build a new IP tunneling circuit 702 and service node SN1. [0123]
This sender attachment system allows the overlay wide area distribution network to be dynamically scaled to reach additional servers. This sender attachment system, when used in conjunction with the client attachment system described above, dynamically maps client-server traffic onto one or more continuous tunneling circuits, with one tunneling endpoint being the closest client. It provides a comprehensive architecture where one tunneling endpoint is the closest server. The mapping is performed in a way that is obvious to the client and server applications and to each access route. [0124]
(Multiple masters) In the various embodiments described herein, only the master AS advertises anycast address blocks via interdomain routing protocols. There are two extensions to this scheme. First, the system can be extended to allow multiple master ASs to coexist by partitioning the anycast address space between multiple master ASs. That is, the plurality of examples of systems described herein work well and do not interfere with each other as long as they use separate address spaces for their anycast blocks. Second, the system can be extended to allow multiple master ASs to advertise the same or partially overlapping address blocks. In this case, the minimum distance anycast routing works at both intradomain and interdomain routing levels. For example, there can be a North American master AS and a European master AS that advertise the same anycast block to the outside world via BGP. Packets sent by any client are then sent to the closest master AS, and once inside the AS, the packets are routed to the nearest service node within that AS, if necessary. Or redirected over a wide area. [0125]
The present invention provides a comprehensive redirection system for content distribution based on a virtual overlay wide area distribution network. It will be apparent to those skilled in the art that the methods and embodiments can be modified or combined without departing from the scope of the invention. Therefore, the disclosure contents and descriptions in the present specification are intended to illustrate the scope of the present invention described in the following claims, and are not intended to limit the scope of the present invention.
[Simple explanation of drawings]
[Figure 1]
Figure 1 shows an example of typical components and interconnects that form part of the connectivity of the Internet. [Figure 2]
Figure 2 shows a typical overlay ISP model. [Fig. 3]
FIG. 3 shows a network portion 300 used to implement the anycast routing scheme according to the present invention. [Fig. 4]
Figure 4 shows the master network and associated networks configured to achieve interdomain anycast routing. [Fig. 5]
FIG. 5 shows how the control and service functions included within the present invention are separated within a particular ISP. [Fig. 6]
FIG. 6 shows a portion of the data network constructed by the present invention to perform failover of an active session. [Fig. 7]
FIG. 7 shows a portion of the data network constructed by the present invention to implement a wide area overflow. [Fig. 8]
FIG. 8 shows the use of IP multicast according to the present invention. [Fig. 9]
FIG. 9 shows an embodiment of the invention that is registered and adapted to connect service equipment to a service broadcast network infrastructure.
Every citation, both ways
| Document | Relation | Office | Cited during |
|---|---|---|---|
| US10009315B2 | Cited by | United States of America | Applicant |
| JP2015501095A | Cited by | Japan | Search report |
| KR20140119090A | Cited by | Republic of Korea | Search report |
| JP2015506523A | Cited by | Japan | Search report |
| US10860384B2 | Cited by | United States of America | Applicant |
| JP2014512739A | Cited by | Japan | Search report |
| JP2013528336A | Cited by | Japan | Search report |
| US10635500B2 | Cited by | United States of America | Applicant |
29 members in 7 offices
Priority claims14
| Document | Office | Kind | Date |
|---|---|---|---|
| 15225799 | United States of America | P | |
| 15225799 | United States of America | P | |
| 60152257 | United States of America | – | |
| 09458216 | United States of America | – | |
| 45821699 | United States of America | A | |
| 45821699 | United States of America | A | |
| 0024047 | United States of America | W | |
| 0024047 | United States of America | W | |
| 1999152257 | – | – | – |
| 1999458216 | – | – | – |
| 200024047 | – | – | – |
| US19990152257P | – | – | – |
| US19990458216 | – | – | – |
| WO2000US24047 | – | – | – |
Members29
| Document | Office | Kind | |
|---|---|---|---|
| WO0118641A1 | World Intellectual Property Organization (WIPO) | A1 | |
| AU7341500A | Australia | A | |
| WO0152497A2 | World Intellectual Property Organization (WIPO) | A2 | |
| AU2282701A | Australia | A | |
| WO0152497A3 | World Intellectual Property Organization (WIPO) | A3 | |
| WO0152497A9 | World Intellectual Property Organization (WIPO) | A9 | |
| KR20020048399A | Republic of Korea | A | |
| US6415323B1 | United States of America | B1 | |
| EP1242870A1 | European Patent Office (EPO) | A1 | |
| EP1250785A2 | European Patent Office (EPO) | A2 | |
| JP2003508996AThis record | Japan | A | |
| US2003105865A1 | United States of America | A1 | |
| AU771353B2 | Australia | B2 | |
| US6785704B1 | United States of America | B1 | |
| US2005010653A1 | United States of America | A1 | |
| EP1242870A4 | European Patent Office (EPO) | A4 | |
| US6901445B2 | United States of America | B2 | |
| KR100524258B1 | Republic of Korea | B1 | |
| JP3807981B2 | Japan | B2 | |
| EP1250785B1 | European Patent Office (EPO) | B1 | |
| DE60036021D1 | Germany | D1 | |
| EP1865684A1 | European Patent Office (EPO) | A1 | |
| DE60036021T2 | Germany | T2 | |
| US7734730B2 | United States of America | B2 | |
| EP2320619A1 | European Patent Office (EPO) | A1 | |
| EP1865684B1 | European Patent Office (EPO) | B1 | |
| EP2320619B1 | European Patent Office (EPO) | B1 | |
| EP2838240A1 | European Patent Office (EPO) | A1 | |
| EP2838240B1 | European Patent Office (EPO) | B1 |
31 legal events, as the office reported them to INPADOC
Over the term
Point at a mark for the eventEvents
| Event | Code | |
|---|---|---|
| Cancellation because of completion of termEXPY | EXPY | |
| Receipt of annual feesJAPANESE INTERMEDIATE CODE: R250R250 | R250 | |
| Receipt of annual feesJAPANESE INTERMEDIATE CODE: R250R250 | R250 | |
| Receipt of annual feesJAPANESE INTERMEDIATE CODE: R250R250 | R250 | |
| Receipt of annual feesJAPANESE INTERMEDIATE CODE: R250R250 | R250 | |
| Receipt of annual feesJAPANESE INTERMEDIATE CODE: R250R250 | R250 | |
| Receipt of annual feesJAPANESE INTERMEDIATE CODE: R250R250 | R250 | |
| Receipt of annual feesJAPANESE INTERMEDIATE CODE: R250R250 | R250 | |
| Receipt of annual feesJAPANESE INTERMEDIATE CODE: R250R250 | R250 | |
| Renewal fee payment (event date is renewal date of database)FPAY | FPAY | |
| Renewal fee payment (event date is renewal date of database)FPAY | FPAY | |
| Receipt of annual feesJAPANESE INTERMEDIATE CODE: R250R250 | R250 | |
| Renewal fee payment (event date is renewal date of database)FPAY | FPAY | |
| Receipt of annual feesJAPANESE INTERMEDIATE CODE: R250R250 | R250 | |
| Renewal fee payment (event date is renewal date of database)FPAY | FPAY | |
| Renewal fee payment (event date is renewal date of database)FPAY | FPAY | |
| Receipt of annual feesJAPANESE INTERMEDIATE CODE: R250R250 | R250 | |
| Renewal fee payment (event date is renewal date of database)FPAY | FPAY | |
| Receipt of annual feesJAPANESE INTERMEDIATE CODE: R250R250 | R250 | |
| Renewal fee payment (event date is renewal date of database)FPAY | FPAY | |
| Written notification of registration of transferJAPANESE INTERMEDIATE CODE: R350R350 | R350 | |
| Written notification of registration of transferJAPANESE INTERMEDIATE CODE: R350R350 | R350 | |
| Renewal fee payment (event date is renewal date of database)FPAY | FPAY | |
| Request for change of ownership or part of ownershipJAPANESE INTERMEDIATE CODE: R313113S111 | S111 | |
| Written request for registration of change of domicileJAPANESE INTERMEDIATE CODE: R313531S531 | S531 | |
| Certificate of patent or registration of utility modelJAPANESE INTERMEDIATE CODE: R150R150 | R150 | |
| Certificate of patent or registration of utility modelJAPANESE INTERMEDIATE CODE: R150R150 | R150 | |
| First payment of annual fees (during grant procedure)JAPANESE INTERMEDIATE CODE: A61A61 | A61 | |
| Written decision to grant a patent or to grant a registration (utility model)JAPANESE INTERMEDIATE CODE: A01A01 | A01 | |
| Decision of grant or rejection writtenTRDD | TRDD | |
| Report on retrievalJAPANESE INTERMEDIATE CODE: A971007A977 | A977 |
Numbers
- Publication
- 2003-508996
- Publication, DOCDB
- 2003508996
- Publication, EPODOC
- JP2003508996
- Application
- 2001522165
- Application, DOCDB
- 2001522165
- Application, EPODOC
- JP20010522165
Titles2
- Japanese
- 【発明の名称】インターネットワークにおけるローバストで拡大縮小可能なサービスノードロケーションの近接ベースのリダイレクトシステム
- English
- INDUSTRIAL APPLICABILITY A proximity-based redirection system for service node locations that is robust and scaleable in an internetwork.
Classification
- CPC, 10
- H04L12/18
- G06F15/173
- H04L45/306
- H04L67/1008
- H04L67/1029
- H04L67/101
- H04L69/329
- H04L67/1001
- H04L45/22
- H04L9/40
- IPC, 3
- H04L12 56
- H04L29 06
- H04L29 08