Chassis controllers for converting universal flows
Abstract
A network control system that generates physical control plane data for managing first and second managed forwarding elements that implement forwarding operations associated with a first logical data path set. The system has a first controller instance that transforms the logical control plane data for the first logical data pathset into universal physical control plane (UPCP) data. The system further has a second controller instance that transforms UPCP data into customized physical control plane (CPCP) data for the first managed forwarding element but not for the second managed forwarding element. The system receives the UPCP data generated by the first controller instance, identifies the second controller instance as the controller instance involved in the generation of CPCP data for the first managed forwarding element, and uses the received UPCP data as the second controller instance. It further has a third controller instance to supply to the controller instance.

Term
6.1 yearsto projected expiry
Projected expiry 25 October 2032, counted from filing; an application has no term until it is granted.
- Priority
- Filed
- Published
- Today
- Projected expiry
19 claims: 4 independent, 15 dependent
- 1第1論理データパスセットと関連付けられたフォワーディング動作を実現する第1及び第2管理対象フォワーディングエレメントを管理する物理コントロールプレーンデータを生成するネットワーク制御システムであって、 a)前記第1論理データパスセット用の論理コントロールプレーンデータを前記第1及び前記第2管理対象フォワーディングエレメントの属性の汎用表現に関して書き込まれたデータタプルを含むユニバーサル物理コントロールプレーン(UPCP)データに変換する第1コントローラインスタンスと、 b)前記第1管理対象フォワーディングエレメントに特有の前記属性の表現に関して前記UPCPデータの前記データタプルをカスタマイズすることにより、前記UPCPデータを前記第1管理対象フォワーディングエレメント用であるが前記第2管理対象フォワーディングエレメント用ではないカスタマイズド物理コントロールプレーン(CPCP)データに変換する第2コントローラインスタンスと、 c)前記第1コントローラインスタンスにより生成された前記UPCPデータを受信し、前記第1管理対象フォワーディングエレメント用の前記CPCPデータの生成に関与するコントローラインスタンスとして前記第2コントローラインスタンスを識別し、前記受信したUPCPデータを前記第2コントローラインスタンスに供給する第3コントローラインスタンスと、 を有することを特徴とするネットワーク制御システム。
- 2a)前記第2管理対象フォワーディングエレメントに特有の前記属性の表現に関してUPCPデータの前記データタプルをカスタマイズすることにより、前記UPCPデータを前記第2管理対象フォワーディングエレメント用のCPCPデータに変換する第4コントローラインスタンスと、 b)前記第1コントローラインスタンスにより生成された前記UPCPデータを受信し、前記第2管理対象フォワーディングエレメント用の前記CPCPデータの生成に関与するコントローラインスタンスとして前記第4コントローラインスタンスを識別し、前記受信したUPCPデータを前記第2コントローラインスタンスに供給する第5のコントローラインスタンスと、 を更に有することを特徴とする請求項1に記載のネットワーク制御システム。
- 3前記第1コントローラインスタンスは、前記第1論理データパスセットに対するマスタコントローラインスタンスであり、 前記第3コントローラインスタンスは、前記第1管理対象フォワーディングエレメントに対するマスタコントローラインスタンスであり、 前記第5のコントローラインスタンスは、前記第2管理対象フォワーディングエレメントに対するマスタコントローラインスタンスである ことを特徴とする請求項2に記載のネットワーク制御システム。
- 4前記第1及び第2管理対象フォワーディングエレメントは、それぞれ第1及び第2装置上で動作するソフトウェアスイッチングエレメントであり、 前記第2コントローラインスタンスは、前記第1装置上で動作するコントローラインスタンスであり、 前記第4コントローラインスタンスは、前記第2装置上で動作するコントローラインスタンスである ことを特徴とする請求項3に記載のネットワーク制御システム。
- 5異なる管理対象フォワーディングエレメントのマスタとして異なるコントローラインスタンスを識別する調整マネージャを更に有することを特徴とする請求項3に記載のネットワーク制御システム。
- 6異なる論理データパスセット及び異なる管理対象フォワーディングエレメントのマスタとして異なるコントローラインスタンスを識別する調整マネージャを更に有することを特徴とする請求項3に記載のネットワーク制御システム。
- 7前記第3コントローラインスタンスは、更に、第2論理データパスセット用の論理コントロールプレーンデータを前記第2論理データパスセット用のUPCPデータに変換するためのものであり、 前記第1コントローラインスタンスは、更に、前記第2論理データパスセット用のUPCPデータを第3管理対象フォワーディングエレメント用のCPCPデータに変換するためのものであることを特徴とする請求項2に記載のネットワーク制御システム。
- 8前記第1コントローラインスタンスは、 論理コントロールプレーンデータを論理フォワーディングプレーンデータに変換する制御モジュールと、 論理フォワーディングプレーンデータをUPCPデータに変換する仮想化モジュールと を有することを特徴とする請求項1に記載のネットワーク制御システム。
- 9テーブルマッピングエンジンと、 論理コントロールプレーンレコード及びフォワーディングプレーンレコードを格納する入力テーブル及び出力テーブルと、 テーブルマッピング規則の集合と、 テーブルマッピング規則の第1部分集合及び前記テーブルマッピングエンジンを含む前記制御モジュールと、 テーブルマッピング規則の第2部分集合及び前記テーブルマッピングエンジンを含む前記仮想化モジュールと を更に有することを特徴とする請求項8に記載のネットワーク制御システム。
- 10第1論理スイッチングエレメントと関連付けられたフォワーディング動作を実現する第1管理対象フォワーディングエレメント及び第2管理対象フォワーディングエレメントを管理するネットワーク制御システムに対する第1コントローラインスタンスであって、 前記第1論理スイッチングエレメント用の論理コントロールプレーンデータからユニバーサル物理コントロールプレーン(UPCP)データを生成した第2コントローラインスタンスから前記UPCPデータを受信するコントローラ間通信インタフェースであって、前記UPCPデータが、前記第1及び第2管理対象フォワーディングエレメントの属性の汎用表現に関して前記第1及び第2管理対象フォワーディングエレメントのフォワーディング挙動を規定するコントローラ間通信インタフェースと、 前記第1管理対象フォワーディングエレメント用であるが前記第2管理対象フォワーディングエレメント用ではないカスタマイズド物理コントロールプレーン(CPCP)データを前記受信したUPCPデータから生成することに関与するコントローラインスタンスとして第3コントローラインスタンスを識別する第1モジュールであって、前記第1管理対象フォワーディングエレメント用の前記CPCPデータが、前記第1管理対象フォワーディングエレメントに特有の前記属性の表現に関して前記第1管理対象フォワーディングエレメントのフォワーディング挙動を規定する第1モジュールと、 前記受信したUPCPデータを前記第3コントローラインスタンスに更に供給する前記コントローラ間通信インタフェースと、 を有することを特徴とする第1コントローラインスタンス。
- 11前記ネットワーク制御システムは、UPCPデータを前記第2管理対象フォワーディングエレメント用のCPCPデータに変換する第4コントローラインスタンスと、前記第2管理対象フォワーディングエレメント用のCPCPデータの生成に関与するコントローラインスタンスとして前記第4コントローラインスタンスを識別する第5のコントローラインスタンスとを有し、 前記第1コントローラインスタンスは、前記第1管理対象フォワーディングエレメントに対するマスタコントローラインスタンスであり、 前記第2コントローラインスタンスは、前記第1論理スイッチングエレメントに対するマスタコントローラインスタンスであり、 前記第5のコントローラインスタンスは、前記第2管理対象フォワーディングエレメントに対するマスタコントローラインスタンスである ことを特徴とする請求項10に記載の第1コントローラインスタンス。
- 12異なる管理対象フォワーディングエレメントのマスタとして異なるコントローラインスタンスを識別する調整マネージャを更に有し、前記第1コントローラインスタンスの前記調整マネージャは、少なくとも1つの他のコントローラインスタンスの調整マネージャと対話して、異なる管理対象フォワーディングエレメントのマスタとして異なるコントローラインスタンスを識別することを特徴とする請求項11に記載の第1コントローラインスタンス。
- 13第2論理スイッチングエレメント用の論理コントロールプレーンデータを前記第2論理スイッチングエレメント用のUPCPデータに変換する第2モジュールを更に有することを特徴とする請求項10に記載の第1コントローラインスタンス。
- 14前記コントローラ間通信インタフェースは、更に、第3管理対象フォワーディングエレメント用のCPCPデータに変換するために、前記第2論理スイッチングエレメント用のUPCPデータを第4コントローラインスタンスに送出するためのものであることを特徴とする請求項13に記載の第1コントローラインスタンス。
- 15前記第2論理スイッチングエレメント用の論理コントロールプレーンデータを論理フォワーディングプレーンデータに変換する制御モジュールと、 論理フォワーディングプレーンデータをユニバーサル物理コントロールプレーンデータに変換する仮想化モジュールと を更に有することを特徴とする請求項13に記載の第1コントローラインスタンス。
- 16第1論理スイッチングエレメントと関連付けられたフォワーディング動作を実現する第1及び第2管理対象フォワーディングエレメントを管理するネットワーク制御システムに対する第1コントローラインスタンスであって、 第3コントローラインスタンスにより前記第1論理スイッチングエレメント用の論理コントロールプレーンデータから生成され、前記第1及び第2管理対象フォワーディングエレメントの属性の汎用表現に関して前記第1及び第2管理対象フォワーディングエレメントのフォワーディング挙動を規定するユニバーサル物理コントロールプレーン(UPCP)データを第2コントローラインスタンスから受信するコントローラ間通信インタフェースと、 前記第1管理対象フォワーディングエレメントに特有の前記属性の表現に関して前記第1管理対象フォワーディングエレメントの前記フォワーディング挙動を規定する前記第1管理対象フォワーディングエレメント用であるが前記第2管理対象フォワーディングエレメント用ではないカスタマイズド物理コントロールプレーン(CPCP)データを前記受信したUPCPデータから生成するモジュールと、 前記生成したCPCPデータを前記第1管理対象フォワーディングエレメントに供給するフォワーディングエレメント通信インタフェースと、 を有することを特徴とする第1コントローラインスタンス。
- 17前記ネットワーク制御システムは、UPCPデータから前記第2管理対象フォワーディングエレメント用のCPCPデータを生成する第4コントローラインスタンスと、前記第3コントローラインスタンスから前記UPCPデータを受信し、前記第2管理対象フォワーディングエレメント用のCPCPデータの生成に関与するコントローラインスタンスとして前記第4コントローラインスタンスを識別し、前記UPCPデータを前記第4コントローラインスタンスに供給する第5のコントローラインスタンスとを有し、 前記第2コントローラインスタンスは、前記第1管理対象フォワーディングエレメントに対するマスタコントローラインスタンスであり、 前記第3コントローラインスタンスは、前記第1論理スイッチングエレメントに対するマスタコントローラインスタンスであり、 前記第5のコントローラインスタンスは、前記第2管理対象フォワーディングエレメントに対するマスタコントローラインスタンスである ことを特徴とする請求項16に記載の第1コントローラインスタンス。
- 18第1論理データパスセットと関連付けられたフォワーディング動作を実現する第1及び第2管理対象フォワーディングエレメントを管理するネットワーク制御システムの第1コントローラインスタンスのコンピュータ読み取り可能な記憶媒体であって、 (i)第3コントローラインスタンスにより前記第1論理スイッチングエレメント用の論理コントロールプレーンデータから生成され、(ii)前記第1及び第2管理対象フォワーディングエレメントの属性の表現に関して書き込まれたデータタプルを含むユニバーサル物理コントロールプレーン(UPCP)データを第2コントローラインスタンスから受信し、 前記第1管理対象フォワーディングエレメントに特有の前記属性の表現に関して前記UPCPデータの前記データタプルをカスタマイズすることにより、前記第1管理対象フォワーディングエレメント用であるが前記第2管理対象フォワーディングエレメント用ではないカスタマイズド物理コントロールプレーン(CPCP)データを、前記受信したUPCPデータから生成し、 前記生成されたCPCPデータを前記第1管理対象フォワーディングエレメントに供給する ための命令セットを格納したコンピュータ読み取り可能な記憶媒体。
- 19前記ネットワーク制御システムは、UPCPデータから前記第2管理対象フォワーディングエレメント用のCPCPデータを生成する第4コントローラインスタンスと、前記第3コントローラインスタンスから前記UPCPデータを受信し、前記第2管理対象フォワーディングエレメント用のCPCPデータの生成に関与するコントローラインスタンスとして前記第4コントローラインスタンスを識別し、前記UPCPデータを前記第4コントローラインスタンスに供給する第5のコントローラインスタンスとを有し、 前記第2コントローラインスタンスは、前記第1管理対象フォワーディングエレメントに対するマスタコントローラインスタンスであり、 前記第3コントローラインスタンスは、前記第1論理スイッチングエレメントに対するマスタコントローラインスタンスであり、 前記第5のコントローラインスタンスは、前記第2管理対象フォワーディングエレメントに対するマスタコントローラインスタンスである ことを特徴とする請求項18に記載のコンピュータ読み取り可能な記憶媒体。
Independent claims19
254 paragraphs, as filed
The present invention relates to a chassis controller for converting universal flow.
Many of today's enterprises have large, sophisticated networks that include switches, hubs, routers, servers, workstations and other network equipment that support a variety of connections, applications and systems. Further sophistication of computer networking, including virtual machine movement, dynamic workload, multi-tenancy, and customer-specific quality of service and security configurations, requires a more appropriate paradigm for network control. Networks have traditionally been managed with low-level settings for individual components. In many cases, the network configuration depends on the underlying network. For example, to block a user's access using an access control list ("ACL") entry, you need to know the user's current IP address. More complex tasks require broader network knowledge. To force a guest user's port 80 traffic to traverse the HTTP proxy, it is necessary to know the current network topology and location of each guest. This process becomes increasingly difficult when the network switching element is shared among a large number of users.
Correspondingly, there is a growing movement towards a new network control paradigm called Software Defined Networking (SDN). In the SDN paradigm, a network controller running on one or more servers in a network controls, maintains and implements control logic that manages the forwarding behavior of shared network switching elements for each user. In many cases, it is necessary to be aware of the network state in order to make network management decisions. To facilitate management decisions, the network controller creates and maintains a view of the network state, on which the management application provides an application programming interface that allows access to the view of the network state.
Some of the main objectives of maintaining a large network (including both data centers and corporate networks) are scalability, mobility and multi-tenancy. Many methods taken to address one of these objectives result in interfering with at least one of the other methods. For example, one method can easily provide network mobility to virtual machines within the L2 domain, but cannot extend the L2 domain. Also, the continued isolation of users makes mobility very complex. Therefore, there is a need for improved solutions that meet the goals of scalability, mobility and multi-tenancy.
Some embodiments of the present invention allow several different logics for different users by one or more shared forwarding elements without even allowing several different users to control or even view each other's transfer logics. A network control system is provided that allows a data path (LDP) set to be specified. Since these shared forwarding elements are managed by the network control system to realize the LDP set, they are hereinafter referred to as managed switching elements or managed forwarding elements.
In some embodiments, a network control system is one or more controllers (hereinafter also referred to as controller instances) that allow the system to accept LDP sets from users and configure switching elements to implement these LDP sets. ). With these controllers, the system is defined by the connections between the shared switching elements in such a way that different users share the same switching element while preventing them from browsing or controlling each other's LDP sets and logical networks. The control of these shared switching elements and logical networks can be virtualized.
In some embodiments, each controller instance translates user input from logical control plane (LCP) data to logical forwarding plane (LFP) data and then LFP data to physical control plane (PCP) data. A device (eg, a general purpose computer) that executes one or more modules. These modules in some embodiments include control modules and virtualization modules. The control module allows the user to specify and populate the logical data path set (LDPS), and the virtualization module implements the specified LDPS by mapping the LDPS to the physical switching infrastructure. The control module and the virtualization module are two separate applications in some embodiments, but are part of the same application in other embodiments.
In some of the embodiments, the controller control module receives LCP data describing the LDPS (eg, data describing the connection associated with the logical switching element) from the user or another source. The control module then converts this data into LFP data that will later be supplied to the virtualization module. Next, the virtualization module generates PCP data from the LFP data. The PCP data is propagated to the managed switching element. In some embodiments, the control module and the virtualization module use the nLog engine to generate LFP data from LCP data and PCP data from LFP data.
Some embodiments of network control systems use different controllers to perform different tasks. For example, in some embodiments, the network control system uses three types of controllers. The first type of controller is an application protocol interface (API) controller. The API controller receives the configuration data and the user query from the user by calling the API, and is involved in the response of the user query. The API controller further disseminates the received configuration data to other controllers. Therefore, the API controller of some embodiments is an interface between the user and the network control system.
The second type of controller is a logical controller involved in the realization of the LDP set by calculating the universal flow entry, which is a general representation of the flow entry for the managed switching element that realizes the LDP set. The logical controller in some embodiments does not interact directly with the managed switching element, but pushes the universal flow entry to the physical controller, which is a third controller type.
Physical controllers in different embodiments have different mechanisms. In some embodiments, the physical controller generates customized flow entries from universal flow entries and pushes these customized flow entries down to the managed switching element. In another embodiment, the physical controller identifies, for a particular managed physical switching element, a chassis controller, which is the type of fourth controller involved in generating customized flow entries for a particular switching element, from the logical controller. Forward the incoming universal flow entry to the chassis controller. The chassis controller then generates customized flow entries from the universal flow entries and pushes these customized flow entries to the managed switching element. In yet another embodiment, the physical controller generates such flow entries for some managed switching elements while instructing the chassis controller to generate customized flow entries for other managed switching elements. ..
The above-mentioned outline of the invention is a brief outline for some embodiments of the present invention and is not intended to be an outline of all the subjects of the present invention disclosed herein. The following brief description of the embodiments for carrying out the invention and the drawings referenced in the embodiments for carrying out the invention further describes the outline of the invention and the embodiments described in other embodiments. Therefore, in order to understand all the embodiments described herein, it is necessary to consider the whole of the outline of the invention, the embodiments for carrying out the invention and the brief description of the drawings. Also, since the claimed subject matter can be implemented in other specific forms without departing from the gist of the subject matter, it is limited by the detailed description in the outline of the invention, the form for carrying out the invention and the brief description of the drawings. It is not specified by the scope of the attached claims.
The novel features of the invention are defined in the appended claims, but for purposes of explanation, some embodiments of the invention will be described with reference to the following drawings.<figref num="1">The figure which shows the virtualized network system in embodiment of this invention.</figref><figref num="2">Diagram showing the switch infrastructure of a multi-user server hosting system.</figref><figref num="3">The figure which shows the network controller which manages an edge switching element.</figref><figref num="4">The figure which shows an example of a plurality of logical switching elements realized over a set of switching elements.</figref><figref num="5">The figure which shows the propagation of the instruction which controls a switching element managed through various processing layers of a controller instance.</figref><figref num="6">The figure which shows the multi-instance, distributed network control system in embodiment.</figref><figref num="7">The figure which shows an example which specifies a master controller instance for a switching element.</figref><figref num="8">The figure which shows the example of the operation of a controller instance.</figref><figref num="9">The figure which conceptually shows the software architecture for an input conversion application.</figref><figref num="10">The figure which shows the control application in embodiment of this invention.</figref><figref num="11">The figure which shows the virtualization application in embodiment of this invention.</figref><figref num="12">The figure which conceptually shows the different tables in the RE output table.</figref><figref num="13">The simplified figure which shows the table mapping operation of the control application and the virtualization application in embodiment of this invention.</figref><figref num="14">The figure which shows an example of the integrated application.</figref><figref num="15">The figure which shows another example of the integrated application.</figref><figref num="16">The figure which conceptually shows an example of the architecture of a network control system.</figref><figref num="17">The figure which conceptually shows an example of the architecture of a network control system.</figref><figref num="18">The figure which shows an example of the architecture for a chassis control application.</figref><figref num="19A">、</figref><figref num="19B">FIG. 5 shows an example of creating a tunnel between two managed switching elements based on universal physical control plane data.</figref><figref num="20">The figure which conceptually shows the process to perform to generate the customized physical control plane data from the universal physical control plane data in embodiment.</figref><figref num="21">The figure which conceptually shows the process which generates the customized tunnel flow instruction and executes to send the customized instruction to the managed switching element in embodiment.</figref><figref num="22A">、</figref><figref num="22B">The figure which conceptually shows an example of the operation of the chassis controller which converts a universal tunnel flow instruction into a customized instruction in seven different stages.</figref><figref num="23">The figure which conceptually shows the electronic system used to realize the embodiment of this invention.</figref>
Hereinafter, various details, examples and embodiments of the present invention will be described in the detailed description of the present invention. However, it will be apparent to those skilled in the art that the present invention is not limited to the embodiments described below and may be practiced without the use of the following specific details and examples.
Some embodiments of the present invention allow for different users by one or more shared forwarding elements without even allowing several different users to control or even view each other's forwarding logic. A network control system is provided that allows several different sets of LDPs to be specified. Shared forwarding elements in some embodiments include virtual or physical network switches, software switches (eg, Open vSwitch), routers and / or other switching devices, and these switches and routers. And and / or any other network element (eg, load balancer, etc.) that establishes a connection with and / or other switching equipment. Such forwarding elements (eg, physical switches or routers) are referred to below as switching elements. Also called elements). In contrast to off-the-shelf switches, a software forwarding element is, in some embodiments, a switching element formed by storing its switching table and switching logic in the memory of a stand-alone device (eg, a stand-alone computer). However, in other embodiments, it is formed by storing the switching table and switching logic in the memory of a hypervisor and a device (eg, a computer) that further executes one or more virtual machines in addition to the hypervisor. It is a switching element.
Since these management / shared switching elements are managed by the network control system to realize the LDP set, they are hereinafter referred to as managed switching elements or managed forwarding elements. In some embodiments, the control system manages them by pushing PCP data to these switching elements, as described further below. In general, a switching element receives data (eg, a data packet), drops the received data packet, for example, passes a packet received from one source device to another destination device, processes the packet to the destination device. Performs one or more processing operations on the data, such as passing to. In some embodiments, the PCP data pushed to the switching element is a data packet received by the switching element (eg, by the switching element's general purpose processor) and by the switching element (eg, the switching element's dedicated switching circuit). Is converted to physical forwarding plane data that specifies how to handle.
In some embodiments, a network control system is one or more controllers (hereinafter also referred to as controller instances) that allow the system to accept LDP sets from users and configure switching elements to implement these LDP sets. ) having a that. With these controllers, the system is by connecting between shared switching elements in such a way that different users share the same managed switching element while preventing them from browsing or controlling each other's LDP sets and logical networks. Control of these defined shared switching elements and logical networks can be virtualized.
In some embodiments, each controller instance is a device (eg, a general purpose computer) that executes one or more modules that convert user input from LCP to LFP and then convert LFP data to PCP data. .. These modules in some embodiments include control modules and virtualization modules. The control module allows the user to specify and populate the LDPS, and the virtualization module implements the specified LDPS by mapping the LDPS to the physical switching infrastructure. In some embodiments, the control module and virtualization module represent data specified or mapped with respect to records written to a relational database data structure. That is, the relational database data structure stores both the logical data path input received via the control module and the physical data to which the logical data path input is mapped by the virtualization module. In some embodiments, the control application and the virtualization application are two separate applications in some embodiments, but are part of the same application in other embodiments.
So far, some examples of network control systems have been described. Hereinafter, a more detailed embodiment will be described. Chapter I describes some embodiments of network control systems. The following Chapter II describes the conversion of the universal forwarding state by the network control system. Chapter III describes the electronic systems used to implement some embodiments of the present invention.
I. Network control system A. Outer layer to push flow to control layer FIG. 1 shows a virtualized network system 100 of some embodiments of the present invention. This system allows a large number of users to create and control a large number of different sets of LDPs on a shared set of network infrastructure switching elements (eg, switches, virtual switches, software switches, etc.). When allowing a user to create and control a set of a user's logical data path (LDP) set (ie, a user's switching logic), the system allows the user to view or modify the switching logic of another user. Do not allow direct access to another user's set of LDP sets. However, the system does not allow such communication to users if different users want to pass packets to each other through virtualization switching logic.
As shown in FIG. 1, the system 100 includes one or more switching elements 105 and a network controller 110. The switching element has N switching devices (N is a number of 1 or more) forming the network infrastructure switching element of the system 100. In some embodiments, network infrastructure switching elements include virtual network switches or physical network switches, software switches (eg, Open vSwitch), routers and / or other switching devices, and these switches and routers. And / or any other network element (eg, load balancer, etc.) that establishes a connection with other switching devices. All such network infrastructure switching elements are referred to below as switching elements or forwarding elements.
The virtual switching device or physical switching device 105 generally includes a control switching logic 125 and a forwarding switching logic 130. In some embodiments, the control switching logic 125 specifies (1) the rules applied to the input packets, (2) the packets to be discarded, and (3) the packet processing methods applied to the input packets. The virtual switching element or the physical switching element 105 uses control logic 125 to populate the table that manages the transfer logic 130. The forwarding logic 130 executes a lookup operation on the input packet and forwards the input packet to the destination address.
As further shown in FIG. 1, the network controller 110 includes a control application 115. By control application 115, switching logic is specified for one or more users (eg, by one or more administrators or users) with respect to the LDP set. The network controller 110 further includes a virtualization application 120 that translates the LDP set into a control switching logic pushed to the switching device 105. In this application, the control application and the virtualization application are referred to as a "control engine" and a "virtualization engine" in some embodiments.
In some embodiments, the virtualization system 100 has two or more network controllers 110. Each network controller has a logic controller involved in designating control logic for a set of switching devices for a particular LDPS. The network controller further comprises a physical controller, and each physical controller pushes control logic to a set of switching elements that it is responsible for managing. In other words, the logical controller specifies the control logic only for the set of switching elements that realize a specific LDPS, and the physical controller controls the switching elements that it manages regardless of the LDP set realized by the switching elements. To push.
In some embodiments, the network controller virtualized application uses a relational database data structure to store a copy of the state of the switch element that is tracked by the virtualized application with respect to a data record (eg, a data tuple). These data records represent graphs of all physical or virtual switching elements and their interconnects within the physical network topology, as well as their forwarding tables. For example, in some embodiments, each switching element in the network infrastructure is represented by one or more data records in a relational database data structure. However, in other embodiments, the relational database data structure for the virtualized application stores state information about only some of the switching elements. For example, as described in more detail below, virtualization applications in some embodiments only track switching elements at the edge of the network infrastructure. In yet another embodiment, the virtualization application stores state information about the edge switching elements in the network and some non-edge switching elements in the network that facilitate communication between the edge switching elements.
In some embodiments, the relational database data structure is at the core of the control model in the virtualized network system 100. Under one method, an application controls a network by reading from a relational database data structure and writing to a relational database data structure. Specifically, in some embodiments, application control logic (1) reads the current state associated with network entity records in a relational database data structure and (2) acts on these records to network. The state can be changed. Under this model, when virtualization application 120 needs to modify records in a table of switching element 105 (eg, a control plane flow table), it first puts one or more records that represent the table into a relational database data structure. Write to. The virtualization application then propagates this change to the table of switching elements.
In some embodiments, the control application further uses a relational database data structure to store the logical configuration and logical state for each user-specified LDPS. In these embodiments, the information in the relational database data structure that represents the state of the actual switching element describes only a subset of all the information stored in the relational database data structure.
In some embodiments, the control and virtualization applications use a secondary data structure to store the logical configuration and logical state for a user-specified LDPS. This secondary data structure in these embodiments serves as a communication medium between different network controllers. For example, if a user specifies a particular LDPS using a logical controller that is not involved in a particular LDPS, then the logical controller is another that is involved in the particular LDPS through the secondary data structure of these logical controllers. Pass the logical configuration for a specific LDPS to the logical controller. In some embodiments, the logical controller that receives the logical configuration for a particular LDPS from the user passes the configuration data to all other controllers in the virtualized network system. As such, the secondary storage structure in all logical controllers includes logical configuration data for all LDP sets for all users in some embodiments.
In some embodiments, the controller instance operating system (not shown) provides different embodiments of control and virtualization applications, as well as different sets of communication structures (not shown) for the switching element 105. For example, in some embodiments, the operating system is used to (1) perform physical switching for any one user, and (2) push switching logic for the user to the switching element. A communication interface (not shown) with the virtualized application 120 to be managed is provided to the managed switching element. In some of these embodiments, a virtualized application is commonly known to specify a set of APIs that allow an external application (eg, a virtualized application) to control the control plane functionality of a switching element. The control switching logic 125 of the switching element is managed through the switch access interface. Specifically, the managed switching element communication interface implements a collection of APIs that allows a virtualized application to send records stored in a relational database data structure to the switching element using the managed switching element communication interface. ..
Two examples of such known switch access interfaces are the OpenFlow interface and the open virtual switch communication interface. They can be searched from the following two documents, McKeown, N. (2008), OpenFlow: Enabling Innovation in Campus Networks (http://www.openflowswitch.org//documents/openflow-wp-latest.pdf). (Yes) and Pettit, J. (2010), Virtual Switching in an Era of Advanced Edges (searchable from http://openvswitch.org/papers/dccaves2010.pdf), respectively. These two documents are incorporated herein by reference.
Note that in the case of these embodiments described above and below where a relational database data structure is used to store data records, a data structure capable of storing object-oriented data object format data can be used instead or connectedly. .. An example of such a data structure is the NIB data structure. Some examples of using NIB data structures are described in US Patent Application 13 / 177,529 and US Patent Application 13 / 177,533, both filed July 6, 2011. U.S. Patent Application No. 13 / 177,529 and U.S. Patent Application No. 13 / 177,533 are incorporated herein by reference.
FIG. 1 conceptually illustrates the use of the switch access API by drawing an enclosure 135 around the control switching logic 125. Through these APIs, virtualization applications can read and write entries in the control plane flow table. The connectivity of virtualized applications to the control plane resources of the switching element (eg, the control plane table) is achieved in-band (ie, with network traffic controlled by the operating system) in some embodiments. However, in other embodiments it is implemented out of band (ie, via an independent physical network). Standard IGP protocols such as IS-IS or OSPF are sufficient when using independent networks, requiring minimal requirements for choices that avoid concentration of failures and basic connectivity to the operating system. is there.
In order to specify a control switching logic 125 for a switching element when the switching element is a physical switching element (as opposed to a software switching element), some virtualization applications have a control of the switching element. Use the open virtual switch protocol to create one or more control tables in the plane. Generally, the control plane is created and executed by a general-purpose CPU of a switching element. When the system creates the control table, the virtualization application uses the OpenFlow protocol to write flow entries to the control table. The general purpose CPU of a physical switching element uses internal logic to convert an entry written to populate one or more forwarding tables in the forwarding plane of the switching element into a control table. The forwarding table is generally created and executed by a dedicated switching chip of the switching element. By performing a flow entry in the forwarding table, the switching chip of the switching element can process and route packets of data it receives.
In some embodiments, the virtualization network system 100 is a chassis controller (chassis) in addition to a logical controller and a physical controller. It has a controller). In these embodiments, the chassis controller implements a switch access API to manage a particular switching element. That is, it is the chassis controller that pushes the control logic to a particular switching element. The physical controller in these embodiments functions as an aggregation point for relaying from the logical controller to the chassis controller interface-connected to the set of switching elements in which the physical controller is involved. The physical controller distributes control logic to the chassis controller that manages the set of switching elements. In these embodiments, the network controller operating system has a communication channel (eg, for example) between the physical controller and the chassis controller so that the physical controller can send control logic stored in the relational database data structure as data records to the chassis controller. A managed switching element communication interface that establishes a remote procedure call (RPC) channel. As a result, the chassis controller pushes control logic to the switching element using the switch access API or other protocol.
Communication structures provided by some embodiments of the operating system transfer data records to another network controller (eg, from a logical controller to another logical controller, from a physical controller to another physical controller, from a logical controller to a physical controller. It also has an exporter (not shown) that can be used by the network controller to send from the physical controller to the logical controller, etc.). Specifically, network controller control and virtualization applications can use exporters to export data records stored in relational database data structures to one or more other network controllers. In some embodiments, the exporter establishes a communication channel (eg, an RPC channel) between the two network controllers so that one network controller can send data records over the channel to another.
The operating system of some embodiments further comprises an importer that can be used by the network controller to receive data records from the network controller. The importer of some embodiments acts as the equivalent of an exporter of another network controller. That is, the importer is on the receiving side of the communication channel established between the two network controllers. In some embodiments, the network controller follows a publish-subscribe model in which the receiving controller subscribes to the channel to receive data only from the network controller that supplies the data of interest.
B. Pushing the flow to the edge switching element As mentioned above, relational database data structures store data for each switching element in the system's network infrastructure in some embodiments, but in other embodiments for switching elements at the edge of the network infrastructure. Stores only status information. 2 and 3 show an example of distinguishing between the two different methods. Specifically, FIG. 2 shows the switch infrastructure of a multi-user server hosting system. In this system, six switching elements are employed to interconnect the six machines of two users A and B. Four of these switching elements, 205-220, are edge switching elements that directly connect the machines 235-260 of users A and B, and two of the switching elements, 225 and 230, connect the edge switching elements to each other. Internal switching elements that connect and connect to each other (ie, non-edge switching elements). All of the illustrated switching elements described above and below may be software switching elements in some embodiments, but in other embodiments a mixture of software switching elements and physical switching elements. For example, edge switching elements 205-220, and non-edge switching elements 225 and 230 are software switching elements in some embodiments. Further, the "machine" described in this application includes a virtual machine such as an arithmetic unit and a physical machine.
FIG. 3 shows a network controller 300 that manages edge switching elements 205-220. The network controller 300 is similar to the network controller 110 described above with reference to FIG. As shown in FIG. 3, the controller 300 includes a control application 305 and a virtualization application 310. The operating system for controller instance 300 maintains a relational database data structure (not shown) containing data records for edge switching elements 205-220 only. Applications 305 and 310 running on the operating system also allow users A and B to change the switching element configuration for the edge switching element they use. The network controller 300 then propagates these changes to the edge switching element as needed. Specifically, in this example, the two edge switching elements 205 and 220 are used by both users A and B machines, while the edge switching element 210 is used only by user A machines 245, edge switching. Element 215 is used only by User B's machine 250. Therefore, FIG. 3 shows a network controller 300 that changes the records of users A and B at the switching elements 205 and 220, but updates only the records of user A at the switching element 210 and the records of user B at the switching element 215.
The controller 300 of some embodiments controls only the edge switching element (ie, maintains only the data in the relational database data structure with respect to the edge switching element) for several reasons. Sufficient to maintain the required separation between machines (eg, arithmetic units), as opposed to maintaining the separation between all unnecessary switching elements by controlling the edge switching elements. Provide means to the controller. The internal switching element transfers data packets between the switching elements. The edge switching element forwards data packets between the machine and other network elements (eg, other switching elements). Thus, since the edge switching element is the last switching element in a line to forward the packet to the machine, the controller can maintain user isolation simply by controlling the edge switching element.
In addition to controlling edge switching elements, network controllers in some embodiments further include non-edge switching elements that are inserted into the switch network hierarchy to simplify and / or facilitate the operation of controlled edge switching elements. Use and control. For example, in some embodiments, the controllers interact with each other in a hierarchical switching architecture in which the switching elements they control have several edge switching elements that are leaf nodes and one or more non-edge switching elements that are non-leaf nodes. Request to be connected. In some such embodiments, each of the edge switching elements is connected to one or more non-leaf switching elements and such non-leaf switching elements to facilitate communication with other edge switching elements. To use.
The above description relates to the control of edge switching elements and non-edge switching elements by network controllers of some embodiments. In some embodiments, edge switching elements and non-edge switching elements (leaf nodes and non-leaf nodes) may be referred to as managed switching elements. This is because these switching elements are managed by the network controller (as opposed to unmanaged switching elements in a network that is not managed by the network controller) in order for the managed switching elements to implement the LDP set. ..
The network controller of some embodiments implements a logical switching element across managed switching elements based on the physical and logical data described above. A logical switching element (also referred to as a "logical forwarding element") may be defined to function in one or more different ways in which the switching element may function (eg, layer 2 switching, layer 3 routing, etc.). The network controller realizes the specified logical switching element by controlling the managed switching element. In some embodiments, the network controller implements a large number of logical switching elements across managed switching elements. This allows a number of different logical switching elements to be implemented across managed switching elements regardless of the network topology of the network.
The managed switching elements of some embodiments may be configured to route network data based on different routing criteria. In this way, the flow of network data through the switching elements in the network can be controlled to implement a large number of logical switching elements across the managed switching elements.
C. Logical switching element and physical switching element FIG. 4 shows an example of a large number of logical switching elements realized over a set of switching elements. In particular, FIG. 4 conceptually shows the logical switching elements 480 and 490 implemented across the managed switching elements 410-430. As shown in FIG. 4, the network 400 includes managed switching elements 410-430 and machines 440-465. As shown in FIG. 4, machines 440, 450 and 460 belong to user A, and machines 445, 455 and 465 belong to user B.
The managed switching elements 410 to 430 of some embodiments route network data (eg, packets, frames, etc.) between network elements in the network connected to the managed switching elements 410 to 430. As shown, the managed switching element 410 routes network data between the machines 440 and 445 and the switching element 420. Similarly, network element 420 routes network data between the machine 450 and managed switching elements 410 and 430, and switching element 430 routes network data between machines 455-465 and switching element 420. ..
Also, each of the managed switching elements 410-430 routes network data, which in some embodiments is in the form of a table, based on the transfer logic of the switch. In some embodiments, the forwarding table determines where network data should be routed (eg, a port on the switch) according to routing criteria. For example, the forwarding table of the layer 2 switching element may determine where network data should be routed based on the MAC address (eg, source MAC address and / or destination MAC address). As another example, the forwarding table of the Layer 3 switching element may determine where network data should be routed based on the IP address (eg, source IP address and / or destination IP address). Many other types of routing criteria are possible.
As shown in FIG. 4, the forwarding table in each of the managed switching elements 410-430 contains several records. In some embodiments, each record specifies an action for routing network data based on routing criteria. Records may be referred to as flow entries in some embodiments because they control the "flow" of data passing through managed switching elements 410-430.
FIG. 4 further shows a conceptual representation of each user's logical network. As shown, user A's logical network 480 includes a logical switching element 485 to which user A's machines 440, 450 and 460 are connected. User B's logical network 490 includes a logical switching element 495 to which User B's machines 445, 455 and 465 are connected. Therefore, from the viewpoint of user A, user A has a switching element to which only the machine of user A is connected, and from the viewpoint of user B, user B has a switching element to which only the machine of user B is connected. .. In other words, for each user, the user has his own network with only the user's machines.
The conceptual flow entry that realizes the flow of the network data to the machine 450 transmitted from the machine 440 and the network data to the machine 460 transmitted from the machine 440 will be described below. The flow entries "A1 to A2" in the forwarding table of the managed switching element 410 instruct the managed switching element 410 to route the network data destined for the machine 450 originating from the machine 410 to the switching element 420. The flow entries "A1 to A2" in the forwarding table of the switching element 420 instruct the switching element 420 to route the network data destined for the machine 450 originating from the machine 410 to the machine 450. Thus, when machine 440 sends network data destined for machine 450, managed switching elements 410 and 420 route network data along the data path 470 based on the corresponding records in the switching element's forwarding table.
Further, the flow entries "A1 to A3" in the forwarding table of the managed switching element 410 instruct the managed switching element 410 to route the network data addressed to the machine 460 transmitted from the machine 440 to the switching element 420. The flow entries "A1 to A3" in the forwarding table of the switching element 420 instruct the switching element 420 to route the network data destined for the machine 460 originating from the machine 440 to the switching element 430. The flow entries "A1 to A3" in the forwarding table of the switching element 430 instruct the switching element 430 to route the network data destined for the machine 460 originating from the machine 440 to the machine 460. Thus, when machine 440 sends network data destined for machine 460, managed switching elements 410-430 route network data along data paths 470 and 475 based on the corresponding records in the switching element's forwarding table. ..
Although the conceptual flow entry for routing network data originating from machine 440 to machine 450 and network data originating from machine 440 to machine 460 has been described above, a similar flow entry is in user A's logical network 480. It will be included in the forwarding table of managed switching elements 410-430 that route network data between other machines. Further similar flow entries will be included in the forwarding table of managed switching elements 410-430 that route network data between machines in User B's logical network 490.
The conceptual flow entry shown in FIG. 4 contains both source and destination information about the managed switching element to estimate the next hop switching element to which the packet should be sent. However, the source information does not need to be in the flow entry because the managed switching element of some embodiments can infer the next hop switching element using only the destination information (eg, context identifier, destination address, etc.). ..
In some embodiments, tunneling protocols (eg, control and provisioning of radio access points (CAPWAP)) (eg, control and provisioning of) to facilitate the implementation of logical switching elements 485 and 495 across managed switching elements 410-430. wireless access point), GRE (generic routing encapsulation), GRE Internet Protocol Security (IPsec) Etc.) may be used. With tunneling, a packet is sent through the switch and router as the payload of another packet. That is, the tunnel packet needs to expose its address (eg, source MAC address and destination MAC address) because the packet is forwarded based on the address contained in the header of the other packet that encapsulates the tunnel packet. There is no. Therefore, tunnel packets can have meaningful addresses in the logical address space, but other packets are forwarded / routed based on the addresses in the physical address space, thus separating the logical address space from the physical address space by tunneling. it can. As such, the tunnel may be viewed as a "logical wire" connecting the managed switching elements in the network to implement the logical switching elements 485 and 495.
By configuring the switching elements in the various ways described above to implement a large number of logical switching elements across a set of switching elements, a large number of users can enjoy an independent network and / or switching element from each user's point of view. In practice, some or all of the same set of switching elements and / or connections are shared between sets of switching elements (eg, tunnels, physical wires).
II. Universal forwarding state A. Controller instance layer FIG. 5 shows the propagation of instructions that control a managed switching element through various processing layers of a controller instance of some embodiments of the present invention. FIG. 5 shows a control data pipeline 500 that transforms and propagates control plane data to managed switching element 525 through four processing layers of the same controller instance or different controller instances. These four layers are an input conversion layer 505, a control layer 510, a virtualization layer 515, and a customization layer 520.
In some embodiments, these four layers are in the same controller instance. However, other configurations of these layers exist in other embodiments. For example, in other embodiments, only the control layer 510 and the virtualization layer 515 are in the same controller instance, but the functionality of propagating customized physical control plane (CPCP) data is not shown in another controller instance (eg, not shown). It is in the customization layer of the chassis controller). In these other embodiments, the universal physical control plane (UPCP) data is the relational database data structure of one controller instance (not shown) before another controller instance generates CPCP data and pushes it to a managed switching element. ) Transfers to the relational database data structures of other controller instances. The former controller instance may be a logical controller that generates UPCP data, and the latter controller instance may be a physical controller or chassis controller that customizes UPCP data into CPCP data.
As shown in FIG. 5, the input conversion layer 505 in some embodiments has an LCP 530 that can be used to represent the output of this layer. In some embodiments, an application (eg, a web-based application (not shown) is provided to the user to provide input for the user to specify an LDP set. This application sends an input in the form of an API call to the input conversion layer 505. The input conversion layer 505 converts the API call into LCP data in a format that can be controlled by the control layer 510. For example, the input is transformed into a set of input events that can be supplied to the control layer's nLog table mapping engine. The nLog table mapping engine and its operation will be further described below.
The control layer 510 in some embodiments has LCP530 and LFP535 that can be used to represent inputs and outputs to this layer. The LCP includes a control layer and a collection of high-level structures that allow its users to specify one or more LDPs within the LCP for one or more users. LFP535 represents a user's LDP set in a format that can be processed by virtualization layer 515. Thus, the two logic planes 530 and 535 are virtualized spatial analogs of the control plane 555 and the forwarding plane 560, which are generally found in the general managed switching element 525, as shown in FIG.
In some embodiments, the control layer 510 defines and exposes the LCP structure. The layer itself or the user of the layer defines a different set of LDPs within the LCP using an LCP structure. For example, in some embodiments, the LCP data 530 includes logical ACL data and the like. Some of this data (eg, logical ACL data) can be specified by the user, while other such data (eg, logical L2 or L3 records) are generated by the control layer and specified by the user. It does not have to be. In some embodiments, control layer 510 responds to certain changes (indicating changes to managed switching elements and management data paths) to relational database data structures detected by control layer 510, such data. And / or specify.
In some embodiments, the LCP data (ie, the LDP set data represented with respect to the control plane structure) takes into account the current operating data from the managed switching element, and this control plane data is the PCP data. Can be specified first without considering how it is converted to. For example, this control plane data may later be transformed into physical control data for three managed switching elements that achieve the desired switching between the five computers, while the LCP data connects the five computers. Control data for one logical switching element may be specified.
The control layer contains a set of modules (not shown) for converting any LDPS in the LCP into LDPS in the LFP535. In some embodiments, control layer 510 uses an nLog table mapping engine to perform this transformation. It will be further described below that the control layer uses the nLog table mapping engine to perform this transformation. The control layer further includes a set of modules (not shown) for pushing the LDP set from the LFP 535 of the control layer 510 to the LFP 540 of the virtualization layer 515.
LFP540 includes one or more LDP sets of one or more users. The LFP540 in some embodiments comprises logical forwarding data for one or more LDP sets of one or more users. Some of this data is pushed to the LFP540 by the control layer, while other such data are in relational database data structures, as further described below for some embodiments. It is pushed to the LFP by the virtualization layer that detects the event.
In addition to LFP540, the virtualization layer 515 includes UPCP545. UPCP545 includes UPCP data for the LDP set. The virtualization layer contains a set of modules (not shown) for converting the LDP set in the LFP 540 into UPCP data in the UPCP 545. In some embodiments, the virtualization layer 515 uses an nLog table mapping engine to perform this transformation. The virtualization layer further includes a set of modules (not shown) for pushing UPCP data from the UPCP 545 of the virtualization layer 515 to the relational database data structure of the customization layer 520.
In some embodiments, the UPCP data sent to the customization layer 515 allows the managed switching element 525 to process the data packet according to the LDP set specified by the control layer 510. However, in contrast to the CPCP data, the UPCP data does not represent the location-specific information of the managed switching element and / or the managed switching element in some embodiments, and thus is the logical data specified by the control layer. Is not completely realized.
The UPCP data must be converted into CPCP data for each of the managed switching elements in order to fully implement the LDP set in the managed switching elements. For example, when an LDP set specifies a tunnel that spans several managed switching elements, the UPCP data uses the specific network address (eg, IP address) of the managed switching element that represents one end of the tunnel. To express. However, each of the other managed switching elements across the tunnel uses a port number local to the managed switching element to refer to the terminal managed switching element with a particular network address. That is, a particular network address must be translated into a local port number for each of the managed switching elements in order to fully implement the LDP set that specifies the tunnel in the managed switching element.
Assuming that the customization layer 520 is operating in a different controller instance than the controller instance that produces the UPCP data, the UPCP data, which is the intermediate data that is converted to the CPCP data, provides the control system of some embodiments. You will be able to change the magnification. This is because the virtualization layer 515 does not need to convert the LFP data that specifies the LDP set into CPCP data for each of the managed switching elements that realize the LDP set. Instead, the virtualization layer 515 converts the LFP data into UPCP data only once for all managed switching elements that implement the LDP set. In this way, the virtualization application saves computational resources that would be spent converting the LDP set to CPCP data as many times as the number of managed switching elements that implement the LDP set.
Customization layer 520 includes UPCP546 and CPCP550 that can be used to represent inputs and outputs to this layer. The customization layer contains a set of modules (not shown) for converting UPCP data in UPCP 546 to CPCP data in CPCP 550. In some embodiments, the customization layer 520 uses the nLog table mapping engine to perform this transformation. The customization layer further includes a set of modules (not shown) for pushing CPCP data from the CPCP 550 of the customization layer 520 to the managed switching element 525.
The CPCP data pushed to each of the managed switching elements is specific to the managed switching element. CPCP data, referred to as "physical" data, allows managed switching elements to perform physical switching operations in both the physical data processing area and the logical data processing area. In some embodiments, the customization layer 520 operates in an independent controller instance for each of the managed switching elements.
In some embodiments, customization layer 520 does not operate on the controller instance. The customization layer 515 in these embodiments is on the controlled switching element 525. Therefore, in these embodiments, the virtualization layer 515 sends the UPCP data to the managed switching element. Each of the managed switching elements customizes the UPCP data to CPCP data specific to the managed switching element. In some of these embodiments, the controller daemon is running on each of the managed switching elements and transforms the universal data into customized data for the managed switching elements. The controller daemon will be further described below.
In some embodiments, the customized physical control plane data propagated to the managed switching element 525 allows the switching element to with respect to network data (eg, packets) based on the logical values defined in the logical domain. You will be able to perform physical forwarding operations. Specifically, in some embodiments, the customized physical control plane data specifies a flow entry that includes a logical value. These logical values include a logical address, a logical port number, and the like used for transferring network data in the logical area. These flow entries bring the logical value to the specified physical value in the physical domain so that the managed switching element can perform the logical forwarding operation on the network data by performing the physical forwarding operation based on the logical value. Further mapping. In this way, physical control plane data facilitates the implementation of logical switching elements across managed switching elements. Some examples of using propagated physical control plane data to achieve logical data processing in controlled switching elements are in US Patent Application No. 13 / 177,535, filed July 6, 2011. Further described. U.S. Patent Application No. 13 / 177,535 is incorporated herein by reference.
The control plane data processed by the layers of the control data pipeline 500 becomes more global as the layers get higher. That is, the logical control plane data in the control layer 510 spans the entire set of managed switching elements that realize the logical switching elements defined by the logical control plane data. In contrast, the customized physical control plane data at customization layer 520 is local and specific to each of the managed switching elements that implement the logical switching elements.
B. Multi-controller instance FIG. 6 shows a multi-instance, distributed network control system 600 of some embodiments. This distributed system uses three controller instances 605, 610 and 615 to control a large number of switching elements 690. In some embodiments, the distributed system 600 allows different controller instances to control the operation of the same switching element or different switching elements. As shown in FIG. 6, each instance has an input module 620, a control module 625, a record 635, a secondary storage structure (eg, PTD) 640, an inter-controller communication interface 645, and managed switching element communication. Includes interface 650 and.
The input module 620 of the controller instance is similar to the input conversion layer 505 described above with reference to FIG. 5 in that it takes input from the user and converts the input into LCP data that the control module 625 understands and processes. To do. As mentioned above, the input is in the form of an API call in some embodiments. The input module 620 sends LCP data to the control module 625.
The control module 625 of the controller instance is similar to the control layer 510 in that it converts the LCP data into LFP data and pushes the LFP data to the virtualization module 630. Further, the control module 625 determines whether the received LCP data belongs to the LDPS managed by the controller instance. If the controller instance is the master of LDPS for the LCP data (ie, the logical controller that manages the LDPS), the virtualization module of the controller instance further processes the data. If the controller instance is not the master of LDPS for LCP data, the control module 625 of some embodiments stores the LCP data in secondary storage 640.
The virtualization module 630 of the controller instance is similar to the virtualization layer 515 in that it converts LFP data into UPCP data. The virtualization module 630 of some embodiments sends UPCP data to another controller instance via the inter-controller communication interface 645 or to a managed switching element via the managed switching element communication interface 650.
If the other controller instance is a physical controller involved in managing at least one of the managed switching elements that implement LDPS, the virtualization module 630 sends UPPC data to another instance. This is the case when the controller instance in which the virtualization module 630 is generating UPCP data is merely a logical controller involved in a particular LDPS, but not a physical controller or chassis controller involved in a managed switching element that implements LDPS. Is.
When the managed switching element is configured to convert the UPCP data into CPCP data specific to the managed switching element, the virtualization module 630 sends the UPCP data to the managed switching element. In this case, the controller instance does not have a customization layer or module that performs the conversion from UPCP data to CPCP data.
In some embodiments, record 635 is a collection of records stored in the relational database data structure of the controller instance. In some embodiments, some or all of the input modules, control modules and virtualization modules use, update and manage records stored in relational database data structures. That is, the inputs and / or outputs of these modules are stored in a relational database data structure.
In some embodiments, the system 600 maintains the same switching element data records in the relational database data structure of each instance, but in other embodiments it is based on the LDPS managed by each controller instance. Allows different sets of switching element data records to be stored in relational database data structures of different instances.
The PTD 640 of some embodiments is a secondary storage structure for storing user-specified network configuration data (eg, LCP data converted from inputs in the form of API calls). In some embodiments, the PTD for each controller instance uses system 600 to store configuration data for all users. Since the controller instance that receives the user input propagates the configuration data to the PTDs of the other controller instances, all the PTDs of all the controller instances have all the configuration data for all the users in these embodiments. However, in other embodiments, the PTD of the controller instance stores only the configuration data for a particular LDPS managed by the controller instance.
By allowing different controller instances to store the same or overlapping configuration data and / or secondary storage structure records, the system can fail any network controller (or relational database data structure and / or secondary storage structure instance. Improve overall recovery by being wary of data loss due to failure). For example, replicating a PTD across a controller instance allows a failed controller instance to quickly reload its PTD from another instance.
An inter-controller communication interface 645 is used to establish a communication channel (eg, RPC channel) with another controller instance (eg, by an exporter not shown). As shown, the inter-controller communication interface facilitates the exchange of data between different controller instances 605-615.
The managed switching element communication interface 650 facilitates communication between the controller instance and the managed switching element, as described above. In some embodiments, a managed switching element communication interface is used to propagate the UPCP data generated by the virtualization module 630 to each of the managed switching elements capable of converting universal data into customized data.
For some or all of the communication between the distributed controller instances, the system 600 uses Coordination Manager (CM) 655. CM655 in each instance allows an instance to coordinate certain behaviors with other instances. Different embodiments use CMs to coordinate sets of different behaviors between instances. Examples of such operations include writing to relational database data structures, writing to PTDs, controlling switching elements, facilitating communication between controllers related to the fault tolerance of controller instances, and the like. Also, CM is used to find the master of LDPS and the master of managed switching elements.
As mentioned above, different controller instances of the system 600 can control the operation of the same switching element or the operation of different switching elements. By distributing control of these behaviors across several instances, the system can more easily scale up to handle additional switching elements. Specifically, the system can distribute the management of different switching elements to different controller instances in order to benefit from the efficiencies that can be achieved by using a large number of controller instances. In such a distributed system, the number of calculations that each controller must perform to generate and distribute flow entries across the switching elements can be reduced by reducing the number of switching elements that each controller instance has under its control. Decrease. In other embodiments, a large number of controller instances can be used to create a scale-out network management system. Calculating the optimal method for distributing network flow tables in a large network is a CPU-intensive task. By partitioning the process by controller instances, the system 600 may use a larger but less sophisticated set of computer systems to create a scale-out network management system that can handle large networks.
In order to distribute the workload and avoid competing with controller instances that behave differently, the system 600 of some embodiments LDPS and / or any predetermined controller instance (eg, 605) within the system 600. It is called the master of the managed switching element (that is, the logical controller or the physical controller). In some embodiments, each master controller instance stores only the data associated with the managed switching element being processed by the master in a relational database data structure.
In some embodiments, as mentioned above, the CM facilitates communication between controllers related to the recovery of controller instances. For example, the CM realizes communication between controllers by the above-mentioned secondary storage device. A controller instance in a control system can fail for many reasons (eg, hardware failure, software failure, network failure, etc.). Different embodiments may use different techniques to determine if the controller instance has failed. In some embodiments, a consensus protocol is used to determine if a controller instance in the control system has failed. Some of these embodiments may use the Apache ZooKeeper to implement the consensus protocol, while other embodiments may implement the consensus protocol in other ways. ..
Some embodiments of CM655 may utilize a defined timeout to determine if the controller instance has failed. For example, if the CM of a controller instance does not respond to communication (eg, sent from another CM of another controller instance in the control system) within a certain length of time (ie, the specified timeout length), it responds. A controller instance that does not is determined to have failed. Other techniques may be utilized to determine if the controller instance has failed in other embodiments.
If the master controller instance fails, a new master for the LDP set and switching elements needs to be determined. Some embodiments of CM655 do so by performing a master election process (eg, to partition the management of the LDP set and / or to partition the management of the switching element) that elects the master controller instance. Make a good judgment. The CM655 of some embodiments may perform a master election process to elect a new master controller instance for both the LDP set and the switching element where the failed controller instance was the master. However, the CM655 of another embodiment includes (1) a master selection process for selecting a new master controller instance for the LDP set in which the failed controller instance was the master, and (2) a failed controller instance. You may perform another master election process to elect a new master controller instance for the switching element for which was the master. In these cases, the CM655 has two different controller instances, one for the LDP set where the failed controller instance was the master and another for the switching element where the failed controller instance was the master. May be determined as a new controller instance.
Alternatively or in connection with, the controller in some embodiments of the cluster executes a consensus algorithm to determine the leader controller as described above. The leader controller separates the tasks involving each controller instance in the cluster by assigning a master controller and possibly a hotstand-by controller to a particular work item to take over in the event of a master controller failure.
In some embodiments, the master election process is also to separate the management of LDP sets and / or the management of switching elements when controller instances are added to the control system. In particular, when the control system 600 detects a change in a member of a controller instance in the control system 600, some embodiments of CM655 perform a master selection process. For example, if the control system 600 detects that a new network controller has been added to the control system 600, the CM655 will transfer some of the LDP set management and / or switching element management from the existing controller instance to the new controller instance. A master selection process may be performed for redistributing. However, in other embodiments, if the control system 600 detects that a new network controller has been added to the control system 600, then part of the management of the LDP set and / or the management of the switching element is from the existing controller instance. Not redistributed to new network controllers. Instead, the control system 600 in these embodiments is an unassigned LDP set and / or switching element (eg, a new LDP set and / or switching element, or an LDP set and / or from a failed network controller. If a switching element) is detected, the unassigned LDP set and / or switching element is assigned to the new controller instance.
C. Separate management of LDP sets and managed switching elements FIG. 7 shows an example of designating a master controller instance for a switching element (ie, a physical controller) in a distributed system 700 similar to the system 600 of FIG. In this example, the two controllers 705 and 710 control three switching elements S1, S2 and S3 for two different users A and B. Through the two control applications 715 and 720, the two users have a large number of records similarly stored by the controller virtualization applications 745 and 750 in the two relational database data structures 755 and 760 of the two controller instances 705 and 710. Specify two different LDP sets 725 and 730 that are converted to.
In the example shown in FIG. 7, both control applications 715 and 720 of both controllers 705 and 710 can change the record of switching element S2 for both users A and B, but only controller 705 is the master of this switching element. Is. This example shows two different scenarios. The first scenario includes a controller 705 that updates record S2b1 at switching element S2 for user B. The second scenario includes a controller 705 in which the control application 720 updates the switching element S2 and the record S2a1 for the user A in the relational database data structure 760 and then updates the record S2a1 in the switching element S2. In the example shown in FIG. 7, the update is routed from the relational database data structure 760 of controller 710 to the relational database data structure 755 of controller 705 and then to the switching element S2.
Different embodiments use different techniques to propagate changes to the relational database data structure 760 of controller instance 710 to the relational database data structure 755 of controller instance 705. For example, to propagate this update, the virtualization application 750 of the controller 710 in some embodiments sends a set of records directly to the relational database data structure 755 (using an inter-controller communication module or exporter / importer). By). In response, the virtualization application 745 sends changes to the relational database data structure 755 to the switching element S2.
The system 700 of some embodiments records the record S2a1 in the switching element S2 in response to a request from the control application 720, rather than propagating changes in the relational database data structure to the relational database data structure of another controller instance. Use other techniques to change. For example, some embodiments of distributed control systems use a secondary storage structure (eg, PTD) as a communication channel between different controller instances. In some embodiments, the PTD is replicated across all instances and some or all of the relational database data structures are pushed from one controller instance to another via the PTD storage layer. Thus, in the example shown in FIG. 7, changes to the relational database data structure 760 are replicated to the PTD of controller 710, from which it can be replicated to the PTD of controller 705 or the relational database data structure 755.
Since some embodiments refer to a controller instance as the master of the switching element in addition to referencing the controller instance as the master of the LDPS, there may be other variations on the sequence of operations shown in FIG. In some embodiments, the different controller instances may be the switching element and the master of the corresponding record for that switching element in the relational database data structure, while in other embodiments the controller instance switches in the relational database data structure. Requires to be the master of all records for the element and its switching elements.
In an embodiment where the system 700 allows the nomination of a master for switching elements and relational database data structure records, the example shown in FIG. 7 shows an example in which the controller instance 710 is the master of the relational database data structure record S2a1. The controller instance 705 is a master for the switching element S2. If a controller instance other than the controller instances 705 and 710 was the master of the relational database data structure record S2a1, the request for modification of the relational database data structure record from the control application 720 must be propagated to the other controller instances. Wouldn't have been. Other controller instances modify relational database data structure records. With this change, the relational database data structure 755 and the switching element S2 update their records by one or more means of propagating the change to the controller instance 705.
In another embodiment, the controller instance 705 may be the master of the relational database data structure record S2a1 or the master of all the records of the switching element S2 and its relational database data structure. In these embodiments, the request for modification of the relational database data structure record from the control application 720 would have to be propagated to the controller instance 705 that modifies the record and switching element S2 in the relational database data structure 755.
As mentioned above, different embodiments employ different techniques to facilitate communication between different controller instances. Also, different embodiments implement controller instances in different ways. For example, in some embodiments, a stack of control applications (eg, 625 or 715 in FIGS. 6 and 7) and virtualization applications (eg, 630 or 745) are installed and run on a single computer. Also, in some embodiments, a large number of controller instances can be installed and run in parallel on a single computer. In some embodiments, the controller instance may also have a stack of components divided among several computers. For example, within an instance, the control application (eg, 625 or 715) may be on the first physical machine or virtual machine, and the virtualization application (eg, 630 or 745) may be on the second physical machine or virtual machine. You can.
FIG. 8 shows an example of the operation of several controller instances that function as controllers for distributing inputs, LDPS master controllers (also called logical controllers), and managed switching element master controllers (also called physical controllers). .. In some embodiments, not all controller instances include a full stack of different modules and interfaces as described above with reference to FIG. Alternatively, not all controller instances perform all full stack functions. For example, none of the controller instances 805, 810 and 815 shown in FIG. 8 has a full stack of modules and interfaces.
The controller instance 805 in this example is a controller instance for distributing the input. That is, the controller instance 805 of some embodiments gets input from the user in the form of an API call. By calling the API, the user queries for requests or information to configure a particular LDPS (eg, to configure a logical switching element or logical router implemented in a set of managed switching elements) (eg, the user's logic). You can specify a request for network traffic statistics for the logical port of the switch. The input module 820 of the controller instance 805 receives calls to these APIs in some embodiments and stores them in the PTD 825 and puts them into a form (eg, a data tuple or record) that can be sent to another controller instance. Convert.
The controller instance 805 in this example then sends these records to another controller instance involved in managing the records for a particular LDPS. In this example, controller instance 810 is involved in the LDPS record. The controller instance 810 receives a record from the PTD 825 of the controller instance 805, and stores the record in the PTD 845 which is a secondary storage structure of the controller instance 810. In some embodiments, the PTDs of different controller instances can communicate directly with each other and therefore do not need to rely on the inter-controller communication interface.
The control application 810 then detects adding these records to the PTD and processes them to generate or modify other records in the relational database data structure 842. In particular, the control application produces LFP data. As a result, the virtualized application detects changes and / or additions to these records in the relational database data structure and modifies and / or generates other records in the relational database data structure. These other records represent UPCP data in this example. These records are then sent via the inter-controller communication interface 850 of the controller instance 810 to another controller instance that manages at least one of the switching elements that implements a particular LDPS.
The controller instance 815 in this example is a controller instance that manages the switching element 855. The switching element implements at least a portion of a particular LDPS. The controller instance 815 receives a record representing UPCP data from the controller instance 810 via the inter-controller communication interface 865. In some embodiments, the controller instance 815 has a control application and a virtualization application for converting UPCP data into CPCP data. However, in this example, the controller instance 815 only identifies the set of managed switching elements to which the UPCP data is sent. In this way, the controller instance 815 functions as an aggregation point for collecting data to be sent to the managed switching element that the controller is responsible for managing. In this example, the managed switching element 855 is one of the managed switching elements by the controller instance 815.
D. Input conversion layer FIG. 9 conceptually shows the software architecture for the input conversion application 900. The input conversion application of some embodiments functions as the input conversion layer 505 described above with reference to FIG. In particular, input conversion applications receive input from user interface applications that allow users to enter input values. The input conversion application translates the input into a request and dispatches the request to one or more controller instances to process the request. The input conversion application operates in the same controller instance in which the control application operates in some embodiments, but operates as an independent controller instance in other embodiments. As shown in FIG. 9, the input conversion application includes an input parser 905, a filter 910, a request generator 915, a request repository 920, a dispatcher 925, a response manager 930, and an inter-controller communication interface 940. ..
In some embodiments, the input conversion application 900 supports a set of LDPs and a set of API calls to specify a query for information. In these embodiments, a user interface application that allows the user to enter an input value is implemented to send an input in the form of an API call to the input conversion application 900. Thus, these API calls specify LDPS (eg, a user-specified logical switching element configuration) and / or query for user information (eg, network traffic statistics for the logical port of the user's logical switching element). To do. Further, the input conversion application 900 may obtain input from another input conversion application of a logical controller, a physical controller and / or another controller instance in some embodiments.
The input parser 905 of some embodiments receives input in the form of API calls from the user interface application. In some embodiments, the input parser extracts the user input value from the API call and passes the input value to the filter 910. Filter 910 excludes input values that do not meet certain requirements. For example, filter 910 excludes input values that specify invalid network addresses for logical ports. In response to an API call containing a non-conforming input value, the response manager 930 sends a response to the user indicating that the input is non-conforming.
The request generator 915 generates a request to be sent to one or more controller instances that process the request to generate a response to the request. These requests may include LDPS data for receiving inquiries about the controller instance to be processed and / or information. For example, the request may ask for statistical information on the logical port of the logical switching element managed by the user. The response to this request will include the requested statistics prepared by the controller instance responsible for managing the LDPS associated with the logical switching element.
Request generators 915 in different embodiments generate requests according to different formats, depending on the implementation of the controller instance that receives and processes the requests. For example, the request generated by the request generator 915 of some embodiments is in the form of a record (eg, a data tuple) suitable for storage in the relational database data structure of the controller instance that receives the request. In some of these embodiments, the receiving controller instance uses the nLog table mapping engine to process the records that represent the request. In another embodiment, the request is in the form of an object-oriented data object that can interact with the NIB data structure of the controller instance that receives the request. In these embodiments, the receiving controller instance processes the data object directly on the NIB data structure without going through the nLog table mapping engine.
The request generator 915 of some embodiments deposits the request in the request repository 920 so that the dispatcher 925 can send the generated request to the appropriate controller instance. The dispatcher 925 identifies the controller instance to which each request should be sent. In some cases, the dispatcher examines the LDPS associated with the request and identifies the controller instance that is the master of that LDPS. In some cases, the dispatcher is specific as a controller instance for sending requests when the request is specifically related to the switching element (for example, when the request relates to logical port statistics that map to the switching element's port). Identify the master of the switching element (ie, the physical controller). The dispatcher sends the request to the identified controller instance. The receiving controller instance returns a response if the request contains a query for information.
The inter-controller communication interface 940 is similar to the inter-controller communication interface 645 described above with reference to FIG. 6 in that it establishes a communication channel (eg, an RPC channel) with another controller instance to which a request can be sent. The communication channel is bidirectional in some embodiments, but unidirectional in other embodiments. When the channel is unidirectional, the inter-controller communication interface establishes a number of channels with another controller instance so that the input transform application can send requests and receive responses over different channels.
When the receiving controller instance receives a request that specifies a query for information, the controller instance processes the request and produces a response containing the queried information. The response manager 930 receives a response from the controller instance that processed the request over the channel established by the inter-controller communication interface 940. In some cases, two or more responses may be returned in response to the requested request sent. For example, a request for statistics from all logical ports of a user-managed logical switching element would return a response from each controller. Responses from multiple physical controller instances to a number of different switching elements whose ports are mapped to logical ports may be sent back to the input conversion application 900 either directly or through the master of the LDPS associated with the logical switch. In such cases, the response manager 930 of some embodiments merges these responses and sends a single merged response to the user interface application.
As described above, the control application running on the controller instance converts the data record representing the LCP data into the data record representing the LFP data by performing the conversion operation. Specifically, in some embodiments, the control application populates the LDP set into an LDPS table (eg, a logical forwarding table) created by the virtualization application.
E. nLog engine The controller instance in some embodiments performs a mapping operation by using an nLog table mapping engine that uses variations in data log table mapping techniques. Data logs are used in the field of database management to map one set of tables to another. Data logs are not a good tool for performing table mapping operations in network control system virtualization applications, as their current implementations are often slow.
Therefore, some embodiments of the nLog engine are custom designed to operate quickly so that LDPS data tuples can be mapped to the data tuples of managed switching elements in real time. This custom design is based on several custom design options. For example, some embodiments compile the nLog table mapping engine from a set of high-level declaration rules represented by the application developer (eg, by the developer of the control application). In some of these embodiments, one custom design option made to the nLog engine is to allow application developers to use only the AND operator to represent declarative rules. It is a thing. By preventing developers from using other operators (eg, OR, XOR, etc.), these embodiments relate to AND operations where the resulting nLog engine rules perform faster at runtime. Guarantee to be represented.
Another custom design option pertains to the join operations performed by the nLog engine. Join operations are a common database operation for creating associations between records in different tables. In some embodiments, performing an outer join operation (also called an outer join operation) is impractical for the real-time operation of the engine because it can be time consuming, so the nLog engine makes the join operation an inner join operation. Limited to (also called inner join operation).
Yet another custom design option is to implement the nLog engine as a distributed table mapping engine run by several different controller instances. Some embodiments implement a distributed nLog engine by partitioning the management of the LDP set. Separating the management of an LDP set involves designating only one controller instance for each of that particular LDPS as the instance responsible for designating the records associated with that particular LDPS. For example, if the control system uses three switching elements to specify five LDP sets for five different users with two different controller instances, one controller instance will be two of the LDP sets. It may be the master for one related record and the other controller instance may be the master for the records for the other three LDP sets.
Separating the management of the LDP set further allocates a table mapping operation for each LDPS to the nLog engine of the controller instance involved in the LDPS in some embodiments. Distributing the nLog table mapping operation across several nLog instances reduces the load on each nLog instance, resulting in an increase in the speed at which each nLog instance can complete the mapping operation. In addition, such distribution reduces the memory size requirement for each machine running the controller instance. Some embodiments classify nLog table mapping operations across different instances by instructing the first join operation performed by each nLog instance to be based on LDPS parameters. Such instructions ensure that if an instance initiates a set of LDPS-related join operations that are not managed by the nLog instance, the join operation for each nLog instance will immediately fail and end. Some examples of using the nLog engine are described in the previously incorporated US Patent Application No. 13 / 177,533.
F. Control layer FIG. 10 shows a control application 1000 of some embodiments of the present invention. The application 1000 is used in some embodiments as the control module 625 of FIG. The application 1000 uses an nLog table mapping engine to map an input table containing an input data tuple representing LCP data to a data tuple representing LFP data. This application is on top of a virtualization application 1005 that receives a data tuple specifying an LDP set from the control application 1000. Virtualization application 1005 maps data tuples to UPCP data.
More specifically, the control application 1000 allows different users to define different sets of LDPs that specify the desired configuration of the logical switching elements that they manage. The control application 1000 transforms the data for each LDPS of each user into a set of data tuples that specify the LFP data for the logical switching element associated with the LDPS by the mapping operation. In some embodiments, the control application runs on the same host on which the virtualization application 1005 runs. The control application and the virtualization application need not run on the same machine in other embodiments.
As shown in FIG. 10, the control application 1000 includes a set of rule engine input tables 1010, a set of function tables and constant tables 1015, an importer 1020, a rule engine 1025, and a set of rule engine output tables 1045. It includes a translator 1050, an exporter 1055, a PTD 1060, and a compiler 1035. Compiler 1035 is one component of an application that runs in a different time instance than the other components of the application. The compiler works when the developer needs to specify a rules engine for a particular control application and / or virtualized environment, while the remaining application modules are LDPs specified by one or more users. Operates at runtime when the application interfaces with the virtualized application to deploy the set.
In some embodiments, the compiler 1035 takes a fairly small set (eg, hundreds of rows) of declarative instructions 1040 specified in a declarative language and uses them in a rule engine 1025 that performs table mapping for the application. Convert to a large set (eg, thousands of lines) of code that specifies the behavior (ie, object code). Therefore, the compiler greatly simplifies the process of control application developers who specify and update control applications. It uses a high-level programming language that allows developers to compactly define complex mapping behaviors for control applications through a compiler, followed by one or more changes (eg, logic supported by the control application). This is because this mapping operation can be updated in response to a change in the networking function, a change in the desired behavior of the control application, etc.). In addition, the compiler eliminates the need for developers to consider the order in which events arrive at the control application when defining mapping behavior.
In some embodiments, the rules engine (RE) input table 1010 is a logical data and / or switching configuration (eg, access control list configuration, dedicated virtual network configuration, port security configuration) specified by the user and / or control application. Etc.) including tables with. The input table 1010 further includes a table containing physical data from the managed switching element by the network control system in some embodiments. In some embodiments, such physical data includes data about managed switching elements and other data about network configurations adopted by network control systems to deploy different sets of LDPs for different users.
The RE input table 1010 partially populates the LCP data provided by the user. The RE input table 1010 also includes LFP data and UPCP data. In addition to the RE input table 1010, the control application 1000 includes other miscellaneous tables 1015 used by the rule engine 1025 to collect inputs for table mapping operations. These tables 1015 include constant tables that store constant values for the constants for which the rules engine 1025 needs to perform table mapping operations. For example, in the constant table 1015, the constant "zero" specified as the value 0, the constant "dispatch_port_no" specified as the value 4,000, and the constant "broadcast_MAC_addr" specified as the value 0xFF: FF: FF: FF: FF: FF And may be included.
When the rule engine 1025 references a constant, the corresponding value specified for the constant is actually searched and used. Further, the values specified for the constants in the constant table 1015 may be changed and / or updated. As such, the constant table 1015 provides the ability for the rule engine 1025 to change a defined value for a referenced constant without having to rewrite or recompile the code that specifies the behavior of the rule engine 1025. Table 1015 further includes a function table that stores the functions that the rule engine 1025 needs to use to calculate the values needed to populate the output table 1045.
The rule engine 1025 performs a table mapping operation that specifies a method for converting LCP data to LFP data. Whenever one of the Rule Engine (RE) input tables changes, the Rule Engine performs a set of table mapping operations that may result in changing one or more data tuples in one or more RE output tables. To do.
As shown in FIG. 10, the rule engine 1025 includes an event processing program 1022, some query plans 1027, and a table processor 1030. Each query plan is a set of rules that specifies a set of join operations to be performed when one of the RE input tables is modified. In the following, such changes will be referred to as input table events. In this example, each query plan is generated by compiler 1035 from one declarative rule in the set of declarations 1040. In some embodiments, two or more query plans are generated from one declarative rule. For example, a query plan is created for each table joined by one declarative rule. That is, if a declarative rule specifies to join four tables, four different query plans are created from that one declaration. In some embodiments, the query plan is defined by using the nLog declarative language.
The event processing program 1022 of the rule engine 1025 detects the occurrence of each input table event. Event processing programs of different embodiments detect the occurrence of input table events in different ways. In some embodiments, the event processor registers a callback in the RE input table to notify changes to the records in the RE input table. In such an embodiment, the event processing program 1022 detects an input table event when one of the records receives a notification from the modified RE input table.
In response to the detected input table event, the event processor 1022 (1) selects the appropriate query plan for the detected table event and (2) tells the table processor 1303 to execute the query plan. Instruct. To execute a query plan, the table processor 1030, in some embodiments, populates one or more input tables 1010 and a miscellaneous table 1015 with one or more records representing one or more sets of data values. Performs the join operation specified by the query plan to generate. Next, the table processor 1030 of some embodiments performs a selection operation to (1) select a subset of data values from the records generated by the join operation, and (2) of the selected data values. Write a subset to one or more RE output tables 1045.
In some embodiments, the RE output table 1045 stores both logical network element data attributes and physical network element data attributes. Table 1045 is called the RE output table to store the output of the table mapping operation of the rule engine 1025. In some embodiments, the RE output table can be grouped into several different categories. For example, in some embodiments, these tables may be RE input tables and / or control application (CA) output tables. If the rule engine detects an input event that requires the execution of a query plan by modifying the table, the table is a RE input table. The RE output table 1045 may be the RE input table 1010 that generates an event that causes the rules engine to execute another query plan. Such an event is called an internal input event and is contrasted with an external input event, which is an event caused by a change in the RE input table by the control application 1000 or the importer 1020.
As will be further described below, if the exporter 1055 exports the changes to the virtualization application 1005 due to table changes, the table is a CA output table. The table of the RE output table 1045 may be both a RE input table, a CA output table, or a RE input table and a CA output table in some embodiments.
The exporter 1055 detects changes to the CA output table of the RE output table 1045. Exporters of different embodiments detect the occurrence of CA output table events in different ways. In some embodiments, the exporter registers a callback in the CA output table to notify changes to the records in the CA output table. In such an embodiment, the exporter 1055 detects an output table event when one of the records receives a notification from the modified CA output table.
In response to the detected output table event, the exporter 1055 gets some or all of the modified data tuples in the modified CA output table and uses the modified data tuples in the input table of the virtualized application 1005. Propagate to (not shown). In some embodiments, instead of the exporter 1055 pushing the data tuple to the virtualization application, the virtualization application 1005 pulls the data tuple from the CA output table 1045 to the input table of the virtualization application. In some embodiments, the CA output table 1045 of the control application 1000 and the input table of the virtualization 1005 may be the same. In yet another embodiment, the CA output table is essentially a virtualization application (VA) input table because the control application and the virtualization application use one set of tables.
In some embodiments, the control application does not retain data for the LDP set that the control application is not responsible for managing in the output table 1045. However, such data can be stored in the PTD and is converted by the translator 1050 into a format stored in the PTD. The PTD of control application 1000 transfers this data to one or more other controls of the other controller instance so that some of the other controller instances responsible for managing the LDP set associated with the data can process the data. Propagate to the application instance.
In some embodiments, the control application further brings the data stored in the output table 1045 (ie, the data held by the control application in the output table) to the PTD for data recovery. In addition, such data is transformed by the translator 1050, stored in the PTD, and propagated to other control application instances of other controller instances. Therefore, in these embodiments, the PTD of the controller instance has all the configuration data for all the LDP sets managed by the network control system. That is, each PTD includes a global view of the configuration of the logical network in some embodiments.
The importer 1020 interfaces with many different sources of input data and uses the input data to modify or create the input table 1010. The importer 1020 of some embodiments receives input data from the input conversion application 1070 via an inter-controller communication interface (not shown). The importer 1020 also interfaces with the PTD 1060 so that the data received via the PTD from other controller instances can be used as input data to modify or create the input table 1010. In addition, the importer 1020 further detects changes by using the RE input table & CA output table of the RE input table and the RE output table 1045.
G. Virtualization layer As mentioned above, the virtualization applications of some embodiments specify how different LDP sets of different users of the network control system can be realized by the managed switching element by the network control system. In some embodiments, the virtualization application specifies an implementation of an LDP set within the infrastructure of a managed switching element by performing a transformation operation. These conversion operations initially store the LDP set data record in a managed switching element and then switch to generate forwarding plane data (eg, flow entry) to define the forwarding behavior of the switching element. Convert to a control data record (eg, UPCP data) used by the element. The conversion operation further generates other data (eg, in a table) that specifies network structures (eg, tunnels, queues, queue collections, etc.) to be defined within and between managed switching elements. The network structure further includes management software switching elements that are dynamically deployed or preset management software switching elements that are dynamically added to the set of managed switching elements.
FIG. 11 shows a virtualization application 1100 of some embodiments of the present invention. The application 1100 is used in some embodiments as the virtualization module 630 of FIG. The virtualization application 1100 uses an nLog table mapping engine to map an input table containing LDPS data tuples that represent UPCP data. This application is under control application 1105 that produces LDPS data tuples. The control application 1105 is similar to the control application 1000 described above with reference to FIG. The virtualization application 1100 is similar to the virtualization application 1005.
As shown in FIG. 11, the virtualization application 1100 includes a set of rule engine input tables 1110, a set of function tables and constant tables 1115, an importer 1120, a rule engine 1125, and a set of rule engine output tables 1145. , Translator 1150, Exporter 1155, PTD 1160, and Compiler 1135. Compiler 1135 is similar to compiler 1035 described above with reference to FIG.
In order for the virtualization application 1100 to map LDPS data tuples to UPPCP data tuples, developers in some embodiments declare that they include instructions to map LDPS data tuples to UPPCP data tuples for some managed switching elements. Specify type instruction 1140 in a declarative language. In some such embodiments, these switching elements include UPCP for converting UPCP data to CPCP data.
For other managed switching elements, provisional Soka application 1100 maps the LDPS data tuple specific CPCP data tuples in each of the managed switching element having no UPCP. In some embodiments, when the virtualization application 1100 receives UPCP data from a virtualized application of another controller instance, it converts the universal physical control plane data tuple to a physical data pathset data tuple. The CPCP data tuples in output table 1140 are further mapped to CPCP data tuples for some managed switching elements that they do not have.
In some embodiments, if there is a chassis controller for converting UPCP tuples to CPCP data specific to a particular managed switching element, the virtualization application 1100 will use the input UPCP data for the particular managed switching element. Do not convert to CPCP data. In these embodiments, the controller instance with the virtualization application 1100 identifies a set of managed switching elements whose master is the controller instance and distributes UPCP data to the set of managed switching elements.
The RE input table 1110 is similar to the RE input table 1010. In addition to the RE input table 1110, the virtualization application 1100 includes other miscellaneous tables 1115 used by the rules engine 1125 to collect inputs for table mapping operations. These tables 1115 are similar to table 1015. As shown in FIG. 11, the rules engine 1125 is a table processor that functions in the same way as the event processing program 1122, some query plans 1127, and the event processing programs 1022, query plan 1027, and table processing 1030. Includes 1130 and.
In some embodiments, the RE output table 1145 stores both logical network element data attributes and physical network element data attributes. Table 1145 is called the RE output table to store the output of the table mapping operation of the rule engine 1125. In some embodiments, the RE output table can be grouped into several different categories. For example, in some embodiments, these tables may be RE input tables and / or virtualization application (VA) output tables. If the rule engine detects an input event that requires the execution of a query plan by modifying the table, the table is a RE input table. The RE output table 1145 may be a RE input table 1110 that generates an event that causes the rule engine to execute another query plan after it has been modified by the rule engine. Such an event is called an internal input event and is contrasted with an external input event, which is an event caused by a change in the RE input table by the control application 1105 via the importer 1020.
When a table change causes the exporter 1155 to export the change to a managed switching element or other controller instance, the table is a VA control table. As shown in FIG. 12, the table of the RE output table 1145 may be both the RE input table 1110, the VA output table 1205 or the RE input table 1110 and the VA output table 1205 in some embodiments.
The exporter 1155 detects changes in the RE output table 1145 to the VA output table 1205. Exporters of different embodiments detect the occurrence of VA output table events in different ways. In some embodiments, the exporter registers a callback in the VA output table to notify changes to the records in the VA output table. In such an embodiment, the exporter 1155 detects an output table event when one of the records receives a notification from the modified VA output table.
In response to the detected output table event, exporter 1155 gets each of the modified data tuples in the modified VA output table and uses this modified data tuple in another controller instance (eg, chassis controller). Propagate to one or more of or one or more of the managed switching elements. In doing so, the export completes deploying the LDPS (eg, one or more logical switching configurations) to one or more managed switching elements as specified by the record.
When the VA output table stores both logical network element data attributes and physical network element data attributes in some embodiments, the PTD 1160 in some embodiments is a logical network element data attribute and physical network element in output table 1145. Stores both logical network element attributes and physical network element attributes that are the same as or derived from the data attributes. However, in other embodiments, the PTD 1160 stores only physical network element attributes that are the same as or derived from the physical network element data attributes in output table 1145.
In some embodiments, the virtualization application does not retain data for the LDP set that the virtualization application is not responsible for managing in the output table 1145. However, such data can be stored in the PTD and is converted by the translator 1150 into a format stored in the PTD. The PTD of virtualization application 1100 transfers this data to one or more of the other controller instances so that some of the other virtualization application instances responsible for managing the LDP set associated with the data can process the data. Propagate to other virtualization application instances.
In some embodiments, the virtualization application further brings the data stored in the output table 1145 (ie, the data held by the virtualization application in the output table) to the PTD for data recovery. In addition, such data is transformed by the translator 1150, stored in the PTD, and propagated to other virtualization application instances of other controller instances. Therefore, in these embodiments, the PTD of the controller instance has all the configuration data for all the LDP sets managed by the network control system. That is, each PTD includes a global view of the configuration of the logical network in some embodiments.
The importer 1120 interfaces with a number of different sources of input data and uses the input data to modify or create the input table 1110. The importer 1120 of some embodiments receives input data from the input conversion application 1170 via the inter-controller communication interface. The importer 1120 also interfaces with the PTD 1160 so that the data received via the PTD from other controller instances can be used as input data to modify or create the input table 1110. Further, the importer 1120 further detects the change by using the RE input table & VA output table of the RE input table and the RE output table 1145.
H. Network controller FIG. 13 is a simplified diagram showing the table mapping operation of the control application and the virtualization application of some embodiments of the present invention. As shown in the upper half of FIG. 13, the control application 1305 maps the LCP data to the LFP data that the virtualization application 1310 of some embodiments then maps to the UPCP data or the CPCP data.
The lower half of FIG. 13 shows the table mapping operation of the control application and the virtualization application. As shown in the lower half, in some embodiments the nLog engine 1320 of the control application to generate LFP data from the input LCP data, along with the data in the constant and function tables (not shown) of all these data. Since collections are used, the control application's input table 1315 stores LCP data, LFP (LFP) data, and UPCP data.
FIG. 13 shows that the importer 1350 receives LCP data from the user (eg, through an input conversion application) and updates the input table 1315 of the control application with the LCP data. FIG. 13 shows that the importer 1350 detects or receives changes in the PTD 1340 (eg, changes in LCP data originating from other controller instances) in some embodiments and in response to such changes the input table 1315. Further indicates that may be updated.
The lower half of FIG. 13 further shows the table mapping operation of the virtualization application 1310. As shown, because LFP data is used by the nLog engine 1320 of the virtualization application in some embodiments to generate UPCP data and / or CPCP data along with data from constant and function tables (not shown). , The input table 1355 of the virtualization application stores LFP data. In some embodiments, the exporter 1370 sends the generated UPCP data to one or more other controller instances (eg, chassis controllers) to convert the UPCP data into CPCP data specific to the managed switching element. Generate CPCP data before pushing this data to a managed switching element or one or more managed switching elements. In another embodiment, the exporter 1370 sends the generated CPCP data to one or more managed switching elements and defines the forwarding behavior of those managed switching elements.
In some embodiments, if there is a chassis controller for converting UPCP data into CPCP data specific to a particular managed switching element, the virtualization application 1310 will use the input UPCP data for the particular managed switching element. Do not convert to CPCP data. In these embodiments, the controller instance with the virtualization application 1310 identifies a set of managed switching elements whose master is the controller instance and distributes UPCP data to the set of managed switching elements.
FIG. 13 shows that the importer 1375 receives the LFP data from the control application 1305 and updates the input table 1355 of the virtualization application with the LFP data. FIG. 13 shows that the importer 1375 detects or receives changes in the PTD 1340 (eg, changes in LCP data originating from other controller instances) in some embodiments and in response to such changes the input table 1355. Further indicates that may be updated. FIG. 13 further shows that the importer 1375 may receive UPCP data from another controller instance.
As mentioned above, some of the logical or physical data that the importer pushes into the input table of the control or virtualization application is related to the data generated by the other controller instance and passed to the PTD. For example, in some embodiments, the logical data for logical structures (eg, logical ports, logical queues, etc.) associated with a large number of LDP sets can change, and the translator (eg, the translator 1380 of the controller instance) can change. This change may be written to the input table. Another example of such logical data generated by another controller instance in a multi-controller instance environment occurs when the user provides LCP data for LDPS on a first controller instance that is not involved in LDPS. This change is added to the PTD of the first controller instance by the translator of the first controller instance. This change is then propagated across the PTDs of other controller instances by the replication process performed by the PTDs. The importer of the second controller instance, which is the master of the LDPS or the logical controller involved in the LDPS, finally changes in this way and writes the changes to the input table of the application (for example, the input table of the control application). Therefore, in some cases, the logical data that the importer writes to the input table may originate from the PTD of another controller instance.
As mentioned above, the control application 1305 and the virtualization application 1310 are two separate applications running on the same machine or different machines in some embodiments. However, other embodiments include these two applications as two modules of one integrated application, namely the control application module 1305 that generates logical data in LFP and the virtualization application that generates physical data in UPCP or CPCP. Realize the two applications.
Yet another embodiment integrates these operations within one integrated application without separating the control and virtualization operations of these two applications into two separate modules. FIG. 14 shows an example of such an integrated application 1400. This application 1400 uses the nLog table mapping engine 1410 to map data from the table input set 1415 to the table output set 1420, as in the embodiment described above. The input set of may include one or more tables. The input set of tables in this integrated application may include LCP data that needs to be mapped to LFP data, or LFP data that needs to be mapped to CPCP or UPCP data. The input set of the table may further include UPCP data that needs to be mapped to CPCP data. The UPCP data is delivered to the set of chassis controllers for the set of managed switching elements without being mapped to the CPCP data. The mapping is that the controller instance running the integrated application 1400 is a logical controller or a physical controller, and the managed switching element of the physical controller has a chassis controller for mapping UPCP data to CPCP data for the managed switching element. It depends on whether you are a master.
In this integrated control / virtualization application 1400, the importer 1430 acquires input data from the user or another controller instance. The importer 1430 further detects or receives changes in the PTD 1440 that are replicated relative to the PTD. Exporter 1425 exports output table records to another controller instance (eg, chassis controller).
When sending an output table record to another controller instance, the exporter has an inter-controller communication interface (non-controller) so that the data contained in the record is sent to the other controller instance over a communication channel (eg, RPC channel). (Shown) is used. When sending an output table record to a managed switching element, the exporter provides a managed switching element communication interface (not shown) so that the data contained in the record is sent to the managed switching element via two channels. use. One channel is established using a switch control protocol (eg, OpenFlow) to control the forwarding plane of the managed switching element, and the other channel uses the configuration protocol to send configuration data. Established.
When sending an output table record to the chassis controller, the exporter 1425 in some embodiments uses a single channel of communication to send the data contained in the record. In these embodiments, the chassis controller accepts data through this single channel but communicates with the managed switching element via two channels. The chassis controller will be described in more detail below with reference to FIG.
FIG. 15 shows another example of such an integrated application 1500. The integrated application 1500 uses a network information base data structure 1510 to store some of the input and output data of the nLog table mapping engine 1410. As mentioned above, the NIB data structure stores data in the form of object-oriented data objects. In the integrated application 1500, the output table 1420 is a primary storage structure. PTD1440 and NIB1510 are secondary storage structures.
The integrated application 1500 uses the nLog table mapping engine 1410 to map data from the table input set 1415 to the table output set 1420. In some embodiments, some of the data in the table output set 1420 is exported by the exporter 1425 to one or more other controller instances, or one or a managed switching element. Such exported data includes UPCP data or CPCP data that define the flow behavior of the managed switching element. These data may be backed up in the PTD by the translator 1435 in the PTD 1440 for data recovery.
A portion of the data in the table output set 1420 is published to NIB 1510 by NIB publisher 1505. These data include configuration information of logical switching elements that the user manages using the integrated application 1500. The data stored in the NIB 1510 is replicated by the Coordination Manager 1520 to other NIBs in other controller instances.
The NIB monitor 1515 receives change notifications from the NIB 1510 and pushes the changes to the input table 1415 via the importer 1430 for some notifications (eg, notifications related to the LDP set for which the integrated application is the master). To do.
The query manager 1525 uses an inter-controller communication interface (not shown) to interface with an input conversion application (not shown) to receive queries (eg, information queries) about configuration data. As shown in FIG. 15, the manager 1525 of some embodiments queries the NIB to provide state information (eg, logical port statistics) about the logical network elements that the user manages. Also connect to the interface. However, in another embodiment, the query manager 1525 queries the output table 1420 to get state information.
In some embodiments, application 1500 uses a secondary storage structure (not shown) other than PTD and NIB. These structures include a persistent non-transactional database (PNTD) and hash table. In some embodiments, these two types of secondary storage structures provide different query interfaces for storing different types of data, storing data in different ways, and / or processing different types of queries. To do.
A PNTD is a persistent database stored on a disk or other non-volatile memory. Some embodiments use this database to store data about one or more switching element attributes or switching element operations (eg, statistics, calculations, etc.). For example, this database is used in some embodiments to store the number of packets routed through a particular port of a particular switching element. Other examples of data types stored in PNTD include error messages, log files, warning messages and billing data.
The PNTD in some embodiments has a database query manager (not shown) capable of processing database queries, but this query manager cannot process complex conditional transaction queries because it is not a transactional database. In some embodiments, access to the PNTD is faster than access to the PTD, but slower than access to the hash table.
Unlike PNTD, a hash table is not a database stored on disk or other non-volatile memory. Instead, the hash table is a storage structure that is stored in volatile system memory (eg, RAM). Hash tables use a hashing technique that uses a hash index to quickly identify the records stored in the table. This structure, combined with the placement of the hash table in system memory, allows access to this table very quickly. To facilitate this rapid access, a simplified query interface is used in some embodiments. For example, in some embodiments, the hash table has only two queries: a Put query to write a value to the table and a Get query to retrieve the value from the table. Some embodiments use hash tables to store rapidly changing data. Examples of such rapidly changing data are network entity status, statistics, status, uptime, link configuration and packet processing information. Also, in some embodiments, the integrated application uses a hash table as a cache to store repeatedly queried information such as flow entries written to a large number of nodes. Some embodiments employ a hash structure in the NIB for quick access to records in the NIB. Therefore, in some of these embodiments, the hash table is part of the NIB data structure.
PTD and PNTD improve controller recovery by storing network data on the hard disk. If the controller system fails, network configuration data is stored on disk in PTD and log file information is stored on disk in PNTD.
I. Network control system hierarchy FIG. 16 conceptually shows an example of the architecture of the network control system 1600. In particular, FIG. 16 shows that different elements of the network control system generate CPCP data from the inputs. As shown, the network control system 1600 of some embodiments includes an input conversion controller 1605, a logical controller 1610, physical controllers 1615 and 1620, and three managed switching elements 1625 to 1635. FIG. 16 further shows five machines 1640 to 1660 connected to managed switching elements (written as "MS" in FIG. 16) and exchanging data between them. The architectural details such as the number of controllers, the number of managed switching elements and machines in each layer in the hierarchy shown in FIG. 16, and the relationship between the controllers, the managed switching elements and the machines are for illustration purposes only. Absent. It will be appreciated by those skilled in the art that many other different combinations of controllers, switching elements and machines are possible for the network control system 1600.
In some embodiments, each controller in a network control system has a full stack of different modules and interfaces described above with reference to FIG. However, each controller does not have to use all the modules and interfaces to perform the functionality given to the controller. Alternatively, in some embodiments, the controller in the system has only the modules and interfaces necessary to perform the functionality given to the controller. For example, the logical controller 1610, which is the master of the LDPS, does not have an input module (eg, an input conversion application) to generate UPCP data from the input LCP data, but a control module and a virtualization module (eg, a control application or). It has a virtualized application or an integrated application).
Also, different combinations of different controllers may be operating on the same machine. For example, the input conversion controller 1605 and the logic controller 1610 may operate in the same arithmetic unit. Also, one controller may function differently for different LDP sets. For example, a single controller may be the master of the first LDPS and the master of the managed switching element that implements the second LDPS.
The input conversion controller 1605 includes an input conversion application (not shown) that generates LCP data from inputs received from a user who specifies a particular LDPS. The input conversion controller 1605 identifies the LDPS master from the configuration data for the system 1605. In this example, the master of the LDPS is the logical controller 1610. In some embodiments, the two or more controllers may be masters of the same LDPS. Also, one logical controller may be the master of two or more LDP sets.
The logical controller 1610 is involved in a particular LDPS. Therefore, the logical controller 1610 generates UPCP data from the LCP data received from the input conversion controller. Specifically, the control module (not shown) of the logical controller 1610 generates LFP data from the received LCP data, and the virtualization module (not shown) of the logical controller 1610 generates UPCP data from the LFP data.
The logical controller 1610 identifies the physical controller that is the master of the managed switching element that implements LDPS. In this example, the logical controllers 1610 identify the physical controllers 1615 and 1620 because the managed switching elements 1625 to 1635 are configured to implement LDPS in this example. The logical controller 1610 sends the generated UPCP data to the physical controllers 1615 and 1620.
Each of the physical controllers 1615 and 1620 may be the master of one or more managed switching elements. In this example, the physical controller 1615 is the master of the two managed switching elements 1625 and 1630, and the physical controller 1620 is the master of the managed switching elements 1635. As a master of a set of managed switching elements, the physical controller of some embodiments generates CPCP data specific to each of the managed switching elements from the received UPCP data. Therefore, in this example, the physical controller 1615 generates customized PCP data for each of the managed switching elements 1625 and 1630. The physical controller 1320 generates customized PCP data for the managed switching element 1635. The physical controller sends CPCP data to a managed switching element whose master is the controller. In some embodiments, many physical controllers may be masters of the same managed switching element.
In addition to sending CPCP data, the physical controller of some embodiments receives data from the managed switching element. For example, the physical controller receives the configuration information of the managed switching element (for example, the VIF identifier of the managed switching element). The physical controller maintains the configuration information and further sends the information to the logical controller so that the logical controller has the configuration information of the managed switching element that realizes the LDP set in which the logical controller is the master.
Each of the managed switching elements 1625 to 1635 generates physical forwarding plane data from the CPCP data received by the managed switching elements. As mentioned above, the physical forwarding plane data defines the forwarding behavior of the managed switching element. In other words, the managed switching element uses CPCP data to populate the forwarding table. Managed switching elements 1625 to 1635 transfer data between machines 1640 to 1660 according to the forwarding table.
FIG. 17 conceptually shows an example of the architecture of the network control system 1700. As shown in FIG. 16, FIG. 17 shows that different elements of the network control system generate CPCP data from the inputs. In contrast to the network control system 1600 of FIG. 16, the network control system 1700 has chassis controllers 1725 to 1735. As shown, the network control system 1700 of some embodiments includes an input conversion controller 1705, a logical controller 1610, a physical controller 1715 and 1720, a chassis controller 1725 to 1735, and three managed switching elements 1740 to. Includes 1750 and. FIG. 17 further shows five machines 1755-1775 connected to managed switching elements 1740 to 1750 and exchanging data between them. The architectural details such as the number of controllers, the number of managed switching elements and machines in each layer in the hierarchy shown in FIG. 17, and the relationship between the controllers, the managed switching elements, and the machines are for illustration purposes only. Absent. It will be appreciated by those skilled in the art that many other different combinations of controllers, switching elements and machines are possible for the network control system 1700.
The input conversion controller 1705 is similar to the input conversion controller 1705 in that it includes an input conversion application that generates LCP data from the input received from a user who specifies a particular LDPS. The input conversion controller 1705 identifies the master of the LDPS from the configuration data for the system 1705. In this example, the master of the LDPS is the logical controller 1710.
The logical controller 1710 is similar to the logical controller 1610 in that it generates UPCP data from the LCP data received from the input conversion controller 1705. The logical controller 1710 identifies the physical controller that is the master of the managed switching element that implements LDPS. In this example, the logical controllers 1710 identify physical controllers 1715 and 1720 because the managed switching elements 1740 to 1750 are configured to implement LDPS in this example. The logical controller 1710 sends the generated UPCP data to the physical controllers 1715 and 1720.
Like the physical controllers 1615 and 1620, each of the physical controllers 1715 and 1720 may be the master of one or more managed switching elements. In this example, the physical controller 1715 is the master of the two managed switching elements 1740 and 1745, and the physical controller 1730 is the master of the managed switching elements 1750. However, the physical controllers 1715 and 1720 do not generate CPCP data for managed switching elements 1740 to 1750. As the master of the set of managed switching elements, the physical controller sends UPCP data to the chassis controllers involved in each of the managed switching elements for which the physical controller is the master. That is, the physical controller of some embodiments identifies a chassis controller that interfaces to a managed switching element on which the physical controller is the master. In some embodiments, the physical controllers identify these chassis controllers by determining if they are subscribed to the channels of the physical controllers.
The chassis controller of some embodiments has a one-to-one relationship with the managed switching element. The chassis controller receives UPCP data from the physical controller that is the master of the managed switching element, and generates CPCP data specific to the managed switching element. An example of the chassis controller architecture will be described in more detail below with reference to FIG. In some embodiments, the chassis controller operates on the same machine on which the managed switching elements managed by the chassis controller operate, whereas in other embodiments, the chassis controller and managed switching elements operate on different machines. To do. In this example, the chassis controller 1725 and the managed switching element 1740 operate in the same arithmetic unit.
Each of the managed switching elements 1740 to 1750, such as the managed switching elements 1625 to 1635, generates physical forwarding plane data from the CPCP data received by the managed switching elements. Managed switching elements 1740 to 1750 use CPCP data to populate their respective forwarding tables. Managed switching elements 1740 to 1750 transfer data between machines 1755 to 1775 according to the flow table.
As described above, the managed switching element may realize two or more LDPS in some cases. In such a case, the physical controller that is the master of such managed switching elements receives the UPCP data for each LDP set. Therefore, the physical controller in the network control system 1700 may function as an aggregation point for relaying UPCP data for different LDP sets to the chassis controller for a specific managed switching element that realizes the LDP set.
Although the chassis controller shown in FIG. 17 is at a level above the managed switching element, the chassis controller is generally a chassis controller because some embodiments are in or adjacent to the managed switching element. Operates at the same level as the controlled switching element operates.
In some embodiments, the network control system may have a hybrid of network control systems 1600 and 1700. That is, in this hybrid network control system, a part of the physical controller generates CPCP data for some of the managed switching elements, and a part of the physical controller generates CPCP data for some of the managed switching elements. Do not generate. For the latter managed switching element, the hybrid system has a chassis controller that produces CPCP data.
As mentioned above, the chassis controller of some embodiments is a controller for managing a single managed switching element. The chassis controller of some embodiments does not have a full stack of different modules and interfaces described above with reference to FIG. One of the modules that a chassis controller has is a chassis control application that generates CPCP data from UPCP data that it receives from one or more physical controllers. FIG. 18 shows an example of the architecture for chassis control application 1800. The application 1800 uses an nLog table mapping engine to map an input table containing an input data tuple representing UPCP data to a data tuple representing LFP data. The application 1800 manages the managed switching element 1885 in this example by exchanging data with the managed switching element 1885. In some embodiments, application 1800 (ie, chassis controller) operates on the same machine on which the managed switching element 1885 is operating.
As shown in FIG. 18, the chassis control application 1800 includes a set of rule engine input tables 1810, a set of function tables and constant tables 1815, an importer 1820, a rule engine 1825, and a set of rule engine output tables 1845. , Exporter 1855, managed switching element communication interface 1865, and compiler 1835. FIG. 18 further shows the physical controller 1805 and the managed switching element 1885.
The compiler 1835 is similar to the compiler 1035 of FIG. In some embodiments, the rules engine (RE) input table 1810 is a UPCP data and / or switching configuration (eg, an access control list) sent by the physical controller 1805, which is the master of the managed switching element 1885, to the chassis control application 1800. Includes tables with configurations, dedicated virtual network configurations, port security configurations, etc.). The input table 1810 further includes a table containing physical data from the managed switching element 1885. In some embodiments, such physical data includes data about the managed switching element 1885 (eg, CPCP data, physical forwarding data) and other data about the configuration of the managed switching element 1885.
The RE input table 1810 is similar to the RE input table 1010. The input table 1810 is partially populated by the UPCP data provided by the physical controller 1805. The physical controller 1805 of some embodiments receives UPCP data from one or more logical controllers (not shown).
In addition to the input table 1810, chassis control application 1800 includes other miscellaneous tables 1815 used by the rules engine 1825 to collect inputs for table mapping operations. These tables 1815 are similar to table 1015. As shown in FIG. 18, the rules engine 1825 is a table processor that functions similarly to the event processor 1822, some query plans 1827, and the event processor 1022, query plan 1027, and table processing 1030. Includes 1830 and.
In some embodiments, the RE output table 1845 stores both logical network element data attributes and physical network element data attributes. Table 1845 is called the RE output table to store the output of the table mapping operation of the rule engine 1825. In some embodiments, the RE output table can be grouped into several different categories. For example, in some embodiments, these tables may be RE input tables and / or chassis control application (CCA) output tables. If the rule engine detects an input event that requires the execution of a query plan by modifying the table, the table is a RE input table. The RE output table 1845 may be a RE input table 1810 that generates an event that causes the rule engine to execute another query plan after it has been modified by the rule engine. Such an event is called an internal input event and is contrasted with an external input event, which is an event caused by a change in the RE input table by the control application 1805 via the importer 1820. When a table change causes the exporter 1855 to export the change to a managed switching element or other controller instance, the table is a CA output table.
The exporter 1855 detects changes in the RE output table 1845 to the CCA output table. Exporters of different embodiments detect the occurrence of CCA output table events in different ways. In some embodiments, the exporter registers a callback in the CCA output table to notify changes to the records in the CCA output table. In such an embodiment, the exporter 1855 detects an output table event when one of the records receives a notification from the modified CCA output table.
In response to the detected output table event, exporter 1855 gets each of the modified data tuples in the modified output table and uses this modified data tuple for another controller instance (eg, a physical controller). Propagate to one or more or managed switching elements 1885. Exporter 1855 uses an inter-controller communication interface (not shown) to send modified data tuples to other controller instances. The inter-controller communication interface establishes a communication channel (eg, RPC channel) with another controller instance.
The exporter 1855 of some embodiments uses the managed switching element communication interface 1865 to send the modified data tuples to the managed switching element 1885. The managed switching element communication interface of some embodiments establishes two channels of communication. The managed switching element communication interface uses a switching control protocol to establish the first channel of the two channels. An example of a switching control protocol is the OpenFlow protocol. The OpenFlow protocol is, in some embodiments, a communication protocol for controlling the forwarding plane (eg, forwarding table) of a switching element. For example, the OpenFlow protocol has a command to add a flow entry to a managed switching element 1885, a command to remove a flow entry from a managed switching element 1885, and a command to change a flow entry in a managed switching element. And provide.
The managed switching element communication interface uses the configuration protocol to send out configuration information to establish a second channel of the two channels. In some embodiments, the configuration information includes information for configuring the managed switching element 1885, such as information for input ports, output ports, QoS configurations for ports, and the like.
The managed switching element communication interface 1865 receives updates in the managed switching element 1885 from the managed switching element 1885 via two channels. If there is a flow entry or configuration change in the managed switching element 1885 that is not initiated by the chassis control application 1800, the managed switching element 1885 in some embodiments sends an update to the chassis control application. Examples of such changes include failure of the machine connected to the port of managed switching element 1885, VM migration to managed switching element 1885, and the like. The managed switching element communication interface 1865 sends updates to the importer 1820 that modifies one or more input tables 1810. If there is an output generated by the rules engine 1825 from these updates, the exporter 1855 sends this output to the physical controller 1805.
J. Generate flow entry FIG. 19A-B shows an example of creating a tunnel between two managed switching elements based on UPCP data. Specifically, FIGS. 19A-B show four different stages of a series of operations performed by different components of the network management system 1900 to establish a tunnel between the two managed switching elements 1925 and 1930. It is shown in 1901-1904. 19A-B further show the logic switching element 1905 and VM1 and VM2. Each of the four stages 1901-1904 shows the network control system 1900 and the managed switching elements 1925 and 1930 at the bottom, and the VM connected to the logical switching element 1905 and the logical switching element 1905 at the top. VMs are shown both at the top and bottom of each stage.
As shown in the first stage 1901, the logical switching element 1905 transfers data between VM1 and VM2. Specifically, the data reaches or originates from VM1 via the logical port 1 of the logical switching element 1905 and reaches or originates from VM2 via the logical port 2 of the logical switching element 1905. The logical switching element 1905 is implemented by the managed switching element 1925 in this example. That is, the logical port 1 is mapped to the port 3 of the managed switching element 1925, and the logical port 2 is mapped to the port 4 of the managed switching element 1925.
The network control system 1900 in this example has a controller cluster 1910 and two chassis controllers 1915 and 1920. The controller cluster 1910 includes an input conversion controller (not shown) that collectively generates UPCP data based on the input received by the controller cluster 1910, a logical controller (not shown), and a physical controller (not shown). The chassis controller receives the UPCP data and customizes the universal data to the PCP data specific to the managed switching element managed by each chassis controller. The chassis controllers 1915 and 1920 have CPCPs so that the managed switching elements 1925 and 1930 can generate the physical forwarding plane data used by the managed switching elements to transfer data between the managed switching elements 1925 and 1930. Data is passed to managed switching elements 1925 and 1930, respectively.
In the second stage 1902, the administrator of the network including the managed switching element 1930 creates the VM3 on the host (not shown) on which the managed switching element 1930 operates. The administrator creates port 5 of the managed switching element 1930 and connects VM3 to the port. When the port 3 is created, the managed switching element 1930 of some embodiments sends information about the newly created port to the controller cluster 1910. In some embodiments, the information may include a port number, a network address (eg, an IP address and a MAC address), a transmission zone to which the managed switching element belongs, a machine connected to the port, and the like. As described above, this configuration information passes through the chassis controller that manages the managed switching element, and then through the physical controller and the logical controller to the user who manages the logical switching element 1905. For this user, a new VM is available to be added to the user-managed logical switching element 1905.
At step 1903, the user in this example uses the VM3 and decides to connect the VM3 to the logical switching element 1905. As a result, the logical port 6 of the logical switching element 1905 is created. Therefore, data that reaches or originates from VM3 passes through logical port 6. In some embodiments, the controller cluster 1910 implements a logical switching element to create a tunnel between each pair of managed switching elements having a pair of ports to which a pair of logical ports of the logical switching element are mapped. Instruct all managed switching elements to do. In this example, the tunnel is the data between logical port 1 and logical port 6 (ie, between VM1 and VM3) and between logical port 2 and logical port 6 (ie, between VM2 and VM3). Can be established between managed switching elements 1925 and 1930 to facilitate the exchange of. That is, the data exchanged between the port 3 of the managed switching element 1925 and the port 5 of the managed switching element 1930 and the exchange between the port 4 of the managed switching element 1925 and the port 5 of the managed switching element 1930. The data to be generated may pass through a tunnel established between the managed switching elements 1925 and 1930.
Since logical port 1 and logical port 2 are mapped to two ports on the same managed switching element 1925, the tunnel between the two managed switching elements is between logical port 1 and logical port 2 (ie, that is). Not required to facilitate the exchange of data (between VM1 and VM2).
The third stage 1903 further indicates that the controller cluster 1910 sends UPCP data specifying instructions to create a tunnel from the managed switching element 1925 to the managed switching element 1930. In this example, the UPCP data is sent to the chassis controller 1915, which customizes the UPCP data to PCP data specific to the managed switching element 1925.
The fourth stage 1904 indicates that the chassis controller 1915 sends PCP data to the tunnel specifying an instruction to create a tunnel and an instruction to forward the packet to the tunnel. The managed switching element 1925 creates a tunnel to the managed switching element 1930 based on the CPCP data. More specifically, the managed switching element 1925 creates a port 7 and establishes a tunnel (eg, a GRE tunnel) to port 8 of the managed switching element 1930. A more detailed operation of creating a tunnel between two managed switching elements will be described below.
FIG. 20 conceptually illustrates the process 2000 that some embodiments perform to generate CPCP data from UPCP data that specifies the creation and use of tunnels between two managed switching elements. In some embodiments, process 2000 is performed by a chassis controller that interfaces with the managed switching element or a physical controller that interfaces directly with the managed switching element.
Process 2000 starts by receiving UPCP data from the logical controller or the physical controller. In some embodiments, the UPCP data have different types. One type of UPCP data is a universal tunnel flow instruction that specifies the creation and use of tunnels in managed switching elements. In some embodiments, universal tunnel flow instructions include instructions for ports created on managed switching elements in the network. This port is the port of the managed switching element to which the user maps the logical port of the logical switching element. This port is also the destination port that the tunneled data needs to reach. Information about the ports can be found in (1) the transmission zone to which the managed switching element with the port belongs and (2) the tunnel used to build the tunnel to the managed switching element with the destination port in some embodiments. Types of tunnels based on protocols (eg, GRE, CAPWAP, etc.) and (3) network addresses (eg, IP addresses) of managed switching elements that have destination ports (eg, act as one end of the tunnel to establish). IP address of the VIF to be used) is included.
Next, process 2000 determines (in 2010) whether the received UPCP data is a universal tunnel flow instruction. In some embodiments, the UPCP data specifies the type so that the process 2000 can determine the type of universal plane data received. If process 2000 determines that the universal data received is not a universal tunnel flow instruction (in 2010), it proceeds to 2015 to generate CPCP data and manage the generated data in a managed switching element managed by process 2000. Process UPCP data so that it can be sent to. After that, the process 2000 ends.
If processing 2000 determines that the received UPCP data is a universal tunnel flow instruction (in 2010), it proceeds to 2020 and parses the data to obtain information about the destination port. Process 2000 then determines (at 2025) whether the managed switching element having the destination port is in the same transmission zone as the managed switching element having the source port. The managed switching element having the source port is a managed switching element managed by the chassis controller or the physical controller that executes the process 2000. In some embodiments, transmission zones include a group of machines that can communicate with each other without the use of second level managed switching elements such as pool nodes.
In some embodiments, the logical controller determines if the managed switching element with the destination port is in the same transmission zone as the managed switching element with the source port. The logical controller considers this determination when preparing a universal tunnel flow instruction to be sent (via the physical controller) to the chassis controller performing process 2000. Specifically, the universal tunnel flow instruction contains different information to create different tunnels. After the description of FIG. 21, examples of these different tunnels will be described below. In these embodiments, process 2000 skips 2025 and proceeds to 2015.
Process 2000 proceeds to 2015 as described above when it is determined (at 2025) that the managed switching element including the source port and the managed switching element including the destination port are not in the same transmission zone. If process 2000 determines (at 2025) that the managed switching element including the source port and the managed switching element including the destination port are in the same transmission zone, proceed to 2030 to customize and customize the universal tunnel flow instruction. The information is sent to a managed switching element that has a source port. Customization of the universal tunnel flow instruction is described in detail below. After that, the process 2000 ends.
In FIG. 21, some embodiments generate customized tunnel flow instructions and manage switching the customized instructions so that the managed switching element creates a tunnel and sends data to the destination through the tunnel. The process to be executed to send to the element is conceptually shown. In some embodiments, process 2100 is performed by a controller instance that interfaces with the managed switching element or a physical controller that interfaces directly with the managed switching element. In some embodiments, the process 2100 is managed by the controller executing the process 2100, receiving the universal tunnel flow instruction, parsing the port information regarding the destination port, and managing the managed switching element having the destination port. It starts when it is determined that it is in the same transmission zone as the target switching element.
Process 2100 begins by generating an instruction (at 2105) to create a tunnel port. In some embodiments, process 2100 generates an instruction to create a tunnel port in a managed switching element managed by the controller based on the port information. The instruction includes the type of tunnel to be established, the IP address of the NIC that is the destination end of the tunnel, and the like. Managed by the controller The tunnel port of the managed switching element becomes the other end of the tunnel.
Processing 2100 then sends the generated instruction to create the tunnel port to the managed switching element managed by the controller (at 2110). As mentioned above, the chassis controller of some embodiments or the physical controller that directly interfaces with the managed switching element uses two channels to communicate with the managed switching element. One channel is a configuration channel that exchanges configuration information with the managed switching element, and the other channel is a switching element control channel for exchanging flow entry and event data with the managed switching element (for example, the OpenFlow protocol). Channels established using). In some embodiments, processing uses a configuration channel to send generated instructions that create a tunnel port to a managed switching element managed by the controller. Upon receiving the generated instruction, the managed switching element of some embodiments creates a tunnel port at the managed switching element using the tunnel protocol specified by the tunnel type, and the tunnel port and destination port. Establish a tunnel to and from the port of the managed switching element that has. Once the tunnel port and tunnel are created and established, the managed switching element of some embodiments returns the value of the tunnel identifier (eg, 4) to the controller instance.
The process 2100 of some embodiments then receives the value of the tunnel port identifier (eg, "tunnel_port = 4") via the configuration channel (at 2115). Process 2100 then uses this received value to change the flow entry contained in the universal tunnel flow instruction. When this flow entry is sent to the managed switching element, it causes the managed switching element to perform an operation. However, because it is universal data, this flow entry identifies the tunnel port by a universal identifier (eg, tunnel_port) rather than the actual port number. For example, this flow entry in the received universal tunnel flow instruction is "If destination = destination machine's UUID, send to tunnel_port (If destination = destination machine's UUID, send to tunnel_port) "may be used. Process 2100 creates a flow entry (at 2120) containing the value of the tunnel port identifier. Specifically, process 2100 replaces the identifier for the tunnel port with the actual value of the identifier that identifies the created port. For example, the modified flow entry would look like "If destination = destination machine's UUID, send to 4".
Process 2100 then sends this flow entry to the managed switching element (at 2125). In some embodiments, the process sends this flow entry to a managed switching element via a switching element control channel (eg, an OpenFlow channel). The managed switching element uses this flow entry to update the flow entry table. After that, the managed switching element transfers the data through the tunnel by sending the data directed to the destination machine to the tunnel port. After that, the process ends.
FIG. 22A-B conceptually shows an example of the operation of the chassis controller 2210 to convert universal tunnel flow instructions into customized instructions received and used by the managed switching element 2215 in seven different stages 2201-2207. The chassis controller 2210 is similar to the chassis controller 1800 described above with reference to FIG. However, for ease of explanation, all components of chassis controller 2210 are not shown in FIGS. 22A-B.
As shown, the chassis controller 2210 includes an input table 2220, a rule engine 2225, and an output table 2230, which are similar to the input table 1820, the rule engine 1825, and the output table 1845. The chassis controller 2210 manages the managed switching element 2215. Two channels 2235 and 2240 are established between the chassis controller and the managed switching element 2215 in some embodiments. Channel 2235 is for exchanging configuration data (eg, data relating to the creation of ports, current status of ports, queues associated with managed switching elements, etc.). Channel 2240 is an OpenFlow channel (OpenFlow control channel) through which flow entries are exchanged in some embodiments.
The first stage 2201 indicates that the chassis controller 2210 has updated the input table 2220 using a universal tunnel flow instruction received from a physical controller (not shown). As shown, the universal tunnel flow instruction includes instruction 2245 to create a tunnel and flow entry 2250. As shown, instruction 2245 includes the type of tunnel created and the IP address of the managed switching element with the destination port. Flow entry 2250 specifies the action to be taken for universal data that is not specific to the managed switching element 2215. The rules engine performs table mapping operations on instruction 2245 and flow entry 2250.
The second stage 2202 shows the result of the table mapping operation performed by the rule engine 2225. Command 2260 is due to command 2245. Instructions 2245 and 2260 may be identical in some embodiments, but may not be identical in other embodiments. For example, the values in instructions 2245 and 2260 representing the type of tunnel may be different. Instruction 2260 includes the IP address and the type of tunnel created, among other things that may be included in instruction 2260. Flow entry 2250 stays in the input table 2220 because it did not trigger any table mapping operation.
The third stage 2203 indicates that the instruction 2260 is pushed to the managed switching element 2215 via the configuration channel 2235. The managed switching element 2215 creates a tunnel port and establishes a tunnel between the managed switching element 2215 and another managed switching element having a destination port. In some embodiments, one end of the tunnel is the created tunnel port and the other end of the tunnel is the port associated with the destination IP address. The managed switching element 2215 of some embodiments uses the protocol specified by the tunnel type to establish the tunnel.
The fourth stage 2204 indicates that the managed switching element 2215 created a tunnel port ("port 1" in this example) and a tunnel 2270. This stage further indicates that the managed switching element returns the actual value of the tunnel port identifier. The managed switching element 2215 sends this information through the OpenFlow channel 2240 in this example. The information enters the input table 2220 as input event data. The fifth step 2205 indicates that the input table 2220 is updated with information from the managed switching element 2215. This update triggers the rules engine 2225 to perform a table mapping operation.
The sixth step 2206 shows the result of the table mapping operation performed in the previous step 2204. The output table 2230 has a flow entry 2275 that specifies what happens here with respect to information specific to the managed switching element 2215. Specifically, flow entry 2275 specifies that the managed switching element 2215 should be sent to the packet through port 1 when the destination of the packet is the destination port. The seventh step 2207 indicates that the flow entry 2275 is being pushed to the managed switching element 2215 that forwards the packet using the flow entry 2275.
The data exchanged between the instruction 2245 and the chassis controller 2210 and the managed switching element 2215 as shown in FIG. 22 is a conceptual display of the universal tunnel flow instruction and the customized instruction, and is an actual expression and an actual expression. It does not have to be in the form.
Further, the example of FIG. 22 will be described with respect to the operation of the chassis controller 2210. This example is also applicable to some embodiments of the physical controller that transform the UPCP data into CPCP data for a managed switching element whose master is the physical controller.
19-22 show a tunnel between two management edge switching elements to facilitate the exchange of data between a pair of machines (eg, VMs) using the two logical ports of the logical switching element. Indicates the creation of. This tunnel covers one of the possible uses of the tunnel. In some embodiments of the invention, many other uses of the tunnel are possible in network control systems. Examples of tunnel applications are (1) a tunnel between a managed switching element and a pool node, and (2) one is an edge switching element and the other provides an L3 gateway service (ie, the network layer (ie, network layer). A tunnel between two managed switching elements (managed switching elements) that are connected to the router to receive routing services in L3), and (3) a logical port and another logical port connected to the L2 gateway service. There is a tunnel between two managed switching elements.
Next, a series of events that create a tunnel will be described in each of the three examples. For a tunnel between a managed switching element and a pool node, the pool node is provided first, followed by the managed switching element. The VM is connected to the port of the managed switching element. This VM is the first VM connected to a managed switching element. This VN is then coupled to the logical port of the logical switching element by mapping the logical port to the port of the managed switching element. When the logical port is mapped to the port of the managed switching element, the logical controller sends a universal tunnel flow instruction to (or to the physical controller) the chassis controller that interfaces with the managed switching element (eg, via the physical controller). hand).
The chassis controller then commands the managed switching element to create a tunnel to the pool node. Once the tunnel is created, another VM that is subsequently provided and connected to the managed switching element exchanges data with the pool node when this new VM is joined to a logical port on the same logical switching element. Share the same tunnel. When a new node is coupled to a logical port on a different logical switch, the logical controller sends out the same universal tunnel flow instruction that was transmitted when the first VM was connected to the managed switching element. However, the universal tunnel flow instruction does not cause a new tunnel to be created to the pool node, for example because the tunnel has already been created and is operational.
If the established tunnel is a one-way tunnel, another one-way tunnel is established from the pool node side. If the logical port to which the first VM is attached is mapped to the port of the managed switching element, the logical controller further sends a universal tunnel flow instruction to the pool node. Based on the universal tunnel flow instruction, the chassis controller that interfaces to the pool node instructs the pool node to create a tunnel to the managed switching element.
In the case of a tunnel between a managed edge switching element and a managed switching element that provides L3 gateway service, a logical switching element containing some VMs of the user is provided and the logical router provides L3 gateway service. It is assumed that it is realized in the transmission node. A logical patch port is created in a logical switching element to link a logical router to a logical switching element. In some embodiments, the order in which the logical patches are created and the VMs are delivered does not make a difference to the tunnel creation. By creating a logical patch port, the logical controller has all the managed switching elements that implement the logical switching element (ie, all managed switching elements each having at least one port to which the logical port of the logical switching element is mapped). ) Is connected to the chassis controller (or physical controller) to send a universal tunnel flow command. Each chassis controller for each of these managed switching elements commands the managed switching elements to create a tunnel to the transmission node. Each of the managed switching elements creates a tunnel to the transmission node, resulting in as many tunnels as there are managed switching elements that implement the logical switching element.
When these tunnels are unidirectional, the transmission node is for creating tunnels to each of the managed switching elements that implement the logical switching elements. When a logical patch port is created and connected to a logical router, the logical switching element pushes universal tunnel flow instructions to the transmission node. The chassis controller interfaced to the transmission node instructs the transmission node to create a tunnel, and the transmission node creates a tunnel to the managed switching element.
In some embodiments, every machine connected to one of the managed switching elements and every machine connected to another managed switching element uses the same logical switching element or the logical ports of two different switching elements. A tunnel established between two managed switching elements, whether or not, can be used to exchange data between these two machines. It is an example of tunneling that allows different users managing different sets of LDPs to share a managed switching element while they are isolated.
Creation of a tunnel between two managed switching elements with a logical port and another logical port connected to the L2 gateway service begins when the logical port is connected to the L2 gateway service. Upon connection, the logical controller sends universal tunnel flow instructions to all managed switching elements that implement the other logical ports of the logical switching element. Based on the instructions, a tunnel is established from these managed switching elements to the managed switching elements that implement the logical ports connected to the L2 gateway service.
III. Electronic system Many of the features and applications described above are implemented as software processing designated as an instruction set recorded on a computer-readable storage medium (also called a computer-readable medium). When these instructions are executed by one or more processing units (eg, one or more processors, processor cores or other processing units), the processing units are made to perform the operations indicated by the instructions. Examples of computer-readable media include, but are not limited to, CD-ROMs, flash drives, RAM chips, hard drives, EPROMs, and the like. Computer-readable media do not include carrier and electronic signals that travel wirelessly or over a wired connection.
As used herein, the term "software" is intended to include firmware or applications stored in magnetic storage that are in read-only memory that can be read into memory for processing by a processor. Also, in some embodiments, many software inventions can be realized as subparts of a larger program, leaving separate software inventions. In some embodiments, numerous software inventions can also be realized as independent programs. Finally, any combination of independent programs that together realize the inventions of the software described herein are within the scope of the invention. In some embodiments, when the software program is installed to operate on one or more electronic systems, it specifies implementations of one or more specific machines that perform and perform the operation of the software program.
FIG. 23 conceptually shows an electronic system 2300 used to realize some embodiments of the present invention. An electronic system 2300 may be used to run any of the control applications, virtualization applications or operating system applications described above. The electronic system 2300 may be a computer (eg, a desktop computer, a personal computer, a tablet computer, a server computer, a mainframe, a blade computer, etc.), a telephone, a PDA, or some other type of electronic device. Such electronic systems include interfaces for various computer-readable media and various other computer-readable media. The electronic system 2300 includes a bus 2305, a processing unit 2310, a system memory 2325, a read-only memory 2330, a fixed storage device 2335, an input device 2340, and an output device 2345.
Bus 2305 collectively represents a chipset bus that communicatively connects all systems, peripherals and many internal devices of the electronic system 2300. For example, the bus 2305 communicatively connects the processing unit 2310 with the read-only memory 2330, the system memory 2325, and the fixed storage device 2335.
From these various memory units, the processing unit 2310 searches for instructions to be executed and data to be processed in order to execute the processing of the present invention. The processing unit may be a single processor or, in different embodiments, a multi-core processor.
The read-only memory (ROM) 2330 stores static data and instructions required by the processing unit 2310 of the electronic system and other modules. On the other hand, the fixed storage device 2335 is a read / write memory element. This device is a non-volatile memory unit that stores instructions and data even when the electronic system 2300 is off. Some embodiments of the present invention use a large capacity storage device (eg, a magnetic disk or optical disk and a corresponding disk drive) as the fixed storage device 2335.
In another embodiment, a removable storage device (eg, floppy disk, flash drive, etc.) is used as the fixed storage device. Like the fixed storage device 2335, the system memory 2325 is a read / write memory element. However, unlike the storage device 2335, the system memory is a volatile read / write memory element such as a random access memory. System memory stores some of the instructions and data that the processor needs at run time. In some embodiments, the processing of the present invention is stored in system memory 2325, fixed storage 2335 and / or read-only memory 2330. From these various memory units, the processing unit 2310 searches for instructions to be executed and data to be processed to execute the processing of some embodiments.
Bus 2305 also connects to input device 2340 and output device 2345. The input device allows the user to communicate information to the electronic system and select commands. The input device 2340 includes an alphanumeric keyboard and a pointing device (also referred to as a "cursor control device"). The output device 2345 displays an image generated by the electronic system. Output devices include printers and display devices such as cathode ray tubes (CRTs) or liquid crystal displays (LCDs). Some embodiments include devices such as touch screens that function as both input and output devices.
Finally, as shown in FIG. 23, the bus 2305 further connects the electronic system 2300 to the network 2365 via a network adapter (not shown). As such, the computer may be part of a network of networks such as a computer network (eg, a local area network ("LAN"), a wide area network ("WAN") or an intranet, or the Internet. Any or all of the 2300 components may be used in connection with the present invention.
Some embodiments are electronic such as microprocessors, storage devices and memory that store computer program instructions on a machine-readable or computer-readable medium (also referred to as a computer-readable storage medium, machine-readable medium or machine-readable storage medium). Includes parts. Some examples of such computer-readable media include RAM, ROM, read-only compact discs (CD-ROMs), writable compact discs (CD-Rs), rewritable compact discs (CD-RWs), and more. Read-only general-purpose discs (eg, DVD-ROM, dual-layer DVD-ROM), various writable / rewritable DVDs (eg, DVD-RAM, DVD-RW, DVD + RW, etc.), flash memory (eg, SD). Cards, mini SD cards, micro SD cards, etc.), magnetic hard drives and / or solid hard drives, read-only and writable Blu-ray (Blu-Ray®) discs, ultra-high density optical discs, and any other optical medium. Alternatively, there are magnetic media and floppy discs. A computer-readable medium may contain a computer program that can be executed by at least one processing unit and includes a set of instructions that perform various operations. Examples of computer programs or computer code include, for example, machine code generated by a compiler, files containing high-level code executed by a computer, electronic component or microprocessor using an interpreter.
The above description has primarily referred to microprocessors or multi-core processors running software, but some embodiments may be one or more, such as application specific integrated circuits (ASICs) or field programmable gate arrays (FPGAs). It is executed by an integrated circuit. In some embodiments, such an integrated circuit executes instructions stored in the circuit itself.
As used herein, the terms "computer," "server," "processor," and "memory" all refer to electronic devices or other technical devices. These terms exclude users or groups of users. For the purposes of this specification, the term "display" means to display on an electronic device. As used herein, the terms "one computer-readable medium," "multiple computer-readable media," and "machine-readable medium" are tangible objects that store information in a computer-readable format. Completely limited to physical objects. These terms exclude any radio signal, wired download signal and any other temporary signal.
Although the present invention has been described with reference to many specific details, it will be appreciated by those skilled in the art that the invention may be practiced in other particular embodiments without departing from the gist of the invention. Also, many drawings (including FIGS. 20 and 21) conceptually illustrate the process. Certain actions of these processes do not have to be performed in the exact order shown and described. A particular action may not be performed in a series of successive actions, and different specific actions may be performed in different embodiments. Further processing can be implemented using some sub-processing or as part of a larger macro processing.
Also, some embodiments have been described above in which the user provides an LDP set for LCP data. However, in other embodiments, the user may provide an LDP set for LFP data. Also, some embodiments have been described above in which the controller instance provides PCP data to the switching element to manage the switching element. However, in other embodiments, the controller instance may provide physical forwarding plane data to the switching element. In such an embodiment, the relational database data structure stores physical forwarding plane data and the virtualization application produces such data.
Also, in some of the above examples, the user specifies one or more logical switching elements. In some embodiments, the user may provide a physical switching element configuration along with such a logical switching element configuration. Also, in some embodiments, controller instances individually formed by several application layers running on an arithmetic unit will be described, such instances being one or more layers of their operation. It will be appreciated by those skilled in the art that it is formed by a dedicated arithmetic unit or other machine in some embodiments to be performed.
Also, some of the above examples show that LDPS is associated with one user. It will then be appreciated by those skilled in the art that the user may be associated with one or more sets of LDP sets in some embodiments. That is, the relationship between the LDPS and the user is not always one-to-one because the user may be associated with a large number of LDP sets. Therefore, it will be appreciated by those skilled in the art that the present invention is not limited by the above exemplary details.
26 sheets
Sheet 1 Sheet 2 Sheet 3 Sheet 4 Sheet 5 Sheet 6 Sheet 7 Sheet 8 Sheet 9 Sheet 10 Sheet 11 Sheet 12 Sheet 13 Sheet 14 Sheet 15 Sheet 16 Sheet 17 Sheet 18 Sheet 19 Sheet 20 Sheet 21 Sheet 22 Sheet 23 Sheet 24 Sheet 25 Sheet 26
Every citation, both waysCites: the store holds 1 of 2
| Document | Relation | Office | Cited during |
|---|---|---|---|
| US12101244B1 | Cited by | United States of America | Applicant |
| JP2019503619A | Cited by | Japan | Search report |
| US12184450B2 | Cited by | United States of America | Applicant |
| US12120088B2 | Cited by | United States of America | Applicant |
| JP2022184934A | Cited by | Japan | Search report |
| US12058102B2 | Cited by | United States of America | Applicant |
| US12301382B2 | Cited by | United States of America | Applicant |
| US12231398B2 | Cited by | United States of America | Applicant |
| JP2015133545A | Cited by | Japan | Search report |
| US12199833B2 | Cited by | United States of America | Applicant |
| KR20180103975A | Cited by | Republic of Korea | Search report |
| US12177124B2 | Cited by | United States of America | Applicant |
| JP2015133545A | Cited by | Japan | Search report |
| US12261746B2 | Cited by | United States of America | Applicant |
| US12197971B2 | Cited by | United States of America | Applicant |
| US12182630B2 | Cited by | United States of America | Applicant |
| US11902245B2 | Cited by | United States of America | Applicant |
| US12267212B2 | Cited by | United States of America | Applicant |
| WO2011013805A1 | Cites | World Intellectual Property Organization (WIPO) | Examiner |
| JPN6015022380; '統合制御プレーン向けネットワークOSの提案とそのOpenFlowコントローラへの適用 A network OS f' 電子情報通信学会技術研究報告 Vol.109 No.448 IEICE Technical Report , 20100225 | Non-patent | – | Search report |
| JPN6015022381; データセンターネットワークにおけるネットワークアプライアンス機能配備の一検討 A Study on Deployment , 20091113 | Non-patent | – | Search report |
| JPN6015022379; '将来のクラウド基盤技術を支える研究開発 (IT/ネットワーク統合制御プレーン向けネットワークOS)' NEC技報 第63巻 第2号 NEC TECHNICAL JOURNAL , 20100423 | Non-patent | – | Examiner |
131 members in 11 offices
Priority claims44
| Document | Office | Kind | Date |
|---|---|---|---|
| 201161551425 | United States of America | P | |
| 201161551425 | United States of America | P | |
| 201161551427 | United States of America | P | |
| 201161551427 | United States of America | P | |
| 201161577085 | United States of America | P | |
| 201161577085 | United States of America | P | |
| 201261595027 | United States of America | P | |
| 201261595027 | United States of America | P | |
| 201261599941 | United States of America | P | |
| 201261599941 | United States of America | P | |
| 201261610135 | United States of America | P | |
| 201261610135 | United States of America | P | |
| 201261647516 | United States of America | P | |
| 201261647516 | United States of America | P | |
| 201213589077 | United States of America | A | |
| 201213589077 | United States of America | A | |
| 201213589078 | United States of America | A | |
| 201213589078 | United States of America | A | |
| 201261684693 | United States of America | P | |
| 201261684693 | United States of America | P | |
| 2012062005 | United States of America | W | |
| 2012062005 | United States of America | W | |
| 13589077 | – | – | – |
| 13589078 | – | – | – |
| 61551425 | – | – | – |
| 61551427 | – | – | – |
| 61577085 | – | – | – |
| 61595027 | – | – | – |
| 61599941 | – | – | – |
| 61610135 | – | – | – |
| 61647516 | – | – | – |
| 61684693 | – | – | – |
| US201161551425P | – | – | – |
| US201161551427P | – | – | – |
| US201161577085P | – | – | – |
| US2012062005 | – | – | – |
| US201213589077 | – | – | – |
| US201213589078 | – | – | – |
| US201261595027P | – | – | – |
| US201261599941P | – | – | – |
| US201261610135P | – | – | – |
| US201261647516P | – | – | – |
| US201261684693P | – | – | – |
| WO2012US62005 | – | – | – |
Members131
| Document | Office | Kind | |
|---|---|---|---|
| US2013103817A1 | United States of America | A1 | |
| US2013103818A1 | United States of America | A1 | |
| CA2849930A1 | Canada | A1 | |
| CA2965958A1 | Canada | A1 | |
| CA3047447A1 | Canada | A1 | |
| WO2013063329A1 | World Intellectual Property Organization (WIPO) | A1 | |
| WO2013063330A1 | World Intellectual Property Organization (WIPO) | A1 | |
| WO2013063332A1 | World Intellectual Property Organization (WIPO) | A1 | |
| US2013114466A1 | United States of America | A1 | |
| US2013117428A1 | United States of America | A1 | |
| US2013117429A1 | United States of America | A1 | |
| US2013208623A1 | United States of America | A1 | |
| US2013211549A1 | United States of America | A1 | |
| US2013212148A1 | United States of America | A1 | |
| US2013212235A1 | United States of America | A1 | |
| US2013212243A1 | United States of America | A1 | |
| US2013212244A1 | United States of America | A1 | |
| US2013212245A1 | United States of America | A1 | |
| US2013212246A1 | United States of America | A1 | |
| US2013219037A1 | United States of America | A1 | |
| US2013219078A1 | United States of America | A1 | |
| WO2013158917A2 | World Intellectual Property Organization (WIPO) | A2 | |
| WO2013158918A1 | World Intellectual Property Organization (WIPO) | A1 | |
| WO2013158920A1 | World Intellectual Property Organization (WIPO) | A1 | |
| WO2013158918A4 | World Intellectual Property Organization (WIPO) | A4 | |
| AU2012328697A1 | Australia | A1 | |
| AU2012328699A1 | Australia | A1 | |
| AU2013249151A1 | Australia | A1 | |
| AU2013249154A1 | Australia | A1 | |
| IL231910A0 | Israel | A0 | |
| IL231910D0 | Israel | D0 | |
| KR20140066781A | Republic of Korea | A | |
| WO2013158917A3 | World Intellectual Property Organization (WIPO) | A3 | |
| CN103891209A | China | A | |
| EP2748706A1 | European Patent Office (EPO) | A1 | |
| EP2748977A1 | European Patent Office (EPO) | A1 | |
| EP2748990A1 | European Patent Office (EPO) | A1 | |
| EP2748993A2 | European Patent Office (EPO) | A2 | |
| EP2748994A1 | European Patent Office (EPO) | A1 | |
| US2014247753A1 | United States of America | A1 | |
| AU2013249152A1 | Australia | A1 | |
| CN104081734A | China | A | |
| CN104170334A | China | A | |
| US2014348161A1 | United States of America | A1 | |
| US2014351432A1 | United States of America | A1 | |
| JP2014535213AThis record | Japan | A | |
| JP2015501109A | Japan | A | |
| JP2015507448A | Japan | A | |
| IN2272CHN2014A | India | A | |
| EP2748977A4 | European Patent Office (EPO) | A4 | |
| EP2748990A4 | European Patent Office (EPO) | A4 | |
| AU2012328697B2 | Australia | B2 | |
| AU2012328697B9 | Australia | B9 | |
| EP2748993B1 | European Patent Office (EPO) | B1 | |
| US9137107B2 | United States of America | B2 | |
| US9154433B2 | United States of America | B2 | |
| RU2014115498A | Russian Federation | A | |
| US9178833B2 | United States of America | B2 | |
| US9203701B2 | United States of America | B2 | |
| AU2013249151B2 | Australia | B2 | |
| AU2013249154B2 | Australia | B2 | |
| AU2015258164A1 | Australia | A1 | |
| EP2955886A1 | European Patent Office (EPO) | A1 | |
| JP5833246B2 | Japan | B2 | |
| US9231882B2 | United States of America | B2 | |
| US9246833B2 | United States of America | B2 | |
| JP5849162B2 | Japan | B2 | |
| US9253109B2 | United States of America | B2 | |
| JP5883946B2 | Japan | B2 | |
| US9288104B2 | United States of America | B2 | |
| US9300593B2 | United States of America | B2 | |
| AU2012328699B2 | Australia | B2 | |
| US9306843B2 | United States of America | B2 | |
| US9306864B2 | United States of America | B2 | |
| US9319336B2 | United States of America | B2 | |
| US9319337B2 | United States of America | B2 | |
| US9319338B2 | United States of America | B2 | |
| AU2013249152B2 | Australia | B2 | |
| JP2016067008A | Japan | A | |
| US9331937B2 | United States of America | B2 | |
| KR101615691B1 | Republic of Korea | B1 | |
| JP2016076959A | Japan | A | |
| KR20160052744A | Republic of Korea | A | |
| US2016197774A1 | United States of America | A1 | |
| US9407566B2 | United States of America | B2 | |
| RU2595540C2 | Russian Federation | C2 | |
| AU2016208326A1 | Australia | A1 | |
| US2016308785A1 | United States of America | A1 | |
| KR101692890B1 | Republic of Korea | B1 | |
| US9602421B2 | United States of America | B2 | |
| AU2015258164B2 | Australia | B2 | |
| RU2595540C9 | Russian Federation | C9 | |
| CN103891209B | China | B | |
| JP6147319B2 | Japan | B2 | |
| CA2849930C | Canada | C | |
| JP6162194B2 | Japan | B2 | |
| CN106971232A | China | A | |
| AU2017204764A1 | Australia | A1 | |
| CN107104894A | China | A | |
| IL231910A | Israel | A |
19 legal events, as the office reported them to INPADOC
Over the term
Point at a mark for the eventEvents
| Event | Code | |
|---|---|---|
| Receipt of annual feesJAPANESE INTERMEDIATE CODE: R250R250 | R250 | |
| Written notification of registration of transferJAPANESE INTERMEDIATE CODE: R350R350 | R350 | |
| Request for change of ownership or part of ownershipJAPANESE INTERMEDIATE CODE: R313111S111 | S111 | |
| Receipt of annual feesJAPANESE INTERMEDIATE CODE: R250R250 | R250 | |
| Written notification of registration of transferJAPANESE INTERMEDIATE CODE: R350R350 | R350 | |
| Written request for registration of change of domicileJAPANESE INTERMEDIATE CODE: R313531S531 | S531 | |
| Receipt of annual feesJAPANESE INTERMEDIATE CODE: R250R250 | R250 | |
| Receipt of annual feesJAPANESE INTERMEDIATE CODE: R250R250 | R250 | |
| Receipt of annual feesJAPANESE INTERMEDIATE CODE: R250R250 | R250 | |
| Receipt of annual feesJAPANESE INTERMEDIATE CODE: R250R250 | R250 | |
| Receipt of annual feesJAPANESE INTERMEDIATE CODE: R250R250 | R250 | |
| Receipt of annual feesJAPANESE INTERMEDIATE CODE: R250R250 | R250 | |
| Certificate of patent or registration of utility modelJAPANESE INTERMEDIATE CODE: R150R150 | R150 | |
| First payment of annual fees (during grant procedure)JAPANESE INTERMEDIATE CODE: A61A61 | A61 | |
| Written decision to grant a patent or to grant a registration (utility model)JAPANESE INTERMEDIATE CODE: A01A01 | A01 | |
| Decision of grant or rejection writtenTRDD | TRDD | |
| Request for written amendment filedJAPANESE INTERMEDIATE CODE: A523A521 | A521 | |
| Notification of reasons for refusalJAPANESE INTERMEDIATE CODE: A131A131 | A131 | |
| Report on retrievalJAPANESE INTERMEDIATE CODE: A971007A977 | A977 |
Numbers
- Publication
- 2014535213
- Publication, DOCDB
- 2014535213
- Publication, EPODOC
- JP2014535213
- Application
- 2014539024
- Application, DOCDB
- 2014539024
- Application, EPODOC
- JP20140539024
Titles2
- Japanese
- ユニバーサルフローを変換するためのシャーシコントローラ
- English
- Chassis controller for converting universal flow
Classification
- CPC, 12
- G06Q10/00
- H04L45/64
- H04L45/02
- G06F15/177
- H04L41/02
- H04L41/0226
- H04L41/042
- H04L41/50
- H04L45/38
- H04L45/42
- H04L45/66
- H04L47/50
- IPC, 5
- H04L45 42
- G06F13 00
- H04L45 02
- H04L12 717
- H04L12 70
Designated states5
- Regional, 4
- Zimbabwe
- Turkmenistan
- Türkiye
- Togo
- National, 1
- Saint Vincent and the Grenadines