A method to extend the physical reach of an infiniband network
41 claims: 19 independent, 22 dependent
- 1インフィニバンド・パケットを、長距離接続を介して搬送する方法であって、 インフィニバンド・パケットを送信インフィニバンド・インターフェースから受信すること、 受信したインフィニバンド・パケットを 長距離接続(WAN) のプロトコル内でカプセル化すること、 前記カプセル化されたインフィニバンド・パケットを、前記送信インフィニバンド・インターフェースと受信インフィニバンド・インターフェースとを接続している前 記W A Nを 介して送信ユニット、すなわち前 記W A Nに よって送信すること、 前記カプセル化されたインフィニバンド・パケットを受信ユニットで受信すること、 前記 受信ユニットで前記カプセル化を除去しかつ前記インフィニバンド・パケットを回復することによって受信されたインフィニバンド・パケットをカプセル化解除すること、 128Ki B(=210byte) を越えてバルク・バッファ・メモリで回復したインフィニバンド・パケットをバッファリングすること、 不 十分なキャパシティゆえに受信バルク・バッファ・メモリによっ てパ ケット が廃棄されないように、前記WANを介してインフィニバンド・パケットのフローをクレジット管理ユニットによって統制する こと、 を含み インフィニバンド物理リンク状態機械とインフィニバンド型のフロー制御 を、 前 記W A Nを 介して維持する方法。
- 2前記プロトコルが、OSI7レイヤ参照モデルであり、レイヤ1、レイヤ2、レイヤ3、レイヤ4からなるグループから選択される、請求項1に記載の方法。
- 3前記 カプセル化されたインフィニバンド ・パケット を送信することが、インフィニバンド・リンク距離を、前記WANを介して、約100kmより大きい距離に延長することをさらに含む、請求項1または2に記載の方法。
- 4リンク距離を増大することが、前記リンク上で告知され るク レジットを、 仮想レーン( VL ) 当た り1 2 8 KiBを越えて増大することをさらに含む、請求項3に記載の方法。
- 5使 用可能な 前記 クレジットを増大することが、告知されるクレジット・ブロック当たりのバイト数を増大することを含む、請求項4に記載の方法。
- 6使 用可能な 前記 クレジットを増大することが、告知当たりのクレジット・ブロックの数を増大することを含む、請求項4に記載の方法。
- 7使 用可能な 前記 クレジットを増大することが、クレジット・ブロックの数および各告知内のブロック当たりのバイト数を共に増大することを含む、請求項4に記載の方法。
- 8不十分なキャパシティゆえに受信バルク・バッファ・メモリによってパケットが廃棄されないように、10ギガビット/秒を超える最大入口帯域幅を有し、同時に10ギガビット/秒を超える出口帯域幅を維持する ことを含む、請求項1から7のいずれか一項に記載の方法。
- 9前記インフィニバンド物理リンク状態機械を維持することが、非インフィニバンド・パケットを前記WAN全体にわたって交換することをさらに含み、非インフィニバンド・パケットを交換することにより、エンド・トゥ・エンドの経路が前記WAN内に存在することが立証され、パケットの前記交換が、PPP LCPパケット、イーサネット(登録商標)ARP交換、TCPセッション初期化、ATM SVCの確立からなるグループから選択される、請求項1から8のいずれか一項に記載の方法。
- 10インフィニバンド型のフロー制御を維持することが、前記WANポート上で受信されたパケットを、128KiBを越えるバッファ・メモリ内でバッファ リング することをさらに含む、請求項1から9のいずれか一項に記載の方法。
- 11論理回路からなる、インフィニバンド・パケットを搬送する装置であって、 インフィニバンド経路指定およびQOSコンポーネントに結合されたインフィニバンド・インターフェースを備え、 前記インフィニバンド経路指定およびQOSブロックのインフィニバンドからWANへの経路が、カプセル化/カプセル化解除コンポーネント(ENCAP)に結合され、 前記ENCAPコンポーネントの インフィニバンド からWANへの経路が、WANインターフェースに結合され、 前記WANインターフェースのWANから インフィニバンド への経路が、ENCAPコンポーネントに結合され、 前記ENCAPコンポーネントのWANから インフィニバンド への経路が、バルク・バッファ・メモリに結合され、 前記バルク・バッファ・メモリが、インフィニバンド・インターフェースの前記WANから インフィニバンド への経路に結合され、 クレジット管理ユニットが、前記WANのためにクレジットを生成し、前記インフィニバンド・インターフェース上にバック・プレッシャを生成し、 前記ENCAPコンポーネントが、クレジット・データをカプセル化およびカプセル化解除するために、前記クレジット管理ユニットに結合され、 管理ブロックが、インフィニバンド・サブネット管理エージェント、WANエンド・トゥ・エンド・ネゴシエーション、および管理サービスを提供する装置。
- 12各方向で同時に、インフィニバンド・パケットの毎秒約1ギガバイトの転送速度を維持することができる、請求項11に記載の装置。
- 13前記インフィニバンド・インターフェースが、WANクロック・ドメインからインフィニバンド・クロック・ドメインに移行するために、追加のフロー制御バッファ用ユニットを含む、請求項11または12に記載の装置。
- 14前記ENCAPコンポーネントが、IPv6、IPv6内のUDP、IPv6内のDCCP、ATM AAL5、またはGFPのいずれかを含む複数のネットワークをサポートすることができる、請求項11、12、または13に記載の装置。
- 15前記WANインターフェースが、 SONET/SDH、10GBASE-R、インフィニバンド、10GBASE-Wのいずれかを含めて、複数のネットワーク・フォーマットをサポートすることができるフレーマ・ユニットと、 SONET/SDH、10GBASE-R、またはインフィニバンドのいずれかをサポートすることができる光サブシステムと をさらに備える、請求項11から14のいずれか一項に記載の装置。
- 16前記光サブシステムがさらに、単独で、またはSONET/SDHマルチプレクサ、光再生器、パケット・ルータ、セル・スイッチなど他の機器と結合されたとき、IBTAインフィニバンド・アーキテクチャのリリース1.2によって指定されている距離より大きい距離に到達することができる、請求項15に記載の装置。
- 17さらに複数のFIFO構造を備え、 前記バルク・バッファ・メモリが、前記パケットが受信された順序と異なる順序で、前記複数のFIFO構造からパケットを取り出すことができる、請求項11から16のいずれか一項に記載の装置。
- 18前記クレジット管理ユニットが 、ク レジット・ブロック・サイズを増大すること、および/または告知当たりのブロックの数を増大することにより、前記インフィニバンド仕様によって規定されているものより多くのクレジットを告知する、請求項11から17のいずれか一項に記載の装置。
- 19前 記管理ブロックが、 汎用プロセッサと、 前記WANインターフェースおよび 前記インフィニバンド・ インターフェースの双方でパケットを送信および受信するための機構と をさらに備える、請求項11から18のいずれか一項に記載の装置。
- 20前記バルク・バッファ・メモリが、 複数のDDR2メモリ・モジュール(DIMMS)をさらに備え、 制御論理が、前記DDR2メモリ内で複数のFIFO構造を維持し、 各FIFO構造が、WANからインフィニバンドへのVLをバッファ リング するために使用され、 確実に、前記インフィニバンド・インターフェースでの輻輳によりパケットが廃棄されないようにするために、前記メモリからの前記パケットの流れが調節される、請求項11から19のいずれか一項に記載の装置。
- 21各方向で同時に、インフィニバンド・パケットの毎秒1ギガバイトの最大転送速度を維持するための、請求項11に記載の装置であって、 WANクロック・ドメインから インフィニバンド・ クロック・ドメインに移行するための、追加のフロー制御バッファ用ユニットをさらに備え、 前記ENCAPコンポーネントが、IPv6、IPv6内のUDP、IPv6内のDCCP、ATM AAL5、またはGFPのいずれかを含む複数のネットワークをサポートすることができ、 SONET/SDH、10GBASE-R、インフィニバンド、10GBASE-Wのいずれかを含めて、複数のネットワーク・フォーマットをサポートすることができるフレーマ・ユニットと、 SONET/SDH、10GBASE-R、またはインフィニバンドのいずれかをサポートすることができる、結合された光サブシステムとをさらに備え、 前記光サブシステムがさらに、単独で、またはSONET/SDHマルチプレクサ、光再生器、パケット・ルータ、セル・スイッチなど他の機器と結合されたとき、IBTAインフィニバンド・アーキテクチャのリリース1.2によって指定されている距離より大きい距離に到達することができ、 前記バルク・バッファ・メモリが、前記パケットが受信された順序と異なる順序で、前記複数のFIFO構造からパケットを取り出すことができ、前記バルク・バッファ・メモリが、 複数のDDR2メモリ・モジュール(DIMMS)をさらに備え、 制御論理が、前記DDR2メモリ内で複数のFIFO構造を維持し、 各FIFO構造が、WANからインフィニバンドへのVLをバッファ リング するために使用され、 確実に、前記インフィニバンド・インターフェースでの輻輳によりパケットが廃棄されないようにするために、前記メモリからの前記パケットの流れが調節され、 前記クレジット管理が 、ク レジット・ブロック・サイズを増大すること、および/または告知当たりのブロックの数を増大することにより 、イ ンフィニバンド仕様によって規定されているものより多くのクレジットを告知し、 前記管理ブロックが、 汎用プロセッサと、 前記WANインターフェースおよび 前記インフィニバンド・ インターフェースの双方でパケットを送信および受信するための機構と をさらに備える装置。
- 22前記ENCAPコンポーネントが、ヌル・カプセル化を実行し、インフィニバンド・パケットを無変更で放出する、請求項11から21のいずれか一項に記載の装置。
- 23前記バルク・バッファ・メモリが、 複数のSRAMメモリ・チップをさらに備え、 制御論理が、前記 SRAMメモリ・チップ 内で複数のFIFO構造を維持し、 各FIFO構造が、WANからインフィニバンドへのVLをバッファ リング するために使用され、 確実に、前記インフィニバンド・インターフェースでの輻輳によりパケットが廃棄されないようにするために、前記メモリからの前記パケットの流れが調節される、請求項11から22のいずれか一項に記載の装置。
- 24前記SRAMメモリ・チップがQDR2 SRAMである、請求項23に記載の装置。
- 25前記バルク・バッファ・メモリが、複数のSRAMメモリ・チップを備え、 制御論理が、前記 SRAMメモリ・チップ 内で複数のFIFO構造を維持し、 各FIFO構造が、WANからインフィニバンドへのVLをバッファ リング するために使用され、 確実に、前記インフィニバンド・インターフェースでの輻輳によりパケットが廃棄されないようにするために、前記メモリからの前記パケットの流れが調節される、請求項21に記載の装置。
- 26前記インフィニバンド・パケットが、IPv6パケット のペ イロード構造内に配置される、請求項11から25のいずれか一項に記載の装置。
- 27前記クレジット・データが、IPv6ヘッダ内の拡張ヘッダ内で符号化される、請求項11から26のいずれか一項に記載の装置。
- 28前記ENCAPコンポーネントが、IEEE802.3ae clause49によって規定されている66/64bコーディング・スキームに適合する形で、インフィニバンド・パケットをフレーム化し、 前記ENCAPコンポーネントが、前記clause49に適合するフレーミングを除去し、前記元のインフィニバンド・パケットを回復することができる、請求項11から 27 のいずれか一項に記載の装置。
- 29前記クレジット・データが、66/64bコードにおいて、順序付けられたセットで符号化される、請求項 28 に記載の装置。
- 30前記インフィニバンド・パケットが、IPv6パケットまたはIPv4パケット内で担持されるUDPまたはDCCPデータグラム のペ イロード構造内に配置される、請求項11から 29 のいずれか一項に記載の装置。
- 31前記インフィニバンド・パケットが、ATMアダプテーション・レイヤ5(AAL5)に従って、ATMセルにセグメント化される、請求項11から 29 のいずれか一項に記載の装置。
- 32前記インフィニバンド・パケットが、ジェネリック・フレーミング・プロトコル・パケット のペ イロード構造内に配置され、SONET/SDHフレーム内に配置される、請求項11から 29 のいずれか一項に記載の装置。
- 33前記クレジット・データが、前記カプセル化 のペ イロード構造内で符号化される、請求項11から 32 のいずれか一項に記載の装置。
- 34第1のデバイスに結合された第1のインフィニバンド・ファブリックと、 第2のデバイスに結合された第1のデバイスと、 第2のインフィニバンド・ファブリックに結合された第2のデバイスとを備え、 前記第1および第2のデバイスが、さらに、 インフィニバンド・パケットを 長距離接続(WAN) のネットワーク・プロトコル内にカプセル化およびカプセル化解除するための論理回路と、 前記インフィニバンド・パケットを 、128KiBを超える受信バルク・バッファ・メモリ内に バッファ リング するための論理回路と、 10kmを超える長距離接続のインフィニバンド・パケットの流れを調節して、不十分なキャパシティゆえに受信バルク・バッファ・メモリによってパケットが廃棄されないように、前記WANを介してインフィニバンド・パケットのフローを調節するための論理回路と、 前記カプセル化されたインフィニバンド・パケットを搬送するネットワーク・インターフェースとで構成されるシステム。
- 35前記第1のデバイスと第2のデバイスが、延長されたWANネットワークを介してさらに間接的に結合され、前記延長されたWANネットワークが、SONET/SDHマルチプレクサ、光再生器、パケット・ルータ、セル・スイッチのうちの1つまたは複数を備える、請求項 34 に記載のシステム。
- 36E NCAPコンポーネントへのパケットの流速が、 各エンド・ デバイスによって、ネットワーク内の条件および管理構成に基づいて、可能な最大速度以下に制限されてもよい、請求項 34 または 35 に記載のシステム。
- 37前記2つのデバイス間の、パケットまたはセルによってスイッチまたは経路指定されるネットワーク出口をさらに備え、 2つより多いデバイスをこのネットワークに接続することができ、 各エンド・デバイスが、パケットをカプセル化し、複数の宛先デバイスにアドレッシングすることができる、請求項 34 、 35 、または 36 に記載のシステム。
- 38請求項21に記載の装置をさらに備える、請求項 34 から 37 のいずれか一項に記載のシステム。
- 39請求項25に記載の装置をさらに備える、請求項 34 から 37 のいずれか一項に記載のシステム。
- 40引き離された ローカル識別子( LID ) アドレス空間および異なるサブネット・プリフィックスを有する2つのインフィニバンド・ファブリックと、 前記デバイスに一体化されたパケット経路指定コンポーネントとをさらに備え、 論理回路が、 グローバル・ルート・ヘッダ( GRH ) 内 の宛 先 グローバル識別子( GID ) を調べることによって、所与のインフィニバンド・パケットの前記LIDアドレスを決定し、 論理回路が、前記GRHからの情報を使用して、前記インフィニバンド・パケットの前記LID、 サービスレベル( SL ) 、VL、または他のコンポーネントを置き換えることができる、請求項 34 から 37 のいずれか一項に記載のシステム。
- 41インフィニバンド・リンク距離を、前記WANを介して、10kmより大きい距離に延長するために、前記バルク・バッファ・メモリが128KiBを越える、請求項11から33のいずれか一項に記載の装置。
Independent claims41
30 paragraphs, as filed
The present invention relates to a method of extending the physical reach of an InfiniBand network beyond what is currently possible within the InfiniBand architecture, and in particular the InfiniBand packet itself, the InfiniBand architecture. Allows transport over networks that are not compatible with. This allows InfiniBand traffic to share the physical network with other standard protocols such as Internet Protocol Version 6 (IPv6) or Asynchronous Transfer Mode (ATM) cells. In addition, the very large flow control buffer in this device is combined with the use of a flow control credit scheme to prevent buffer overflows, which allows the present invention to combine large amounts of data over a wide area network. Allows you to be on the move within a WAN), while still ensuring that packets are not lost due to improper buffering resources at the receiver. To help ensure that packets do not drop in the WAN, the device also responds to back pressure and injects data into the WAN. It can include several quality of service (QOS) features that act to limit rate). The present invention also allows packets to be routed to allow multiple devices to be connected to a WAN, thus using a minimal number of devices to create an InfiniBand network. Allows extension to more than two physical locations. The processor contained within this device can handle management functions such as InfiniBand Subnet Management Agent and device management.
10 Gigabit InfiniBand is at most 128KiB per virtual lane (VL)<u style="single">(= 210byte)</u>It is known that it is only possible to reach about 10km due to the restrictions given in the InfiniBand architecture of the credits given. This constraint sets an upper limit on the amount of data that can be moved at one time. This is because standard InfiniBand transmitters do not transmit without available credits. Furthermore, it is known that limiting the amount of data that can be in motion to less than the bandwidth latency product of the network path directly limits the maximum data transfer rate that can be obtained. Has been done.
For example, a 10 Gigabit InfiniBand link with a round-trip latency of 130 microseconds has a bandwidth latency product of 128 KiB, which is the maximum amount of credit that can be given for a single VL within the InfiniBand link. Is.
In general, an InfiniBand link will have multiple inlet and exit VLs (up to 15), and the InfiniBand architecture buffers and flows each of these VLs individually, head-of-line blocking. And the flow control deadlock must be prevented. In some embodiments, the InfiniBand interface includes an additional flow control buffer unit for transitioning from the WAN clock domain to the InfiniBand clock domain.
Due to physical limitations, data travels through fiber optics at speeds slower than the speed of light. Considering a fiber as a conduit for carrying bits, it is clear that one long fiber can contain several megabits of data in motion. For example, if the speed of light in a particular fiber carrying a 10 Gigabit data stream is 5 ns / m and the fiber is 100 km long, then the fiber should contain 5 megabits of data in each direction. become. Many WAN paths also include additional latency from in-band devices such as playback devices, optical multiplexes, add / drop multiplexers, routers, and switches. This extra device adds additional latency and further extends the bandwidth latency product of the path.
InfiniBand electrical and optical signaling protocols are not compatible with and are not suitable for use in traditional WAN environments, as specified by the InfiniBand architecture. A typical WAN environment uses a synchronous optical network (SONET) over long-range optical fibers.
It is also desirable to perform routing on InfiniBand packets, as described in the InfiniBand Architecture, to facilitate management of the InfiniBand network. Routing means that each remote site maintains local control over parts of a larger InfiniBand network of these sites without imposing substantive policies on all other participants. To enable.<nplcit num="1"><text>InfiniBand Trade Association (2005), The InfiniBand Architecture release 1.2</text></nplcit><nplcit num="2"><text>Internet Engineering Task Force (1998) RFC2460-Internet Protocol, Version 6 (IPv6) Specification</text></nplcit><nplcit num="3"><text>Internet Engineering Task Force (19989, RFC2615-PPP over SONET / SDH</text></nplcit><nplcit num="4"><text>The ATM Forum (1994), ATM User-Network Interface Specification Version 3.1</text></nplcit><nplcit num="5"><text>International Telecommunication Union, ITU-T Recommendation I.432.1 General Characteristics</text></nplcit><nplcit num="6"><text>Open Systems Interconnection (OSI)-Basic Reference Model: The basic Model (1994), ISO7498-1: 1994</text></nplcit><nplcit num="7"><text>IEEE802.3ae clause49; 66 / 64b coding scheme</text></nplcit>
<p> A very well-managed and very large buffer when considering the physical limitations of the fiber, the need to have a buffer capacity that exceeds the capacity of the bandwidth latency product of the path, and the capabilities of multiple VLs within the InfiniBand architecture. It becomes clear that devices that require memory should extend the InfiniBand network to cross-continental distances.</p>
<p> Some of the features of this device are the more suitable number of credits advertised on the local short InfiniBand link, typically 8KiB, typically 512MiB per VL.<u style="single">(= 220byte)</u>Is to extend to. This is done using a first-in first-out buffer (FIFO), which is emptied when local InfiniBand credits are available and filled as incoming data arrives. This device periodically informs other remote devices of how much space is available in the FIFO via a credit notification packet for each active VL, which uses this information to use the FIFO. Never send more data than can be accepted. This is the same basic flow control mechanism (end-to-end credit information exchange) used with the InfiniBand architecture, but to handle gigabytes of buffers and to be more suitable for WAN environments. Scaled up to. In this way, InfiniBand flow control semantics are maintained over long distances to ensure that packets are not dropped due to congestion.</p><p> In other embodiments, the bulk buffer memory can retrieve packets from multiple FIFO structures in a different order than the packets were received.</p><p> If credit is craving on the local InfiniBand port, the FIFO is filled, but credit packets sent over the WAN send the transmitter before the FIFO can overflow. Will be stopped. Credit packets may be inserted into the IPv6 payload structure or embedded in the IPv6 extension header to improve efficiency.</p><p> Credit information or data is encoded in an ordered set in 66 / 64b code. Infiniband packets can be placed within the payload structure of UDP or DCCP datagrams carried within IPv6 or IPv4 packets.</p><p> To achieve compatibility with WAN standards, InfiniBand packets are encapsulated by this device within other protocols, such as IPv6 within packet over SONET (POS), for transmission over the WAN. To. As stated in the IBTA, the full-duplex independent transmit / receive data path is controlled by the link state machine. Infiniband physical link state machines can be maintained by exchanging non-infiniband packets across the WAN, thereby demonstrating that end-to-end routes exist within the WAN. This packet exchange is a PPP LCP packet (according to RFC1661), an Ethernet® ARP exchange (according to RFC826 and RFC2461 (IPv6 Neighbor Discovery)), a TCP session initialization (according to RFC793), and an ATM Forum Private Network Network Interface specification. ) Includes establishing an ATM SVC, or initiating any other form of session.</p><p> After the packet is encapsulated, it is sent over the WAN and the receiver performs a decapsulation step to remove any data added during encapsulation, thereby removing the original InfiniBand packet. Recover.</p><p> This encapsulation serves two purposes. The first purpose is to change the optical signaling format to something that can be originally carried over a WAN connection, such as SONET. This allows the device to be directly connected to SONET optical equipment that is part of a larger SONET topology and to persist to a single remote destination. SONET protocols, such as the Generic Framing Protocol (GFP), are designed for this type of encapsulation task. The encapsulation component can support multiple networks, including either IPv6, UDP in IPv6, DCCP in IPv6, ATM AAL5, or GFP.</p><p> It also allows the device to interface with intelligent devices in the WAN that can route individual packets or cells. It also allows InfiniBand traffic to share the WAN infrastructure with traffic from other sources by relying on the WAN to aggregate, route, and / or switch a large number of connections. To.</p><p> Protocols such as ATM Adaptation Layer 5 (AAL5), IPv6 over POS, and IPv6 over Ethernet® are designed to make this possible.</p><p> In order to establish and maintain an end-to-end route across the WAN, telecommunications equipment needs to exchange non-Infiniband packets in addition to encapsulated InfiniBand packets.</p><p> Numerous encapsulations include ATM AAL5, IPv6 over POS, IPv6 over Ethernet®, DCCP in IPv6, UDP in IPv6, IPv6 over GMPLS (generic multi-protocol label switching), GFP, etc. Is possible. Similarly, it can support a number of WAN signaling standards and speeds, including SONET, Ethernet LAN-PHY, and Ethernet WAN-PHY. A single device can support multiple encapsulation and signaling standards, allowing the user to choose which one to use during installation.</p><p> For shorter distances less than 10 km, encapsulation is eliminated and simply uses the optical signaling specified by the InfiniBand architecture in combination with a very large flow control buffer to create an InfiniBand architecture. Extends the reach of regular InfiniBand equipment while being fully compliant. In this case, the encapsulation step is reduced to null encapsulation, leaving the InfiniBand packet unchanged. The number and / or credit block size of credit blocks can be increased to extend the range beyond 10 km, while still adhering to otherwise unchanged InfiniBand communication protocols.</p><p> Multiple MFP: When this device is connected to an intelligent WAN that can be addressed using the encapsulation protocol, it is possible to have more than two devices communicate. This allows devices located at multiple physical sites to share the same WAN connection and the same device while extending and linking their InfiniBand network to a larger mesh.</p><p> In this mode of operation, the device is required to examine each incoming local InfiniBand packet, determine to which remote device it should be sent, and then form the proper encapsulation for delivery. .. It looks up the local identifier (LID) in the InfiniBand packet and either by using the switching infrastructure specified by the InfiniBand specification or by looking up the global identifier (GID) in the InfiniBand packet. This can be done by routing based on the longest prefix match of the subnet prefix.</p><p> Each device must also reserve a separate part of its flow control buffer for each possible remote device. This further increases the demand for buffer memory by N-1 times, where N is the number of devices in the mesh.</p><p> When multicast InfiniBand packets are received, the device either maps them to the appropriate WAN multicast addresses or performs packet replication and joins multiple copies of the packets to the multicast group. It will be sent to each remote device.</p><p> As specified by Release 1.2 of the InfiniBand architecture, InfiniBand routing behavior is to be transmitted over the local InfiniBand network using the Global Root Header (GRH) by this device. Requires converting a 128-bit IPv6 GID to a local InfiniBand route description, a 16-bit local identifier (LID), a 24-bit partition key, and a 4-bit service level.</p><p> When this device is used on an intelligent network rather than in a point-to-point configuration, quality of service issues within the intelligent network become important. This device only ensures that InfiniBand packets are never dropped due to inadequate buffering and does not provide any guarantee that the intelligent network will not drop packets due to internal congestion or otherwise.</p><p> The main means of minimizing packet loss in a network is by having this device control the amount of packets injected into the network. The device does this by inserting a delay between the packets as they are sent into the WAN towards the receiving unit.</p><p> The injection volume may be set by the user or dynamically controlled by dialogue and protocol between the device and the intelligent network. There are numerous protocols and methods for this type of dynamic control.</p><p> The second method is for the device to specially tag the packets so that the intelligent network can minimize the loss. This technique can be used with injection volume control.</p><p> The management software within the management block of this device is responsible for running any protocol and method that may be required to establish service quality assurance using general purpose processors throughout the WAN network.</p><p> 1 to 3 are data flow diagrams showing the routes that packets can take in the system. Each box with right-angled corners represents a buffer, conversion process, or decision point. Larger, rounded boxes represent related feature groups. Arrows indicate the direction of packet flow.</p>
Those skilled in the art understand that various standards and resources are inherent in the formalization of traditional digital data. Some of the standards and principles of operation referred to herein as known in the art can be found with reference to: -InfiniBand Trade Association (2005), The InfiniBand Architecture release 1.2 (also known as IBTA) -Internet Engineering Task Force (1998) RFC2460-Internet Protocol, Version 6 (IPv6) Specification -Internet Engineering Task Force (19989, RFC2615-PPP over SONET / SDH The ATM Forum (1994), ATM User-Network Interface Specification Version 3.1 International Telecommunication Union, ITU-T Recommendation I.432.1 General Characteristics -Open Systems Interconnection (OSI)-Basic Reference Model: The basic Model (1994), ISO7498-1: 1994 -IEEE802.3ae clause49; 66 / 64b coding scheme
The data flow in the prototype device will be described with reference to FIG. The device includes six main blocks: InfiniBand interface, management block, packet routing, encapsulation / decapsulation component (ENCAP), WAN interface, and bulk buffer memory. There are various techniques and techniques that can be used to implement each of these blocks. These blocks are identified as logical functions in the data flow diagram, and certain implementations choose to extend these functions between various physical blocks in order to achieve a more optimal implementation. be able to. The device can maintain a transfer rate of about 1 gigabyte per second for InfiniBand packets at the same time in each direction.
InfiniBand interface is local<u style="single">InfiniBand</u>Achieve LAN connection to the fabric. For clarity, the InfiniBand interface includes two small flow control buffers that handle the data transfer rate from other connected blocks.
The management block is a variety of high-level management and control protocols required by the various standards that this device can follow, such as<u style="single">InfiniBand</u>Subnet management agent, implementation of point-to-point (PPP) protocol for packet over SONET, ATM operation and maintenance cell (OAM), neighbor discovery caching / query for Ethernet® Implement. Typically, this block is implemented using some form of general purpose microprocessor in combination with specialized logic for any low latency or high frequency management packet, such as some type of OAM cell. become.
The packet routing block implements the functionality required by the multi-device InfiniBand routing and quality of service (QOS) features described above. It also sends WAN credit packets, as discussed in the context of distance extension. In addition, the packet routing block can identify packets to be delivered to the management block for special processing.
The encapsulation / decapsulation block implements the encapsulation process discussed in the context of protocol encapsulation above. In one embodiment, the protocol is an OSI7 layer reference model (as defined in ISO 7498-1: 1994), Layer 1 (Physical), Layer 2 (Data Link), Layer 3 (Network), Layer 4 ( Selected from a group consisting of (Transport).
This prototype shows several possible different schemes. The encapsulation block relies on additional data from the routing block to determine the exact form of encapsulation. Decapsulation is the original from the encapsulated data<u style="single">InfiniBand</u>Restore the packet. Some packets, if identified as management packets, may be routed to the management block and not sent through the decapsulation block.
A WAN interface is a generic interface to a WAN port. WAN interfaces are possible, including optical subsystems, but using electrical signaling, as shown here. The framer unit or function takes the packet data from the encapsulation block and formats it according to the selected WAN protocol. For example, the Ethernet® specification would refer to the framer as a combination of media access controller (MAC), physical coding sublayer (PCS), and physical media attachment (PMA). The framer also does the reverse and extracts the packet passed to the decapsulation block from the WAN interface.
Supported framing formats include SONET / SDH, 10GBASE-R, InfiniBand, 10GBASE-W, and the 66 / 64b coding scheme specified by IEEE 802.3ae clause 49-10GBASE-R.
Bulk buffer memory implements credit management with a description of distance extension. The exact nature of the underlying memory can vary from implementation to implementation. FIG. 2 describes a data flow within a preferred embodiment for a long-range configuration of the present invention. This embodiment of the invention includes a system-on-chip mounted within a field programmable gate array (FPGA), CX4 copper wire 4 x InfiniBand connector, SONET / Ethernet® framer / mapper. , 2-slot Registered Double Data Rate 2 (DDR2) Synchronous Dynamic Random Access Memory (SDRAM), Network Search Engine, Management Processor Support Elements, and Mutual Compliant with MSA-300 Specifications Consists of a printed circuit board (PCB) assembly that includes a replaceable WAN optical module.
This FPGA provides unique functionality for the device, while the remaining components are industry standard components. This FPGA is the main data path connected to the CX4 connector, 2.5 Gigabit 4x InfiniBand, 266MHz DDR2 SDRAM used for FIFO, SPI-4.2 to connect to framer / mapper, and network search. Implement four electrical interfaces to the LA-1 connected to the engine.
FIFO buffers are implemented using standard DDR2 SDRAM. Time-division multiplex access to memory provides effective dual-port RAM with a maximum ingress bandwidth of over 10 Gbit / sec, while maintaining an exit bandwidth of over 10 Gbit / sec at the same time. This makes it possible to use inexpensive general-purpose memory for FIFO buffers. The control logic in the FPGA divides the SDRAM into multiple VLs and operates the SDRAM memory bus to provide FIFO functionality.
Access to the WAN is provided using components that comply with the specifications specified by the OIF (Optical Internetworking Forum). Specifically, it uses an SFP-4.1 interface and connects to the optical module via the connector specified by the MSA-300 specification. This same interface can also be converted to an IEEE802.3ae XSBI interface on the fly for use with the 10G Ethernet® LAN PHY. Interchangeable modules allow this device to be OC-192 SONET, 10G Ethernet over several types of fibers with varying emission power and receiver sensitivity, depending on user requirements and installed optical modules. It can support (registered trademark) LAN PHY and 10G Ethernet (registered trademark) LAN PHY.
The device can communicate directly across the optical WAN or indirectly through additional standard networking equipment such as SONET / SDH multiplexers, optical regenerators, packet routers, and cell switches.
The SFI-4.1 / XSBI interface is connected to a framer / mapper that internally handles aspects of the low-level signaling protocol (MAC / PCS / PMA functionality in Ethernet terminology). The FPGA communicates the entire packet (or cell in the case of ATM) with the framer / mapper via the SPI-4.2 interface, which is then translated by the framer / mapper into the desired WAN signaling protocol. This conversion is governed by standards published by the ITU (International Telecommunication Union), IETF (Internet Engineering Task Force), ATM Forums, IEEE (the Institute of Electrical and Electronic Engineers), and OIF (Optical Internetworking Forum).
The final component is the Network Search Engine (NSE) connected by LA-1. NSE is used as part of the InfiniBand routing feature to translate incoming IPv6 addresses into local InfiniBand routing descriptions. The FPGA will extract the IPv6 packet from the packet arriving from the WAN and pass it to the NSE, which will then quickly search the internal table to find a match and then the relevant data (match).<u style="single">Infiniband</u>The route) will be returned to the FPGA. If necessary, the management processor in the FPGA will update the NSE table with the new data when new data becomes available.
A second embodiment of the present invention is shown in FIG. This embodiment is a cost-reduced version of the same prototype device shown in FIG. The main goal of this embodiment is to allow distance extension up to 10 km using only 4x data transfer rate (QDR) 1 x InfiniBand, as defined by the InfiniBand architecture.
This implementation consists of an FPGA, a CX4 connector, a single-chip QDR SRAM, and an XFP optical module. The FPGA directly interfaces with both 10 Gigabit 4x InfiniBand (local) and 10Gigabit 1x InfiniBand (WAN).
As in the long-distance embodiment, the interchangeable modules provide an optical WAN interface. However, this module complies with the XFP specification (specified by the XFP MSA group) rather than the MSA-300 interface and communicates directly with the FPGA over the 10 Gigabit XFI bus. This allows the user to select the XFP module that best suits their local environment.
FIFO buffers are implemented using QDR (or QDR2) SRAM, which is a form of memory optimally designed for small dual-port memory. The controller in the FPGA divided the memory into multiple VLs and managed the operation of the FIFO.
<figref num="1">It is a data flow diagram about a prototype device. The main blocks relating to one embodiment of the present invention are shown.</figref><figref num="2">A data flow diagram for a particular long-range implementation designed for use with a number of WAN signaling standards and protocols. It shares many of the functional blocks outlined in Figure 1.</figref><figref num="3">A data flow diagram for a specific, reduced functionality short-range implementation that shows how the InfiniBand architecture can be used as a WAN protocol.</figref>
3 sheets
Sheet 1 Sheet 2 Sheet 3
Every citation, both ways
| Document | Relation | Office |
|---|---|---|
| JP2005505037A | Cites | Japan |
| JP2005064901A | Cites | Japan |
| JP2004056728A | Cites | Japan |
23 members in 11 offices
Priority claims9
| Document | Office | Kind | Date |
|---|---|---|---|
| 59557605 | United States of America | P | |
| 59557605 | United States of America | P | |
| 60595576 | United States of America | – | |
| 2006001161 | Canada | W | |
| 2006001161 | Canada | W | |
| 2005595576 | – | – | – |
| 2006001161 | – | – | – |
| US20050595576P | – | – | – |
| WO2006CA01161 | – | – | – |
Members23
| Document | Office | Kind | |
|---|---|---|---|
| CA2552414A1 | Canada | A1 | |
| US2007014308A1 | United States of America | A1 | |
| AU2006272404A1 | Australia | A1 | |
| WO2007009228A1 | World Intellectual Property Organization (WIPO) | A1 | |
| GB0801276D0 | United Kingdom | D0 | |
| GB2441950A | United Kingdom | A | |
| EP1905207A1 | European Patent Office (EPO) | A1 | |
| KR20080031397A | Republic of Korea | A | |
| AU2006272404A2 | Australia | A2 | |
| CN101258719A | China | A | |
| JP2009502059A | Japan | A | |
| RU2008105890A | Russian Federation | A | |
| GB2441950B | United Kingdom | B | |
| US7843962B2 | United States of America | B2 | |
| AU2006272404B2 | Australia | B2 | |
| JP4933542B2This record | Japan | B2 | |
| RU2461140C2 | Russian Federation | C2 | |
| EP1905207A4 | European Patent Office (EPO) | A4 | |
| CA2552414C | Canada | C | |
| CN101258719B | China | B | |
| KR101290413B1 | Republic of Korea | B1 | |
| IL188772A | Israel | A | |
| EP1905207B1 | European Patent Office (EPO) | B1 |
24 legal events, as the office reported them to INPADOC
Over the term
Point at a mark for the eventEvents
| Event | Code | |
|---|---|---|
| Cancellation because of no payment of annual feesLAPS | LAPS | |
| Receipt of annual feesJAPANESE INTERMEDIATE CODE: R250R250 | R250 | |
| Receipt of annual feesJAPANESE INTERMEDIATE CODE: R250R250 | R250 | |
| Receipt of annual feesJAPANESE INTERMEDIATE CODE: R250R250 | R250 | |
| Receipt of annual feesJAPANESE INTERMEDIATE CODE: R250R250 | R250 | |
| Receipt of annual feesJAPANESE INTERMEDIATE CODE: R250R250 | R250 | |
| Renewal fee payment (event date is renewal date of database)FPAY | FPAY | |
| Certificate of patent or registration of utility modelJAPANESE INTERMEDIATE CODE: R150R150 | R150 | |
| Certificate of patent or registration of utility modelJAPANESE INTERMEDIATE CODE: R150R150 | R150 | |
| First payment of annual fees (during grant procedure)JAPANESE INTERMEDIATE CODE: A61A61 | A61 | |
| Written decision to grant a patent or to grant a registration (utility model)JAPANESE INTERMEDIATE CODE: A01A01 | A01 | |
| Written decision to grant a patent or to grant a registration (utility model)JAPANESE INTERMEDIATE CODE: A01A01 | A01 | |
| Decision of grant or rejection writtenTRDD | TRDD | |
| Written amendmentJAPANESE INTERMEDIATE CODE: A523A521 | A521 | |
| Written permission of extension of timeJAPANESE INTERMEDIATE CODE: A602A602 | A602 | |
| Written request for extension of timeJAPANESE INTERMEDIATE CODE: A601A601 | A601 | |
| Written permission of extension of timeJAPANESE INTERMEDIATE CODE: A602A602 | A602 | |
| Written request for extension of timeJAPANESE INTERMEDIATE CODE: A601A601 | A601 | |
| Written permission of extension of timeJAPANESE INTERMEDIATE CODE: A602A602 | A602 | |
| Written request for extension of timeJAPANESE INTERMEDIATE CODE: A601A601 | A601 | |
| Notification of reasons for refusalJAPANESE INTERMEDIATE CODE: A131A131 | A131 | |
| Report on retrievalJAPANESE INTERMEDIATE CODE: A971007A977 | A977 | |
| Written amendmentJAPANESE INTERMEDIATE CODE: A523A521 | A521 | |
| Written request for application examinationJAPANESE INTERMEDIATE CODE: A621A621 | A621 |
Numbers
- Publication
- 4933542
- Publication, DOCDB
- 4933542
- Publication, EPODOC
- JP4933542B
- Application
- 2008521757
- Application, DOCDB
- 2008521757
- Application, EPODOC
- JP20080521757
Titles2
- Japanese
- インフィニバンド・ネットワークの物理的な到達距離を延長する方法
- English
- How to extend the physical reach of an InfiniBand network
Classification
- CPC, 8
- H04L12/4633
- H04L69/14
- H04L47/10
- H04L47/39
- H04L49/90
- H04L67/08
- H04L2212/00
- H04L12/28
- IPC, 4
- H04L12 46
- G06F13 00
- H04L12 56
- H04L69 14
