Gigabit ethernet adapter supporting the iscsi and ipsec protocols
79 claims: 16 independent, 63 dependent
- 1ネットワークプロトコルを復号および符号化しデータを処理する統合ネットワークアダプタであって、 ストリーミングデータを処理するための、配線接続されたデータパスと、 パケットを送受信しパケットを符号化および復号化するための、配線接続されたデータパスと、 複数の並列の、配線接続されたプロトコル状態マシーンと、 配線接続された転送オフロード・エンジン(TOE)と、 前記TOEと一体化したプロセッサと、 プログラム可能なポート範囲に入るUDPまたはTCPのパケットをすべて例外パスに転送するポート範囲レジスタによるネットワークアドレス変換(NAT)、IPマスカレード、及びポート転送のうちのいずれかについて最適化されたハードウェアサポートを提供するモジュールと、 トラフィックに基づいて共有リソースをスケジュールする手段と、を備え、 前記複数のプロトコル状態マシーンは特定のネットワークプロトコル用に最適化され、 前記プロトコル状態マシーンは並列で処理を実行し、 前記ポート範囲レジスタはある範囲のポートをネットワーク制御動作および前記ポート転送において使用可能とする、ことを特徴とする統合ネットワークアダプタ。
- 2単一の集積回路で構成される統合ネットワークアダプタであって、 配線接続された転送オフロードエンジン(TOE)と、 前記TOEと一体化したプロセッサと、 物理層モジュール(PHY)と、 メディアアクセス層モジュール(MAC)と、 前記TOEと一体化したIPsec処理エンジンと、 前記TOEと一体化した、オフロード処理のための上位レベルプロトコル(ULP)と、 プログラム可能なポート範囲に入るUDPまたはTCPのパケットをすべて例外パスに転送するポート範囲レジスタによるネットワークアドレス変換(NAT)、IPマスカレード、及びポート転送のうちのいずれかについて最適化されたハードウェアサポートを提供するモジュールと、を備え、 前記ポート範囲レジスタはある範囲のポートをネットワーク制御動作および前記ポート転送において使用可能とする、ネットワークアダプタ。
- 3前記ULPはiSCSIプロトコルを実行する、請求項2に記載のネットワークアダプタ。
- 4前記ULPは送受信のためにiSCSI CRCの計算をオフロードする、請求項3に記載のネットワークアダプタ。
- 5前記ULPは送信のために固定インターバルのマーカー(FIM)を使用してiSCSIフレーミングを実行する、請求項3に記載のネットワークアダプタ。
- 6前記TOEはiSCSIヘッダセグメント及びiSCSIデータセグメントをホストiSCSIドライバから取得し、送信のためにiSCSI PDUを準備する、請求項3に記載のネットワークアダプタ。
- 7ホストコンピュータに存在するホストiSCSIドライバを更に備え、 前記ホストiSCSIドライバは前記TOEと通信する、請求項3に記載のネットワークアダプタ。
- 8前記TOEはiSCSIプロトコルデータユニット(PDU)を受信し、iSCSI CRCを計算し、当該iSCSI CRCを前記ホストiSCSIドライバへ送る、請求項7に記載のネットワークアダプタ。
- 9前記ホストiSCSIドライバは算出されたiSCSI CRC値を、iSCSI命令ブロック(IB)のiSCSI CRCシードフィールドを使用してシードする、請求項8に記載のネットワークアダプタ。
- 10前記ホストiSCSIドライバはホストメモリで完全なiSCSIプロトコルデータユニット(PDU)ヘッダを組立て、iSCSI命令ブロック(IB)を生成し、当該iSCSI IBを前記TOEへ送信する、請求項7に記載のネットワークアダプタ。
- 11iSCSI IBは、ホストコンピュータメモリのバッファのリンクリストに対応し、転送ブロックとして知られている、アドレス及び長さの対のセットを含む、請求項10に記載のネットワークアダプタ。
- 12前記ホストiSCSIドライバは、iSCSIデータを受信するときに最終転送ブロックのバッファサイズを調節してCRCバイトを記憶し、iSCSIヘッダとデータセグメントとを正確に分離する、請求項11に記載のネットワークアダプタ。
- 13対応する基本ヘッダセグメント(BHS)と、任意の追加ヘッダセグメント(AHS)と、任意のデータセグメントとを含むiSCSIプロトコルデータユニットが、iSCSI命令ブロック(IB)を用いて前記ホストiSCSIドライバと前記TOEとの間で伝送される、請求項7に記載のネットワークアダプタ。
- 14前記ホストiSCSIドライバは、iSCSI PDUヘッダに対する正確なサイズの受信バッファをポストし、且つ、データセグメントが存在する場合にはiSCSI PDUデータセグメントの正確なサイズの受信バッファをポストすることにより、受信の際にiSCSIプロトコルデータユニット(PDU)ヘッダとデータセグメントを分割する、請求項7に記載のネットワークアダプタ。
- 15前記ホストiSCSIドライバは、命令ブロックを用いて受信された任意の追加ヘッダセグメント(AHS)の正確に寸法付けされたバッファをポストする、請求項14に記載のネットワークアダプタ。
- 16前記TOEと前記ホストiSCSIドライバはiSCSI PDUレベルでインターフェースする請求項7に記載のネットワークアダプタ。
- 17前記TOEは、プロトコルデータユニット(PDU)ヘッダを前記プロセッサまたは前記ホストコンピュータに直接メモリアクセス(DMA)しPDUデータセクションを前記ホストコンピュータに直接メモリアクセス(DMA)することにより、前記ホストコンピュータのメモリの付加的なメモリコピーを必要とせずにiSCSIプロトコルデータユニット(PDU)のヘッダとデータセグメントとを分離する、請求項7に記載のネットワークアダプタ。
- 18前記TOEはセキュリティ・アソシエーション(SA)毎にIPsecのアンチリプレイサポートを行う、請求項2に記載のネットワークアダプタ。
- 19前記TOEはIPSecゼロ、DES、3DESアルゴリズム、およびAES128ビットアルゴリズムを暗号ブロック連鎖(CBC)モードで実行する、請求項2に記載のネットワークアダプタ。
- 20前記TOEはIPSecゼロ、SHA-1、およびMD-5認証アルゴリズムを実行する、請求項2に記載のネットワークアダプタ。
- 21前記TOEはIPSec可変長暗号キーを実行する、請求項2に記載のネットワークアダプタ。
- 22前記TOEはIPSec可変長認証キーを実行する、請求項2に記載のネットワークアダプタ。
- 23前記TOEはIPSecジャンポフレームサポートを実行する、請求項2に記載のネットワークアダプタ。
- 24前記TOEは時間と転送される全データとに基づいてセキュリティ・アソシエーション(SA)の満了のIPSec自動処理を実行する、請求項2に記載のネットワークアダプタ。
- 25前記TOEはIPSecポリシー施行を実行する、請求項2記載のネットワークアダプタ。
- 26前記TOEは例外パケットの生成と状態通知とを含むIPSec例外処理を実行する、請求項2記載のネットワークアダプタ。
- 27統合ネットワークアダプタであって、 パケットを送受信しパケットを符号化および復号化するための、配線接続されたデータパスと、 少なくとも一つの配線接続されたプロトコル状態マシーンと、 少なくとも一つの、前記ネットワークアダプタとホストコンピュータとの間の通信チャンネルと、 トラフィックに基づいて共有リソースをスケジュールするスケジューラと、 配線接続された転送オフロード・エンジン(TOE)と、 前記TOEと一体化したプロセッサと、 プログラム可能なポート範囲に入るUDPまたはTCPのパケットをすべて例外パスに転送するポート範囲レジスタによるネットワークアドレス変換(NAT)、IPマスカレード、及びポート転送のうちのいずれかについて最適化されたハードウェアサポートを提供するモジュールと、を備え、 前記ポート範囲レジスタはある範囲のポートをネットワーク制御動作および前記ポート転送において使用可能とする、統合ネットワークアダプタ。
- 28前記少なくとも一つの通信チャンネルは、命令ブロック(IB)と状態メッセージ(SM)とを使用してデータおよび制御情報を転送する、請求項27に記載のネットワークアダプタ。
- 29前記少なくとも一つの通信チャンネルを経由する通信を制御するための少なくとも一つの閾値タイマを更に備え、 データは選択された閾値インターバルで伝送される、請求項27に記載のネットワークアダプタ。
- 30前記閾値タイマは、前記ネットワークアダプタと前記ホストコンピュータとの間の割り込みの数を減少させ、データの処理能力を増加させるための割込み集約機構を備える、請求項29に記載のネットワークアダプタ。
- 31前記少なくとも一つの通信チャンネルを経由する通信を制御するための少なくとも一つのデータ閾値を設定するモジュールを更に備え、 データレベルが選択された閾値に達したときにデータが伝送される、請求項27に記載のネットワークアダプタ。
- 32前記プロセッサおよび前記TOEのデータ処理能力を最適化する割込み集約機構を更に備える請求項27に記載のネットワークアダプタ。
- 33TCP選択肯定応答(SACK)用に最適化されたハードウェアサポートを提供するモジュールを更に備え、 TCPはデータパケットの紛失を応答し、当該紛失したデータパケットだけを再送信する、請求項27に記載のネットワークアダプタ。
- 34TCP高速再送信用に最適化されたハードウェアサポートを提供するモジュールを更に備える、請求項27に記載のネットワークアダプタ。
- 35前記TCP高速再送信は、順序違いのセグメントが受信されたときに、標準的なアイムアウトを待つ代わりに、送信機が穴を迅速に埋めることを可能にするためにACKを直ちに生成する、請求項34に記載のネットワークアダプタ。
- 36受信機が重複する三つのACKを受信したときに、前記TCP高速再送信が呼び出され、 前記TCP高速再送信が呼び出されたときに、送信機が穴を埋めようとし、 ACKとセグメントのウィンドウ広告値とが相互に一致した場合に、重複するACKが重複と判定される、請求項34に記載のネットワークアダプタ。
- 37TCPウィンドウスケーリング用に最適化されたハードウェアサポートを提供するモジュールを更に備える請求項27に記載のネットワークアダプタ。
- 38ウィンドウスケーリング動作が、ウィンドウスケールを可能にする少なくとも一つのビットと、スケーリングファクタを設定するための少なくとも一つのビットと、スケーリング値を決定するためのパラメータとを備える3つの変数に基づいている、請求項27に記載のネットワークアダプタ。
- 39iSCSIヘッダとデータCRCの生成およびチェックとのために最適化されたハードウェアサポートを提供するモジュールを更に備える請求項27に記載のネットワークアダプタ。
- 40iSCSIの固定インターバルのマーカー(FIM)を生成するために最適化されたハードウェアサポートを提供するモジュールを更に備える請求項27に記載のネットワークアダプタ。
- 41診断プログラムおよびパケット監視プログラムをサポートするTCPダンプモード用に最適化されたハードウェアサポートを提供するモジュールを更に備える請求項27に記載のネットワークアダプタ。
- 42TCPダンプモードが使用可能になった場合に、受信されたパケットのすべてが例外として前記ホストに送信され、ハードウェアスタックから来る出力TCP/UDPパケットのすべてが例外パケットとしてループバックされる、請求項41に記載のネットワークアダプタ。
- 43前記例外パケットをネットワークモニタ用にコピーし、RXパケットを再注入するためにTXパケットを生のイーサネット(登録商標)フレームとして送信するドライバを更に備える請求項42に記載のネットワークアダプタ。
- 44ホストACKモード用に最適化されたハードウェアサポートを提供するモジュールを更に備え、 前記ホストコンピュータと前記ネットワークアダプタとの間をデータが通過したときに当該データが破損した可能性がある場合にデータの完全性を保証するために、前記ホストがTCPセグメントからデータを受信したときだけTCP ACKが送信される、請求項27に記載のネットワークアダプタ。
- 45前記ホストACKモードは、ACKを送信する前に、データセグメントを含んでいるMTXバッファの直接メモリアクセス(DMA)が完了するのを待つ、請求項44に記載のネットワークアダプタ。
- 46TCPがラウンドトリップ時間測定(RTTM)を良好に計算し、ラップシーケンスに対する保護(PAWS)をサポートすることを可能にするために、TCPタイムスタンプ用に最適化されたハードウェアサポートを提供するモジュールを更に備える請求項27に記載のネットワークアダプタ。
- 47古い重複セグメントによるTCP接続の破壊を防止するためにTCP PAWS用に最適化されたハードウェアサポートを提供するモジュールを更に備える請求項27に記載のネットワークアダプタ。
- 48前記ネットワークアダプタ内のバッファからではなくホストメモリのバッファから直接にデータを再送信させるために、TCPホスト再送信モード用に最適化されたハードウェアサポート提供するモジュールを更に備える請求項27記載のネットワークアダプタ。
- 49ランダムな初期シーケンス番号用に最適化されたハードウェアサポートを行うモジュールを更に備える請求項27記載のネットワークアダプタ。
- 50デュアルスタックモード用に最適化されたハードウェアサポートを行うモジュールと、 前記ホストのソフトウェアTCP/IPスタックと協動および関連して動作する前記ネットワークアダプタに組み込まれたハードウェアTCP/IPスタックと、を更に備え、 前記ネットワークアダプタは、前記ネットワークアダプタと同一のIPアドレスを使用して並列に動作する前記ソフトウェアTCP/IPスタックの共存をサポートする、請求項27に記載のネットワークアダプタ。
- 51同期(SYN)状態メッセージモードをサポートするモジュールを更に備え、 受信された任意のSYNは状態メッセージを前記ホストへ戻し、 SYN/ACKは、前記ホストが適切な命令ブロックを前記ネットワークアダプタへ戻すまで前記ネットワークアダプタにより生成されず、 前記SYN状態メッセージモードが前記ネットワークアダプタで使用可能にならない場合には、SYN/ACKが前記ネットワークアダプタにより自動的に生成され、SYN受信状態メッセージは生成されない、請求項50に記載のネットワークアダプタ。
- 52ネットワークアダプタ制御ブロックデータベースに一致しないTCPパケットが受信されたときに、前記ネットワークアダプタからのリセット(RST)メッセージの抑制をサポートするモジュールを更に備え、 前記ネットワークアダプタは、自動的にRSTを生成する代わりにパケットを例外パケットとして前記ホストへ送信することで、前記ホストのTCP/IPスタックが当該パケットを例外パケットとして処理できるようにする、請求項50に記載のネットワークアダプタ。
- 53前記ホストおよび前記ネットワークアダプタがIP IDをオーバーラップすることなくIPアドレスを共有できるようにするために、IP ID分割用に最適化されたハードウェアサポートを提供するモジュールを更に備える請求項27に記載のネットワークアダプタ。
- 54あるタイプのパケットに対する特別な動作を制限し、受け付け、または実行するために、データパケットのフィルタリング用に最適化されたハードウェアサポートを提供するモジュールを更に備える請求項27に記載のネットワークアダプタ。
- 55前記フィルタリングは、プログラムされたユニキャストアドレスの受付、放送パケットの受付、マルチキャストパケットの受付、ネットマスクにより特定される範囲内のアドレスの受付、および全パケットを受け付ける無差別モードのうちの任意の特性を取ることができる、請求項54に記載のネットワークアダプタ。
- 56仮想構内網(VLAN)用に最適化されたハードウェアサポートを提供するVLANモジュールを更に備える、請求項27に記載のネットワークアダプタ。
- 57前記VLANモジュールは、入力パケットからVLANヘッダを除去する要素、VLANのタグが付された出力パケットを生成する要素、入力SYNフレームからVLANパラメータを生成する要素、および例外パケットとUDPパケットのVLANタグ情報を通す要素のいずれかを備える、請求項56に記載のネットワークアダプタ。
- 58ジャンボフレーム用に最適化されたハードウェアサポートを提供するモジュールを更に備える請求項27に記載のネットワークアダプタ。
- 59シンプルネットワーク管理プロトコル(SNMP)用に最適化されたハードウェアサポートを提供するモジュールを更に備える請求項27に記載のネットワークアダプタ。
- 60管理情報ベース(MIB)用に最適化されたハードウェアサポートを提供するモジュールを更に備える請求項27に記載のネットワークアダプタ。
- 61レガシーモードでのネットワークアダプタの動作用に最適化されたハードウェアサポートを提供するモジュールを更に備え、 すべてのネットワークトラフィックはトラフィックタイプにかかわりなく前記ホストに送信され、 前記ネットワークアダプタはハードウェアTCP/IPスタックが該アダプタに存在しないかのように動作する、請求項27に記載のネットワークアダプタ。
- 62ハードウェアおよびソフトウェアのいずれかでIP分割が処理されることを可能にする最適化されたハードウェアサポートを提供するモジュールを更に備え、 例外パケットとして伝送されソフトウェアドライバで再構築されるIP分割されたパケットがIP注入モードにより前記ネットワークアダプタへ再注入されて戻される、請求項27に記載のネットワークアダプタ。
- 63IPパケットが前記ネットワークアダプタのTCP/IPスタックに注入されることを可能にするIP注入用に最適化されたハードウェアサポートを提供するIP注入モジュールを更に備える請求項27に記載のネットワークアダプタ。
- 64前記IP注入モジュールは、IPパケットを前記ネットワークアダプタのTCP/IPスタックへ注入するための1以上の注入制御レジスタを備え、 前記1以上の注入制御レジスタは前記ホストがIPパケットを前記ネットワークアダプタのTCP/IPスタックへ注入することを可能にする、請求項63に記載のネットワークアダプタ。
- 65複数のIPアドレス用に最適化されたハードウェアサポートを提供するモジュールを更に備える請求項27に記載のネットワークアダプタ。
- 66デバッグモード用に最適化されたハードウェアサポートを提供するモジュールを更に備え、 試験および制御ビットが前記ネットワークアダプタ中で使用可能になった場合に、すべてのIPパケットが例外として前記ホストへ送信される、請求項27に記載のネットワークアダプタ。
- 67TCP時間待機状態用に最適化されたハードウェアサポートを提供するモジュールを更に備える請求項27に記載のネットワークアダプタ。
- 68可変の接続数に対して最適化されたハードウェアサポートを行うモジュールを更に備える請求項27に記載のネットワークアダプタ。
- 69前記ネットワークアダプタがネットワークアダプタの最大容量に等しい接続を容認する場合に、次の同期(SYN)が、前記ホストが当該接続を処理できるように、例外パケットとして前記ホストへ通される、請求項27に記載のネットワークアダプタ。
- 70ユーザデータグラムプロトコル(UDP)用に最適化されたハードウェアサポートを提供するモジュールを更に備える請求項27記載のネットワークアダプタ。
- 71IPパケットの寿命を選択されたホップ数に限定するためにTTL(生存時間)に対して最適化されたハードウェアサポートを提供するモジュールを更に備える請求項27に記載のネットワークアダプタ。
- 72キープ・アライブ・パケットをリンクにおいて周期的に送信することによりアイドル状態のTCP接続を維持しタイムアウトさせないことを可能にするために、TCPのキープアライブ用に最適化されたハードウェアサポートを提供するモジュールを更に備える請求項27に記載のネットワークアダプタ。
- 73IPパケットを優先させるためにルータにより使用されるサービスのTCPタイプ(TOS)用に最適化されたハードウェアサポートを提供するモジュールを更に備える請求項27に記載のネットワークアダプタ。
- 74統合ネットワークアダプタであって、 パケットを送受信しパケットを符号化および復号化するための、配線接続されたデータパスと、 少なくとも一つの配線接続されたプロトコル状態マシーンと、 少なくとも一つの、前記ネットワークアダプタとホストコンピュータとの間の通信チャンネルと、 トラフィックに基づいて共有リソースをスケジュールするスケジューラと、 TCPスロースタート用に最適化されたハードウェアサポートを行うためのモジュールと、を備え、 前記スロースタートは、 肯定応答(ACK)を予測する前に、最大セグメントサイズ(MSS)の2倍の現在のウィンドウ(cwnd)に対応する二つのデータセグメントが流れることだけをまず許可し、 更に一つのセグメントを流すために、前記cwndが受信機の広告ウィンドウと等しくなるまで、成功ACKが受信される度に一つのMSS分だけcwndを増加させることにより、一度に流れるデータセグメント数を徐々に増加させる、統合ネットワークアダプタ。
- 75前記スロースタートは新規のデータ接続で常に開始され、 前記スロースタートはデータトラフィックの渋滞が発生したときに接続の中間で起動される、請求項74に記載のネットワークアダプタ。
- 76統合ネットワークアダプタであって、 パケットを送受信しパケットを符号化および復号化するための、配線接続されたデータパスと、 少なくとも一つの配線接続されたプロトコル状態マシーンと、 少なくとも一つの、前記ネットワークアダプタとホストコンピュータとの間の通信チャンネルと、 トラフィックに基づいて共有リソースをスケジュールするスケジューラと、 配線接続された転送オフロード・エンジン(TOE)と、 前記TOEと一体化したプロセッサと、 フレキシブルでプログラム可能なメモリエラー検査および補正(ECC)用に最適化されたハードウェアサポートを提供するECCモジュールと、を備え、 前記ECCモジュールは、少なくとも一つの追加ビットを用いて、暗号化されたECCコードをデータと共にパケットに記憶し、 前記データがメモリに書き込まれたときに前記ECCコードも記憶され、 前記データが読み出されたときに、前記記憶されたECCコードは当該データが書き込まれたときに生成されたECCコードと比較され、 前記ECCコードが一致しない場合には、前記データ中のどのビットがエラーであるかについての判定が実行され、 エラーになっているビットが反転され、メモリ制御装置が当該補正されたデータを解放し、 エラーはオンザフライで補正され、補正されたデータは前記メモリに戻されず、 同一の破損データが再度読み取られたならば、前記ECCモジュールの動作が繰り返される、統合ネットワークアダプタ。
- 77統合ネットワークアダプタであって、 パケットを送受信しパケットを符号化および復号化するための、配線接続されたデータパスと、 少なくとも一つの配線接続されたプロトコル状態マシーンと、 少なくとも一つの、前記ネットワークアダプタとホストコンピュータとの間の通信チャンネルと、 トラフィックに基づいて共有リソースをスケジュールするスケジューラと、 配線接続された転送オフロード・エンジン(TOE)と、 前記TOEと一体化したプロセッサと、 TCPのサービス品質(QoS)用に最適化されたハードウェアサポートを提供するモジュールと、を備え、 TCP送信データフローはソケット問い合わせモジュールで開始し、当該データフローは送信データ有効ビットセットを有するエントリを探す送信データ有効ビットテーブルを通過し、 前記ソケット問い合わせモジュールは、前記エントリを発見した場合に、ソケットのユーザ優先順位レベルに従って当該エントリを複数の待ち行列のうちの一つに置く、統合ネットワークアダプタ。
- 78統合ネットワークアダプタであって、 パケットを送受信しパケットを符号化および復号化するための、配線接続されたデータパスと、 少なくとも一つの配線接続されたプロトコル状態マシーンと、 少なくとも一つの、前記ネットワークアダプタとホストコンピュータとの間の通信チャンネルと、 トラフィックに基づいて共有リソースをスケジュールするスケジューラと、 配線接続された転送オフロード・エンジン(TOE)と、 前記TOEと一体化したプロセッサと、 フェイルオーバー用に最適化されたハードウェアサポートを提供するフェイルオーバーモジュールと、を備え、 前記フェイルオーバーモジュールは、接続を開始しようとすることなくソケットが生成されることを可能にするNO_SYNモードを有し、 前記ネットワークアダプタ内のソケットおよび当該ソケットに関連するすべてのデータ構造が接続を生成することなく生成され、 前記NO_SYNモードは別のカードからの、またはソフトウェアTCP/IPスタックから前記ネットワークアダプタへの接続移行からのフェイルオーバーをサポートする、統合ネットワークアダプタ。
- 79統合ネットワークアダプタであって、 パケットを受信するための、配線接続されたデータパスと、 複数の並列の、配線接続されたプロトコル状態マシーンと、 トラフィックに基づいて共有リソースをスケジュールするスケジューラと、 配線接続された転送オフロード・エンジン(TOE)と、 前記TOEと一体化したプロセッサと、 プログラム可能なポート範囲に入るUDPまたはTCPのパケットをすべて例外パスに転送するポート範囲レジスタによるネットワークアドレス変換(NAT)、IPマスカレード、及びポート転送のうちのいずれかについて最適化されたハードウェアサポートを提供するモジュールと、を備え、 前記プロトコル状態マシーンは並列で処理を実行し、 前記ポート範囲レジスタはある範囲のポートをネットワーク制御動作および前記ポート転送において使用可能とする、統合ネットワークアダプタ。
Independent claims79
862 paragraphs, as filed
The present invention relates to telecommunications, in particular methods and devices for processing data with communication protocols used to transmit and receive data.
Computer networks require provisions for various communication protocols for transmitting and receiving data. Typically, a computer network comprises a system of devices such as computers, printers, and other computer peripherals that are connected so that they can communicate with each other. Data is transferred between each of these devices by data packets communicated over the network using communication protocol standards. Many different protocol standards are currently in use. Examples of popular protocols are Internet Protocol (IP), Internet Protocol Packet Exchange (IPX), Sequenced Packet Exchange (SPX), Transmission Control Protocol (TCP), and Point-to-Point Protocol (PPP). Each network device contains a combination of hardware and software that translates protocols and processes data.
One example is a computer mounted on a local area network (LAN) system, where the network device uses hardware to handle the link layer protocol and software to manage network, transfer, communication protocols and information data processing. Network devices typically consist of one link-layer protocol in hardware, limiting the attached computer to that particular LAN protocol. Higher-order protocols, such as network, transfer and communication protocols, are configured as software programs with data handlers and process the data once they are sent through the hardware of the network device to system memory. The advantage of this configuration is that general purpose equipment such as computers can be used in many different network setups to support the required optional network applications. However, as a result of this configuration, the system has a high processor overhead and a large amount of system memory to coordinate different software protocols and data handlers that communicate with the computer's operating system (OS) and the computer, and to the network hardware. Complex configuration setup is required on the computer user side.
This high overhead required for processing time is shown in Schrier's US Pat. No. 5,485,460, filed January 16, 1996, which provides a large number of software protocol stacks running the same protocol on the device. Shows how it works. This type of configuration is used by Microsoft Windows for disk operating system (DOS) -based machine operation. During normal operation, once the hardware verifies the forwarding or link layer protocol, the resulting data packet is forwarded to the software layer, which determines the packet frame format and decomposes any particular frame header. The packet is then sent to a different protocol stack, where it is evaluated for a particular protocol. However, a packet can be forwarded to several protocol stacks before it is accepted or rejected. The time delay generated by the software protocol stack prevents audio and video transmissions from being processed in real time, and the data must be buffered before it can be played. The amount of processing overhead required to process the protocol is very high and extremely cumbersome, and it is clear that it is suitable for applications with powerful central processing units (CPUs) and large amounts of memory.
There are computer products on the market that do not fit the traditional network equipment model. Some examples of these products include pagers, cellular phones, gaming machines, smart phones and televisions. Most of these products have a small footprint, an 8-bit controller, limited memory, or require a very limited form of factor. Consumer products like these are simple, inexpensive and require low power consumption. The protocol configuration described above requires a great deal of hardware and processor power to meet these requirements. Due to the complexity of such configurations, it is difficult for prices to be integrated into consumer products in an efficient manner. These products can access network services such as the Internet if network access can be facilitated to be easily manufactured in inexpensive, low power, low form factor equipment.
Communication networks use protocols for sending and receiving data. A communication network typically includes a collection of network devices, also called nodes, such as computers, printers, storage devices, and other computer peripherals that are connected so that they can communicate with each other. Data is transferred between each of these network devices using data packets transmitted over the communication network using the protocol. Many different protocols are currently in use. Examples of popular protocols are Internet Protocol (IP), Internet Protocol Packet Exchange (IPX) Protocol, Sequenced Packet Exchange (SPX) Protocol, Transmission Control Protocol (TCP), Point-to-Point Protocol (PPP), and Under Development. Includes other similar new protocols in. Network devices include a combination of hardware and software that processes protocols and data packets.
In 1978, the International Organization for Standardization (ISO), the standards-setting body, created a network reference model known as the Open Systems Interconnection (OSI) model. This OSI model contains seven conceptual layers. That is, 1) the physical (PHY) layer that defines the physical components that connect the network device to the network, and 2) the data that controls the movement of data in a discrete form known as a frame containing data packets. A link layer, a network layer that constructs data packets according to a specific protocol, a transfer layer that ensures reliable transfer of data packets, and a session that enables two-way communication between network devices. There is a layer, 6) a presentation layer that controls the method of representing data and confirms that the data is in the correct form, and 7) an application layer that shares files, processes messages, prints, and so on. Sometimes the session and presentation layers are omitted from this model. For an explanation of how modern telecommunications networks and the Internet relate to ISO's seven-tier model, see the text Internetworking, for example. with TCP / IP , Chapter 11 of Douglas E. Corner (Volume 1, 4th Edition, ISBN 020 1633469) and the text TCP / IP Illustrated , W. Richard Stevens (Volume 1, ISBN 0130183806). See Chapter 1.
An example of a network device is a computer mounted on a premises network (LAN), where the network device uses the hardware of the host computer to manage the physical and data link layers, network, transfer, session, presentation, Use software running on the host computer to manage the application layer. The network, transfer, session, and presentation layers are performed using protocol processing software, also known as the protocol stack. The application layer is executed using the application software that processes the data once the data is transferred through the hardware of the network device and the protocol processing software. The advantage of this software-based protocol processing structure is that it allows general purpose computers to be used in many different types of communication networks and supports any application required. However, as a result of this software-based protocol processing structure, the overhead of protocol processing software running on the host computer's central processing unit (CPU) to process the network, transfer, session, and presentation layers is very high. Software-based protocol processing structures also require large amounts of memory on the host computer, as the data must be copied and moved when the software processes the data. The high overhead required by protocol processing software is described in Schrier's US Pat. No. 5,485,460, filed January 16, 1996, which teaches how to operate a large number of software protocol stacks. This type of software-based protocol processing structure is used, for example, in Microsoft Windows, where computers run.
During normal operation of the network device, the network software hardware extracts data packets that are then sent to the host computer's protocol processing software. The protocol processing software runs on the host computer, which is not optimized for the tasks performed by the protocol processing software. The combination of protocol processing software and a general purpose host computer is not optimized for protocol processing, which limits its performance. Performance limitations in protocol processing, such as the time delay generated by the execution of protocol processing software, are detrimental, preventing, for example, audio and video transmissions from being processed in real time, or the full speed and capacity of the communication network. Prevents the use of. The amount of overhead of the host computer CPU required to process the protocol is very high and extremely annoying, and it is clear that the host computer requires the use of CPU and large amounts of memory.
New consumer and industrial products that do not fit the traditional network equipment model are entering the market, while network speeds continue to increase. Examples of these consumer products are internet enable cell phones, internet enable TVs, and internet devices. Examples of industrial products are network interface cards (NICs), internet routers, internet switches, and internet storage servers. Software-based protocol processing configurations are inefficient in meeting the demands of these new consumer and industrial products. Software-based protocol processing configurations are complex and difficult to include in consumer products at an efficient price. Software-based protocol processing configurations require processing power and are difficult to integrate into high-speed industrial products. If the processing of the protocol can be simplified and optimized so that it is easily manufactured in an inexpensive, low power, high performance integrated small form factor device, these consumer and industrial products Data can be read and written over any communication network such as the Internet.
As opposed to software-based, Internet tuners with hardware-based protocol processing configurations are J. Minami, R. Koyama, M. Johnson, M. Shinohara, T. Poff, and D. Burkes' Multiple network protocol encoder / decoder and data. processor, described in US Pat. No. 6,034,963 (March 7, 2000) (hereinafter referred to as '963). This internet tuner provides the basic technology for processing protocols.
<p> It is useful to provide a Gigabit Ethernet® adapter that provides a hardware solution to high communication network speeds. It is useful to provide Gigabit Ethernet® adapters that are compatible with many more communication protocols.</p>
<p> The present invention is practiced with Gigabit Ethernet® adapters. The system according to the invention provides a compact hardware solution for handling high network communication speeds. In addition, the present invention fits a number of communication protocols due to its modular structure and design. A preferred embodiment of the present invention provides an integrated network adapter that decodes and encodes a network protocol and processes the data. This network adapter has a wired data path for processing streaming data, a wired data path for receiving and sending packets and encoding and decoding packets, and multiple parallel, parallel data paths. It comprises a wired and connected protocol state machine, where each protocol state machine is optimized for a special network protocol, the protocol state machine performs processing in parallel, and further shares shared resources based on traffic. It has the means to schedule.</p>
<figref num="1">Schematic block diagram of the NIC card structure according to the present invention.</figref><figref num="2">Schematic block diagram of an interface of a device mounted on a network according to the present invention.</figref><figref num="3">Level block diagram of a system according to the present invention.</figref><figref num="4">High-level block diagram of a Gigabit Ethernet® adapter according to the present invention.</figref><figref num="5">The block diagram which showed the outline of the I / O used in the MAC interface module according to this invention.</figref><figref num="6">Schematic block diagram of an Ethernet® interface according to the present invention.</figref><figref num="7">Schematic block diagram of an address filter and packet type parser / module according to the present invention.</figref><figref num="8">The timing diagram which showed the operation of the address filter and the packet type parser / module according to this invention.</figref><figref num="9">Schematic block diagram of a data aligner module according to the present invention.</figref><figref num="10">Schematic block diagram of an ARP module according to the present invention.</figref><figref num="11">Schematic block diagram of an ARP cache according to the present invention.</figref><figref num="12">The figure which shows the transmission queue entry format according to this invention.</figref><figref num="13">The figure which shows the search table entry format according to this invention.</figref><figref num="14">The figure which shows the ARP cache entry format according to this invention.</figref><figref num="15">The flow chart which shows the ARP search process according to this invention.</figref><figref num="16">Schematic block diagram of an IP module according to the present invention.</figref><figref num="17">Block diagram of an IP generator according to the present invention.</figref><figref num="18">The block diagram which shows the data flow by the injector according to this invention.</figref><figref num="19">Top-level block diagram of a TCP module according to the present invention.</figref><figref num="20">The figure which shows the TCP reception data flow according to this invention.</figref><figref num="21">The flow diagram of the control block search solution of the VSOCK / Rcv state handler according to this invention.</figref><figref num="22">A basic data flow diagram according to the present invention.</figref><figref num="23">The flow diagram of the socket received data according to this invention.</figref><figref num="24">The figure which shows the socket transmission flow according to this invention.</figref><figref num="25">A flow chart of data according to the present invention.</figref><figref num="26">Block diagram of a module according to the present invention.</figref><figref num="27">The figure which shows the algorithm according to this invention.</figref><figref num="28">Block diagram of the entire algorithm shown in Figure 27.</figref><figref num="29">The figure which shows the logical means according to this invention.</figref><figref num="30">Format diagram of options according to the present invention.</figref><figref num="31">Another optional format diagram according to the present invention.</figref><figref num="32">Another optional format diagram according to the present invention.</figref><figref num="33">Still another optional format diagram according to the present invention.</figref><figref num="34">Still another optional format diagram according to the present invention.</figref><figref num="35">Schematic block diagram of an IP router according to the present invention.</figref><figref num="36">Format diagram of the IP route entry according to the present invention.</figref><figref num="37">An explanatory diagram of signaling used to request and receive a route according to the present invention.</figref><figref num="38">Schematic block diagram of an exception handler according to the present invention.</figref><figref num="39">The figure which shows the M1 memory map according to this invention.</figref><figref num="40">The figure which shows the sample memory map according to this invention.</figref><figref num="41">A block diagram of a data flow according to the present invention.</figref><figref num="42">Block diagram of the mtxarb subunit according to the present invention.</figref><figref num="43">A block diagram of a data flow according to the present invention.</figref><figref num="44">Block diagram of the mcbarb subunit according to the present invention.</figref><figref num="45">The figure which shows the default memory map of the network stack according to this invention.</figref><figref num="46">The figure which shows the default setting according to this invention.</figref><figref num="47">FIG. 5 showing matching IB and SB queues forming channels together according to the present invention.</figref><figref num="48">FIG. 5 is a flow chart of processing an instruction block queue according to the present invention.</figref><figref num="49">A block diagram showing a data flow of an ar state block passing between a network stack, an on-chip processor, and a host according to the present invention.</figref><figref num="50">Block diagram of iSCSI transmission data path according to the present invention.</figref><figref num="51">The figure which shows the iSCSI transmission flowchart according to this invention.</figref><figref num="52">Explanatory drawing of the use of a 4-byte buffer according to the present invention.</figref><figref num="53">Block diagram of iSCSI receive data path according to the present invention.</figref><figref num="54">Explanatory drawing which shows the transfer divided into two requests according to this invention.</figref><figref num="55">The figure which shows the DMA transfer to the host which was divided into separate requests according to this invention.</figref><figref num="56">The figure which shows the SA block flow according to this invention.</figref><figref num="57">The figure which shows the format of the TX AH transfer SA block by this invention.</figref><figref num="58">The figure which shows the format of the TX ESP-1 transfer SA block by this invention.</figref><figref num="59">The figure which shows the format of the TX ESP-2 transfer SA block by this invention.</figref><figref num="60">The figure which shows the format of the TX AH tunnel SA block by this invention.</figref><figref num="61">The figure which shows the format of the TX AH tunnel SA block by this invention.</figref><figref num="62">The figure which shows the format of the TX ESP-2 tunnel SA block by this invention.</figref><figref num="63">The figure which shows the format of the RX AH SA block by this invention.</figref><figref num="64">The figure which shows the format of the RX ESP-1 SA block by this invention.</figref><figref num="65">The figure which shows the format of the RX ESP-2 SA block by this invention.</figref><figref num="66">A block diagram showing the overall flow for IPSEC logic according to the present invention.</figref><figref num="67">Schematic block diagram of the data flow according to the present invention.</figref><figref num="68">Block diagram of the flow of the data path for the received IPSEC packet according to the present invention.</figref><figref num="69">The flow chart which shows the IPSEC anti-relay algorithm according to this invention.</figref>
The present invention is practiced with Gigabit Ethernet® adapters. The system according to the present invention provides a compact hardware solution for handling high network communication speeds. Moreover, the present invention is compatible with a number of communication protocols through modular construction and design.
[Introduction] [General description] The present invention includes an architecture (hereinafter referred to as IT10G) used in a high-speed hardware network stack. The description here defines the theory and timing of data paths and flows, registers, applications. Combined with other system blocks, IT10G provides the core of line speed TCP / IP processing.
[Definition] The following terms used here have corresponding meanings: 10Gbps 10 Gigabit (10,000,000,000 bits per second) ACK acknowledgment AH authentication header AHS additional header segment ARP address resolution protocol BHS basic header segment CB control block CPU central processing unit CRC Cyclic Redundancy Check DAV Available data DDR double data rate DIX Digital Intel Xerox DMA direct memory access Denial of DOS service DRAM dynamic RAM EEPROM Electrically erasable PROM ESP Encapsulation Security Payload FCIP Fiber channels across IP FIFO first in first out FIM Fixed interval marker FIN end Gb Gigabit (10,000,000,000 bits per second) HDMA host DMA HO half open HR host retransmission HSU header storage IB instruction block ICMP Internet Control Message Protocol ID identification IGMP internet group management protocol IP Internet protocol IPsec IP security IPX Internet packet switching IQ instruction block queue iSCSI Internet Small Computer System Interface ISN initial sequence number LAN premises network LDMA local DMA LIP local IP address LL linked list LP local port LSB least significant bit LUT search table MAC media access controller MCB CB memory MDL memory descriptor list MIB Management Information Base MII media independent interface MPLS multiprotocol label switching MRX receive memory MSB most significant bit MSS maximum segment size MTU maximum transmission unit MTX TX (transmit) memory NAT network address translation NIC network interface card NS network stack OR or logical function PDU protocol data unit PIP peer IP address PP peer port PROM Programmable ROM PSH push PV valid pointer QoS service quality RAM random access memory RARP Reverse Address Resolution Protocol Rcv reception RDMA remote DMA ROM read-only memory RST reset RT round trip RTO retransmission timeout RTT round trip time RX reception SA Security Association SB state block SEQ sequence SM status message SNMP Simple Network Management Protocol SPI security parameter index Stagen state generator SYN sync TCP transport control protocol TOE transfer offload engine Type of TOS service TTL survival time TW standby time TX transmission UDP user datagram protocol URG emergency VLAN virtual LAN VSOCK virtual socket WS window scaling XMTCTL transmission control XOR Exclusive Or
[Application Overview] [Overview] As bandwidth continues to increase, the ability to handle TCP / IP communications will exceed the overhead of the system processor. Many sources consume nearly 100% of the host computer's CPU bandwidth when the Ethernet rate reaches gigabit (Gbps) per second, and when the rate increases further to 10 Gbps. , Tells us that the entire TCP / IP protocol processing must be offloaded to a dedicated subsystem. The IT10G described here configures TCP and IP as a series of state machines with related protocols including, for example, ARP, RARP, and IP host routing. The IT10G core constitutes an accelerator or engine, also known as a transfer offload engine (TOE). Hooks are provided to allow connected on-chip processors to be used to extend the characteristics of the network stack, but the IT10G core does not use a processor or software.
[Sample application] An example use of the IT10G core is an intelligent network interface card (NIC). In a typical application, the NIC is plugged into a computer server and handles TCP / UDP / IP packets natively.
FIG. 1 is a block diagram illustrating the NIC of the present invention. In Figure 1, the IT10G core 10 is coupled with a processor 11, system peripherals 12, and a system bus interface 13 to a single-chip NIC controller. The single-chip NIC controller is integrated with Ethernet® PHY14 and is coupled with configuration EEPROM15, an optional external memory in the network stack to form a low chip count NIC. The processor memory 16 (both ROM and RAM) may be located inside or outside the integrated chip.
Another use of the IT10G core 10 is to act as an interface to network-mounted devices such as storage devices, printers, cameras, etc. In such cases, a custom application socket (or interface) 17 can be provided on the IT10G to handle Layer 6 and 7 protocols and facilitate the movement of special data for that application. Examples include custom data paths for protocols such as streaming media, bulk data movement, iSCSI and FCIP.
FIG. 2 is a schematic block diagram of an interface to a device mounted on a network according to the present invention. IT10G is designed to support 10Gbps line speed processing, and the same architecture and logic can also be used at lower speeds. In these cases, there is only a difference between Ethernet (registered trademark) MAC21 and PHY14. Benefits of using this architecture with low line speeds include, for example, low power consumption.
[Challenge] The challenge for high-speed bandwidth lies in processing TCP / IP packets at wire-line speeds. This is shown in the table below.<tables num="1"><img file="JP4875126B2_D0001.tif" /></tables>
Note: 1 This assumes an average packet size of 500 bytes. 2 This assumes 500 instruction overhead per packet and 1 instruction per byte.
The numbers in the table above are very conservative and do not consider, for example, the dual characteristics of networking. If dual operation is taken into account, the processing power requirement can easily be doubled. In any case, starting at the gigabit level, it is clear that TCP / IP processing overhead is the main drain in host computer processing power and another solution is needed.
Bandwidth limit IT10G solves the processing power limitation of the host computer by implementing various architectures. The implementation includes the following characteristics: -On-the-fly (streaming) processing of input and output data, -Ultra-wide data path (64-bit in current structure), -Parallel operation of protocol state machines, · Intelligent scheduling of shared resources, -Minimized memory copy.
[System overview] [Overview] This section describes the upper level of preferred embodiments. It provides a block-level description of the system and a theory of behavior for different data paths and transfer types.
Embodiments of the invention include an IT10G network stack that combines processor cores and system components to provide a fully networked subsystem for different applications. A block-level diagram of the system is shown in Figure 3.
[Clock request] A preferred embodiment of the present invention is a chip designed to operate in different clock domains. The table below lists all clock domains for both 1 Gbps and 10 Gbps operation.<tables num="2"><img file="JP4875126B2_D0002.tif" /></tables>
[Protocol processor] [Overview] This section gives an overview of the internal protocol processor.
[Process core] The chips described here use an internal (or on-chip) processor for programming power and flexibility. This processor is also equipped with all the peripherals needed to complete the operating system. Under normal operating conditions, the on-chip processor controls network slack.
[Memory architecture] The on-chip processor has the ability to address up to 4GB of memory. Within this address space are all its peripherals, its RAM, ROM and network slack.
[Network Slack Architecture] [Overview] This section outlines the IT10G architecture. Subsequent sections here go into details about individual modules. IT10G takes network slack hardware protocol processing capabilities and adds enhancements to enable scaling up to 10Gbps rates. The main additions to previous versions are to extend the data path, parallel operation of state machines, and intelligent scheduling of shared resources. In addition, other previously unsupported protocols will be added support for protocols such as RARP, ICMP, and IGMP. Figure 4 is a high-level block diagram of the IT10G.
[Operation theory] The socket connection must be initialized before using IT10G to transfer any data. This can be done by using command blocks or by programming the TCP socket registers directly. The characteristics that must be programmed for each socket include the destination IP address, destination port number, and type of connection (eg TCP or UDP, server or client). Optional parameters include settings such as QoS level, source port, TTL, and TOS settings. Once these parameters are entered, the socket can be activated. For UDP sockets, data can begin to be sent or received immediately. For the TCP client, the socket connection must be set up first, and for the TCP server, the SYN packet must be received from the client and then the socket connection must be set up. All such operations can be fully performed by IT10G hardware.
[Send packet] When a TCP packet needs to be sent, the application running on the host computer first writes the data to a socket (fixed socket or virtual socket-virtual sockets are supported by the IT10G architecture). If the current send buffer is empty, a partially working checksum is maintained when the data is written to memory. The partial checksum is used as a start seed for the checksum calculation, eliminating the need for a TCP layer in the IT10G network stack to read through the data again before sending it out. Data can be written to the socket buffer in 32-bit or 64-bit chunks. Up to 4 valid_byte signals are used to indicate valid bytes. When writing to the socket buffer, the data must be packed with only the last word with possible invalid bytes. This step also applies to UDP packets where there is an option not to calculate the data checksum.
Once all the data has been written, the SEND command can be issued by an application running on the host computer. At this point, the TCP / UDP engine calculates the packet length, checkssums it, and creates a TCP / IP header. This TCP / IP header is pre-pending to the socket data section. The packet's buffer pointer is then placed on the send queue with the socket QoS level.
The transmit scheduler observes all sockets with pending packets and selects the packet with the highest QoS level. This transmit scheduler observes all types of packets that need to be transmitted. These packets can include, for example, TCP, UDP, ICMP, ARP, RARP, raw packets. The minimum bandwidth algorithm is used to ensure that the sockets are not fully coupled. When a socket packet is selected for transmission, the socket buffer pointer is forwarded to the MAC TX interface. MAC The TX interface operates to read data from the socket buffer and send that data to the MAC. Buffer retransmission collision or other reason an Ethernet (registered trademark) is used to store the packets is output when signal is required. Once packet data is sent from the original socket buffer, the data buffer is released. When a valid transmit state is received from the MAC, the data buffer is flushed and the next packet can be subsequently transmitted. If an invalid transmit state is received from the MAC, the last packet stored in the data buffer is retransmitted.
[Receive packet] When a packet is received from the MAC, the Ethernet (R) header is analyzed to determine if the packet is destined for this network stack. The MAC address filter can be programmed to receive unicast addresses, unicast addresses, broadcast addresses, or multicast addresses that fall within the programmed mask. In addition, the encapsulation protocol is also determined. If the 16-bit TYPE field in the Ethernet header indicates an ARP (0x0806) or RARP (0x0835) packet, the ARP / RARP module is allowed to process further packets. If the TYPE field is decrypted to IPv4 (0x0800), the IP module is allowed to process further packets. A complete list of supported TYPE fields for the example is shown in the table below. If the TYPE field is decrypted to any other value, the packet is optionally buffered and the host computer is notified that an unknown Ethernet® packet is being received. In this last case, the application reads the packet and determines the appropriate course of operation. This data path configuration allows IT10G to indirectly support any protocol that is not directly supported by hardware, such as IPX.<tables num="3"><img file="JP4875126B2_D0003.tif" /></tables>
Note: IPv6 packets are managed as an exception at the Ethernet tier.
[ARP / RARP packet] If the packet received is an ARP or RARP packet, the ARP / RARP module becomes available. It inspects the OP field of the packet to determine if it is a request or an answer. If it is a request, an external entity is polling the information. If the address being poled is for IT10G, reply_req will be sent to the ARP / RARP answer module. If the packet received is an ARP or RARP reply, the result, ie the MAC and IP address, is sent to the ARP / RARP request module.
In another embodiment, the ARP and / or RARP function is processed in the host computer using IT10G's dedicated and optimized hardware to transmit ARP / RARP packets to the host over the exception path. To.
[IP packet] If the received packet is an IP packet, the IP module is available. The IP module checks the version field in the IP header to determine if the first packet received is an IPv4 packet.
The IP module parses the embedded protocol of the received packet. Based on the protocol being decrypted, the received packet is sent to the appropriate module. Protocols directly supported by the hardware of this embodiment of the invention include, for example, TCP and UDP. Other protocols such as RDMA can be supported by other optimized processing modules. All unknown protocols are handled using exception handlers.
[TCP packet] If the TCP packet is received by IT10G, the socket information will be parsed and the corresponding socket will be available. The socket state information is retrieved and the socket state is updated accordingly based on the type of packet received. The packet data payload (if applicable) is stored in the socket data buffer. If an ACK packet needs to be generated, the TCP state module will generate the ACK packet and schedule it to be sent. If a TCP packet that does not correlate with the open socket is received, the TCP state module spawns a RST packet and the RST packet is scheduled to be sent.
[UDP packet] If a UDP packet is received, the socket information is parsed and the data stored in the socket receives the data buffer. If the open socket does not exist, the UDP packet is implicitly dropped.
In another embodiment, UDP packets can be processed by the host computer using exception handlers.
[Network stack register] The IT10G hardware network stack is configured to appear as a peripheral to the on-chip processor. The base address of the network stack is programmed through the on-chip processor's NS_Base_Add register. This architecture allows on-chip processors to place the network stack in various locations in its memory or I / O space.
Ethernet (registered trademark) MAC interface [Overview] The following description describes an Ethernet® MAC interface module. The function of the Ethernet MAC Interface Module is to abstract the Ethernet MAC from the core of the IT10G. This allows, for example, the IT10G network stack core to be coupled to MACs of different speeds and / or MACs from various sources without changing the IT10G core architecture. This section describes the interface requirements for communication with the IT10G core.
[Module I / O] Figure 5 is a block diagram schematically showing the I / O used in the MAC interface module.
Ethernet Interface [Overview] This section describes Ethernet interface modules. Ethernet interface modules communicate with low-end Ethernet (registered trademark) MAC interfaces and blocks such as high-end ARP and IP modules. The Ethernet® interface module processes both receive and transmit path data. On the transmitting side, the Ethernet interface module operates to schedule packets for transmission, set up DMA channels for transmission, and communicate via Ethernet MAC interface transmission signals. On the receiving side, the Ethernet interface module parses the Ethernet header, determines if the packet should be received based on the address filter settings, and then based on the TYPE field in the packet header. Encapsulated protocol enable, works to align data to start at 64-bit boundaries with respect to upper layer protocols. FIG. 6 is a schematic block diagram of the Ethernet® interface 40.
[Description of submodule block] [Send Scheduler] The transmit scheduler block 60 acts to take transmit requests from ARP, IP, TCP, and the raw transmit module and determine which packet should be transmitted next. The transmission scheduler determines the transmission order by comparing the QoS levels of each transmission request. Along with the QoS level, each send request contains a pointer to the starting memory block of the packet along with the packet length. The transmission scheduler has the ability to be programmed to weight the transmission priority of some packet types more heavily than others. For example, QoS level 5 from a TCP module can be made more valuable than a level 5 request from an IP module. The outbound scheduler allows a large number of modules to work in parallel in a shared way based on outbound data traffic. The following description is the algorithm currently used to determine packet scheduling.
Check if the packet channel has reached a deficiency state. This is a programmable level per channel type, namely TCP, IP, ARP, raw buffer, which indicates how many times the channel was passed before the packet was sent, with the scheduler disabling the QoS level. .. If two or more packets reach the deficiency state at the same time, the channel with the higher weight is given priority. Other packets are then scheduled to be sent next. If the packets have the same priority weighting, they are sent in turn according to the following order: TCP / UDP, then ARP, then IP, then Raw Ethernet®.
If the channel does not have deficient packets, then the channel with the highest combined QoS level and channel weight is transmitted.
If only one channel has a packet to be transmitted, it will be transmitted immediately.
Once the packet channel is selected to transmit, the channel memory pointer, packet length, and type are forwarded to the DMA engine. When the transfer is complete, the DMA engine then returns the signal to the transmit scheduler. At this point, the scheduler sends the packet parameters to the DMA engine.
[DMA engine] The DMA engine 61 receives packet parameters from the transmit scheduler. Packet parameters include packet type, packet length, and start memory pointer. The DMA engine uses the packet length to determine how many data bytes are transferred from memory. The packet type tells the DMA engine from which memory buffer to retrieve the data, and the start memory pointer tells where to start reading the data. Since the output packet covers a large number of memory blocks, the DMA engine needs to understand how large each memory block used in the channel packet is. The DMA engine receives 64 bits of data from the memory controller at one time and supplies 64 bits of data to the transmitter interface at one time.
[Transmitter interface] The transmitter interface 62 takes output from the DMA engine and generates the macout_lock, macout_rdy, macout_eof, and macout_val_byte signals of the Ethernet (registered trademark) MAC interface. The 64-bit macout_data bus connects directly from the DMA engine to the Ethernet (registered trademark) MAC interface.
[Receiver interface] The receiver interface 63 operates to interface with an Ethernet® MAC interface. The receiver interface takes the data and feeds it along with the state count information to the address filter and packet type parser block.
Address Filter and Packet Type Parser The address filter and packet type parser 64 parses Ethernet (registered trademark) headers and performs two main functions: -Determine whether the packet is for the premises network stack. · Analyze the encapsulated packet type to determine where to send the remaining packets.
[Address filtering] The network stack can be programmed with the following filter choices: Reception of programmed unicast addresses, Reception of broadcast packets, Acceptance of multicast packets, Reception of addresses within the range specified by the net mask, -Promiscuous mode (accepts all packets).
All of these parameters can be set by the host computer via registers.
Supported packet types The following packet types are known by IT10G hardware and are uniquely supported. -IPv4 packets with type = 0x8000, ARP packets with type = 0 × 0806, -RARP packets with type = 0x8035.
The packet type parser also handles cases where a parameter with a length of 802.3 is included in the TYPE field. This case is detected when the value is less than or equal to 1500 (decimal). When this condition is detected, the type parser should send the encapsulated packet with an 802_frame signal claim to both the ARP and IP receiving modules and decode the packet with the recognition that the module is not really intended. Each subsequent module approves that.
Note: IPv6 packets are treated as exception packets by the Ethernet (registered trademark) layer.
FIG. 7 is a schematic block diagram of the address filter and the packet type parser module, and FIG. 8 is a timing diagram showing the operation of the address filter and the packet type parser module. At the I / O timing, the signal indicating the packet type is in the claimed state until the claim of the macin_lock signal corresponding to the packet is released. The All_packet signal also triggers only if the destination MAC address is acceptable.
The address filter and packet type parser module parses packets that it does not understand, and if unsupported types of characteristics become available, the packets are forwarded to an exception handler for storage and further processing. Will be done.
[Data Aligner] The data aligner 65 operates to align the data bytes of the subsequent packet processing layer. A data aligner is required because the Ethernet (registered trademark) header is not an even multiple of 64-bit. Depending on the presence of the VLAN tag, the data aligner directs the 64-bit data to the upper processing layer so that the data is MSB aligned. In this way, the payload section of an Ethernet® frame is always aligned on even 64-bit boundaries. The data aligner also operates to generate a ready signal on the next layer. The ready signal activates 2 or 3 operable cycles after macin_rdy is claimed. FIG. 9 is a schematic block diagram of the configuration data alignment module.
Ethernet (registered trademark) packet format IT10G receives both 802.3 (SNAP) and DIX format packets from the network, but only sends DIX format packets. In addition, when 802.3 packets are received, they are first converted to DIX format and then processed by an Ethernet® filter. Therefore, all Ethernet® exception packets are stored in DIX format.
[ARP Protocol and ARP Cache Module] [Overview] The following is a detailed description of the ARP protocol and the ARP cache module. In one embodiment of the IT10G architecture, the ARP protocol module also supports the RARP protocol, but does not include the ARP cache itself. This common resource is isolated from this ARP module because each module that can send the packet queries the ARP cache earlier than the others. The ARP protocol and the ARP cache module send updates to the ARP cache based on the packet type received.
ARP characteristic list: -It is possible to respond to an ARP request by generating an ARP reply. -A ARP request can be generated in response to the ARP cache. -It is possible to give ARP answers to multiple IP addresses (multi-homed host / ARP surrogate). -A target (unicast) ARP request can be generated. -Filter illegal addresses. -Transfer the aligned ARP data to the processor. Free ARP can be performed. -The CPU may bypass the automatic ARP reply generation and dump the ARP data to the exception handler. -The CPU may generate a custom ARP answer (in bypass mode). -Has a variable priority of ARP packets according to the network status.
RARP Characteristic List: -Request an IP address. -Request a specific IP address. -RARP requests are handed off to the exception handler. · Manage irregular RARP responses. -Send the aligned RARP data to the processor. -The CPU can generate custom RARP requests and replies.
ARP cache characteristics: Dynamic ARP table size, -Automatically updated ARP entry information, -Interrupt when the sender's hardware address changes, Ability to collect ARP data in a congested manner, -Duplicate IP address detection and interrupt generation, ARP request capability via ARP module, · Support for static ARP entries, An option that allows static ARP entries to be replaced by dynamic ARP data, ARP proxy support, · The configurable expiration time of the ARP entry. (The CPU can be either the host computer CPU or the on-chip processor in this context.)
[ARP module block diagram] FIG. 10 is a schematic block diagram of one configuration of the ARP module.
[ARP cache module block diagram] FIG. 11 is a schematic block diagram of one configuration of the ARP cache block.
[AR P module theory of operation] [Packet parsing] ARP module 100 processes only ARP and RARP packets. The module listens for a ready signal received from the Ethernet® receiving module. When the signal is received, the frame type of the incoming Ethernet® frame is checked. If the frame type is not ARP / RARP, the packet is ignored. Otherwise the module will start parsing.
The data is read from the Ethernet® interface in 64-bit words. ARP packets take 3.5 words. The first word of an ARP type packet contains almost static information. The first 48 bits of the first word of an ARP type packet contain the hardware type, protocol type, hardware address length, and protocol address length. These received values are compared with the values expected in the ARP request for IPv4 on Ethernet®. If the received values do not match, the data is sent to the exception handler for further processing. Otherwise, the ARP module continues the analysis. The last 16 bits of the first word of an ARP type packet contain the arithmetic code. The ARP module stores the operation code and checks whether it is valid, that is, 1, 2 or 4. If the arithmetic code is not valid, the data is sent to the exception handler for further processing. Otherwise, the ARP module will continue the analysis.
The second word of the ARP type packet contains the source Ethernet® address and half the source IP address. The ARP module stores the first 48 bits in the Source Ethernet® address register. The ARP module then checks if this field is a valid Source Ethernet® address. The address must not be the same as the address in the IT10G network stack. If the source address is not valid, the packet is dropped. The last 16 bits of the packet are then stored in the upper half of the source IP address register.
The third word of the ARP type packet contains the second half of the source IP address and the target Ethernet® address. The ARP module stores the first 16 bits in the lower half of the source IP address register and checks if the stored value is a valid source IP address. The address must not be the same as the IT10G hardware address or broadcast address. Also, the source address should be on the same subnet. If the source address is not valid, the ARP module drops the packet. If the packet is an ARP / RARP answer, compare the target hardware address with the Ethernet address. If the addresses do not match, the ARP module drops the packet. Otherwise, the ARP module continues the analysis.
Only the first 32 bits of the last word of an ARP type packet contain data (target IP address). The ARP module stores the target IP address in a register. If the packet is an ARP packet (as opposed to an ARP request or RARP packet), compare the target IP address with the IP address. If the addresses do not match, this packet is dropped. Otherwise, if this packet is an ARP request, it will generate an ARP reply. If this is a RARP answer, forward the assigned IP address to the RARP handler.
Once all address data is verified, the source address is transferred to the ARP cache.
[Outgoing packet] The ARP module sends packets internally from three sources: the ARP cache 110 (ARP request), the parser / FIFO buffer (of the ARP answer), and the system controller (of the custom ARP / RARP) or the host computer. Can receive requests for. Because of this condition, the priority-queue type is required to schedule the transmission of ARP / RARP packets.
Send requests are queued on a first-come, first-served basis, except when two or more entries want to send. In that case, the next request in the queue will follow its priority. RARP requests usually have the highest priority, followed by ARP requests. ARP replies usually have the lowest priority. The use of priorities allows resources to be shared according to data traffic.
There is one state in which the ARP response has the highest priority. This happens when the ARP answer FIFO buffer is full. When the FIFO buffer is full, incoming ARP requests begin to be discarded, so ARP replies should have the highest priority at that time to avoid having the ARP request retransmitted.
When the send queue is full, no further requests are made until one or more send requests are fulfilled and F (disappeared from the queue). When the ARP module detects a full queue, it requests an increase in priority from the transmit arbiter. This request signal can be single bit, as there are only two states in the queue, one that is full or one that is not.
When a transmit arbiter allows an ARP module to be transmitted, ARP / RARP packets are dynamically generated according to the type of packet being transmitted. The type of packet is determined by the instruction code, which is stored at each entry in the queue. Figure 12 shows the send queue entry format.
[Bypass mode] The ARP module has the option of bypassing the automatic processing of incoming packet data. When the bypass flag is set, incoming ARP / RARP data is transferred to the exception handler buffer. The CPU then accesses the buffer and processes the data. When in bypass mode, the CPU itself generates an ARP reply and transfers the data to the transmit scheduler. The fields that can be customized in the output ARP / RARP packet are the source IP address, source Ethernet® address, target IP address, and instruction code. All other fields match the standard values used by Ethernet® for ARP / RARP in IPv4, and the Source Ethernet® address is set to the address of the Ethernet® interface. (The CPU may be the host computer or on-chip processor in this context.) Note: If these other ARP / RARP fields need to be changed, the CPU must generate the raw Ethernet® frame itself.
[ARP cache behavior theory] [Adding ARP entry to ARP cache] ARP is generated when a targeted ARP request and answer (dynamic) is received, or when requested by the CPU (static). (The CPU may be the host computer or on-chip processor in this context.) A dynamic entry is an ARP entry that is generated when an ARP request or answer is received for one of the interface IP addresses. Dynamic entries typically exist for a limited time of 5 to 15 minutes when identified by an application program running on the user or host computer. Static entries are user-generated ARP entries that normally do not end.
The new ARP data comes from two sources: the CPU and the ARP packet parser via ARP registers. Dynamic ARP entries have priority when both sources request to add ARP entries at the same time because it is necessary to process incoming ARP data as quickly as possible.
Once the ARP data source is selected, it is necessary to determine where in the IT10G hardware memory the ARP entry will be stored. To do this, a look-up table (LUT) is used to map a given IP address into memory location. The search table contains 256 entries. Each entry is 16 bits wide and contains a memory pointer and a pointer valid (PV) bit. The PV bit is used to determine if the pointer points to a valid address, the start address of the memory block allocated by the ARP cache. Figure 13 shows the search table entry format.
Use an 8-bit index to determine where in the LUT the pointer should be retrieved. This index is taken from the last octet of the 32-bit IP address. The reason for using the last octet is that in a local area network (LAN), this is the part of the IP address that changes the most between hosts.
Once the slot to be used by the LUT is determined, it is checked whether the slot contains a valid pointer (PV = 1). If a valid pointer exists, it means that there is a block of memory allocated to this index, and the target IP address can be found in that block. At this point, the directed block of memory is searched and the target IP address is searched. If the LUT does not contain a valid pointer in this slot, memory must be allocated from internal memory, malloc1. Once memory is allocated, the address of the first word of the allocated memory is stored in the pointer field of the LUT entry.
After allocating memory and storing the pointer in the LUT, it is necessary to store the required ARP data. This ARP data contains the IP address needed to determine if this is an exact entry during a cache search. A set of control fields is also used. The retry counter is used to keep track of multiple ARP request attempts made at a given IP address. The type field indicates the type of cache entry (000 = dynamic entry; 001 = static entry; 010 = surrogate entry; 011 = ARP check entry). The resolution flag indicates that this IP address is properly resolved to an Ethernet® address. A valid flag indicates that this ARP entry contains valid data. Note; The entry is valid and unresolved while the first ARP request is being made. The src field shows the source of the ARP entry (00 = dynamic addition, 01 = system interface, 10 = IP router, 11 = both system interface and IP router). The interface field allows the use of many Ethernet® interfaces, but defaults to a single interface (0). Subsequent control fields are link addresses pointing to the following ARP entries: The most significant bit (MSB) of the link address is actually the flag link_valid. The link_valid bit indicates that there is another ARP entry following it. The last two fields are the Ethernet address where the IP address is resolved and the time stamp. The time stamp indicates when the ARP address was generated and is used to determine if the entry has expired. Figure 14 shows an example of the ARP cache entry format.
On LANs with more than 256 hosts or multiple subnets, conflicts between different IP addresses occur in the LUT. In other words, two or more IP addresses can be mapped to the same LUT index. This is for two or more hosts that have a given value in the last octet of their IP address. To deal with conflicts, the ARP cache uses chains, which are described below.
A LUT search is performed, and when it is found that an entry already exists in that slot, an ARP entry pointed to from memory is searched. Check the IP address of the ARP entry and compare it with the target IP address. If the IP addresses match, you can simply update the entry. However, if the addresses do not match, observe the Link_Valid flag and the last 16 bits of the ARP entry. The last 16 bits contain a link address pointing to another ARP entry that maps to the same LUT index. If the Link_Valid bit is claimed, look up the ARP entry pointing to the link address field. The IP address in the entry is compared with the target IP address again. If a match exists, the entry is updated, otherwise the search process continues until a match is found or the Link_Valid bit is no longer claimed (on subsequent links in the chain).
When the end of the chain is reached and no match is found, a new ARP entry is generated. Generating a new ARP entry requires memory allocation by the malloc1 memory controller. Each block of memory is 128 bytes in size. Therefore, each block can accommodate 8 ARP entries. When the end of the block is reached, a new memory block must be requested from malloc1.
As mentioned earlier, the user (or application running on the host computer) has the option of generating static or permanent ARP entries. The user has the option of allowing dynamic ARP data to replace static entries. In other words, when ARP data is received for an IP address for which a static ARP entry has already been generated, that static entry can be replaced with the received data. The advantage of this arrangement is that the static entries are out of date, allowing dynamic data to be overwritten on the static data, resulting in a newer ARP table. This update capability is disable if the user is confident that the IP-to-Ethernet address mapping is constant, for example remembering the IP and Ethernet (registered trademark) addresses of the router interface. Will be done. The user can also choose to keep static entries to minimize the number of ARP broadcasts on the LAN. Note: ARP surrogate entries can never be overwritten with dynamic ARP data.
Search for cache entries Searching for ARP cache entries follows a process similar to the ARP entry generation process. The search is initiated by using a LUT to determine if memory is allocated to a given index. If memory is allocated, memory is searched until an entry is found (a cache hit occurs) or an entry with the link_valid flag set to zero (cache loss) is encountered.
If there is a cache loss, an ARP request will be made. This involves creating a new ARP entry for the cache and a new LUT entry, if necessary. The new ARP entry remembers the target IP address, the resolution bit is set to zero, and the significant bit is set to 1. The request counter is set to zero as well. The entry is then time stamped and the ARP request is sent to the ARP module. If no answer is received after 1 second, the request counter is incremented and another request is sent. After sending three requests and not receiving an answer, the attempt to resolve the target IP is abandoned. Note: The retry interval and the number of request retries can be configured by the user.
When a cache loss occurs, the request module is notified of the loss. This gives the CPU or IP router the opportunity to decide to wait for an ARP reply to the current target IP address, or to start a new search for another IP address to place the current IP address at the back of the queue. Be done. This minimizes the impact of cache loss in the configuration of many connections. FIG. 15 is a flow chart showing the ARP search process.
If a matching entry is found (cache hit), the resolved Ethernet (registered trademark) address is returned to the module requesting the ARP search. Otherwise, if the target IP address is not found in the cache and all ARP request attempts time out, the request module is notified that the target IP address has not been resolved.
Note: When an ARP discovery request from an IP router fails, the router must wait at least 20 seconds before starting further discovery of that address.
[Cache initialization] Some components are reset when the ARP cache is initialized. The look-up table (LUT) is cleared by setting all PV bits to zero. All currently used memory is deallocated and released back to the malloc1 memory controller. The ARP expiration timer is also set to zero.
No ARP request is made during the initialization period. Also, attempts to generate an ARP entry from the CPU (static entry) or from received ARP data (dynamic entry) are ignored or discarded.
[ARP entry expired] Dynamic ARP entries can only exist in the ARP cache for a limited amount of time. This prevents any IP-to-Ethernet address mapping from being revoked. Old address mapping occurs if the LAN uses DHCP to assign an IP address or if the device's Ethernet interface is modified during a communication session.
A 16-bit counter is used to keep track of time. A counter operating at a clock frequency of 1 Hz is used to track the number of seconds elapsed. Each ARP entry contains a 16-bit timestamp taken from this counter. This time stamp is taken when the IP address is properly resolved.
The expiration of the ARP entry occurs when the ARP cache is idle, that is, when no search or request is currently processed. At this time, the 8-bit counter is used to cycle through and search the LUT. Each slot in the LUT is checked to see if it contains a valid pointer. If the pointer is valid, the pointing memory block is searched. Each entry in the block is then checked to see if the difference between its timestamp and the current time is greater than or equal to the maximum lifetime of the ARP entry. If other memory blocks are unchained from the first memory block, the entries contained in these blocks are also checked. Once all entries related to a given LUT index have been checked, the next LUT slot is checked.
If it is found that the entry has expired, the valid bits in the entry are set to zero. If no other entry exists in the same memory block, the block is deallocated and returned to malloc1. If the deassigned block is the only block associated with a given LUT slot, then the PV bit for that slot is also set to zero.
[Run ARP Deputy] The ARP cache supports surrogate ARP entries. ARP surrogate is used when this device acts as a router for LAN traffic or when there is a device on the LAN that cannot respond to ARP queries.
When ARP delegation is enabled, the ARP module passes requests for IP addresses that do not belong to that host to the ARP cache. The ARP cache then searches to find the target IP address. If it finds a match, it checks the type field of the ARP entry to determine if it is a surrogate entry. If it is a surrogate entry, the ARP cache returns the corresponding Ethernet® address to the ARP module. The ARP module then uses the Ethernet (registered trademark) address found in the surrogate entry as the source Ethernet (registered trademark) address to generate the ARP answer. Note: ARP surrogate search is only done for incoming ARP requests.
[Duplicate IP address detection (ARP check)] When the system (host computer and IT10G hardware) first connects to the network, the user or application running on the host computer uses one of the IP addresses assigned to that interface by any other device on the network. You should make a free ARP request to test if you have. If two devices on the same LAN use the same IP address, this causes problems with routing packets on the two hosts. A free ARP request is a request for a host-specific IP address. If no answer to the query is received, it can be assumed that no other host is using the IP address on the LAN.
The ARP check is started in a manner similar to that of performing an ARP search. The only difference is that the cache is destroyed once the ARP request is complete. If no answer is received, the entry is removed. If an answer is received, an interrupt is generated to inform the host computer that the IP address is being used by another device on the LAN, and the entry is removed from the cache.
[Cache access priority] Different tasks have different priorities for accessing ARP cache memory. The search for surrogate entries has the highest priority because it requires a quick response to the ARP request. The second priority is to cache dynamic entries, incoming ARP packets should be received at a very high rate and processed as quickly as possible to avoid retransmissions. The ARP search from the IP router has the next highest priority, followed by the search by the host computer. Manual generation of ARP entries has a second lowest priority, expired cache entries have the lowest priority, and occurs whenever the cache does not process ARP searches or generate new entries.
[IP module] [Overview] The UT10G uniquely supports IPv4 packets that automatically parse all types of received packets.
[IP module block diagram] FIG. 16 is a schematic block diagram of one configuration of the IP module block.
[Description of IP submodule] [IP parser] The IP parser module 161 analyzes the received IP packet and operates to determine where to send the packet. Each received IP packet can be sent to a TCP / UDP module or exception handler.
[Analysis of IP header fields] Only IPv4 is received and parsed by the IP module, therefore this field must be 0x4 processed. If an IPv6 packet is detected, it is managed and handled as an exception by the exception handler. Any packet with a version smaller than 0x4 is considered malformed (illegal) and the packet is dropped.
[IP header length] The IP header length field is used to determine if any IP option is present. This field must be 5 or greater. If it is smaller than that, the packet is considered malformed and dropped.
[IP TOS] This field is not parsed or maintained in received packets.
[Packet Len] This field is used to determine the total number of bytes in a packet received and to indicate where the end of that data section is in the next level protocol. After this count expires, all data bytes received before the ip_packet signal breaks the claim are assumed to be embedded bytes and are implicitly discarded.
[Packet ID, flag, fragmentation offset] These fields are used to defragment packets. Fragmented IP packets can be handled by dedicated hardware or treated as exceptions and handled by exception handlers.
[TTL] This field is not parsed or maintained in received packets.
[PROT] This field is used to determine the protocol to be encapsulated next. The following protocols are fully supported (or partially supported in other embodiments) in hardware.<tables num="4"><img file="JP4875126B2_D0004.tif" /></tables>
Packets can be sent to the host computer if any other protocol is received and the unsupport_prot characteristic is enabled. A protocol filter can selectively allow a protocol to be received. Otherwise, the packet is implicitly dropped.
[Checksum] This field is not parsed or maintained. This is simply used to ensure that the checksum is correct. If the checksum turns out to be inadequate, a bad_checksum signal is claimed that goes to all the next layers. This is a state of claim until it is responded.
[Source IP address] This field is parsed and sent to the TCP / UDP layer.
[Destination IP address] This field is parsed and checked against a valid IP address that the local stack should respond to. This takes more than one clock cycle, in which case the analysis must continue. If the packet turns out to be misguided, the bad_ip_add signal is claimed. This is a state of claim until it is responded.
[IP ID generation algorithm] The on-chip processor can set the IP ID seed value by writing any 16-bit value to the IP_ID_Start register. The ID generator takes this value and performs a 16-bit mapping to generate the IP ID used by different requesters. On-chip processors, TCP modules, and ICMP echo response generators can all request an IPID. A block diagram of one structure of the ID generator is shown in Figure 17.
The IP ID seed register is incremented each time a new IP ID is requested. The bitmapper block rearranges the IP ID_Reg value so that the IP ID_Out bus is not a simple increment value.
[IP injector module] The IP injector module is used to inject packets from the on-chip processor into the IP and TCP modules. The IP injector control registers are located in the IP module register space and these registers are programmed by the on-chip processor. A block diagram showing the data flow of the IP injector is shown in Figure 18.
As you can see from the figure, the IP injector can insert data under the IP module. To use IP injection, the on-chip processor programs the IP injector module with the starting address of the memory in which the packet resides, the length of the packet, and the source MAC address. The injector module interrupts when it completes sending a packet from the on-chip processor's memory to the stack.
[TCP / UDP module] [Overview] This section describes the TCP module, which manages TCP and UDP transport protocols. The TCP module is divided into four main sections: socket send interface, TCP send interface, TCP receive interface, and socket receive interface.
[Feature list] The following is a list of TCP characteristics of IT10G. 64K socket support Inappropriate support for TCP Slow Start High Speed Retransmission / High Speed Recovery Selectable Nagle algorithm window scaling Selective ACK (SACK) Protection against wrapped sequence numbers (PAWS) Timestamp support keepalive timer [Window Scaling] IT10G supports window scaling. Window scaling is an extension that expands the TCP window specification to 32 bits. This is specified in RFC 1323 Section 2 (see http://www.rfc-editor.org/rfc/rfc1323.txt). Window scale behavior is based on three variables. The first is the TCP_Control1 SL_Win_En bit (which enables window scale), the second is the TCP_Control3 sliding window scale (which sets the scaling factor), and the last is the value to be scaled. The WCLAMP parameter that determines.
Without the SL_Win_En bit, the hardware will not attempt to negotiate window scaling via TCP window scale options during a TCP-3 direction handshake.
[TCP dump mode] TCP Dump Mode is an IT10G hardware mode that allows support for widely used diagnostic programs such as TCP Dump and other packet monitoring programs. When TCP dump mode is enabled, all received packets are sent to the host as exceptions, and all output TCP / UDP packets coming from the hardware stack are looped back as exception packets.
The driver makes a copy of these packets to the network monitor, reinjects the received packet, and sends the transmitted packet as a raw Ethernet® frame.
[Host ACK mode] Host ACK mode is an IT10G hardware mode that sends only a TCP ACK when the host computer receives data from a TCP segment. Host ACK mode waits for the DMA in the MTX buffer containing the data segment to complete before sending the ACK. If data is corrupted when passing between the host computer and the integrated network adapter or vice versa, host ACK mode guarantees additional data integrity.
[Time stamp] IT10G supports time stamps. Timestamp is an enhancement to TCP as specified in RFC1323 (see http://www/rfc-editor.org/rfc/rfc1323txt). Timestamps allow TCP to better calculate RTT (Round Trip Time) measurements and are used to support PAWS (Protection against Wrapped Sequences).
[PAWS] PAWS (protection against wrapped sequences) is specified in RFC1323 (see http://www/rfc-editor.org/rfc/rfc1323txt). PAWS is protection against old duplicate segments that break TCP connections. This is an important feature for high speed links.
[TCP host retransmission mode] TCP host retransmission mode allows data to be retransmitted directly from the host's memory buffer rather than from a buffer located on the integrated network adapter. This allows the amount of memory required by the integrated network adapter to be reduced.
[Initial sequence number generation] The occurrence of the initial sequence number needs to be kept secret. RFC1948 points out the weaknesses of RFC793's original initial sequence number specification and recommends several alternatives. The integrated network adapter uses an optimized method that works according to RFV1948, but is efficient to configure in hardware.
[Double stack mode] The dual stack mode allows the hardware TCP / IP stack integrated on the network adapter to work together with the host computer's TCP / IP stack.
Dual stack mode allows an integrated network adapter to support the coexistence of software stacks operating in parallel using the same IP address.
Dual stack mode requires two basic hardware features of the integrated network adapter. The first hardware feature is the SYN state message mode. In SYN state message mode, any received SYN raises a state message to the host computer, and SYN / ACK is the integrated network adapter hardware until the host computer returns the appropriate instruction block to the integrated network adapter hardware. Not caused by hardware. If SYN status message mode is not enabled on the integrated network adapter, SYN / ACK is automatically generated by the integrated network adapter and no status message is generated.
The second hardware feature required by dual stack mode is suppression of RST messages from the integrated network adapter hardware when TCP packets that do not match the integrated network adapter control block database are received. In this case, instead of automatically raising a RST, the integrated network adapter hardware hosts the packet as an exception packet to allow the host computer's software TCP / IP stack to treat this packet as an exception packet. Must be sent to the computer.
[IP ID split] IP (Internet Protocol) ID (Internet Identification) partitioning is part of the dual stack support package. IP ID splitting allows host computers and integrated network adapters to share IP addresses without IP ID overlap. When IP ID splitting is switched off, the integrated network adapter uses the entire 16-bit ID range (0-255). When IP ID splitting is switched on, the IP ID bit [15] is set to 1 to allow the host computer software's TCP / IP stack to use half of the IP ID range (ie, integrated network adapter). Uses 128-255 and the host computer software TCP / IP stack uses 0-127 as the IP ID).
[Custom Filtering] There are several places in the hardware that have custom filters that can be used to limit, approve, or perform special behavior on certain types of packets. Ethernet (R) filtering can take the following attributes: -Receiving a programmed unicast address, Receive broadcast packets, -Receiving multicast packets, Receiving addresses within the range specified by the netmask, -Enables indiscriminate mode (accepts all packets).
[VLAN support] VLAN support consists of several optimized hardware elements. One hardware element removes the VLAN header from the input packet, the second optimized hardware element generates a VLAN-tagged output packet, and the third optimized hardware element is the input SYN. The VLAN parameter is generated from the frame, and the fourth optimized hardware element passes the VLAN tag information of the exception packet and UDP packet.
[Jumbo frame support] Jumbo frames are larger than regular 1500 byte size Ethernet® frames. Jumbo frames smaller than the network bandwidth used for header information allow for increased data processing capacity in the network. The integrated network adapter uses hardware optimized to support jumbo frames up to 9 kbytes.
[SNMP Support] SNMP is a form of higher-level protocol that allows remote monitoring of network and hardware adapter performance statistics.
[MIB support] The integrated network adapter includes optimized hardware support for a large number of statistical counters. These statistical counters are specified by the standard Management Information Base (MIB). Each of these SNMP MIB counters tracks what happens (received, transmitted, and dropped packets) on the network and integrated network adapters.
[Memory check] Memory error checking and correction (ECC) is similar to partial checking. However, ECC can detect and correct single-bit memory errors and detect double-bit errors even if the parity check can only detect single-bit errors.
Single-bit errors are most common and are characterized by a single bit of data that is inaccurate when reading a complete byte or word. A multi-bit error is the result of an error in two or more bits within the same byte or word.
ECC memory uses additional bits to store the encrypted ECC code along with the data. When the data is written to memory, the ECC code is also stored. When the data is read back, the stored ECC code is compared to the ECC code that was generated when the data was written. If the ECC codes do not match, you can determine which bit in the data is in error. The error bit is "inverted" and the memory controller releases the corrected data. Errors are corrected on-the-fly and the corrected data is not returned to memory. If the same corrupted data is read again, the correction process is repeated.
The integrated network adapter allows ECC to be programmed and selected in a flexible way to protect both packet data and control information within the adapter.
[Heritage Mode] Heritage mode allows all network traffic to be sent to the host computer regardless of traffic type. These modes allow the integrated network adapter to act as if the hardware TCP / IP stack does not exist on the adapter, and the modes are often referred to as dump NICs (network interface cards).
[IP Fragmentation] Reassembly of IP fragmented packets is not handled by the integrated network adapter. IP fragmented packets are passed as exception packets, reassembled in the driver, and then "reinjected" into the integrated network adapter via IP injection mode.
[IP injection] IP injection mode allows IP packets (eg, reassembled IP fragments or IPsec packets) to be injected into the hardware TCP / IP stack of the integrated network adapter.
The injection control register is used to inject IP packets into the TCP / IP stack in the integrated network adapter. A feature of the injection control register allows the host computer to control the injection and inject IP packets into the hardware TCP / IP stack in the integrated network adapter. The injection control register therefore allows injection of SYN, IPSec, or other packets into the integrated network adapter. The injection control register is also part of the TCP dump mode function.
[NAT, IP masquerade, port forwarding] NAT, IP masquerading, and port forwarding are supported on integrated network adapters via a port range register that forwards all packets of a particular type of UDP or TCP that fall within the programmable range of a port to an exception path. Port registers allow a range of ports to be used in network control operations such as NAT, IP masquerading, and port forwarding.
[Multiple IP addresses] The integrated network adapter hardware supports a range of up to 16 IP addresses that can respond as its unique IP address. These address ranges are accessible as a masked IP address base. This allows IP addresses to be extended to the range of IP addresses. This allows the integrated network adapter to perform multihoming or the ability to respond to multiple IP addresses.
[IP debug mode] When the test and control bits are available on the integrated network adapter, all IP packets are sent to the host computer with the exception. This mode is designed for diagnostic purposes.
[Waiting for time] Time wait is the final state of a passive mode TCP connection, and the time wait state and its behavior are described in RFC793, see http://www.rfc-editor.org/rfc/rfc793.txt. I want to.
[Virtual Socket] The integrated network adapter hardware supports a variable number of sockets or connections, and the current configuration supports sockets up to 645535.
The integrated network adapter provides optimized hardware support (integrated with the TOE) to transfer the connection between the adapter and the host computer.
When the integrated network adapter hardware approves a connection equal to its maximum capacity, the next SYN is passed to the host as an exception packet and the host can handle this connection. When to the host to open this connection, use part of the dual stack mode that allows TCP packets that do not match the hardware database to be forwarded to the host as exception packets instead of answering RST. It should be noted that the hardware needs to be configured in.
[Survival time] TTL or time-to-live is an IP address parameter that limits the lifetime of IP packets on the network to the number of hops (hops are jumps across Layer 3 devices such as routers). The integrated network adapter hardware sets this value for the output frame to limit the lifetime of the transmitted packet. For more information, see RFC791 section Lifetime at http://www.faqs.org/rfcs/rfc791.html.
[Keepalive] Keepalives are described in Section 4.2.3.6 of RFC1122. See http://www.faqs.org/rfcs/rfc1122.html. Keep-alive allows idle TCP connections to remain connected and not time out by periodically sending keep-alive packets across the link.
[ToS] The TOS or service type is an IP address parameter that can be used by a router to prioritize IP packets. TOS parameters need to be adjustable in the system and socket layer. At the time of transmission, the integrated network adapter hardware sets the TOS through several registers. For more information, see http://www.faqs.org/rfcs/rfc791.html for service types in the RFC791 section.
[QoS] The integrated network adapter hardware supports four transmit queues to allow QoS of transmitted data. Each socket can be assigned a QoS value. This value can also be used to map to the VLAN priority level.
The transmission operation can be summarized as follows. The TCP transmit data flow begins with the socket query module, which traverses the transmit data valid bit table and looks for entries with those transmit data valid bit sets. When it finds such an entry, the socket query module places it in one of four queues according to the socket's user priority level. Sockets with priority level 7 or 6 are placed in queue list 3, levels 5 and 4 are placed in queue list 2, levels 3 and 2 are placed in queue list 1, and levels 1 and 0 are waiting. Placed in queuing list 0.
These QoS characteristics, which use priority queues, allow the parallel use of large numbers of hardware modules according to outgoing traffic.
[Failover] Failover between network adapters is supported by NO_SYN mode. NO_SYN mode allows a socket to be created without attempting to initiate a connection. This allows the integrated network adapter hardware socket and all its associated data structures to be generated without a connection. NO_SYN mode allows failover support from the migration of another card or connection from the software TCP / IP stack to the integrated network adapter.
[Top level block diagram] Figure 19 shows a top-level block diagram of one TCP module configuration. The Socket Control Block (CB) 191 contains information, state, and parameter settings specific to each socket connection and forms the most important or key part of the Virtual Socket (VSOCK) architecture. The locking mechanism is installed so that one module works on the CB while the other module does not.
[TCP receive submodule] Figure 20 shows the TCP receive 200 data flow.
[Overview] For normal IP traffic, packets are received over a 64-bit TCP Rx data path. The packet header proceeds to the TCP parser module, and the packet data is routed to the received data memory controller 201. In IP fragmented traffic, data is received via memory blocks and header information is received via normal paths. This makes the memory block from IP fragmentation look similar to the data block written by the received data memory controller. CPU data also uses memory blocks to inject received data through the received data memory controller. This architecture provides maximum flexibility in handling normal and fragmented traffic, which allows performance to be optimized for normal traffic while supporting fragmented traffic.
The receive TCP parser 202 parses the header information and operates to pass the parameters to VSOCK 203 and the receive state handler 204. If the incoming TCP parser does not know what to do with the data, it is forwarded to exception handler 205. In addition, the incoming TCP parser can be programmed to send all data to the exception handler.
The VSOCK module takes local and remote IP and port addresses and returns the pointer to the control block.
The NAT & IP masquerade module 206 determines whether the received packet is a NAT or IP masquerade packet. If yes, the packet is passed to the host system in raw (complete packet) format.
The receive state handler keeps track of the state of each connection and therefore updates its control block.
[Incoming TCP parser] The receive TCP parser registers the packet header information with other modules in the TCP receive part of the IT10G network stack and sends it. The receive TCP parser module also contains registers that are required to inject data from the CPU (on-chip processor or host computer) into the receive stream. The CPU must set up a memory block and then program the information in the receive TCP parser register. The incoming TCP parser generates a partial checksum in the TCP header, attaches this partial checksum to the partial checksum from the received data memory controller, and adds the resulting full checksum to the TCP header. Compare with checksum. For fragmented packets, the incoming TCP parser checks the checksum of the TCP header against the checksum sent by IP fragmentation in the last fragment.
Note: The IP module must set the IP fragmentation bit and insert the first memory block pointer, last memory block pointer, index, and partial checksum into the data stream of the appropriate packet fragment. Also, TCP reception requires IP protocol information to calculate pseudo-headers.
[Received data memory controller] The received data memory controller transfers data from the 64-bit bus between the IP and TCP modules to the data memory block of RxDRAM. There are two transfer modes. Normal mode is used to store TCP data in memory blocks. The raw mode is used to store the entire packet in a memory block. Raw mode is used for NAT / IP masquerade. In exception handling, normal mode is used with registers to transfer CB data to the CPU.
[VSOCK] The VSOCK module passes local and remote IP and port addresses from the packet and returns a socket open or TIME_WAIT (TW) control block (CB) pointer back to the receive state handler. VSOCK performs hash calculations on IP and port addresses and produces a hash value that acts as an index to the open / TW CB lookup table (LUT) 207. The LUT entry at that position holds a pointer to the open 208 or TW209 control block.
The pointer from the open / TW CB LUT points to the first control block in the linked list of zero or more CBs, each with a different IP and port address, but with the same hash number (resulting from a hash collision). Produces. VSOCK goes through this chain and compares the packet IP and port address with the entries in the chained CB until a match is found or the end of the chain is reached. If a match is found, the pointer to the CB is passed to the receive state handler. When the end of the chain is reached, VSOCK notifies the TCP parser of the error.
The chain of CBs connected to the open / TW CB LUT entry contains the open CB and the TIME_WAIT CB. The open CB is the first in the chain. There is a maximum number of open CBs as determined by the incoming TCPMaxOpenCBperChain. The TW CB is chained after the open CB. There is also a maximum number of TW CBs per chain. The open CB is generated when the three-way handshake is completed, and the HO CB is moved to the open CB by the receive state handler. The TW CB is generated from the open CB by the receive state handler when the last ACK is sent in the FIN sequence. If there is no more room in either case, the error is returned to the receive state handler.
The CB cache of the open CB is composed of open CBs instead of presetting the number of links from the LUT entry. The CB bit is set when it is in the cache. The cache is searched in parallel for hash / LUT operations. Figure 21 shows the control block search resolution flow of the VSOCK / reception status handler.
[Rcv state handler] If a SYN is received, a 12-bit hash is done in addition to the VSOCK call (doing a 17-bit hash to free or search for the TIME_WAIT control block), and the destination port is checked against the allowed port list. Will be done. If the port is enabled and VSOCK does not find a matching open / TW CB, the hash result will be used as an index into the HO CB table. If VSOCK finds an open or TIME_WAIT CB, a Dup CB error can be sent to the host computer and the SYN will be dropped. If an entry with a different IP and port address already exists in the HO CB table, the new packet information will be overwritten with the old information. This allows resources to be held in a SYN flood denial for service (DOS) attacks. Overwriting is HO Eliminate the need for CB table aging. Connections that have already been SYN / ACKed can be implicitly dropped. The pointer to the CB is transferred to the receive status handler. Only connections opened by the remote side (the local side receives SYN instead of SYN / ACK) enter the HO CB table. Connections opened by the local side are tracked by the open CB.
If an ACK is received, a 12-bit hash is done and VSOCK is called. A three-way handshake for the connection if there is a hit in the HO CB via a 12-bit hash, but VSOCK does not find an open or TW CB, and if the sequence and ACK number are valid. Is complete and the CB is transferred to the open CB table by the receive state handler. If VSOCK finds an open or TW CB, but there is no match with the 12-bit hash, the ACK is checked by the receive state handler for a valid sequence and ACK number, and a duplicate ACK.
Once the VSOCK module finds the correct socket CB, other appropriate information is read and updated by the receive state handler. TCP data is stored in large (2 kbytes) or small (128 bytes) memory buffers. A single segment can span a memory buffer. When one size buffer runs out, the other size is used. When data is received in a given socket, its Data_Avail bit in the socket hash LUT is also set.
If the receive state handler determines that an RST packet is needed, the receive state handler forwards the appropriate parameters to the RST generator 210. If SYN / ACK or ACK is required, the receive state handler sends a CB handle to the receive / send FIFO buffer 211.
[RST generator] The RST generator takes the MAC address, four socket parameters, and the sequence number received in the packet requiring the RST response to form the RST packet. The RST generator first requests a block from MTX memory to create a packet. RST packets are always 40 bytes long, so these packets fit in MTX blocks of any size. The RST generator always requests the smallest block available, typically a 128-byte block. RST packets have their IP IDs fixed at 0x0000 and their DF bits set in the IP header (not fragment bits).
After the RST generator configures the RST packet, the RST generator stores the start address of the MTX block containing the RST packet in the RST send queue. This queue is formed in M1 memory. A block of M1 memory is requested and used until it is full. The last entry in each M1 block points to the address of the next M1 block to be used. Therefore, the RST queue can grow dynamically. Since the MTX block address is only 26 bits, the RST generator accesses 32 bits of M1 memory at a time. The length of this queue can grow to the size that M1 memory is available. If no more memory is available for the queue, the RST generator implicitly drops the RST packet request from the RCV state handler. This has a network effect similar to dropping RST packets in transmission. This does not have a serious impact on performance as there are no connections.
The output of the RST send queue is given to the TCP send packet scheduler. When the transmit scheduler indicates to the RST generator that an RST packet has been sent, the MTX block used for the RST packet is released. The M1 memory block is released when all entries for the M1 memory block are sent and the link access to the next M1 block is read. The basic data flow is shown in Figure 22.
[Receive / Send FIFO Buffer] The receive / transmit FIFO buffer is used to queue the SYV / ACK and ACK that the receive state handler has determined to be required to transmit in response to the received packet. The reception status handler passes the following information to the receive / transmit FIFO buffer. CB address (16 bits) including socket information, CB type (2 bits; 00 = half open, 01 = open, 10 = Time_Wait), -The message type to be sent (1 bit, 0 = SYN / ACK, 1 = ACK).
The receive / transmit FIFO buffer entry is 4 bytes long and is stored in various memory buffers. Currently, this buffer is allocated 4K bytes, which provides a receive / send FIFO buffer with a depth of 1K entry. The output of the receive / transmit FIFO buffer is given to the SYN / ACK generator.
[SYN / ACK generator] The SYN / ACK generator 212 takes the information output from the receive / transmit FIFO buffer, searches for other appropriate information from the identified CB of half-open, open or Time_Wait, and the desired packet of SYN / ACK or ACK. To create. The SYN / ACK generator first requests a block from MTX memory to make a packet. SYN / ACK and ACK packets are always 40 bytes long, so these packets always fit into MTX blocks of any size. The SYN / ACK generator always requests the smallest block available, typically a 128-byte block.
After the SYN / ACK generator creates a SYN / ACK or ACK packet, the SYN / ACK generator puts the starting MTX block address in a 16-depth queue, which then feeds the TCP transmit packet scheduler. To do. If the receive / transmit FIFO buffer passes through a high programmable watermark, the transmit packet scheduler is notified of the status and increases the transmit priority of these packets.
[NAT and IP masquerade] The NAT and IP masquerade block 206 operates in parallel with the VSOCK module. The NAT and IP masquerade blocks decrypt incoming packets and see if they are in a pre-specified NAT or IP masquerade port range. If yes, a signaling mechanism is used to indicate to VSOCK that this is a NAT packet. When this happens, the entire packet is stored as if it were in the receive memory buffer. The packet is then forwarded to the host system at several points. The host system driver then performs the routing function, replacing the header parameters and sending them to the appropriate network interface.
[Exception handler] The exception handler sends the packet to the host computer that is not handled by the IT10G core.
[Support circuit] [rset occurrence] The rst signal is simultaneously generated from the top level rset signal.
[occurrence of tcp_in_rd] tcp_in_rd is used to freeze the input data from the IP. This is claimed when the DRAM write request is not immediately approved. It also occurs when DRAM runs out of small or large memory blocks.
[Word counter] The word counter, word_cnt [12: 0], is zero for resets and software resets. This increments when tcp_in_rd is active and ip_in_eof is not active. It is zero in the first word of ip_in_data [63: 0] and one in the second word. This is zero in the first word after ip_in_eof and remains zero between valid ip_in_data words.
[Memory block control circuit] [Holding memory block] The memory block control circuit keeps small and large memory blocks available at all times. This ensures that there is little delay when the data has to be written to a block of memory. The memory block control circuit also makes a block request in parallel with writing data. These reserved memory blocks are initialized from reset.
Initialization and memory block size selection TCP or UDP segment parameters are initialized. The size of the memory block used is determined by the TCP length information from the IP and the TCP header length information from the parser. If the data size (TCP length-TCP header length) fits a small memory block, the small memory block on hold is used and another small memory block is requested to replenish the hold. Otherwise, a large block of memory on hold is used and another large block of memory is requested to replenish the hold. If a small memory block is not available, a large memory block is used. However, if a large memory block is required but not available, then a small memory block is not used.
[Write aligned TCP data to a memory block] If there are an odd number of optional halfwords (32 bits wide each) in the TCP header, the data in the TCP packet is aligned and the data starts on a 64-bit boundary. If the TCP packet data is aligned, the TCP packet data can be placed directly in the memory block when the data is sent from the IP block. The address of the second block for the segment is sent to the TCP state machine. Similar to the data left in the TCP segment, the count is maintained in the space left in the block. If the memory block is already full, the record must also be maintained. If the previous block was full when the end of the TCP segment was reached, that block must be linked to the current block. Also, the link in the current block header is cleared, and the data length and the operating checksum of the data are written in the block header. The data length is a function of the number of bytes in the last 64-bit word, as determined by the bits in ip_in_bytes_val. If there is no room in the block before the end of the TCP segment, the data length and working checksum are written in the block header and flagged to indicate that the block is finished. The rest of the data in the TCP segment is used to determine whether large or small reserved memory blocks are used. When the block size is exhausted, the same rules as described above are used. The address of the last memory block must be sent to the TCP state machine.
[Write unaligned TCP data to a memory block] To store the first lo32-bit halfword from the IP block if the data in the TCP segment is unaligned (ip_in_data [63: 0] contains data written to two different memory blocks) First, there must be an extra cycle, which allows the data to be written as high 32-bit halfwords in memory blocks. During the next bus cycle from the IP block, the high 32-bit halfword is written as the low 32-bit halfword in the same cycle as the stored halfword. Count and checksum calculations must also be adjusted to handle this condition. Otherwise, the unaligned data is processed in the same way as the aligned data and has the same end cases as described above.
[Write UDP data to memory block] UDP data is always aligned, so UDP data is processed in the same way as TCP aligned data. The same end case applies.
[Checksum calculation] The checksum is calculated as described in REC1071. In the checksum calculation block, the checksum is calculated only for the data. The parser calculates the header checksum, and the TCP state machine combines the two checksums to determine how to handle packets with checksum errors.
[Analysis of fixed header fields] [Word counter] The word counter, word_cnt [12: 0], is zero for resets and software resets. The counter increments when tcp_in_rd is active and ip_in_eof is not active. The counter is zero in the first word of ip_in_data [63: 0] and one in the second word. The counter is zero in the first word after ip_in_eof and remains zero between valid ip_in_data words.
[Latch of remote_jp_add [31: 0] and local_ip_indx [3: 0]] These latches are zero for resets and software resets. remote_jp_add is latched from src_ip_add [31: 0] when src_ip_add_valid, ip_in_dav, tcporudp_packet are claimed. Local_jp_indx is latched from dest_ip_index [3: 0] when dest_ip_add_valid, ip_in_dav, tcporudp_packet are claimed. Both are valid until reloaded in the next packet.<tables num="5"><img file="JP4875126B2_D0005.tif" /></tables>
Note: The field is cleared or initialized with the first word from the reset and IP (in parentheses). Note: Latch is also eligible by tcp_packet, udp_packet & ip_in_dav.
[Checksum, error checking, optional header field parsing] [Checksum] The RFC1071 checksum is a header length field with an IP source address, an IP destination address (not an index), an 8-bit zero-leading (= 0006 hexadecimal) 8-bit type field, and a pseudo-header containing the TCP length. Calculated in the header as specified in.
[Header length check] Once the rx_tcp_hdr_len field is parsed (word count> 0001), the rx_tcp_hdr_len field is checked. The rx_tcp_hdr_len field cannot be less than 5 or greater than MAX_HDR_LEN set to 10. If so, the error rx_tcp_hdr_len signal is claimed and remains in that state until the start of the next packet.
[TCP length check] The IP payload length from IP is checked. The IP payload length cannot be less than 20 bytes (used when rx_tcp_header_len is not yet valid) and cannot be less than rx_tcp_header_len. If so, the error rx_tcp_hdr_len signal is claimed and remains in that state until the start of the next packet.
[Analysis of optional header fields] The field is cleared with a reset and the first word from the IP block. Only timestamps (10 bytes + padding), MSS (4 bytes), and window scales (3 bytes + padding) are searched. There can only be 20 bytes of options, so the header length is only 10 32-bit half-words (word count = 0004).
Options are parsed in a direct way. The option starts with the second half word count = 2 and is aligned to 32 bits. The first byte of the option identifies the option type.
If the timestamp option is detected, rx_tcp_timestamp, rx_tcp_tx_echo and rx_tcp_timestamp_val are loaded.
If the MSS option is detected, rx_tcp_remote_mss and rx_tcp_remote_mss_val are loaded.
If the window scale option is detected, rx_tcp_remote_win_scl and rx_tcp_rem_win_scl_val are loaded.
These remain valid until the start of the next packet.
[Socket receiving submodule] The socket receive submodule handles the interface between the IT10G and the system of received data (host computer). FIG. 23 shows the socket received data flow.
The socket receive submodule process starts at receive logical unit 230 and sets 1 bit in the socket receive DAV bitmap table 231. The socket receive DAV bitmap table has 1 bit associated with each 64K socket (hence the table is 8K bytes). By knowing the position of CB, the appropriate bit is set.
Socket_DAV query module 232 is a block that continuously scans the socket receive DAV bitmap table. When the Socket_DAV query module finds a set of bits, it generates the corresponding CB address and checks the CB structure to see if it contains a valid link_list block. A valid link_list block consists of a 64-bit memory address and a 16-bit length. If the CB has a valid link_list block, the CB address and valid link_list information will be DMA via a two-stage pipeline register pair. Passed to Prep module 233. The Socket_DAV module also clears the corresponding bits of the CB at that time. If the CB does not contain a valid link_list block, a status message is raised for the socket to inform the host that data is available on the socket but no valid forwarding block information exists on that socket. To. In this case, the corresponding bits in the bitmap table have not yet been cleared. The CB can also be updated in this case to know that it has already sent a status message asking the link_list block to the host. This step is necessary because it does not send many status messages to the same CB. If a valid link_list block exists, the next step is to send the CB and transfer information to the DMA prep module. The DMA prep module operates to read data from the socket data buffer and transfer that data to one of the DMA engine's two ping-pong transfer FIFO buffers 234. DMA when this data transfer ends The prep module sends a request to the transmit DMA engine 235 that the data to be transferred exists. The link_list information is also transferred to the transmit DMA engine.
When the transmit DMA engine gets a request from the DMA prep module, the transmit DMA engine notifies the main DMA engine 236 that it wants to make a DMA transfer to the host. When the bus is allowed, the DMA engine reads the data from the ping-pong buffer and sends them to the host computer. When the DMA transfer ends, the CB for the socket is updated with a status message indicating that data is being sent to the host computer.
The status message generator 237 is a module that operates to generate a status message and write the status message to a status message block in memory (1 Kbytes). The request to generate a status message comes from the outgoing DMA engine, the socket DAV query module, or the CPU.
[Socket transmission submodule] The socket transmission submodule handles the interface between the IT19G and the system for data transmission. FIG. 24 shows the socket transmission flow.
The socket transmission flow starts with the receipt of the instruction block list from the host. The command block list is received via DMA transfer and located at command list 241. Blocks are extracted from this and analyzed by the command parser module 242. Commands understood by the parser are executed, and those that are not understood are transferred to the on-chip processor.
If the command transfers data, the link_list information is extracted from the command block along with the CB address and placed in transfer queue 243.
The receive DMA engine 244 unentries this transfer queue and performs a data transfer from the host computer memory. Data is placed in a pair of ping-pong FIFO buffers 245. The CB address associated with the just received data is transferred to socket Xmt data control module 246.
The socket Xmt data control module collects data from the ping-pong FIFO buffer and feeds that data to the transmit socket data memory 248. The socket Xmt data control module looks up the block address from the malloctx memory allocation device 247. The socket Xmt data control module also queries socket CB for the socket priority level. When all the data has been transferred to the data buffer, the socket Xmt data control module assigns the CB address to one of the four priority queues. The socket Xmt data control module also updates the socket CB with new data transmission count information.
When data is transferred from the DMA receive FIFO buffer to the socket data memory, a running checksum is performed at that time. Checksums are calculated on a block-by-block basis. This helps reduce transmission latency when the data does not need to be read again later.
[CB LUT] [Overview] TCP receive logical devices use LUTs to discover open socket connections. The LUT is 18 bits deep at 128K.
[CB LUT and DRAM interface] This section describes the interface between the miscmem module and the NS DDR arbitration module. It describes the data flow, lists the interface signals, and details the required timing.
[data flow] The CB LUT provides only a single read / write access to the data DRAM. In the configuration of the present invention, the DRAM is external, but the memory is provided with an on-chip instead. DRAM access is in DWORD. Since each LUT entry has only 18 bits, only the bottom 18 bits of the DWORD are sent to the CB LUT memory interface.
[TCP transmission submodule] [Overview] The TCP transmit submodule acts to determine the next socket to be serviced for data transmission and therefore update the socket CB block. The data flow is shown in Figure 25.
The TCP send data flow starts with the socket query module. The socket query module scans the XMT_DAV bit table for entries that have a valid bit set of their transmitted data. When the socket query module finds a valid bit set of configured transmit data, the socket query module places the entry in one of four queues according to the socket's Use_Priority level. Sockets with priority level 7 or 6 are located in queue list 3, levels 5 and 4 are located in queue list 2, levels 3 and 2 are located in queue list 1, and levels 1 and 0 are waiting. Located in queuing list 0.
All of these lists are supplied to the packet scheduler 251. The packet scheduler works in a non-intrusive way to remove packets from the priority queue. The packet scheduler also arbitrates between outgoing data packets and SYN_ACK and RST packets originating from a half-open support module.
When the packet scheduler determines the next packet to be sent, the packet scheduler forwards this information to the socket send handler 252. The socket transmission handler module reads the socket CB information, generates a packet header, updates the CB, and transmits the packet transmission information to the transmission queue 253. All packet headers are then generated in a separate memory buffer that is prepending to the data buffer. This also applies if the data being transmitted starts in the middle of the data buffer. In this case, the point from the packet header data buffer points to the first byte of the data being transmitted. The locking mechanism is used because this module does not modify the same socket CB that another module can operate at the same time.
The transmit queue acts to form a queue of packets as it is sent to the master transmit arbitrator.
[Packet scheduler] The packet scheduler acts to determine the next packet to be sent. A block diagram of this module in one configuration is shown in Figure 26.
The packet scheduling process is started by the comparison device 260 by taking the queue number in the current state and checking if there is anything in the queue to be transmitted. The queue number can represent either a queue list or a TCP RCV packet. If there is a packet of that type waiting, that entry is pulled and scheduled as the next packet to be sent. If there are no packets in that queue, the state counter is incremented and the state of the next queue is checked. This continues until the queue number matches the queue list (or TCP received packet) with packets ready to be sent, or the last bit of the state entry is set. If the last bit is set, the status counter resets to zero (0).
The queue arbitration sequence is programmable. The application can set the queuing arbitration sequence by first setting the Queue_State register to 0x00 and then writing the queue number and last bit to the Queue_Entry register. There are two built-in arbitration sequences that can be set by claiming flat or steep bits in the Queue_State register.
[Flat sequence] The flat sequence is the default sequence state used by the scheduler after any reset. The flat sequence is also set to 01 by writing the seq_prog field to the T sequence register. The sequence is shown below. 3-2-3-2-3X-1-3-2-3-X-2-3-1-3-X-2-3-2-3-DX <Repetition> Here, 3, 2, 1 and 0 are queue lists, respectively, and X is a packet from the TCP receive logical device. In this scheme, each list has the following bandwidth allocation:<tables num="6"><img file="JP4875126B2_D0006.tif" /></tables>
[Steep sequence] An alternative to programmed flat sequences is steep sequences. A steep sequence further weights the high-priority queue, which is useful when many high-priority applications are running simultaneously. The steep sequence is set to 10 by writing the seq_prog field to the T sequence register. The steep sequence is as follows. 3-3-2-3-3-2-X-3-3-1-3-3-2-3-X-3-2-3-3-1-3-3-X-2-3- 3-2-3-3-0-X The percentage bandwidth for this sequence is shown below.<tables num="7"><img file="JP4875126B2_D0007.tif" /></tables>
The use of the packet scheduler allows sharing of resources with hardware modules according to incoming data traffic.
[Hash algorithm] The hash algorithm used combines the socket's local and peer ports with local and peer IP addresses to form a single 17-bit hash value. The hash algorithm is designed to be simple, thereby producing the result of a single clock cycle, as well as a diffuse spectrum sufficient to minimize hash LUT collisions. The hash equation is shown below. LP = local port, PP = peer port, LI = local IP address, PIP = peer IP address Hash [16: 0] = {({LP [7: 0], LP [15: 8] ^ PP}, 1'b0) ^ {1'b0, ((LIP [15: 8] ^ LIP [7:: 0]) ^ (PIP [15: 8] ^ PIP [7: 0]))}; The following are the steps performed on the hash. -Bytes on the local port are swapped. -The swapped local port of the new byte is XORed by the peer port to form a 16-bit port product. -The MSB of the local IP address is XOR-processed with the local IP address and LSB. -The MSB of the peer IP address is XOR-processed by the LSB of the peer IP address. -The product of two IP addresses is XORed to form the product of IP words. The port product is shifted left by 1 bit and padded with 0s to form a 17-bit value. -A leading 0 is attached to the IP word product to form a 17-bit product. -The new 17-bit port product and IP word product are XORed to form the final hash value.
This hash algorithm is shown in Figure 27.
The following is an example of a hash algorithm.
LP = 1024, PP = 0080, LI = 45.C3.E0.19, PIP = 23.D2.3F.A1 Port product = 0 × 2410 ^ 0 × 0080 = 0 × 2490 IP product = (0 × 45 C3 ^ 0 × E019) ^ (0 × 23 D2 ^ 0 × 3 FA1) = 0 × A5 DA ^ 0 × 1 C73 = 0 × B9A9 Final hash = {0 × 2490,0} ^ {0,0 × 79A9} = 0 × 04920 ^ 0 × 0B9A9 = 0 × 0F089 The half-open control block LUT uses a 12-bit hash. This 12-bit hash value is obtained from the above equation as follows. Hash 12 [11: 0] = Hash 17 [16: 5] ^ Hash 17 [11: 0] Here, hash 17 is the formula specified above.
[ISN algorithm] [Operation theory] The ISN algorithm used in IT10G is similar to the algorithm described in RFC1948, with a 4-microsecond-based timer, a random boot value configurable by the host system, and four socket parameters (port and IP address). doing. The function has the following form. ISN = timer + F (boot_value, src_jp, dest_ip, src_port, dest_port) Here the timer is a 4-microsecond based 32-bit up counter. The F () function is based on the function FC () and is defined as follows. FC (hv32, data32) = [(prev32 << B) ^ (prev32 >> 7)] ^ data32 First, the value of hv32 is set to a random boot_value from the system (host computer). After that, for each requested ISN, hv32 is calculated as follows. hv32 = FC (prev32, src_ip) hv32 = FC (prev32, dest_ip) v hv32 = FC (prev32, {src_port, dest_port}) A block diagram of the entire algorithm is shown in Figure 28.
The current architecture takes 4 clock cycles to calculate the ISN. In the first cycle, the source IP address (src_ip) is fed through to generate hv32. In the second clock cycle, the destination IP address is fed through, and in the third clock cycle, the port information is fed through. The fourth clock cycle is used to add an incrementing timer value to the final function value. Also note that register A and register B are not clocked in the fourth clock cycle.
[Test mode] The ISN test mode is given as a way to set the ISN to a given number for diagnostic purposes. This test mode is used as follows. Write 0x00 to 0x1A06 (TCP TX read index register), Write the ISN bit [7: 0] to 0x1A07, Write 0x01 to 0x1A06, Write the ISN bit [15: 8] to 0x1A07, Write 0x02 to 0x1A06, Write the ISN bit [23:16] to 0x1A07, Write 0x03 to 0x1A06, Write the ISN bit [31:24] to 0x1A07, Write 0x04 to 0x1A06, Write any value to 0x1A07 (enable test mode).
Writing in step # 10 above enables ISN test mode. The ISN identified now will be used the next time the ISN is requested. To clear the test mode, write 0x05 to 0x1A06 and write any value to 0x1A07.
[Socket control block structure] [Overview] The socket control block or CB contains the operation information for each socket and exists in the CB memory space. The CB contains all state information, memory pointers, configuration settings, and timer settings for sockets. These parameters can be updated by various TCP submodules. The locking mechanism ensures that only one submodule can change the socket CB at a time. IT10G uses different CB structures for sockets that are half-open, configured, and closed.
[Main TCP / UDP CB structure of configured socket] The table below lists all the fields in the main TCP / UDP CB structure of the configured socket memory. There is also an accompanying TCP CB structure defined below.<tables num="8"><img file="JP4875126B2_D0008.tif" /></tables>
[Main CB field specifications for configured sockets] [Remote IP address (address 0 x 00, 32-bit)] This 32-bit field represents the remote IP address for the connection. For client sockets, this field is set by the application. On server sockets, this field is filled with the IP address received in the SYN packet or the IP address of the received UDP packet.
[Remote port (address 0 x 01 [31:16], 16 bits)] This field represents the remote port number for the connection. For client sockets, this field is always specified by the application. On server sockets, this field is always filled with the port number received in SYN or UDP packets.
[Local port (address 0 × 01 [15: 0], 16 bits)] This field represents the local port number for the connection. For client sockets, this field is specified by the application or automatically generated by the network stack. On server sockets, this field is always specified by the application.
[IP index (address 0 × 02 [31:28], 4 bits)] This field represents the index of the network stack IP address table of the host interface's IP address for the socket and is filled by the network stack hardware.
[ConnState (address 0 × 02 [27:24], 4 bits)] This field shows the current state of the connection and decrypts it as follows:<tables num="9"><img file="JP4875126B2_D0009.tif" /></tables>
[AX (address 0 × 02 [23], 1 bit)] This bit indicates that the received ACK status message mode is enabled. In this mode, a status message is raised when an ACK is received to approve all prominent data. This mode is used with the Nagle algorithm.
[SA (address 0 × 02 [22], 1 bit)] This bit indicates that the selective ACK option (SACK) should be used with this socket (0 = SACK must not be used, 1 = SACK should be used).
[TS (address 0 × 02 [21], 1 bit)] This bit indicates that the timestamp option should be used by this socket (0 = timestamp option must not be used, 1 = timestamp option should be used).
[WS (address 0 × 02 [20], 1 bit)] This flag indicates that the sliding window option is properly negotiated for the TCP socket (0 = WS option should not be used, 1 = WS option will be available). This bit is not used on UDP sockets.
[ZW (address 0 × 02 [19], 1 bit)] This bit indicates that the peer window of the socket is 0x0000 and the socket is in the window probe state of zero.
[AR (address 0 × 02 [18], 1 bit)] This bit is set when an ACK packet is sent on a particular socket. This is cleared when an ACK packet is sent.
[CF (address 0 × 02 [17], 1 bit)] This bit indicates that the current CB is in the receive / transmit FIFO buffer queue and therefore cannot be decremented. Once an ACK is sent to the socket, the TCP send block moves from the Open CB to the Time_Wait CB.
[CV (address 0 × 02 [16], 1 bit)] This bit indicates that the CB contains valid information. This bit is always cleared before any CB decrement.
[RD (address 0 × 02 [15], 1 bit)] This bit indicates that the CB retransmission time has expired but cannot be queued for retransmission because the priority queue is full. When the CB Polar finds this bit claimed in the CB, it ignores the retransmission time field and processes the CB for retransmission. This bit is cleared when the CB enters the priority send queue.
[CBVer (address 0 × 02 [14:12], 3 bits)] These bits indicate the version of the CB and are primarily used by on-chip processors to distinguish the type of CB in future hardware versions. This field is currently specified as 0x1.
[CB interface (address 0 x 02 [11: 8], 4 bits)] These bits are used to identify a particular physical interface of the socket and are used in multiport architectures. In a single port architecture, this field should be left at 0x0.
[SY (address 0 × 02 [7], 1 bit)] These bits indicate that the socket has received a SYN packet while in the TW state. If this bit is set and then the socket receives a RST and kill_tw_mode is set to 0x00, the CB is immediately ignored.
[ST (address 0 × 02 [6], 1 bit)] This bit indicates that the TX left and right SACK fields are valid. When the TCP receive logical device notices a hole in the received data, the TCP receive logical device positions the start and end addresses of the hole in the TX left and right SACK fields, and the next packet sent contains the SACK option. Set this bit to indicate that it should be. This bit is cleared when a packet containing these parameters is sent.
[SR (address 0 × 02 [5], 1 bit)] This bit indicates that the RX left and right SACK fields are valid. When the TCP receive logical unit receives the SACK option, the TCP receive logical unit positions the first pair of SACK values in the RX left and right SACK fields and TCP sends that this section needs to be retransmitted. Set this bit to indicate to the machine. This bit is cleared when the section is retransmitted.
[KA (address 0 × 02 [4], 1 bit)] For TCP sockets, this bit indicates that keepalive timers should use this socket (0 = keepalive timers must not be used, 1 = keepalive timers should be used). For UDP sockets, this bit indicates whether the checksum should be checked (1 = enable checksum, 0 = ignore checksum).
[DF (address 0 × 02 [3], 1 bit)] This bit represents the state of the DF bit in the IP header for the socket. When this bit is set, the packet is not fragmented by any hop. This bit is used to find the path MTU.
[VL (address 0 × 02 [2], 1 bit)] This bit indicates whether or not it is included in the Ethernet (registered trademark) frame output by the VLAN tag. If this bit is claimed, 4 bytes of VLAN tag information is created by cb_tx_vian_priority and cb_tx_vid, and the VLAN tag identification (fixed value) is included in the subsequent Ethernet® address field.
[RE (address 0 × 02 [1], 1 bit, TCP CB)] This flag is used to indicate a timeout condition and that a retransmission is needed. Cleared when the data is retransmitted. This bit is specified only for TCP CB.
[UP (address 0 × 02 [1], 1 bit, UDP CB)] This flag is used to indicate whether UDP ports are dynamically assigned or pre-identified. If assigned dynamically, it will be deallocated when the CB is replicated. Otherwise, if pre-specified, no port action will be taken when the CB is replicated. This bit is specified only for UDP CB.
[AD (Address 0 × 02 [0], 1 bit)] This bit indicates that the ACK is queued for delayed ACK transmissions on the socket. This bit is set on the TCP RX logical unit and cleared when the ACK is queued by the CB polar for transmission.
[TX ACK number (address 0 x 03,32 bits)] This is the TCP connection operation ACK number. This represents the expected SEQ number of the received TCP PSH packet. When a TCP data packet is received, the SYN number is checked against this number. Data is accepted if they match or if the received SEQ number + packet length covers this number. This number is automatically updated by the network stack and is not used for UDP connections .
[SEQ number (address 0 x 4,32 bits)] This is the operating SEQ number of the TCP connection. It represents the SEQ number used in TCP packets and is automatically updated by the network stack. This field is not used for UDP connections.
[Keepalive time 0 x 05 [31:24], 8 bits] This field represents the future time when the keepalive trigger is claimed. This time is reset each time a packet is sent on a socket or a pure ACK packet is received. The time is expressed here in minutes.
[LWin scale (address 0 × 05 [23:20], 4 bits)] This field represents the local sliding window scale factor used for the connection. Valid values are 0x00 to 0x0E.
[RWin scale (address 0 × 05 [19:16], 4 bits)] This field represents the sliding window scale factor when requested by the remote end in a TCP SYN packet. Valid values are 0x00 to 0x0E.
[Remote MSS (address 0 x 05 [15: 0], 16 bits)] This field represents the MSS received in the SYN packet for the TCP connection. This indicates the maximum packet size that the remote end can receive. If no MSS option is received, this field defaults to 536 (0x0218). This field is not used for UDP connections.
[PA (address 0 x 06 [31], 1 bit)] This bit is used to indicate that the CB port has been auto-allocated. If this bit is set, when it is time to replicate the CB, this bit is checked to see if the port needs to be deallocated.
[Priority (address 0 x 06 [30:28], 3 bits)] This field represents the user priority level used for the VLAN tag. This field also represents the service level of the socket during the send schedule. The higher the number, the higher the priority (7 is the highest, 0 is the lowest). The default value for this field is 0x0 and can be set by software.
[VID (address 0 x 06 [27:16], 12 bits)] This field represents the VLAG identification used in the VLAG tag frame. The default value for this field is 0x000 and can be set by software. For peer-initiated connections, this field is set by the VID received in the opening SYN packet.
[Remote MAC address (address 0 × 06 [15: 0]-0 × 07, 48 bits in total)] This field represents the destination MAC address of the packet being sent on this socket. The ARP cache is queried at this address when the socket is first configured. After this is resolved, the address will be stored here and further ARP cache queries will be avoided. If the CB is generated as a server socket, this address is taken from the destination MAC address contained in the SYN socket. Address bits [47:32] are stored at CB address 0x6.
[Local IP address (address 0 x 08,32 bits)] This field represents the local IP address for the socket connection. On client sockets, this field is specified by the application. On server sockets, this field is filled when the CB is transferred from half-open to open.
[RX ACK number (address 0 x 09, 32 bits)] This field represents the most recent ACK number received from the peer for this socket. This is used to determine if any retries are required on the socket. It should be noted that this is a different number than the sequence number used in the transmitted packet.
[Congested window (cwnd) (address 0 × 0A [31: 0], 32 bits)] This field tracks the cwnd parameter. The cwnd parameter is used in the congestion avoidance and slow start algorithm and is initialized to 1MSS.
[HA (host ACK) (address 0 × 0B [31], 1 bit)] This bit indicates that host_ACK mode is active on this socket. In this mode, the data ACK is triggered by the host and is not automatically sent when the data is received.
[SS (Sent RX DAV state) (address 0 × 0B [30], 1 bit)] This bit indicates that an RX DAV status message for the CB has been sent to the on-chip processor. The bit clears when the data is DMAd to the host.
[Socket type (address 0 x 0B [29:27], 3 bits)] This field indicates the type of socket represented by the control block according to the table below.<tables num="10"><img file="JP4875126B2_D0010.tif" /></tables>
All other decryptions not shown are pending for future use. In the raw UDP mode, the application is given the remote IP address and UDP header information along with the UDP data. In normal UDP mode, only the data portion is given to the application.
[MI (address 0 × 0B [26], 1 bit)] This bit is used to indicate that the data does not currently exist in the socket receive memory block.
[RX end memory block pointer (address 0 × 0B [25: 0], 26 bits)] This field represents the address of the last MRX buffer in which the received data was written. This is used to link the next block of memory used for the socket.
[RX start memory block pointer (address 0 × 0C [25: 0], 26 bits)] This field represents the address of the next MRX buffer sent to the host.
[HD (address 0 x 0C [26], 1 bit)] This bit indicates that the data to be DMAed to the host remains and the RX DMA status message has not yet been sent. The TCP RX logical unit sets this bit when scheduling an RX DMA operation and is cleared by the TCP RX logical unit when requesting an RX DMA status message. This bit is used to facilitate FIN vs. status message timing issuance.
[ISPEC mode (address 0 × 0C [30:27], 4 bits)] These bits indicate the IPSEC mode enabled for packets sent on this socket connection. Decryption is shown in the table below.<tables num="11"><img file="JP4875126B2_D0011.tif" /></tables>
If there are no bits set, then there is no IPSEC used in the socket.
[AB (ACK Req pending) (address 0 × 0C [31], 1 bit)] This bit indicates that an ACK request from the TCP receive logical device to the TCP transmit logical device is pending. It is set when a TCP data packet is received and cleared when an ACK of data is sent. If another data packet is received while this bit is still set, another ACK is not requested and the number of received ACKs is updated.
[HR (Host Retransmit Socket) (Address 0 x 0D [30], 1 bit)] This bit indicates that this socket is an iSCSI application.
[AS (address 0 × 0D [30], 1 bit)] This bit indicates that this socket is an on-chip processor socket application.
[DA (address 0 × 0D [29], 1 bit)] This bit indicates that a retransmission was performed due to a duplicate ACK. This is set when the retransmitted packet is sent. The received ACK pointer is advanced. Cleared when an ACK is acquired. Set cwnd = ssthresh (part of the fast recovery algorithm) when getting an ACK for new data after duplicate ACK retransmissions. This bit is word 0x14 and is used with the DS bit. Retransmission of the lost segment is done only once, so it requires two bits.
[R0 (address 0 × 0D [28], 1 bit)] This bit indicates that the retransmission time field is valid (0 = invalid, 1 = valid).
[RV (address 0 × 0D [27], 1 bit)] This bit indicates that the retransmission time field is valid (0 = invalid, 1 = valid).
[TP (address 0 × 0D [26], 1 bit)] This bit indicates that a timed packet is currently in progress. Used only if the RTO option is not available.
[Next TX memory block pointer (address 0 x 0D [25: 0], 26 bits)] This field represents the address of the next MTX buffer to be sent.
[MSS_S (2 bits, word 0 × 0E [31:30])] These bits are used to identify the MSS size used in the SYN packet. The valid settings are shown below.<tables num="12"><img file="JP4875126B2_D0012.tif" /></tables>
[Series window size (address 0 x 0E [29: 0], 30 bits)] This field represents the operating window size when tracked by TCP RX hardware. This window maintains the actual window size and is transferred to the actual window only when it is larger than the MSS.
[RX DMA count (address 0 × 0F [31:16], address 0 × 12 [31:24] 24 bits)] This field represents the number of bytes sent to the host via RX DMA during the last RX DMA state message transition. The MSB of the count is stored at address 0x12.
[Host buffer offset pointer (address 0 x 0F [15: 0], 16 bits)] This field represents the offset to the current host memory buffer where RX DMA is used.
[Window clamp (address 0 x 10 [31: 0], 32 bits)] This field represents the largest advertised window enabled by a socket connection.
[RX Link List Address (Address 0 × 11 [31: 0], 32 bits)] This is the RX linked list on-chip processor memory address used for RX DMA operation.
[RX transmission limit (address 0 × 12 [15: 0], 16 bits)] This is the transfer limit for RX DMA transfers. The RX DMA status message is generated when this limit is reached. For CBs that require a status message after each DMA transfer, this limit should be set to 0x0001.
[RX linked list entry (address 0 × 12 [23:16], 8 bits)] This is the number of entries in the RX DMA linked list.
[K (Keepalive Triggered) (Address 0 × 13 [31], 1 Bit)] This bit is used to indicate that the current CB is in keepalive state. This bit is set when the keepalive timer expires on the CB, when the CB is deallocated because the keepalive discovers that the other side is gone (disappeared), or the keepalive probe (ACK packet). Cleared by TCP-RX when receiving a response to.
[DupAck (address 0 × 13 [30:28], 3 bits)] This field keeps track of how many duplicate ACKs have been received. This parameter is used in the fast retransmission algorithm.
[Retry / probe (address 0 × 13 [27:24], 4 bits)] This field keeps track of the number of retries sent for a particular packet. It is also used to keep track of the number of window probes sent. The latter number is needed so that the appropriate window probe transmission interval can be used.
[Available TX data (address 0 × 13 [23: 0], 24-bit)] This field represents the total amount of data available to be sent on the socket.
[Socket channel number (address 0 x 14 [31:24], 8 bits)] This is the socket channel number. When a status message is sent back to the host, this channel identifies the queue to be used.
[SX (address 0 × 14 [23], 1 bit)] This bit indicates that SACK retransmission is required. This is set and cleared by the retransmission module when a SACK retransmission is sent.
[FX (address 0 × 14 [23], 1 bit)] This bit indicates that a FIN packet is being sent on this socket. This is set by the TCP TX logical unit.
[DA (address 0 × 14 [19], 1 bit)] This bit indicates that the socket is in a duplicate ACK state. This bit is set when the duplicate ACK threshold is reached. Cleared when peer ACK is new data.
[UM (address 0 × 14 [18], 1 bit)] This bit indicates that the advertised window of the peer has more than doubled the MSS. If this bit is set, the data will be stored in the MTX with a maximum MSS size. If no bits are set, the data is stored at a maximum of a quarter of the MSS.
[UC (address 0 × 14 [17], 1 bit)] This bit indicates that the control block type for the next CB link field is UDP CB.
[TW (address 0 × 14 [16], 1 bit)] This bit indicates that the control block type for the next CB link field is TW CB.
[Next CB link (address 0 × 14 [15: 0], 16 bits)] This field represents the CB memory address of the next linked CB. This CB has the same hash value as this socket. The VSOCK submodule fills this field when it concatenates CBs to the same hash value.
[RX window size (address 0 x 15 [31:16], 16 bits)] This field represents an advertised window at the remote end that is not adjusted by the sliding window scale factor.
[IP TTL (address 0 x 15 [15: 8], 8 bits)] This field represents the TTL used in the socket's IP header.
[IP TOS (address 0 x 15 [7: 0], 8 bits)] This field represents the TOS field setting used in the IP header for socket connections. This is an optional parameter that can be set by the application. If no TOS parameter is specified, this field defaults to 0x00.
[RX Timestamp / Timed Sequence Number (Address 0 x 16, 32 Bits)] When timing one segment at a time, this field is used to store the sequence number of the timed packet. When an ACK covering this sequence number is received, the RTT can be obtained for that packet. The time stamp when this packet was sent is stored in the local time stamp of the last send field. When the timestamp option is enabled, this field is used to remember the timestamp received in TCP packets.
[Smooth average derivation (address 0 × 17, 32 bits)] This field represents a smoothed mean derivation of round-trip time measurements as calculated using the Van Jacobson algorithm specified in RFC793.
[Slow start threshold (ssthresh) (address 0 × 18, 32 bits)] This field tracks the ssthresh parameter. The ssthresh parameter is used for the congestion avoidance algorithm and is initialized to 0x0000FFFF.
[Smooth RTT (address 0x19, 32 bits)] This field represents a smoothed round-trip time value as calculated using the Van Jacobson algorithm identified in RFC793.
[Resend time stamp (address 0 x 1B, 16 bits)] This field is used to represent the future time when retransmissions are needed on the socket.
[TX Left SACK (Address 0 x 1C, 32-bit)] This field represents the sequence number (lowest sequence number) of the first byte of the first island of the out-of-sequence data received after the in-sequence data.
[TX Right SACK (Address 0 x 1D, 32-bit)] This field represents the sequence number (highest sequence number) of the last byte of the first island of the out-of-sequence data received after the in-sequence data plus one.
[RX Left SACK (Address 0 × 1E, 32-bit)] This field represents the sequence number of the first byte of the first island of out-of-sequence data reported by the SACK option of the received packet.
[RX Right SACK (Address 0 x 1F, 32-bit)] This field represents the sequence number of the last byte of the first island of out-of-sequence data reported by the SACK option of the received packet plus one.
[CB structure with socket attached]] The table below lists all the fields in the attached CB structure of the configured socket memory.<tables num="13"><img file="JP4875126B2_D0013.tif" /></tables>
[CB field definition attached to the set socket] [RX left SACK memory address (address 0 × 0 [23: 0], 24-bit)] This field represents the first address of the first MRX memory block in the linked list of blocks on the SACK island.
[RX right SACK memory address (address 0 × 1 [23: 0], 24-bit)] This field represents the first address of the last MRX memory block in the linked list of blocks on the SACK island.
[SACK block count (address 0 × 0-0 × 1 [31:24], 16 bits)] This field represents the number of memory blocks used by the socket.
[Written RX bytes (address 0x2-0x3, 64-bit)] This field represents the total number of bytes written to the socket's MRS memory. The least significant digit word is stored at address 0x2, and the most significant digit word is stored at address 0x3.
[Sent TX bytes (address 0x4-0x5, 64-bit)] This field represents the total number of bytes sent on the socket. The least significant digit word is stored at address 0x4, and the most significant digit word is stored at address 0x5.
[iS (address 0 x 6 [31], 1 bit)] This bit indicates that the CB is being used as an iCSI socket.
[MD (address 0 x 6 [30], 1 bit)] This bit indicates that the CB is using the MDL for the received buffer list.
[DS (address 0 x 6 [29], 1 bit)] This bit indicates that the socket received a duplicate ACK. This is set by a logical device that handles duplicate ACKs. If a duplicate ACK is received and this bit is already set, the lost segment will not be retransmitted. It is used with the word 0x0D DA bits of the main open CB structure. The state table of these two bits is shown below.<tables num="14"><img file="JP4875126B2_D0014.tif" /></tables>
[Non-updated iSCSI seed (IN) (address 0x6 [28], 1 bit)] This bit is set by the davscan module when an iSCSI memory block is processed and cleared when the block's seed is saved. Bits are claimed to indicate to the host that the hardware has not updated the local copy of the CRC seed. This bit must be probed first when the host wants to reset the CRC seed. If the bit is claimed, the socket must be set to the Seed_Clrd_by_Host bit. The hardware clears this bit when updating the local copy of the CRC seed with a new value written by the host.
[Host cleared (CH) iSCSI seed (address 0x6 [27], 1 bit)] This bit is set by the host and the IN bit is set when writing a new value to the CB AND's CB iSCSI seed. The davscan module clears this bit in the iSCSI seed and CB when observing that this bit is set.
[MDL transfer length (address 0 x 6 [15: 0], 16 bits)] This field is used to indicate how long the MDL DMA transfer should be. Usually the same length as reported in the RX_DAV status message that generated the RX_MDL IB.
[IPSEC hash (address 0x7 [31: 0], 31 bits)] This field represents the hash value of all SA entries used in the socket connection.
[Last unapproved sequence number (address 0x8, 32-bit)] This field represents the last unapproved sequence number from the peer. This relative is used in the timestamp to remember the calculation.
[iSCSI FIM Interval (Address 0 x 9 [31:16], 16 Bit)] This field represents the FIM interval for iSCSI connections. This is stored as the number of bytes.
[iSCSI FIM offset (address 0 × 9 [15: 0], 16 bits)] This field represents the number of bytes until the next FIM insertion.
[iSCSI CRC seed (address 0 x A, 32-bit)] This field represents the CRC seed of the received iSCSI data.
[IPSEC overhead (address 0 x B [7: 0], 8 bits)] This field represents the overhead for the number of double words (4 bytes) that the IPSEC header (and excess IP header) occupy in a packet. This information is used to determine the amount of data that can be placed in a packet.
[iSCSI FIM Enable [FE] (Address 0 x B [8], 1 bit)] This bit indicates whether FIM support is required on the CB (0 = FM disable, 1 = FIM enable).
[MRX buffer (address 0 x B [31:16], 16 bits)] This field represents the number of MRX buffers currently in use by the socket. When this number reaches the MRX buffer limit (general setting), no more data packets can be received by the socket.
[DAV buffer length (address 0 x C [15: 0], 16 bits)] This field indicates the buffer length used for DAV status messages.
[RX transmission FIN status message [SM] (address 0 x C [16], 1 bit)] This bit indicates that the RX FIN receive status message should be sent on the CB.
[RX transmission DAV status message [SV] (address 0 x C [17], 1 bit)] This bit indicates that an RX DAV status message should be sent to the CB.
[RX DMA status message pending [DP] (address 0 x C [18], 1 bit)] This bit indicates that the status message made by RX DMA is pending in CB.
[Last DMA [LD] (address 0 x C [19], 1 bit)] This bit indicates that the last DMA of the host buffer segment is being transmitted. Upon completion of this DMA, an RX DMA status message is generated.
[FIN status message pending [FP] (address 0 x C [20], 1 bit)] This bit indicates that the FIN status message is pending. This is set when data remains in MRX memory but there is no DMA currently in progress. This is cleared when the DMA starts and the SV bit is set.
[Reset Pending [RP] (Address 0 x C [21], 1 bit)] This bit indicates that the reset state message is pending and should be sent after the RX DMA state message. This is set when data remains in MRX memory but there is no DMA currently in progress. This is cleared when DMA starts and the SV bit is set.
[Send reset message [RM] (address 0 x C [22], 1 bit)] This bit receives RST packets and is only used between the tcprxsta and davscan modules.
[Open to TW [OT] (address 0 x C [23], 1 bit)] This bit indicates that the socket has been transferred to the wait time state and has been opened to the wait time transfer process. This bit is set by davscan and read by rxcbupd.v.
[Zero window timestamp (address 0 x C [31:24], 16 bits)] This field shows the future time stamp to send the next zero window probe.
[TX tunnel AH handle (address 0 × D [15: 0], 16 bits)] This field represents the TX tunnel AH SA (security related) handle associated with the socket.
[TX Tunnel ESP Handle (Address 0 x D [31:16], 16 Bit)] This field represents the TX tunnel AH ESP SA (security related) associated with the socket.
[TX transfer AH handle (address 0 x E [15: 0], 16 bits)] This field represents the TX transfer AH SA (security related) handle associated with the socket.
[TX transfer ESP handle (address 0 x E [31:16], 16 bits)] This field represents the TX forwarding ESP SA (security related) associated with the socket.
[iSCSI byte 0-1-2 (address 0 x F [23: 0], 8 bits each)] These bytes represent the unaligned bytes in the iSCSI CRC calculation.
[SBVal (address 0 x F [25:24], 2 bits)] This field is used to indicate a valid iSCSI byte [2: 0].
[Max RX window (address 0 × 10 [15: 0], 16 bits)] This field represents the largest window advertised by the peer. This is used to determine the parameters used in the MTX data packing limit.
[Local MSS (address 0 × 10 [31:16], 16 bits)] This field represents the local MSS value used in SYN or SYN / ACK packets. This is used to adjust the window size.
[Last CWND update (address 0x11 [31:16], 16 bits)] This field represents the last timestamp when cwnd was updated and this information is used to increment cwnd when it is greater than ssthresh.
[DMA ID (address 0 × 12 [3: 0], 4 bits)] This field represents the DMA ID for the CB. This field is incremented each time a TCP TX DMA is requested. When the xmtcbwr module finishes packing the DMA transfer, the DMA_Pending bit is cleared (by xmtcbwr) if the MDA ID matches that in the CB.
[DMA pending (address 0 × 12 [4], 1 bit)] This field indicates that TX DMA is pending on this socket. This bit is set by sockregs when requesting TX DMA and is cleared by xmtcbwr.v when the TX DMA ID exactly matches the TX_DMA of the CB.
[Combined CB structure] In CB memory, socket CB is stored as one contiguous block of memory. The formats are shown in the table below.<tables num="15"><img file="JP4875126B2_D0015.tif" /></tables>
[Main CB structure of half-open socket] The table below defines the main CB structures for half-open socket memory.<tables num="16"><img file="JP4875126B2_D0016.tif" /></tables>
[Definition of main CB fields for half-open sockets] [Remote IP address (32-bit, word 0x0 [31: 0])] This is the IP address of the remote end of the connection. Bits [31:16] of the IP address are stored in word 0x0, and bits [15: 0] are stored in word 0x1.
[Remote port (16-bit, word 0x0 [47:32])] This is the port address of the remote end of the connection.
[Local port (16 bits, word 0x0 [63:48])] This is the port address of the local end of the connection.
[Local IP index (4 bits, word 0x1 [15:12])] This is the index of the local IP address used for this connection. This value is resolved by the IP module to the full IP address and the corresponding MAC address.
[Connection status (4 bits, word 0 × 1 [11: 8])] This is the current state of the connection and is decrypted as follows:<tables num="17"><img file="JP4875126B2_D0017.tif" /></tables>
[RxACK (1 bit, word 0x1 [7])] This bit indicates that the RX ACK status message mode is enabled. In this mode, a status message is raised when an ACK is received to approve all prominent data.
[SACK (1 bit, word 0 × 1 [6])] This bit indicates that the remote end sent the SACK option in the SYN packet.
[RxTS (1 bit, word 0x1 [5])] This bit indicates that the remote end sends a timestamp option and the received timestamp field is valid in a half-open attachment control block.
[WinSc (1 bit, word 0x1 [4])] This bit indicates that the remote end has transmitted a window scale option.
[RxMSS (1 bit, word 0x1 [3])] This bit indicates that the remote end sends the MSS option and the remote MSS field is valid.
[ACKRq (1 bit, word 0 × 1 [2])] This bit indicates that ACK is requested for this socket. This means that similar bits are duplicated in the configuration control block, but are not used here.
[CBinFF (1 bit, word 0 × 1 [1])] This bit indicates that this CB cannot be extinguished yet because it is in the Rx to TX FF queue.
[CBVal (1 bit, word 0 × 1 [0])] This bit indicates that a half-opened CB is valid.
[CB version (4 bits, word 0x1 [31:28])] These bits indicate the CB version. This field is currently specified as 0x1.
[TX interface (4 bits, word 0x1 [27:24])] These bits indicate the interface on which the outgoing SYN comes.
[IPSEC (1 bit, word 0 × 1 [23])] This bit indicates that HO CB is used in IPSEC protected packets. In this case, the received SYN that does not match this HO CB is discarded.
[MSS size (2 bits, word 0x1 [22:21])] These bits are used to identify the MSS size used in SYN / ACK packets. The valid settings are shown below.<tables num="18"><img file="JP4875126B2_D0018.tif" /></tables>
[KA (1 bit, word 0 × 1 [20])] This bit indicates that a keepalive timer should be used on this socket.
[DF (1 bit, word 0 × 1 [19])] This bit is used to identify the bits that should not be fragmented in the IP header.
[VLANV (1 bit, word 0 × 1 [18])] This bit indicates that the VLAN field is valid.
[Resend (1 bit, word 0 × 1 [17])] This bit indicates that retransmission is required on this socket.
[RxURG (1 bit, word 0x1 [16])] This bit indicates that the URG bit was set in the received packet.
[ACK (32-bit, word 0 × 1 [63:32])] This is the number of ACKs used in the transmitted packet.
[SEQ (32-bit, word 0 × 2 [31: 0])] This is the operation sequence number used in the transmitted packet.
[SA overhead (8 bits, word 0x2 [47:40])] This field is used to store the CB's IPSEC overhead (in double words).
[Local WinScale (4 bits, word 0x2 [39:36])] This field represents the local sliding window scale factor used for the connection. Valid values are 0x00 to 0x0E.
[Remote WinScale (4 bits, word 0x2 [35:32])] This field represents the sliding window scale factor as requested by the remote end of the TCP SYN packet. Valid values are 0x00 to 0x0E.
[Remote MSS (16-bit, word 0x2 [63:48])] This field represents the MSS received in SYN packets for TCP connections. This indicates the maximum packet size that the remote end can receive. If no MSS option is received, this field defaults to 536 (0x0218). This field is not used for UDP connections.
[VLAN priority (3 bits, word 0x3 [14:12])] This field represents the user priority level used for the VLAN tag. It also represents the level of service for sockets in the transmission schedule. The higher the number, the higher the priority (7 is the highest and 0 is the lowest). The default value for this field is 0x0 and can be set by software.
[VLAN VID (12-bit, word 0 × 3 [11: 0])] This field represents the VLAN identification used in the VLAN tag frame. The default value for this field is 0x000 and can be set by software. For peer-initiated connections, this field is set by the VID received in the opening SYN packet.
[Remote MAC address (48 bits, word 0x3 [63:16])] These fields represent the destination MAC address for packets being sent on this socket. The ARP cache is queried for this address when the socket is first configured. After resolution, the address is stored here, thereby preventing further ARP cache queries. If the CB is generated as a server socket, this address is taken from the target MAC address contained in the SYN packet.
[Received timestamp (32-bit, word 0x4 [31: 0])] This is the time stamp contained in the received packet.
[Local Timestamp (32-bit, word 0x4 [63:32])] This is the time stamp of the last packet sent to this socket.
[TTL (8 bits, word 0x5 [15: 8])] This is the TTL used for the connection. When SYN / ACK occurs, this parameter is given by the IP router and stored in this position. This information is passed to the open control block as the socket transitions to its configured state.
[TOS (8 bits, word 0 × 5 [7: 0])] This is the TTL used for the connection. When SYN / ACK is occurring, this parameter is given by the IP router and stored in this position. This information is passed to the open control block as the socket transitions to its configured state.
[SYN / ACK retry count (4 bits, word 0x5 [19:16])] This field keeps track of the number of SYN / ACK retries sent for a particular socket. If the number of retries reaches the programmed maximum, the socket is replicated.
[IPSEC mode (4 bits, word 0x5 [31:28])] This field is used to indicate the active IPSEC mode for the socket. Decoding of this field is shown below.<tables num="19"><img file="JP4875126B2_D0019.tif" /></tables>
[Local IP address (32-bit, word 0x5 [63:32])] This field is used to remember the destination IP address received in the SYN packet. This represents the local IP address for the socket connection.
[TX Tunnel AH SA Handle (16 Bit, Word 0x6 [15: 0])] This field is used to store the TX tunnel AH SA handle for the CB (if applicable).
[TX Tunnel ESP SA Handle (16 Bit, Word 0x6 [31:16])] This field is used to store the TX tunnel ESP SA handle for the CB (if applicable).
[SA hash (31 bits, word 0x6 [63:32])] This field is used to store the SA hash for the CB.
[TX Transfer AH SA Handle (16 Bit, Word 0x7 [15: 0])] This field is used to store the TX transfer AH SA handle for the CB (if applicable).
[TX Transfer ESP SA Handle (16 Bit, Word 0x7 [31:16])] This field is used to store the TX transfer ESP SA handle for the CB (if applicable).
[Local MSS (16-bit, word 0x7 [47:32])] This field is used to remember the local MSS value used in the connection. This is needed to adjust the window size to avoid the series-window syndrome.
[Waiting time CB structure] The following table defines the CB structure of the memory of the socket in the Time_Wait state.<tables num="20"><img file="JP4875126B2_D0020.tif" /></tables>
[Waiting time CB field definition] [Remote IP address (address 0 x 00, 32-bit)] This 32-bit field represents the remote IP address for the connection. For client sockets, this field is set by the application. On server sockets, this field is filled with the IP address received in the SYN packet or the IP address of the received UDP packet.
[Remote port (address 0 x 01 [31:16], 16 bits)] This field represents the remote port number for the connection. For client sockets, this field is always specified by the application. On server sockets, this field is filled with the port number received in the SYN or UDP packets.
[Local port (address 0 × 01 [15: 0], 16 bits)] This field represents the local port number for the connection. For client sockets, this field is specified by the application or automatically generated by the network stack. On server sockets, this field is always specified by the application.
[IP index (address 0 × 02 [31:28], 4 bits)] This field represents the index of the network stack IP address table of IP addresses on the socket's host interface and is filled with network stack hardware.
[ConnState (address 0 × 02 [27:24], 4 bits)] This field shows the current connection status and is decrypted as follows.<tables num="21"><img file="JP4875126B2_D0021.tif" /></tables>
[RX (address 0 × 02 [23], 1 bit)] This bit indicates that the RX ACK status message mode is enabled. This mode has no meaning for TW CB.
[SA (address 0 × 02 [22], 1 bit)] This bit indicates that a selective ACK option (SACK) should be used on this socket (0 = SACK must not be used, 1 = SACK is used).
[TS (address 0 × 02 [21], 1 bit)] This bit indicates that the timestamp option should be used on this socket (0 = timestamp option must not be used, 1 = timestamp option should be used).
[WS (address 0 × 02 [20], 1 bit)] This flag indicates that the sliding window option is negotiating properly for the TCP socket (0 = WS option should not be used, 1 = WS option will be available). This bit is not used in UDP sockets.
[MS (address 0 × 02 [19], 1 bit)] This bit indicates that the CB's remote MSS field is valid. MSS is located in word 5 [15: 0].
[AR (address 0 × 02 [18], 1 bit)] This bit is set when an ACK packet is sent on a particular socket. This is cleared when an ACK packet is being sent.
[CF (address 0 × 02 [17], 1 bit)] This bit indicates that the current CB is in the receive / send FIFO buffer queue so that it cannot be replicated yet. Once the socket ACK is sent, the TCP send block moves from the open CB to the Time_Wait CB.
[CVI (address 0 × 02 [16], 1 bit)] This bit indicates that the CB contains valid information. This bit is always cleared before deplicating any CB.
[CBVer (address 0 × 02 [15:12], 4 bits)] These bits indicate the version of the CB and will be used by the on-chip processor to discriminate between CBs in future versions of the hardware. This field is currently specified as 0x1.
[CB interface (address 0 x 02 [11: 8], 4 bits)] This bit is used to identify a particular physical interface to a socket and is used in multiport architectures. In a single port architecture, this fee should be left at 0x0.
[KA (address 0 × 02 [4], 1 bit)] This bit indicates that the keepalive timer should be used with this socket (0 = keepalive timer should not be used, 1 = keepalive timer should be used).
[DF (address 0 × 02 [3], 1 bit)] This bit represents the state of the DF bit in the IP header for the socket. When set, the packet is not fragmented by any hop. This bit is used to find the path MTU.
[VL (address 0 × 02 [2], 1 bit)] This bit indicates whether or not it is included in the Ethernet (registered trademark) frame output by the VLAN tag. If so, it contains 4 bytes of VLAN tag information created by cb_tx_vlan_priority and cb_tx_vid and a VLAN tag identification (fixed value), followed by an Ethernet® address field.
[RE (address 0 × 02 [1], 1 bit)] This flag is used to indicate a timeout condition and the need for retransmission. It is cleared when the data is retransmitted.
[UR (address 0 × 02 [0], 1 bit)] This bit indicates that urgent data has been received. This is claimed until the application reads (or indicates that it has read) this received emergency data pointer.
[ACK number (address 0 x 03, 32 bits)] This is the operation ACK number for TCP connection. This represents the expected number of SEQs of TCP PSH packets received. When a TCP data packet is received, the SYN number is checked against this number. If this matches, or if the number of SEQs received + packet length covers this number, then the data is received. This number is automatically updated by the network stack and is not used for UDP connections.
[SEQ number (address 0 x 4,32 bits)] This is the number of operating SEQs for a TCP connection. This represents the number of SEQs used in TCP packets and is automatically updated by the network stack. This field is not used for UDP sockets.
[IPSEC mode (address 0 x 5 [30:27], 4 bits)] This field is used to indicate the active IPSEC mode on the socket. The decoding of this field is shown below.<tables num="22"><img file="JP4875126B2_D0022.tif" /></tables>
[Socket type (address 0 x 05 [26:24], 3 bits)] This field indicates the type of socket represented by the control block according to the table below.<tables num="23"><img file="JP4875126B2_D0023.tif" /></tables>
[LWinScale (address 0 × 05 [23:20], 4 bits)] This field represents the local sliding window scale factor used for the connection.
[RWinScale (address 0 × 05 [19:16], 4 bits)] This field represents the sliding window scale factor when requested by the remote end of a TCP SYN packet.
[Local window (address 0 × 05 [15: 0], 16 bits)] This field represents the 16 bits of the window advertised in the TCP Local Sliding Window field.
[Priority (address 0 x 06 [30:28], 3 bits)] This field represents the user priority level used for the VLAN tag. It also represents the service level for sockets that are being scheduled for transmission. The higher the number, the higher the priority (7 is the highest and 0 is the lowest). The default value for this field is 0x0 and can be set by software.
[VID (address 0 x 06 [27:16], 12 bits)] This field represents the VLAN identification used in the VLAN tag frame. The default value for this field is 0x000 and can be set by software. For peer-initiated connections, this field is set by the VLAN identification received in the opening SYN packet.
[Remote MAC address (address 0 × 06 [15: 0]-0 × 07, 48 bits in total)] This field represents the destination MAC address of the packet being sent on this socket. The ARP cache is queried for this address when the socket is first configured. After resolution, the address is stored here to avoid further ARP cache queries. If the CB is generated as a server socket, this address is taken from the destination MAC address contained in the SYN packet. Address bits [47:32] are stored at CB address 0x6.
[Local IP address (address 0 x 8, 32-bit)] This field represents the local IP address of the socket connection.
[Remote time stamp (address 0 x 09, 32 bits)] This is the time stamp contained in the received packet.
[TW generation time stamp (address 0 x A, 32 bits)] This is the time stamp when the CB enters the time standby state.
[Tunnel AH handle (address 0 x B [31:24], address 0 x C [31:24], 16 bits)] This is the tunnel AH SA handle for CB. This is also valid when the tunnel AH (TA) bit of the CB is set.
[Forward link (address 0 x B [23: 0], 24-bit)] The next CB link in the search chain.
[Tunnel ESP handle (address 0 x D [31:24], address 0 x E [31:24], 16 bits)] This is the tunnel ESP SA handle for the CB. This is also valid when the tunnel ESP (TE) bit of the CB is set.
[Forward age link (address 0 x D [23: 0], 24-bit)] A link to the next CB by the age of the search chain.
[Reverse age link (address 0 x E [23: 0], 24-bit)] This is a link to the previous CB by the age of the search chain.
[Transfer AH handle (address 0 x F [15: 0], 16 bits)] This is the transfer AH SA handle for the CB. This is also valid when the transfer AH (NA) bit of the CB is set.
[Transfer ESP handle (address 0 x F [31:16], 16 bits)] This is the transfer ESP SA handle for the CB. This is also valid when the CB's transfer ESP (NE) bit is set.
[TCP congestion control support] [Overview] IT 10G performs slow start, congestion prevention, fast retransmissions, fast recovery algorithms, window scaling, and hardware out-of-order packet reordering. In addition, IT 10G supports a round-trip time TCP option that allows more than one segment to be timed at one time. These features are needed in high speed networks.
[Measurement of round trip time] IT 10G can measure round trip time (RTT) in two ways. In the traditional method, the time measurement is done from the TCP PSH packet until the packet ACK is received. The sequence number of the timed packet is stored in the sequence number of the timed packet field in the CB, and the packet time stamp is stored in the time stamp of the last transmit field of the CB. When an ACK of a timed packet is received, the delta between the current time stamp and the stored time stamp is the round-trip time RTT. When an ACK is received, the RTO [1] bit in the socket CB is cleared to indicate that the next packet is timed.
When RT options are negotiated in the opening TCP handshake, RT measurements can be taken from each ACK received.
Regardless of the method used to make round trip time measurements. The logical flow that takes that value and determines the retransmission timeout (RTO) value is the same. This logic is shown in Figure 29.
Scaled and smoothed RTTs, mean deviations and RTOs are all stored in the socket CB.
[Slow start algorithm] Slow Start is the TCP congestion control mechanism first described in RFC1122, as well as RFC2001 http; //www.rfc-editor.org/rfc/rfc1122.txt and http; //www.rfc-editor.org. It is described in /rfc/rfc2001.txt.
Slow start slowly raises many data segments in transit at once. First the slow start must send only two data segments (corresponding to twice the maximum segment size or the current window cwnd of 2 * MSS) before predicting an acknowledgment (ACK). Upon receiving each successful ACK, the transmitter can increase cwnd by 1 MSS until cwnd equals the receiver's advertising window (thus allowing more segments to be transmitted than 1).
Slow start always starts with a new data connection and is sometimes triggered in the middle of a connection when data traffic is congested. Slow start is used to get things done again.
The network stack supports a slow start algorithm for each TCP connection. This algorithm uses the congestion window parameter (cwnd), which is initialized to 1MSS when the socket is first configured.
The slow start algorithm indicates when the socket is first configured, indicating that only one packet can be sent and no more data can be sent until an ACK of the packet is received. When an ACK is received, the cwnd is incremented by 1 MSS, which allows up to two packets to be sent. Each time an ACK is received, cwnd is incremented by 1 MSS.
This continues until cwnd exceeds the advertised window size from the peer. The network stack always sends the smallest cwnd and advertised windows.
If the network stack receives an ICMP source disappearance message, reset cwnd to 1MSS. However, the slow start threshold variable (ssthresh) is maintained at which same value.
[Congestion prevention algorithm] The network stack keeps the minimum cwnd and advertised window launches from the peer. The congestion prevention algorithm also uses a slow start threshold variable (ssthresh), which is initialized to 0 × FFFF.
When congestion is detected by a timeout, ssthresh is set to half of the current outbound window (minimum cwnd and peer ad window). If this value is less than twice the MSS, then this value is used instead. Also, cwnd is set to 1MSS.
When new data is acknowledged, cwnd is incremented by 1 MSS until it is greater than ssthresh (hence the name). After that, cwnd is increased by 1 / cwnd. This is the congestion prevention phase.
[High-speed retransmission and high-speed recovery algorithm] Fast retransmissions were first described in RFC112 as an experimental protocol, formalized in RFC2001, http; //www.rfc-editor.org/rfc/rfc1122.txt, and http; //www.rfc-editor.org. You can refer to /rfc/rfc2001.txt.
Fast retransmissions generate an ACK immediately when an out-of-order segment is received to allow the sender to fill holes quickly instead of waiting for a standard timeout.
Fast retransmission is called when the receiver receives three duplicate ACKs. When fast retransmission is called, the sender tries to fill the hole. Duplicate ACKs are considered to overlap when the segment ACKs and window ad values match each other.
It is strongly indicated that packets will be dropped when the network stack receives a duplicate ACK. When n duplicate packets are received, the dropped segment is immediately retransmitted even if the retransmission timer has not yet expired. This is a fast retransmission algorithm. The number of duplicate ACKs n that must be received before retransmission is set via the TCP_Dup_ACK register (0x36) and defaults to 3.
When identified and a large number of duplicate ACKs are received, ssthresh is again set to half the current window size, as in the case of the congestion prevention algorithm, but then cwnd is set to ssthresh + (3 * MSS). The algorithm. This ensures that after receiving duplicate ACKs, it returns to the congestion prevention algorithm instead of slow start. Each time another duplicate ACK is received, cwnd is incremented by 1 MSS. This is a high speed recovery algorithm.
When an ACK for new data is received, cwnd is set to ssthresh.
[Retransmission theory of operation] Logical retransmission support exists in three positions: the position where data is transmitted, the position where the received ACK is processed, and the position in the CB polar.
[Data transmission] When transmitting data, the transmit logical unit observes the retransmit valid bits of the associated CB. If the bit is not set, it implies that no valid retransmission time is remembered, and the currently remembered RTO (in the CB) is added to the current timestamp. The resulting time is the retransmission time for the packet. This time is stored in the CB, and the CB retransmission valid bit is set.
If the retransmission valid bit is already set in the CB, it means that there is currently prominent data in the socket and the retransmission time does not need to be updated here.
If the data buffer being sent is due to a timeout and therefore a retransmission, then the CB's retransmission time is constantly updated and the RTO stored in the CB is also adjusted (doubled). ..
[ACK processing] When the TCP RX module receives an ACK, it sends a notification to the TX module with the ACK timestamp (if included). The TX module uses this information to update the RTO. If the ACK responds to the retransmitted packet, this update will not occur. An indication as to whether the packet has been retransmitted is maintained in the data buffer header in MTX memory.
After it is calculated, if a new RTO is needed, it will be stored in the CB. If any prominent data is detected in the socket, the new retransmission time is also calculated by adding a new RTO to the current timestamp, which is also stored in the CB.
[CB POLA] CB POLA circulates all CBs looking for an active block. When CB POLA discovers an active block, CB POLA checks the retransmit valid bit. If the retransmission valid bit is set, the CB POLA compares the retransmission times. If the CB polar finds that the resender time has expired, the CB polar knows that the CB's retransmission bit (ie how the data transmitter knows that the packet is a retransmitted packet). And put the CB in one of the priority queues.
[Select MSS] [Overview] This section gives an overview of how MSS options are obtained.
[Initial setup] Before enabling TCP transactions, the host should set up the following parameters and settings:
· The default non-local MSS used in register 0x1A4A-0x1A4B, The default local MSS used by register 0x1A4C-0x1A4D.
[Selection algorithm] When choosing the MSS value to be used for any of the two connections, the TCP engine queries the IP router. If the destination route goes through the gateway, a non-local MSS is used.
[TCP Options] [Overview] The following is an overview of the supported TCP options and their formats. The four supported options are MSS, Window scaling, ·Time stamp, SACK.
[MSS Options] This option is always sent. The actual MSS value used is determined according to the algorithm already described. The format of the MSS option is shown in Figure 30.
Window Scaling Options This option is always sent in SYN packets as long as the Sl_Win_En bit is set in the TCP_Control register. This option is sent in a SYN / ACK packet only if it is included in a SYN packet that raises a SYN / ACK response. The format of this option is shown in Figure 31. Note that this option is always preceded by NOP bytes so that it is aligned on a 4-byte boundary.
Timestamp Options This option is always sent in a SYN packet if the timestamp option is enabled, and if the option is included in a SYN packet that raises a SYN / ACK response, regardless of whether the timestamp option is enabled or not. Is sent as a SYN / ACK packet only to. The optional format is shown in Figure 32. Note that this option is always preceded by two NOP bytes so that it is aligned on a 4-byte boundary.
[Selective ACK (SACK) option] A selective acknowledgment (SACK) allows TCP to approve data containing holes (lost data packets) and retransmits lost data rather than all data after the drop. This feature is described in RFC 2018.
This option is always sent in SYN and SYN / ACK packets as long as the SACK_En bit is set in the TCP_Control register. SACK uses two different TCP option types. One option type is used in SYN packets and the other option type is used in data packets. The optional formats are shown in Figures 33 and 34.
[MTX buffer header format] [Overview] This section describes the header format used at the start of each MTX data buffer. The header contains information about the data contained in the buffer, such as checksum, sequence number, next link, etc. The header format is different for TCP and UDP.
[TCP keep-alive timer] [Overview] This section describes the keepalive timer used for TCP connections. This section contains behavioral theories about how this feature is implemented in the IT 10G network stack.
[Keepalive parameters] The following parameters are used in the keepalive characteristics of the network stack.<tables num="24"><img file="JP4875126B2_D0024.tif" /></tables>
[Operation theory] By default, keepalive timers are disabled in sockets. When a socket is created, the host has the option to set the Keep_Alive bit in the Socket_Configuration register. If this bit is set and a socket parameter is given, the keepalive property will be available on that socket.
Once the socket is set, the keep_alive timestamp is kept in the CB (word 0x05). This time stamp represents the future time when the socket is considered idle. This timestamp is updated each time a packet is sent on the socket or when the packet arrives on the socket. The updated timestamp is the sum of the currently running counter and the keep_alive time (as entered by the host via the Keep_Alive_Timer register (0x1A3A)).
The CB polar module checks the keepalive time stamp once each CB with keepalive characteristics becomes available. If the CB Polar module finds a socket with a time stamp that matches the current free-running minute counter, the CB Polar module schedules a keepalive probe sent on that socket. The CB polar module does this by queuing the CB in the data transmission queue and setting the K bit (CB word 0x0B keepalive retry bit). The CB Polar module also sets the CB keepalive retry field (word 0x13) to 0x1. If the send queue is full, the CB polar module increments the CB keep_alive timestamp to trigger the CB polar again during the next minute check.
When the data packet generator encounters a CB with a set K bit, the data packet generator sends a 1-byte PSH packet with a sequence number whose sequence number is set to a value one lower than normally set. Send out. The data generator also updates the CB keepalive timestamp with the current free-running minute count plus keepalive retry interval. If this is still active, it causes the peer to return an ACK with the exact predicted sequence number.
If the peer is active and sends a response, the TCP-RX module clears the K bit in the CB, resets the retry count to 0x0, and updates the keepalive timestamp.
If the peer does not send a response within the keepalive retry interval, CB POLA again encounters a socket from the previous minute keepalive timestamp check. When observing that the K bit is already set, increment the keepalive retry counter by 1 and reschedule the CB for data packet transmission.
If the CB POLA discovers that the maximum number of retries has been reached, the CB POLA sends a notification to the host computer and replicates the socket.
During the keepalive probe's sending course, the network stack receives a RST packet, which is likely to mean that the peer has been rebooted (reset or restarted) without properly closing the socket. In this case, the socket is replicated.
Another possibility is that the network stack receives an ICMP error message in response to a keepalive probe, eg, a keepalive probe for an unreachable network. This condition occurs when the intermediate router goes down. In this case, the ICMP message is sent to the on-chip processor. After receiving multiple of these ICMP messages, the on-chip processor can read the socket CB and observe that it is in a keepalive retry state. The on-chip processor can then initiate the CB deplication process.
[TCP ACK mode] [Overview] The following is a detailed description of the TCP ACK mode enabled in the IT10G network stack. Of the four modes, there are two bits that determine the mode in which the stack is operating. These are the Dly_ACK bit in the TCP_Control1 register and the Host_ACK bit in the Socket_configuration register. The arithmetic matrix is shown in the table below.<tables num="25"><img file="JP4875126B2_D0025.tif" /></tables>
[Normal ACK mode] In this mode, the data are ACKed as soon as they are received by the NRX DRAM. The TCP receive logical device schedules an ACK by positioning it in the RX-to-TX packet FIFO buffer. The AR bit on the socket CB is not used in this mode. This is the default mode for TCP modules and can be enabled by declaiming the Dly_ACK bit in the TCP_Control1 register. This is a general TCP setting that applies to all sockets.
[Delayed ACK mode] In this mode, the data is not ACKed immediately. Instead, when data arrives at the socket, the TCP receive logical device sets the AR bit on the socket CB. The CB Polar module then periodically tests all active CBs to observe CBs with an AR bit set. When the CB POLA discovers the set bit, the CB POLA schedules an ACK for that socket. The polling interval to look for the AR bit can be set via the Del_ACK_Wait register and is specified for 2ms timer increments. The default interval is 250ms. This mode is enabled by claiming the Dly_ACK bit in the TCP_Control1 register. It is also a general TCP setting to apply to all sockets.
[Host-Normal ACK mode] In this mode, the data is not ACKed when received by the MRX DRAM. Instead, an ACK is sent only after the data has been sent to the host computer, and the host computer approves the receipt of the data by issuing the data ACK command. When the IT 10G hardware receives the ACK command, it writes the CB, the socket handle, to the Host_ACK_CB register. This operation schedules an ACK packet via the RX-TX packet FIFO buffer. The AR bit on the socket CB is not used in this mode. This mode can be enabled according to socket base by claiming the Host_ACK bit in the Socket_Configuration register.
[Host delayed ACK mode] In this mode, the data is not ACKed immediately. Instead, the data is sent to the host computer after it is received, and the host computer approves the reception of the data by issuing the data ACK command. IT When the 10G hardware receives the ACK command, it writes the CB, the socket handle, to the Host_ACK_CB register. This operation sets the AR bit in the CB of the socket. When the CB polar module checks all AR bits in the next ACK polling cycle, the CB polar schedules an ACK to be sent to that socket. The CB polling cycle identified via the Del_ACK_Wait register only identifies the polling interval for the AR bits and takes into account the delay caused by first sending the data to the host computer and waiting for the host computer to respond to that data. It should be noted that it does not. Therefore, if this mode is used, it is suggested that the ACK polling interval be set to some value lower than the default 250ms. The optimal interval depends on the turnaround time of the host computer that issues the ACK command.
[Host retransmission mode] [Overview] The IT 10G network stack is designed to retransmit TCP data from its local data buffer or from host memory. This description details the behavior when the network stack is operating in the latter mode. The advantage of retransmitting from host memory is that the amount of local transmit buffer memory can be kept to a minimum. This is because the MX data buffer block is released as soon as the data is wire-transmitted. The disadvantage is that it takes more host memory, host CPU cycles, and latency to support host memory retransmissions.
[Host retransmission mode enable] Host retransmission mode can be enabled according to socket base by claiming the Host_Retrans bit in the Socket_Configuration2 register. This mode can be used with any ACK operating mode.
[Host retransmission mode operation theory] When operating in host retransmission mode, the host CPU behaves to keep track of sequence numbers. In each data section DMAged from host memory, the first 128 bits form the header for the data block.
The host must write to this header as the first 128 bits of each data transfer. Header bit [127: 64] should be DMAd to the first 64-bit word, followed by header bit [63: 0], followed by the first data byte.
When the IT10G hardware spawns a data packet, it takes the sequence number of the alternate data buffer header from the socket CB. The IT10G hardware then forms packets as usual and wires the data packets. After forwarding the packet, the MTX data buffer is immediately freed.
When an ACK of data is received from the peer, the host computer is notified via a status message about the CB that received the data and the received SCK number. This information allows the host computer to determine when it will be cleared to free the data from the host computer memory.
The retransmission timer is still used in the socket CB under normal conditions. When the retransmission timer expires, a status message is sent to the host computer with the corresponding socket handle. The host computer then schedules a DMA transfer of the oldest data present in unACKed memory.
[Pier Zero Window Case] [Overview] Sometimes the peer side of a TCP connection runs out of receive buffer space. In this case, the peer advertises a window size of 0x0000. The peer window is always written to the socket CB by the TCP receive section. When the receiving logical device detects a window size of 0x0000, it sets the ZW bit in the CB flag word (word 0x02 bits [19]). The next time data is sent on this socket, the TCP transmit logical device discovers the configured ZW bit and instead of sending the data, sends a 1-byte data packet with an incorrect sequence number. The sequence number is one less than the correct sequence number. This data packet looks like a keepalive probe packet. In response to this incorrect data packet, the peer must send an ACK packet with an updated window size. After sending a 1-byte data packet, the CBzero_window probe count is incremented by 1 (word 0x13 bits [27:24]) and CB when the next zero_window probe should be sent. zero_win_timestamp (word 0x0C bits [31:26]) is written.
[Zero_Window Timestamp] This field is used to determine when the next zero_window probe should be sent. Every time a zero_window probe is sent, the probe count is incremented and the timestamp is updated. The transmit logical unit shifts the zero_window_interval count by the number of probes transmitted and adds this time to the current free-running second timer to determine the next zero_window timestamp. The resulting timestamp represents the future time when the next probe will be sent. Therefore, with a default value of 3 seconds, the probe is delivered at intervals of 0, 3, 6, 12, 24, 48, 60 seconds. Once the interval reaches 60 seconds, it will be capped.
In the CB polar module, the logic device poles all active CB ZW bits every X seconds, where X is the zero_window_interval count. When the CB polar finds a CB with the ZW bit set, the CB polar reads the zero_window timestamp and compares that timestamp with a free-running second timer. If they match, the zero-window probe is queued for transmission. Unlike keepalive probes, there is no limit to the number of zero-window probes that can be sent.
[Open window] When the peer's window finally opens, the peer sends an arbitrary ACK (window notification), or the peer indicates that the window with the ACK open responds to the zero window probe. The TCP receive logical unit then updates the CB peer window. The next time the CB appears on the zero_window probe, the CB POLA observes that the window is open. CB Polar then clears the ZW bit and schedules regular data transfers. The number of zero_window retries is also reset to 0x0.
[Nuggle operation] [Overview] This section describes IT10G hardware support for the Nagle algorithm. This algorithm states that when a TCP connection has unACKed data during the transfer, small segments cannot be transmitted until the data is approved.
[Naguru Send IB] Separate transmit instruction blocks are given using the Nagle algorithm. These instruction blocks (IBs) are similar to regular non-nagle IBs, except that they do not have the claimed IB code msb. The 32-bit address nagle transmission IB is 0x81 and the 64-bit address code is 0x82. The nagle transmit IB operates in the same way as a normal transmit IB, except that the R × ACK bit of the socket CB is claimed when they are parsed by the HW. Normal transmission IB processing clears this bit if it is set.
[RxACK bit] This is a bit that exists in the socket CB and can be set via the application by claiming bit [3] in the Socket_Confg1 register. When set, this bit raises an RX_ACK status message that all prominent data has been acknowledged when an ACK is received. This state is satisfied when the received ACK number matches the SEQ number used by the socket. The status message code is 0x16 for the received ACK.
The RxACK status bits are cleared when the regular transmit IB is processed by the IT10G hardware, or when an RxACK status message is transmitted, or by an application that manually clears the bits in the socket CB. The last case is done for diagnostic purposes only.
[MTX data storage] [Overview] This section describes the algorithm that the TOE uses to determine the size that data can be stored in the MTX buffer. This affects the size of TCP packets sent when one MTX buffer corresponds to one TCP packet.
MSS vs. 1/4 MSS (or up to half peer window) If the advertised window of the peer is more than twice the MSS, then the socket MSS size (as stored at CB position 0x05) is a socket about how many bytes can be stored in the MTX buffer. Used as a determinant. If the peer's advertised window does not reach twice the MSS level, then the maximum number of bytes stored in the MTX buffer is (depending on the configuration) a quarter of the MSS or the peer's largest advertised window. It is half. The bit that maintains this comparison is the CB word 0x14 bit [18].
[Packet overhead] Fixed byte overhead is deducted from the MSS (depending on the configuration), or a quarter of the MSS or half the maximum window value of the peer. This overhead includes packet header bytes and IPSEC overhead.
[CB size limit vs. buffer size] After deciding (depending on the configuration) whether to use the MSS, or a quarter of the MSS or half of the peer's maximum window value, and deducting the packet overhead, the logical unit has the type of buffer available. Observe. This limits the storage size if only a 128-byte buffer is available. Also, after the above calculation, if the data size is still larger than the large MTX buffer size, the packet size can be limited again by the MTX buffer size.
[IP router] [IP router features] -Provide a default route capability. -Provide routes for a large number of host IP addresses. -Provide host-specific and network-specific routes. -Dynamic update of routes after ICMP re-guidance. Process IP broadcast addresses (limited, subnet-guided, network-guided broadcasts). -Process IP loopback addresses. -Process IP multicast addresses.
[Module block diagram] FIG. 35 is a schematic block diagram of the IP router.
[IP router operation theory] [Route request] When a local host attempts to send an IP packet, it must decide where to send the packet: another host on the private network, an external network, or the local host itself. This is the task of the IP router to direct the output IP packet to the appropriate host.
When the module requests a route, the sending module sends the destination IP address of the packet to the IP router. The IP router compares the target IP address with the list of destinations stored in the IP route list. If a match is found, the IP router will try to resolve the appropriate Ethernet address. The IP router resolves this by requesting an ARP search for the destination IP address in the ARP cache. Destination If the Ethernet (registered trademark) address is resolved, the Ethernet (registered trademark) address is returned to the transmitting module and used as the destination of the Ethernet (registered trademark) frame that outputs this Ethernet (registered trademark) address. ..
Route information is provided by three separate components: the default route register 351, the custom route list 352, and the non-routable address cache 353. All of these components are queued at the same time when a route request is given.
[Default route] The destination of the packet can be described as local or external. The local destination is attached to the same private network as the sending host. The external destination belongs to a network different from the sending host's premises network.
When the egress packet destination IP address is found to belong to a host attached to the private network, the IP router attempts to set the resolved destination Ip address to its corresponding Ethernet® address. Use ARP. Once it is determined that the destination IP address belongs to the external network, the IP router must determine the gateway host to use to relay the outgoing packets to the external network. Once the gateway hosts are selected, the output IP packets use the gateway host's Ethernet® address as their destination Ethernet® address.
If the route cannot be found for the packet destination IP address, the packet must use the gateway host specified by the default route. The default route is used only when no other route can be found for a given destination IP address.
To minimize the number of accesses to the ARP cache, the IP router caches the default gateway Ethernet® address when the default route is configured. The default gateway Ethernet® address is cached for a maximum amount of time equal to the amount of time that allows dynamic entries in the ARP cache to be cached.
[Broadcast and Multicast Destinations] ARP search is required when the destination IP address is a broadcast or multicast IP address. Instead, the destination Ethernet® address is generated depending on the type of IP address. Packets with a destination IP address set for the IP broadcast address (255.255.255.255) are sent to the Ethernet (registered trademark) broadcast address (FF: FF: FF: FF: FF: FF). Packets with a destination IP address set to a multicast IP address (244.xxx) have their destination Ethernet® address calculated from the multicast IP address.
[Static route] In addition to the default route, the IP router allows the generation of static routes to map the destination IP address to a special Ethernet interface or gateway host. The IP route entry contains the destination IP address, netmask, and gateway IP address. The netmask is used to match the range of destination IP addresses with the destination IP addresses stored in the IP route entry. Netmasks also allow discrimination between special host routes and network routes. The gateway IP address is used when solving the destination Ethernet® address via ARP.
Since it is possible to have many routes in an IP route list, IP route entries are stored in dynamically allocated memory (called m1 memory in this structure). Each IP route entry uses 128 bits. The last 32 bits of each entry do not store any data and are used as padding to align the IP route engine along the 64-bit boundaries. The format of each IP route entry is shown in Figure 36.
The IP route list is organized as a categorized linked list. As IP routes are added to the IP route list, they are ordered according to their netmasks, the top specific IP routes appear in front of the list, and the IP routes with the lowest specific netmasks are listed. Placed at the end of. The route pointer field of an IP route entry contains the m1 memory address, where the next IP route entry can be found in m1 memory. The first (most significant) bit of the route pointer field is used as a flag to determine if the m1 memory address is valid and there is a route following the current one. If the pointer valid bit of the route pointer field is not claimed, then there are no more IP routes in the IP route list and the end of the IP route list is reached.
If the destination IP address is not determined to be a broadcast or multicast IP address, the IP route list is searched for matching IP route entries. If no match is found in the IP route list, the default route is used to provide gateway information.
IP routers also allow the use of numerous physical and loopback interfaces. The interface ID field of the IP route entry allows the IP router to direct the outgoing packet to a specific Ethernet® interface on the IT10G. The interface ID field is also used to direct ARP requests to the appropriate Ethernet® interface.
[Loopback address] If the destination IP address belongs to IT10G or is a loopback IP address (127.xxx), it is expected that the output packet will be fed back to IT10G. The IP router for the loopback destination is stored in the IP route list. IP addresses that are not assigned to IT10G can also be configured as loopback addresses. The interface ID field of the IP route entry must be set to 0x8 to allow this local redirection. Otherwise, the interface ID field of the IP route entry must be set to one of the Ethernet® interfaces (0x0, 0x1, 0x2, etc.).
[Generate route] The new IP route comes from the system interface (eg host computer). The IP routes generated by the system interface are static routes, which means they remain in the table until they are removed by the system interface. The system interface adds and removes routes via the IP router module register interface.
ICMP reintroduction messages are sent when IP packets are being sent to an incorrect gateway host. ICMP redirect messages usually contain the correct gateway host information to be used for inaccurately derived IP packets. When an ICMP reintroduction message is received, the message is processed by the system interface. The system interface is engaged in updating the route list through the IP router's register interface, updating existing IP routes, or creating new IP routes.
Routing to hosts on the local network An IP route with an IT10G subnet mask must be generated to direct the packet directly to other hosts on the Local Ethernet® network. Instead of identifying another host as the gateway for this route, the gateway IP address must be set to 0.0.0.0 to indicate that this route makes a direct connection across the local network.
[Route Request Signaling] Each transmit module has its own interface to the IP router to request a route. The signaling used to request and receive routes is shown in Figure 37.
When a module requests an IP route, the requesting module claims a route request signal (eg TCP_Route_Req) and provides the destination IP address (TCP_Trgt_IP in Figure 21) to the IP router. Once the IP router finds a route for the supplied IP address, the IP router claims a root-done signal and outputs the destination Ethernet® address. The route_valid signal is used to indicate to the transmit module if an IP route is properly discovered. If the route_valid signal is claimed when the route-done signal is claimed, a valid route has been found. If no route_valid signal is claimed, this means that the route was unsuccessful. Does this routing failure have a default route set? This is due to several causes, such as the gateway supplied by the matching IP route entry being down and not responding to the ARP request. In the event of a route failure, the transmit module waits and later either tries to resolve the route again or engages in aborting the current connection attempt.
When an IP route requires an ARP lookup to solve an Ethernet® address to a host or gateway's IP address, delays can occur if the IP address is not found in the ARP cache. When there is a cache loss, that is, when the target IP address has no entry in the ARP cache, the ARP cache notifies the IP router of the loss. The IP router then signals the sending module that requested the IP route that an ARP cache loss has occurred. At this point, the transmit module chooses to delay the current connection configuration or attempts to configure the next connection in the connection queue to request another route. Even if the transmitting module has a route request to the IP router, the ARP search continues. If the target of the ARP search is active and responds to an ARP request from IT10G, the resolved Ethernet® address of the target IP address is added to the ARP cache for possible later use. .. Note: IP routers can have a large number of prominent ARP requests.
[View individual routes] After generating a static route, the user can read back the entries stored in the route table in two ways. If the user knows the target IP address of a given route, the Show_Route command code can be used to display the netmask and gateway for that route.
The Show_Index command can be used to display all entries in the route table. Using the Route_Index register, the system interface may access routes in a specific order. More special (host) routes are displayed first, followed by less specific (network) routes. For example, an IP route entry with route_index 0x001 is the most special route in the IP route list. Note: The default is stored at index zero (0x0000). The Route_Found register is claimed when the route is found properly, and the route data is Route_Trgt, Route_Mask. Stored in the Route_Gw register.
[Unresolvable destination caching] When the IP router is unable to resolve the Ethernet® address to the destination host / gateway, the router caches the destination IP address for 20 seconds. During that time, if the router receives a request for one of these cached unresolvable destinations, the IP router immediately responds to the module requesting the route with the route failure. Caching of this unresolvable destination aims to reduce the number of accesses to the shared m1 memory, where ARP cache entries are stored. Unresolvable destination caching also works to avoid redundant ARP requests. Note: The amount of time unresolved addresses are cached is user configurable via the Unres_Cash_Time register.
[System Exception Handler] [Overview] The following description is the details of the system exception handler. This module is called whenever there is data that the IT10G hardware cannot handle directly. This may be an unknown Ethernet type packet, IGMP packet, TCP or IP option, and so on. For each of these exceptions, the primary parser makes the system exception handler operational when detecting an exception case. The system exception handler module stores data, notifies the system interface (typically an application running on the host computer) that there is exception data to be processed, and acts to send the data to the system interface.
[Exception handler block diagram] FIG. 38 is a schematic block diagram of one configuration of an exception handler.
[Exception handler behavior theory] Exception memory is part of the on-chip processor memory. Each source module capable of generating exception packets has its own exception buffer start address and exception buffer length register. This module puts an exception packet in the FIFO buffer and queues it for storage in on-chip processor memory. The exception_fifo_full signal is fed back to each call module to indicate that the exception FIFO buffer is full. This module can serve only one exception packet at a time. If a subsequent exception packet is received, it is retained until the previous packet is stored in memory.
Exception packets are stored in on-chip processor memory, so host computer access is not granted through this module.
[Memory Allocator 1] [Overview] This section describes the IP module, ARP cache, route table, and memory allocator (malloc1) used to serve the on-chip processor. The allocator first splits the M1 memory into discrete blocks, allocates them on request, and works to put the freed blocks back on the stack.
[Operation theory] malloc1 must have two parameters entered before it can start its operation. These are the overall size of the M1 memory blocks and the size of each memory block. Only one memory size is supported by this allocator.
After entering these parameters, the system claims the M1_Enable bit in the M1_Control register. When this happens, the allocator starts at the top of the M1 memory block and begins embedding the block address. That is, if the M1 memory block is 4 Kbytes deep in total and the block size is 512 bytes, then the M1 memory map looks like shown in FIG.
Four addresses are maintained for each M1 address position relative to the M1 block address. In addition to keeping the starting block address in memory, Malloc1 also contains a 16-entry cache. At initialization, the first 16 addresses are kept in their cache. When blocks are requested, they are taken out of the cache. When the number of caches reaches zero, four addresses (one memory read) are read from memory. Similarly, whenever the cache is full of addresses. The four addresses are written back to memory. This is only done after the allocator first reads the address from M1 memory.
[TX / RX / CB / SA Memory Allocator] [Overview] This section describes the memory allocators used in socket send (malloctx), socket receive (mallocrx), control block (malloccb), and SA (mallcsa) memory. These allocators act to allocate memory blocks on request, return the freed blocks to the stack, and arbitrate to use the memory.
[Operation theory] malloc must have some parameters entered before the start of its operation. These are bitmaps that represent the start and end address pointer positions in MP memory space and each available block in each memory space. Two size blocks, 128 bytes and 2 kbytes, are available in socket data memory. The CB memory has a fixed 128-byte block. All allocators also utilize an 8-entry cache for block addresses of each memory size.
After entering these parameters, the system claims the enable bits in the control register. The allocator can then allocate memory blocks and initiate deallocation.
[Default memory map] Figure 40 shows a sample memory map assuming 256 Mbytes of integrated network stack data memory, that is, MRX and MTX share a common physical memory bank. Current configurations use external DDR DRAM for both MRX and MTX, but in other configurations it is possible to place these memories on-chip or use different types of memory (eg SRAM).
[Network stack data DDR block diagram] The network stack data DDR DRAM is shared between the transmit and receive data buffers and the on-chip processor. The flow of the data block diagram is shown in Figure 41.
The data DDR DRAM arbitrator 411 operates to arbitrate access to DDR DRAM shared between different resources. The overall DDR DRAM 412 is also memory that maps to the memory space of the on-chip processor.
[MTX DRAM Interface and Data Flow] [Overview] The following is an overview of the interface between the malloctx module and the data DDR arbitration module. Describe the data flow, list the interface signals, and explain the required timing details.
[data flow] Three different access types are supported by MTX DRAM. These are burst write, burst read, and single access. malloctx413 mediates requests from different sources for each of these cycle types, but all three cycles are requested simultaneously by the data DDR arbitrator. A block diagram of the mtxarb subunit is shown in Figure 42.
The data written to the MTX memory is first written to the FIFO buffer and then burst written to the DDR DRAM.
[RX DRAM interface and data flow] [Overview] This section describes the interface between mallocrx414 and the receive socket data DRAM controller. It describes the data flow, lists the interface signals, and gives the required timing details.
[data flow] In the receive DRAM, the highest priority is given to the data written in memory. This is because the network stack must keep up with the data being received from the network interface. Writing to the DRAM is first done to the FIFO buffer. When all the data is written to the FIFO buffer, the controller burst writes it to DRAM. In jumbo frames, the TCP receive logical unit splits burst write requests into 2K sized chunks.
For data burst read from DRAM, the DRAM controller reads the memory and writes the requested data to a pair of ping-pong FIFO buffers that supply the PCI controller. These FIFO buffers are needed to transfer the data rate from the DRAM clock to the PCI clock domain.
[Network stack DDR block diagram] The network stack DDR is shared between CB, miscellaneous memory, and SA memory. The flow of the block diagram of the data is shown in Figure 43.
The NS DDR Arbitrator acts to arbitrate access to DDR shared between different resources.
[MCB DRAM Interface and Data Flow] [Overview] This section describes the interface between the malloccb module and the NS DDR arbitration module. It shows the data flow, lists the interface signals, and details the required timing.
[data flow] Two different access types are supported for MCB DRAM. These are burst reads, single access. malloccb arbitrates requests from different sources for each of these cycle types, but both cycles can be requested simultaneously by the NS DDR arbitrator. A block diagram of the mcbarb subunit is shown in Figure 44.
[Network Stack Memory Map] Figure 45 shows the default memory map for the network stack. The address shown is a byte address.
[Miscellaneous memory] [Overview] The following description is the details of the 512K byte miscellaneous memory bank. This memory is used for the purposes listed below. Half-open control block (main), -TCP port authentication table, -UDP source port usage table, -TCP source port usage table, -Waiting time control block allocation table, -Set control block allocation table, -TX memory block allocation table (both 128 and 2K bit blocks), -RX memory block allocation table (both 128 and 2K bit blocks), -FIFO buffer for packets from TCP RX to TX, · A valid bitmap of socket data, -Server port information, -SA entry allocation table.
[Memory organization and performance] Miscellaneous memory is shared with CB memory. Most resources access 256-bit word data to minimize access.
[Definition of stored data] [Half open control block] These are control blocks for half-open TCP connections. Each control block is 64 bytes in size, and there are 4K control blocks in total. Therefore, the number of bytes required for a half-open control block is 4K x 64 = 256K bytes.
[TCP port authentication table] This table keeps track of TCP ports that are authenticated to accept connections. Keep one bit of each of the 64K possible ports. Therefore, this table uses 64K / 8 = 8K bytes. In an alternative configuration, TCP port authentication tables and UDP and TCP source port usage tables can be maintained on the host computer.
[UDP source port usage table] This table keeps track of the UDPP ports available for the source port used for locally initialized connections. Keep one bit of each of the 64K possible ports. Therefore, this table uses 64K / 8 = 8K bytes. This table should not contain local UDP service ports and ports that can be used.
[TCP source port usage table] This table keeps track of the number of available ports on the source port used for locally initialized connections. Keep one bit of each of the 64K possible ports. Therefore, this table uses 64K / 8 = 8K bytes.
[Wait time control block allocation table] This is the allocation table for the wait time control block. Keep 1 bit of each of the 32K wait time control blocks. Therefore, this allocation table uses 32K / 8 = 4K bytes. This module uses a total of 16-bit data buses.
[Set control block allocation table] This is the configured control block allocation table. Keep one bit of each of the 64K control blocks. Therefore, this allocation table uses 64K / 8 = 8K bytes.
TX Socket Data Buffer Block Allocation Table This table is made up of a 2 Kbyte block allocation table and a 128 Kbyte block allocation table, which are used in the dynamically allocated transmit data buffer memory. The number of blocks for each type is configurable, but the size of both of the joined allocation tables is fixed at 72 Kbytes. This allows up to 475K 128-byte blocks. At this level, the number of 2K byte blocks is 98K.
[RX Socket Data Buffer Block Allocation Table] This table is made up of a 2 Kbyte block allocation table and a 128 Kbyte block allocation table, which are used in the dynamically allocated receive data buffer memory. The number of blocks for each type is configurable, but the size of both of the joined allocation tables is fixed at 72 Kbytes. This allows up to 475K 128-byte blocks. At this level, the number of 2K byte blocks is 98K.
[TCP RX FIFO buffer] This FIFO buffer is used to keep track of packet transmission requests from the TCP receive logical unit to the TCP transmit logical unit. Each TCP RX FIFO buffer entry is made up of several control flags and control block addresses with a total of 4 bytes (4 flags, 26-bit address, 2 unused bits). This TCP RX FIFO buffer is 1024 words deep and therefore requires 1024 x 4 = 4K bytes.
[Available Bitmap of Socket Data] This bitmap represents a 64K socket with data ready to be sent to the host system. Keep 1 bit on each socket. Therefore, this bitmap uses 64K / 8 = 4K bytes.
SA Entry Allocation Table This is the SA entry allocation table. Keep 1 bit for each 64K SA entry. Therefore, this allocation table uses 64K / 8 = 4K bytes.
[Server port information] This database is used to store parameter information for TCP ports that are open in the LISTEN state. Port-specific parameters are maintained in this area because these server ports do not have a CB associated with them until they are opened. Each port entry consists of 2 bytes and there are 64K possible ports. Therefore, this database requires 64K x 2 = 128K bytes.
[Miscellaneous memory map] The memory map used for miscellaneous memory is configurable. The default settings are shown in Figure 46. Blocks are not scaled.
[Miscellaneous memory operation theory] [Module initialization] The CPU rarely needs to be initialized before starting the miscellaneous memory arbitrator. If the default memory map is used, the CPU can enable the arbitrator by claiming the MM_Enable bit in the Misc_Mem_Control register.
If a non-default memory map is used, all base address registers must be initialized before the arbitrator is available. It is the responsibility of the host computer software to ensure that the programmed base address does not generate any overlapping memory areas. No hardware is given to check this.
[CPU access] The CPU can access any location in miscellaneous memory. This is done by first programming at the address to the MM_CPU_Add register (0x1870-0x1872) and then reading bytes or writing to the MM_CPU_Data register (0x1874). The address register is automatically incremented each time the data register is accessed.
[MIB support] [Overview] This section describes MIB support built into the IT10G network stack. It contains a definition of registers, a theory of behavior, and an overview of what the MIB is used for.
[Introduction to SNMP and MIB] SNMP is a management protocol that allows the SNMP management station to obtain statistics and other information about networked devices and allows the devices to be configured. The software that runs on the device is called an agent. The agent runs on top of the UDP socket and handles requests from the SNMP management station. The agent can also send a trap to the SNMP management station when an event occurs on the device. The SNMP RFC documents a standard set of information objects called the Management Information Base (MIB) when grouped together. Vendors can also specify the unique MIBs they support, in addition to the applicable standard MIBs specified in the RFC.
A standard MIB defines the information that most devices of a particular type must provide. Some information can be fully processed by the SNMP agent software, and the rest require some level of support from the operating system, drivers and equipment. SNMP support is provided for all three major transferables, hardware, embedded software, and drivers.
In most cases, hardware support is limited to collecting statistical information about certain counters associated with the hardware networking layer. Interrupts can also be triggered by an event. Embedded software development also supports counters for layers executed by software such as ICMP, exception handling, etc. In addition, the embedded software can query the hardware-generated database to report on things like ARP table entries, TCP / UDP socket status, and other non-counter statistics. Driver support focuses on interfacing embedded platforms into APIs that allow related objects to be returned to the SNMP agent, allowing some level of configuration of the SNMP agent.
[MMU and timer module] [Overview] The following description provides an overview of MMUs used in general purpose timers and systems.
[CPU timer] There are four general purpose 32-bit timers that are cascaded or independent of the previous timer. All timers can be operated in single shot or loop mode. In addition, a clock prescaler is provided that can split the main core clock before it is used by each timer. This allows minimal code changes for different core clock frequencies.
[Timer test mode] Each individual timer can be put into test mode by claiming the Timer_Test bit in the corresponding timer control register. When this mode is activated, the 32-bit counter increments at different rates according to the three lsbs in the 8-bit clock divider setting.
[On-chip processor MMU] The on-chip processor uses an MMU to divide the processor's 4GB memory space into different areas. Each area can be individually identified by a base address and an address mask. The areas may also overlap, but in this case the data read back from the shared memory area is unpredictable.
[On-chip processor DRAM interface and data flow] [Overview] This section describes the interface between the on-chip processor and the network stack (NS) data DDR arbitration module. It shows the data flow, lists the interface signals, and details the required timing.
[data flow] Three different access types are supported on the on-chip processor DRAM. These are burst write, burst read, and single access. In addition, locks can be applied so that consecutive single accesses are made by the on-chip processor without abandoning the DRAM. The on-chip processor memory arbitrator arbitrates requests from different sources for each of these cycle types, but all three cycles are requested simultaneously by the NS DDR arbitrator. It is assumed that only the on-chip processor itself uses the locking property when it needs to be enabled via the register bit.
[Instruction / state block behavior theory] [Overview] The following description, the host, the on-chip processor, the operation management of the instruction block passing between the network stack (IB) and status block (SB) is the outline of the theory. It also covers the concept of MDL.
The individual IBs and SBs are concatenated into a queue. There are IB and SB queues. Each queue has a host computer component and an on-chip processor component. Also, each IB queue has an associated SB queue. Matching IB and SB queues together form a channel. This is shown in Figure 47.
Note: All states and instruction queue lengths must be multiples of 16 bytes and are DWORD aligned.
The preferred embodiment supports four channels. For each queue (both IB and SB), there is a queue descriptor. These descriptors detail the queue length, where the queue resides, and the read and write pointers. Each queue is treated as a circular FIFO buffer. The parameters that define each queue are divided into registers in the host register interface for the parameters that make up host memory usage and registers in the general network stack register space that defines on-chip processor memory usage. The format of the queue descriptor is shown in the table below. The address offsets listed here relate to the start of the queue descriptor register.
The processing flow of the instruction block queue is shown in Figure 48. It also shows its association with the IB parser module.
[SB Passing] A block diagram showing the data flow of state blocks passing between the network stack, the on-chip processor, and the host computer is shown in FIG.
The SB timer threshold is the main interrupt aggregation mechanism in the integrated network adapter hardware. Interrupt aggregation reduces the number of interrupts between the integrated network adapter and the host computer, and reducing or a set of interrupts increases data processing capacity.
The HW can send a general SB to an on-chip processor and a socket-specific SB directly to the host computer. Host computer status messages include socket configuration, RX DMA, CB_Create, and Socket_RST status messages.
[SB from HW to host] When a socket-specific event occurs, the owner of the CB is determined by inspecting the HS bits in the socket CB structure. If the socket belongs to the host computer, the CB channel number is also read from the CB. The channel number is sent to the HW status message generator (statgen module) 491 together with the event notification.
The statgen module forms the appropriate SB and places it in one of the four DMA FIFO buffers 492. These FIFO buffers are physically located in the on-chip processor memory and are identified by the on-chip processor SB FIFO buffer address and the on-chip processor SB FIFO buffer length.
[SB from HW to on-chip processor] When a non-socket specific event, an exception Ethernet® packet, occurs, the statgen module generates the appropriate SB and sends it to one state message queue defined in the on-chip processor memory. Interrupts can be generated for the on-chip processor at this time if statgen_int is enabled. There is only one state message queue defined from the hardware to the on-chip processor, and all SBs go to this one queue.
The SB queue actually resides in on-chip processor memory. When entries are written to this memory area, they are automatically drawn into the on-chip processor's register FIFO buffer for reading through the STAT_READ_DATA register (network stack general register 0x005C-0x005F). The on-chip processor can also pole to see how much data is available by reading the STAT_FIFO_FILLED register at 0x005A. This register returns the number of double words available in the queue.
[SB from on-chip processor to host computer] When the on-chip processor needs to send an SB to one of the host's SB queues, it sends it through the statgen hardware module. In this case, the on-chip processor generates an SB in its memory and programs the start address of this SB, the channel number of the SB, and the length of the CB sent to the statgen module. The HW then forwards the SB (or multiple SBs) to the appropriate SB queue to the host. In this way the on-chip processor and the SB from the HW are mixed in the same SB queue supplying the SB DMA engine 493.
The on-chip processor can be notified that the SB has been transferred to the appropriate SB queue via interrupts or by polling the state bits. Before writing any parameters for this feature, the on-chip processor should ensure that the statgen module has completed the transfer of any destination SB request, i.e. the on-chip processor has the SB bit statgen. You should make sure that it is not claimed in the command register.
[MDL processing] The MDL (Memory Descriptor List) is first DMAed from the host computer memory to the DDR DRAM on-chip processor memory. The on-chip processor returns the MDL handle to the host based on the MDL DDR address.
The SendMDL32 / 64IB, which includes the MDL handle, offset to MDL, and transfer length, is analyzed by the IB parser and performs the following operations.
Read the first entry in the MDL to understand if the identified offset is within this entry. This determination is made by comparing the offset with the length of the first MDL entry.
If it does not fit within the first MDL entry, read the next MDL entry. When it finally finds the correct MDL entry, it uses the address in the MDL to queue the TX DMA transfer. The transfer extends from the point at which the offset begins to a large number of MDL entries.
When the on-chip processor acquires the IB, the on-chip processor programs the offset into the socket CB. When the CB acquires this information, it can DMA the data. Instead, if the host computer knows the offset unexpectedly quickly, the offset can be programmed to the CB unexpectedly quickly.
Transfer thresholds that specify the number of bytes that must be DMAd before the RX DMA state message is generated are still applicable in MDL mode.
[Instruction block and status message] [Overview] The following description details the instruction block and status message formats. The host uses the IB to transfer commands to the IT10G hardware, and the IT10G hardware uses state blocks to transfer information to the on-chip processor and the host computer.
[Instruction block structure] The only IB that the hardware parses directly is the various forms of Send_Command.
[HW parsed instruction] These instructions are used to send data and request the IT10G hardware to close or post the receive parameters of a given socket. There are four forms of the SEND instruction to handle TCP, UDP, 32-bit and 64-bit host computer memory addressing.
[ISCSI support] One embodiment of IT10G uses the host computer to assemble the iSCSI PDU header, while another embodiment uses IT10G hardware and an on-chip processor to assemble and process the iSCSI PDU. Both embodiments will be described below.
[ISCSI support] [Overview] This section describes the IT10G iSCSI hardware support that uses the host computer to assemble the iSCSI PDU header. The IT10G hardware offloads the iSCSI CRC calculation for transmission and reception, and the hardware uses fixed interval markers (FIM) for transmission for iSCSI framing. Framing is not currently supported for reception.
[Operation theory] The iSCSI protocol is specified in the IETF iSCSI Internet Draft, a normative document for the definition of hardware features. The iSCSI protocol encapsulates SCSI commands in a protocol data unit (PDU) that is carried over a TCP / IP byte stream (SCSI commands are command descriptor blocks or CDBs). SCSI commands are documented in several standards. The hardware receives the iSCSI header segment and the iSCSI data segment from the host iSCSI driver and prepares to send the iSCSI PDU. The hardware receives the iSCSI PDU, calculates the iSCSI CRC, and sends the result to the host iSCSI driver.
The hardware is primarily related to the external format of the iSCSI PDU, which contains the iSCSI header segment and the iSCSI data segment. An iSCSI PDU consists of the required basic header segment (BHS), followed by zero or more additional header segments (AHS), followed by zero or more data segments. The iSCSI CRC is optional and is included in the iSCSI PDU as a header digest and data digest. The configuration described allows the iSCSI headers and data to be separated and copied to the host computer's memory without the need for an additional memory copy to the host computer.
[ISCSI transmission] [Overview] A block diagram of the iSCSI transmit data path is shown in Figure 50.
[iSCSI control module] The iSCSI control module 501 receives control signals from the TX DMA engine module 502, the CB access module 503, and the Statgen module 504.
The iSCSI control module provides control signals to the CB access module 503, the iSCSI CRC calculation module 505, and the FIM insertion module 506. The iSCSI control module calculates the iSCSI PDU length containing any iSCSI CRC word and any marker, and provides the iSCSI PDU length containing any iSCSI CRC word and any marker to the XMTCTL MUX module.
The host iSCSI driver assembles the complete iSCSI PDU header in host memory. The host iSCSI driver then generates a TCP transmit 32 iSCSI IB or a YCP transmit 64 iSCSI IB, depending on whether a 32-bit address or a 64-bit address is required. The host iSCSI then sends a TCP transmit iSCSI IB, which is received by the on-chip processor. The effect of the host iSCSI driver sending a TCP transmit iSCSI IB to the on-chip processor is to instruct the iSCSI control module to initiate a DMA transfer of the iSCSI PDU to the concatenated list of buffers in the host computer memory. ..
An iSCSI PDU containing BHS, any AHS, or any data segment is forwarded using TCP transmit 32 iSCSI IB or TCP transmit 64 iSCSI IB.
It is difficult to design an iSCSI CRC compute module that must insert iSCSI at the end of the iSCSI PDU header segment or at the end of the iSCSI PDU data segment if the iSCSI PDU is not included in a completely single DMA transfer. is there.
When the iSCSI PDU is as small as 48 bytes (BHS), the overhead of using a single IB for each iSCSI PDU merges the iSCSI PDUs for a single DMA transfer and then separates the iSCSI PDUs for processing again. No more than the overhead to do.
The first transfer block must be one complete iSCSI PDU header or multiple headers, i.e. BHS followed by 0, 1, or more AHS, and all headers are in the first transfer block. included.
The first transfer block must be one complete iSCSI PDU header or multiple headers so that the CRC compute module can insert the iSCSI header CRC at the end of the iSCSI PDU header.
The iSCSI control module should check that the TCP transmit 32 iSCSI IB or TCP transmit 64 iSCSI IB contains at least one transfer block. This is a necessary condition for correct operation. If this condition is not met, it is an error and the module is ineligible.
The iSCSI control module initiates the DMA transfer of the iSCSI PDU by requesting the DMA engine module to allow the DMA transfer. When the iSCSI control module's request to perform a DMA transfer is accepted, the iSCSI control module is locked to access to the DMA engine. The DMA engine is locked until all forwarding blocks in the iSCSI IB have been serviced. All iSCSI IB transfer blocks are serviced when all iSCSI headers and iSCSI data information are transferred from host memory to hardware by DMA. The DMA engine signals the iSCSI control module as all iSCSI headers and iSCSI data information are transferred from the host memory to the hardware by DMA.
The iSCSI control module facilitates FIM and CRC operations and locks access to the DMA engine to take into account the iSCSI login phase.
2 bits of iSCSI IB, iSCSI CRC selection bit 1 and iSCSI CRC selection bit 0 are iSCSI control if the iSCSI CRC calculation module should calculate the iSCSI header and / or iSCSI data and / or other iSCSI CRC. Set by the host iSCSI driver to signal the module. These two iSCSI CRC selection bits may take any combination of 1 and 0.
The iSCSI IB corresponds to a linked list (LL) of buffers in host memory and contains a set of address and length pairs known as forwarding blocks. A linked list of buffers in host memory stores iSCSI headers and iSCSI data information. The transport block contained in the TCP Send iSCSI IB provides iSCSI control module information about where to find the iSCSI header and iSCSI data via a linked list of buffers. The first transfer block in each iSCSI IB points to the iSCSI header in host memory, and the remaining transfer blocks in the iSCSI IB point to the iSCSI data in host memory.
The iSCSI control module takes an iSCSI PDU length that contains CRC words but excludes markers, and calculates the length of the iSCSI PDU that contains any CRC words and any markers. The XMTCTL MUX module requires the length of the iSCSI PDU containing CRC words and markers to provide the XMTCTL module with the length of data that will be input to the XMTCTL module.
The XMTCTL MUX module requires a PDU length containing any CRC word and any marker before storing the iSCSI PDU in the MTX buffer.
The iSCSI control module is included in the TCP transmit 32 iSCSI IB or TCP transmit 64 iSCSI IB and calculates the length of the iSCSI PDU by using the iSCSI PDU length that contains any CRC words but excludes any markers.
Use the following information to calculate the PDU length containing any CRC word and any marker. FIM interval, marker interval in bytes, -Current FIM count, number of bytes since the last marker was inserted and stored in socket CB, · PDU length included in TCP transmit 32 iSCSI IB or TCP transmit 64 iSCSI IB, including CRC words but excluding markers.
Using two counters, the PDU length counter and the marker counter, the calculation of the PDU length containing any CRC word and any marker is equal to: 1. Initialize the marker counter to zero. 2. Initialize the PDU length counter to the current FIM count. 3. Check the PDU length counter for PDU lengths that contain any CRC word but exclude any markers. 4. Exit the loop if the PDU length counter contains any CRC word but is greater than the PDU length excluding any markers. 5. Add the FIM interval to the PDU length counter and increment the marker counter by 1. 6. Proceed to step 3. 7. Calculate the PDU length containing any CRC word and any marker by adding the marker length calculated from the marker count.
[Another FIM algorithm] 1. Initialize total_marker_length to zero. 2. Initialize length_counter to the next marker (FIM_interval-current_FIM_count). 3. If length_counter> PDU_length, go to step 7. 4. Increment total_marker_len by marker_size. 5. length_counter = length_counter + FIM_interval. 6. Proceed to step 3. 7. xmitctl gets total_marker_len + PDU size.
Note that the final marker must be positioned if the marker must be placed exactly after the PDU to ensure that it handles the case with the current FIM count = 0 correctly.
In the above calculation The FIM interval comes from the CB. Current_FIM_count comes from CB. total_marker_length is a time variable for the FIM insert module. length_counter is the time variable of the FIM insertion module. marker_size is the total size of the FIM marker in bytes.
FIG. 51 shows an iSCSI transmission flowchart.
[iSCSI CRC calculation module] The DMA TX module provides an iSCSI PDU to the iSCSI CRC calculation module, which consists of an iSCSI header segment and an iSCSI data segment.
The output from the iSCSI CRC compute module, which is a PDU containing any inserted CRC words, is given to the FIM insert module.
The iSCSI CRC calculation module calculates the iSCSI CRC value for both the iSCSI PDU header and the iSCSI PDU data.
The CRC calculation module inserts an iSCSI CRC at the end of the iSCSI PDU header and / or at the end of the iSCSI PDU data, or both. The calculation of the iSCSI CRC is controlled and instructed by the iSCSI CRC selection bit 1 and the iSCSI selection bit 0 of the TCP transmission iSCSI IB.
The input to the CRC compute module is a BHS followed by 0, 1, or more AHS in the first transfer block and a BHS followed by a 0 or 1 data segment in the second transfer block.
The DMA engine must send a signal to the iSCSI compute module at the boundary of the transfer block. The iSCSI compute module must know where the first transfer block starts and ends and where the last transfer block ends. The iSCSI header CRC always inserts the iSCSI PDU at the end of the first transfer block when the header CRC becomes available. The iSCSI data CRC always inserts the iSCSI PDU at the end of the last transfer block when the data CRC becomes available.
It is critical that the first transfer block of the IB is a complete iSCSI header (BHS plus 0, 1 or more AHS).
[FIM Insert Module] The FIM Insert module inserts a fixed interval marker (FIM) into an iSCSI PDU, which is equivalent to inserting a fixed interval marker into an iSCSI transmit stream.
When the FIM insert module is active, the FIM insert module inserts only markers. The FIM insertion module is active only when FIM is configured on the CB corresponding to the current iSCSI socket (iSCSI socket CB).
The FIM insertion module must be able to keep track of the 4-byte word count of each iSCSI stream containing any CRC bytes that can be inserted by the iSCSI CRC calculator for each iSCSI socket. The FIM insertion module maintains a 4-byte word number tracking by storing two FIM values in the iSCSI socket CB, namely the FIM interval and the current FIM count. Each of these fields, the FIM interval and the current FIM count are measured in 4-byte words.
The FIM interval and current FIM count can be stored in 16-bit fields.
If the FIM interval or the current FIM count overflows, then a qualified failure mechanism should exist.
If a marker is used to enable an iSCSI connection setup that includes an iSCSI login phase, marker insertion will only begin at the first marker interval after the end of the iSCSI login phase. However, to allow the marker inclusion and exclusion mechanism to operate without knowing the length of the iSCSI login phase, the first marker is in the iSCSI PDU stream as if the unmarked interval contained the marker. Be placed. Therefore, all markers appear in the iSCSI PDU stream at the byte position given by the following equation. [(MI + 8) * n-8] Here, MI = FIM (marker) interval, and n = integer.
As an example, if the marker interval is 512 bytes and the iSCSI login phase ends at byte 1003 (the first iSCSI positioned byte is 0), then the first marker is after byte 1031 of the stream. Will be inserted.
The iSCSI PDU stream is specified in the iSCSI Internet Draft, but the term used in the Internet Draft is TCP Stream. FIM uses payload byte stream counting, which contains any bytes located by iSCSI in the TCP stream except the marker itself. This excludes any bytes that TCP counts but is not started by iSCSI.
The host iSCSI driver initializes the FIM interval and initial FIM count of the iSCSI socket CB by generating the iSCSI FIM interval command block. The on-chip processor receives the iSCSI FIM interval command block. The on-chip processor writes the FIM interval to the socket-specific TX_FIM_Interval register and the FIM interval field of the iSCSI socket CB. The on-chip processor writes the initial FIM count to the socket-specific TX_FIM_COUNT register and the current FIM count of the iSCSI socket CB.
The host iSCSI driver controls the FIM status of the iSCSI socket CB by generating the iSCSI set FIM status command block. The on-chip processor receives the iSCSI set FIM state command block. The on-chip processor then writes the FIM state to the FIMON bit in the iSCSI socket CB.
After initializing the FIM interval, initial FIM count, and FIM state, the FIM insert module is in the initialization state waiting for the iSCSI IB.
Attempts to change the FIM state, FIM count, or FIM interval for the current iSCSI socket while the FIM is active are in error.
When the FIM insert module is waiting for an iSCSI IB and an iSCSI IB is received, the FIM insert module reads the FIMON bit in socket CB and the current FIM count. If the FIMON bit in the iSCSI socket CB is set, the FIM insert module is active. If the FIM insert module is active, the FIM insert module sets the FIM interval counter to the current FIM count from the iSCSI socket CB. If the FIM insert module is not active, the FIM insert module returns to the state of waiting for the iSCSI IB.
The FIM insert module is now ready to count each 4-byte word of data received from the CRC calculator. After each 4-byte word in the iSCSI PDU is received by the FIM insert module containing any CRC bytes inserted by the iSCSI CRC calculator, the FIM insert module decrements the FIM interval counter by one. The FIM insert module then returns to counting each 4-byte word.
When the FIM insert module completes scanning data from the iSCSI socket while counting each 4-byte word, the FIM insert module saves the current FIM interval counter value to the current FIM count of the current iSCSI socket CB. To do. The FIM insert module completes scanning data from the iSCSI socket when all transfer blocks in the iSCSI IB have been serviced. At this point, the FIM insert module is in a state of waiting for the iSCSI IB.
While in a state of counting each 4-byte word, the FIM interval insert module inserts a marker in the iSCSI PDU when the FIM interval counter reaches zero. After inserting the marker in the iSCSI PDU, the FIM insertion module resets the FIM interval counter to the FIM interval, reads it from the socket CB, and then returns to counting 4-byte words.
The marker contains the next iSCSI PDU start pointer equal to the number of bytes skipped in the iSCSI stream to the next iSCSI PDU header. The FIM insertion module must calculate the next iSCSI PDU start pointer for each marker. To calculate the next iSCSI PDU start pointer, the FIM insert module must count the number of bytes in the current PDU and measure from the start of the current iSCSI PDU to the marker insertion point, which is the current iSCSI PDU byte count. .. To count the number of bytes in the current iSCSI PDU, the FIM insert module uses the iSCSI PDU byte counter. In addition to the current iSCSI PDU byte count, the FIM insert module must know the start position of the next iSCSI PDU header with respect to the start of the current PDU header. The difference between the start of the next iSCSI PDU header and the start of the current iSCSI PDU header is equal to the current iSCSI PDU length. The FIM insert module calculates the next iSCSI PDU start pointer as follows: Next iSCSI PDU Start Pointer = Current iSCSI PDU Length-Current iSCSI PDU Byte Count The next iSCSI PDU start pointer is 32 bits long.
The iSCSI Internet Draft specifies the maximum value MaxRecvPDUDataSize negotiated for the PDU data segment size from 512 bytes to ((224 **) -1) bytes (16777216-1 bytes or about 16 Mbytes).
The current iSCSI PDU length (measured in bytes) should be at least 25 bits.
There are several ways to give the FIM insert module an iSCSI PDU length that contains arbitrary CRC words but excludes arbitrary markers. One way to provide the PDU length to the FIM module is for the iSCSI host driver to calculate the PDU length.
The iSCSI host driver calculates the current iSCSI PDU length excluding the iSCSI PDU header CRC or iSCSI PDU data CRC as the sum of 48 bytes (BHS length) + total AHS length + data segment length.
The total AHS length and data segment length must already be calculated in software when they are included in the BHS.
The host iSCSI driver then calculates the current iSCSI PDU length that contains any CRC word but excludes markers. The current iSCSI PDU length that contains any CRC word but excludes markers is equal to the iSCSI PDU length that excludes the iSCSI PDU header CRC or the iSCSI PDU data CRC, +4 bytes if the iSCSI PDU header CRC is available. , Equal to +4 bytes if iSCSI data CRC is available.
The host iSCSI driver inserts the current iSCSI PDU length into the TCP transmit 32 iSCSI IB or TCP transmit 64 iSCSI IB that contains any CRC word but excludes markers. Please refer to the explanation of TCP transmission 32iSCSI IB and TCP transmission 64iSCSI IB.
[XMTCTL Mux Module] The XMTCTL mux module receives TCP data from the DMA Tx FIFO buffer and iSCSI data from the FIM insert module.
The XMTCTL mux module provides data from the DMA TX FIFO buffer or FIM insert module to the XMTCTL module.
Function: The XMTCTL mux module acts to multiplex the data input to the XMTCTL module.
The XMTCTL mux module provides the data length to the XMTCTL module.
The length of the iSCSI data must include any CRC bytes inserted by the CRC calculation module.
The length of the iSCSI data must include any markers inserted by the FIM insertion module.
The XMTCTL module uses the size of the input data to separate the data into the appropriate dimensions of the TCP packet, based on MSS etc. The XMTCTL module generates packets of appropriate size stored in the MTX memory buffer. Each MTX buffer corresponds to a separate TCP packet.
The XMTCTL mux module provides data to the XMTCTL module in 128-bit wide format. The data is sent to the XMTCTL module using the dav / grant handshake.
[Receive iSCSI] [Overview] The receive DMA engine includes the iSCSI receive data path. The receive DMA data path calculates the iSCSI CRC and passes the last iSCSI CRC to the host iSCSI driver via the RX DMA state with the iSCSI CRC message.
[Operation theory of iSCSI reception] The hardware offloads the iSCSI CRC calculation. The host iSCSI driver controls the incoming data path to ensure the correct operation of the iSCSI CRC mechanism.
iSCSI CRC information is transferred between the hardware and the host iSCSI driver in two places. The final iSCSI CRC is part of the RX DMA state with the iSCSI CRC message sent from the hardware to the host iSCSI driver. The iSCSI CRC seed is part of both the TCP receive iSCSI 32IB and the TCP receive iSCSI 64IB sent from the host iSCSI driver to the hardware.
The host iSCSI driver must post a buffer of the correct size during TCP receive iSCSI 32IB and TCP receive iSCSI 64IB to keep the iSCSI CRC calculator working properly.
During the iSCSI recording phase, the host iSCSI driver treats the iSCSI socket in the same way as a normal TCP socket. When the iSCSI recording phase is complete, the iSCSI connection enters the full feature phase.
When an iSCSI connection enters the full feature phase, the connection is ready to transfer an iSCSI PDU that may contain iSCSI CRC information in the iSCSI header and / or iSCSI data.
The RX DMA engine sends the last iSCSI CRC calculated via the RX DMA state with the iSCSI CRC message from the scattered set list identified by TCP receive iSCSI 32IB and TCP receive iSCSI 64IB.
Once the iSCSI recording phase is complete, the host iSCSI CRC driver posts a receive buffer of the correct size for the iSCSI PDU header. The host iSCSI driver then parses the iSCSI PDU header. The host iSCSI driver then posts a receive buffer of the correct size for the iSCSI PDU data segment if the data segment exists. Each host iSCSI driver continues to post the correct size buffer for the next iSCSI PDU data segment. The DMA engine calculates the iSCSI CRC for each iSCSI PDU header segment and iSCSI PDU data segment.
For iSCSI PDU headers, the host iSCSI driver posts the sum of the iSCSI base header segment (BHS) size (48 bytes) and the iSCSI PDU CRC (4 bytes) if the header CRC negotiates during the iSCSI login phase. .. The BHS is then forwarded by the DMA, including the subsequent iSCSI PDU header, if present.
When the BHS DMA transfer is complete, the RX DMA state with the iSCSI CRC message contains the rest of the final iSCSI CRC calculated by the hardware. The host iSCSI driver receives an RX DMA state with an iSCSI CRC message containing the rest of the final iSCSI CRC calculated by the hardware. The host iSCSI driver then checks the last iSCSI CRC in the RX DMA state to ensure that its value matches the expected remainder of this iSCSI polynomial. If the header CRC is not negotiated during the iSCSI recording phase, the host iSCSI driver ignores the rest of the last iSCSI CRC returned in the RX DMA state with the iSCSI CRC message.
BHS is followed by an additional header segment (AHS). Byte 4 of BHS contains Total AHS Length, that is, the total length of AHS. For iSCSI PDU headers that contain AHS and are therefore longer than 48 bytes, the host iSCSI driver will generate an additional TCP receive iSCSI 32IB or TCP receive iSCSI 64IB for the AHS.
If iSCSI CRC is enabled, the host iSCSI driver seeds the calculated iSCSI CRC value using the iSCSI CRC seed field in TCP receive iSCSI 32IB or TCP receive iSCSI 64IB. The use of the iSCSI CRC seed field allows the rest of the iSCSI PDU header CRC to be checked accurately.
If no continuation of the iSCSI CRC calculation is required, the iSCSI CRC seed in TCP receive iSCSI 32IB or TCP receive iSCSI 64IB must be set to zero.
When requesting an iSCSI CRC for the data section of an iSCSI PDU, the host iSCSI driver must identify an extra 4-byte data buffer to store the iSCSI CRC. This 4-byte data buffer is inserted as the final transfer block and efficiently adds an extra 4 bytes to the end of the linked list with TCP receive iSCSI 32IB or TCP receive iSCSI 64IB. The 4-byte data buffer allows the iSCSI CRC of the iSCSI PDU data segment to be transferred outside the iSCSI PDU data stream. The use of the 4-byte buffer is shown in Figure 52.
[RX DMA operation theory] The host driver knows that there is RX data available when receiving the RX DAV status message from the hardware. The notification of this RX DAV status message is the same as a normal TCP connection. The CB handle in the RX DAV status message points to the CB associated with iSCSI. The host driver then allocates a buffer to store the RX data, and the total linked list buffer size must be equal to the amount of data the host driver expects to receive plus CRC if expected (see Figure 52). .. Also, since iSCSI PDUs are expected to be in that format, the total buffer size must always be aligned on the dword boundary. The host driver must also initialize the CRC seed to 32'hffffffff.
When the RX DMA operation is performed, the hardware first extracts the iSCSI CRC seed from the socket CB attachment. The hardware also extracts the host retransmission (HR) bits from the main CB structure. The HR bit indicates to other hardware whether this is an iSCSI socket. The hardware then calculates a new CRC while transferring the data.
The data stream may or may not contain the expected iSCSI CRC value for the PDU header or data segment. If the data stream does not contain an iSCSI CRC digest, the iSCSI CRC calculated by the hardware is stored back in the CB attachment as a working iSCSI CRC. The RX-DMA status message may or may not be generated. If an RX DMA status message is generated, the status message contains this working iSCSI CRC. If the data stream contains an iSCSI CRC digest at its end, the iSCSI CRC hardware computes the iSCSI CRC across the PDU segment containing the iSCSI CRC digest, which is the expected iSCSI CRC value. If the PDU segment is not corrupted and the iSCSI CRC digest is accurate, the final iSCSI CRC value calculated by the hardware is always equal to the remainder of the fixed iSCSI CRC.
Once all posted buffer lists are filled, the Statgen module raises an RX DMA state with an iSCSI CRC message containing the remainder of the last iSCSI CRC.
The on-chip processor operates to handle TCP receive iSCSI 32IB or TCP receive iSCSI 64IB. During IB processing, the on-chip processor writes the iSCSI CRC seed, which may be 32'hffffffff, to indicate the start of a new CRC calculation, or the iSCSI CRC seed from the previous RX DMA status message to the socket CB attachment. Include. In this way, writing the iSCSI CRC seed sets the correct value for the iSCSI CRC.
For other non-iSCSI mechanisms that use standard receive 32 and receive 64 lbs and follow the RX DMA engine, the CRC mechanism is ignored and the CRC seed is meaningless and can be skipped.
Before the RX DMA operation takes place, the peer sends one or more TCP packets containing the data payload that will eventually be transferred by DMA to host memory. Depending on how the peer's network stack behaves, this payload can be split into non-dword aligned boundaries, even if iSCSI requires all PDUs to be themselves aligned at the dword boundary. The DMA operation does not directly match the payload size of each TCP packet. If the payload is non-dword aligned, then the DMA operation is non-dword aligned and it is necessary to calculate the CRC through this non-dword aligned data stream. The problem is that the iSCSI CRC hardware predicts that all the data will be dword aligned.
When iSCSI CRC hardware encounters this predicament, it continues to compute iSCSI CRC through an unaligned data stream. When it finally reaches, it stores 1-3 bytes of non-double word aligned to the CB attachment. It also remembers the number of valid bytes. For example, if the data stream is 15 bytes long, the hardware will remember the last 3 bytes and the number of valid bytes is 3. If the data stream is 32 bytes long, then this data stream is dword-aligned and the number of valid bytes is 0. The information stored in the CB attachment is retrieved and inserted before the next DMA transfer.
[iSCSI instruction block] TCP Send iSCSI Instruction Block (IB) is used to send an iSCSI PDU using an iSCSI socket. The length of the TCP transmit iSCSI 32 IB is variable because this IB can contain one or more transfer blocks used in DMA transfers. A transfer block consists of a pair of address (header or data) and transfer length.
The iSCSI CRC selection bit 1 and the iSCSI CRC selection bit 0 determine the iSCSI CRC calculation by the iSCSI control module. If the iSCSI CRC selection bit 0 = 1, the iSCSI control module calculates the iSCSI CRC across the iSCSI header. If the iSCSI CRC selection bit = 1, the iSCSI control module calculates the iSCSI CRC over the iSCSI data. The iSCSI CRC calculation can be performed by the iSCSI CRC calculation engine for both the iSCSI header and the iSCSI data, the iSCSI data only, the iSCSI header only, or the rest.
The total DMA length of iSCSI is the total length of the data that needs to be DMAed by the iSCSI transmit block. This does not include the CRC length that needs to be calculated.
The pair of address and transfer length of the first transfer block of TCP transmit iSCSI 32IB points to the iSCSI header in host memory. The second and optional additional transfer blocks of the TCP transmit iSCSI 32IB point to the iSCSI data in host memory.
An iSCSI PDU containing a BHS and any AHS and any data segment is forwarded using TCP transmit 32 iSCSI IB or TCP transmit 64 iSCSI IB.
The first transport block of a TCP transmit iSCSI 32IB must be a complete single or multiple iSCSI PDU header, a BHS followed by 0, 1, or more AHS, and all single or multiple headers must be a TCP transmit iSCSI 32IB Must be included in the first transfer block of.
To say the same thing in a different way to emphasize, the host iSCSI driver must include any iSCSI data segment with an iSCSI header segment in TCP transmit iSCSI 32IB. The iSCSI data segment must be in a transport block separate from the iSCSI header segment in the TCP transmit iSCSI 32IB.
The host memory address of each transfer block of TCP transmission iSCSI 32IB is 32 bits. If 64-bit addresses are required, TCP outbound iSCSI 64IB should be used.
The format of the TCP transmit iSCSI 32IB with three clearly illustrated transfer blocks is described below.
The CRC selection bit 0 and the CRC selection bit 1 are stored in separate words in the IB instead of the unused part of the first transfer block in order to process each transfer block in the same way.
TCP Receive iSCSI32IB is used to post a buffer to be used to receive on an iSCSI socket.
The length of the TCP receive iSCSI 32 IB is variable because this IB consists of an optional number of transfer blocks used for DMA transfers. The TCP reception iSCSI 32IB transfer block consists of a pair of address and transfer length.
The iSCSI CRC seed is used to seed the iSCSI CRC engine before the DMA and iSCSI CRC calculations begin. The iSCSI CRC seed must be set to 32'hffffffff if an iSCSI CRC is requested. If the iSCSI CRC is not requested, the seed value set is incorrect.
Description: If no iSCSI CRC calculation is required, then the iSCSI CRC seed does not need to be set to zero.
The address length of each transfer block in TCP receive iSCSI 32IB is 32 bits. For 64-bit addressing, TCP receive iSCSI 64IB must be used.
[Another iSCSI configuration] The above description details the operation of the IT10G, which uses the host computer to assemble the iSCSI PDU header. Another embodiment allows IT10G hardware and on-chip processors to assemble and process iSCSI PDUs. This other embodiment will be described below.
This section describes iSCSI hardware support using on-chip processors. From a high level, the hardware offloads the CRC and fixed interval markers (FIM) of the transmitted packets. In the received data path, the ability to DMA the PDU header to the on-chip processor or host computer and the data section of the PDU to the host is supported with CRC checking. This allows the iSCSI header and data to be separated and copied to the host computer's memory without the need for an additional memory copy on the host computer.
[Header Storage (HSU)] For iSCSI PDUs sent, the on-chip processor acts to build the packet header in its memory. The on-chip processor receives an iSCSI instruction block (IB) from the host computer to indicate the type of iSCSI PDU that is generated. If SCSI data associated with the PDU exists, the linked list of buffers is also given via the IB.
Once the on-chip processor generates a PDU header, it uses the HSU module to transfer the header from the on-chip processor memory to the MTX memory (transmission data buffer). The HSU also calculates the CRC of the header if needed.
If there is no SCSI data associated with the PDU, it will begin transferring data from the on-chip processor memory to the MTX memory as soon as the HSU module becomes available. If necessary, the CRC is also calculated and the FIM is inserted.
If there is SCSI forwarded in the PDU, then two modes of DMA are given: one-shot and linked list. If linked list mode is used, the HSU looks up the first linked list (LL) entry and requests a host DMA transfer. When the DMA engine indicates to the HSU that the first transfer has finished, the HSU begins transferring the header from the ARM memory to the MTX memory. The just DMA data is attached at the end of the header. Also during this time, the next LL entry is searched and a DMA is requested. Once the HSU is approved for host DMA forwarding, it locks access to the DMA engine (at least on the TX side) until all entries in the LL are serviced. This makes it easier for FIM and CRC logical devices to run in the data path.
[CRC calculation module] This module works to calculate the CRC values for both the PDU header and the data section. The CRC is then attached to the corresponding field. The data that feeds this module can come from on-chip processor memory (for header data) via HSU, or from DMA TX FIFO (for SCSI data). The output from the CRC calculator is given to the FIM module.
[FIM Insert Module] This module works to insert the FIM into the PDU. It receives the initial offset, interval count and PDU length from the HSU unit. Some repacking of the data may be required when inserting the FIM. Another function of this module is to determine the overall length of the PDU sent. This length is the programmed length of PDU + any CRC bytes + any FIM. This length is then sent to the XMTCTL module, which acts to split the PDU into properly sized TCP packets.
[XMTCTL Mux Module] This module works to multiplex the data that goes into the XMTCTL module. The data is direct from the DMA TX FIFO buffer or FIM insert module. With either path, the multiplexing logical unit indicates the full length of the packet and also provides data in a 128-bit wide format. The data is sent to the XMTCTL module using the dav / grant handshake.
[ISCSI reception support] [Overview] A block diagram of the iSCSI receive data path is shown in Figure 53.
[TCP RX module] This module 531 acts to parse received TCP packets. When an iSCSI packet arrives, it is stored in MRX memory in a 128-byte or 2K-byte buffer, so it is treated first like any other TCP data. The CB bit (WORD 0 × D, bit [30]) of the socket indicates whether this socket is owned by the host or on-chip processor. If the socket is owned by an on-chip processor, the data will not be automatically DMAd to the host. Instead, a status message is generated and sent to the on-chip processor over the normal network stack status message queue.
On-chip processor access to MRX memory When the on-chip processor is notified that the socket is receiving iSCSI data, it can read the data in the MRX via the CPU address and the data registers of the mallocrx register set. It is expected that the on-chip processor will mainly read the MRX memory 532 to determine the PDU type received from the header in it. If the on-chip processor decides to move the header to its memory, it uses the LDMA module. If you decide to DMA and PDU the data to the host, use HDMA module 533.
[RXI SCSI module] Use the RXI SCSI module when the on-chip processor wants to move data from MRX memory to its own local memory. The CRC of the data is selectively checked during this transfer. When the operation is complete, an interrupt or status message is generated.
If the data spans a large number of MRX buffers, the transfer must be split into two requests. An example of this is shown in FIG.
In this case, the first transfer is programmed with Add1 as the MRX source address and Length1 as the transfer length. The LAST_BLK bit is also not set for the first transfer. When the operation is complete, the partial CRC result is also returned in the status message. The on-chip processor then programs the local DMA (LDAM) with Add2 and Length2 and sets the LAST__BLK bit to end the header transfer. The CRC seed does not need to be programmed if the operation is continuous. Otherwise, the partial result of the CRC returned in the status message should not be programmed as a CRC seed for the second transfer. If only CRC bytes remain in the second buffer, a zero length transfer should be used in the CRC check enabled.
[HDMA started on-chip processor] When the on-chip processor wants to send the received SCSI data to the host, it programs the host DMA (HDMA) engine with the transfer length into both MRX and host memory at the start address. Instead, the on-chip processor can locate in its memory where the linked list of the buffer is located.
The linked list can contain up to 255 entries. During the DMA transfer, the HDMA module can be optionally programmed to check the CRC value. The DMA transfer length must not include CRC bytes if this option is enabled. Also, an optional CRC start seed can be programmed when a CRC check is requested.
If SCSI data is split across a large number of MRX buffers, then DMA transfers to the host must be split into separate requests. This state is shown in Figure 55.
In this case, HDMA must first be programmed with Add1 as the starting MRX memory address and Length1 as the transfer length. The Last_Host_Blk bit must also not be set for the first transfer. A status message is generated when the DMA operation is complete. If a CRC check is also requested, the status message also returns a partial checksum value. The on-chip processor then programs the HDMA engine with Add2 and Length2 and sets the LAST_Host_Blk bit to end the data transfer. If these two transfers are consecutive, the CRC seed value does not need to be programmed. Otherwise, the partial CRC result returned in the status message is programmed as part of the second DMA transfer request. If only CRC bytes remain in the second buffer, a transfer of length = 0 and CRC_En = 1 must be used.
[Release MRX buffer of on-chip processor] In a socket owned by the on-chip processor, the on-chip processor operates to release an MRX buffer that is no longer in use to the MRX buffer deallocation device. This is done by writing the base address for the block to be released to the MRX_128_Block_Add or MRX_2K_Block_Add registers and then issuing a release command to the corresponding MRX_Block_Command register.
IPSec Support Architecture [Overview] The following description details the hardware support for IPSEC. This configuration assumes separate modules to handle the computer characteristics of protocol encryption, decryption, and authentication functions, all of which are well known and understood. The configuration also assumes that any well-known, well-understood and used key exchange protocol, such as IKE, will be processed as an application on the host computer. Of course, it is also possible to integrate the key exchange function.
This description splits IPSEC support into transmit and receive sections. This is done because the two modules operate independently of each other, in addition to sharing a common security-related (SA) block.
IPSec characteristics: Anti-replay support for each SA, · Zero, DES, 3DES algorithms and AES 128-bit algorithms in Cryptographic Block Chaining (CBC) mode, Zero, SHA-1, MD-5 authentication algorithm, -Variable length encryption key up to 192 bits, -Variable length authentication key up to 160 bits, Jumbo frame support, -Automatic processing of SA expiration based on time and total data transferred, IPsec policy enforcement, -IPsec exception handling including exception packet generation and status notification.
IPsec protocol and mode: Transfer AH, Transfer ESP, Transfer ESP + AH, Tunnel AH, Tunnel ESP, Tunnel ESP + AH, Transfer AH + Tunnel AH, Transfer AH + Tunnel ESP, Transfer AH + Tunnel ESP + AH, Transfer ESP + tunnel AH, Transfer ESP + Tunnel ESP, Transfer ESP + Tunnel ESP + AH, Transfer ESP + AH + Tunnel AH, Transfer ESP + AH + Tunnel ESP, -Transfer ESP + AH + Tunnel ESP + AH.
[Security Related Block Format (SA)] A dedicated memory structure is used to store information for each IPSEC connection. There are separate blocks for the AH and ESP protocols, and for both RX and TX SA (RX SA cover data is received and TX SA cover data is transmitted). Therefore, socket connections that use both AH and ESP for both sending and receiving data require a large number of SA blocks. AH requires 1 SA block and ESP requires 2 blocks (ESP-1 and ESP-2). The total number of secret-protected socket connections follows the total amount of memory given to the SA block.
The Tx data path can support tunnel and transfer modes with a single transfer. In the worst case scenario, if both AH and ESP are used in both tunnel and forwarding modes, 6 SA blocks can be concatenated together. For data transmitted, the socket CB contains a pointer to the tunnel / transfer TX AH or tunnel / transfer TX ESP-1 SA. When both protocols are used in the transmitted data, the CB contains a link to the TX AH SA, which contains a link to the TX ESP-1 SA.
The Rx data path does not support AH and ESP decoding on a single transfer. The Rx data path does not support tunnel and transfer modes with a single transfer. The encrypted packet is iteratively decrypted. In the received data, RX_SA_LUT contains a pointer to RX AH or TX ESP-1 SA.
Figure 56 shows the SA block flow.
[Generate client socket] When an application needs to generate an IPSEC-protected client socket, the following sequence must be performed. -Write all applicable SA parameters to the IPSEC SA specific register. This requires the generation of a large number of write commands. The SA handle is inserted into the RX SA LUT. -IPSEC returns the new SA handle. -Write this SA handle to the socket specific register. -Write all applicable socket parameters including the IPSEC bit setting of the socket structure 2 register to the CP socket specific register. -Generate the commit_socket command.
The link to TX SA is stored in the open CB. The socket is now ready for use.
[SA generation of server socket] When an application needs to create an IPSEC-protected server socket, the following sequence must be performed. -Write all applicable SA parameters to the IPSEC SA specific register. This requires the generation of a large number of write commands. The SA handle is inserted into the RX SA LUT. -IPSEC returns the new SA handle. -Write this SA handle to the socket specific register. -Write all applicable socket parameters including the IPSEC bit setting of the socket structure 2 register to the CP socket specific register. -Generate the commit_socket command.
The commit socket command generates a HO CB in addition to generating a server port information table entry. The HO CB has an IPSEC bit set in it. This bit prevents the HO CB from being reused by another incoming SYN that accidentally decrypts to the same HASH. The TX SA handle is stored in the HO CB. The socket is now ready for use.
Note: In this case, as mentioned above, the HO CB will not be overwritten when a SYN with the same HASH is received. The HO CB is replicated when the socket is transferred to the configured state and an open CB is created. However, if the socket cannot reach the configured state, the host behaves to manually replicate the HO CB. This is done by entering the HO CB handle into the HO_CB handle register and then issuing the Deprecate_HOCB command. Hosts in this state must replicate the associated SA block.
[SA connection] The SA block has a link valid bit and a link field, which points to another SA. The CPU can use this field to link a large number of SAs if desired. The CPU must execute the following sequence to concatenate the SA blocks. -The SA_Link register is composed of the concatenated SA blocks. -Set the Link_Val bit in the CFG1 register. -The SA_Handle register is composed of the concatenated SA blocks. -Generate the "Update SA" command.
[SA Deplication / Unassignment] The Tx or Rx SA block is not automatically replicated when the CB is replicated. The CPU issues a "SA disable" command to keep track of unused SA blocks and to replicate / deallocate these SA blocks.
When the SA expires, the HW will generate an interrupt if it is not masked. Expired SA blocks are not deallocated from memory. This event also raises a status message containing the SA handle of the expired SA. CPU This handle can be used to update this SA, or the CPU can issue a "SA disable" command to replicate / deallocate that SA block.
[TX AH Transfer SA Block Format] Figure 57 shows the TX AH transfer SA block format.
[TX ESP-1 Transfer SA Block Format] Figure 58 shows the TX ESP-1 transfer SA block format.
[TX ESP-2 Transfer SA Block Format] Figure 59 shows the TX ESP-2 transfer SA block format.
[TX AH Tunnel SA Block Format] Figure 60 shows the TX AH tunnel SA block format.
[TX ESP-1 Tunnel SA Block Format] Figure 61 shows the TX ESP-1 tunnel SA block format.
[TX ESP-2 Tunnel SA Block Format] Figure 62 shows the TX ESP-2 tunnel SA block format.
[RX AH SA Block Format] Figure 63 shows the RX AH SA block format.
[RX ESP-1 SA block format] Figure 64 shows the RX ESP-1 SA block format.
[RX ESP-2 SA block format] Figure 65 shows the RX ESP-2 SA block format.
[Definition of security-related block fields] [SA type] These bits are used to identify the type of SA block that the block represents and are given for diagnostic support. Decryption is shown in the table below.<tables num="26"><img file="JP4875126B2_D0026.tif" /></tables>
All other decryptions not shown in the table above are reserved for future use.
[SA version] These bits identify the version number of SA and are given for diagnostic purposes. The current version of all SA block types is 0x1.
[XV, RV (Valid Send / Receive SA)] These bits indicate that the SA block is valid and can be used.
[XA, RA (Available Send / Receive Authentication)] These bits indicate that authentication is available for this protocol and should be used in packets in this socket. The corresponding authentication algorithm and key field are also valid. For TX / RX AH SA, this bit should always be set. For TX / RX ESP SA, this is optional.
[XA_ALG, RA_ALG (Send / Receive Authentication Algorithm)] These bits indicate the algorithm used for authentication. The possible options are listed in the table below.<tables num="27"><img file="JP4875126B2_D0027.tif" /></tables>
All decryptions not shown are pending for future use.
[XE, RE (Send / Receive Encryption Enable)] These bits indicate that encryption should be used for packets in this socket and that the encryption key and encryption algorithm fields are valid. This bit is specified only for TX / RX ESP SA.
[XE_ALG, RE_ALG (send / receive algorithm)] These bits indicate the algorithm used for encryption. The possible options are listed in the table below.<tables num="28"><img file="JP4875126B2_D0028.tif" /></tables>
All other decryptions not shown are pending for future use.
[RAR (Receive Anti-Replay Enable)] This bit indicates that the anti-replay algorithm should be used in the received packet.
[XTV, RTV (send / receive time stamp valid)] These bits indicate that the timestamp field is valid in SA and that the SA block should be considered out of date when the timestamp expires.
[XBV, RBV (send / receive byte count enabled)] These bits indicate that the byte count limited field is valid in SA and that the SA block should be considered out of date when the byte count reaches this value.
[XSV, RSV (valid for transmit / receive sequence only)] These bits indicate that the SA block should be considered old when the sequence number field wraps 0xFFFFFFFF.
Note: If all xTV, xBV or xSV bits are not set, the SA block is considered persistent and never expires.
[RDC (Receive Destination IP Check Enable)] This bit indicates that the destination IP address of any packet must match the destination IP address field of this SA.
[RSC (Receive SPI Check Enable)] This bit indicates that the SPI field in the IPSEC header must match the destination IP address field for this SA.
[RTR (Limited Watermark of Reached / Disabled Receiving Timestamps)] If this bit is not set, a time stamp watermark check can be done. When this raises a status message, this bit is set to prevent further watermark status messages from being sent.
[LV (link enabled)] This bit indicates that the SA block is linked to this SA. When this bit is valid, the link field contains the SA handle of the next SA block.
[SPI number] This is the protocol-related SPI. In TX SA, this SPI is included in the protocol header. In RX SA, this SPI is compared against the corresponding field in the received data packet.
[Sequence number] This is the sequence number associated with the protocol. The sequence numbers are different between the two protocols because AH and ESP SA expire to different vice ministers. In TX SA, this sequence number is reset to 0x00000000 when the SA is generated and incremented by 1 for each packet sent using the SA block. For received packets, this field represents the last sequence number received on the socket and is used to check in packet replays.
[Authentication key] This parameter is the key used to authenticate the protocol. If the key is less than 160 bits, it should be lsb correct.
[Encryption key] This parameter is the key used to encrypt IPSEC packets and is specified only for TX / RX ESP SAs. For algorithms that use bits smaller than 192 bits, the key should be lsb correct.
[Time stamp] This is a future timestamp when the SA block is considered out of date. Initialized when SA is generated. When a packet is sent or received, the current free-running millisecond timestamp is compared to this time. If the time matches or exceeds this time stamp, the SA block is considered out of date.
[Byte count] This parameter is set when the SA is initialized. Determine the maximum number of bytes that can be sent using the SA block. HW uses this as the initial number for the decrement counter. The HW decrements this byte count when a packet is sent or received on this SA. When the byte count reaches zero, the SA is considered out of date. If this limit is used, the xBV bit should be set.
[Link] This field points to the next SA block associated with the socket.
[IPsec module] IPsec is divided into the following four main modules. IPSECX, IPSECR, IPSECREGS, -IPSEC memory.
IPSECX contains IPSEC transmit datapath logic. Data and control packets are first forwarded to internal IPSECX memory. The encryption engine reads these packets, encrypts them, and writes these packets back to IPSECX memory. These packets are then forwarded to an Ethernet® transmitter for transmission.
IPSECR contains IPSEC receive data path logic. Any incoming packet is parsed for IPSEC type packets. If the packet is an IPSEC packet, it will be forwarded to the internal IPSECR memory. The decryption engine decrypts these packets and writes them back to IPSECR memory. This packet is read, injected into the network stack and returned. At this time, since this packet is decrypted, the parser does not identify it as an IPSEC packet, which is identified as a normal TCP / IP packet.
IPSECREGS contains programmable registers for the CPU. By programming these registers, the CPU creates, updates or erases SAs. This module provides indirect memory access to internal memory for diagnostic purposes.
IPSEC uses two SRAMs as shown in the figure below. Both memories can hold 9KB size packets. The IPSEC memory size is shown below. -IPSECX memory: 592 x 32 bits x 4 dual ports (9472 bytes), -IPSECR memory: 1168 x 32 bits x 2 dual ports (9344 bytes).
FIG. 66 is a block diagram showing the overall flow of the IPSEC logical device.
[IPSEC Outgoing Data Path (IPSECX)] [Overview] The following description is the details of the outbound IPSEC data path. The IPSEC encryption / authentication engine operates on packets that occur before it is called. A block diagram outlining the data flow is shown in Figure 67.
[Data flow of packets sent by IPSEC] [TXDATCB / TCPACK] Data transmitted to an IPSEC-protected socket is first stored in the MTX DRAM 671 in the same way as a regular socket. The TXDATCB / TCPACK module 672 operates to form Ethernet, IP, TCP, and any IPSEC headers for packets. This module reads information from the appropriate SA block to determine which header and the required header format. Each socket has a CB structure associated with it. The CB contains old information for each socket. Within the CB structure, there is a pointer to the SA block used. For sockets that require a large number of SA blocks, the CB contains pointers to all AH and ESP1 SA blocks.
The TXDATCB / TCPACK module writes the packet header to a buffer in the MTX DRAM. When the TXDATCB / TCPACK module finds that a packet requires IPSEC processing, it notifies IPSEC XIF module 673 that the packet is ready to be processed. If the TXDATCB / TCPACK module notices that the IPSECXIF block is full at the start of the header generation process, it returns the packet to the send queue for later processing.
[TCPACK] In addition to data packets, TCP utility packets (ACK, SYN, FIN) can also be encrypted. These packets come from the TCPACK module. Although not shown in the aforementioned data path in Figure 67, this module also accessed CB and SA memory.
Note: RST packets generated because local sockets are not available are not IPSEC protected. The only protected RST is a packet that is used to terminate censoring or is received in response to a SYN in response to an illegal SYN / ACK. Both of these RST packets are formed by the TCPACK module rather than the normal RST packet queue.
[IPSECXIF] This module operates to forward packets from MTX DRAM or TCPACK (via TCPDATCB) to IPSECX internal memory. This internal memory is organized as 576 x 128 bits. When a packet is obtained from TXDATCB or TCPACK, this module begins transferring data to IPSEC memory. It also gets it from the SA register to send the information applicable to the SA to the crypto engine. The read pointer is fed back from IPSECTX to prevent this module from overwriting the current packet in IPSECX memory.
[Cryptography / Authentication Engine] This module 674 works to add encryption and authentication to the packet. It takes the parameters provided by the IPSECXIF module and processes the packets stored in IPSECX memory 675. If both authentication and encryption are required on the packet, this module is assumed to be before both characteristics before forwarding the packet to the IPSECTX module. The encrypted data should be written to the same memory location as the source packet. When the process is finished, it informs the IPSECTX module that the completed packet is ready to be sent. The encryption engine runs globally in the dram_clk domain.
This module consists of two parallel identical cryptographic engines. In this case, they are serviced in alternating order. Also, when the encrypt_rdy indicator is sent to IPSECTX, the packets given must be in the same order they were given to the encryption engine.
[IPSECTX] This module 676 operates to retrieve processed packets from the crypto engine and schedule them for transmission. When the preparation completion instruction is received from the encryption engine, the start address and packet length information are registered. The transmit request is then transmitted to the Ethernet® transmit arbitrator. The Ethernet® transmit arbitrator reads the packet information directly from the IPSECX memory and sends it to the MAC's transmit buffer. This strobes ipsectx_done when the entire packet is read. Upon receiving this instruction, the IPSECTX module updates its ipsectx_rd_add. This bus is returned to IPSECXIF to indicate that more memory is freed.
[IPSECX Memory Arbitrator (IPSECX MEMARB)] This module 677 acts to arbitrate access to the IPSECX memory bank. This module runs in the dram_clk domain.
[IPSECX Memory Interface (IPSECX MEMIF)] This module 678 provides an interface to dual port RAM. This module connects SRAM with other logical devices that access RAM. This module needs to be changed when the RAM model changes. This module works globally in the dram_clk domain.
[IPSEC Received Data Path (IPSECR)] [Overview] The block diagram of FIG. 68 shows the data path flow of the received IPSEC packet. This forms the basis of the following explanation.
[Data flow of packets received by IPSEC] [IPIN] The received data in the IPSEC protected socket is first parsed by the IPIN module 687. Inside the outermost IP header is a protocol field that indicates the protocol type of the next header. If it detects that the protocol is 0x50 (ESP) or 0x51 (AH), it completes the packet containing the outermost IP header for the IPSECRIF module.
If the received packet is fragmented, it is treated like any other fragmented IP packet and sent to the exception processor. When the packet is complete, it is re-injected to the bottom of the IP stack via the IPINARB module 682.
[IPSECRIF] When the IPIN indicates to this module 683 that an IPSEC packet has been received, it begins to store the packet in IPSECR memory 684. The packet begins to be stored in the header, followed shortly after which the outermost IP header is stored. The IPSECRIF module also analyzes the SPI, source IP address, and AH / ESP settings to perform LUT discovery to discover the correct RX SA block. When acquiring the LUT value, read the SA block and check if its parameters match the parameters of the received packet. If they match, you will have the appropriate SA. If they do not match, the packet is dropped and the event is recorded.
When the exact SA block is found, the received SA parameters are read and stored. If the anti-replay property becomes available for this SA, the sequence number is checked to see if it is valid. If this is good, the SA parameter is forwarded to decryption module 685 along with the starting memory address in the packet's IPSECR memory. If the sequence number is not valid, the packet is dropped and the event is recorded. However, the sequence number and sequence bitmap are not updated at this time. In this respect, the purpose of the anti-replay check is to prevent bad packets from being unnecessarily decrypted and authenticated. Anti-replay updates are processed in IPSECRX module 686 after the packet is authenticated.
[Encryption / Authentication] This module 685 works to decrypt and authenticate IPSEC packets. When a packet is available, the IPSECRIF module indicates this to the decoding engine by claiming the ipsecrif_rdy signal. The IPSECRIF module also provides the starting address of the packet. If the packet fails to authenticate, it is discarded and the event is recorded. If the authentication passes, the packet is decrypted if necessary. The decrypted packet is written back to the same IPSECR memory location as the original packet. The decryption engine then indicates to the IPSECRX module that the packet is ready to be received by claiming decrypt_rdy. If the packet requires both authentication and decryption, it is assumed that both functions are completed before processing the packet off to the IPSECRX module.
This module consists of two parallel identical decoding engines. In this case, they are serviced in an alternating order. Also, when the decrypt_rdy indicator is sent to IPSECRX, the packets given must be in the same order as they were given to the decryption engine.
Note: The decryption / authentication engine operates off the DRAM clock.
[IPSECRX] This module 686 works to take a decrypted / authenticated packet and re-inject it onto the stack. In tunnel mode packets, processed packets can be injected and returned directly via IPINARB. For forwarded mode packets, an IP header is generated and prepended at the beginning of the packet. By doing this, it is possible to reuse the TCP checksum and interface logic inside the IPIN for all packets.
This module also acts to update the receive SA block of the packet. The new sequence number and bitmap for the SA is transferred from the decryption engine to this module along with the SA handle. This module updates the received timestamp and byte count. If it finds that any parameters have reached their limits, the SA is disabled and a status message is sent to the exception processor.
[IPSECR Memory Arbitrator (IPSECRMRMARB)] This module 687 acts to arbitrate access to the IPSECR memory bank. This module runs in the dram_clk domain.
[IPSECR Memory Interface (IPSECR MEMIF)] This module provides an interface to dual port RAM. This module connects SRAM with other logical devices that access RAM. This module needs to be changed when the RAM model changes. This module works globally in the dram_clk domain.
[IPSEC LUT] This LUT is organized as 128K with 17 bits. Each word consists of a 16-bit SA handle and a 1-bit valid indicator. Physically the LUT is contained within the MTX DRAM bank.
[IPARB] This module arbitrates traffic coming from the Ethernet (R) input parser and IPSE CRX.
[IPSEC Anti-Replay Algorithm] This is the algorithm used to check for anti-replay checks. This algorithm uses a 32-bit window to check the last 32 sequence numbers. The bitmap msb represents the oldest sequence number and the lsb represents the current sequence number. This algorithm is used for both the AH and ESP protocols.
LAST_SEQ: Last received sequence number. This is stored in the SA block. SEQ: Sequence number of the received packet. BITMAP: A 32-bit bitmap that represents the last 32 sequential sequence numbers.
FIG. 69 is a flow diagram showing the IPSEC anti-replay algorithm.
[IPSEC Register (IPSECREGS)] [Overview] This module provides an interface to the CPU and contains programmable registers. This module also generates, updates, or disables SAs when the CPU issues the appropriate commands. It provides indirect access to the following memory for diagnostic purposes: IPSECX, IPSECR, RX SA LUT, SA block.
[IPSEC SA Status Message]] For both inbound and outbound SA blocks, a status message is issued to the on-chip processor whenever the block becomes invalid. Blocks can become invalid when the sequence number reaches 0xFFFF, or when the timestamp or byte count reaches those limits.
[SA block and data flow for DRAM interface] [Overview] The following description is an overview of the interface for SA block access via the NS DDR arbitration module. It shows the data flow, lists the interface signals, and details the required timing.
[data flow] Three different access types are supported in SA blocks. These are burst write, burst read, and single read. The ipsecsarb module arbitrates requests from different sources to access SA memory.
[SA LUT] [Overview] The IPSEC receive logic device uses a LUT to find the appropriate SA block. The LUT is 16 bits x 32K deep. Each LUT block contains an SA block handle. If the SA handle has a value of 0, this is considered invalid.
[SA LUT to DRAM interface] The following description outlines the interface between the SA LUT and the data DDR arbitration module. It shows the data flow, lists the interface signals, and details the required timing.
[data flow] The SA LUT has only a single read and write access to the DRAM. Its access is always in WORD. Since each LUT block has only 16 bits, only the bottom 16 bits of the WORD are transferred to the SA LUT memory interface. Reads and writes are not done in the same LUT access cycle.
Although the present invention has been described herein with reference to preferred embodiments, one of ordinary skill in the art will readily recognize that the applications described herein can be replaced by other applications without departing from the technical scope of the invention. Will. Therefore, the present invention should be limited only by the claims.
2 sheets
Sheet 1 Sheet 2
Every citation, both ways
| Document | Relation | Office |
|---|---|---|
| US20030165160A1 | Cites | United States of America |
| US20030061505A1 | Cites | United States of America |
40 members in 9 offices
Priority claims10
| Document | Office | Kind | Date |
|---|---|---|---|
| 38692402 | United States of America | P | |
| 38692402 | United States of America | P | |
| 60386924 | United States of America | – | |
| 10456871 | United States of America | – | |
| 45687103 | United States of America | A | |
| 45687103 | United States of America | A | |
| 2002386924 | – | – | – |
| 2003456871 | – | – | – |
| US20020386924P | – | – | – |
| US20030456871 | – | – | – |
Members40
| Document | Office | Kind | |
|---|---|---|---|
| CA2265692A1 | Canada | A1 | |
| WO9819412A1 | World Intellectual Property Organization (WIPO) | A1 | |
| AU4595297A | Australia | A | |
| EP0935855A1 | European Patent Office (EPO) | A1 | |
| CN1237295A | China | A | |
| US6034963A | United States of America | A | |
| EP0935855A4 | European Patent Office (EPO) | A4 | |
| AU723724B2 | Australia | B2 | |
| JP2001503577A | Japan | A | |
| CA2265692C | Canada | C | |
| WO02086674A2 | World Intellectual Property Organization (WIPO) | A2 | |
| AU2002258974A1 | Australia | A1 | |
| US2003165160A1 | United States of America | A1 | |
| WO02086674A3 | World Intellectual Property Organization (WIPO) | A3 | |
| WO03105011A1 | World Intellectual Property Organization (WIPO) | A1 | |
| AU2003251419A1 | Australia | A1 | |
| EP1382145A2 | European Patent Office (EPO) | A2 | |
| US2004062267A1 | United States of America | A1 | |
| CN1154268C | China | C | |
| JP2005502225A | Japan | A | |
| EP1525535A1 | European Patent Office (EPO) | A1 | |
| JP2005529523A | Japan | A | |
| USRE39501E | United States of America | E | |
| JP2007133902A | Japan | A | |
| JP3938599B2 | Japan | B2 | |
| US2007253430A1 | United States of America | A1 | |
| EP1382145A4 | European Patent Office (EPO) | A4 | |
| EP1525535A4 | European Patent Office (EPO) | A4 | |
| JP2008259238A | Japan | A | |
| EP0935855B1 | European Patent Office (EPO) | B1 | |
| DE69739159D1 | Germany | D1 | |
| US7535913B2 | United States of America | B2 | |
| JP2010063110A | Japan | A | |
| EP1382145B1 | European Patent Office (EPO) | B1 | |
| AT493821T | Austria | T | |
| ATE493821T1 | Austria | T1 | |
| DE60238751D1 | Germany | D1 | |
| JP4875126B2This record | Japan | B2 | |
| JP4916482B2 | Japan | B2 | |
| US8218555B2 | United States of America | B2 |
17 legal events, as the office reported them to INPADOC
Over the term
Point at a mark for the eventEvents
| Event | Code | |
|---|---|---|
| Cancellation because of completion of termEXPY | EXPY | |
| Receipt of annual feesJAPANESE INTERMEDIATE CODE: R250R250 | R250 | |
| Receipt of annual feesJAPANESE INTERMEDIATE CODE: R250R250 | R250 | |
| Receipt of annual feesJAPANESE INTERMEDIATE CODE: R250R250 | R250 | |
| Receipt of annual feesJAPANESE INTERMEDIATE CODE: R250R250 | R250 | |
| Receipt of annual feesJAPANESE INTERMEDIATE CODE: R250R250 | R250 | |
| Receipt of annual feesJAPANESE INTERMEDIATE CODE: R250R250 | R250 | |
| Receipt of annual feesJAPANESE INTERMEDIATE CODE: R250R250 | R250 | |
| Receipt of annual feesJAPANESE INTERMEDIATE CODE: R250R250 | R250 | |
| Receipt of annual feesJAPANESE INTERMEDIATE CODE: R250R250 | R250 | |
| Renewal fee payment (event date is renewal date of database)FPAY | FPAY | |
| Certificate of patent or registration of utility modelJAPANESE INTERMEDIATE CODE: R150R150 | R150 | |
| Certificate of patent or registration of utility modelJAPANESE INTERMEDIATE CODE: R150R150 | R150 | |
| First payment of annual fees (during grant procedure)JAPANESE INTERMEDIATE CODE: A61A61 | A61 | |
| Written decision to grant a patent or to grant a registration (utility model)JAPANESE INTERMEDIATE CODE: A01A01 | A01 | |
| Written decision to grant a patent or to grant a registration (utility model)JAPANESE INTERMEDIATE CODE: A01A01 | A01 | |
| Decision of grant or rejection writtenTRDD | TRDD |
Numbers
- Publication
- 4875126
- Publication, DOCDB
- 4875126
- Publication, EPODOC
- JP4875126B
- Application
- 217450
- Application, DOCDB
- 2009217450
- Application, EPODOC
- JP20090217450
Titles2
- Japanese
- ISCSIおよびIPSECプロトコルをサポートするギガビットイーサネットアダプタ
- English
- Gigabit Ethernet adapter that supports ISCSI and IPSEC protocols
Classification
- IPC, 1
- H04L12 28
