Network congestion management using aggressive timers
Claim Score by NHIP
Abstract
A network system includes links and end stations coupled between the links. Types of end stations include endnodes which originate or consume frames and routing devices which route frames between the links. At least one end station includes an aggressive timer adapted to respond to an occurrence of at least one condition of delayed frame transmission progress, to provide a frame delay indication when the at least one condition exists for a duration that exceeds a variable timing threshold. The variable timing threshold is configurable based on at least one network system attribute.

Term
Term ended
Projected expiry passed 23 May 2020, 6.3 years ago.
- Priority
- Filed
- Published
- Projected expiry
- Today
72 claims: 4 independent, 68 dependent
- 1A network system comprising:links;end stations coupled between the links, wherein types of end stations include endnodes which originate or consume frames and routing devices which route frames between the links, wherein at least one end station includes: an aggressive timer adapted to respond to an occurrence of at least one condition of delayed frame transmission progress to provide a first frame delay indication when the at least one condition exists for a duration that exceeds a variable timing threshold, wherein the variable timing threshold is configurable based on at least one network system attribute.
- 40A network system comprising:links;end stations coupled between the links, wherein types of end stations include endnodes which originate or consume frames and routing devices which route frames between the links, wherein at least one end station includes: a forward progress timer adapted to provide a first frame delay indication when at least one delay in frame transmission exceeds a first timing threshold;and an aggressive timer adapted to respond to an occurrence of at least one condition of delayed frame transmission progress, to provide a second frame delay indication when the at least one condition exists for a duration that exceeds a second timing threshold, wherein the second timing threshold is configurable based on at least one network system attribute.
- 51Broadest claimClaim Score 75, broad(NHIP)An end station comprising:an aggressive timer adapted to monitor network traffic, and respond to an occurrence of at least one condition of delayed frame transmission progress to provide a first frame delay indication when the at least one condition exists for a duration that exceeds a variable timing threshold, wherein the variable timing threshold is configurable based on at least one network system attribute.
- 64A method of detecting potential network system bandwidth underutilization, the method comprising:configuring a timing threshold that defines when a continuous existence of at least one condition of delayed frame transmission progress constitutes an underutilization of network system bandwidth, wherein the configuring is based on at least one network system attribute;and monitoring transmission progress of a frame, wherein the monitoring includes: responding to any indicated existence of the at least one condition of delayed frame transmission progress;timing the duration of the existence of the at least one condition;and comparing the measured duration against the configured timing threshold.
Independent claims4
193 paragraphs in 6 sections, as filed
CROSS REFERENCE TO RELATED APPLICATIONS
P-0001[0001] This patent application is a Continuation-in-part of U.S. patent application Ser. No. 09/578,019, entitled “RELIABLE MULTICAST,” filed May 24, 2000, and having Attorney Docket No. HP PDNO 10991834-2, which is herein incorporated by reference. U.S. patent application Ser. No. 09/578,019 is a Continuation-in-Part Application of U.S. patent application Ser. No. 09/578,155, filed May 23, 2000, entitled “RELIABLE DATAGRAM” having Attorney Docket No. HP PDNO 10991833-1 which is herein incorporated by reference. U.S. patent application Ser. No. 09/578,019 also claimed the benefit of the filing date of U.S. Provisional Patent Applications Serial No. 60/135,664, filed May 24, 1999 and having Attorney Docket No. HP PDNO 10991654-1; and Ser. No. 60/154,150, filed Sep. 15, 1999 and having Attorney Docket No. HP PDNO 10992562-1, both of which are herein incorporated by reference.
THE FIELD OF THE INVENTION
P-0002[0002] The present invention generally relates to communication in network systems and more particularly to network system congestion management.
BACKGROUND OF THE INVENTION
P-0003[0003] A traditional network system, such as a computer system, has an implicit ability to communicate between its own local processors and from the local processors to its own I/O adapters and the devices attached to its I/O adapters. Traditionally, processors communicate with other processors, memory, and other devices via processor-memory buses. I/O adapters communicate via buses attached to processor-memory buses. The processors and I/O adapters on a first computer system are typically not directly accessible to other processors and I/O adapters located on a second computer system.
P-0004[0004] In conventional distributed computer systems, distributed processes, which are on different nodes in the distributed computer system, typically employ transport services to communicate. A source process on a first node communicates messages to a destination process on a second node via a transport service. A message is herein defined to be an application-defined unit of data exchange, which is a primitive unit of communication between cooperating sequential processes. Messages are typically packetized into frames for communication on an underlying communication services/fabrics. A frame is herein defined to be one unit of data encapsulated by a physical network protocol header and/or trailer.
P-0005[0005] Messages communicated over the underlying communication services/fabrics can often experience congestion for various reasons, such as head of line blocking. There are conventional congestion control mechanisms. Congestion control mechanisms typically fall into three categories which include congestion detection mechanisms; congestion reporting mechanisms; and congestion response mechanisms. Congestion reporting mechanisms report the occurrence of congestion provided from congestion detection mechanisms possibly for short term use in alleviating congestion and possibly for long term network management. The congestion response mechanisms attempt to alleviate or remove congestion. Congestion in large distributed computer systems is a significant problem today, especially in infrastructures of remote computer systems having congestion resulting from message traffic over an internet or intranet coupling the remote computer systems.
P-0006[0006] Certain conventional distributed computer systems employ various means for addressing network traffic congestion. One known method is to drop frames that fail to make forward progress over a defined period of time. One realization of the method involves the use of forward progress timers implemented to prevent deadlock, or abnormal congestion, of a transmission port. Forward progress timers are typically implemented to expire 100-500 milliseconds after their initiation. Thus, relative to network transmission rates, a forward progress timer is quite coarse, as its purpose is to relieve severe deadlock.
P-0007[0007] On expiration of the forward progress timer, one or more frames are typically purged from the port. For example, the current stalled frame can be dropped from the queue. Alternatively, all packets in a queue targeting the congested port can be dropped. Typically, the sender of a dropped frame will become aware of the transmission failure via a protocol layer, such as by nonreceipt of an acknowledgment frame (ACK), and will respond to the inferred congestion assumed to be the cause of the transmission failure. Example sender responses include re-sending of the frame, and reducing of the transmission rate.
P-0008[0008] In practice, a plurality of senders are often targeting a congestion-prone port. As a result, frame dropping at times of congestion can lead to a problem of synchronization, where sources simultaneously respond to perceived congestion by reducing their respective transmission rates. As a result, the previously-congested port becomes under-utilized for some time. Hence, a cycle of congestion and under-utilization can develop, resulting in reduced overall throughput and efficiency.
P-0009[0009] One solution to the problem of synchronization is the use of random frame dropping, whereby only randomly-selected frames are dropped upon the expiration of a forward progress timer. Senders of the randomly-dropped frames perceive the failure of transmission as congestion, and respond to the perceived congestion. Consequently, the random frame dropping can dampen the aggregate global response to a port's congestion, resulting in greater port utilization, throughput, and efficiency.
P-0010[0010] A further improvement involves weighted random algorithms, which treat frames of a higher priority more favorably while employing randomization to avoid synchronization. Random and weighted-random-based techniques have been implemented in early detection schemes, such as disclosed in <i>Weighted Random Early Detection on the Cisco </i>12000 <i>Series Router, </i>available at http://www.cisco.com/univercd/cc/td/doc/product/software/ios112/ios112p/gsr/w red_gs.pdf. Early detection is premised on anticipating congestion and responding before it becomes severe. To address the problem of a forward progress timer's coarse granularity, some anticipatory methods use a queue size measuring technique for estimating a congestion level. One drawback of such queue size-based approaches to congestion management, is their inability to provide a time-based service guarantee for high priority frames that require a minimum-delay service priority, such as multimedia or real-time applications.
P-0011[0011] Network utilization and associated congestion continues to grow as a result of the increasing scale of distributed computer systems, variety of network elements, and increasingly-complex applications. For example, the expanding use of multimedia and real-time applications, and their associated bandwidth requirements, is fueling the need for specialized congestion management techniques and policies.
P-0012[0012] For reasons stated above and for other reasons presented in the Description of the Preferred Embodiments section of the present specification, there is a need for a congestion management technique that improves network utilization, throughput, and efficiency.
SUMMARY OF THE INVENTION
P-0013[0013] One aspect of the present invention provides a network system having links and end stations coupled between the links. Types of end stations include endnodes which originate or consume frames and routing devices which route frames between the links. At least one end station includes an aggressive timer adapted to respond to an occurrence of at least one condition of delayed frame transmission progress, to provide a frame delay indication when the at least one condition exists for a duration that exceeds a variable timing threshold. The variable timing threshold is configurable based on at least one network system attribute.
BRIEF DESCRIPTION OF THE DRAWINGS
P-0014[0014]FIG. 1 is a diagram of a distributed computer system.
P-0015[0015]FIG. 2 is a diagram of an example host processor node for the computer system of FIG. 1.
P-0016[0016]FIG. 3 is a diagram of a portion of a distributed computer system employing a reliable connection service to communicate between distributed processes.
P-0017[0017]FIG. 4 is a diagram of a portion of distributed computer system employing a reliable datagram service to communicate between distributed processes.
P-0018[0018]FIG. 5 is a diagram of an example host processor node for operation in a distributed computer system.
P-0019[0019]FIG. 6 is a diagram of a portion of a distributed computer system illustrating subnets in the distributed computer system.
P-0020[0020]FIG. 7 is a diagram of a switch for use in a distributed computer system.
P-0021[0021]FIG. 8 is a diagram of a portion of a distributed computer system.
P-0022[0022]FIG. 9A is a diagram of a work queue element (WQE) for operation in the distributed computer system of FIG. 8.
P-0023[0023]FIG. 9B is a diagram of the packetization process of a message created by the WQE of FIG. 9A into frames and flits.
P-0024[0024]FIG. 10A is a diagram of a message being transmitted with a reliable transport service illustrating frame transactions.
P-0025[0025]FIG. 10B is a diagram illustrating a reliable transport service illustrating flit transactions associated with the frame transactions of FIG. 10A.
P-0026[0026]FIG. 11 is a diagram of a layered architecture.
P-0027[0027]FIG. 12 is a diagram illustrating a method for recognizing and managing network system congestion using congestion detection, reporting, and responding mechanisms.
P-0028[0028]FIG. 13A is a simplified diagram illustrating one embodiment of a network system that includes an end station comprising an aggressive timer according to one embodiment of the present invention.
P-0029[0029]FIG. 13B is a diagram illustrating an exemplary end station of FIG. 13A that includes an aggressive timer adapted to monitor the end station's port according to one embodiment of the present invention.
P-0030[0030]FIG. 13C is a diagram of an exemplary end station of FIG. 13A that runs an application program according to one embodiment of the present invention.
P-0031[0031]FIG. 13D is a diagram of an exemplary end station of FIG. 13A that includes a traffic congestion manager according to the present invention.
P-0032[0032]FIG. 14 is a diagram of an example of one embodiment of a routing element having a network traffic congestion manager utilizing aggressive timers.
P-0033[0033]FIG. 15 is a flow diagram illustrating an aggressive timer's and traffic congestion manager's operation in the exemplary routing element of FIG. 14.
P-0034[0034]FIG. 16 is a diagram of one embodiment of an end station utilizing an aggressive timer and a forward progress timer.
DESCRIPTIONS OF THE PREFERRED EMBODIMENTS
P-0035[0035] In the following detailed description of the preferred embodiments, reference is made to the accompanying drawings which form a part hereof, and in which is shown by way of illustration specific embodiments in which the invention may be practiced. It is to be understood that other embodiments may be utilized and structural or logical changes may be made without departing from the scope of the present invention. The following detailed description, therefore, is not to be taken in a limiting sense, and the scope of the present invention is defined by the appended claims.
P-0036[0036] One aspect of the present invention is directed to a method and apparatus providing an aggressive timer-based congestion manager. This aspect facilitates the application of various criteria for selective treatment of frames based on frame or network system attributes as part of congestion detection, reporting, and/or response mechanisms. In one embodiment, the congestion manager supports the enforcement of policies set by network fabric management. In one embodiment, the congestion manager is employed on a network entity, including a network fabric element or end station, such as a switch, router, or endnode.
P-0037[0037] An example embodiment of a distributed computer system is illustrated generally at <b>30</b> in FIG. 1. Distributed computer system <b>30</b> is provided merely for illustrative purposes, and the embodiments of the present invention described below can be implemented on network systems of numerous other types and configurations. For example, network systems implementing the present invention can range from a small server with one processor and a few input/output (I/O) adapters to massively parallel supercomputer systems with hundreds or thousands of processors and thousands of I/O adapters.
P-0038[0038] Furthermore, the present invention can be implemented in an infrastructure of remote computer systems connected by an internet or intranet.
P-0039[0039] Distributed computer system <b>30</b> includes a system area network (SAN) <b>32</b> which is a high-bandwidth, low-latency network interconnecting nodes within distributed computer system <b>30</b>. A node is herein defined to be any device attached to one or more links of a network and forming the origin and/or destination of messages within the network. In the example distributed computer system <b>30</b>, nodes include host processors <b>34</b><i>a</i>-<b>34</b><i>d; </i>redundant array independent disk (RAID) subsystem <b>33</b>; and I/O adapters <b>35</b><i>a </i>and <b>35</b><i>b. </i>The nodes illustrated in FIG. 1 are for illustrative purposes only, as SAN <b>32</b> can connect any number and any type of independent processor nodes, I/O adapter nodes, and I/O device nodes. Any one of the nodes can function as an endnode, which is herein defined to be a device that originates or finally consumes messages or frames in the distributed computer system.
P-0040[0040] A message is herein defined to be an application-defined unit of data exchange, which is a primitive unit of communication between cooperating sequential processes. A frame is herein defined to be one unit of data encapsulated by a physical network protocol header and/or trailer. The header generally provides control and routing information for directing the frame through SAN <b>32</b>. The trailer generally contains control and cyclic redundancy check (CRC) data for ensuring frames are not delivered with corrupted contents.
P-0041[0041] SAN <b>32</b> is the communications and management infrastructure supporting both I/O and interprocess communication (IPC) within distributed computer system <b>30</b>. SAN <b>32</b> includes a switched communications fabric (SAN FABRIC) allowing many devices to concurrently transfer data with high-bandwidth and low latency in a secure, remotely managed environment. Endnodes can communicate over multiple ports and utilize multiple paths <b>30</b> through the SAN fabric. The multiple ports and paths through SAN <b>32</b> can be employed for fault tolerance and increased bandwidth data transfers.
P-0042[0042] SAN <b>32</b> includes switches <b>36</b> and routers <b>38</b>. A switch is herein defined to be a device that connects multiple links <b>40</b> together and allows routing of frames from one link <b>40</b> to another link <b>40</b> within a subnet using a small header destination ID field. A router is herein defined to be a device that connects multiple links <b>40</b> together and is capable of routing frames from one link <b>40</b> in a first subnet to another link <b>40</b> in a second subnet using a large header destination address or source address.
P-0043[0043] In one embodiment, a link <b>40</b> is a full duplex channel between any two network fabric elements, such as endnodes, switches <b>36</b>, or routers <b>38</b>. Example suitable links <b>40</b> include, but are not limited to, copper cables, optical cables, and printed circuit copper traces on backplanes and printed circuit boards.
P-0044[0044] Endnodes, such as host processor endnodes <b>34</b> and I/O adapter endnodes <b>35</b>, generate request frames and return acknowledgment frames. By contrast, switches <b>36</b> and routers <b>38</b> do not generate and consume frames. Switches <b>36</b> and routers <b>38</b> simply pass frames along. In the case of switches <b>36</b>, the frames are passed along unmodified. For routers <b>38</b>, the network header is modified slightly when the frame is routed. Endnodes, switches <b>36</b>, and routers <b>38</b> are collectively referred to as end stations.
P-0045[0045] In distributed computer system <b>30</b>, host processor nodes <b>34</b><i>a</i>-<b>34</b><i>d </i>and RAID subsystem node <b>33</b> include at least one system area network interface controller (SANIC) <b>42</b>. In one embodiment, each SANIC <b>42</b> is an endpoint that implements the SAN <b>32</b> interface in sufficient detail to source or sink frames transmitted on the SAN fabric. The SANICs <b>42</b> provide an interface to the host processors and I/O devices. In one embodiment the SANIC is implemented in hardware. In this SANIC hardware implementation, the SANIC hardware offloads much of CPU and I/O adapter communication overhead. This hardware implementation of the SANIC also permits multiple concurrent communications over a switched network without the traditional overhead associated with communicating protocols. In one embodiment, SAN <b>32</b> provides the I/O and IPC clients of distributed computer system <b>30</b> zero processor-copy data transfers without involving the operating system kernel process, and employs hardware to provide reliable, fault tolerant communications.
P-0046[0046] As indicated in FIG. 1, router <b>38</b> is coupled to wide area network (WAN) and/or local area network (LAN) connections to other hosts or other routers <b>38</b>.
P-0047[0047] The host processors <b>34</b><i>a</i>-<b>34</b><i>d </i>include central processing units (CPUs) <b>44</b> and memory <b>46</b>.
P-0048[0048] I/O adapters <b>35</b><i>a </i>and <b>35</b><i>b </i>include an I/O adapter backplane <b>48</b> and multiple I/O adapter cards <b>50</b>. Example adapter cards <b>50</b> illustrated in FIG. 1 include an SCSI adapter card; an adapter card to fiber channel hub and FC-AL devices; an Ethernet adapter card; and a graphics adapter card. Any known type of adapter card can be implemented. I/O adapters <b>35</b><i>a </i>and <b>35</b><i>b </i>also include a switch <b>36</b> in the I/O adapter backplane <b>48</b> to couple the adapter cards <b>50</b> to the SAN <b>32</b> fabric.
P-0049[0049] RAID subsystem <b>33</b> includes a microprocessor <b>52</b>, memory <b>54</b>, read/write circuitry <b>56</b>, and multiple redundant storage disks <b>58</b>.
P-0050[0050] SAN <b>32</b> handles data communications for I/O and IPC in distributed computer system <b>30</b>. SAN <b>32</b> supports high-bandwidth and scalability required for I/O and also supports the extremely low latency and low CPU overhead required for IPC. User clients can bypass the operating system kernel process and directly access network communication hardware, such as SANICs <b>42</b> which enable efficient message passing protocols. SAN <b>32</b> is suited to current computing models and is a building block for new forms of I/O and computer cluster communication. SAN <b>32</b> allows I/O adapter nodes to communicate among themselves or communicate with any or all of the processor nodes in distributed computer system <b>30</b>. With an I/O adapter attached to SAN <b>32</b>, the resulting I/O adapter node has substantially the same communication capability as any processor node in distributed computer system <b>30</b>.
P-0051[0051] Channel and Memory Semantics
P-0052[0052] In one embodiment, SAN <b>32</b> supports channel semantics and memory semantics. Channel semantics is sometimes referred to as send/receive or push communication operations, and is the type of communications employed in a traditional I/O channel where a source device pushes data and a destination device determines the final destination of the data. In channel semantics, the frame transmitted from a source process specifies a destination processes' communication port, but does not specify where in the destination processes' memory space the frame will be written. Thus, in channel semantics, the destination process pre-allocates where to place the transmitted data.
P-0053[0053] In memory semantics, a source process directly reads or writes the virtual address space of a remote node destination process. The remote destination process need only communicate the location of a buffer for data, and does not need to be involved with the transfer of any data. Thus, in memory semantics, a source process sends a data frame containing the destination buffer memory address of the destination process. In memory semantics, the destination process previously grants permission for the source process to access its memory.
P-0054[0054] Channel semantics and memory semantics are typically both necessary for I/O and IPC. A typical I/O operation employs a combination of channel and memory semantics. In an illustrative example I/O operation of distributed computer system <b>30</b>, host processor <b>34</b><i>a </i>initiates an I/O operation by using channel semantics to send a disk write command to I/O adapter <b>35</b><i>b. </i>I/O adapter <b>35</b><i>b </i>examines the command and uses memory semantics to read the data buffer directly from the memory space of host processor <b>34</b><i>a. </i>After the data buffer is read, I/O adapter <b>35</b><i>b </i>employs channel semantics to push an I/O completion message back to host processor <b>34</b><i>a. </i>
P-0055[0055] In one embodiment, distributed computer system <b>30</b> performs operations that employ virtual addresses and virtual memory protection mechanisms to ensure correct and proper access to all memory. In one embodiment, applications running in distributed computed system <b>30</b> are not required to use physical addressing for any operations.
P-0056[0056] Queue Pairs
P-0057[0057] An example host processor node <b>34</b> is generally illustrated in FIG. 2. Host processor node <b>34</b> includes a process A indicated at <b>60</b> and a process B indicated at <b>62</b>. Host processor node <b>34</b> includes SANIC <b>42</b>. Host processor node <b>34</b> also includes queue pairs (QPs) <b>64</b><i>a </i>and <b>64</b><i>b </i>which provide communication between process <b>60</b> and SANIC <b>42</b>. Host processor node <b>34</b> also includes QP <b>64</b><i>c </i>which provides communication between process <b>62</b> and SANIC <b>42</b>. A single SANIC, such as SANIC <b>42</b> in a host processor <b>34</b>, can support thousands of QPs. By contrast, a SAN interface in an I/O adapter <b>35</b> typically supports less than ten QPs.
P-0058[0058] Each QP <b>64</b> includes a send work queue <b>66</b> and a receive work queue <b>68</b>. A process, such as processes <b>60</b> and <b>62</b>, calls an operating-system specific programming interface which is herein referred to as verbs, which place work items, referred to as work queue elements (WQEs) onto a QP <b>64</b>. A WQE is executed by hardware in SANIC <b>42</b>. SANIC <b>42</b> is coupled to SAN <b>32</b> via physical link <b>40</b>. Send work queue <b>66</b> contains WQEs that describe data to be transmitted on the SAN <b>32</b> fabric. Receive work queue <b>68</b> contains WQEs that describe where to place incoming data from the SAN <b>32</b> fabric.
P-0059[0059] Host processor node <b>34</b> also includes completion queue <b>70</b><i>a </i>interfacing with process <b>60</b> and completion queue <b>70</b><i>b </i>interfacing with process <b>62</b>. The completion queues <b>70</b> contain information about completed WQEs. The completion queues are employed to create a single point of completion notification for multiple QPs. A completion queue entry is a data structure on a completion queue <b>70</b> that describes a completed WQE. The completion queue entry contains sufficient information to determine the QP that holds the completed WQE. A completion queue context is a block of information that contains pointers to, length, and other information needed to manage the individual completion queues.
P-0060[0060] Example WQEs include work items that initiate data communications employing channel semantics or memory semantics; work items that are instructions to hardware in SANIC <b>42</b> to set or alter remote memory access protections; and work items to delay the execution of subsequent WQEs posted in the same send work queue <b>66</b>.
P-0061[0061] More specifically, example WQEs supported for send work queues <b>66</b> are as follows. A send buffer WQE is a channel semantic operation to push a local buffer to a remote QP's receive buffer. The send buffer WQE includes a gather list to combine several virtual contiguous local buffers into a single message that is pushed to a remote QP's receive buffer. The local buffer virtual addresses are in the address space of the process that created the local QP.
P-0062[0062] A remote direct memory access (RDMA) read WQE provides a memory semantic operation to read a virtually contiguous buffer on a remote node. The RDMA read WQE reads a virtually contiguous buffer on a remote endnode and writes the data to a virtually contiguous local memory buffer. Similar to the send buffer WQE, the local buffer for the RDMA read WQE is in the address space of the process that created the local QP. The remote buffer is in the virtual address space of the process owning the remote QP targeted by the RDMA read WQE.
P-0063[0063] A RDMA write WQE provides a memory semantic operation to write a virtually contiguous buffer on a remote node. The RDMA write WQE contains a scatter list of locally virtually contiguous buffers and the virtual address of the remote buffer into which the local buffers are written.
P-0064[0064] A RDMA FetchOp WQE provides a memory semantic operation to perform an atomic operation on a remote word. The RDMA FetchOp WQE is a combined RDMA read, modify, and RDMA write operation. The RDMA FetchOp WQE can support several read-modify-write operations, such as Compare and Swap if equal.
P-0065[0065] A bind/unbind remote access key (RKey) WQE provides a command to SANIC hardware to modify the association of a RKey with a local virtually contiguous buffer. The RKey is part of each RDMA access and is used to validate that the remote process has permitted access to the buffer.
P-0066[0066] A delay WQE provides a command to SANIC hardware to delay processing of the QP's WQEs for a specific time interval. The delay WQE permits a process to meter the flow of operations into the SAN fabric.
P-0067[0067] In one embodiment, receive work queues <b>68</b> only support one type of WQE, which is referred to as a receive buffer WQE. The receive buffer WQE provides a channel semantic operation describing a local buffer into which incoming send messages are written. The receive buffer WQE includes a scatter list describing several virtually contiguous local buffers. An incoming send message is written to these buffers. The buffer virtual addresses are in the address space of the process that created the local QP.
P-0068[0068] For IPC, a user-mode software process transfers data through QPs <b>64</b> directly from where the buffer resides in memory. In one embodiment, the transfer through the QPs bypasses the operating system and consumes few host instruction cycles. QPs <b>64</b> permit zero processor-copy data transfer with no operating system kernel involvement. The zero processor-copy data transfer provides for efficient support of high-bandwidth and low-latency communication.
P-0069[0069] Transport Services
P-0070[0070] When a QP <b>64</b> is created, the QP is set to provide a selected type of transport service. In one embodiment, a distributed computer system implementing the present invention supports four types of transport services.
P-0071[0071] A portion of a distributed computer system employing a reliable connection service to communicate between distributed processes is illustrated generally at <b>100</b> in FIG. 3. Distributed computer system <b>100</b> includes a host processor node <b>102</b>, a host processor node <b>104</b>, and a host processor node <b>106</b>. Host processor node <b>102</b> includes a process A indicated at <b>108</b>. Host processor node <b>104</b> includes a process B indicated at <b>110</b> and a process C indicated at <b>112</b>. Host processor node <b>106</b> includes a process D indicated at <b>114</b>.
P-0072[0072] Host processor node <b>102</b> includes a QP <b>116</b> having a send work queue <b>116</b><i>a </i>and a receive work queue <b>116</b><i>b; </i>a QP <b>118</b> having a send work queue <b>118</b><i>a </i>and receive work queue <b>118</b><i>b; </i>and a QP <b>120</b> having a send work queue <b>120</b><i>a </i>and a receive work queue <b>120</b><i>b </i>which facilitate communication to and from process A indicated at <b>108</b>. Host processor node <b>104</b> includes a QP <b>122</b> having a send work queue <b>122</b><i>a </i>and receive work queue <b>122</b><i>b </i>for facilitating communication to and from process B indicated at <b>110</b>. Host processor node <b>104</b> includes a QP <b>124</b> having a send work queue <b>124</b><i>a </i>and receive work queue <b>124</b><i>b </i>for facilitating communication to and from process C indicated at <b>112</b>. Host processor node <b>106</b> includes a QP <b>126</b> having a send work queue <b>126</b><i>a </i>and receive work queue <b>126</b><i>b </i>for facilitating communication to and from process D indicated at <b>114</b>.
P-0073[0073] The reliable connection service of distributed computer system <b>100</b> associates a local QP with one and only one remote QP. Thus, QP <b>116</b> is connected to QP <b>122</b> via a non-sharable resource connection <b>128</b> having a non-sharable resource connection <b>128</b><i>a </i>from send work queue <b>116</b><i>a </i>to receive work queue <b>122</b><i>b </i>and a non-sharable resource connection <b>128</b><i>b </i>from send work queue <b>122</b><i>a </i>to receive work queue <b>116</b><i>b. </i>QP <b>118</b> is connected to QP <b>124</b> via a non-sharable resource connection <b>130</b> having a non-sharable resource connection <b>130</b><i>a </i>from send work queue <b>118</b><i>a </i>to receive work queue <b>124</b><i>b </i>and a non-sharable resource connection <b>130</b><i>b </i>from send work queue <b>124</b><i>a </i>to receive work queue <b>118</b><i>b. </i>QP <b>120</b> is connected to QP <b>126</b> via a non-sharable resource connection <b>132</b> having a non-sharable resource connection <b>132</b><i>a </i>from send work queue <b>120</b><i>a </i>to receive work queue <b>126</b><i>b </i>and a non-sharable resource connection <b>132</b><i>b </i>from send work queue <b>126</b><i>a </i>to receive work queue <b>120</b><i>b. </i>
P-0074[0074] A send buffer WQE placed on one QP in a reliable connection service causes data to be written into the receive buffer of the connected QP. RDMA operations operate on the address space of the connected QP.
P-0075[0075] The reliable connection service requires a process to create a QP for each process which is to communicate with over the SAN fabric. Thus, if each of N host processor nodes contain M processes, and all M processes on each node wish to communicate with all the processes on all the other nodes, each host processor node requires M<sup>2</sup>×(N−1) QPs. Moreover, a process can connect a QP to another QP on the same SANIC.
P-0076[0076] In one embodiment, the reliable connection service is made reliable because hardware maintains sequence numbers and acknowledges all frame transfers. A combination of hardware and SAN driver software retries any failed communications. The process client of the QP obtains reliable communications even in the presence of bit errors, receive buffer underruns, and network congestion. If alternative paths exist in the SAN fabric, reliable communications can be maintained even in the presence of failures of fabric switches or links.
P-0077[0077] In one embodiment, acknowledgements are employed to deliver data reliably across the SAN fabric. In one embodiment, the acknowledgement is not a process level acknowledgment, because the acknowledgment does not validate the receiving process has consumed the data. Rather, the acknowledgment only indicates that the data has reached its destination.
P-0078[0078] A portion of a distributed computer system employing a reliable datagram service to communicate between distributed processes is illustrated generally at <b>150</b> in FIG. 4. Distributed computer system <b>150</b> includes a host processor node <b>152</b>, a host processor node <b>154</b>, and a host processor node <b>156</b>. Host processor node <b>152</b> includes a process A indicated at <b>158</b>. Host processor node, <b>154</b> includes a process B indicated at <b>160</b> and a process C indicated at <b>162</b>. Host processor node <b>156</b> includes a process D indicated at <b>164</b>.
P-0079[0079] Host processor node <b>152</b> includes QP <b>166</b> having send work queue <b>166</b><i>a </i>and receive work queue <b>166</b><i>b </i>for facilitating communication to and from process A indicated at <b>158</b>. Host processor node <b>154</b> includes QP <b>168</b> having send work queue <b>168</b><i>a </i>and receive work queue <b>168</b><i>b </i>for facilitating communication from and to process B indicated at <b>160</b>. Host processor node <b>154</b> includes QP <b>170</b> having send work queue <b>170</b><i>a </i>and receive work queue <b>170</b><i>b </i>for facilitating communication from and to process C indicated at <b>162</b>. Host processor node <b>156</b> includes QP <b>172</b> having send work queue <b>172</b><i>a </i>and receive work queue <b>172</b><i>b </i>for facilitating communication from and to process D indicated at <b>164</b>. In the reliable datagram service implemented in distributed computer system <b>150</b>, the QPs are coupled in what is referred to as a connectionless transport service.
P-0080[0080] For example, a reliable datagram service <b>174</b> couples QP <b>166</b> to QPs <b>168</b>, <b>170</b>, and <b>172</b>. Specifically, reliable datagram service <b>174</b> couples send work queue <b>166</b><i>a </i>to receive work queues <b>168</b><i>b, </i><b>170</b><i>b, </i>and <b>172</b><i>b. </i>Reliable datagram service <b>174</b> also couples send work queues <b>168</b><i>a, </i><b>170</b><i>a, </i>and <b>172</b><i>a </i>to receive work queue <b>166</b><i>b. </i>
P-0081[0081] The reliable datagram service permits a client process of one QP to communicate with any other QP on any other remote node. At a receive work queue, the reliable datagram service permits incoming messages from any send work queue on any other remote node.
P-0082[0082] In one embodiment, the reliable datagram service employs sequence numbers and acknowledgments associated with each message frame to ensure the same degree of reliability as the reliable connection service. End-to-end (EE) contexts maintain end-to-end specific state to keep track of sequence numbers, acknowledgments, and time-out values. The end-to-end state held in the EE contexts is shared by all the connectionless QPs communicating between a pair of endnodes. Each endnode requires at least one EE context for every endnode it wishes to communicate with in the reliable datagram service (e.g., a given endnode requires at least N EE contexts to be able to have reliable datagram service with N other endnodes).
P-0083[0083] The reliable datagram service greatly improves scalability because the reliable datagram service is connectionless. Therefore, an endnode with a fixed number of QPs can communicate with far more processes and endnodes with a reliable datagram service than with a reliable connection transport service. For example, if each of N host processor nodes contain M processes, and all M processes on each node wish to communicate with all the processes on all the other nodes, the reliable connection service requires M<sup>2</sup>×(N−1) QPs on each node. By comparison, the connectionless reliable datagram service only requires M QPs+(N−1) EE contexts on each node for exactly the same communications.
P-0084[0084] A third type of transport service for providing communications is a unreliable datagram service. Similar to the reliable datagram service, the unreliable datagram service is connectionless. The unreliable datagram service is employed by management applications to discover and integrate new switches, routers, and endnodes into a given distributed computer system. The unreliable datagram service does not provide the reliability guarantees of the reliable connection service and the reliable datagram service. The unreliable datagram service accordingly operates with less state information maintained at each endnode.
P-0085[0085] A fourth type of transport service is referred to as raw datagram service and is technically not a transport service. The raw datagram service permits a QP to send and to receive raw datagram frames. The raw datagram mode of operation of a QP is entirely controlled by software. The raw datagram mode of the QP is primarily intended to allow easy interfacing with traditional internet protocol, version <b>6</b> (IPv6) LAN-WAN networks, and further allows the SANIC to be used with full software protocol stacks to access transmission control protocol (TCP), user datagram protocol (UDP), and other standard communication protocols. Essentially, in the raw datagram service, SANIC hardware generates and consumes standard protocols layered on top of IPv6, such as TCP and UDP. The frame header can be mapped directly to and from an IPv6 header. Native IPv6 frames can be bridged into the SAN fabric and delivered directly to a QP to allow a client process to support any transport protocol running on top of IPv6. A client process can register with SANIC hardware in order to direct datagrams for a particular upper level protocol (e.g., TCP and UDP) to a particular QP. SANIC hardware can demultiplex incoming IPv6 streams of datagrams based on a next header field as well as the destination IP address.
P-0086[0086] SANIC and I/O Adapter Endnodes
P-0087[0087] An example host processor node is generally illustrated at <b>200</b> in FIG. 5. Host processor node <b>200</b> includes a process A indicated at <b>202</b>, a process B indicated at <b>204</b>, and a process C indicated at <b>206</b>. Host processor <b>200</b> includes a SANIC <b>208</b> and a SANIC <b>210</b>. As discussed above, a host processor endnode or an I/O adapter endnode can have one or more SANICs. SANIC <b>208</b> includes a SAN link level engine (LLE) <b>216</b> for communicating with SAN fabric <b>224</b> via link <b>217</b> and an LLE <b>218</b> for communicating with SAN fabric <b>224</b> via link <b>219</b>. SANIC <b>210</b> includes an LLE <b>220</b> for communicating with SAN fabric <b>224</b> via link <b>221</b> and an LLE <b>222</b> for communicating with SAN fabric <b>224</b> via link <b>223</b>. SANIC <b>208</b> communicates with process A indicated at <b>202</b> via QPs <b>212</b><i>a </i>and <b>212</b><i>b. </i>SANIC <b>208</b> communicates with process B indicated at <b>204</b> via QPs <b>212</b><i>c</i>-<b>212</b><i>n. </i>Thus, SANIC <b>208</b> includes N QPs for communicating with processes A and B. SANIC <b>210</b> includes QPs <b>214</b><i>a </i>and <b>214</b><i>b </i>for communicating with process B indicated at <b>204</b>. SANIC <b>210</b> includes QPs <b>214</b><i>c</i>-<b>214</b><i>n </i>for communicating with process C indicated at <b>206</b>. Thus, SANIC <b>210</b> includes N QPs for communicating with processes B and C.
P-0088[0088] An LLE runs link level protocols to couple a given SANIC to the SAN fabric. RDMA traffic generated by a SANIC can simultaneously employ multiple LLEs within the SANIC which permits striping across LLEs. Striping refers to the dynamic sending of frames within a single message to an endnode's QP through multiple fabric paths. Striping across LLEs increases the bandwidth for a single QP as well as provides multiple fault tolerant paths. Striping also decreases the latency for message transfers. In one embodiment, multiple LLEs in a SANIC are not visible to the client process generating message requests. When a host processor includes multiple SANICs, the client process must explicitly move data on the two SANICs in order to gain parallelism. A single QP cannot be shared by SANICS. Instead a QP is owned by one local SANIC.
P-0089[0089] The following is an example naming scheme for naming and identifying endnodes in one embodiment of a distributed computer system according to the present invention. A host name provides a logical identification for a host node, such as a host processor node or I/O adapter node. The host name identifies the endpoint for messages such that messages are destine for processes residing on an endnode specified by the host name. Thus, there is one host name per node, but a node can have multiple SANICs.
P-0090[0090] A globally unique ID (GUID) identifies a transport endpoint. A transport endpoint is the device supporting the transport QPs. There is one GUID associated with each SANIC.
P-0091[0091] A local ID refers to a short address ID used to identify a SANIC within a single subnet. In one example embodiment, a subnet has up 2<sup>16 </sup>endnodes, switches, and routers, and the local ID (LID) is accordingly 16 bits. A source LID (SLID) and a destination LID (DLID) are the source and destination LIDs used in a local network header. A LLE has a single LID associated with the LLE, and the LID is only unique within a given subnet. One or more LIDs can be associated with each SANIC.
P-0092[0092] An internet protocol (IP) address (e.g., a 128 bit IPv6 ID) addresses a SANIC. The SANIC, however, can have one or more IP addresses associated with the SANIC. The IP address is used in the global network header when routing frames outside of a given subnet. LIDs and IP addresses are network endpoints and are the target of frames routed through the SAN fabric. All IP addresses (e.g., IPv6 addresses) within a subnet share a common set of high order address bits.
P-0093[0093] In one embodiment, the LLE is not named and is not architecturally visible to a client process. In this embodiment, management software refers to LLEs as an enumerated subset of the SANIC.
P-0094[0094] Switches and Routers
P-0095[0095] A portion of a distributed computer system is generally illustrated at <b>250</b> in/FIG. 6. Distributed computer system <b>250</b> includes a subnet A indicated at <b>252</b> and a subnet B indicated at <b>254</b>. Subnet A indicated at <b>252</b> includes a host processor node <b>256</b> and a host processor node <b>258</b>. Subnet B indicated at <b>254</b> includes a host processor node <b>260</b> and host processor node <b>262</b>. Subnet A indicated at <b>252</b> includes switches <b>264</b><i>a</i>-<b>264</b><i>c. </i>Subnet B indicated at <b>254</b> includes switches <b>266</b><i>a</i>-<b>266</b><i>c. </i>Each subnet within distributed computer system <b>250</b> is connected to other subnets with routers. For example, subnet A indicated at <b>252</b> includes routers <b>268</b><i>a </i>and <b>268</b><i>b </i>which are coupled to routers <b>270</b><i>a </i>and <b>270</b><i>b </i>of subnet B indicated at <b>254</b>. In one example embodiment, a subnet has up to 2<sup>16 </sup>endnodes, switches, and routers.
P-0096[0096] A subnet is defined as a group of endnodes and cascaded switches that is managed as a single unit. Typically, a subnet occupies a single geographic or functional area. For example, a single computer system in one room could be defined as a subnet. In one embodiment, the switches in a subnet can perform very fast worm-hole or cut-through routing for messages.
P-0097[0097] A switch within a subnet examines the DLID that is unique within the subnet to permit the switch to quickly and efficiently route incoming message frames. In one embodiment, the switch is a relatively simple circuit, and is typically implemented as a single integrated circuit. A subnet can have hundreds to thousands of endnodes formed by cascaded switches.
P-0098[0098] As illustrated in FIG. 6, for expansion to much larger systems, subnets are connected with routers, such as routers <b>268</b> and <b>270</b>. The router interprets the IP destination ID (e.g., IPv6 destination ID) and routes the IP like frame.
P-0099[0099] In one embodiment, switches and routers degrade when links are over utilized. In this embodiment, link level back pressure is used to temporarily slow the flow of data when multiple input frames compete for a common output. However, link or buffer contention does not cause loss of data. In one embodiment, switches, routers, and endnodes employ a link protocol to transfer data. In one embodiment, the link protocol supports an automatic error retry. In this example embodiment, link level acknowledgments detect errors and force retransmission of any data impacted by bit errors. Link-level error recovery greatly reduces the number of data errors that are handled by the end-to-end protocols. In one embodiment, the user client process is not involved with error recovery no matter if the error is detected and corrected by the link level protocol or the end-to-end protocol.
P-0100[0100] An example embodiment of a switch is generally illustrated at <b>280</b> in FIG. 7. Each I/O path on a switch or router has an LLE. For example, switch <b>280</b> includes LLEs <b>282</b><i>a</i>-<b>282</b><i>h </i>for communicating respectively with links <b>284</b><i>a</i>-<b>284</b><i>h. </i>
P-0101[0101] The naming scheme for switches and routers is similar to the above-described naming scheme for endnodes. The following is an example switch and router naming scheme for identifying switches and routers in the SAN fabric. A switch name identifies each switch or group of switches packaged and managed together. Thus, there is a single switch name for each switch or group of switches packaged and managed together.
P-0102[0102] Each switch or router element has a single unique GUID. Each switch has one or more LIDs and IP addresses (e.g., IPv6 addresses) that are used as an endnode for management frames.
P-0103[0103] Each LLE is not given an explicit external name in the switch or router. Since links are point-to-point, the other end of the link does not need to address the LLE.
P-0104[0104] Virtual Lanes
P-0105[0105] Switches and routers employ multiple virtual lanes within a single physical link. As illustrated in FIG. 6, physical links <b>272</b> connect endnodes, switches, and routers within a subnet. WAN or LAN connections <b>274</b> typically couple routers between subnets. Frames injected into the SAN fabric follow a particular virtual lane from the frame's source to the frame's destination. At any one time, only one virtual lane makes progress on a given physical link. Virtual lanes provide a technique for applying link level flow control to one virtual lane without affecting the other virtual lanes. When a frame on one virtual lane blocks due to contention, quality of service (QoS), or other considerations, a frame on a different virtual lane is allowed to make progress.
P-0106[0106] Virtual lanes are employed for numerous reasons, some of which are as follows. Virtual lanes provide QoS. In one example embodiment, certain virtual lanes are reserved for high priority or isonchronous traffic to provide QoS.
P-0107[0107] Virtual lanes provide deadlock avoidance. Virtual lanes allow topologies that contain loops to send frames across all physical links and still be assured the loops won't cause back pressure dependencies that might result in deadlock.
P-0108[0108] Virtual lanes alleviate head-of-line blocking. With virtual lanes, a blocked frames can pass a temporarily stalled frame that is destined for a different final destination.
P-0109[0109] In one embodiment, each switch includes its own crossbar switch. In this embodiment, a switch propagates data from only one frame at a time, per virtual lane through its crossbar switch. In another words, on any one virtual lane, a switch propagates a single frame from start to finish. Thus, in this embodiment, frames are not multiplexed together on a single virtual lane.
P-0110[0110] Paths in SAN Fabric
P-0111[0111] Referring to FIG. 6, within a subnet, such as subnet A indicated at <b>252</b> or subnet B indicated at <b>254</b>, a path from a source port to a destination port is determined by the LID of the destination SANIC port. Between subnets, a path is determined by the IP address (e.g., IPv6 address) of the destination SANIC port.
P-0112[0112] In one embodiment, the paths used by the request frame and the request frame's corresponding positive acknowledgment (ACK) or negative acknowledgment (NAK) frame are not required to be symmetric. In one embodiment employing oblivious routing, switches select an output port based on the DLID. In one embodiment, a switch uses one set of routing decision criteria for all its input ports. In one example embodiment, the routing decision criteria is contained in one routing table. In an alternative embodiment, a switch employs a separate set of criteria for each input port.
P-0113[0113] Each port on an endnode can have multiple IP addresses. Multiple IP addresses can be used for several reasons, some of which are provided by the following examples. In one embodiment, different IP addresses identify different partitions or services on an endnode. In one embodiment, different IP addresses are used to specify different QoS attributes. In one embodiment, different IP addresses identify different paths through intra-subnet routes.
P-0114[0114] In one embodiment, each port on an endnode can have multiple LIDs. Multiple LIDs can be used for several reasons some of which are provided by the following examples. In one embodiment, different LIDs identify different partitions or services on an endnode. In one embodiment, different LIDs are used to specify different QoS attributes. In one embodiment, different LIDs specify different paths through the subnet.
P-0115[0115] A one-to-one correspondence does not necessarily exist between LIDs and IP addresses, because a SANIC can have more or less LIDs than IP addresses for each port. For SANICs with redundant ports and redundant conductivity to multiple SAN fabrics, SANICs can, but are not required to, use the same LID and IP address on each of its ports.
P-0116[0116] Data Transactions
P-0117[0117] Referring to FIG. 1, a data transaction in distributed computer system <b>30</b> is typically composed of several hardware and software steps. A client process of a data transport service can be a user-mode or a kernel-mode process. The client process accesses SANIC <b>42</b> hardware through one or more QPs, such as QPs <b>64</b> illustrated in FIG. 2. The client process calls an operating-system specific programming interface which is herein referred to as verbs. The software code implementing the verbs intern posts a WQE to the given QP work queue.
P-0118[0118] There are many possible methods of posting a WQE and there are many possible WQE formats, which allow for various cost/performance design points, but which do not affect interoperability. A user process, however, must communicate to verbs in a well-defined manner, and the format and protocols of data transmitted across the SAN fabric must be sufficiently specified to allow devices to interoperate in a heterogeneous vendor environment.
P-0119[0119] In one embodiment, SANIC hardware detects WQE posting and accesses the WQE. In this embodiment, the SANIC hardware translates and validates the WQEs virtual addresses and accesses the data. In one embodiment, an outgoing message buffer is split into one or more frames. In one embodiment, the SANIC hardware adds a transport header and a network header to each frame. The transport header includes sequence numbers and other transport information. The network header includes the destination IP address or the DLID or other suitable destination address information. The appropriate local or global network header is added to a given frame depending on if the destination endnode resides on the local subnet or on a remote subnet.
P-0120[0120] A frame is a unit of information that is routed through the SAN fabric. The frame is an endnode-to-endnode construct, and is thus created and consumed by endnodes. Switches and routers neither generate nor consume request frames or acknowledgment frames. Instead switches and routers simply move request frames or acknowledgment frames closer to the ultimate destination. Routers, however, modify the frame's network header when the frame crosses a subnet boundary. In traversing a subnet, a single frame stays on a single virtual lane.
P-0121[0121] When a frame is placed onto a link, the frame is further broken down into flits. A flit is herein defined to be a unit of link-level flow control and is a unit of transfer employed only on a point-to-point link. The flow of flits is subject to the link-level protocol which can perform flow control or retransmission after an error. Thus, flit is a link-level construct that is created at each endnode, switch, or router output port and consumed at each input port. In one embodiment, a flit contains a header with virtual lane error checking information, size information, and reverse channel credit information.
P-0122[0122] If a reliable transport service is employed, after a request frame reaches its destination endnode, the destination endnode sends an acknowledgment frame back to the sender endnode. The acknowledgment frame permits the requester to validate that the request frame reached the destination endnode. An acknowledgment frame is sent back to the requester after each request frame. The requestor can have multiple outstanding requests before it receives any acknowledgments. In one embodiment, the number of multiple outstanding requests is determined when a QP is created.
P-0123[0123] Example Request and Acknowledgment Transactions
P-0124[0124]FIGS. 8, 9A, <b>9</b>B, <b>10</b>A, and <b>10</b>B together illustrate example request and acknowledgment transactions. In FIG. 8, a portion of a distributed computer system is generally illustrated at <b>300</b>. Distributed computer system <b>300</b> includes a host processor node <b>302</b> and a host processor node <b>304</b>. Host processor node <b>302</b> includes a SANIC <b>306</b>. Host processor node <b>304</b> includes a SANIC <b>308</b>. Distributed computer system <b>300</b> includes a SAN fabric <b>309</b> which includes a switch <b>310</b> and a switch <b>312</b>. SAN fabric <b>309</b> includes a link <b>314</b> coupling SANIC <b>306</b> to switch <b>310</b>; a link <b>316</b> coupling switch <b>310</b> to switch <b>312</b>; and a link <b>318</b> coupling SANIC <b>308</b> to switch <b>312</b>.
P-0125[0125] In the example transactions, host processor node <b>302</b> includes a client process A indicated at <b>320</b>. Host processor node <b>304</b> includes a client process B indicated at <b>322</b>. Client process <b>320</b> interacts with SANIC hardware <b>306</b> through QP <b>324</b>. Client process <b>322</b> interacts with SANIC hardware <b>308</b> through QP <b>326</b>. QP <b>324</b> and <b>326</b> are software data structures. QP <b>324</b> includes send work queue <b>324</b><i>a </i>and receive work queue <b>324</b><i>b. </i>QP <b>326</b> includes send work queue <b>326</b><i>a </i>and receive work queue <b>326</b><i>b. </i>
P-0126[0126] Process <b>320</b> initiates a message request by posting WQEs to send work queue <b>324</b><i>a. </i>Such a WQE is illustrated at <b>330</b> in FIG. 9A. The message request of client process <b>320</b> is referenced by a gather list <b>332</b> contained in send WQE <b>330</b>. Each entry in gather list <b>332</b> points to a virtually contiguous buffer in the local memory space containing a part of the message, such as indicated by virtual contiguous buffers <b>334</b><i>a</i>-<b>334</b><i>d, </i>which respectively hold message 0, parts 0, 1, 2, and 3.
P-0127[0127] Referring to FIG. 9B, hardware in SANIC <b>306</b> reads WQE <b>330</b> and packetizes the message stored in virtual contiguous buffers <b>334</b><i>a</i>-<b>334</b><i>d </i>into frames and flits. As illustrated in FIG. 9B, all of message 0, part 0 and a portion of message 0, part 1 are packetized into frame 0, indicated at <b>336</b><i>a. </i>The rest of message 0, part 1 and all of message 0, part 2, and all of message 0, part 3 are packetized into frame 1, indicated at <b>336</b><i>b. </i>Frame 0 indicated at <b>336</b><i>a </i>includes network header <b>338</b><i>a </i>and transport header <b>340</b><i>a. </i>Frame <b>1</b> indicated at <b>336</b><i>b </i>includes network header <b>338</b><i>b </i>and transport header <b>340</b><i>b. </i>
P-0128[0128] As indicated in FIG. 9B, frame 0 indicated at <b>336</b><i>a </i>is partitioned into flits 0-3, indicated respectively at <b>342</b><i>a</i>-<b>342</b><i>d. </i>Frame 1 indicated at <b>336</b><i>b </i>is partitioned into flits 4-7 indicated respectively at <b>342</b><i>e</i>-<b>342</b><i>h. </i>Flits <b>342</b><i>a </i>through <b>342</b><i>h </i>respectively include flit headers <b>344</b><i>a</i>-<b>344</b><i>h. </i>
P-0129[0129] Frames are routed through the SAN fabric, and for reliable transfer services, are acknowledged by the final destination endnode. If not successfully acknowledged, the frame is retransmitted by the source endnode. Frames are generated by source endnodes and consumed by destination endnodes. The switches and routers in the SAN fabric neither generate nor consume frames.
P-0130[0130] Flits are the smallest unit of flow control in the network. Flits are generated and consumed at each end of a physical link. Flits are acknowledged at the receiving end of each link and are retransmitted in response to an error. For example, controlling retransmission or abortion of flit transmission can be accomplished with each sender maintaining a per-port link retry timer to monitor the time between flit transmission and receipt acknowledgment. If the link retry timer expires, the sender attempts to retry transmission of the outstanding flit. A second timer, called a link kill timer, is active while the sender is operating in a retry mode. On expiration of the link kill timer, transmission is aborted.
P-0131[0131] Referring to FIG. 10A, the send request message 0 is transmitted from SANIC <b>306</b> in host processor node <b>302</b> to SANIC <b>308</b> in host processor node <b>304</b> as frames 0 indicated at <b>336</b><i>a </i>and frame 1 indicated at <b>336</b><i>b. </i>ACK frames <b>346</b><i>a </i>and <b>346</b><i>b, </i>corresponding respectively to request frames <b>336</b><i>a </i>and <b>336</b><i>b, </i>are transmitted from SANIC <b>308</b> in host processor node <b>304</b> to SANIC <b>306</b> in host processor node <b>302</b>.
P-0132[0132] In FIG. 10A, message 0 is being transmitted with a reliable transport service. Each request frame is individually acknowledged by the destination endnode (e.g., SANIC <b>308</b> in host processor node <b>304</b>).
P-0133[0133]FIG. 10B illustrates the flits associated with the request frames <b>336</b> and acknowledgment frames <b>346</b> illustrated in FIG. 10A passing between the host processor endnodes <b>302</b> and <b>304</b> and the switches <b>310</b> and <b>312</b>. As illustrated in FIG. 10B, an ACK frame fits inside one flit. In one embodiment, one acknowledgment flit acknowledges several flits.
P-0134[0134] As illustrated in FIG. 10B, flits <b>342</b><i>a</i>-<i>h </i>are transmitted from SANIC <b>306</b> to switch <b>310</b>. Switch <b>310</b> consumes flits <b>342</b><i>a</i>-<i>h </i>at its input port, creates flits <b>348</b><i>a</i>-<i>h </i>at its output port corresponding to flits <b>342</b><i>a</i>-<i>h, </i>and transmits flits <b>348</b><i>a</i>-<i>h </i>to switch <b>312</b>. Switch <b>312</b> consumes flits <b>348</b><i>a</i>-<i>h </i>at its input port, creates flits <b>350</b><i>a</i>-<i>h </i>at its output port corresponding to flits <b>348</b><i>a</i>-<i>h, </i>and transmits flits <b>350</b><i>a</i>-<i>h </i>to SANIC <b>308</b>. SANIC <b>308</b> consumes flits <b>350</b><i>a</i>-<i>h </i>at its input port. An acknowledgment flit is transmitted from switch <b>310</b> to SANIC <b>306</b> to acknowledge the receipt of flits <b>342</b><i>a</i>-<i>h. </i>An acknowledgment flit <b>354</b> is transmitted from switch <b>312</b> to switch <b>310</b> to acknowledge the receipt of flits <b>348</b><i>a</i>-<i>h. </i>An acknowledgment flit <b>356</b> is transmitted from SANIC <b>308</b> to switch <b>312</b> to acknowledge the receipt of flits <b>350</b><i>a</i>-<i>h. </i>
P-0135[0135] Acknowledgment frame <b>346</b><i>a </i>fits inside of flit <b>358</b> which is transmitted from SANIC <b>308</b> to switch <b>312</b>. Switch <b>312</b> consumes flits <b>358</b> at its input port, creates flit <b>360</b> corresponding to flit <b>358</b> at its output port, and transmits flit <b>360</b> to switch <b>310</b>. Switch <b>310</b> consumes flit <b>360</b> at its input port, creates flit <b>362</b> corresponding to flit <b>360</b> at its output port, and transmits flit <b>362</b> to SANIC <b>306</b>. SANIC <b>306</b> consumes flit <b>362</b> at its input port. Similarly, SANIC <b>308</b> transmits acknowledgment frame <b>346</b><i>b </i>in flit <b>364</b> to switch <b>312</b>. Switch <b>312</b> creates flit <b>366</b> corresponding to flit <b>364</b>, and transmits flit <b>366</b> to switch <b>310</b>. Switch <b>310</b> creates flit <b>368</b> corresponding to flit <b>366</b>, and transmits flit <b>368</b> to SANIC <b>306</b>.
P-0136[0136] Switch <b>312</b> acknowledges the receipt of flits <b>358</b> and <b>364</b> with acknowledgment flit <b>370</b>, which is transmitted from switch <b>312</b> to SANIC <b>308</b>. Switch <b>310</b> acknowledges the receipt of flits <b>360</b> and <b>366</b> with acknowledgment flit <b>372</b>, which is transmitted to switch <b>312</b>. SANIC <b>306</b> acknowledges the receipt of flits <b>362</b> and <b>368</b> with acknowledgment flit <b>374</b> which is transmitted to switch <b>310</b>.
P-0137[0137] Architecture Layers and Implementation Overview
P-0138[0138] A host processor endnode and an I/O adapter endnode typically have quite different capabilities. For example, an example host processor endnode might support four ports, hundreds to thousands of QPs, and allow incoming RDMA operations, while an attached I/O adapter endnode might only support one or two ports, tens of QPs, and not allow incoming RDMA operations. A low-end attached I/O adapter alternatively can employ software to handle much of the network and transport layer functionality which is performed in hardware (e.g., by SANIC hardware) at the host processor endnode.
P-0139[0139] One embodiment of a layered architecture for implementing the present invention is generally illustrated at <b>400</b> in diagram form in FIG. 11. The layered architecture diagram of FIG. 11 shows the various layers of data communication paths, and organization of data and control information passed between layers.
P-0140[0140] Host SANIC endnode layers are generally indicated at <b>402</b>. The host SANIC endnode layers <b>402</b> include an upper layer protocol <b>404</b>; a transport layer <b>406</b>; a network layer <b>408</b>; a link layer <b>410</b>; and a physical layer <b>412</b>.
P-0141[0141] Switch or router layers are generally indicated at <b>414</b>. Switch or router layers <b>414</b> include a network layer <b>416</b>; a link layer <b>418</b>; and a physical layer <b>420</b>.
P-0142[0142] I/O adapter endnode layers are generally indicated at <b>422</b>. I/O adapter endnode layers <b>422</b> include an upper layer protocol <b>424</b>; a transport layer <b>426</b>; a network layer <b>428</b>; a link layer <b>430</b>; and a physical layer <b>432</b>.
P-0143[0143] The layered architecture <b>400</b> generally follows an outline of a classical communication stack. The upper layer protocols employ verbs to create messages at the transport layers. The transport layers pass messages to the network layers. The network layers pass frames down to the link layers. The link layers pass flits through physical layers. The physical layers send bits or groups of bits to other physical layers. Similarly, the link layers pass flits to other link layers, and don't have visibility to how the physical layer bit transmission is actually accomplished. The network layers only handle frame routing, without visibility to segmentation and reassembly of frames into flits or transmission between link layers.
P-0144[0144] Bits or groups of bits are passed between physical layers via links <b>434</b>. Links <b>434</b> can be implemented with printed circuit copper traces, copper cable, optical cable, or with other suitable links.
P-0145[0145] The upper layer protocol layers are applications or processes which employ the other layers for communicating between endnodes.
P-0146[0146] The transport layers provide end-to-end message movement. In one embodiment, the transport layers provide four types of transport services as described above which are reliable connection service; reliable datagram service; unreliable datagram service; and raw datagram service.
P-0147[0147] The network layers perform frame routing through a subnet or multiple subnets to destination endnodes.
P-0148[0148] The link layers perform flow-controlled, error controlled, and prioritized frame delivery across links.
P-0149[0149] The physical layers perform technology-dependent bit transmission and reassembly into flits.
P-0150[0150] Congestion Management
P-0151[0151] If a network is properly provisioned and the topology is well understood and controlled, network traffic will make forward progress under most operating workloads (albeit at reduced operating efficiency) assuming that network end nodes are consuming their inbound frames, and thus generally avoiding back pressure. However, at very high speeds, or in a network with variable speed link elements and/or topologies, the overall network efficiency and application effective throughput can drop to the point where the network appears to be down, even while network traffic is not in a condition of deadlock per se.
P-0152[0152] An example embodiment of a method for recognizing and managing network system congestion using congestion detection, reporting, and responding mechanisms is illustrated in FIG. 12. The method involves monitoring network traffic conditions for signs of inefficiency, or network bandwidth underutilization. At <b>435</b>, monitored network conditions <b>436</b> are tested for the presence of certain conditions indicative of delayed frame transmission progress, defined at <b>437</b>. The result of the testing at <b>435</b> is indicated at <b>438</b>. At <b>439</b>, based on test result <b>438</b>, the duration of monitored conditions <b>436</b> is timed and compared to a configured timing threshold <b>440</b>. At <b>441</b>, an indication is provided if, at <b>439</b>, the measured duration of monitored conditions <b>436</b> exceeds configured timing threshold <b>440</b>. A response <b>442</b> is made based on indication <b>441</b> and on response configuration <b>443</b>.
P-0153[0153] One embodiment of measurement mechanism <b>439</b> is implemented in a network end station, such as an end node or routing device, that includes one or more aggressive timers for indicating a delay in frame transmission. An aggressive timer is herein defined as a timing device or mechanism that monitors the transmission of information in a network system or network system element, and can be configured to provide an indication of whether and when the transmission of information fails to meet at least one variable timing threshold. An aggressive timer can be implemented in hardware or software.
P-0154[0154] One embodiment of a network system is illustrated generally at <b>449</b> in FIG. 13A. Network system <b>449</b> includes an end station <b>450</b> comprising aggressive timer <b>452</b>. End station <b>450</b> interfaces with other end stations of SAN <b>451</b> through links <b>484</b><i>a </i>and <b>484</b><i>b. </i>Network traffic internal to end station <b>450</b> is indicated at <b>454</b>. Aggressive timer <b>452</b> monitors network traffic <b>454</b> via a monitoring mechanism <b>456</b>.
P-0155[0155] Aggressive timer <b>452</b>'s variable timing threshold is configurable and re-configurable as a function of network system attributes, as indicated at <b>453</b>. An attribute of an entity is herein defined as including, but not limited to, at least one aspect, characteristic, condition, configuration, essence, parameter, property, quality, setting, or status, of the entity to which it refers. Aggressive timer <b>452</b> sufficiently monitors network traffic to permit its variable timing threshold to be configured to detect or anticipate the onset of congestion. Aggressive timer <b>452</b>'s variable timing threshold is configurable and re-configurable to accommodate different network circumstances, such as variations in network system attributes, including frame attributes and operating conditions within the network system.
P-0156[0156] The variable timing threshold of the aggressive timer <b>452</b> refers to a duration of time against which one or more defined conditions of delayed frame transmission progress can be compared. The conditions of delayed frame transmission progress can be defined in a number of ways according to network traffic management policy. For example, in one embodiment, an expiration of a link retry timer is a condition of delayed frame transmission progress. In one embodiment, the variable timing threshold is varied from one configuration to another by setting a starting count, ending count and/or counting rate. In one embodiment, the variable timing threshold is varied by adjusting a timing interval having a configurable ending time relative to a starting time. In one embodiment, the variable timing threshold is configurable to enable or disable the aggressive timer. When the variable timing threshold is exceeded, numerous suitable responses can be taken as part of enforcing a congestion management policy.
P-0157[0157]FIG. 13B illustrates an exemplary end station <b>450</b> according to one embodiment of the present invention, which includes an aggressive timer <b>452</b>. The Aggressive timer <b>452</b> is adapted to monitor one or more ports <b>455</b> of end station <b>450</b>. In one embodiment, aggressive timer <b>452</b> measures whether the transmission of a frame has stalled for a time that exceeds the variable timing threshold. In one embodiment, a variable timing threshold's duration is only slightly longer than the time needed for transmitting a given frame out of a port <b>455</b> under normal operating conditions. Such a short time limit facilitates fast recognition of transmission delay, which permits anticipation of network congestion. In one embodiment, the aggressive timer has microsecond granularity.
P-0158[0158] In one embodiment, aggressive timer <b>452</b> is implemented in end station <b>450</b> as a count-down timer set to a specific counting value and counting rate, which together, represent the configured variable timing threshold. In this embodiment, count down aggressive timer <b>452</b> counts down during one or more defined conditions of delayed frame transmission progress, such as a period of transmission delay associated with a port <b>455</b>. In one form of this embodiment, the frame transmission delay is defined by an indication of the output frame's failure to make forward progress. If the frame makes forward progress, count-down aggressive timer <b>452</b> is reset. In another form of this embodiment, the count-down aggressive timer <b>452</b> is set to count while any part of a frame remains in a transmit register, regardless of forward progress status.
P-0159[0159] In one embodiment, the variable timing threshold is configured by an appropriate authority, such as network fabric management. In an alternative embodiment, the variable timing threshold is dynamically configurable, either manually or by program, based on potentially changing circumstances or decision-making criteria relating to the network system. Examples of potentially changing network circumstances include network system attributes, such as a congestion level of an end station's port, a policy for managing network congestion, or various attributes of a frame to be transmitted.
P-0160[0160] An exemplary potentially changing decision-making criteria relating to the network system, is a congestion management policy instituted by network fabric management. Network circumstances and decision-making criteria can also be interrelated, such that a congestion management policy, for example, can change in response to changing network system conditions. The dynamic characteristics of embodiments of the present invention can be employed to implement potentially changing network decision-making criteria, measuring potentially changing network circumstances, or both.
P-0161[0161] One embodiment of aggressive timer <b>452</b> is configurable based on one or more network system attributes. In this embodiment, configuration of the variable timing threshold can be made based on one or more predetermined network system attributes and/or based on measured network system status. Predetermined network system attributes herein refers to information characterizing at least one aspect of the network system and its contents that was determined at some time before the period surrounding the setting of the variable timing threshold, or over a period of time prior to, and potentially including, the decision-making period. Measured network system status, on the other hand, refers to at least one aspect of the network system and its contents that is determined in the period surrounding the setting of the variable timing threshold. The variable timing threshold can be dynamically configurable in embodiments of aggressive timer <b>452</b> having a variable timing threshold based on predetermined or measured network system attributes.
P-0162[0162] In some embodiments, aggressive timer <b>452</b> is configurable based on one or more attributes of at least one port <b>455</b>. In one such embodiment, the aggressive timer <b>452</b>'s variable timing threshold is adjustable as a function of a measured presence of back pressure from a port. Back pressure herein refers to network operation that is indicative of insufficient buffer space. In another such embodiment, the variable timing threshold is based on historical data of a port's back pressure occurrences over a period of time. In one example operation of end station <b>450</b> of FIG. 13B, port <b>455</b><i>a </i>tends to be congested, and end station <b>450</b> accordingly institutes a more stringent variable timing threshold for congestion management.
P-0163[0163] In one embodiment, the variable timing threshold is configurable based on the type of workload of applications utilizing the network system. FIG. 13C illustrates an end station <b>450</b>, which runs an application <b>458</b>. Aggressive timer <b>452</b>'s variable timing threshold is configurable based on the type of workload of application <b>458</b>. Workload herein refers to the information exchanged over the network for a given application. Example types of workloads include digital video information for video applications, and file transfer protocol (FTP). Since video workloads are less tolerant of inconsistent transmission rates, an exemplary congestion management policy can impose variable timing thresholds with shorter limits for video applications than for web page browsing applications. In one embodiment, a network system which predominately carries video application data, comprises end stations <b>450</b> with aggressive timers <b>452</b> configured with stringent variable timing thresholds for early detection and prevention of congestion.
P-0164[0164] In one embodiment, the variable timing threshold is configurable based on an amount of time one or more frames fails to make forward progress. In one embodiment, end station <b>450</b> employs aggressive timer <b>452</b> to measure congestion experienced by a first frame targeting a particular port <b>455</b>. In one example of this embodiment, the aggressive timer <b>452</b> is set to expire after a time period that is approximately the time needed to transmit the first frame to its targeted port <b>455</b> in the absence of congestion. In one embodiment, end station <b>450</b> is further configured to track the number of frame delay indications of aggressive timer <b>452</b>, as applied to the first frame. If the frame is delayed due to congestion, aggressive timer <b>452</b> provides at least one frame delay indication. In this embodiment, the tracked number of frame delay indications represents an amount of time during which the frame fails to make forward progress. In this embodiment, the variable timing threshold for identifying a level of congestion that requires a response, is represented by a number of aggressive timer <b>452</b> expirations. The variable timing threshold for a second frame targeting the same port <b>455</b> can be adjusted based on the number.
P-0165[0165] In one embodiment, the variable timing threshold is configurable based on the type of network system architecture. In some embodiments, an end station <b>450</b> transmitting a frame employs predetermined information about at least one end station along the frame's intended transmission path to configure the variable timing threshold for the frame. In one such embodiment, the predetermined information includes the at least one end station's role in overall system performance. For example, if a first frame's routing path includes a switch that is a major network hub and potential bottleneck, the variable timing threshold for the first frame can be set to be more stringent.
P-0166[0166] In another embodiment, the variable timing threshold is configurable based on historical data of the congestion status of the end stations along a frame's routing path. In another embodiment, the variable timing threshold is based on a measured congestion status of an end station along the routing path. In another embodiment, the variable timer threshold is configurable based on a predetermined or measured transmission bandwidth of at least one downstream end station. The bandwidth might be restricted due to the end stations' capacity or congestion status. In an example of such an embodiment, a variable timing threshold of an aggressive timer <b>452</b> used with a port <b>455</b> is a function of the associated link hop speed.
P-0167[0167] In some embodiments, the variable timing threshold is based on a frame attribute. In one such embodiment, the frame upon which the threshold is based is examined to ascertain its relevant attributes. In this embodiment, the end station <b>450</b> includes a hardware or software mechanism for examining a frame's protocol header and/or trailer. In another embodiment, the variable timing threshold is based on the size of a frame. In one form of this embodiment, the aggressive timer <b>452</b> is configured to a limit that is proportional to the size of the frame in the transmit register of port <b>455</b>. A limit that is only slightly greater than the time for transmitting the frame, provides a high sensitivity for detecting transmission delay.
P-0168[0168] In another embodiment, the variable timing threshold is configurable based on the output frame's information type, such as whether the frame carries data or control information. In one such embodiment, the information type is determined using a mechanism for parsing out a message's opcode. Depending on a policy for managing congestion in a distributed computer system, control frames can be given a higher or lower priority of service than data-bearing frames, such that the aggressive timer <b>452</b> is configured to allow higher priority frames more time to make forward progress before a congestion management response is taken with respect to the frames.
P-0169[0169] In some embodiments, the variable timing threshold is dynamically configurable based on a frame's source or destination end station. In one such embodiment, the aggressive timer <b>452</b> is configured based on the frame's final destination end station. In this embodiment, the end station <b>450</b> transmitting the frame examines the frame to determine its final destination, and configures the variable timing threshold of aggressive timer <b>452</b> accordingly. In one example embodiment, a frame addressed to an inherently slower device, such as a disk drive, is subject to a more stringent timing threshold (i.e., a shorter time limit) than a frame addressed to a faster device, such as a video output device. In this embodiment, the transmitting end station <b>450</b> has a mechanism to recognize the final destination end station type with respect to its role on system performance. Therefore, an exemplary transmitting end station, utilizing aggressive timing configured based on final frame destination, can enforce a priority policy that gives service precedence to frames destined for end station types that have more time-critical roles.
P-0170[0170] In one embodiment, the variable timing threshold is based on the source end station of a given frame. Network system policy may have assigned a higher or lower priority for frames originating from particular end nodes. Thus, an exemplary transmitting end station <b>450</b> along the frame's routing path examines the frame header to determine the originating end node, looks up the variable timing threshold for the end node, and applies the variable timing threshold as a time limit for the aggressive timer <b>452</b>.
P-0171[0171] In another embodiment where the variable timing threshold for an aggressive timer is configurable based on one or more frame attributes, the frame being transported is examined for its group identification, which is indicative of the type of communication service of the frame. Example communication service types include, but are not limited to: multicasting, unicasting, and broadcasting. An exemplary embodiment of this type implements one or more policies for congestion management or priority servicing, where the policy dictates various priorities for respective communication services. In one such embodiment, the variable timing threshold is applied frame-by-frame according to each frame's communication service. For example, in a network system where a policy provides higher priority to multicast transmissions, end stations <b>450</b> along a routing path for a given multicast frame, apply less stringent timing thresholds (longer aggressive timer limits) to the multicast frame.
P-0172[0172] In one type of embodiment, the variable timing threshold is configurable based on a frame's assigned service level. In one such embodiment, the end station <b>450</b> examines the frame to be transmitted to determine its service level flag values. Depending on the frame's priority, the variable timing threshold is configured to be more or less stringent. In an exemplary embodiment, an aggressive timing limit for a low-priority frame is set to a lower, more stringent level, to cause the low-priority frame to not be transmitted in favor of preserving transmission bandwidth for higher-priority frames.
P-0173[0173] Upon expiration of the aggressive timer <b>452</b>, an exemplary end station <b>450</b> takes one or more actions in response thereto. In one embodiment, as illustrated in FIG. 13D, the end station <b>450</b> includes a traffic congestion manager <b>460</b> for responding to the aggressive timer <b>452</b>'s indication of its expiration. Traffic congestion manager <b>460</b>'s response is denoted at <b>462</b>; aggressive timer <b>452</b>'s indication is denoted at <b>464</b>. The traffic congestion manager <b>460</b> can be realized in hardware or software. In one configuration, a traffic congestion manager and at least one aggressive timer are realized as a single functional unit <b>466</b>, which performs the functions of both the aggressive timer and traffic congestion manager.
P-0174[0174] In some embodiments, the traffic congestion manager <b>460</b> is configurable to respond to the aggressive timer <b>452</b>'s indication in a number of ways. In one such embodiment, the traffic congestion manager responds to an expiration of the aggressive timer <b>452</b> by dropping at least one frame. In one exemplary embodiment, the traffic congestion manager <b>460</b> drops the frame currently being transmitted. In another exemplary embodiment, the traffic congestion manager <b>460</b> drops frames randomly, or by a weighted random algorithm. In another exemplary embodiment, the traffic congestion manager <b>460</b> drops one or more frames in a port's transmission buffer based on the frames' attributes. For example, traffic congestion manager <b>460</b> drops all frames in the output buffer targeting the same intermediate or final destination end station that the current frame targets when the aggressive timer <b>452</b> measuring the current frame's transmission progress expires. In another example, the traffic congestion manager <b>460</b> clears the entire transmit buffer of the port in response to aggressive timer <b>452</b>'s expiration. In one type of embodiment, the frame dropping is performed over a certain period of time in order to affect future frames as well as the current frame.
P-0175[0175] In an embodiment of a traffic congestion manager <b>460</b>, upon expiration of the aggressive timer <b>452</b>, congestion manager <b>460</b> truncates the current frame by discarding the frame's untransmitted flits. In another embodiment, the traffic congestion manager <b>460</b> responds to an expiration of aggressive timer <b>452</b> by generating one or more reporting frames to be used by one or more end stations or fabric management for congestion management purposes. An exemplary reporting frame includes data characterizing the nature of the aggressive timer's expiration, or circumstances surrounding the aggressive timer's expiration. In one embodiment, the reporting frame has data containing information about the delayed frame during the transmission of which the aggressive timer <b>452</b> expired, such as the frame's size, destination, or service level.
P-0176[0176] In another type of embodiment, the traffic congestion manager <b>460</b> responds to aggressive timer <b>452</b>'s expiration by tagging the current frame with information indicative of the congestion experienced by the frame. In one such embodiment, the current frame and all subsequent frames thereto are tagged for a period of time. Tagging can thus be used as a means for communicating traffic congestion information to other end stations in the network. The other end stations can include fabric management agents, neighboring routing elements, and source endnodes.
P-0177[0177] In another type of embodiment, the traffic congestion manager <b>460</b> responds to the aggressive timer <b>452</b>'s expiration by logging the expiration occurrence and its surrounding circumstances. One such embodiment of traffic congestion manager <b>460</b> maintains a log of aggressive timer <b>452</b> expirations, wherein the log contains information useful for characterizing the end station <b>450</b>'s congestion status over a period time. In one embodiment, this characterization is employed to re-configure at least one aggressive timer <b>452</b>'s variable timing threshold.
P-0178[0178] In one type of embodiment, the traffic congestion manager <b>460</b> is dynamically configurable. In one such embodiment, the response to the aggressive timer <b>452</b> is selectable by network system or fabric management. In other embodiments, the response is dynamically configurable based on predetermined or measured changing circumstances, as described above for the dynamically configurable aggressive timer.
P-0179[0179] In another embodiment, the response <b>462</b> is configurable based on one or more frame delay indications <b>464</b> provided by the aggressive timer <b>452</b>. For example, in an implementation where the aggressive timer is configured to indicate multiple levels of delay (such as with a plurality of frame transmission delay indications), the traffic congestion manager <b>460</b> can take different responses as the frame delay continues. The responses <b>462</b> can progressively increase in their efficacy as a measured delay becomes longer. In an exemplary configuration, the initial response <b>462</b> is a type of event logging; the next response <b>462</b> is communication of congestion status (such as frame tagging or sending congestion management packets); the next response <b>462</b> is a type of selective frame dropping; finally, the most drastic response <b>462</b> to aggressive timer expiration is a flushing of all buffered frames for a period of time.
P-0180[0180]FIG. 14 illustrates an exemplary routing element <b>500</b> interfacing with SAN fabric <b>509</b> through a link indicated generally at <b>584</b>. Routing element <b>500</b> includes ports <b>502</b><i>a, </i><b>502</b><i>b </i>and <b>502</b><i>c. </i>Ports <b>502</b> connect via links <b>584</b><i>a, </i><b>584</b><i>b </i>and <b>584</b><i>c, </i>respectively, with SAN fabric <b>509</b>. Each port <b>502</b> has a receive register <b>504</b> and transmit queue <b>506</b>. Transmit queues <b>506</b> include head-of-line transmit registers <b>508</b>, and back-of-line queue input registers <b>509</b>. Receive registers <b>504</b> and transmit registers <b>508</b> connect to link <b>284</b>.
P-0181[0181] Bus <b>510</b> provides an interface for receive registers <b>504</b> and queue input registers <b>509</b> to exchange frames within the routing element <b>500</b>. Receive registers <b>504</b> interface with a bus <b>510</b> that facilitates communication with the two transmit queues <b>506</b> from the two other ports. Each queue input register <b>509</b> interfaces with the two receive registers <b>504</b> from the two other ports via bus <b>510</b>.
P-0182[0182] To illustrate a switching operation of routing element <b>500</b>, consider an exemplary incoming frame to routing element <b>500</b> arriving to port <b>502</b><i>a, </i>and which is to be transmitted out through port <b>502</b><i>c. </i>The incoming frame arrives to port <b>502</b><i>a</i>'s receive register <b>504</b><i>a </i>via link <b>584</b>. The frame then passes into transmit queue <b>506</b><i>c </i>through queue input register <b>509</b><i>c </i>via bus <b>510</b>. As transmit queue <b>506</b><i>c </i>transmits previous frames out through transmit register <b>508</b><i>c, </i>the frame advances until it reaches transmit register <b>508</b><i>c. </i>Finally, the frame is transmitted from transmit register <b>508</b><i>c </i>via link <b>584</b> into SAN fabric <b>309</b> on its way to its final destination.
P-0183[0183] Exemplary routing element <b>500</b> also includes per-port aggressive timers <b>512</b><i>a, </i><b>512</b><i>b </i>and <b>512</b><i>c, </i>and traffic congestion manager <b>516</b>. Aggressive timers <b>512</b> and traffic congestion manager <b>516</b> are components of a local network traffic congestion management system for implementing traffic congestion management policies. Each aggressive timer <b>512</b> monitors the transmission status <b>514</b> of each corresponding transmission register <b>508</b>. Transmission status <b>514</b> is indicative of frame transmission progress, such as whether flits of the transmitted frame are being sent. In another configuration, transmission status <b>514</b> is simply an indicator of when a new frame enters transmit register <b>508</b>. Timers <b>512</b> also monitor transmission demand <b>515</b> for each corresponding port <b>502</b>. The transmission demand <b>515</b> of each port <b>502</b> is indicative of each port's pending outgoing traffic volume. For example, transmission demand <b>515</b> provides a signal whenever the corresponding queue input register <b>509</b> receives a new frame.
P-0184[0184] Each aggressive timer <b>512</b> measures the presence of any transmission delay of frames in its corresponding transmit register <b>508</b>. The measurement is based on one or more variable timing threshold(s) with which each aggressive timer <b>512</b> is configured. As discussed above, the variable timing thresholds themselves can be based on a variety of attributes, parameters, or attributes relating to the network system, including one or more frames. In the present example, a variable timing threshold is provided to each aggressive timer <b>512</b> by traffic congestion manager <b>516</b> via a timer configuration signal, indicated at <b>522</b>. Timer configuration signal <b>522</b> can be supplied continuously, periodically, or occasionally to dynamically configure the aggressive timers <b>512</b>. Each aggressive timer <b>512</b> can have a unique configuration signal, as indicated at <b>522</b><i>a, </i><b>522</b><i>b </i>and <b>522</b><i>c. </i>
P-0185[0185] If a measured transmission delay exceeds the configured variable timing threshold of an aggressive timer <b>512</b>, the aggressive timer <b>512</b> will provide a timer expiration indication <b>518</b> to traffic congestion manager <b>516</b>. Traffic congestion manager <b>516</b> can then take an appropriate response according to the traffic management policy instituted. Responses include a variety of actions, some of which are described above. One such action is dropping one or more frames from a congested port <b>502</b>. Frames can be dropped selectively, based on one or more frame attributes. In the present example, a clear register signal is illustrated for each port at <b>520</b>.
P-0186[0186]FIG. 15 illustrates an example operation of the aggressive timer <b>512</b> of one port <b>502</b> of routing element <b>500</b>. As indicated at <b>600</b>, aggressive timer <b>512</b> monitors transmission demand <b>515</b> of transmit queue <b>506</b>. As indicated at <b>602</b>, aggressive timer <b>512</b> monitors transmission status <b>514</b>. Under non-congested, or unexceptional operating conditions, frames are either making forward progress or not arriving into transmit queue <b>506</b>. Therefore, under unexceptional circumstances, aggressive timer <b>512</b> remains in a reset state or is in a state receptive to configuration signal <b>522</b>, as indicated at <b>604</b>.
P-0187[0187] If an exception occurs at a time when frames are arriving and when the output frame in transmit register <b>508</b> does not make forward progress, then aggressive timer <b>512</b> starts counting at <b>606</b>. At <b>608</b>, an inquiry and decision is made as to whether aggressive timer <b>512</b> has expired. If aggressive timer <b>512</b> has not expired and the exceptional conditions remain, then aggressive timer <b>512</b> continues applying the configured variable timing threshold, as indicated at <b>606</b>. If the aggressive timer <b>512</b> has not expired but the exceptional circumstances no longer exist, then aggressive timer <b>512</b> is rest at <b>604</b> and the normal non-exceptional operating mode described above resumes. If, however, aggressive timer <b>512</b> expires, then traffic congestion manager <b>516</b> takes an appropriate action and response to the aggressive timer <b>512</b>′ expiration, such as dropping one or more frames from transmit queue <b>506</b> and/or receive register <b>504</b>. Aggressive timer <b>512</b> is then reset and/or reconfigured, as indicated at <b>604</b>.
P-0188[0188]FIG. 16 illustrates an exemplary embodiment of an end station <b>750</b> having an aggressive timer <b>752</b> and a forward progress timer <b>770</b>. End station <b>750</b> transmits or receives network traffic <b>754</b> via links <b>784</b><i>a </i>and <b>784</b><i>b. </i>
P-0189[0189] Aggressive timer <b>752</b> provides timewise monitoring of network traffic <b>754</b> via a monitoring mechanism <b>756</b>, and according to aggressive timer configuration <b>753</b>. Upon expiration, aggressive timer <b>752</b> provides a timer expiration indication <b>764</b>. Traffic congestion manager <b>760</b> receives expiration indication <b>764</b>, and provides a selected response <b>762</b> according to network traffic management policy for aggressively-timed network traffic. Forward progress timer <b>770</b> monitors network traffic <b>754</b> for an occurrence of a severe congestion indication <b>772</b>, and according to preselected timing criteria. Forward progress timer <b>770</b> provides a response action <b>774</b> according to network traffic management policy for forward progress-timed network traffic.
P-0190[0190] Forward progress timer <b>770</b> is generally employed to facilitate a solution for severe, abnormal congestion. Forward progress timer <b>770</b> receives severe congestion indication <b>772</b> when end station <b>750</b> cannot send a frame after a long time, such as for tens or hundreds of milliseconds. In response, forward progress timer <b>770</b> takes one or more actions <b>774</b> for relieving the congestion.
P-0191[0191] Conversely, aggressive timer <b>752</b> is generally employed to facilitate implementation of a congestion management policy by recognizing and responding to early indications of congestion or bandwidth underutilization. Thus, exemplary aggressive timer <b>752</b> is implemented with finer granularity than the forward progress timer <b>770</b>. In one embodiment, the aggressive timer <b>752</b> has a counting rate that is a multiple of forward progress timer <b>770</b>'s counting rate. In one embodiment, aggressive timer <b>752</b> has microsecond resolution.
P-0192[0192] The following Pseudo-Code I demonstrates the timing and congestion management functionality of an exemplary end station having a forward progress timer (FT) and an aggressive timer (AT). <tables id="TABLE-US-00001" num="1"><table frame="none" colsep="0" rowsep="0"><tgroup align="left" colsep="0" rowsep="0" cols="1"><colspec colname="1" colwidth="217PT" align="center" /><thead><row><entry /></row><row><entry namest="1" nameend="1" align="center" rowsep="1" /></row><row><entry>PSEUDO-CODE I</entry></row><row><entry namest="1" nameend="1" align="center" rowsep="1" /></row></thead><tbody valign="top"><row><entry /></row></tbody></tgroup><tgroup align="left" colsep="0" rowsep="0" cols="1"><colspec colname="1" colwidth="217PT" align="left" /><tbody valign="top"><row><entry>while (frames are arriving to buffer) {</entry></row></tbody></tgroup><tgroup align="left" colsep="0" rowsep="0" cols="2"><colspec colname="OFFSET" colwidth="14PT" align="left" /><colspec colname="1" colwidth="203PT" align="left" /><tbody valign="top"><row><entry /><entry>if (output port frame cannot make forward progress) {</entry></row></tbody></tgroup><tgroup align="left" colsep="0" rowsep="0" cols="2"><colspec colname="OFFSET" colwidth="28PT" align="left" /><colspec colname="1" colwidth="189PT" align="left" /><tbody valign="top"><row><entry /><entry>start AT on output port</entry></row><row><entry /><entry>start FT on output port</entry></row></tbody></tgroup><tgroup align="left" colsep="0" rowsep="0" cols="2"><colspec colname="OFFSET" colwidth="14PT" align="left" /><colspec colname="1" colwidth="203PT" align="left" /><tbody valign="top"><row><entry /><entry>} else</entry></row></tbody></tgroup><tgroup align="left" colsep="0" rowsep="0" cols="2"><colspec colname="OFFSET" colwidth="28PT" align="left" /><colspec colname="1" colwidth="189PT" align="left" /><tbody valign="top"><row><entry /><entry>reset AT and FT</entry></row></tbody></tgroup><tgroup align="left" colsep="0" rowsep="0" cols="1"><colspec colname="1" colwidth="217PT" align="left" /><tbody valign="top"><row><entry>}</entry></row><row><entry>Exception Processing:</entry></row></tbody></tgroup><tgroup align="left" colsep="0" rowsep="0" cols="2"><colspec colname="OFFSET" colwidth="14PT" align="left" /><colspec colname="1" colwidth="203PT" align="left" /><tbody valign="top"><row><entry /><entry>if (AT expires and FT has not) {</entry></row></tbody></tgroup><tgroup align="left" colsep="0" rowsep="0" cols="2"><colspec colname="OFFSET" colwidth="28PT" align="left" /><colspec colname="1" colwidth="189PT" align="left" /><tbody valign="top"><row><entry /><entry>reset AT</entry></row><row><entry /><entry>respond according to congestion management policy</entry></row></tbody></tgroup><tgroup align="left" colsep="0" rowsep="0" cols="2"><colspec colname="OFFSET" colwidth="14PT" align="left" /><colspec colname="1" colwidth="203PT" align="left" /><tbody valign="top"><row><entry /><entry>} else if (FT expires) {</entry></row></tbody></tgroup><tgroup align="left" colsep="0" rowsep="0" cols="2"><colspec colname="OFFSET" colwidth="28PT" align="left" /><colspec colname="1" colwidth="189PT" align="left" /><tbody valign="top"><row><entry /><entry>reset AT and FT</entry></row><row><entry /><entry>discard frames for a period of time defined by architecture</entry></row></tbody></tgroup><tgroup align="left" colsep="0" rowsep="0" cols="2"><colspec colname="OFFSET" colwidth="14PT" align="left" /><colspec colname="1" colwidth="203PT" align="left" /><tbody valign="top"><row><entry /><entry>}</entry></row><row><entry /><entry>Update link-level flow control as required</entry></row><row><entry /><entry namest="OFFSET" nameend="1" align="center" rowsep="1" /></row></tbody></tgroup></table></tables>
P-0193[0193] Although specific embodiments have been illustrated and described herein for purposes of description of the preferred embodiment, it will be appreciated by those of ordinary skill in the art that a wide variety of alternate and/or equivalent implementations calculated to achieve the same purposes may be substituted for the specific embodiments shown and described without departing from the scope of the present invention. Those with skill in the chemical, mechanical, electro-mechanical, electrical, and computer arts will readily appreciate that the present invention may be implemented in a very wide variety of embodiments. This application is intended to cover any combinations, adaptations or variations of the preferred embodiments discussed herein. Unless otherwise described, no single embodiment is exclusive of any other described embodiment. Therefore, it is manifestly intended that this invention be limited only by the claims and the equivalents thereof.
Contents6
19 sheets
Sheet 1 Sheet 2 Sheet 3 Sheet 4 Sheet 5 Sheet 6 Sheet 7 Sheet 8 Sheet 9 Sheet 10 Sheet 11 Sheet 12 Sheet 13 Sheet 14 Sheet 15 Sheet 16 Sheet 17 Sheet 18 Sheet 19
Every citation, both ways
| Document | Relation | Office | Cited during |
|---|---|---|---|
| US8903968B2 | Cited by | United States of America | Search report |
| US8364849B2 | Cited by | United States of America | Applicant |
| US2008059554A1 | Cited by | United States of America | Pre-grant |
| US8885467B2 | Cited by | United States of America | Applicant |
| US10630590B2 | Cited by | United States of America | Search report |
| US2006098681A1 | Cited by | United States of America | Pre-grant |
| US8185896B2 | Cited by | United States of America | Applicant |
| US7769891B2 | Cited by | United States of America | Applicant |
| US7793158B2 | Cited by | United States of America | Search report |
| US2016261339A1 | Cited by | United States of America | Pre-grant |
| US2011264758A1 | Cited by | United States of America | Pre-grant |
| US2013135992A1 | Cited by | United States of America | Pre-grant |
| US2022191767A1 | Cited by | United States of America | Search report |
| US9674092B2 | Cited by | United States of America | Applicant |
| US7480298B2 | Cited by | United States of America | Applicant |
| US7904590B2 | Cited by | United States of America | Applicant |
| US2003074469A1 | Cited by | United States of America | Pre-grant |
| US7769892B2 | Cited by | United States of America | Applicant |
| US10437492B1 | Cited by | United States of America | Search report |
| CN107005484A | Cited by | China | Search report |
| US7478138B2 | Cited by | United States of America | Search report |
| US8140731B2 | Cited by | United States of America | Applicant |
| US7958182B2 | Cited by | United States of America | Applicant |
| US8055818B2 | Cited by | United States of America | Search report |
| US2023236992A1 | Cited by | United States of America | Pre-grant |
| US2002059451A1 | Cited by | United States of America | Pre-grant |
| US2017339071A1 | Cited by | United States of America | Search report |
| US9391899B2 | Cited by | United States of America | Search report |
| US2018019947A1 | Cited by | United States of America | Search report |
| US2009063815A1 | Cited by | United States of America | Pre-grant |
| US2008104659A1 | Cited by | United States of America | Pre-grant |
| US8037218B2 | Cited by | United States of America | Applicant |
| US7958183B2 | Cited by | United States of America | Applicant |
| US7760752B2 | Cited by | United States of America | Search report |
| US7564869B2 | Cited by | United States of America | Applicant |
| US9444758B2 | Cited by | United States of America | Applicant |
| US7813369B2 | Cited by | United States of America | Applicant |
| US7830793B2 | Cited by | United States of America | Applicant |
| US9419912B2 | Cited by | United States of America | Applicant |
| US9674091B2 | Cited by | United States of America | Applicant |
| US2016056886A1 | Cited by | United States of America | Pre-grant |
| US7827428B2 | Cited by | United States of America | Applicant |
| US7984196B2 | Cited by | United States of America | Search report |
| US7406092B2 | Cited by | United States of America | Search report |
| US7953085B2 | Cited by | United States of America | Applicant |
| US12218855B2 | Cited by | United States of America | Search report |
| US8259720B2 | Cited by | United States of America | Applicant |
| US9307053B2 | Cited by | United States of America | Search report |
| US2006242304A1 | Cited by | United States of America | Pre-grant |
| WO2006057730A2 | Cited by | World Intellectual Property Organization (WIPO) | International search |
| US7801125B2 | Cited by | United States of America | Applicant |
| US8612536B2 | Cited by | United States of America | Search report |
| US7921316B2 | Cited by | United States of America | Applicant |
| US10200116B2 | Cited by | United States of America | Search report |
| US8023417B2 | Cited by | United States of America | Applicant |
| US10320929B1 | Cited by | United States of America | Applicant |
| US2009063816A1 | Cited by | United States of America | Pre-grant |
| US2007248111A1 | Cited by | United States of America | Pre-grant |
| US11617124B2 | Cited by | United States of America | Search report |
| US7023811B2 | Cited by | United States of America | Search report |
| US2004184481A1 | Cited by | United States of America | Pre-grant |
| US2008091868A1 | Cited by | United States of America | Pre-grant |
| US9729467B2 | Cited by | United States of America | Search report |
| WO2016105446A1 | Cited by | World Intellectual Property Organization (WIPO) | International search |
| US2006098681A1 | Cited by | United States of America | Pre-grant |
| US2010293275A1 | Cited by | United States of America | Pre-grant |
| US7969971B2 | Cited by | United States of America | Search report |
| US8532099B2 | Cited by | United States of America | Applicant |
| US10757039B2 | Cited by | United States of America | Search report |
| US7602720B2 | Cited by | United States of America | Applicant |
| US7522597B2 | Cited by | United States of America | Applicant |
| US7779148B2 | Cited by | United States of America | Applicant |
| US2009063814A1 | Cited by | United States of America | Pre-grant |
| US8014387B2 | Cited by | United States of America | Applicant |
| WO2006057730A3 | Cited by | World Intellectual Property Organization (WIPO) | International search |
| US2012230212A1 | Cited by | United States of America | Pre-grant |
| US2009063811A1 | Cited by | United States of America | Pre-grant |
| US7430615B2 | Cited by | United States of America | Applicant |
| US7822889B2 | Cited by | United States of America | Applicant |
| US8108545B2 | Cited by | United States of America | Applicant |
| US2015139024A1 | Cited by | United States of America | Pre-grant |
| US2017339071A1 | Cited by | United States of America | Search report |
| US2006171318A1 | Cited by | United States of America | Pre-grant |
| US8238347B2 | Cited by | United States of America | Applicant |
| US9973270B2 | Cited by | United States of America | Search report |
| US2007294393A1 | Cited by | United States of America | Pre-grant |
| US2009070617A1 | Cited by | United States of America | Pre-grant |
| US2006047867A1 | Cited by | United States of America | Pre-grant |
| US8953442B2 | Cited by | United States of America | Search report |
| US2009198958A1 | Cited by | United States of America | Pre-grant |
| US7809970B2 | Cited by | United States of America | Applicant |
| US8761020B1 | Cited by | United States of America | Search report |
| US2007041383A1 | Cited by | United States of America | Pre-grant |
| US7346702B2 | Cited by | United States of America | Search report |
| US2009064139A1 | Cited by | United States of America | Pre-grant |
| US2014317165A1 | Cited by | United States of America | Pre-grant |
| US8982688B2 | Cited by | United States of America | Applicant |
| US9491101B2 | Cited by | United States of America | Applicant |
| US8077602B2 | Cited by | United States of America | Applicant |
| US11297006B1 | Cited by | United States of America | Search report |
29 members in 4 offices; this record represents the family
Priority claims4
| Document | Office | Kind | Date |
|---|---|---|---|
| 13566499 | United States of America | P | |
| 15415099 | United States of America | P | |
| 57815500 | United States of America | A | |
| 57801900 | United States of America | A |
Members29
| Document | Office | Kind | |
|---|---|---|---|
| WO0072142A1 | World Intellectual Property Organization (WIPO) | A1 | |
| WO0072158A1 | World Intellectual Property Organization (WIPO) | A1 | |
| WO0072159A1 | World Intellectual Property Organization (WIPO) | A1 | |
| WO0072169A1 | World Intellectual Property Organization (WIPO) | A1 | |
| WO0072170A1 | World Intellectual Property Organization (WIPO) | A1 | |
| WO0072421A1 | World Intellectual Property Organization (WIPO) | A1 | |
| WO0072487A1 | World Intellectual Property Organization (WIPO) | A1 | |
| WO0072575A2 | World Intellectual Property Organization (WIPO) | A2 | |
| AU5041900A | Australia | A | |
| AU5159900A | Australia | A | |
| AU5284500A | Australia | A | |
| AU5285200A | Australia | A | |
| AU5286400A | Australia | A | |
| AU5288800A | Australia | A | |
| AU5289000A | Australia | A | |
| AU5292100A | Australia | A | |
| WO0072575A3 | World Intellectual Property Organization (WIPO) | A3 | |
| US2002133620A1 | United States of America | A1 | |
| JP2003500962A | Japan | A | |
| US2003195983A1 | United States of America | A1 | |
| US6769021B1 | United States of America | B1 | |
| US7010607B1 | United States of America | B1 | |
| US7016971B1 | United States of America | B1 | |
| US7103626B1 | United States of America | B1 | |
| US7171484B1 | United States of America | B1 | |
| US7318102B1 | United States of America | B1 | |
| US7346699B1 | United States of America | B1 | |
| US2008177890A1 | United States of America | A1 | |
| US7904576B2 | United States of America | B2 |
96 transactions on the USPTO file
Abandoned after 3 non-final rejections, 2 final rejections and 2 appeals.
- Non-final rejections
- 3
- Final rejections
- 2
- RCEs
- 0
- Appeals
- 2
Over time
Point at a mark for the transactionTransactions
| Event | Code | |
|---|---|---|
| Email NotificationEML_NTR | EML_NTR | |
| Mailing of Abandonment after Board of AppealsAbandonedMABN10 | MABN10 | |
| Abandonment after Board of AppealsAbandonedABN10 | ABN10 | |
| Electronic ReviewELC_RVW | ELC_RVW | |
| Email NotificationEML_NTF | EML_NTF | |
| Mail BPAI Decision on Appeal - AffirmedMAPDA | MAPDA | |
| BPAI Decision - Examiner AffirmedAPDA | APDA | |
| Case Docketed to Examiner in GAUDOCK | DOCK | |
| Email NotificationEML_NTR | EML_NTR | |
| Docketing Notice Mailed to AppellantAP_DK_M | AP_DK_M | |
| Assignment of Appeal NumberAPAS | APAS | |
| Email NotificationEML_NTR | EML_NTR | |
| Mail Reply Brief Noted by ExaminerMRBNE | MRBNE | |
| Appeal Awaiting BPAI DocketingAPWD | APWD | |
| Reply Brief Noted by ExaminerRBNE | RBNE | |
| Reply Brief FiledAPRB | APRB | |
| Date Forwarded to ExaminerFWDX | FWDX | |
| Exam. Ans. Review CompletePACC | PACC | |
| Reply Brief FiledAPRB | APRB | |
| Rejection- New GroundsRJ.NG | RJ.NG | |
| Electronic ReviewELC_RVW | ELC_RVW | |
| Email NotificationEML_NTF | EML_NTF | |
| Mail Supplemental Examiner's AnswerMAPE2 | MAPE2 | |
| 2nd or Subsequent Examiner's Answer to Appeal BriefAPE2 | APE2 | |
| Return of Undocketed appeal to the TCTCRD | TCRD | |
| Exam. Ans. Review CompletePACC | PACC | |
| Rejection- New GroundsRJ.NG | RJ.NG | |
| Electronic ReviewELC_RVW | ELC_RVW | |
| Email NotificationEML_NTF | EML_NTF | |
| Mail Examiner's AnswerMAPEA | MAPEA | |
| Examiner's Answer to Appeal BriefAPEA | APEA | |
| Appeal Brief Review CompleteAPBR | APBR | |
| Date Forwarded to ExaminerFWDX | FWDX | |
| Appeal Brief FiledAP.B | AP.B | |
| Notice of Appeal FiledN/AP | N/AP | |
| Electronic ReviewELC_RVW | ELC_RVW | |
| Email NotificationEML_NTF | EML_NTF | |
| Mail Final Rejection (PTOL - 326)Final rejectionMCTFR | MCTFR | |
| Final RejectionFinal rejectionCTFR | CTFR | |
| Date Forwarded to ExaminerFWDX | FWDX | |
| Response after Non-Final ActionA... | A... | |
| Electronic ReviewELC_RVW | ELC_RVW | |
| Email NotificationEML_NTF | EML_NTF | |
| Mail Non-Final RejectionNon-final rejectionMCTNF | MCTNF | |
| Non-Final RejectionNon-final rejectionCTNF | CTNF | |
| Date Forwarded to ExaminerFWDX | FWDX | |
| Response after Non-Final ActionA... | A... | |
| Case Docketed to Examiner in GAUDOCK | DOCK | |
| Electronic ReviewELC_RVW | ELC_RVW | |
| Email NotificationEML_NTF | EML_NTF | |
| Mail Non-Final RejectionNon-final rejectionMCTNF | MCTNF | |
| Non-Final RejectionNon-final rejectionCTNF | CTNF | |
| Appeal Brief Review CompleteAPBR | APBR | |
| Date Forwarded to ExaminerFWDX | FWDX | |
| Appeal Brief FiledAP.B | AP.B | |
| Email NotificationEML_NTR | EML_NTR | |
| Notice -- Defective Appeal BriefAPBD | APBD | |
| Appeal Brief Review CompleteAPBR | APBR | |
| Date Forwarded to ExaminerFWDX | FWDX | |
| Defective / Incomplete Appeal Brief FiledAPBI | APBI | |
| Appeal Brief FiledAP.B | AP.B | |
| Email NotificationEML_NTR | EML_NTR | |
| Notice -- Defective Appeal BriefAPBD | APBD | |
| Appeal Brief Review CompleteAPBR | APBR | |
| Date Forwarded to ExaminerFWDX | FWDX | |
| Defective / Incomplete Appeal Brief FiledAPBI | APBI | |
| Appeal Brief FiledAP.B | AP.B | |
| Notice of Appeal FiledN/AP | N/AP | |
| Mail Final Rejection (PTOL - 326)Final rejectionMCTFR | MCTFR | |
| Final RejectionFinal rejectionCTFR | CTFR | |
| Date Forwarded to ExaminerFWDX | FWDX | |
| Response after Non-Final ActionA... | A... | |
| Mail Non-Final RejectionNon-final rejectionMCTNF | MCTNF | |
| Non-Final RejectionNon-final rejectionCTNF | CTNF | |
| Reference capture on IDSRCAP | RCAP | |
| Information Disclosure Statement (IDS) FiledM844 | M844 | |
| Information Disclosure Statement (IDS) FiledWIDS | WIDS | |
| Miscellaneous Incoming LetterLET. | LET. | |
| IFW TSS Processing by Tech Center CompleteTSSCOMP | TSSCOMP | |
| Information Disclosure Statement consideredIDSC | IDSC | |
| Reference capture on IDSRCAP | RCAP | |
| Information Disclosure Statement (IDS) FiledM844 | M844 | |
| Information Disclosure Statement (IDS) FiledWIDS | WIDS | |
| Information Disclosure Statement consideredIDSC | IDSC | |
| Information Disclosure Statement (IDS) FiledM844 | M844 | |
| Information Disclosure Statement (IDS) FiledWIDS | WIDS | |
| Information Disclosure Statement consideredIDSC | IDSC | |
| Information Disclosure Statement (IDS) FiledM844 | M844 | |
| Information Disclosure Statement (IDS) FiledWIDS | WIDS | |
| Workflow incoming amendment IFWWAMD | WAMD | |
| Correspondence Address ChangeC.ADB | C.ADB | |
| Case Docketed to Examiner in GAUDOCK | DOCK | |
| Application Dispatched from OIPEOIPE | OIPE | |
| Application Is Now CompleteCOMP | COMP | |
| IFW Scan & PACR Auto Security ReviewSCAN | SCAN | |
| Initial Exam Team nnIEXX | IEXX |
2 legal events, as the office reported them to INPADOC
Over the term
Point at a mark for the eventEvents
| Event | Code | |
|---|---|---|
| Information on status: application discontinuationABANDONED -- AFTER EXAMINER'S ANSWER OR BOARD OF APPEALS DECISIONSTCB | STCB | |
| AssignmentAS | AS |
Numbers
- Application
- 44240103
Titles
- English
- Network congestion management using aggressive timers
Classification
- CPC, 5
- H04L47/12
- H04L47/26
- H04L47/28
- H04L49/356
- H04L49/505
- IPC, 1
- G06F15 173