Pipelined configurable processor
31 claims: 24 independent, 7 dependent
- 1複数のスレッドを同時に取り扱うことが可能な構成可能処理回路であって、 スレッドデータストアと、 複数の構成可能実行ユニットと、 前記スレッドデータストアを前記実行ユニットに接続する構成可能ルーティングネットワークと、 構成インスタンスを記憶し、該構成インスタンスのそれぞれがルーティングネットワークの構成及び前記複数の実行ユニットのうち1つ以上の構成を規定している、構成データストアと、 前記実行ユニット、前記ルーティングネットワーク及び前記スレッドデータストアから形成されるとともに複数のパイプラインセクションを備え、該複数のパイプラインセクションが、各クロックサイクルにおいて各スレッドが一つのパイプラインセクションから次のパイプラインセクションに伝播するように構成されている、パイプラインとを備え、 (i)各スレッドを構成インスタンスと関連付け、 (ii)各クロックサイクルにおいて、前記複数のパイプラインセクションのそれぞれを、そのクロックサイクル中にそのパイプラインセクション内を伝播する各スレッドと関連付けられた構成インスタンスに従うように構成し、 (iii)前記 スレッド データストア内のどの位置から読み出すか、及び、前記 スレッド データストア内のどの位置に前記実行ユニットが書き込むかを、前記構成インスタンスに基づき選択するように構成された回路。
- 2各構成インスタンスを構成識別子と関連付けるように構成された、請求項1に記載の構成可能処理回路。
- 3スレッドと関連付けられた構成識別子を、該スレッドと協調して前記パイプライン内を伝播させるように構成された、請求項2に記載の構成可能処理回路。
- 4前記構成データストアが、複数のメモリを備え、 前記構成インスタンスを前記複数のメモリに亘って分割して、各メモリが特定のパイプラインセクションに適用可能な前記構成インスタンスの部分を記憶するように構成された、請求項1~3のいずれか一項に記載の構成可能処理回路。
- 5各パイプラインセクションが、それに適用可能な前記構成インスタンスの部分を記憶する前記メモリにアクセスすることにより構成インスタンスにアクセスするように構成されている、請求項4に記載の構成可能処理回路。
- 6前記パイプラインの各セクションが、スレッドと関連付けられた前記構成識別子を使用して、前記構成データストア内のそのスレッドと関連付けられた前記構成インスタンスにアクセスするように構成されている、請求項2、3、請求項2に従属する請求項4又は5のいずれか一項に記載の構成可能処理回路。
- 7前記複数のスレッドが独立している、請求項1~6のいずれか一項に記載の構成可能処理回路。
- 82つ以上のスレッドを同一の構成識別子と関連付けるように構成された、請求項1~7のいずれか一項に記載の構成可能処理回路。
- 9前記構成可能ルーティングネットワークが、複数のネットワーク入力及び複数のネットワーク出力を備え、各ネットワーク入力をネットワーク出力に接続するように構成可能である、請求項1~ 8 のいずれか一項に記載の構成可能処理回路。
- 10前記構成可能ルーティングネットワークが、任意のネットワーク入力を任意のネットワーク出力に接続することが可能である、請求項 9 に記載の構成可能処理回路。
- 11前記構成可能ルーティングネットワークが、任意のネットワーク入力を前記ネットワーク出力のうちの任意の1つ以上に接続することが可能である、請求項 9 又は1 0 に記載の構成可能処理回路。
- 12前記構成可能ルーティングネットワークの出力が、前記実行ユニットの入力に接続されている、請求項1~1 1 のいずれか一項に記載の構成可能処理回路。
- 13前記構成可能ルーティングネットワークが、マルチステージスイッチを備える、請求項1~1 2 のいずれか一項に記載の構成可能処理回路。
- 14前記マルチステージスイッチが、各ステージに1つ以上のスイッチを備え、各スイッチが、複数のスイッチ入力及び複数のスイッチ出力を有し、各スイッチ入力をスイッチ出力に接続するように構成可能である、請求項1 3 に記載の構成可能処理回路。
- 15前記マルチステージスイッチの各ステージにおけるスイッチが、同じ数のスイッチ入力及びスイッチ出力を備える、請求項1 4 に記載の構成可能処理回路。
- 16前記マルチステージスイッチの1つのステージに備えられた前記スイッチが、他のステージに備えられた前記スイッチとは異なる数のスイッチ入力及びスイッチ出力を備える、請求項1 4 に記載の構成可能処理回路。
- 171つのパイプラインセクションが、前記マルチステージスイッチの1つ以上のステージに備えられた前記スイッチから形成されている、請求項1 3 ~1 6 のいずれか一項に記載の構成可能処理回路。
- 18前記マルチステージスイッチの内側ステージにおけるスイッチから形成されたパイプラインセクションが、該マルチステージスイッチにおける、前記マルチステージスイッチの外側ステージに備えられたスイッチから形成されたパイプラインセクションとは異なる数のステージからのスイッチを備える、請求項1 7 に記載の構成可能処理回路。
- 19前記構成可能ルーティングネットワークが、Closネットワークを備える、請求項1~ 18 のいずれか一項に記載の構成可能処理回路。
- 20前記構成可能ルーティングネットワークが、1つ以上のクロスバースイッチを備える、請求項1~ 19 のいずれか一項に記載の構成可能処理回路。
- 21前記構成可能ルーティングネットワークが、非ブロッキングである、請求項1~2 0 のいずれか一項に記載の構成可能処理回路。
- 22前記構成可能ルーティングネットワークが、完全に構成可能である、請求項1~2 1 のいずれか一項に記載の構成可能処理回路。
- 23前記構成可能ルーティングネットワークが、部分的に構成可能である、請求項1~2 1 のいずれか一項に記載の構成可能処理回路。
- 24各実行ユニットのために専用のオンチップメモリを備える、請求項1~2 3 のいずれか一項に記載の構成可能処理回路。
- 25前記スレッドデータストア内に記憶されたデータが有効であることをチェックするチェックユニットを備える、請求項1~2 4 のいずれか一項に記載の構成可能処理回路。
- 26前記スレッドデータストア内の位置が、2つの有効ビットと関連付けられている、請求項1~2 5 のいずれか一項に記載の構成可能処理回路。
- 27前記構成可能ルーティングネットワークが、前記スレッドデータストアから読み出されたデータを運ぶためのマルチプルビットワイドであるデータ経路を備える、請求項1~ 26 のいずれか一項に記載の構成可能処理回路。
- 282つの構成可能ルーティングネットワークを備え、前記構成可能ルーティングネットワークの一方が、他方よりも広いデータ経路を備える、請求項1~ 27 のいずれか一項に記載の構成可能処理回路。
- 29フラクチャブル実行ユニットを備える、請求項1~ 28 のいずれか一項に記載の構成可能処理回路。
- 30動的再構成が可能な、請求項1~ 29 のいずれか一項に記載の構成可能処理回路。
- 31スレッドデータストアと、複数の構成可能実行ユニットと、前記スレッドデータストアを前記実行ユニットに接続する構成可能ルーティングネットワークと、前記実行ユニット、前記ルーティングネットワーク及び前記スレッドデータストアから形成され、複数のパイプラインセクションを備えるパイプラインとを備える構成可能処理回路において、複数のスレッドを同時に取り扱う方法であって、 各スレッドを、前記ルーティングネットワークの構成と、前記複数の実行ユニットのうちの1つ以上の構成と、前記データストア内のどの位置から読み出すか、及び、前記データストア内のどの位置に前記実行ユニットが書き込むかと、を規定する構成インスタンスと関連付けることと、 各クロックサイクルで、各スレッドを、一つのパイプラインセクションから次のパイプラインセクションに伝播させることと、 各クロックサイクルにおいて、前記複数のパイプラインセクションのそれぞれを、そのクロックサイクル中にそのパイプラインセクション内を伝播する各スレッドと関連付けられた前記構成インスタンスに応じるように構成することとを含む方法。
Independent claims31
13 paragraphs, as filed
The present invention relates to processor design for integrated circuits.
Integrated circuits typically include a number of functional units connected to each other by interconnect circuits. In some cases, functional units and interconnect circuits can be configured. This means that the functional unit can be programmed to adopt a particular behavior and the interconnect circuit can be programmed to connect different parts of the circuit. A well-known example of a configurable circuit is an FPGA (Field Programmable Gate Array), which is user-programmable to perform a wide range of different functions. Other examples of configurable integrated circuits are described in Patent Documents 1, 2 and 3.
In many configurable circuits, there is a trade-off between speed and flexibility. For maximum flexibility, it is desirable to be able to connect as many different combinations of functional units as possible to each other. This requires a long interconnect path if the execution units are separated on the chip. In general, integrated circuits cannot be clocked faster than the longest operation that can be done in a single clock cycle. In many cases, the delay caused by the interconnect has a greater effect than the delay caused by the functional unit, so the time it takes to transmit data over a long interconnect path ultimately limits the clock speed of the entire circuit. It becomes a possible constraint.
One option for setting an upper limit on time delays in integrated circuits is to limit the length of all interconnect paths across a single clock cycle. This can be achieved by pipeline processing the data as it travels through the integrated circuit. An example is described in Patent Document 4, in which an input to a switch cell in an interconnect network has a latch that pipelines the data as it is sent through the interconnect network. The problem with this approach is that the user's design may need to be modified to incorporate the required latches.
<p><patcit num="1"><text>U.S. Pat. No. 7,276,933</text></patcit><patcit num="2"><text>U.S. Pat. No. 8,493,090</text></patcit><patcit num="3"><text>U.S. Pat. No. 6,282,627</text></patcit><patcit num="4"><text>U.S. Pat. No. 6,940,308</text></patcit></p>
<p> Therefore, there is a demand for improved and flexible processing circuits.</p>
<p> According to one embodiment, it is a configurable processing circuit that can handle a plurality of threads at the same time, and connects a thread data store, a plurality of configurable execution units, and a position in the thread data store to the execution unit. A configurable routing network and a configuration datastore, an execution unit, and a routing network that store the configuration instances and each of the configuration instances defines the configuration of the routing network and the configuration of one or more of the execution units. And is formed from thread datastores and has multiple pipeline sections, which are configured so that each thread propagates from one pipeline section to the next in each clock cycle. Has a pipeline, (i) associates each thread with a configuration instance, and (ii) for each clock cycle, each of multiple pipeline sections, within that pipeline section during that clock cycle. Circuits are provided that are configured to follow the configuration instance associated with each propagating thread.</p><p> The circuit may be configured to associate each configuration instance with a configuration identifier.</p><p> The circuit may be configured to propagate the configuration identifier associated with a thread in the pipeline in cooperation with that thread.</p><p> The configuration data store may have multiple memories so that the circuit divides the configuration instances across multiple memories so that each memory stores the portion of the configuration instance that is applicable to a particular pipeline section. It may be configured in.</p><p> Each pipeline section may be configured to access a configuration instance by accessing memory that stores a portion of the configuration instance applicable to it.</p><p> Each section of the pipeline may be configured to use the configuration identifier associated with a thread to access the configuration instance associated with that thread in the configuration data store.</p><p> Multiple threads may be independent.</p><p> The circuit may be configured to associate two or more threads with the same configuration identifier.</p><p> The circuit may be able to change the configuration identifier associated with the thread so that the thread can follow a configuration on one path in the circuit that is different from the second subsequent path in the circuit. ..</p><p> The circuit may be configured to change the configuration identifier based on the output produced by one of the execution units as it operates on the input associated with the thread.</p><p> A configurable routing network may include multiple network inputs and multiple network outputs, and may be configurable to connect each network input to a network output.</p><p> The configurable routing network may be capable of connecting any network input to any network output.</p><p> A configurable routing network may be capable of connecting any network input to any one or more of the network outputs.</p><p> The output of the configurable routing network may be connected to the input of the execution unit.</p><p> The configurable routing network may include multistage switches.</p><p> A multi-stage switch may include one or more switches in each stage, each switch having multiple switch inputs and multiple switch outputs, which can be configured to connect each switch input to the switch output. There may be.</p><p> The switches in each stage of the multistage switch may have the same number of switch inputs and outputs.</p><p> A switch on one stage of a multistage switch may have a different number of switch inputs and outputs than switches on the other stage.</p><p> The pipeline section may be formed from switches on one or more stages of a multistage switch.</p><p> The pipeline section formed from the switches on the inner stage of the multistage switch is a switch from a different number of stages than the pipeline section formed from the switches on the outer stage of the multistage switch in the multistage switch. May be provided.</p><p> The configurable routing network may include a Clos network.</p><p> The configurable routing network may include one or more crossbar switches.</p><p> The configurable routing network may be non-blocking.</p><p> The configurable routing network may be fully configurable.</p><p> The configurable routing network may be partially configurable.</p><p> The circuit may include dedicated on-chip memory for each execution unit.</p><p> The circuit may include a check unit that checks that the data stored in the thread data store is valid.</p><p> The check unit may be configured to temporarily stop the execution unit from writing to the thread data store when it recognizes invalid data, and / or acts on the thread from which they read the invalid data. You may perform a memory access operation while you are.</p><p> The circuit may be configured to associate the thread reading the invalid data with the same state on the next path in the circuit.</p><p> A position in the thread data store may be associated with two significant bits.</p><p> The configurable routing network may include a data path that is multiple bit wide to carry the data read from the threaded data store.</p><p> The circuit may include two configurable routing networks, one of which may have a wider data path than the other.</p><p> The circuit may include a fractable execution unit.</p><p> The circuit may include an execution unit configured with interchangeable inputs. A configurable routing network may be configured to connect a thread data store to exchangeable inputs for execution units and non-exchangeable inputs for execution units, with the outermost stage of the configurable routing network. , A first number of switches configured to connect the thread data store to the execute unit to the exchangeable input, and a second number of switches configured to connect the thread data store to the execute unit to the non-exchangeable input. A number of switches may be provided, and the first number may be less than the second number per connected input.</p><p> The circuit may be dynamically reconfigurable.</p><p> According to a second embodiment of the present invention, it is formed from a thread data store, a plurality of configurable execution units, a configurable routing network that connects the thread data store to the execution unit, and an execution unit, a routing network, and a thread data store. A method of handling multiple threads simultaneously in a configurable processing circuit with a pipeline having multiple pipeline sections, each thread being one or more of a routing network configuration and multiple execution units. In each clock cycle, each thread propagates from one pipeline section to the next, and in each clock cycle, each of the multiple pipeline sections. Is provided, including configuring it to accommodate the configuration instance associated with each thread propagating within its pipeline section during its clock cycle.</p><p> Hereinafter, the present invention will be described with reference to the accompanying drawings by way of examples.</p>
<figref num="1">An example of a configurable processing circuit is shown.</figref><figref num="2">An example of a routing network is shown.</figref><figref num="3">An example of a crossbar switch is shown.</figref><figref num="4">An example of the execution unit is shown.</figref><figref num="5">An example of an execution unit configured as an adder is shown.</figref><figref num="6">An example of an execution unit configured as a pipelined ALU is shown.</figref><figref num="7">An example of a long latency execution unit is shown.</figref><figref num="8">An example of an execution unit for setting a configuration instance identifier for a thread is shown.</figref><figref num="9">An example of a fractable execution unit is shown.</figref><figref num="10">Two examples of optimized lookup tables are shown.</figref>
The configurable processing circuit can preferably handle a plurality of threads at the same time. The circuit comprises a thread data store, one or more configurable routing networks, and several configurable execution units. Values from the data store are read and sent to the execution unit over the routing network. The execution unit operates on these values and delivers new values from their output. The output of the execution unit is written back to the data store.
The circuit also includes a pipeline. The pipeline consists of data stores, routing networks and execution units. It has multiple pipeline sections, with each thread propagating from one pipeline section to the next in each clock cycle. The circuits are preferably arranged so that each clock cycle is adapted to the threads that the pipeline sections are currently handling. A thread configuration can be thought of as "clocking" the circuit "clocking" in that thread so that the data in each thread is directed to its own particular path within the processing circuit.
The circuit also includes on-chip memory that holds multiple configuration instances. The circuit selects from which position in the data store to read from and to which position in the data store the execution unit writes, based on the configuration instance. In addition, the circuit is configured to set a route obtained via a routing network and control the behavior of an execution unit using a configuration instance. Each configuration instance can be independently referenced by the configuration instance identifier. The circuit may be configured to select which configuration instance to use for that thread by associating a thread with a particular configuration instance identifier.
With the advent of GPUs (graphics processing units), programmers have become accustomed to solving computational problems using a large number of threads that interact less with each other. These generally independent threads are ideally suitable for processing by the multithreaded reconfigurable processors described herein. GPUs are often built from multiple identical processors and are referred to as homogeneous computing. Unlike GPUs, the circuits described herein allow for multiple different execution units and are a form of heterogeneous computing. The number and capabilities of execution units in a particular instance of a circuit can be selected to suit a particular class of problem. This makes it possible to implement any given task more efficiently than with a GPU.
<p><u style="single">Circuit overview</u> Figure 1 shows an example of a configurable processing circuit. This circuit comprises a configurable routing network (implemented as two routing networks 111,112 in this example). The circuit also includes several configurable execution units (115,116). The circuit is pipelined and is represented by the dotted line 102 in the figure. In the illustrated example, the pipeline consists of eight stages, as indicated by the numbers along the bottom of the figure. The boundaries between pipeline sections are appropriately selected to limit the maximum time required in any pipeline section to accommodate the maximum clock speed.</p><p> The following description assumes that it is the rising clock edge that triggers thread propagation through the pipeline. It should be understood that this is for illustration purposes only and that falling clock edges may be used as well. Similarly, a mixture of rising and falling edges may be used throughout the pipeline. Each pipeline stage may have its own clock (although those clocks are synchronized so that the clock edges occur simultaneously in all pipeline stages).</p><p> The circuit is configured to handle multiple threads at the same time. A thread in hardware is generally considered to be a set of actions that are executed independently of other threads. Also, a thread often has some state that is available only to that thread. Threads are usually contained within a process. A process can include multiple threads. Threads existing in the same process can share resources such as memory.</p><p> The thread counter 101 introduces a new thread into the circuit at each clock cycle. In some situations, this new thread may be a iteration of the thread that has just completed propagation through the pipeline. Thread numbers may be propagated from one pipeline section to the next in each clock cycle. One option for propagating thread numbers is to have in each pipeline section a register 108 for storing the thread numbers of the threads currently in that pipeline section.</p><p> The thread counter may itself be configurable. Typically, the thread counter is configured by an external processor, eg, changing the sequence and / or sequence length.</p><p> Each configuration instance can contain thousands of bits. In this example, each instance is associated with an identifier that consists of a large number of bits smaller than the configuration instance and acts as a convenient shorthand. The first stage in the pipeline is configured to look up the configuration instance identifier used by the current thread from the register store (103). The configuration instance identifier is propagated through the pipeline using register (105). The configuration instance identifier is used in each pipeline stage to look up the portion of the configuration instance required for that pipeline stage. This can be achieved by splitting the configuration instance into separate on-chip memory for each pipeline stage (104). The pipeline stage gets the configuration instance it needs for a particular thread by looking up the configuration identifier for that thread in that particular section of memory. As each thread travels through the pipeline, it experiences only the configuration instance associated with its configuration instance identifier.</p><p> The on-chip memory containing the configuration instance is shared between threads so that any thread can use any configuration instance. One thread can use the same configuration instances as others. Threads can also use different configuration instances. In many instances, the thread may use a completely different configuration instance in the circuit than its predecessor thread. Therefore, it is possible (and likely) that multiple configuration instances are active in the circuit at any given time. The execution of a thread may change which configuration instance identifier it uses (and thus which configuration instance it uses) on the next path in the circuit.</p><p> Use the thread number and some configuration instance bits to access the value from the data store. In this example, this is done by register store (106) for convenience. In one embodiment of the invention, a thread does not have access to values in the register store used by other threads. The register store value is input to the data routing network 111 in subsequent clock cycles. Data routing networks can send values to specific execution units. Data routing networks are configurable, but at least some of the switching over the routing network can be changed from one clock cycle to the next. The switching experienced by each input as it propagates from one pipelined stage of the routing network to the next is determined by the configuration instances derived from the configuration instance identifiers that follow it in the network.</p><p> Data routing The data path within the network is preferably multiple bit wide. The exact width of the data path can be adjusted for a particular application. Within any given routing network, not all data paths need to have the same width. For example, some of the data paths may accommodate wider inputs than others. This can limit routing flexibility in some situations. Input must be sent through a sufficiently wide path of the data path, which can limit the routes available for other inputs in the thread. The input does not have to utilize the full width of the data path, but the output of the network must be able to accommodate some bits equal to the widest path in the data routing network.</p><p> In some embodiments of the invention, it may be more convenient to have several separate routing networks than a single monolithic routing network. In one embodiment of the invention, the control and data values are separated, each with its own set of register stores (106 and 107) and routing networks (111 and 112). In one example, a routing network (111) may have a data path with a width of only 1 bit for a control value, and another routing network (112) may have a data path with a width of 32 bits for a data value. May be provided. The size of the routing network depends on the number of inputs and outputs, and such different routing networks may require different pipeline depths. The routing network in Figure 1 is shown only with one or two pipeline stages. In practice, a routing network can typically have around 12 pipeline stages.</p><p> The input selection unit connects each output from the routing network to the input of the execution unit (115). Execution units can be configured such that their actions on the input are determined by the bits from the configuration instance. The action performed by the execution unit may be determined by one or more bits from the thread data (eg, control values contained in the data). Execution units typically form a single section of the pipeline, but some execution units are configured to perform longer operations that require more than one clock cycle (116). These execution units may form two or more pipeline sections. Similarly, execution units may be connected to each other at the ends of the pipeline such that threads propagate from one execution unit to another (not shown).</p><p> Each execution unit may write the resulting value to a register store (117) to which it can be written. Only one execution unit writes to each register store. The execution unit may write to two or more register stores. Some execution units can read and write to common shared resources (eg, external memory). Reads and writes to shared resources (on-chip or external) tend to be variable latency operations that can be longer than a single clock cycle.</p><p> The position of some registers in some register stores may be associated with a significant bit that asserts whether the data stored at that position is valid. Typically, only the register store associated with the variable latency execution unit needs to have an extra bit to mark each position as valid or invalid. Other register stores may be considered to always hold valid values.</p><p> The valid bit may be set to "disabled" at the start of the write operation and return to "valid" only when the write operation is complete. The circuit may incorporate means to ensure that the position of the register that the thread wants to read is valid before the thread reaches the execution unit (110). These means can be efficiently placed in the same pipeline section as the routing network. This role may be performed by a check unit configured to read the appropriate valid bit for a thread before it is submitted to the execution unit. When a thread is submitted to an execution unit that operates on invalid data, the check unit stops all of those execution units (or at least stops writing to their memory and writing to the register store). May be good. This prevents the results of the actions performed on the "invalid" data from being written to registers or other memory.</p><p> In one example, two significant bits are assigned to each register store location that requires them. Data stored at a register store location is invalid if its two significant bits are different, and may be considered valid if the two bits are the same (or vice versa). Having two significant bits allows two different pipeline stages to write them at the same time. Typically, the pipeline stage that wants to invalidate the data in the register store is configured to flip one of the significant bits, while the other pipeline stage that activates the data in the register store is the other of the significant bits. Is configured to flip.</p><p> The execution unit (118) can also change the configuration instance used by a particular thread on other paths in the circuit by changing the configuration instance identifier associated with that thread (119). The new configuration instance identifier is used for that thread in the next path in the circuit.</p><p> Execution units sometimes need to act on the results of the thread's previous execution. One example is the accumulator operation. The circuit may include one or more dedicated units that perform such operations. One example is the Accumulate Register Store. These register stores (eg: 114) do not have to move within the routing network and can reduce the size of the routing network required.</p><p> Execution units typically have no feedback within them. Feedback is provided on a whole circuit basis by the execution of a thread that modifies the data stored in register store or external memory and / or modifies the configuration instance identifier of the thread.</p><p><u style="single">Register store</u> Each register store contains multiple locations that store separate values. The circuit may select the position using the register store address. In one embodiment of the invention, the thread accesses a separate set of locations within each register store. This can be implemented by ensuring that some of the read and write addresses to the register store are based on the thread number (in the appropriate pipeline stage) and zero or more configuration instance bits. In this embodiment, the thread cannot access the values held in the register store associated with other threads.</p><p> Since a register store is usually read and written in different pipeline stages, the read and write addresses for the register store in any predetermined clock cycle are often different. Therefore, it is advantageous to implement the register store in on-chip memory that can perform separate read and write operations in one clock cycle.</p><p><u style="single">Routing network</u> A routing network is basically a switch that connects multiple inputs to multiple outputs. Inputs can be connected to a single output or multiple outputs. The routing network is preferably configurable such that at least a portion of its switching can be configured on a clock cycle basis by bits from the configuration instance.</p><p> The routing network may be connectable to any input or any output (and, in some embodiments, two or more outputs). The routing network may be non-blocking so that the inputs can be connected to the outputs in any combination.</p><p> A crossbar switch is an example of a suitable switch for implementing a configurable routing network. The term "crossbar switch" is sometimes used to refer to a completely flexible switch, but a switch that has the ability to connect each and every input to one (and only one) output. Also used to refer. Clos networks may be appropriate for large switches. Clos networks are multi-stage switches. One option is to build a Clos network from multiple crossbar switches. Clos networks typically allow each and every input to be connected to one output without limitation. Also, although this is not always possible, it is possible to connect inputs to multiple outputs, depending on the connectivity required.</p><p> Figure 2 shows an example of a switch suitable for implementing a routing network. This figure shows an NxN Clos network, where at least two outer stages of the network are implemented by 2x2 crossbar switches (201). The inner part of the network has been shown to be implemented by two N / 2 crossbar switches (203). These larger crossbar switches may be "nested" and may themselves be a Clos network implemented by multiple stages of the crossbar switch (or any other switch). .. The switch is pipelined, as indicated by register 202. Registers are configured to hold thread data from one clock cycle to the next.</p><p> The advantage of pipelined routing networks is that it allows long data paths to be divided into smaller sections. These smaller sections can be moved more quickly and propagation along them can accommodate a single clock cycle, even at faster clocks. One option is to have registers at all levels of the nested multistage switch (so that each stage of the switch represents a section of the pipeline). However, in practice this is not necessary if the distance in the inner stage of the switch tends to be much shorter and the clock speed is unlikely to be suppressed. Therefore, a single pipeline section may include two or more inner stages of a multistage switch so that no registers are required for each stage of the switch.</p><p> An example of a 2 × 2 crossbar switch is shown in FIG. This switch is arranged to receive two inputs 301 and output two outputs 304. The switch comprises two multiplexers 302. Each multiplexer receives each of the two inputs and selects one as the output. Each multiplexer is controlled by configuration instance bit 303 to select a particular one of its inputs as its output. Therefore, the configuration instance controls the mapping from input 301 to output 304. A 2x2 crossbar is a simple example, but by stacking layers of 2x2 crossbars, you can build a flexible routing network that can take multiple inputs and send them to multiple outputs. Is possible. This allows the input to be delivered to the appropriate location on the circuit for further processing.</p><p> The 2x2 crossbar is just one example of a crossbar switch, and crossbars of other sizes can be used (eg, 3x3, 4x4 or more). Multistage switches may use crossbars of different sizes on different stages.</p><p><u style="single">Execution unit</u> The execution unit can be designed to be capable of performing a series of operations including, but not limited to, calculation operations, logical operations or shift operations, or memory read or write operations. The execution unit can use bits from the configuration instance in addition to bits from its data input (eg, thread control values), which allows it to determine what action to perform on a particular thread. Some execution units may have unique capabilities that differ from other execution units. For example, it may be possible to perform an operation that cannot be performed by another execution unit. The number and capabilities of execution units can be modified to suit a particular application.</p><p> An example of the execution unit is shown in Fig. 4. Execution unit 401 can be configured to perform operations based on configuration instance bit 407. It is the configuration instance bits that determine how the execution unit behaves with respect to the data. The execution unit also includes a data input 405. Typically, these inputs receive thread data sent by the data routing circuit to the execution unit. Some of these inputs also affect how the execution unit behaves. The execution unit also receives the clock signal 402 and thread number 403. The clock signal controls the pipeline. The thread number identifies the thread currently being processed by the execution unit. The final input 404 allows the register to be written, which will be described in more detail below.</p><p> The execution unit outputs the data for writing to its dedicated register store (408,409). The output data represents the result of the action performed by the execution unit on its input. Each data output 412 is preferably provided with two attached outputs, namely a write enable 410 and a write address 411. The write enable 410 is set by an input 404 that allows the register to be written. Data may be written to the register only if the write enable is held at the appropriate value (typically 1 or 0). The write operation is stopped if the write enable is not an appropriate value. This can be used if a register position is found to be invalid and all register writes can be disabled until the register position is valid again (this is in the "Pipeline" section below). Will be explained in more detail). The write address 411 is usually determined by the thread number and some configuration instance bits.</p><p> Some examples of specific execution units are shown in Figures 5-10.</p><p> Figure 5 shows an execution unit configured as a simple adder. The execution unit includes an input 501 for enabling writing and an input 502 for identifying the thread number. The execution unit also includes inputs 503,504 for the data to be added. In this example, the execution unit can only add (507). One configuration instance bit extracted from the configuration instance determines whether the adder result is written to the register store (505). Inputs that allow high-driven register writes when none of the current thread values are valid can also prevent the adder results from being written. Execution unit output 508 outputs write data, write addresses, and write enable for register stores.</p><p> Figure 6 shows an execution unit configured as a pipelined ALU. In this more complex example, the execution unit can perform several different actions. Some configuration instance bits and bits from input 601 control what the ALU does (603). For example, in one configuration the ALU acts as a multiplexer, in the other configuration the ALU uses the control input as a carry bit to execute Add, and in the other configuration the control input is whether the ALU executes Add or Subtract. Can be selected. Register 602 is provided to pipeline the other inputs to fit the ALU pipeline. The ALU produces output 604 as well as data values. As an example, this 1-bit output is high when the ALU result is zero.</p><p> FIG. 7 shows an execution unit with a long latency operator 701. In this example, a plurality of registers 702 are provided to pipeline the write enable and write address values. These values propagate through registers at each clock cycle, allowing new threads to be submitted to the execution unit to initiate long latency operations at each clock cycle. The number of registers preferably matches the latency of operation. Operator 701 may or may not be pipelined. Some operators do not require pipelined because some of the long latency operations are not performed by the operators. For example, this operation is a read or write operation to external memory, in which latency is associated with access to that memory via the system bus.</p><p> Figure 8 shows an example of an execution unit that modifies the configuration instance identifier used by a thread on the next path in the circuit. In this example, the selection is controlled by three control bits 801 that select one of the eight configuration instance identifiers 802. The execution unit has output 803 for storing the selected configuration instance identifier.</p><p> The execution unit may be fractable. That is, it may be separable into smaller separate executable units, depending on the requirements of the thread. An example is shown in FIG. The execution unit of FIG. 9 has a 64-bit ALU901 that can be split into two 32-bit ALUs. Inputs 902,903 can be used for two pairs of 32-bit values or one pair of 64-bit values. The configuration instance bits set whether the ALU operates as a 64-bit ALU or as two 32-bit ALUs. One of the advantages of a fractable execution unit is that it is cheaper to implement than two or more separate units.</p><p> Traditionally, some execution units need to be presented in a particular order. Preferably, the execution unit is configured so that the order of inputs does not matter, if possible. Two examples include the look-up table shown in Figure 10. These lookup tables are configured to be independent of the order in which the inputs are presented, which allows some switches to be removed from the routing network.</p><p><u style="single">pipeline</u> The time it takes for an instruction to complete is the maximum number of pipeline stages (indicated by "p") between a register store read and the corresponding register store write, and the processor clock frequency ("f"). ). And the latency per instruction is p / f. However, the pipeline can process instructions from different threads of p or more per clock cycle. Threads circulate continuously and are issued to the pipeline, one for each clock cycle.</p><p> Whenever a value read from a register store is considered invalid, the thread is prevented from writing to any register store or changing its configuration instance identifier. This makes it impossible for a thread to change its visible state to that thread and resumes from the same state when it is reissued to the pipeline. Preferably, the circuit is configured such that each thread only accesses its own register store (as described above). And all other threads are in the pipeline unaffected by the condition that their read values are valid, regardless of whether any of the other threads hit the invalid value. proceed. The invalid register value results from an execution unit with variable latency, so the invalid register value eventually becomes valid and can be done by a thread that was previously blocked from updating its state. Become. In this way, individual threads can be considered "stalled" even though the pipeline itself continues to propagate values.</p><p> The user has no visibility of the pipeline registers. This allows the program to operate on different circuits designed according to the principles described herein without deformation, even if the circuits have different pipelines. The only difference is the length of time it takes to complete each instruction.</p><p><u style="single">Configuration instance</u> On-chip memory has a set of configuration instances. In one embodiment of the invention, the configuration instance memory can be accessed by an external processor. Individual configuration instances can be loaded by writing to configuration instance memory. If the configuration memory can be read and written in the same clock cycle, the thread can continue to travel in the pipeline while the configuration instance is loaded. Configuration instances used by any thread in the pipeline should not be loaded. This can be achieved by an operating system or some additional hardware that monitors all configuration instance identifiers in use.</p><p> One configuration instance may not be able to modify register stores or access memory. This "null" configuration instance can be used for slots in the pipeline when the thread is inactive (eg at startup).</p><p> In one embodiment of the invention, the circuitry for a particular execution unit or portion of routing may be dynamically modified. The operating system must be configured to ensure that no threads use circuits that are undergoing dynamic changes. FPGA is an example of a technology that can dynamically change a circuit. Typically, this type of reprogramming involves downloading a program file from off-chip and reconfiguring all or part of the circuit. This process is typically done on the millisecond order (as opposed to the configuration of the circuit for each thread on the nanosecond order). Delays are justified if the circuit needs to perform some specialized processing at once, such as encryption or some other intensive processing operation. The circuits described herein are particularly suitable for this type of dynamic reconstruction because they contain an execution unit within them. They can be modified without the need to modify the structure of the surrounding circuitry.</p><p> The particular examples described above can be modified in various ways within the scope of the present invention. For example, the circuit described above allows a thread to change its configuration instance by changing the configuration instance identifier it uses in the next path in the circuit, or by writing control data. .. Other possibilities that can be implemented in the future include having threads change the configuration instance identifiers that apply to other threads, or having threads write directly to configuration instance memory.</p><p> Applicants separately disclose the individual features described herein and any combination of any two or more of such features, the disclosure of which such features and combinations are herein. Such features and combinations can be implemented in light of the general wisdom of those skilled in the art, whether or not they resolve disclosure issues, and without any limitation to the claims. It is done to the extent that. Applicants point out that aspects of the invention may consist of any such individual feature or combination of features. In view of the above description, it will be apparent to those skilled in the art that various modifications may be made within the scope of the present invention.</p>
10 sheets
Sheet 1 Sheet 2 Sheet 3 Sheet 4 Sheet 5 Sheet 6 Sheet 7 Sheet 8 Sheet 9 Sheet 10
Every citation, both ways
| Document | Relation | Office |
|---|---|---|
| JP2010146102A | Cites | Japan |
| US20120284379A1 | Cites | United States of America |
| JP2011522317A | Cites | Japan |
| JP2008539485A | Cites | Japan |
| JP2007504536A | Cites | Japan |
| JP2006502507A | Cites | Japan |
| WO2006109835A1 | Cites | World Intellectual Property Organization (WIPO) |
| JP2006040254A | Cites | Japan |
| US07149996B1 | Cites | United States of America |
| US20120159062A1 | Cites | United States of America |
| US20050177703A1 | Cites | United States of America |
| US20040150422A1 | Cites | United States of America |
| 天野英晴,外4名,動的リコンフィギュラブルプロセッサMuCCRAのコンフィギュレーション機構,電子情報通信学会技術研究報告,日本,社団法人電子情報通信学会,2006年 9月 8日,Vol.106,No.247,(RECONF2006-27~37),Pages:19~24 | Non-patent | – |
22 members in 7 offices
Members22
| Document | Office | Kind | |
|---|---|---|---|
| GB201319279D0 | United Kingdom | D0 | |
| GB2519813A | United Kingdom | A | |
| WO2015063466A1 | World Intellectual Property Organization (WIPO) | A1 | |
| WO2015063466A4 | World Intellectual Property Organization (WIPO) | A4 | |
| GB201513909D0 | United Kingdom | D0 | |
| GB2526018A | United Kingdom | A | |
| GB2519813B | United Kingdom | B | |
| CN105830054A | China | A | |
| EP3063651A1 | European Patent Office (EPO) | A1 | |
| KR20160105774A | Republic of Korea | A | |
| US2016259757A1 | United States of America | A1 | |
| JP2016535913A | Japan | A | |
| US9658985B2 | United States of America | B2 | |
| GB201801471D0 | United Kingdom | D0 | |
| US2018089140A1 | United States of America | A1 | |
| GB2555363A | United Kingdom | A | |
| GB2555363B | United Kingdom | B | |
| GB2526018B | United Kingdom | B | |
| CN105830054B | China | B | |
| US10275390B2 | United States of America | B2 | |
| US2020026685A1 | United States of America | A1 | |
| JP6708552B2This record | Japan | B2 |
17 legal events, as the office reported them to INPADOC
Over the term
Point at a mark for the eventEvents
| Event | Code | |
|---|---|---|
| Receipt of annual feesJAPANESE INTERMEDIATE CODE: R250R250 | R250 | |
| Receipt of annual feesJAPANESE INTERMEDIATE CODE: R250R250 | R250 | |
| Receipt of annual feesJAPANESE INTERMEDIATE CODE: R250R250 | R250 | |
| Receipt of annual feesJAPANESE INTERMEDIATE CODE: R250R250 | R250 | |
| First payment of annual fees (during grant procedure)JAPANESE INTERMEDIATE CODE: A61A61 | A61 | |
| Certificate of patent or registration of utility modelJAPANESE INTERMEDIATE CODE: R150R150 | R150 | |
| Written decision to grant a patent or to grant a registration (utility model)JAPANESE INTERMEDIATE CODE: A01A01 | A01 | |
| Decision of grant or rejection writtenTRDD | TRDD | |
| Transfer to examiner for re-examination before appeal (zenchi)AppealJAPANESE INTERMEDIATE CODE: A911A911 | A911 | |
| Request for written amendment filedJAPANESE INTERMEDIATE CODE: A523A521 | A521 | |
| Decision of refusalJAPANESE INTERMEDIATE CODE: A02A02 | A02 | |
| Request for written amendment filedJAPANESE INTERMEDIATE CODE: A523A521 | A521 | |
| Written request for extension of timeJAPANESE INTERMEDIATE CODE: A601A601 | A601 | |
| Notification of reasons for refusalJAPANESE INTERMEDIATE CODE: A131A131 | A131 | |
| Report on retrievalJAPANESE INTERMEDIATE CODE: A971007A977 | A977 | |
| Request for written amendment filedJAPANESE INTERMEDIATE CODE: A523A521 | A521 | |
| Written request for application examinationJAPANESE INTERMEDIATE CODE: A621A621 | A621 |
Numbers
- Publication
- 6708552
- Application
- 2016551066
Titles2
- Japanese
- パイプライン化構成可能プロセッサ
- English
- Pipelined configurable processor
Classification
- CPC, 14
- G06F15/7878
- G06F9/3897
- G06F15/7867
- G06F9/3851
- G06F9/3869
- G06F9/3873
- G06F15/7875
- G06F15/7892
- G06F9/3867
- G06F13/4022
- H04L49/1546
- H04L49/1569
- H04L49/101
- G06F1/10
- IPC, 3
- G06F9 38
- G06F9 46
- G06F15 78
