Thread optimized multiprocessor architecture
Abstract
In one aspect, the invention is a system comprising (a) multiple parallel processors on a single chip and (b) computer memory located on the chip and accessible by each of the processors. Each of the processors can operate to handle the de minimis instruction set, and each of the processors has its own local cache for each of at least three specific registers in the processor. In another aspect, the invention is a system comprising (a) multiple parallel processors on a single chip and (b) computer memory located on the chip and accessible by each of the processors. Each processor can operate to process an instruction set optimized for thread-level parallelism, and each processor accesses the internal data bus of computer memory on the chip and is internal. The data bus is the width of one line of memory.

Term
Projected expiry 27 June 2028.
- Priority
- Filed
- Published
- Today
- Projected expiry
74 claims: 10 independent, 64 dependent
- 1メモリモジュール上に搭載された複数の並列プロセッサと、 外部メモリコントローラと、 汎用中央演算処理装置とを備えていて、 前記並列プロセッサの各々は、スレッドレベルの並列処理のために最適化された命令セットを処理するように動作可能であることを特徴とするシステム。
- 2前記並列プロセッサの各々は、de minimis命令セットを処理するように動作可能であることを特徴とする請求項1に記載のシステム。
- 3メモリモードレジスタに割り当てられる1以上のビットは、前記並列プロセッサのうちの1つ以上をイネーブルまたはディスエーブルにするように動作可能であることを特徴とする請求項1に記載のシステム。
- 4前記メモリモジュールは、デュアルインラインメモリモジュールであることを特徴とする請求項1に記載のシステム。
- 5前記プロセッサの各々は、単一のスレッドを処理するように動作可能であることを特徴とする請求項1に記載のシステム。
- 6複数のスレッドが、共有メモリを通してデータを共有することを特徴とする請求項5に記載のシステム。
- 7複数のスレッドが、1つ以上の共有変数を通してデータを共有することを特徴とする請求項5に記載のシステム。
- 8前記メモリモジュールは、DRAM、SRAM、およびフラッシュメモリのうちの1つ以上であることを特徴とする請求項1に記載のシステム。
- 9少なくとも一つの前記並列プロセッサがマスタプロセッサとみなされ、他の前記並列プロセッサはスレーブプロセッサとみなされることを特徴とする請求項1に記載のシステム。
- 10各プロセッサは、クロック速度を有していて、前記マスタプロセッサ以外の各プロセッサは、性能または電力消費を最適化するように調整された前記プロセッサのクロック速度を有するように動作可能であることを特徴とする請求項9に記載のシステム。
- 11各プロセッサは、マスタプロセッサまたはスレーブプロセッサとみなされるように動作可能であることを特徴とする請求項9に記載のシステム。
- 12前記マスタプロセッサは、いくつかのスレーブプロセッサによる処理を要求し、前記いくつかのスレーブプロセッサからの出力を待ち、かつ前記出力を結合することを特徴とする請求項9に記載のシステム。
- 13前記マスタプロセッサは、前記出力が前記いくつかのプロセッサの各々から受信されるとき、前記いくつかのプロセッサからの出力を結合することを特徴とする請求項12に記載のシステム。
- 14停止されるべき前記並列プロセッサのうちの1つ以上をイネーブルにすることによって、低電力消費が提供されることを特徴とする請求項1に記載のシステム。
- 15前記並列プロセッサの各々は、プログラムカウンタを伴っていて、前記並列プロセッサが伴っているプログラムカウンタに全て1を書き込むことによって停止されるように動作可能であることを特徴とする請求項14に記載のシステム。
- 16ダイナミックランダムアクセスメモリ(DRAM)のダイに埋め込まれた複数の並列プロセッサを備えていて、 前記複数の並列プロセッサは、外部メモリコントローラおよび外部プロセッサと通信し、 前記並列プロセッサの各々は、スレッドレベルの並列処理のために最適化された命令セットを処理するように動作可能であることを特徴とするシステム。
- 17前記ダイは、DRAMピン配列を有するパッケージに入れられていることを特徴とする請求項16に記載のシステム。
- 18前記並列プロセッサは、デュアルインラインメモリモジュール上に搭載されていることを特徴とする請求項16に記載のシステム。
- 19前記システムは、前記プロセッサがDRAMモードレジスタを通してイネーブルにされる時以外は、DRAMとして動作することを特徴とする請求項16に記載のシステム。
- 20前記外部プロセッサは、関連する永久記憶装置から前記DRAMにデータおよび命令を転送するように動作可能であることを特徴とする請求項16に記載のシステム。
- 21前記永久記憶装置は、フラッシュメモリであることを特徴とする請求項20に記載のシステム。
- 22前記外部プロセッサは、前記並列プロセッサと外部装置との間の入出力インターフェースを提供するように動作可能であることを特徴とする請求項16に記載のシステム。
- 23単一のチップ上の複数のプロセッサと、 前記チップ上に配置されていて、前記プロセッサの各々によってアクセス可能なコンピュータメモリとを備えていて、 前記プロセッサの各々は、de minimis命令セットを処理するように動作可能であり、かつ 前記プロセッサの各々は、前記プロセッサ内の少なくとも3つの特定のレジスタの各々専用のローカルキャッシュを有していることを特徴とするシステム。
- 24前記ローカルキャッシュの各々のサイズは、前記チップ上のランダムアクセスメモリの1行に等しいことを特徴とする請求項23に記載のシステム。
- 25各前記プロセッサは、前記チップ上のランダムアクセスメモリの内部データバスにアクセスし、前記内部データバスは、ランダムアクセスメモリの1行の幅を有していることを特徴とする請求項23に記載のシステム。
- 26前記内部データバスの幅は、1024、2048、4096、8192、16328、または32656ビットであることを特徴とする請求項25に記載のシステム。
- 27前記内部データバスの幅は、1024ビットの整数倍であることを特徴とする請求項25に記載のシステム。
- 28前記プロセッサ内の少なくとも3つの特定のレジスタの各々専用のローカルキャッシュは、1メモリ読出し又は書込みサイクルの中で満たされるか又は消去されるように動作可能であることを特徴とする請求項23に記載のシステム。
- 29前記de minimis命令セットは、基本的に7つの基本命令から成ることを特徴とする請求項23に記載のシステム。
- 30前記基本命令セットは、ADD、XOR、INC、AND、STOREACC、LOADACC、およびLOADI命令を含むことを特徴とする請求項29に記載のシステム。
- 31前記de minimis命令セット内の各命令は、長さが最長でも8ビットであることを特徴とする請求項23に記載のシステム。
- 32前記de minimis命令セットは、プロセッサ上での命令シーケンスの実行を最適化するための複数の命令拡張を有していて、更に、このような命令拡張は、基本的に20未満の命令から成ることを特徴とする請求項23に記載のシステム。
- 33各命令拡張は、長さが最長でも8ビットであることを特徴とする請求項23に記載のシステム。
- 34前記de minimis命令セットは、前記チップ上の複数のプロセッサを選択的に制御するための一組の命令を有していることを特徴とする請求項23に記載のシステム。
- 35各プロセッサ制御命令は、長さが最長でも8ビットであることを特徴とする請求項34に記載のシステム。
- 36複数のプロセッサは、モノリシックメモリデバイスのために設計された半導体製造プロセスを用いて前記チップ上に配置されるコンピュータメモリと共に前記チップ上に製造されることを特徴とする請求項23に記載のシステム。
- 37半導体製造プロセスは、4層未満のメタル相互接続を用いることを特徴とする請求項36に記載のシステム。
- 38半導体製造プロセスは、3層未満のメタル相互接続を用いることを特徴とする請求項36に記載のシステム。
- 39複数のプロセッサのコンピュータメモリ回路内への集積化は、チップダイサイズの30%未満の増加という結果をもたらすことを特徴とする請求項23に記載のシステム。
- 40複数のプロセッサのコンピュータメモリ回路内への集積化は、チップダイサイズの20%未満の増加という結果をもたらすことを特徴とする請求項23に記載のシステム。
- 41複数のプロセッサのコンピュータメモリ回路内への集積化は、チップダイサイズの10%未満の増加という結果をもたらすことを特徴とする請求項23に記載のシステム。
- 42複数のプロセッサのコンピュータメモリ回路内への集積化は、チップダイサイズの5%未満の増加という結果をもたらすことを特徴とする請求項23に記載のシステム。
- 43250,000個未満のトランジスタが、前記チップ上の各プロセッサを作成するために用いられることを特徴とする請求項23に記載のシステム。
- 44チップは、4層未満のメタル相互接続を用いる半導体製造プロセスを用いて製造されることを特徴とする請求項23に記載のシステム。
- 45前記プロセッサの各々は、単一のスレッドを処理するように動作可能であることを特徴とする請求項23に記載のシステム。
- 46アキュムレータは、インクリメント命令を除く、あらゆる基本命令のためのオペランドであることを特徴とする請求項29に記載のシステム。
- 47各基本命令のための宛先は、常にオペランドレジスタであることを特徴とする請求項29に記載のシステム。
- 483つのレジスタは自動インクリメントであり、かつ3つのレジスタは自動デクリメントであることを特徴とする請求項23に記載のシステム。
- 49各基本命令は、完了するために1クロックサイクルのみを必要とすることを特徴とする請求項29に記載のシステム。
- 50前記命令セットは、分岐命令およびジャンプ命令を有していないことを特徴とする請求項29に記載のシステム。
- 51単一のマスタプロセッサが、前記並列プロセッサの各々を管理する役割を担っていることを特徴とする請求項23に記載のシステム。
- 52単一のチップ上の複数の並列プロセッサと、 前記チップ上に配置されていて、前記プロセッサの各々によってアクセス可能なコンピュータメモリとを備えていて、 前記プロセッサの各々は、スレッドレベルの並列処理のために最適化された命令セットを処理するように動作可能であり、かつ 各前記プロセッサは、前記チップ上のコンピュータメモリの内部データバスにアクセスし、前記内部データバスは、メモリの1行より幅が広くないことを特徴とするシステム。
- 53前記プロセッサの各々は、de minimis命令セットを処理するように動作可能であることを特徴とする請求項52に記載のシステム。
- 54前記プロセッサの各々は、前記プロセッサ内の少なくとも3つの特定のレジスタの各々専用のローカルキャッシュを有していることを特徴とする請求項52に記載のシステム。
- 55前記ローカルキャッシュの各々のサイズは、前記チップ上のコンピュータメモリの1行に等しいことを特徴とする請求項54に記載のシステム。
- 56少なくとも3つの特定のレジスタは、命令レジスタ、ソースレジスタ、および宛先レジスタを含むことを特徴とする請求項54に記載のシステム。
- 57前記de minimis命令セットは、基本的に7つの基本命令から成ることを特徴とする請求項53に記載のシステム。
- 58前記基本命令セットは、ADD、XOR、INC、AND、STOREACC、LOADACC、およびLOADI命令を含むことを特徴とする請求項57に記載のシステム。
- 59前記命令セット内の各命令は、長さが最長でも8ビットであることを特徴とする請求項52に記載のシステム。
- 60前記プロセッサの各々は、単一のスレッドを処理するように動作可能であることを特徴とする請求項52に記載のシステム。
- 61単一のマスタプロセッサが、前記並列プロセッサの各々を管理する役割を担っていることを特徴とする請求項52に記載のシステム。
- 62前記de minimis命令セットは、プロセッサ上での命令シーケンスの実行を最適化するための複数の命令拡張を有していて、更に、このような命令拡張は、20未満の命令を有していることを特徴とする請求項53に記載のシステム。
- 63各命令拡張は、長さが最長でも8ビットであることを特徴とする請求項62に記載のシステム。
- 64前記de minimis命令セットは、前記チップ上の複数のプロセッサを選択的に制御するための一組の命令を有していることを特徴とする請求項53に記載のシステム。
- 65各プロセッサ制御命令は、長さが最長でも8ビットであることを特徴とする請求項64に記載のシステム。
- 66複数のプロセッサは、モノリシックメモリデバイスのために設計された半導体製造プロセスを用いて前記チップ上に配置されるコンピュータメモリと共に前記チップ上に製造されることが可能であることを特徴とする請求項52に記載のシステム。
- 67単一のチップ上の複数の並列プロセッサ、マスタプロセッサ、およびコンピュータメモリを利用するスレッドレベルの並列処理の方法において、前記複数のプロセッサの各々は、de minimis命令セットを処理し、かつ単一のスレッドを処理するように動作可能であり、(a)ローカルキャッシュを前記複数のプロセッサの各々の中の3つの特定のレジスタの各々に割り当てるステップと、(b)単一のスレッドを処理するために複数のプロセッサのうちの1つを割り当てるステップと、(c)前記プロセッサによって各々の割り当てられたスレッドを処理するステップと、(d)前記プロセッサによって処理された各スレッドからの結果を処理するステップと、(e)スレッドが処理された後に、前記複数のプロセッサのうちの1つの割り当てを解除するステップとを有していることを特徴とする方法。
- 68de minimis命令セットは、基本的に7つの基本命令から成ることを特徴とする請求項67に記載の方法。
- 69前記基本命令は、ADD、XOR、INC、AND、STOREACC、LOADACC、およびLOADI命令を有していることを特徴とする請求項68に記載の方法。
- 70de minimis命令セットは、複数のプロセッサを選択的に制御するための一組の命令を有していることを特徴とする請求項67に記載の方法。
- 71各プロセッサ制御命令は、長さが最長でも8ビットであることを特徴とする請求項70に記載の方法。
- 72各プロセッサが前記メモリの内部データバスを用いてコンピュータメモリにアクセスするステップを更に有していて、内部データバスは、前記チップ上のメモリの1行の幅であることを特徴とする請求項52に記載の方法。
- 73de minimis命令セット内の各命令は、長さが最長でも8ビットであることを特徴とする請求項67に記載の方法。
- 74メモリデバイスのための電子工業規格デバイスのパッケージングおよびピンレイアウトと互換性があるメモリチップの中に埋め込まれた複数のプロセッサを備えていて、 プロセッサのうちの1つ以上は、メモリチップのメモリモードレジスタに送信される情報によって起動することができ、メモリチップは、プロセッサのうちの1つ以上が前記メモリモードレジスタによって起動する場合を除き、工業規格メモリデバイスの動作と機能的に互換性があることを特徴とするシステム。
Independent claims74
160 paragraphs, as filed
The present invention relates to a thread-optimized multiprocessor architecture.
Computer speed can be increased using two common methods. Increase instruction execution speed or execute more instructions in parallel. When instruction execution speed approaches the limit of electron mobility in silicon, parallelism is the best alternative to speeding up computers.
Previous attempts at the parallel method include: 1. Overlapping the fetching of the next instruction with the execution of the current instruction. 2. Pipelining instructions. The instruction pipeline breaks down each instruction into as many parts as possible and then attempts to map the sequential instructions into parallel execution units. The best improvements in theory are the inefficiency of multi-step instructions, the inability of many software programs to provide sufficient sequential instructions to keep parallel execution units filled, and branching, looping, Or when the case syntax encounters a situation that requires replenishment of the execution unit, it is rarely achieved due to the considerable amount of time penalty paid. 3. Single instruction multiple data) That is SIMD. This type of technology is found in the Intel SSE instruction set, as implemented in Intel Pentium® 3 and other processors. In this technique, a single instruction is executed on multiple datasets. This technique is only useful for special applications, such as video graphics rendering applications. 4. Hyper Cube. This technique uses large two-dimensional arrays of processors and local memory, sometimes three-dimensional arrays. The communications and interconnects required to support these arrays of processors essentially limit them to highly specialized applications.
A pipeline is an instruction execution unit consisting of multiple sequential stages that execute a part of the execution of one instruction, such as fetch, decode, execute, and store in succession. Since several pipelines can be arranged in parallel, program instructions are supplied to each pipeline one after another until all the pipelines are executing instructions. The filled instructions are then repeated in the first pipeline. When N pipelines are filled with instructions and executions, the effect on performance is theoretically the same as increasing execution speed by N times for a single execution unit.
Successful pipelined depends on: 1. Instruction execution must be able to be defined as several contiguous states. 2. Each instruction must have the same number of states. 3. The number of states per instruction determines the maximum number of parallel execution units.
Pipelines can be achieved because pipelined can achieve performance gains based on the number of parallel pipelines, and the number of parallel pipelines is determined by the number of states in one instruction. , Facilitates instructions with complex multiple states.
Severely pipelined computers rarely achieve performance close to the theoretical performance improvements expected from parallel pipeline execution units. Some reasons for this pipeline penalty include: 1. A software program does not consist only of sequential instructions. Various studies have shown that changes in execution flow occur every 8-10 instructions. Branches that change the flow of a program capsize the pipeline. Attempts to minimize pipeline capsizing tend to be complex and incomplete in their mitigation. 2. Forcing all instructions to have the same number of states often leads to an execution pipeline that meets the requirements of the lowest common denominator (ie, the slowest and most complex) instruction. Because of the pipeline, all instructions are put in the same number, regardless of whether they need it or not. For example, logical operations (eg AND or OR) are performed orders of magnitude faster than ADD, but both are often allocated the same amount of time to perform. 3. Pipelines facilitate complex instructions with multiple states. Instructions that may require two states are usually extended to satisfy 20 states. Because that is the depth of the pipeline. (Intel Pentium® 4 uses a 20-state pipeline.) 4. The time required for each pipeline state must be due to propagation delays due to logic circuits and associated transistors, as well as design margins or tolerances for a particular state. 5. Arbitration for access to pipeline registers and other resources often degrades performance due to the propagation delay of transistors in the arbitration logic. 6. There is an upper limit to the number of states in which an instruction can be split, before adding states actually slows down execution rather than speeding it up. Some studies have shown that the pipeline architecture of recent generations of Digital Equipment Corporation's alpha processors goes beyond that point and is actually slower than previous shorter pipeline versions of the processor.
Dividing the pipeline separately One way of thinking about rethinking CPU design is to think of pipelined execution units, which are divided into multiple (N) simplified processors. (Registers and some other logic may need to be duplicated in such a design.) Each of the N simplified processors has the following advantages over the pipeline architecture described above: There is. 1. There is no delay due to the pipeline. No branch prediction is required. 2. Not all instructions can be assigned the same execution time as the slowest instruction, but can take as much time as needed for the instruction. 3. Instructions are simplified by reducing the required execution state, which can reduce pipeline penalties. 4. Propagation delay can be reduced by each state removed from the pipeline And the design margin required for this condition can be removed. 5. Register arbitration can be excluded.
In addition, a system with N simplified processors has the following advantages over pipelined CPUs: 1. There is no limit on maximum pipeline parallelism. 2. Unlike pipelined processors, multiple stand-alone processors can be selectively turned off to reduce power consumption when not in use.
Other challenges with current approaches to parallelism Many embodiments of the parallel method yield to the limitations of Amdahl's law. Acceleration by the parallel method is limited by the overhead due to the non-serializable part of the task. In essence, as the amount of parallelism increases, the communication required to support it overwhelms the enhancements that result from parallelism.
Stoplight on the Redline Another inefficiency of current processors is the inability to scale computing power to meet immediate computational demands. Most computers spend most of their time waiting for something to happen. They are waiting for I / O, the next instruction, memory access, or sometimes the human interface. This wait is an inefficient waste of computing power. In addition, computer time spent on standby often results in increased power consumption and heat generation.
Exceptions to the wait principle are applications such as engine controllers, signal processors, and firewall routers. These applications are good candidates for accelerating parallel computing because of the predetermined nature of the set of problems and the set of solutions. Problems that require the product of N independent multipliers can be solved faster with N multipliers.
The recognized performance of a general purpose computer is actually its peak performance. Recently, general-purpose computers have started to get busy by running video games with fast screen refreshes, compiling large source files, or searching databases. In the optimal world, video rendering takes into account special purposes, shading, conversion, and rendering hardware. One way to think about programming for such special purpose hardware is to use "threads".
Threads are independent programs, self-contained, and rarely transmit data to other threads. A common use for threads is to collect data from slow, real-time movements and provide organized results. Threads may also be used to draw changes on the display. A thread can transition through thousands or millions of states before requiring further interaction with other threads. Independent threads provide an opportunity for performance enhancements through the parallel method.
Many software compilers support thread creation and management to take into account the software design process. The same consideration supports parallel processing of multiple CPUs through thread-level parallelism techniques implemented within a thread-optimized microprocessor (TOMI) in a preferred embodiment.
Thread level parallel method Threading is a well-understood technique for considering software programs on a single CPU. Thread-level parallelism can achieve program acceleration through the use of TOMI processors.
One important advantage of TOMI processors over other parallel methods is that TOMI processors require minimal changes to current software programming techniques. No new algorithm needs to be developed. Many existing programs may need to be recompiled, but virtually not need to be rewritten.
An effective TOMI computer architecture should be built for a large number of simplified processors. Different architectures can be used for different types of computing tasks.
Basic computer behavior For general purpose computers, the most common infrequent actions are load and store, sequencing, and math and logic.
Load and store Load and store parameters are source and destination. Load and store power is in the source and destination range (for example, 4 gigabytes is a stronger range than 256 bytes). The locality associated with the current source and destination is important for many datasets. Plus 1 and minus 1 are the most useful. The greater the offset from the current source and destination, the less useful it becomes.
Loads and stores can also be affected by the memory hierarchy. Loading from storage is the slowest operation the CPU can perform.
Sequencing Branches and loops are basic sequencing instructions. Instructions that change sequences based on tests are the way computers make decisions.
Math and logic Math and logic actions are the least used of the three actions. Logic operation is the fastest operation that a CPU can perform and requires as little delay as the delay of a single logic gate. Mathematical movements are more complex. This is because the high-order bit depends on the result of the operation of the low-order bit. 32-bit ADD requires a delay of at least 32 gates, even with carry forward control. MULTIPLY using the shift and add method requires 32 ADD equivalents.
Instruction size trade-off The complete instruction set consists of arithmetic codes, which are large enough to select an infinite number of possible sources, destinations, arithmetics, and the next instruction. Unfortunately, the math code for the complete instruction set is infinitely wide, so the instruction bandwidth is zero.
Computer design for high instruction bandwidth requires the creation of an instruction set with the most common sources, destinations, operations, and operation codes that can efficiently define the next instruction with the fewest operation code bits. To do.
Wide math code leads to high instruction bus bandwidth requirements, the resulting architecture is rapidly limited by the von Neumann bottleneck, and computer performance is limited by the speed at which instructions are fetched from memory.
If the memory bus is 64-bit wide, a single 64-bit instruction, two 32-bit instructions, four 16-bit instructions, or eight 8-bit instructions can be fetched within each memory cycle. 32-bit instructions should be twice as useful as 16-bit instructions. Because it cuts half the instruction bandwidth.
The main purpose of instruction set design is to reduce instruction redundancy. In general, an optimized and efficient instruction set takes advantage of instruction and data locality. The simplest instruction optimization was done long ago. For most computer programs, the most promising next instruction is the sequential next instruction in memory. Therefore, instead of any instruction with the following instruction field, most instructions assume that the next instruction is the current instruction + 1. It is possible to build an architecture with 0 bits for the source and 0 bits for the destination.
Stack architecture Stack architecture Computers are also known as zero operand architectures. The stack architecture performs all operations based on the contents of the pushdown stack. Two-operand operation requires that both operands be present on the stack. When performing an operation, both operands are popped off the stack, the operation is performed, and the result is pushed back onto the stack. A stack architecture computer can have very short arithmetic code. This is because the source and destination are implied to be on the stack.
Most programs require the contents of global registers that are not always available on the stack when needed. Attempts to minimize this occurrence include stack indexing, which allows access to operands other than those at the top of the stack. Stack indexing results in larger instructions with additional arithmetic code bits, or requires additional action to place the stack index value on the stack itself. Sometimes one or more additional stacks are defined. A better but less optimal solution is the combination stack / register architecture.
Stack architecture behavior is a method that goes against obvious optimizations and is often redundant. For example, each POP and PUSH operation can cause time-consuming memory operations when the stack is manipulated in memory. In addition, stack operations can consume operands that may be needed immediately for the next operation, which requires duplicate operands due to the possibility of yet another memory operation. For example, multiplying all the elements of a one-dimensional array by 15.
In one stack architecture, this is done by: 1. Push the start address of the array 2. DUPLICATE the address (so we have an address to store the results in the array) 3. DUPLICATE the address (so we have an address to read from the array) 4. PUSH INDIRECT (PUSH the contents of the array position indicated by the top of the stack) 5.PUSH 15 6. MULTIPLY (15 times the contents of the array we read in the third row) 7. SWAP (get the first array address on the stack for the next instruction) 8.POP INDIRECT (POP the multiplication result, return it to the array and store it.) 9. INCREMENT (Points for the next array item.) 10. Go to step 2 until the array is complete. The loop counter on line 9 requires additional parameters. In some architectures, this parameter is stored on other stacks.
In a hypothetical register / accumulator architecture, this example is realized by: 1. STORE POINTER the start address of the array 2.READ POINTER (Read the contents of the address shown in the accumulator.) 3.MULTIPLY 15 4.STORE POINTER (Store the result in the indicated address.) 5. INCREMENT POINTER 6. Go to the second row until the array is complete.
In the above example, compare 9 steps for a stack architecture vs. 5 steps for a register architecture. Furthermore, the stack operation has at least three possible opportunities due to extra memory access due to the stack operation. Loop control in a hypothetical register / accumulator architecture can be easily handled within a register.
The stack is useful for evaluating expressions and is used in this way by most compilers. Stacks are also useful for nested actions, such as function calls. Most C compilers make function calls on the stack. However, without supplementing general purpose storage, the stack architecture requires a lot of additional data transfer and manipulation. For optimization, stack PUSH and POP operations must also be distinguished from mathematical and logic operations. However, as can be seen from the above example, the stack is particularly inefficient when loading and storing iterative data. This is because the array address is consumed by PUSH INDIRECT and POP INDIRECT.
<p> In one aspect, the invention is a system comprising (a) multiple parallel processors on a single chip and (b) computer memory located on the chip and accessible by each of the processors. Each of the processors can operate to handle the de minimis instruction set, and each of the processors has its own local cache for each of at least three specific registers in the processor.</p><p> In various embodiments, (1) each size of the local cache is equal to one row of random access memory on the chip. (2) At least three specific registers with associated caches include an instruction register, a source register, and a destination register. (3) The de minimis instruction set consists of seven instructions. (4) Each of the processors can operate to process a single thread. (5) The accumulator is an operand for all instructions except the increment instruction. (6) The destination for each instruction is always the operand register. (7) Three registers are auto-increment and three registers are auto-decrement. (8) Each instruction requires only one clock cycle to complete. (9) The instruction set does not include BRANCH and JUMP instructions. (10) Each instruction has a maximum length of 8 bits. (11) A single master processor plays a role in managing each of the parallel processors.</p><p> In another aspect, the invention is a system comprising (a) multiple parallel processors on a single chip and (b) computer memory located on the chip and accessible by each of the processors. Thus, each of the processors can operate to process an instruction set optimized for thread-level parallelism.</p><p> In various embodiments, (1) each of the processors can operate to handle the de minimis instruction set. (2) Each processor has its own local cache for at least three specific registers in the processor. (3) Each size of the local cache is equal to one row of random access memory on the chip. (4) At least three specific registers include an instruction register, a source register, and a destination register. (5) The de minimis instruction set consists of seven instructions. (6) Each of the processors can operate to process a single thread. (7) A single master processor plays a role in managing each of the parallel processors. (8) The de minimis instruction set contains a minimum set of instruction extensions to optimize the operation of the processor and promote the efficiency of the software compiler.</p><p> In another aspect, the invention is a thread-level parallel processing method using multiple parallel processors, master processors, and computer memory on a single chip, each of which processes a de minimis instruction set. It can operate to handle a single thread, (a) assigning a local cache to each of the three specific registers in each of multiple processors, and (b) a single thread. To process, the step of allocating one of multiple processors, (c) the step of processing each assigned thread by the processor, and (d) processing the result from each thread processed by the processor. It has steps to (e) deallocate one of the processors after the thread has been processed, and (f) the de minimis instruction set optimizes processor management. Contains a minimal set of instructions for.</p><p> In various embodiments, the de minimis instruction set consists of seven basic instructions, and the instructions in the de minimis instruction set are up to 8 bits in length. The de minimis instruction set may contain a set of extended instructions beyond the seven basic instructions. This helps optimize the internal behavior of the TOMI CPU and the execution of software program instructions executed by the TOMI CPU, and optimizes the behavior of the software compiler for the TOMI CPU. Embodiments of the present invention with multiple TOMI CPU cores may have a limited set of processor management instructions used to manage multiple CPU cores.</p><p> In another aspect, the invention is a system comprising: (a) Multiple parallel processors mounted on a memory module, (b) an external memory controller, and (c) a general-purpose central processing unit. Here, each of the parallel processors can operate to process an instruction set optimized for thread-level parallelism.</p><p> In various embodiments, (1) each of the parallel processors de It can operate to handle the minimis instruction set, and (2) one or more bits allocated in the memory mode register can operate to enable or disable one or more of the parallel processors. (3) The memory module is a dual inline memory module, (4) each of the processors can operate to handle a single thread, and (5) multiple threads are through shared memory. Data is shared, (6) multiple threads share data by one or more shared variables, (7) memory modules are one or more of DRAM, SRAM, and flash memory, and (8) ) At least one parallel processor is considered a master processor, the other parallel processors are considered slave processors, (9) each processor has a clock speed, and each processor other than the master processor has performance or power. It can operate to have a processor clock speed tuned to optimize consumption, (10) each processor can operate to be considered a master processor or slave processor, and (11) a master processor. Requests processing by some slave processors, waits for output from some slave processors, and combines the outputs, (12) the master processor receives output from each of several processors. And by combining the outputs from several processors, (13) allowing one or more parallel processors to be shut down, low power consumption is provided, and (14) each of the parallel processors, It is accompanied by a program counter and can be operated to be stopped by writing all 1s to the program counters associated with parallel processors.</p><p> In another aspect, the invention comprises a plurality of parallel processors embedded in a die of dynamic random access memory (DRAM), the plurality of parallel processors communicating with and parallel to an external memory controller and an external processor. Each of the processors is a system characterized in that it can operate to process an instruction set optimized for thread-level parallelism.</p><p> In various other embodiments, (1) the die is packaged with a DRAM pinout, (2) the parallel processor is mounted on a dual inline memory module, and (3) the system is a processor. It operates as a DRAM unless enabled by a DRAM mode register, (4) an external processor can operate to transfer data and instructions from the associated permanent storage device to the DRAM, and (5) permanent storage. The device is a flash memory, and (6) the external processor can operate to provide an input / output interface between the parallel processor and the external device.</p><p> In another aspect, the present invention is a system such as: It has (a) multiple processors on a single chip and (b) computer memory located on the chip and accessible by each of the processors, each of which has a de minimis instruction set. It is operational to handle, and each processor has its own local cache for each of at least three specific registers in the processor.</p><p> In various other embodiments, (1) each size of the local cache is equal to one row of random access memory on the chip, and (2) each processor is the internal data of the random access memory on the chip. Accessing the bus, the internal data bus has the width of one line of random access memory, (3) the width of the internal data bus is 1024, 2048, 4096, 8192, 16328, or 32656 bits, (4) The width of the internal data bus is an integral multiple of 1024 bits, and (5) the local cache dedicated to each of at least three specific registers in the processor is either filled or filled within one memory read or write cycle. It can be operated to be erased, (6) the de minimis instruction set basically consists of seven basic instructions, and (7) the basic instruction set consists of ADD, XOR, INC, AND, STOREACC, LOADACC, and Each instruction in the (8) de minimis instruction set, including the LOADI instruction, is up to 8 bits in length and (9) de The minimis instruction set has multiple instruction extensions to optimize the execution of the instruction sequence on the processor, and moreover, such instruction extensions consist essentially of less than 20 instructions (10). ) Each instruction extension has a maximum length of 8 bits, (11) de The minimis instruction set has a set of instructions for selectively controlling multiple processors on the chip, (12) each processor control instruction is at most 8 bits in length, and (13). ) Multiple processors are manufactured on the chip with computer memory placed on the chip using a semiconductor manufacturing process designed for monolithic memory devices. (14) The semiconductor manufacturing process is a metal-to-layer less than four layers. Using connections, (15) semiconductor manufacturing processes use less than three layers of metal interconnects, and (16) integration of multiple processors into computer memory circuits is less than 30% of the chip die size. (17) Integration of multiple processes into the computer memory circuit resulted in an increase of less than 20% of the chip die size, (18) Into the computer memory circuit of multiple processes. Integration results in less than 10% increase in chip die size, and (19) integration of multiple processes into computer memory circuits results in less than 5% increase in chip die size. (20) Less than 250,000 transistors are used to make each processor on the chip, and (21) the chip is manufactured using a semiconductor manufacturing process with less than four layers of metal interconnection, (22). Each of the processors can operate to handle a single thread, (23) the accumulator is an operand for all basic instructions except increment instructions, and (24) for each basic instruction. The destination is always the operand register, (25) three registers are auto-increment, and three registers are auto-decrement, (26) each basic instruction requires only one clock cycle to complete. , (27) The instruction set does not have BRANCH and JUMP instructions, and (28) a single master processor is responsible for managing each of the parallel processors.</p><p> In another aspect, the present invention is a system such as: It has (a) multiple parallel processors on a single chip and (b) computer memory located on the chip and accessible by each of the processors, each of which is thread-level parallel. It can operate to process an instruction set optimized for processing, and each processor accesses the internal data bus of computer memory on the chip, which is wider than one line of memory. Not wide.</p><p> In various embodiments, (1) each of the processors is capable of operating to handle the de minimis instruction set, and (2) each of the processors is local to each of at least three specific registers in the processor. It has a cache, (3) each size of the local cache is equal to one line of computer memory on the chip, (4) at least three specific registers are instruction registers, source registers, and destination registers. Including, (5) the de minimis instruction set basically consists of seven basic instructions, (6) the basic instruction set contains ADD, XOR, INC, AND, STOREACC, LOADACC, and LOADI instructions, and (7) Each instruction in the instruction set is up to 8 bits in length, (8) each of the processors can operate to handle a single thread, and (9) a single master processor, It is responsible for managing each of the parallel processors, (10) de The minimis instruction set has multiple instruction extensions to optimize the execution of the instruction sequence on the processor, and such instruction extensions have less than 20 instructions (11). Each instruction extension is up to 8 bits in length, and (12) the de minimis instruction set has a set of instructions to selectively control multiple processors on the chip, (12) 13) Each processor control instruction can be up to 8 bits in length, and (14) multiple processors, along with computer memory placed on the chip using a semiconductor manufacturing process designed for monolithic memory devices. It can be manufactured on a chip.</p><p> In another aspect, the invention is a method of thread-level parallel processing utilizing multiple parallel processors, master processors, and computer memory on a single chip, each of which is a de minimis instruction set. It can operate to process a single thread and has the following steps: (a) the step of allocating the local cache to each of the three specific registers in each of the multiple processors, and (b) the step of allocating one of the multiple processors to handle a single thread. , (C) the step of processing each assigned thread by the processor, (d) the step of processing the result from each thread processed by the processor, and (e) multiple processors after the thread has been processed. The step of unassigning one of them.</p><p> In various embodiments, (1) the de minimis instruction set basically consists of seven basic instructions, and (2) the basic instructions have ADD, XOR, INC, AND, STOREACC, LOADACC, and LOADI instructions. (3) The de minimis instruction set has a set of instructions for selectively controlling multiple processors, and (4) each processor control instruction has a maximum length of 8 bits. Yes, method (5) further has steps for each processor to access computer memory using the memory's internal data bus, which is the width of one line of memory on the chip, ( 6) Each instruction in the de minimis instruction set has a maximum length of 8 bits.</p><p> In another aspect, the present invention is a system such as: It has multiple processors embedded in a memory chip that are (a) compatible with the packaging and pin layout of electronic industry standard devices for memory devices, and (b) one or more of the processors. It can be booted by the information sent to the memory mode registers of the memory chip, where the memory chips work with the behavior of industrial standard memory devices, unless one or more of the processors boot through the memory mode registers. Functionally compatible.</p>
<figref num="1">An exemplary TOMI architecture in one embodiment is shown.</figref><figref num="2">An exemplary instruction set is shown.</figref><figref num="2A">Indicates a forward branch in operation.</figref><figref num="2B">An instruction map for an exemplary TOMI instruction set is shown.</figref><figref num="2C">An exemplary set of multiple TOMI processor management instruction extensions is shown.</figref><figref num="2D">Shows the clock programming circuit for the TOMI processor.</figref><figref num="2E">An exemplary set of instruction extensions is shown.</figref><figref num="3">It shows the valid addresses of various addressing modes.</figref><figref num="4">It shows how a data path can be easily created from 4 to 32 bits.</figref><figref num="5">An exemplary local cache is shown.</figref><figref num="6">An exemplary cache management status is shown.</figref><figref num="7A">An embodiment of an additional processing function configured to take advantage of the wide system RAM bus is shown.</figref><figref num="7B">An exemplary circuit for interleaving data lines for two memory banks accessed by a TOMI processor is shown.</figref><figref num="8">An exemplary memory map is shown.</figref><figref num="9">An exemplary processor utilization table is shown.</figref><figref num="10">It shows the three components of processor allocation.</figref><figref num="10A">An embodiment of a plurality of TOMI processors on a DIMM package is shown.</figref><figref num="10B">It shows an embodiment of a plurality of TOMI processors on a DIMM package interfaced with a general purpose CPU.</figref><figref num="10C">An exemplary TOMI processor initialization for an embodiment of multiple TOMI processors is shown.</figref><figref num="10D">Shows the use of memory mode register 5 for TOMI processor initialization.</figref><figref num="10E">An exemplary TOMI processor operating phase diagram is shown.</figref><figref num="10F">An exemplary processor-to-processor communication circuit design is shown.</figref><figref num="10G">An exemplary hardware implementation is shown to identify the TOMI processor to perform the work.</figref><figref num="10H">An exemplary processor arbitration diagram is shown.</figref><figref num="11">Illustrative factoring is shown.</figref><figref num="12">An exemplary system RAM is shown.</figref><figref num="13A">An exemplary plan view is shown for a monolithic array of 64 processors.</figref><figref num="13B">Other exemplary floor plans for a monolithic array of TOMI processors are shown.</figref><figref num="14A">An exemplary plan view for the TOMI Peripheral Controller Chip (TOMIPCC) is shown.</figref><figref num="14B">It shows an exemplary design for a mobile phone language translation routine application using TOMIPCC.</figref><figref num="14C">It demonstrates an exemplary design for memory-centric database applications using TOMIPCC and multiple TOMI DIMMs.</figref><figref num="15A">The highest level wiring diagram for an exemplary embodiment of a 32-bit TOMI processor is shown.</figref><figref num="15B">The highest level wiring diagram for an exemplary embodiment of a 32-bit TOMI processor is shown.</figref><figref num="15C">The highest level wiring diagram for an exemplary embodiment of a 32-bit TOMI processor is shown.</figref><figref num="15D">The highest level wiring diagram for an exemplary embodiment of a 32-bit TOMI processor is shown.</figref><figref num="15E">The explanation of the signal for the wiring diagram shown in FIGS. 15A to 15D is shown.</figref>
The TOMI architecture of at least one embodiment of the present invention preferably uses minimal logic capable of operating as a general purpose computer. Priority is given to the most common actions. Most behaviors are visible, regular, and available for compiler optimization.
In one embodiment, as shown in Figure 1, the TOMI architecture is a variant on the accumulator, register, and stack architecture. In this embodiment 1. Similar to the accumulator architecture, the accumulator is always one of the operands, except for the increment instruction. 2. Similar to the register architecture, the destination is always one of the operand registers. 3. Accumulators and program counters are also in register space and can therefore be manipulated. 4. Three special registers are auto-increment and auto-decrement, which help create I / O stacks and streams. 5. All instructions are 8 bits long, simplifying instruction decoding and increasing speed. 6. There are no branch (BRANCH) or jump (JUMP) instructions. 7. As shown in Figure 2, there are only seven instructions that allow the selection of 3-bit operators from 8-bit instructions.
Some advantages of the preferred embodiments include: 1. All movement is not suppressed by the equivalent required by the pipeline, but at the maximum speed allowed by the logic. Logical operations are the fastest. Mathematical operations are the next fastest. The slowest operation requires memory access. 2. The architecture is proportional to any data width limited only by package pins, adder carry times, and usefulness. 3. The architecture is close to the minimum possible functionality required to perform all the operations of a general purpose computer. 4. The architecture is very transparent and very regular, and most of the behavior is available in the optimizing compiler.
The architecture is designed to be easy enough to be replicated multiple times on a single monolithic chip. One embodiment embeds multiple copies of memory and a monolithic CPU. For 32-bit CPUs, most gates can be realized with less than 1,500 gates that define registers. Almost 1,000 TOMI CPUs in a preferred embodiment can be implemented using as many transistors as a single Intel Pentium® 4.
The TOMI CPU's reduced instruction set is taken into account to perform the operations required for a general purpose computer. The smaller the instruction set for one processor, the more efficiently it works. TOMI CPUs are designed with a very small number of instructions compared to modern processor architectures. For example, an Intel Pentium processor has 286 instructions, an Intel Itanium Montecito processor has 195 instructions, and a Strong ARM processor has 127 instructions, compared to one embodiment of the TOMI CPU having 25 instructions. It has instructions, and the IBM Cell processor has more than 400 instructions.
The basic set of instructions for TOMI CPUs is simplified and designed to run within a single system clock cycle, as opposed to the 30 clock cycles required by the latest generation Pentium processors. ing. The TOMI CPU architecture is a "pipelineless" architecture. This architecture and single clock cycle instruction execution significantly reduces or eliminates stalls, dependencies, and wasted clock cycles found in other parallel processing or pipeline architectures. Because the basic instruction only requires a single clock cycle to execute, and the clock speed increases (clock cycle time decreases), the execution result is due to complex mathematical instructions (eg ADD). The time required to convey through the transistor gate of the circuit can reach the limit of a single clock cycle. In such cases, it may be optimal to allow two clock cycles for the execution of a particular instruction so as not to slow down the execution of faster instructions. This depends on optimizations for the system clock speed, manufacturing process, and circuit layout of the CPU design.
TOMI's simplified instruction set allows 32-bit TOMI CPUs to be built with less than 5,000 transistors (without cache). A top-level schematic of one embodiment of a single 32-bit TOMI CPU is shown in Figures 15A-15D, and a description of the signals is shown in Figure 15E. Including the cache and associated decoding logic, a 32-bit TOMI CPU can be built with 40,000 to 200,000 transistors (depending on the size of the CPU cache). In comparison, the latest generation of Intel Pentium microprocessor chips requires 250,000,000 transistors. Traditional microprocessor architectures (Intel Pentium, Itanium, IBM Cell, and StrongARM, to name a few) required a huge and growing number of transistors to achieve increased processing power. The TOMI CPU architecture denies the progress of this industry by using a very small number of transistors per CPU core. TOMI The small number of transistors for the CPU offers a number of advantages.
Due to the compact size of TOMI CPUs, multiple CPUs can be built on the same silicon chip. This also allows multiple CPUs to be built on the same chip as the main memory, eg DRAM, at a small additional manufacturing cost, in addition to the manufacturing cost of the DRAM chip itself. Therefore, multiple TOMI CPUs can be placed on a single chip for parallel processing with only a minimal increase in die size and manufacturing costs for the DRAM chip. For example, 512MB of DRAM has nearly 700,000,000 transistors. 64 TOMI CPUs (assuming 200,000 transistors for a single TOMI CPU) only add 12,800,000 transistors for any DRAM design. For 512MB of DRAM, 64 TOMI CPUs increase the die size by less than 5%.
TOMI CPUs are designed to be manufactured using existing inexpensive commodity memory manufacturing processes, for example for DRAM, SRAM, and flash memory devices. The small number of transistors for a TOMI CPU is that the CPU is a large microprocessor chip that utilizes metal interconnects of eight or more layers (eg Intel). It can be made in a small area and silicon using an inexpensive semiconductor manufacturing process with two layers of metal interconnection, rather than the complex and expensive manufacturing process or other logical process used to manufacture Pentium). It means that they can be easily interconnected within. Modern DRAM and other essentials memory chips are more due to the metal interconnection of fewer layers (eg 2) to achieve lower manufacturing costs, more product production, and higher product yields. It utilizes a simple, cheaper cost semiconductor manufacturing process. Essentials Semiconductor manufacturing processes for memory devices are typically characterized by the operation of devices with low leakage current. On the other hand, the processes used to build modern microprocessors strive for high speed and high performance characteristics rather than low transistor level leakage current values. The TOMI CPU's capabilities, efficiently realized by the same manufacturing process used for DRAM and other memory devices, are TOMI. The CPU can be embedded in an existing DRAM chip (or other memory chip) to take advantage of the low cost, high yield memory chip manufacturing process. This is also because TOMI CPUs have the same packaging and device pin layout (eg, conforming to JEDEC standards for memory devices), manufacturing equipment, test equipment, and DRAM and other memory chips currently in industry. It provides the advantage that it can be manufactured using the test vector used in. Conversely, embedding DRAM memory in a conventional microprocessor chip works in the opposite direction. Because microprocessor chips expose memory circuits to the high levels of electrical noise and heat generated by the operation of microprocessors, they also expose expensive and complex logic manufacturing processes with eight or more layers of metal interconnects. This is because it is manufactured by using it. This in turn affects the type, dimensions, and functionality of the memory embedded in the processor chip. The result is higher cost, lower yield, higher power consumption, smaller memory, and ultimately lower performance microprocessors.
Another advantage of the preferred embodiment is that the TOMI CPUs are small enough (and require very little power) so that they can physically reside next to the DRAM (or other memory) circuit. It also allows the CPU to access the ultra-wide internal DRAM data bus. In modern DRAM, this bus is 1024, 4096, or 8192 bits wide (or an integral multiple thereof), which also typically corresponds to the width of one row of data in a databank in a DRAM design. (By comparison, the Intel Pentium data bus is 64-bit and the Intel Itanium bus is 128-bit wide.) The internal cache of the TOMI CPU can be sized to fit the row size of the DRAM, so the CPU cache is It can be filled (or erased) within a single DRAM memory read or write cycle. The TOMI CPU uses an ultra-wide internal DRAM data bus as the data bus for the TOMI CPU. TOMI CPU cache is TOMI It may be designed to reflect the design of DRAM row and / or column latch circuits for efficient layout and circuit operation, including data transfer to the CPU cache.
Another advantage of the preferred embodiment is the low level of electrical noise produced by the TOMI CPU due to the small number of transistors. This is because the CPU uses an ultra-wide internal DRAM data bus to access memory rather than the constantly driven I / O circuits that access off-chip memory for data. The on-chip CPU cache takes into account immediate access to the data for processing that minimizes the need for off-chip memory access.
The design objective of the processor architecture is to maximize processing capacity and speed, while minimizing the power required to achieve that processing speed. The TOMI CPU architecture is a fast processor with extremely low power consumption. The power consumption of the processor is directly related to the number of transistors used in the design. The small number of transistors for a TOMI CPU minimizes its power consumption. The simplified and efficient instruction set also allows the TOMI CPU to reduce its power consumption. In addition, access to the TOMI CPU cache and on-chip memory using the wide internal DRAM data bus eliminates the need to constantly drive I / O circuits to access off-chip memory. A single TOMI CPU operating at a clock speed of 1GHz consumes nearly 20 to 25 milliwatts of power. In contrast, the Intel Pentium 4 processor requires 130 watts at 2.93GHz and Intel The Itanium processor requires 52 watts at 1.6 GHz, the StrongARM processor requires 1 watt at 200 MHz, and the IBM Cell processor requires 100 watts at 3.2 GHz. It is well known that the generation of heat within a processor is directly related to the amount of power required by the processor. The extremely low power TOMI CPU architecture eliminates the need for fans, large heat sinks, and new cooling mechanisms found in current microprocessor architectures. At the same time, the low power TOMI CPU architecture enables new low power battery and solar energy applications.
Instruction set Seven of the instruction sets are shown in Figure 2 with their bit mappings. Each instruction preferably consists of a single 8-bit word.
Addressing mode Figure 3 shows the valid addresses for the various addressing modes.
The addressing mode is as follows. immediate register Register indirect Register indirect auto increment Register indirect automatic decrement
Special case Both register 0 and register 1 point to the program counter (PC). All operations with register 0 (PC) as the operand are conditional on the accumulator carry bit (C) being equal to 1. If C = 1, the old value of the PC is swapped to the accumulator (ACC). All operations that have register 1 (PC) as an operand are unconditional.
In an alternative embodiment, the write operation with register 0 (PC) as the destination is conditional on the carry bit (C) being equal to 0. If C = 1, no operation is performed. If C = 0, the accumulator (ACC) value is written to the PC and program control shifts to the new PC address. The write operation with register 1 (PC) as the destination is unconditional. The accumulator (ACC) value is written to the PC and program control shifts to the new PC address.
A read operation with register 0 as the source loads the value of PC + 2. In this way, the address at the beginning of the loop can be read and stored for later use. In most cases, the loop address is pushed onto the stack (S). A read operation with register 1 as the source loads the value pointed to by the next full word addressed by the PC. In this way, 32-bit immediate operands can be loaded. The 32-bit immediate operand must be word aligned, but the LOADACC instruction may be at any byte position in the 4-byte word immediately preceding the 32-bit immediate operand. Following the read execution, the PC is incremented so that it addresses the first word-aligned instruction following the 32-bit immediate operand.
No branch Branching and jumping behavior is usually a challenge for CPU designers. Because they require many bits of valuable arithmetic code space. The branch function can be triggered by loading the desired branch address into the ACC using LOADACC, xx and then performing the branch using the STOREACC, PC instruction. Branching is done depending on the state of C when it is saved in register 0.
skip Skipping can be caused by running INC, PC. Execution requires two cycles, one to complete the current PC increment cycle and one for INC. Skipping is done depending on the state of C when register 0 is incremented.
Relative branch Relative branching can be triggered by loading the desired offset into the ACC and then executing the ADD, PC instructions. Relative branching is done depending on the state of C when it is added to register 0.
Branch forward Forward branching is more useful than backward branching. This is because the position of the backward branch required for the loop is easily captured by saving the PC during the first program step through the beginning of the loop.
A more efficient forward branch than a relative branch can be triggered by loading the least significant bit of the branch endpoint into the ACC and then storing it on the PC. Since the PC can be accessed both conditionally and unconditionally, depending on the use of register 0 or register 1, forward branching also depends on the selection of the PC register as the destination operand (register 0 or register 1). Depending, it can be conditional or unconditional.
For example LOADI, # 1C STOREACC, PC
If the most significant bit of ACC is zero, only the least significant 6 bits are transferred to the PC register. If the least significant bit of the current PC register is less than the ACC value to be loaded, the most significant bit of the register remains immutable. If the least significant 6 bits of the current PC register are greater than the ACC value to be loaded, the current PC register is incremented and starts at bit 7.
This effectively allows branching up to 31 instructions ahead. This method of forward branching should be used whenever possible. This is because not only does it require only two instructions for every three instructions for relative branching, but it also does not require a path through the adder, which is one of the slowest operations. Figure 2A shows a forward branch in operation.
loop The beginning of the loop can be saved using LOADACC and PC. The pointer to the beginning of the resulting loop syntax is either stored in a register or pushed to one of the autoindexing registers. At the end of the loop, the pointer is retrieved by LOADACC, EA and restored to the PC using STOREACC, PC, which causes a backward loop. The loop is made depending on the state of C by saving to register 0, which causes a conditional backward loop.
Self-modifying code It is possible to write self-modification code using STOREACC and PC. The instruction is triggered or fetched into the ACC and stored on the PC to be executed as the next instruction. This technique can be used to create CASE syntax.
Suppose an in-memory jump table array consisting of N JUMPTABLE addresses and a base address. For convenience, JUMPTABLE is in low memory 20, so its address can be generated by LOADI or LOADI followed by one or more right shift ADDs, ACCs.
Suppose the index to the jump table is in ACC, and the base address of the jump table is in a general-purpose register named JUMPTABLE. Add the ADD, JUMPTABLE index to the base address of the jump table. LOADACC, (JUMPTABLE) Load the indexed address Perform STOREACC, PC jump.
If low memory starting at 0000 is allocated for system calls Each system call is executed as follows. Where SPECIAL_FUNCTION is the name of the immediate operand 0-63. LOADI, SPECIAL_FUNCTION Load system call number LOADACC, (ACC) Load the address of the system call STOREACC, jump to PC function
Shift right The basic architecture does not assume right shift operations. If such an operation is required, the solution of the preferred embodiment is to specify one of the general purpose registers as a "right shift register". STOREACC, RIGHTSHIFT stores an ACC that shifts a single position to the "right shift register" to the right. Here, the value can be read by LOADACC and RIGHTSHIFT.
Architectural scalability The TOMI architecture preferably features 8-bit instructions, but the data width need not be limited. Figure 4 shows how easily data paths of any width, 4 to 32 bits, can be created. Performing wider data processing only requires increasing the width of the register set, internal data path, and ALU for the desired width. The upper limit of the data path is only limited by the carry propagation delay of the adder and the budget of the transistor.
A preferred TOMI architecture is implemented as a von Neumann memory configuration for simplicity, but can also be implemented by a Harvard architecture (with separate data and instruction buses).
Common mathematical operations Two's complement mathematics can be done in several ways. All general-purpose registers are preset as "1s" and named ALLONES. The operand is assumed to be in a register named OPERAND. LOADACC, ALLONES XOR, OPERAND INC, OPERAND The complement of 2s remains in OPERAND.
Common compiler structure Most computer programs are generated by a compiler. Therefore, a practical computer architecture should conform to a common compiler structure.
The C compiler usually maintains a stack for passing parameters to function calls. You can use the S, X, or Y registers as stack pointers. Function calls use, for example, STOREACC, (X) + to push parameters to one of the auto-indexing registers acting as a stack. When you enter a function, the parameters are popped into general purpose registers for use.
Stack relative addressing There may be more elements that have passed a function call than can be conveniently adapted to a general purpose register. For the following example, assume that the stack push action decrements the stack. If S is used as a stack register, to read the Nth item relative to the top of the stack, LOADI, N STOREACC, X LOADACC, S ADD, X LOADACC, (X)
Indexing to the array When you enter the array function, the array base address is placed in a general-purpose register named ARRAY. To read the Nth element in the array LOADI, N STOREACC, X LOADACC, ARRAY ADD, X LOADACC, (X)
Indexing to N-word element array From time to time, arrays are assigned to N-word wide elements. The base address of the array is placed in a general purpose register named ARRAY. To access the first word of the Nth element in a 5-word wide array LOADI, N STOREACC, X Store in temporary register Apply ADD, ACC 2 ADD, ACC Multiply 2 again = 4 ADD, X plus 1 = 5 LOADACC, ARRAY Add the base address of the ADD, X array LOADACC, (X)
Instruction set extension Other embodiments of the present invention include extensions to the seven basic instructions shown in FIG. The instruction set extension shown in Figure 2E helps further optimize the internal operation of the TOMI processor, software program instructions, and the software compiler for the TOMI processor.
SAVELOOP-This instruction pushes the current value of the program counter onto the stack. Saveloop is most likely to be executed at the beginning of the loop syntax. At the end of the loop, the stored program counter value is copied from the stack and stored in the program counter to perform a reverse jump to the beginning of the loop.
SHIFTLOADBYTE-This instruction shifts the 8 bits left in the ACC to the left, reads the 8-bit byte following the instruction, and puts it in the least significant 8 bits of the ACC. In this way, long immediate operands can be loaded with a series of instructions. For example, to load a 14-bit immediate operand LOADI, # 14 \\ Load the most significant 6 bits of the 14-bit operand. SHIFTLOADBYTE \\ Shifts the 8-bit position of 6 bits to the left and loads the next 8-bit value. CONSTANT # E8 \\ 8-bit immediate operand. The hexadecimal value as a result of ACC is 14E8.
LOOP-This instruction copies the top of the stack to the program counter. Loop is most likely to be executed at the end of the loop syntax following the execution of Saveloop to store the program counter at the beginning of the loop. When the loop is executed, the saved program counter is copied from the stack, stored in the program counter, and a reverse jump to the beginning of the loop is executed.
LOOP_IF-This instruction copies the top of the stack to the program counter. It executes a conditional loop based on the value of C. Loop_if is most likely executed at the end of the loop syntax following the execution of Saveloop to store the program counter at the beginning of the loop. If C = 0 when executing Loop_if, the saved program counter is copied from the stack, stored in the program counter, and a reverse jump to the beginning of the loop is executed. When C = 1, the program counter is incremented to point to the next sequential instruction.
NOTACC-Performs complement operation on each bit of ACC. If ACC = 0, set C to 1. Otherwise, set C to 0.
Rotate the ROTATELE FT8-ACC 8 bits to the left. At each rotation step, the MSB shifted from the ACC is shifted to the LSB of the ACC.
OR Performs a logical sum of ACC and the top value of the stack. Put the result in ACC. If ACC = 0, set C to 1. Otherwise, set C to 0.
OR OR ACK + -ACC and OR on the top value of the stack. Put the result in ACC. After the logical operation, the stack pointer S is incremented. If ACC = 0, set C to 1. Otherwise, set C to 0.
RIGHTSHIFT ACC-Shifts ACC to the right by a single bit. The LSB of ACC is shifted to C.
SETMSB-Sets the most significant bit of ACC. There is no change for C. This instruction is used in performing a signed comparison.
Local TOMI caching A cache is smaller in size and faster in access than main memory. The reduced access time and locality of program and data access allows for cached operations and increases the performance of TOMI processors suitable for many operations. From another perspective, the cache increases parallel processing performance by increasing the independence of the TOMI processor from the main memory. The relative performance of the cache to main memory and the number of cycles that the TOMI processor can perform before requesting a load or store from the cache to other main memory is an increase in performance due to the TOMI processor parallel method. Determine the amount of.
The TOMI local cache enhances performance gains with the TOMI processor parallel method. As shown in Figure 5, each TOMI processor preferably has three associated local caches. Instructions-related to PC Source-related to the X register Destination-related to Y register
Since the cache is associated with a particular register rather than a "data" or "instruction" fetch, the cache control logic is simplified and cache latency is significantly reduced. The optimal size of these caches depends on the application. A typical embodiment requires 1024 bytes for each cache. In other words, there are 1024 instructions and 256 32-bit words for the source and destination. At least two factors determine the optimal size of the cache. The first is the number of states that the TOMI processor can repeat before another cache load or store operation is required. The second is the cost of loading or storing cache from main memory associated with the number of TOMI processor execution cycles possible during main memory operation.
Embedding the TOMI processor in RAM In one embodiment, the wide bus connects a large embedded memory to the cache so that a load or store operation on the cache can occur faster. With a TOMI processor embedded in RAM, all cache loads or stores consist of a single memory cycle for a column of RAM. In one embodiment, the embedded memory is responding to the requests of 63 TOMI processors, so the cache load or store response time for one TOMI processor is the cache load or store for the other TOMI processors. Can be extended while is completed.
As shown in Figure 6, the cache is stored and loaded based on changes in the associated memory addressing registers X, Y, PC. For example, the full width of a PC register can be 24 bits. If the PC cache is 1024 bytes, the lower 10 bits of the PC define access in the PC cache. A cache load cycle is required when the PC is written so that there is a change in the upper 14 bits. The TOMI CPU associated with that PC cache will stop executing until the cache load cycle is complete, and the indicated instructions may be fetched from the PC cache.
Cache double buffering The secondary cache can be loaded in anticipation of a cache load request. The two caches are identical and are alternately selected and deselected based on the contents of the upper 14 bits of the PC. In the above example, when the upper 14 bits of the PC change to match that of the data pre-stored in the secondary cache, the secondary cache will be selected as the primary cache. The old primary cache then becomes the secondary cache. Since most computer programs grow linearly in memory, one embodiment of the invention always fetches the contents of the cache, the contents of the main memory indicated by the upper 14 bits of the current PC plus 1. Has a cache.
Adding a secondary cache reduces the amount of time the TOMI processor has to wait for memory data to be fetched from main memory when moving outside the boundaries of the current cache. Adding a secondary cache almost doubles the complexity of the TOMI processor. If the complexity is doubled for the optimal system, it must be offset by doubling the performance of the corresponding TOMI processor. Otherwise, two simpler TOMI processors without a quadratic cache could be realized with the same number of transistors.
Fast multiplication, floating point arithmetic, additional features Integer multiplication and all floating point operations require many cycles to perform, even with special purpose hardware. Therefore, these operations can take into account other processors rather than including them in the basic TOMI processor. However, a simple 16-bit x 16-bit multiplier can be added to the TOMI CPU (using less than 1000 transistors) to provide additional functionality and versatility to the TOMI CPU architecture.
Digital signal processing (DSP) operations often use highly pipelined multipliers that produce results on a cycle-by-cycle basis, even though full multiplication can require many cycles. For signal processing applications that repeat the same algorithm over and over, such a multiplier architecture is ideal and can be incorporated as a peripheral processor to the TOMI processor. But even if it was built directly into the TOMI processor, it would probably increase complexity and reduce overall performance. Figure 7A shows an example of an additional processing function configured to utilize the wide system RAM bus.
Access to adjacent memory banks The physical layout design of memory circuits within a memory chip is often designed so that the memory transistors are laid out in a large bank of memory cells. Banks are usually configured as equally sized rectangular areas and are arranged in two or more rows on the chip. The layout of memory cells in a large bank of cells can be used to speed up memory read and / or write access.
In one embodiment of the invention, one or more TOMI processors may be located between two rows of memory cell banks in a memory chip. Using the logic shown in Figure 7B, by enabling Select A or Select B, the row data lines of the two memory banks allow the TOMI processor to access memory bank A or memory bank B. Can be interleaved. In this way, the memory that can be addressed directly by a particular TOMI processor in the memory chip can be doubled.
TOMI interrupt strategy An interrupt is an external event to the normal sequential operation of a processor, which forces the processor to immediately change its operation sequence. An example of an interrupt is the completion of operation by an external device or an error condition by some hardware. Traditional processors quickly stop normal sequential operation, save the state of the current operation, and start performing some special operation to handle any event that triggered the interrupt, and the special operation. Do whatever it takes to restore the previous state and continue the sequential operation when is completed. Response time is a major measure of interrupt processing quality.
Interrupts pose some challenges for traditional processors. They make the execution time uncertain. They waste processor cycles to store and then restore state. They can complicate processor design and introduce delays that slow down any processor operation.
Immediate interrupt response is unnecessary for most processors, except for those that are directly linked to error handling and real-world activity.
In one embodiment of a multiprocessor TOMI system, only one processor has the main interrupt function. All other processors run uninterrupted until they complete some assigned work and stop on their own. Or it works until they are stopped by the coordinating processor.
Input / output (I / O) In one embodiment of the TOMI processor environment, a single processor is responsible for all interfaces to the outside world.
Direct memory access (DMA) control In one embodiment, the immediate response to the outside world in the TOMI processor system occurs via the DMA controller. The DMA controller transfers data from the external device to the internal data bus for writing to the system RAM when requested by the external device. The same controller also transfers data from the system RAM to an external device when requested. DMA requests have the highest priority for internal bus access.
Organizing an array of TOMI processors The TOMI processor of the preferred embodiment of the present invention is designed to be coupled with a significant number of replicated and additional processing functions on a monolithic chip, a very wide internal bus, and system memory. An exemplary memory map for such a system is shown in Figure 8.
The memory map for each processor spends the first 32 positions (1F in hexadecimal) relative to the local registers for that processor (see Figure 3). The rest of the memory map can be addressed by all processors through their cache registers (see Figure 6). The addressing capability of system RAM is limited only by the width of the three registers PC, X, and Y associated with the local cache. If the register is 24-bit wide, the total addressing capability is 4 megabytes, but there is no upper limit.
In one embodiment, 64 TOMI processors are implemented monolithically with memory. A single master processor is responsible for managing the other 63. When one of the slave processors is idle and the clock is not running, it consumes very little power and produces very little heat. At initialization, only the master processor is available. The master initiates fetching and executing instructions until the time the thread should start. Each thread is precompiled and loaded into memory. To start a thread, the master assigns this thread to one of the TOMI CPUs.
Processor operation Coordination of the TOMI processor's operation to do a good job is handled by the processor operation table shown in FIG. The coordinated (master) processor is preferably capable of performing the following functions: 1. Push the parameters calling for the slave processor onto the stack, including, but not limited to, the thread's execution address, source memory, and destination memory. 2. Start the slave processor. 3. Respond to slave processor thread completion events by responding to polls or interrupts.
Processor request The coordinated processor can request the processor from the working table. Returns the number of lowest processor with available_flag set to "0". The coordinating processor then sets the available_flag for the available processor to "1", which activates the slave processor. If the processor is not available, the request returns an error. As an alternative, processors can be assigned by the coordinating processor based on the priority level for the requested work to be performed. Techniques for allocating resources based on a priority scheme are well known in the prior art. Figure 10 shows three suitable components of processor allocation. The operation of starting the coordinate processor, the operation of the slave processor, and the result processing of the coordinate processor by the interrupt response.
Starting the slave processor in stages 1. The coordinating processor pushes parameters to the thread to run on its own stack. The parameters include: Thread start address, source memory for thread, destination memory for thread, and last parameter count. 2. The coordinated processor requires an available processor. 3. The processor allocation logic returns the numerically lowest slave processor number, or error, that sets its associated available_flag and clears its associated done_flag. 4. If an error is returned, the coordinating processor either retries the request until a slave processor is available or performs some special action to handle the error. 5. When the available processor number is returned, the coordinating processor clears the available_flag for the indicated processor. This action pushes the parameter_count number of stack parameters onto the stack of the selected slave processor. done_flag is cleared to zero. 6. The slave processor searches for the first stack item and transfers it to the slave processor's program counter. 7. The slave processor then fetches the memory column indicated by the program counter into the instruction cache. 8. The slave processor starts executing instructions from the beginning of the instruction cache. The first instruction is probably to find the parameter you are calling from the stack. 9. The slave processor executes a thread from the instruction cache. When the thread completes, it checks the status of its associated done_flag. If done_flag is set, wait until done_flag is cleared. This indicates that the coordinating processor has processed any previous results. 10. If the interrupt vector for the slave processor is set to -1, no interrupt will occur even if done_flag is set. Therefore, the coordinating processor polls so that done_flag is set.
When the coordinating processor detects that done_flag is set, it processes the result of the slave processor and probably reassigns the slave processor to do a new job. When the slave processor results are processed, the associated coordinating processor clears the associated done_flag.
If the interrupt vector for the slave processor is not equal to -1, setting the associated done_flag causes the coordinate processor to be interrupted and starts executing the interrupt handler at the interrupt vector address.
If the associated available_flag is also set, the coordinating processor can also read the return parameters pushed onto the slave processor's stack.
The interrupt handler processes the result of the slave processor and probably reassigns the slave processor to do a new job. When the result of the slave processor is processed, the interrupt handler running on the coordinate processor clears the associated done_flag.
11. When done_flag is cleared, the slave processor sets its associated done_flag and saves a new start_time. The slave processor may continue to work or may return to an available state. To return to the available state, the slave processor pushes a return parameter onto the stack, followed by setting the stack count and its available_flag.
TOMI processor management using memory mode registers One technique for implementing and managing multiple TOMI processors is to mount the TOMI processors on dual in-line memory modules (DIMMs), as shown in Figure 10A. TOMI / DIMMs can be included in a system consisting of an external memory controller and a general purpose CPU, such as a personal computer. FIG. 10B shows such a configuration. Mode registers are commonly found in DRAM, SRAM and flash memory. A mode register is a set of latches that can be written by an external memory controller that has nothing to do with memory access. Bits in memory mode registers are often used to identify parameters such as timing, refresh control, and output burst length.
One or more bits can be allocated in the memory mode register to enable or disable the TOMI CPU. For example, when a TOMI CPU is disabled by a mode register, the memory containing the TOMI CPU acts as normal DRAM, SRAM or flash memory. When the mode register enables TOMI CPU initialization, the sequence is executed as described in Figure 10C. Within this embodiment, a single processor is defined as the master processor. It is this processor that always boots first following the reset operation. After the initialization is complete, the master processor operates at full speed to execute the desired application program. DRAM, SRAM, or flash memory is inaccessible while the TOMI CPU is running. From time to time, the memory mode register may be instructed by an external memory controller to stop execution of the TOMI CPU. TOMI When the CPU is stopped, the contents of DRAM, SRAM, or flash can be read by an external memory controller connected to a general purpose CPU. In this way, the result can be passed to a general purpose CPU. Then, additional data or executable files may be written to DRAM, SRAM, or flash memory.
When the general purpose CPU completes reading or writing DRAM, SRAM, or flash memory, the external memory controller writes a stop-to-operation bit to the mode register, and the TOMI CPU continues execution from where it left off. Figure 10D shows a typical memory mode register from DRAM, SRAM, or flash memory. It then shows how that register is modified to control the TOMI CPU.
Processor clock speed adjustment The processor clock speed determines the processor power consumption. The TOMI architecture enables low power consumption by being able to shut down all but one processor. Furthermore, each processor other than the master processor can have its clock speed tuned to optimize performance or power consumption using the logic shown in Figure 2D.
Other embodiments of TOMI processor management Some computer software algorithms are cyclical. In other words, the primary function of the algorithm is to call itself. A class of algorithms known as "divide and conquer" is often implemented using cyclical techniques. Division and acquisition are applicable to data retrieval and classification, and to specific mathematical functions. It is possible to parallel such algorithms with those available with multiple processors, such as the TOMI architecture. In order to execute such an algorithm, one TOMI CPU must be able to pass work to and receive results from that CPU. Other embodiments of the TOMI processor allow any processor to be the master processor and any other available TOMI processor to be considered a slave processor. Starting and stopping TOMI processors, communication between processors, and management of independent and dependent threads are supported within the processor management of this embodiment.
Stop TOMI CPU The TOMI CPU can be stopped by writing all 1s to its PC. When the TOMI CPU is stopped, its clock is not running and it is not consuming power. Any TOMI CPU can write all 1s to its own PC.
Start TOMI CPU The TOMI CPU can start execution when all values other than 1 are written to the PC. The master processor has a value of 0 written to its PC when it is reset by a mode register as shown in Figure 10D.
Independent processor thread When multiple threads run on a single general purpose processor, they can rarely communicate and be very loosely coupled. Instead of running, returning results, and stopping, some threads can deliver results continuously and forever, forever. An example of such a thread is a network communication thread or a thread that reads a mouse device. The mouse thread operates continuously to deliver mouse position and click information to a shared memory area that can be polled or a callback routine that is immediately indicated.
Such independent threads are primarily used to simplify programming rather than accelerate performance. Similar threads can run on multiprocessor systems such as TOMI. The result can be delivered to the shared memory area. In some cases, communication can be achieved by shared variables.
In the TOMI architecture, shared variables can be more efficient than shared memory. This is because variables can avoid the need for a memory RAS cycle to load all rows into X_cache or Y_cache. An example of the use of variables is an input pointer to the receive buffer for the TOMI CPU that monitors network traffic. The network monitoring CPU increments the variable when data is received. The data consuming CPU sometimes reads variables and, when there is enough data, performs an operation to load a row of memory into X_cache or Y_cache. The received network data can then be read from the cache up to the value indicated by the shared variable.
Dependent processor thread Some algorithms (eg, those classified as split and acquisition) can achieve parallelism by running parts of the algorithm simultaneously on several processors and combining the results. In such a design, a single master processor requests work from several slave processors and waits while the slaves perform work in parallel. Thus, the master processor depends on the work completed by the slave processor.
When the slave processor's work is completed, the master processor reads the partial results and combines them into the final result. This feature allows the TOMI architecture to efficiently handle a class of algorithms known as "split and acquisition". Some of the more common and simple division and acquisition algorithms are search and classification.
Multiprocessor management instruction set extension A set of extensions to the basic TOMI instruction set allows independent and dependent thread management. These instructions are executed using some of the available NOOP code shown in Figure 2B. These management extension instructions are summarized in Figure 2C.
GETNEXTPROCESSOR-This instruction queries the processor availability table and loads an ACC with the number associated with the next available processor.
SELECTPROCESSOR-This instruction writes the ACC to the processor selection register. The processor selection register selects which processor is evaluated by TESTDONE and READSHAREDVARIABLE.
STARTPROCESSOR-This instruction writes to the PC of the processor selected by the processor selection register. This instruction is most likely to be executed when the master processor wants to start a stopped slave processor. The slave processor is stopped if its PCs are all 1. By writing the value to the slave processor's PC, the master processor causes the slave processor to start executing the program at the location of the written PC. If the operation is successful, the ACC contains the PC value written to the selected processor. If the operation is unsuccessful, the ACC contains -1. The most likely reason for the failed operation is that the selected processor was not available.
TESTDONE-This instruction tests the PC of the processor selected by the processor selection register and sets the C bit of the calling processor to "1" if PC = all 1. A loop for testing a PC in this way can be created as follows: LOADI, processorNumber SELECT PROCESSOR LOADACC, LOOPADDRESS TOP TESTDONE STOREACC, PC_COND // Loop to TOP until the selected processor stops at PC = all 1.
TESTAVAILABLE-This instruction tests the "available" bits in the processor allocation table bits for the processor selected by the processor selection register and calls if the selected processor is available. Set the C bit of the processor. A loop for testing availability for further work in this way can be created as follows. LOADI, processorNumber SELECT PROCESSOR LOADACC, LOOPADDRESS TOP TESTAVAILABLE STOREACC, PC_COND // Loop to TOP until the selected processor is available.
SETAVAILABLE-This instruction sets the "available" bits in the processor allocation table for the processor selected by the processor selection register. This instruction is most likely to be executed by one processor that requires the other processor to do the job as shown in Figure 10E. When a working processor completes its work, it stops by setting all its PCs to 1. The requesting processor periodically performs a TEST DONE on the working processor. When the working processor completes its work, the requesting processor reads the result by shared memory location or shared variable. When the result is read, the processor that was doing the work is available to reassign other work. And since the requesting processor uses SETAVAILABLE, other processors may require it to do more work.
READSHAREDVARIABLE-This instruction reads a processor shared variable selected by the processor select register. This instruction is most likely to be executed by one processor that has requested the other processor to do the job. Shared variables can be read by any processor to determine the progress of assigned work. For example, a working processor may be assigned to process data received from a high-speed network. Shared variables indicate the amount of data that was read and available to other processors. Each processor contains shared variables. The shared variable can be read by any other processor, but the shared variable can only be written by its own processor.
STORESHAREDVARIABLE-This instruction writes the value of ACC to a shared variable of the processor executing the instruction. Shared variables can be read by any other processor. It is then used to convey state and results to other processors.
Interprocessor communication with a data ready latch Figure 10F shows one possible hardware implementation of communication between TOMI processors. One TOMI processor can use the SELECT PROCESSOR command to establish a connection between itself and another TOMI processor. This instruction establishes a logical connection that allows the selected and selected processors to exchange data using shared registers and the READSHAREDVARIABLE and STORE SHAREDVARIABLE commands.
The upper half of Figure 10F shows the logic for a processor to send data to another processor controlled by the status register. The lower half of Figure 10F shows the logic for a processor that receives data from other processors controlled by the status register.
The state of the shared register can be read into the C bits of the selected or selected processor. During operation, the processor writes data to its shared registers. This sets the relevant data ready flag. The connected processor reads its shared register until the C bit indicates that the associated data ready flag has been set. The read operation clears the ready flag, so the process can be repeated.
Arbitration of processor allocation As mentioned above, a TOMI processor can delegate work to another TOMI processor whose availability has been determined by GETNEXTPROCESSOR.
GETNEXTPROCESSOR determines if a processor is available. Available processors are those that are not currently performing work and do not retain the results of previous work that has not yet been retrieved by READSHAREDVARIABLE.
Figure 10G shows one hardware implementation to identify the available processors to which work can be delegated. Figure 10H shows an exemplary processor arbitration event. The process is as follows. 1. The requesting processor executes the GETNEXT PROCESSOR instruction. This pulls the "next available processor request" line down into the arbitration logic. 2. The arbitration logic generates the number corresponding to the TOMI CPU on the "next available processor" line. 3. The requesting processor executes the SELECT PROCESSOR instruction. It stores the number in the PSR (processor selection register) of the requesting processor. 4. The requesting processor then executes the START PROCESSOR instruction. It writes to the PC of the processor selected by the processor selection register. If the operation is successful, the number of selected processors is further stored in the arbitration logic to indicate that the selected processors are no longer available to be assigned to do the work. If the operation is unsuccessful, the reason is probably that the selected processor is not available. The requesting processor executes another GETNEXT PROCESSOR to find another available processor. 5. When the selected processor executes STORE SHAREDVARIABLE to make the result available, the arbitration logic is informed that the selected processor has the result waiting to be read. To. 6. When the selected processor is stopped by writing -1 to its PC, the arbitration logic is notified that the selected processor is available. 7. When the result of the selected processor is retrieved by READSHAREDVARIABLE, the arbitration logic is notified that the result of the selected processor has been read.
Memory lock The TOMI processor reads and writes system memory through those caches. Fully cached columns are read and written at once. Any processor can read any part of the system memory. Individual processors can lock columns of memory for their exclusive writes. This locking mechanism avoids memory write conflicts between processors.
Proposed application The parallel method effectively accelerates applications that are considered to be an independent part of the work for individual processors. One well-thought-out application is image manipulation for robot vision. Image manipulation algorithms include correlation, equalization, edge identification, and other behaviors. Many are performed by matrix operations. As shown in Figure 11, this algorithm is very often well thought out.
The image map illustrated in FIG. 11 shows 24 processors, including a processor assigned to manipulate image data for a rectangular subset of the entire image map.
FIG. 12 shows how the TOMI system RAM can be allocated in one embodiment. One block of system RAM holds the pixels of the image capture and the other block holds the processed result.
During operation, the coordinating processor allocates a DMA channel to transfer image pixels from the external source to the internal system RAM at regular intervals. The typical speed of image capture is 60 images per second.
The coordinating processor then pushes the address of the image map to be used by the X register, the address of the processed image to be used by the Y register, the parameter count of 2, and the address of the image processing algorithm. Enable 1. The coordinating processor then enables processors 2 to 25 as well. The processors continue to execute in parallel until the image processing algorithm is completed.
When the algorithm is complete, each processor sets its associated done_flag in the processor utilization table. The result is processed by the coordinating processor. It is either polling for completion or responding to interrupts on the "done" event.
Figures 13A and 13B are exemplary plan views for a monolithic array of 64 TOMI processors implemented using on-chip system memory. The floor plan can vary depending on the circuit design and memory circuit layout and overall chip architecture.
TOMI Peripheral Controller Chip (TOMIPCC) One embodiment of the TOMI architecture embeds one or more TOMI CPUs in a standard DRAM die. The resulting die is placed in a standard DRAM package with a standard DRAM pinout. Packaged components can be mounted on standard DRAM DIMMs (dual in-line memory modules).
In operation, this embodiment behaves like a standard DRAM except when the embedded TOMI CPU is enabled by a DRAM mode register. When the TOMI CPUs are enabled and running, they run programs loaded into DRAM by an external processor. The result of the TOMI CPU calculation is provided to the external processor through the shared DRAM.
In some applications, the external processor is provided by the PC. In other applications, specialized processors may be provided to perform the following functions for the TOMI DRAM chip: FIG. 14A shows an embodiment of such a processor, a TOMI peripheral controller chip (TOMIPCC). The functions of this processor are as follows. 1. Provide a mechanism for transferring data and instructions from the associated permanent storage device to the TOMI chip DRAM for use. In many systems, the permanent storage device can be flash RAM. 2. Provides an input / output interface between the TOMI chip and real-world equipment such as displays and networks. 3. Perform a small set of operating system features needed to coordinate TOMI CPU behavior.
Figure 14B shows a very small system using TOMIPCC. An example of a cellphone language translator consists of only three chips. That is, flash RAM, TOMIPCC, and a single TOMI CPU / DRAM. In this minimal application, a single TOMI CPU / DRAM communicates through a standard 8-bit DRAM I / O labeled D0-D7.
Flash RAM contains a dictionary of phonemes and syntax that defines the vocal language, in addition to the instructions needed to interpret and translate the vocal language from one format to the other.
TOMIPCC receives an analog speech language (or its equivalent), converts it into a digital representation, and submits it to TOMI DRAM for interpretation and translation. The resulting digitized speech is returned from the TOMI DRAM to the TOMIPCC, converted to an analog speech representation, and then output to the cellphone user.
Figure 14C shows a very large system using TOMIPCC. An example of such a system is a memory-centric database device (or MCDB). The MCDB system runs on all databases in fast memory instead of a very slow disk or its paging part in storage. Fast retrieval and classification is possible with the TOMI architecture by using so-called split and acquisition algorithms that run on the TOMI CPU on the same chip as the memory-centric database.
Such a system would probably be built with TOMI DIMMs (Dual Inline Memory Modules) that incorporate multiple TOMI CPU chips. The data path for standard 240-pin DIMMs is 64-bit. Therefore, TOMIPCC in memory-centric database applications drives 64-bit wide databases, as indicated by D0-D63.
It goes without saying that the present invention has been described for illustration purposes only with reference to the accompanying drawings and is not limited to the particular embodiments described herein. As those skilled in the art will appreciate, improvements and modifications can be made to the invention and exemplary embodiments described herein without departing from the scope or spirit of the invention.
39 sheets
Sheet 1 Sheet 2 Sheet 3 Sheet 4 Sheet 5 Sheet 6 Sheet 7 Sheet 8 Sheet 9 Sheet 10 Sheet 11 Sheet 12 Sheet 13 Sheet 14 Sheet 15 Sheet 16 Sheet 17 Sheet 18 Sheet 19 Sheet 20 Sheet 21 Sheet 22 Sheet 23 Sheet 24 Sheet 25 Sheet 26 Sheet 27 Sheet 28 Sheet 29 Sheet 30 Sheet 31 Sheet 32 Sheet 33 Sheet 34 Sheet 35 Sheet 36 Sheet 37 Sheet 38 Sheet 39
Every citation, both ways
| Document | Relation | Office | Cited during |
|---|---|---|---|
| JP2020181407A | Cited by | Japan | Search report |
| JP2000020450A | Cites | Japan | Search report |
| JP2000207248A | Cites | Japan | Examiner |
| JP2002108691A | Cites | Japan | Examiner |
| WO2007092528A2 | Cites | World Intellectual Property Organization (WIPO) | Examiner |
| JP2009525545A | Cites | Japan | Examiner |
| JPH06231092A | Cites | Japan | Search report |
| JPH1049428A | Cites | Japan | Examiner |
| JPH11110215A | Cites | Japan | Examiner |
45 members in 12 offices
Priority claims9
| Document | Office | Kind | Date |
|---|---|---|---|
| 12147332 | United States of America | – | |
| 14733208 | United States of America | A | |
| 14733208 | United States of America | A | |
| 2008068566 | United States of America | W | |
| 2008068566 | United States of America | W | |
| 2008147332 | – | – | – |
| 2008068566 | – | – | – |
| US20080147332 | – | – | – |
| WO2008US68566 | – | – | – |
Members45
| Document | Office | Kind | |
|---|---|---|---|
| AU2007212342A1 | Australia | A1 | |
| CA2642022A1 | Canada | A1 | |
| US2007192568A1 | United States of America | A1 | |
| WO2007092528A2 | World Intellectual Property Organization (WIPO) | A2 | |
| WO2007092528A3 | World Intellectual Property Organization (WIPO) | A3 | |
| EP1979808A2 | European Patent Office (EPO) | A2 | |
| KR20080109743A | Republic of Korea | A | |
| US2008320277A1 | United States of America | A1 | |
| CN101395578A | China | A | |
| EP1979808A4 | European Patent Office (EPO) | A4 | |
| WO2007092528A9 | World Intellectual Property Organization (WIPO) | A9 | |
| JP2009525545A | Japan | A | |
| HK1127414A | Hong Kong, China | A | |
| HK1127414A1 | Hong Kong, China | A1 | |
| CA2684753A1 | Canada | A1 | |
| WO2009157943A1 | World Intellectual Property Organization (WIPO) | A1 | |
| AU2008355072A1 | Australia | A1 | |
| EP2154607A2 | European Patent Office (EPO) | A2 | |
| RU2008135666A | Russian Federation | A | |
| KR20100032359A | Republic of Korea | A | |
| EP2154607A3 | European Patent Office (EPO) | A3 | |
| CN101796484A | China | A | |
| JP2010532905AThis record | Japan | A | |
| EP2288988A1 | European Patent Office (EPO) | A1 | |
| AU2007212342B2 | Australia | B2 | |
| RU2009145519A | Russian Federation | A | |
| RU2427895C2 | Russian Federation | C2 | |
| EP1979808B1 | European Patent Office (EPO) | B1 | |
| AT536585T | Austria | T | |
| ATE536585T1 | Austria | T1 | |
| KR101120398B1 | Republic of Korea | B1 | |
| KR101121606B1 | Republic of Korea | B1 | |
| RU2450339C2 | Russian Federation | C2 | |
| AU2008355072B2 | Australia | B2 | |
| EP2288988A4 | European Patent Office (EPO) | A4 | |
| JP4987882B2 | Japan | B2 | |
| AU2008355072C1 | Australia | C1 | |
| CN101395578B | China | B | |
| BRPI0811497A2 | Brazil | A2 | |
| CN101796484B | China | B | |
| US8977836B2 | United States of America | B2 | |
| US8984256B2 | United States of America | B2 | |
| CN104536723A | China | A | |
| US2015234777A1 | United States of America | A1 | |
| US9934196B2 | United States of America | B2 |
10 legal events, as the office reported them to INPADOC
Over the term
Point at a mark for the eventEvents
| Event | Code | |
|---|---|---|
| Decision of refusalJAPANESE INTERMEDIATE CODE: A02A02 | A02 | |
| Written amendmentJAPANESE INTERMEDIATE CODE: A523A521 | A521 | |
| Written permission of extension of timeJAPANESE INTERMEDIATE CODE: A602A602 | A602 | |
| Written request for extension of timeJAPANESE INTERMEDIATE CODE: A601A601 | A601 | |
| Written permission of extension of timeJAPANESE INTERMEDIATE CODE: A602A602 | A602 | |
| Written request for extension of timeJAPANESE INTERMEDIATE CODE: A601A601 | A601 | |
| Written permission of extension of timeJAPANESE INTERMEDIATE CODE: A602A602 | A602 | |
| Written request for extension of timeJAPANESE INTERMEDIATE CODE: A601A601 | A601 | |
| Notification of reasons for refusalJAPANESE INTERMEDIATE CODE: A131A131 | A131 | |
| Written amendmentJAPANESE INTERMEDIATE CODE: A523A521 | A521 |
Numbers
- Publication
- 2010532905
- Publication, DOCDB
- 2010532905
- Publication, EPODOC
- JP2010532905
- Application
- 2010518258
- Application, DOCDB
- 2010518258
- Application, EPODOC
- JP20100518258
Titles2
- Japanese
- スレッドに最適化されたマルチプロセッサアーキテクチャ
- English
- Thread-optimized multiprocessor architecture
Classification
- CPC, 21
- G06F9/30032
- G06F9/3885
- G06F15/80
- G06F9/30043
- G06F9/30134
- G06F9/30138
- G06F9/30145
- G06F9/30163
- G06F9/30167
- G06F9/322
- G06F9/325
- G06F9/34
- G06F9/38
- G06F9/3851
- G06F12/0848
- G06F12/0875
- G06F15/8007
- G06F15/7846
- G06F13/28
- Y02D10/00
- G06F9/323
- IPC, 4
- G06F9 38
- G06F9 48
- G06F15 167
- G06F9 34
Designated states4
- Regional, 4
- Zimbabwe
- Turkmenistan
- Türkiye
- Togo