EP2154607A2

Thread optimized multiprocessor architecture

Abstract

In one aspect, the invention comprises a system comprising: (a) a plurality of parallel processors on a single chip; and (b) computer memory located on the chip and accessible by each of the processors; wherein each of the processors is operable to process a de minimis instruction set, and wherein each of the processors comprises local caches dedicated to each of at least three specific registers in the processor. In another aspect, the invention comprises a system comprising: (a) a plurality of parallel processors on a single chip; and (b) computer memory located on the chip and accessible by each of the processors, wherein each of the processors is operable to process an instruction set optimized for thread-level parallel processing.

EP2154607A2, drawing sheet 1
Sheet 1 of 15

Term

Projected expiry 5 February 2027.

  1. Priority
  2. Filed
  3. Published
  4. Today
  5. Projected expiry

15 claims: 4 independent, 11 dependent

  1. 1
    A system comprising:a plurality of parallel processors on a single chip;and computer memory located on said chip and accessible by each of said processors, wherein each of said processors is operable to process a de minimis instruction set, and wherein each of said processors comprises local caches dedicated to each of at least three specific registers in said processor.
  2. 5
    A system as in claim. 1, wherein three registers auto-increment and three registers auto-decrement.
  3. 9
    A system comprising;a plurality of parallel processors on a single chip;and computer memory located on said chip and accessible by each of said processors, wherein each of said processors is operable to process an instruction set optimized for thread-level parallel processing.
  4. 12
    A method of thread-level parallel processing utilizing a plurality of parallel processors, a master processor, and a computer memory on a single chip, wherein each of said plurality of processors is operable to process a de minimis instruction set and to process a single thread, comprising:(a) allocating local caches to each of three specific registers in each of said plurality of processors;(b) allocating one of the plurality of processors to process a single thread;(c) processing each allocated thread by said processors;(d) processing the results from each thread processed by said processors;and (e) de-allocating one of said plurality of processors after a thread has been processed.