Nova Patents
EP0485690A2

Parallel associative processor system.

Abstract

Multiprocessor parallel computing systems and a byte serial SIMD processor parallel architecture is used for parallel array processing with a simplified architecture adaptable to chip implementation in an air cooled environment. The array provided is an N dimensional array of byte wide processing units each coupled with an adequate segment of byte wide memory and control logic. A partitionable section of the array containing several processing units are contained on a silicon chip arranged with "Picket"s, an element of the processing array preferably consisting of combined processing element with a local memory for processing bit parallel bytes of information in a clock cycle. A Picket Processor system (or Subsystem) comprises an array of pickets, a communication network, an I/O system, and a SIMD controller consisting of a microprocessor, a canned routine processor, and a microcontroller that runs the array. The Picket Architecture for SIMD includes set associative processing, parallel numerically intensive processing, with physical array processing similar to image processing. a military picket line analogy fits quite well. Pickets, having a bit parallel processing element, with local memory coupled to the processing element for the parallel processing of information in an associative way where each picket is adapted to perform one element of the associative process. We have provided a way for horizontal association with each picket. The memory of the picket units is arranged in an array. The array of pickets thus arranged comprises a set associative memory. The set associative parallel processing system on a single chip permits a smaller set of `data' out of a larger set to be brought out of memory where an associative operation can be performed on it. This associative operation, typically an exact compare, is performed on the whole set of data in parallel, utilizing the Picket's memory and execution unit.

EP0485690A2, drawing sheet 1
Sheet 1 of 9

Term

Term ended

Projected expiry passed 15 June 2011, 15.3 years ago.

  1. Priority
  2. Filed
  3. Published
  4. Projected expiry
  5. Today

20 claims: 3 independent, 17 dependent

  1. 1
    A parallel processing system comprising a plurality of picket units, each picket unit having a bit parallel processing element combined with a local memory coupled to the processing element for the parallel processing of information in all picket units in an associative way where each picket unit is adapted to perform one element of the association process.
  2. 4
    A parallel processing system according to one of claims 1 to 3 wherein there is provided a picket processor array memory chip with a plurality of picket units and memory and having data flow paths between each picket's processing element and memory and between picket units.
  3. 5
    A parallel processing system according to one of claims 1 to 4 where there is provided a plurality of picket units arranged in an array and paths for data flow between the picket units having one-on-one memory with each processing element of the array is across, left or right to or from an adjacent picket unit neighbor and with slide means providing for a slide communication with picket units farther away.
  4. 6
    A parallel processing system according to one of claims 1 to 5 where each picket unit has a processor which has access to its own coupled local memory, and wherein character wide, or character multiples wide data and instructions flow between picket units in one clock cycle of the system.
  5. 7
    A parallel processing system according to one of claims 1 to 6 wherein each picket chip has its own local memory providing character wide set associative storage in an array of picket units with at least 32 Kbytes storage provided for each local memory and there are sixteen picket units provided as nodes of a linear sub-array.
  6. 8
    A parallel processing system according to one of claims 1 to 7 wherein there is provided a broadcast bus for communication between picket units which is multi-byte wide, so that all pickets can see the same data at the same time, and wherein there is provided means for passing picket control and address propagation transfers on this broadcast bus.
  7. 9
    A parallel processing system according to one of claims 1 to 8 wherein local memory of a picket unit could be DRAM CMOS memory in a memory array and which supports row-column access by deleting the column demultiplexing on the back of the memory array, and which provides a row address that reads out a row of the memory array to cause data flows in parallel.
  8. 10
    A parallel processing system according to one of claims 1 to 9 wherein the memory, in addition to data, contains "tri-bits" or "trit", so that there are three states recognized by the logic, signifying, either logic 1, logic 0, or don't care, and a trit is contained in successive storage locations in the storage array of the set associative memory.
  9. 11
    A parallel processing system according to one of claims 1 to 10 wherein there is provided picket unit control means for providing a control function for individual operation by a picket unit.
  10. 12
    A parallel processing system according to one of claims 1 to 11 wherein the local memory has a multibit binary reference storage address.
  11. 13
    A parallel processing system according to one of claims 1 to 12 wherein is provided an external control store and control means for control functions which within an individual picket unit suspend operations in a picket unit which has a status output which meets a specific condition, which control functions provide a doze function, and an inhibit function, and an enable write to memory based on conditions in the picket unit and control functions which are provided to the picket unit after retrieval from said external control store.
  12. 14
    A parallel processing system according to one of claims 1 to 13 wherein each local memory and processing element of a picket unit are provided with byte transfer means, and means for input of data and a mask for location of information in memory, and wherein means are provided for addressing memory in order to perform a search, and wherein there is means providing for input of data and a mask for a search in order to locate a word among N words in memory, and wherein matching locations raise a match line, and a separate operation reads or selects a first match, and wherein there are broadcast means for a broadcast operation in which a raised select line indicates participation and broadcast data is copied to all selected word locations.
  13. 15
    A parallel processing system according to one of claims 1 to 14 wherein a processing unit comprises an ALU, mask registers, and a latch, as well as status registers (SR) and data flow registers (DF) which are coupled to a one-on-one local memory for each processing unit.
  14. 16
    A parallel processing system according to one of claims 1 to 15 wherein local memory is a multi-bit wide DRAM and logic of each picket processing unit is formed on the same silicon chip substrate as the DRAM local memory, and wherein there is a direct one-on-one coupling between the local memory and its processing element, said local memory having cells which have a a multi-bit address.
  15. 17
    A parallel processing system according to one of claims 1 to 16, wherein each processing unit is provided with mask registers and a latch, said latch being placed to function as a coupling port for each processing unit along a communication line which is common to said plurality of processing units.
  16. 18
    A parallel processing system according to one of claims 1 to 17 having an external control sequencer and local control register means for controlling the status of individual picket processing units of the system.
  17. 19
    In a parallel processing system having an external controller and a plurality of processing units with logical gates and registers and memory contained therein, a process for keeping a description of a digital system comprising the steps of:assigning each gate description of the digital system as a list of signals that the described gate accepts as inputs and naming the signal it generates, requiring that each time a signal changes, its name is broadcast to all processing units and is compared in parallel with the names of an expected input signal, continuing to determine if a match is found, and recording record in the a processing unit a new value of the signal in a dataflow register, and continuing until all signal changes have been recorded, and then causing all processing units to read out in parallel a control word which tells their data flow how to use the current set of inputs to compute the output, causing these computations to be performed in parallel, with the results compared with the old value from the local gate, and recording in a dataflow status register all of those gates of the the processing units whose outputs change, and causing the external external controller to interrogate all the processing units and asking for the next gate that changed and then broadcasting the appropriate signal name and value from the processing unit to all other processing units and repeating the cycle until no more signal changes occur or the process is stopped.
  18. 20
    A process for using a parallel processing system having a plurality of processing units and local memories coupled for the transfer of data and controls there between comprising, storing dictionary names in a processing unit memory such that the first letter of all names can be compared with that of a desired broadcast name broadcast to a plurality of processing units of the parallel processing system, and wherein all processing units without a match are turned off with the control characteristic provided, and then comparing the second letter of the name and the compare and turnoff procedure is repeated for successive letters (characters) until no active processing units remain or the end of the word has been reached, and then providing a query for said processing units and then causing the index of the desired data to be read out of the processing units.
Independent claims18