EP3662474A2

A memory-based distributed processor architecture

Abstract

This record has no abstract on file.

Term

11.8 yearsto projected expiry

Projected expiry 30 July 2038, counted from filing; an application has no term until it is granted.

  1. Priority
  2. Filed
  3. Published
  4. Today
  5. Projected expiry

131 claims: 22 independent, 109 dependent

  1. 1
    Claims of equivalent WO 2019025864 A2 WHAT IS CLAIMED IS:1. A distributed processor, comprising: a substrate;a memory array disposed on the substrate, the memory array including a plurality of discrete memory banks;a processing array disposed on the substrate, the processing array including a plurality of processor subunits, each one of the processor subunits being associated with a corresponding, dedicated one of the plurality of discrete memory banks;a first plurality of buses, each connecting one of the plurality of processor subunits to its corresponding, dedicated memory bank;and a second plurality of buses, each connecting one of the plurality of processor subunits to another of the plurality of processor subunits.
  2. 11
    A memory chip, comprising:a substrate;a memory array disposed on the substrate, the memory array including a plurality of discrete memory banks;a processing array disposed on the substrate, the processing array including a plurality of logic portions each including an address generator, wherein each one of the address generators is associated with a corresponding, dedicated one of the plurality of discrete memory banks;and a plurality of buses, each connecting one of the plurality of address generators to its corresponding, dedicated memory bank.
  3. 18
    The memory chip of 17, wherein the memory interface comprises an interface compliant with at least one Joint Electron Device Engineering Council (JEDEC) standard or its variants.
  4. 20
    A distributed processor, comprising:a substrate;a memory array disposed on the substrate, the memory array including a plurality of discrete memory banks, wherein each of the discrete memory banks has a capacity greater than one megabyte;and a processing array disposed on the substrate, the processing array including a plurality of processor subunits, each one of the processor subunits being associated with a corresponding, dedicated one of the plurality of discrete memory banks.
  5. 28
    A distributed processor, comprising:a substrate;a memory array disposed on the substrate, the memory array including a plurality of discrete memory banks;and a processing array disposed on the substrate, the processing array including a plurality of processor subunits, each one of the processor subunits being associated with a corresponding, dedicated one of the plurality of discrete memory banks;and a plurality of buses, each one of the plurality of buses connecting one of the plurality of processor subunits to at least another one of the plurality of processor subunits, wherein the plurality of buses are free of timing hardware logic components such that data transfers between processor subunits and across corresponding ones of the plurality of buses are uncontrolled by timing hardware logic components.
  6. 38
    A distributed processor on a memory chip, comprising:a substrate;a memory array disposed on the substrate, the memory array including a plurality of discrete memory banks;and a processing array disposed on the substrate, the processing array including a plurality of processor subunits, each one of the processor subunits being associated with a corresponding, dedicated one of the plurality of discrete memory banks;and a plurality of buses, each one of the plurality of buses connecting one of the plurality of processor subunits to a corresponding, dedicated one of the plurality of discrete memory banks, wherein the plurality of buses are free of timing hardware logic components such that data transfers between a processor subunit and a corresponding, dedicated one of the plurality of discrete memory banks and across a corresponding one of the plurality of buses are uncontrolled by timing hardware logic components.
  7. 39
    A distributed processor, comprising:a substrate;a memory array disposed on the substrate, the memory array including a plurality of discrete memory banks;and a processing array disposed on the substrate, the processing array including a plurality of processor subunits, each one of the processor subunits being associated with a corresponding, dedicated one of the plurality of discrete memory banks;and a plurality of buses, each one of the plurality of buses connecting one of the plurality of processor subunits to at least another one of the plurality of processor subunits,wherein the plurality of processor subunits are configured to execute software that controls timing of data transfers across the plurality of buses to avoid colliding data transfers on at least one of the plurality of buses.
  8. 40
    A distributed processor on a memory chip, comprising:a substrate;a plurality of processor subunits disposed on the substrate, each processor subunit being configured to execute a series of instructions independent from other processor subunits, each series of instructions defining a series of tasks to be performed by a single processor subunit;a corresponding plurality of memory banks disposed on the substrate, each one of the plurality processor subunits being connected to at least one dedicated memory bank not shared by any others of the plurality of processor subunits;and a plurality of buses, each of the plurality of buses connecting one of the plurality of processor subunits to at least one other of the plurality of processor subunits, wherein data transfers across at least one of the plurality of buses are predefined by the series of instructions included in a processor subunit connected to the at least one of the plurality of buses.
  9. 46
    A distributed processor on a memory chip, comprising:a plurality of processor subunits disposed on the memory chip;a plurality of memory banks disposed on the memory chip, wherein each one of the plurality of memory banks is configured to store data independent from data stored in other ones of the plurality of memory banks, and wherein each one of the plurality of processor subunits is connected to at least one dedicated memory bank from among the plurality of memory banks;and a plurality of buses, wherein each one of the plurality of buses connects one of the plurality of processor subunits to one or more corresponding, dedicated memory banks from among the plurality of memory banks, wherein data transfers across a particular one of the plurality of buses are controlled by a corresponding processor subunit connected to the particular one of the plurality of buses.
  10. 51
    A distributed processor on a memory chip, comprising:a plurality of processor subunits disposed on the memory chip;a plurality of memory banks disposed on the memory chip, wherein each one of the plurality of processor subunits is connected to at least one dedicated memory bank from among the plurality of memory banks, and wherein each memory bank of the plurality of memory banks is configured to store data independent from data stored in other ones of the plurality of memory banks, and wherein at least some of the data stored in one particular memory bank from among the plurality of memory banks comprises a duplicate of data stored in at least another one of the plurality of memory banks;and a plurality of buses, wherein each one of the plurality of buses connects one of the plurality of processor subunits to one or more corresponding, dedicated memory banks from among the plurality of memory banks, wherein data transfers across a particular one of the plurality of buses are controlled by a corresponding processor subunit connected to the particular one of the plurality of buses.
  11. 56
    A non-transitory computer-readable medium storing instructions for compiling a series of instructions for execution on a memory chip comprising a plurality of processor subunits and a plurality of memory banks, wherein each processor subunit from among the plurality of processor subunits is connected to at least one corresponding, dedicated memory bank from among the plurality of memory banks, the instructions causing at least one processor to:divide the series of instructions into a plurality of groups of sub-series instructions, the division comprising: assigning tasks associated with the series of instructions to different ones of the processor subunits, wherein the processor subunits are spatially distributed among the plurality of memory banks disposed on the memory chip;generating tasks to transfer data between pairs of the processor subunits of the memory chip, each pair of processor subunits being connected by a bus, and grouping the assigned and generated tasks into the plurality of groups of sub-series instructions, wherein each of the plurality of groups of sub-series instructions corresponds to a different one of the plurality of processor sub-units;generate machine code corresponding to each of the plurality of groups of subs-series instructions;and assign the generated machine code corresponding to each of the plurality of groups of subs-series instructions to a corresponding one of the plurality of processor subunits in accordance with the division.
  12. 61
    A memory chip, comprising:a plurality of memory banks, each memory bank having a bank row decoder, a bank column decoder, and a plurality of memory sub-banks, each memory sub-bank having a sub-bank row decoder and a sub-bank column decoder for allowing reads and writes to locations on the memory sub-bank, each memory sub-bank comprising: a plurality of memory mats, each memory mat having a plurality of memory cells, wherein the sub-bank row decoders and the sub-bank column decoders are connected to the bank row decoder and the bank column decoder.
  13. 68
    A memory chip, comprising:a plurality of memory banks, each memory bank having a bank controller and a plurality of memory sub-banks, each memory sub-bank having a sub-bank row decoder and a sub-bank column decoder for allowing reads and writes to locations on the memory sub-bank, each memory sub-bank comprising: a plurality of memory mats, each memory mat having a plurality of memory cells, wherein the sub-bank row decoders and the sub-bank column decoders process read and write requests from the bank controller.
  14. 74
    A memory chip, comprising:a plurality of memory banks, each memory bank having a having a bank controller for processing reads and writes to locations on the memory bank, each memory bank comprising: a plurality of memory mats, each memory mat having a plurality of memory cells and having a mat row decoder and a mat column decoder, wherein the mat row decoders and the mat column decoders process read and write requests from the sub-bank controller.
  15. 79
    A memory chip, comprising:a plurality of memory banks, each memory bank having a bank controller, a row decoder, and a column decoder for allowing reads and writes to locations on the memory bank;and a plurality of buses connecting each controller of the plurality of bank controllers to at least one other controller of the plurality of bank controllers.
  16. 86
    A memory device, comprising:a substrate;a plurality of memory banks on the substrate;a plurality of primary logic blocks on the substrate, each of the plurality of primary logic blocks being connected to at least one of the plurality of memory banks;a plurality of redundant blocks on the substrate, each of the plurality of redundant blocks being connected to at least one of the memory banks, each of the plurality of redundant blocks replicating at least one of the plurality of primary logic blocks;and a plurality of configuration switches on the substrate, each one of the plurality of the configuration switches being connected to at least one of the plurality of primary logic blocks or to at least one of the plurality of redundant blocks;wherein upon detection of a fault associated with one of the plurality of primary logic blocks: a first configuration switch of the plurality of configuration switches is configured to disable the one of the plurality of primary logic blocks, and a second configuration switch of the plurality of configuration switches is configured to enable one of the plurality of redundant blocks that replicates the one of the plurality of primary logic blocks.
  17. 105
    A distributed processor on a memory chip, comprising:a substrate;an address manager on the substrate;a plurality of primary logic blocks on the substrate, each of the plurality of primary logic blocks being connected to at least one of the plurality of memory banks;a plurality of redundant blocks on the substrate, each of the plurality of redundant blocks being connected to at least one of the plurality of memory banks, each of the plurality of redundant blocks replicating at least one of the plurality of primary logic blocks;and a bus on the substrate connected to each of the plurality of primary logic blocks, each of the plurality of redundant blocks, and the address manager, wherein the processor is configured to: assign running ID numbers to blocks in the plurality of primary logic blocks that pass a testing protocol;assign illegal ID numbers to blocks in the plurality of primary logic blocks that do not pass the testing protocol;and assign running ID numbers to blocks in the plurality of redundant blocks that pass the testing protocol.
  18. 109
    A method for configuring a distributed processor on a memory chip, comprising:testing each one of a plurality of primary logic blocks on the substrate of the memory chip for at least one circuit functionality;identifying at least one faulty logic block in the plurality of primary logic blocks based on the testing results, the at least one faulty logic block being connected to at least one memory bank disposed on the substrate of the memory chip;testing at least one redundant block on the substrate of the memory chip for the at least one circuit functionality, the at least one redundant block replicating the at least one faulty logic block and being connected to the at least one memory bank;disabling the at least one faulty logic block by applying an external signal to a deactivation switch, the deactivation switch being connected with the at least one faulty logic block and being disposed on the substrate of the memory chip;and enabling the at least one redundant block by applying the external signal to an activation switch, the activation switch being connected with the at least one redundant block and being disposed on the substrate of the memory chip.
  19. 110
    A method for configuring a distributed processor on a memory chip, comprising:enabling a plurality of primary logic blocks and a plurality of redundant blocks on the substrate of the memory;testing each one of the plurality of primary logic blocks on the substrate of the memory chip for at least one circuit functionality;identifying at least one faulty logic block in the plurality of primary logic blocks based on the testing results, the at least one faulty logic block being connected to at least one memory bank disposed on the substrate of the memory chip;testing at least one redundant block on the substrate of the memory chip for the at least one circuit functionality, the at least one redundant block replicating the at least one faulty logic block and being connected to the at least one memory bank;disabling at least one redundant block by applying the external signal to an activation switch, the activation switch being connected with the at least one redundant block and being disposed on the substrate of the memory chip.
  20. 111
    1 11. A processing device, comprising:a substrate;a plurality of memory banks on the substrate;a memory controller on the substrate connected to each one of the plurality of memory banks;and a plurality of processing units on the substrate, each one of the plurality of processing units being connected to the memory controller, the plurality of processing units comprising a configuration manager;wherein the configuration manager is configured to: receive a first indication of a task to be performed, the task requiring at least one computation;signal at least one selected processing unit from the plurality of processing units based upon a capability of the selected processing unit for performing the at least one computation;and transmit a second indication to the at least one selected processing unit, and wherein the memory controller is configured to: route data from at least two memory banks to the at least one selected processing unit using at least one communication line, the at least one communication line being connected to the at least two memory banks and the at least one selected processing unit via the memory controller.
  21. 129
    A method performed for operating a distributed memory device comprising:compiling, by a compiler, a task for the distributed memory device, the task requiring at least one computation, the compiling comprising: determining a number of words that are required simultaneously to perform the task, and providing instructions for writing words that need to be accessed simultaneously in a plurality of memory banks disposed on the substrate when a number a number of words that can be accessed simultaneously from one of the plurality of memory banks is lower than the number of words that are required simultaneously;receiving, by a configuration manager disposed on the substrate, an indication to perform the task;and in response to receiving the indication, configuring a memory controller disposed in the substrate to: within a first line access cycle: access at least one first word from a first memory bank from the plurality of memory banks using a first memory line, send the at least one first word to at least one processing unit, and open a first memory line in the second memory bank to access a second address from the second memory bank from the plurality of memory banks, and within a second line access cycle: access at least one second word from the second memory bank using the first memory line, send the at least one second word to at least one processing unit, and access a third address from the first memory bank using a second memory line in the first bank.
  22. 131
    A non-transitory computer-readable medium that stores instructions that, when executed by at least one processor, cause the at least one processor to:determine a number of words that are required simultaneously to perform a task, the task requiring at least one computation;write words that need to be accessed simultaneously in a plurality of memory banks disposed on the substrate when a number a number of words that can be accessed simultaneously from one of the plurality of memory banks is lower than the number of words that are required simultaneously;transmit an indication to perform the task to a configuration manager disposed on the substrate;and transmit instructions to configure a memory controller disposed on the substrate to, within a first line access cycle: access at least one first word from a first memory bank from the plurality of memory banks using a first memory line, send the at least one first word to at least one processing unit, and open a first memory line in the second memory bank to access a second address from the second memory bank from the plurality of memory banks, and within a second line access cycle: access at least one second word from the second memory bank using the first memory line, send the at least one second word to at least one processing unit, and access a third address from the first memory bank using a second memory line in the first bank.
Independent claims22