Nova Patents
CA3051990C

Accelerated deep learning

Abstract

Techniques in advanced deep learning provide improvements in one or more of accuracy, performance, and energy efficiency, such as accuracy of learning, accuracy of prediction, speed of learning, performance of learning, and energy efficiency of learning. An array of processing elements performs flow-based computations on wavelets of data. Each processing element has a respective compute element and a respective routing element. Each compute element has processing resources and memory resources. Each router enables communication via wavelets with at least nearest neighbors in a 2D mesh. Stochastic gradient descent, mini-batch gradient descent, and continuous propagation gradient descent are techniques usable to train weights of a neural network modeled by the processing elements. Reverse checkpoint is usable to reduce memory usage during the training.

CA3051990C, drawing sheet 1
Sheet 1 of 33

Term

11.4 yearsleft in the term

Expires 23 February 2038.

  1. Priority
  2. Filed
  3. Granted
  4. Today
  5. Expires

47 claims: 3 independent, 44 dependent

  1. 1
    A system comprising:a fabric of processor elements, each processor element comprising a fabric router and a compute engine collectively enabled to perform processing comprising dataflow-based and instruction-based processing;wherein each processor element selectively communicates fabric packets with others of the processor elements;wherein each compute engine selectively performs the processing in accordance with a virtual channel specifier and a task specifier of at least some of the fabric packets the compute engine receives;and wherein the dataflow-based processing is in accordance with the virtual channel specifier identifying in part one or more communication pathways between a plurality of the processor elements, and the instruction-based processing is in accordance with the task specifier identifying in part a starting address for fetching instructions executable by one or more of the computing engines.
  2. 12
    A method comprising:in each of a fabric of processor elements, selectively communicating fabric packets with others of the processor elements, each processor element comprising a fabric router and a compute engine collectively enabled to perform processing comprising dataflow-based and instruction-based processing;in each compute engine, selectively performing the processing in accordance with a virtual channel specifier and a task specifier of at least some of the fabric packets the compute engine receives;and wherein the dataflow-based processing is in accordance with the virtual channel specifier identifying in part one or more communication pathways between a plurality of the processor elements, and the instruction-based processing is in accordance with the task specifier identifying in part a starting address for fetching instructions executable by one or more of the computing engines.
  3. 23
    A system comprising:in each of a fabric of processor elements, means for selectively communicating fabric packets with others of the processor elements, each processor element comprising a fabric router and a compute engine collectively enabled to perform processing comprising dataflow-based and instruction-based processing;-151 Date Reçue/Date Received 2020-07-08 in each compute engine, means for selectively performing the processing in accordance with a virtual channel specifier and a task specifier of at least some of the fabric packets the compute engine receives;and wherein the dataflow-based processing is in accordance with the virtual channel specifier identifying in part one or more communication pathways between a plurality of the processor elements, and the instruction-based processing is in accordance with the task specifier identifying in part a starting address for fetching instructions executable by one or more of the computing engines.