Special purpose neural network training chip
20 claims: 3 independent, 17 dependent
- 1専用ハードウェアチップを使用してニューラルネットワークをトレーニングする方法であって、前記専用ハードウェアチップのベクトルプロセッサが、活性化入力の2次元行列を表すデータを受信することを含み、前記活性化入力の2次元行列は、ニューラルネットワークのそれぞれのネットワーク層に対する入力活性化行列の活性化入力の一部を含み、前記ベクトルプロセッサは、2次元構造に配置された複数のベクトル処理ユニットと、前記活性化入力の2次元行列を格納するよう構成されるベクトルレジスタとを含み、前記方法はさらに、行列乗算ユニットが、前記2次元行列について乗算結果を生成することを含み、前記生成することは、複数のクロックサイクルの各クロックサイクルにおいて、前記ベクトルプロセッサの前記ベクトルレジスタから前記専用ハードウェアチップの前記行列乗算ユニットに、前記活性化入力の2次元行列における活性化入力のそれぞれの行の少なくとも一部をロードすることと、前記複数のクロックサイクルの終了前に、前記行列乗算ユニットに重み値をロードすることとを含み、前記重み値は、前記活性化入力の2次元行列における前記活性化入力に対応する形状に配置され、前記生成することはさらに、前記行列乗算ユニットが、前記活性化入力の2次元行列および重み値を乗算して、前記乗算結果を得ることを含み、前記方法はさらに、前記乗算結果に基づいて、逆伝播を通じて前記ニューラルネットワークの重み値を更新することを含む、方法。
- 2前記行列乗算ユニットに、前記活性化入力の2次元行列における前記活性化入力のそれぞれの行の少なくとも前記一部をロードすることは、さらに、前記活性化入力のそれぞれの行の少なくとも前記一部を、第1の浮動小数点フォーマットから、前記第1の浮動小数点フォーマットよりも小さいビットサイズを有する第2の浮動小数点フォーマットに変換することを含む、請求項1に記載の方法。
- 3前記乗算結果を前記ベクトルレジスタに格納することをさらに含む、請求項1 または2 に記載の方法。
- 4前記乗算結果は2次元行列であり、前記乗算結果を格納することは、別のセットのクロックサイクルのクロックサイクルごとに、前記乗算結果の値のそれぞれの行を、前記行列乗算ユニットにおける先入れ先出しキューに一時的に格納することを含む、請求項1 ~3のいずれか に記載の方法。
- 5前記ベクトルプロセッサが、前記専用ハードウェアチップのベクトルメモリから前記活性化入力の2次元行列を表す前記データを受信することをさらに含み、前記ベクトルメモリは、高速のプライベートメモリを前記ベクトルプロセッサに提供するよう構成される、請求項1 ~4のいずれか に記載の方法。
- 6前記専用ハードウェアチップの転置ユニットが、前記活性化入力の2次元行列の転置演算を実行することをさらに含む、請求項1 ~5のいずれか に記載の方法。
- 7前記専用ハードウェアチップの削減ユニットが、前記乗算結果の数値の削減を実行することと、前記専用ハードウェアチップの置換ユニットが、前記複数のベクトル処理ユニットに格納された数値を置換することとをさらに含む、請求項1 ~6のいずれか に記載の方法。
- 8前記乗算結果を前記専用ハードウェアチップの高帯域幅メモリに格納することをさらに含む、請求項1 ~7のいずれか に記載の方法。
- 9前記専用ハードウェアチップの疎計算コアによって、予め構築されたルックアップテーブルを使用して、疎な高次元データを密な低次元データにマッピングすることをさらに含む、請求項1 ~8のいずれか に記載の方法。
- 10前記専用ハードウェアチップのチップ間相互接続を使用して、前記専用ハードウェアチップのインターフェイスまたはリソースを他の専用ハードウェアチップまたはリソースに接続することをさらに含む、請求項1 ~9のいずれか に記載の方法。
- 11前記チップ間相互接続は、前記インターフェイスおよび前記専用ハードウェアチップの高帯域幅メモリを他の専用ハードウェアチップに接続する、請求項10に記載の方法。
- 12前記インターフェイスは、ホストコンピュータへのホストインターフェイス、またはホストコンピュータのネットワークへの標準ネットワークインターフェイスである、請求項10に記載の方法。
- 13前記ベクトルレジスタは32個のベクトルレジスタを含む、請求項1 ~12のいずれか に記載の方法。
- 14前記複数のベクトル処理ユニットの各ベクトル処理ユニットは、各クロックサイクルにおいて、2つのそれぞれの算術論理ユニット(ALU)命令、それぞれのロード命令、およびそれぞれのストア命令を実行するよう構成される、請求項1 ~13のいずれか に記載の方法。
- 15前記複数のベクトル処理ユニットにおける各ベクトル処理ユニットは、各クロックサイクルにおいてそれぞれのロード命令およびストア命令を実行するためにそれぞれのオフセットメモリアドレスを演算するよう構成される、請求項14に記載の方法。
- 16前記複数のベクトル処理ユニットの前記2次元構造は、複数のレーンと、前記複数のレーンの各々のための複数のサブレーンとを備え、それぞれのベクトル処理ユニットは、前記複数のサブレーンの各々に位置し、同じレーンに位置する前記複数のベクトル処理ユニットのうちのベクトル処理ユニットは、それぞれのロード命令およびストア命令を介して互いと通信するよう構成される、請求項1 ~15のいずれか に記載の方法。
- 171つ以上のコンピュータと1つ以上のストレージデバイスとを備えるシステムであって、前記1つ以上のストレージデバイスには、前記1つ以上のコンピュータによって実行されると前記1つ以上のコンピュータに専用ハードウェアチップを使用してニューラルネットワークをトレーニングするための動作を実行させるよう動作可能である命令が格納され、前記動作は、前記専用ハードウェアチップのベクトルプロセッサが、活性化入力の2次元行列を表すデータを受信することを含み、前記活性化入力の2次元行列は、ニューラルネットワークのそれぞれのネットワーク層に対する入力活性化行列の活性化入力の一部を含み、前記ベクトルプロセッサは、2次元構造に配置された複数のベクトル処理ユニットと、前記活性化入力の2次元行列を格納するよう構成されるベクトルレジスタとを含み、前記動作はさらに、行列乗算ユニットが、前記2次元行列について乗算結果を生成することを含み、前記生成することは、複数のクロックサイクルの各クロックサイクルにおいて、前記ベクトルプロセッサの前記ベクトルレジスタから前記専用ハードウェアチップの前記行列乗算ユニットに、前記活性化入力の2次元行列における活性化入力のそれぞれの行の少なくとも一部をロードすることと、前記複数のクロックサイクルの終了前に、前記行列乗算ユニットに重み値をロードすることとを含み、前記重み値は、前記活性化入力の2次元行列における前記活性化入力に対応する形状に配置され、前記生成することはさらに、前記行列乗算ユニットが、前記活性化入力の2次元行列および重み値を乗算して、前記乗算結果を得ることを含み、前記動作はさらに、前記乗算結果に基づいて、逆伝播を通じて前記ニューラルネットワークの重み値を更新することを含む、システム。
- 18前記行列乗算ユニットに、前記活性化入力の2次元行列における前記活性化入力のそれぞれの行の少なくとも前記一部をロードすることは、さらに、前記活性化入力のそれぞれの行の少なくとも前記一部を、第1の浮動小数点フォーマットから、第1の浮動小数点フォーマットよりも小さいビットサイズを有する第2の浮動小数点フォーマットに変換することを含む、請求項17に記載のシステム。
- 191つ以上のコンピュータによって実行されると、前記1つ以上のコンピュータに、専用ハードウェアチップを使用してニューラルネットワークをトレーニングするための動作を実行させる命令 を含むプログラム であって、前記動作は、前記専用ハードウェアチップのベクトルプロセッサが、活性化入力の2次元行列を表すデータを受信することを含み、前記活性化入力の2次元行列は、ニューラルネットワークのそれぞれのネットワーク層に対する入力活性化行列の活性化入力の一部を含み、前記ベクトルプロセッサは、2次元構造に配置された複数のベクトル処理ユニットと、前記活性化入力の2次元行列を格納するよう構成されるベクトルレジスタとを含み、前記動作はさらに、行列乗算ユニットが、前記2次元行列について乗算結果を生成することを含み、前記生成することは、複数のクロックサイクルの各クロックサイクルにおいて、前記ベクトルプロセッサの前記ベクトルレジスタから前記専用ハードウェアチップの前記行列乗算ユニットに、前記活性化入力の2次元行列における活性化入力のそれぞれの行の少なくとも一部をロードすることと、前記複数のクロックサイクルの終了前に、前記行列乗算ユニットに重み値をロードすることとを含み、前記重み値は、前記活性化入力の2次元行列における前記活性化入力に対応する形状に配置され、前記生成することはさらに、前記行列乗算ユニットが、前記活性化入力の2次元行列および重み値を乗算して、前記乗算結果を得ることを含み、前記動作はさらに、前記乗算結果に基づいて、逆伝播を通じて前記ニューラルネットワークの重み値を更新することを含む、 プログラム 。
- 20前記行列乗算ユニットに、前記活性化入力の2次元行列における前記活性化入力のそれぞれの行の少なくとも前記一部をロードすることは、さらに、前記活性化入力のそれぞれの行の少なくとも前記一部を、第1の浮動小数点フォーマットから、前記第1の浮動小数点フォーマットよりも小さいビットサイズを有する第2の浮動小数点フォーマットに変換することを含む、請求項19に記載の プログラム 。
Independent claims20
74 paragraphs, as filed
FIELD OF THE DISCLOSURE This specification relates to performing neural network computations in hardware. Neural networks are machine learning models, each of which uses one or more layers of the model to generate an output, such as a classification, for a received input. Some neural networks include one or more hidden layers in addition to an output layer. The output of each hidden layer is used as the input to the next layer in the network, i.e., the next hidden or output layer of the network. Each layer of the network generates an output from the received input according to the current values of a respective set of parameters.
Overview This document describes technology for a dedicated hardware chip that is a programmable linear algebra accelerator optimized for machine learning workloads, specifically the training phase.
In general, one innovative aspect of the subject matter described herein can be embodied in a specialized hardware chip.
Other embodiments of this aspect include corresponding computer systems, devices, and computer programs recorded on one or more computer storage devices, each configured to perform the actions of the method. When one or more computer systems are configured to perform a particular operation or action, it means that software, firmware, hardware, or a combination thereof is installed on the system that, during operation, causes the system to perform such operation or action. When one or more computer programs are configured to perform a particular operation or action, it means that one or more programs contain instructions that, when executed by a data processing device, cause the data processing device to perform such operation or action.
Each of the above and other embodiments can optionally include one or more of the following features, either alone or in combination. In particular, one embodiment includes all of the following features in combination:
A dedicated hardware chip for training a neural network may comprise a scalar processor configured to control computational operations of the dedicated hardware chip and a vector processor configured to have a two-dimensional array of vector processing units, all executing the same instructions in a single-instruction, multiple-data manner and communicating with each other through load and store instructions of the vector processor, and the dedicated hardware chip may further comprise a matrix multiplication unit coupled to the vector processor and configured to multiply at least one two-dimensional matrix with a second one-dimensional vector or two-dimensional matrix to obtain a multiplication result.
The vector memory may be configured to provide a high speed private memory to the vector processor. The scalar memory may be configured to provide a high speed private memory to the scalar processor. The transpose unit may be configured to perform a matrix transpose operation. The reduction and permutation unit may be configured to perform a reduction on numbers and permutation of numbers between different lanes of the vector array. The high bandwidth memory may be configured to store data for a dedicated hardware chip. The dedicated hardware chip may comprise a sparse computation core.
The special purpose hardware chips may include interfaces and inter-chip interconnects that connect the interfaces or resources on the special purpose hardware chips to other special purpose hardware chips or resources.
The special-purpose hardware chip may include a high-bandwidth memory. The inter-chip interconnect may connect the interface and the high-bandwidth memory to other special-purpose hardware chips. The interface may be a host interface to a host computer. The interface may be a standard network interface to a network of the host computer.
The subject matter described in this specification can be implemented in certain embodiments to achieve one or more of the following advantages: A dedicated hardware chip includes a processor that is optimized for 32-bit or less precision computations for machine learning, yet natively supports higher dimensional tensors (i.e., 2 or more dimensions) in addition to traditional 0- and 1-dimensional tensor computations.
The details of one or more embodiments of the subject matter of this specification are set forth in the accompanying drawings and the detailed description below. Other features, aspects, and advantages of the subject matter will become apparent from the detailed description, the drawings, and the claims.
<figref num="1">1 illustrates an example topology of high speed connections connecting an example collection of dedicated hardware chips connected in a circular topology on a board.</figref><figref num="2">1 shows a high-level diagram of an exemplary dedicated hardware chip for training neural networks.</figref><figref num="3">1 shows a high level example of a compute core.</figref><figref num="4">1 shows a more detailed diagram of a chip that performs training for a neural network.</figref>
Like reference numbers and designations in the various drawings indicate like elements.
DETAILED DESCRIPTION A neural network having multiple layers can be trained and used to compute inference. Typically, some or all of the layers of a neural network have parameters that are adjusted during training of the neural network. For example, some or all of the layers can multiply a matrix of parameters, also referred to as weights, for that layer by the input to that layer as part of generating the layer output. The values of the parameters in the matrix are adjusted during training of the neural network.
In particular, during training, the training system executes a training procedure for the neural network to adjust the values of the parameters of the neural network, e.g., to determine the trained values of the parameters from their initial values. The training system uses backpropagation of errors, known as backpropagation, in combination with optimization methods to calculate the gradient of the objective function with respect to each parameter of the neural network and use the gradients to adjust the values of the parameters.
A trained neural network can compute inference using forward propagation, i.e., processing an input through the layers of the neural network to generate a neural network output for that input.
For example, given an input, a neural network can compute an inference for that input. The neural network computes this inference by processing the input through each layer of the neural network. In some implementations, the layers of the neural network are arranged in a sequence.
Thus, to compute an inference from a received input, a neural network takes the input and processes it through each neural network layer in sequence to generate an inference, with the output from one neural network layer being given as the input to the next neural network layer. The data input to a neural network layer, e.g., the input to the neural network, or the output to a neural network layer of a layer below it in the sequence, can be referred to as the activation input to that layer.
In some implementations, the layers of a neural network are arranged in a directed graph, meaning that any particular layer can receive multiple inputs, multiple outputs, or both. The layers of a neural network can also be configured such that the output of one layer can be sent back as an input to a previous layer.
One exemplary system is a high-performance multi-chip tensor computing system optimized for matrix multiplication and other multi-dimensional array computations. These operations are important for training neural networks and, optionally, for computing inferences using neural networks.
In one exemplary system, multiple dedicated chips are arranged to distribute operations so that the system can efficiently perform training and inference calculations. In one implementation, there are four chips on a board, and in larger systems many boards may be next to each other in a rack or otherwise in data communication with each other.
FIG. 1 illustrates an exemplary topology of high-speed connections connecting an exemplary collection of dedicated hardware chips 101a-101d connected in a circular topology on a board. Each chip includes two processors (102a-102h). The topology is a one-dimensional (1D) torus, where each chip is directly connected to two adjacent chips. As shown, in some implementations, the chips include a microprocessor core that is programmed with software or firmware instructions to perform operations. In FIG. 1, all chips are on a single module 100. The lines between the processors shown in the figure represent high-speed data communication links. The processors are advantageously fabricated on one integrated circuit substrate, but can also be fabricated on multiple substrates. Beyond chip boundaries, the links are inter-chip network links, and processors on the same chip communicate via intra-chip interface links. The links may be half-duplex links, where only one processor can send data at a time, or full-duplex links, where data can be sent in both directions simultaneously. Parallel processing using this exemplary topology and more is described in detail in U.S. patent application Ser. No. 62/461,758, entitled "PARALLEL PROCESSING OF REDUCTION AND BROADCAST OPERATIONS ON LARGE DATASETS OF NON-SCALAR DATA," filed Feb. 21, 2017, and incorporated herein by reference.
2 shows a high-level view of an exemplary dedicated hardware chip for training neural networks. As shown, the single dedicated hardware chip includes two independent processors (202a, 202b). Each processor (202a, 202b) includes two different cores: (1) a compute core, e.g., a very long instruction word (VLIW) machine (203a, 203b), and (2) a sparse computation core, i.e., an embedded layer accelerator (205a, 205b).
Each core (203a, 203b) is optimized for dense linear algebra problems. A single very long instruction word controls several compute cores in parallel. The compute cores are described in more detail with reference to Figures 3 and 4.
Exemplary sparse computation cores (205a, 205b) map very sparse high-dimensional data into dense low-dimensional data, allowing the remaining layers to process densely packed input data. For example, the sparse computation cores can perform computations for the embedding layer of a neural network during training.
To perform this sparse-to-dense mapping, the sparse computing core uses a pre-built lookup table, which is an embedding table. For example, given a series of query words as user input, each query word is converted into a hash identifier or one-hot encoded vector. Using the identifier as a table index, the embedding table returns a corresponding dense vector, which can become the input activation vector for the next layer. The sparse computing core can also perform a reduce operation across the search query words to create one dense activation vector. The sparse computing core performs an efficient sparse, distributed lookup because the embedding table can be huge and does not fit into the limited capacity high bandwidth memory of one of the dedicated hardware chips. Details regarding the sparse computing core functionality are described in U.S. Patent Application No. 15/016,486, entitled "MATRIX PROCESSING APPARATUS," filed February 5, 2016, which is incorporated herein by reference.
3 shows a high-level example of a compute core (300). A compute core can be a machine that controls multiple compute units in parallel, i.e., a VLIW machine. Each compute core (300) includes a scalar memory (304), a vector memory (308), a scalar processor (303), a vector processor (306), and an extended vector unit (i.e., a matrix multiplication unit (MXU) (313), a transpose unit (XU) (314), and a reduction and permute unit (RPU) (316)).
The exemplary scalar processor executes a fetch/execute loop of VLIW instructions and controls the compute cores. After fetching and decoding an instruction bundle, the scalar processor itself only executes the instructions found in the scalar slots of the instruction bundle using a number of multi-bit registers, i.e., 32 32-bit registers, of the scalar processor (303) and scalar memory (304). The scalar instruction set includes regular arithmetic operations used in address calculations, load/store instructions, branch instructions, etc. The remaining instruction slots encode instructions for the vector processor (306) or other extended vector units (313, 314, 316). The decoded vector instructions are forwarded to the vector processor (306).
Along with vector instructions, the scalar processor (303) can transfer the values of up to three scalar registers to other processors and units to perform operations. The scalar processor can also get the computation results directly from the vector processor. However, in some implementations, the exemplary chip has a low bandwidth communication path from the vector processor to the scalar processor.
The vector instruction dispatcher sits between the scalar and vector processors. It receives decoded instructions from non-scalar VLIW slots and broadcasts them to the vector processors (306). The vector processors (306) consist of a two-dimensional array, i.e., a 128x8 array, of vector processing units that execute the same instructions in a single instruction, multiple data (SIMD) fashion. The vector processing units are described in more detail with reference to FIG. 4.
The exemplary scalar processor (303) accesses a small, fast, private scalar memory (304), which is backed by a much larger, slower high bandwidth memory (HBM) (310). Similarly, the exemplary vector processor (306) accesses a small, fast, private vector memory (306), which is also backed by the HBM (310). Word granularity accesses occur between the scalar processor (303) and the scalar memory (304) or between the vector processor (306) and the vector memory (308). The granularity of loads and stores between the vector processor and the vector memory is a vector of 128 32-bit words. Direct memory accesses occur between the scalar memory (304) and the HBM (310), and between the vector memory (306) and the HBM (310). In some implementations, memory transfers from the HBM (310) to the processors (303, 306) can only be performed via scalar or vector memory. Furthermore, direct memory transfers between scalar and vector memories may not occur.
An instruction may specify an extended vector unit operation. In addition to each vector unit instruction executed, there are two dimensional, or 128x8, vector units, each of which can send one register value as an input operand to the extended vector unit. Each extended vector unit receives an input operand, performs a corresponding operation, and returns the result to the vector processor (306). The extended vector unit is described below with reference to FIG. 4.
4 shows a more detailed view of a chip that performs training for a neural network. As shown and described above, the chip includes two compute cores (480a, 480b) and two sparse computation cores (452a, 452b).
The chip has a shared area that includes an interface to a host computer (450) or multiple host computers. This interface can be a host interface to the host computer or a standard network interface to a network of host computers. The shared area may also have a stack of high bandwidth memory (456a-456d) along the bottom, and an inter-chip interconnect (448) that connects the interface to the memory, as well as data from other chips. The interconnect may also connect the interface to computational resources on the hardware chip. Multiple stacks of high bandwidth memory, namely two stacks (456a-456b, 456c-456d), are associated with each compute core (480a, 480b).
The chip stores data in a high bandwidth memory (456c-456d) and processes the data by loading and unloading the data in a vector memory (446). The compute core (480b) itself includes a vector memory (446), which is an on-chip S-RAM partitioned in two dimensions. The vector memory has an address space that holds 128 numbers whose addresses are floating point numbers, i.e. 32 bits each. The compute core (480b) also includes a computation unit that computes the values, and a scalar unit that controls the computation unit. The computation unit may include a vector processor, and the scalar unit may include a scalar processor. The compute core, which may form part of a dedicated chip, may further include another extension computation unit, such as a matrix multiplication unit, or a transpose unit (422) that performs transpose operations on matrices, i.e. 128x128 matrices, as well as a reduction and substitution unit.
The Vector Processor (306) consists of a two-dimensional array of vector processing units, i.e. 128x8, all of which execute the same instructions in a Single Instruction Multiple Data (SIMD) fashion. The Vector Processor has lanes and sublanes, i.e. 128 lanes and 8 sublanes. Within a lane, the vector units communicate with each other via load and store instructions. Each vector unit can access one 4-byte value at a time. Vector units that do not belong to the same lane cannot communicate directly. These vector units must use a reduce/replace unit, described below.
The computation unit includes vector registers (440), i.e., 32 registers, that can be used for both floating-point and integer operations in the vector processing unit. The computation unit includes two arithmetic logic units (ALUs) (406c-406d) for performing computations. One ALU (406c) performs floating-point addition, and the other ALU (406d) performs floating-point multiplication. Both ALUs (406c-406d) can perform various other operations such as shifts, masks, and comparisons. For example, a compute core (480b) may want to add a vector register V1 to a second vector register V2 and place the result in a third vector register V3. To compute this addition, the compute core (480b) performs multiple operations in one clock cycle. Using these registers as operands, each vector unit can simultaneously execute two ALU instructions and one load and one store instruction per clock cycle. The base address of a load or store instruction can be calculated by the scalar processor and forwarded to the vector processor. Each of the vector units in each sublane can calculate its own offset address using various methods such as strides or special indexed address registers.
The compute unit also includes an Extended Unary Pipeline (EUP) (416) that performs operations such as square root and reciprocal. The compute core (480b) takes three clock cycles to perform these operations because they are computationally complex. Because EUP operations take more than one clock cycle, there is first-in, first-out data storage to store the results. When the operation is finished, the results are stored in a FIFO. The compute core can later pull the data from the FIFO and store it in a vector register with another instruction. The random number generator (420) allows the compute core (480b) to generate multiple random numbers per cycle, i.e. 128 random numbers per cycle.
As noted above, each processor, which may be implemented as part of a dedicated hardware chip, has three extended arithmetic units: a matrix multiplication unit (448) that performs matrix multiplication operations; a transposition unit (422) that performs transposition operations on matrices, i.e., 128×128 matrices; and a reduction and permutation unit (shown as separate units 424, 426 in FIG. 4 ).
The matrix multiplication unit performs matrix multiplication between two matrices. The matrix multiplication unit (438) takes in data because the compute core needs to read a set of numbers that are the matrices to be multiplied. As shown, the data comes from vector registers (440). Each vector register contains 128×8 numbers, i.e., 32-bit numbers. However, sending the data to the matrix multiplication unit (448) to change the numbers to a smaller bit size, i.e., from 32-bit to 16-bit, may result in a floating-point conversion. The parallel-to-serializer (440) ensures that when the numbers are read from the vector registers, the two-dimensional array, i.e., the 128×8 matrix, is read as a set of 128 numbers and sent to the matrix multiplication unit (448) for each of the next eight clock cycles. After the matrix multiplication has completed its calculations, the result is deserialized (442a, 442b), which means that the result matrix is held for a certain number of clock cycles. For example, for a 128x8 array, 128 numbers are held for each of 8 clock cycles and then pushed into the FIFO, so that a 2-dimensional array of 128x8 numbers can be obtained in one clock cycle and stored in the vector register (440).
Over a period of multiple cycles, i.e., 128 cycles, the weights are shifted into the matrix multiplication unit (448) as numbers to multiply the matrix. Once the matrices and weights are loaded, the compute core (480) can send a set of numbers, i.e., 128×8, to the matrix multiplication unit (448). Each line of the set can be multiplied by the matrix to produce a number of results, i.e., 128, every clock cycle. While the compute core is performing the matrix multiplication, it also shifts in a new set of numbers that will become the next matrix in the background so that the next matrix is available for the compute core to multiply when the calculation process of the previous matrix is completed. The matrix multiplication unit (448) is described in 16113-8251001, entitled LOW MATRIX MULTIPLY UNIT COMPOSED OF MULTI-BIT CELLS and in MATRIX MULTIPLY UNIT WITH NUMERICS OPTIMIZED FOR NEURAL NETWORK This is described in more detail in US Pat. No. 6,16113-8252001 entitled "NUMERICALLY OPTIMIZED MATRIX MULTIPLICATION UNIT FOR NEURAL NETWORK APPLICATIONS," both of which are incorporated herein by reference.
The transpose unit transposes matrices. The transpose unit (422) takes numbers and transposes them so that the numbers across a lane are transposed with the numbers in the other dimension. In some implementations, the vector processor includes a 128x8 vector unit. Thus, to transpose a 128x128 matrix, 16 separate transpose instructions are required for the complete matrix transposition. Once the transposition is finished, the transposed matrix is available. However, an explicit instruction is required to move the transposed matrix to the vector register file.
The reduction/permutation units (or units 424, 426) address the cross-lane communication problem by supporting various operations such as permutation, lane rotation, rotate permutation, lane reduction, permuted lane reduction, and segmented permuted lane reduction. As shown, these calculations are separate, but the compute core can use one or the other, or the other chained to one. The reduction unit (424) adds all the numbers in each line of numbers and provides them to the permutation unit (426). The permutation unit moves data between different lanes. The transposition unit, reduction unit, permutation unit, and matrix multiplication unit each take one or more clock cycles to complete. Thus, each unit has a FIFO associated with it, and can push the results of a calculation into the FIFO and later execute another instruction to pull the data from the FIFO into a vector register. By using the FIFO, the compute core does not need to reserve multiple vector registers for lengthy operations. As shown, each unit gets data from a vector register (440).
The compute cores control the computational units using a scalar unit. The scalar unit has two main functions: (1) performing loop counting and addressing, and (2) generating direct memory address (DMA) requests for the DMA controller to move data between the high bandwidth memory (456c-456d) and the vector memory (446) in the background, and then to the inter-chip connections (448) to other chips in the exemplary system. The scalar unit includes an instruction memory (404), an instruction decode and issue (402), a scalar processing unit (408) that includes a scalar register, i.e., 32 bits, a scalar memory (410), and two ALUs (406a, 406b) that perform two operations per clock cycle. The scalar unit can pass operands and immediate values to the vector operations. Each instruction can be sent from the instruction decode and issue (402) as an instruction bundle that includes the instruction to be executed in the vector register (440). Each instruction bundle is a very long instruction word (VLIW), where each instruction is a certain number of bits wide and is divided into a certain number of instruction fields.
The chip 400 can be used to perform at least a portion of the training of a neural network. In particular, when training a neural network, the system receives labeled training data from a host computer using a host interface (450). The host interface can also receive instructions including parameters for the neural network computation. The parameters can include at least one or more of the number of layers to be processed, a corresponding set of weight inputs for each layer, an initial set of activation inputs, i.e., training data that is input to the neural network on which inference is computed or trained, the corresponding input and output sizes of each layer, a stride value for the neural network computation, and the type of layer to be processed, e.g., a convolutional layer or a fully connected layer.
The set of weight inputs and the set of activation inputs can be sent to a matrix multiplication unit of the compute core. Before sending the weight inputs and activation inputs to the matrix multiplication unit, other components in the system may perform other calculations on the inputs. In some implementations, there are two ways to send activations from the sparse compute core to the compute core. First, the sparse compute core can send communication via a high-bandwidth memory. For large amounts of data, the sparse compute core can store the activations in the high-bandwidth memory using a direct memory address (DMA) instruction, which updates a target synchronization flag in the compute core. The compute core can wait for this synchronization flag using a synchronization instruction. When the synchronization flag is set, the compute core uses a DMA instruction to copy the activations from the high-bandwidth memory to the corresponding vector memory.
Then, the sparse compute core can send the communication directly to the compute core vector memory. If the amount of data is not large (i.e., it fits into the compute core vector memory), the sparse compute core can directly store the activation into the compute core's vector memory using a DMA instruction while notifying the compute core with a synchronization flag. The compute core can wait for this synchronization flag and then perform the calculation that depends on the activation.
The matrix multiplication unit can process the weight inputs and activation inputs and provide a vector or matrix of outputs to the vector processing unit. The vector processing unit can store the processed vector or matrix of outputs. For example, the vector processing unit can apply a non-linear function to the output of the matrix multiplication unit to generate the activation values. In some implementations, the vector processing unit generates normalized values, pooled values, or both. The vector of processed outputs can be used as activation inputs to the matrix multiplication unit, for example, for use in a subsequent layer in the neural network.
Once the vector of processed outputs for a batch of training data has been computed, the outputs can be compared to the expected outputs of the labeled training data to determine the error. The system can then perform backpropagation to propagate the error through the neural network to train the network. The gradient of the loss function is computed on-chip using the arithmetic logic unit of the vector processing unit.
In one exemplary system, activation gradients are required to perform backpropagation through a neural network. To send activation gradients from a compute core to a sparse computation core, an exemplary system can use compute core DMA instructions to store the activation gradients in a high-bandwidth memory while notifying the target sparse computation core with a synchronization flag. The sparse computation core can wait for the synchronization flag and then perform computations that depend on the activation gradients.
The matrix multiplication unit performs two matrix multiplication operations for backpropagation. One matrix multiplication applies the backpropagated error to weights from a previous layer in the network along a backward pass through the network to adjust the weights to determine new weights for the neural network. The second matrix multiplication applies the error to the original activations as feedback to a previous layer in the neural network. The original activations may be generated during the forward pass and saved for use during the backward pass. The computations may use general purpose instructions in the vector processing unit, including floating point addition, subtraction, and multiplication. General purpose instructions may also include comparisons, shifts, masks, and logical operations. While matrix multiplications may be highly accelerated, the arithmetic logic unit of the vector processing unit performs common computations at a rate of 128x8x2 operations per core per cycle.
The subject matter and functional operations described herein may be implemented in digital electronic circuitry, tangibly embodied computer software or firmware, computer hardware, or one or more combinations thereof, including the structures disclosed herein and their structural equivalents. The subject matter described herein may be implemented as one or more computer programs, i.e., one or more modules of computer program instructions encoded on a tangible, non-transitory storage medium for execution by or to control the operation of a data processing device. The computer storage medium may be a machine-readable storage device, a machine-readable storage substrate, a random or serial access memory device, or one or more combinations thereof. Alternatively, or in addition, the program instructions may be encoded on an artificially generated propagated signal, such as, for example, a machine-generated electrical, optical, or electromagnetic signal, that is generated to encode information for transmission to a receiving device suitable for execution by a data processing device.
The term "data processing apparatus" refers to data processing hardware and encompasses all kinds of apparatus, devices and machines for processing data, including, by way of example, a programmable processor, a computer, or multiple processors or computers. The apparatus may also be or further include special purpose logic circuitry, such as, for example, an FPGA (field programmable gate array) or an ASIC (application specific integrated circuit). In addition to hardware, the apparatus may optionally include code that creates an execution environment for a computer program, such as, for example, code that constitutes a processor firmware, a protocol stack, a database management system, an operating system, or one or more combinations thereof.
A computer program, which may also be referred to or described as a program, software, software application, application, module, software module, script or code, may be written in any form of programming language, including compiled or interpreted languages, or declarative or procedural languages, and may be deployed in any form as a stand-alone program or as a module, component, subroutine, or other unit suitable for use in a computing environment. A program may, but need not, correspond to a file in a file system. A program may be stored in a single file dedicated to that program, or in multiple coordinated files. The program may be stored in a file (e.g., a file that stores one or more modules, subprograms, or portions of code) that holds other programs or data (e.g., one or more scripts stored in a markup language document). A computer program may be deployed to be executed on one computer or on multiple computers that are located at one site or distributed across multiple sites and interconnected by a data communications network.
The processes and logic flows described herein may be performed by one or more programmable computers executing one or more computer programs to perform functions by operating on input data and generating output. The processes and logic flows may also be performed by special purpose logic circuitry, such as, for example, an FPGA or an ASIC, or a combination of special purpose logic circuitry and one or more programmed computers.
A computer suitable for executing a computer program may be based on a general-purpose or special-purpose microprocessor or both, or on any kind of central processing unit. In general, the central processing unit receives instructions and data from a read-only memory or a random access memory or both. The essential elements of a computer are a central processing unit for executing instructions and one or more memory devices for storing instructions and data. The central processing unit and the memory may be supplemented by or incorporated in special-purpose logic circuitry. In general, the computer further includes one or more mass storage devices for storing data, for example magnetic disks, magneto-optical disks, or optical disks, and is operatively coupled to receive data from or transfer data to the one or more mass storage devices, or both. However, a computer need not have such devices. Furthermore, the computer may be embedded in another device, for example, a mobile phone, a personal digital assistant (PDA), a mobile audio or video player, a game console, a global positioning system (GPS) receiver, or a portable storage device (for example, a universal serial bus (USB) flash drive).
Computer-readable media suitable for storing computer program instructions and data include, by way of example, all forms of non-volatile memory, media and memory devices including semiconductor memory devices such as EPROM, EEPROM and flash memory devices; magnetic disks such as internal hard disks or removable disks; magneto-optical disks; and CD-ROM and DVD-ROM disks.
To provide for user interaction, embodiments of the subject matter described herein may be implemented on a computer having a display device, such as a CRT (cathode ray tube) or LCD (liquid crystal display) monitor, for displaying information to the user, and a keyboard and pointing device, such as a mouse or trackball, for allowing the user to provide input to the computer. Other types of devices may be used to provide for user interaction as well; for example, feedback provided to the user may be any form of sensory feedback, such as visual feedback, auditory feedback, or tactile feedback; input from the user may be received in any form, including acoustic input, speech input, or tactile input. In addition, the computer may interact with the user by sending documents to and receiving documents from a device used by the user, for example, by sending a web page to a web browser on the user's device in response to a request received from the web browser. The computer may also interact with the user by sending text messages or other forms of messages to a personal device, such as a smartphone, running a messaging application, and receiving a response message from the user.
An embodiment of the subject matter described herein may be implemented in a computing system that includes a back-end component, e.g., as a data server, a middleware component, e.g., an application server, a front-end component, e.g., a client computer having a graphical user interface, a web browser, or an app by which a user can interact with an implementation of the subject matter described herein, or any combination of one or more such back-end, middleware, or front-end components. The components of the system may be interconnected by any form or medium of digital data communication, e.g., a communications network. Examples of communications networks include local area networks (LANs) and wide area networks (WANs), e.g., the Internet.
A computing system may include clients and servers. Clients and servers are generally remote from each other and typically interact through a communications network. The relationship of clients and servers arises by virtue of computer programs running on the respective computers and having a client-server relationship to each other. In some embodiments, the server sends data, e.g., HTML pages, to a user device for the purpose of displaying the data to and receiving user input from a user interacting with the user device acting as a client. Data generated at the user device, e.g., results of user interaction, can be received at the server from the user device.
Embodiment 1 is a dedicated hardware chip for training a neural network, comprising: a scalar processor configured to control computational operations of the dedicated hardware chip; and a vector processor configured to have a two-dimensional array of vector processing units, all of which execute the same instructions in a single-instruction, multiple-data manner and communicate with each other through load and store instructions of the vector processor; the dedicated hardware chip further comprises a matrix multiplication unit coupled to the vector processor and configured to multiply at least one two-dimensional matrix with a second one-dimensional vector or two-dimensional matrix to obtain a multiplication result.
A second embodiment is the special-purpose hardware chip of the first embodiment, further comprising a vector memory configured to provide a high-speed private memory to the vector processor.
[0023] Example 3 is the special-purpose hardware chip of Example 1 or 2, further comprising a scalar memory configured to provide a high-speed private memory to the scalar processor.
A fourth embodiment is the special-purpose hardware chip of any one of the first to third embodiments, further comprising a transposition unit configured to perform a matrix transposition operation.
Embodiment 5 is the dedicated hardware chip of any one of embodiments 1 to 4, further comprising a reduction and replacement unit configured to perform reduction on numerical values and replace numerical values between different lanes of the vector array.
A sixth embodiment is the dedicated hardware chip of any one of the first to fifth embodiments, further comprising a high-bandwidth memory configured to store data of the dedicated hardware chip.
A seventh embodiment is the dedicated hardware chip of any one of the first to sixth embodiments, further including a sparse computation core.
An eighth embodiment is the special purpose hardware chip of any one of the first to seventh embodiments, further comprising an interface and an inter-chip interconnect that connects an interface or resource on the special purpose hardware chip to another special purpose hardware chip or resource.
A ninth embodiment is the special-purpose hardware chip of any one of the first to eighth embodiments, further comprising a plurality of high-bandwidth memories, and the inter-chip interconnect connects the interface and the high-bandwidth memories to other special-purpose hardware chips.
A tenth embodiment is the dedicated hardware chip of any one of the first to ninth embodiments, wherein the interface is a host interface to a host computer.
An eleventh embodiment is the dedicated hardware chip of any one of the first to tenth embodiments, wherein the interface is a standard network interface to the network of the host computer.
Although the specification contains many specific implementation details, these should not be construed as limitations on the scope of any invention or what may be claimed, but rather as descriptions of features that may be specific to a particular embodiment of a particular invention. Certain features described herein in the context of separate embodiments may also be implemented in combination in a single embodiment. Conversely, various features described in the context of a single embodiment may also be implemented in multiple embodiments separately or in any suitable subcombination. Furthermore, although features may be described above as acting in a combination, and may even be initially claimed as such, one or more features from a claimed combination may in some cases be deleted from the combination, and the claimed combination may be directed to a subcombination or a variation of the subcombination.
Similarly, although operations are shown in a particular order in the figures, it should not be understood that such operations need to be performed in the particular order shown, or sequential order, to achieve desirable results, or that all of the shown operations need to be performed. In some situations, multitasking and parallel processing may be advantageous. Furthermore, it should be understood that the separation of various system modules and components in the above-described embodiments does not require such separation in all embodiments, and that the program components and systems described may generally be integrated into a single software product or packaged into multiple software products.
Specific embodiments of the subject matter have been described. Other embodiments are within the scope of the following claims. For example, the actions recited in the claims can be performed in a different order and still achieve desirable results. By way of example, the processes depicted in the accompanying figures do not necessarily require the particular order shown, or sequential order, to achieve desirable results. In some cases, multitasking and parallel processing may be advantageous.
4 sheets
Sheet 1 Sheet 2 Sheet 3 Sheet 4
Every citation, both ways
| Document | Relation | Office |
|---|---|---|
| US20160342891A1 | Cites | United States of America |
| US20190354862A1 | Cites | United States of America |
| JP04290155A | Cites | Japan |
| JP2017138966A | Cites | Japan |
| 下川 勝千,ディジタルニューロコンピュータMULTINEURO,東芝レビュー 第46巻 第12号,第46巻 第12号,日本,株式会社東芝,1991年12月,pp.931~934,【ISSN】0372-0462 | Non-patent | – |
| 斎藤 康毅,ゼロから作るDeep Learning 初版,株式会社オライリー・ジャパン,第1版,日本,株式会社オライリー・ジャパン ,2016年09月,pp.123-165,第1版 ISBN: 978-4-87311-758-4 | Non-patent | – |
38 members in 9 offices
Priority claims4
| Document | Office | Kind | Date |
|---|---|---|---|
| 62507771 | United States of America | – | |
| 201762507771 | United States of America | P | |
| 2019549507 | Japan | A | |
| 2021142529 | Japan | A |
Members38
| Document | Office | Kind | |
|---|---|---|---|
| US2018336456A1 | United States of America | A1 | |
| WO2018213598A1 | World Intellectual Property Organization (WIPO) | A1 | |
| TW201908965A | Taiwan Province of China | A | |
| KR20190111132A | Republic of Korea | A | |
| EP3568756A1 | European Patent Office (EPO) | A1 | |
| CN110622134A | China | A | |
| JP2020519981A | Japan | A | |
| TWI728247B | Taiwan Province of China | B | |
| TW202132978A | Taiwan Province of China | A | |
| JP6938661B2 | Japan | B2 | |
| KR102312264B1 | Republic of Korea | B1 | |
| KR20210123435A | Republic of Korea | A | |
| JP2022003532A | Japan | A | |
| US11275992B2 | United States of America | B2 | |
| TWI769810B | Taiwan Province of China | B | |
| EP3568756B1 | European Patent Office (EPO) | B1 | |
| US2022261622A1 | United States of America | A1 | |
| DK3568756T3 | Denmark | T3 | |
| EP4083789A1 | European Patent Office (EPO) | A1 | |
| KR102481428B1 | Republic of Korea | B1 | |
| KR20230003443A | Republic of Korea | A | |
| TW202311939A | Taiwan Province of China | A | |
| CN110622134B | China | B | |
| JP7314217B2 | Japan | B2 | |
| TWI812254B | Taiwan Province of China | B | |
| CN116644790A | China | A | |
| JP2023145517A | Japan | A | |
| TW202414199A | Taiwan Province of China | A | |
| KR102661910B1 | Republic of Korea | B1 | |
| KR20240056801A | Republic of Korea | A | |
| EP4361832A2 | European Patent Office (EPO) | A2 | |
| EP4083789B1 | European Patent Office (EPO) | B1 | |
| FI4083789T3 | Finland | T3 | |
| DK4083789T3 | Denmark | T3 | |
| EP4361832A3 | European Patent Office (EPO) | A3 | |
| TWI857742B | Taiwan Province of China | B | |
| JP7561925B2This record | Japan | B2 | |
| US12524658B2 | United States of America | B2 |
7 legal events, as the office reported them to INPADOC
Over the term
Point at a mark for the eventEvents
| Event | Code | |
|---|---|---|
| Certificate of patent or registration of utility modelJAPANESE INTERMEDIATE CODE: R150R150 | R150 | |
| First payment of annual fees (during grant procedure)JAPANESE INTERMEDIATE CODE: A61A61 | A61 | |
| Written decision to grant a patent or to grant a registration (utility model)JAPANESE INTERMEDIATE CODE: A01A01 | A01 | |
| Decision of grant or rejection writtenTRDD | TRDD | |
| Report on retrievalJAPANESE INTERMEDIATE CODE: A971007A977 | A977 | |
| Request for written amendment filedJAPANESE INTERMEDIATE CODE: A523A521 | A521 | |
| Written request for application examinationJAPANESE INTERMEDIATE CODE: A621A621 | A621 |
Numbers
- Publication
- 7561925
- Application
- 114361
Titles2
- Japanese
- 専用ニューラルネットワークトレーニングチップ
- English
- Dedicated Neural Network Training Chip
Classification
- CPC, 14
- G06N3/063
- G06F17/16
- G06N3/084
- G06F9/3887
- G06F9/30036
- G06F9/3001
- G06F9/30032
- G06F9/30141
- G06F9/3885
- Y02D10/00
- G06N3/0464
- G06N3/09
- G06N3/0495
- G06N3/08
- IPC, 7
- G06F9 38
- G06F9 30
- G06F15 173
- G06F15 80
- G06F17 16
- G06F17 10
- G06N3 063
