Programmable output block for simulating neural memory in deep learning artificial neural network
Abstract
Many embodiments of programmable output blocks for use with VMM arrays within artificial neural networks are disclosed. In one embodiment, the gain of the output block is configurable via a configuration signal. In another embodiment, the resolution of the ADC in the output block is configurable via a configuration signal.

Term
15 yearsto projected expiry
Projected expiry 5 October 2041, counted from filing; an application has no term until it is granted.
- Priority
- Filed
- Published
- Today
- Projected expiry
14 claims: 3 independent, 11 dependent
- 1一种用于生成神经网络存储器阵列的输出的可编程神经元输出块,包括: 一个或多个输入节点,所述一个或多个输入节点用于从神经网络存储器阵列接收电流;和增益配置电路,所述增益配置电路用于接收增益配置信号并且响应于所述增益配置信号而将增益因子应用于所接收的电流以生成输出。
- 2根据权利要求1所述的可编程神经元输出块,其中所述增益配置信号包括模拟信号。
- 3根据权利要求1所述的可编程神经元输出块,其中所述增益配置信号包括数字位。
- 4根据权利要求1所述的可编程神经元输出块,其中所述增益配置电路包括由所述增益配置信号控制的可变电阻器。
- 5根据权利要求1所述的可编程神经元输出块,其中所述增益配置电路包括由所述增益配置信号控制的可变电容器。
- 6根据权利要求1所述的可编程神经元输出块,其中所述增益配置信号取决于从其接收所述电流的所述神经网络存储器阵列中启用的数目或行或列。
- 7一种用于生成神经网络存储器阵列的输出的可编程神经元输出块,包括: 模数转换器,所述模数转换器用于接收来自所述神经网络存储器阵列的电流以及控制信号并且生成数字输出,其中所述数字输出的分辨率由所述控制信号确定。
- 8根据权利要求7所述的可编程神经元输出块,其中所述控制信号包括模拟信号。
- 9根据权利要求7所述的可编程神经元输出块,其中所述控制信号包括数字位。
- 10一种用于生成神经网络存储器阵列的输出的可编程神经元输出块,包括: 混合模数转换器,所述混合模数转换器用于将所述神经网络存储器阵列的所述输出转换为数字输出,所述混合模数转换器包括: 第一模数转换器,所述第一模数转换器用于生成所述数字输出的第一部分;和第二模数转换器,所述第二模数转换器用于生成所述数字输出的第二部分。
- 11根据权利要求10所述的可编程神经元输出块,其中所述第一模数转换器包括逐次逼近寄存器模数转换器。
- 12根据权利要求11所述的可编程神经元输出块,其中所述第二模数转换器包括串行模数转换器。
- 13根据权利要求10所述的可编程神经元输出块,其中所述第一模数转换器包括算法模数转换器。
- 14根据权利要求13所述的可编程神经元输出块,其中所述第二模数转换器包括串行模数转换器。
Independent claims14
354 paragraphs, as filed
Programmable output blocks for simulating neural memory in deep learning artificial neural networks
[0001] Priority statement
This application is submitted on March 14, 2019 and is titled System for Converting Neuron Current Into Neuron Current-Based Time Pulses in an Analog Neural Memory in a Deep Learning Artificial Neural Network "US Patent Application No. 16/353,830 Continuation-in-part application, the U.S. patent application filed on March 6, 2019 and titled "System for Converting Neuron Current Into Neuron Current-Based Time Pulses in an Analog Neural Memory in a Deep Learning Artificial Neural Network" Application No. 62/814,813 and filed on January 18, 2019 and titled System for Converting Neuron Current Into Neuron Current - Bas ed Time Pulses Priority is granted to U.S. Provisional Application No. 62/794,492 in an Analog Neural Memory in a Deep Learning Artificial Neural Network, all of which are incorporated herein by reference.
Technical field
[0003] A number of implementations of programmable output blocks for use with vector-matrix multiplication (VMM) arrays within artificial neural networks are disclosed.
Background technique
[0004] Artificial neural networks simulate biological neural networks (the central nervous system of animals, especially the brain) and are used to estimate or approximate functions that may depend on a large number of inputs and are often unknown. Artificial neural networks typically include layers of interconnected "neurons" that exchange messages with each other.
[0005] Figure 1 illustrates an artificial neural network, where circles represent inputs or layers of neurons. Connections (called synapses) are represented by arrows and have numerical weights that can be adjusted empirically. This allows the neural network to adapt to the input and learn. Typically, neural networks include layers with multiple inputs. There are typically one or more intermediate layers of neurons, and an output layer of neurons that provide the output of the neural network. Neurons at each level individually or collectively make decisions based on data received from synapses.
[0006] One of the major challenges in developing artificial neural networks for high performance information processing is the lack of adequate hardware technology. In fact, real neural networks rely on a large number of synapses to achieve high connectivity between neurons, that is, very high computational parallelism. In principle, such complexity could be achieved with digital supercomputers or clusters of dedicated graphics processing units. However, in addition to high cost, these methods are also mediocre in energy efficiency compared to biological networks, which consume less energy mainly due to their execution of low-precision simulation calculations. CMOS analog circuits have been used in artificial neural networks, but most CMOS implementations of synapses are too bulky due to the large number of neurons and synapses required.
[0007] Applicants previously disclosed in U.S. Patent Application No. 15/594,439 (published as U.S. Patent Publication 2017/0337466) an artificial (analog) neural network utilizing one or more non-volatile memory arrays as synapses, This patent application is incorporated herein by reference. The non-volatile memory array operates as simulated neuromorphic memory. A neural network device includes a first plurality of synapses configured to receive a first plurality of inputs and generate a first plurality of outputs therefrom, and a first plurality of neurons configured to receive a first plurality of outputs. The first plurality of synapses includes a plurality of memory cells storing therein
Each of the memory cells includes: a spaced source region and a drain region formed in a semiconductor substrate, wherein a channel region extends between the source region and the drain region; a floating gate over and insulated from the first portion; and a non-floating gate disposed over and insulated from the second portion of the channel region. Each memory cell of the plurality of memory cells is configured to store a weight value corresponding to a plurality of electrons on the floating gate. A plurality of memory units are configured to multiply a first plurality of inputs by the stored weight values to generate a first plurality of outputs.
[0008] Each non-volatile memory cell used in a simulated neuromorphic memory system must be erased and programmed to maintain a very specific and precise amount of charge (ie, the number of electrons) in the floating gate. For example, each floating gate must hold one of N different values, where N is the number of different weights that can be indicated by each cell. Examples of N include 16, 32, 64, 128, and 256.
[0009] One challenge faced by systems using VMM arrays is the ability to accurately measure the output of the VMM array and transmit that output to another stage (such as the input block of another VMM array). Many methods are known, but each method has certain disadvantages, such as current leakage leading to information loss.
[0010] What is needed is an improved output block for receiving output current from a VMM array and converting the output current into a form more suitable for transmission to another stage of electronics.
Contents of the invention
[0011] Many embodiments of programmable output blocks for use with VMM arrays within artificial neural networks are disclosed. In one embodiment, the gain of the output block is configurable via a configuration signal. In another embodiment, the resolution of the ADC in the output block is configurable via a configuration signal.
Description of the drawings
[0012] FIG. 1 is a schematic diagram showing a prior art artificial neural network.
[0013] Figure 2 illustrates a prior art split-gate flash memory cell.
FIG. 3 illustrates another prior art split-gate flash memory cell.
[0015] FIG. 4 illustrates another prior art split-gate flash memory cell.
[0016] FIG. 5 illustrates another prior art split-gate flash memory cell.
[0017] FIG. 6 illustrates another prior art split-gate flash memory cell.
[0018] Figure 7 illustrates a prior art stacked gate flash memory cell.
[0019] FIG. 8 is a schematic diagram illustrating different levels of an exemplary artificial neural network utilizing one or more non-volatile memory arrays.
[0020] Figure 9 is a block diagram illustrating a vector-matrix multiplication system.
[0021] FIG. 10 is a block diagram illustrating an exemplary artificial neural network using one or more vector-matrix multiplication systems.
[0022] Figure 11 shows another embodiment of a vector-matrix multiplication system.
[0023] Figure 12 shows another embodiment of a vector-matrix multiplication system.
[0024] Figure 13 shows another embodiment of a vector-matrix multiplication system.
[0025] Figure 14 shows another embodiment of a vector-matrix multiplication system.
[0026] Figure 15 shows another embodiment of a vector-matrix multiplication system.
[0027] Figure 16 illustrates a prior art long and short term memory system.
[0028] Figure 17 illustrates example cells used in a long short term memory system.
[0029] Figure 18 illustrates one embodiment of the exemplary unit of Figure 17.
[0030] Figure 19 illustrates another embodiment of the exemplary unit of Figure 17.
[0031] Figure 20 shows a prior art gated recursive unit system.
[0032] Figure 21 illustrates an exemplary cell used in a gated recursive cell system.
[0033] Figure 22 illustrates one embodiment of the exemplary unit of Figure 21.
[0034] Figure 23 illustrates another embodiment of the exemplary unit of Figure 21.
[0035] Figure 24 illustrates another embodiment of a vector-matrix multiplication system.
[0036] Figure 25 shows another embodiment of a vector-matrix multiplication system.
[0037] Figure 26 shows another embodiment of a vector-matrix multiplication system.
[0038] Figure 27 shows another embodiment of a vector-matrix multiplication system.
[0039] Figure 28 illustrates another embodiment of a vector-matrix multiplication system.
[0040] Figure 29 shows another embodiment of a vector-matrix multiplication system.
[0041] Figure 30 shows another embodiment of a vector-matrix multiplication system.
[0042] Figure 31 illustrates another embodiment of a vector-matrix multiplication system.
[0043] Figure 32 shows a VMM system.
[0044] Figure 33 illustrates a flash memory simulated neural memory system.
[0045] Figure 34A shows an integrating type analog-to-digital converter.
[0046] FIG. 34B shows the voltage characteristics of the integrating type analog-to-digital converter of FIG. 34A.
[0047] Figure 35A shows an integrating type analog-to-digital converter.
[0048] FIG. 35B shows the voltage characteristics of the integrating type analog-to-digital converter of FIG. 35A.
[0049] FIGS. 36A and 36B show waveforms of an operation example of the analog-to-digital converter of FIGS. 34A and 35A.
[0050] Figure 36C shows a timing control circuit.
[0051] Figure 37 shows a pulse-to-voltage converter.
[0052] Figure 38 shows a current-to-voltage converter.
[0053] Figure 39 shows a current-to-voltage converter.
[0054] Figure 40 shows a current-to-log voltage converter.
[0055] Figure 41 shows a current-to-log voltage converter.
[0056] Figure 42 shows a digital data to voltage converter.
[0057] Figure 43 shows a digital data to voltage converter.
[0058] Figure 44 shows a reference array.
[0059] Figure 45 shows a digital comparator.
[0060] Figure 46 shows the converter and digital comparator.
[0061] Figure 47 shows an analog comparator.
[0062] Figure 48 shows the converter and analog comparator.
[0063] Figure 49 shows the output circuit.
[0064] Figure 50 shows one aspect of an output that is activated after digitization.
[0065] Figure 51 shows one aspect of an output that is activated after digitization.
[0066] Figure 52 shows a charge summer circuit.
[0067] Figure 53 shows a current summer circuit.
[0068] Figure 54 shows a digital summer circuit.
[0069] Figures 55A and 55B illustrate digital bit-to-pulse row converters and waveforms, respectively.
[0070] Figure 56 illustrates a power management method.
[0071] Figure 57 illustrates another power management method.
[0072] Figure 58 illustrates another power management method.
[0073] Figure 59 shows a programmable neuron output block.
[0074] Figure 60 shows a programmable analog-to-digital converter used in a neuron output block.
[0075] Figures 61A, 61B, and 61C illustrate a hybrid output conversion block.
[0076] Figure 62 shows a configurable serial analog-to-digital converter.
[0077] Figure 63 shows a configurable neuron SAR (successive approximation register) analog-to-digital converter.
[0078] Figure 64 shows a pipelined SAR ADC circuit.
[0079] Figure 65 shows a hybrid SAR and serial ADC circuit.
[0080] Figure 66 shows the algorithm ADC output block.
Detailed ways
[0081] The artificial neural network of the present invention utilizes a combination of CMOS technology and non-volatile memory arrays.
[0082] Non-volatile memory unit
[0083] Digital non-volatile memories are well known. For example, U.S. Patent 5,029,130 ("the '130 Patent"), which is incorporated herein by reference, discloses a split-gate non-volatile memory cell array, which is one type of flash memory cell. Such a memory unit 210 is shown in FIG. 2 . Each memory cell 210 includes a source region 14 and a drain region 16 formed in the semiconductor substrate 12 with a channel region 18 therebetween. Floating gate 20 is formed over and insulates from (and controls the conductivity of) a first portion of channel region 18 and over a portion of source region 14 . Wordline terminal 22 (which is typically coupled to the wordline) has a first portion disposed over and insulated from (and controls the conductivity of) a second portion of channel region 18 , and upwardly A second portion extending above floating gate 20 . Floating gate 20 and word line terminal 22 are insulated from substrate 12 by gate oxide. Bit line 24 is coupled to drain region 16 .
[0084] Memory cell 210 is erased (where electrons are removed from the floating gate) by placing a high positive voltage on word line terminal 22, which causes the electrons on floating gate 20 to tunnel via Fowler-Nordheim Tunneling through the intervening insulator from floating gate 20 to word line terminal 22.
[0085] Memory cell 210 is programmed by placing a positive voltage on word line terminal 22 and a positive voltage on source region 14 (where electrons are placed on the floating gate). Electron current will flow from source region 14 to drain region 16 . When electrons reach the gap between word line terminal 22 and floating gate 20, they will accelerate and become hot. Due to the electrostatic attraction from floating gate 20, some heated electrons will be injected onto floating gate 20 through the gate oxide.
[0086] Memory cell 210 is read by placing a positive read voltage across drain region 16 and word line terminal 22 (which turns on the portion of channel region 18 below the word line terminal). If floating gate 20 is positively charged (i.e., electrons are erased), the portion of channel region 18 below floating gate 20 is also turned on, and current will flow through channel region 18 which is sensed Tested as erased state or "1" state. If floating gate 20 is negatively charged (ie, programmed by electrons), the portion of the channel region below floating gate 20 is mostly or completely turned off, and no (or very little) current will flow. Through channel region 18, the channel region is sensed as a programmed state or "0" state.
Table 1 illustrates terminals that may be applied to memory cell 110 for performing read operations, erase operations, and programming operations.
Typical voltage range for operation:
[0088]
Table 1: Operation of flash memory unit 210 of Figure 2
<td></td><td>wL</td><td>BL</td><td>SL</td>
<td>read</td><td>2-3V</td><td>0.6-2V</td><td>0V</td>
<td>Erase</td><td>About 11-13v</td><td>0V</td><td>0V</td>
<td>programming</td><td>1-2V</td><td>1-3μΆ</td><td>9-10V</td>
[0090] FIG. 3 shows a memory cell 310, which is similar to the memory cell 210 of FIG. 2, but with the addition of a control gate (CG).
28. Control gate 28 is biased at a high voltage (e.g., 10V) during programming, at a low or negative voltage (e.g., 0v/-8V) during erasing, and at a low voltage during reading. Low or medium voltage (for example, 0v/2.5V). The other terminals are offset similar to Figure 2.
4 illustrates a four-gate memory cell 410 including a source region 14, a drain region 16, a floating gate 20 over a first portion of channel region 18, a second portion of channel region 18 Overlying select gate 22 (generally coupled to word line WL), control gate 28 over floating gate 20 , and erase gate 30 over source region 14 . This configuration is described in US Patent 6,747,310, which is incorporated herein by reference for all purposes. Here, with the exception of floating gate 20, all gates are non-floating gates, which means that they are electrically connected or capable of being electrically connected to a voltage source. Programming is performed by heated electrons from channel region 18 injecting themselves into floating gate 20 . Erasing is performed by electrons tunneling from floating gate 20 to erase gate 30 .
[0092] Table 2 shows typical voltage ranges that may be applied to the terminals of memory cell 310 for performing read operations, erase operations, and programming operations:
Table 2: Operation of flash memory unit 410 of Figure 4
[0094]
<td></td><td>WL/SG</td><td>BL</td><td>CG</td><td>EG</td><td>SL</td>
<td>read</td><td>1.0-2V</td><td>0.6-2V</td><td>0-2.6V</td><td>0-2.6V</td><td>0V</td>
<td>Erase</td><td>-0.5V/0V</td><td>0V</td><td>0V/-8V</td><td>8-12V</td><td>0V</td>
<td>programming</td><td>1V</td><td>1μΆ</td><td>8-11V</td><td>4.5-9V</td><td>4.5-5V</td>
FIG. 5 shows a memory cell 510 that is similar to the memory cell of FIG. 4 except that it does not include the erase gate EG.
[0095] Yuan 410 is similar. Erasing is performed by biasing substrate 18 to a high voltage and control gate CG 28 to a low or negative voltage. Alternatively, erasing is performed by biasing word line 22 to a positive voltage and control gate 28 to a negative voltage. Programming and reading are similar to those in Figure 4.
[0096] FIG. 6 shows a tri-gate memory cell 610, which is another type of flash memory cell. Memory cell 610 is the same as memory cell 410 of FIG. 4 except that memory cell 610 does not have a separate control gate. The erase operation (thereby erasing is performed using the erase gate) and the read operation are similar to those of Figure 4 except that no control gate bias is applied. The programming operation is also accomplished without control gate bias, and as a result, a higher voltage must be applied to the source line during the programming operation to compensate for the lack of control gate bias.
[0097] Table 3 shows typical voltage ranges that may be applied to the terminals of memory cell 610 for performing read operations, erase operations, and programming operations:
[0098]
Table 3: Operation of flash memory unit 610 of Figure 6
<td></td><td>WL/SG</td><td>BL</td><td>EG</td><td>SL</td>
<td>read</td><td>0.7-2.2V</td><td>0.6-2V</td><td>0-2.6V</td><td>0V</td>
<td>Erase</td><td>-0.5V/0V</td><td>0V</td><td>11.5V</td><td>0V</td>
<td>programming</td><td>1V</td><td>2-3μA</td><td>4.5V</td><td>7-9V</td>
[0100] FIG. 7 shows a stacked gate memory cell 710, which is another type of flash memory cell. Memory cell 710 is similar to memory cell 210 of FIG. 2 except that floating gate 20 extends over the entire channel region 18 and control gate 22 (which here will be coupled to the word line) extends over floating gate 20 by The insulation layer (not shown) separates. Erase, program, and read operations operate in a similar manner as previously described for memory cell 210.
[0101] Table 4 shows typical voltage ranges that may be applied to the terminals of memory cell 710 and substrate 12 for performing read operations, erase operations, and programming operations:
Table 4: Operation of flash memory unit 710 of Figure 7
<td></td><td>CG</td><td>BL</td><td>SL</td><td>substrate</td>
<td>read</td><td>2-5V</td><td>0.6-2V</td><td>0V</td><td>0V</td>
<td>Erase</td><td>-8 to -10V/0V</td><td>FLT</td><td>FLT</td><td>8-10V/15-20V</td>
<td>programming</td><td>8-12V</td><td>3-5V</td><td>0V</td><td>0V</td>
[0104] In order to utilize a memory array including one of the types of non-volatile memory cells described above in an artificial neural network, two modifications were made. First, the circuitry is configured so that each memory cell can be programmed, erased, and read individually without adversely affecting the memory state of other memory cells in the array, as explained further below. Second, it provides continuous (analog) programming of memory cells.
[0105] Specifically, the memory state (i.e., the charge on the floating gate) of each memory cell in the array can continuously change from a fully erased state to a fully erased state independently and with minimal interference to other memory cells. programming status. In another embodiment, the memory state (i.e., charge on the floating gate) of each memory cell in the array can be continuously changed from a fully programmed state to a fully erased state independently and with minimal disruption to other memory cells. except status and vice versa. This means that the cell memory device is analog, or at least can store one discrete value out of many discrete values (such as 16 or 64 different values), which allows for very precise and individual tuning of all cells in the memory array, And this makes the memory array ideal for storing and fine-tuning the synaptic weights of neural networks.
[0106] The methods and apparatus described herein may be applied to other non-volatile memory technologies such as SONOS (silicon-oxide-nitride-oxide-silicon, charge trapped in nitride), MONOS (metal-oxide Material - nitride - oxide - silicon, metal charges trapped in nitride), ReRAM (resistive ram), PCM (phase change memory), MRAM (magnetic ram), FeRAM (ferroelectric ram), OTP (double layer or multi-layer one-time programmable) and CeRAM (associated electronic ram), etc. The methods and apparatus described herein can be applied to volatile memory technologies for neural networks, such as but not limited to SRAM, DRAM and/or volatile Sexual synaptic unit.
[0107] Neural Network Employing Non-Volatile Memory Cell Arrays
[0108] Figure 8 conceptually illustrates a non-limiting example of a neural network utilizing a non-volatile memory array of this embodiment. This example uses a non-volatile memory array neural network for a facial recognition application, but any other suitable application can be implemented using a non-volatile memory array based neural network.
[0109] For this example, S0 is the input layer, which is a 32 x 32 pixel RGB image with 5 bits of precision (i.e., three 32 x 32 pixel arrays, one for each color R, G, and B, each Pixels are 5-bit precision). Synapse CB1 from input layer S0 to layer C1 applies a different set of weights in some cases and shared weights in other cases, and scans the input image with a 3 X 3 pixel overlapping filter (kernel), shifting the filter 1 pixel (or more than 1 pixel as indicated by the model). Specifically, the values of 9 pixels in a 3 After the outputs of the multiplications are summed, the sum is determined by the first synapse of CB1
Provides a single output value for the pixels of one of the layers C1 used to generate the feature map. The 3 The 9 pixel values in the detector are provided to synapse CB1, where they are multiplied by the same weight and a second single output value is determined by the associated synapse. Continue this process until the 3 X 3 filter scans all three colors and all bits (precision values) over the entire 32 X 32 pixel image of input layer S0. The process is then repeated using different sets of weights to generate different feature maps for C1 until all feature maps for layer C1 are calculated.
[0110] At layer C1, in this example, there are 16 feature maps, each with 30 × 30 pixels. Each pixel is a new feature pixel extracted from the product of the input and the kernel, so each feature map is a 2D array, so in this example layer C1 consists of 16 layers of 2D arrays (remember, as quoted in this article The relationship between layers and arrays is logical, not necessarily physical, i.e. the arrays do not have to be oriented in a physical two-dimensional array). Each of the 16 feature maps in layer C1 is generated by one of sixteen different sets of synaptic weights applied to the filter sweep. C1 feature maps may all relate to different aspects of the same image feature, such as boundary recognition. For example, a first map (generated using a first set of weights, shared for all scans used to generate it) identifies rounded edges, and a second map (generated using a second set of weights that is different from the first set of weights) ) can identify rectangular edges, or the aspect ratio of certain features, and so on.
[0111] Before going from layer C1 to layer S1, an activation function P1 (pooling) is applied, which pools values from contiguous non-overlapping 2 × 2 regions in each feature map. The purpose of the pooling function is to average the neighboring locations (or the max function can also be used) to, for example, reduce the dependence on edge locations and reduce the data size before entering the next stage. At layer S1, there are sixteen 15X 15 feature maps (ie, sixteen different arrays of 15X 15 pixels per feature map). Synapse CB2 from layer S1 to layer C2 scans the map in S1 with a 4X4 filter, where the filter is shifted by 1 pixel. At layer C2, there are 22 12X12 feature maps. Before going from layer C2 to layer S2, an activation function P2 (pooling) is applied, which pools values from consecutive non-overlapping 2 X 2 regions in each feature map. At layer S2, there are 22 6 X 6 feature maps. Apply an activation function (pooling) to synapse CB3 from layer S2 to layer C3, where each neuron in layer C3 passes through CB3 The corresponding synapse of is connected to each map in layer S2. At layer C3, there are 64 neurons. Synapse CB4 from layer C3 to output layer S3 completely connects C3 to S3, i.e. every neuron in layer C3 is connected to every neuron in layer S3. The output at S3 consists of 10 neurons, with the highest output neuron determining the class. For example, the output may indicate identification or classification of the content of the original image.
[0112] The synapses of each layer are implemented using an array or a portion of an array of non-volatile memory cells.
[0113] Figure 9 is a block diagram of an array that may be used for this purpose. Vector-matrix multiplication (VMM) array 32 includes non-volatile memory cells and serves as a synapse between one layer and the next (such as CB1, CB2, CB3 and CB4 in Figure 6). Specifically, the VMM array 32 includes a non-volatile memory cell array 33, an erase gate and word line gate decoder 34, a control gate decoder 35, a bit line decoder 36 and a source line decoder 37, which are The corresponding input to the non-volatile memory cell array 33 is decoded. Inputs to VMM array 32 may come from erase gate and word line gate decoders 34 or from control gate decoders 35. In this example, source line decoder 37 also decodes the output of non-volatile memory cell array 33. Alternatively, bit line decoder 36 may decode the output of non-volatile memory cell array 33.
[0114] Non-volatile memory cell array 33 serves two purposes. First, it stores the weights to be used by the VMM array 32. Second, the non-volatile memory cell array 33 effectively multiplies the input with the weights stored in the non-volatile memory cell array 33 and each output line (source line or bit line) sums them to produce an output, This output will serve as input to the next layer or to the final layer. By performing multiply and add functions, the non-volatile memory cell array 33 eliminates the need for separate multiply and add logic circuits, and is also power efficient due to its in-situ memory computing.
The output of the non-volatile memory cell array 33 is provided to a differential summer (such as a summing operational amplifier or a summing current mirror) 38, which The summation is performed to create a single value for this convolution. The difference summer 38 is arranged for performing the summation of positive and negative weights.
[0116] The output values of the differential summer 38 are then summed and provided to the activation function circuit 39, which corrects the output. The activation function circuit 39 can provide sigmoid, tanh or ReLU functions. The modified output values of the activation function circuit 39 become elements of the feature map of the next layer (e.g., layer C1 in Figure 8) and are subsequently applied to the next synapse to produce the next feature map layer or final layer. . Thus, in this example, the non-volatile memory cell array 33 constitutes a plurality of synapses (which receive their input from existing neuron layers or from an input layer such as an image database), and the summer 38 and activation function circuits 39 constitute multiple neurons.
The inputs to VMM array 32 in Figure 9 (WLx, EGx, CGx and optionally BLx and SLx) may be analog levels, binary levels, digital pulses (in which case pulse-analog may be required Converter PAC to convert the pulses to the appropriate input analog level) or digital bits (in this case, a DAC is provided to convert the digital bits to the appropriate input analog level); the output can be analog levels, binary levels levels, digital pulses, or digital bits (in this case, an output ADC is provided to convert the output analog levels into digital bits).
[0118] FIG. 10 is a block diagram illustrating the use of a multi-layer VMM array 32 (labeled here as VMM arrays 32a, 32b, 32c, 32d, and 32e). As shown in Figure 10, the input (denoted as Inputx) is converted from digital to analog by a digital-to-analog converter 31 and provided to the input VMM array 32a. The converted analog input can be voltage or current. Input D/A conversion of the first layer may be accomplished using a function or LUT (look-up table) that maps the input Inputx to the appropriate analog levels of the matrix multiplier of the input VMM array 32a. Input conversion may also be accomplished by an analog-to-analog (A/A) converter to convert external analog inputs into mapped analog inputs to input VMM array 32a. Input conversion may also be accomplished by a digital-to-digital pulse (D/P) converter to convert the external digital input into one or more digital pulses mapped to the input VMM array 32a.
The output produced by the input VMM array 32a is provided as an input to the next VMM array (hidden level 1) 32b, which in turn generates the output provided as an input to the next VMM array (hidden level 2) 32c, And so on. The layers of the VMM array 32 serve as different layers of synapses and neurons of a convolutional neural network (CNN). Each VMM array 32a, 32b, 32c, 32d, and 32e may be an independent physical non-volatile memory array, or multiple VMM arrays may utilize different portions of the same non-volatile memory array, or multiple VMM arrays may utilize Overlapping portions of the same physical non-volatile memory array. Each VMM array 32a, 32b, 32c, 32d, and 32e may also be time multiplexed for different portions of its array or neurons. The example shown in Figure 10 contains five layers (32a, 32b, 32c, 32d, 32e): one input layer (32a), two hidden layers (32b, 32c) and two fully connected layers (32d, 32e). One of ordinary skill in the art will appreciate that this is exemplary only and that, instead, the system may include more than two hidden layers and more than two fully connected layers.
[0120] Vector-Matrix Multiplication (VMM) Array
[0121] Figure 11 shows a neuron VMM array 1100 that is particularly suitable for use with the memory cell 310 shown in Figure 3 and serves as a synapse and component of the neurons between the input layer and the next layer. VMM array 1100 includes a memory array 1101 of non-volatile memory cells and a reference array 1102 of non-volatile reference memory cells (at the top of the array). Alternatively, another reference array can be placed at the bottom.
[0122] In VMM array 1100, control gate lines (such as control gate lines 1103) extend in the vertical direction (so reference array 1102 is orthogonal to control gate lines 1103 in the row direction), and erase gate lines (such as erase gate lines) The gate lines 1104) extend in the horizontal direction. Here, the input to the WMM array 1100 is provided on the control gate lines (CG0, CG1, CG2, CG3) and the output of the VMM array 1100 appears on the source lines (SL0, SL1). In one embodiment, only even rows are used, and in another
In this implementation, only odd-numbered rows are used. The current placed on each source line (SL0, SL1 respectively) performs a summation function of all currents from the memory cells connected to that particular source line.
[0123] As described herein for neural networks, the non-volatile memory cells of VMM array 1100 (ie, the flash memory of VMM array 1100) are preferably configured to operate in the sub-threshold region.
[0124] Biasing non-volatile reference memory cells and non-volatile memory cells described herein in weak inversion:
[0125] Ids = lo*e<sup>(Vg</sup><sup>m)</sup> '<sup>htK</sup>=w*Io*e<sup>(Vg)/kVt</sup>,
[0126] Wherein w=e<sup>(</sup>-<sup>Vth)/kVt</sup>
[0127] For a I to V logarithmic converter that uses a memory cell (such as a reference memory cell or a peripheral memory cell) or a transistor to convert input current to input voltage:
[0128] Vg = k*Vt*10g[Ids/wp*Io]
[0129] Here, wp is w of the reference memory unit or peripheral memory unit.
[0130] For a memory array used as a vector matrix multiplier VMM array, the output current is:
[0131] Iout=wa*Io*e<sup>(Vg)/kVt</sup>,Right now
[0132]
[0133]
Iout= (wa/wp)*Iin=W*Iin
W - e (Vthp-Vtha)/kVt
[0134] Here, wa = w for each memory cell in the memory array.
[0135] A word line or control gate may be used as an input to a memory cell for input voltage.
[0136] Alternatively, the flash memory cells of the VMM arrays described herein may be configured to operate in the linear region:
Ids=P*(Vgs-Vth)*Vds; P=u*Cox*W/L
[0138] W=a (Vgs-Vth)
[0139] Word lines or control gates or bit lines or source lines may be used as inputs to memory cells operating in the linear region. Bit lines or source lines can be used as the output of the memory cell.
[0140] For an I to V linear converter, a memory cell (such as a reference memory cell or a peripheral memory cell) or a transistor or resistor operating in the linear region can be used to linearly convert the input/output current into the input/output voltage.
[0141] Other embodiments of VMM array 32 of Figure 9 are described in U.S. Patent Application No. 15/826,345, which is incorporated herein by reference. As described herein, source lines or bit lines can be used as neuron outputs (current summation outputs). Alternatively, the flash memory cells of the VMM array described herein may be configured to operate in the saturation region:
[0142] Ids=a1/2*p*(Vgs-Vth)<sup>2</sup>;p = u*Cox*W/L
[0143] W=a (Vgs-Vth)<sup>2</sup>
[0144] Word lines, control gates, or erase gates may be used as inputs to memory cells operating in the saturation region. Bit lines or source lines can be used as the output of the output neuron.
[0145] Alternatively, the flash memory cells of the VMM arrays described herein may be used in all regions or combinations thereof (sub-threshold, linear or saturation regions).
[0146] Figure 12 shows a neuron VMM array 1200 that is particularly suitable for use with the memory cell 210 shown in Figure 2 and serves as a synapse between the input layer and the next layer. VMM array 1200 includes a memory array 1203 of non-volatile memory cells, a reference array 1201 of first non-volatile reference memory cells, and a reference array 1202 of second non-volatile reference memory cells. Reference arrays 1201 and 1202 arranged in the column direction of the array serve to convert current inputs flowing into terminals BLR0, BLR1, BLR2 and BLR3 into voltage inputs WL0, WL1, WL2 and WL3. In practice, the first non-volatile reference memory cell and the second non-volatile reference memory cell are diode-connected via multiplexer 1214 (only partially shown), where
Current input flows into it. The reference unit is tuned (eg, programmed) to the target reference level. The target reference level is provided by a reference microarray matrix (not shown).
[0147] Memory array 1203 serves two purposes. First, it stores the weights that the VMM array 1200 will use on their corresponding memory cells. Second, memory array 1203 effectively converts the inputs (i.e., current inputs provided in terminals BLR0, BLR1, BLR2, and BLR3) to reference arrays 1201 and 1202 into input voltages for supply to word lines WLO, WL1 , WL2 and WL3) are multiplied by the weights stored in the memory array 1203 and then all the results (memory cell currents) are summed to produce an output on the corresponding bit line (BLO-BLN) which will go to the next layer input to or to the final layer. By performing multiply and add functions, memory array 1203 eliminates the need for separate multiply logic circuits and add logic circuits, and is also highly power efficient. Here, voltage inputs are provided on word lines (WLO, WL1, WL2 and WL3) and outputs appear on the corresponding bit lines (BLO-BLN) during read (inference) operations. The current placed on each of the bit lines BLO-BLN performs a summation function of the currents from all non-volatile memory cells connected to that particular bit line.
[0148] Table 5 shows the operating voltages for the VMM array 1200. The columns in the table indicate the placement of word lines for selected cells, word lines for unselected cells, bit lines for selected cells, bit lines for unselected cells, sources for selected cells. The voltage on the pole line and the source line for unselected cells. Rows indicate read, erase, and program operations.
Table 5: Operation of VMM array 1200 of Figure 12:
<td></td><td>wL</td><td>WL-not selected</td><td>BL</td><td>BL-not selected</td><td>SL</td><td>SL moxuan</td>
<td>read</td><td>1-3.5V</td><td>-0.5V/OV</td><td>0.6-2V(Ineuron)</td><td>0.6V-2V/0V</td><td>ον</td><td>ον</td>
<td>Erase</td><td>About 5-13V</td><td>0V</td><td>0V</td><td>ον</td><td>ον</td><td>ον</td>
<td>programming</td><td>1-2V</td><td>-0-5V/0V</td><td></td><td>Vicih about 2.5V</td><td>4-10V</td><td>0-1V/FLT</td>
[0151] Figure 13 shows a neuron VMM array 1300 that is particularly suitable for use with the memory cell 210 shown in Figure 2 and serves as a synapse and component of the neurons between the input layer and the next layer. VMM array 1300 includes a memory array 1303 of non-volatile memory cells, a reference array 1301 of first non-volatile reference memory cells, and a reference array 1302 of second non-volatile reference memory cells. Reference arrays 1301 and 1302 extend in the row direction of VMM array 1300. The VMM array is similar to VMM 1000, except that in VMM array 1300, the word lines extend in the vertical direction. Here, the inputs are set on the word lines (WLAO, WLBO, WLA1, WLB2, WLA2, WLB2, WLA3, WLB3) and the outputs appear on the source lines (SLO, SL1) during the read operation. The current placed on each source line performs a summation function of all currents from the memory cells connected to that particular source line.
[0152] Table 6 shows the operating voltages for the VMM array 1300. The columns in the table indicate the placement of word lines for selected cells, word lines for unselected cells, bit lines for selected cells, bit lines for unselected cells, sources for selected cells. The voltage on the pole line and the source line for unselected cells. Rows indicate read, erase, and program operations.
Table 6: Operation of VMM array 1300 of Figure 13
<td></td><td>wL</td><td>WL-not selected</td><td>BL</td><td>BL-not selected</td><td>SL</td><td>SL-not selected</td>
<td>read</td><td>135V</td><td>-0.5V/0V</td><td>0.6-2V</td><td>0.6V-2V/0V</td><td>-0.3-1V(Incuron)</td><td>ον</td>
<td>Erase</td><td>About 5-13V</td><td>0V</td><td>0V</td><td>0V</td><td>0V</td><td>SL-Suppression (approx. 4-8 V)</td>
<td>programming</td><td>1-2V</td><td>-0.5V/0V</td><td>0.1-3uA</td><td>Vinh about 2.5V</td><td>4-10V</td><td>0-1V/FLT</td>
[0155] Figure 14 shows a neuron VMM array 1400 that is particularly suitable for use with the memory cell 310 shown in Figure 3 and serves as a synapse and component of the neurons between the input layer and the next layer. VMM array 1400 includes a memory array 1403 of non-volatile memory cells, a reference array 1401 of first non-volatile reference memory cells, and a reference array 1402 of second non-volatile reference memory cells. Reference arrays 1401 and 1402 are used to convert current inputs flowing into terminals BLRO, BLR1, BLR2 and BLR3 into voltage inputs CGO, CGI, CG2 and CG3. In practice, the first non-volatile reference memory cell and the second non-volatile reference memory cell are diode-connected via a multiplexer 1412 (only partially shown), where the current inputs are via BLRO, BLR1, BLR2 and BLR3 flow into it. Multiplexers 1412 each include a corresponding multiplexer 1405 and cascode transistor 1404 to ensure first non-volatile reference memory cells and second non-volatile reference memory cells during read operations The voltage on the bit line of each of them (such as BLRO) is constant. Tune the reference unit to the target reference level.
[0156] Memory array 1403 serves two purposes. First, it stores the weights that will be used by the VMM array 1400. Second, memory array 1403 effectively provides inputs (current inputs to terminals BLRO, BLR 1, BLR2, and BLR3) and reference arrays 1401 and 1402 convert these current inputs into input voltages to provide to control gates (CGO, CG1, CG2, and CG3)) is multiplied by the weights stored in the memory array and then all the results (cell currents) are summed to produce an output which appears at the BLO-BLN and will be the input to the next layer or the input to the final layer. By performing multiply and add functions, the memory array eliminates the need for separate multiply and add logic circuits and is also highly power-efficient. Here, inputs are provided on the control gate lines (CGO, CG1, CG2, and CG3) and outputs appear on the bit lines (BLO-BLN) during read operations. The current placed on each bit line performs a summation function of all currents from the memory cells connected to that particular bit line.
[0157] VMM array 1400 implements unidirectional tuning of non-volatile memory cells in memory array 1403. That is, each non-volatile memory cell is erased and then partially programmed until the desired charge on the floating gate is reached. This can be performed, for example, using precision programming techniques described below. If too much charge is placed on the floating gate (so that the wrong value is stored in the cell), the cell must be erased and the sequence of partial programming operations must begin again. As shown, two rows sharing the same erase gate (such as EGO or EG1) need to be erased together (it is called page erase) and thereafter, each cell is partially programmed until all the values on the floating gate are reached. Requires charge.
[0158] Table 7 shows the operating voltages for the VMM array 1400. The columns in the table indicate the controls placed on word lines for selected cells, word lines for unselected cells, bit lines for selected cells, bit lines for unselected cells, and controls for selected cells. gate, control gate for unselected cells in the same sector as the selected cell, control gate for unselected cells in a different sector than the selected cell, erase gate for selected cells, for The erase gates of unselected cells, the source lines for selected cells, and the voltage lines on the source lines for unselected cells indicate read, erase, and program operations.
Table 7: Operation of VMM array 1400 of Figure 14
[0160]
<td></td><td>wL</td><td>WL-not selected</td><td>BL</td><td>BL-not selected</td><td>CG</td><td>CG mo select the same sector</td><td>CG-Not selected</td><td>EG</td><td>EG-not selected</td><td>SL</td><td>SL is not selected</td>
<td>read</td><td>1.0-2V</td><td>-0.5V/0V</td><td>0.6-2V (Ineuron)</td><td>ον</td><td>0-2.6V</td><td>0-2.6V</td><td>0-2.6V</td><td>0-2.6V</td><td>0-2.6V</td><td>ον</td><td>ον</td>
<td>Erase</td><td>0V</td><td>0V</td><td>ον</td><td>ον</td><td>ον</td><td>0-2.6V</td><td>0-2.6V</td><td>5-12V</td><td>0-2.6V</td><td>ον</td><td>ον</td>
<td>programming</td><td>0.7-1V</td><td>-0.5V/0V</td><td>0.1-luA</td><td>Vinh(l2V)</td><td>4-1IV</td><td>0-2.6V</td><td>0-2.6V</td><td>4.5-5V</td><td>0-2.6V</td><td>4.5-5V</td><td>0-1V</td>
[0161] Figure 15 shows a neuron VMM array 1500 that is particularly suitable for use with the memory cell 310 shown in Figure 3 and serves as a synapse and component of the neurons between the input layer and the next layer. VMM array 1500 includes a memory array 1503 of non-volatile memory cells, a reference array 1501 of first non-volatile reference memory cells, and a reference array 1502 of second non-volatile reference memory cells.<sub>o</sub> The EG lines EGRO, EGO, EG1 and EGR1 extend vertically, while the CG lines CGO, CGI, CG2 and CG3 and the SL lines WLO, WL1, WL2 and WL3 extend horizontally. The VMM array 1500 is similar to the VMM array 1400, except that the VMM array 1500 implements bi-directional tuning, where each individual cell can be fully erased, partially programmed, and partially erased as needed to float on or off due to the use of separate EG lines. The desired amount of charge is achieved on the gate. As shown, reference arrays 1501 and 1502 convert input currents in terminals BLRO, BLR1, BLR2, and BLR3 into control gate voltages CGO, CG1, CG2, and CG3 to be applied to the memory cells in the row direction (by multiplexing action of the diode-connected reference unit using device 1514). The current outputs (neurons) are in bit lines BLO-BLN, where each bit line sums all currents from the non-volatile memory cells connected to that particular bit line.
[0162] Table 8 shows the operating voltages for the VMM array 1500. The columns in the table indicate the controls placed on word lines for selected cells, word lines for unselected cells, bit lines for selected cells, bit lines for unselected cells, and controls for selected cells. gate, control gate for unselected cells in the same sector as the selected cell, control gate for unselected cells in a different sector than the selected cell, erase gate for selected cells, for The erase gates of unselected cells, the source lines for selected cells, and the voltage lines on the source lines for unselected cells indicate read, erase, and program operations.
Table 8: Operation of VMM array 1500 of Figure 15
<td></td><td>wL</td><td>WL-not selected</td><td>BL</td><td>BL is not selected</td><td>CG</td><td>CG-Same sector not selected</td><td>CG moxuan</td><td>EG</td><td>EG-not selected</td><td>SL</td><td>SL moxuan</td>
<td>read</td><td>1.0-2V</td><td>-0.5V/0V</td><td>0.6-2V (Ineur on)</td><td>OV</td><td>0-2.6V</td><td>0-2.6V</td><td>0-2.6V</td><td>0-2.6V</td><td>0-2.6V</td><td>OV</td><td>ον</td>
<td>Erase</td><td>0V</td><td>ον</td><td>OV</td><td>OV</td><td>OV</td><td>4-9V</td><td>0-2.6V</td><td>5-12V</td><td>0-2.6V</td><td>OV</td><td>ον</td>
<td>programming</td><td>0.7IV</td><td>-0,5V/OV</td><td>OJ-luA</td><td>Vinh(12V)</td><td>4-1IV</td><td>026V</td><td>0-2.6V</td><td>4.5-5V</td><td>0-2,6V</td><td>4.55V</td><td>0-1V</td>
[0165] Figure 24 shows a neuron VMM array 2400 that is particularly suitable for use with the memory cell 210 shown in Figure 2 and serves as a synapse and component of the neurons between the input layer and the next layer. In VMM array 2400, enter
INPUT<sub>0</sub>.....INPUTN is on bit line BL respectively<sub>0</sub>.....Receive on BLN, and output OUTPUT^ OUTPUT2,
OUTPUT^DOUTPUT<sub>4</sub> respectively on the source line SL<sub>0</sub>,SL"SL<sub>2</sub> and SL<sub>3</sub>±generated.
[0166] Figure 25 shows a neuron VMM array 2500 that is particularly suitable for use with the memory cell 210 shown in Figure 2 and serves as a synapse and component of the neurons between the input layer and the next layer. In this example, enter INPUT. , INPUT], INPUT<sub>2</sub>and INPUT<sub>3</sub>respectively on the source line SL<sub>0</sub>,SL"SL<sub>2</sub>and SL<sub>3</sub>Receive on, and output OUTPUT<sub>0</sub>....... OUTPUT<sub>N</sub>Current line BL<sub>0</sub>,...,BL<sub>N</sub>generated on.
[0167] Figure 26 shows a neuron VMM array 2600 that is particularly suitable for use with the memory cell 210 shown in Figure 2 and serves as a synapse and component of the neurons between the input layer and the next layer. In this example, the inputs INPUT0.....INPUTM are on word line WL respectively.<sub>0</sub>........Receive on WLM, and output OUTPUT<sub>0</sub>.......
OUTPUT<sub>N</sub>Current line BL<sub>0</sub>........BL<sub>N</sub>generated on.
[0168] Figure 27 shows a neuron VMM array 2700 that is particularly suitable for use with the memory cell 310 shown in Figure 3 and serves as a synapse and component of the neurons between the input layer and the next layer. In this example, the inputs INPUT0.....INPUTM are on word line WL respectively.<sub>0</sub>........Receive on WLM, and output OUTPUT<sub>0</sub>........
OUTPUT<sub>N</sub>Current line BL<sub>0</sub>........BL<sub>N</sub>generated on.
[0169] Figure 28 shows a neuron VMM array 2800 that is particularly suitable for use with the memory cell 410 shown in Figure 4 and serves as a synapse and component of the neurons between the input layer and the next layer. In this example, input INPUT0.....INPUT<sub>n</sub>respectively on bit line BL<sub>0</sub>........BL<sub>N</sub>Receive on, and output OUTPUT^DOUTPUT<sub>2</sub>On the erase gate line EG<sub>0</sub>and generated on EG1.
[0170] Figure 29 shows a neuron VMM array 2900 that is particularly suitable for use with the memory cell 410 shown in Figure 4 and serves as a synapse and component of the neurons between the input layer and the next layer. In this example, input INPUT0.....INPUT<sub>n</sub>Received on the gates of bit line control gates 2901-1, 2901-2...2901-(N-1) and 2901-N respectively, which gates are respectively coupled to bit line BL<sub>0</sub>........BL<sub>N</sub>,. Example outputs OUTPUT and OUTPUT2 on erase gate line SL<sub>0</sub>and generated on SL1.
[0171] FIG. 30 shows a neuron VMM array 3000 that is particularly suitable for use with the memory unit 310 shown in FIG. 3, the memory unit 510 shown in FIG. 5, and the memory unit 710 shown in FIG. 7, and with Synapses and components of neurons between the input layer and the next layer. In this example, enter INPUT<sub>0</sub>.....INPUT<sub>M</sub>on word line WL<sub>0</sub>........Receive on WLM, and output OUTPUT<sub>0</sub>.....OUTPUT<sub>N</sub>respectively on bit line BL<sub>0</sub>.......BL<sub>N</sub>generated on.
[0172] Figure 31 shows a neuron VMM array 3100 that is particularly suitable for use with the memory unit 310 shown in Figure 3, the memory unit 510 shown in Figure 5, and the memory unit 710 shown in Figure 7, and with Synapses and components of neurons between the input layer and the next layer. In this example, input INPUT0.....INPUT<sub>M</sub>on the control gate line
CG<sub>0</sub>.......CG<sub>M</sub>receive on. OUTPUT<sub>0</sub>.....OUTPUT<sub>N</sub>respectively on the source line SL<sub>0</sub>.....generated on SLN, where each source line SLj is coupled to the source line terminals of all memory cells in column i.
32 illustrates a VMM system 32000. The VMM system 3200 includes a VMM array 3201 (which may be based on any of the previously discussed VMM designs, such as VMMs 900, 1000, 1100, 1200, and 1320, or other VMM designs), low Voltage Row Decoder 3202, High Voltage Row Decoder 3203, Reference Cell Low Voltage Column Decoder 3204 (shown in the column direction, which means that it provides input to output conversion in the row direction), Bit Line Multiplexing 3205, control logic 3206, analog circuit 3207, neuron output block 3208, input VMM circuit block 3209, predecoder 3210, test circuit 3211, erase-program control logic EPCTL 3212, analog and high voltage generation circuit 3213, bit Line PE driver 3214, redundant arrays 3215 and 3216, NVR sector 3217, and reference sector 3218. The input circuit block 3209 serves as an input terminal for input from the outside to the memory array.
Take . Neuron output block 3208 serves as an interface for output from the memory array to external interfaces.
[0174] The low voltage row decoder 3202 provides bias voltages for read operations and programming operations, and provides the decode signal for the high voltage row decoder 3203. High voltage row decoder 3203 provides high voltage bias signals for program operations and erase operations. Reference cell low voltage column decoder 3204 provides decoding functionality for the reference cell. Bit line PE driver 3214 provides control functions for the bit lines during program, verify and erase operations. Analog and high voltage generation circuitry 3213 is a shared bias block that provides multiple voltages required for various program, erase, program verify, and read operations. Redundant arrays 3215 and 3216 provide array redundancy for replacing defective array portions. NVR (Non-Volatile Register, also known as Information Sector) Sector 3217 is a sector that serves as an array sector for storing user information, device ID, passwords, security keys, trim bits, configuration bits, and manufacturing information, etc. district.
[0175] Figure 33 illustrates a simulated neural memory system 3300. Analog neural memory system 3300 includes macroblocks 3301a, 3301b, 3301c, 3301d, 3301e, 330If, 3301g, and 3301h; neuron output blocks (such as summer circuits and sample and hold S/H circuits) 3302a, 3302b, 3302c, 3302d , 3302e, 3302f, 3302g and 3302h; and input circuit blocks 3303a, 3303b, 3303c, 3303d, 3303e, 3303f, 3303g and 3304h. Each of macroblocks 3301a, 3301b, 3301c, 3301d, 3301e, and 3301f is a VMM subsystem that includes a VMM array that includes rows and columns of non-volatile memory cells, such as flash memory cells. Neural memory subsystem 3333 includes macroblocks 3301, input blocks 3303, and neuron output blocks 3302. Neural memory subsystem 3333 may have its own digital control block.
[0176] The analog neural memory system 3300 also includes a system control block 3304, an analog low voltage block 3305, a high voltage block 3306, and a timing control circuit 3670, which components are discussed in further detail below with reference to FIG. 36.
[0177] System control block 3304 may include one or more microcontroller cores such as ARM/MIPS/RISC_V cores to handle general control functions and arithmetic operations. The system control block 3304 may also include a SIMD (Single Instruction Multiple Data) unit to operate on multiple data with a single instruction. The system control block may include a DSP core. The system control block may include hardware for performing functions such as, but not limited to, pooling, average, min, max, softmax, addition, subtraction, multiplication, division, logarithm, antilog, ReLu, sigmoid, tanh, and data compression or software. The system control block may include hardware or software that performs functions such as activating the approximator/quantizer/normalizer. The system control block may include the ability to perform functions such as input data approximator/quantizer/normalizer. The system control block may include hardware or software that performs activation approximator/quantizer/normalizer functions. The control blocks of neural memory subsystem 3333 may include similar elements of system control block 3304, such as microcontroller cores, SIMD cores, DSP cores, and other functional units.
[0178] In one embodiment, neuron output blocks 3302a, 3302b, 3302c, 3302d, 3302e, 3302f, 3302g, and 3302b each include a buffered (e.g., operational amplifier) low impedance output type circuit that can drive a long and Configurable interconnectors. In one embodiment, input circuit blocks 3303a, 3303b, 3303c, 3303d, 3303e, 3303f, 3303g, and 3303b each provide a summed high impedance current output. In another embodiment, neuron output blocks 3302a, 3302b, 3302c, 3302d, 3302e, 3302f, 3302g, and 3302h each include activation circuitry, in which case additional low impedance buffers are required to drive the outputs.
[0179] In another embodiment, neuron output blocks 3302a, 3302b, 3302c, 3302d, 3302e, 3302f, 3302g, and 3302b each include an analog-to-digital conversion block that outputs digital bits rather than analog signals. In this embodiment, input circuit blocks 3303a, 3303b, 3303c, 3303d, 3303e, 3303f, 3303g, and 3303h each include a digital-to-analog conversion that receives a digital bit from a corresponding neuron output block and converts the digital bit to an analog signal. piece.
[0180] Accordingly, neuron output blocks 3302a, 3302b, 3302c, 3302d, 3302e, 3302f, 3302g, and 3302h receive output currents from macroblocks 3301a, 3301b, 3301c, 3301d, 3301e, and 3301f, and optionally convert the output currents Convert
An analog voltage, a digital bit, or one or more digital pulses in which the width of each pulse or the number of pulses varies in response to the value of the output current. Similarly, input circuit blocks 3303a, 3303b, 3303c, 3303d, 3303e, 3303f, 3303g, and 3303h optionally receive analog current, analog voltage, digital bits, or digital pulses, where the width of each pulse or the number of pulses is responsive to the output current changes, and supplies analog current to the macro blocks 3301a, 3301b, 3301c, 3301d, 3301e, and 3301f. Input circuit blocks 3303a, 3303b, 3303c, 3303d, 3303e, 3303f, 3303g, and 3303h optionally include voltage-to-current converters for determining the number of digital pulses in the input signal or the length of the width of the digital pulses in the input signal. Analog or digital counter, or digital-to-analog converter that performs counting.
[0181] Optionally, when neuron output blocks 3302a, 3302b, 3302c, 3302d, 3302e, 3302f, 3302g, and 3302h convert output currents to analog voltages, digital bits, or one or more digital pulses, they may apply Programming gain. This can be called a programmable neuron.
[0182] Figure 59 shows an example of a programmable neuron output block 5900 that receives an output neuron current Ineu from a VMM array, gain configuration 5901, and generates an output 5902 that represents a The output neuron current Ineu, where the value of G is set in response to the gain configuration 5901. Gain configuration 5901 can be an analog signal or a digital bit. In one embodiment, programmable neuron output block 5900 includes gain control circuit 5903, which in turn includes variable resistor 5904 or variable capacitor 5905 controlled by gain configuration 5901 to generate gain G. Output 5902 may be an analog voltage, an analog current, a digital bit, or one or more digital pulses. In some embodiments, the gain configuration 5901 is used to trim the gain G of the programmable neuron output block 5900 to compensate for undesirable phenomena such as current leakage.
[0183] Optionally, each programmable neuron output block 5900 can be provided with a different gain configuration 5901. This would allow, for example, different gains at different layers of a neural network (such as for scaling array outputs).
[0184] In another embodiment, the gain configuration 5901 depends in part on the input size, meaning, for example, how many rows are enabled to generate the output neuron current Ineu.
[0185] In another embodiment, the gain configuration 5901 depends in part on the values of all rows input to the VMM array. For example, for an 8-bit row input of a VMM array, the maximum value as input for one row is 256 (2~8), for 4 rows the maximum value is 1024, etc. For example, if 256 rows are enabled, a determination is made regarding the total value of those rows and the gain configuration 5901 is modified in response to that value.
[0186] In another embodiment, the gain configuration 5901 depends on the output neuron range. For example, if the output neuron current Ineu is within a first range, a first gain G1 is applied through the gain configuration 5901; if the output neuron current Ineu is within a second range, a second gain G2 is applied through the gain configuration 5901. Although this has been described with respect to ranges, those skilled in the art will recognize that a greater number of ranges may be implemented without limitation.
[0187] Long and short term memory
[0188] Existing technology includes a concept known as long short-term memory (LSTM). LSTM units are commonly used in neural networks. LSTM allows a neural network to remember information over a predetermined arbitrary time interval and use that information in subsequent operations. A conventional LSTM unit consists of a unit, an input gate, an output gate and a forget gate. These three gates regulate the flow of information into and out of the unit and the time interval during which information is remembered in the LSTM. VMM is particularly useful in LSTM units.
[0189] Figure 16 shows an example LSTM 1600. LSTM 1600 in this example includes units 1601, 1602, 1603, and 1604. Unit 1601 receives the input vector x<sub>0</sub>and generate the output vector h<sub>0</sub>and unit state vector c<sub>0</sub>. Unit 1602 receives the input vector x1, the output vector (hidden state) h from unit 1601<sub>0</sub>and unit status c from unit 1601<sub>0</sub>, and generate the output vector h1 and the unit state vector %. Unit 1603 receives the input vector x<sub>2</sub>, the output vector (hidden state) from unit 1602 is
and the cell state c1 from cell 1602, and generates the output vector h<sub>2</sub>and unit state vector c<sub>2</sub>. Unit 1604 receives the input vector x<sub>3</sub>, the output vector (hidden state) h from unit 1603<sub>2</sub>and unit status c from unit 1603<sub>2</sub>, and generate the output vector h<sub>3</sub>. Additional units can be used, and the LSTM with four units is just an example.
[0190] FIG. 17 shows an exemplary implementation of an LSTM unit 1700 that may be used with units 1601, 1602, 1603, and 1604 in FIG. 16. LSTM unit 1700 receives the input vector kd), the unit state vector c(t-1) from the previous unit, and the output vector h(t-1) from the previous unit, and generates the unit state vector c(t) and the output vector h(t).
[0191] The LSTM unit 1700 includes sigmoid function devices 1701, 1702, and 1703, each of which applies a number between 0 and 1 to control how much of each component in the input vector is allowed to pass to the output vector. The LSTM unit 1700 also includes tanh devices 1704 and 1705 for applying the hyperbolic tangent function to the input vector, multiplier devices 1706, 1707 and 1708 for multiplying the two vectors together and adding the two vectors to Together with the adding device 1709. The output vector h(t) can be provided to the next LSTM unit in the system, or it can be accessed for other purposes. [0192] FIG. 18 shows an LSTM unit 1800, which is an example of a specific implementation of the LSTM unit 1700. For the convenience of the reader, the same numbering is used in LSTM unit 1800 as in LSTM unit 1700. Sigmoid function devices 1701, 1702 and 1703 and tanh device 1704 each include a plurality of VMM arrays 1801 and activation circuit blocks 1802. Therefore, it can be seen that VMM arrays are particularly useful in LSTM cells used in some neural network systems. The multiplier devices 1706, 1707 and 1708 and the adding device 1709 are implemented digitally or analogously. Activation function block 1802 may be implemented digitally or analogously.
[0193] An alternative form of LSTM unit 1800 (and another example of a specific implementation of LSTM unit 1700) is shown in Figure 19. In Figure 19, the sigmoid function devices 1701, 1702 and 1703 and the tanh device 1704 share the same physical hardware (VMM array 1901 and activation function block 1902) in a time division multiplexed manner. The OLSTM unit 1900 also includes the multiplication of the two vectors together. Multiplier device 1903, adding device 1908 to add two vectors together, tanh device 1705 (which includes activation circuit block 1902), register 1907 to store value i(t) when it is output from sigmoid function block 1902 , the register 1904 that stores the value f (t)*c(t-1) when it is output from the multiplier device 1903 through the multiplexer 1910, when the value i(t)*u (t) passes through the multiplexer 1910 A register 1905 that stores the value when the value o(t) *c~(t) is output from the multiplier device 1903 through the multiplexer 1910, a register 1906 that stores the value when the value o(t) *c~(t) is output from the multiplier device 1903 through the multiplexer 1910, and multiplexer 1909.
The LSTM unit 1800 includes multiple sets of VMM arrays 1801 and corresponding activation function blocks 1802. The LSTM unit 1900 only includes a set of VMM arrays 1901 and activation function blocks 1902, which are used to represent multiple implementations of the LSTM unit 1900. layer. LSTM unit 1900 will require less space than LSTM 1800 because compared to LSTM unit 1800, LSTM unit 1900 only requires 1/4 of its space for VMM and activation function blocks.
[0195] It will also be appreciated that an LSTM cell will typically include multiple VMM arrays, each of which will require certain circuit blocks outside of the VMM array (such as summer and activation circuit blocks and high voltage generation blocks) Features provided. Providing separate circuit blocks for each VMM array would require a large amount of space within the semiconductor device and would be somewhat inefficient.
[0196] Gated Recursive Unit
[0197] Analog VMM implementations may be used in gated recursive unit (GRU) systems. GRU is a gate control mechanism in recurrent neural networks. GRU is similar to LSTM, except that GRU units generally contain fewer components than LSTM units.
[0198] Figure 20 illustrates an example GRU 2000. GRU 2000 in this example includes units 2001, 2002, 2003, and 2004. Unit 2001 receives the input vector x<sub>0</sub>and generate the output vector h<sub>0</sub>. Unit 2002 receives the input vector K from unit
The output vector h of 2001<sub>0</sub>and generates the output vector ear. Unit 2003 receives the input vector x<sub>2</sub>and the output vector (hidden state) h from unit 2002 and generate the output vector h<sub>2</sub>. Unit 2004 receives the input vector x<sub>3</sub>and the output vector (hidden state) h from unit 2003<sub>2</sub>and generate the output vector h<sub>3</sub>. Additional units may be used, and the GRU with four units is only an example.
[0199] FIG. 21 shows an exemplary implementation of a GRU unit 2100 that may be used in units 2001, 2002, 2003, and 2004 of FIG. 20. The GRU unit 2100 receives the input vector x(t) and the output vector h(t-1) from the previous GRU unit, and generates the output vector h(t). The GRU unit 2100 includes sigmoid function devices 2101 and 2102, each sigmoid function device Applies a number between 0 and 1 to the components from the output vector h(t-1) and the input vector x(t). The GRU unit 2100 also includes a tanh device 2103 for applying the hyperbolic tangent function to the input vector, a plurality of multiplier devices 2104, 2105 and 2106 for multiplying the two vectors together, and a tanh device 2106 for multiplying the two vectors together. An adding device 2107 that adds together, and a complementary device 2108 that subtracts the input from 1 to generate the output.
[0200] FIG. 22 shows a GRU unit 2200, which is an example of a specific implementation of the GRU unit 2100. For the convenience of the reader, the same numbering is used in GRU unit 2200 as in GRU unit 2100. As shown in Figure 22, sigmoid function devices 2101 and 2102 and tanh device 2103 each include a plurality of VMM arrays 2201 and activation function blocks 2202. Therefore, it can be seen that VMM arrays are particularly useful in GRU units used in some neural network systems. The multiplier devices 2104, 2105 and 2106, the adding device 2107 and the complementary device 2108 are implemented in digital or analog form. Activation function block 2202 may be implemented digitally or analogously.
[0201] An alternative form of GRU unit 2200 (and another example of a specific implementation of GRU unit 2300) is shown in Figure 23. In Figure 23, GRU unit 2300 utilizes a VMM array 2301 and an activation function block 2302, which when configured as a sigmoid function applies a number between 0 and 1 to control how much of each component in the input vector is allowed to pass Arrive at the output vector. In Figure 23, sigmoid function devices 2101 and 2102 and tanh device 2103 share the same physical hardware (VMM array 2301 and activation function block 2302) in a time division multiplexing manner. The GRU unit 2300 also includes a multiplier that multiplies the two vectors together. Device 2303, adding device 2305 that adds two vectors together, complementary device 2309 that subtracts the input from 1 to generate the output, multiplexer 2304, when the value h(t- 1)*r(t) is passed through multiple The register 2306 holds the value when the value h (t-1) *z (t) is output from the multiplier device 2303 through the multiplexer 2304. Register 2307, and register 2308 that holds the value hyt)*(1-z(t)) as it is output from the multiplier device 2303 through the multiplexer 2304.
[0202] The GRU unit 2200 includes multiple sets of VMM arrays 2201 and activation function blocks 2202, and the GRU unit 2300 includes only one set of VMM arrays 2301 and activation function blocks 2302, which are used to represent multiple layers in the implementation of the GRU unit 2300. GRU unit 2300 will require less space than GRU unit 2200 because GRU unit 2300 only requires 1/3 of the space for VMM and activation function blocks compared to GRU unit 2200.
[0203] It will also be appreciated that a GRU system will typically include multiple VMM arrays, each of which will require certain circuit blocks external to the VMM array, such as summer and activation circuit blocks and high voltage generation areas. block) provides functionality. Providing separate circuit blocks for each VMM array would require a large amount of space within the semiconductor device and would be somewhat inefficient.
The inputs to the VMM array may be analog levels, binary levels, or digital bits (in which case a DAC is required to convert the digital bits to the appropriate input analog levels), and the output may be analog levels, Binary levels or digital bits (in this case, an output ADC is required to convert the output analog levels into digital bits).
[0205] For each memory cell in the VMM array, each weight w can be represented by a single memory cell or by a differential cell.
Or implemented by two hybrid memory cells (average of 2 cells). In the case of differential units, two memory units are required to implement the weight w as a differential weight (w=w+-w-). In two hybrid memory cells, two memory cells are required to implement the weight w as the average of the two cells.
[0206] Output circuit
[0207] Figure 34A shows application to an output neuron to convert the output neuron current I<sub>NEU</sub>The integrating dual hybrid slope analog-to-digital converter (ADC) 3400 converts the 3406 to digital pulses or digital output bits.
[0208] In one embodiment, ADC 3400 converts the analog output current in a neuron output block (such as neuron output blocks 3302a, 3302b, 3302c, 3302d, 3302e, 3302f, 3302g, and 3302h in Figure 32) to A digital pulse whose width varies in proportion to the magnitude of the analog output current in the neuron output block. The integrator including the integrating operational amplifier 3401 and the integrating capacitor 3402 versus the reference current IREF 3407 versus the memory array current I<sub>neu</sub>3406 (which is the output neuron current) is integrated.
[0209] Optionally, iref 3407 may include a temperature coefficient of zero or a temperature coefficient tracking neuron current of I<sub>NEU</sub>3406 bandgap filter. The latter temperature coefficient can optionally be obtained from a reference array containing values determined during the test phase.
[0210] Optionally, a calibration step can be performed with the circuit at or above operating temperature to offset any current leakage present in the array or control circuit, and this bias can then be subtracted from Ineu in Figure 34B or Figure 35B Shift value.
[0211] During the initialization phase, switch 3408 is closed. Then, the input to Vout 3403 and the negative terminal of op amp 3401 will become VREF. Thereafter, as shown in Figure 34B, switch 33408 is opened, and during a fixed period of time tref, the neuron current I<sub>NEU</sub>3406 upward points. During a fixed period of time tref, Vout rises, and its slope changes as the neuron current changes. Thereafter, the constant reference current IREF is integrated downward during the time period tmeas (during which Vout falls), where tmeas is the time required to integrate Vout downward to VREF.
[0212] When Vout>VREFV, the output EC 3405 will be high, otherwise it will be low. Therefore, the EC3405 generates a pulse whose width reflects the time period tmeas, which in turn is related to the current I<sub>NEU</sub>3406 is directly proportional. In Figure 34B, EC 3405 is shown as waveform 3410 in the example of tmeas = Ineu1 and as waveform 3412 in the example of tmeas = Ineu2. Therefore, the output neuron current I<sub>NEU</sub>3406 is converted into a digital pulse EC 3405, where the width of the digital pulse EC 3405 is related to the output neuron current I<sub>NEU</sub>The magnitude of 3406 changes proportionally.
<sup>[0213]</sup>Current I<sub>neu</sub>3406 =tmeas/tref*IREF. For example, for a required 10-bit output bit resolution, tref is equivalent to a period of 1024 clock cycles. According to I<sub>NEU</sub>The value of 3406 and the value of Iref, the time period tmeas varies from equal to 0 to 1024 clock cycles. Figure 34B shows I<sub>NEU</sub>Example of two different values for 3406, one of which is I<sub>NEU</sub>3406 = Ineu1, while the other I<sub>NEU</sub>3406 = Ineu2. Therefore, the neuronal current I<sub>NEU</sub>3406 affects the rate and slope of charging.
[0214] Optionally, the output pulses EC 3405 may be converted into a series of pulses with a uniform period for transmission to the next stage of the circuit, such as the input block of another VMM array. At the beginning of time period tmeas, output EC 3405 is input into AND gate 3440 together with reference clock 3441. During the time period when Vout>VREF, the output will be a pulse train 3442 (where the frequency of the pulses in pulse train 3442 is the same as the frequency of clock 3441). The number of pulses is proportional to the time period tmeas, which is proportional to the current I<sub>NEU</sub>3406 is directly proportional.
[0215] Optionally, pulse sequence 3443 may be input to a counter 3420 that will count the number of pulses in pulse sequence 3442 and will generate a count value 3421 that is a digit of the number of pulses in pulse sequence 3442 count, the number count is related to the neuron current I<sub>NEU</sub>3406 is directly proportional. The count value 3421 includes a set of digital bits. In another embodiment, the integrating dual slope ADC 3400 can convert the neuron current I<sub>NEU</sub>3407 is converted to a pulse, where the width of the pulse is related to the neuronal current I<sub>NEU</sub>The magnitude of 3407 is inversely proportional. This inversion can be done digitally or analogously and is converted into a series of
Pulses or digital bits for output to a follower circuit.
[0216] Figure 35A shows the application to output neuron I<sub>NEU</sub>The 3504 is an integrating dual hybrid slope ADC 3500 that converts cell current into digital pulses of varying widths or a series of digital output bits. For example, ADC 3500 may be used to convert analog output currents in neuron output blocks (such as neuron output blocks 3302a, 3302b, 3302c, 3302d, 3302e, 3302f, 3302g, and 3302h in Figure 32) to a set of digital output bits. An integrator including an integrating operational amplifier 3501 and an integrating capacitor 3502 changes the neuron current I with respect to the reference current IREF 3503<sub>NEU</sub>3504 for points. Switch 3505 can be closed to reset Vout.
[0217] During the initialization phase, switch 3505 is closed and Vout is charged to voltage V<sub>B1As</sub>. Thereafter, as shown in FIG. 35B, the switch 3505 is turned on, and during the fixed time tref, the unit current I<sub>NEU</sub> 3504 upward points. Thereafter, the reference current IREF 3503 is integrated downward for a period of tmeas until Vout drops to zero. Current I<sub>NEU</sub>3504 = tmeas Ineu/treU IREF. For example, for a required output bit resolution of 10 bits, tref is equivalent to a period of 1024 clock cycles. According to I<sub>NEU</sub>For the values of 3504 and Iref, the time period tmeas varies from equal to 0 to 1024 clock cycles. Figure 35B shows an example of two different Ineu values, one with current Ineul and the other with current Ineu2. Therefore, the neuronal current I<sub>NEU</sub>3504 affects the rate and slope of charge and discharge.
[0218] When Vout>VREF, output 3506 will be high, otherwise it will be low. Therefore, the output 3506 generates a pulse whose width reflects the time period tmeas, which in turn is related to the current I<sub>NEU</sub>3404 is directly proportional. In Figure 35B, output 3506 is shown as waveform 3512 in the example of tmeas = Ineu1 and as waveform 3515 in the example of tmeas = Ineu2. Therefore, the output neuron current I<sub>NEU</sub>3504 is converted into a pulse that is output 3506, where the width of the pulse is related to the output neuron current I<sub>NEU</sub>The magnitude of 3504 changes proportionally.
[0219] Optionally, the output 3506 can be converted into a series of pulses with a uniform period for transmission to the next stage of the circuit, such as the input block of another VMM array. At the beginning of time period tmeas, output 3506 is input into AND gate 3508 along with reference clock 3507. During the period of time when Vout > VREF, the output will be a pulse train 3509 (where the frequency of the pulses in the pulse train 3509 is the same as the frequency of the reference clock 3507). The number of pulses is proportional to the time period tmeas, which is related to the current I<sub>NEU</sub>3504 is directly proportional.
[0220] Optionally, pulse sequence 3509 may be input to a counter 3510 that will count the number of pulses in pulse sequence 3509 and will generate a count value 3511 that is a digit of the number of pulses in pulse sequence 3509 Count, the digital count as shown in waveforms 3514 and 3517 is related to the neuron current I<sub>NEU</sub>3504 is directly proportional. Count value 3511 consists of a set of digital digits.
[0221] In another embodiment, the integrating dual slope ADC 3500 can convert the neuronal current I<sub>NEU</sub>3504 is converted to a pulse, where the width of the pulse is related to the neuronal current I<sub>NEU</sub>The magnitude of 3504 is inversely proportional. This inversion can be done digitally or analogously and is converted into one or more pulses or digital bits for output to a follower circuit.
Figure 35B shows respectively I<sub>NEU</sub>The two neuron current values of 3504, Ineu1 and Ineu2, have a count value of 3511 (digital bits).
[0223] Figures 36A and 36B illustrate waveforms associated with example methods 3600 and 3650 performed in a VMM during operation. In each of methods 3600 and 3650, word lines WLO, WL1, and WL2 receive a variety of different inputs that optionally can be converted into analog voltage waveforms for application to the word lines. In these examples, voltage VC represents the voltage across the integrating capacitor 3402 or 3502 in FIGS. 34A and 35A respectively in the ADC 3400 or 3500 in the output block of the first VMM; OT pulse (= "1") represents the time period during which the neuron's output (which is proportional to the neuron's value) is captured using an integrating dual-slope ADC 3400 or 3500. As shown with reference to Figure 34 and Figure 35, the output of the output block may be a width equal to the first
The output neuronal current of the VMM varies in proportion to varying pulses, or it can be a series of pulses with uniform width, where the number of pulses varies in proportion to the neuronal current of the first VMM. Those pulses can then be applied as input to the second VMM.
[0224] During method 3600, the series of pulses (such as pulse sequence 3442 or pulse sequence 3509) or an analog voltage derived from the series of pulses is applied into the word lines of the second VMM array. Alternatively, the series of pulses, or an analog voltage derived from the series of pulses, may be applied to the control gates of cells within the second VMM array. The number of pulses (or clock cycles) corresponds directly to the magnitude of the input. In this particular example, the input on WL1 is 4 times larger than on WL0 (4 pulses vs. 1 pulse).
[0225] During method 3650, a single pulse of varying width (such as EC 3405 or output 3506) or an analog voltage derived from a single pulse is applied to the word lines of the second VMM array, but with a variable pulse width. . Alternatively, the pulse or an analog voltage derived from the pulse may be applied to the control gate. The width of a single pulse corresponds directly to the magnitude of the input. For example, the magnitude of the input on WL 1 is 4 times that of WL0 (the pulse width of WL1 is 4 times the pulse width of WL0).
[0226] In addition, referring to FIG. 36C, the timing control circuit 3670 may be used to manage the power of the VMM system by managing the output interface and the input interface of the VMM array and sequentially splitting the conversion of various outputs or various inputs. Figure 56 illustrates a power management method 5600. The first step: receive multiple inputs for the vector-matrix multiplication array (step 5601); the second step: organize the multiple inputs into multiple sets of inputs (step 5602); the third step: combine the multiple sets of inputs into Each group of is provided to the array in turn (step 5603).
[0227] An implementation of the power management method 5600 is as follows. Inputs may be applied sequentially over time to a VMM system, such as a word line or control gate of a VMM array. For example, for a VMM array with 512 word line inputs, the word line inputs can be divided into 4 groups: Sichuan 10-12751128-25551256-383 and Sichuan 1383-511. Each group can be enabled at different times, and an output read operation (converting the neuron current into digital bits) can be performed on one of the four groups of word lines, such as through the output integration type in Figures 34 through 36 circuit. Then, after reading each of the four groups in sequence, the output digital bit results are combined together. This operation can be controlled by timing control circuit 3670.
[0228] In another embodiment, timing control circuitry 3670 performs power management in a vector-matrix multiplication system such as analog neural memory system 3300 in Figure 33. Timing control circuitry 3670 may cause inputs to be applied to the VMM subsystem 3333 sequentially over time, such as by enabling input circuit blocks 33032, 3303, 3303, 3303, 3303, 33038, and 3303 at different times. Similarly, timing control circuit 3670 may cause output from VMM subsystem 333 to be read sequentially over time, such as by enabling neuron output blocks 33022, 3302, 3302, 33028, and 3302 at different times.
[0229] Figure 57 illustrates a power management method 5700. Step 1: Receive multiple outputs from the vector-matrix multiplication array (step 5701). Step 2: Organize multiple outputs from the array into groups of outputs (step 5702). Step 3: Provide each of the multiple sets of outputs to the converter circuit in turn (step 5703).
[0230] An implementation of the power management method 5700 is as follows. Power management may be achieved by timing control circuitry 3670 by sequentially reading groups of neuron outputs at different times, i.e., by multiplexing an output circuit (such as an output ADC circuit) across multiple neuron outputs (bit lines). The bit lines can be placed into different groups, and the output circuits operate on one group at a time in sequence under the control of timing control circuit 3670.
[0231] Figure 58 illustrates a power management method 5800. Step 1: Receive multiple inputs in a vector-matrix multiplication system involving multiple arrays. Step 2: Sequentially enable one or more arrays of the plurality of arrays to receive some or all of the plurality of inputs (step 5802).
[0232] An implementation of the power management method 5800 is as follows. Timing control circuitry 3670 may operate on one neural network layer at a time. For example, if one neural network layer is represented in a first VMM array and a second neural network layer is represented in a second VMM array, output read operations (such as neuron outputs being transformed) can be performed sequentially on one VMM array at a time. (in the case of digital bits) to manage the power of the VMM system.
[0233] In another embodiment, the timing control circuit 3670 may operate by sequentially enabling multiple neural memory subsystems 3333 or multiple macros 3301 as shown in FIG. 33 .
[0234] In another embodiment, the timing control circuit 3670 may be enabled sequentially by sequentially enabling multiple neural memory subsystems 3333 or multiple macros 3301 as shown in FIG. period between) releasing the array bias (e.g., the bias on word line WL and/or bit line BL for control gate CG as input and bit line BL as output, or control gate CG and/or bit line The bias on BL is used to operate word line WL as input and bit line BL as output). This is to save power from unnecessary discharging and charging of array biases that are used multiple times during one or more read operations (e.g., during inference or classification operations).
[0235] Figures 37-44 illustrate circuits that can be used in VMM input blocks (such as input circuit blocks 3303a, 3303b, 3303c, 3303d, 3303e, 3303f, 3303g, and 3303h in Figure 33) or neuron output blocks (such as in Figure 33 Various circuits used in neuron output blocks 3302a, 3302b, 3302c, 3302d, 3302e, 3302f, 3302g, and 3302h).
[0236] Figure 37 shows a pulse-to-voltage converter 3700, which optionally can be used to convert the digital pulses generated by the integrating dual-slope ADC 3400 or 3500 into a signal that can be used as an input to a VMM memory array (e.g., in WL or CG line) applied voltage. Pulse-to-voltage converter 3700 includes a reference current generator 3701 that generates a reference current IREF, a capacitor 3702, and a switch 3703. Input is used to control switch 3703. When a pulse is received on the input, the switch closes and charge accumulates on capacitor 3702 such that the voltage on capacitor 3702 after the input signal is complete will indicate the number of pulses received. The capacitor optionally can be a word line or control gate capacitor.
[0237] Figure 38 shows a current-to-voltage converter 3800 that optionally can be used to convert neuron output current into a voltage that can be applied as an input to a VMM memory array (e.g., on the WL or CG lines), for example. . Current-to-voltage converter 3800 includes a current generator 3801, here representing the received neuronal current Ineu (or Iin), and a variable resistor 3802. The output Vout will increase as the neuron current increases. Variable resistor 3802 can be adjusted as needed to increase or decrease the maximum range of Vout.
[0238] Figure 39 shows a current-to-voltage converter 3900 that optionally can be used to convert neuron output current into a voltage that can be applied as an input to a VMM memory array (e.g., on the WL or CG lines), for example. . Current-to-voltage converter 3900 includes an operational amplifier 3901, a capacitor 3902, a switch 3903, a switch 3904, and a current source 3905, here represented as the neuronal current ICELL. During operation, switch 3903 will be open and switch 3904 will be closed. The amplitude of the output Vout will increase proportionally to the magnitude of the neuron current ICELL 3905.
[0239] Figure 40 illustrates a current-to-voltage converter 4000 that optionally may be used to convert neuron output current into a signal that may be applied as an input to a VMM memory array (e.g., on the WL or CG lines). logarithmic voltage. Current-log-to-voltage converter 4000 includes a memory cell 4001, a switch 4002 that selectively connects the word line terminal of memory cell 4001 to a node that generates Vout, and a current source 4003, here represented as neuron current Iin. During operation, switch 4002 will be closed and the amplitude of output Vout will increase proportionally to the magnitude of neuron current iIN.
[0240] Figure 41 illustrates a current-to-voltage converter 4100 that optionally may be used to convert neuron output current to a signal that may be applied as an input to a VMM memory array (e.g., on the WL or CG lines). logarithmic voltage. The current to logarithmic voltage converter 4100 includes a memory unit 4101, a switch 4102 (which selects the control gate terminal of the memory unit 4101
is electrically connected to the node that generates Vout) and a current source 4103, here represented as the neuron current Iin. During operation, switch 4102 will close and the amplitude of output Vout will increase in proportion to the magnitude of neuron current Iin.
[0241] Figure 42 illustrates a digital data-to-voltage converter 4200 that optionally can be used to convert digital data (i.e., for 0s and 1s) to, for example, input to a VMM memory array (e.g., , the voltage applied on the WL or CG line). Digital data to voltage converter 4200 includes a capacitor 4201, an adjustable current source 4202 (here current from a reference array of memory cells), and a switch 4203. Digital data control switch 4203. For example, switch 4203 may be closed when the digital data is "1" and open when the digital data is "0". The voltage accumulated on capacitor 4201 will be the output OUT and will correspond to the value of the digital data. Optionally, the capacitor may be a word line or control gate capacitor.
[0242] Figure 43 illustrates a digital data-to-voltage converter 4300 that optionally may be used to convert digital data (i.e., data for 0s and 1s) into data that may, for example, be input to a VMM memory array (e.g., in a WL or CG line) applied voltage. Digital data to voltage converter 4300 includes a variable resistor 4301, an adjustable current source 4302 (here current from a reference array of memory cells), and a switch 4303. Digital data controls switch 4303. For example, switch 4303 may be closed when the digital data is "1" and open when the digital data is "0". The output voltage will correspond to the value of the digital data.
[0243] FIG. 44 illustrates a reference array 4400 that may be used to provide reference currents for the adjustable current sources 4202 and 4302 of FIGS. 42 and 43.
[0244] Figures 45-47 illustrate components for verifying that a flash memory cell in a VMM contains the appropriate charge corresponding to the value of w expected to be stored in the flash memory cell after a programming operation.
[0245] Figure 45 shows a digital comparator 4500 that receives as digital inputs a set of reference w values and sensed w digital values from a plurality of programmed flash memory cells. If a mismatch exists, the digital comparator 4500 generates a flag, which will indicate that one or more flash memory cells have not been programmed with the correct value.
[0246] FIG. 46 shows the digital comparator 4500 of FIG. 45 in cooperation with a converter 4600. The sensed value of w is provided by multiple instantiations of converter 4600. Converter 4600 receives the cell current ICELL from the flash memory cell and converts the cell current into digital data, which may be provided to digital comparator 4500 using one or more of the aforementioned converters, such as ADC 3400 or 3500.
[0247] FIG. 47 illustrates an analog comparator 4700 that receives as analog inputs a set of reference w values and sensed w analog values from a plurality of programmed flash memory cells. If a mismatch exists, the analog comparator 4700 generates a flag, which will indicate that one or more flash memory cells have not been programmed with the correct value.
[0248] FIG. 48 shows the analog comparator 4700 of FIG. 47 in cooperation with a converter 4800. The sensed value of w is provided by converter 4800. Converter 4800 receives the digital values of the sensed w values and converts them into an analog signal, which may be performed using a previously described converter such as pulse-to-voltage converter 3700, digital data-to-voltage converter 4200, or digital data - one or more of the voltage converters 4300) are provided to the analog comparator 4700.
[0249] Figure 49 shows an output circuit 4900. It will be appreciated that if the output of the neuron is digitized (such as by using an integrating dual slope ADC 3400 or 3500 as previously described), it may still be necessary to perform activation function operations on the neuron output. Figure 49 illustrates an embodiment in which activation occurs before the neuron output is converted into variable width pulses or pulse trains. The output circuit 4900 includes an activation circuit 4901 and a current-to-pulse converter 4902. The activation circuit receives Ineuron values from various flash memory cells and generates Ineuron_act, which is the sum of the received Ineuron values. Current-to-pulse converter 4902 then converts Ineuron_act into a series of digital pulses and/or digital data representing a count of the series of digital pulses. Other converters previously described (such as the integrating dual slope ADC 3400 or 3500) may be used in place of converter 4902.
[0250] In another embodiment, activation may occur after digital pulse generation. In this embodiment, the digital output bits are mapped to a new set of digital bits using an activation mapping table or function implemented by activation mapping unit 5010. Examples of such mappings are shown graphically in Figures 50 and 51. Activation number mapping can simulate sigmoid, tanh, ReLu or any activation function. Additionally, the activation number map quantifies the output neurons.
[0251] Figure 52 shows an example of a charge summer 5200 that can be used to sum the outputs of a VMM during a verify operation following a programming operation to obtain a single analog value that represents the output and This can then optionally be converted to a numeric bit value. Charge summer 5200 includes a current source 5201 and a sample-and-hold circuit including a switch 5202 and a sample-and-hold (S/H) capacitor 5203 . As shown in the example for a 4-bit digital value, there are 4 S/H circuits to hold the values from the 4 evaluation pulses, where these values are added at the end of the process. The S/H capacitor 5203 is selected to have proportions associated with the 2"n*DINn bit positions of the S/H capacitor; for example, C_DIN3 = x8 Cu, C_DIN2 = x4 Cu, C_DIN1 = x2 Cu, DIN0 = x1 Cu. Current source 5201 is also scaled accordingly.
[0252] Figure 53 illustrates a current summer 5300 that may be used to sum the outputs of a VMM during a verify operation following a programming operation. Current summer 5300 includes current source 5301, switch 5302, switches 5303 and 5304, and switch 5305. As shown in the example for a 4-bit digital value, there is a current source circuit to hold the values from the 4 evaluation pulses, where these values are added at the end of the process. Current sources are scaled based on 2~n*DINn bit positions; for example, I_DIN3 = x8 Icell units, I_DIN2 = x4 Icell units, I_DIN1 = x2 Icell units, I_DIN0 = x1 Icell units.
[0253] Figure 54 shows a digital summer 5400 that receives a plurality of digital values, adds them together and generates an output DOUT that represents the sum of the inputs. Digital summer 5400 may be used during verification operations following programming operations. As shown in the example for a 4-bit digital value, there are digital output bits to hold the values from the 4 evaluation pulses, where these values are added at the end of the process. The digital output is digitally scaled based on 2 ~ n*DINn bit positions such as DOUT3 = x8 DOUT0, _DOUT2 = x4 DOUT1, I_DOUT1=x2 DOUT0, I_DOUT0 = DOUT0.
[0254] Figures 55A and 55B illustrate a digital bit-to-pulse width converter 5500 for use within an input block, row decoder, or output block. The pulse width output from digital bit-to-pulse width converter 5500 is proportional to the values described above with respect to Figure 36B. The digital bit-to-pulse width converter includes a binary counter 5501. The state Q[N:0] of the binary counter 5501 can be loaded by serial or parallel data in the load sequence. Row control logic 5510 outputs voltage pulses with a pulse width proportional to the value of the digital data input provided from a block such as the integrating ADC in Figures 34 and 35.
[0255] Figure 55B shows a waveform of an output pulse width having a width proportional to its digital bit value. First, the data in the received digital bits is inverted and the inverted digital bits are loaded into the counter 5501 either serially or in parallel. The row pulse width is then generated by row control logic 5510 as shown in waveform 5520 by counting in binary until it reaches the maximum counter value.
[0256] Optionally, a pulse train-to-pulse converter may be used to convert an output comprising a pulse train (such as signal 3411 or 3413 in Figure 34B and signal 3513 or 3516 in Figure 35B) into pulses in the pulse train with a width equal to A sequence of pulses that vary proportionally to a number of single pulses (such as signals WL0, WL1, and WLe in Figure 36B) serves as an input to the VMM array to be applied to a word line or control gate within the VMM array. An example of a pulse train-to-pulse converter is a binary counter with control logic.
[0257] An example of 4-digit digital input is shown in Table 9:
[0258] Table 9: Digital input bit to output pulse width
[0259]
<td>DIN<3:0></td><td>count</td><td>Inverted DIN<3:0> loaded into counter</td><td>Output pulse width=#clks</td>
<td>0000</td><td>0</td><td>1111</td><td>0</td>
<td>0001</td><td>1</td><td>1110</td><td>1</td>
<td>0010</td><td>2</td><td>1101</td><td>2</td>
<td>0011</td><td>3</td><td>1100</td><td>3</td>
<td>0100</td><td>4</td><td>1011</td><td>4</td>
<td>0101</td><td>5</td><td>1010</td><td>5</td>
<td>0110</td><td>6</td><td>1001</td><td>6</td>
<td>0111</td><td>7</td><td>1000</td><td>7</td>
<td>1000</td><td>8</td><td>0111</td><td>8</td>
<td>1001</td><td>9</td><td>0110</td><td>9</td>
<td>1010</td><td>10</td><td>0101</td><td>10</td>
<td>1011</td><td>11</td><td>0100</td><td>11</td>
<td>1100</td><td>12</td><td>0011</td><td>12</td>
<td>1101</td><td>13</td><td>0010</td><td>13</td>
<td>1110</td><td>14</td><td>0001</td><td>14</td>
<td>1111</td><td>15</td><td>0000</td><td>15</td>
[0260] Another embodiment uses an upward binary counter and digital comparison logic. That is, the output pulse width is generated by counting up a binary counter until the digital output of the binary counter is the same as the digital input bit.
[0261] Another embodiment uses a downward binary counter. First, a down binary counter is loaded serially or in parallel with the digital data input pattern. The output pulse width is then generated by counting down a binary counter until the binary counter's digital output reaches a minimum value (i.e., a "0" logic state).
[0262] In another embodiment, the resolution of the analog-to-digital converter is configurable through a control signal. Figure 60 shows a programmable ADC 6000. Programmable ADC 6000 receives an analog signal, such as the output neuron current Ineu, and converts the analog signal into an output 6002 that includes a set of digital bits. Programmable ADC 600 receives configuration 6001, which may be an analog control signal or a set of digital control bits. In one example, the resolution of output 6002 is determined by configuration 6001. For example, if configuration 6001 has a first value, the output may be a set of 4 bits, but if configuration 6001 has a second value, the output may be a set of 8 bits.
[0263] Coarse level sensing circuitry (not shown) may be used to sample multiple array output currents, and based on the value of this current, a gain (scaling factor) may be configured.
[0264] Gains are configurable for each specific neural network, and gains can be set during neural network training for optimal performance.
[0265] In another embodiment, the ADC may be a hybrid of the above architectures. For example, the first ADC may be a hybrid of a SAR ADC and a slope ADC; the second ADC may be a hybrid of a SAR ADC and a ramp ADC; and the third ADC may be a hybrid of an algorithmic ADC and a slope ADC; etc.
[0266] Figure 61A illustrates a hybrid output conversion block 6100. Output block 6100 receives differential signals Iw+ and w-. Successive approximation register ADC 6101 receives the differential signals Iw+ and IW- and determines the higher order digital bits that best correspond to the analog values represented by Iw+ and IW- (e.g., the most significant bits B7-B4 in the 8-bit digital representation). Once the SAR ADC 6101 determines those higher order bits, an analog signal representing the signals Iw+ and IW- minus the value represented by the higher order bits is provided to the serial ADC block 6102 (such as a slope ADC or ramp ADC), which The row ADC block then determines the lower order bit corresponding to that difference (e.g., an 8-bit digital representation of
The least significant bits in B3-B0). The higher order bits and lower order bits are then combined in a serial fashion to produce a digital output representing the input signals Iw+ and IW-.
[0267] Figure 61B shows output block 6110. Output block 6110 receives differential signals Iw+ and IW-. Algorithm ADC 6103 determines the higher-order bits corresponding to Iw+ and IW- (e.g., bits B7-B4 in the 8-bit digital representation), and the serial ADC block 6104 then determines the lower-order bits (e.g., bits B7-B4 in the 8-bit digital representation) bits B3-B0).
[0268] Figure 61C shows output block 6120. Output block 6120 receives differential signals Iw+ and IW-. Output block 6120 includes a hybrid ADC that converts differential signals Iw+ and IW- into digital bits by combining different conversion schemes (such as in Figures 61A and 61B) into one circuit.
[0269] Figure 62 illustrates a configurable serial ADC 6200. It includes an integrator 6270 that integrates the output neuron current Ineu into an integrating capacitor 6202 (Cint).
[0270] In one embodiment, VRAMP 6250 is provided to the inverting input of comparator 6204. In this case, IREF 6251 shuts down. The digital output (count value) 6221 is produced by ramping up VRAMP 6250 until comparator 6204 switches polarity, where counter 6220 counts clock pulses from the ramp up until comparator 6204 switches polarity, at which time counter 6220 provides a digital output (count value) 6221.
[0271] In another embodiment, VREF 6255 is provided to the inverting input of comparator 6204. VOUT 6203 ramps down by ramp current 6251 (IREF) until VOUT 6203 reaches VREF 6255, at which time the EC 6205 signal disables the counting of the counter 6220, and the counter 6220 provides a digital output (count value) 6221. The (n-bit) ADC 6200 can be configured with lower accuracy (less than n bits) or higher accuracy (greater than n bits), depending on the target application. Configurability of accuracy is achieved by configuring the capacitance of capacitor 6202, the current 6251 (IREF), the ramp rate of VRAMP 6250, or the clock frequency of clock 6241, etc.
[0272] In another embodiment, the ADC circuitry of an array of VMMs is configured to have an accuracy of less than n bits, while the ADC circuitry of another VMM array is configured to have a high accuracy above n bits.
[0273] In another embodiment, one instance of the serial ADC circuit 6200 of one neuron circuit is configured to be combined with another instance of the serial ADC circuit 6200 of the next neuron circuit, such as by combining the serial Two examples of ADC circuit 6200 integrate capacitor 6202 to produce an ADC circuit with greater than n-bit accuracy.
[0274] Figure 63 illustrates a configurable neuron SAR (successive approximation register) ADC 6300. The circuit is based on a successive approximation converter using charge redistribution of binary capacitors, which converts the voltage input Vin into a digital output 6306 based on the reference voltage VREF. The circuit includes a binary capacitor DAC (CDAC) 6301, op amp/comparator 6302, SAR logic, and register 6303. As shown in the figure, GndV 6304 is a low voltage reference level, such as ground. SAR logic sum register 6303 provides digital output 6306. Other non-binary capacitor structures can be implemented with weighted reference voltages or with output correction.
[0275] Figure 64 shows a pipelined SAR ADC circuit 6400 that can be used in combination with the next SAR ADC to increase the number of bits in a pipelined manner. SAR ADC circuit 6400 includes binary CDAC 6401, operational amplifier/comparator 6402, operational amplifier/comparator 6403, and SAR logic sum register 6404. As shown in the figure, GndV is a low voltage reference level, such as ground level. SAR logic sum register 6404 provides digital output 64060Vin is the input voltage and VREF is the ground voltage. Vresidue is generated by capacitor 6405 and provided as input to the next stage of the SAR ADC conversion sequence.
[0276] Figure 65 shows a hybrid SAR + serial ADC circuit 65000 that can be used to increase the number of bits in a hybrid manner. The SAR ADC circuit 6500 includes a binary CDAC 6501, an operational amplifier/comparator 6502, and a SAR logic sum register 6503. As shown in the picture
As shown, GndV is a low voltage reference level, for example, the ground level during SAR ADC operation. SAR logic and register 6503 provide digital output. Vin is the input voltage and VREF is the ground voltage. VREFRAMP is used as a reference ramp voltage during serial ADC operation to replace the GndV input to op amp/comparator 6502.
[0277] Other implementations of hybrid ADC architectures are SAR ADC plus ΣΔ ADC, flash ADC plus serial ADC, pipeline ADC plus serial ADC, serial ADC plus SAR ADC, and other architectures.
[0278] Figure 66 illustrates an algorithmic ADC output block 6600. Output block 6600 includes sample and hold circuit 6601, 1-bit analog-to-digital converter 6602, digital-to-analog converter 6603, summer 6604, operational amplifier 6605, and switches 6606 and 6607 configured as shown.
[0279] In another embodiment, sample and hold circuits are used for the input to each row in the VMM array. For example, if the input includes a DAC, the DAC may include sample and hold circuitry.
[0280] It should be noted that, as used herein, the terms "on" and "on" both inclusively include "directly on" (without intervening materials, elements or spaces disposed therebetween) and "indirectly on" "On" (with intermediate materials, components or spaces between them). Similarly, the term "adjacent" includes "directly adjacent" (with no intervening materials, elements, or spaces disposed therebetween) and "indirectly adjacent" (with intervening materials, elements, or spaces disposed therebetween), and "mounted to" includes "Directly mounted to" (without intervening materials, components or spaces between them) and "Indirectly mounted to" (with intermediate materials, components or spaces between them), and "Electrically coupled to" includes "directly electrically coupled to" " (without intervening materials or elements electrically connecting the elements together) and "indirectly electrically coupled to" (with intervening materials or elements electrically connecting the elements together). For example, forming an element "over" a substrate may include forming the element directly on the substrate with no intervening materials/elements in between, as well as forming the element directly on the substrate with one or more intervening materials/elements in between. Components are formed indirectly on the substrate.
1 sheet
Sheet 1
180 members in 7 offices
Priority claims3
| Document | Office | Kind | Date |
|---|---|---|---|
| 17367542 | United States of America | – | |
| 202117367542 | United States of America | A | |
| 2021053644 | United States of America | W |
Members180
| Document | Office | Kind | |
|---|---|---|---|
| US2019164617A1 | United States of America | A1 | |
| WO2019108334A1 | World Intellectual Property Organization (WIPO) | A1 | |
| US2019237136A1 | United States of America | A1 | |
| US2019237142A1 | United States of America | A1 | |
| TW201933361A | Taiwan Province of China | A | |
| US2019286976A1 | United States of America | A1 | |
| WO2019177698A1 | World Intellectual Property Organization (WIPO) | A1 | |
| TW201941209A | Taiwan Province of China | A | |
| TWI694448B | Taiwan Province of China | B | |
| KR20200060739A | Republic of Korea | A | |
| US10699779B2 | United States of America | B2 | |
| CN111386572A | China | A | |
| EP3676840A1 | European Patent Office (EPO) | A1 | |
| US10720217B1 | United States of America | B1 | |
| US2020233482A1 | United States of America | A1 | |
| US2020234111A1 | United States of America | A1 | |
| US2020234758A1 | United States of America | A1 | |
| WO2020149886A1 | World Intellectual Property Organization (WIPO) | A1 | |
| WO2020149887A1 | World Intellectual Property Organization (WIPO) | A1 | |
| WO2020149889A1 | World Intellectual Property Organization (WIPO) | A1 | |
| WO2020149890A1 | World Intellectual Property Organization (WIPO) | A1 | |
| US2020242453A1 | United States of America | A1 | |
| US2020243139A1 | United States of America | A1 | |
| WO2020159579A1 | World Intellectual Property Organization (WIPO) | A1 | |
| WO2020159580A1 | World Intellectual Property Organization (WIPO) | A1 | |
| WO2020159581A1 | World Intellectual Property Organization (WIPO) | A1 | |
| TW202030649A | Taiwan Province of China | A | |
| TW202030735A | Taiwan Province of China | A | |
| US10748630B2 | United States of America | B2 | |
| KR20200102506A | Republic of Korea | A | |
| TW202032561A | Taiwan Province of China | A | |
| TW202032562A | Taiwan Province of China | A | |
| TWI705389B | Taiwan Province of China | B | |
| TWI705390B | Taiwan Province of China | B | |
| US10803943B2 | United States of America | B2 | |
| CN111886804A | China | A | |
| TW202042116A | Taiwan Province of China | A | |
| TW202042117A | Taiwan Province of China | A | |
| TW202042118A | Taiwan Province of China | A | |
| TW202046183A | Taiwan Province of China | A | |
| TWI716222B | Taiwan Province of China | B | |
| EP3766178A1 | European Patent Office (EPO) | A1 | |
| TWI717703B | Taiwan Province of China | B | |
| JP2021504868A | Japan | A | |
| TWI719757B | Taiwan Province of China | B | |
| TWI732414B | Taiwan Province of China | B | |
| KR20210090275A | Republic of Korea | A | |
| JP2021517705A | Japan | A | |
| KR20210095712A | Republic of Korea | A | |
| US11087207B2 | United States of America | B2 | |
| TW202131337A | Taiwan Province of China | A | |
| EP3676840A4 | European Patent Office (EPO) | A4 | |
| TWI737078B | Taiwan Province of China | B | |
| TWI737079B | Taiwan Province of China | B | |
| CN113302629A | China | A | |
| KR20210105428A | Republic of Korea | A | |
| CN113316793A | China | A | |
| CN113330461A | China | A | |
| KR20210107100A | Republic of Korea | A | |
| KR20210107101A | Republic of Korea | A | |
| KR20210107796A | Republic of Korea | A | |
| CN113366503A | China | A | |
| CN113366504A | China | A | |
| CN113366505A | China | A | |
| CN113366506A | China | A | |
| KR20210110354A | Republic of Korea | A | |
| TWI740487B | Taiwan Province of China | B | |
| US2021334639A1 | United States of America | A1 | |
| US2021342682A1 | United States of America | A1 | |
| EP3912100A1 | European Patent Office (EPO) | A1 | |
| EP3912101A1 | European Patent Office (EPO) | A1 | |
| EP3912102A1 | European Patent Office (EPO) | A1 | |
| EP3912103A1 | European Patent Office (EPO) | A1 | |
| KR102331445B1 | Republic of Korea | B1 | |
| EP3918532A1 | European Patent Office (EPO) | A1 | |
| EP3918533A1 | European Patent Office (EPO) | A1 | |
| EP3918534A1 | European Patent Office (EPO) | A1 | |
| EP3766178A4 | European Patent Office (EPO) | A4 | |
| US2021407588A1 | United States of America | A1 | |
| KR102350213B1 | Republic of Korea | B1 | |
| KR102350215B1 | Republic of Korea | B1 | |
| JP7008167B1 | Japan | B1 | |
| JP2022514111A | Japan | A | |
| US11270763B2 | United States of America | B2 | |
| US11270771B2 | United States of America | B2 | |
| JP2022517810A | Japan | A | |
| JP2022519041A | Japan | A | |
| JP2022519494A | Japan | A | |
| JP2022522987A | Japan | A | |
| JP2022523291A | Japan | A | |
| JP2022523292A | Japan | A | |
| CN113366503B | China | B | |
| TWI764503B | Taiwan Province of China | B | |
| KR102407363B1 | Republic of Korea | B1 | |
| US11409352B2 | United States of America | B2 | |
| EP3676840B1 | European Patent Office (EPO) | B1 | |
| CN113366504B | China | B | |
| EP3912100B1 | European Patent Office (EPO) | B1 | |
| CN113366505B | China | B | |
| US11500442B2 | United States of America | B2 |
2 legal events, as the office reported them to INPADOC
Over the term
Point at a mark for the eventEvents
| Event | Code | |
|---|---|---|
| Entry into force of request for substantive examinationSE01 | SE01 | |
| PublicationPB01 | PB01 |
Numbers
- Publication
- 117581300
- Application
- 801000172
Titles2
- Chinese
- 深度学习人工神经网络中模拟神经存储器的可编程输出块
- English
- Programmable output blocks for simulating neural memory in deep learning artificial neural networks
Classification
- CPC, 6
- G11C11/54
- G06N3/065
- G06N3/0442
- G06N3/0464
- G06N3/048
- H03M1/12
- IPC, 1
- G11C11 54