3D processor having stacked integrated circuit die
Summary by NHIP
Vertically stacked 3D processor
The 3D processor circuit vertically mounts a second IC die containing a cache over a first die with a processor core. A hybrid bond featuring directly bonded metal contact pads and non-conductive regions electrically connects the core to the cache within their overlapping area. Z-axis connections between these blocks have a center-to-center pitch less than 5 microns.
Claim Score by NHIP
Abstract
Some embodiments of the invention provide a three-dimensional (3D) circuit that is formed by vertically stacking two or more integrated circuit (IC) dies to at least partially overlap. In this arrangement, several circuit blocks defined on each die (1) overlap with other circuit blocks defined on one or more other dies, and (2) electrically connect to these other circuit blocks through connections that cross one or more bonding layers that bond one or more pairs of dies. In some embodiments, the overlapping, connected circuit block pairs include pairs of computation blocks and pairs of computation and memory blocks. The connections that cross bonding layers to electrically connect circuit blocks on different dies are referred to below as z-axis wiring or connections. This is because these connections traverse completely or mostly in the z-axis of the 3D circuit, with the x-y axes of the 3D circuit defining the planar surface of the IC die substrate or interconnect layers. These connections are also referred to as vertical connections to differentiate them from the horizontal planar connections along the interconnect layers of the IC dies.

Term
11.7 yearsleft in the term
Expires 24 June 2038, including 263 days of term adjustment.
- Priority
- Filed
- Granted
- Today
- Expires
26 claims: 3 independent, 23 dependent
- 1Broadest claimClaim Score 65, broad(NHIP)A three dimensional (3D) processor circuit, comprising:a first integrated circuit (IC) die comprising a first processor core;and a second IC die vertically mounted on the first IC die, the second IC die comprising a first cache for the first processor core, wherein the first cache vertically overlaps with at least a portion of the first processor core at an overlapping area, and wherein the first IC die is communicatively coupled with the second IC die through a hybrid bond comprising a plurality of directly bonded metal contact pads and directly bonded non-conductive regions, and wherein the first processor core is electrically connected to the first cache through at least one of the plurality of directly bonded metal contact pads disposed within the overlapping area.
- 15An electronic device, comprising:a three-dimensional (3D) processor circuit, comprising: a first integrated circuit (IC) die comprising a first processor core;a second IC die vertically mounted on the first IC die, the second IC die comprising a first cache for the first processor core, wherein the first IC die is communicatively coupled with the second IC die through a hybrid bond comprising a plurality of directly bonded metal contact pads and directly bonded non-conductive regions, wherein the first cache vertically overlaps with at least a portion of the first processor core at an overlapping area, and wherein the first processor core is electrically connected to the first cache through at least one of the plurality of directly bonded metal contact pads disposed within the overlapping area;and a board on which the 3D processor circuit is mounted.
- 21A three dimensional (3D) processor circuit, comprising:a first integrated circuit (IC) die comprising a processor core;a second IC die vertically stacked on the first IC die, the second IC die comprising a cache for the processor core;and interconnect layers interconnecting the first IC die and the second IC die through a hybrid bond comprising a plurality of directly bonded metal contact pads and directly bonded non-conductive regions, wherein the processor core is electrically connected to the cache through at least one of the plurality of directly bonded metal contact pads disposed within an area defined by overlapping circuit blocks on the first and second IC dies.
Independent claims3
148 paragraphs in 5 sections, as filed
CLAIM OF BENEFIT
0001This application claims benefit to U.S. Provisional Patent Application 62/678,246 filed May 30, 2018, U.S. Provisional Patent Application 62/619,910 filed Jan. 21, 2018, U.S. Provisional Patent Application 62/575,221 filed Oct. 20, 2017, U.S. Provisional Patent Application 62/575,184 filed Oct. 20, 2017, U.S. Provisional Patent Application 62/575,240 filed Oct. 20, 2017, and U.S. Provisional Patent Application 62/575,259 filed Oct. 20, 2017. This application is a continuation-in-part of U.S. Non-Provisional patent application Ser. Nos. 15/859,546, 15/859,548, 15/859,551, and 15/859,612, all of which were filed on Dec. 31, 2017, and all of which claim benefit of U.S. Provisional Patent Application 62/541,064, filed on Aug. 3, 2017. This application is also a continuation-in-part of U.S. Non-Provisional patent application Ser. No. 15/976,809, filed on May 10, 2018, which claim benefit of U.S. Provisional Patent Application 62/619,910 filed Jan. 21, 2018. This application is also a continuation-in-part of U.S. Non-Provisional patent application Ser. No. 15/250,030, filed on Oct. 4, 2016, which claim benefit of U.S. Provisional Patent Application 62/405,833 filed Oct. 7, 2016. U.S. Provisional Patent Applications 62/678,246, 62/619,910, 62/575,221, 62/575,184, 62/575,240, 62/575,259, and 62/405,833 are incorporated herein by reference. U.S. Non-Provisional patent application Ser. Nos. 15/859,546, 15/859,548, 15/859,551, 15/859,612, 15/976,809, and 15/250,030 are incorporated herein by reference.
BACKGROUND
0002Electronic circuits are commonly fabricated on a wafer of semiconductor material, such as silicon. A wafer with such electronic circuits is typically cut into numerous dies, with each die being referred to as an integrated circuit (IC). Each die is housed in an IC case and is commonly referred to as a microchip, “chip,” or IC chip. According to Moore's law (first proposed by Gordon Moore), the number of transistors that can be defined on an IC die will double approximately every two years. With advances in semiconductor fabrication processes, this law has held true for much of the past fifty years. However, in recent years, the end of Moore's law has been prognosticated as we are reaching the maximum number of transistors that can possibly be defined on a semiconductor substrate. Hence, there is a need in the art for other advances that would allow more transistors to be defined in an IC chip.
BRIEF SUMMARY
0003Some embodiments of the invention provide a three-dimensional (3D) circuit that is formed by vertically stacking two or more integrated circuit (IC) dies to at least partially overlap. In this arrangement, several circuit blocks defined on each die (1) overlap with other circuit blocks defined on one or more other dies, and (2) electrically connect to these other circuit blocks through connections that cross one or more bonding layers that bond one or more pairs of dies. The 3D circuit in some embodiments can be any type of circuit such as a processor, like a CPU (central processing unit), a GPU (graphics processing unit), a TPU (tensor processing unit), etc., or other kind of circuits, like an FPGA (field programmable gate array), AI (artificial intelligence) neural network chip, encrypting/decrypting chips, etc.
0004The connections in some embodiments cross the bonding layer(s) in a direction normal to the bonded surface. In some embodiments, the overlapping, connected circuit block pairs include pairs of computation blocks and pairs of computation and memory blocks. The connections that cross bonding layers to electrically connect circuit blocks on different dies are referred to below as z-axis wiring or connections. This is because these connections traverse completely or mostly in the z-axis of the 3D circuit, with the x-y axes of the 3D circuit defining the planar surface of the IC die substrate or interconnect layers. These connections are also referred to as vertical connections to differentiate them from the horizontal planar connections along the interconnect layers of the IC dies.
0005The preceding Summary is intended to serve as a brief introduction to some embodiments of the invention. It is not meant to be an introduction or overview of all inventive subject matter disclosed in this document. The Detailed Description that follows and the Drawings that are referred to in the Detailed Description will further describe the embodiments described in the Summary as well as other embodiments. Accordingly, to understand all the embodiments described by this document, a full review of the Summary, Detailed Description, the Drawings and the Claims is needed.
BRIEF DESCRIPTION OF THE DRAWINGS
0006The novel features of the invention are set forth in the appended claims. However, for purposes of explanation, several embodiments of the invention are set forth in the following figures.
0007<figref idref="DRAWINGS">FIG. <b>1</b></figref> illustrates a 3D circuit of some embodiments of the invention.
0008<figref idref="DRAWINGS">FIG. <b>2</b></figref> illustrates an example of a high-performance 3D processor that has a multi-core processor on one die and an embedded memory on another die.
0009<figref idref="DRAWINGS">FIG. <b>3</b></figref> illustrates how multi-core processors are commonly used today in many devices.
0010<figref idref="DRAWINGS">FIG. <b>4</b></figref> illustrates an example of a 3D processor that is formed by vertically stacking three dies.
0011<figref idref="DRAWINGS">FIG. <b>5</b></figref> illustrates three vertically stacked dies with the backside of the second die thinned through a thinning process after face-to-face bonding the first and second dies but before face-to-back mounting the third die to the second die.
0012<figref idref="DRAWINGS">FIGS. <b>6</b>-<b>9</b></figref> illustrate other 3D processors of some embodiments.
0013<figref idref="DRAWINGS">FIG. <b>10</b></figref> illustrates that some embodiments place on different stacked dies two compute circuits that perform successive computations.
0014<figref idref="DRAWINGS">FIG. <b>11</b></figref> illustrates an example of a high-performance 3D processor that has overlapping processor cores on different dies.
0015<figref idref="DRAWINGS">FIG. <b>12</b></figref> illustrates another example of a high-performance 3D processor that has a processor core on one die overlap with a cache on another die.
0016<figref idref="DRAWINGS">FIG. <b>13</b></figref> illustrates an example of a 3D processor that has different parts of a processor core on two face-to-face mounted dies.
0017<figref idref="DRAWINGS">FIG. <b>14</b></figref> shows a compute circuit on a first die that overlaps a memory circuit on a second die, which is vertically stacked over the first die.
0018<figref idref="DRAWINGS">FIG. <b>15</b></figref> shows two overlapping compute circuits on two vertically stacked dies.
0019<figref idref="DRAWINGS">FIG. <b>16</b></figref> illustrates an array of compute circuits on a first die overlapping an array of memories on a second die that is face-to-face mounted with the first die through direct bonded interconnect (DBI) boding process.
0020<figref idref="DRAWINGS">FIG. <b>17</b></figref> illustrates a traditional way of interlacing a memory array with a compute array.
0021<figref idref="DRAWINGS">FIGS. <b>18</b> and <b>19</b></figref> illustrates two examples that show how high density DBI connections can be used to reduce the size of an arrangement of compute circuit that is formed by several successive stages of circuits, each of which performs a computation that produces a result that is passed to another stage of circuits until a final stage of circuits is reached.
0022<figref idref="DRAWINGS">FIG. <b>20</b></figref> presents a compute circuit that performs a computation (e.g., an addition or multiplication) on sixteen multi-bit input values on two face-to-face mounted dies.
0023<figref idref="DRAWINGS">FIG. <b>21</b></figref> illustrates a device that uses a 3D IC.
0024<figref idref="DRAWINGS">FIG. <b>22</b></figref> provides an example of a 3D chip that is formed by two face-to-face mounted IC dies that are mounted on a ball grid array.
0025<figref idref="DRAWINGS">FIG. <b>23</b></figref> illustrates a manufacturing process that some embodiments use to produce the 3D chip.
0026<figref idref="DRAWINGS">FIGS. <b>24</b>-<b>27</b></figref> show two wafers at different stages of the fabrication process of <figref idref="DRAWINGS">FIG. <b>23</b></figref>.
0027<figref idref="DRAWINGS">FIG. <b>28</b></figref> illustrates an example of a 3D chip with three stacked IC dies.
0028<figref idref="DRAWINGS">FIG. <b>29</b></figref> illustrates an example of a 3D chip with four stacked IC dies.
0029<figref idref="DRAWINGS">FIG. <b>30</b></figref> illustrates a 3D chip that is formed by face-to-face mounting three smaller dies on a larger die.
DETAILED DESCRIPTION
0030In the following detailed description of the invention, numerous details, examples, and embodiments of the invention are set forth and described. However, it will be clear and apparent to one skilled in the art that the invention is not limited to the embodiments set forth and that the invention may be practiced without some of the specific details and examples discussed.
0031Some embodiments of the invention provide a three-dimensional (3D) circuit that is formed by vertically stacking two or more integrated circuit (IC) dies to at least partially overlap. In this arrangement, several circuit blocks defined on each die (1) overlap with other circuit blocks defined on one or more other dies, and (2) electrically connect to these other circuit blocks through connections that cross one or more bonding layers that bond one or more pairs of dies. In some embodiments, the overlapping, connected circuit block pairs include pairs of computation blocks and pairs of computation and memory blocks.
0032In the discussion below, the connections that cross bonding layers to electrically connect circuit blocks on different dies are referred to below as z-axis wiring or connections. This is because these connections traverse completely or mostly in the z-axis of the 3D circuit (e.g., because these connections in some embodiments cross the bonding layer(s) in a direction normal or nearly normal to the bonded surface), with the x-y axes of the 3D circuit defining the planar surface of the IC die substrate or interconnect layers. These connections are also referred to as vertical connections to differentiate them from the horizontal planar connections along the interconnect layers of the IC dies.
0033The discussion above and below refers to different circuit blocks on different dies overlapping with each other. As illustrated in the figures described below, two circuit blocks on two vertically stacked dies overlap when their horizontal cross sections (i.e., their horizontal footprint) vertically overlap (i.e., have an overlap in the vertical direction).
0034<figref idref="DRAWINGS">FIG. <b>1</b></figref> illustrates an example of such a 3D circuit. Specifically, it illustrates a 3D circuit <b>100</b> that is formed by vertically stacking two IC dies <b>105</b> and <b>110</b> such that each of several circuit blocks on one die (1) overlaps at least one other circuit block on the other die, and (2) electrically connects to the overlapping die in part through z-axis connections <b>150</b> that cross a bonding layer that bonds the two IC dies. In this example, the two dies <b>105</b> and <b>110</b> are face-to-face mounted as further described below. Also, although not shown in <figref idref="DRAWINGS">FIG. <b>1</b></figref>, the stacked first and second dies in some embodiments are encapsulated into one integrated circuit package by an encapsulating epoxy and/or a chip case.
0035As shown, the first die <b>105</b> includes a first semiconductor substrate <b>120</b> and a first set of interconnect layers <b>125</b> defined above the first semiconductor substrate <b>120</b>. Similarly, the second IC die <b>110</b> includes a second semiconductor substrate <b>130</b> and a second set of interconnect layers <b>135</b> defined below the second semiconductor substrate <b>130</b>. In some embodiments, numerous electronic components (e.g., active components, like transistors and diodes, or passive components, like resistors and capacitors) are defined on the first semiconductor substrate <b>120</b> and on the second semiconductor substrate <b>130</b>.
0036The electronic components on the first substrate <b>120</b> are connected to each other through interconnect wiring on the first set of interconnect layers <b>125</b> to form numerous microcircuits (e.g., Boolean gates, such as AND gates, OR gates, etc.) and/or larger circuit blocks (e.g., functional blocks, such as memories, decoders, logic units, multipliers, adders, etc.). Similarly, the electronic components on the second substrate <b>130</b> are connected to each other through interconnect wiring on the second set of interconnect layers <b>135</b> to form additional microcircuits and/or larger circuit block.
0037In some embodiments, a portion of the interconnect wiring needed to define a circuit block on one die's substrate (e.g., substrate <b>120</b> of the first die <b>105</b>) is provided by interconnect layer(s) (e.g., the second set interconnect layers <b>135</b>) of the other die (e.g., the second die <b>110</b>). In other words, the electronic components on one die's substrate (e.g., the first substrate <b>120</b> of the first die <b>105</b>) in some embodiments are also connected to other electronic components on the same substrate (e.g., substrate <b>120</b>) through interconnect wiring on the other die's set of interconnect layers (e.g., the second set of interconnect layers <b>135</b> of the second die <b>110</b>) to form a circuit block on the first die.
0038As such, the interconnect layers of one die can be shared by the electronic components and circuits of the other die in some embodiments. The interconnect layers of one die can also be used to carry power, clock and data signals for the electronic components and circuits of the other die, as described in U.S. patent application Ser. No. 15/976,815 filed May 10, 2018, which is incorporated herein by reference. The interconnect layers that are shared between two dies are referred to as the shared interconnect layers in the discussion below.
0039Each interconnect layer of an IC die typically has a preferred wiring direction (also called routing direction). Also, in some embodiments, the preferred wiring directions of successive interconnect layers of an IC die are orthogonal to each other. For example, the preferred wiring directions of an IC die typically alternate between horizontal and vertical preferred wiring directions, although several wiring architectures have been introduced that employ 45 degree and 60 degree offset between the preferred wiring directions of successive interconnect layers. Alternating the wiring directions between successive interconnect layers of an IC die has several advantages, such as providing better signal routing and avoiding capacitive coupling between long parallel segments on adjacent interconnect layers.
0040To form the 3D circuit <b>100</b> of <figref idref="DRAWINGS">FIG. <b>1</b></figref>, the first and second dies are face-to-face stacked so that the first and second set of interconnect layers <b>125</b> and <b>135</b> are facing each other. The top interconnect layers <b>160</b> and <b>165</b> are bonded to each other through a direct bonding process that establishes direct-contact metal-to-metal bonding, oxide bonding, or fusion bonding between these two sets of interconnect layers. An example of such bonding is copper-to-copper (Cu—Cu) metallic bonding between two copper conductors in direct contact. In some embodiments, the direct bonding is provided by a hybrid bonding technique such as DBI® (direct bond interconnect) technology, and other metal bonding techniques (such as those offered by Invensas Bonding Technologies, Inc., an Xperi Corporation company, San Jose, CA). In some embodiments, DBI connects span across silicon oxide and silicon nitride surfaces.
0041The DBI process is further described in U.S. Pat. Nos. 6,962,835 and 7,485,968, both of which are incorporated herein by reference. This process is also described in U.S. patent application Ser. No. 15/725,030, which is also incorporated herein by reference. As described in U.S. patent application Ser. No. 15/725,030, the direct bonded connections between two face-to-face mounted IC dies are native interconnects that allow signals to span two different dies with no standard interfaces and no input/output protocols at the cross-die boundaries. In other words, the direct bonded interconnects allow native signals from one die to pass directly to the other die with no modification of the native signal or negligible modification of the native signal, thereby forgoing standard interfacing and consortium-imposed input/output protocols.
0042Direct bonded interconnects allow circuits to be formed across and/or to be accessed through the cross-die boundary of two face-to-face mounted dies. Examples of such circuits are further described in U.S. patent application Ser. No. 15/725,030. The incorporated U.S. Pat. Nos. 6,962,835, 7,485,968, and U.S. patent application Ser. No. 15/725,030 also describe fabrication techniques for manufacturing two face-to-face mounted dies.
0043A DBI connection between two dies terminates on electrical contacts (referred to as pads in this document) on each die's top interconnect layer. Through interconnect lines and/or vias on each die, the DBI-connection pad on each die electrically connects the DBI connection with circuit nodes on the die that need to provide the signal to the DBI connection or to receive the signal from the DBI connection. For instance, a DBI-connection pad connects to an interconnect segment on the top interconnect layer of a die, which then carries the signal to a circuit block on the die's substrate through a series of vias and interconnect lines. Vias are z-axis structures on each die that carry signals between the interconnect layers of the die and between the IC die substrate and the interconnect layers of the die.
0044As shown in <figref idref="DRAWINGS">FIG. <b>1</b></figref>, the direct bonding techniques of some embodiments allow a large number of direct connections <b>150</b> to be established between the top interconnect layer <b>165</b> of the second die <b>110</b> and top interconnect layer <b>160</b> of the first die <b>105</b>. For these signals to traverse to other interconnect layers of the first die <b>105</b> or to the substrate <b>120</b> of the first die <b>105</b>, the first die in some embodiments uses other IC structures (e.g., vias) to carry these signals from its top interconnect layer to these other layers and/or substrate. In some embodiments, more than 1,000 connections/mm<sup>2</sup>, 10,000 connections/mm<sup>2</sup>, 100,000 connections/mm<sup>2</sup>, 1,000,000 connections/mm<sup>2 </sup>or less, etc. can be established between the top interconnect layers <b>160</b> and <b>165</b> of the first and second dies <b>105</b> and <b>110</b> in order to allow signals to traverse between the first and second IC dies.
0045The direct-bonded connections <b>150</b> between the first and second dies are very short in length. For instance, based on current manufacturing technologies, the direct-bonded connections can range from a fraction of a micron to a single-digit or low double-digit microns (e.g., 2-10 microns). As further described below, the short length of these connections allows the signals traversing through these connections to reach their destinations quickly while experiencing no or minimal capacitive load from nearby planar wiring and nearby direct-bonded vertical connections. The planar wiring connections are referred to as x-y wiring or connections, as such wiring stays mostly within a plane defined by an x-y axis of the 3D circuit. On the other hand, vertical connections between two dies or between two interconnect layers are referred to as z-axis wiring or connections, as such wiring mostly traverses in the z-axis of the 3D circuit. The use of “vertical” in expressing a z-axis connection should not be confused with horizontal or vertical preferred direction planar wiring that traverses an individual interconnect layer.
0046In some embodiments, the pitch (distance) between two neighboring direct-bonded connections <b>150</b> can be extremely small, e.g., the pitch for two neighboring connections is between 0.5 μm to 15 μm. This close proximity allows for the large number and high density of such connections between the top interconnect layers <b>160</b> and <b>165</b> of the first and second dies <b>105</b> and <b>110</b>. Moreover, the close proximity of these connections does not introduce much capacitive load between two neighboring z-axis connections because of their short length and small interconnect pad size. For instance, in some embodiments, the direct bonded connections are less then 1 or 2 μm in length (e.g., 0.1 to 0.5 μm in length), and facilitate short z-axis connections (e.g., 1 to 10 μm in length) between two different locations on the two dies even after accounting for the length of vias on each of the dies. In sum, the direct vertical connections between two dies offer short, fast paths between different locations on these dies.
0047Through the z-axis connections <b>150</b> (e.g., DBI connections), electrical nodes in overlapping portions of the circuit blocks on the first and second dies can be electrically connected. These electrical nodes can be on the IC die substrates (e.g., on the portions of the substrates that contain node of electronic components of the circuit blocks) or on the IC die interconnect layers (e.g., on the interconnect layer wiring that form the circuit block). When these electrical nodes are not on the top interconnect layers that are connected through the z-axis connections, vias are used to carry the signals to or from the z-axis connections to these nodes. On each IC die, vias are z-axis structures that carry signals between the interconnect layers and between the IC die substrate and the interconnect layers.
0048<figref idref="DRAWINGS">FIG. <b>1</b></figref> illustrates numerous z-axis connections <b>150</b> between overlapping regions <b>181</b>-<b>185</b> in the top interconnect layers <b>160</b> and <b>165</b>. Each of these regions corresponds to a circuit block <b>171</b>-<b>175</b> that is defined on one of the IC die substrates <b>120</b> and <b>130</b>. Also, each region on the top interconnect layer of one die connects to one or more overlapping regions in the top interconnect layer of the other die through numerous z-axis connections. Specifically, as shown, z-axis connections connect overlapping regions <b>181</b> and <b>184</b>, regions <b>182</b> and <b>184</b>, and regions <b>183</b> and <b>185</b>. Vias are used to provide signals to these z-axis connections from the IC die substrates and interconnect layers. Also, vias are used to carry signals from the z-axis connections when the electrical nodes that need to receive these signals are on the die substrates or the interconnect layers below the top layer.
0049When the z-axis connections are DBI connections, the density of connections between overlapping connected regions can be in the range of 1,000 connections/mm<sup>2 </sup>to 1,000,000 connections/mm<sup>2</sup>. Also, the pitch between two neighboring direct-bonded connections <b>150</b> can be extremely small, e.g., the pitch for two neighboring connections is between 0.5 μm to 15 μm. In addition, these connections can be very short, e.g., in the range from a fraction of a micron to a low single-digit microns. These short DBI connections would allow very short signal paths (e.g., single digit or low-double digit microns, such as 2-20 microns) between two electrically connected circuit nodes on the two substrates of the IC dies <b>105</b> and <b>110</b> even after accounting for interconnect-layer vias and wires.
0050In the example illustrated in <figref idref="DRAWINGS">FIG. <b>1</b></figref>, each top interconnect-layer region <b>181</b>-<b>185</b> corresponds to a circuit block region <b>171</b>-<b>175</b> on an IC die substrate <b>120</b> or <b>130</b>. One of ordinary skill will realize that a circuit block's corresponding top interconnect-layer region (i.e., the region that is used to establish the z-axis connections for that circuit block) does not have to perfectly overlap the circuit block's region on the IC substrate. Moreover, in some embodiments, all the z-axis connections that are used to connect two overlapping circuit blocks in two different dies do not connect one contiguous region in the top-interconnect layer of one die with another contiguous region in the top-interconnect layer of the other die.
0051Also, in some embodiments, the z-axis connections connect circuits on the two dies that do not overlap (i.e., do not have any of their horizontal cross section vertically overlap). However, it is beneficial to use z-axis connections to electrically connect overlapping circuits (e.g., circuit blocks <b>173</b> and <b>175</b>, circuit blocks <b>171</b> and <b>174</b>, etc.) on the two dies <b>105</b> and <b>110</b> (i.e., circuits with horizontal cross sections that vertically overlap) because such overlaps dramatically increase the number of candidate locations for connecting the two circuits. When two circuits are placed next to each other on one substrate, the number of connections that can be established between them is limited by the number of connections that can be made through their perimeters on one or more interconnect layers. However, by placing the two circuits in two overlapping regions on two vertically stacked dies, the connections between the two circuits are not limited to periphery connections that come through the perimeter of the circuits, but also include z-axis connections (e.g., DBI connections and via connections) that are available through the area of the overlapping region.
0052Stacking IC dies in many cases allows the wiring for delivering the signals to be much shorter, as the stacking provides more candidate locations for shorter connections between overlapping circuit blocks that need to be interconnected to receive these signals. For instance, in <figref idref="DRAWINGS">FIG. <b>1</b></figref>, circuit blocks <b>173</b> and <b>175</b> on dies <b>105</b> and <b>110</b> share a data bus <b>190</b> on the top interconnect layer of the second die <b>110</b>. This data bus carries data signals to both of these circuits.
0053Direct-bonded connections are used to carry signals from this data bus <b>190</b> to the circuit block <b>175</b> on the first die <b>105</b>. These direct-bonded connections are much shorter than connections that would route data-bus signals on the first die about several functional blocks in order to reach the circuit block <b>175</b> from this block's periphery. The data signals that traverse the short direct-bonded connections reach this circuit <b>175</b> on the first die very quickly (e.g., within 1 or 2 clock cycles) as they do not need to be routed from the periphery of the destination block. On a less-congested shared interconnect layer, a data-bus line can be positioned over or near a destination circuit on the first die to ensure that the data-bus signal on this line can be provided to the destination circuit through a short direct-bonded connection.
0054Z-axis connection and the ability to share interconnect layers on multiple dies reduce the congestion and route limitations that may be more constrained on one die than another. Stacking IC dies also reduces the overall number of interconnect layers of the two dies because it allows the two dies to share some of the higher-level interconnect layers in order to distribute signals. Reducing the higher-level interconnect layers is beneficial as the wiring on these layers often consumes more space due to their thicker, wider and coarser arrangements.
0055Even though in <figref idref="DRAWINGS">FIG. <b>1</b></figref> the two dies are face-to-face mounted, one of ordinary skill will realize that in other embodiments two dies are vertically stacked in other arrangements. For instance, in some embodiments, these two dies are face-to-back stacked (i.e., the set of interconnect layers of one die is mounted next to the backside of the semiconductor substrate of the other die), or back-to-back stacked (i.e., the backside of the semiconductor substrate of one die is mounted next to the backside of the semiconductor substrate of the other die).
0056In other embodiments, a third die (e.g., an interposer die) is placed between the first and second dies, which are face-to-face stacked, face-to-back stacked (with the third die between the backside of the substrate of one die and the set of interconnect layers of the other die), or back-to-back stacked (with the third die between the backsides of the substrates of the first and second dies). Also, as further described by reference to <figref idref="DRAWINGS">FIG. <b>4</b></figref>, the vertical stack of dies in some embodiments includes three or more IC dies in a stack. While some embodiments use a direct bonding technique to establish connections between the top interconnect layers of two face-to-face stacked dies, other embodiments use alternative connection schemes (such as through silicon vias, TSVs, through-oxide vias, TOVs, or through-glass vias, TGVs) to establish connections between face-to-back dies and between back-to-back dies.
0057In <figref idref="DRAWINGS">FIG. <b>1</b></figref>, the overlapping circuit blocks <b>171</b>-<b>175</b> on the two dies <b>105</b> and <b>110</b> are different types of blocks in different embodiments. Examples of such blocks in some embodiments include memory blocks that store data, computational blocks that perform computations on the data, and I/O blocks that receive and output data from the 3D circuit <b>100</b>. To provide more specific examples of overlapping circuit blocks, <figref idref="DRAWINGS">FIGS. <b>2</b>, <b>4</b>, and <b>6</b></figref> illustrate several different overlapping memory blocks, computational blocks, and/or I/O blocks architectures of some embodiments. Some of these examples illustrate high performance 3D multi-core processors. <figref idref="DRAWINGS">FIGS. <b>10</b>-<b>11</b></figref> then illustrate several examples of overlapping computation blocks, including different cores of a multi-core processor being placed on different IC dies. <figref idref="DRAWINGS">FIG. <b>13</b></figref> illustrates an example of overlapping functional blocks of a processor core.
0058<figref idref="DRAWINGS">FIG. <b>2</b></figref> illustrates an example of a high-performance 3D processor <b>200</b> that has a multi-core processor <b>250</b> on one die <b>205</b> and an embedded memory <b>255</b> on another die <b>210</b>. As shown in this figure, the horizontal cross section of the multi-core processor has a substantially vertical overlaps with the horizontal cross section of the embedded memory. Also, in this example, the two dies <b>205</b> and <b>210</b> are face-to-face mounted through a direct bonding process, such as the DBI process. In other embodiments, these two dies can be face-to-back or back-to-back mounted.
0059As shown in <figref idref="DRAWINGS">FIG. <b>2</b></figref>, numerous z-axis connections <b>150</b> cross a direct bonding layer that bonds the two IC dies <b>205</b> and <b>210</b> in order to establish numerous signal paths between the multi-core processor <b>250</b> and the embedded memory <b>255</b>. When the DBI process is used to bond the two dies <b>205</b> and <b>210</b>, the z-axis connections can be in the range of 1,000 connections/mm<sup>2 </sup>to 1,000,000 connections/mm<sup>2</sup>. As such, the DBI z-axis connections allow a very large number of signal paths to be defined between the multi-core processor <b>250</b> and the embedded memory <b>255</b>.
0060The DBI z-axis connections <b>150</b> also support very fast signal paths as the DBI connections are typically very short (e.g., are 0.2 μm to 2 μm). The overall length of the signal paths is also typically short because the signal paths are mostly vertical. The signal paths often rely on interconnect lines (on the interconnect layers) and vias (between the interconnect layers) to connect nodes of the processor <b>250</b> and the embedded memory <b>255</b>. However, the signal paths are mostly vertical as they often connect nodes that are in the same proximate z-cross section. Given that the DBI connections are very short, the length of a vertical signal path mostly accounts for the height of the interconnect layers of the dies <b>205</b> and <b>210</b>, which is typically in the single digit to low-double digit microns (e.g., the vertical signal paths are typically in the range of 10-20 μm long).
0061As z-axis connections provide short, fast and plentiful connections between the multi-core processor <b>250</b> and the embedded memory <b>255</b>, they allow the embedded memory <b>255</b> to replace many of the external memories that are commonly used today in devices that employ multi-core processors. In other words, robust z-axis connections between vertically stacked IC dies enable next generation system on chip (SoC) architectures that combine the computational power of the fastest multi-core processors with large embedded memories that take the place of external memories.
0062To better illustrate this, <figref idref="DRAWINGS">FIG. <b>3</b></figref> illustrates how multi-core processors are commonly used today in many devices. As shown, a multi-core processor <b>350</b> in a device <b>305</b> typically communicates with multiple external memories <b>310</b> of the device <b>305</b> through an external I/O interface <b>355</b> (such as a double data rate (DDR) interface). As further show, the multi-core processor has multiple general processing cores <b>352</b> and one or more graphical processing cores <b>354</b> that form a graphical processing unit <b>356</b> of the processor <b>350</b>.
0063Each of the processing cores has its own level 1 (L1) cache <b>362</b> to store data. Also, multiple level 2 (L2) caches <b>364</b> are used to allow different processing cores to store their data for access by themselves and by other cores. One or more level 3 (L3) caches <b>366</b> are also used to store data retrieved from external memories <b>310</b> and to supply data to external memories <b>310</b>. The different cores access the L2 and L3 caches through arbiters <b>368</b>. As shown, I/O interfaces <b>355</b> are used to retrieve data for L3 cache <b>366</b> and the processing cores <b>352</b> and <b>354</b>. L1 caches typically have faster access times than L2 caches, which, in turn, have faster access times often than L3 caches.
0064The I/O interfaces consume a lot of power and also have limited I/O capabilities. Often, I/O interfaces have to serialize and de-serialize the output data and the input data, which consumes power and also restricts the multi-core processors input/output. Also, the architecture illustrated in <figref idref="DRAWINGS">FIG. <b>3</b></figref> requires enough wiring to route the signals between the various components of the multi-core processor and the I/O interfaces.
0065The power consumption, wiring and processor's I/O bottleneck is dramatically improved by replacing the external memories with one or more embedded memories <b>255</b> that are vertically stacked with the multi-core processor <b>250</b> in the same IC package. This arrangement dramatically reduces the length of the wires needed to carry signals between the multi-core processor <b>250</b> and its external memory (which in <figref idref="DRAWINGS">FIG. <b>2</b></figref> is the embedded memory <b>255</b>). Instead of being millimeters in length, this wiring is now in the low microns. This is a 100-1000 times improvement in wirelength.
0066The reduction in wirelength allows the 3D processor <b>200</b> of <figref idref="DRAWINGS">FIG. <b>2</b></figref> to have much lower power consumption than the present day design of <figref idref="DRAWINGS">FIG. <b>3</b></figref>. The 3D processor's stacked design also consumes much less power as it foregoes the low throughput, high power consuming I/O interface between the external memories <b>310</b> and the multi-core processor <b>350</b> with plentiful, short z-axis connections between the embedded memory <b>255</b> and the multi-core processor <b>250</b>. The 3D processor <b>200</b> still needs an I/O interface on one of its dies (e.g., the first die <b>205</b>, the second die <b>210</b> or another stacked die, not shown), but this processor <b>200</b> does not need to rely on it as heavily to input data for consumption as a large amount of data (e.g., more than 200 MB, 500 MB, 1 GB, etc.) can be stored in the embedded memory <b>255</b>.
0067The stacked design of the 3D processor <b>200</b>, <figref idref="DRAWINGS">FIG. <b>2</b></figref>, also reduces the size of the multi-core processor by requiring less I/O interface circuits and by placing the I/O interface circuits <b>257</b> on the second die <b>210</b>. In other embodiments, the I/O interface circuits <b>257</b> are on the first die <b>205</b>, but are fewer and/or smaller circuits. In still other embodiments, the I/O interface circuits are placed on a third die stacked with the first and second dies, as further described below.
0068The stacked design of the 3D processor <b>200</b> also frees up space in the device that uses the multi-core processor as it moves some of the external memories to be in the same IC chip housing as the multi-core processor. Examples of memories that can be embedded memories <b>255</b> stacked with the multi-core processor <b>250</b> include any type of memory, such as SRAM (static random access memory), DRAM (dynamic random access memory), MRAM (magnetoresistive random access memory), TCAM (ternary content addressable random access memory), NAND Flash, NOR Flash, RRAM (resistive random access memory), PCRAM (phase change random access memory), etc.
0069Even though <figref idref="DRAWINGS">FIG. <b>2</b></figref> illustrates one embedded memory on the second die <b>210</b>, multiple embedded memories are defined on the second die <b>210</b> in some embodiments, while multiple embedded memories are defined on two or more dies that are vertically stacked with the first die <b>205</b> that contains the multi-core processor <b>250</b>. In some embodiments that use multiple different embedded memories, the different embedded memories all are of the same type, while in other embodiments, the different embedded memories are different types (e.g., some are SRAMs while others are NAND/NOR Flash memories). In some embodiments, the different embedded memories are defined on the same IC die, while in other embodiments, different embedded memories are defined on different IC dies.
0070<figref idref="DRAWINGS">FIG. <b>2</b></figref> illustrates that in some embodiments the multi-core processor <b>250</b> has the similar components (e.g., multiple general processing cores <b>270</b>, L1, L2, and L3 caches <b>272</b>, <b>274</b>, and <b>276</b>, cache arbiters <b>278</b> and <b>280</b>, graphical processing core <b>282</b>, etc.) like other multi-processor cores. However, in the 3D processor <b>200</b>, the I/O interface circuits <b>257</b> for the multi-core processor <b>250</b> are placed on the second die <b>205</b>, as mentioned above.
0071The I/O circuits <b>257</b> write data to the embedded memory <b>255</b> from external devices and memories, and reads data from the embedded memory <b>255</b> for the external devices and memories. In some embodiments, the I/O circuit <b>255</b> can also retrieve data from external devices and memories for the L3 cache, or receive data from the L3 cache for external devices and memories, without the data first going through the embedded memory <b>255</b>. Some of these embodiments have a direct vertical (z-axis) bus between the L3 cache and the I/O circuit <b>257</b>. In these or other embodiments, the first die <b>205</b> also includes I/O circuits as interfaces between the I/O circuit <b>255</b> and the L3 cache <b>276</b>, or as interfaces between the L3 cache <b>276</b> and the external devices/memories.
0072Instead of, or in conjunction with, placing I/O circuits on a different die than the rest of the multi-core processor, some embodiments place other components of a multi-core processor on different IC dies that are placed in a vertical stack. For instance, <figref idref="DRAWINGS">FIG. <b>4</b></figref> illustrates an example of a 3D processor <b>400</b> that is formed by vertically stacking three dies <b>405</b>, <b>410</b> and <b>415</b>, with the first die <b>405</b> including multiple processor cores <b>422</b> and <b>424</b> of a multi-core processor, the second die <b>410</b> including L1-L3 caches <b>426</b>, <b>428</b> and <b>430</b> for the processing cores, and the third die <b>415</b> including I/O circuits <b>435</b>. In this example, the first and second dies <b>405</b> and <b>410</b> are face-to-face mounted (e.g., through a direct bonding process, such as a DBI process), while the second and third dies <b>410</b> and <b>415</b> are back-to-face mounted.
0073In this example, the processor cores are in two sets of four cores <b>432</b> and <b>434</b>. As shown, each core on the first die <b>405</b> overlaps (1) with that core's L1 cache <b>426</b> on the second die <b>410</b>, (2) with one L2 cache <b>428</b> on the second die <b>410</b> that is shared by the three other cores in the same four-core set <b>432</b> or <b>434</b>, and (3) with the L3 cache <b>430</b> on the second die <b>410</b>. In some embodiments, numerous z-axis connections (e.g., DBI connections) establish numerous signal paths between each core and each L1, L2, or L3 cache that it overlaps. These signal paths are also established by interconnect segments on the interconnect layers, and vias between the interconnect layers, of the first and second dies.
0074In some embodiments, some or all of the cache memories (e.g., the L2 and L3 caches <b>428</b> and <b>430</b>) are multi-ported memories that can be simultaneously accessed by different cores. One or more of the cache memories in some embodiments include cache arbiter circuits that arbitrate (e.g., control and regulate) simultaneous and at time conflicting access to the memories by different processing cores. As shown, the 3D processor <b>400</b> also includes one L2 cache memory <b>436</b> on the first die <b>405</b> between the two four-core sets <b>432</b> and <b>434</b> in order to allow data to be shared between these sets of processor cores. In some embodiments, the L2 cache memory <b>436</b> includes a cache arbiter circuit (not shown). In other embodiments, the 3D processor <b>400</b> does not include the L2 cache memory <b>436</b>. In some of these embodiments, the different processor core sets <b>432</b> and <b>434</b> share data through the L3 cache <b>430</b>.
0075The L3 cache <b>430</b> stores data for all processing cores <b>422</b> and <b>424</b> to access. Some of this data is retrieved from external memories (i.e., memories outside of the 3D processor <b>400</b>) by the I/O circuit <b>435</b> that is defined on the third die <b>415</b>. The third die <b>415</b> in some embodiments is face-to-back mounted with the second die. To establish this mounting, TSVs <b>460</b> are defined through the second die's substrate, and these TSVs electrically connect (either directly or through interconnect segments defined on the back side of the second die) to direct bonded connections that connect the backside of the second die to the front side of the third die (i.e., to the top interconnect layer on the front side of the third die). As shown in <figref idref="DRAWINGS">FIG. <b>5</b></figref>, the backside of the second die is thinned through a thinning process after face-to-face bonding the first and second dies but before face-to-back mounting the third die to the second die. This thinning allows the TSVs through the second die's substrate to be shorter. The shorter length of the TSVs, in turn, allows the TSVs to have smaller cross sections and smaller pitch (i.e., smaller center-to-center distance to neighboring TSVs), which thereby improves their density.
0076Most of the signal paths between the second and third dies <b>410</b> and <b>415</b> are very short (e.g., typically in the range of 10-20 μm long) as they mostly traverse in the vertical direction through the thinned second die's substrate and third die's interconnect layers, which have relatively short heights. In some embodiments, a large number of short, vertical signal paths are defined between the L3 cache <b>430</b> on the second die <b>410</b> and the I/O circuit <b>435</b> on the third die <b>415</b>. These signal paths use (1) direct-bonded connections between the top interconnect layer of the third die <b>415</b> and the backside of the second die <b>410</b>, (2) TSVs <b>460</b> through the second die's substrate, and (3) vias between the interconnect layers, and interconnect segments on the interconnect layers, of the second and third die. The number and short length of these signal paths allow the I/O circuit to rapidly write to and read from the L3 cache.
0077The signal paths between the first and second dies <b>405</b> and <b>410</b> use (1) direct-bonded connections between the top interconnect layers of the first and second dies <b>405</b> and <b>410</b>, and (2) vias between the interconnect layers, and interconnect segments on the interconnect layers, of the first and second dies <b>405</b> and <b>410</b>. Most of these signal paths between the first and second dies <b>405</b> and <b>410</b> are also very short (e.g., typically in the range of 10-20 μm long) as they mostly traverse in the vertical direction through the first and second dies' interconnect layers, which have relatively short heights. In some embodiments, a large number of short, vertical signal paths are defined between the processing cores on the first die <b>405</b> and their associated L1-L3 caches.
0078In some embodiments, the processor cores use these fast and plentiful signal paths to perform very fast writes and reads of large data bit sets to and from the L1-L3 cache memories. The processor cores then perform their operations (e.g., their instruction fetch, instruction decode, arithmetic logic, and data write back operations) based on these larger data sets, which in turn allows them to perform more complex instruction sets and/or to perform smaller instruction sets more quickly.
0079<figref idref="DRAWINGS">FIG. <b>6</b></figref> illustrates another 3D processor <b>600</b> of some embodiments. This processor <b>600</b> combines features of the 3D processor <b>200</b> of <figref idref="DRAWINGS">FIG. <b>2</b></figref> with features of the 3D processor <b>400</b> of <figref idref="DRAWINGS">FIG. <b>4</b></figref>. Specifically, like the processor <b>400</b>, the processor <b>600</b> places multiple processor cores <b>422</b> and <b>424</b> on a first die <b>605</b>, L1-L3 caches <b>426</b>, <b>428</b> and <b>430</b> on a second die <b>610</b>, and I/O circuits <b>435</b> on a third die <b>615</b>. However, like the processor <b>200</b>, the processor <b>600</b> also has one die with an embedded memory <b>622</b>. This embedded memory is defined on a fourth die <b>620</b> that is placed between the second and third dies <b>610</b> and <b>615</b>.
0080In <figref idref="DRAWINGS">FIG. <b>6</b></figref>, the first and second dies <b>605</b> and <b>610</b> are face-to-face mounted (e.g., through a direct bonding process, such as a DBI process), the fourth and second dies <b>620</b> and <b>610</b> are face-to-back mounted, and the third and fourth dies <b>615</b> and <b>620</b> are face-to-back mounted. To establish the face-to-back mounting, TSVs <b>460</b> are defined through the substrates of the second die and third dies. The TSVs through the second die <b>610</b> electrically connect (either directly or through interconnect segments defined on the back side of the second die) to direct bonded connections that connect the backside of the second die <b>610</b> to the front side of the fourth die <b>620</b>, while the TSVs through the fourth die <b>620</b> electrically connect (either directly or through interconnect segments defined on the back side of the second die) to direct bonded connections that connect the backside of the fourth die <b>620</b> to the front side of the third die <b>615</b>.
0081To allow these TSVs to be shorter, the backside of the second die is thinned through a thinning process after face-to-face bonding the first and second dies but before face-to-back mounting the fourth die <b>620</b> to the second die <b>610</b>. Similarly, the backside of the fourth die <b>620</b> is thinned through a thinning process after face-to-back mounting the fourth and second dies <b>620</b> and <b>610</b> but before face-to-back mounting the third die <b>615</b> to the fourth die <b>620</b>. Again, the shorter length of the TSVs allows the TSVs to have smaller cross sections and smaller pitch (i.e., smaller center-to-center distance to neighboring TSVs), which thereby improves their density.
0082As in <figref idref="DRAWINGS">FIG. <b>4</b></figref>, the L3 cache <b>430</b> in <figref idref="DRAWINGS">FIG. <b>6</b></figref> stores data for all processing cores <b>422</b> and <b>424</b> to access. However, in <figref idref="DRAWINGS">FIG. <b>6</b></figref>, the L3 cache does not connect to the I/O circuits <b>435</b> but rather connects to the embedded memory <b>622</b> on the fourth die through vertical signal paths. In this design, the embedded memory <b>622</b> connects to the I/O circuits <b>435</b> on the third die <b>615</b> through vertical signal paths. In some embodiments, the vertical signal paths between the second and fourth dies <b>610</b> and <b>620</b> and between the fourth and third dies <b>620</b> and <b>615</b> are established by z-axis direct bonded connections and TSVs, as well as interconnect segments on the interconnect layers and vias between the interconnect layers. Most of these signal paths are very short (e.g., typically in the range of 10-20 μm long) as they are mostly vertical and the height of the thinned substrates and their associated interconnect layers is relatively short.
0083Like the embedded memory <b>255</b> of <figref idref="DRAWINGS">FIG. <b>2</b></figref>, the embedded memory <b>622</b> is a large memory (e.g., is larger than 200 MB, 500 MB, 1 GB, etc.) in some embodiments. As such, the embedded memory in some embodiments can replace one or more external memories that are commonly used today in devices that employ multi-core processors. Examples of the embedded memory <b>622</b> include SRAM, DRAM, MRAM, NAND Flash, NOR Flash, RRAM, PCRAM, etc. In some embodiments, two or more different types of embedded memories are defined on one die or multiple dies in the stack of dies that includes one or more dies on which a multi-core processor is defined.
0084Through numerous short, vertical signal paths, the embedded memory <b>622</b> receives data from, and supplies data to, the I/O circuit <b>435</b>. Through these signal paths, the I/O circuit <b>435</b> writes data to the embedded memory <b>622</b> from external devices and memories, and reads data from the embedded memory <b>622</b> for the external devices and memories. In some embodiments, the I/O circuit <b>435</b> can also retrieve data from external devices and memories for the L3 cache, or receive data from the L3 cache for external devices and memories, without the data first going through the embedded memory <b>622</b>. Some of these embodiments have a direct vertical (z-axis) bus between the L3 cache and the I/O circuit <b>435</b>. In these or other embodiments, the second die <b>610</b> and/or fourth die <b>620</b> also include I/O circuits as interfaces between the I/O circuit <b>435</b> and the L3 cache <b>430</b>, or as interfaces between the L3 cache <b>430</b> and the external devices/memories.
0085<figref idref="DRAWINGS">FIG. <b>7</b></figref> illustrates yet another 3D processor <b>700</b> of some embodiments. This processor <b>700</b> is identical to the processor <b>600</b> of <figref idref="DRAWINGS">FIG. <b>6</b></figref>, except that it only has two layers of caches, L1 and L2, on a second die <b>710</b> that is face-to-face mounted on a first die <b>705</b> that has eight processor cores <b>722</b>. As shown, each L1 cache <b>726</b> overlaps just one core <b>722</b>. Unlike the L1 caches <b>726</b>, the L2 cache <b>728</b> is shared among all the cores <b>722</b> and overlaps each of the cores <b>722</b>. In some embodiments, each core connects to each L1 or L2 cache that it overlaps through (1) numerous z-axis DBI connections that connect the top interconnect layers of the dies <b>705</b> and <b>721</b>, and (2) the interconnects and vias that carry the signals from these DBI connections to other metal and substrate layers of the dies <b>705</b> and <b>710</b>. The DBI connections in some embodiments allow the data buses between the caches and the cores to be much wider and faster than traditional data buses between the caches and the cores.
0086In some embodiments, L1 caches are formed by memories that can be accessed faster (i.e., have faster read or write times) than the memories that are used to form L2 caches. Each L1 cache <b>726</b> in some embodiments is composed of just one bank of memories, while in other embodiments it is composed of several banks of memories. Similarly, the L2 cache <b>728</b> in some embodiments is composed of just one bank of memories, while in other embodiments it is composed of several banks of memories. Also, in some embodiments, the L1 caches <b>726</b> and/or L2 cache <b>728</b> are denser than traditional L1 and L2 caches as they use z-axis DBI connections to provide and receive their signals to and from the overlapping cores <b>722</b>. In some embodiments, the L1 and L2 caches <b>726</b> and <b>728</b> are much larger than traditional L1 and L2 caches as they are defined on another die than the die on which the cores are defined, and hence face less space restrictions on their placement and the amount of space that they consume on the chip.
0087Other embodiments use still other architectures for 3D processors. For example, instead of using just one L2 cache <b>728</b>, some embodiments use two or four L2 caches that overlap four cores (e.g., the four left cores <b>726</b> and the four right cores <b>726</b>) or two cores (e.g., one of the four pairs of vertically aligned cores <b>722</b>). <figref idref="DRAWINGS">FIG. <b>8</b></figref> illustrates another 3D processor <b>800</b> of some embodiments. This processor <b>800</b> is identical to the processor <b>700</b> of <figref idref="DRAWINGS">FIG. <b>7</b></figref>, except that it does not have the L2 cache <b>728</b>. In place of this L2 cache, the processor <b>900</b> has a network on chip (NOC) <b>8028</b> on the die <b>810</b>, which is face-to-face mounted to the die <b>705</b> through a DBI bonding process.
0088In some embodiments, the NOC <b>828</b> is an interface through which the cores <b>722</b> communicate. This interface includes one or more buses and associated bus circuitry. The NOC <b>828</b> in some embodiments also communicatively connects each core to the L1 caches that overlap the other cores. Through this NOC, a first core can access data stored by a second core in the L1 cache that overlap the second core. Also, through this NOC, a first core in some embodiments can store data in the L1 cache that overlaps a second core. In some embodiments, an L1 and L2 cache overlaps each core <b>722</b>, and the NOC <b>828</b> connects the cores to L2 caches of other cores, but not to the L1 caches of these cores. In other embodiments, the NOC <b>828</b> connects the cores to both L1 and L2 caches that overlap other cores, as well as to the other cores.
0089<figref idref="DRAWINGS">FIG. <b>9</b></figref> illustrates yet another 3D processor <b>900</b> of some embodiments. This processor <b>900</b> is identical to the processor <b>400</b> of <figref idref="DRAWINGS">FIG. <b>4</b></figref>, except that it only has one L1 cache <b>932</b> on a die <b>910</b> for each of six CPU (central processing unit) cores <b>922</b> and one L1 cache <b>934</b> for each of two GPU (graphical processing unit) cores <b>924</b> that are defined on a die <b>905</b> that is face-to-face mounted to the die <b>910</b> through a DBI bonding process. The processor <b>900</b> does not use layers 2 and 3 caches as it uses large L1 caches for its CPU and GPU cores. The L1 caches can be larger than traditional L1 caches as they are defined on another die than the die on which the cores are defined, and hence face less space restrictions on their placement and the amount of space that they consume on the chip.
0090In <figref idref="DRAWINGS">FIG. <b>9</b></figref>, the processor <b>900</b> has its I/O interface defined on a third die <b>415</b> that is face-to-back mounted on the die <b>910</b>. In other embodiments, the processor <b>900</b> does not include the third die <b>415</b>, but just includes the first and second dies <b>905</b> and <b>910</b>. In some of these embodiments, the I/O interface of the processor <b>900</b> is defined on the first and/or second dies <b>905</b> and <b>910</b>. Also, in other embodiments, one L1 cache <b>932</b> is shared across multiple CPU cores <b>922</b> and/or multiple GPU cores <b>924</b>.
0091<figref idref="DRAWINGS">FIG. <b>10</b></figref> illustrates that some embodiments place on different stacked dies two compute circuits that perform successive computations. A compute circuit is a circuit that receives a multi-bit value as input and computes a multi-bit value as output based on the received input. In <figref idref="DRAWINGS">FIG. <b>10</b></figref>, one compute circuit <b>1015</b> is defined on a first die <b>1005</b> while the other compute circuit <b>1020</b> is defined on a second die <b>1010</b>.
0092The first and second dies are face-to-face mounted through a direct bonding process (e.g., a DBI process). This mounting defines numerous z-axis connections between the two dies <b>1005</b> and <b>1010</b>. Along with interconnect line on the interconnect layers, and vias between the interconnect layers, of the two dies, the z-axis connections define numerous vertical signal paths between the two compute circuits <b>1015</b> and <b>1020</b>. These vertical signal paths are short as they mostly traverse in the vertical direction through the die interconnect layers, which are relatively short. As they are very short, these vertical signal paths are very fast parallel paths that connect the two compute circuit <b>1015</b> and <b>1020</b>.
0093In <figref idref="DRAWINGS">FIG. <b>10</b></figref>, the first compute circuit <b>1015</b> receives a multi-bit input value <b>1030</b> and computes a multi-bit output value <b>1040</b> based on this input value. In some embodiments, the multi-bit input value <b>1030</b> and/or output value <b>1040</b> are large bit values, e.g., 32 bits, 64 bits, 128 bits, 256 bits, 512 bits, 1024 bits, etc. Through the vertical signal paths between these two compute circuits, the first compute circuit <b>1015</b> provides its multi-bit output value <b>1040</b> as the input value to the compute circuit <b>1020</b>. Based on this value, the compute circuit <b>1020</b> computes another multi-bit output value <b>1045</b>.
0094Given the large number of vertical signal paths between the first and second compute circuits <b>1015</b> and <b>1020</b>, large number of bits can be transferred between these two circuits <b>1015</b> and <b>1020</b> without the need to use serializing and de-serializing circuits. The number of the vertical signal paths and the size of the exchanged data also allow many more computations to be performed per each clock cycle. Because of the short length of these vertical signal paths, the two circuits <b>1015</b> and <b>1020</b> can exchange data within one clock cycle. When two computation circuits are placed on one die, it sometimes can take 8 or more clock cycles for signals to be provided from one circuit to another because of the distances and/or the congestion between the two circuits.
0095In some embodiments, the two overlapping computation circuits on the two dies <b>1005</b> and <b>1010</b> are different cores of a multi-core processor. <figref idref="DRAWINGS">FIG. <b>11</b></figref> illustrates an example of a high-performance 3D processor <b>1100</b> that has overlapping processor cores on different dies. In this example, two dies <b>1105</b> and <b>1110</b> are face-to-face mounted through a direct bonding process (e.g., the DBI process). The first die <b>1105</b> includes a first processor core <b>1112</b>, while the second die <b>1110</b> includes a second processor core <b>1114</b>.
0096The first die <b>1105</b> also includes an L1 cache <b>1116</b> for the second core <b>1114</b> on the second die <b>1110</b>, and L2 and L3 caches <b>1122</b> and <b>1126</b> for both cores <b>1112</b> and <b>1114</b>. Similarly, the second die <b>1110</b> also includes an L1 cache <b>1118</b> for the first core <b>1112</b> on the first die <b>1105</b>, and L2 and L3 caches <b>1124</b> and <b>1128</b> for both cores <b>1112</b> and <b>1114</b>. As shown, each core completely overlaps its corresponding L1 cache, and connects to its L1 cache through numerous vertical signal paths that are partially defined by z-axis connections between the top two interconnect layers of the dies <b>1105</b> and <b>1110</b>. As mentioned above, such vertical signal paths are also defined by (1) vias between interconnect layers of each die, and/or (3) interconnect segments on interconnect layers of each die.
0097Each core on one die also overlaps with one L2 cache and one L3 cache on the other die and is positioned near another L2 cache and another L3 cache on its own die. Each L2 and L3 cache <b>1122</b>-<b>826</b> can be accessed by each core <b>1112</b> or <b>1114</b>. Each core accesses an overlapping L2 or L3 cache through numerous vertical signal paths that are partially defined by z-axis connections between the top two interconnect layers of the dies <b>1105</b> and <b>1110</b> and by (1) vias between interconnect layers of each die, and/or (3) interconnect segments on interconnect layers of each die.
0098Each core can also access an L2 or L3 cache on its own die through signal paths that are defined by visa between interconnect layers, and interconnect segments on interconnect layers, of its own die. In some embodiments, when additional signal paths are needed between each core and an L2 or L3 cache on its own die, each core also connects to such L2 or L3 cache through signal paths that are not only defined by vias between interconnect layers, and interconnect segments on interconnect layers, of its own die, but by vias between interconnect layers and interconnect segments on interconnect layers of the other die.
0099Other embodiments, however, do not use signal paths that traverse through the other die's interconnect layers to connect a core with an L2 cache or an L3 cache on its own die, because such signal paths might have different delay (i.e., a greater delay) than signal paths between this core and this cache that only use the interconnect layers of the core's own die. On the other hand, given the very short length of the z-axis connections, other embodiments use signal paths defined through the other die's interconnect layers (e.g., through its top interconnect layers) when the difference in the signal path delay is very small (as compared to the speed of the signal paths that only use the interconnect layers of the core's die).
0100The 3D architecture illustrated in <figref idref="DRAWINGS">FIG. <b>11</b></figref> dramatically increases the number of connections (through vertical signal paths) between each core <b>1112</b> or <b>1114</b> and its corresponding L1, L2 and L3 caches. With this increase, each core <b>1112</b> or <b>1114</b> retrieves much larger sets of data bits and performs more complex operations faster with such larger sets of data bits. In some embodiments, each core uses wider instruction and data buses in its pipelines as it can retrieve wider instructions and data from overlapping memories. In these or other embodiments, each core has more pipelines that perform more operations in parallel as the core can retrieve more instruction and data bits from the overlapping memories.
0101In some embodiments, each core on one die only uses the L2 cache or L3 cache on the other die (i.e., only uses the L2 or L3 cache that vertically overlaps the core) in order to take advantage of the large number of vertical signal paths between it and the overlapping L2 cache. Each core in some of these embodiments stores a redundant copy of each data, which it stores in its own overlapping cache (e.g., its own overlapping L2 cache) in the corresponding cache (e.g., in the other L2 cache) that is defined on the core's own die, so that the data is also available for the other core. In some of these embodiments, each core reaches the cache on its own die through signal paths that are not only defined through the interconnect lines and vias on the core's die, but also defined through interconnect lines and vias of the other die.
0102<figref idref="DRAWINGS">FIG. <b>12</b></figref> illustrates another example of a high-performance 3D processor <b>1200</b> that has a processor core on one die overlap with a cache on another die. In this example, two dies <b>1205</b> and <b>1210</b> are face-to-face mounted through a direct bonding process (e.g., the DBI process). The first die <b>1205</b> includes a first processor core <b>1212</b>, while the second die <b>1210</b> includes a second processor core <b>1214</b>. The first die <b>1205</b> includes an L1 cache <b>1216</b> for the second processor core <b>1214</b> defined on the second die <b>1210</b>, while the second die <b>1210</b> includes an L1 cache <b>1218</b> for the first processor core <b>1212</b> defined on the first die <b>1205</b>.
0103In this example, the cross-section of each L1 cache on one die completely overlaps the cross-section of the corresponding core on the other die. This ensures the largest region for defining z-axis connections (e.g., DBI connections) in the overlapping regions of each core and its corresponding L1 cache. These z-axis connections are very short and hence can be used to define a very fast bus between each core and its corresponding L1 cache. Also, when high density z-axis bonding is used (e.g., when DBI is used), this z-axis bus can be wide and it can be defined wholly within the x-y cross-section of the core and its L1 cache, as further described below. By being wholly contained within this cross-section, the z-axis bus would not consume routing resources around the core and its L1 cache. Also, the speed and width of this bus allows the bus to have a very high throughput bandwidth, which perfectly complements the high speed of the L1 cache.
0104As shown in <figref idref="DRAWINGS">FIG. <b>12</b></figref>, the 3D processor <b>1200</b> defines an L2 cache for each core on the same die on which the core is defined. In some embodiments, each core can access the other core's L2 cache through z-axis connections established through the face-to-face bonding of the two IC dies. Also, due to the size of the L1 cache, the 3D process <b>1200</b> in some embodiments does not use an L3 cache.
0105In some embodiments, different components of a processor core of a multi-processor core are placed on different dies. <figref idref="DRAWINGS">FIG. <b>13</b></figref> illustrates an example of a 3D processor <b>1300</b> that has different parts of a processor core on two face-to-face mounted dies <b>1305</b> and <b>1310</b>. In this example, the first die <b>1305</b> includes multiple pipeline <b>1390</b>, with each pipeline having an instruction fetch (IF) unit <b>1312</b>, an instruction decode unit <b>1314</b>, an execution unit <b>1316</b> and a write-back unit <b>1318</b>. The second die includes the instruction memory <b>1322</b> and data registers and memories <b>1324</b>.
0106As shown, the instruction memory <b>1322</b> on the second die overlaps with the IF units <b>1312</b> on the first die <b>1305</b>. Also, the data registers and memories <b>1324</b> on the second die overlap with the execution units <b>1316</b> and the write-back units. Numerous vertical signal paths are defined between overlapping core components by the z-axis connections between the top two interconnect layers of the dies <b>1305</b> and <b>1310</b>, and by (1) vias between interconnect layers of each die <b>1305</b> or <b>1310</b>, and/or (2) interconnect segments on interconnect layers of each die.
0107Through vertical signal paths, each IF unit <b>1312</b> retrieves instructions from the instruction memory and provides the retrieved instructions to its instruction decode unit <b>1314</b>. This decode unit decodes each instruction that it receives and supplies the decoded instruction to its execution unit to execute. Through the vertical signal paths, each execution unit receives, from the data registers and memories <b>1324</b>, operands that it needs to execute a received instruction, and provides the result of its execution to its write-back unit <b>1318</b>. Through the vertical signal paths, each write-back unit <b>1318</b> stores the execution results in the data registers and memories <b>1324</b>. Other embodiments use other architectures to split a processor core between two different dies. For instance, some embodiments place the instruction decode and execution units <b>1314</b> and <b>1316</b> on different layers than the instruction fetch and write back units <b>1312</b> and <b>1318</b>. Still other embodiments use other arrangements to split a processor core between different dies. These or other embodiments put different ALUs, or different portions of the same ALU, of a processor core on different vertically stacked dies (e.g., on two dies that are face-to-face mounted through a DBI bonding process).
0108As mentioned above, it is advantageous to use DBI connections to connect overlapping connected regions on two dies that are vertically stacked because DBI allows for far greater density of connections than other z-axis connection schemes. <figref idref="DRAWINGS">FIG. <b>14</b></figref> presents an example that illustrates this. This figure shows a compute circuit <b>1415</b> on a first die <b>1405</b> that overlaps a memory circuit <b>1420</b> on a second die <b>1410</b>, which is vertically stacked over the first die <b>1405</b>. The compute circuit can be any type of compute circuit (e.g., processor cores, processor pipeline compute units, neural network neurons, logic gates, adders, multipliers, etc.) and the memory circuit can be any type of memory circuit (e.g., SRAMs, DRAMs, non-volatile memories, caches, etc.).
0109In this example, both circuits <b>1415</b> and <b>1420</b> occupy a square region of 250 by 250 microns on their respective dies <b>1405</b> and <b>1410</b> (only the substrate surfaces of which are shown in <figref idref="DRAWINGS">FIG. <b>14</b></figref>). Also, in this example, a 100-bit z-axis bus <b>1425</b> is defined between these circuits, with the term bus in this example referring to the data and control signals exchanged between these two circuits <b>1415</b> and <b>1420</b> (in other examples, a bus might only include data signals). <figref idref="DRAWINGS">FIG. <b>14</b></figref> illustrates that when TSVs are used to define this z-axis bus <b>1425</b>, this bus will consume on each die a region <b>1435</b> that is at least 2.5 times as large as the size of either circuit on that die. This is because TSVs have a 40 micron pitch. For the TSV connections, the two dies <b>1405</b> and <b>1410</b> will be front-to-back mounted with the TSVs going through the substrate of one of the two dies.
0110On the other hand, when the two dies are face-to-face bonded through mounted, and DBI connections are used to define the 100-bit z-axis bus <b>1425</b>, the cross-section <b>1430</b> of the DBI bus can be contained within the footprint (i.e., the substrate region) of both circuits <b>1415</b> and <b>1420</b> on their respective dies. Specifically, assuming that the DBI connections have a 2 micron pitch, the 100 DBI connections can be fit in as little as 20-by-20 micron square, as the 100 connections can be defined as a 10-by-10 array with each connection having a minimum center-to-center spacing of 2 microns with its neighboring connections. By being contained within the footprint of the circuits <b>1415</b> and <b>1420</b>, the DBI connections would typically not consume any precious routing space on the dies <b>1405</b> and <b>1410</b> beyond the portion already consumed by the circuits. In some embodiments, DBI connections can have a pitch ranging from less than 1 micron (e.g., 0.2 or 0.5 microns) to 5 microns.
0111As the numbers of bits increase in the bus <b>1425</b>, the difference between the amount of space consumed by the TSV connections and the space consumed by DBI connections becomes even more pronounced. For instance, when 3600 bits are exchanged between the two circuits <b>1415</b> and <b>1420</b>, a 60-by-60 TSV array would require a minimum 2400-by-2400 micron region (at a 40-micron DBI pitch), while a 60-by-60 DBI array would require a minimum 120-by-120 region (at a 2-micron DBI pitch). In other words, the TSVs would have a footprint that is at least 400 times greater than the footprint of the DBI connections. It is quite common to have a large number of bits exchanged between a memory circuit and a compute circuit when performing computations (e.g., dot product computations) in certain compute environments (e.g., machine-trained neural networks). Moreover, the density of DBI connections allows for very large bandwidth (e.g., in the high gigabytes or in the terabytes range) between overlapping compute and memory circuits.
0112The density of DBI connection is also advantageous in connecting overlapping circuit regions on two dies that are vertically stacked. <figref idref="DRAWINGS">FIG. <b>15</b></figref> presents an example that illustrates this. This figure shows two overlapping compute circuits <b>1515</b> and <b>1520</b> on two vertically stacked dies <b>1505</b> and <b>1510</b>. Each of the circuits occupies a 250-by-250 micron square on its corresponding die's substrate, and can be any type of compute circuit (e.g., processor cores, processor pipeline compute units, neural network neurons, logic gates, adders, multipliers, etc.).
0113Like the example in <figref idref="DRAWINGS">FIG. <b>14</b></figref>, the example in <figref idref="DRAWINGS">FIG. <b>15</b></figref> shows that when DBI connections are used (i.e., when the two dies <b>1505</b> and <b>1510</b> are face-to-face mounted through DBI), and the DBI connections have a 2-micron pitch, a 100-bit bus <b>1525</b> between the two circuits <b>1515</b> and <b>1520</b> can be contained in a region <b>1530</b> that is 20-by-20 micron square that can be wholly contained within the footprints of the circuits. On the other hand, when TSV connections are used (e.g., when the two dies are face-to-back mounted and connected using TSVs), the 100-bit bus <b>1525</b> would consume at a minimum a 400-by-400 micron square region <b>1535</b>, which is larger than the footprint of the compute circuits <b>1515</b> and <b>1520</b>. This larger footprint would consume additional routing space and would not be as beneficial as the smaller footprint that could be achieved by the DBI connections.
0114High density DBI connections can also be used to reduce the size of a circuit formed by numerous compute circuits and their associated memories. The DBI connections can also provide this smaller circuit with very high bandwidth between the compute circuits and their associated memories. <figref idref="DRAWINGS">FIG. <b>16</b></figref> presents an example that illustrates these benefits. Specifically, it illustrates the reduction in the size of an array <b>1600</b> of compute circuits <b>1615</b> on a first die, by moving the memories <b>1620</b> for these circuits to a second die <b>1610</b> that is face-to-face mounted with the first die <b>1605</b> through DBI boding process. In this example, a 6-by-10 array of compute circuits is illustrated, but in other examples, the array can have larger number of circuits (e.g., more than 100 circuits, more than 1000 circuits). Also, in other embodiments, the compute circuits and their associated memory circuits can be organized in an arrangement other than an array.
0115The compute circuits <b>1615</b> and the memory circuits <b>1620</b> can be any type of computational processing circuits and memory circuits. For instance, in some embodiments, the circuit array <b>1600</b> is part of an FPGA that has an array of logic circuits (e.g., logic gates and/or look-up tables, LUTs) and an array of memory circuits, with each memory in the memory array corresponding to one logic circuit in the circuit array. In other embodiments, the compute circuits <b>1615</b> are neurons of a neural network or multiplier-accumulator (MAC) circuits of neurons. The memory circuits <b>1620</b> in these embodiments store the weights and/or the input/output data for the neurons or the MAC circuits. In still other embodiments, the compute circuits <b>1615</b> are processing circuits of a GPU, and the memory circuits store the input/output data from these processing circuits.
0116As shown in <figref idref="DRAWINGS">FIG. <b>17</b></figref>, a memory array is typically interlaced with the circuit array in most single die implementations today. The combined length of the two interlaced arrays is X microns in the example in <figref idref="DRAWINGS">FIG. <b>17</b></figref>. To connect two circuits in the same column in the array, the wiring would have to be at least X microns. But by moving the memory circuits onto the second die <b>1610</b>, as shown in <figref idref="DRAWINGS">FIG. <b>16</b></figref>, two circuits in the same column can be connected with a minimum wiring length of X/2 microns.
0117Moreover, each memory circuit can have a higher density of storage cells as less space is consumed for defining shared peripheral channels for outputting signals to the circuits, as these output signals can now traverse in the z-axis. Also, by moving the memory circuits onto the second IC die <b>1610</b>, more routing space is available in the open channels <b>1650</b> (on the substrate and metal layers) between the compute circuits in the compute array <b>1600</b> on the first die <b>1605</b>, and between the memory circuits in the memory array <b>1602</b> on the second die <b>1610</b>. This additional routing space makes it easier to connect the outputs of the compute circuits. In many instances, this extra routing space allows these interconnects to have shorter wire lengths. It also makes it easier for compute circuits in some embodiments to read or write data from the memory circuits of other compute circuits. The DBI connections are also used in some embodiments to route signals through the metal layers of the second die <b>1610</b> in order to define the signal paths (i.e., the routes) for connecting the compute circuits <b>1615</b> that are defined on the first die <b>1605</b>.
0118The higher density of DBI connections also allow a higher number of z-axis connections to be defined between corresponding memory and compute circuits that are wholly contained within the footprints (i.e., within the substrate regions occupied by) of a pair of corresponding memory and compute circuits. As mentioned before, these DBI connections connect the top interconnect layer of one die with the top interconnect layer of the other die, while the rest of the connection between a pair of memory and compute circuits is established with interconnects and vias on these dies. Again, such an approach would be highly beneficial when the compute circuits need wide buses (e.g., 128 bit buses, 256 bit buses, 512 bit buses, 1000 bit buses, 4000 bit buses, etc.) to their corresponding memory circuits. One such example would be when the array of compute circuits are arrays of neurons that need to access a large amount of data from their corresponding memory circuits.
0119<figref idref="DRAWINGS">FIGS. <b>18</b> and <b>19</b></figref> illustrates two examples that show how high density DBI connections can be used to reduce the size of an arrangement of compute circuit that is formed by several successive stages of circuits, each of which performs a computation that produces a result that is passed to another stage of circuits until a final stage of circuits is reached. In some embodiments, such an arrangement of compute circuits can be an adder tree, with each compute circuit in the tree being an adder. In other embodiments, the circuits in the arrangement are multiply accumulate (MAC) circuits, such as those used in neural networks to compute dot products.
0120The examples in <figref idref="DRAWINGS">FIGS. <b>18</b> and <b>19</b></figref> both illustrate one implementation of a circuit <b>1800</b> that performs a computation (e.g., an addition or multiplication) based on eight input values. In some embodiments, each input value is a multi-bit value (e.g., a thirty-two bit value). The circuit <b>1800</b> has three stages with the first stage <b>1802</b> having four compute circuits A-D, the second stage <b>1804</b> having two compute circuits E and F, and the third stage <b>1806</b> having a compute circuit G. Each compute circuit in the first stage <b>1802</b> performs an operation based on two input values. In the second stage <b>1804</b>, the compute circuit E performs a computation based on the outputs of compute circuits A and B, while the compute circuit F performs a computation based on the outputs of compute circuits C and D. Lastly, the compute circuit G in the third stage <b>1806</b> performs a computation based on the outputs of compute circuits E and F.
0121<figref idref="DRAWINGS">FIG. <b>18</b></figref> illustrates a prior art implementation of the circuit <b>1800</b> on one IC die <b>1805</b>. In this implementation, the compute circuits A-G are arranged in one row in the following order: A, E, B, G, C, F and D. As shown, the first stage compute circuits A-D (1) receive their inputs from circuits (e.g., memory circuits or other circuits) that are above and below in the planar y-axis direction, and (2) provide their results to the compute circuit E or F. The compute circuits E and F provide the result of their computations to compute circuit G in the middle of the row. The signal path from the compute circuits E and F is relatively long and consumes nearby routing resources. The length and congestion of interconnects become worse as the size of the circuit arrangement (e.g., the adder or multiplication tree) grows. For instance, to implement an adder tree that adds 100 or 1000 input values, numerous adders are needed in numerous stages, which quickly results in long, big data buses to transport computation results between successive stages of adders.
0122<figref idref="DRAWINGS">FIG. <b>19</b></figref> illustrates a novel implementation of the circuit <b>1800</b> that drastically reduces the size of the connections needed to supply the output of compute circuits E and F to the compute circuit G. As shown, this implementation defines the compute circuits A, B, E and G on a first die <b>1910</b>, while defining the compute circuits C, D and F on a second die <b>1905</b> that is face-to-face mounted on the first die <b>1905</b> through a DBI boding process. The compute circuits A, B, E, and G are defined in a region on the first die <b>1910</b> that overlaps with a region on the second die <b>1905</b> in which the compute circuits C, D and F are defined.
0123In this implementation, the compute circuit G is placed below the compute circuit E in the planar y-direction. At this location, the compute circuit G receives the output of the compute circuit E through a short data bus defined on the die <b>1910</b>, while receiving the output of the compute circuit F through (1) z-axis DBI connections that connect overlapping locations <b>1950</b> and <b>1952</b> on the top interconnect layers of the dies <b>1905</b> and <b>1910</b>, and (2) interconnects and vias on these dies that take the output of circuit F to the input of circuit G. In this implementation, the interconnects that provide the inputs to the compute circuit G are very short. The computation circuits E and G are next to each other and hence the signal path just includes a short length of the interconnect and vias between the circuit E and G. Also, the length of interconnects, vias, and z-axis DBI connections needed to provide the output of the compute circuit F to the compute circuit G is very small.
0124Hence, by breaking up the arrangement of the circuit <b>1800</b> between two dies <b>1905</b> and <b>1910</b>, successive compute circuits can be placed closer to each other (because an additional dimension, i.e., the z-axis, is now available for placing circuits near each other), which, in turn, allows shorter interconnects to be defined between compute circuits in successive stages. Also, the high density of DBI connections makes it easier to define larger number of z-axis connections (that are needed for larger z-axis data buses) within the cross section of the regions that are used to define successive compute circuits.
0125Compute circuit arrangements can have more than three stages. For example, large adder or MAC trees can have many more stages (e.g., 8 stages, 10 stages, 12 stages, etc.). To implement such circuit arrangements, some embodiments (1) divide up the compute circuits into two or more groups that are then defined on two or more vertically stacked dies, and (2) arrange the different groups of circuits on these dies to minimize the length of interconnects needed to connect compute circuits in successive stages.
0126<figref idref="DRAWINGS">FIG. <b>20</b></figref> presents an example to illustrate this point. This example shows one implementation of a compute circuit <b>2000</b> that performs a computation (e.g., an addition or multiplication) on sixteen multi-bit input values. This circuit includes two versions <b>2012</b> and <b>2014</b> of the compute circuit <b>1800</b> of <figref idref="DRAWINGS">FIGS. <b>18</b> and <b>19</b></figref>. The compute circuits in the second version are labeled as circuits H-N. Each of these versions has three stages. The outputs of these two versions are provided to a fourth stage compute circuit O that performs a computation based on these outputs, as shown.
0127To implement the four stage circuit <b>2000</b>, the two versions <b>2012</b> and <b>2014</b> have an inverted layout. This is because the compute circuits A, B, and E (that operate on the first four inputs of the first version <b>2012</b>) are defined on IC die <b>2010</b> while the compute circuits H, I and L (that operate on the first four inputs of the second version) are defined on the IC die <b>2005</b>. Similarly, the compute circuits C, D, and F (that operate on the second four inputs of the first version <b>2012</b>) are defined on IC die <b>2005</b> while the compute circuits J, K, and M (that operate on the second four inputs of the second version) are defined on the IC die <b>2010</b>. Also, the third stage circuit G of the first version is defined on the IC die <b>2010</b>, while the third stage circuit N is defined on the IC die <b>2005</b>. The fourth stage aggregating circuit O is also defined on IC die <b>2010</b>. Lastly, the second version <b>2014</b> is placed to the right of the first version in the x-axis direction.
0128This overall inverted arrangement of the second version <b>2014</b> with respect to the first version ensures that the length of the interconnect needed to provide the output of the third stage compute circuits G and N to the fourth stage compute circuit O is short. This is because, like compute circuits E, F and G, the compute circuits L, M, and N are placed in nearby and/or overlapping locations, which allows these three circuits L, M and N to be connected through short DBI connections, and mostly vertical signal paths facilitated by small planar interconnects plus several via connections. This arrangement also places compute circuits G, N and O in nearby and/or overlapping locations, which again allows them to be connected through short DBI connections, and mostly vertical signal paths facilitated by small planar interconnects plus several via connections.
0129<figref idref="DRAWINGS">FIG. <b>21</b></figref> illustrates a device <b>2102</b> that uses a 3D IC <b>2100</b> (like any of the 3D IC <b>210</b>, <b>200</b>, <b>400</b>, <b>600</b>-<b>900</b>). In this example, the 3D IC <b>2100</b> is formed by two face-to-face mounted IC dies <b>2105</b> and <b>2110</b> that have numerous direct bonded connections <b>2115</b> between them. In other examples, the 3D IC <b>2100</b> includes three or more vertically stacked IC dies. As shown, the 3D IC die <b>2100</b> includes a cap <b>2150</b> that encapsulates the dies of this IC in a secure housing <b>2125</b>. On the back side of the die <b>2110</b> one or more TSVs and/or interconnect layers <b>2106</b> are defined to connect the 3D IC to a ball grid array <b>2120</b> (e.g., a micro bump array) that allows this to be mounted on a printed circuit board <b>2130</b> of the device <b>2102</b>. The device <b>2102</b> includes other components (not shown). In some embodiments, examples of such components include one or more memory storages (e.g., semiconductor or disk storages), input/output interface circuit(s), one or more processors, etc.
0130In some embodiments, the first and second dies <b>2105</b> and <b>2110</b> are the first and second dies shown in any of the <figref idref="DRAWINGS">FIGS. <b>1</b>-<b>2</b>, <b>4</b>, <b>6</b>-<b>16</b>, and <b>19</b>-<b>20</b></figref>. In some of these embodiments, the second die <b>2110</b> receives data signals through the ball grid array, and routes the received signals to I/O circuits on the first and second dies through interconnect lines on the interconnect layer and vias between the interconnect layers. When such data signals need to traverse to the first die, these signals traverse through z-axis connections crossing the face-to-face bonding layer.
0131<figref idref="DRAWINGS">FIG. <b>22</b></figref> provides another example of a 3D chip <b>2200</b> that is formed by two face-to-face mounted IC dies <b>2205</b> and <b>2210</b> that are mounted on a ball grid array <b>2240</b>. In this example, the first and second dies <b>2205</b> and <b>2210</b> are face-to-face connected through direct bonded connections (e.g., DBI connections). As shown, several TSVs <b>2222</b> are defined through the second die <b>2210</b>. These TSVs electrically connect to interconnects/pads on the backside of the second die <b>2210</b>, on which multiple levels of interconnects are defined.
0132In some embodiments, the interconnects on the backside of the second die <b>2210</b> create the signal paths for defining one or more system level circuits for the 3D chip <b>2200</b> (i.e., for the circuits of the first and second dies <b>2205</b> and <b>2210</b>). Examples of system level circuits are power circuits, clock circuits, data I/O signals, test circuits, etc. In some embodiments, the circuit components that are part of the system level circuits (e.g., the power circuits, etc.) are defined on the front side of the second die <b>2210</b>. The circuit components can include active components (e.g., transistors, diodes, etc.), or passive/analog components (e.g., resistors, capacitors (e.g., decoupling capacitors), inductors, filters, etc.
0133In some embodiments, some or all of the wiring for interconnecting these circuit components to form the system level circuits are defined on interconnect layers on the backside of the second die <b>2210</b>. Using these backside interconnect layers to implement the system level circuits of the 3D chip <b>2200</b> frees up one or more interconnect layers on the front side of the second die <b>2210</b> to share other types of interconnect lines with the first die <b>2205</b>. The backside interconnect layers are also used to define some of the circuit components (e.g., decoupling capacitors, etc.) in some embodiments. As further described below, the backside of the second die <b>2210</b> in some embodiments can also connect to the front or back side of a third die.
0134In some embodiments, one or more of the layers on the backside of the second die <b>2210</b> are also used to mount this die to the ball grid array <b>2240</b>, which allows the 3D chip <b>2100</b> to mount on a printed circuit board. In some embodiments, the system circuitry receives some or all of the system level signals (e.g., power signals, clock signals, data I/O signals, test signals, etc.) through the ball grid array <b>2240</b> connected to the backside of the third die.
0135<figref idref="DRAWINGS">FIG. <b>23</b></figref> illustrates a manufacturing process <b>2300</b> that some embodiments use to produce the 3D chip <b>2200</b> of <figref idref="DRAWINGS">FIG. <b>22</b></figref>. This figure will be explained by reference to <figref idref="DRAWINGS">FIGS. <b>24</b>-<b>27</b></figref>, which show two wafers <b>2405</b> and <b>2410</b> at different stages of the process. Once cut, the two wafers produce two stacked dies, such as dies <b>2205</b> and <b>2210</b>. Even though the process <b>2300</b> of <figref idref="DRAWINGS">FIG. <b>23</b></figref> cuts the wafers into dies after the wafers have been mounted and processed, the manufacturing process of other embodiments performs the cutting operation at a different stage at least for one of the wafers. Specifically, some embodiments cut the first wafer <b>2405</b> into several first dies that are each mounted on the second wafer before the second wafer is cut into individual second dies.
0136As shown, the process <b>2300</b> starts (at <b>2305</b>) by defining components (e.g., transistors) on the substrates of the first and second wafers <b>2405</b> and <b>2410</b>, and defining multiple interconnect layers above each substrate to define interconnections that form micro-circuits (e.g., gates) on each die. To define these components and interconnects on each wafer, the process <b>2300</b> performs multiple IC fabrication operations (e.g., film deposition, patterning, doping, etc.) for each wafer in some embodiments. <figref idref="DRAWINGS">FIG. <b>24</b></figref> illustrates the first and second wafers <b>2405</b> and <b>2410</b> after several fabrication operations that have defined components and interconnects on these wafers. As shown, the fabrication operations for the second wafer <b>2410</b> defines several TSVs <b>2412</b> that traverse the interconnect layers of the second wafer <b>2410</b> and penetrate a portion of this wafer's substrate <b>2416</b>.
0137After the first and second wafers have been processed to define their components and interconnects, the process <b>2300</b> face-to-face mounts (at <b>2310</b>) the first and second wafers <b>2205</b> and <b>2210</b> through a direct bonding process, such as a DBI process. <figref idref="DRAWINGS">FIG. <b>25</b></figref> illustrates the first and second wafers <b>2405</b> and <b>2410</b> after they have been face-to-face mounted through a DBI process. As shown, this DBI process creates a number of direct bonded connections <b>2426</b> between the first and second wafers <b>2405</b> and <b>2410</b>.
0138Next, at <b>2315</b>, the process <b>2300</b> performs a thinning operation on the backside of the second wafer <b>2410</b> to remove a portion of this wafer's substrate layer. As shown in <figref idref="DRAWINGS">FIG. <b>26</b></figref>, this thinning operation exposes the TSVs <b>2412</b> on the backside of the second wafer <b>2410</b>. After the thinning operation, the process <b>2300</b> defines (at <b>2320</b>) one or more interconnect layers <b>2430</b> the second wafer's backside. <figref idref="DRAWINGS">FIG. <b>27</b></figref> illustrates the first and second wafers <b>2405</b> and <b>2410</b> after interconnect layers have been defined on the second wafer's backside.
0139These interconnect layers <b>2430</b> include one or more layers that allow the 3D chip stack to electrically connect to the ball grid array. In some embodiments, the interconnect lines/pads on the backside of the third wafer also produce one or more redistribution layers (RDL layers) that allow signals to be redistributed to different locations on the backside. The interconnect layers <b>2430</b> on the backside of the second die in some embodiments also create the signal paths for defining one or more system level circuits (e.g., power circuits, clock circuits, data I/O signals, test circuits, etc.) for the circuits of the first and second dies. In some embodiments, the system level circuits are defined by circuit components (e.g., transistors, etc.) that are defined on the front side of the second die. The process <b>2300</b> in some embodiments does not define interconnect layers on the backside of the second wafer to create the signal paths for the system level circuits, as it uses only the first and second dies' interconnect layers between their two faces for establishing the system level signal paths.
0140After defining the interconnect layers on the backside of the second wafer <b>2410</b>, the process cuts (at <b>2325</b>) the stacked wafers into individual chip stacks, with each chip stack include two stacked IC dies <b>2205</b> and <b>2210</b>. The process then mounts (at <b>2330</b>) each chip stack on a ball grid array and encapsulates the chip stack within one chip housing (e.g., by using a chip case). The process then ends.
0141In some embodiments, three or more IC dies are stacked to form a 3D chip. <figref idref="DRAWINGS">FIG. <b>28</b></figref> illustrates an example of a 3D chip <b>2800</b> with three stacked IC dies <b>2805</b>, <b>2810</b> and <b>2815</b>. In this example, the first and second dies <b>2805</b> and <b>2810</b> are face-to-face connected through direct bonded connections (e.g., DBI connections), while the third and second dies <b>2815</b> and <b>2810</b> are face-to-back connected (e.g., the face of the third die <b>2815</b> is mounted on the back of the second die <b>2810</b>). In some embodiments, the first and second dies <b>2805</b> and <b>2810</b> are the first and second dies shown in any of the <figref idref="DRAWINGS">FIGS. <b>1</b>-<b>2</b>, <b>4</b>, <b>6</b>-<b>16</b>, and <b>19</b>-<b>20</b></figref>.
0142In <figref idref="DRAWINGS">FIG. <b>28</b></figref>, several TSVs <b>2822</b> are defined through the second die <b>2810</b>. These TSVs electrically connect to interconnects/pads on the backside of the second die <b>2810</b>, which connect to interconnects/pads on the top interconnect layer of the third die <b>2815</b>. The third die <b>2815</b> also has a number of TSVs that connect signals on the front side of this die to interconnects/pads on this die's backside. Through interconnects/pads, the third die's backside connects to a ball grid array <b>2840</b> that allows the 3D chip <b>2800</b> to mount on a printed circuit board.
0143In some embodiments, the third die <b>2815</b> includes system circuitry, such as power circuits, clock circuits, data I/O circuits, test circuits, etc. The system circuitry of the third die <b>2815</b> in some embodiments supplies system level signals (e.g., power signals, clock signals, data I/O signals, test signals, etc.) to the circuits of the first and second dies <b>2805</b> and <b>2810</b>. In some embodiments, the system circuitry receives some or all of the system level signals through the ball grid array <b>2840</b> connected to the backside of the third die.
0144<figref idref="DRAWINGS">FIG. <b>29</b></figref> illustrates another example of a 3D chip <b>2900</b> with more than two stacked IC dies. In this example, the 3D chip <b>2900</b> has four IC dies <b>2905</b>, <b>2910</b>, <b>2915</b> and <b>2920</b>. In this example, the first and second dies <b>2905</b> and <b>2910</b> are face-to-face connected through direct bonded connections (e.g., DBI connections), while the third and second dies <b>2915</b> and <b>2910</b> are face-to-back connected (e.g., the face of the third die <b>2915</b> is mounted on the back of the second die <b>2910</b>) and the fourth and third dies <b>2920</b> and <b>2915</b> are face-to-back connected (e.g., the face of the fourth die <b>2920</b> is mounted on the back of the third die <b>2915</b>). In some embodiments, the first and second dies <b>2905</b> and <b>2910</b> are the first and second dies shown in any of the <figref idref="DRAWINGS">FIGS. <b>1</b>-<b>2</b>, <b>4</b>, <b>6</b>-<b>16</b>, and <b>19</b>-<b>20</b></figref>.
0145In <figref idref="DRAWINGS">FIG. <b>29</b></figref>, several TSVs <b>2922</b> are defined through the second, third and fourth die <b>2910</b>, <b>2915</b> and <b>2920</b>. These TSVs electrically connect to interconnects/pads on the backside of these dies, which connect to interconnects/pads on the top interconnect layer of the die below or the interconnect layer below. Through interconnects/pads and TSVs, the signals from outside of the chip are received from the ball grid array <b>2940</b>.
0146Other embodiments use other 3D chip stacking architectures. For instance, instead of face-to-back mounting the fourth and third dies <b>2920</b> and <b>2915</b> in <figref idref="DRAWINGS">FIG. <b>29</b></figref>, the 3D chip stack of another embodiment has these two dies face-to-face mounted, and the second and third dies <b>2910</b> and <b>2915</b> back-to-back mounted. This arrangement would have the third and fourth dies <b>2915</b> and <b>2920</b> share a more tightly arranged set of interconnect layers on their front sides.
0147While the invention has been described with reference to numerous specific details, one of ordinary skill in the art will recognize that the invention can be embodied in other specific forms without departing from the spirit of the invention. For instance, in the examples illustrated in <figref idref="DRAWINGS">FIGS. <b>1</b>-<b>2</b>, <b>4</b>, <b>6</b>-<b>16</b>, and <b>19</b>-<b>20</b></figref>, a first IC die is shown to be face-to-face mounted with a second IC die. In other embodiments, the first IC die is face-to-face mounted with a passive interposer that electrically connects the die to circuits outside of the 3D chip or to other dies that are face-to-face mounted or back-to-face mounted on the interposer. Some embodiments place a passive interposer between two faces of two dies. Some embodiments use an interposer to allow a smaller die to connect to a bigger die.
0148Also, the 3D circuits and ICs of some embodiments have been described by reference to several 3D structures with vertically aligned IC dies. However, other embodiments are implemented with a myriad of other 3D structures. For example, in some embodiments, the 3D circuits are formed with multiple smaller dies placed on a larger die or wafer. <figref idref="DRAWINGS">FIG. <b>30</b></figref> illustrates one such example. Specifically, it illustrates a 3D chip <b>3000</b> that is formed by face-to-face mounting three smaller dies <b>3010</b><i>a</i>-<i>c </i>on a larger die <b>3005</b>. All four dies are housed in one chip <b>3000</b> by having one side of this chip encapsulated by a cap <b>3020</b>, and the other side mounted on a micro-bump array <b>3025</b>, which connects to a board <b>3030</b> of a device <b>1935</b>. Some embodiments are implemented in a 3D structure that is formed by vertically stacking two sets of vertically stacked multi-die structures.
Contents5
28 sheets
Sheet 1 Sheet 2 Sheet 3 Sheet 4 Sheet 5 Sheet 6 Sheet 7 Sheet 8 Sheet 9 Sheet 10 Sheet 11 Sheet 12 Sheet 13 Sheet 14 Sheet 15 Sheet 16 Sheet 17 Sheet 18 Sheet 19 Sheet 20 Sheet 21 Sheet 22 Sheet 23 Sheet 24 Sheet 25 Sheet 26 Sheet 27 Sheet 28
Every citation, both ways
| Document | Relation | Office | Cited during |
|---|---|---|---|
| US10121743B2 | Cites | United States of America | Applicant |
| US10180692B2 | Cites | United States of America | Applicant |
| US10241150B2 | Cites | United States of America | Applicant |
| US10255969B2 | Cites | United States of America | Applicant |
| US10262911B1 | Cites | United States of America | Applicant |
| US10269586B2 | Cites | United States of America | Applicant |
| US10289604B2 | Cites | United States of America | Applicant |
| US10354975B2 | Cites | United States of America | Search report |
| US10373657B2 | Cites | United States of America | Applicant |
| US10446207B2 | Cites | United States of America | Applicant |
| US10446601B2 | Cites | United States of America | Applicant |
| US10522352B2 | Cites | United States of America | Applicant |
| US10580735B2 | Cites | United States of America | Applicant |
| US10580757B2 | Cites | United States of America | Applicant |
| US10580817B2 | Cites | United States of America | Applicant |
| US10586786B2 | Cites | United States of America | Applicant |
| US10593667B2 | Cites | United States of America | Applicant |
| US10600691B2 | Cites | United States of America | Applicant |
| US10600735B2 | Cites | United States of America | Applicant |
| US10600780B2 | Cites | United States of America | Applicant |
| US10607136B2 | Cites | United States of America | Applicant |
| US10672663B2 | Cites | United States of America | Applicant |
| US10672743B2 | Cites | United States of America | Applicant |
| US10672744B2 | Cites | United States of America | Applicant |
| US10672745B2 | Cites | United States of America | Applicant |
| US10719762B2 | Cites | United States of America | Applicant |
| US10762420B2 | Cites | United States of America | Applicant |
| CN1776905A | Cites | China | Search report |
| US2001017418A1 | Cites | United States of America | Applicant |
| US2002008309A1 | Cites | United States of America | Applicant |
| US2003102495A1 | Cites | United States of America | Applicant |
| US2004178819A1 | Cites | United States of America | Search report |
| US2005127490A1 | Cites | United States of America | Applicant |
| US2006036559A1 | Cites | United States of America | Applicant |
| US2006087013A1 | Cites | United States of America | Applicant |
| US2006097402A1 | Cites | United States of America | Search report |
| US2007220207A1 | Cites | United States of America | Applicant |
| US2008150088A1 | Cites | United States of America | Search report |
| US2008225493A1 | Cites | United States of America | Search report |
| US2008237310A1 | Cites | United States of America | Applicant |
| US2008237891A1 | Cites | United States of America | Search report |
| US2009037658A1 | Cites | United States of America | Search report |
| US2009070727A1 | Cites | United States of America | Applicant |
| US2009138688A1 | Cites | United States of America | Search report |
| US2009224388A1 | Cites | United States of America | Search report |
| US2010140750A1 | Cites | United States of America | Applicant |
| JP2010250511A | Cites | Japan | Search report |
| US2010255262A1 | Cites | United States of America | Search report |
| US2010258890A1 | Cites | United States of America | Search report |
| US2010261159A1 | Cites | United States of America | Applicant |
| US2010308455A1 | Cites | United States of America | Search report |
| US2011026293A1 | Cites | United States of America | Applicant |
| US2011078412A1 | Cites | United States of America | Search report |
| US2011102066A1 | Cites | United States of America | Search report |
| US2011131391A1 | Cites | United States of America | Applicant |
| US2011248396A1 | Cites | United States of America | Search report |
| US2012092062A1 | Cites | United States of America | Applicant |
| US2012098140A1 | Cites | United States of America | Search report |
| US2012119357A1 | Cites | United States of America | Applicant |
| US2012124532A1 | Cites | United States of America | Search report |
| US2012136913A1 | Cites | United States of America | Applicant |
| US2012170345A1 | Cites | United States of America | Applicant |
| US2012171818A1 | Cites | United States of America | Search report |
| US2012201068A1 | Cites | United States of America | Applicant |
| US2012242346A1 | Cites | United States of America | Applicant |
| US2012256653A1 | Cites | United States of America | Applicant |
| US2012262196A1 | Cites | United States of America | Applicant |
| US2012286431A1 | Cites | United States of America | Applicant |
| US2012313263A1 | Cites | United States of America | Applicant |
| US2013021866A1 | Cites | United States of America | Applicant |
| US2013032950A1 | Cites | United States of America | Applicant |
| US2013051116A1 | Cites | United States of America | Applicant |
| US2013070436A1 | Cites | United States of America | Search report |
| US2013144542A1 | Cites | United States of America | Applicant |
| US2013187292A1 | Cites | United States of America | Applicant |
| US2013207268A1 | Cites | United States of America | Applicant |
| US2013242500A1 | Cites | United States of America | Applicant |
| US2013275823A1 | Cites | United States of America | Applicant |
| US2013283008A1 | Cites | United States of America | Search report |
| US2013321074A1 | Cites | United States of America | Applicant |
| US2014011324A1 | Cites | United States of America | Search report |
| US2014022002A1 | Cites | United States of America | Applicant |
| US2014176187A1 | Cites | United States of America | Search report |
| US2014177626A1 | Cites | United States of America | Search report |
| US2014269022A1 | Cites | United States of America | Search report |
| US2014285253A1 | Cites | United States of America | Applicant |
| US2014323046A1 | Cites | United States of America | Applicant |
| US2014369148A1 | Cites | United States of America | Applicant |
| US2014379975A1 | Cites | United States of America | Search report |
| KR20150137970A | Cites | Republic of Korea | Applicant |
| US2015032962A1 | Cites | United States of America | Search report |
| US2015121052A1 | Cites | United States of America | Search report |
| US2015162220A1 | Cites | United States of America | Search report |
| US2015199997A1 | Cites | United States of America | Search report |
| US2015228584A1 | Cites | United States of America | Applicant |
| US2015262902A1 | Cites | United States of America | Applicant |
| US2015348943A1 | Cites | United States of America | Search report |
| US2015370754A1 | Cites | United States of America | Search report |
| US2016111386A1 | Cites | United States of America | Search report |
| US2016155725A1 | Cites | United States of America | Search report |
105 members in 6 offices
Priority claims16
| Document | Office | Kind | Date |
|---|---|---|---|
| 201662405833 | United States of America | P | |
| 201762541064 | United States of America | P | |
| 201715725030 | United States of America | A | |
| 201762575184 | United States of America | P | |
| 201762575240 | United States of America | P | |
| 201762575259 | United States of America | P | |
| 201762575221 | United States of America | P | |
| 201715859546 | United States of America | A | |
| 201715859612 | United States of America | A | |
| 201715859551 | United States of America | A | |
| 201715859548 | United States of America | A | |
| 201862619910 | United States of America | P | |
| 201815976809 | United States of America | A | |
| 201862678246 | United States of America | P | |
| 201816159705 | United States of America | A | |
| 202016833341 | United States of America | A |
Members105
| Document | Office | Kind | |
|---|---|---|---|
| US2018102251A1 | United States of America | A1 | |
| WO2018067719A2 | World Intellectual Property Organization (WIPO) | A2 | |
| WO2018067719A3 | World Intellectual Property Organization (WIPO) | A3 | |
| TW201834083A | Taiwan Province of China | A | |
| US2018330992A1 | United States of America | A1 | |
| US2018330993A1 | United States of America | A1 | |
| US2018331037A1 | United States of America | A1 | |
| US2018331038A1 | United States of America | A1 | |
| US2018331072A1 | United States of America | A1 | |
| US2018331094A1 | United States of America | A1 | |
| US2018331095A1 | United States of America | A1 | |
| US2018350775A1 | United States of America | A1 | |
| US2019042377A1 | United States of America | A1 | |
| US2019042912A1 | United States of America | A1 | |
| US2019042929A1 | United States of America | A1 | |
| US2019043832A1 | United States of America | A1 | |
| US2019123022A1 | United States of America | A1 | |
| US2019123023A1 | United States of America | A1 | |
| US2019123024A1 | United States of America | A1 | |
| WO2019079625A1 | World Intellectual Property Organization (WIPO) | A1 | |
| WO2019079631A1 | World Intellectual Property Organization (WIPO) | A1 | |
| KR20190053275A | Republic of Korea | A | |
| TW201931551A | Taiwan Province of China | A | |
| CN110088897A | China | A | |
| TW201933578A | Taiwan Province of China | A | |
| US10522352B2 | United States of America | B2 | |
| US10580735B2 | United States of America | B2 | |
| US10580757B2 | United States of America | B2 | |
| US10586786B2 | United States of America | B2 | |
| US10593667B2 | United States of America | B2 | |
| US10600691B2 | United States of America | B2 | |
| US10600735B2 | United States of America | B2 | |
| US10600780B2 | United States of America | B2 | |
| US10607136B2 | United States of America | B2 | |
| TWI691037B | Taiwan Province of China | B | |
| US10672663B2 | United States of America | B2 | |
| US10672743B2 | United States of America | B2 | |
| US10672744B2 | United States of America | B2 | |
| US10672745B2 | United States of America | B2 | |
| US2020194262A1 | United States of America | A1 | |
| US2020203318A1 | United States of America | A1 | |
| TW202025426A | Taiwan Province of China | A | |
| US2020219771A1 | United States of America | A1 | |
| CN111418060A | China | A | |
| US2020227389A1 | United States of America | A1 | |
| US10719762B2 | United States of America | B2 | |
| CN111492477A | China | A | |
| EP3698401A1 | European Patent Office (EPO) | A1 | |
| EP3698402A1 | European Patent Office (EPO) | A1 | |
| US2020273798A1 | United States of America | A1 | |
| US10762420B2 | United States of America | B2 | |
| US2020293872A1 | United States of America | A1 | |
| US2020294858A1 | United States of America | A1 | |
| US10832912B2 | United States of America | B2 | |
| US2020357641A1 | United States of America | A1 | |
| US10886177B2 | United States of America | B2 | |
| US10892252B2 | United States of America | B2 | |
| US10950547B2 | United States of America | B2 | |
| US10970627B2 | United States of America | B2 | |
| US2021104436A1 | United States of America | A1 | |
| US10978348B2 | United States of America | B2 | |
| TWI725771B | Taiwan Province of China | B | |
| US2021202387A1 | United States of America | A1 | |
| US2021202445A1 | United States of America | A1 | |
| TWI737832B | Taiwan Province of China | B | |
| US11152336B2 | United States of America | B2 | |
| TWI745626B | Taiwan Province of China | B | |
| US11176450B2 | United States of America | B2 | |
| US2022068890A1 | United States of America | A1 | |
| US11289333B2 | United States of America | B2 | |
| US2022108161A1 | United States of America | A1 | |
| KR102393946B1 | Republic of Korea | B1 | |
| KR20220060559A | Republic of Korea | A | |
| US2022238339A1 | United States of America | A1 | |
| US11557516B2 | United States of America | B2 | |
| KR102512017B1 | Republic of Korea | B1 | |
| KR20230039780A | Republic of Korea | A | |
| US2023137580A1 | United States of America | A1 | |
| US11790219B2 | United States of America | B2 | |
| US11823906B2 | United States of America | B2 | |
| US11824042B2 | United States of America | B2 | |
| US11881454B2 | United States of America | B2 | |
| KR102647767B1 | Republic of Korea | B1 | |
| KR20240036154A | Republic of Korea | A | |
| US2024152743A1 | United States of America | A1 | |
| US2024234320A1 | United States of America | A1 | |
| US2024265305A1 | United States of America | A1 | |
| US2024266325A1 | United States of America | A1 | |
| CN110088897B | China | B | |
| US12142528B2 | United States of America | B2 | |
| CN119028957A | China | A | |
| CN119028958A | China | A | |
| US12218059B2 | United States of America | B2 | |
| US12248869B2 | United States of America | B2 | |
| US2025142942A1 | United States of America | A1 | |
| US12293993B2 | United States of America | B2 | |
| US2025174561A1 | United States of America | A1 | |
| US12362182B2 | United States of America | B2 | |
| US2025252299A1 | United States of America | A1 | |
| US12401010B2This record | United States of America | B2 |
157 transactions on the USPTO file
Allowed after 6 non-final rejections, 2 final rejections, 1 RCE and 1 appeal.
- Non-final rejections
- 6
- Final rejections
- 2
- RCEs
- 1
- Appeals
- 1
Over time
Point at a mark for the transactionTransactions
| Event | Code | |
|---|---|---|
| Recordation of Patent Grant MailedPGM/ | PGM/ | |
| Email NotificationEML_NTR | EML_NTR | |
| Mail Patent eGrant NotificationMEPG_NTF | MEPG_NTF | |
| Patent eGrant NotificationEPG_NTF | EPG_NTF | |
| Recordation of Patent eGrantEPG/ | EPG/ | |
| Patent Issue Date Used in PTA CalculationAllowedPTAC | PTAC | |
| Email NotificationEML_NTR | EML_NTR | |
| Issue Notification MailedAllowedWPIR | WPIR | |
| Dispatch to FDCD1935 | D1935 | |
| Application Is Considered Ready for IssuePILS | PILS | |
| Issue Fee Payment VerifiedN084 | N084 | |
| Issue Fee Payment ReceivedIFEE | IFEE | |
| Mail Notice of AllowanceAllowedMN/=. | MN/=. | |
| Notice of Allowance Data Verification CompletedAllowedN/=. | N/=. | |
| Date Forwarded to ExaminerFWDX | FWDX | |
| Response after Non-Final ActionA... | A... | |
| Electronic ReviewELC_RVW | ELC_RVW | |
| Email NotificationEML_NTF | EML_NTF | |
| Mail Non-Final RejectionNon-final rejectionMCTNF | MCTNF | |
| Non-Final RejectionNon-final rejectionCTNF | CTNF | |
| Information Disclosure Statement consideredIDSC | IDSC | |
| Date Forwarded to ExaminerFWDX | FWDX | |
| Disposal for a RCE / CPA / R129AbandonedABN9 | ABN9 | |
| Electronic Information Disclosure StatementEIDS. | EIDS. | |
| Request for Continued Examination (RCE)RCEX | RCEX | |
| Information Disclosure Statement (IDS) FiledWIDS | WIDS | |
| Workflow - Request for RCE - BeginBRCE | BRCE | |
| Email NotificationEML_NTR | EML_NTR | |
| Mail Examiner Interview Summary (PTOL - 413)MEXIN | MEXIN | |
| Interview Summary - Applicant Initiated - TelephonicEXAT | EXAT | |
| Interview Summary RecordEXIN | EXIN | |
| Electronic ReviewELC_RVW | ELC_RVW | |
| Email NotificationEML_NTF | EML_NTF | |
| Mail Final Rejection (PTOL - 326)Final rejectionMCTFR | MCTFR | |
| Final RejectionFinal rejectionCTFR | CTFR | |
| Date Forwarded to ExaminerFWDX | FWDX | |
| Response after Non-Final ActionA... | A... | |
| Email NotificationEML_NTR | EML_NTR | |
| Mail Examiner Interview Summary (PTOL - 413)MEXIN | MEXIN | |
| Interview Summary - Applicant Initiated - TelephonicEXAT | EXAT | |
| Interview Summary RecordEXIN | EXIN | |
| Electronic ReviewELC_RVW | ELC_RVW | |
| Email NotificationEML_NTF | EML_NTF | |
| Mail Non-Final RejectionNon-final rejectionMCTNF | MCTNF | |
| Non-Final RejectionNon-final rejectionCTNF | CTNF | |
| Appeal Brief Review CompleteAPBR | APBR | |
| Date Forwarded to ExaminerFWDX | FWDX | |
| track 1 OFFT1OFF | T1OFF | |
| Appeal Brief FiledAP.B | AP.B | |
| Email NotificationEML_NTR | EML_NTR | |
| Notice -- Defective Appeal BriefAPBD | APBD | |
| Appeal Brief Review CompleteAPBR | APBR | |
| Date Forwarded to ExaminerFWDX | FWDX | |
| Email NotificationEML_NTR | EML_NTR | |
| Change in Power of Attorney (May Include Associate POA)PA.. | PA.. | |
| Email NotificationEML_NTR | EML_NTR | |
| Filing Receipt - ReplacementFLRCPT.R | FLRCPT.R | |
| track 1 OFFT1OFF | T1OFF | |
| Defective / Incomplete Appeal Brief FiledAPBI | APBI | |
| Appeal Brief FiledAP.B | AP.B | |
| Notice of Appeal FiledN/AP | N/AP | |
| Electronic ReviewELC_RVW | ELC_RVW | |
| Electronic ReviewELC_RVW | ELC_RVW | |
| Electronic ReviewELC_RVW | ELC_RVW | |
| Email NotificationEML_NTF | EML_NTF | |
| Mail Non-Final RejectionNon-final rejectionMCTNF | MCTNF | |
| Non-Final RejectionNon-final rejectionCTNF | CTNF | |
| Information Disclosure Statement consideredIDSC | IDSC | |
| Date Forwarded to ExaminerFWDX | FWDX | |
| Information Disclosure Statement (IDS) FiledM844 | M844 | |
| Response after Non-Final ActionA... | A... | |
| Electronic Information Disclosure StatementEIDS. | EIDS. | |
| Electronic Information Disclosure StatementEIDS. | EIDS. | |
| Information Disclosure Statement (IDS) FiledWIDS | WIDS | |
| Email NotificationEML_NTR | EML_NTR | |
| Mail Examiner Interview Summary (PTOL - 413)MEXIN | MEXIN | |
| Interview Summary - Applicant Initiated - TelephonicEXAT | EXAT | |
| Interview Summary RecordEXIN | EXIN | |
| Electronic ReviewELC_RVW | ELC_RVW | |
| Electronic ReviewELC_RVW | ELC_RVW | |
| Email NotificationEML_NTF | EML_NTF | |
| Mail Non-Final RejectionNon-final rejectionMCTNF | MCTNF | |
| Non-Final RejectionNon-final rejectionCTNF | CTNF | |
| Date Forwarded to ExaminerFWDX | FWDX | |
| Supplemental ResponseSA.. | SA.. | |
| Date Forwarded to ExaminerFWDX | FWDX | |
| Response after Non-Final ActionA... | A... | |
| Email NotificationEML_NTR | EML_NTR | |
| Mail Examiner Interview Summary (PTOL - 413)MEXIN | MEXIN | |
| Interview Summary - Applicant Initiated - TelephonicEXAT | EXAT | |
| Interview Summary RecordEXIN | EXIN | |
| Electronic request for Examiner InterviewM865E | M865E | |
| Electronic ReviewELC_RVW | ELC_RVW | |
| Electronic ReviewELC_RVW | ELC_RVW | |
| Email NotificationEML_NTF | EML_NTF | |
| Mail Non-Final RejectionNon-final rejectionMCTNF | MCTNF | |
| Non-Final RejectionNon-final rejectionCTNF | CTNF | |
| Date Forwarded to ExaminerFWDX | FWDX | |
| Response after Final ActionA.NE | A.NE | |
| Request for Extension of Time - GrantedXT/G | XT/G |
22 legal events, as the office reported them to INPADOC
Over the term
Point at a mark for the eventEvents
| Event | Code | |
|---|---|---|
| Information on status: patent grantGrantedPATENTED CASESTCF | STCF | |
| Information on status: patent application and granting procedure in generalPUBLICATIONS -- ISSUE FEE PAYMENT VERIFIEDSTPP | STPP | |
| Information on status: patent application and granting procedure in generalPUBLICATIONS -- ISSUE FEE PAYMENT RECEIVEDSTPP | STPP | |
| Information on status: patent application and granting procedure in generalNOTICE OF ALLOWANCE MAILED -- APPLICATION RECEIVED IN OFFICE OF PUBLICATIONSSTPP | STPP | |
| Information on status: patent application and granting procedure in generalRESPONSE TO NON-FINAL OFFICE ACTION ENTERED AND FORWARDED TO EXAMINERSTPP | STPP | |
| Information on status: patent application and granting procedure in generalNON FINAL ACTION MAILEDSTPP | STPP | |
| Information on status: appeal procedureAppealAPPEAL BRIEF (OR SUPPLEMENTAL BRIEF) ENTERED AND FORWARDED TO EXAMINERSTCV | STCV | |
| AssignmentAS | AS | |
| Information on status: appeal procedureAppealNOTICE OF APPEAL FILEDSTCV | STCV | |
| AssignmentAS | AS | |
| Information on status: appeal procedureAppealNOTICE OF APPEAL FILEDSTCV | STCV | |
| Information on status: patent application and granting procedure in generalNON FINAL ACTION MAILEDSTPP | STPP | |
| Information on status: patent application and granting procedure in generalRESPONSE TO NON-FINAL OFFICE ACTION ENTERED AND FORWARDED TO EXAMINERSTPP | STPP | |
| Information on status: patent application and granting procedure in generalNON FINAL ACTION MAILEDSTPP | STPP | |
| Information on status: patent application and granting procedure in generalRESPONSE TO NON-FINAL OFFICE ACTION ENTERED AND FORWARDED TO EXAMINERSTPP | STPP | |
| Information on status: patent application and granting procedure in generalNON FINAL ACTION MAILEDSTPP | STPP | |
| Information on status: patent application and granting procedure in generalRESPONSE AFTER FINAL ACTION FORWARDED TO EXAMINERSTPP | STPP | |
| Information on status: patent application and granting procedure in generalFINAL REJECTION MAILEDSTPP | STPP | |
| AssignmentAS | AS | |
| Information on status: patent application and granting procedure in generalDOCKETED NEW CASE - READY FOR EXAMINATIONSTPP | STPP | |
| AssignmentAS | AS | |
| Fee payment procedureENTITY STATUS SET TO UNDISCOUNTED (ORIGINAL EVENT CODE: BIG.); ENTITY STATUS OF PATENT OWNER: LARGE ENTITYFEPP | FEPP |
Numbers
- Publication
- 12401010
- Application
- 17474917
Titles
- English
- 3D processor having stacked integrated circuit die
Patent term adjustment
- A delay
- +58 daysthe office missed an examination deadline
- B delay
- +277 dayspendency past three years
- Applicant delay
- −72 days
- Net adjustment
- 263 days
Classification
- CPC, 47
- H01L25/18
- H10W90/00
- H10K19/201
- H10W72/019
- H01L24/92
- H10W90/732
- H01L25/0657
- H01L24/03
- H10W90/792
- H01L24/08
- H10W72/07254
- H10W72/247
- H01L24/11
- H10W90/722
- H01L24/16
- H01L24/32
- H10W90/724
- H01L24/80
- H10W80/312
- H01L25/043
- H10W80/327
- H01L25/074
- H10W72/012
- H01L25/0756
- H01L25/105
- H10W90/20
- H01L25/117
- H10W90/26
- H01L25/16
- H10W90/297
- H01L2224/08145
- H10W99/00
- H01L2224/16145
- H01L2224/16225
- H10B41/20
- H01L2224/17181
- H10B51/20
- H01L2224/32145
- H10B53/20
- H01L2224/80895
- H10B63/84
- H01L2224/80896
- H10D88/00
- H01L2224/9202
- H01L2225/06503
- H01L2225/06548
- H10W72/823
- IPC, 15
- H10B41 20
- H01L23 00
- H01L25 065
- H01L25 18
- H01L25 04
- H01L25 07
- H01L25 075
- H01L25 10
- H01L25 11
- H01L25 16
- H10B51 20
- H10B53 20
- H10B63 00
- H10D88 00
- H10K19 00