Method and apparatus for dual issue of program instructions to symmetric multifunctional execution units
Summary by NHIP
Dual Issue Control Unit
The instruction issue control unit issues undecoded program instructions to multiple symmetrical multifunctional processing logic units. Each unit contains identical sets of execution circuits, including floating point adders, dividers, ALUs, and integer multipliers.
Claim Score by NHIP
Abstract
A microprocessor capable of processing at least two program instructions at the same time and capable of issuing the two program instructions to two symmetrical multifunctional program execution units. The microprocessor includes a plurality of registers which store a plurality of operands and an instruction issue control which controls issuance of program instructions to the two symmetrical multifunctional program execution units. The instruction issue control issues the two program instructions (e.g. first and second) without decoding them in order to determine the processing functions required to be performed in response to the two program instructions.

Term
Term ended
Expired 6 March 2020, 6.6 years ago.
- Priority
- Filed
- Granted
- Expired
- Today
37 claims: 4 independent, 33 dependent
- 1Broadest claimClaim Score 79, broad(NHIP)An instruction issue control unit comprising logic to issue an instruction that is not decoded to a first multifunctional processing logic unit of a plurality of multifunctional processing logic units that each contain a set of independent processing logic circuits capable of performing a matching set of functions.
- 13A microprocessor comprising:a plurality of registers to store operands;an instruction cache to store instructions;an instruction issue control unit coupled with the instruction cache to receive instructions from the cache and to issue instructions;a first multifunctional execution unit coupled with the instruction issue control unit and with the plurality of registers and containing a first plurality of different execution units that are capable of providing a first set of functions on operands received from the cache according to instructions received from the instruction issue control unit;a first multiplexer coupled between the plurality of registers and the first multifunctional execution unit to receive a first operand from the registers and provide the first operand to the first multifunctional execution unit, wherein the first multiplexer has a single input to receive a result of an executed instruction that has been executed by the first multifunctional execution unit;a second multifunctional execution unit coupled with the instruction issue control unit and with the plurality of registers and containing a second plurality of different execution units that are capable of providing a second set of functions on operands received from the cache according to instructions received from the instruction issue control unit, wherein the first set of functions and the second set of functions match;and a second multiplexer coupled between the plurality of registers and the second multifunctional execution unit to receive a second operand from the registers and provide the second operand to the second multifunctional execution unit, wherein the second multiplexer has only one input to receive a result of an executed instruction that has been executed by the second multifunctional execution unit.
- 16A microprocessor comprising:a plurality of registers to store at least one operand;an instruction cache to store at least one un-decoded instruction;an instruction issue control unit coupled with the instruction cache to receive the instruction and comprising logic to issue the un-decoded instruction;and a plurality of multifunctional processing pipelines that are each coupled with the plurality of registers to receive the operand, that are each coupled with the instruction issue control unit to receive the issued instruction, and that each comprise a set of independent processing logic circuits that are capable of providing a matching set of processing functions, wherein each set of independent processing logic circuits contains at least two different independent processing logic circuits including an independent processing logic circuit that is capable of processing the operand according to the issued instruction.
- 28An integrated circuit comprising:a plurality of registers to store at least one operand;an instruction cache to store at least one un-decoded instruction;instruction issue means coupled with the instruction cache for issuing the un-decoded instruction received from the instruction cache;and a plurality of multifunctional processing pipelines that are each coupled with the plurality of registers to receive the operand, that are each coupled with the instruction issue means to receive the issued un-decoded instruction, and that each comprise a set of independent processing logic circuits that are capable of providing a matching set of processing functions, wherein each set of independent processing logic circuits contains at least two different independent processing logic circuits including an independent processing logic circuit that is capable of processing the operand according to the issued instruction.
Independent claims4
40 paragraphs in 4 sections, as filed
This is a continuation of application Ser. No. 08/883,147, filed on Jun. 27, 1997, now U.S. Pat. No. 6,035,388.
BACKGROUND OF THE INVENTION
The present invention relates generally to the field of microprocessors which are capable of processing at least two program instructions at the same time.
Modern microprocessors, including superscalar microprocessors, have improved performance due to the capability of processing at least two program instructions at the same time. This capability arises from having a first group of execution units which can receive a program instruction for execution, and a second group of execution units which can receive a second program instruction for execution.
FIG. 1 shows a typical microprocessor of the prior art which uses a dual issue mechanism wherein two program instructions may be issued to two groups of execution units. The register file <b>30</b> and the issue control and bypass control unit <b>12</b> support the issuance of two instructions, one instruction going to the execution units <b>14</b>, <b>16</b>, and <b>18</b> (issue left) and another instruction going to the execution units <b>20</b>, <b>22</b>, and <b>24</b> (issue right). Within each group of execution units, there are three specialized functional units. In particular, the execution units on the issue left side include a floating point division execution unit <b>14</b>, a floating point multiplier execution unit <b>16</b>, and an ALU execution unit <b>18</b>. In the group of execution units on the issue right side, there is a floating point adder execution unit <b>20</b>, an ALU execution unit <b>22</b>, and an integer multiplier unit <b>24</b>. Each of these execution units is coupled to the issue control and bypass control unit <b>12</b> by a bi-directional link which provides instructions from the issue control unit <b>12</b> to the particular execution unit and which provides a signal indicating the execution unit is busy to the issue control and bypass control unit <b>12</b>. In this manner, the issue control and bypass control unit <b>12</b> can determine the status of each execution unit (e.g. is the particular execution unit busy executing an instruction previously provided?) and can provide instructions for execution if the particular execution unit is not busy. These links are shown as <b>31</b>A-<b>31</b>F in FIG. <b>1</b>. The issue control and bypass unit <b>12</b> is coupled to an instruction cache <b>10</b> through a bus <b>11</b>. It will be appreciated that the issue control unit <b>12</b> provides read commands to the instruction cache <b>10</b> to cause the instruction cache <b>10</b> to deliver one or two instructions at a time to the issue control unit <b>12</b>.
Each execution unit within a group of execution units is coupled to an output of a multiplexer in order to receive operands which are processed according to the instruction being executed in the particular execution unit. These operands are received from either the register file <b>30</b> or from a bypass pathway in which an output from a prior executed instruction is used as an operand for a current instruction. The multiplexer <b>26</b> receives an output <b>30</b><i>b </i>from the register file and also receives an output from each of the six execution units and provides a selected output to the three execution units <b>14</b>, <b>16</b>, and <b>18</b> in the issue left group of execution units. The multiplexer <b>28</b> receives an output <b>30</b><i>a </i>from the register file <b>30</b> and also receives outputs from each of the six execution units, and provides an output which is selected by the control select line <b>15</b>. This output is provided to the three execution units <b>20</b>, <b>22</b>, and <b>24</b> in the issue right execution group. The six outputs <b>32</b><i>a</i>, <b>32</b><i>b</i>, <b>32</b><i>c</i>, <b>32</b><i>d</i>, <b>32</b><i>e</i>, and <b>32</b><i>f </i>from the six execution units <b>14</b>, <b>16</b>, <b>18</b>, <b>20</b>, <b>22</b>, and <b>24</b> are provided to both multiplexers <b>26</b> and <b>28</b> and also provided to the register file <b>30</b> as inputs to the register file <b>30</b>. It will be appreciated that the register file <b>30</b> may be configured to provide dual port reads such that operand outputs <b>30</b><i>a </i>and <b>30</b><i>b </i>can be provided based upon the addresses provided over address bus <b>17</b> from the control unit <b>12</b>. Moreover, the register file <b>30</b> may support multiple writes, such as six multiple write ports from the six outputs. It will also be appreciated that in typical operation of the microprocessor shown in FIG. 1, only two of the write ports will be active at once since normally only one result of a computation is provided from the issue left side and only one execution result is provided from the issue right side.
The operation of the microprocessor shown in FIG. 1 will now be described. The issue control unit <b>12</b> receives two instructions from the instruction cache <b>10</b>. The issue control unit <b>12</b> then decodes each instruction to determine the resources or functions to be performed as required by the particular instruction. For example, if an instruction requires floating point division or floating point multiplication, then this instruction must be steered into the issue left group of execution units. Similarly, if a decoded instruction reveals that a floating point addition or integer multiplication is required by the instruction, then it must be issued to the issue right group of execution units. Thus, decoding in the issue control and bypass control unit <b>12</b> is required in order to determine whether an instruction goes to issue left or to issue right.
The issue control and bypass control unit <b>12</b> must also perform the resolution of execution unit conflicts before issuing an instruction. The following table shows an example of the stall logic in the issue control unit <b>12</b> in order to resolve execution unit conflicts. If there is an execution unit conflict indicated by a “X”, then the issue control will stall the issue of the instruction.
<tables><table frame="none" colsep="0" rowsep="0"><tgroup align="left" colsep="0" rowsep="0" cols="2"><colspec colname="1" colwidth="35pt" align="left" /><colspec colname="2" colwidth="182pt" align="center" /><thead><row><entry namest="1" nameend="2" rowsep="1">TABLE A</entry></row></thead><tbody valign="top"><row><entry namest="1" nameend="2" align="center" rowsep="1" /></row><row><entry>Issue</entry><entry /></row><row><entry>Instruction</entry><entry>Instruction in Unit:</entry></row></tbody></tgroup><tgroup align="left" colsep="0" rowsep="0" cols="7"><colspec colname="1" colwidth="35pt" align="left" /><colspec colname="2" colwidth="35pt" align="center" /><colspec colname="3" colwidth="28pt" align="center" /><colspec colname="4" colwidth="28pt" align="center" /><colspec colname="5" colwidth="28pt" align="center" /><colspec colname="6" colwidth="35pt" align="center" /><colspec colname="7" colwidth="28pt" align="center" /><tbody valign="top"><row><entry>Type</entry><entry>FP Div</entry><entry>FP Mult</entry><entry>FP Add</entry><entry>ALU 0</entry><entry>ALU 1</entry><entry>Int Mult</entry></row><row><entry namest="1" nameend="7" align="center" rowsep="1" /></row><row><entry>FP Div</entry><entry>X</entry><entry /><entry /><entry /><entry /><entry /></row><row><entry>FP Mult</entry><entry /><entry>X</entry><entry /><entry /><entry /><entry /></row><row><entry>FP Add</entry><entry /><entry /><entry>X</entry><entry /><entry /><entry /></row><row><entry>ALU 1</entry><entry /><entry /><entry /><entry>X</entry><entry>X</entry><entry /></row><row><entry>Int Mult</entry><entry /><entry /><entry /><entry /><entry /><entry>X</entry></row><row><entry namest="1" nameend="7" align="center" rowsep="1" /></row></tbody></tgroup></table></tables>
For example, if the issue instruction is of the type “FP Div” (i.e., the instruction is for a floating point division), the instruction will stall if there is an instruction in the floating point division unit <b>14</b> which is currently being executed by the floating point execution unit <b>14</b>.
The issue control and bypass unit <b>12</b> also stalls the issuance of instructions in order to resolve register conflicts.
<tables><table frame="none" colsep="0" rowsep="0"><tgroup align="left" colsep="0" rowsep="0" cols="2"><colspec colname="1" colwidth="42pt" align="left" /><colspec colname="2" colwidth="175pt" align="center" /><thead><row><entry namest="1" nameend="2" rowsep="1">TABLE B</entry></row></thead><tbody valign="top"><row><entry namest="1" nameend="2" align="center" rowsep="1" /></row><row><entry>Issue</entry><entry /></row><row><entry>Instruction</entry><entry>Instruction in Unit:</entry></row></tbody></tgroup><tgroup align="left" colsep="0" rowsep="0" cols="7"><colspec colname="1" colwidth="42pt" align="left" /><colspec colname="2" colwidth="28pt" align="center" /><colspec colname="3" colwidth="28pt" align="center" /><colspec colname="4" colwidth="35pt" align="center" /><colspec colname="5" colwidth="28pt" align="center" /><colspec colname="6" colwidth="28pt" align="center" /><colspec colname="7" colwidth="28pt" align="center" /><tbody valign="top"><row><entry>Registers</entry><entry>FP Div</entry><entry>FP Mult</entry><entry>FP Add</entry><entry>ALU 0</entry><entry>ALU 1</entry><entry>Int Mult</entry></row><row><entry namest="1" nameend="7" align="center" rowsep="1" /></row><row><entry>Operand 1</entry><entry>X</entry><entry>X</entry><entry>X</entry><entry>X</entry><entry>X</entry><entry>X</entry></row><row><entry>Operand 2</entry><entry>X</entry><entry>X</entry><entry>X</entry><entry>X</entry><entry>X</entry><entry>X</entry></row><row><entry>Destination</entry><entry>X</entry><entry>X</entry><entry>X</entry><entry>X</entry><entry>X</entry><entry>X</entry></row><row><entry namest="1" nameend="7" align="center" rowsep="1" /></row></tbody></tgroup></table></tables>
If there is a register match between any of the registers indicated by “X”, then the issue control unit <b>12</b> will stall the issue on the instruction. For example, if for a particular instruction which is yet to be issued, if the first operand for the instruction is to be stored in the same register as the destination register for a floating point division operation which is currently being executed, then the yet to be issued instruction will be stalled.
The issue control unit <b>12</b> also resolves resource conflicts at the register file <b>30</b> which arise because different instructions have different processor cycle times. This is shown by way of example in Table C below.
<tables><table frame="none" colsep="0" rowsep="0"><tgroup align="left" colsep="0" rowsep="0" cols="5"><colspec colname="1" colwidth="21pt" align="center" /><colspec colname="2" colwidth="49pt" align="left" /><colspec colname="3" colwidth="49pt" align="left" /><colspec colname="4" colwidth="49pt" align="left" /><colspec colname="5" colwidth="49pt" align="left" /><thead><row><entry namest="1" nameend="5" rowsep="1">TABLE C</entry></row><row><entry namest="1" nameend="5" align="center" rowsep="1" /></row><row><entry>Cycle</entry><entry>ALU0</entry><entry>ALU1</entry><entry>FP Mult</entry><entry>FP Add</entry></row><row><entry namest="1" nameend="5" align="center" rowsep="1" /></row></thead><tbody valign="top"><row><entry>1</entry><entry>—</entry><entry>—</entry><entry>Issue</entry><entry>—</entry></row><row><entry>2</entry><entry>—</entry><entry>—</entry><entry>Execute</entry><entry>Issue</entry></row><row><entry>3</entry><entry>—</entry><entry>—</entry><entry>Execute</entry><entry>Execute</entry></row><row><entry>4</entry><entry>Issue</entry><entry>Issue</entry><entry>Execute</entry><entry>Execute</entry></row><row><entry>5</entry><entry>Write Result</entry><entry>Write Result</entry><entry>Write Result</entry><entry>Write Result</entry></row><row><entry namest="1" nameend="5" align="center" rowsep="1" /></row></tbody></tgroup></table></tables>
As can be seen from Table C, in Cycle 5 there are four results that are produced. Unless there are four write ports into the register file which contains a plurality of registers, the issue of instructions for the ALU0 and ALU1 in Cycle 4 may need to be stalled.
As can be seen from the foregoing description, the issue control unit must perform a variety of control operations in order to resolve various conflicts and yet attempt to issue two program instructions substantially concurrently if possible. It will be appreciated that such control, such as the decoding of program instructions in order to steer the instruction into the appropriate group of execution units, requires considerable circuitry and also requires considerable time in designing such a control unit for this type of microprocessor.
In many circumstances, it will be desirable to provide a microprocessor which requires less complicated issue control.
SUMMARY OF THE INVENTION
A microprocessor which is capable of processing at least two program instructions at the same time and which is capable of issuing the two program instructions to two symmetrical multifunctional program execution units is described.
In a typical embodiment, the microprocessor is a superscalar microprocessor which includes a plurality of registers which store a plurality of operands and further includes an instruction issue control unit which controls the issuance of program instructions to the two symmetrical multifunctional program execution units. The instruction issue control issues the two program instructions, such as a first and second program instruction, without decoding the instructions in order to determine the processing functions required to be performed in response to the two program instructions.
The symmetrical multifunctional program execution units each consist of a set of independent processing logic circuits which are capable of performing a first set of functions. Typically, the independent processing logic circuits in each of the symmetrical multifunctional program execution units are identical to the extent of the functionality required in response to execution of program instructions issued by the issue control logic.
An example of a method according to the present invention issues two program instructions substantially concurrently to two multifunctional digital logic processing units. It will be appreciated that this method is capable of issuing these two program instructions substantially concurrently although it need not do so in every instance depending on stalls asserted in response to resource conflicts. The issue control logic is capable of substantially concurrently issuing two program instructions to the two multifunctional digital logic processing units and it can do so without first decoding a first and a second program instruction in order to determine the logical function required to be performed in response to the first and second program instructions.
BRIEF DESCRIPTION OF THE DRAWINGS
FIG. 1 shows a prior art implementation for a microprocessor which supports dual issue of program instructions to two groups of execution units.
FIG. 2 illustrates an example of one embodiment of a microprocessor according to the present invention.
FIG. 3 represents another embodiment of a microprocessor according to the present invention.
FIG. 4 illustrates register dependency comparisons performed in one implementation of a microprocessor according to the present invention.
FIG. 5 shows a flowchart representing one method for performing stall and bypass logic processing for a high priority instruction according to one embodiment of the present invention.
FIG. 6 shows a flowchart illustrating a method according to one embodiment of the present invention for determining whether to issue an instruction for execution.
DETAILED DESCRIPTION
The present invention relates to a microprocessor, which is typically implemented on a single semiconductor integrated circuit and which is capable of supporting concurrent processing of at least two program instructions through at least two “pipelines.” The present invention will be described by using various examples which are shown in the accompanying figures and described below. It will be appreciated that various alternative implementations may be utilized in accordance with the present invention.
FIG. 2 shows an example of a microprocessor which executes program instructions and which supports concurrent processing of at least two program instructions. Typically, the microprocessor shown in FIG. 2 will be formed in a monocrystalline silicon semiconductor integrated circuit, although other implementations will be appreciated by those skilled in the art. The microprocessor will include, at least in certain embodiments, numerous well-known supporting circuitry, such as clock generating circuitry, input and output buffer circuitry, data and address buses, and other well-known supporting components. FIG. 2 illustrates the core of a microprocessor according to the present invention. This includes an instruction cache <b>101</b> which is coupled by a bus to an issue control and bypass control unit <b>102</b>. In turn, the issue control and bypass control <b>102</b> unit is coupled to two symmetrical multifunctional units <b>103</b> and <b>104</b>. It will be appreciated that the symmetric multifunctional unit <b>103</b> may support a plurality of different processing functions, such as addition, division, logical operations, and multiplications. The particular group of functions supported will depend on the particular implementation. The multifunctional execution unit <b>104</b> also represents a group of execution units providing a plurality of processing functions. The processing functions provided by the unit <b>104</b> will match the processing functions provided by the unit <b>103</b> such that each unit <b>103</b> and <b>104</b> is capable of performing instructions issued by the issue control and bypass control unit <b>102</b>. Operand inputs for the multifunctional processing unit <b>103</b> is provided by the multiplexer <b>105</b>, and operands for the multifunctional processing unit <b>104</b> is provided by the multiplexer <b>106</b>. These operands are received from either the register file <b>107</b> or from the output of a prior executed instruction from either unit <b>103</b> or unit <b>104</b>.
The register file <b>107</b> includes a plurality of registers each having a specific address or number identifying the register. Data may be stored into the registers from the outputs <b>114</b><i>a </i>and <b>114</b><i>b </i>from the two groups of multifunctional execution units <b>103</b> and <b>104</b>. Moreover, data may be stored into the various registers of the register file <b>107</b> through a data and address bus (not shown) which is coupled to the register file <b>107</b> to provide data to and from the register file <b>107</b>. As shown in FIG. 2, the register file <b>107</b> includes two write ports and two read ports such that it is fully dual ported. Alternatively, the register file <b>107</b> may include an additional read and write port for providing data to and from a data bus on the microprocessor. The read ports <b>115</b><i>a </i>and <b>115</b><i>b </i>provide operand data to the multiplexers <b>105</b> and <b>106</b> respectively. Normally, if a bypass is not required, the operands from the read ports <b>115</b><i>a </i>and <b>115</b><i>b </i>are selected under control of selection line <b>111</b> to be outputted at the output of the multiplexers <b>105</b> and <b>106</b>. These outputs are provided as operand inputs to the multifunctional execution units <b>103</b> and <b>104</b> respectively. The addresses for retrieving the various operands are provided over the address and control bus <b>112</b> from the issue control and bypass control unit <b>102</b>. Each port <b>115</b><i>a </i>and <b>115</b><i>b </i>may each provide a plurality of operands to each execution unit. It will be appreciated that operands may also be obtained from other sources (e.g. a data cache or a bus).
The issue control and bypass control <b>102</b> receives signals over interconnections <b>110</b><i>a </i>and <b>110</b><i>b </i>which indicate the status of the multifunctional execution units <b>103</b> and <b>104</b> respectively. In particular, the issue control and bypass control unit <b>102</b> determines whether each of the multifunctional execution units <b>103</b> and <b>104</b> is busy executing a prior instruction. This is typically performed during each processor cycle under control of a processor clock. If a particular unit is not busy, then the issue control and bypass control unit will issue the next in order program instruction to the multifunctional execution unit which is not busy over the interconnection <b>110</b><i>a </i>or <b>110</b><i>b </i>as appropriate. Using the program instruction provided by the issue control and bypass control <b>102</b>, and using the operands provided through the particular multiplexer, the multifunctional execution unit will perform the operation required by the program instruction and will provide an output which is then provided to an input of both multiplexers <b>105</b> and <b>106</b> and also to the register file <b>107</b>.
In this discussion, it is assumed that the processor executes program instructions in order rather than out of order. In order means that the program instructions are executed in the order of receipt which is typically determined by the order in which the compiler generates and stores executable program instructions. In an alternative embodiment according to the present invention, a microprocessor may implement out of order issuing of program instructions by utilizing reservation stations which are well known in the art.
FIG. 3 shows another embodiment of a microprocessor according to the present invention. This embodiment is similar to the microprocessor shown in FIG. 2 except that a data cache <b>209</b> is shared between the two multifunctional processing logic units <b>203</b> and <b>204</b>. An issue control unit <b>202</b> controls the issue of program instructions as well as controlling the bypass operation and also controls the operation of the data cache <b>209</b>. The issue control <b>202</b> receives instructions, usually two at a time, from the instruction cache <b>201</b>. The issue control unit <b>202</b> is coupled to the two groups of multifunctional processing logic units <b>203</b> and <b>204</b> by the interconnects <b>210</b><i>a </i>and <b>210</b><i>b </i>respectively. These interconnects <b>210</b><i>a </i>and <b>210</b><i>b </i>provide instructions to the units <b>203</b> and <b>204</b> and receive status indicators from these units, such as a busy status. Input operands for units <b>203</b> and <b>204</b> are received from the outputs of multiplexers <b>205</b> and <b>206</b> respectively, and the results of the executed instructions are provided at the outputs <b>214</b><i>a </i>and <b>214</b><i>b </i>respectively of the multifunctional processing logic units <b>203</b> and <b>204</b>. These outputs are routed back as inputs to each of the multiplexers <b>205</b> and <b>206</b> and also as inputs to the register file <b>207</b>. Read ports <b>215</b><i>a </i>and <b>215</b><i>b </i>provide operand inputs from the register file <b>207</b> which are selected when bypassing is disabled. Address and control bus <b>212</b> from the issue control unit <b>202</b> provides address and control signals to the register file <b>207</b> to retrieve operands and to store execution results in the register file <b>207</b>. Select line <b>211</b> controls the bypass or no bypass status of the multiplexers <b>205</b> and <b>206</b>, and this status is controlled by the bypass control unit which is part of the issue control unit <b>202</b>.
As with the example shown in FIG. 2, the microprocessor of FIG. 3 includes two symmetrical multifunctional processing logic units each of which provide the same set of processing operations or functions which are capable of performing various operations or functions as required by the various program instructions issued by the issue control unit <b>202</b>. For example, if multifunctional processing logic unit <b>203</b> includes a floating point adder and a floating point divider and an ALU, then the multifunctional processing logic unit <b>204</b> will include logic which provides the same processing functions. Thus the issue control unit <b>202</b> will not need to decode program instructions in order to determine the function specified by the program instructions which are to be executed in either group of execution units.
It will be appreciated that the data cache <b>209</b> represents one example of a shared resource which may be used in a microprocessor in accordance with the present invention. Other types of shared resources will also be understood to be available to be used by those of ordinary skill in the art. Inputs, such as address and/or data inputs to the data cache are multiplexed by the multiplexer <b>208</b> which is controlled by the cache control <b>202</b> through the select line <b>216</b>. The inputs <b>208</b><i>a </i>and <b>208</b><i>b </i>may be addresses and/or data. The output from data cache <b>209</b> is provided simultaneously over buses <b>209</b><i>a </i>and <b>209</b><i>b </i>to units <b>203</b> and <b>204</b> respectively. In the case of a read of data cache <b>209</b>, an address is provided by the execution unit which is controlling the data cache <b>209</b> over either input bus <b>208</b><i>a </i>or <b>208</b><i>b </i>and this address causes the data cache <b>209</b> to retrieve data which is provided over both output buses <b>209</b><i>a </i>and <b>209</b><i>b</i>. The particular execution unit which is controlling the operation of the data cache <b>209</b> will receive and utilize the data retrieved from the data cache <b>209</b> and the other execution unit will merely ignore the data. When the issue control unit <b>202</b> receives two instructions which are to be issued, it determines whether a shared resource will be required by both instructions. If both instructions require the shared resource, then the issue control <b>202</b> will only issue the high priority instruction and will stall the low priority instruction. It will be appreciated that the high priority instruction in the case of an in order microprocessor is the first program instruction in the order of the executed program. The data cache <b>209</b> may be coupled to a data and address bus (and input/output buffers) in order to exchange data between the cache <b>209</b> and systems (e.g. system RAM) which are separate from the microprocessor.
FIGS. 4, <b>5</b>, and <b>6</b> illustrate the various control operations which are performed by the issue control unit in one example of the present invention. In the following discussion, it will be assumed that the particular example of the issue control unit is the issue control and the bypass control unit <b>102</b> of the example shown in FIG. <b>2</b>.
In the example shown in FIG. 4, there are fifteen comparisons which are performed to determine whether there are any matches between the addresses for various registers which are to be used for the low priority and the high priority instructions that are yet to be issued and the destination registers for instructions that are currently in the left and right groups of execution units. In particular, there are fifteen comparison operations <b>321</b>-<b>335</b>, and the results of these comparison operations determine register dependency stalls and bypass operations. For example, the address or identification of the destination register <b>304</b> for the high priority instruction is compared in comparison <b>322</b> to the address or identification of the destination register <b>301</b> for the low priority instruction (yet to be issued). Similarly, the address or identification of the destination register <b>304</b> for the high priority instruction which is yet to be issued is compared against the address or identification of the first and second operands <b>302</b> and <b>303</b> for the low priority instruction yet to be issued in comparisons <b>321</b> and <b>323</b>. The remainder of the comparisons shown in FIG. 4 determine whether the address or identification of either destination register for both instructions currently in both execution unit pipelines matches the address or identification for the registers <b>301</b>-<b>306</b> of both instructions which are yet to be issued.
FIG. 5 shows the control processing operations performed by an issue control unit of the present invention. This flowchart shows the control processes for the high priority instruction; it will be appreciated that a similar set of control processing operations is performed for the low priority instruction. The flowchart of FIG. 5 may be interpreted to show a sequence in time of these control processes; however, it will be appreciated that these control processes may be performed in parallel. For example, steps <b>401</b>, <b>405</b>, <b>409</b>, and <b>413</b> may be performed in parallel such that the processes are performed substantially concurrently (at the same time). Other methods for performing the processes in parallel will be appreciated by those skilled in the art. In steps <b>401</b>, <b>405</b>, <b>409</b>, and <b>413</b>, the issue control logic determines whether the address for the register for either operand for the high priority instruction matches the address of the destination registers being used (or to be used) by the currently executed instructions in the left and right pipelines. It will be appreciated that the term “pipeline” refers to one group of execution units such that the microprocessor shown in FIG. 2 has two pipelines represented by the multifunctional units <b>103</b> and <b>104</b> respectively. Steps <b>401</b>, <b>405</b>, <b>409</b>, and <b>413</b> represent the comparisons <b>330</b>, <b>333</b>, <b>332</b>, and <b>335</b> respectively of FIG. <b>4</b>. If all four of these comparisons reveal that there are no matches, then in step <b>417</b>, the high priority instruction is issued to a non-busy group of execution units. If any one of the comparisons results in a match then further processing is performed to determine whether to stall the issuance of the high priority instruction or to bypass the result contained in a destination register for use as the operand input for the high priority instruction to be issued. These additional control processing steps are shown as steps <b>402</b>-<b>404</b>, <b>406</b>-<b>408</b>, <b>410</b>-<b>412</b>, and <b>414</b>-<b>416</b>. For example, if the comparison in step <b>401</b> indicates that there is a match between the first operand for the high priority instruction and the destination register for the left pipeline, then step <b>402</b> determines whether the data is available in that destination register (e.g. due to the fact that the execution of the prior instruction has been completed and the result of that execution has been stored in the destination register for the left pipeline). If this register is not available because it does not contain the result of the prior execution then in step <b>404</b> the issue control logic stalls the issuance of the high priority instruction. If the data is available then processing proceeds to step <b>403</b> in which the result from the prior executed instruction is used as the operand <b>1</b> input for the yet to be issued high priority instruction. This is typically implemented by causing the output from the left pipeline to be provided to the multiplexer which is used to route operand inputs to the particular group of execution units. After determining that a bypass is required in step <b>403</b>, processing proceeds to step <b>405</b> (in the embodiment where steps <b>401</b> and <b>405</b> are not performed in parallel).
FIG. 6 shows one example of the control operations performed by the particular example of an issue control logic according to the present invention. The issue control logic normally receives the first and second program instructions which are kept in order; this is shown in step <b>450</b>. Then in step <b>452</b>, the issue control unit determines whether both groups of execution units are busy by monitoring the status lines from each group of execution units. It should be noted that the present invention may be used where there are more than two groups of execution units; for example, three groups of symmetrical multifunctional execution units may be used with the present invention. In step <b>454</b>, the various register dependent stalls and bypasses are checked for the high priority instruction. These register dependent stalls and bypass checks may be similar to those described and shown in FIGS. 4 and 5. In step <b>456</b>, if a group of execution units is not busy and if there are no stalls asserted for the high priority instruction, then the high priority instruction will issue to a group of execution units. Concurrently, this group of execution units will receive operands from its operand input port and will perform the instruction on the operands and provide an instruction result at the result output port of the group of execution units. Step <b>458</b> shows that the next program instruction is received from the instruction cache and the previously low priority instruction will now become the high priority instruction and processing recycles back to step <b>452</b>. It will be appreciated that, as an alternative to steps <b>454</b> and <b>456</b> shown in FIG. 6, the issue control logic may concurrently check register dependent stalls and bypasses for both the high and low priority instructions and then issue concurrently both the high and low priority instructions and then receive the next two program instructions and then recycle back to step <b>452</b>.
The present invention has been described in the context of several examples which have assumed certain specific architectures. It will be appreciated that the present invention may be employed in other architectures. For example, more than two groups of execution units each being symmetrical and multifunctional may be employed with the present invention. Moreover, the groups of execution units may share a shared resource such as a data cache. The present invention will allow simpler control logic to be used to control the issuance of instructions to the groups of execution units. There will be no need to decode program instructions for the purpose of determining the functions or processing required by each program instruction. Thus, instruction steering is eliminated as a requirement since each group of execution units provides identical functionality such that program instructions may be directed to any of the groups for execution. This design also simplifies other control operations as described above. While the foregoing invention has been described with respect to the above examples, it will be appreciated that the scope of the invention is limited only by the scope of the following claims.
Contents4
7 sheets
Sheet 1 Sheet 2 Sheet 3 Sheet 4 Sheet 5 Sheet 6 Sheet 7
Every citation, both ways
| Document | Relation | Office | Cited during |
|---|---|---|---|
| US7730284B2 | Cited by | United States of America | Search report |
| US2007089102A1 | Cited by | United States of America | Pre-grant |
| US7647513B2 | Cited by | United States of America | Applicant |
| US6845456B1 | Cited by | United States of America | Search report |
| US2006004990A1 | Cited by | United States of America | Pre-grant |
| US6889332B2 | Cited by | United States of America | Applicant |
| US7254721B1 | Cited by | United States of America | Applicant |
| US2006212686A1 | Cited by | United States of America | Pre-grant |
| US7840783B1 | Cited by | United States of America | Applicant |
| US2007283176A1 | Cited by | United States of America | Pre-grant |
| US7441106B2 | Cited by | United States of America | Applicant |
| US3943494A | Cites | United States of America | Search report |
| US4442484A | Cites | United States of America | Applicant |
| US5450607A | Cites | United States of America | Applicant |
| US5467476A | Cites | United States of America | Applicant |
| US5530816A | Cites | United States of America | Applicant |
| US5546597A | Cites | United States of America | Applicant |
| US5559976A | Cites | United States of America | Applicant |
| US5564056A | Cites | United States of America | Applicant |
| US5574928A | Cites | United States of America | Applicant |
| US5574942A | Cites | United States of America | Applicant |
| US5613080A | Cites | United States of America | Applicant |
| US5628021A | Cites | United States of America | Applicant |
| US5664136A | Cites | United States of America | Applicant |
| US5671382A | Cites | United States of America | Applicant |
| US5689720A | Cites | United States of America | Applicant |
| US5742791A | Cites | United States of America | Search report |
| US5790827A | Cites | United States of America | Applicant |
| US5898849A | Cites | United States of America | Search report |
| US5922068A | Cites | United States of America | Applicant |
| US5974522A | Cites | United States of America | Search report |
| US6035388A | Cites | United States of America | Search report |
| Meriam-Webster, "Meriam-Webster's Collegiate dictionary", 1997, Meriam-Webster, Inc., 10th ed., pp. 544, and 1144.* | Non-patent | – | Search report |
| Bauman et al, UltraSparc: The Next Generation Superscalar 64-Bit Sparc, IEEE Computer Society Press, pp. 1-10. | Non-patent | – | Applicant |
| Blanck et al, The Super Sparc Microprocessor, IEEE Computer Society Press, pp. 1-6. | Non-patent | – | Applicant |
3 members in 1 office
Priority claims6
| Document | Office | Kind | Date |
|---|---|---|---|
| 88314797 | United States of America | A | |
| 88314797 | United States of America | A | |
| 51952400 | United States of America | A | |
| 08883147 | – | – | – |
| US19970883147 | – | – | – |
| US20000519524 | – | – | – |
Members3
| Document | Office | Kind | |
|---|---|---|---|
| US6035388A | United States of America | A | |
| US2002199084A1 | United States of America | A1 | |
| US6594753B2This record | United States of America | B2 |
62 transactions on the USPTO file
Allowed after 3 non-final rejections and 2 final rejections.
- Non-final rejections
- 3
- Final rejections
- 2
- RCEs
- 0
- Appeals
- 0
Over time
Point at a mark for the transactionTransactions
| Event | Code | |
|---|---|---|
| Expire PatentEXP. | EXP. | |
| Correspondence Address ChangeC.ADB | C.ADB | |
| Change in Power of Attorney (May Include Associate POA)PA.. | PA.. | |
| Correspondence Address ChangeC.AD | C.AD | |
| Correspondence Address ChangeC.AD | C.AD | |
| Change in Power of Attorney (May Include Associate POA)PA.. | PA.. | |
| Correspondence Address ChangeC.AD | C.AD | |
| Correspondence Address ChangeC.AD | C.AD | |
| Recordation of Patent Grant MailedPGM/ | PGM/ | |
| Patent Issue Date Used in PTA CalculationAllowedPTAC | PTAC | |
| Issue Notification MailedAllowedWPIR | WPIR | |
| Receipt into PubsR1021 | R1021 | |
| Application Is Considered Ready for IssuePILS | PILS | |
| Issue Fee Payment VerifiedN084 | N084 | |
| Issue Fee Payment ReceivedIFEE | IFEE | |
| Receipt into PubsR1021 | R1021 | |
| Workflow - File Sent to ContractorSENT | SENT | |
| Receipt into PubsR1021 | R1021 | |
| Dispatch to PublicationsD1220 | D1220 | |
| Mail Notice of AllowanceAllowedMN/=. | MN/=. | |
| Notice of Allowance Data Verification CompletedAllowedN/=. | N/=. | |
| Case Docketed to Examiner in GAUDOCK | DOCK | |
| Date Forwarded to ExaminerFWDX | FWDX | |
| Response after Non-Final ActionA... | A... | |
| Mail Non-Final RejectionNon-final rejectionMCTNF | MCTNF | |
| Non-Final RejectionNon-final rejectionCTNF | CTNF | |
| Date Forwarded to ExaminerFWDX | FWDX | |
| Disposal for a RCE / CPA / R129AbandonedABN9 | ABN9 | |
| Continuing Prosecution Application - Continuation (ACPA)ACPA | ACPA | |
| Workflow - Request for CPA - BeginBCPA | BCPA | |
| Mail Advisory Action (PTOL - 303)MCTAV | MCTAV | |
| Advisory Action (PTOL-303)CTAV | CTAV | |
| Date Forwarded to ExaminerFWDX | FWDX | |
| Response after Final ActionA.NE | A.NE | |
| Workflow - Drawings FinishedDRWF | DRWF | |
| Mail Final Rejection (PTOL - 326)Final rejectionMCTFR | MCTFR | |
| Final RejectionFinal rejectionCTFR | CTFR | |
| Workflow - Drawings Matched with File at ContractorDRWM | DRWM | |
| New or Additional Drawing FiledC614 | C614 | |
| Date Forwarded to ExaminerFWDX | FWDX | |
| Response after Non-Final ActionA... | A... | |
| Mail Non-Final RejectionNon-final rejectionMCTNF | MCTNF | |
| Non-Final RejectionNon-final rejectionCTNF | CTNF | |
| Date Forwarded to ExaminerFWDX | FWDX | |
| Response after Final ActionA.NE | A.NE | |
| Mail Final Rejection (PTOL - 326)Final rejectionMCTFR | MCTFR | |
| Final RejectionFinal rejectionCTFR | CTFR | |
| Date Forwarded to ExaminerFWDX | FWDX | |
| Response after Non-Final ActionA... | A... | |
| Request for Extension of Time - GrantedXT/G | XT/G | |
| Information Disclosure Statement (IDS) FiledM844 | M844 | |
| Information Disclosure Statement (IDS) FiledWIDS | WIDS | |
| Affidavit(s) (Rule 131 or 132) or Exhibit(s) ReceivedAF/D | AF/D | |
| Case Docketed to Examiner in GAUDOCK | DOCK | |
| Mail Non-Final RejectionNon-final rejectionMCTNF | MCTNF | |
| Non-Final RejectionNon-final rejectionCTNF | CTNF | |
| Preliminary AmendmentA.PE | A.PE | |
| Case Docketed to Examiner in GAUDOCK | DOCK | |
| Application Dispatched from OIPEOIPE | OIPE | |
| Correspondence Address ChangeC.AD | C.AD | |
| IFW Scan & PACR Auto Security ReviewSCAN | SCAN | |
| Initial Exam Team nnIEXX | IEXX |
25 legal events, as the office reported them to INPADOC
Over the term
Point at a mark for the eventEvents
| Event | Code | |
|---|---|---|
| AssignmentAS | AS | |
| AssignmentAS | AS | |
| AssignmentAS | AS | |
| AssignmentAS | AS | |
| AssignmentAS | AS | |
| Lapsed due to failure to pay maintenance feeLapsedFP | FP | |
| Information on status: patent discontinuationPATENT EXPIRED DUE TO NONPAYMENT OF MAINTENANCE FEES UNDER 37 CFR 1.362STCH | STCH | |
| Lapse for failure to pay maintenance feesLapsedLAPS | LAPS | |
| AssignmentAS | AS | |
| AssignmentAS | AS | |
| Maintenance fee reminder mailedREMI | REMI | |
| AssignmentAS | AS | |
| Fee paymentFPAY | FPAY | |
| AssignmentAS | AS | |
| AssignmentAS | AS | |
| Fee payment procedurePAYOR NUMBER ASSIGNED (ORIGINAL EVENT CODE: ASPN); ENTITY STATUS OF PATENT OWNER: LARGE ENTITYFEPP | FEPP | |
| Fee payment procedurePAYER NUMBER DE-ASSIGNED (ORIGINAL EVENT CODE: RMPN); ENTITY STATUS OF PATENT OWNER: LARGE ENTITYFEPP | FEPP | |
| AssignmentAS | AS | |
| Fee payment procedurePAYOR NUMBER ASSIGNED (ORIGINAL EVENT CODE: ASPN); ENTITY STATUS OF PATENT OWNER: LARGE ENTITYFEPP | FEPP | |
| AssignmentAS | AS | |
| Fee paymentFPAY | FPAY | |
| Surcharge for late paymentSULP | SULP | |
| Maintenance fee reminder mailedREMI | REMI | |
| AssignmentAS | AS | |
| AssignmentAS | AS |
Numbers
- Publication, DOCDB
- 6594753
- Publication, EPODOC
- US6594753
- Application
- 9519524
- Application, DOCDB
- 51952400
- Application, EPODOC
- US20000519524
Titles
- English
- Method and apparatus for dual issue of program instructions to symmetric multifunctional execution units
Patent term adjustment
- Applicant delay
- −4 days
- Net adjustment
- 0 days
Classification
- CPC, 3
- G06F9/3891
- G06F9/3836
- G06F9/3885
- IPC, 1
- G06F9 38
- USPC, 4
- 712214000
- 712216000
- 712E09049
- 712E09071