Backward-compatible computer architecture with extended word size and address space
Abstract
This record has no abstract on file.
Term
Term ended
Expired 11 March 2012, 14.5 years ago.
- Priority
- Filed
- Granted
- Expired
- Today
17 claims: 15 independent, 2 dependent
- 1In a processing unit of a computer system consisting of a processing unit and a memory subsystem, a means for identifying an operation corresponding to an instruction fetched from the memory subsystem, a large number of m-bit load instructions, a large number of logical operation instructions, and the like. It has a large number of m-bit shift instructions and a large number of m-bit addition instructions, and a first subset instruction called an m-bit instruction, a large number of N-bit load instructions, a large number of N-bit shift instructions, and a large number of N-bit addition instructions. It has a second subset of instructions called N-bit instructions, a load circuit that performs the specified load operation to search for an entity with a maximum N-bit length from the memory subsystem, and the searched m-bits. To perform controllable load code extension means to code-extend an entity to N bits, a register file with a set of N-bit registers, and a logical operation specified between two N-bit operands. The logic circuit of, the shift circuit that performs the shift operation specified for the N-bit operand, the shift code extension means that can control the shift operation result to code-extend from m bits to N bits, and a pair of N bits. An adder for performing an addition operation specified between the operands, an addition code extension means capable of controlling operation to extend the operation result in the adder from m bits to N bits, and m in the adder. Means for detecting bit overflow, means for defining an N-bit data path from each output port of the load circuit, the logic circuit, the shift circuit, and the adder to the register file, and the logic from the register file. At least one means of defining an N-bit data path to each input port of the circuit, said shift circuit and said adder, and means of activating the load code extension means in response to at least one m-bit load instruction. A means for activating the shift code extension means in response to one m-bit shift instruction and the addition code extension means in response to the occurrence of at least one m-bit addition instruction and m-bit overflow.The means for activating the load code extension means, the means for activating the shift code extension means, and the means for activating the additive code extension means are N-bit loads. , A processing unit that does not respond to shift and add instructions. 処理ユニットとメモリサブシステムとからなるコンピュータシステムの処理ユニットにおいて、前記メモリサブシステムからフェッチされた命令に応答し対応する動作を特定する手段と、多数のmビットロード命令、多数の論理演算命令、多数のmビットシフト命令及び多数のmビット加算命令を有しておりmビット命令と呼ばれる第一サブセットの命令並びに多数のNビットロード命令、多数のNビットシフト命令、多数のNビット加算命令を有しておりNビット命令と呼ばれる第二サブセットの命令と、前記メモリサブシステムから最大でNビット長のエンティティを検索するために指定されたロード動作を実施するロード回路と、検索したmビットのエンティティをNビットへ符号拡張するために制御動作可能なロード符号拡張手段と、一組のNビットレジスタを持ったレジスタファイルと、二つのNビットオペランドの間で特定された論理演算を実施するための論理回路と、Nビットオペランドに関して特定されたシフト動作を実施するシフト回路と、シフト動作の結果をmビットからNビットへ符号拡張すべく制御動作可能なシフト符号拡張手段と、一対のNビットオペランドの間で特定された加算演算を実施するための加算器と、前記加算器における演算結果をmビットからNビットへ符号拡張すべく制御動作可能な加算符号拡張手段と、前記加算器におけるmビットオーバーフローを検知する手段と、前記ロード回路、前記論理回路、前記シフト回路及び前記加算器の各々の出力ポートから前記レジスタファイルへのNビットデータ経路を画定する手段と、前記レジスタファイルから前記論理回路、前記シフト回路及び前記加算器の各々の入力ポートへNビットデータ経路を画定する手段と、少なくとも一つのmビットロード命令に応答して前記ロード符号拡張手段を活性化させる手段と、少なくとも一つのmビットシフト命令に応答して前記シフト符号拡張手段を活性化させる手段と、少なくとも一つのmビット加算命令及びmビットオーバーフローの発生に応答して前記加算符号拡張手段を活性化させる手段と、を有しており、前記ロード符号拡張手段を活性化させる手段、前記シフト符号拡張手段を活性化させる手段、及び前記加算符号拡張手段を活性化させる手段がNビットのロード、シフト及び加算命令に応答するものでないことを特徴とする処理ユニット。
- 3In claim 1, further sign extension or zero of the searched entries less than m bits associated with the load circuit and the load sign extension means to define the searched entity of m bits. A processing unit characterized by being provided with means for expansion. 請求項1において、更に、前記ロード回路及び前記ロード符号拡張手段と関連しており更にmビットの検索されたエンティティを画定するためにmビット未満の検索されたエントリをmビットへ符号拡張又はゼロ拡張させる手段が設けられていることを特徴とする処理ユニット。
- 5In claim 1, the m-bit addition instruction has ADD and SUD instructions that are trapped by m-bit overflow and do not generate a result, and are not trapped by m-bit overflow and generate a result. A processing unit having ADDU and SUBU instructions, and the means for activating the addition sign extension means responds to the ADDU and SUBU instructions and does not respond to the ADD and SUBU instructions. 請求項1において、前記mビット加算命令がmビットオーバーフローでトラップし且つ結果を発生させることのないADD及びSUD命令を有しており、且つmビットオーバーフローでトラップすることがなく且つ結果を発生するADDU及びSUBU命令を有しており、且つ前記加算符号拡張手段を活性化させる手段が前記ADDU及びSUBU命令に応答し且つ前記ADD及びSUB命令に応答することがないことを特徴とする処理ユニット。
- 6In claim 1, further, the DADD and DSUB instructions in which the N-bit overflow detecting means is provided in the adder and the N-bit addition instruction is trapped by the N-bit overflow and does not generate a result. A processing unit that has DADDU and DSUBU instructions that generate results without being trapped by N-bit overflow. 請求項1において、更に、前記加算器内にNビットオーバーフロー検知手段が設けられており、且つ前記Nビット加算命令が、Nビットオーバーフローでトラップし且つ結果を発生することのないDADD及びDSUB命令を有しており且つNビットオーバーフローでトラップすることがなく且つ結果を発生するDADDU及びDSUBU命令を有していることを特徴とする処理ユニット。
- 7In the adder in the execution unit of the data processor, which can operate to combine two N-bit operands, the maximum digit bit is specified as bit (N-1) and the minimum digit bit is specified as bit (0), N A means of adding a pair of input operands in response to a first operation code named bit addition and supplying the addition result, adding a pair of input operands in response to a second operation code named m-bit addition and Initial result of bits (N-1) to bits (m) before feeding the result, at least if the bit (m) of the initial result is not equal to the bit (m-1) of the initial result. An adder, characterized in that it has a means of setting to the value of bits (m-1) of. 二つのNビットオペランドを結合すべく動作可能であり最大桁ビットはビット(N-1)と指定され且つ最小桁ビットがビット(0)として指定されるデータプロセサの実行ユニットにおける加算器において、Nビット加算と命名された第一オペコードに応答して一対の入力オペランドを加算し且つその加算結果を供給する手段、mビット加算と命名された第二オペコードに応答し一対の入力オペランドを加算し且つ少なくとも初期的な結果のビット(m)が初期的な結果のビット(m-1)に等しくない場合にはその結果を供給する前にビット(N-1)乃至ビット(m)を初期的結果のビット(m-1)の値にセットする手段、を有していることを特徴とする加算器。
- 8N-bit logic in the shift unit in the execution unit of the data processor that can operate to shift the N-bit operand, the maximum digit bit is specified as bit (N-1) and the minimum digit bit is specified as bit (0). A means of shifting the input operand to the right by a specified number of bit positions in response to the first operation code named right shift, and filling the empty bit positions with 0 if the specified number is not zero. In response to the second operation code named m-bit logical right shift, the input operand is shifted to the right by the specified number of bit positions, and if the specified number is not zero, the empty bit position and the next A shift unit characterized by having a means for filling the Nm maximum digit bit positions of the above with 0. Nビットオペランドをシフトすべく動作可能であり最大桁ビットはビット(N-1)として指定され且つ最小桁ビットはビット(0)として指定されるデータプロセサの実行ユニットにおけるシフトユニットにおいて、Nビット論理右シフトと命名された第一オペコードに応答して入力オペランドを指定した数のビット位置だけ右側へシフトさせ且つその指定した数がゼロでない場合には空きとされたビット位置を0で充填させる手段、mビット論理右シフトと命名された第二オペコードに応答して入力オペランドを指定した数のビット位置だけ右側へシフトさせ且つその指定した数がゼロでない場合には空きとされたビット位置及び次のN-m個の最大桁ビット位置を0で充填させる手段、を有していることを特徴とするシフトユニット。
- 9It is operational to shift the N-bit operand, the maximum digit bit is specified as bit (N-1) and the minimum digit bit is specified as bit (0) in the shift unit in the execution unit of the data processor. In response to the first operation code named as left shift, (1) the input operand is shifted to the left by the specified number of bit positions, and if the specified number is non-zero, it is (2) empty. A means of filling bit positions with 0, in response to a second operation code named m-bit left shift (1) shifts the input operand to the left by a specified number of bit positions and the specified number is zero. If not, (2) fill the empty bit positions with 0 and (3) set the resulting bits (N-1) to bits (m) to the resulting bit (m-1) value. A shift unit characterized by having means to make it. Nビットオペランドをシフトすべく動作可能であり最大桁ビットはビット(N-1)として指定され且つその最小桁ビットはビット(0)として指定されるデータプロセサの実行ユニットにおけるシフトユニットにおいて、Nビット左シフトとして命名された第一オペコードに応答して(1)入力オペランドを指定された数のビット位置だけ左側へシフトさせ且つその指定された数がゼロでない場合には(2)空きとされたビット位置を0で充填させる手段、mビット左シフトと命名された第二オペコードに応答して(1)入力オペランドを指定された数のビット位置だけ左側へシフトさせ且つその指定された数がゼロでない場合には(2)空きとされたビット位置を0で充填させ且つ(3)その結果のビット(N-1)乃至ビット(m)をその結果のビット(m-1)の値へセットさせる手段、を有していることを特徴とするシフトユニット。
- 10In claim 1, the shift circuit, the shift code extension means, and the means for activating the shift code extension means all provide an N-bit operand in response to a first operation code designated as an N-bit logical right shift. A means of shifting the bit position of the specified number S to the right and filling the bit position of the bit (N-1) to the bit (NS) with 0 when the specified number is not zero, with m-bit logical right shift. In response to the specified second operation code, the N-bit operand is shifted to the right by the specified number of S bit positions, and if the specified number is not zero, the bits of bits (N-1) to (mS). Means to fill the position with 0, in response to the third operation code specified as N-bit left shift (1) Shift the N-bit operand to the left by the specified number of S bit positions, and the specified number is not zero. In some cases, (2) means to fill the bit positions of bits (S-1) to (0) with 0, and (1) specify the N-bit operand in response to the fourth operation code specified as m-bit left shift. If only the bit position of the number S is shifted to the left and the specified number is not zero, (2) the bit positions of bits (S-1) to (0) are filled with 0 to give the initial result. A processing unit characterized in that (3) means for setting bits (N-1) to bits (m) to the value of the initial result bit (m-1). 請求項1において、前記シフト回路、前記シフト符号拡張手段、及び前記シフト符号拡張手段を活性化させる手段が、共に、Nビット論理右シフトと指定される第1オペコードに応答してNビットオペランドを指定した数Sのビット位置だけ右へシフトさせ且つその指定した数がゼロでない場合にはビット(N-1)乃至ビット(N-S)のビット位置を0で充填する手段、mビット論理右シフトと指定される第2オペコードに応答してNビットオペランドを指定した数Sのビット位置だけ右へシフトさせ且つその指定した数がゼロでない場合にはビット(N-1)乃至ビット(m-S)のビット位置を0で充填する手段、Nビット左シフトと指定される第3オペコードに応答して(1)Nビットオペランドを指定した数Sのビット位置だけ左へシフトさせ且つその指定した数がゼロでない場合には(2)ビット(S-1)乃至ビット(0)のビット位置を0で充填する手段、mビット左シフトと指定される第4オペコードに応答して(1)Nビットオペランドを指定した数Sのビット位置だけ左へシフトさせ且つその指定した数がゼロでない場合には(2)初期的結果を与えるためにビット(S-1)乃至ビット(0)のビット位置を0で充填し且つ(3)ビット(N-1)乃至ビット(m)を前記初期的結果のビット(m-1)の値へ設定する手段、を有していることを特徴とする処理ユニット。
- 11In the processing unit in the computer system consisting of the processing unit and the memory subsystem, a decoder that identifies the corresponding operation in response to the instruction fetched from the memory subsystem is provided, and the operation can be performed. A bit entity has a maximum digit bit designated as bit (m-1) and a minimum digit bit designated as bit (0), and an N-bit entity capable of performing an operation is a bit (N). It has a maximum digit bit designated as -1) and a minimum digit bit designated as bit (0), and the instructions are the first subset of instructions called m-bit instructions and the N-bit instruction. Containing a second subset of instructions called, where N is greater than m, the m-bit instructions include a large number of m-bit load instructions, a large number of logical operation instructions, and a large number of operations. It includes an m-bit shift instruction and a large number of m-bit addition instructions, and the N-bit instruction includes a large number of N-bit load instructions, a large number of N-bit shift instructions, and a large number of N-bit addition instructions. A load circuit is provided to perform a specified load operation to search for an entity with a maximum N-bit length from the memory subsystem, and the load circuit is provided with bits (m-) of the m-bit entity. N-bit result bits (N-1) to bits (Nm) equal to 1) and N-bit result bits (m-1) to bits (0) equal to m-bit entity bits (m-1) to bits (0) It has a set of related load code extension circuits that can control and operate in response to at least one m-bit load instruction to convert the retrieved m-bit entity to an N-bit result that has 0). A register file having the N-bit register of is provided, a logic circuit is provided to perform a specified logical operation between two N-bit operands and give a result, and a shift operation specified with respect to the N-bit operand is performed. A shift circuit is provided to perform and give the result, and the result of the shift operation is a bit (m) of the result.A shift code extension circuit that can be controlled and operated in response to at least one m-bit shift instruction is provided to change to an N-bit entity having bits (N-1) to bits (Nm) equal to -1). , An adder is provided that performs and gives a result of a specific adder operation between a pair of N-bit operands, which is from the bit position of bit (m-1) and bit (m-2). A condition in which a carry-out is provided and the carry-out from the bit position of the bit (m-1) and the bit (m-2) is different is called an m-bit overflow, and at least one m-bit adder instruction is executed and the carry-out is executed. Controllable operation is possible when an m-bit overflow occurs and the operation result in the adder is changed to an N-bit entity having bits (N-1) to bits (Nm) equal to the bit (m-1) of the result. An adder code extension circuit is provided, and an N-bit data path from each output port of the load circuit, the logic circuit, the shift circuit, and the adder to the register file is provided, and the register file is provided. An N-bit data path is provided from each input port of the logic circuit, the shift circuit, and the adder, and the load code extension circuit, the shift code extension circuit, and the adder code extension circuit are N bits. A processing unit that does not respond to load instructions, N-bit shift instructions, and N-bit adder instructions.M-bit addition instruction is executed and m-bit overflow occurs, and the operation result in the adder is an N-bit entity having bits (N-1) to bits (Nm) equal to the bit (m-1) of the result. An addition code extension circuit that can be controlled and operated when changing to is provided, and an N-bit data path from each output port of the load circuit, the logic circuit, the shift circuit, and the adder to the register file. Is provided, and an N-bit data path is provided from the register file to each input port of the logic circuit, the shift circuit, and the adder, and the load code extension circuit, the shift code extension circuit, and the like. A processing unit characterized in that the addition code extension circuit does not respond to an N-bit load instruction, an N-bit shift instruction, and an N-bit addition instruction.M-bit addition instruction is executed and m-bit overflow occurs, and the operation result in the adder is an N-bit entity having bits (N-1) to bits (Nm) equal to the bit (m-1) of the result. An addition code extension circuit that can be controlled and operated when changing to is provided, and an N-bit data path from each output port of the load circuit, the logic circuit, the shift circuit, and the adder to the register file. Is provided, and an N-bit data path is provided from the register file to each input port of the logic circuit, the shift circuit, and the adder, and the load code extension circuit, the shift code extension circuit, and the like. A processing unit characterized in that the addition code extension circuit does not respond to an N-bit load instruction, an N-bit shift instruction, and an N-bit addition instruction. 処理ユニットとメモリサブシステムとからなるコンピュータシステムにおける処理ユニットにおいて、メモリサブシステムからフェッチされた命令に応答して対応する動作を特定するデコーダが設けられており、動作を実施することが可能なmビットエンティティはビット(m-1)として指定される最大桁ビットとビット(0)として指定される最小桁ビットとを有しており且つ動作を実施することが可能なNビットエンティティはビット(N-1)として指定される最大桁ビットとビット(0)として指定される最小桁ビットとを有しており、前記命令は、mビット命令と呼称される命令の第一サブセットと、Nビット命令と呼称される命令の第二サブセットとを包含しており、尚Nはmよりも大きいものであり、前記mビット命令は、多数のmビットロード命令と、多数の論理演算命令と、多数のmビットシフト命令と、多数のmビット加算命令とを包含しており、前記Nビット命令は、多数のNビットロード命令と、多数のNビットシフト命令と、多数のNビット加算命令とを包含しており、前記メモリサブシステムから最大でNビット長のエンティティを検索するために指定されたロード動作を実施するロード回路が設けられており、前記ロード回路は、mビットエンティティのビット(m-1)に等しいNビット結果のビット(N-1)乃至ビット(N-m)及びmビットエンティティのビット(m-1)乃至ビット(0)に等しいNビット結果のビット(m-1)乃至ビット(0)を有しているNビット結果へ検索したmビットエンティティを変換させるために少なくとも一つのmビットロード命令に応答して制御動作可能な関連するロード符号拡張回路を有しており、一組のNビットレジスタを有するレジスタファイルが設けられており、二つのNビットオペランドの間で特定した論理演算を実施し且つ結果を与える論理回路が設けられており、Nビットオペランドに関して特定したシフト動作を実施し且つ結果を与えるシフト回路が設けられており、シフト動作の結果を該結果のビット(m-1)に等しいビット(N-1)乃至ビット(N-m)を有するNビットエンティティへ変化させるために少なくとも一つのmビットシフト命令に応答して制御動作可能なシフト符号拡張回路が設けられており、一対のNビットオペランドの間で特定した加算演算を実施し且つ結果を与える加算器が設けられており、前記加算器はビット(m-1)及びビット(m-2)のビット位置からのキャリーアウトを有しており、ビット(m-1)及びビット(m-2)のビット位置からのキャリーアウトが異なる条件をmビットオーバーフローと呼称し、少なくとも一つのmビット加算命令が実行され且つmビットオーバーフローが発生して前記加算器における演算結果を前記結果のビット(m-1)に等しいビット(N-1)乃至ビット(N-m)を有するNビットエンティティへ変化させる場合に制御動作可能な加算符号拡張回路が設けられており、前記ロード回路、前記論理回路、前記シフト回路、及び前記加算器の各々の出力ポートから前記レジスタファイルへのNビットデータ経路が設けられており、前記レジスタファイルから前記論理回路、前記シフト回路、及び前記加算器の各々の入力ポートへNビットデータ経路が設けられており、前記ロード符号拡張回路、前記シフト符号拡張回路、及び前記加算符号拡張回路がNビットロード命令、Nビットシフト命令、及びNビット加算命令に応答するものでない、ことを特徴とする処理ユニット。
- 12Claim 11A processing unit characterized in that N = 64 and m = 32. 請求項11において、N=64及びm=32であることを特徴とする処理ユニット。
- 13Claim 11The processing unit is characterized in that the additive sign extension circuit operates only when an m-bit overflow occurs. 請求項11において、前記加算符号拡張回路は、mビットオーバーフローが発生した場合にのみ動作することを特徴とする処理ユニット。
- 14Claim 11In the above, the m-bit addition instruction includes an ADD and SUB instruction that is trapped by an m-bit overflow and does not generate a result, and an ADDU and SUBU instruction that is not trapped by an m-bit overflow and produces a result. The processing unit is characterized in that the means for activating the additive sign extension circuit responds to the ADDU and SUBU instructions, but does not respond to the ADD and SUB instructions. 請求項11において、前記mビット加算命令は、mビットオーバーフローでトラップし且つ結果を発生することのないADD及びSUB命令と、mビットオーバーフローでトラップすることが無く且つ結果を発生するADDU及びSUBU命令とを包含しており、前記加算符号拡張回路を活性化させる手段が前記ADDU及びSUBU命令に応答するが前記ADD及びSUB命令には応答するものではない、ことを特徴とする処理ユニット。
- 15Claim 11Further, a means for detecting an N-bit overflow in the adder is provided, and the N-bit adder instruction is a DADD and DSUB instruction that is trapped by the N-bit overflow and does not generate a result, and N. A processing unit characterized by including DADDU and DSUBU instructions that do not trap due to bit overflow and generate a result. 請求項11において、更に、前記加算器におけるNビットオーバーフローを検知する手段が設けられており、且つ前記Nビット加算命令が、Nビットオーバーフローでトラップし且つ結果を発生することの無いDADD及びDSUB命令と、Nビットオーバーフローでトラップすることが無く且つ結果を発生するDADDU及びDSUBU命令とを包含していることを特徴とする処理ユニット。
- 16In a shift unit in a processing unit of a computer system that performs the shift operation specified with respect to the N-bit operand and gives an N-bit result, the processing unit includes a first subset of instructions called an m-bit shift instruction and an N-bit. It executes instructions that include a second subset of instructions, called shift instructions, and each of the input operands and results is specified as the maximum digit bit and bit (0), which are specified as bits (N-1). N-bit input, which receives the N-bit entity to be shifted, N-bit output which gives the N-bit entity generated from the shift operation, and N-bit output for the shift operation. A control input that specifies the bit position and direction of the specified number S, a first shift input terminal that receives a value for replacing the bit (N-1) for the right shift, and a bit (0) for the left shift. An N-bit shifter provided with a second input terminal that receives a value to be replaced is provided, and a first unit that is coupled to receive bits (N-1) to bits (m) of the N-bit operand. It comprises a (Nm) bit input, a second input that receives (Nm) 0s, and an output that is coupled to the bit positions of the input bits (N-1) to (m) of the shifter. The first data selector is provided, and the first data selector passes bits (N-1) to bits (m) of the N-bit operand to the N-bit shift instruction, and the designation When the number of bit positions is not zero, the (Nm) number of 0s are passed for the instruction specified as m-bit logical right shift, and the input bits (m-1) to bits of the shifter are passed. The bit positions of (0) are combined to receive the bits (m-1) to bits (0) of the N-bit operand, and are combined to receive the bit (N-1) of the N-bit input operand. A second data selector that includes a first input, a second input that receives 0, and an output that is coupled to the first shift input of the shifter.The second data selector is provided so that the bit (N-1) of the N-bit input operand is passed through the instruction designated as m-bit and N-bit operation right shift, and the m-bit and N-bit are provided. The 0 is passed for an instruction designated as a logical right shift, and the second shift input is coupled to receive 0, and the output bits (N-1) to bits (m) of the shifter are combined. The first input coupled to the bit position, the second input coupled to receive (Nm) copies of the output bits (m-1) of the shifter, and the N-bit result bits ( A third data selector is provided that includes an output that gives N-1) to bits (m), and the third data selector is a bit of the output of the shifter in response to the N-bit shift instruction. The (Nm) copies of the output bits (m-1) of the shifter for an instruction that passes through the bit positions of (N-1) to bits (Nm) and is designated as m-bit left shift. The bit positions of the output bits (m-1) to bits (0) of the shifter are combined to give the N-bit result bits (m-1) to bits (0). A shift unit featuring.Passing through the bit position of -m) and passing the (Nm) copies of the output bits (m-1) of the shifter for an instruction designated as m-bit left shift of the shifter. A shift unit characterized in that the bit positions of the output bits (m-1) to bits (0) are combined to give the N-bit result bits (m-1) to bits (0).Passing through the bit position of -m) and passing the (Nm) copies of the output bits (m-1) of the shifter for an instruction designated as m-bit left shift of the shifter. A shift unit characterized in that the bit positions of the output bits (m-1) to bits (0) are combined to give the N-bit result bits (m-1) to bits (0). Nビットオペランドに関して特定されたシフト動作を実施し且つNビット結果を与えるコンピュータシステムの処理ユニットにおけるシフトユニットにおいて、前記処理ユニットは、mビットシフト命令と呼称される第一サブセットの命令と、Nビットシフト命令と呼称される第二サブセットの命令とを包含している命令を実行し、入力オペランド及び結果の各々はビット(N-1)として指定される最大桁ビットとビット(0)として指定される最小桁ビットとを有しており、Nはmよりも大きく、シフトされるべきNビットエンティティを受取るNビット入力と、シフト動作から発生するNビットエンティティを与えるNビット出力と、シフト動作に対する指定した数Sのビット位置及び方向を特定する制御入力と、右シフト用にビット(N-1)を置換させるための値を受取る第1シフト入力端子と、左シフト用にビット(0)を置換させるための値を受取る第2入力端子とを具備しているNビットシフタが設けられており、前記Nビットオペランドのビット(N-1)乃至ビット(m)を受取るべく結合されている第1(N-m)ビット入力と、(N-m)個の0を受取る第2入力と、前記シフタの前記入力のビット(N-1)乃至ビット(m)のビット位置へ結合されている出力とを具備している第1データセレクタが設けられており、前記第1データセレクタは、前記Nビットシフト命令に対して前記Nビットオペランドのビット(N-1)乃至ビット(m)を通過させ、且つ前記指定した数のビット位置がゼロではない場合に、mビット論理右シフトと指定される命令に対して前記(N-m)個の0を通過させ、前記シフタの前記入力のビット(m-1)乃至ビット(0)のビット位置は前記Nビットオペランドのビット(m-1)乃至ビット(0)を受けとるべく結合されており、前記Nビット入力オペランドのビット(N-1)を受取るべく結合されている第1入力と、0を受取る第2入力と、前記シフタの前記第1シフト入力へ結合されている出力とを具備している第2データセレクタが設けられており、前記第2データセレクタは、mビット及びNビット演算右シフトと指定される命令に対して前記Nビット入力オペランドのビット(N-1)を通過させ、且つmビット及びNビット論理右シフトと指定される命令に対して前記0を通過させ、前記第2シフト入力は0を受取るべく結合されており、前記シフタの前記出力のビット(N-1)乃至ビット(m)のビット位置へ結合されている第1入力と、前記シフタの前記出力のビット(m-1)の(N-m)個のコピーを受取るべく結合されている第2入力と、前記Nビット結果のビット(N-1)乃至ビット(m)を与える出力とを具備している第3データセレクタが設けられており、前記第3データセレクタは、前記Nビットシフト命令に対して前記シフタの前記出力のビット(N-1)乃至ビット(N-m)のビット位置を通過させ、且つmビット左シフトと指定される命令に対して前記シフタの前記出力のビット(m-1)の前記(N-m)個のコピーを通過させ、前記シフタの前記出力のビット(m-1)乃至ビット(0)のビット位置は前記Nビット結果のビット(m-1)乃至ビット(0)を与えるべく結合されている、ことを特徴とするシフトユニット。
- 17In a shift unit that performs the shift operation specified for the N-bit operand and gives an N-bit result to the processing unit of the computer system, the processing unit is a first subset of instructions called an m-bit shift instruction and an N-bit shift instruction. Executes an instruction that includes a second subset of instructions called, the input operand and each of the results is the maximum digit bit specified as bit (N-1) and the minimum specified as bit (0). Has digit bits, where N is greater than m and is specified for the N-bit input that receives the N-bit entity to be shifted, the N-bit output that gives the N-bit entity that arises from the shift operation, and the shift operation. Replace the control input that specifies the bit position and direction of the number S, the first shift input terminal that receives the value for replacing the bit (N-1) for the right shift, and the bit (0) for the left shift. An N-bit shifter provided with a second input terminal for receiving a value to be made is provided, and (a) the N-bit to the bit positions of the input bits (N-1) to (m) of the shifter. Bits (N-1) to bits (m) and (b) of the N-bit operand for a shift instruction For an instruction specified as m-bit logical right shift when the specified number of bit positions are not zero. A means for communicating (Nm) 0s is provided, and the bit position of the input bit (m-1) to bit (0) of the shifter is the bit (m-1) to bit of the N bit operand. It is combined to receive (0), and to the first shift input of the shifter, (a) m-bit and N-bit operation The bit (N-) of the N-bit input operand for the instruction specified as right shift. 1) and (b) m-bit and N-bit A means for communicating 0 to an instruction designated as a logical right shift is provided, and the second shift input is combined to receive 0, and the N As the bit result bits (N-1) to bits (m), (a) the output bits (N-1) to bits (N-1) to the shifter for the N-bit shift instruction.Means are provided to provide (Nm) copies of the output bits (m-1) of the shifter for instructions designated as m) bit positions and (b) m bits left shift. A shift unit characterized in that the bit positions of the output bits (m-1) to bits (0) are combined to give the N-bit result bits (m-1) to bits (0). Nビットオペランドに関して特定したシフト動作を実施し且つコンピュータシステムの処理ユニットにNビット結果を与えるシフトユニットにおいて、前記処理ユニットはmビットシフト命令と呼称される命令の第一サブセットと、Nビットシフト命令と呼称される命令の第二サブセットとを包含する命令を実行し、前記入力オペランド及び前記結果の各々はビット(N-1)として指定される最大桁ビットとビット(0)として指定される最小桁ビットとを有しており、Nはmより大きく、シフトされるべきNビットエンティティを受取るNビット入力と、シフト動作から発生するNビットエンティティを与えるNビット出力と、シフト動作に対して指定した数Sのビット位置及び方向を指定する制御入力と、右シフト用にビット(N-1)を置換させるための値を受取る第1シフト入力端子と、左シフト用にビット(0)を置換させる値を受取る第2入力端子とを具備しているNビットシフタが設けられており、前記シフタの前記入力のビット(N-1)乃至ビット(m)のビット位置へ、(a)前記Nビットシフト命令のための前記Nビットオペランドのビット(N-1)乃至ビット(m)及び(b)前記指定した数のビット位置がゼロでない場合にmビット論理右シフトと指定される命令に対して(N-m)個の0を通信する手段が設けられており、前記シフタの前記入力のビット(m-1)乃至ビット(0)のビット位置が前記Nビットオペランドのビット(m-1)乃至ビット(0)を受取るべく結合されており、前記シフタの前記第1シフト入力へ、(a)mビット及びNビット演算右シフトと指定される命令に対して前記Nビット入力オペランドのビット(N-1)及び(b)mビット及びNビット論理右シフトと指定される命令に対して0を通信する手段が設けられており、前記第2シフト入力は0を受取るべく結合されており、前記Nビット結果のビット(N-1)乃至ビット(m)として、(a)前記Nビットシフト命令に対する前記シフタの前記出力のビット(N-1)乃至ビット(m)のビット位置及び(b)mビット左シフトと指定される命令に対する前記シフタの前記出力のビット(m-1)の(N-m)個のコピーを与える手段が設けられており、前記シフタの前記出力のビット(m-1)乃至ビット(0)のビット位置は前記Nビット結果のビット(m-1)乃至ビット(0)を与えるべく結合されている、ことを特徴とするシフトユニット。
Independent claims15
227 paragraphs, as filed
【0001】
[Industrial application field]
The present invention relates broadly to computer architecture, and more specifically to techniques for extending instruction set architecture to have larger word dimensions and address spaces.
【0002】
[Conventional technology]
The largest programs increase the need for their address space by about 0.5 to 1 bit per year. Soon, 32-bit address space will be inadequate for such programs, just as 16-bit addressing was inadequate in the 1970s. One example of such a phenomenon is the IBM / 370, which adopted 31-bit addressing because the 24-bit addressing on the IBM / 360 turned out to be inadequate.
【0003】
Computer manufacturers tend to move to a larger address space in new computers by moving to new instruction sets. Moving to a new instruction set can have potentially fatal consequences for both users and manufacturers. From the user's point of view, this means that programs written for the old machine will not work on the new machine. Users who have made significant investments in software are faced with the unpleasant decision to either pay for the conversion or replacement of the software or give up the benefits of the various advanced ones built into the new machine. .. From the manufacturer's point of view, such a transition is likely to buy user resentment and stall the initial sales of such new machines.
【0004】
One designer uses a technique called segmentation to extend instruction sets. Many 32-bit architectures have already offered segmented configurations to extend the address space. Examples are IBM / 370 (ESA mode), IBM Power and HP Precision.
【0005】
Segmentation is widely used, for example in Intel 80286 microprocessors, but it is not always satisfactory. With a few exceptions, such as the Multics system on DE-Honeywell hardware, segmentation turned out to be inefficient and visible to the programmer. A full address becomes a multiword object, which requires multiple instructions and / or cycles to perform access and computation. Moreover, most segmentation methods do not allow access to a single data object that is larger than the segment size, which is one of the main uses for larger address spaces. ..
【0006】
Another technique used in the past is to bring both traditional (eg, 16-bit) architecture and new (eg, 32-bit) architecture on the same hardware. For example, the DEC VAX 11/780 had a mode in which it was possible to run PDP-11 programs, with some restrictions. This technique is primarily applicable in the case of microcoded execution, in which case traditional architecture is merely executed as additional microcode. In the case of hard-wired configurations, designers are basically forced to use two CPUs, or at least handle more complex ones, the complexity of which significantly impacts performance. In some cases. Both approaches are likely to consume significant chip die area, which is an important consideration when configuring a single chip.
【0007】
In any case, compatibility mode, or compatibility mode, can be costly, especially when traditional and new architectures differ in aspects beyond address dimensions. This is an extra baggage that I've been carrying in compatibility mode for a while but throwing it away later. For example, later versions of the VAX family do not support PDP-11 emulation.
【0008】
[Means for solving problems]
The present invention provides an efficient technique for extending an existing architecture in such a manner that the hardware for the extended architecture also supports the existing architecture. Backward compatibility, or inverse compatibility, requires a minimal amount of additional hardware and has minimal impact on operating speed. Moreover, the programming model for extended architecture is a simple extension, not a radical change that requires extensive software redesign.
【0009】
Data word dimensions for integer operations determine machine registers and data paths<u style="single">m</u>When expanding from a bit to an N bit and loading it in a register, the entity of bits m or less is extended from the m bit to the N bit by sign-extending to the N bit. In this case, the term "sign-extending" refers to the maximum digit bit (ie, sign) bit of an m-bit entity as the (Nm) maximum digit bit position of the N-bit container (otherwise). Means to write in (indefinite).
【0010】
The first subset of the extended architecture instruction set contains instructions from traditional architecture. These instructions (ie, instructions) are called m-bit instructions and have been redefined to work with N-bit entities that can be N-bit sign extensions of m-bit (or less) entities. There is. For compatibility, i.e., said<u style="single">m</u>Bit instructions<u style="single">m</u>When operating with an N-bit sign extension of a bit entity, it must generate an N-bit result, which is an N-bit sign extension version of the correct m-bit result.
【0011】
The second subset contains instructions that are not defined in traditional architecture. These instructions are called N-bit instructions and usually work with N-bit entities that are not sign-extended versions of m-bit entities. Whether or not an absolute N-bit instruction is required corresponds to code-extended and non-sign-extended entities.<u style="single">m</u>It depends on the operation of the bit instruction.
【0012】
For example, logical operation<u style="single">m</u>Some of the bit instructions<u style="single">m</u>Of course, when working with N-bit sign extensions of bit entities, N bits<u style="single">To</u>Sign-extended correct<u style="single">m</u>Generates an N-bit result that corresponds to the bit result. Therefore, compatibility does not require further definitions of these instructions and works correctly for N-bit entities that are not sign-extended. Therefore, these<u style="single">m</u>Bi<u style="single">To</u>It is not necessary to provide another N-bit instruction corresponding to the instruction.
【0013】
Others such as some shift instructions etc.<u style="single">m</u>In the case of bit instructions, the situation is different. If these instructions work for an N-bit sign-extended version of an m-bit entity, then the N-bit result is correct with the N-bit sign-extended.<u style="single">m</u>It may not correspond to the bit result. In the case of instructions that do not naturally guarantee a sign-extended result, compatibility requires that these instructions be defined to guarantee a sign-extended result. This means that these instructions tend not to produce correct results for non-sign-extended N-bit operands. Therefore, these<u style="single">m</u>Another N-bit instruction corresponding to the bit instruction is required.
【0014】
Addition operation is one of the instructions that does not guarantee a sign-extended result, and therefore the result<u style="single">m</u>It requires an extra circuit for sign extension of the bit part. However, addition tends to set a lower limit with respect to cycle time. Therefore, performing a sign extension for all addition operations requires a longer cycle time, or the addition needs to be performed in two cycles. According to one aspect of the invention, sign extension for addition is only performed when necessary, when an overflow of m-bit 2's complement is detected. In the case of a pipelined configuration, for one instruction per cycle, this would stall or stop the pipeline, perform sign extension, and set the correct value during the restart period of the pipeline's sequence. Achieved by inserting into the pipeline.
【0015】
This extended architecture performs error checking and address translation for various parts of the N-bit virtual address field, including the part above bit (m-1). Traditional architecture m-bit addresses are all the extra high-order bits required by address translation and error checking mechanisms.<u style="single">To</u>Is accepted by giving the value of the sign bit (bit (m-1)) to.
【0016】
In the specific configuration, N-bit addressing and<u style="single">m</u>Bit addressing shares a common address generation circuit. In this configuration, this extended architecture is as an N-bit entity in sign-extended form.<u style="single">m</u>By generating and storing bit addresses and requiring that the results of address calculations for these entities be in sign-extended form.<u style="single">m</u>Supports bit architecture addressing.
【0017】
Support for expanded virtual address spaces is provided, in part, by extending the virtual address data path, which includes the data address adder, the branch adder, and the program counter (PC) to N bits. ing. In one embodiment, the N-bit virtual space is divided into a number of regions distinguished by higher-order virtual address bits. For example, in certain embodiments where N = 64 and m = 32, the virtual address bit (63..62) specified as VA (63..62) has the address 0,2.<sup>62</sup>,2×2<sup>62</sup>,3×2<sup>62</sup>It presents four areas starting with. This extended architecture has a maximum of 2<sup>VSIZE</sup> It gives a large number of uniform virtual subspaces of bytes, and VSIZE depends on its configuration. In this particular embodiment, VSIZE is constrained to the range 36-62. Some parts of these areas are available, depending on whether the machine is in user mode, supervisor mode, or kernel mode. User mode starts at address 0 2<sup>VSIZE</sup> It is possible to address a flat space of bytes. If VA (61..VSIZE) is not converted by TLB and VA (63..VSIZE) is not all 0, an address error will occur. Sparvisor mode is user mode space, address 2<sup>62</sup>Start with 2<sup>VSIZE</sup> 2 in byte space and near the top of the 4th region<sup>29</sup>It is possible to address byte space. Kernel mode is the space in the first and second regions, many unmapped spaces in the third region, address 3x2.<sup>62</sup>Start with (2)<sup>VSIZE</sup> -2<sup>31</sup>) Byte space, and 2 at the top of the fourth region<sup>31</sup>It is possible to address byte space. 2<sup>64</sup>2 at the top and bottom of the bite space<sup>31</sup>The byte space is called the compatibility space. This is because those addresses are in the form of 32-bit addresses with a 64-bit sign extension, and therefore access to 32-bit addressing is possible.
【0018】
In certain configurations, traditional architecture-enabled user addresses (MSB = 0) were allowed to arise from a two's complement overflow. To handle this special case in the Extended Architecture, the machine is running in an m-bit program (ie, written for a traditional architecture) or in an N-bit program (ie, in the Extended Architecture). It is necessary to supply the address mode in the status register of the machine to identify whether it is running on (written for). In m-bit user mode, it is necessary to sign-extend the address when 2's complement overflow occurs.
【0019】
A simple approach for m-bit mode is to provide sign-extension hardware in the virtual address path to guarantee sign-extension output in the case of N-bit two's complement overflow. However, if timing constraints adversely affect such sign extensions, it is sufficient to force (Nm) maximum digit bits to zero. Therefore, in one embodiment, a zeroing circuit is provided in the path to the address translation unit. This circuit is evoked for m-bit user mode when the maximum digit (Nm) is passed invariant in m-bit kernel mode and N-bit mode.
【0020】
Using this sign extension property provides an elegant way to extend data word and virtual address dimensions without requiring an unreasonable amount of extra hardware to support traditional architecture. Giving. Using virtual addressing for VSIZE bits with error exceptions has two advantages. First, it only requires that the TLB for a given processor be the required length for an existing VSIZE, and will provide virtual address space in future implementations or configurations. It is possible to increase. Second, the use of unmapped virtual address bits for other purposes is prohibited, so programs written for a processor with a certain VSIZE will run on a later processor with a larger VSIZE. That's what it means.
【0021】
[Example]<u style="single">Introduction and definition</u>The present invention provides a technique for extending an existing architecture to a new architecture characterized by one or both of larger data word dimensions and / or larger virtual address space.
【0022】
The particular traditional 32-bit architecture described below is RISC (Reduced Instruction Computer) architecture, which is known as R2000, R3000, R6000 manufactured by Mips Computer Systems, Inc. in Sunny Bell, California. It is a RISC architecture realized on a processor. A comprehensive description of this architecture can be found in Gerry Kane, MIPS RISC Architecture, Prentice Hall Publishing Co., Ltd., 1988 (Library of Congress No. 88-060290). The 64-bit extension gives a number of new 64-bit instructions, but includes virtually all traditional 32-bit instructions.
【0023】
In terms of technical terms, the terms "double word" and "half word" usually mean 64-bit and 16-bit entities, respectively. The term "word" is sometimes used generically and occasionally refers to a 32-bit entity. Bits are usually numbered with bit (0) as the least significant (rightmost) bit. The bytes in a word or halfword can be ordered in a big-endian (byte 0) leftmost) or little-endian (byte 0) rightmost) system.
【0024】
The term "sign-extension" means an action or operation performed when a data entity is smaller than the size of the container in which it is stored. In such a case, the maximum digit bit (ie, sign bit) is repeated within the free bit position on the left. For example, a code extension for a 32-bit entity that should be stored in a 64-bit container is that the bit position of the 64-bit container stores that 32-bit entity at (31..0) and the bit position of that container (63 .. It is necessary to store the value of bit (31) of the 32-bit entity in all of 32).
【0025】
The term "zero-extension" means the behavior or operation when 0 is used to fill the bit position to the left of a data entity smaller than the container. The term "extension" is used in some cases (eg, depending on the definition of a particular instruction) to mean either sign extension or zero extension.
【0026】
The term "true 64-bit entity" refers to a 64-bit entity when the data content of the 64-bit entity exceeds 32 bits, that is, it is not a 32-bit or less entity sign-extended to 64-bit. Occasionally used to mean.
【0027】
The term "two's complement overflow" refers to the situation where the carryouts from the two maximum digit bits are different. This happens when the addition of two positive or two negative numbers produces an unacceptable result. For 64-bit numbers, the range is -2<sup>63</sup>~ (2<sup>63</sup>-1). An overflow that does not cause a carryout from bit (63) is, for example, (2).<sup>63</sup>Occurs when trying to add -1) to 1. The correct result is 2<sup>63</sup>However, the calculated result is -2<sup>63</sup>Is. Carry-out without overflow occurs, for example, when adding -1 (which is 64 positions in 2's complement binary form) to 1. The correct result is 0, but the calculated result is 0 with a carryout from bit (63).
【0028】
The term "32-bit overflow" or "32-bit 2's complement overflow" refers to the situation where the carryouts from bit (31) and bit (30) are different. This happens in two contexts. In the context of adding 32-bit entities, it has the same meaning and effect as described above. In the context of adding two 32-bit entities, each sign-extended to 64-bit, the result of such an overflow is a 64-bit entity with bits (32) different from bits (31). That is, it is a 64-bit entity that is no longer in sign extension form.
【0029】
The relationship between this overflow and sign extension is described in the simpler context of sign-extending an 8-bit entity to 16 bits. First, consider the case of sign-extending the result of adding sign-extended entities, for example, the case of the sum of hexadecimal numbers 7F and 80. As an 8-bit entity, its sum is FF. The sum of these sign extensions, that is, the sum of 007F and FF80, is FFFF, which is the sign extension version of the sum FF. Next, consider the case where the result of adding the sign-extended entities is not the sign-extended one, for example, the case of the sum of 7F and 01. As an 8-bit entity, its sum is 80. It represents a two's complement overflow. The sum of these sign extensions, that is, the sum of 007F and 0001, is 0080. However, the result is not a sign-extended entity. This is because the proper sign-extended entity is FF80.
【0030】<u style="single">System overview</u>FIG. 1 is a block diagram of a single-chip processor 10. With a few exceptions noted below, the schematic description of the system applies to conventional processors and processors under development incorporating the extended architecture of the present invention. The main difference at this high level is that traditional processors are characterized by 32-bit word dimensions and virtual addresses, while this extended architecture is characterized by 64-bit word dimensions and virtual addresses up to 64-bit. That is the point. The functional configurations and pipelines described below are compatible with the R2000 processor.
【0031】
The processor 10 has six synchronized functional units, namely the master pipeline control unit (MPC) 12, the execution unit (EU) 15, the address unit (AU) 17, and the translation. It has a (translation) lookaside buffer (TLB) 20, a system coprocessor 22, and an external interface controller (EIC) 25. These functional units communicate with each other via a number of internal buses, including a data / instruction bus 30, a virtual address bus 32, and a physical address bus 33. Off-chip communication takes place via data, address and tag bus 35.
【0032】
Instructions occur in a five-stage pipeline at a peak rate of one instruction per cycle. The MPC12 has an instruction decoding circuit 37 for decoding an instruction field latched from the data / instruction bus 30. When the instruction is decoded, the decoding MPC supplies an appropriate control signal to other functional units. It also has defect handling logic 38 that controls the pipeline when some anomalous conditions occur. For example, in the event of a cache miss, the MPC stalls or shuts down the pipeline. If another operation, such as address translation, cannot be completed without interference, the MPC shuts down the pipeline and transfers control to the operating system. The MPC also ensures that simultaneous exceptions can be serialized and execution can be precisely resumed after the exception service.
【0033】
EU15 is described in detail below. For this initial explanation, the EU has a register file 42 containing a large number of general purpose registers, an ALU 45 for performing logic, shift and addition operations, and a multiplication / division unit 47. It is enough to point out. The register file 42 has 32 registers, and the register (0) is hardwired to the value 0. In addition, there are special registers HI and LO used to multiply and divide the instructions. In traditional architecture and hardware configurations, the data paths in registers, ALUs and execution units are 32 bits wide, and in this extended architecture they are 64 bits wide.
【0034】
The EU installs two source operands from the register in every cycle. The operand is delivered to the ALU or to the AU17 or EIC25. At the same time, the EU writes a single result from the ALU, AU, or memory back into the register. The bypass operation allows the ALU or memory reference to take its source operand from the previous operation, even if the result was not written into the register file. The multiplication / division section 47 operates autonomously from the rest of the processor, so it can operate in parallel with other ALU operations.
【0035】
The AU17 will be described in detail below. In this initial explanation, it is sufficient to point out that the AU has a program counter (PC) 50 and shares the register file 42 with the EU15. In conventional architecture and hardware configurations, the PC and data paths in the AU are 32 bits wide, whereas in this extended architecture and hardware configurations they are 64 bits wide. In this extended architecture and hardware configuration, the PC and data path are 64-bit wide. The AU generates one instruction or data virtual address on each of the two clock phases per cycle. It generates instruction addresses from the current PC, from a branch offset from the PC, or from a jump address that comes directly from the EU. In a subroutine call, the AU also passes the PC to the EU as a return link. The AU generates data addresses from instruction offsets and base registers provided by the EU.
【0036】
The TLB20 is completely associative and receives instructions and data virtual addresses in alternating clock phases for mapping to physical addresses. Each translation combines the virtual address with the current processing identifier. Therefore, the TLB does not need to be cleared by a context switch between processors. The system coprocessor 22 translates virtual addresses into physical addresses and manages translations and exceptions between kernel and user states. It also controls the cache subsystem and provides diagnostic control and error recovery capabilities. One of the system coprocessors called coprocessors (0) and TLB20 can both be referred to as memory management units. This system coprocessor has a number of special registers including a status register 51 with bits indicating kernel / user mode, interrupt enable, processor diagnostic state, and 32-bit mode and supervisor mode in extended architecture. There is.
【0037】
The EIC25 manages separate instructions and data caches, main memory and processor interfaces with external coprocessors. It generates and tests data and address-tag parity for all cache operations to aid system reliability. The EIC also monitors external and internal software interrupts.
【0038】<u style="single">Traditional Architecture-Instruction Set Overview</u>All processor instructions consist of a single 32-bit word. Table 1 shows the three processor instruction types (immediate value, jump, register) and the instruction format for the coprocessor instruction. Immediate value type instructions include load, store, ALU immediate values, and branch instructions. Jump type instructions include direct jump instructions. Register type instructions include ALU3 operands (addition, subtraction, set and logic), shifts, multiplication / division, indirect jumps, and exception instructions.
【0039】
The immediate instruction identifies two registers and a 16-bit immediate field. In the case of load and store instructions, the registers are the base register and the source / destination register, and the immediate field contains the address displacement (offset) that is sign-extended and added to the contents of the base register. doing. The resulting virtual address is translated and data is transferred between the addressed memory location and the source / destination register. In the case of a computational (ALU immediate) instruction, the immediate field is expanded, combined with the contents of the source register, and the result is stored or stored in the destination register. In the case of a branch or branch instruction, the immediate field is sign-extended and added to the PC to form the target address.
【0040】
The register instruction identifies up to three registers and one numeric field. In the case of addition, subtraction, AND, OR, XOR, NOR instructions, the two source registers are combined and the result is stored in the destination register. In the case of a set-on-less-than instruction, the two source registers are compared and, depending on their relative value, the destination register is a value of 1 or 0. Is set to. In the case of a shift instruction, is the content of one source or source register shifted and sign-extended by the number defined by the lower bits of the content of the source register or by the specified number? Or it is zero-extended and stored or stored in the destination register. In the case of a multiplication instruction, the contents of the two source or source registers are multiplied and the double result is stored in the LO and HI special registers. In the case of a division or division instruction, the contents of one source or source register are divided or divided by the contents of the other register, and their quotients and remainders are stored in the LO and HI registers. The LO and HI registers can also be written and read by a move or move instruction.
【0041】
Many of the instructions (load, set on less than, multiply, divide) have an "unsigned" counterpart, in which case the operand is an unsigned integer instead of a two's complement integer. Will be handled. Addition and subtraction instructions also have a corresponding "unsigned", but the terms have different meanings. Normal addition and subtraction instructions are trapped by overflow (different carryouts from bit (30) and bit (31)), but unsigned ones are not trapped by overflow. The definition of a separate instruction that traps and does not trap represents the chosen alternative to using a condition code that identifies the overflow.
【0042】<u style="single">Traditional execution unit</u>FIG. 2A is a block diagram showing the configuration and data path in a conventional example of EU15 corresponding to the R2000 processor. FIG. 2 is a schematic schematic diagram showing only relevant parts for understanding the present invention. For example, this figure is actually a latch-based configuration using a two-phase clock, but is shown in a register-based display. In addition, hardware for RF, ALU, MEM, and WB stages is shown, but not for IF stages. It should be understood that the control signals obtained from instruction decoding by MPC12 are transmitted to the various elements in this figure. With a few exceptions, these control signals are shown roughly as "CTL", where CTLs show different signals in different places.
【0043】
In this embodiment, the register file 42, ALU45 and all other registers and data paths are 32 bits wide. The ALU45 is shown to consist of a shift unit 52, a logic unit 53, an adder 55, and an ALU multiplexer 57. A conditional branch circuit is provided in connection with the ALU, which has a comparison circuit 58 that compares two data operands and a branch decision logic 60. The branch decision logic relies on a particular branch instruction to make a branch decision based on the result of the comparison and / or one of the sign bits (bit (31)) of the operand, and the control signal is AU17. Send to. Zero determination is made by comparing with register (0).
【0044】
The multiplexer 57 supplies the ALU output selected from the outputs of the shift unit, the logic unit, and the adder. The shift unit 52 receives a single operand, but each of the logical unit 53 and the adder 55 receives two operands. The second operand may be register data or immediate data extended by the extension circuit 63 as selected by the operand multiplexer 64.
【0045】
The adder 55 executes addition, subtraction and set-on-less-zan instructions (in the latter case, subtraction is involved). The adder has a circuit that monitors 32-bit overflow, revealing as different carryouts from bit (30) and bit (31). As mentioned above, some add and subtract instructions are trapped in such overflows.
【0046】
Pipeline registers 65a-b, 67,68 are inserted at various points along the data path. These registers are clocked in every cycle unless MPC12 specifies the stall condition used to disable each clock input in the register. The bypass multiplexer 70 supplies the operand from any one of the pipeline registers 65a, 67, 68 to the ALU. This allows the ALU to access the loaded or processed data before it is loaded into the register file. The ALU also receives an operand from pipeline register 65b.
【0047】
All instructions follow the same sequence of five pipeline stages during execution. These stages are instruction fetch (IF), source operand fetch from register file (RF), ALU operation or data operand address generation (ALU), data memory reference (MEM), write back to register file (WB). is there. During the IF cycle, the processor translates the instruction virtual address into the instruction physical address and sends it to the instruction cache. The processor chip receives the instruction and decodes it during the RF cycle. The source operand transitions to the appropriate operation, logic or address unit during the ALU cycle. If the instruction makes a memory reference, the data cache receives the translated data address during the MEM cycle period and returns the data for register file writing during the WB cycle period. Writing for ALU operations is done in the same pipeline stage. Bypassing between instructions in the pipeline allows the latency of branching and memory references to be maintained in one cycle and the ALU results to be used in subsequent instructions.
【0048】
During the RF cycle, data is read from register file 42 and clocked into one or both of pipeline registers 65a-b. During the ALU cycle, the data from the selected pipeline register is processed by the ALU and the result is clocked into pipeline register 67. Many conditional branch instructions depend on the sign or sign of the data entity. For this, the bypass multiplexer output bits (31) (sign bits) are sent to the sign bit test logic associated with the AU17. During the MEM cycle, the contents of pipeline register 67 or the result of fetching from memory are clocked into pipeline register 68. The load multiplexer 72 determines the source, that is, the source or source, based on the signal from the instruction decoder. During the WB cycle, the contents of pipeline register 68 are fed to the register file and clocked into the appropriate registers.
【0049】
Data from memory (typically cache memory) is sent to load logic 75, which receives a signal from the instruction decoder. For load instructions that load less than the full 32-bit word from memory, the load logic performs sign-extension or zero-extension to 32-bit, depending on the particular instruction. The result of the sign extension is that the leftmost bit of the byte or halfword from memory is repeated within the open bit position on the left. The result of zero expansion is that the bit positions that were not filled by the loaded data entity are written at zero. FIG. 2B is an enlarged block diagram of the 32-bit shifter 52. As will be understood, this shifter has the ability to shift input 0 for a left shift and either 0 for a logical right shift or a sign bit for an arithmetic right shift (bit (31)). A multiplexer 76 for shifting any of the above is provided.
【0050】<u style="single">Traditional Architecture-Addressing Overview</u>In traditional architectures, sometimes referred to as 32-bit architectures, virtual addresses are 32-bit entities. The virtual memory system gives a logical expansion of the physical memory space of the machine by translating the address composed of the 32-bit virtual address space into the physical space of the machine. The number of bits in the physical address is specified as PSIZE. For R2000 and R3000 processors, PSIZE = 32 and virtual address mapping uses 4096 bytes (4KB) pages. Therefore, mapping through the TLB affects only the maximum 20 bits of a 32-bit virtual address, the virtual page number (VPN), and the remaining 12 bits, referred to as offsets, are passed unchanged. .. The number of bits in the offset is specified as OSIZE. For the R6000 processor, OSIZE = 14, PSIZE = 36, and the page is 16384 bytes (16KB) with an 18-bit VPN. This virtual address is extended with an address space identifier (ASID). The ASID field is 6 bits for the R2000 and R3000 processors and 8 bits for the R6000 processor.
【0051】
In the following description, the given bit (i) or given bit (j..k) of the virtual address is referred to as VA (i) or VA (j..k). Similarly, bits at physical addresses are referred to as PA (i) or PA (j..k). The bit output from the TLB is referred to as the TLB (j..k), whose bit position refers to the corresponding bit position at the physical address. Therefore, if the offset bits are not converted, the TLB output bits are TLB ((PSIZE-1) .. OSIZE) or TLB (31..12) for R2000 and R3000 processors and TLB for R6000 processors. (35..14).
【0052】
Traditional processors (R2000, R3000, R6000) that support 32-bit architectures can operate in user mode or kernel mode, as determined by one or more bits in the machine's status register.
【0053】
Figure 3A is an address map for the user-mode virtual address space. The processor that operates in user mode is 2<sup>31</sup>Gives a single, uniform, mapped virtual address space of bytes (2GB). In this context, the term "mapped" means that the virtual address has been translated by TLB20, while the term "unmapped" means that the address has not been translated by TLB20. And it means that the physical address bit is taken directly from the virtual address. All valid user mode virtual addresses have bit (31) = 0, and attempts to translate addresses with bit (31) = 1 or fetch from such addresses will result in an address error exception. generate.
【0054】
The calculation of the data address requires adding an immediate offset to the contents of the base register. It is possible to start with the contents of a base register with bit (31) = 1 that is not a valid user mode address, and by adding a negative offset, bit (31) = 0 and 2 It is possible to give a data address with a complement overflow. Furthermore, it is possible to start with the contents of the base register with bit (31) = 0, and by adding the offset, the data with 2's complement overflow and with bit (31) = 1. It is possible to give an address. The 32-bit architecture ignores the overflow in all cases, but an address error exception occurs when bit (31) = 1.
【0055】
The calculation of the instruction address requires adding an immediate offset (ie 1) to the contents of the PC. In this case, the PC cannot have an address with bit (31) = 1, so it is not possible to form a valid user address with a two's complement overflow.
【0056】
Figure 3B is an address map for the kernel mode virtual address space. When operating in kernel mode, four separate virtual address spaces can be used at the same time, distinguished by higher-order bits of the virtual address. If VA (31) = 0, the selected virtual address space covers the full 2GB of the current user address space. If VA (31..29) = 100, the selected virtual address space is 2<sup>29</sup>Bytes (0.5GB) Cached and unmapped kernel physical address space. If VA (31..29) = 101, the selected virtual address space is 0.5GB of uncached, unmapped kernel physical address space. If VA (31..30) = 11, the selected virtual address space is 2<sup>30</sup>A mapped kernel virtual space of bytes (1GB). Virtual addresses for unmapped kernel space (addresses with VA (31..30) = 10) do not pass through the TLB, but they are 0- (2).<sup>29</sup>-Mapped as a block for physical addresses within the range of 1) (in a limited sense). That is, they have a physical address with PA (31..29) = 000. Kernel address behavior is constrained, so the base address register must point to the same space as the result.
【0057】<u style="single">Traditional address generation and translation</u>FIG. 3C is a block diagram showing a configuration and an address path in a conventional embodiment of the AU17 and a related address translation circuit corresponding to the R2000 processor. In this embodiment, the PC50 and all other registers and data paths are 32 bits wide. The virtual data address for the load and store instructions is calculated by the data address adder 77. The adder 77 combines the base address from one of the registers with the offset derived from the 16-bit immediate field in the instruction. The sign extension circuit 78 signs the 16-bit offset to 32 bits before the offset is coupled in the adder. Depending on whether the instruction is a load or a store, the data entity is read from an addressed memory location and loaded into a destination register, or source or source or source. The data entity in the register is written to the addressed memory location.
【0058】
The instruction address portion of the AU has a PC 50 and the next address multiplexer 80, which loads the PC at one of the four inputs selected. The first input selected for normal sequence operation is the address of the next sequence. This is given by the incrementer 82, which receives the contents of the PC and increments it. The second input is the one chosen in the case of a jump instruction, which is the jump address. It is either sourced from that instruction or sourced from a register file. The third input is selected in the case of a branch instruction taken, i.e. a branch instruction, which is the branch target address given by the branch adder 85. The branch adder combines the contents of the PC and the offset derived from the 16-bit immediate field in the branch instruction. The fourth input is the one chosen in the case of exceptions, that is, the exception vector. It is supplied by an exception multiplexer that receives a fixed exception vector at its input end.
【0059】
The virtual addresses from the data address adder 77 and the PC 50 are sent to the address translation circuit that generates the physical address. The 12-bit offset (VA (11..0)) bypasses TLB20 for all addresses and defines PA (11..0). The VA (31..12) is sent to the TLB and the VA (29..12) bypasses the TLB. A set of multiplexers 83 and 85 are controlled to TLB (31..12) for mapped space or three 0s and VA (28..12) preceding unmapped kernel space. select. As mentioned above, the unmapped kernel space has VA (31..30) = 10, and it is this condition that is used to control the multiplexer.
【0060】
In the R2000 and R3000 processors, the TLB20 is fully associated with an on-chip TLB with 64 entries, all of which are checked simultaneously for matching or matching with the extended virtual address. .. The TLB entry is defined as a 64-bit entry, but only 50 bits are stored, namely the 20-bit VPN, the 6-bit ASID, the 20-bit page frame number (PFN), and the page cached. Four bits that identify the cache algorithm for the page with respect to whether it is, whether the page is dirty or dirty, and whether the entry is valid. In the R6000 processor, this TLB is two sets of associative cache TLBs.
【0061】<u style="single">Instruction set for extended architecture</u>The extended architecture employs a hardware configuration in which registers, ALUs and other data paths within the EU are 64-bit wide. The instruction set or instruction set has all conventional (ie, 32-bit) instructions as well as a large number of 64-bit instructions for handling doublewords. Therefore, this extended architecture is an upper set of conventional architectures. However, it should be understood that 32-bit instructions actually process 64-bit entities, but in some cases 64-bit entities have their actual data content less than or equal to 32 bits (words, halfwords, and). It is a sign extension or zero extension of an entity that is a byte).
【0062】
As will be understood from the following description, many 32-bit instructions work correctly for true 64-bit entities, and also for 32-bit entities that are sign-extended to 64-bit. The instruction decoding circuit in the MPC provides the EU with sufficient information to identify whether the instruction is one of a 32-bit instruction or one of a doubleword instruction. Is possible. An important feature of this extended architecture is that the hardware configuration that supports the extended architecture has only 32-bit instructions and is capable of running programs that process 32-bit or smaller entities. It is possible to produce exactly the same results as those produced by traditional processors that support bit architecture. This is achieved by requiring a 32-bit instruction to operate or perform an operation on a 64-bit sign-extended version of a 32-bit data entity and generate a 64-bit sign-extended version of the correct 32-bit result. Has been done. In the following description, which of the 32-bit instructions naturally operates or performs an operation or operation on the sign extension entity to produce a sign extension result, and which of them operates or performs an operation or operation on the sign extension entity. Explain whether sign extension is required. An extra sign extension circuit is needed if the 32-bit instructions that operate on the sign extension entity do not naturally produce a sign extension result. This means that the instruction generally has no effect on true 64-bit entities. In these cases, a separate 64-bit (doubleword) instruction is added to the instruction set.
【0063】
Instructions are listed in a set of tables, the description of the instruction and its opcode (ie, the instruction code), and in the case of 32-bit instructions, sign extension of the result, if required, or so. If not, an indication of what type of extension is needed to get a well-defined amount, and in the case of 64-bit instructions, that the instruction is unique to the extension architecture. The display is specified. The table, ie Table 3A-3D, shows most immediate instructions. As mentioned above, the immediate instruction identifies an opcode, a pair of registers, and an immediate field that is either data or an offset for address calculation.
【0064】
Table 3A shows the load instructions. Byte and halfword load instructions (LB, LBU, LH, LHU) produce the result of sign extension or zero extension to 32 bits when executed on a 32-bit machine. Note that unsigned load instructions (LBU, LHU) use zero extension from 8 or 16 bits to 32 bits. Therefore, zero extension to 64-bit is equivalent to sign extension, and these instructions operate on 64-bit machines with sign extension to 64-bit. Similarly, wordload instructions (LW) and special inconsistent wordload instructions (LWL, LWR) are defined in a 64-bit architecture for sign extension from bit (31) and therefore require new instructions. There is no such thing. Actions or operations on true 64-bit entities are analogical to four new doubleword load instructions and unsigned bytes and halfword load instructions (LBU, LHU) that fill the left bit (63..32) with zeros Requires an unsigned wordload instruction (LWU) that operates on.
【0065】
Table 3B shows the store instructions. For bytes, halfwords and words stored as sign-extended 64-bit entities, existing instructions do the same for 32-bit and 64-bit operations. This is because the instruction ignores at least the upper 32 bits of the 64-bit register. New doubleword store instructions (SD, SDL, SDR, SCD) are required for 64-bit architectures.
【0066】
Table 3C shows the ALU immediate instruction. Bit-by-bit logical immediate instructions (ANDI, ORI, XORI) generate sign extension results when supplied with sign extension input operands. It should be noted that a 16-bit immediate value is zero-extended to 32-bit in a 32-bit architecture and zero-extended to 64-bit in a 64-bit architecture before being combined with a register value. Therefore, these instructions work for both 32-bit and 64-bit operations.
【0067】
Additive instructions (ADDI and ADDIU) require additional 64-bit instructions for different reasons. ADDI is defined in 32-bit architectures to trap on a 32-bit 2's complement overflow and not write the result in the destination register. Therefore, when the result of adding two sign extension entities is written, the result is in sign extension form and therefore no sign extension of the result is required. However, another 64-bit instruction (DADDI) trapped on a 64-bit two's complement overflow is required for operations on true 64-bit entities. Because 32-bit overflow is irrelevant.
【0068】
ADDIU does not trap on 32-bit overflows, so it is possible to produce unsign-extended results when operating on sign-extended inputs. Therefore, the sign extension of the result is required, and the operation or operation on the 64-bit entity requires a new instruction (DADDIU) that does not sign-extend the result.
【0069】
Immediate load (LUI) requires sign extension when performing memory load. The set-on-less-zan instructions (SLTI and SLTIU) load 1 or 0, and its 64-bit version is a sign-extended (and zero-extended) 32-bit version. No separate sign extension circuit is needed, so no new instructions are needed.
【0070】
Table 3D shows conditional branch instructions. Branch instructions (BEQ, BNE, BEQL, BNEL) that test the identity of two entities make 64-bit bit-by-bit comparisons and therefore operate in the same manner with respect to sign-extended input operands and true 64-bit operands. Perform the operation. Therefore, no new instructions are required. The remaining conditional branch instructions test the sign bit, or bit (63), but for sign extension input operands, this is the same as the 32-bit operation or bit (31) for operation. Therefore, no new instructions are needed.
【0071】
Table 4A-4C shows most of the register instructions. As mentioned above, the register instruction identifies up to three registers. Table 4A shows the ALU3 operand register instructions, where the contents of the two registers are processed and the result or the value representing the result is stored in the third register. Addition and subtraction instructions (ADD, ADDU, SUB, SUBU) require new 64-bit instructions for ADDI and ADDIU for the same reasons as described above. ADDs and SUBs are defined to trap on 32-bit overflows and therefore require DADD and DSUB instructions to trap on 64-bit overflows. ADDU and SUBU are not guaranteed to produce a sign extension result and therefore require a sign extension of that result. Therefore, new instructions (DADDU and DSUBU) are needed for double word addition and subtraction without sign extension.
【0072】
Set-on-less Zan instructions (SLT and SLTU) give zero extension values of 0 and 1 in 32 bits and thus a sign extension version in 64 bits. Therefore, no new instructions are needed. Bit-by-bit logical instructions (AND, OR, XOR, NOR), given a sign-extension input operand, naturally produce a sign-extension result, so no new instructions are needed.
【0073】
Table 4B shows the shift instructions. Shifting a sign extension entity usually does not produce a sign extension result, and therefore a sign extension circuit is required. Special logic for logical right shift operations (SRL and SRLV) is also required to shift 0 into bit (31) before sign extension (assuming some shift occurs). Therefore, an instruction for shifting more than 32 bits and an additional double word shift instruction are given.
【0074】
Shift write arithmetic instructions (SRA and SRAV) sign-extend from the left and thus give correct results for sign-extended entities and true 64-bit entities. Nevertheless, another doubleword instruction (DSRA and DSRAV) has been given. This is necessary for variable shift. This is because the SRA uses the specified register bits (4..0) to determine the shift amount, while the bits (5..0) are used to identify the shift amount for a true 64-bit entity. Because it is necessary. This is not a problem for SRA where the shift amount is a 5-bit field within the instruction. Nevertheless, it is convenient to define new 64-bit instructions.
【0075】
Table 4C shows the multiplication and division instructions. In the multiplication instruction, the low-order (lower) word of the double result is loaded in the LO special register, and the higher (higher) word of the double result is loaded in the HI special register. Separate sign extensions are required to fill the 64-bit LO and HI registers with 32-bit results. Similarly, for division or division instructions, the quotient and remainder are loaded into the LO and HI special registers, respectively, and a separate sign extension is required in this case as well. Therefore, separate doubleword multiplication and division instructions are required. Instructions that transfer content between special and general purpose registers work equally well for sign-extended and true 64-bit entities, and no additional instructions are required.
【0076】
Tables 5A and 5B show direct (immediate) and indirect jump (register) instructions. In a direct jump instruction in a 32-bit architecture, the 26-bit target address is shifted to the 2-bit shift left, that is, to the left and combined with the PC bit (31..28) and the result is jumped. In the extended architecture, the bond is a bit of a PC (63..28). If the address is sign-extended, the correct operation or operation is performed without any additional instructions. An indirect jump, or jump to the contents of a register, naturally occurs when a 32-bit address is sign-extended to 64-bit.
【0077】
Tables 6A and 6B show exception instructions. No new instruction is required for the trap instruction for the same reasons as described above for the conditional branch instruction.
【0078】<u style="single">Execution unit for extended architecture</u>FIG. 4 is a schematic block diagram showing a configuration and a data path in an embodiment of an execution unit that supports an extended architecture. The elements in FIG. 4 are approximately 64-bit wide and the elements in FIG. 2 are 32-bit wide, but the elements corresponding to the elements in FIG. 2 are given the same reference numbers in FIG. Only relevant pipeline stages are shown, as shown in Figure 2.
【0079】
It is important to make a distinction between the architecture defined by the instruction set and the hardware configuration. For example, the hardware configurations described above in connection with traditional 32-bit architectures internally use a 5-stage pipeline, with each cycle divided into two phases, the R2000 and R3000. It is related to the processor. The traditional R6000 processor has a slightly different 5-stage pipeline, while the processor currently under development to support 64-bit architecture has an 8-stage pipeline. Hardware that supports the extended instruction set architecture is represented as an extension of the R2000 / R3000 hardware configuration. This simplifies the description of the extended architecture and how to form the hardware to support it. However, it should be understood that the hardware configuration to support this extended architecture does not have to be the same as that of the conventional processor.
【0080】
In order to support the extension architecture described above, an extra extension circuit is inserted in the data path in which the operation related to the sign-extended entity, that is, the operation result does not guarantee the sign extension result. For this purpose, the extension circuit 120 is inserted in the data path between the shift unit 52 and the ALU multiplexer 57. This extension circuit is activated only for certain 32-bit shift instructions, the fact of which is outlined by a control input with a set of signals designated as 32S. This 32S signal identifies, among other things, whether sign extension or zero extension is required.
【0081】
Similarly, the load logic 125, which supplies the data entity from the memory subsystem, must in this case provide 64-bit sign extension or zero extension for 32-bit words and halfwords and bytes. This load logic is shown to receive the general control signal CTL and an additional signal 32L that identifies an extension to the 32-bit load. This load logic can be conceptually thought of as a 32-bit entity sign-extended to 64-bit, in which case the 32-bit entity itself is a sign-extension of an entity less than 32 bits or It may be zero-extended.
【0082】
Adder 55 has a circuit that monitors the overflow of 64-bit 2's complement, which causes traps for DADD and DSUB instructions. For compatibility, the adder also monitors a 32-bit overflow that causes traps for ADD and SUB instructions. This is also necessary for ADDU and SUBU instructions that do not trap on 32-bit overflow but require sign extension in that case.
【0083】
Therefore, at least sometimes it is necessary to sign-extend the result of combining two 32-bit signed integers, each sign-extended to 64 bits, in order for the result to be in sign-extended form. Exists. A direct extension to the hardware is to provide a sign extension circuit between the adder 55 and the ALU multiplexer 57 (this is exactly what would be done if an extension circuit 120 were provided for shift operation). Is). However, addition is an operation or operation that takes more time than shift and logical operations, and adding a sign extension circuit following an adder results in overall performance degradation. Since the time to perform an addition operation tends to set a lower bound on the cycle time, a concise solution requires two cycles for every addition or more for every operation or operation. Either it requires a long cycle.
【0084】
Since 32-bit 2's complement overflow occurs relatively rarely, it is possible to use the hardware configuration shown in Figure 4A. In this configuration, the output from adder 55 does not travel through the ALU multiplexer 57 and pipeline register 67, but instead travels through another pipeline register 130 and is directed to the load multiplexer 72. Will be done. The output of pipeline register 130 is also fed to the input end of bypass multiplexer 70 (this is similar to the output of pipeline register 67). Therefore, the pipeline register 130 can be seen as a parallel extension of the pipeline register 67, and most of the addition results are introduced directly into the load multiplexer, not through the ALU multiplexer.
【0085】
However, an additional data path must be provided to handle the case where a 32-bit two's complement overflow occurs in the addition of sign-extended operands. This is accomplished by a sign extension circuit 140, which is coupled to the output end of pipeline register 130, but is outside the data path from pipeline register 130 to the load multiplexer 72 and bypass multiplexer 70. An extra multiplexer, called the overflow multiplexer 145, is inserted in the path between the ALU multiplexer 57 and the pipeline register 67, allowing the output from the sign extension circuit 140 to be selected.
【0086】
When a 32-bit overflow occurs on the ADDU or SUBU instruction, MPC12 stops the pipeline and controls the overflow multiplexer 145 with the signal SE STALL to select the output from the sign extension circuit 140. This is schematically shown in Figure 4B. This sign-extended output represents the correct result of the addition. Therefore, the load multiplexer 72 and the bypass multiplexer 70 can be provided with the correct result to be inserted into the pipeline during the pipeline restart sequence. Therefore, the sign extension circuit 140 is in the critical data path only in the case of the overflow in which it is required, and for the addition of 64-bit or 32-bit sign extension operands where overflow does not occur. Exists outside the critical path.
【0087】
FIG. 4C is an enlarged schematic block diagram of the 64-bit shifter 52 and its associated expansion circuit 120. The shifter has the ability to shift input 0 for all rest shifts, i.e. shifts to the left, and has a sign bit (bit (63)) or logic write for arithmetic light shifts (shifts to the right). A multiplexer 152 is provided to shift any of the 0's to the shift. Compatibility for 32-bit shifts is provided by multiplexers 153 and 155. For 32-bit logical light shift, the multiplexer 153 loads a 0 in the upper section of the shifter and thus shifts into an empty position where the 0 starts at bit (31). The multiplexer 155 selects 32 copies of the upper or lower section of the shifter bit (31) for the resulting bit (63..32). This is necessary to give sign extension for 32-bit left shift. It is not strictly necessary for 32-bit light shift. This is because the multiplexers 152 and 153 can ensure a sign-extended result when their input is sign-extended.
【0088】<u style="single">Address behavior for extended architecture</u>In this extended architecture, the virtual address is a 64-bit entity. The main purpose of extended architecture addressing, or addressing, is to provide an extended flat (unsegmented) secondary space while maintaining traditional architectural addressing as a subset. For this purpose, the address space of the conventional architecture is transferred in a sign-extended form. Therefore, 32-bit addresses in traditional architectures are stored and processed in sign-extended form.
【0089】
The processor for the extended architecture is determined by the bits in the status register of the machine and can operate in user mode, supervisor mode, or kernel mode. The status register also has bits that identify whether the machine is in 32-bit mode or 64-bit mode. In 32-bit mode, the address is stored and processed as a 64-bit sign extension entity.
【0090】
As mentioned above, the typical address dimension PSIZE depends on its processor. The processor currently under development called R4000 to support this 64-bit architecture has PSIZE = 36. This extended architecture is not characterized by a fixed virtual address, but is intended for processor-dependent virtual address dimension VSIZE within the range of 32-62 bits. As explained below. This is 2<sup>VSIZE</sup> Gives a virtual space of the number of bytes. The practical implication of using VSIZE is that the TLB does not need to convert the VA (61..VSIZE). Therefore, assuming that the VSIZE is chosen to reflect a valid and foreseeable need, the TLB does not need to be larger than it is needed. The VSIZE can be increased in later processor configurations when the need for a larger virtual address space arises, and it is only at this time that a larger TLB burden must be borne.
【0091】
The R4000 processor has VSIZE = 40. This means that it gives 1024GB of virtual address space, which seems appropriate for the time being. From a compatibility point of view, the important point is to maintain the compatibility of user programs. Therefore, the extended architecture of the present invention provides compatibility with user programs by maintaining addressing compatibility in 32-bit user mode. However, kernel programs tend to be incompatible with each processor, even within the same architecture. Therefore, this extended architecture does not guarantee compatibility in 32-bit kernel mode.
【0092】
To maintain compatibility in 32-bit user mode, address calculations with the result of having bit (31) = 0 are considered valid user addresses even if a 32-bit 2's complement overflow occurs. I have to. This requires that the result be sign-extended, or equivalent in this case, to force the bit (63..32) to zero. Thus, in 32-bit user mode, all bits (63..31) are 0, and the address refers to the 2GB user address space, which appears as a subset of the extended user address space. Otherwise, an address exception will occur. The need for 32-bit mode bits in the status register is the fact that two's complement overflow is allowed in traditional architectures. However, this 32-bit mode is useful in itself in that it can be used to select the TLB refill vector in case of a TLB miss.
【0093】
In 32-bit kernel mode, the virtual address space is 4GB, which is divided into five areas, which are distinguished by the high-order or high-order bits of the 32-bit portion of the virtual address. In 32-bit mode, it is assumed that any address supplied to the address translation mechanism is in sign extension form. If VA (31) = 0, then VA (63..32) is all 0 and the selected virtual address space covers the full 2GB of the current user address space. The kernel and supervisor addresses have VA (31) = 1, so VA (63..32) are all 1. Therefore, four 0.5GB spaces are 2<sup>64</sup>It exists at the localized top of the byte space, i.e. (2)<sup>64</sup>-2<sup>32</sup>+ 8000000H) and (2<sup>64</sup>It has an address between -1) and. If VA (31..29) = 100, the selected virtual address space is the unmapped kernel physical address space of the 0.5GB cache. If VA (31..29) = 101, the selected virtual address space is 0.5 GB of uncached and unmapped kernel physical address space. For VA (31..29) = 110, the selected virtual address space is 0.5 GB of mapped supervisor virtual address space. If VA (31..29) = 111, the selected virtual address space is 0.5GB of mapped kernel virtual address space. The first two of these 0.5GB spaces correspond to the 0.5GB space in a 32-bit architecture on a 32-bit processor. Providing a supervisor virtual address space represents a further distinction of the 1GB mapped kernel virtual space of the 32-bit architecture.
【0094】
Figure 5A shows an address map for a 64-bit user-mode virtual address space. In 64-bit user mode, the processor is 2 with 62 VSIZE 36.<sup>VSIZE</sup> Gives a single uniform virtual address space of bytes. Different processor configurations can achieve different virtual address space dimensions as long as they are trapped at addresses larger than the configured ones. If the VSIZE bits are configured, all valid user mode virtual addresses have VA (63..VSIZE) all 0s, and attempts to refer to addresses with all nonzero bits are addresses. Raises an error exception. Making these bits illegal for other uses ensures that the user program will run in subsequent processor configurations characterized by a larger value for VSIZE.
【0095】
Virtual address calculations that require the addition of 64-bit base registers and 16-bit offsets sign-extended to 64-bit must not overflow from bit (61..0) into bit (63..62). The virtual address is extended with the contents of the address space identifier field to form a unique virtual address. In these extended virtual address mappings, the physical addresses do not have to be one-to-one, but two virtual addresses are allowed to map to the same physical address.
【0096】
Figure 5B shows an address map for a 64-bit supervisor mode virtual address space. 2 in 64-bit supervisor mode<sup>VSIZE</sup> Two virtual address spaces of bytes and 2<sup>29</sup>There is a byte (0.5GB) space. If VA (63..62) = 00, the virtual address space selected is 2 of the current user address space.<sup>VSIZE</sup> It is a byte. If VA (63..62) = 01, the virtual address space selected is address 2.<sup>62</sup>2 of the current supervisor address space starting at<sup>VSIZE</sup> It is a byte. If VA (63..62) = 11 and VA (31..29) = 110, the address is (2).<sup>64</sup>-2<sup>32</sup>Refers to a 0.5GB supervisor address space with a starting address of + C0000000H).
【0097】
Figure 5C shows an address map for a 64-bit kernel-mode virtual address space. In 64-bit kernel mode, it is possible to use four separate virtual address space areas at the same time, which are distinguished by the higher bits of the virtual address, or VA (63..62). If VA (63..62) = 00, the virtual address space selected is 2 of the current user address space.<sup>VSIZE</sup> It is a byte. If VA (63..62) = 01, the virtual address selected is 2 in the current supervisor address space.<sup>VSIZE</sup> It is a byte. If VA (63..62) = 10, the virtual address space selected is the address range 2x2.<sup>62</sup>~ (3 × 2<sup>62</sup>-1) 8 2s located in<sup>PSIZE</sup> One of the unmapped kernel physical spaces of bytes. A particular unmapped or unmapped space depends on the value of VA (61..59), which is address 2<sup>59</sup>Start part-time away. (2<sup>59</sup>-2<sup>PSIZE</sup> ) Addresses in the byte gap (addresses with all non-zero VA (58..VSIZE)) will cause an address error. When VA (63..62) = 11, the selected virtual address space is address 3 × 2 when VA (61..VSIZE) is all 0.<sup>62</sup>Start with (2)<sup>VSIZE</sup> -2<sup>31</sup>) Bytes Kernel Virtual address space or 2GB area compatible with 32-bit kernel mode space when VA (61..31) is all 1.
【0098】
Therefore, as is understood, the 32-bit addressing region is a subset of extended addressing. A 32-bit address with a 64-bit sign extension is 2<sup>64</sup>Map to the upper and lower 2GB parts of the byte virtual address space. The 32-bit architecture virtual addressing and 64-bit addressing 32-bit modes can be associated more closely by viewing the address as a two's complement signed address. When viewed in this way, the 32-bit space is -2.<sup>31</sup>From -2<sup>31</sup>Kernel address between and -1 and 0 and (2)<sup>31</sup>Have a user address between -1) and (2)<sup>31</sup>-1) Extend to. Similarly, 64-bit space is -2<sup>63</sup>From 0 and (2<sup>63</sup>-1) with user and supervisor addresses and -2<sup>63</sup>Has a kernel address between and -1 (2)<sup>63</sup>-1) Extend to. Therefore, the 32-bit address space is -2.<sup>31</sup>32-bit mode kernel and supervisor address between -1 and 0 and (2)<sup>31</sup>It can be thought of as a central subset of the 64-bit address space with 32-bit mode user addresses between -1).
【0099】<u style="single">Address generation and translation for extended architecture</u>FIG. 5D is a block diagram showing the configuration and address path in an embodiment of an address translation circuit that supports AU17 and its associated extended addressing of 64-bit architectures. The data address generation circuit and the instruction address generation circuit differ from the prior art in that various elements are 64-bit wide. For example, the sign extension circuit 78 extends the 16-bit offset to 64 bits and the exception vector is stored in a sign-extended form.
【0100】
This address translation circuit is similar to the traditional architecture in that it accepts 64-bit addresses that include a portion that provides a flat (ie, unsegmented) extended virtual address space while maintaining 32-bit addressing as a subset. It's different. As in traditional architectures, the 12-bit offset VA (11..0) bypasses TLB20 and defines PA (11..0).
【0101】
VA (63..62) and VA ((VSIZE-1) ..12) are sent to TLB20, and VA (63..29) is sent to address test and control logic 170. The multiplexer 172 is inserted in the path and substitutes 0 for VA (63..32) in 32-bit user mode. Zeroing the higher order bits ensures that the valid user address (VA (31) = 0) that overflowed the two's complement is in sign-extended form.
【0102】
A set of multiplexers 180,182,185 provide PA ((PSIZE-1) .. 12) by selecting either the TLB bits for the mapped space or the virtual address bits for the unmapped kernel space. In the case of mapped space, the multiplexer selects TLB ((PSIZE-1) ..12), which is combined with VA (11..0) to give the physical address.
【0103】
As mentioned above, the unmapped space has VA (63..62) = 10 2<sup>PSIZE</sup> Byte space and VA (63..31) all 1 and two 2 with VA (30) = 0<sup>29</sup>Contains byte space. The logical OR of these conditions is used to control the multiplexers 180,182,185, so they select the VA bit instead of the TLB bit for the physical address. In particular, when VA (62) = 1, a non-compatible unmapped space is displayed and 0 is assigned to VA ((PSIZE-1) .. 29) by a set of multiplexers 190 and 192. The resulting physical address is 0 and (2)<sup>29</sup>-1) Ensure that it is between. It should be noted that the space is perfect 2<sup>PSIZE</sup> It is only necessary to test VA (62) on unmapped space to determine if it is bytes or just 0.5GB. Therefore, it is possible to support 32-bit addressing as a subset of extended addressing with almost no extra hardware required.
【0104】
The TLB for the R4000 processor is a fully associative on-chip TLB with 40 entries that are all checked for matches with extended virtual addresses. Each TLB entry maps an even-odd page pair. Page dimensions are controlled on a per-entry basis by a bitmask that specifies which virtual address bits should be ignored by the TLB. Page dimensions range from 4KB (VA (11..0) bypasses the TLB) to 16MB (VA (11..0) bypasses the TLB and VA (23..12) is ignored by the TLB) It is possible to change with.
【0105】
Table 7 shows the TLB format for the R4000 processor for VSIZE = 40 and PSIZE = 36. The nominal TLB entry is 256 bits, which allows for VSIZE and PSIZE extensions for future processor configurations. The fields indicated by "-" are not stored.
【0106】
Figure 5E is a schematic block diagram showing the address test and control logic 170 for the R4000 processor. It is this logic that allows the processor to access a given address rather than anything else when in a given mode (user, supervisor, kernel). References in this figure for VA (36) and VA (35) can be generalized to VA (PSIZE) and VA (PSIZE-1), while references to VA (40) and VA (39). Can be generalized to VA (VSIZE) and VA (VSIZE-1). The names of many signals consist of two two-digit numbers meaning NZ or AO followed by the bit position. The NZ signal means that if it is true, all the bits in the range are not 0, and the AO signal means that if it is true, the bits are all 1. doing. For example, NZ3532 is asserted when VA (35..32) is not all 0, while AO3532 is asserted when VA (35..32) is all 1. These signals are used to test that the virtual address satisfies the various constraints described above.
【0107】
As described in detail above, the present invention provides an effective technique for extending a computer architecture while maintaining inverse compatibility without requiring significant hardware overhead. The dilated virtual addressing technique of the present invention provides a sufficient address space that can afford to grow in realizing future processors.
【0108】
Although the specific embodiments of the present invention have been described in detail above, the present invention is not limited to these specific examples, and various modifications can be made without departing from the technical scope of the present invention. Of course it is possible. For example, the above description describes a processor that performs 64-bit integer operations with 32-bit integer operations as a subset and 64-bit addressing with 32-bit addressing as a subset, but both extensions are 64-bit. It doesn't have to be. It is also possible to extend the integer operation without extending the addressing, or vice versa. Furthermore, although it was explained that the 32-bit addressing and the 64-bit addressing share a common address generation circuit, it is also possible to simply share only the address conversion and the error check circuit, which is also remarkable in that case. It is possible to achieve various effects.
【0109】<img he="216" id="000002" wi="152" file="2_0003554342.tif" img-format="tif" img-content="drawing" /><img he="216" id="000003" wi="144" file="3_0003554342.tif" img-format="tif" img-content="drawing" /><img he="216" id="000004" wi="144" file="4_0003554342.tif" img-format="tif" img-content="drawing" /><img he="134" id="000005" wi="152" file="5_0003554342.tif" img-format="tif" img-content="drawing" /> 【0110】<img he="134" id="000006" wi="152" file="6_0003554342.tif" img-format="tif" img-content="drawing" /> 【0111】<img he="119" id="000007" wi="152" file="7_0003554342.tif" img-format="tif" img-content="drawing" /> 【0112】<img he="75" id="000008" wi="152" file="8_0003554342.tif" img-format="tif" img-content="drawing" /> 【0113】<img he="97" id="000009" wi="152" file="9_0003554342.tif" img-format="tif" img-content="drawing" /> 【0114】<img he="75" id="000010" wi="114" file="10_0003554342.tif" img-format="tif" img-content="drawing" /> 【0115】<img he="149" id="000011" wi="152" file="11_0003554342.tif" img-format="tif" img-content="drawing" />[Simple explanation of drawings]
FIG. 1 is a schematic block diagram showing a processor incorporating either a conventional architecture or an extended architecture of the present invention.
FIG. 2A is a schematic block diagram showing a conventional execution unit.
FIG. 2B is an enlarged schematic block diagram showing a shifter in a conventional execution unit.
FIG. 3A is a schematic diagram showing an address map for a conventional user mode virtual address space.
FIG. 3B is a schematic diagram showing an address map for a conventional kernel-mode virtual address space.
FIG. 3C is a schematic block diagram showing a conventional address unit and an address translation circuit.
FIG. 4A is a schematic block diagram showing an execution unit for an extended architecture.
FIG. 4B is a schematic block diagram showing a part of a master pipeline control unit.
FIG. 4C is an enlarged schematic block diagram showing a shifter in an execution unit for an extended architecture.
FIG. 5A is a schematic diagram showing an address map for a user-mode virtual address space of an extended architecture.
FIG. 5B is a schematic diagram showing an address map for a virtual address space in supervisor mode of an extended architecture.
FIG. 5C is a schematic diagram showing an address map for the kernel mode virtual address space of an extended architecture.
FIG. 5D is a schematic block diagram showing an address unit and an address translation circuit for an extended architecture.
FIG. 5E is a schematic block diagram showing address testing and control logic.
[Explanation of symbols]
Ten Single-chip processor 12 Master pipeline control 15 Execution unit 17 Address unit 20 Translation lookaside buffer 22 System coprocessor 25 External interface controller 30 Data / instruction bus 32 Virtual address bus 35 Data / address / tag bus
Every citation, both waysCites: the store holds 5 of 6
| Document | Relation | Office |
|---|---|---|
| JP01220026A | Cites | Japan |
| JP54075252A | Cites | Japan |
| JP60134937A | Cites | Japan |
| JP61177539A | Cites | Japan |
| JP63055638A | Cites | Japan |
15 members in 4 offices
Priority claims5
| Document | Office | Kind | Date |
|---|---|---|---|
| 668275 | United States of America | – | |
| 66827591 | United States of America | A | |
| 66827591 | United States of America | A | |
| 1991668275 | – | – | – |
| US19910668275 | – | – | – |
Members15
| Document | Office | Kind | |
|---|---|---|---|
| EP0503514A2 | European Patent Office (EPO) | A2 | |
| EP0503514A3 | European Patent Office (EPO) | A3 | |
| US5420992A | United States of America | A | |
| JPH08106416A | Japan | A | |
| US5568630A | United States of America | A | |
| EP0871108A1 | European Patent Office (EPO) | A1 | |
| EP0503514B1 | European Patent Office (EPO) | B1 | |
| DE69227604D1 | Germany | D1 | |
| DE69227604T2 | Germany | T2 | |
| EP0871108B1 | European Patent Office (EPO) | B1 | |
| DE69231451D1 | Germany | D1 | |
| DE69231451T2 | Germany | T2 | |
| JP2004094959A | Japan | A | |
| JP3554342B2This record | Japan | B2 | |
| JP3657949B2 | Japan | B2 |
14 legal events, as the office reported them to INPADOC
Over the term
Point at a mark for the eventEvents
| Event | Code | |
|---|---|---|
| Cancellation because of completion of termEXPY | EXPY | |
| Renewal fee payment (event date is renewal date of database)FPAY | FPAY | |
| Renewal fee payment (event date is renewal date of database)FPAY | FPAY | |
| Renewal fee payment (event date is renewal date of database)FPAY | FPAY | |
| Renewal fee payment (event date is renewal date of database)FPAY | FPAY | |
| Renewal fee payment (event date is renewal date of database)FPAY | FPAY | |
| Renewal fee payment (event date is renewal date of database)FPAY | FPAY | |
| Receipt of annual feesR250 | R250 | |
| Certificate of patent or registration of utility modelR150 | R150 | |
| First payment of annual fees (during grant procedure)A61 | A61 | |
| Written decision to grant a patent or to grant a registration (utility model)A01 | A01 | |
| Decision of grant or rejection writtenTRDD | TRDD | |
| Request for written amendment filedA521 | A521 | |
| Notification of reasons for refusalA131 | A131 |
Numbers
- Publication
- 3554342
- Publication, DOCDB
- 3554342
- Publication, EPODOC
- JP3554342B
- Application
- 5285892
- Application, DOCDB
- 5285892
- Application, EPODOC
- JP19920052858
Titles2
- Japanese
- 拡張ワード寸法及びアドレス空間を有する逆互換性コンピュータアーキテクチュア
- English
- Reverse compatible computer architecture with extended word dimensions and address space
Classification
- CPC, 9
- G06F9/342
- G06F9/30014
- G06F9/30036
- G06F9/30167
- G06F9/30189
- G06F9/30192
- G06F9/324
- G06F9/30181
- G06F9/30038
- IPC, 7
- G06F9 30
- G06F9 302
- G06F9 305
- G06F9 318
- G06F9 34
- G06F9 355
- G06F12 02