Risc microprocessor architecture implementing fast trap and exception state
4 claims: 2 independent, 2 dependent
- 1A data processor for handling exceptions in a microprocessor for use with instruction sources and exception handlers, which stores the main buffer that stores instructions in the main instruction stream and the instructions in the target branch instruction stream. A prefetch buffer unit with a group of buffers consisting of a target instruction to be used and a procedure buffer that stores the instructions of an alternative instruction stream used to realize the operation specified by the procedure instruction, and a register sequence. The prefetch buffer, which manages the operation of the prefetch buffer unit and the execution means for executing the given instruction by loading the instruction into the register sequence, shifting the instruction, and sequentially sending the instruction. -The instruction from the source of the instruction is stored in any one of the buffers in the unit, the instruction prefetched from any one of the buffers is transferred to the prefetch buffer output bus, and the prefetch instruction is specified. The prefetch buffer control means is provided to the execution means in the order of, and the prefetch buffer control means is at least the main buffer, the target buffer, and each instruction prefetch buffer in the emulation buffer. -Corresponds to the prefetched instruction ID consisting of the status storage location logically corresponding to the location, the instruction reservation information, the status register array that stores and manages the valid information of the instruction, and the instruction given to the execution means. Then, it has an instruction means for instructing whether or not a synchronization exception has occurred with respect to the predetermined instruction, and the execution means has generated a synchronization exception with respect to the instruction given to the execution means by the prefetch means. Is a data processing device comprising a calling means for calling an exception handler corresponding to the instruction means. 命令のソースおよび例外ハンドラといっしょに用いるため、マイクロプロセッサ内で例外を取扱うためのデータ処理装置であって、メイン命令ストリームの命令をストアするメイン・バッファと、ターゲット・ブランチ命令ストリームの命令をストアするターゲット・バッファと、プロシージャ命令で指定されたオペレーションを実現するために使用される代替命令ストリームの命令をストアするプロシージャ・バッファからなる複数のバッファのグループを有するプリフェッチ・バッファ・ユニットと、レジスタ列を有し、該レジスタ列に前記命令をロードしてシフトし順次送出することにより与えられた命令を実行するための実行手段と、前記プリフェッチ・バッファ・ユニットのオペレーションを管理し、前記プリフェッチ・バッファ・ユニット内のバッファの任意の1つへ前記命令のソースからの命令をストアし、前記バッファの任意の1つからプリフェッチした命令を、プリフェッチ・バッファ出力バスへ転送して、前記プリフェッチ命令を特定の順序で前記実行手段に与えるプリフェッチ・バッファ制御手段とを有すると共に、さらに前記プリフェッチ・バッファ制御手段は、少なくとも前記メイン・バッファ、前記ターゲット・バッファ、および前記エミュレーション・バッファ内の各命令プリフェッチ・バッファ・ロケーションに論理的に対応する状況記憶ロケーションからなりプリフェッチされている命令ID、命令の予約情報、命令の有効情報をストアし管理する状況レジスタアレイ、および前記実行手段に与えられた前記命令に対応して、前記所定の命令に関して同期例外が発生したかどうかを指示するための指示手段を有し、前記実行手段は、前記プリフェッチ手段によって前記実行手段に与えられた命令に関して同期例外が発生したときは、前記指示手段に対応して例外ハンドラを呼出すための呼出し手段を有することを特徴とするデータ処理装置。
- 3A data processor for handling traps in a microprocessor, the main buffer that stores the instructions in the main instruction stream, the target buffer that stores the instructions in the target branch instruction stream, and the operations specified by the procedure instructions. A prefetch buffer unit that has a group of multiple buffers consisting of procedure buffers that store instructions of an alternative instruction stream used to realize the above, and the operation of the prefetch buffer unit that manages the operation of the prefetch request. The buffer that receives the control signal and stores the instruction is determined to be the main buffer, the target buffer, or the emulation buffer, and to any one of the buffers in the prefetch buffer unit from the source of the instruction. It has a prefetch buffer control means for storing the instructions of, and further, the prefetch buffer control means is at least the main buffer, the target buffer, and each instruction prefetch buffer in the emulation buffer. -It has a status register array consisting of status storage locations that logically correspond to the location, and is characterized by storing and managing the instruction ID, instruction reservation information, and instruction valid information prefetched in the status register array. Data processing device. マイクロプロセッサにおいてトラップを取扱うためのデータ処理装置であって、メイン命令ストリームの命令をストアするメイン・バッファと、ターゲット・ブランチ命令ストリームの命令をストアするターゲット・バッファと、プロシージャ命令で指定されたオペレーションを実現するために使用される代替命令ストリームの命令をストアするプロシージャ・バッファからなる複数のバッファのグループを有するプリフェッチ・バッファ・ユニットと、前記プリフェッチ・バッファ・ユニットのオペレーションを管理し、プリフェッチ要求の制御信号を受けて命令をストアするバッファを前記メイン・バッファ、前記ターゲット・バッファ、または前記エミュレーション・バッファに決定して前記プリフェッチ・バッファ・ユニット内のバッファの任意の1つへ前記命令のソースからの命令をストアするためのプリフェッチ・バッファ制御手段とを有し、さらに、前記プリフェッチ・バッファ制御手段は、少なくとも前記メイン・バッファ、前記ターゲット・バッファ、および前記エミュレーション・バッファ内の各命令プリフェッチ・バッファ・ロケーションに論理的に対応する状況記憶ロケーションからなる状況レジスタアレイを有し、該状況レジスタアレイにプリフェッチされている命令ID、命令の予約情報、命令の有効情報をストアし管理することを特徴とするデータ処理装置。
Independent claims2
431 paragraphs in 1 section, as filed
【0001】
[Technical field to which the invention belongs]
The present invention relates to a microprocessor architecture, and more specifically to handling interrupts and exceptions in a microprocessor.
【0002】
The description of the present specification is based on the description of the US patent application No. 07 / 726,942, which is the basis of the priority of the present application, and the US patent application is referred to by referring to the number of the US patent application. The contents of this specification shall constitute a part of this specification.
【0003】
[Conventional technology]
Related Patent Applications Mutual Reference The US patent applications listed below are pending and pending at the same time as the Patent Application, but are disclosed in these US patent applications and are filed accordingly. The matters disclosed in the patent application in Japan shall constitute a part of the present specification by quoting the application number in the present specification. 1. Title of the Invention "High-Performance RISC Microprocessor Architecture" SMOS 7984 MCF / GBR, US Patent Application No. 07 / 727,006, filed July 8, 1991, inventor Le T. Nguyen et al., And the corresponding Japanese Patent Application No. 5-502150 (Japanese Patent Publication No. 6-501122). 2. Title of the Invention "Extensible RISC Microprocessor Architecture" SMOS 7985 MCF / GBR, U.S. Patent Application No. 07 / 727,058, filed July 8, 1991, inventor Le T. Nguyen et al. .. 3. RISC Microprocessor Architecture with isolated Architectural Dependencies SMOS 7987 MCF / GBR, US Patent Application No. 07 / 726,744, filed July 8, 1991, invention Le T. Nguyen et al., And the corresponding Japanese Patent Application No. 5-502152 (Japanese Patent Publication No. 6-502034). 4. Title of the invention "RISC Microprocessor Architecture Implementing Multiple Typed Register Sets" SMOS 7988 MCF / GBR / RCC, US Patent Application No. 07 / 726,773, filed July 8, 1991, inventor Sanjiv Garg et al. .. 5. Title of the Invention "Single Chip Page Printer Controller" SMOS 7991 MCF / GBR, US Patent Application No. 07 / 726,929, filed July 8, 1991, inventor Derek J .Lentze et al., And the corresponding Japanese Patent Application No. 5-502149 (Japanese Patent Publication No. 6-501586). 6. Title of the Invention "Microprocessor Architecture Capable of Supporting Multiple Heterogeneous Processors" SMOS 7992 MCF / WMB, US Patent Application No. 07 / 726,893, filed July 8, 1991, inventor Derek J. Lentze et al. ..
【0004】
Description of Related Techniques In a typical microprocessor, instructions are usually executed in order unless instructions that change the flow of control appear or exceptions occur. For exceptions, there is a built-in mechanism to change the flow of control when a particular event occurs, regardless of whether the event was related to a particular instruction in the instruction stream. For example, a microprocessor may contain an interrupt request (IRQ) read (1ead), and when this read is activated by an external device, the microprocessor displays the address of the next instruction to execute. Save some information about the current state of the machine, including, and then immediately interrupt the interrupt handler (1nterrupt). When the control right (control) is passed to the handler), the interrupt handler starts from a certain predetermined address. In another example, if an exception error such as divide-by-zero occurs during the execution of a particular instruction, the microprocessor saves information about the current state of the machine before giving control. It is supposed to be passed to the exception handler. In another example, some microprocessors provide a "software trap" instruction for each instruction set. Again, the microprocessor saves information about the current state of the machine before passing control to the exception handler. The terms interrupt, trap, fault and exception used herein are used interchangeably.
【0005】
In some microprocessors, when an interrupt occurs externally, the microprocessor passes control to the same interrupt handler entry point. If there are multiple external devices and they can activate the interrupt request read, the interrupt handler first determines which device caused the interrupt and then gives control to the part of the code that handles that particular device. Must be handed over. For example, the Intel 8048 microprocessor contains an INT input, which, when activated, allows the microcontroller to pass control to absolute memory location 3. The 8048 also includes a RESET input, which, when activated, causes the microcontroller to pass control to absolute memory location 0. It also includes an internal timer / counter, which transfers control to absolute memory location 7 when it triggers an interrupt.
【0006】
Other microprocessors include "interrupt level" reads in addition to interrupt request reads. In these microprocessors, when an external device activates an interrupt request read, it also sends a trap number unique to that particular device on the interrupt level line. In response, the microprocessor's internal hardware passes control, or "vector," to one of multiple interrupt handlers, each corresponding to a different trap number. Similarly, some microprocessors have a pre-defined entry point for every routine written to handle an internally raised exception, while others raise it. It provides a mechanism for automatically passing a vector to a routine that depends on a trap number defined for each particular type of possible internal exception.
【0007】
Conventionally, when a vector is passed to an interrupt handler and an exception handler, some methods have been used to determine the entry point of the corresponding handler. In one technique, an address table is created starting from a specific table base address, and the base address is fixed or user-definable. Each entry in the table was as long as the address, for example, 2 or 4 bytes long, and contained the entry point for the corresponding trap number. When an interrupt or exception occurs, the microprocessor first determines the base address of the table, then adds the trap number multiplied by m (where m is the number of bytes in each entry) and then asks. The information stored at the address was loaded into the program counter (PC), and control was passed to the routine starting from the address specified in the table entry.
【0008】
Other microprocessors stored the entire branch instruction in each entry in the table, rather than just storing the handler's address. The number of bytes in each entry was equal to the number of bytes in the branch instruction. Upon receiving an interrupt or exception, the microprocessor first determined the table-based address, added the trap number multiplied by m, and loaded the result into the program counter. The first instruction to be executed after that was a branch instruction in the table, and finally control was passed to the corresponding exception handler.
【0009】
[Problems to be Solved by the Invention]
In both of the above methods of passing a vector to the handler, there is a delay (waiting time) because the operation part of the handler cannot start execution until the preliminary operation is executed. With the first method, the entry point address could only be loaded into the program counter after it was first retrieved from the table. In the second method, the important part of the handler could not start executing until the preliminary branch instruction was fetched and executed. The delay in addition that occurs during the calculation of adding the trap number multiplied by m to the table base address concatenates the trap numbers with the high-order bit from the base address as the low-order bit, and then logs.<sub>2 </sub>It can be avoided by just continuing m zero bits, but the delay caused by the preliminary operation described above remained. Such delays can be a hindrance in systems where response time is important for processing certain types of interrupts.
【0010】
Another problem with exception handling in traditional microprocessors is the state of the machine when the trap handler returns to the main instruction flow. It is related to the amount of information that needs to be stored so that the machine) can be restored. There is a trade-off between the need to store as much information as possible and the need to minimize the delay in dispatching to trap handlers. In particular, for on-chip data registers, do not store any on-chip data registers, let the handler temporarily store the data in each register, and then leave it to the handler. Techniques were used to allow handlers to use them for their own purposes. In this case, the handler had to replace the data in the register before returning. In this case, these registers need to be stored and restored, which can significantly slow down the operation of the handler. Another approach is for the hardware to automatically store the contents of the registers on the stack before passing control to the handler. This method is also inappropriate because it complicates the hardware and can significantly delay the transfer of control to the handler. Therefore, in the case of the vector method described above, the delay caused by the existing method for protecting the contents of the register when the trap handler is called is unacceptable in a high-performance microprocessor.
【0011】
[Means for solving problems]
According to the present invention, a microprocessor architecture that solves many of the above-mentioned problems in known systems is adopted. In particular, the "fast trap" exception dispatch method is adopted. This technique allows the entire handler to be stored in a single vector address table entry. Each table entry has enough space to hold at least two, and preferably more, instructions, so when a fast trap occurs, the microprocessor uses the trap number multiplied by m as the base address. By concatenating, it is only necessary to branch to the obtained address. The time delay required to fetch the entry point address from the table or fetch and execute a preliminary branch instruction is eliminated. The microprocessor can also incorporate other time-inefficient vector techniques for less important types of traps.
【0012】
In another embodiment of the invention, when a trap appears, the processor goes into an interrupted state, automatically shifting some shadow registers to the foreground and the corresponding foreground register set in the background. Shift to. The register contents are not transferred, instead only the shadow register is made available in place of the regular register. Therefore, the handler will have a set of registers that can be used immediately without worrying about corrupting the data needed for the main instruction stream.
【0013】
In the above-mentioned US patent application No. 07 / 727,006 (PCT / JP92 / 00868), the title of the invention "High-Performance RISC Microprocessor Architecture" (Japanese Patent Publication No. 6-501122) Allows an advanced microprocessor that prefetches instructions before its execution to handle out-of-order returns of instruction prefetch requests, can execute more than one instruction during the same execution, and for a sequence of instructions in the instruction stream. It is explained that instructions can be executed out of order. Another embodiment of the invention incorporates a mechanism for maintaining the accuracy of synchronization exceptions raised for an instruction before and during the execution of the instruction.
【0014】
The microprocessor architecture described in the patent application also incorporates a mechanism for processing another procedure instruction flow called from a procedure instruction or emulation instruction in the main instruction flow. The transfer of control to the procedural instruction flow is done by a separate emulation prefetch queue without flushing any instructions already prefetched in the main instruction flow. According to another embodiment of the invention, the interrupt state remains available regardless of whether the processor is running from the main instruction stream or from the procedural instruction stream. It has a marker indicating which instruction stream to return to when returning from the trap. In addition, separate prefetch program counters are maintained for the main instruction stream and the emulation instruction stream, and the processor stores only the prefetch PC from the current instruction stream when the trap handler is called and when the handler returns. Restore the prefetch PC to the correct prefetch program counter.
【0015】
BEST MODE FOR CARRYING OUT THE INVENTION
Hereinafter, embodiments of the present invention will be described with reference to the drawings. I. Microprocessor Architecture Overview II. Instruction Fetch Unit A) IFU Data Path B) IFU Control Path C) IFU / IEU Control Interface D) PC Logic Unit Details 1) PF and ExPC Control / Data Unit Details 2) Details of PC control architecture E) Interrupt and exception handling 1) Overview 2) Asynchronous interrupt 3) Synchronous exception 4) Handler dispatch and return 5) Nest 6) Trapper list III. Instruction Execution Unit A) IEU Data Path Details 1) Register Array Unit Details 2) Integer Data Path Details 3) Floating Point Data Path Details 4) Boolean Register Data Path Details B) Load / Store Control Unit C) IEU control path details 1) E-decode unit details 2) Carry checker unit details 3) Data dependency checker unit details 4) Register renaming unit details 5) Instruction issuing unit details 6 ) Details of the completion control unit 7) Details of the backup control unit 8) Details of the control flow control unit 9) Details of the bypass control unit IV. Virtual memory control unit V. Cache control unit VI. Summary and conclusion.
【0016】
I. Overview of the microprocessor architecture Figure 1 shows an overview of the architecture 100 of the present invention. The instruction fetch unit (IFU) 102 and the instruction execution unit (IEU) 104 are the core functional elements of Architecture 100. The virtual memory unit (VMU) 108, cache control unit (CCU) 106, and memory control unit (MCU) 110 are intended to directly support the functions of IFU 102 and IEU 104. In addition, the memory array unit (MAU) 112 is for operating the architecture 100 as a basic element. However, MAU112 does not directly exist as an integral component of Architecture 100. That is, in a preferred embodiment of the invention, the IFU102, IEU104, VMU108, CCU106, and MCU110 are mounted on a single silicon chip utilizing a traditional 0.8 micron design rule low power CMOS process, with approximately 1,200,000 pieces. It is composed of transistors. The clock speed of a standard processor or system in Architecture 100 is 40 MHz. However, according to a preferred embodiment of the present invention, the internal clock speed of the processor is 160 MHz.
【0017】
The basic role of the IFU102 is to fetch an instruction, buffer the instruction while it is pending execution by IEU104, and typically calculate the next virtual address to be used when fetching the next instruction. That is.
【0018】
In a preferred embodiment of the invention, each instruction has a fixed length of 32 bits. An instruction set, or four-instruction "bucket," is simultaneously fetched by the IFU 102 from the instruction cache 132 in the CCU 106 via the 128-bit wide instruction bus 114. The instruction set transfer is coordinated by the control signal sent via control line 116 and takes place between IFU 102 and CCU 106. The virtual address of the fetched instruction set is output from IFU102 via IFU arbitration, control and address sharing bus 118, and further sent out on the arbitration, control and address sharing bus 120 that connects IEU104 and VMU108. Arbitration of access to VMU108 is done because both IFU102 and IEU104 use VMU108 as a common shared resource. In a preferred embodiment of the invention, the low-order bits that define the address within the physical page of the virtual address are transferred directly from the IFU 102 via the control line 116 to the cache control unit (CCU) 106. The virtual high-order bits of the virtual address given by IFU102 are sent to VMU108 by the address portion of the arbitration, control and address sharing buses 118 and 120, where they are translated into the corresponding physical page addresses. On the IFU102, this physical page address is sent directly from the VMU108 to the cache control unit (CCU) 106 via control line 122 during half the internal processor clock cycle after the translation request is issued to the VMU108. Transferred.
【0019】
The instruction stream fetched by IFU102 is passed to IEU104 via instruction stream bus 124. Control signals are exchanged between IFU102 and IEU104 via control line 126. In addition, some instruction fetch addresses, such as those that require access to the register array unit residing in IEU104, are sent back to IFU102 via the target address return bus in control line 126. Is done.
【0020】
The IEU104 stores data to and from the data cache 134 provided in the CCU 106 through the 80-bit wide bidirectional data bus 130 and retrieves the data. The entire physical address when IEU104 accesses data is passed to CCU106 by the address portion of control bus 128. In addition, control signals for managing data transfer can be exchanged between the IEU 104 and the CCU 106 through the control bus 128. IEU104 uses VMU108 as a resource to change the virtual data address to a suitable physical data address to pass to CCU106. The virtualized portion of the data address is passed to VMU108 via arbitration, control and address sharing bus 120. Unlike operations on IFU102, VMU108 returns the corresponding physical address to IEU104 via arbitration, control and address sharing bus 120. In a preferred embodiment of Architecture 100, IEU104 uses physical addresses to ensure that load / store operations are performed in the correct program stream order.
【0021】
The CCU 106 has a conventional high-level function for determining whether a data request defined by a physical address can be satisfied from either the instruction cache 132 or the data cache 134, whichever is applicable. If the access request is correctly satisfied by accessing the instruction cache 132 or the data cache 134, the CCU 106 coordinates the data transfer via the buses 114 and 130 to perform the transfer.
【0022】
If the data access request defined by the physical address is not satisfied by either the instruction cache 132 or the data cache 134, the CCU 106 passes the corresponding physical address to the MCU 110 and the MAU 112 sends the CCU 106 source or destination for each request. Sufficient control information to identify whether the caches 132, 134 are requesting read or write access, and to relate the final data request issued by IFU102 or IEU104 to the request operation. Pass the additional identification information of.
【0023】
The MCU110 preferably comprises a port switch unit 142, which is connected to the instruction cache 132 of the CCU106 by the unidirectional data bus 136 and to the data cache 134 by the bidirectional data bus 138. Has been done. The port switch unit 142 is basically a large multiplexer, and the physical address obtained from the control bus 140 is applied to multiple ports (P).<sub>0 </sub>-P<sub>n </sub>)146<sub>0-n </sub>It enables sending to any of the ports and bidirectional transfer of data from the port to data buses 136, 138. Each memory access request processed by the MCU 110 is for the purpose of arbitrating access to the main system memory bus 162, which is requested when accessing the MAU 112, on port 146.<sub>0-n </sub>Associated with one of. When the data transfer connection is established, the MCU 110 passes the control information to the CCU 106 via the control bus 140 to port 146.<sub>0-n </sub>Initiates the transfer of data between the instruction cache 132 or the data cache 134 and the MAU 112 via the corresponding one of them. In a preferred embodiment of Architecture 100, MCU110 does not actually store or latch data in the process of transferring between CCU106 and MAU112. This is done to minimize transfer latency and avoid tracking or managing only one piece of data on the MCU110.
【0024】
II. Instruction fetch unit The main elements of the instruction fetch unit 102 are shown in Figure 2. In order to make it easier to understand the operations and interrelationships of these elements, the cases where these elements are involved in the IFU data path and control path will be described below.
【0025】
A) IFU data path The IFU data path begins on instruction bus 114, which receives an instruction set and temporarily stores it in the prefetch buffer unit 260. The instruction set from the prefetch buffer unit 260 is passed through the I decode unit 262 to the IFIFO unit 264. The instruction set stored in the last two stages of the IFIFO unit 264 can be continuously retrieved and used by the IEU104 through the output buses 278 and 280.
【0026】
The prefetch buffer unit 260 receives one instruction set from instruction bus 114 at a time. The complete 128-bit wide instruction set is typically written in parallel to one of the four 128-bit wide prefetch buffer locations in the 188 portion of the main buffer (MBUF) of prefetch buffer unit 260. Up to four additional instruction sets can be added to the prefetch buffer location of two 128-bit wide target buffers (TBUF) 190, or to the prefetch buffer of two 128-bit wide procedure buffers (EBUF) 192. -It is possible to write to the location. In preferred architecture 100, an instruction set located at either the prefetch buffer location in MBUF188, TBUF190 or EBUF192 can be forwarded to the prefetch buffer output bus 196. Further, the flow-through bus 194 is for bypassing the MBUF188, TBUF190 and EBUF192 by connecting the instruction bus 114 directly to the prefetch buffer output bus 196.
【0027】
In preferred architecture 100, MBUF188 is utilized to buffer the instruction set nominally or in the main instruction stream. The TBUF190 is used to buffer the prefetched instruction set from the trial target branch instruction stream. As a result, both instruction streams that may be placed after the conditional branch instruction can be prefetched through the prefetch buffer unit 260. This feature eliminates at least subsequent access latency to the CCU106, even if the MAU112 waits longer, so it is conditional regardless of which instruction stream is ultimately selected when resolving a conditional branch instruction. The correct next set of instructions placed after the attached branch instruction can be obtained and executed. In the preferred architecture 100 of the present invention, the presence of MBUF188 and TBUF190 allows the instruction fetch unit 102 to prefetch both possible instruction streams, as described below in connection with the instruction execution unit 104. As you can, you can continue to execute the instruction stream that was supposed to be correct. If the correct instruction stream is prefetched and put into MBUF188 when the conditional branch instruction is resolved, the instruction set remaining in TBUF190 is only invalidated. On the other hand, if the instruction set of the correct instruction stream exists in the TBUF190, these instruction sets are transferred directly from the TBUF190 and in parallel to their respective buffer locations in the MBUF188 through the instruction prefetch buffer unit 260. The instruction set previously stored in MBUF188 is effectively invalidated by overwriting the instruction set transferred from TBUF190. If there is no TBUF instruction set to forward to an MBUF location, that location will only be marked invalid.
【0028】
Similarly, EBUF192 is another alternative prefetch route via the prefetch buffer unit 260. The EBUF192 is preferably used to prefetch a single instruction appearing in the MBUF188 instruction stream, an alternative instruction stream used to implement the operation specified in the "procedure" instruction. In this way, complex or extended instructions can be implemented through software routines or procedures and processed through the prefetch buffer unit 260 without disturbing the instruction stream already prefetched and placed in the MBUF188. be able to. In general, according to the invention, the procedure instruction that first appeared in TBUF190 can be processed, but the prefetch of the procedure instruction stream is deferred and all previously appearing pending conditional branch instruction streams are deferred. It will be resolved. This ensures that conditional branch instructions appearing in the procedural instruction stream are processed consistently through the use of TBUF190. Therefore, if the procedure stream is branched, the target instruction set has already been prefetched and placed in TBUF190 and can be forwarded in parallel to EBUF192.
【0029】
Finally, each of the MBUF188, TBUF190, and EBUF192 is connected to the prefetch buffer output bus 196 to send the instruction set stored by the prefetch buffer unit 260 onto the prefetch buffer output bus 196. Further, the flow-through bus 194 is for transferring the instruction set directly from the instruction bus 114 to the prefetch buffer output bus 196.
【0030】
In preferred architecture 100, the prefetch buffers in MBUF188, TBUF190, and EBUF192 do not directly form a FIFO structure. Instead, any buffer location is connected to the prefetch buffer output bus 196, which gives a great deal of freedom in the prefetch order of the instruction set fetched from the instruction cache 132. That is, it is common for the instruction fetch unit 102 to determine and request an instruction set in the order of instructions arranged in a fixed order in the instruction stream. However, the order in which the instruction set is returned to IFU102 is tailored to the case where one instruction set requested is available and accessible only from CCU106, while other instruction sets require access to MAU112. , It is also possible to appear out of order.
【0031】
Even though the instruction set may not be returned to the prefetch buffer unit 260 in a certain order, the sequence of instruction sets output on the prefetch buffer output bus 196 is generally the instruction set request issued by the IFU 102. Must follow the order of. This is because the in-order instruction stream sequence is affected, for example, by the trial execution of the target branch stream.
【0032】
The I-decode unit 262 receives an instruction set from the prefetch buffer output bus 196, usually at a rate of one per cycle, as space in the IFIFO unit 264 allows. Each set of four instructions that make up one instruction set is decoded in parallel by the I-decoding unit 262. The contents of the instruction set are not modified by I-decode unit 262 while the relevant control flow information is extracted from line 318 for the control path portion of IFU 102.
【0033】
The instruction set from the I-decode unit 262 is sent over the 128-bit wide input bus 198 of the IFIFO unit 264. Internally, the IFIFO unit 264 consists of columns of master / slave registers 200, 204, 208, 212, 216, 220, 224. Each register is connected to its successor register and the contents of master registers 200, 208, 216 are transferred to slave registers 204, 212, 220 during the first half of the internal processor cycle of the FIFO operation, and then during the second half of the operation cycle. It is to be transferred to the next subsequent master register 208, 216, 224. Input bus 198 is now connected to the inputs of master registers 200, 208, 216, and 224, and the instruction set is now loaded directly from the I-decode unit 262 into the master register during the second half cycle of FIFO operation. There is. However, loading the master register from input bus 198 does not have to be done at the same time as the FIFO shift of the data in the IFIFO unit 264. As a result, regardless of the current depth of the instruction set stored in IFIFO unit 264, and independent of FIFO shifting of data within IFIFO unit 264, IFIFO unit 264 continuously from input bus 198. You can put it in.
【0034】
Each of the master / slave registers 200, 204, 208, 212, 216, 220, and 224 can store all bits of a 128-bit wide instruction set in parallel, as well as some bits of control information in their respective control registers. It can also be stored in 202, 206, 210, 214, 218, 222, 226. Preferably, the set of control bits consists of exception miss and exception modify (VMU), no memory (MCU), branch bias, stream, and offset (IFU). This control information is generated from the control path portion of the IFU 102 at the same time that the new instruction set is loaded from the input bus 198 into the IFIFO master register. The control register information is then shifted in parallel within the IFIFO unit 264 in parallel with the instruction set.
【0035】
Finally, in preferred architecture 100, the instruction set output from IFIFO unit 264 is obtained simultaneously from the last two master registers 216, 224, and I_Bucket.<u style="single"></u>0 and I<u style="single"></u>Bucket<u style="single"></u>It is sent on the output buses 278 and 280 of one instruction set. In addition, the corresponding control register information is sent on the output buses 282, 284 of the IBASV0 and IBASV1 control fields. These output buses 278, 282, 280, and 284 are all instruction stream buses 124 leading to IEU104.
【0036】
B) IFU control path The IFU102 control path directly supports the operations of prefetch buffer unit 260, I decode unit 262 and IFIFO unit 264. The prefetch control logic unit 266 mainly manages the operation of the prefetch buffer unit 260. Prefetch control logic units 266 and IFU 102 typically receive system clock signals from clock line 290 to synchronize IFU operations with IEU104, CCU106, and VMU108 operations. Control signals for selecting an instruction set and writing to MBUF188, TBUF190 and EBUF192 are sent over control line 304.
【0037】
A large number of control signals are sent over control line 316 and to prefetch control logic unit 266. Specifically, the fetch request control signal is sent to initiate a prefetch operation. Other control signals sent over control line 316 specify whether the requested prefetch operation targets MBUF188, TBUF190, or EBUF192. Upon receiving the prefetch request, the prefetch control logic unit 266 generates an ID value and determines whether the prefetch request can be notified to the CCU 106. The ID value is generated using a circular 4-bit counter.
【0038】
The use of 4-bit counters is important in three ways: The first is that up to nine instruction sets can be activated at once by the prefetch buffer unit 260. That is, a set of 4 instructions on the MBUF188, a set of 2 instructions on the TBUF190, a set of 2 instructions on the EBUF192, and a set of 1 instruction that is passed directly to the I decode unit 262 via the flow-through bus 194. The second is that the instruction set consists of four instructions, each of which is 4 bytes. As a result, the least significant 4 bits of any address that selects the instruction to fetch are extra. Finally, the prefetch request ID can be easily associated with the prefetch request by inserting it as the least significant 4 bits of the prefetch request address. This will reduce the total number of addresses required to interface with the CCU 106.
【0039】
In order to allow the instruction set to be returned from the CCU106 out of order with respect to the order of the prefetch requests issued by the IFU102, architecture 100 now returns the ID request value along with the instruction set returned from the CCU106. It has become. However, according to the out-of-order instruction set return function, 16 unique IDs may be used up. If the conditional instruction combination is executed out of order, the ID value can be reused because of the additional prefetch and instruction set that was requested but not yet returned. Therefore, it is preferable to keep the 4-bit counter, and the prefetch request of the instruction set after that is not issued. In that case, the next ID value is the fetch request that remains unprocessed. And then associated with another instruction set held in the prefetch buffer unit 260.
【0040】
The prefetch control logic unit 266 directly manages the status register array (array) 268, which consists of status storage locations that logically correspond to the instruction set prefetch buffer locations in MBUF188, TBUF190, and EBUF192. ing. Prefetch control logic unit 266 can scan, read, and write data to status register array 268 through selection and data line 306. Within the status register array 268, the main buffer register 308 has four 4-bit ID values (MB ID), four 1-bit reserved flags (MB RES), and four 1-bit valid flags (MB VAL). Each of these is associated with each instruction set storage location in the MBUF180 by logical bit position. Similarly, the target buffer register 310 and the extended buffer register 312 each have two 4-bit ID values (TBID, EB ID) and two 1-bit reserved flags (TB RES, EB). It is for storing RES) and two 1-bit valid flags (TB VAL, EB VAL). Finally, the flow-through status register 314 stores one 4-bit ID value (FT TD), one reserved flag bit (FT RES), and one valid flag bit (FT VAL). Is for.
【0041】
The status register array 268 is first scanned and, if applicable, updated by prefetch control logic unit 266 each time a prefetch request is issued to CCU106, and then scanned and updated each time an instruction set is returned. To. Specifically, upon receiving a prefetch request signal from control line 316, the prefetch control logic unit 266 increments the current circular counter generation ID value and scans the status register array 268 to determine the available ID values. Is it possible to determine if the type of prefetch buffer location specified in the prefetch request signal is available, check the state of control line 300 in CCU IBUSY, and allow CCU 106 to accept the prefetch request? If it is acceptable, the CCU IREAD control signal on the control line 298 is affirmed, and the incremented ID value is connected to the CCU 106. Send on the output line 294 of the ID. The prefetch storage location can be used when both the corresponding reservation status flag and the validity status flag are false. The prefetch ID is written to the ID storage location in status register array 268, which corresponds to the target storage location in MBUF188, TBUF190, or EBUF192, in parallel with the request being issued to CCU106. In addition, the corresponding reservation status flag is truly set.
【0042】
When the CCU 106 can return the previously requested instruction set to the IFU 102, the CCU IREADY signal is affirmed on the control line 302 and the corresponding instruction set ID is sent on the output line 296 of the CCU ID. The prefetch control logic unit 266 scans the ID value and the reserved flag in the status register array 268 to determine the target destination of the instruction set in the prefetch buffer unit 260. Only one match is possible. Once determined, the instruction set is written to the appropriate location in prefetch buffer unit 260 via instruction bus 114, and when determined to be a flow-through request, it is passed directly to I-decode unit 262. .. In both cases, the valid status flag contained in the corresponding status register array 268 is set to true.
【0043】
The PC logic unit 270 examines the entire IFU102 to find the virtual addresses of the MBUF188, TBUF190, and EBUF192 instruction streams, as described in detail below. When performing this function, the PC logic unit 270 controls the I-decode unit 262 and operates from it at the same time. Specifically, the instruction portion decoded by the I-decoding unit 262 and which may be related to the change in the flow of the instruction stream of the program is sent to the control flow detection unit 274 via the bus 318 and directly. Is sent to the PC logic unit 270. The control flow detection unit 274 sets each instruction that constitutes a control flow instruction including a conditional branch instruction and an unconditional branch instruction, a call type instruction, a software trap procedure instruction, and various return instructions in a decoded instruction set. Determine from the inside. The control flow detection unit 274 sends a control signal to the PC logic unit 270 via line 322. This control signal indicates the location and type of control flow instructions within the instruction set present in I-decode unit 262. In response to this, the PC logic unit 270 generally determines the target address of the control flow instruction from the data put into the instruction and transferred to the PC logic unit 270 via line 318. For example, if a branch logic bias is selected to execute a conditional branch instruction first, PC logic unit 270 will instruct it to prefetch the instruction set from the conditional branch instruction target address. , Start tracking separately. Therefore, the next affirmation of the prefetch request on control line 316 is assumed that the PC logic unit 270 further affirms the control signal via control line 316 and the preceding prefetch instruction set is sent to MBUF188 or EBUF192. Then, the prefetch destination is selected as TBUF190. Priff If the prefetch control logic unit 266 determines that the etch request can be passed to the CCU 106, the prefetch control logic unit 266 again sends an enable signal to the PC logic unit 270 via control line 316. It allows the page offset portion of the target address (CCU PADDR [13: 4]) to be passed directly to CCU 106 via address line 324. At the same time, the PC logic unit 270 further sends the VMU request signal via control line 328 to the virtualized portion of the target address (VMU VADDR) if a new virtual page to physical page conversion is required. [13:14]) is passed to VMU108 via address line 326 and translated into a physical address. If no page conversion is required, no VMU108 operation is required. Instead, the previous conversion results are stored in the output latch connected to control line 122 and are used immediately by the CCU 106. 14]) is passed to VMU108 via address line 326 and converted to a physical address. If no page conversion is required, no VMU108 operation is required. Instead, the previous conversion results are stored in the output latch connected to control line 122 and are used immediately by the CCU 106. 14]) is passed to VMU108 via address line 326 and converted to a physical address. If no page conversion is required, no VMU108 operation is required. Instead, the previous conversion results are stored in the output latch connected to control line 122 and are used immediately by the CCU 106.
【0044】
Operational errors in the VMU 108 during the virtual-to-physical conversion requested by the PC logic unit 270 are reported through VMU exceptions and VMU miss control lines 332, 334. VMU mismatch control line 334 reports a translation lookaside buffer (TLB) mismatch. The VMU exception control signal on VMU exception line 332 occurs when other exceptions occur. In either case, the PC logic unit 270 stores the current execution location in the instruction stream and then receives it to diagnose the error condition, just as an unconditional branch was made. Dedicated exception handling routine for processing Error conditions are processed by prefetching the instruction stream. The VMU exception and mismatch control signals indicate the type of exception that occurred so that the PC logic unit 270 can determine the prefetch address of the corresponding exception handling routine.
【0045】
IFIFO control logic unit 272 is for direct support of IFIFO unit 264. Specifically, the PC logic unit 270 outputs the control signal via the control line 336 and the IFIFO control logic unit 272 indicates that the instruction set is available from the I decode unit 262 via the input bus 198. Notify to. The IFIFO control logic unit 272 is responsible for selecting the innermost available master registers 200, 208, 216, 224 to receive the instruction set. The outputs of the master control registers 202, 210, 218, and 226 are passed to the IFIFO control logic unit 272 via the control bus 338. The control bits stored by each master control register consist of a 2-bit buffer address (IF_Bx_ADR), a single stream indicator bit (IF_Bx_STRM), and a single valid bit (IF_Bx_VLD). The 2-bit buffer address specifies the first valid instruction in the corresponding instruction set. That is, the instruction set returned by CCU106 may not be bound so that, for example, the target instruction for a branch operation is placed at the first instruction location in the instruction set. Therefore, the buffer address value is given to uniquely indicate the first instruction in the instruction set to be considered for execution.
【0046】
The stream bit indicates the location of the instruction set containing the conditional control flow instructions and is based on being used as a marker to cause potential control flow changes in the stream of instructions through IFIFO unit 264. The main instruction stream is generally processed through MBUF188 when the stream bit value is 0. For example, when a relative conditional branch instruction appears, the corresponding instruction set is marked with a stream bit value of 1. The conditional instruction set is detected by the I decode unit 262. Up to four conditional control flow instructions can exist in an instruction set. It is then stored in the innermost available master register of IFIFO unit 264 when the instruction is set.
【0047】
To determine the target address of a conditional branch instruction, the current IEU104 execution point address (DPC), the relative location of the instruction set containing the conditional instruction specified by the stream bit, control flow detection unit 274 The conditional instruction location offset in the instruction set obtained from is combined with the relative branch offset value obtained from the corresponding branch instruction field through control line 318. The result is the virtual address of the branch target, which is stored by PC Logic Unit 270. The first instruction set in the target instruction stream can be prefetched into TBUF190 using this address. Depending on the branch bias preselected for PC logic unit 270, IFIFO unit 264 will continue to load from MBUF188 or TBUF190. When a second instruction set appears that contains one or more conditional flow instructions, the instruction set is marked with a 0 for the stream bit value. The second target stream cannot be fetched, so the target address is calculated and stored by PC logic unit 270, but not prefetched. In addition, subsequent instruction sets cannot be processed through the I-decode unit 262. At the very least, no instruction set found to contain conditional flow control instructions is processed.
【0048】
In a preferred embodiment of the invention, the PC logic unit 270 can manage up to eight conditional flow instructions appearing in up to two instruction sets. Each target address in the two instruction sets marked with a change in stream bit is stored in an array of four address registers, and the target address is for the location of the corresponding conditional flow instruction in the instruction set. Is placed in a logical position.
【0049】
When the branch result of the first in-order conditional flow instruction is resolved, PC logic unit 270 now forwards the contents of TBUF190 to MBUF188 and marks the contents of TBUF190 as invalid when the branch occurs. , Instruct prefetch control logic unit 266 by a control signal on control line 316. If the IFIFO unit 264 has an incorrect instruction stream, that is, an instruction set from the target stream if it is not branched, or from the main instruction stream if it is branched, it is cleared from IFIFO unit 264. If a second or subsequent conditional flow control instruction is in the instruction set marked with the first stream bit, the instruction is processed in a unified way. That is, the instruction set from the target stream is prefetched, the instruction set from MBUF188 or TBUF190 is processed through I-decode unit 262 according to the branch bias, and correctly when the conditional flow instruction is finally resolved. No stream instruction set is cleared from IFIFO unit 264.
【0050】
When an incorrect stream instruction is cleared from IFIFO unit 264, the second conditional flow instruction remains in IFIFO unit 264 and the first conditional flow instruction set does not contain any subsequent conditional flow instructions. Then, the target address of the instruction set marked with the second stream bit is promoted to the first array of address registers. In either case, the next set of instructions containing conditional flow instructions can be evaluated through the I-decode unit 262. Therefore, using the stream bit as a toggle is used for the purpose of calculating the branch target address, and if the branch bias is later determined to be incorrect for a particular conditional flow control instruction. Potential control flow changes can be marked and tracked through IFIFO unit 264 for the purpose of marking instruction set locations that should be cleared above.
【0051】
Instead of actually clearing the instruction set from the master register, IFIFO control logic unit 272 simply resets the significant bit flag in the control register of the corresponding master register of IFIFO unit 264. This clear operation is initiated by the PC logic unit 270 with a control signal sent to line 336. Each input of master control registers 202, 210, 218, 226 can be accessed directly by IFIFO control logic unit 272 through status bus 230. In Architecture 100 of the preferred embodiment, the bits in these master control registers 202, 210, 218, 226 are set by the IFIFO control logic unit 272 in parallel with or independently of the data shift operation by the IFIFO unit 264. It is possible to do. This feature allows the instruction set to be written to one of the master registers 200, 208, 216, 224 and the corresponding status information to the master control registers 202, 210, 218, 226 asynchronously with the IEU104 operation. it can.
【0052】
Finally, an additional control line on the control and status bus 230 enables and directs IFIFO operation of IFIFO unit 264. The IFIFO shift is performed by the IFIFO unit 264 in response to the shift request control signal output from the PC logic unit 270 through the control line 336. The IFIFO control logic unit 272 sends a control signal to the prefetch control logic unit 266 via control line 316 to prefetch when the master registers 200, 208, 216, 224 that accept the instruction set are available. -Requests the transfer of the next corresponding instruction set from the buffer unit 260. When the instruction set is transferred, the corresponding valid bits in status register array 268 are reset.
【0053】
C) IFU / IEU control interface The control interface connecting IFU102 and IEU104 is provided by control bus 126. This control bus 126 is connected to the PC logic unit 270 and consists of a plurality of controls, addresses and special data lines. By passing the interrupt request and the reception confirmation control signal via the control line 340, the IFU 102 can notify the interrupt operation and synchronize with the IEU 104. The interrupt signal generated externally is sent to the PC logic unit 270 via the interrupt line 292. In response to this, when an interrupt request control signal is sent on line 340, IEU104 cancels the trially executed instruction. Information about the contents of the interrupt is exchanged through the interrupt information line 341. When the IEU104 is ready to start receiving prefetched instructions from the address of the interrupt service routine determined by the PC logic unit 270, the IEU104 affirms the interrupt reception confirmation control signal on the line 340. The interrupt service routine prefetched by IFU102 is then started.
【0054】
The IFIFO read (IFIFO RD) control signal is output from the IEU 104 onto control line 342, notifying that the instruction set in the innermost master register 224 has completed execution and that the next instruction set is needed. .. Upon receiving this control signal, the PC logic unit 270 instructs the IFIFO control logic unit 272 to perform an IFIFO shift operation on the IFIFO unit 264.
【0055】
A PC increment request and size value (PC INC / SIZE) are sent over control line 344 to instruct the PC logic unit 270 to update the current program counter value by the corresponding size of the instruction. This allows the PC logic unit 270 to maintain an execution program counter (DPC) at the exact location of the first in-order execution instruction in the current program instruction stream.
【0056】
The target address (TARGET ADDR) is returned to PC logic unit 270 via address line 346. This target address is the virtual target address of the branch instruction, which is determined by the data stored in the register array of IEU104. Therefore, IEU104 operations are required to calculate the target address.
【0057】
Control Flow Result (CF RESULT) The control signal is sent via control line 348 to PC Logic Unit 270 to resolve the currently pending conditional branch instruction and whether the result is due to the branch. , Indicates that it does not depend on the branch. Based on these control signals, the PC logic unit 270 must cancel any of the instruction sets located in the prefetch buffer unit 260 and the IFIFO unit 264 as a result of executing conditional flow instructions. Can be determined.
【0058】
Several IEU instruction return type control signals (IEU returns) are sent over control line 350 to notify IFU102 that an instruction has been executed by IEU104. These instructions include returns from procedural instructions, returns from traps, and returns from subroutine calls. The return instruction from the trap is used in the hardware interrupt processing routine and the software trap processing routine in the same way. Returns from subroutine calls are also used with jumps and linked calls. In each case, the return control signal is sent to inform the IFU102 to resume the instruction fetch operation for the previously interrupted instruction stream. By issuing these signals from IEU104, the accurate operation of Architecture 100 can be maintained. The "interrupted" instruction stream is restarted from the point where the return instruction is executed.
【0059】
The current instruction execution PC address (currently IF_PC) is sent to IEU104 via the address bus 352. This address value (DPC) specifies the exact instruction executed by IEU104. That is, IEU104 is the current IF<u style="single"></u>While the instruction that passed the PC address is being tried first, this address is used for interrupts, exceptions, and other events that require the exact state of the machine to be known. Must be held for precise control. When IEU104 determines that it is possible to advance the exact state of the machine in the currently executing instruction stream, a PC Inc / Size signal is sent to IFU102, which is immediately reflected in the current IF_PC address value.
【0060】
Finally, the address and bidirectional data bus 354 is for transferring data in special registers. This data can be programmed by IEU104 to be put into or read from a special register in IFU102. Special register data is generally loaded or calculated by IEU104 for use by IFU102.
【0061】
D) Details of the PC logic unit A detailed view of the PC logic unit 270 including the PC control unit 362, the interrupt control unit 363, the prefetch PC control unit 364 and the execution PC control unit 366 is shown in FIG. PC control unit 362 receives control signals from prefetch control logic unit 266 and IFIFO control logic unit 272 through control line 316 and from IEU104 through interface bus 126 to prefetch and execute PC control units 364, 366. On the other hand, timing control is performed. The interrupt control unit 363 is responsible for accurate management of interrupts and exceptions, including determining the prefetch trap address offset and selecting the appropriate processing routine to handle each trap type. The prefetch PC control unit 364, in particular, of the program counters required to support buffers 188, 190, 192, including storing return addresses for trap processing and procedural routine instruction flow. Responsible for management. To support this operation, the prefetch PC control unit 364 is responsible for generating a prefetch virtual address that includes the CCU PADDER address on the physical address bus line 324 and the VMU VMADDR address on the address line 326. As a result, the prefetch PC control unit 364 is responsible for holding the current prefetch PC virtual address value.
【0062】
The prefetch operation is generally initiated by IFIFO control logic unit 272 through a control signal sent over control line 316. In response to this, the PC control unit 362 generates some control signals and outputs them on the control line 372, operates the prefetch PC control unit 364, and needs the PADDR address on the address lines 324 and 326. Generate a VMADDR address according to. Increment signals with values from 0 to 4 may also be sent on control line 374, which means that PC control unit 362 is re-fetching the instruction set from the current prefetch address or a series. It depends on whether you are aligning to the second request in the prefetch request or selecting the next full sequential instruction set for prefetch. Finally, the current prefetch address PF_PC is sent over bus 370 and passed to execution PC control unit 366.
【0063】
New prefetch addresses come from several sources. The main source of the address is the current IF sent from the running PC control unit 366 via bus 352.<u style="single"></u>It is a PC address. In principle, the IF_PC address yields the return address, which is later used by the prefetch PC control unit 364 when an initial call, trap, or procedure instruction appears. The IF_PC address is stored in a register in the prefetch PC control unit 364 each time these instructions appear. In this way, when the PC control unit 362 receives an IEU return signal through control line 350, it simply has to select the return address register in the prefetch PC control unit 364 to retrieve the new prefetch virtual address. Resume the original program instruction stream.
【0064】
Another source of prefetch addresses is the target address value sent from the execution PC control unit 366 via the relative target address bus 382 or from the IEU 104 via the absolute target address bus 346. is there. The relative target address is an address that can be calculated directly by the execution PC control unit 366. Absolute target addresses need to be generated by IEU104 as these target addresses depend on the data contained in the IEU register array. The target address is sent to the prefetch PC control unit 364 via the target address bus 384 and used as the prefetch virtual address. When calculating the relative target address, the operand portion of the corresponding branch instruction is also sent from the I decode unit 262 via the operand displacement portion of bus 318.
【0065】
Another source of prefetch virtual addresses is the execution PC control unit 366. The return address bus 352'is for forwarding the current IF_PC value (DPC) to the prefetch PC control unit 364. This address is used as the return address where control flow instructions such as interrupts, traps, and other calls appear in the instruction stream. Prefetch PC control unit 364 is released to prefetch a new instruction stream. The PC control unit 362 receives an IEU return signal from IEU 104 via line 350 when the corresponding interrupt or trap processing routine or subroutine is executed. On the other hand, the PC control unit 362 selects the register containing the current return virtual address based on the ID of the return instruction sent and executed through one of the PFPC signals on line 372 and over line 350. To do. This address is then used to continue the prefetch operation by PC logic unit 270.
【0066】
Finally, another source from which the prefetch virtual address is retrieved is the special register address and data bus 354. The address value calculated or loaded by IEU104, or at least the base address value, is transferred as data to the prefetch PC control unit 364 via bus 354. The base address contains the addresses of the trap address table, the fast trap table, and the base procedure instruction dispatch table. Many of the registers in the prefetch and execute PC control units 364 and 366 can also be read through bus 354, allowing the corresponding aspects of machine state to be processed through IEU104.
【0067】
The execution PC control unit 366 is controlled by the PC control unit 362 and is currently IF.<u style="single"></u>Its main role is to calculate the PC address value. In this role, the execution PC control unit 366 receives the control signal sent from the PC control unit 362 via the ExPC control line 378 and the increment / size control signal sent via the control line 380. , Adjust the IF_PC address. These control signals are mainly generated upon receiving the IFIFO read control signal sent via line 342 and the PC increment / size value sent from IEU104 via control line 344.
【0068】
1) Detailed view of PF and ExPC control / data units FIG. 4 is a detailed block diagram of prefetch and execution PC control units 364 and 366. These units consist primarily of registers, incrementers and other similar components, selectors and adder blocks. Control to manage data transfer between these blocks is provided by the PC control unit 362 through PFPC control lines 372, ExPC control lines 378 and increment control lines 374, 380. For the sake of clarity, the block diagram of FIG. 4 does not show these individual control lines. However, it goes without saying that these control signals are sent to these blocks, as described below.
【0069】
Central to the prefetch PC control unit 364 is the prefetch selector (PF_PC SEL) 390, which acts as the central selector for the current prefetch virtual address. This current prefetch address is sent from the prefetch selector 390 to the incremental unit 394 via the output bus 392 to generate the next prefetch address. This next prefetch address is sent through the incremental output bus 396 to a parallel array of registers MBUF PFnPC398, TBUF PFnPC400, and EBUF PFnPC402. These registers 398, 400, 402 effectively store the following instruction prefetch addresses, but according to a preferred embodiment of the invention, separate prefetch addresses are held in MBUF188, TBUF190, and EBUF192. Has been done. MBUF, TBUF and EBUF Prefetch addresses stored in PFnPC registers 398, 400, 402 are passed from address buses 404, 408, 410 to prefetch selector 390. Therefore, the PC control unit 362 can instruct the immediate switching of the prefetch instruction stream only by instructing the prefetch selector 390 to select another one of the prefetch registers 398, 400, 402. When the address value is incremented by the incremental 394 to prefetch the next set of instructions in the stream, the value is returned to the corresponding register of the prefetch addresses 398, 400, 402. The other parallel register array is shown as a single special register block 412 for brevity, but this array is for storing some special addresses. Register block 412 contains a trap return address register, a procedure instruction return address register, a procedure instruction dispatch table-based address register, a trap routine dispatch table-based address register, and a fast trap. -Consists of a routine-based address register. Under the control of PC control unit 362, these return address registers can accept the current IF_PC execution address through bus 352. The return in register block 412 and the address value stored in the base address register can be read and written independently of IEU104. A register is selected and the value is transferred via the special register address and data bus 354.
【0070】
The selector in the special register block 412 is controlled by the PC control unit 362, and the address stored in the register of the register block 412 can be sent on the special register output bus 416 and passed to the prefetch selector 390. The return address is passed directly to the prefetch selector 390. The base address value is combined with the offset value sent from the interrupt control unit 363 via the interrupt offset bus 373. The special address passed from the source to the prefetch selector 390 via bus 373 is used as the initial address of the new prefetch instruction stream, and then at the address through incrementer 394 and one of the prefetch registers 398, 400, 402. The increment loop can continue.
【0071】
Another source of addresses sent to the prefetch selector 390 is the register array in the target address register block 414. The target register in block 414 stores eight potential branch target addresses, according to a preferred embodiment. These eight storage locations logically correspond to the eight potentially executable instructions held in the lowest two master registers 216 and 224 of IFIFO unit 264. Target register block 414 stores the precomputed target address because any of these instructions, and potentially all, can be conditional branch instructions, so the target instruction stream through TBUF190. Can be made to wait to be used to prefetch. In particular, if a conditional branch bias is set such PC control unit 362 initiates a prefetch target instruction stream immediately, the target address target register via the address bus 418 from the data block 414 Sent to prefetch selector 390. After being incremented by incrementer 394, the address is returned to the TBUF PFnPC400 for storage and later used in prefetching the target instruction stream. When another branch instruction appears in the target instruction stream, the target address of that second branch is calculated and stored in the target register array 414 until the first conditional branch instruction is resolved and used. ing.
【0072】
The calculated target address stored in the target register block 414 is the absolute target address from the target address calculation unit in the execution PC control unit 366 via address line 382 or from IEU104. Transferred via bus 346.
【0073】
The address value transferred through the prefetch selector (PF_PC SEL) 390 is a complete 32-bit virtual address value. In a preferred embodiment of the present invention, the page size is fixed at 16 Kbytes and corresponds to the maximum page offset address value [13: 0]. Therefore, if the current prefetch virtual page address [27:14] does not change, no VMU page translation is required. The comparator in the prefetch selector 390 detects this. The VMU conversion request signal (VMXLAT) is sent to the PC via line 372 when the virtual page address changes, either because the increment has crossed the page boundary or because the control flow has branched to another page address. It is sent to the control unit 362. On the other hand, the PC logic unit 270 has a VM in addition to the CCU PADDR on line 324. The VADDR address is sent from the buffer unit 420 onto the address line 326 and the corresponding control signal is sent over the VMU control line 328, instructing it to obtain a translation from the VMU virtual page to the physical page. If no page translation is required, the current physical page address [31:14] is held by the output-side latch of VMU108 on control line 122.
【0074】
The virtual address sent on the bus 370 receives the signal sent from the increment control line 374 and is incremented by the incrementer 394. Incrementer 394 increments by the value representing the instruction set (4 instructions or 16 bytes) to select the next instruction set. The lower 4 bits of the prefetch address passed to CCU106 are zero. Therefore, the actual target address instruction in the first branch target instruction set may not be located at the first instruction location. However, since the lower 4 bits of the address are sent to the PC control unit 362, the IFU 102 can determine the location of the first branch instruction. The detection and processing to return the low-order bits [3: 2] of the target address as a 2-bit buffer address and select the correct first instruction to execute from the unaligned target instruction set is a new instruction. Only done during the first prefetch of the stream, that is, the first non-sequential instruction set address in the instruction stream. The non-alignment relationship between the address of the first instruction in the instruction set and the prefetch address used when prefetching the instruction set can be ignored for the life of the current sequential instruction stream. It will be ignored.
【0075】
The remaining part of the functional block shown in Fig. 4 constitutes the execution PC control unit 366. According to a preferred embodiment of the present invention, the execution PC control unit 366 independently includes a program counter incrementer that functions independently. At the heart of this feature is the Execution Selector (DPC SEL) 430. The address output from the execution selector 430 on the address bus 352'is the current execution address (DPC) of architecture 100. This execution address is sent to the addition unit 434. The increment / size control signal sent on line 380 specifies an instruction increment value from 1 to 4, which is added to the address obtained from the execution selector 430 by the addition unit 434. Each time the addition unit 434 performs the output latching function, the next incremented execution address is returned directly to the execution selector 430 via address line 436 for use in the next instruction increment cycle.
【0076】
The initial execution address and all subsequent new stream addresses are obtained from the new stream register unit 438 via address line 440. The new stream register unit 438 also passes the new current prefetch address sent from the prefetch selector 390 via the PFPC address bus 370 directly to the address bus 440, or stores it for later use. You can also keep it. That is, if the prefetch PC control unit 364 decides to start prefetching from the new virtual address, the new stream address is temporarily stored by the new stream register unit 438. By participating in both the prefetch and execution increment cycles, the PC control unit 362 registers the new stream address in the new stream register until the execution address reaches the program execution point corresponding to the control flow instruction that started the new instruction stream. Leave it at 438. The new stream address is then output from the new stream register unit 438 and sent to the execution selector 430 to start independently generating the execution address in the new instruction stream.
【0077】
According to a preferred embodiment of the present invention, the new stream register unit 438 has a function of buffering two control flow instruction target addresses. By extracting the new stream address immediately, the execution PC control unit 366 can be switched from the generation of the current execution address string to the generation of the new execution address stream string with almost no waiting time.
【0078】
Finally, the IF_PC selector (IF_PC SEL) 442 is for finally sending the current IF_PC address over address bus 352 and sending it to IEU104. The input to IF_PC selector 442 is the output address obtained from execution selector 430 or new stream register unit 438. In most cases, the IF_PC selector 442 receives an instruction from the PC control unit 362 and selects the execution address output from the execution selector 430. However, to further reduce the latency of switching to the new virtual address used to start the execution of the new instruction stream, bypass the selected address from the new stream register unit 438 and via bus 440. It can be sent directly to the IF_PC selector 442 and obtained as the current IF_PC execution address.
【0079】
Execution PC control unit 366 has a function to calculate all relative branch target addresses. The current execution point address and the address obtained from the new stream register unit 438 are passed to the control flow selector (CF_PC) 446 via the address buses 352'and 440. As a result, the PC control unit 362 can select the exact initial address on which the target address calculation is based, with great flexibility. This initial address, the base address, is sent over address bus 454 to target address ALU450. Another input value to the target ALU450 is sent from the control flow displacement calculation unit 452 via bus 458. The relative branch instruction contains a displacement value in the form of an immediate mode constant with a new relative target address, according to a preferred embodiment of Architecture 100. The control flow displacement calculation unit 452 receives the operand displacement value obtained for the first time from the operand output bus 318 of the I decoding unit. Finally, the offset register value is sent over line 456 to target address ALU450. The offset register 448 receives the offset value from the PC control unit 362 via the control line 378'. The magnitude of the offset value is determined by the PC control unit 362 based on the address offset from the base address sent over address line 454 to the address of the current branch instruction when calculating the relative target address. That is, the PC control unit 362 is currently processing the instruction at the current execution point address (requested by CP_PC) and the I decode unit 262 by controlling the IFIFO control logic unit 272, and thus the PC logic unit. The number of instructions separating the instructions being processed by the 270 is tracked to determine the target address of the instructions.
【0080】
When the relative target address is calculated by the target address ALU450, the target address is written to the corresponding target register 414 through address bus 382.
【0081】
2) Details of PC control algorithm 1. Processing of main instruction stream: MBUF PFnPCl.1 The address of the next main flow prefetch instruction is stored in MBUF PFnPC. 1.2 In the absence of control flow instructions, the 32-bit incremental adjusts the address value contained in the MBUF PFnPC by 16 bytes (x16) for each prefetch cycle. 1.3 When an unconditional control flow instruction is I-decoded, all prefetch data fetched following the instruction set is flushed and the MBUF PFnPC has the target register unit, PF.<u style="single"></u>The new main instruction stream address is loaded through the PC selector and incremental. The new address is also stored in the new stream register. 1.3.1 The target address of the relative unconditional control flow is calculated by the IFU from the register data held by the IFU and from the operand data placed after the control flow instruction. 1.3.2 The target address of an absolutely unconditional control flow is finally calculated by the IEU from register reference values, base register values, and index register values. 1.3.2.1 The instruction prefetch cycle is stopped until the target address is returned from the IEU for the absolute address control flow instruction. The instruction execution cycle continues. 1.4 The address of the next main flow prefetch instruction obtained from the unconditional control flow instruction is bypassed, sent via the target address register unit, PF_PC selector and incremental, and finally stored in MBUFPFnPC. And prefetching continues from 1.2.
【0082】
2. Processing the procedure instruction stream: EBUF PFnPC2.1 Procedure instructions are prefetched in the main or branch target instruction stream. If fetched in the target stream, stop prefetching the procedure stream until the conditional control flow instruction is resolved and the procedure instruction is transferred to the MBUF. This allows TBUF to be used when processing conditional control flows that appear in the procedural instruction stream. 2.1.1 Procedure instructions must not be placed in the procedure instruction stream. That is, procedure instructions must not be nested. Upon returning from the procedure instruction, execution returns to the main instruction stream. Another dedicated return from the nested procedure instruction is required to allow nesting. The architecture can easily support this type of instruction, but the ability to nest procedural instructions is unlikely to improve the performance of the architecture. 2.1.2 In the main instruction stream, a procedural instruction stream containing an instruction set containing first and second conditional control flow instructions is resolved by the conditional control flow instructions in the first instruction set and the second conditional control. Stop prefetching for the second conditional control flow instruction set until the flow instruction set is transferred to the MBUF. 2.2 The procedure instruction indicates the starting address of the procedure routine by the relative offset included as the instruction's immediate mode operand field. 2.2.1 The offset value obtained from the procedure instruction is combined with the value stored in the procedure base address (PBR) register maintained in the IFU. This PBR register is readable and writable through the special address and data bus when the special register move instruction is executed. 2.3 When a procedure instruction appears, the next main instruction stream IF<u style="single"></u>The PC address is stored in the DPC return address register and the procedure-in-progress bit in the processor status register (PSR) is set. 2.4 The starting address of the procedure stream is sent from the PBR register (plus the procedure instruction operand offset value) to the PF_PC selector. 2.5 The starting address of the procedure stream is sent to the new stream register unit and increment at the same time, incrementing by (x16). The incremented address is then stored in the EBUF PFnPC. 2.6 In the absence of control flow instructions, the 32-bit incrementer adjusts the address value contained in the EBUF PFnPC by (x16) for each procedure instruction prefetch cycle. 2.7 When the unconditional control flow instruction is I-decoded, all prefetch data fetishized after the branch instruction is flushed and EBUF The PFnPC is loaded with the new procedure instruction stream address. 2.7.1 The target address of a relative unconditional control flow instruction is calculated by the IFU from the register data held in the IFU and from the operand data contained in the immediate mode operand field of the control flow instruction. 2.7.2 The target address of an absolutely unconditional branch is calculated by the IEU from register reference values, base register values and index register values. 2.7.2.1 The instruction prefetch cycle is stopped for the absolute address branch until the target address is returned by the IEU. The execution cycle continues. 2.8 The address of the next set of procedure prefetch instructions is stored in the EBUF PFnPC and prefetching continues from 1.2. 2.9 When the return from the procedure instruction is I decoded, prefetching continues from the address stored in the uPC register, then incremented by (x16), and then MBUF for prefetching later. Returned to the PFnPC register.
【0083】
3. Branch instruction stream processing: TBUF PFnPC3.1 When the conditional control flow instruction that appears in the first instruction set in the MBUF instruction stream is I-decoded, the target address is relative to the current address. If it is an address, it is judged by IFU, and if it is an absolute address, it is judged by IEU. 3.2 For branching bias: 3.2.1 If the branch is to an absolute address, stop the instruction prefetch cycle until the target address is returned from the IEU. The execution cycle continues. 3.2.2 PF<u style="single"></u>Load the branch target address into the TBUF PFnPC by forwarding via the PC selector and incremental. 3.2.3 The target instruction stream is prefetched and placed in TBUF before being sent to IFIFO for execution. Stop prefetching when IFIFO and TBUF are full. 3.2.4 32-bit Incrementer Adjust the address value contained in the TBUF PFnPC by (x16) for each prefetch cycle. 3.2.5 When the conditional control flow instruction appearing in the second instruction set in the target instruction stream is I-decoded, the prefetch operation is resolved and all the conditional branch instructions in the first (main) set are resolved. Stop until (but go ahead, calculate the relative target address and store it in the target register). 3.2.6 If the conditional branch in the first instruction set is interpreted as "doing": 3.2.6.1 If the source of the branch is the EBUF instruction set determined from the procedure in progress bits, flush the instruction set placed after the first conditional flow instruction set contained in MBUF or EBUF. 3.2.6.2 Transfer the TBUFPFnPC value to the MBUF PFnPC or EBUF based on the state of the procedure in progress bit. 3.2.6.3 Transfer the prefetched TBUF instruction to MBUF or EBUF based on the state of the procedure in progress bit. 3.2.6.4 If the second conditional branch instruction set is not I-decoded, continue the MBUF or EBUF prefetch operation based on the state of the procedure in progress bit. 3.2.6.5 If the second conditional branch instruction is I-decoded, start processing that instruction (go to step 3.3.1. 3.2.7 When interpreted as "do not perform" conditional control over the instructions in the first conditional instruction set: 3.2.7.1 Flush the instruction set and instruction IFIFO and IEU from the target instruction stream. 3.2.7.2 Continue MBUF or EBUF prefetch operation. 3.3 "Bias not branching": 3.3.1 Stop prefetching instructions into MBUF. Continue the execution cycle. 3.3.1.1 If the conditional control flow instructions in the first conditional instruction set are relative, calculate the target address and store it in the target register. 3.3.1.2 If the conditional control flow instructions in the first conditional instruction set are absolute, wait until the IEU calculates the target address and returns that address in the target register. 3.3.1.3 When the conditional control flow instruction in the second instruction set is I-decoded, the prefetch operation is stopped until the conditional control flow instruction in the first conditional instruction set is resolved. 3.3.2 TBUF when the target address of the first conditional branch is calculated Loads into the PFnPC and starts prefetching instructions into the TBUF in parallel with the execution of the main instruction stream. The target instruction set is not loaded (thus branch target instructions are provided when each conditional control flow instruction in the first instruction set is resolved). 3.3.3 If the conditional control flow instructions in the first set are interpreted as "performed": 3.3.3.1 When the state of the procedure in progress bit determines that the source of the branch is an EBUF instruction stream, Flush the MBUF or EBUF and flush the IFIFO and IEU of instructions from the main instruction stream placed after the first conditional branch instruction set. 3.3.3.2 Transfer the TBUF PFnPC value to MBUF or EBUF as determined from the state of the procedure in progress bit. 3.3.3.3 Transfer the prefetched TBUF instruction to MBUF or EBUF as judged from the state of the procedure in progress bit. 3.3.3.4 Continue the MBUF or EBUF prefetch operation as determined by the state of the procedure in progress bit. 3.3.4 If the conditional control flow instructions in the first set are parsed as "not done": 3.3.4.1 Flush the instruction set TBUF from the target instruction stream. 3.3.4.2 If the second conditional branch instruction is not I-decoded, continue the MBUF or EBUF prefetch operation as determined by the state of the procedure in progress bit. 3.3.4.3 If the second conditional branch instruction is I-decoded, start processing the instruction (go to step 3.4.1).
【0084】
4. Interrupts, exceptions and trap instructions 4.1 Traps consist of the following in a broad sense. 4.1.1 Hardware interrupt 4.1.1.1 Asynchronous (external) event, internal or external. 4.1.1.2 Occurs and persists at any time. 4.1.1.3 Receive services in priority order between atomic (normal) instructions and suspend procedure instructions. 4.1.1.4 The start address of the interrupt handler is determined as the vector number offset to the predefined table of the trap handler entry point. 4.1.2 Software Trap Instruction 4.1.2.1 Asynchronous (external) generated instruction. 4.1.2.2 Software instructions executed as an exception. 4.1.2.3 The trap handler start address is determined from the trap number offset combined with the base address value stored in the TBR or FTB register. 4.1.3 Exception 4.1.3.1 Event that occurs in synchronization with the instruction. 4.1.3.2 Processed when the instruction is executed. 4.1.3.3 The instruction expected by the result of the exception and all subsequent execution instructions are cancelled. 4.1.3.4 The start address of the exception handler is determined from the trap number offset to the predefined table of the trap handler entry point. 4.2 The trap instruction stream operation is executed inline with the instruction stream currently being executed. 4.3 Traps can be nested, provided that the trap handling routine saves the xPC address before the next interruptible trap. Otherwise, if a trap appears before the current trap operation is complete, the state of the machine will be corrupted.
【0085】
5. Processing of trap instruction stream: xPC5.1 When a trap appears: 5.1.1 When an asynchronous interrupt occurs, the execution of the instruction currently being executed is suspended. 5.1.2 When a synchronization exception occurs, the trap is handled when the instruction that caused the exception is executed. 5.2 When a trap is processed: 5.2.1 Interrupts are disabled. 5.2.2 Current IF<u style="single"></u>The PC address is stored in the xPC trap state return address register. 5.2.3 IFIFO and MBUF prefetch buffers at the IF_PC address and subsequent addresses are flushed. 5.2.4 The executed instruction of address IF_PC and subsequent addresses and the result of that instruction are flushed from the IEU. 5.2.5 MBUF PFnPC is loaded with the address of the trap handler routine. 5.2.5.1 The trap source addresses the TBR or FTB registers, depending on the trap type determined by the trap number in the special register group. 5.2.6 Instructions are prefetched and put into IFIFO for normal execution. 5.2.7 The instructions of the trap routine are then executed. 5.2.7.1 The trap handling routine has the ability to save the xPC address to a given location, enabling interrupts again. The xPC register is a special register move instruction and is read and written through the special register address and data bus. 5.2.8 It is necessary to get out of the trap state by executing the return from the trap instruction. If saved prior to 5.2.8.1, the xPC address must be restored from its predefined location before returning from the trap instruction. 5.3 When a return from a trap instruction is executed: 5.3.1 Interrupts are enabled. 5.3.2 The xPC address is returned to the current instruction stream register MBUF or EBUF PFnPC, as determined from the state of the procedure in progress bit, and prefetching continues from that address. 5.3.3 xPC address is restored to IF_PC register through new stream register.
【0086】
E) Handling interrupts and exceptions 1) Overview Interrupts and exceptions are handled regardless of whether the processor is running from the main instruction stream or the procedural instruction stream, as long as they are enabled. Interrupts and exceptions are serviced in priority order and persist until cleared. The trap handler start address is determined as a vector number offset to the trap handler's predefined table, as described below.
【0087】
There are basically two types of interrupts and exceptions in this embodiment. That is, one that is triggered synchronously with a specific instruction in the instruction stream and one that is triggered asynchronously with a specific instruction in the instruction stream. The terms interrupt, exception, trap and fault are used interchangeably herein. Asynchronous interrupts are caused by on-chip or off-chip hardware that is not operating in sync with the instruction stream. For example, interrupts triggered by on-chip timers / counters are hardware interrupts or non-maskable interrupts triggered by off-chips. Like interrupt) (NMI), it is asynchronous. When an asynchronous interrupt is triggered, the processor context is frozen, all traps are interrupt disabled, some processor status information is stored, and the processor sends a vector to the interrupt handler that corresponds to the particular interrupt it receives. Turn. When the interrupt handler completes its processing, program execution continues from the instruction placed after the last completed instruction in the stream that was being executed when the interrupt occurred.
【0088】
A synchronization exception is an exception that is triggered in synchronization with an instruction in the instruction stream. These exceptions are raised in connection with a particular instruction and are suspended until the instruction in question is executed. In a preferred embodiment, synchronization exceptions are triggered during prefetch, instruction decoding, or instruction execution. Prefetch exceptions include, for example, TLB mismatches and other VMU exceptions. A decode exception is raised, for example, if the instruction being decoded is an illegal instruction or does not match the current privilege level of the processor. Execution exceptions are caused by arithmetic errors, such as division by zero. When these exceptions occur, in a preferred embodiment, the specific instruction that caused the exception is associated with the exception, and the state is maintained until the instruction is retired. At that point, all previously completed instructions are saved, and any trial results from the instruction that caused the exception are flushed, similar to the trial results of subsequent instructions that were tried and executed. Control is then passed to the exception handler that corresponds to the highest priority exception raised by the instruction.
【0089】
Software trap instructions are detected by the control flow detection unit (CF_DET) 274 (Figure 2) at the I-decode stage and are treated like unconditional call instructions and other synchronous traps. That is, the target address is calculated and prefetching continues up to the current prefetching queue (EBUF or MBUF). At the same time, the exception is recorded in association with the instruction and processed when the instruction is saved. All other types of synchronization exceptions are recorded, accumulated, and processed at run time in association with the particular instruction that caused the exception.
【0090】
2) Asynchronous interrupt Asynchronous interrupt is notified to PC logic unit 270 through interrupt line 292. As shown in Figure 3, these lines are for notifying the interrupt control unit 363 in the PC logic unit 270 and consist of an NMI line, an IRQ line and a set of interrupt level lines (LVL). .. The NMI line notifies non-maskable interrupts and originates from an external source. This is the highest priority interrupt except for a hardware reset. The IRQ line also originates from an external source and notifies when an external device requests a hardware interrupt. In a preferred embodiment, the user can define up to 32 externally generated hardware interrupts, and the specific external device that requested the interrupt has an interrupt number (0-31) on the interrupt level line (LVL). Send out. The memory error line is activated by the MCU110 to signal various types of memory errors. Other asynchronous interrupt lines (not shown) are also provided to notify the interrupt control unit 363. These have lines for requesting timer / counter interrupts, memory input / output (I / O) error interrupts, machine check interrupts, and performance monitor interrupts. Each asynchronous interrupt is associated with a corresponding predefined trap number, similar to the synchronization exception described below. 32 of these trap numbers are associated with 32 hardware interrupt levels. A table of these trap numbers is maintained in interrupt control unit 363. In general, the higher the trap number, the higher the priority of the trap.
【0091】
When one of the asynchronous interrupts is notified to the interrupt control unit 363, the interrupt control unit 363 sends an interrupt request to the IEU 104 via the INT REQ / ACK line 340. Further, the interrupt control unit 363 transmits a prefetch pause signal to the PC control unit 362 via the line 343, and causes the PC control unit 362 to stop prefetching the instruction. IEU104 cancels all instructions in progress at that time and either aborts all trial results or completes some or all instructions. In the preferred embodiment, the response to the asynchronous interrupt is speeded up by canceling all the instructions being executed at that time. In either case, the DPC in the execution PC control unit 366 is updated to correspond to the last completed and saved instruction before IEU104 confirms the reception of the interrupt. All other instructions that have been prefetched and placed in MBUF, EBUF, TBUF and IFIFO unit 264 are also canceled.
【0092】
The IEU104 sends an interrupt reception confirmation signal back to the interrupt control unit 363 via the INT REQ / ACK line 340 only when it is ready to receive an interrupt from the interrupt handler. Upon receiving this signal, the interrupt control unit 363 dispatches to the corresponding trap handler, as described below.
【0093】
3) Synchronous exception In the case of synchronous exception, the interrupt control unit 363 has a set of four internal exception indicator bits (not shown) for each instruction set, and each bit is associated with each instruction in the set. ing. The interrupt control unit 363 also maintains a trap number to notify when it is found by each instruction.
【0094】
If a VMU notifies a TLB mismatch or another VMU exception while a particular instruction set is being prefetched, this information goes to the PC logic unit 270, especially to the interrupt control unit 363, via the VMU control lines 332, 334. Will be sent. Upon receiving this signal, the interrupt control unit 363 notifies the PC control unit 362 via the line 343 to suspend the subsequent prefetch. At the same time, the interrupt control unit 363 is the VM associated with the prefetch buffer to which the instruction set is sent.<u style="single"></u>Set either the Miss or the VM_Excp bit, whichever is applicable. After that, the interrupt control unit 363 sets all four internal exception indicator bits corresponding to the instruction set because none of the instructions in the instruction set are valid, and 4 in the instruction set that caused the problem. Stores the trap number of the specific exception received corresponding to each instruction. Shifting and executing instructions prior to the problematic instruction continues normally until the problematic instruction set reaches the lowest level within IFIFO unit 264.
【0095】
Similarly, if another synchronization exception is detected while shifting instructions through prefetch buffer unit 260, I-decode unit 262, or IFIFO unit 264, this information is also sent to interrupt control unit 363 for interrupts. The control unit 363 sets the internal exception indicator bit corresponding to the instruction that caused the exception and stores the trap number corresponding to the exception. As with prefetch synchronization exceptions, the shift and execution of instructions prior to the problematic instruction continues normally until the problematic instruction set reaches the lowest level within IFIFO unit 264.
【0096】
In a preferred embodiment, only one type of software trap instruction is detected while shifting instructions through the prefetch buffer unit 260, the I decode unit 262, or the IFIFO unit 264. The software trap instruction is detected at the I decoding stage by the control flow detection unit (CF_DET) 274. In some embodiments, other forms of synchronization exceptions are detected in the I-decode stage, but detection of other synchronization exceptions preferably waits until the instruction arrives at the instruction execution unit 104. This prevents certain exceptions that occur when processing privileged instructions from being signaled based on processor states that may change before the instructions are effectively executed in order. To. Processor state-independent exceptions, such as illegal instructions, can be detected at the I-decode stage, but at a minimum, all pre-execution synchronization exceptions (apart from VMU exceptions) should be detected with the same logic. Hardware will suffice. Also, since handling of such exceptions is rarely time-sensitive, there is no wasted time waiting for an instruction to reach the instruction execution unit 104.
【0097】
As mentioned above, the software trap instruction is detected by the control flow detection unit (CF_DET) 274 at the I decoding stage. The internal exception indicator bit corresponding to that instruction in the interrupt control unit 363 is set, and the software trap number that can be specified in the immediate mode field of the software trap instruction with a number from 0 to 127 is associated with the trap instruction. Will be stored. However, unlike prefetch synchronization exceptions, software traps are treated as synchronization exceptions as well as control flow instructions, so interrupt control unit 363 should suspend prefetch when a software trap instruction is detected. Do not notify control unit 362. Instead, IFU102 prefetches the trap handler into the MBUF instruction stream buffer at the same time that the instruction is shifted through IFIFO unit 264.
【0098】
When the instruction set reaches the lowest level of IFIFO unit 264, interrupt control unit 363 sets the exception indicator bit of that instruction set to a 4-bit vector and SYNCH.<u style="single"></u>INT<u style="single"></u>It is sent to IEU104 via INFO line 341, and if there is an instruction in the instruction set that has already been determined to be the source of the synchronization exception, it is notified which instruction it is. IEU104 does not respond immediately, allowing all instructions in the instruction set to be scheduled in the usual way. Other exceptions, such as integer arithmetic exceptions, may be triggered at run time. Exceptions that depend on the current state of the machine, such as exceptions caused by the execution of privileged instructions, are also detected at this point so that the state of the machine is up to date for all previous instructions in the instruction stream. In order to do so, all instructions that may affect the PSR (such as special moves and returns from trap instructions) are forced to be executed in order. The interrupt control unit 363 is notified that an exception has occurred only when the instruction that is the source of some synchronization instruction is about to be saved.
【0099】
IEU104 saves all instructions that appear in the instruction stream that precedes the first instruction that was tried to execute and caused a synchronization exception, and then trials from the instructions that appeared later in the instruction stream and were executed on a trial basis. Flush the results. The particular instruction that caused the exception is usually re-executed when returning from the trap, so this instruction is also flushed. After that, IF_PC in the execution PC control unit 366 is updated to correspond to the last instruction actually saved, and the exception is notified to the interrupt control unit 363.
【0100】
When the instruction that is the source of the exception is saved, the IEU104 will indicate which instruction, if any, in the saved instruction set (register 224) has caused a synchronization exception. The vector is returned to interrupt control unit 363 via SYNCH_INT_INFO line 341 with information indicating the source of the first exception in the instruction set. The information contained in the 4-bit exception vector returned from IEU104 is the accumulation of the 4-bit exception vector passed from interrupt control unit 363 to IEU104 and the exceptions raised by IEU104. If there is information already stored in interrupt control unit 363 due to an exception detected during prefetch or I decoding, the rest of the information returned from IEU104 to interrupt control unit 363 along with that information is interrupt control. Unit 363 is sufficient to determine the content of the highest priority synchronization exception and its trap number.
【0101】
4) After the handler dispatch and return interrupt reception confirmation signal is received from IEU104 via line 340, or a non-zero exception vector is received via line 341, the current DPC is the return address of the special register. It is temporarily stored in the xPC register, which is one of blocks 412 (Fig. 4). The current processor status register (PSR) is also stored in the previous PSR (PPSR) register, and the current status comparison register (CSR) is saved in the old status comparison register (PCSR) in the special register block 412.
【0102】
The trap handler address is calculated as the trap base register address plus an offset. The PC logic unit 270 has two base registers for trapping, both of which are part of the special register block 412 (Figure 4), which is initialized by a previously executed special move instruction. For most traps, the base register used to compute the handler's address is the trap base register TBR.
【0103】
The interrupt control unit 363 determines the currently pending highest priority interrupt or exception and, through the look-up table, determines the trap number associated with it. This is passed to the prefetch PC control unit 364 via a set of INT_OFFSET lines 373 as an offset to the selected base register. The vector address has the advantage that it can be obtained simply by concatenating the offset bit as the low-order bit to the high-order bit obtained from the TBR register. Therefore, the delay of the adder is prevented. (In this specification, the 2'bit is the i'th bit.) For example, if the trap number is from 0 to 255 and this is represented by an 8-bit value, the handler address is 8 bits. It is required to concatenate the trap number at the end of the 22-bit TBR store value. Adding the two-digit low-order bit to the trap number ensures that the trap handler address is always on the word boundary. The concatenated handler address created in this way is a prefetch selector (PF) as one of inputs 373.<u style="single"></u>It is sent to PC SEL) 390 (Figure 4), selected as the next address, from which the instructions are prefetched. All trap vector handler addresses using the TBR register are separated by one word. Therefore, the instruction at the trap handler address must be a preliminary branch instruction to the lengthened trap handling routine. However, there are some traps that need to be handled with care to prevent system performance degradation. For example, TLB traps need to run fast. For that reason, preferred embodiments incorporate a fast trap mechanism that allows small trap handlers to be called without paying for the preliminary branch. In addition, fast trap handlers can be placed independently in memory, for example in on-chip ROM, eliminating memory system problems related to RAM location.
【0104】
In a preferred embodiment, the only trap that becomes a high-speed trap is the VMU exception described above. Fast trap numbers are distinguished from other traps and range from 0 to 7. However, the priority is the same as the MMU exception. When the interrupt control unit 363 recognizes that the fast trap is the highest priority pending at that time, it selects the fast trap base register (FTB) from the special register (FTB) and combines it with the trap offset. Send on line 416. Prefetch selector (PF) via line 373'<u style="single"></u>The resulting vector address sent to PC SEL) 390 is a concatenation of the upper 22 bits from the FTB register, followed by 3 bits representing the fast trap number, followed by 7 zeros. The bit continues. Therefore, each fast trap address is 128 bytes, or 32 words apart. When called, the processor branches to the starting word and causes the program to run within a block or branch out of it. Execution of a small program, such as a standard TLB processing routine that can be achieved with 32 or less instructions, is faster than a normal trap because it avoids a preliminary branch to the actual execution processing routine. ..
【0105】
In a preferred embodiment, all instructions have the same 4-byte length (that is, they occupy 4 address locations), but it is worth noting that even a microprocessor with variable instructions will have a fast trap. The mechanism is available. In this case, of course, there should be enough space between the fast trap vector addresses to accommodate at least two, preferably 32, average size instructions of the shortest length that the microprocessor can use. .. Of course, if the microprocessor has a return instruction from the trap, there must be enough space between the vector addresses to allow at least one other instruction in the handler to be placed on that instruction. There is.
【0106】
Also, when dispatched to a trap handler, the processor enters kernel mode and interrupt state. In parallel, a copy of the status register (CSR) is placed in the previous carry state register (PCSR) and a copy of the PSR is stored in the previous PSR (PPSR). The kernel and interrupt state modes are represented by bits in the processor status register (PSR). When the interrupt state bit of the current PSR is set, the shadow register or trap registers RT [24] to RT [31] become visible as shown above and in FIG. 7 (B). An interrupt handler can exit kernel mode by simply writing a new mode to the PSR, but the only way to exit an interrupt state is to execute a return (RTT) instruction from the trap.
【0107】
When the IEU104 executes the RTT instruction, the PCSR is restored to the CSR register and the PPSR register is restored to the PSR register, so that the interrupt state bits in the PSR are automatically cleared. The prefetch selector (PF_PC SEL) 390 selects the address stored in the special register xPC in the special register block 412 as the address to be prefetched from it. The xPC is restored to either the MBUF PFnPC or the EBUF PFnPC through the incremental 394 and the bus 396. The decision whether to restore xPC to EBUF PFnPC or MBUF PFnPC is made according to the "procedure in progress" bits of the PSR after it has been restored.
【0108】
It should be noted that the processor does not use the same special register xPC to store the return addresses of both trap and procedural instructions. The return address of the trap is stored in the special register xPC as described above, but the address to be returned after the procedure instruction is stored in another special register uPC. Therefore, the interrupt state remains available while the processor is executing the emulation stream called by the procedural instruction. Exception handling routines, on the other hand, must not contain any procedural instructions, as there is no special register to store the address to return to the exception handler after the emulation stream completes.
【0109】
5) Nesting Certain processor status information is automatically dispatched to trap handlers, especially CSR, PSR, return PCs, and in a sense the "A" register set ra [64] to ra [31]. It is backed up, but other contextual information is not protected. For example, the contents of the Floating Point Status Register (FSR) are not automatically backed up. In order for the trap handler to modify these registers, it must perform its own backup.
【0110】
Trap nesting is not automatic because the backups that are automatically performed when dispatching to the trap handler are limited. The trap handler needs to back up the required registers, clear the interrupt conditions, read the information required for trap processing from the system registers, and process that information appropriately. Interrupts are automatically disabled when dispatched to a trap handler. At the end of the process, the handler can restore the backed up registers, enable interrupts again, and execute RTT instructions to return from the interrupts.
【0111】
To enable nested traps, the trap handler must be split into a first part and a second part. In the first part, while interrupts are disabled, you need to use a special register move instruction to copy the xPC and push it onto the stack maintained by the trap handler. Next, you need to use a special register move instruction to move the start address of the second part of the trap handler to xPC and execute the return instruction (RTT) from the trap. RTT removes the interrupt state (by restoring PPSR to PSR) and transfers control to the address in xPC. xPC contains the address of the second part of the handler. The second part allows interrupts at this point and allows exception handling to continue in interruptable mode. It should be noted that the shadow registers RT [24] to RT [31] can only be seen in the first part of this handler, not in the second part. Therefore, in the second part, the handler needs to reserve the value of the "A" register if it is likely to be changed by the handler. When the trapping routine is finished, it restores all the backed up registers, pops the original xPC from the trap handler stub, returns it to the xPC special register using the special register move instruction, and another. You need to run RTT. This returns control to the corresponding instruction in the main or emulation instruction stream.
【0112】
6) Trap list The following [Table 1] shows the trap numbers, priorities, and processing modes of the traps recognized in the preferred embodiments.
【0113】
[Table 1] Trap number Processing mode Asynchronous / synchronous Trap name 0-127 Normal synchronous Trap instruction 128 Normal synchronous FP Exception 129 Normal synchronous Integer arithmetic operation Exception 130 Normal synchronous MMU (excluding TLB mismatch or correction) 135 Normal synchronous unaligned memory Address 136 Normal Synchronous Illegal Instruction 137 Normal Synchronous Privileged Instruction 138 Normal Synchronous Debug Exception 144 Normal Asynchronous Performance Monitor 145 Normal Asynchronous Timer / Counter 146 Normal Asynchronous Memory I / O Error 160-191 Normal Asynchronous Hardware Interrupt 192-253 Reserved 254 Normal Asynchronous Machine Check 255 Normal Asynchronous NMI0 Fast Trap Synchronous Fast MMU TLB Mismatch 1 fast trap sync High Speed MMU TLB Correction 2-3 High Speed Trare Synchronous High Speed (Reserved) 4-7 High Speed Trap Synchronous High Speed (Reserved) [0114]
III. Instruction Execution Unit Figure 5 shows the control path part and data path part of IEU104. The main data path begins with the instruction / operand data bus from IFU102. As a data bus, immediate operands are sent to operand alignment unit 470 and passed to register array (REG ARRAY) unit 472. Register data is from register array unit 472 via register array output bus 476, bypass unit 474, and functional unit distribution bus 480 to functional compute elements (FU).<sub>0-n </sub>) Functional unit 478<sub>0-n </sub>Is sent to a parallel array of. Functional unit 478<sub>0-n </sub>The data generated by is sent back to bypass unit 474 and / or register array unit 472 via output bus 482.
【0115】
Load / store unit 484 completes the data path portion of IEU104. Load / store unit 484 is responsible for managing the data transfer between IEU104 and CCU106. Specifically, the load data fetched from the data cache 134 of the CCU 106 is transferred by the load / store unit 484 to the register array unit 472 via the load data bus 486. The data stored in the data cache of CCU106 is received from the functional unit distribution bus 480.
【0116】
The control path portion of the IEU104 is responsible for sending, managing, and processing information through the IEU data path. In a preferred embodiment of the present invention, the IEU control path has a function of managing the parallel execution of a plurality of instructions, and the IEU data path has a function of independently performing a plurality of data transfers between almost all data path elements of the IEU 104. It has. When the IEU control path receives an instruction via the instruction stream bus 124, it operates accordingly. Specifically, the instruction set is received by the E-decoding unit 490. In a preferred embodiment of the invention, the E-decoding unit 490 receives and decodes both instruction sets held in IFIFO master registers 216 and 224. The results of decoding all eight instructions are: carry checker (CRY CHKR) unit 492, dependency checker (DEP CHKR) unit 494, register renaming unit (REG RENAME) 496, instruction issuer (ISSUER) unit 498, and save control unit ( RETIRE CLT) Sent to 500.
【0117】
Carry checker unit 492 receives decoding information from the E-decoding unit 490 via control line 502 for the eight pending instructions. The function of the carry checker unit 492 is to identify the instructions on hold that affect the carry bit of the processor status word or depend on the state of the carry bit. This control information is sent to the instruction issuing unit 498 via the control line 504.
【0118】
The decoding information indicating the registers of register array unit 472 used by the eight pending instructions is sent directly to register renaming unit 496 via control line 506. This information is also sent to dependency checker unit 494. The function of dependency checker unit 494 determines which of the pending instructions references a register as the destination of data, and which instruction, if any, depends on any of these destination registers. That is. Register-dependent instructions are identified by a control signal sent to register renaming unit 496 via control line 506.
【0119】
Finally, the E-decoding unit 490 sends control information identifying the specific contents and functions of each of the eight pending instructions to the instruction issuing unit 498 via the control line 510. The instruction issuing unit 498 is responsible for determining which functional units can be used to execute data path resources, especially pending instructions. According to a preferred embodiment of architecture 100, the instruction issuing unit 498 executes any of the eight hold state instructions out of order, subject to the availability of data path resources and the constraints of carry and register dependencies. It can be so. The register renaming unit 496 sends a bitmap of the instruction, which is appropriately unconstrained so that it can be executed, to the instruction issuing unit 498 via the control line 512. Instructions that have already been executed (completed) and that depend on registers or carry are logically excluded from the bitmap.
【0120】
Required functional unit 478<sub>0-n </sub>The instruction issuing unit 498 can start executing multiple instructions in each system clock cycle, depending on whether is available. Functional unit 478<sub>0-n </sub>The status of is sent to the instruction issuing unit 498 via the status bus 514. The control signal for starting the execution of the instruction and performing the execution management after the start is sent from the instruction issuing unit 498 to the register renaming unit 496 via the control line 516, and selectively the functional unit 478.<sub>0-n </sub>Will be sent to. Upon receiving the control signal, the register renaming unit 496 sends a register selection signal onto the register array access control bus 518. Which register is made interruptible by the control signal sent over bus 518 is determined by selecting the in-execution instruction and by the register renaming unit 496 determining the register referenced by that particular instruction. Will be done.
【0121】
The bypass control unit (BYPASS CTL) 520 generally controls the operation of the bypass unit 474 through a control signal on control line 524. Bypass control unit 520 is functional unit 478<sub>0-n </sub>Monitor each situation and transfer data from register array unit 472 to functional unit 478 in relation to the register reference sent from register renaming unit 496 via control line 522.<sub>0-n </sub>Should be sent to, or functional unit 478<sub>0-n </sub>The data output from is immediately sent to the functional unit distribution bus 480 via bypass unit 474 to determine if it can be used to execute the newly issued instruction selected by the instruction issuing unit 498. In both cases, the instruction issuing unit 498 is the functional unit 478.<sub>0-n </sub>From functional unit distribution bus 480 to functional unit 478 by selectively enabling specific register data for each of<sub>0-n </sub>Directly control the sending of data.
【0122】
The remaining units in the IEU control path include the evacuation control unit 500, the control flow control (CF CTL) unit 528, and the completion control (DONE CTL) unit 540. The evacuation control unit 500 operates to invalidate or confirm the execution of instructions executed out of order. When an instruction is executed out of order, the instruction can be confirmed or saved if all preceding instructions are also saved. When the identification information of which of the eight pending instructions in the current set is executed is sent on the control line 542, the save control unit 500 is sent on the control line 534 connected to the bus 518 based on the identification information. The control signal is sent to the register array unit 472 to effectively confirm the result data stored in the register array unit 472 as the result of the pre-execution of the instructions executed out of order.
【0123】
When saving each instruction, the save control unit 500 sends a PC increment / size control signal to the IFU 102 via the control line 344. Since a plurality of instructions can be executed out of order and therefore can be put in a ready state for saving at the same time, the save control unit 500 determines the size value based on the number of instructions saved at the same time. Finally, if all the instructions in the IFIFO master register 224 have been executed and saved, the save control unit 500 sends the IFIFO read control signal to the IFU 102 via the control line 342 to shift the IFIFO unit 264. By initiating the operation, the E-decode unit 490 is given four additional instructions as execution pending instructions.
【0124】
The control flow control unit 528 has a specific function of detecting the logical branch result of each conditional branch instruction. Control Flow Control unit 528 receives the 8-bit vector ID of the currently pending conditional branch instruction from E-decoding unit 490 via control line 510. The 8-bit vector instruction completion control signal is similarly received from the completion control unit 540 via the control line 542. With this completion control signal, the control flow control unit 528 can determine when the conditional branch instruction is completed to a point sufficient to determine the conditional control flow status. The control flow status result of a pending conditional branch instruction is stored by control flow control unit 528 at its execution. The data needed to determine the outcome of the conditional control flow instruction is obtained from the temporary status register in register array unit 472 via control line 530. When each conditional control flow instruction is executed, the control flow control unit 528 sends a new control flow result signal to the IFU 102 via the control line 348. In a preferred embodiment, this control flow result signal contains two 8-bit vectors, which are the status results by bit position of each of the eight control flow instructions that may be held. Is known, and also defines the corresponding situation-result-state obtained by bit-position mapping.
【0125】
Finally, the completion control unit 540 is the functional unit 478.<sub>0-n </sub>It is for monitoring the execution status of each operation of. Functional unit 478<sub>0-n </sub>When any of the instructions is notified that the instruction execution operation is completed, the completion control unit 540 sends a corresponding completion control signal on the control line 542 to register the register renaming unit 496, the instruction issuing unit 498, the save control unit 500, and the bypass control. Alert unit 520.
【0126】
Functional unit 478<sub>on </sub>By adopting a parallel array configuration, the consistency of control of IEU104 is improved. Individual functional units 478 to correctly recognize instructions and schedule them for execution<sub>on </sub>It is necessary to inform the instruction issuing unit 498 of the characteristics of. Functional unit 478<sub>on </sub>Is responsible for determining and executing the specific control flow operations required to perform the required function. Therefore, except for the instruction issuing unit 498, it is not necessary to independently inform the IEU control unit of the instruction control flow processing. Command issuing unit 498 and functional unit 478<sub>on </sub>Jointly inform the remaining control flow management units 496, 500, 520, 528, 540 of the functions to be performed at the required control signal prompt. Therefore, functional unit 478<sub>on </sub>Changes to certain control flow operations in IEU104 do not affect control operations in IEU104. In addition, the existing functional unit 478<sub>on </sub>If you want to enhance the functionality of, or another functional unit such as extended precision floating point multiplication unit, extended precision floating point ALU, fast Fourier calculation function unit, trigonometric function calculation unit 478<sub>on </sub>If you want to add one or more, you only need to make minor changes to the command issuing unit 498. To make the necessary changes, it recognizes a particular instruction based on the corresponding instruction field isolated by the E-decode unit 490, and that instruction and the required functional unit 478.<sub>on </sub>Need to be related. Controlling register data selection, data routing, instruction completion and save are functional units 478.<sub>on </sub>It is consistent with the processing of all other instructions executed for all other functional units.
【0127】
A) Details of the IEU data path The central element of the IEU data path is the register array unit 472. However, according to the present invention, within the IEU data pathway, there are several parallel data pathways optimized for individual functions. There are two main data paths, integer and floating point. Within each parallel data path, a portion of register array unit 472 is designed to support data manipulation performed within that data path.
【0128】
1) Detailed view of the register array unit FIG. 6 (A) is a schematic diagram of the preferred architecture of the data path register file 550 in the register array unit 472. The data path register file 550 contains a temporary buffer 552, a register array 554, an input selector 559, and an output selector 556. A typical example is that the data finally sent to the register array 554 is first received by the temporary buffer 552 via the combined data input bus 558'. That is, all data sent to the data path register file 550 is multiplexed by the input selector 559 and sent from the plurality of input buses 558 (preferably two) onto the input bus 558'. The register selection and enable control signals sent over control bus 518 select the register location of the received data in the temporary buffer 552. When the instruction that generated the data stored in the temporary buffer 552 is saved, the control signal sent to the control bus 518 again is sent from the temporary buffer 552 to the logically associated register in the register array 554. -Allow data to be transferred via bus 560. However, before the instruction is saved, the data stored in the temporary buffer 552 is sent to the output selector 556 via the bypass portion of the data bus 560 to send the data stored in the temporary buffer to the output selector 556 for subsequent instructions. It can be used at runtime. The output selector 556, controlled by a control signal sent via control bus 518, selects either data from the registers of the temporary buffer 552 or data from the registers of the register array 554. The resulting data is sent over the register array output bus 564. Also, if the instruction being executed is saved as soon as it is completed, that is, if the instruction is executed in order, the result data is sent directly to the register array 554 via the bypass extension portion 558 ". Can be instructed.
【0129】
According to a preferred embodiment of the present invention, each data path register file 550 can perform two register operations at the same time. Therefore, two full register width data values can be written to the temporary buffer 552 through the input bus 558. Internally, the temporary buffer 552 is a multiplexer array, so input data can be sent to any two registers in the temporary buffer 552 at the same time. Similarly, an internal multiplexer allows the data to be output onto bus 560 by selecting any 5 registers in temporary buffer 552. The register array 554 also has a human output multiplexer, so you can select two registers to receive their data from bus 560 at the same time, or you can select five registers and send them over bus 562. You can also do it. Finally, the output selector 556 is preferably implemented so that any five of the 10 register data values received from the buses 560 and 562 are simultaneously output onto the register array output bus 564.
【0130】
The register set in the temporary buffer 552 is outlined in Figure 6 (B). Register set 552'consists of eight single word (32-bit) registers I0RD, I1RD ... I7RD. Register set 552'can also be used as a set of four double word registers I0RD, I0RD + 1 (I0RD4), I1RD, I1RD + 1 (ISRD ... I3RD, I3RD + 1 (I7RD)). is there.
【0131】
According to a preferred embodiment of the present invention, instead of duplicating each register in the register array 554, the registers in the temporary buffer register set 552 are in the two IFIFO master registers 216 and 224, respectively. Referenced by register renaming unit 496 based on the relative location of the instruction in. Each instruction implemented in the Architecture 100 can refer to up to two registers or one double word register as output and be the destination of the data generated by the execution of the instruction. As a typical example, an instruction refers to only one output register. Therefore, as shown in Figure 6 (C), instruction 2 (I) refers to the output register of one of the eight pending instructions.<sub>2 </sub>In the case of), the data destination register I2RD is selected and accepts the data generated by the execution of the instruction. Command I<sub>2 </sub>The data generated by is followed by an instruction, eg I<sub>5 </sub>When used by, the data stored in the I2RD register is transferred via bus 560 and the resulting data is sent back to temporary buffer 552 and stored in the register indicated by I5RD. In particular, instruction I<sub>5 </sub>Is the instruction I<sub>2 </sub>Because it depends on the instruction I<sub>5 </sub>Is I<sub>2 </sub>It cannot be executed until the result data from is obtained. But as you can see, the instruction I<sub>5 </sub>Needs input data in instruction I of register set 552'<sub>2 </sub>Obtained from the data location of, instruction I<sub>2 </sub>It is possible to execute before saving.
【0132】
Finally, instruction I<sub>2 </sub>When is saved, the data from register I2RD is determined from the logical position of the instruction at the save location and written to the register location in register array 554. That is, the save control unit 500 determines the address of the destination register in the register array from the register reference field data given by the E-decode unit 490 via the control line 510. Command I<sub>0-3 </sub>When is saved, the value contained in I4RD-I7RD is shifted at the same time as the shift of IFIFO unit 264 and transferred to I0RD-I3RD.
【0133】
Command I<sub>2 </sub>It becomes even more complicated if the double word result value is obtained from. According to a preferred embodiment of the present invention, the combination of location I2RD and I6RD is command I.<sub>2 </sub>Is used to store the result data obtained from the instruction until it is evacuated or otherwise canceled. In a preferred embodiment, instruction I<sub>4-7 </sub>Execution of instruction I<sub>0-3 </sub>If a doubleword output reference by any of the above is detected by register renaming unit 496, it is deferred. This allows the entire register set 552'to be used as the single rank of a double word register. Command I<sub>0-3 </sub>Once saved, register set 552'can be reused as the second rank of a single word register. In addition, either instruction I<sub>4-7 </sub>Execution of the instruction corresponds to the I if a double word output register is required.<sub>0-3 </sub>Holds until shifted to.
【0134】
The logical organization of the register array 554 is shown in FIGS. 7 (A) to 7 (B). According to a preferred embodiment of the present invention, the register array 554 for an integer data path is composed of 40 32-bit wide registers. This register set constitutes register set "A", a top set consisting of base register set ra [0..23] 565, general purpose register ra [24..31] 566, and eight general purpose registers. It is organized as a shadow register set consisting of trap registers rt [24..31] 567. In normal operation, registers ra [0..31] 565, 566 make up the active "A" register set of the register array for integer data paths.
【0135】
As shown in Figure 7 (B), if the trap register rt [24..31] 567 is swapped and moved to the active register set A, the active register ra [0..23] 565 is active. -Can be accessed with the base set. This configuration of the "A" register set is selected when interrupt reception is confirmed or an exception trap handling routine is executed. This state of register set A is maintained until the state shown in FIG. 7 (A) is explicitly returned by the execution of the interrupt enable instruction or the return instruction from the trap.
【0136】
In a preferred embodiment of the invention realized by Architecture 100, the floating point data path uses a floating point register set 572 as outlined in Figure 8. The floating-point register set 572 consists of 32 registers rf [0..31], each 64-bit wide. The floating point register set 572 can also be logically referenced as the "B" set of the integer register rb [0..31]. In architecture 100, this "B" set of registers corresponds to the lower 32 bits of each of the floating-point registers rf [0..31].
【0137】
A Boolean register set 574 is provided to represent the third data path, as shown in FIG. It stores the logical result of the Boolean operation. This Boolean register set 574 consists of 32 1-bit registers rc [0..31]. The operation of the Boolean register set 574 is unique in that the result of the Boolean operation can be sent to any instruction selection register of the Boolean register set 574. This is in contrast to using a single processor status word register that stores 1-bit flags that represent conditions such as equal, unequal, greater than, and other simple Boolean status values.
【0138】
Floating-point register set 572 and Boolean register set 574 are both complemented by a temporary buffer of the same architecture as the temporary buffer 552'shown in Figure 6 (B). The fundamental difference is that the register width of the temporary buffer is defined to be the same as the width of the complementary register sets 572, 574. In a preferred embodiment, the widths are 64 bits and 1 bit, respectively.
【0139】
A number of additional special registers are at least logically present in register array unit 472. As shown in Figure 7 (C), the registers physically present in the register array unit 472 are the kernel stack pointer 568, the processor state register (PSR) 569, and the old processor state register (PPSR) 570. It consists of an array of eight temporary processor state registers (tPSR [0..7]) 571. The remaining special registers are distributed throughout Architecture 100. The special address and data bus 354 is for selecting data and transferring it between special registers and the "A" and "B" register sets. The special register move instruction is for selecting a register from the "A" or "B" register set, selecting the transfer direction, and specifying the address ID of the special register.
【0140】
The kernel stack pointer register and processor status register are different from other special registers. Kernel stack pointers can be accessed while in the kernel state by executing standard register-to-register movement instructions. The temporary processor status register is not directly accessible. Instead, this register array propagates the value of the processor state register so that it can be used by instructions that are executed out of order. Used to implement mechanism). The initial propagation value is the value of the processor status register. That is, it is the value obtained from the last saved instruction. This initial value is propagated forward from the temporary processor status register, allowing instructions executed out of order to access the value in the temporary processor status register at the corresponding location. The condition code bits that an instruction depends on and can be changed are defined by the characteristics of the instruction. If it is determined by the dependency checker unit 494 and the carry checker checker 492 that the instruction is not constrained by a dependency, register or condition code, the instruction can be executed out of order. Changes to the condition code bits of the processor status register are directed to the logically corresponding temporary processor status register. Specifically, only the bits that are subject to change are applied to the values in the temporary processor status register and propagated to all higher temporary processor status registers. As a result, all out-of-order instructions are executed from the processor status register value appropriately modified by the intervening PSR modification instruction. When the instruction is saved, only the corresponding temporary processor status register value is transferred to the PSR register 569. Other special registers are described in [Table 2].
【0141】<img he="216" id="000002" wi="152" file="2_0003552995.tif" img-format="tif" img-content="drawing" /><img he="201" id="000003" wi="152" file="3_0003552995.tif" img-format="tif" img-content="drawing" /> 【0142】
2) Details of Integer Data Path The integer data path of IEU104 constructed according to the preferred embodiment of the present invention is shown in FIG. For convenience of explanation, many control paths connected to the integer data path 580 are not shown in the figure. These connection relationships are as described with reference to FIG.
【0143】
The input data for data path 580 is obtained from alignment units 582, 584 and integer load / store units 586. Integer immediate data values are initially given as instruction embedded data fields and are obtained from operand unit 470 via bus 588. The alignment unit 582 isolates the integer data value and sends the resulting value to the multiplexer 592 via the output bus 590. Another input to the multiplexer 592 is the special register address and data bus 354.
【0144】
Immediate operands from the instruction stream are also obtained from operand unit 570 via data bus 594. These values are again right-justified by the alignment unit 584 before being delivered onto the output bus 596.
【0145】
Integer load / store unit 586 communicates bidirectionally with CCU106 through external data bus 598. Inbound data to IEU104 is transferred from integer load / store unit 586 to input latch 602 via input data bus 600. The output data from the multiplexer 592 and latch 602 is sent on the multiplexer input buses 604 and 606 of the multiplexer 608. Data from the functional unit output bus 482'is also sent to the multiplexer 608. In a preferred embodiment of Architecture 100, the multiplexer 608 includes two passages that simultaneously send data to the output multiplexer bus 610. In addition, the data transfer through the multiplexer 608 can be completed within each half cycle of the system clock. Most instructions implemented in this architecture 100 utilize one destination register, so up to four instructions can send data to the temporary buffer 612 during each system clock cycle.
【0146】
Data from the temporary buffer 612 can be transferred to the integer register array 614 via the temporary register output bus 616 or to the output multiplexer 620 via the alternate temporary buffer register bus 618. The integer register array output bus 622 can transfer integer register data to the multiplexer 620. The output bus connected to the temporary buffer 612 and the integer register array 614 each allows five register values to be output simultaneously. In other words, it is possible to issue two instructions at the same time that refer to a total of up to five source registers. Temporary buffer 612, integer register array 614, and multiplexer 620 allow the transfer of outbound register data to occur every semi-system clock cycle. Therefore, up to four integer and floating point instructions can be issued during each clock cycle.
【0147】
The multiplexer 620 serves to select outbound register data values directly from the integer register array 614 or from the temporary buffer 612. This allows the IEU104 to execute out-of-order instructions that depend on previously executed out-of-order instructions. This maximizes the execution throughput capacity of the IEU integer data path by executing pending instructions out of order, and accurately delivers out-of-order data results from the data results obtained from the executed and saved instructions. The two goals of separation can be easily achieved. According to the present invention, the data value existing in the temporary buffer 612 can be easily cleared in the event of an interrupt or other exception condition that requires the restoration of the exact state of the machine. Therefore, the integer register array 614 remains accurate in the data values obtained only by executing the saved instruction, which was completed before the interrupt or other exception condition occurred.
【0148】
Up to five register data values selected during each half-system cycle operation of the multiplexer 620 are sent to the bypass unit 626 via the multiplexer output bus 624. The bypass unit 626 basically consists of an array of parallel multiplexers that can send data appearing at any of its inputs to any of its outputs. The input of the bypass unit 626 is a special register addressing data value or an immediate integer value sent from the multiplexer 592 via the output bus 604, up to five register data values sent over the bus 624, and an integer load / Load operand data from store unit 586 via double integer bus 600, immediate operand value obtained from alignment unit 584 via its output bus 596, and finally bypass data from functional unit output bus 482. It consists of routes. This bypass path and data bus 482 can simultaneously transfer four register values per system clock cycle.
【0149】
Data is output from bypass unit 626 onto an integer bypass bus 628 connected to a floating point data bus, with two operand data buses capable of simultaneously transferring up to five register data values. Is sent to the store data bus 632, which is used to send data to the integer load / store unit 586.
【0150】
Router unit 634 is implemented by a parallel multiplexer array that allows the five register values received from its input to be sent to a functional unit provided in an integer data path. Specifically, the router unit 634 has five register data values sent from the bypass unit 626 via bus 630, and the current IF sent via address bus 352.<u style="single"></u>The PC address value, the control flow offset value determined by the PC control unit 362 and sent on line 378'is received. Router unit 634 can also optionally receive operand data values retrieved from bypass unit 626 within the floating-point data path via data bus 636.
【0151】
Register data values received by router unit 634 are forwarded over special register addresses and data bus 354 and sent to functional units 640, 642, 644. Specifically, the router unit 634 has a function of sending up to three register operand values to each of the functional units 640, 642, and 644 via the router output buses 646, 648, and 650. According to the general architecture of this architecture 100, up to two instructions can be issued to functional units 640, 642, 644 at the same time. According to a preferred embodiment of the present invention, each of the three dedicated integer function units can have a programmable shift function and two arithmetic logic unit functions.
【0152】
The ALU0 functional unit 644, the ALU1 functional unit 642, and the shifter functional unit 640 send their respective output register data onto the functional unit bus 482I. Output data obtained from ALU0 and shifter function units 644 and 640 is also sent on the shared integer function unit bus 650 connected to the floating-point data path. A similar floating-point functional unit output value data bus 652 is provided from the floating-point data path to the functional unit output bus 482'.
【0153】
ALU0 functional unit 644 is also used to generate virtual address values to support both IFU102 prefetch operations and integer load / store unit 586 data operations. The virtual address value calculated by the ALU0 functional unit 644 is sent on the output bus 654 connected to both the target address bus 346 and the CCU 106 of the IFU 102 to obtain the physical address (EX PADDR) of the execution unit. Latch 656 is for storing the virtualized portion of the address generated by the ALU0 functional unit 644. This virtualized portion of the address is sent over output bus 658 and sent to VMU108.
【0154】
3) Details of floating-point data path Next, Fig. 11 shows the floating-point data path. Initial data is again received from multiple sources, including the immediate integer operand bus 588, the immediate operand bus 594, and the special register address data bus 354. The final source of external data is the floating point load / store unit 662 connected to CCU106 through the external data bus 598.
【0155】
The immediate integer operand is received by the alignment unit 664, which acts to right justify the integer data field before passing it through the alignment output data bus 668 to the multiplexer 666. The multiplexer 666 also receives the special register address data bus 354. The immediate operand is sent to the second alignment unit 670, right-justified, and then sent onto the output bus 672. Inbound data from floating point load / store unit 662 data) is received by latch 674 from load data bus 676. Data from the multiplexer 666, latch 674 and functional unit data return bus 482 is received from the input of the multiplexer 678. The multiplexer 678 has a selectable data path and two register data values are the system clock. Allows the temporary buffer 680 to be written to the temporary buffer 680 via the multiplexer output bus 682 every half cycle of. The temporary buffer 680 has the same set of registers as the temporary buffer 552 shown in FIG. 6 (B). The temporary buffer 680 also reads up to five register data values from the temporary buffer 680, via the data bus 686, via the floating-point register array 684, and via the output data bus 690. The multiplexer 688 also receives up to five register data values from the floating-point register array 684 over the data bus 692. The multiplexer 688 has up to five. Selects the register data values up to and simultaneously transfers them to the bypass unit 694 via the data bus 696. The bypass unit 694 is the output data bus from the data bus 672 and the multiplexer 666. It also receives the immediate operand value given by the alignment unit 670 via the bypass extension of the 698, load data bus 676 and functional unit data return bus 482 . Bypass unit 694 simultaneously selects up to five register operand data values, bypass unit output bus 700, store data bus 702 connected to floating-point load / store unit 662, and integers. It works to output on the floating point bypass bus 636 connected to router unit 634 on data path 580.
【0156】
Floating-point router unit 704 simultaneously connects bypass unit output bus 700 and integer data path bypass bus 628 to functional unit input buses 706, 708, 710 connected to their respective functional units 712, 714, 716. It has a function that allows you to select a data path. Each of the input buses 706, 708, 710 according to the preferred embodiment of Architecture 100 can simultaneously transfer up to three register operand data values to each of the functional units 712, 714, 716. The output buses of these functional units 712, 714, 716 are coupled to the functional unit data return bus 482 "to return data to the register array input multiplexer 678. Integer data path functional unit output bus. The 650 can also be provided to connect to the functional unit data return bus 482'. According to the architecture 100 of the present invention, the function unit data return bus 482 of the integer data path 500 via the function unit output bus of the multiplexer function unit 712 and the floating point ALU714 via the floating point data path function unit bus 652. It is possible to connect to.
【0157】
Four) Details of the Boolean register data path The Boolean operation data path 720 is shown in Figure 12. This data path 720 is basically used to support the execution of two types of instructions. The first type is the operand comparison instruction, in which the two operands selected from the integer register set and the floating point register set, or given as immediate operands, are one of the ALU functional units and are integers. Is compared by subtracting the floating point data path. This comparison is made by subtraction by one of the ALU functional units 642, 644, 714, 716, and the resulting sign and zero status bits are sent to the input selector and comparison operator coupling unit 722. When this unit 722 receives an instruction with a control signal from the E-decode unit 490, it selects the output of the ALU functional units 642, 644, 714, 716, combines the sign and zero bits, and Booleans the comparison result value. Is extracted. The result of the comparison operation can be simultaneously transferred to the input multiplexer 726 and the bypass unit 742 through the output bus 723. Like integer and floating point data paths, bypass unit 742 is implemented as a parallel multiplexer array, allowing multiple data paths to be chosen between the inputs of bypass unit 742 and connected to multiple outputs. .. The other manpower of the bypass unit 742 consists of a Boolean result return data bus 724 and two Boolean operands on the data bus 744. Bypass unit 742 can transfer up to two Boolean operands representing currently executing Boolean instructions to Boolean function unit 746 via operand bus 748. In addition, the bypass unit 742 can simultaneously transfer up to two single-bit Boolean operand bits (CF0, CF1) via the control flow result control lines 750 and 752.
【0158】
The rest of the Boolean operation data path includes an input multiplexer 726 that receives the comparison and Boolean operation result values sent on the comparison result bus 723 and the Boolean result bus 724 as its inputs. This bus 724 can transfer up to two Boolean result bits to the multiplexer 726 at the same time. In addition, up to two comparison result bits can be transferred to the multiplexer 726 via bus 723. The multiplexer 726 can transfer any two signal bits that appear at the input end of the multiplexer via the output end of the multiplexer to the Boolean temporary buffer 728 during each half cycle of the system clock. Temporary buffer 728 is logically the same as register set 552'shown in Figure 6 (B), except that there are two important differences. The first difference is that each register entry in the temporary buffer 728 consists of a single bit. The second difference is that each of the eight pending instruction slots has only one register. This is because the entire result of the Boolean operation is defined by one result bit by definition.
【0159】
Temporary buffer 728 outputs up to four output operand values at the same time. This allows two Boolean instructions, each requiring access to two source registers, to be executed simultaneously. The four Boolean register values are sent over the operand bus 736 every half cycle of the system clock and transferred to the multiplexer 738 or via the Boolean operand data bus 734 to the Boolean register array 732. can do. The Boolean register array 732 is a single 32-bit wide data register, as logically shown in Figure 9, which allows up to four any combination of single bit locations from temporary buffer 728. It can be modified with the data in, read from the Boolean register array 732 every half cycle of the system clock, and sent out on the output bus 740. The multiplexer 738 sends any pair of Boolean operands received from its output end via buses 736 and 740 onto the operand output bus 744 and forwards them to the bypass unit 742.
【0160】
Boolean Calculator Unit 746 has the ability to perform a wide range of Boolean operations on two source values. In the case of a comparison instruction, the source value is the operand of the pair obtained from either an integer or floating point register set and any immediate operand sent to IEU104, and in the case of a Boolean instruction, the Boolean register operand. Any two. [Table 3] and [Table 4] show logical comparison operations in a preferred embodiment of the architecture 100 of the present invention. [Table 5] shows the direct Boolean operation in a preferred embodiment of the architecture 100 of the present invention. The instruction condition code and function code shown in [Table 2]-[Table 5] represent the corresponding instruction segment. The instruction also specifies a pair of source operand registers and a destination Boolean register to store the corresponding Boolean operation result.
【0161】<img he="112" id="000004" wi="110" file="4_0003552995.tif" img-format="tif" img-content="drawing" /> 【0162】<img he="142" id="000005" wi="135" file="5_0003552995.tif" img-format="tif" img-content="drawing" /> 【0163】<img he="157" id="000006" wi="110" file="6_0003552995.tif" img-format="tif" img-content="drawing" /> 【0164】
B) Load / store control unit Figure 13 shows an example of load / store unit 760. Although shown separately in data paths 580 and 660, load / store units 586 and 662 are preferably implemented as one shared load / store unit 760. Interfaces from the respective data paths 580 and 660 go through address bus 762 and load and store data buses 764 (600, 676), 766 (632, 702).
【0165】
The address used by load / store unit 760 is a physical address, as opposed to the virtual address used by the rest of IFU102 and IEU104. IFU102 operates on virtual addresses and relies on coordination between CCU106 and VMU108 to generate physical addresses, while IEU104 requires load / store unit 760 to operate directly in physical address mode. This requirement is required when there are instructions that cause physical address data and store operations to overlap because they are executed out of order, and the order from CCU 106 to load / store unit 760. This is to maintain data integrity in the presence of external data returns. For data integrity, the load / store unit 760 buffers the data obtained from the store instruction until the store instruction is saved by IEU104. As a result, only one store data buffered by load / store unit 760 can exist in load / store unit 760. A load instruction that refers to the same physical address as a store instruction that has been executed but has not been saved is delayed in execution until the store instruction is actually saved. At that point, store data can be transferred from load / store unit 760 to CCU106 and immediately loaded back by performing a CCU data load operation.
【0166】
Specifically, the entire physical address is sent from the VMU108 over the load / store address bus 762. The load address is generally the load address register 768.<sub>3-0 </sub>Stored in. The store address is the store address register 770<sub>3-0 </sub>Latched to. The load / store control unit 774 operates by receiving the control signal received from the instruction issuing unit 498, and registers the load address and the store address in the register 768.<sub>3-0 </sub>、770<sub>3-0 </sub>Adjust to latch to. The load / store control unit 774 sends a control signal for latching the load address onto the control line 778 and a control signal for latching the store address on the control line 780. Store data is store data register set 782<sub>3-0 </sub>The store address is latched at the same time as the logically corresponding slot of. 4x4x32 bit wide address comparison unit 772 has load and store address registers 768<sub>3-0 </sub>、770<sub>3-0 </sub>Each of the contained addresses is entered at the same time. The execution of a full matrix address comparison during each half cycle of the system clock is controlled by load / store control unit 774 via control line 776. The existence and logical location of a load address that matches the store address is sent via control line 776 to load / store control unit 774.
【0167】
If the load address is given by VMU108 and there are no pending stores, the load address is bypassed directly from bus 762 to address selector 786 at the start of the CCU load operation. However, if store data is pending, the load address is an available load address latch 768.<sub>3-0 </sub>Latched to. Upon receiving a control signal from the save control unit 500 that the corresponding store data instruction is saved, the load / store control unit 774 initiates a CCU data transfer operation and arbitrates access to the CCU 106 through control line 784. .. When the CCU 106 notifies ready, the load / store control unit 774 instructs the address selector 786 to send the CCU physical address over the CCUPADDR address bus 788. This address is the corresponding store register 770 via address bus 790.<sub>3-0 </sub>Obtained from. Corresponding store data register 782<sub>3-0 </sub>Data from is sent on the CCU data bus 792.
【0168】
When a load instruction is issued from instruction issuing unit 498, load / store control unit 774 is loaded address latch 768.<sub>3-0 </sub>Allows one of the to latch the requested load address. Specific latch 768 selected<sub>3-0 </sub>Logically corresponds to the position of the load instruction in the related instruction set. The instruction issuing unit 498 passes a 5-bit vector to the load / store control unit 774 indicating a load instruction in either of the two possible instruction sets that may be pending. If the comparator 772 does not indicate a matching store address, the load address is sent via address bus 794 to address selector 786 and output on CCU PADDR address bus 788. The address is provided according to the CCU request and ready control signal exchanged between the load / store control unit 774 and the CCU 106. The execution ID value (ExID value) is also prepared by the load / store control unit 774 and issued to CCU106, which identifies the load request when it subsequently returns the request data containing the ExID value. This ID value consists of a 4-bit vector, each load address latch 768 that issued the current load request.<sub>3-0 </sub>Is specified by a unique bit. The fifth bit is used to identify the instruction set that contains the load instruction. This ID value is therefore the same as the bit vector sent by the instruction issuing unit 498 with the load request.
【0169】
When the CCU 106 notifies the load / store control unit 774 that the preceding request load data is available, the load / store control unit 774 receives the data from the alignment unit 798 and loads it. -Allow to send on bus 764. The alignment unit 798 serves to right justify the load data.
【0170】
As soon as the data is returned from CCU106, the load / store control unit 774 receives the ExID value from CCU106. On the other hand, the load / store control unit 774 sends a control signal to the instruction issuing unit 498 notifying that the load data is sent on the load data bus 764, and further, the load data is returned for which load instruction. Returns a bit vector indicating whether it will be done.
【0171】
C) Details of the IEU control route With reference to FIG. 5 again, the operation of the IEU control route will be described in relation to the timing diagram shown in FIG. The execution timing of the instruction shown in FIG. 14 exemplifies the operation of the present invention, and of course, it can be changed in various modes.
【0172】
The timing diagram in FIG. 14 shows the processor system clock cycle P.<sub>0-6 </sub>The sequence of is shown. Each processor cycle is an internal T cycle T. start from. In Architecture 100 according to a preferred embodiment of the present invention, each processor cycle consists of two T cycles.
【0173】
During processor cycle 0, IFU102 and VMU108 operate to generate physical addresses. This physical address is sent to CCU106 and the instruction cache access operation is initiated. If the requested instruction set is in the instruction cache 132, the instruction set is returned to IFU 102 approximately halfway through processor cycle 1. The IFU 102 then manages the transfer of the instruction set through the prefetch buffer unit 260 and the IFIFO unit 264, and the transferred instruction set is first passed to the IEU 104 for execution.
【0174】
1) Details of the E-decoding unit The E-decoding unit 490 receives the entire instruction set in parallel and decodes it before processor cycle 1 is completed. The E-decoding unit 490 is realized in the preferred architecture 100 as a logic block based on the permutation combination theory having the function of directly decoding all valid instructions received via the instruction stream bus 124 in parallel. .. The instructions recognized by Architect 100 are shown in [Table 6] for each type, along with instructions, register requirements and required resource specifications.
【0175】<img he="216" id="000007" wi="152" file="7_0003552995.tif" img-format="tif" img-content="drawing" /><img he="157" id="000008" wi="152" file="8_0003552995.tif" img-format="tif" img-content="drawing" /> 【0176】
The E-decoding unit 490 decodes each instruction in the instruction set in parallel. The resulting instruction identification, instruction functionality, register references and functional requirements are obtained from the output of E-decode unit 490. This information is regenerated and latched by the E-decode unit 490 for each half cycle of the processor cycle until all instructions in the instruction set are saved. Therefore, information about all eight pending instructions is constantly available from the output of the E-decoding unit 490. This information is displayed in the form of an 8-element bit vector, where the bits or subfields of each vector logically correspond to the physical location of the corresponding instruction in the two pending instruction sets. Therefore, eight vectors are sent to the carry checker unit 492 via the control line 502. In this case, each vector specifies whether the corresponding instruction is acting on or depends on the carry bit of the processor status word. Eight vectors are sent via control line 510 to indicate the specific content and functional unit requirements of each instruction. Eight vectors are sent via control line 506, specifying the register reference used by each of the eight pending instructions. These vectors are sent before the end of processor cycle 1.
【0177】
2) Details of carry checker unit Carry checker unit 492 operates in parallel with dependency checker unit 494 during the data dependency phase period of the operation shown in FIG. The carry checker unit 492 is realized as logic based on the permutation combination theory in the preferred architecture 100. Therefore, at each iteration of the operation by carry checker unit 492, all eight instructions are considered as to whether the instruction changed the carrier flag in the processor status register. This is needed to allow instructions that depend on the status of the carry bits set by the previous instruction to be executed out of order. The control signal sent over control line 504 allows the carry checker unit 492 to identify a particular instruction that depends on the execution of the preceding instruction for the carry flag.
【0178】
In addition, carry checker unit 492 has a temporary copy of the carry bit for each of the eight pending instructions. For instructions that have not changed the carry bit, the carry checker unit 492 passes the carry bit to the next instruction in the order of the program instruction stream. Therefore, it is possible to execute an instruction that is executed out of order and changes the carry bit, and further, a subsequent instruction that depends on an instruction that is executed out of order is also an instruction that changes the carry bit. It can be executed even if it is placed later. In addition, the carry bit is maintained by the carry checker unit 492, so if an exception occurs prior to saving these instructions, the carry checker unit simply clears the internal temporary carry bit register. The good thing is that it's easier to do out of order. As a result, the processor status register is unaffected by the execution of instructions that are executed out of order. The temporary carry bit register maintained by carry checker unit 492 is updated upon completion of each out-of-order instruction. When an instruction executed out of order is saved, the carry bit corresponding to the last saved instruction in the program instruction stream is transferred to the carry bit location of the processor status register.
【0179】
3) Details of the data dependency checker unit Dependency checker unit 494 receives eight register reference identification vectors from the E-decoding unit 490 via the control line 506. Each register reference identifies a 5-bit value suitable for identifying 32 registers at a time and a register bank located in the "A", "B" or Boolean register set. It is indicated by a 2-bit value. Floating-point register sets are also known as "B" register sets. Each instruction can have up to three register reference fields. Two source register fields and one destination register field. Some instructions, especially inter-register transfer instructions, may specify a destination register, but the instruction bit field recognized by the E-decode unit 490 has no output data actually created. May mean. Rather, instruction execution is only intended to determine changes in the value of the processor status register.
【0180】
Dependency checker unit 494 is also implemented in the preferred architecture 100 with pure pure combinatorial logic, which precedes the source register references of instructions that appear later in the program instruction stream. It operates to determine the dependency between the instruction's destination register reference at the same time. The bit array is created by the dependency checker unit 494, which not only identifies which instructions depend on other instructions, but also on which registers each dependency originated. The carry-register data dependency is determined shortly after the start of the second processor cycle.
【0181】
4) Details of the register renaming unit The register renaming unit 496 receives the IDs of the register references of all eight pending instructions via the control line 506 and the register dependencies via the control line 536. A matrix from eight elements is also received via control line 542. These elements indicate which instructions were executed (completed) in the current set of pending instructions. From this information, the register renaming unit 496 sends an 8-element array of control signals to the instruction issuing unit 498 via control line 512. The control information sent in this way is a register renaming unit 496 that, when the data dependency of the current set is determined, indicates which of the currently pending instructions that has not yet been executed can be executed. Reflects the judgment made by. The register renaming unit 496 receives a selection control signal identifying up to six instructions issued simultaneously for execution via line 516. That is, two integer instructions, two floating point instructions and two Boolean instructions.
【0182】
The register renaming unit 496 has another function of selecting the source register to access when executing the identified instruction through the control signal sent to the register array unit 472 via bus 518. ing. The destination registers of instructions executed out of order are selected as being placed in the temporary buffers 612, 680, 728 of the corresponding data path. Instructions executed in order are saved upon completion and the resulting data is stored in register arrays 614, 684, 732. The choice of source register depends on whether the register was previously selected as the destination and the corresponding previous instruction has not yet been saved. In such cases, the source register is selected from the corresponding temporary buffers 612, 680, 728. If the previous instruction was saved, the corresponding register array 614, 684, 732 registers are selected. As a result, the register renaming unit 496 operates to effectively replace the register array reference with the temporary buffer register reference in the case of instructions that are executed out of order.
【0183】
According to architecture 100, temporary buffers 612, 680, 728 do not overlap the register structure of the corresponding register array. Rather, there is one destination register slot for each of the eight hold instructions. As a result, the replacement of the temporary buffer destination register reference is determined by the location of the corresponding instruction in the hold register set. Subsequent source register references are identified by dependency checker unit 494 for the instruction for which the source dependency has occurred. Therefore, the destination slot in the temporary buffer register can be easily determined by the register renaming unit 496.
【0184】
5) Details of the instruction issuing unit The instruction issuing unit 498 determines the set of instructions that can be issued based on the output of the register renaming unit 496 and the functional requirements of the instructions identified by the E-decoding unit 490. Command issuing unit 498 is a functional unit 478 reported via control line 514.<sub>0-n </sub>Make this judgment based on each situation of. Therefore, the instruction issuing unit 498 starts an operation when it receives a usable instruction set to be issued from the register renaming unit 496. Given that access to the register array is required to execute each instruction, the instruction issuing unit 498 is the functional unit 478 currently executing the instruction.<sub>0-n </sub>Expect to be available. In order to minimize the delay in determining the instruction to be issued to the register renaming unit 496, the instruction issuing unit 498 is realized by a dedicated combination logic.
【0185】
Upon determining which instruction to issue, register renaming unit 496 initiates access to the register array, which is in third processor cycle P.<sub>2 </sub>Continues until the end. Processor cycle P<sub>3 </sub>When is started, the instruction issuing unit 498 has one or more functional units 478 as indicated by "Execute 0".<sub>0-n </sub>Starts the operation by, receives and processes the source data sent from the register array unit 472.
【0186】
As a typical example, most instructions processed by Architecture 100 are executed through functional units in one processor cycle. However, some instructions require multiple processor cycles to complete simultaneous instructions, as indicated by "Execute 1". The Execute 0 and Execute 1 instructions can be, for example, executed by the ALU and the floating point multiplication function unit, respectively. As shown in FIG. 14, the ALU functional unit generates output data within one processor cycle, and this output data can be simply latched in the fifth processor cycle P.<sub>4 </sub>Sometimes it can be used to execute another instruction. The floating point multiplication function unit is preferably an internal pipelined function unit. Therefore, another floating point instruction can be issued in the next processor cycle. However, the result of the first instruction cannot be used for the number of processor cycles that depend on the data. The instructions shown in FIG. 14 require 3 processor cycles to complete processing in the functional unit.
【0187】
During each processor cycle, the function of instruction issuing unit 498 is repeated. As a result, the current pending instruction set status and functional unit 478<sub>0-n </sub>The availability of the entire set of is reassessed during each processor cycle. Therefore, under optimal conditions, the preferred architecture 100 can execute up to 6 instructions per processor cycle. However, the total average number of instructions executed from a typical instruction mix is 1.5 to 2.0 per processor cycle.
【0188】
The final consideration in the function of the instruction issuing unit 498 is that this unit is involved in the processing of trap conditions and the execution of specific instructions. Trap conditions in order to generate the matter is, there is a need to clear the instruction of Te to pair that has not yet been retracted from the IEU104. Such a situation occurs in response to an arithmetic error in functional unit 478.<sub>0-n </sub>Responds to receiving an external interrupt from the E-decode unit 490 when decoding an illegal instruction and relaying it to IEU104 via the interrupt request / receive confirmation control line 340. And it can happen. When a trap condition occurs, the instruction issuing unit 498 is responsible for suspending or invalidating all non-evacuated instructions currently held by IEU104. All instructions that cannot be saved at the same time are invalidated. This result is essential for the accurate generation of interrupts as opposed to the traditional method of executing a program instruction stream in sequence. When the IEU104 is ready to start executing the trap processing program routine, the instruction issuing unit 498 confirms the reception of the interrupt by the return control signal via the control line 340. Also, to prevent the possibility of recognizing exception conditions for an instruction based on the processor state bits changed before the instruction was executed in a conventional purely in-order routine, the instruction issuing unit 498 Responsible for ensuring that all instructions that may change the PSR (such as special moves and returns from traps) are executed strictly in order.
【0189】
Certain instructions that change the flow of program control are not discriminated by the I-decode unit 262. Instructions of this type include subroutine returns, returns from procedural instructions, and returns from traps. The command issuing unit 498 sends a discriminant control signal to the IFU 102 via the IEU return control line 350. The corresponding one of the special register blocks 412 is selected to output the IF_PC execution address that existed when the call instruction was executed, when a trap occurred, or when the procedure instruction appeared.
【0190】
6) Details of completion control unit Completion control unit 540 is functional unit 478<sub>0-n </sub>To check the completion status of the current operation. In the preferred architecture 100, the completion control unit 540 anticipates the completion of the operation by each functional unit and provides a completion vector indicating the execution status of each instruction in the currently pending instruction set to the functional unit 478.<sub>0-n </sub>It is sent to the register renaming unit 496, the bypass control unit 520, and the save control unit 500 about half a processor cycle before the instruction is executed by. This allows the instruction issuing unit 498 to consider the functional unit that completes execution as a resource available for the next instruction issuing cycle through the register renaming unit 496. The bypass control unit 520 can prepare to bypass the data output from the functional unit so as to pass through the bypass unit 474. Finally, the evacuation control unit 500 has a functional unit 478.<sub>0-n </sub>It operates to transfer data from the register array unit 472 to the register array unit 472 and at the same time save the corresponding instruction.
【0191】
7) Details of the save control unit In addition to the instruction completion vector sent from the completion control unit 540, the save control unit 500 monitors the oldest instruction set output from the E-decode unit 490. When each instruction in the instruction stream sequence is marked as complete by the completion control unit 540, the save control unit 500 registers from the temporary buffer slot through the control signal sent over control line 534. The corresponding instruction in array unit 472 indicates that data should be transferred to the specified register location. When one or more instructions are saved at the same time, a PC Inc / Size control signal is sent on the control line 344. Up to 4 instructions can be saved for each processor cycle. When the entire instruction set is saved, an IFIFO read control signal is sent over control line 342 to advance the IFIFO unit 264.
【0192】
8) Detailed control flow control unit The control flow control unit 528 constantly provides the IFU 102 with information specifying whether the control flow instructions in the current pending instruction set have been resolved and, as a result, a branch has occurred. Acts to give. The control flow control unit 528 acquires the identification information of the control flow branch instruction by the E decoding unit 490 via the control line 510. The current set of register dependencies is sent from dependency checker unit 494 to control flow control unit 528 via control line 536, so that control flow control unit 528 is constrained by the results of branch instructions. You can determine if it is or is known. Register references sent from register renaming unit 496 via bus 518 are monitored by control flow control unit 528 to determine the Boolean register that defines the branch decision. Therefore, the branch determination can be determined even before the execution of the control flow instructions out of order.
【0193】
Simultaneously with the execution of the control flow instruction, the register array unit 472 controls the result of the control flow via the control line 530 consisting of the control lines 750 and 752 of the control flow 0 and the control flow 1 by the bypass control unit 520. Instructed to send to control unit 528. Finally, the control flow control unit 528 continuously sends two vectors, each of which is 8-bit, to the IFU 102 via control line 348. These vectors define whether the instruction placed at the logical location corresponding to the bits in the vector has been resolved and, as a result, a branch has been made.
【0194】
In the preferred architecture 100, the control flow control unit 528 is realized as a combination logic that operates continuously by receiving an input control signal to the control unit 528.
【0195】
9) Bypass control unit details Command issuing unit 498 works closely with bypass control unit 520 to register array unit 472 and functional unit 478.<sub>0-n </sub>Controls the routing of data between. The bypass control unit 520 operates in association with the register array access, output and store phases of the operation shown in FIG. During register array access, bypass control unit 520 recognizes access to destination registers in register array unit 472 that are being written during the output phase of instruction execution through control line 522. can do. In this case, the bypass control unit 520 instructs to select the data transmitted on the functional unit output bus 482 so as to bypass and return to the functional unit distribution bus 480. Control over the bypass control unit 520 is performed by command issuing unit 498 through control line 542.
【0196】
IV. The interface definition of the virtual memory control unit VMU108 is shown in Fig. 15. VMU108 is mainly VMU control logic unit 800 and content address (content) Addressable) Consists of memory (CAM) 802. The general functions of the VMU108 are shown in a block diagram in Figure 16. In the figure, the representation of the virtual address is the space ID (sID [31:28]), the virtual page number (VADDR [27:14]), the page offset (PADDR [13: 4]), and the request ID (rID). It is divided into [3: 0]). The algorithm for generating the physical address uses the space ID to select one of the 16 registers in the space table 842. The contents of the selected space register are combined with the virtual page number to be used as the address when accessing the table index buffer (TLB) 844. The 34-bit address acts as a content address tag and is used to specify the corresponding buffer register in buffer 844. If a match is found for the tag, the 18-bit wide register value is obtained as the upper 18 bits of physical address 846. The page offset and request ID are obtained as the lower 14 bits of physical address 846.
【0197】
A VMU mismatch is reported if no matching tag is found in the table index buffer 844. In this case, it is necessary to execute the VMU high-speed trap processing routine that employs the conventional hash algorithm 848 that accesses the full page table data structure maintained in MAU112. This page table 850 contains entries for all memory pages currently in use by architecture 100. The hash algorithm 848 determines the page table entries needed to satisfy the current virtual page conversion operation. These page table entries are loaded from MAU112 into the trap register of register set "A" and then transferred to the table index buffer 844 by a special register move instruction. Upon returning from the exception handling routine, the instruction that caused the VMU mismatch exception is re-executed by IEU104. The virtual address to physical address translation operation should complete without exception.
【0198】
The VMU control logic 800 will be a dual interface with the IFU 102 and IEU 104. The readiness signal is sent via control line 822 to IEU104, notifying that VMU108 is available for address translation. In a preferred embodiment, the VMU 108 is always ready to accept the conversion request of the IFU 102. Both IFU102 and IEU104 can present requests via control lines 328 and 804. In the preferred architecture 100, the IFU can preferentially access the VMU108. As a result, only one busy control line 820 is output to IEU104.
【0199】
Both IFU102 and IEU104 send space ID and virtual page number fields to VMU control logic 800 via address lines 326 and 808, respectively. In addition, the IEU104 outputs a read / write control signal on control line 806. This control signal optionally defines whether the address should be used for load operations or store operations in order to change the memory access protection attributes of the referenced virtual memory. The space ID of the virtual address and the virtual page field are passed to CAM unit 802 for the actual translation operation. The page offset and ExID fields are finally sent directly from IEU104 to CCU106. The physical page and request ID field are sent to CAM unit 802 via address line 836. If a match is found in the table index buffer, the VMU control logic unit 800 is notified via the hit line and control output line 930. The resulting 18-bit long physical address is output on the address output line 824.
【0200】
Upon receiving a hit and control output control signal from output line 930, the VMU control logic unit 800 outputs a virtual memory mismatch and virtual memory exception control signal on lines 334 and 332. A virtual memory conversion mismatch means that it did not match the page table ID in the table index buffer 844. All other conversion errors are reported as virtual memory exceptions.
【0201】
Finally, the data table in CAM unit 802 can be modified by the IEU104 executing a special register-to-register transfer instruction. Read / write, register select, reset, load and clear control signals are output from IEU104 via control lines 810, 812, 814, 816 and 818. The data to be written to the CAM unit register is received by the VMU control logic unit 800 from the IEU 104 via the address bus 808 connected to the special address data bus 354. This data is transferred to the CAM unit 802 via the bus 836 at the same time as the control signals for initial setting, register selection, and read / write control signals. As a result, the data registers in CAM unit 802 are dynamic operations of architecture 100, including reads for stores that are needed when processing context switches defined in higher level operating systems. Can be instantly exported as needed during.
【0202】
V. Control of the cache control unit CCU106 over the data interface is shown in Figure 17. Again, separate interfaces are provided for IFU 102 and IEU 104. In addition, a logically separate interface is provided on the CCU 106, which is connected to the MCU 110 for instruction and data transfer.
【0203】
The IFU interface is from the physical page address sent on address line 324, the VMU translated page address sent on address line 824, and the request ID forwarded separately on the ID output lines 294 and 296. It has become. The unidirectional data transfer bus (instruction bus) 114 is for transferring the entire instruction set in parallel with the IFU 102. Finally, read / in-use and ready control signals are sent to CCU106 via control lines 298, 300, 302.
【0204】
Similarly, the entire physical address is sent to IEU104 via the physical address bus 788. The request ExID is passed separately to the IEU104 load / store unit via control line 796. The 80-bit wide unidirectional data bus is output from CCU106 to IEU104. However, in a preferred embodiment of Architecture 100, only the lower 64 bits are used by IEU104. All 80-bit data transfer buses are available and supported within CCU106 to support subsequent execution of the Architecture 100 by modifying the floating-point data path 660. Supports floating point operations according to IEEE standard 754.
【0205】
The IEU control interface is established through request, in use, preparation, read / write, and through control signal 784, and is substantially the same as the corresponding control signal used by IFU 102. The exception is that read / write control signals are provided to distinguish between load and store operations. The width control signal specifies the number of bytes transferred when the IEU104 accesses each CCU106. In contrast, all access to the instruction cache 132 is a fixed 128-bit width data fetch operation.
【0206】
The CCU 106 has almost the same cache control function as the conventional one for the instruction cache 132 and the data cache 134. In the preferred architecture 100, the instruction cache 132 is a high-speed memory with the ability to store 256 128-bit wide instruction sets. The data cache 134 has the ability to store 1024 32-bit wide word data. Instruction requests and data requests that are not immediately satisfied from the contents of the instruction cache 132 and the data cache 134 are passed to the MCU 110. If the instruction cache fails, a 28-bit wide physical address is passed to the MCU 110 via address bus 860. The request ID and additional control signals to coordinate the operation of CCU106 and MCU110 are sent on control line 862. When the MCU110 adjusts the required read access of the MAU112, two consecutive 64-bit wide data transfers are made directly from the MAU112 to the instruction cache 132. Two transfers are required because the data bus 136 is a 64-bit wide bus in the preferred architecture 100. When the requested data is returned through the MCU110, the request ID that was held while the request operation was pending is also returned to the CCU106 via control line 862.
【0207】
The data transfer operation between the data cache 134 and the MCU 110 is almost the same as the instruction cache transfer operation. Since data load and store operations can refer to a single byte, a full 32-bit wide physical address is sent to the MCU 110 via address bus 864. The interface control signal and request ExID are transferred via control line 866. Bi-directional 64-bit width data transfer is performed via the data bus 138 of the data cache.
【0208】
VI. Summary and Conclusion The microprocessor architecture based on high performance RISC is as described above. According to the architecture of the present invention, instructions can be executed out of order, prefetch instruction transfer routes for the main and target instruction streams can be provided separately, and procedure instruction recognition and dedicated prefetch routes can be provided. The optimized instruction execution unit allows multiple optimized data processing paths to support integer, floating-point, and Boolean operations, and each has its own temporary register array for ease of use. Out-of-order execution and instruction revocation can be easily performed while accurately maintaining the machine state status set to.
【0209】
Therefore, although the above description discloses preferred embodiments of the present invention, it is of course possible for those skilled in the art to make various changes and improvements within the scope of the present invention.
[Simple explanation of drawings]
FIG. 1 is a simplified block diagram showing a microprocessor architecture of a preferred embodiment of the present invention.
FIG. 2 is a detailed block diagram showing an instruction fetch unit constructed according to the present invention.
FIG. 3 is a block diagram showing a program counter logic unit constructed according to the present invention.
FIG. 4 is another detailed block diagram showing program counter data and control path logic.
FIG. 5 is a simplified block diagram showing an instruction execution unit of the present invention.
FIG. 6 (A) is a simplified block diagram showing a register file architecture used in a preferred embodiment of the present invention, and FIG. 6 (B) is a temporary buffer register used in a preferred embodiment of the present invention. The figure which shows the storage register format of a file graphically, (C) is the figure which shows the primary and secondary instruction set when it exists in the last two stages of the instruction FIFO unit of this invention.
FIG. 7 is a graphic representation of a reconfigurable state of a first-order integer register set provided according to a preferred embodiment of the present invention.
FIG. 8 is a graphic representation of a reconfigurable floating point and quadratic integer register set provided according to a preferred embodiment of the present invention.
FIG. 9 is a graphic representation of a cubic Boolean register set provided in a preferred embodiment of the present invention.
FIG. 10 is a detailed block diagram showing a first-order integer processing data path portion of an instruction execution unit configured according to a preferred embodiment of the present invention.
FIG. 11 is a detailed block diagram showing a first-order floating-point data path portion of an instruction execution unit configured according to a preferred embodiment of the present invention.
FIG. 12 is a detailed block diagram showing a Boolean operation data transition portion of an instruction execution unit configured according to a preferred embodiment of the present invention.
FIG. 13 is a detailed block diagram showing a load / store unit configured according to a preferred embodiment of the present invention.
FIG. 14 is a timing diagram showing a preferable operation sequence of a preferred embodiment of the present invention when executing a plurality of instructions according to the present invention.
FIG. 15 is a simplified block diagram showing a virtual memory control unit configured according to a preferred embodiment of the present study.
FIG. 16 is a graphic representation of a virtual memory control algorithm used in a preferred embodiment of the present invention.
FIG. 17 is a simplified block diagram showing a cache control unit used in a preferred embodiment of the present invention.
[Explanation of symbols]
100 ... Architecture, 102 ... Instruction Fetch Unit (IFU), 104 ... Instruction Execution Unit (IEU), 106 ... Cache Control Unit (CUU), 108 ... Virtual Memory Unit (VMU) ), 112 ... Memory Array Unit (MAU)
24 sheets
Sheet 1 Sheet 2 Sheet 3 Sheet 4 Sheet 5 Sheet 6 Sheet 7 Sheet 8 Sheet 9 Sheet 10 Sheet 11 Sheet 12 Sheet 13 Sheet 14 Sheet 15 Sheet 16 Sheet 17 Sheet 18 Sheet 19 Sheet 20 Sheet 21 Sheet 22 Sheet 23 Sheet 24
Every citation, both ways
| Document | Relation | Office |
|---|---|---|
| JP02234248A | Cites | Japan |
| JP03018932A | Cites | Japan |
| JP03137729A | Cites | Japan |
| JP63316131A | Cites | Japan |
| JP64021629A | Cites | Japan |
24 members in 8 offices
Priority claims5
| Document | Office | Kind | Date |
|---|---|---|---|
| 07726942 | United States of America | – | |
| 72694291 | United States of America | A | |
| 72694291 | United States of America | A | |
| 1991726942 | – | – | – |
| US19910726942 | – | – | – |
Members24
| Document | Office | Kind | |
|---|---|---|---|
| WO9301547A1 | World Intellectual Property Organization (WIPO) | A1 | |
| EP0547240A1 | European Patent Office (EPO) | A1 | |
| KR930702719A | Republic of Korea | A | |
| JPH06502035A | Japan | A | |
| US5448705A | United States of America | A | |
| US5481685A | United States of America | A | |
| EP0945787A2 | European Patent Office (EPO) | A2 | |
| HK1014783A1 | Hong Kong, China | A1 | |
| EP0547240B1 | European Patent Office (EPO) | B1 | |
| AT188786T | Austria | T | |
| ATE188786T1 | Austria | T1 | |
| DE69230554D1 | Germany | D1 | |
| DE69230554T2 | Germany | T2 | |
| JP2001022583A | Japan | A | |
| JP2001022584A | Japan | A | |
| JP2001027949A | Japan | A | |
| JP2001067220A | Japan | A | |
| KR100294276B1 | Republic of Korea | B1 | |
| JP3333196B2 | Japan | B2 | |
| JP2003330708A | Japan | A | |
| JP3552995B2This record | Japan | B2 | |
| JP3750743B2 | Japan | B2 | |
| JP3879812B2 | Japan | B2 | |
| EP0945787A3 | European Patent Office (EPO) | A3 |
12 legal events, as the office reported them to INPADOC
Over the term
Point at a mark for the eventEvents
| Event | Code | |
|---|---|---|
| Written notification of registration of transferR350 | R350 | |
| Request for change of ownership or part of ownershipS111 | S111 | |
| Cancellation because of no payment of annual feesLAPS | LAPS | |
| Renewal fee payment (event date is renewal date of database)FPAY | FPAY | |
| Renewal fee payment (event date is renewal date of database)FPAY | FPAY | |
| Renewal fee payment (event date is renewal date of database)FPAY | FPAY | |
| Renewal fee payment (event date is renewal date of database)FPAY | FPAY | |
| Renewal fee payment (event date is renewal date of database)FPAY | FPAY | |
| Certificate of patent or registration of utility modelR150 | R150 | |
| First payment of annual fees (during grant procedure)A61 | A61 | |
| Written decision to grant a patent or to grant a registration (utility model)A01 | A01 | |
| Decision of grant or rejection writtenTRDD | TRDD |
Numbers
- Publication
- 3552995
- Publication, DOCDB
- 3552995
- Publication, EPODOC
- JP3552995B
- Application
- 175145
- Application, DOCDB
- 2000175145
- Application, EPODOC
- JP20000175145
Titles2
- Japanese
- データ処理装置
- English
- Data processing device
Classification
- CPC, 18
- G06F9/3814
- G06F9/3865
- G06F9/3017
- G06F9/32
- G06F9/322
- G06F9/3802
- G06F9/3804
- G06F9/3836
- G06F9/3861
- G06F9/3885
- G06F9/462
- G06F9/4812
- G06F15/7842
- G06F9/384
- G06F9/3858
- G06F9/38
- G06F9/3854
- G06F9/323
- IPC, 9
- G06F9 30
- G06F9 318
- G06F9 32
- G06F9 38
- G06F9 42
- G06F9 455
- G06F9 46
- G06F9 48
- G06F15 78
