Multiprocessor system with multiple instruction sources
19 claims: 10 independent, 9 dependent
- 1Digitale Verarbeitungsvorrichtung mit:einem Satz miteinander verbundener Verarbeitungselemente (58, 60, 82, 84), wobei der Satz von Verarbeitungselementen ein erstes Verarbeitungselement umfaßt, das erste Verarbeitungselement (58) aufweist: eine Heranholeinrichtung für das Heranholen von Befehlen von einer ersten Befehlsquelle, wobei Befehle bzw. Anweisungen, die von der ersten Befehlsquelle herangeholt werden, einen Anweisungs- bzw. Befehlsstrom bilden, und eine Ausführungseinrichtung für das Ausführen der Befehlen, wobei zumindest ein weiteres der Verarbeitungselemente Eingabe-/Ausgabeverarbeitungseinrichtungen (60) für das Verarbeiten von Signalen umfaßt, die von einem peripheren Gerät empfangen oder zu diesem übermittelt werden, wobei die Eingabe-/Ausgabeverarbeitungseinrichtung eine Einfügungseinrichtung aufweist und die Einfügungseinrichtung in der Weise betreibbar ist, daß sie auf ausgewählte Signale von dem peripheren Gerät reagiert, um Direktspeicherzugriffs-, \DMA, -steuerbefehle zu erzeugen und diese DMA-Steuerbefehle als eingefügte Befehle für das erste Verarbeitungselement bereitzustellen, damit sie durch diese verarbeitet werden, um zumindest DMA- Übertragungen mit dem peripheren Gerät auszulösen, wobei die Ausführungseinrichtung sowohl die Befehle, die von der ersten Befehlsquelle herangeholt wurden, als auch die eingefügten Befehle verarbeitet und wobei die Ausführungseinrichtung die eingefügten Befehle in derselben Art und Weise ausführt wie die von der ersten Befehlsquelle herangeholten Befehle, ohne deren Abfolge zu beeinflussen.
- 2Vorrichtung nach Anspruch 1, welche weiterhin aufweist:eine Befehlspipelineeinrichtung für das Verbinden der Verarbeitungselemente und der ersten Befehlsquelle für das Ausführen der Befehle, und wobei die Einfügungseinrichtung eine Einrichtung zum Anwenden einer oder mehrerer der eingefügten Befehle auf die Befehlspipelineeinrichtung aufweist.
- 3Vorrichtung nach Anspruch 1 oder 2, wobei die eingefügten Befehle dasselbe Format haben wie die Befehle aus der ersten Befehlsquelle.
- 4Vorrichtung nach Anspruch 3, wobei das Format eine ausgewählte Anzahl digitaler Befehlsbits aufweist, wobei zumindest ein erster Satz der Befehlsbits ein erstes Befehlsfeld bilden.
- 5Vorrichtung nach Anspruch 3, wobei das Format einen ersten Satz digitaler Befehlsbits aufweist, um ausgewählte Adreßsignale anzugeben, und einen zweiten Satz von digitalen Befehlsbits aufweist für das Angeben ausgewählter Befehlssignale.
- 6Vorrichtung nach einem der vorstehenden Ansprüche, wobei die erste Befehlsquelle ein Speicherelement aufweist.
- 7Vorrichtung nach Anspruch 6, wobei das Speicherelement ein Befehlscacheelement zum Speichern digitaler Werte aufweist, welche Befehlen entsprechen.
- 8Vorrichtung nach Anspruch 7, wobei das erste Verarbeitungselement eine Ausführungseinheit (CEU) aufweist, wobei die CEU umfaßt eine Einrichtung zum Ausgeben von Signalen an das Befehlscacheelement, um zu bewirken, daß Befehle von dem Befehlscacheelement zu der CEU übertragen werden.
- 9Vorrichtung nach Anspruch 7, wobei die Befehle Cacheverwaltungsbefehle umfassen, die durch das Befehlscacheelement eingefügt werden.
- 10Vorrichtung nach Anspruch 7, wobei das Befehlscacheelement Einrichtungen zum Speichern von Befehlen umfaßt, die Programmschritten entsprechen.
- 11Vorrichtung nach einem der vorstehenden Ansprüche, wobei die Eingabe- /Ausgabeverarbeitungseinrichtung eine periphere Schnittstelleneinheit (XIU) für das Steuern von Signalen aufweist, die von einem Peripheriegerät empfangen und durch dieses gesendet werden.
- 12Vorrichtung nach einem der vorstehenden Ansprüche, wobei die Eingabe- /Ausgabeverarbeitungseinrichtung eine Grafiksteuereinrichtung aufweist, um Signale zu steuern, die zu einer Anzeigeeinrichtung übermittelt werden.
- 13Vorrichtung nach einem der vorstehenden Ansprüche, wobei die Eingabe- /Ausgabeverarbeitungseinrichtung eine Textsucheinrichtung zum Suchen von Datenstrukturen aufweist, welche Text entsprechen.
- 14Vorrichtung nach einem der vorstehenden Ansprüche, wobei zumindest ein ausgewähltes der Verarbeitungselemente zumindest ein erstes Registerelement aufweist, welches diesem Verarbeitungselement zugeordnet ist, um digitale Werte zu speichern, die Daten entsprechen, sowie eine Einfügungseinrichtung aufweist, die Einrichtungen zum Erzeugen und Anwenden von Befehlen auf das ausgewählte Verarbeitungselement umfaßt, um die Bewegung von Daten in zumindest das erste Registerelement und aus diesem heraus zu kontrollieren.
- 15Vorrichtung nach einem der vorstehenden Ansprüche, wobei zumindest ein ausgewähltes der Verarbeitungselemente zumindest ein erstes Registerelement aufweist, welches diesem Verarbeitungselement zugeordnet ist, um digitale Werte zu speichern, die Daten entsprechen, und eine Einfügungseinrichtung aufweist, welche Einrichtungen zum Erzeugen und Anwendungen von Befehlen auf das ausgewählte Verarbeitungselement umfaßt, um die Ausführung von ausgewählten Logikoperationen mit ausgewählten Datenwerten zu bewirken, die in zumindest dem ersten Registerelement gespeichert sind.
- 16Vorrichtung nach einem der vorstehenden Ansprüche, wobei zumindest eines der Verarbeitungselemente eine Abfangeinrichtung aufweist, der auf ein Abfangsignal reagiert, um eine Abfangsequenz auszulösen, wobei die Abfangsequenz ausgewählte Programmschritte umfaßt, die in Reaktion auf ein Abfangsignal ausgeführt werden sollen, und mit einer Einfügungseinrichtung, welche Einrichtungen zum Erzeugen und Anwenden von Befehlen auf das Verarbeitungselement, um ein Signal für ein Abfangereignis zu erzeugen.
- 17Vorrichtung nach einem der vorstehenden Ansprüche, wobei zumindest eines der Verarbeitungselemente eine Unterbrechungseinrichtung aufweist, welche auf ein Unterbrechungssignal (Interruptsignal) reagiert, um eine Unterbrechungssequenz auszulösen, wobei die Unterbrechungssequenz ausgewählte Programmschritte umfaßt, die in Reaktion auf ein Unterbrechungssignal ausgeführt werden sollen, und mit einer Einfügungseinrichtung, die Einrichtungen zum Erzeugen und Anwenden von Befehlen auf das Prozessorelement umfaßt, um eine Unterbrechungssequenz auszulösen.
- 18Vorrichtung nach Anspruch 17, wobei die Unterbrechungseinrichtung Einrichtungen zum Erzeugen eines Abfangsignales in Reaktion auf ein Unterbrechungssignal aufweist.
- 19Verfahren zum Betreiben eines digitalen Datenprozessors einschließlich eines Satzes von miteinander verbundenen Verarbeitungselementen (58, 60, 82, 84), wobei das Verfahren die Schritte aufweist:Heranholen von Befehlen in Form eines Befehlsstromes von einer ersten Befehlsquelle durch ein erstes Verarbeitungselement (58), wobei zumindest ein weiteres der Verarbeitungselemente Eingabe-/Ausgabeverarbeitungseinrichtungen (60) zum Verarbeiten von Signalen aufweist, die von einem peripheren Gerät empfangen und an ein solches übermittelt werden, die Eingabe-/Ausgabeverarbeitungseinrichtung eine Einfügungseinrichtung aufweist, die Einfügungseinrichtung auf ausgewählte Signale von dem peripheren Gerät reagiert, um Steuerbefehle für einen direkten Speicherzugriff (DMA) zu erzeugen, und diese DMA-Steuerbefehle als eingefügte Befehle auf das erste Verarbeitungselement anwendet, damit sie durch diese verarbeitet werden, um zumindest DMA-Übertragungen mit der peripheren Einrichtung auszulösen, das erste Verarbeitungselement sowohl die Befehle, die von der ersten Befehlsquelle herangeholt werden als auch die eingefügten Befehle verarbeitet, wobei die eingefügten Befehle in derselben Art und Weise wie die von der ersten Befehlsquelle herangeholten Befehle ausgeführt werden und ohne die Abfolge der letzteren zu beeinflussen.
Independent claims19
223 paragraphs in 7 sections, as filed
Background of the invention
The present invention relates generally to data processing methods and apparatus, and more particularly to digital multiprocessor computer systems having distributed memory systems.
Multiprocessor computer systems provide multiple independent central processing units (CPUs) that can be coherently interconnected. Recent efforts in the multiprocessor field have focused on multiprocessor systems in which each of a plurality of processors is provided with an associated Random Access Memory (RAM) or with a cache memory unit. These multiprocessors typically communicate with each other via a common system bus structure or through signaling within a shared memory address space. Multiprocessors using a common bus are referred to as shared bus systems, while those using a shared memory area are referred to as shared address space systems.
To minimize bottlenecks in transmission, some distributed storage systems connect individual processing units to local storage elements to form semi-autonomous processing cells. To achieve the advantages of multi-processor processing, some such systems provide cell communications via use of hierarchical architectures. For example, U.S. Patent No. 4,622,631 to Frank et al. a multiprocessor system in which a plurality of processors, each having an associated private memory or cache, share data shared in a main memory element. Data within the shared memory is divided into blocks, each of which may be owned by any of the multiple processors or main memory. The current owner of a data block is said to have the correct data for that block.
In addition, in recent years, a variety of methods and apparatus have been proposed or developed to interconnect the processors of a shared bus system for multiple processors.
Such a shared bus multiprocessing computer system is disclosed in United Kingdom Patent Application No. 2,178,205, published February 4, 1987. The device disclosed therein comprises a plurality of processors, each of which has its associated cache memory. The system's cache memories are interconnected via a shared bus structure.
However, conventional bus sharing systems lack adequate bandwidth to provide multiple processors with short effective access times during times of high bus competition or heavy bus operation. Although a number of cache schemes have been proposed and developed for the purpose of reducing bus contention, the speed and size of many multiprocessor computers is still limited by bus saturation.
In addition, the processing speed of a conventional bus structure is limited by the bus length. In particular, when additional processors are connected to a typical shared bus system, the bus length increases, as does the time required for signal transmission and processing.
Another class of interconnect systems, known as crossbar networks, avoids some of the limitations of conventional shared bus systems. However, in a crossbar network, the path taken by a given signal can not be uniquely specified. In addition, the system cost increases proportionally to the square of the number of interconnected processors. These features make crossbar networks generally unsuitable for multiprocessor systems.
International Patent Application WO 84/016 describes a digital processing apparatus which includes a target CPU, a data cache, an instruction cache, and an emulator support processor (EAP). If a predetermined field within the source instruction advances one step and accesses a portion of controlled information from memory, and if controlled information designates a field-to-field mapping, then a skeletal target instruction can be filled by either selecting the fields of the Source command or otherwise be calculated. When the allocation is performed by an intervening independent processor, the overlap of such a conversion increases throughput, and the independent processor converts multi-field instructions for a first type CPU into multi-field instructions for a second type CPU, without the logic stream or interrupt the execution of any streams of source or drain target commands.
It is therefore an object of the invention to provide multiprocessing methods and apparatuses having flexible connection configurations which allow for increased processing speed.
Summary of the invention
According to a first aspect of the invention there is provided a digital processing apparatus comprising a set of interconnected processing elements, the set of processing elements comprising a first processing element and the first processing element having prefetch means for prefetching instructions from a first instruction source; prefetched from the first instruction source, forming a command stream, and execution means for executing instructions, wherein at least one other of the processing elements comprises input / output processing means for processing signals received from and forwarded to a peripheral device, the input / output processing means having insertion means in that the inserter is operable to respond to selected signals from the peripheral device, that it generates direct memory access (DMA) control directives and inserts these DMA control statements as inserted instructions for the first work item to be processed, thereby causing at least DMA transfers to the peripheral device, the execution means including both the first command source prefetched commands or Instructions as well as the inserted instructions, the execution means executing the inserted instructions in the same way as the instructions prefetched from the first instruction source without affecting the sequence thereof.
According to a second aspect of the invention, there is provided a method of operating a digital data processor comprising a set of interconnected processing elements, the method comprising the steps of having a first processing element prefetch instructions in a instruction stream from a first instruction source, at least one other the processing elements comprise input / output processing means for processing signals, which are received from a peripheral device and sent to a peripheral device, the input / output processing means comprising insertion means, the insertion means responsive to selected signals from the peripheral device to generate direct memory access (DMA) control instructions; submit these DMA control statements as inserted instructions for the first processing element to be processed, thereby at least triggering DMA transfers to the peripheral device, the first processing element processing both the instructions fetched in advance from the first instruction source and the inserted instructions, the inserted instructions being executed in the same manner as that of first command source prefetched commands and without affecting their sequence.
The first instruction source may be a memory element including an instruction cache element for storing digital values representing instructions and program steps, or an execution unit (CEU) comprising elements for providing signals to the instruction cache element to cause Commands are transmitted from the instruction cache element to the CEU.
An embodiment of the invention may include a command pipeline to interconnect the processors and to transport the commands. The insert elements can insert the inserted instructions into the instruction pipeline.
The inserted instructions may be of the same format as the instructions from the first instruction source, including a first set of digital instruction bits for inputting selected address signals and a second set of digital instruction bits for specifying selected instruction signals. Inserted commands that have this format may include cache management instructions inserted by the command cache element.
The I / O processors may include a peripheral interface unit (XIU) for controlling signals received from and transmitted through a peripheral device, and may further include a graphics controller for controlling signals sent to a peripheral device Display devices can be transmitted, and they may have text search elements to search data structures that are representative of text.
Selected processors may further include a register element for storing digital data values representing data. In accordance with this aspect of the invention, the inserts may provide inserted instructions to control movement of data into and out of register elements associated with the selected processors.
The inserted instructions may be arranged to effect execution of selected logical operations on the digital values stored in the register elements.
In addition, the processors may include capture elements that trigger a capture sequence in response to an applied capture signal. The insert elements may include elements for generated inserted instructions to generate the intercept signal, and the resulting intercept sequence may comprise any set of selected program steps. The processors may further include interrupt elements that respond to an interrupt signal to trigger an interrupt sequence. This interrupt sequence may comprise any set of selected program steps analogous to the capture sequence. The insertion elements may include elements to generate inserted instructions designed to trigger the interrupt sequence or to generate a trap signal in response to an interrupt signal.
Brief description of the figures
For a fuller understanding of the nature and objects of the invention, reference is made to the following detailed description and the accompanying drawings, in which:
Fig. 1 is a schematic diagram showing a multiprocessor structure used in connection with a preferred embodiment of the invention,
FIG. 2 is a block diagram of an exemplary processing cell as shown in FIG. 1; FIG.
Fig. 3 shows another embodiment of a processing cell constructed according to the invention,
4 shows single-clock instructions according to the invention,
Figures 5 and 6 show examples of instruction sequences that violate source register restrictions,
Figures 7-10 show resource usage and timing for example commands;
Fig. 11 shows an example of overlapping instructions associated with an intercept sequence;
Fig. 12 shows an exemplary branch instruction according to the invention,
13-19 show examples of program code which uses branching features according to the invention,
Fig. 20 shows an example of a program code for a remote execution,
Figs. 21-23 show features of interception events, errors and interrupts according to the invention, and Figs
Figures 24-32 show examples of program code associated with intercept sequences.
Description of the illustrated embodiments
Fig. 1 shows a multiprocessor structure 10 which may be used in connection with an embodiment of the invention. A structure of this kind is further described in Canadian Patent No. 1,320,003 assigned to the same assignee on November 8, 1988 for a multiprocessor digital data processing system incorporated herein by reference. The illustrated multiprocessor structure is exemplified and the following invention may be advantageously practiced in conjunction with digital processing structures and systems that differ from that illustrated in FIG.
The illustrated multiprocessor structure 10 comprises three information transfer domains: domain (0), domain (1) and domain (2). Each information transfer domain includes one or more domain segments characterized by a bus element and a plurality of cell interface elements. In particular, the domain (0) of the illustrated system 10 includes six segments, designated as 12A, 12B, 12C, 12D, 12E, and 12F, respectively. Similarly, domain (1) comprises segments 14A and 14B, while domain (2) comprises segment 16.
Each segment of the domain (0), that is, the segments 12A, 12B ... 12F, has a plurality of processing segments. For example, as shown, segment 12A includes cells 18A, 18B, and 18C, segment 12B includes cells 18D, 18E, and 18F, and so on. Each of these cells includes a central processing unit and a storage element disposed along an intracellular processor bus (not shown) are interconnected. In accordance with the preferred practice of the invention, the memory element contained in each of the cells stores all of the control and data signals used by its associated central processing unit.
As further illustrated, each domain (0) segment may be characterized as having a bus element providing a communication path for transmitting information representative signals between the cells of the segment. Thus, the illustrated segment 12A is characterized by the bus 20A, the segment 12B by 20B, the segment 12C by 20C, etc. As in Canadian Patent Application Serial No. 582,560 of the same assignee, filed on the 8th. In this regard, and which is incorporated herein by reference in more detail, information representative signals are passed between cells 18A, 18B and 18C of exemplary segment 12A using the memory elements associated with each of these cells. Special interfaces between these memory elements and bus 20A are provided by cell interface units 22A, 22B and 22C, as shown. Similar direct communication paths are provided in segments 12B, 12C and 12D between their respective cells 18D, 18E ... 18R as represented by cell interface units 22D, 22E ... 22R.
As shown in the diagram and mentioned above, the remaining information transfer domains, that is the domain (1) and the domain (2) each comprise one or more corresponding domain segments. The number of segments in each subsequent segment is less than the number of segments in the previous one. Thus, the number of the two segments 14A and 14B of the domain (1) is less than the number of the six (segments) 12A, 12B ... 12F of the domain (0), whereas the domain (2) only has the segment 16 and thus has the least number of all. Each of the segments in domain (1) and domain (2), that is, the "higher" domains, comprises a bus element for transmitting information representative signals within the respective segments. In the illustration, the segments 14A and 14B of the domain (1) comprise bus elements 24A and 24B, respectively, while the segment 16 of the domain (2) comprises the bus element 26.
The segment busses serve to transfer information between the component elements of each segment, that is, between the multiple domain routing elements of the segment. The routing elements in turn provide a mechanism for transferring information between associated segments of successive domains. Routing elements 28A, 28B, and 28C provide, for example, means for transmitting information to and from segment 14A of domain (1) and each of segments 12A, 12B, and 12C, respectively, of domain (0). Similarly, routing elements 28D, 28E, and 28F provide means for transferring information to and from segment 14B of domain (1) and segments 12D, 12E, and 12F, respectively, of domain (0). Further, domain routing elements 30A and 30B provide an information transfer path 16 of domain (2) and segments 14A and 14B of domain (1), as shown.
The domain routing elements interface with their respective segments via interconnects on the bus elements. Thus, the domain routing element 28A interfaces the bus elements 20A and 24A to the cell interface units 32A and 34A, respectively, while the element 28B interfaces the bus elements 20B and 24B to the cell interface units 32B and 34B, and so on. Similarly, routing elements 30A and 30B interfaces to their respective buses, that is, 24A, 24B, and 26 at cell interface units 36A, 36B, 38A, and 38B, as shown.
Figure 1 further shows a preferred mechanism connecting remote domains and cells in a digital data processing system constructed in accordance with the invention. The cell 18R located at a point spatially distant from the bus segment 20F may be connected to this bus and its associated cells (18P and 18O) via a fiber optic transmission line indicated by a dashed line. A remote interface unit 19 establishes a physical interface between the cell interface 22R and the remote cell 18R. Remote cell 18R is constructed and operated in a manner similar to the other illustrated cells and includes a remote interface unit for connection to the fiber optic port on its remote earth.
Similarly, domain segments 12F and 14B may be interconnected via a fiber optic link from their parent segments. As shown, the respective domain routing units 28F and 30B each have two remotely connected parts. With regard to the domain routing unit 28F, for example, a first part is directly connected to the cell interface 34F of the segment 14B via a standard bus connection, while a second part is directly connected to the cell interface unit 32F of the segment 12F. These two parts, which are identically constructed, are coupled together via a fiber optic link indicated by a dashed line. As above, a physical interface between the parts of the domain routing unit and the fiber optic medium is provided by a remote interface unit (not shown).
Fig. 2 shows an embodiment of the processing cells 18A, 18B ... 18R of Fig. 1. The illustrated processing cell 18A comprises a central processing unit 58 connected to an external device interface 60, a data sub-cache 62 and an instruction sub-cache 64 via the processor bus 66 and the command bus 68 is connected. The interface 60, the communications between external devices, e.g. B. Hard disks, provided via the external device bus, are constructed in a conventional manner.
The processor 58 may be any of a number of commercially available processors, such as the Motorola 68000 CPU, which is adapted to interface sub-caches 62 and 64 and operates under the control of a sub-cache co-executive unit via data and address control lines 69A and 69B in a conventional manner is designed to execute memory commands, as described below. The processing cells are further described in Canadian Patent Application 1,320,003 of the same assignee, filed on November 8, 1988 for a "Multiprocessor Digital Data Processing System", which is incorporated herein by reference.
Processing cell 18A further includes data storage units 72A and 72B connected to cache bus 76 via cache controllers 74A and 74B. The cache controllers 74C and 74D, in turn, provide communication between the cache bus 76 and the processing and data buses 66 and 68. As indicated in Figure 2, the bus 78 provides a connection between the cache bus 76 and the bus segment 20A of the domain (0) associated with the illustrated cell. Preferred embodiments for cache controllers 74A, 74B, 74C and 74D are disclosed in Canadian Patent No. 1,320,003 entitled "Multiprocessor Digital Data Processing System" filed on November 8, 1988 and Canadian Patent 2,019,300 filed on the same date as the present application for discussed an "Improved Multiprocessor System". The teachings of both applications are incorporated herein by reference.
In a preferred embodiment, data caches 72A and 72B include dynamic random access memory (DRAM) devices, each capable of storing up to 16 Mbytes of data. Subcaches 62 and 64 are static random access memory (SRAM) devices, the former being capable of storing up to 256K of data and storing the latter up to 256K of instruction information. As shown, cache and processor buses 76 and 64 provide 64-bit transfer paths, while the command bus 68 provides a 64-bit transfer path. One preferred construction of a cache bus 76 is provided in Canadian Patent Application 1,320,003, filed on November 8, 1988, entitled "Multiprocessor Digital Data Processing System", which is incorporated herein by reference.
Those skilled in the art will understand that the illustrated CPU 58 may be a conventional central processing unit and, in general, any device capable of issuing memory requests, e.g. B. may be an I / O controller or other processing element for a specific purpose.
The instruction execution of a processing cell described herein differs from conventional digital processing systems in several significant ways. The processing cell - for example 18A - has multiple processing cells or functional units - e.g. B. 58, 60 - which can execute commands in parallel. In addition, the functional units operate on the "pipeline" method to allow multiple instructions to advance simultaneously by overlapping their execution. This pipelining is further described in Canadian Patent Application Serial No. 582,560, filed on November 8, 1988 for a "Multiprocessor Digital Data Processing System", which is incorporated herein by reference. Further description of the commands discussed herein - including LOADS, STORES, MOVOUT, MOVB, FDIV, and others - can be found in Canadian Patent Application 2,019,300, filed the same day as the present and hereby incorporated by reference Reference is made.
A processing cell of one embodiment of the invention executes a sequence of instructions prefetched from the memory. The context of the execution may be partly defined by the architecture and partly by software. The architectural scope of the execution context may consist of a context address space, a privilege level, general registers, and a set of program counters. The context address space and privilege level determine which data in the storage system can reference. General registers that can be constructed according to known design practice are used for the calculation. These features are further described in Canadian Application Serial No. 582,560, which is incorporated herein by reference. The program counters define which part of the instruction stream has already been executed and what will be executed next, as will be described in more detail below.
Two units of time can be used to specify the timing of instructions. These units are referred to herein as "clocks" or "cycles". A clock is a unit of real time that has a duration defined by the system hardware. The processor executes a pickup with each cycle. A cycle takes one clock unless a "stall" occurs, in which case one cycle may include a larger integer number of clocks. Execution of instructions is described in terms of cycles and is data independent.
Pipeline demolition or pipeline stops may be due to an overhead load of subcache and cache management. Most load (LOAD) and store (STORE) operations proceed without a halt, however, any load, store, or memory control instruction may cause a halt to allow the system to retrieve data from or to the local cache retrieves distant cells. These delays are referred to herein as "stopping" or "stops". During a stop, the execution of other commands does not proceed and no new commands are prefetched. Stops do not refer to the command itself, but to the proximity of its data. Stops are measured in units of bars, and each stop is an integer number of bars. Although a CEU can stop while receiving data from the local cache, the programming model (expressed in cycles) remains constant.
As shown in Figure 3, a processing cell 18.1 in accordance with the invention may include four processing elements, also referred to herein as "functional units: the CEU 58, the IPU 84, the FPU 82, and the XIU 60. While Figure 3 illustrates a processing cell 18.1, which has four processing elements, those skilled in the art will recognize that the invention may be practiced in conjunction with a processing cell having more or fewer processing elements.
In particular, the CEU (central execution unit) prefetches all instructions, controls data fetching and storing (FETCH and STORE, referred to herein as LOADS and STORES), controls the instruction stream (branches), and performs arithmetic operations that are required for address calculations. The IPU (integer processing unit) executes integer arithmetic and logical instructions. The FPU (Floating Point Processing Unit) executes floating point instructions. The XIU (External I / O Unit) is a common execution unit that provides the interface for external devices.The XIU performs DMA (direct memory operations) and programmed I / O (inputs and outputs) and contains timer registers, and executes several commands to control a programmed I / O (input / output).
The processing cell 18.1 thus comprises a set of interconnected processors 58, 60, 82 and 84, including a CEU 58 for normal processing of an instruction stream including instructions from the instruction cache 64. The stream of instructions from the instruction cache 64 becomes indicated by dashed lines 86 in Fig. 3.
As shown in FIG. 3, at least one of the processors - in the illustrated example, the FPU 82 and the XIU 60 - may issue instructions, referred to herein as "inserted instructions" or as "insertion instructions", executed by the CEU 58 can. The stream of inserted instructions from the FPU 82 to the CEU 58 is indicated by dashed lines 88 in FIG. In an analogous manner, the movement of inserted instructions from the XIU 60 to the CEU 58 is indicated by dashed lines 90.
Moreover, as will be explained in greater detail below, these inserted instructions may be executed by the CEU in the same manner as the instructions from the instruction cache 64 and without affecting the execution order thereof. Furthermore, as will be explained below, the inserted instructions may have the same format as the instructions from the first instruction source, including a first set of digital command bits for specifying selected address signals and a second set of digital command bits for specifying selected instruction signals.
Inserted instructions having this format may include cache management instructions that may be inserted by the instruction cache 64 or by the cache controller 74D illustrated in FIG.
While FIG. 3 shows an instruction cache 64 as the source of instructions, alternatively, the source of instructions may also be a processor or an execution unit - including, in some circumstances, the CEU 58 - configured to provide signals to the instruction cache element to cause commands from the instruction cache element to be transmitted to the CEU 58.
As discussed above, the processing cell 18.1 may include a command pipeline having a command bus 68 for connecting the processors and executing the commands. In turn, the processors may include hardware and software elements to insert the inserted instructions into the instruction pipeline.
The XIU 60 illustrated in FIG. 3 may include input / output (I / O) modules for handling signals 70 received from and sent to peripheral devices, also referred to herein as external ones Devices are called. These I / O modules may include direct memory access (DMA) elements that respond to selected signals from a peripheral device to insert DMA commands that can be processed by the CEU 58 in the same manner as the commands of the first command source without affecting the processing sequence of the same. These processing sequences will be discussed in more detail below. Thus, the XIU 60 may include graphics control circuitry constructed in accordance with known design practice to control signals that are transmitted to a display device or include conventional text search elements to search for data structures that are typical of text.
Each processor 58, 60, 82, 84 shown in Figure 3 may have registers for storing digital values representing data and processor states in a manner which will be discussed in more detail below. The inserted instructions control the movement of data into and out of the registers and cause the execution of selected logical operations with values stored in the registers.
In a preferred embodiment of the invention, the processors shown in FIG. 3 may trigger a capture sequence in response to an applied capture signal, as discussed in more detail below. The interception sequence can be triggered by selected inserted commands. Similarly, the processors of cell 18.1 shown in FIG. 3 may include elements for initiating an interrupt sequence. Interrupt sequence and the inserted instructions may cause entry into an interrupt sequence or trigger a intercept signal in response to an interrupt signal. These features of the invention, including special instruction codes for triggering intercept and interrupt sequences, are explained below.
The four functional units shown in Fig. 3 operate in parallel. The cell pipeline can send two instructions with each cycle. Some commands, such as B. FMAD (Floating Point Multiplication and Addition) performs more than one operation. Others, such as Eg LD64 (load 64 bytes) will generate more than one result. Each can execute one command independently of the others.
Program instructions may be stored in memory in instruction pairs. Each pair consists of a command for the CEU or XIU and a command for the FPU or IPU. The former becomes the CX command and the latter is called the FI command.
The CEU may have three program counters (PCs) called PC0, PC1 and PC2. PC2 is also referred to here as the "pick-up PC". From the programmer's point of view, the processing element executes the instruction pair pointed to by PC0, will next execute the instruction pair designated by PC1 and will pre-fetch the instruction pair designated by PC2. When a command is executed, PC0 assumes the previous value of PC1, PC1 takes the previous value of PC2, and PC2 is renewed according to the CX command just executed. If this instruction was not a branch instruction or was a conditional branch instruction whose condition was not met, then PC2 is set to the value of PC2 + 8. If this value is not in the same segment as the previous value of PC2, the result is undefined. If this command was a selected branch, PC2 is updated to the branch destination.
In each cycle, the processor logically fetches the instruction pair from the memory designated by PC2 and begins executing both instructions in parallel in the pair designated by PC0. Thus, a single command pair can trigger work in the CEU and IPU, the CEU and FPU, the XIU and IPU, or the XIU and FPU. Those skilled in the art will recognize that since the functional units operate in accordance with the pipelined procedure, each unit in each cycle may begin to execute a new instruction, regardless of the number of cycles required to execute a command. However, there are limitations to the use of processor element or functional unit resources which affect the ordering of instructions by the compiler or programmer.
Certain orders have effects on more than one unit. For example, load and store instructions concern the CEU and the unit containing the source or destination registers. However, the processor may submit a load or save (LOAD or STORE) for the FPU or IPU on the same cycle as it is dispatching an execute instruction for the same unit.
The MOVB (shift between units) command shifts data between the registers of two units. Most data shifts between units require a single command; Moving data between the FPU and the fPU requires specification of MOVIN and MOVOUT instructions in a single instruction pair.
When the value of PC2 changes, the processor gets this instruction pair. The instructions are entered into the processor pipeline and occupy pipeline states in the order in which they occurred. Even if a command can not be removed from a pipeline, it can still be marked as "rejected" or "discarded". According to the invention, there are two types of discard, which is referred to herein as "discard result" and "discard start". The discard of a result occurs during "interception events". A trapping event or "trap" is an operating sequence triggered by the interception mechanism used to transfer control to privileged software in the case of interrupts and "exceptions". One exception, which will be described in more detail below, is a condition that occurs when a command executing in the FPU or IPU reports a "trap" and any operating command for the same unit in the cycles between the dispatch and the current one Cycle was sent. An exception is indicated when any error is detected as a direct result of prefetching or executing an instruction in the instruction stream. Exceptions include overflow of a data type, access violations, parity errors, and page faults.
An interception event can be triggered in two basic ways: by an error or an interrupt. An error is explicitly associated with the executing command stream. An interrupt is an event in the system that is not directly related to the instruction stream. Polling events (traps), errors and interrupts are described in more detail below.
Commands that are executed at the time of a trap can be discarded as a result. An instruction that is discarded in its result has been dispatched and processed by the functional unit, but does not affect the register or memory state, except for the fact that it reports the state in one or more special intercept status registers, which will be described below ,
A command that is discarded for launch is handled in a manner similar to that used for no_operation commands (NOP). A start-discarded command can only generate traps for prefetching this command. All other effects of a command discarded with respect to the startup are deleted. If a command is discarded at the time it reaches the PC0 state, it will not be started and will not use any resource normally used by the command. Start discarding is linked to the three execution PCs. It is possible to individually control the start discard for the PC0 CX and FI commands and to control the start discard for the PC1 instruction pair. The system software and hardware can individually change all three controllers of the discard. A trap results in a start discard of certain instructions in the pipeline. In addition, conditional branching systems allow the program to discard the two pairs of instructions that follow in the pipeline. This is called branch discard and causes the processor to discard the instructions in the branch delay with respect to the start. These features will be described in more detail below.
If PC1 is copied to PC0 when prefetching a command, this sets a start discard for both the CX and FI commands, depending on the old startup discard state of PC1. If the CX instruction just processed was a conditional branch that indicated branch discard and if no trap occurred, start discard for PC0 CX and FI and PC1 will be set after the PCs have been updated. An instruction typically causes a processing element to read one or more source operands, process them in a particular manner, and provide a result operand. According to the invention, "execution class" instructions can be classified into three groups according to how they read their operands and output results. The first group causes a functional unit to read the source operands immediately, calculate and output the result immediately. The result can be used by the next instruction pair. The second group causes a functional unit to read the source operands immediately, calculate them, and output the result after a certain delay. The result may be used by the Nth pair of instructions following the instruction, where N varies according to the instruction. The third group causes the function unit to read any source operands immediately, compute a portion of the result, after a certain delay, read other source operands, and output the result after a certain delay. The result may be used by the Nth pair of instructions following the instruction, where N varies according to the instruction.
Load and Store commands have several significant properties. All load instructions use the source address immediately and provide one or more results after a certain delay. In addition, all load instructions use their CEU index register source immediately. If stored in a CEU or XIU register, this value is obtained immediately. If stored in an FPU or IPU register, this result is obtained after a certain delay. The STORE-64BYTE (ST64) instruction uses its CEU index register source for the duration of the instruction and receives the various FPU and IPU source data after different delays.
At each cycle, the processor elements and functional units examine the appropriate instruction of the instruction pair addressed by the program counter (PC). An instruction within the instruction pair may be an instruction to one or both corresponding units (CEU / XIU or FPU / IPU), or indicates that there is no new work for any of the units. The latter case is indicated by non-action command codes, CXNOP and FINOP. As used herein, an action command is a command that is not FINOP or CXNOP and is not discarded. If there is an action, the appropriate unit will cancel (or start) that command. When the command execution is completed, the functional unit sets the command to "sleep". In general, the result of an instruction for the instruction pair is available following the instruction retirement, as shown in FIG.
FIG. 4 shows single-cycle instructions, defined herein as an instruction that is quiesced before the next instruction pair is considered for start-up and has a "result delay" of zero. All other commands are called multi-cycle instructions and have a non-fading result delay. The result delay corresponds to the number of instruction pairs that must exist between a particular instruction and the instruction that uses the result. All other timings are expressed by cycles from the dispatch time of the instruction. The first cycle is numbered 0.
Many commands can use interception events to indicate that the command did not execute successfully. The system disclosed herein provides users with considerable control over arithmetic intercept events or traps. Other traps may be used by the system software to implement features such as: A virtual memory as described in Canadian Application Serial No. 582,560, which is incorporated herein by reference. As will be described in more detail below, commands report traps at well-defined trap points expressed in cycles completed since the command was dispatched.
Each command reads its source registers at a specified time. All single cycles and many multi-cycle instructions read all their sources in cycle 0 of execution (ie with a delay of 0). Certain multi-cycle commands read one or more sources at a later time.
When a trap occurs, the software may perform a corrective action (eg, make the page available) and restart the user's program command stream. The program is generally not allowed to change the source registers during the time during which the instruction could be affected by an error. This property is called the source register restriction. Fig. 5 shows an example of a command sequence that violates this restriction.
Each functional unit uses a selected set of source registers. For example, the CEU {A, B} source register is used during all CEU instructions. It provides the index register used by a load or store, using the source operands of instructions of one execution class. The FPU {A, B} source register is used during FPU execution class instructions. It provides the first or first and second source operands used by execution class instructions. The FPU {C} source is used during FPU execution class triad instructions. It returns the third operand used by these commands. It is also used when the CEU accesses an FPU register with a memory type or MOVB instruction.
In addition, the IPU {A, B} source is used during IPU execution class instructions. It provides the first or first and second source operands used by execution class instructions. The IPU {C} source is used when the CEU accesses an IPU register with a memory type or MOVB instruction. The XIU {A, B} source is used during XIU execution class instructions. It provides the first or first and second source operands used by execution class instructions. It is also used when the CEU accesses an XIU register with a storage class command or MOVB instruction.
As described above, each instruction that produces a result has a result delay that indicates how many cycles pass before the result is available. During the result delay, the result registers are undefined. Programs can not rely on the old value of a result register of a command during the result delay of this command. This is called the result register restriction. If an exception occurs, all started commands are allowed to execute before the system software handler is called. Thus, it is possible that the result of a multi-cycle instruction will be issued before the defined result delay has expired. An instruction which uses the result register of a multi-cycle instruction during the result delay of this instruction, indefinitely receives one of (at least two) values of that register. FIG. 6 shows a sequence which violates this restriction. The FNEG instruction attempts to use the value that% f2 had before the FADD instruction. The FADD command describes% f2 in time for the FSUB command to read it. If the LD8 instruction gets a page fault or if an interrupt is displayed before FNEG is fetched, FADD is completed before FNEG is dispatched. This program therefore produces unpredictable results.
Each of the functional units has a number of internal resources that are used to execute instructions. These resources may only work for one command at a given time. At each point in time, each resource must be either idle or at most one command in use. This is called the resource restriction. Different functional units can detect resource constraint violations and cause a trap.
The CEU has only one resource exposed to conflict. This is the load / store resource used by all LOAD, STORE, MOVB, MOVOUT, and memory system commands. All commands except LD64 and ST64 (load and store 64 bytes) only use this resource during its third cycle (that is, with a delay of 2). The LD64 and ST64 commands use the load / store resource during the third through ninth cycles (delay 2-8). The resource usage of LD and MOVB instructions is shown in Figure 7, while Figure 8 shows the resource usage. The timing of an LD64 instruction is shown in FIG. 9, and that of an ST64 instruction is reproduced in FIG.
The IPU resources include a multiplier resource used by the MUL and MULH commands. Resources that belong to the FPU include result, divide, add, and multiply resources. The result resource is used by all FX commands to store results in registers. This resource is not used by certain CX commands - LD, ST, LD64, ST64, MOVOUT and MOVB - that work with FPU registers. It is used by MOVIN for% f registers.
The IPU divider resource is used in FDIV instructions, the IPU adder resource is used in many floating point computation instructions, and the IPU multiplier resource is used in many floating point computation instructions. No resource conflicts are possible in the XIU.
In the description of commands given here and in co-pending European application EP-A-0 404 560, the resource usage is indicated by the name of the resource, the number of cycles of the delay before using the resource and then the number the cycles during which it is used, specified in a table format. Thus, the timing of an LD instruction would be described as follows:
The timing for sources is a triple indicating [delay, cycles, source restriction]. "Delay" is the number of cycles until the resources are used, it is counted from 0, starting with the sending of the command. "Cycles" is the number of cycles during which the source is used after the delay has expired. "Source constraint" is the number of cycles during which the source should not be changed counted after the delay has expired. "Result Delay" is the number of instructions that must occur between the instruction pair and the first instruction that references the result.
Since some instructions require several cycles to be executed or report a state of exception, the CPU maintains a common execution PC for the FPU and for the IPU. If an exception occurs, then the trap handler may need to examine the co-execution PC to determine the current address of the faulty instruction, as described in more detail below. The CEU performs a similar function with load / memory type instructions so that exceptions of an ST64 instruction can be resolved.
When a command is intercepted (trap), there must be no action commands for the same unit in the command sections between the containing command pair and the command pair where the trap was reported. This is called the trap PC limitation. It is possible to place an action command in the command pair where the trap is reported or in any command pair thereafter. The application of this restriction depends on the requirements of the operating system and the user application.
These coding practices ensure that a command sequence generates deterministic results and that any exception that occurs is resolved by the system software or forwarded to the user program for analysis. In all cases, it is possible to determine exactly what was happening right now to create a temporary state such as For example, to correct a missing page, change data, and finally start the calculation again. The program must not override the result register restriction or any resource limitation. The program can ensure that data-dependent errors do not occur during FI commands, either because it knows the data or by using the modifier for the no trap command. In the latter case, the program may decide to examine different condition codes (such as @IOV) to determine if an arithmetic error has occurred or not. If no errors can occur, it is possible to violate the source register restriction and violate the trap PC restriction for FI statements. It is also possible to violate this restriction even if traps occur, unless exact knowledge of the trap command is required. Whether the CEU source register restriction is violated or not depends on the system software, but typical implementations do not guarantee the results of such violations. Figure 11 shows an example of overlapping instructions that obey the rules for exact traps.
As described above, the CEU has three PCs that define the current instruction stream. A branch instruction alters the prefetch PC (PC2) to the target value of the branch. A branch instruction may be a conditional branch (B ** instruction), an unconditional branch (JMP or RTT instruction) or an unconditional subroutine jump (JSR instruction). Conditional branches allow the program to compare two CEU registers or a CEU register and a constant, or to examine a CEU condition code. The prefetch PC is changed when the branch condition is met and is simply incremented if the branch condition is not met.
In order to track the instruction pairs executed by a program, it is necessary to keep track of the values of the three PCs as the program progresses. A program may specify branch instructions in a branch delay. This technique is referred to herein as remote instruction execution and will be described in more detail below. Any JMP, JSR, or RTT command that significantly changes the segment portion of PC2 may not have a "PC-related" branch in that branch delay. A PC-related branch is defined as any conditional branch or unconditional branch which specifies the program counter at its index register.
Branching is always followed by two instructions in the processor pipeline. These instructions are called branch delay instructions. The branch delay is actually a special case of the result register delay, where the result register of a branch is randomly PC0. For unconditional branches, these commands are always executed. For conditional branches, their execution is governed by the branch discard option of the branch instruction. Since branch instructions may occur in the branch delay slots of another branch, the control by the branch discard option does not necessarily mean that the two instruction pairs following a branch in the program memory in turn are fetched or executed become.
This feature will be explained in more detail below.
There is no source registry restriction, branch registry restriction or branch restriction resource restriction. This is because the fetch PC is changed by the branch instruction, and since every exception relating to the new prefetch PC is reported at the time the value arrived at PC0 and the instruction pair is sent. For optimal performance, branch delays may be filled with instructions that logically belong to the branch, but do not affect or are influenced by the branch itself. If no such commands are available, the delay sections can be filled with NOPS.
A typical branch instruction is shown in FIG. The JMP command is prefetched together with its partner. The partner starts the execution. The two delay pairs are then prefetched and begin execution. Then, the instruction pair is prefetched and executed at the destination address.
The programmer or compiler may fill the branch delay of an unconditional branch instruction with instructions that precede or follow the branch itself. The branch delay of conditional branches may be harder to fill. In the best case, instructions preceding the branch may be inserted in the branch delay. These must be executed regardless of whether the branch is executed or not. However, instructions from the location before the branch are not always available for shifting into the branch delay. Filling the branch delay of conditional branches is facilitated by branch discarding. In particular, conditional branch instructions allow the programmer to specify whether the branch delay instructions should be executed based on the result of the branch decision. The branch instruction may specify fulfillment on fulfillment if the instructions are to be branched out when the branch is taken or discard if not satisfied, if they are to be "forked", if the branch is not taken, and it can never be discarded "if the commands should always be executed. The mnemonic of the conditional assembly branch uses the letters QT, QF or QN to indicate which branching semantics is required. The branching warp results in start-up discard when the instructions arrive at the branch delay at PC0 and PC1.
If instructions are to be used before the branch in the branch delay, a "never discard" is specified. If such instructions are not available, the programmer may fill the delay with commands from the target route and select discard on default or select from below the branch and select discard upon fulfillment. The decision on which source to refill depends on which commands can simply be shifted and the prediction at the time of code generation whether the branch is likely to be taken. Examples are shown in Figs. 13-19.
Figs. 13-15 show an example of a padded branch delay. In this example, code is shifted from a position before a branch to the branch delay, thereby removing two NOPS from the instruction stream. In particular, Fig. 13 shows the original code sequence with NOPS in the branch delay. The commands executed are FI_INSA0 / CX_INSA0, FI_INSA1 / CX_INSA1, FI_INSA2 / CX_INSA2, FI_INSA3 / jmp, FI_NOP / CXNOP, FI_NOP / CXNOP, FI_INSB4 / CX_INSB4, FI_INSB5 / CX_INSB5. This sequence leads to the waste of two cycles.
Alternatively, the optimized code sequence with filled branch delay may be used, as shown in FIG. As shown there, the commands FI_INSA1 / CX_INSA1 and FI_INSA2 / CX_INSA2 are shifted to the branch delay, saving two instruction cycles. The commands executed are FI_INSA0 / CX_INSA0, FI_INSA3 / jmp, FI_INSA1 / CX_INSA1, FI_INSA2 / CX_INSA2, FI_INSB4 / CX_INSB4, FI_INSB5 / CX_INSB5, which means that no cycles are wasted. It is also possible to arrange the FI commands independently of the rearrangement of the CX commands, as shown in FIG.
Certain programming constructions, such. For example, the loop makes it likely that a branch will be taken. If the branch is taken most likely, the first two instructions from the target branch may be placed in the branch delay. Failure branching is used to produce correct results in case the branch is not to be taken. If the branch is indeed taken, the instruction cycles are preserved. If not, the cycles are forked and thus the correctness of the program is preserved. Fig. 16 shows a code sequence using NOPS in a branch delay, while Fig. 17 shows an optimized code sequence with destination branch instructions in the branch delay and branch discard. According to FIG. 16 If the branch is not taken, the executed instructions are FI_INSA0 / CX_INSA0, FI_INSA1 / CX_INSA1, FI_INSA2 / CX_INSA2, ..., FI_INSA7 / CBR.QN, FINOP / CXNOP, FINOP / CXNOP, FI_INSC0 / CX_INSC0, resulting in wasted two Cycles leads. When the branch is taken, the executed instructions are FI_INSA0 / CX_INSA0, FI_INSA1 / CX_INSA1, FI_INSA2 / CX_INSA2, FI_INSA7 / CBR.QN, FINOP / CXNOP, FINOP / CXNOP, FI_INSC0 / CX_INSC0, resulting in two wasted cycles.
Fig. 17 shows that, for padding the branch delay, the user can copy the two commands FI_INSA0 / CX_INSA0 and FI_INSA1 / CX_INSA1 into the branch delay and can select branch discard if not satisfied (if the branch is taken) and set the destination branch. If the branch is not taken, the executed commands FI_INSA0 / CX_INSA0, FI_INSA1 / CX_INSA1, FI_INSA2 / CX_INSA2, ..., FI_INSA7 / CBR.QF are branched, branched, FI_INSC0 / CX_INSC0, resulting in two wasted cycles. If the branch is taken, the commands executed are FI_INSA0 / CX_INSA0, FI_INSA1 / CX_INSA1, FI_INSA2 / CX_INSA2, ..., FI_INSA7 / CBR.QF, FI_INSA0.1 / CX_INSA0.1, FI_INSA1 / CX_INSA1.1, FI_INSA2 / CX_INSA2, so that in the most probable case no cycles are wasted.
In some programs, some branches are most likely to be skipped or not taken. Such a branch is a test of a very rarely occurring condition, such. B. the arithmetic overflow. If the branch is likely to be skipped, the first two instructions after branching may be inserted into the branch delay. It uses branch fulfillment on fulfillment to produce correct results if branching should be taken. If the branch is indeed not taken, two instruction cycles are saved. If this does not happen, the two cycles are forked and the execution time is not improved. A corresponding example is shown in FIGS. 18 and 19.
Fig. 18 is a code sequence having NOPS in the branch delay. If the branch is not taken, the executed instructions are FI_INSA0 / CX_INSA0, FI_INSA1 / CBR.QN, FINOP / CXNOP, FINOP / CXNOP, FI_INSB0 / CX_INSB0, FI_INSB1 / CX_INSB1, FI_INSB2 / CX_INSB2, resulting in two wasted cycles. If the branch is taken, the executed instructions are FI_INSA0 / CX_INSA0, FI_INSA1 / CBR.QN, FINOP / CXNOP, FINOP / CXNOP, FI_INSC0 / CX_INSC0, FI_INSC1 / CX_INSC1, FI_INSC2 / CX_INSC2, resulting in two wasted cycles.
Figure 19 shows an optimized code sequence with branched instructions in branch delay and branch discard. As shown in FIG. 19, for the branch delay fill, the user may shift the INSA1 and INSA2 instructions to the branch delay and select branch discard if satisfied, saving two instruction cycles if the branch is actually not taken. When the branch is taken, the executed instructions are FI_INSA0 / CX_INSA0, FI_INSA1 / CBR.QT, branched, branched, FI_INSC0 / CX_INSC0, FI_INSC1 / CX_INSC1, FI_INSC2 / CX_INSC2, resulting in two wasted cycles. If the branch is not taken, the executed instructions are FI_INSA0 / CX_INSA0, FI_INSA1 / CBR.QT, FI_INSB0 / CX_INSB0, FI_INSB1 / CX_INSB1, FI_INSB2 / CX_INSB2, so that in the most probable case there are no wasted cycles.
Because of the three PCs used to determine the instruction stream, it is possible to "remotely" execute one or two instructions that are not associated with the linear stream of a program. These operations may be performed in the manner of a sequence as shown in FIG. The program sequence of FIG. 20 executes the instruction pair at addresses 000, 008, 010, 100, 018, 020 and so forth. By shifting the JMP from address 008 to address 0x10, two remote command pairs (at 100 and 108) are executed. These special sequences do not support remote commands that contain branches as CX commands.
The transmission of interrupts and DMA inserts instructions into the processor pipeline between successive instructions of the instruction stream. These commands are referred to herein as insertion instructions. The CEU controls the "right" to insert commands and occasionally ignores or discards an insertion command.
The architecture allows any arbitrary instruction to be inserted, however, the functional units may be designed to use only a limited portion of the instruction set. These insertion commands do not change the PCs. Insertion instructions use cycles and allow the pipelines of all processing elements or functional units to proceed, just as an ordinary instruction does.
The effect of insert instructions on the programming model is that an insert instruction can cause a result to appear sooner than expected. This is because the insertion instruction has a physical pipeline stage and a hidden cycle appears. If the program is doing the result register restriction, there is no change in the logical execution of the program, only in terms of the time required to do it. Insertion instructions can not be discarded by the state of branch discard and start discard belonging to the logical pipeline (PC0, PC1, PC2), but they may be discarded in their result or, in the event of an exception, be dispatched in the physical pipeline.
The following examples show how the CCU and XIU can use inserted commands, as further described in the concurrently filed appendix. The XADDR, XCACHE, XNOP, and XDATA instructions and subpages, subblock, and other memory operations presented in the following examples are further described in Canadian Patent 2,019,300, filed concurrently herewith, and US Pat Canadian Patent 1,320,003, both of which are incorporated herein by reference. The CCUs and the XIU supply the CX portion of a command pair and the CEU logically supplies a FINOP command. The CCUs and the XIUs handle the processor busses at the same time that they insert a command to supply the operands of the command. The CCU and the XIU insert two or more contiguous commands.
Rinse a bottom out of the subcache
XADDR
XADDR
xnop
xcache
xcache
Loading or saving data is not in the pipeline
XADDR
xnop
xnop
xdata
Load or save two objects (each 8 bytes or less) pipelined
XADDR
XADDR
xnop
xdata
xdata
Loading or saving a subblock
XADDR
XADDR
xnop
xdata
xnop
xnop
xnop
xnop
xnop
xnop
xnop
Requesting an interrupt
xtrap
xnop
The insertion instructions may be encoded as part of a program by diagnostic software. In a preferred embodiment of the invention, the CEU implements the FI instruction which accompanies the CX instruction. The program must initiate a special process to supply or extract data as required. This can be done, for example, by using MOVIN or MOVOUT commands.
In a preferred embodiment of the invention, a trap mechanism is used to transfer control to privileged software in the event of interrupts and exceptions. The taxonomy of traps is shown in FIG. As shown there, a trap can be triggered in two basic ways: by an error or by an interrupt. An error is explicitly associated with the executed instruction stream and occurs when certain combinations of data, states and a command occur. An interrupt is an event in the system that is not directly related to the instruction stream.
Errors are still classified as software errors or hardware errors. Software errors are those errors that are part of the expected execution of the program and that may be caused by user or system software as part of the implementation of a calculation model. Hardware failures can occur when hardware detects unexpected errors while it is working. Preferably, the processor handles errors immediately, but sometimes can also refrain from handling interrupts.
The most important property of the trap sequence is its ability to suspend execution and maintain the execution state of the processor so that the software can resume execution in a manner that is transparent - that is, "invisible" - to the original program is. Such sequences become possible through the configuration of processor registers and the limitations described in Canadian Patent Application Serial No. 582,560, which is incorporated herein by reference. However, a program which violates the applicable constraints may suffer from indeterminate results or the inability to (again) resume instruction flow after trap handling. The trap with the highest priority is referred to herein as a RESET. A RESET can not be masked.
Between three and six PC values are required to specify the instructions in execution at the time of a trap event. As described more fully in Canadian Application Serial No. 582,560, which is incorporated herein by reference, the CEU pipeline is described by PC0, PC1 and PC2. During a trap event, these PCs are backed up in the CEU registers% TR0,% TR1, and% TR2 (also referred to as% C0,% C1, and% C2). The CEU keeps track of the addresses of the most recent FPU and IPU commands. These addresses are called the co-execution PCs.
The co-execution PCs for a given functional unit will display the PC value of the last working instruction sent by that unit as long as that instruction is not yet discarded as a result of a previous instruction reporting an exception in any functional unit. This mechanism, in conjunction with the Trap PC limitation, allows software to determine the exact command PC that is responsible for an exception, regardless of the result time of the command.
The execution point of the XIU is always described by PC0 at the time of the interception event because the XIU does not have overlapping execution. During a trap event, the co-executing PCs are saved in! PC_IPU and! PC_FPU, as indicated in FIG. The CEU also provides! PC_SCEU to help the system software handle errors that result from an ST64 command. The CEU and the common execution PCs are collectively referred to as the execution PCs and are shown in FIG.
When a command executed in the FPU or the IPU reports a trap event and any operation instruction for the same unit has been started in the cycles between the start and the current cycle, this unit reports an "imprecise exception". Otherwise the exception is called "precise". According to the invention, the instruction pair indicated at PC0 may contain a command for the same unit without affecting the precision of the exception reported by the previous instruction.
An exception is marked as "imprecise" if the processor does not have enough information to accurately state the state of the calculation. If there is an instruction in the pipeline after the instruction reporting the exception, then there is no PC information for the intercept instruction because the CEU has already updated the co-execution PC. Such calculations can not be meaningfully restarted and the "imprecise_exception" indicator is set to 1 in! I_TRAP and / or! F_TRAP, as appropriate.
The trap mechanism stores the trap state values in various registers. These registers include the following:
% TR0 saves the PC of the instruction at the trap point.
% TR1 saves the PC of the first command after the trap point.
% TR2 saves the PC of the command ready for pickup (the second behind the trap point).
! CONTEXT stores the context register of the attached instruction stream.
! TRAP stores the trap register, which records the causes of the trap.
! PC_SCEU stores the PC of the last LD or ST command that was sent and reported a trap, or the last LD64 or ST64 command that was not dispatched and was not result-discarded by any other exception. If an STT or memory system error is displayed in! TRAP, this register will contain the PC of the conflicting command.
! PC_FPU stores the PC of the last working, submitted FPU command that may have produced the present exception. This register is valid only if! TRAP indicates an FPU exception and "F_TRAP" indicates that the exception was accurate.
! F_TRAP stores the FPU trap register, which records any FPU exceptions.
! PC_IPU stores the PC of the last started IPU commands that might have generated the current exception. This register is only valid if! TRAP indicates an IPU exception and! I_TRAP indicates that the exception was precise.
! I_TRAP stores the details of the IPU exception when an IPU exception is indicated in! TRAP.
! X_TRAP stores details of the XIU exception when an XIU exception is displayed in! TRAP.
Upon entering the trap handling software, the state of execution is specified by these registers. In addition, the causes of the trap are indicated by the contents of these registers, which is more fully described in U.S. Patents 5,055,999 and 5,251,308, which are incorporated herein by reference.
Gaps in the instruction stream may occur when a multi-cyclic instruction indicates an exception after cycle 0 of execution. A command is started if its address exists in PC0. In the next cycle, the execution PCs are updated to describe the next three commands to be executed. If this multi-cycle command reports a precise exception, its address will be in the co-execution PC (! PC_FPU, or! PC_IPU) or! PC_SCEU. The address of the instruction is lost if the program starts another operation instruction for the same unit within the result delay of that instruction.
After a trap (interception event) has occurred, the system software may signal an error to the program or resolve or eliminate the trap cause. To restart a command stream without gaps, the kernel executes a simple sequence that restores the execution PCs and register state. The user or system software must complete all "pending" instructions before the instruction stream can be restarted, as discussed in greater detail below.
A CEU vulnerability may exist if an ST64 command reports an exception in its end cycle of execution. This is only the case if! PC_SCEU is valid (an STT or memory system exception has occurred), but is not equal to% TR0. The current instruction was started seven cycles before the instruction pair designated by PC0 when the trap occurs.
If multiple commands are executed in the IPU or FPU while a trap is occurring, the trap state of this unit is imprecise. An imprecise state can not be meaningfully analyzed so that the system software typically signals an error to the user process and does not allow the previous instruction stream to be restarted. If the trap state is accurate, it is possible that the trap was caused by the command at the trap point (PCO /% TR0) or by a command started before the trap point.
When the processor signals a trap, it detects a trap point. The trap point is one of the PCs in the sequence of instruction pairs that are executed by the program. All pairs of commands at the trap point are treated specifically according to the sources of the trap and the existing commands.
For single-cycle instructions that indicate exceptions, the trap point is the PC of the intercepted command. Some multi-cycle instructions report exceptions in the zero cycle of execution or at a later time. In many cases, the later trap point is the cycle before the result becomes available. The CEU attains a steady state, assures the execution state, and enters trap handling, as described below.
When a trap is displayed, the processor stops fetching a command, refuses to allow inserted instructions, and waits for all co-execution units to hibernate all expiring instructions. If any of these commands report exceptions, then each exception is included as part of the trap information. Each co-execution unit can be put at rest by successfully completing its actions or by reporting a state of emergency and discarding its results. If a command reports no state of emergency while it is being completed, no further action is required. If an idle command started before the command pair at PC0 reports an exception condition, this command represents a gap in the command stream before the trap point. Its state and address must be secured for the software for use in filling in the gap.
The CEU handles the instruction pair at PC0 (the intercept point) corresponding to the start discard state of the instruction stream, the intercept source, and the CX instruction at PC0. For example, interrupts are generated when the XIU or a CCU inserts an XTRAP instruction into the instruction stream. An inserted command does not affect the program PCs; the XTRAP appears before starting the command pair at PC0. Thus, if the trap was triggered by an interrupt (regardless of whether any functional unit reported a trap as part of reaching a steady state), the command pair will not start at PC0. The commands at PC0, PC1 and PC2 are discarded as a result.
When the CEU updates the execution PCs (PC0, PC1, PC2), it tries to fetch the command displayed by PC2. It is possible that an error is signaled during an address translation (STT violation) or even while the CEU receives the command sub-block (eg page_error). The error state is assigned to the instruction pair and follows it through the pipeline. If the command pair is not result-discarded, the exception is not reported. Otherwise, the exception is reported and the commands at PC0, PC1 and PC2 are result-discarded.
If a trap is reported by the CEU or XIU, the CX instruction will result-rooke at PC0. A service request is treated like any other CEU command reporting a trap in cycle 0. If the FI command was not already discarded at PC0, it will be discarded. The commands at PC1 and PC2 are result-discarded.
The trap sequence causes the FI command to be rejected. If the CX instruction at PC0 is not a memory-type instruction, it is result-discarded. If the CX command in PC0 CX is a memory type command, it can be executed. The memory type command can be normal or report a trap. In the former case PC0 CX is marked as start-discarded. If the instruction reports an exception of the memory type, it becomes part of the trap state; the state of the start discarding is not changed. This behavior ensures that a memory-type command is executed only once.
The commands at PC1 and PC2 are result-discarded. The cause or causes of the trap are stored in the trap registers. The CEU sets its trap register,! TRAP, to indicate the causes and sources of the trap. Each co-execution unit reporting an exception also sets its trap register -! F_TRAP,! I trap, or! X-trap - to include further details of the exception it has detected.
Fig. 23 shows the instruction execution model and the occurrence of a trap. When a program uses conditional branch discard, it is important that its discard state be preserved as part of the trap state. The state of branch discard affects the start discard state. When an inserted XTRAP instruction causes a trap, the trap occurs before or after the conditional branch instruction. In the first case, the trap causes a start discard of the conditional branch; when the instruction stream is restarted, the conditional branch is retrieved and started. In the second case, the state of branch discard causes a start discard to be set for PC0 CX / FI and PC1 CX / FI, and then the inserted instruction (which is logically unassigned PC0) is executed and causes a trap. Thus, the stored state of start discarding indicates that the two instruction pairs should be discarded when the instruction stream is restarted.
If an instruction prior to the conditional branch or before the FI instruction paired with a conditional branch indicates a trap, the conditional branch instruction is result-discarded and start discard is not affected. When the instruction stream is restarted, the conditional branch instruction pair is restarted and branch discard occurs when the pipeline PCs are updated.
The trap sequence stores the state of the instruction stream in the processor registers. The contents of these registers are described in U.S. Patents 5,055,999 and 5,251,308, which are incorporated herein by reference. To protect these register values from being destroyed by another trap, the trap sequence shuts down further traps. Trap handling software will re-enable traps if the registers are safely stored in memory. In particular, to save the execution state, the hardware trap sequence shuts down further traps by setting! CONTEXT.TE = 0, stores PC0 (the intercept point) in trap register 0 (! TR0), PC1 (the next PC) in Trap Register 1 (% TR1) stores, PC2 (the Command Fetch PC) stores in Trap Register 2 (% TR2), the context registers,! CONTEXT, modifies the privilege level,! CONTEXT.PL, at the old privilege level, ! CONTEXT.OP, secure, copies the state of the start discard in! CONTEXT.QSH, and saves the current co-execution PC and! PC_SCEU. The validity of! PC_FPU and! PC_SCEU depends on the execution state reported by the individual functional units or processor elements.
The PCs stored in% TR0,% TR1 and% TR2 and the information about the start discard information stored in! CONTEXT define the command stream to be resumed. The Trap Register! TRAP indicates whether the command pair has caused an exception on PC0 (% TR0). The PCs stored in% TR1 and% TR2 are not related to the cause of the trap.
The co-execution unit PCs (! PC_FPU,! PC_IPU, and! Pc_xiu) held in the CEU are valid only if the! TRAP control register indicates that the corresponding co-execution unit has reported an exception. Finally, the processor must gather the information describing the causes of the trap and store it in the trap registers,! TRAP,! F_TRAP,! X_TRAP, and! I_TRAP.
In the third stage of the trap sequence, the processor begins execution of the trap handler and changes the privilege level of the processor to the highest privilege by setting! CONTEXT.pl = 0, whereby the state of the start discard is cleared, so that no Commands are discarded and the PCs are set to effect a sequential execution beginning with the context address zero.
Except for the above, the context of the previous instruction stream is taken from the trap handler. The system software must ensure that the context address 0 is assigned by the ISTT of each execution context. Trap handling may decide to save the state and then change it to some other context. Since trap handling works at privilege level 0, it has access to the general core registers,% C0-% C3.
Since the trap handler takes over the context address space of everything that was executed when the trap occurred, each context address space must map the code and data segments that the trap handler needs to start. The data mappings may be hidden from the user command stream by restricting access to level 0 only. The trap sequence claims the number of clocks required to quiesce any shared commands in progress plus three instruction cycles. Interrupts are not accepted during these cycles.
Errors are traps that are directly related to the instruction stream being executed. For example, the KSR command is used to request an operating system service or a debugging breakpoint. The system software defines an interface through which a program passes information that more accurately reflects the particular nature of its request. A service request has the same trap properties as any other CX instruction that is faulty in cycle 0 of execution. This is shown separately because restarting the command stream requires explicit activity of the system software.
The KSR instruction is defined as a single-cycle instruction and is intercepted in cycle 0 of execution. A KSR command never completes normally. The address of the KSR is recorded in% TR0. The interception status indicates the service request and also indicates whether the paired FI command is faulty. If the instruction stream must be restarted, the system software must change the discard state so that the CX instruction is discarded. Note that this process causes the FI instruction to complete after the service call is completed.
An exception is signaled when any error is detected as a direct result of fetching or executing an instruction in the instruction stream. Exceptions include overflow of a data type, access violations, parity errors, and page faults. The causes of exceptions are described by the registers! TRAP,! F_TRAP,! I_TRAP and! X_TRAP.
Since multiple instructions are executed in parallel in the co-execution units, more than one exception can be signaled in the same cycle. When a trap (a trap event) is signaled, the software must examine all source indications in! TRAP to determine the sources of the traps. Individual units report the additional status in their private trap registers.
When a CX instruction signals an exception in cycle 0 of a design, it is discarded and the corresponding FI instruction is result-discarded. If the FI instruction or both instructions in a pair signal an exception in their first execution cycle (cycle 0), then the instruction pair is discarded and the trap point is the instruction pair with the exception of an ST or ST64 instruction partner of an FPU or IPU command indicating an exception.
Thus, the stored state is as it was before the exception occurred. The address of the command that caused the exception is stored in% TR0.
In the example shown in Fig. 24, the add8 instruction has a result delay of 0 and reports an overflow in cycle 0 of the embodiment. The register value of% TR0 is 0,% TR1 is 8,% TR2 is 0x10. In addition,! PC_IPU is 0 and the exception is precise.
As described above, an exception signaled by an instruction after cycle 0 of the execution results in a gap in the instruction stream, which is indicated by the fact that the corresponding PC register is not equal to% TR0. If the exception is inaccurate, the PC register may or may not be different from% TR0 and does not indicate that the instruction is signaling an exception.
In the example of a command sequence shown in Fig. 25, the FMUL command has a result delay of 2 and may report a trap in cycle 0 or cycle 2 of execution. If the exception is reported in cycle 0 then% TR0 is 0,% TR1 is 8,% TR2 is 0x10. The value of! PC_FP0 is 0 and the exception is precise.
The example of an overlapping embodiment shown in Fig. 26 is similar to that of Fig. 25 of the previous example but with data causing the FMUL command in cycle 2 to be erroneous. In this case% TR0 is 0 · 10,% TR1 is 0 · 18,% TR2 is 0 · 20,! PC_FPU is 0. This exception is accurate.
In the example shown in Fig. 27, FMUL again reports an exception in cycle 2. Regardless of whether the instruction reports an exception at 0x10,% TR0 is 0x10,% TR1 0x18,% TR2 0x20 ,! PC_FP0 is 0. This exception is accurate.
In the example of the instruction sequence of Fig. 28, the FMUL instruction again reports an exception in cycle 2. If the FADD instruction reports an exception in cycle 0, then% TR0 is 8,% TR1 is 0x10,% TR2 0 · 18,! PC_FP0 is 8, the exception is imprecise. Otherwise,% TR0 is 0x10,% TR1 is 0x18,% TR is 0x20, and PC_FP0 is 8, and the exception is imprecise.
Fig. 29 shows a command sequence in which data is such that the FMUL command is not intercepted. If the FADD instruction reports an exception in cycle 0, then TR0 is 8,% TR1 is 0x10,% TR2 is 0x18,! PC_FPU is 8, the exception is precise. If the FADD instruction reports an exception in Cycle 2 then% TR0 is 0x18,% TR1 is 0x20,% TR2 is 0x28. If the FADD instruction at 0x10 is an FPU operation instruction, so the FADD exception is imprecise and% PC_FPU is 0x10. Otherwise, the FADD exception is precise and! PC_FPU is 8.
In the example shown in Fig. 30, the FMUL instruction has data that does not cause an error. The CX instruction at 0 and 8 experiences a trap in cycle 0 (page_fault). The FPU discards its started commands and the result from FMUL is delivered to% f2. % TR0 is 8,% TR1 is 0 · 10,% TR2 is 0 · 18,! PC_FPU is not valid. The CEU exception is precise and! PC_SCEU is 8, indicating that an ST64 command was not the cause of the memory system error.
The instruction sequence shown in Fig. 31 takes advantage of the fact that memory-type instructions have only one cycle delay before reading the source. This code sequence will only produce correct results if no trap can occur if the store instruction is addressed by PC0.
Similarly, if the result delay for a LOAD instruction is two cycles, it is similarly possible to compress the sequence if it is known that no error can occur if the store instruction is addressed by PC0. The sequence shown in Fig. 32 is precise and can be restarted even if a CX error occurs at the address 0 or 0x105.
All LD, LD64, and ST commands capture exceptions in cycle 0 of a run. Thus, if an STT or memory system error (e.g., missing_segment, missing_page) is reported at TR0 and! PC_SCEU is set to the address of this instruction, the ST64 instruction may fail in cycle 0 (relative to STT) or cycle 7 (detected by the storage system) report. Errors that are not program-technical in nature (such as Parity errors) can occur at any time and the value of% TR0 is unpredictable.
The XIU and the storage system can use inserted instructions to request interrupts and perform direct memory access (DMA). In a preferred embodiment of the invention, these commands do not cause a trap. Instead, every inserted command reports an error condition to its source. The source can then inform the CEU of the error with an interrupt. Inserted commands can be started if any previous command causes a trap.
However, as described above, interrupts are events that are not associated with the main instruction stream but require the attention of the processor. Interrupts can be generated by the storage system or the XIU while performing asynchronous activities. The generator supplies the interrupt to the CEU by inserting an XTRAP command. The CEU will only accept one interrupt at a given time and may sometimes reject all interrupts. Interrupt sources are responsible for maintaining interrupts until the CEU accepts them. The! TRAP control register indicates the source of the interrupt.
Interrupts may include memory system interrupts, inter-cell interrupts, and XIU interrupts. A memory system interrupt is an interrupt generated by the memory system. A cache generates interrupts whenever it detects errors in asynchronous operations that it performs, in the data it holds, or in its view of the storage system. The priority of the memory interrupts is determined by the configuration position of the cell that detects it.
An intercell interrupt is a special case of a memory system interrupt and occurs only as a result of a write to the CTL $ CCU_CELL_INT control position of a cell. Because of the hierarchical layout of the SPA space, processors can route interrupts to specific processors or to groups of processors at one level of the hierarchy.
An XIU interrupt is caused by the timing of an I / O execution. This aspect of I / O operations is more fully described in the appended application, which is incorporated herein by reference.
When an XTRAP (interrupt request) instruction is inserted into the instruction before any instruction that causes an exception, the interrupt is accepted and the following instructions are start-discarded. Moreover, if the XTRAP instruction is inserted in the pipeline and any previous instruction causes a trap before the XTRAP is started, the XTRAP is ignored so that the interrupt will be rejected in the result. Accordingly, interrupt requests do not cause a double trap reset. When this happens, the response to the asynchronous instruction requesting the interrupt indicates that it has been rejected.
When an interrupt is received, the normal trap sequence is triggered. This causes all co-execution unit instructions to be executed and report their exception status, if any. If any concurrently executing command reports an exception, the interrupt and exception status are merged and reported in! TRAP.
In addition, when the trap sequence is completed, a new instruction stream is started at the context address 0. This code, which is executed at privilege level 0, is the software trap handling that completes the trap mechanism. Their job is to save the trap state stored in the registers, send control to the appropriate software to handle the trap, and later resume or discard the suspended instruction stream.
The trap sequence turns off traps. A processor will receive a double trap reset if another error occurs before the traps are turned on. However, XTRAP instructions inserted by the CCUs or the XIU to signal an interrupt will not create a trap while traps are turned off. If traps are turned back on before the machine state is safely stored, this condition can be overridden by another trap, precluding a reboot analysis. Therefore, trap handling of the system software preferably first secures the trap state and then releases the traps as fast as possible. This minimizes the amount of system software that must be coded to avoid errors. Trap handling must! TRAP examine and determine which other registers are valid.
Since the trap handler routine expires in the context of the previous intercepted instruction stream, it must also back up any registers that might be interfering with it, such as the. Eg! CONTEXT,! I_CONTEXT,! F_CONTEXT and certain CIU / IPU / FPU general registers.
Certain traps require that the system respond to a condition and later resume the suspended instruction stream as if the trap had not occurred. Others cause the instruction stream to be eliminated or restarted at a location other than where the trap occurred. These reactions are collectively referred to herein as "resumption of the instruction stream."
Trap handling starts at privilege level 0, where it must re-enter and then act on the special trap. The system software can handle the trap at privilege level 0 and then resume the command stream. The trap state may also be passed to less privileged code by invoking a new instruction stream. This software handling can have corrective action and then issue a service request for the kernel to restart the intercepted command stream. The system software or less privileged code may further decide that the intercepted instruction stream is abandoned and a new instruction stream is started.
An important aspect in handling a trap involves filling gaps in the instruction stream left by FPU, IPU or ST64 instructions which have reported exceptions. The need to fill gaps is the basis for the source register restriction described above. To handle these gaps, the software must "manually" execute the pending commands. In some cases, the instruction is effectively executed by changing its result register or memory. For example, a calculation that had an overflow could be handled by setting the result register to the largest valid value.
It is also possible to change source values or the machine state and to execute the faulty command again. An example of such modification and re-execution involves changing an arithmetic operation or making a page available. The system software may provide a special context that starts the pending command at its actual context address and immediately re-invokes the core.
In one specific context example, PC0 is the address of the pending command and PC1 and PC2 are the addresses of a KSR instruction (with a special operand code) in the text space of the system software. The command paired with the pending command specifies the start discard and PC1 has cleared the start discard. This context starts the desired command. If the pending command reports an exception in cycle 0 of execution, a trap immediately appears. Otherwise, the KSR command is executed and causes a trap; if the pending command was a single cycle command, it has been successfully completed. If the pending command is a multi-cycle command, it can still report an exception when the processor reaches steady state, or it can be completed normally.
When the core re-enters, it examines the trap state. When the pending command completes successfully, the original cued command stream can be restarted. Otherwise, the system software must handle the new error or discard the instruction stream. If there are multiple pending instructions in the original intercepted instruction stream, they can be resolved in sequence using the above technique. The system software must take precautions to ensure that users do not attempt to execute the special KSR instruction at inappropriate times.
Most of the context of an intercepted instruction stream can be recovered while traps are still turned off. For example, all FPU and IPU general registers, the! F_context register, and most CEU registers are not used by trap handling while traps are off. Assuming that the trap handling software implements a suitable recursive model, any trap that occurs during the recovery of that state would eventually restore any state that it has changed. The system software normally runs with traps on, but must disable traps as the final part of resuming an intercepted instruction stream. If trap handling has been initially invoked, this is required to prevent a recursive trap from destroying the state. Next, the! CONTEXT register is restored. Finally, the trap PCs in% TR0,% TR1 and% TR2 are reloaded and the following code is executed:
RTT 0 (% TR0) / * Unlock Traps, set privilege level! CONTEXT.OPL restores. Replace Discard! CONTEXT.QSH (with two Delay Commands). Branch to the command pair at the interception point, which is identified by% TR0. * /
JMP 0 (% TR1) / * Jump to the first command after the interception point. * /
JMP 0 (% TR2) / * Jump to the second command after the interception point. * /
This sequence restores the status of the suspended instruction stream and begins execution at the trap point as if no trap had occurred. The use of three consecutive branch instructions is indeed a means of remote control engineering as described above. The privilege level changes and trap enable by the RTT instruction take effect when executed at% TR0. The two JMP instructions have already been retrieved by the segment containing this code. All subsequent command retrievals use the restored values of! CONTEXT.PL to detect privilege violations. The processor state is therefore restored as well as the suspended code resumes execution. The conditions stored by the trap are explicitly saved again before the return sequence and are not changed by the sequence. The start discard information, which is restored by the RTT instruction, controls an individual discard of the first CX and FI instruction and the discarding of the second instruction pair. This capability is necessary to allow interrupts to occur between a conditional branch and the commands it discards and to allow the system software to take control of the first pair of instructions which is restarted.
The system software does not need to take any special precautions regarding the ISTT or memory system to ensure that the addresses at% TR0,% TR1 or% TR2 are accessible. This is because any exception relating to the fetching of these commands is reported during the trap phase of this command. If z. For example, if the page containing the addresses specified by% TR0 is missing, the command page error occurs at that address.
If the system software calls a less privileged error handler, signals a user program or begins a new process, the software must start a new instruction stream. This can be accomplished by providing information equivalent to that stored by the trap handling software and then resuming execution of this interrupted instruction stream. This is the preferred technique for switching from the core mode to the user mode.
An embodiment of the invention efficiently achieves the objects described above, in addition to those which have become apparent from the foregoing description. In particular, the invention provides multiprocessing methods and apparatus in which each processor can optionally issue instructions to other processing elements, thereby increasing the parallelism of execution and speeding up the processing speed.
It will be understood that changes may be made in the above construction and sequences of operation without departing from the scope of the invention. For example, the invention may be practiced in conjunction with multiprocessor structures other than those shown in FIG. It is therefore intended that all matter contained in the above description or illustrated in the accompanying drawings should be interpreted as illustrative and exemplary and not in a limiting sense.
Contents7
12 sheets
Sheet 1 Sheet 2 Sheet 3 Sheet 4 Sheet 5 Sheet 6 Sheet 7 Sheet 8 Sheet 9 Sheet 10 Sheet 11 Sheet 12
115 members in 8 offices
Priority claims5
| Document | Office | Kind | Date |
|---|---|---|---|
| 37032589 | United States of America | A | |
| 37032589 | United States of America | A | |
| 37032589 | United States of America | – | |
| 370325 | – | – | – |
| US19890370325 | – | – | – |
Members115
| Document | Office | Kind | |
|---|---|---|---|
| EP0322116A2 | European Patent Office (EPO) | A2 | |
| EP0322117A2 | European Patent Office (EPO) | A2 | |
| JPH01281555A | Japan | A | |
| JPH022451A | Japan | A | |
| EP0322117A3 | European Patent Office (EPO) | A3 | |
| EP0322116A3 | European Patent Office (EPO) | A3 | |
| CA2019299A1 | Canada | A1 | |
| CA2019300A1 | Canada | A1 | |
| EP0404559A2 | European Patent Office (EPO) | A2 | |
| EP0404560A2 | European Patent Office (EPO) | A2 | |
| JPH03116234A | Japan | A | |
| JPH03129454A | Japan | A | |
| WO9115088A1 | World Intellectual Property Organization (WIPO) | A1 | |
| US5055999A | United States of America | A | |
| CA2042610A1 | Canada | A1 | |
| CA2042290A1 | Canada | A1 | |
| EP0458552A2 | European Patent Office (EPO) | A2 | |
| EP0458553A2 | European Patent Office (EPO) | A2 | |
| CA2042291A1 | Canada | A1 | |
| EP0468542A2 | European Patent Office (EPO) | A2 | |
| EP0473777A1 | European Patent Office (EPO) | A1 | |
| US5119481A | United States of America | A | |
| EP0404560A3 | European Patent Office (EPO) | A3 | |
| EP0458552A3 | European Patent Office (EPO) | A3 | |
| EP0468542A3 | European Patent Office (EPO) | A3 | |
| JPH04230146A | Japan | A | |
| EP0458553A3 | European Patent Office (EPO) | A3 | |
| JPH05501041A | Japan | A | |
| CA2078311A1 | Canada | A1 | |
| EP0534665A2 | European Patent Office (EPO) | A2 | |
| CA1320003C | Canada | C | |
| US5226039A | United States of America | A | |
| EP0534665A3 | European Patent Office (EPO) | A3 | |
| US5251308A | United States of America | A | |
| JPH05298190A | Japan | A | |
| EP0404559A3 | European Patent Office (EPO) | A3 | |
| US5282201A | United States of America | A | |
| JPH0619784A | Japan | A | |
| US5297265A | United States of America | A | |
| EP0593100A2 | European Patent Office (EPO) | A2 | |
| US5335325A | United States of America | A | |
| US5341483A | United States of America | A | |
| CA1333727C | Canada | C | |
| EP0473777A4 | European Patent Office (EPO) | A4 | |
| EP0593100A3 | European Patent Office (EPO) | A3 | |
| US5761413A | United States of America | A | |
| JP2780032B2 | Japan | B2 | |
| JP2782521B2 | Japan | B2 | |
| US5822578A | United States of America | A | |
| EP0534665B1 | European Patent Office (EPO) | B1 | |
| AT178417T | Austria | T | |
| ATE178417T1 | Austria | T1 | |
| DE69228784D1 | Germany | D1 | |
| EP0936778A1 | European Patent Office (EPO) | A1 | |
| US5960461A | United States of America | A | |
| DE69228784T2 | Germany | T2 | |
| EP0322116B1 | European Patent Office (EPO) | B1 | |
| AT189070T | Austria | T | |
| ATE189070T1 | Austria | T1 | |
| DE3856394D1 | Germany | D1 | |
| DE3856394T2 | Germany | T2 | |
| EP1016971A2 | European Patent Office (EPO) | A2 | |
| EP1016977A2 | European Patent Office (EPO) | A2 | |
| EP1016978A2 | European Patent Office (EPO) | A2 | |
| EP1016979A2 | European Patent Office (EPO) | A2 | |
| EP1020799A2 | European Patent Office (EPO) | A2 | |
| EP0404560B1 | European Patent Office (EPO) | B1 | |
| AT195382T | Austria | T | |
| ATE195382T1 | Austria | T1 | |
| DE69033601D1 | Germany | D1 | |
| JP3103581B2 | Japan | B2 | |
| CA2042291C | Canada | C | |
| CA1341154C | Canada | C | |
| EP0322117B1 | European Patent Office (EPO) | B1 | |
| AT198673T | Austria | T | |
| ATE198673T1 | Austria | T1 | |
| DE3856451D1 | Germany | D1 | |
| DE69033601T2 | Germany | T2 | |
| CA2019300C | Canada | C | |
| DE3856451T2 | Germany | T2 | |
| US6330649B1 | United States of America | B1 | |
| CA2042290C | Canada | C | |
| CA2019299C | Canada | C | |
| EP1182544A2 | European Patent Office (EPO) | A2 | |
| EP0458553B1 | European Patent Office (EPO) | B1 | |
| AT214175T | Austria | T | |
| ATE214175T1 | Austria | T1 | |
| DE69132945D1 | Germany | D1 | |
| EP0404559B1 | European Patent Office (EPO) | B1 | |
| AT218225T | Austria | T | |
| ATE218225T1 | Austria | T1 | |
| US2002078310A1 | United States of America | A1 | |
| DE69033965D1 | Germany | D1 | |
| CA2042610C | Canada | C | |
| ES2173075T3 | Spain | T3 | |
| DE69132945T2 | Germany | T2 | |
| DE69033965T2This record | Germany | T2 | |
| EP0458552B1 | European Patent Office (EPO) | B1 | |
| EP0593100B1 | European Patent Office (EPO) | B1 | |
| AT231259T | Austria | T |
1 legal event, as the office reported them to INPADOC
Events
| Event | Code | |
|---|---|---|
| Ceased/non-payment of the annual feeCeased8339 | 8339 |
Numbers
- Publication
- 69033965
- Publication, DOCDB
- 69033965
- Publication, EPODOC
- DE69033965T
- Application
- 69033965
- Application, DOCDB
- 69033965
- Application, EPODOC
- DE1990633965T
Titles2
- German
- Multiprozessorsystem mit mehrfachen Befehlsquellen
- English
- Multiprocessor system with multiple command sources
Classification
- CPC, 17
- G06F9/3865
- G06F9/30003
- G06F9/30032
- G06F9/30043
- G06F9/3005
- G06F9/30072
- G06F9/3828
- G06F9/383
- G06F9/3836
- G06F9/3851
- G06F9/3877
- G06F9/3885
- G06F11/004
- G06F11/0724
- G06F11/073
- G06F11/0793
- G06F9/38585
- IPC, 6
- G06F9 30
- G06F9 38
- G06F11 00
- G06F11 07
- G06F15 16
- G06F15 173
