Method of parallel processing of commands and arrangement of a machine operating with a scaled set of commands
Abstract
Described is a scalable compound instruction set machine and method which provides for processing a set of instructions or program to be executed by a computer to determine statically which instructions may be combined into compound instructions which are executed in parallel by a scalar machine. Such processing looks for classes of instructions that can be executed in parallel without data-dependent or hardware-dependent interlocks. Without regard to their original sequence the individual instructions are combined with one or more other individual instructions to form a compound instruction which eliminates interlocks. Control information is appended to identify information relevant to the execution of the compound instructions. The result is a stream of scalar instructions compounded or grouped together before instruction decode time so that they are already flagged and identified for selective simultaneous parallel execution by execution units. The compounding does not change the object code results and existing programs realize performance improvements while maintaining compatibility with previously implemented systems for which the original set of instructions was provided. <IMAGE>
Term
No projected expiry on record.
- Priority
- Filed
- Granted
- Today
2 claims: 1 independent, 1 dependent
- 1Patent claims Zastrzeżenia patentowe 1. A method of processing an electrical pulse stream associated with an instruction stream in a computer system in which the electrical pulse stream consists of successive sequences of electrical impulses associated with successive instructions of the instruction stream, and each of the electrical impulse sequences designates an instruction combined with one of many instructions with different categories, characterized by that subsequent sequences of electrical pulses are tested and the ability of the instruction category from adjacent, sequentially combined instructions to execute is determined, and additional electrical pulses are generated from subsequent sequences of electrical pulses, depending on the result of the test stage. 1. Sposób przetwarzania strumienia impulsów elektrycznych, powiązanego ze strumieniem instrukcji w układzie komputera, w którym strumień impulsów elektrycznych składa się z kolejnych sekwencji impulsów elektrycznych, powiązanych z kolejnymi instrukcjami strumienia instrukcji, a każda z sekwencji impulsów elektrycznych wyznacza instrukcję połączoną z jedną z wielu instrukcji o różnych kategoriach, znamienny tym, że bada się kolejne sekwencje impulsów elektrycznych i określa się zdolność kategorii instrukcji z przyległych, połączonych kolejno instrukcji do wykonywania równoległego, oraz generuje się dodatkowe impulsy elektryczne, z kolejnych sekwencji impulsów elektrycznych, zależnie od wyniku etapu badania.
226 paragraphs in 14 sections, as filed
The present invention relates to a method for processing an electrical pulse stream associated with a stream of instructions processed in a computer system.
There is a known way to improve the performance of computer systems by executing commands in parallel. Parallel operation consists of the existence of separate functional blocks capable of performing two or more of the same or different commands simultaneously.
Another known way to improve the performance of your computer's layout is so-called piping. The pipe consists in the separation of the functions given by the computer system into independent sub-functions and the allocation of separate system parts and stages to perform individual sub-functions. Each stage is adapted to work in one basic machine cycle. Piping provides some kind of parallel processing because it allows you to perform many functions independently. Ideally, one new instruction is inserted into the pipeline during the cycle, and each instruction in the pipeline is at a different stage of execution. The work is analogous to the operation of a factory assembly tape with many positions on which the product is located at various stages of assembly.
However, the benefits of parallel execution or piping often fail to be achieved due to delays caused by, for example, data addiction and addiction caused by systems. An example of data dependence is the so-called input-output dependency, where the first order must write its result before the second order can read and then use it. An example of systemic dependence is when one command uses a specific fragment of the system and the other command must also use the same fragment of the system.
One method previously used to prevent addiction, sometimes called pipe hazard, is so-called dynamic scheduling. Dynamic scheduling is based on the fact that by using specialized systems it is possible to reorganize the sequence of instructions after they are downloaded to the pipe for execution.
Attempts have been made to improve performance by using so-called static scheduling, which is done before the instruction stream is retrieved for execution. Static scheduling is done by moving codes and resetting the order string to be executed. As a result of this re-setting, an equivalent instruction stream is generated, allowing more full use of system capabilities through parallel processing.
165 524
Such static scheduling usually occurs during complications. However, the set orders remain in their original form, and parallel processing still requires some dynamic decision before execution itself to determine whether the next two orders should be executed sequentially or in parallel.
This scheduling method can improve the overall performance of a piped computer system, but it cannot meet the ever-increasing performance requirements. Therefore, many other general-purpose processing methods relate to the use of simultaneous processing at the command level, regardless of the use of pipelining. ' For example, a further increase in the degree of simultaneity can be achieved directly - by issuing multiple commands in one machine cycle in so-called super-scale machines, instead of indirectly, by dynamic ordering of individual instructions, or by using vector machines. The name superscalar machine is used to refer to a machine that issues multiple orders in one machine cycle to distinguish it from scalar machines that issue one order per cycle.
In a typical superscalar machine, machine codes in the downloaded instruction stream are decoded and dynamically analyzed by the instruction issuing logic to determine if the instructions can be executed in parallel. The criteria for such fast scheduling are specific to each type of structure of the instruction set, as well as to the use of this structure in a given processor. Its effectiveness is therefore limited by the complexity of the logic that determines which combinations of commands can be executed in parallel and what is likely to increase cycle time. The increase in cycle time and the degree of systemic complexity occurs even in super-scale machines with hundreds of different orders.
There are other shortcomings with dynamic scheduling, static scheduling, or a combination of both. For example, it is necessary to review each scalar order again each time it is downloaded for execution to determine its parallel processing capability. It is not possible to identify and mark previously those scalar orders that have the ability to perform in parallel.
Another disadvantage of dynamic scheduling used in super-scale machines is the way the instructions are checked for parallel processing. Superscalar machines check scalar orders based on the description of their machine codes, and it is not possible - taking into account the system implementation. Because orders are issued on a FIFO basis (first in - first out), it is therefore impossible to group them selectively to eliminate or reduce addiction. There are several ways to look for opportunities to consider system requirements for parallel processing. In one of these systems working with static scheduling, called a machine with a very long command word, a high-capacity compiler is used, which sets orders so that dynamic scheduling is much simpler. With this approach, the compiler must be more complex than standard compilers, so that a larger window can be used to search for simultaneity in the command stream. However, the resulting commands do not have to be machine codes compatible with the given machine structure, thus solving one problem creates new additional problems. The problems also increase significantly due to the frequent occurrence of program jumps limiting the simultaneous processing.
Thus, none of the known approaches to parallel processing are satisfactorily versatile to minimize all possible addictions, while avoiding the essential redesign of a set of instructions and the use of complex logic circuits for dynamic decoding of downloaded instructions.
Therefore, it is necessary to improve the data processing process, which would allow the execution of existing machine orders in parallel, in order to increase processor performance. Because the number of commands executed in a second depends on the basic cycle time of the processor and the average number of cycles needed to execute the order, there is a need to find a solution that takes into account both of these parameters. More precisely, a mechanism is needed to reduce the number of cycles needed to execute the command at a given structure. In addition, improvement is needed to reduce the complexity of the system needed
165 524 to ensure command execution and minimize cycle time increase. It is also very desirable for the proposed improvement to ensure compatibility with known structures by introducing simultaneous processing at the level of the order of both new and existing machine code.
The essence of the method of processing the electric pulse stream proposed according to the invention, associated with the instruction stream in a computer system, in which the electric pulse stream consists of successive sequences of electric impulses, associated with successive instructions of the instruction stream, and each of the electric impulses sequences determines the instruction connected with one of many instructions in different categories, that successive sequences of electrical impulses are tested and the ability of instruction categories from adjacent, sequentially connected instructions to perform parallel are determined, and additional electrical impulses are generated from subsequent sequences of electrical impulses, depending on the result of the test stage. Preferably, according to the invention, in the stage of testing subsequent sequences of electrical impulses, the sequence of electrical impulses of the first instruction and the sequence of electrical impulses of the second instruction are examined for the possibility of their parallel execution and then the sequence of electrical impulses of the second instruction and the sequence of electrical impulses of the third instruction are examined parallel execution.
An advantage of the inventive solution is to provide static analysis, before decoding and executing orders, of the sequence of existing instructions to generate compound instructions created by grouping adjacent commands suitable for parallel execution. Another advantage of this solution is the addition of control information to the instruction stream, including information indicating where the compound instruction starts, as well as indicating the number of existing instructions included in each of the compound instructions. In addition, the method of the invention enables pre-processing of the stream to create compound instructions so that they can be used programmatically or systemically at multiple points in the computer system prior to decoding and execution. The coupled orders created are suitable for use in compound instruction structures with variable lengths and instructions mixed with data, and can be used in RISC structures in which instructions are usually of a fixed length and not mixed with data.
The solution according to the invention used in the Scalable Compound Instruction Set Machine (SCISM) enables coupling of a scalar stream, i.e. grouping together before the decoding of the order, so that they are already provided with markers and identified for selective parallel execution by relevant execution blocks. Because such coupling does not change machine codes, existing programs can be implemented with increased efficiency while maintaining compatibility with already implemented computer systems.
The invention will be explained in more detail in the example of its implementation based on the attached drawing, in which Fig. 1 shows a high-level processing diagram of the invention, Fig. 2 - a time diagram of a single-processor implementation showing a parallel execution of some non-dependent orders that can be grouped selectively in coupled instruction stream, fig. 3 - time plot of microprocessor implementation with the representation of parallel execution of scalar and conjugated, non-addictive commands, Fig. 4A and Fig. 4B - an example of a possible selective categorization of some orders executed by a known scalar machine, Fig. 5 - a typical path taken by the program from the source code to the actual implementation, Fig. 6 - a flowchart showing the creation of a set of instructions coupled from a program in symbolic language, Fig. 7 - flowchart showing the execution of the program contained in the coupled instruction set, Fig. 8 - analytical map of instruction stream texts with reference points identifying orders, Fig. 9 - analytical map of instruction stream texts with different length of instructions without reference points, illustrating the relative arrangement of possible identifier bits couplings, fig. 10 - an illustration of the logical implementation enabling the coupling of instructions for operating the text of the instruction stream from Fig. 9, Fig. 11 - the flowchart for coupling an instruction stream equipped with reference markers for identifying points denoting instruction boundaries, Fig. 12 - reference compound control field, Fig. 13 - a flowchart of developing and applying coupling rules used in a particular configuration of computer system blocks being a partially compiled set of instructions, Fig. 14 - the way in which different groups of non-dependent instruction pairs form multiple conjugate instructions for the target sequential execution or execution with jumps, Fig. 15 - the manner in which various clusters of valid non-interrelated triples of orders form multiple compound instructions intended for the target sequential execution or execution with a jump, Fig. 16A and Fig. 16B - a map of actions for coupling a stream of instructions similar to that in Fig. 9, including orders of varying lengths without reference points marking the boundaries, and fig. 17 - a map of instruction group operations for part of the System / 370 instruction set shown in Fig. 4.
The invention relates to the processing of a stream of electrical impulses associated with a stream of instructions processed in a computer system, where in the further content of the description the processing of instructions is described for its simplicity. A set of instructions intended for execution by a computer system is pre-processed to statically determine which non-reciprocal instructions can be combined into compound instructions, and to attach control information to identify such compound instructions. This determination is based on coupling rules that are developed for a set of instructions with a specific configuration. Existing scalar orders are grouped based on the analysis of their operands, the use of machine systems and its functions, so that grouping orders by coupling to avoid indelible additions is based on comparing order groups instead of comparing specific orders.
As shown generally in Figure 1, instruction pair 20 takes the binary scalar instructions stream 21, with or without data contained therein, and selectively groups some of the adjacent scalar instructions to form coded compound instructions. The resulting stream of 22 compound instructions thus combines scalar instructions not suitable for parallel operation and conjugate instructions formed by groups of scalar instructions suitable for parallel execution. When a scalar instruction is given to the instruction processing unit 24, it is passed to the appropriate functional unit for serial execution. When a coupled instruction is fed to the processing unit 24, each of its scalar components arrives at the respective functional block or addiction removal block for simultaneous parallel execution. Typical functional blocks include the arithmetic logic unit (ALU) 26, 28, the floating point arithmetic unit (FP) 30 and the memory address generation unit (AU) 32.
As shown in Figure 2, the invention may be implemented in a single processor environment in which each execution unit executes a scalar (S) order or a coupled scalar (CS) order. The instruction stream 33 containing the sequence of scalar and conjugated scalar instructions contains control marks (T) assigned to each compound instruction. The first scalar order 4 can be executed individually by functional unit A in cycle 1, three conjugated scalar orders contained in the triple conjugate command 36 identified by the tag T3 can be executed by functional units A, C and B in cycle 2, another conjugate command 38 identified by the T2 tag it can contain a pair of commands executed in parallel by functional units A and B in cycle 3, the second scalar instruction 40 can be executed individually · by functional unit C in cycle 4, four coupled scalar orders of large group compound instruction 42 can be executed by functional units AD in cycle 5, and the third scalar 44 can be executed by functional unit A in cycle 6.
During implementation it is important that such multiple conjugate orders are suitable for parallel execution in some computer system configurations. The invention could potentially be implemented in a multiprocessor environment, such as shown in Figure 3, where the coupled instruction is treated as a block for parallel processing by one of the CPU (central processing units). As shown in the figure, the same instruction stream could only be processed in two cycles as follows. In the first cycle, CPU 1 1 executes the first scale command 34, CPU # 2 functional units execute the triple compound instruction 36, and CPU 0 3 functional units execute the two compound scalar instructions in the compound instruction 38. In the second cycle, CPU # 1 executes the second scalar instruction 40, the CPU # 2 functional units execute four conjugated scalar orders in the compound instruction 42, and the CPU functional unit 3 executes the third scalar instruction 44.
One example of a computer architecture that can be adapted to work with compound instructions is the IBM / 370 system level command architecture, in which multiple scalar commands can be issued for execution in each machine cycle. In this context, the machine cycle means each of the steps, or piping stages, needed to execute the scalar command. Scalar commands operate on operands representing individual values. When the instruction stream is coupled, adjacent scalar orders are grouped selectively for simultaneous or parallel execution. Generally speaking, enabling command coupling is provided for classes of instructions that can be executed in parallel and that provide no dependence between the elements of a compounded command that would prevent them from being executed in the system structure. When a compatible instruction sequence is found, a compound instruction is created.
More specifically, the System / 370 instruction set can be divided into categories of orders that can be executed in parallel in a particular computer system configuration. Orders in some of these configurations from these categories can be combined, i.e., combined with orders of the same category or with orders of some other categories to combine a compound order. For example, part of the System / 370 instruction set may be divided into categories shown in figure 4. The basis for this categorization are the basic requirements for System / 370 commands and their systemic use in a typical system configuration. Other System / 370 commands are not particularly important for coupling in this example implementation. This does not exclude the possibility of their coupling according to the invention. It should be noted that the system structures require that the execution of the compound instruction could be a control of a horizontal microcode enabling, in order to increase efficiency, the use of processing simultaneity for other commands not included in the coupling and not included in the categories of Fig. 4.
One of the most frequently occurring sequences in System / 370 programs is execution of COMPARES TM or RX commands (C, CH, Cl, C1.I, CLM), the result of which is used to control the execution of the following BRANCH conditional jump commands immediately after them (BC, BCR). You can improve performance by using COMPARE and BRANCH commands in parallel, and this is sometimes used dynamically for high-performance command processors. Some difficulties relate to quickly identifying all the various COMPARE class elements and all BRANCH instruction class elements during the decoding process in a typical structure. This is one of the reasons why superscalar machines usually only take into account a small number of specific scalar commands suitable for parallel execution. In contrast, this limited grouping based only on the current comparison of two specific commands in the described example is ignored, because the analysis of all elements of the classes takes place in advance for a certain time, in order to determine the appropriate coupling rules when creating a compound order guaranteeing some work.
The problem resulting from dynamic grouping of individual orders after they are downloaded is illustrated by the example of such a two-way coupling of fifty-seven single orders from Fig. 4, forming a 57 X 57 matrix with over three thousand possible combinations. This is in sharp contrast to the 10 X 10 matrix of Fig. 17 for the same number of instructions taken into account from the possible combinations of categories as envisaged according to the invention. Multiple classes of orders can be executed in parallel, depending on which system structure they are intended for. In addition to the couplable COMPARE and BRANCH pairs described above, many other couplable combinations (see Fig. 17) are suitable for parallel execution, such as LOADS (category 7) coupled with the commands of the format RR (category 1), BRANCHES (category 3-5) coupled with LOAD ADDRESS (category 8) and the like.
165 524
In some cases, the sequential order may affect the capabilities of parallel execution, and therefore decides which two adjacent orders can be coupled. Because of this, the header of line 45 identifies the category of the first instruction in the byte stream, and the header of column 47 defines the category of the next instruction. For example, BRANCHES (category 3-5) as always coupled 49 to any of the following SHIFTS (category 2) orders, while SHIFTS (category 2) follows BRANCHES (category 3 - 5) are only "sometimes coupling 51. Status" sometimes marked "S" on the map of Fig. 17 can often be changed to "always" marked on the map by "A" by adding new system blocks to the computer system configuration. For example, one can consider a two-way coupling configuration that has no mooring shifting unit for addiction reduction, and instead has a conventional arithmetic technology unit (ALU) and a separate shift block. In other words, when operating with addicted ADD and SHIFT commands, no system blocks are used to eliminate addiction. Consider the following sequence of orders:
ARR1, R2 SRL R2 through D2
It is clear that this pair of commands can be coupled for parallel execution. In some cases, such orders may not be clutchable due to their indelible addiction, as shown in the following sequence of instructions:
ARR1, R2 SRLR1 via D2
Therefore, the map indicates that category 1 (AR) orders after category 2 (SRL) orders are sometimes closable 53. By attaching an ALU that eliminates some dependency, such as the addiction and addiction relationship shown above, the designation on the map in Fig. 17 can change from S to A. Accordingly, the coupling rules must be updated to reflect changes made to a particular computer system configuration.
As another example, consider orders that belong to category 1 and are combined with orders from the same category in the following order sequence:
ARR1, R2 SR R3, R4
This sequence is free from data gambling addiction and gives the following result of two independent System / 370 instructions.
R1 = R1 + R2 R3 = R3-R4
Performing such a sequence would require two independent and working in parallel in a two-on-one ALU system adapted to the structure of the command level. It is understood that in a system configuration equipped with two ALUs, these two commands can be grouped to form a compound instruction. This example of combining scalar orders can be generalized to all pairs of instructions in the sequence, free from addictions between data as well as from systemic addictions.
Any specific instruction processor will have an upper limit on the number of individual instructions that can be covered by a compound instruction. This upper limit is specific to a system or program unit that creates compound instructions, so that compound instructions will not contain more separate commands, e.g. double, triple, quadruple groups, than the maximum capacity of the corresponding system structure. This upper limit is a strict consequence of the system implementation in a specific configuration of the computer system - it neither limits the total number of commands taken into account as coupling candidates nor the length of the group window in a given code sequence analyzed for coupling.
In general, the greater the length of the group window analyzed for coupling, the greater the result of favorable combinations for coupling, the degree of processing simultaneity can be achieved. Consider the order sequence in this respect in Table 1 below.
165 524
Table 1
<td>XI</td><td>any joinable command</td>
<td>X2</td><td>any joinable command</td>
<td>LOADR1, (X)</td><td>load R1 from memory location X</td>
<td>ADDR3, R1</td><td>R3 = R3 + R2</td>
<td>SUB R1, R2</td><td>R1 = R1-R2</td>
<td>COMPR1, R3</td><td>compare R1 with R3</td>
<td>X3</td><td>any joinable command</td>
<td>X4</td><td>any joinable command</td>
If the system structure imposes an upper join limit of two, i.e. at most two commands can be executed in parallel in the same cycle, there are several ways to join this sequence of instructions, depending on the scope of the connector's operation.
If the scope of action is four, the connector considers it together (X1, X2, LOAD, ADD) and then shifts, one order at a time and examines (X2, LOAD, ADD, SUB), then (LOAD, ADD, SUB , COMP), then (ADD, SUB, COMP, X3) and (SUB, COMP, X3, X4) to make the following optimal pairing as candidates for a combined order:
[-X1] [X2 LOAD] [ADD SUB] [COMP X3] [X4-]
This optimal pairing completely frees you from addictions, between LOAD and ADD, and between SUB and COMP, and gives you the option of joining X1 with the preceding order and X4 - with the order following it.
On the other hand, a superscalar machine dynamically combining orders in pairs in its logic circuits to issue orders working on the strict FIFO principle, will give as the candidates for parallel execution only the following pairs:
[X1 X2] [LOAD ADD] [SUB COMP] [X3 X4]
This inflexible pairing causes adverse effects on some addicted commands and only partial parallel processing benefits are achieved.
The map of Figure 13 illustrates the various steps involved in determining which of the existing adjacent commands in the byte stream belong to categories or classes eligible for grouping to form a conjugate order for a particular computer system configuration.
There are many possible places in a computer system where coupling can occur, both in the software and in the system structure. Each of them is characterized by individual advantages and disadvantages. As shown in Figure 5, there are several stages in which the program passes from the source code to the current execution. In the complication phase, the source program is translated into machine code and saved to disk storage 46. In the program execution phase, it is read from the disk memory 46 and loaded into the main memory 48 of the specific configuration of the computer system 50, where the instructions are executed by the appropriate instruction processing units 52, 54, 56. Coupling can take place anywhere on this path. In general, if the interface unit is closer to the instruction processing unit or CPU, time considerations prevail. If the interconnection assembly is further away from the CPU, then more instructions can be viewed in a large instruction stream window for grouping before coupling to achieve increased performance. However, such early coupling increases the requirements for building the rest of the system, necessitating additional expansion and increasing costs.
According to the invention, it is possible to execute existing programs written in existing higher-level languages or existing programs in symbolic language by programming means that can identify sequences of adjacent orders suitable for parallel execution by individual functional units.
The flow chart of FIG. 6 shows the generation of a program consisting of a set of compound instructions including a set of customized coupling rules 58 reflecting the system structure. This symbolic language program is fed as input to the programmable coupling device 59 which the program with compound instructions produces. They are analyzed by a software device coupling subsequent blocks of a predetermined length. The length of each block 60, 62, 64 in a stream of bytes containing groups of instructions considered together before coupling depends on the complexity of the coupling device. As shown in Fig. 6, this particular coupling device is intended to control a two-way coupler for a number of "m" fixed length commands in each block. The first main step is to check whether the first and second orders form a coupling pair, and then whether the third and fourth orders form a coupling pair and continue until the end of the block.
It is highly desirable that, after identifying the various possible Cl-C5 coupling pairs, the next step is to determine the optimal product to perform parallel conjugate orders formed by adjacent scalar orders. In the example shown, the following different sequences of compounded commands are possible (assuming no jumps):
II, C2, ICI, ^, C3, Ci ^, n (^; ^ 1, <C2 ^, IC I5, I6, C4, I9, I10; C1,13,14,15, C3, C5, I10; C1 , 13,14,15, 16, C4, 09,110.
Based on the specific system configuration, the coupling device may select the preferred sequence of compound instructions and use markers, i.e. identification bits, to determine the optimal sequence of compound instructions. If there is no optimal sequence, all coupled adjacent scalar orders can be marked, so that any of these paired conjugates can be used by jumping to an instruction selected from different conjugated instructions (see Figure 14). If multiple coupling units are available, multiple consecutive instruction stream blocks can be coupled at the same time.
The specific structure of the software coupling device will not be discussed here because the details are specific to the given instruction set architecture and the implementation being considered. Although the design of such coupling programs is in principle somewhat similar to the design of modern compilers for order scheduling and other optimizations based on a particular machine structure, the criteria for carrying out such coupling are specific to the present invention, which is best seen in the action map of Figure 13. In both cases, the output program is generated based on the given input program and the description of the instruction set and the architecture of the component part, i.e. structural aspects of the implementation. For a modern compiler, the end product is an optimized new sequence of existing commands.
In the case of the present invention, the output is a series of compound instructions, each of which is formed by a group of adjacent scalar commands suitable for parallel execution, with the displacement of instructions combined with non-coupled scalar instructions and forming part of the final control product necessary for the execution of compound instructions. . Processing the initial instruction stream to create compound instructions is easier if known reference points already exist, indicating where the instructions begin. Here, reference points indicate a certain marked field or other indicator that gives information about the location of command boundaries. In many computer systems, these reference points are exactly known only by the compiler during complications and only by the CPU when downloading instructions. Such reference points are not known between compilation and command retrieval, except when the layout is adapted to use special reference marks.
If the coupling takes place after compilation, the compiler could mark with reference marks (see figure 11) which bytes contain the first bytes of instructions and which contain the data. This additional information increases the efficiency of the coupling device due to accurate knowledge of the location of the order. Of course, the compiler could otherwise mean pleasure and distinguish orders from data to provide the coupling device with detailed information indicating the location of the boundaries. If such instruction boundary information is known, then the generation of the corresponding bits of the coupling identifier takes place in the usual way, based on the coupling rules developed for the specific architecture and system configuration (see figure 8). If there is no such information about the order boundaries and the orders have different lengths, the problem is more complex (see Figures 9 and 16).
165 524
The drawing figures presented here are based on the preferred coding scheme, described in more detail in Table 2A below, according to which the command is provided with the "1" tag bit if the order is coupled with the next one, and with the "0" tag bit if two-way coupling it is linked to the next order. The control bits in the control field added by the coupling device contain information relating to the execution of compound instructions and may contain a minimum of information considered effective in a given implementation. An exemplary 8-bit control field is shown in Figure 12. Although in the simplest implementation only the first control bit is needed to indicate the start of a compound instruction, the remaining control bits give additional optional information regarding the execution of the instructions.
In another coding scheme that can be used for both 2-way coupling and large group coupling, the first control bit is set to "1" to indicate that the corresponding instruction is the beginning of the compounded instruction. All other elements of the compound instruction will have their first bits set to "0". At the same time, it will not be possible to associate the given order with others, so that the given order will appear as a combined order of length one. That is, the first control bit will be set to "1", but the first control bit of the next command will also be set to "1". According to this alternative coding scheme, the system will be able to determine how many commands make up the compound instruction - by monitoring all identification bits instead of monitoring the same identification bit of the start of the compound instruction, as is the case with the preferred coding scheme shown below in Table 2A -2C.
The flow chart of FIG. 7 shows a typical implementation for a program consisting of a set of compound instructions, generated by a system preprocessor 66 or program preprocessor 67. The stream of bytes containing compound instructions flows into the buffer of 68 instructions (CI), which is a fast buffer enabling quick access to compound instructions. . Logic circuits 69 retrieve compound instructions from the CI buffer and issue their individual instructions combined to the appropriate functional units for parallel execution. It should be emphasized that compound instruction execution units (CI, EU) 71 such as ALUs in a computer system working with compound instructions are able to execute at once either alone one scalar instruction or conjugated scalar commands in parallel with other conjugated scalar orders. Such parallel execution can also be performed in other types of execution units, such as ALU units, floating point units (FP) 73, address storage and generation units (AU) 75 or multiple units of the same type (FP1, FP2, etc.), respectively to the computer architecture and specific configuration of the computer system. Thus, the system configurations suitable for carrying out the present invention are scalable to a virtually unlimited number of execution units to achieve maximum parallel processing efficiency. By combining several existing commands into a single compounded instruction, it is possible to provide efficient decoding and execution of these instructions coupled in parallel in one or more processing units without the delay that occurs in conventional computer systems with parallel processing.
In the simplest sample coding schemes, a minimum coupling information of one bit for every two bytes of text (instructions and data) is added to the instruction stream. In general, a tag containing control information may be added to each instruction in the byte stream, that is, to each unconjugated scalar instruction containing a pair, a triplet, or a larger joined group. Identification bits are those bits that refer to the part of the tag that is specifically used to identify and distinguish scalar orders forming a conjugate group from other non-conjugated scalar instructions. Such non-linked scalar instructions remain in a program of compound instructions and are executed one after the other. In a system with all 4-byte instructions aligned to four-byte boundaries, one tag is assigned to every four bytes of text.
In the embodiment shown here, all System / 370 instructions are aligned to the half-word boundary (two bytes) with each of them being two or four, or six
165 524 bytes, and one tag with identification bits is needed for each halfword. In the example of creating small groups for coupled pairs and adjacent commands, the identification bit - "1" means that the order starting in the considered byte is coupled with the next instruction, while "0" means that the instruction starting in the considered byte does not is coupled. The identification bit assigned to the half-word that does not contain the first byte of the instruction is skipped. The identification bit of the first byte of the second instruction in the conjugate pair is also skipped, however, in some jump situations, these identification bits are not skipped. It follows that with this identification bit coding procedure, in the simplest case of a two-way coupling, the CPU needs only one information bit to identify the compound instruction when it is executed.
When grouping more than two scalar commands into a compound instruction, additional identification bits may be needed to provide appropriate control information. However, there is another one for reducing the minimum number of information bits needed to control the coupling information path format. For example, even for large groups it is possible to achieve the ratio of one information bit per instruction, with the following coding: the value "1" means coupling with the next instruction, and the value "0" means no coupling with the next instruction. A compound instruction created by a group of four individual instructions could have a sequence of coupling identification bits (1.1, 1.0).
As described in the execution of other compound instructions, the coupling identification bits assigned to non-command half words and therefore having no machine codes assigned are skipped at execution time. In the preferred coding scheme described in detail below, the minimum number of identification bits needed for. providing additional information about the current number of scalar instructions actually coupled is equal to the logarithm at base 2 (rounded up to the nearest integer) of the maximum number of scalar commands that can fall into a group when creating a compound order. For example, if the maximum is two, then one identification bit is needed for each compound instruction. If the maximum is five, six, seven or eight, then three identification bits are needed for each conjugate. This coding scheme is shown below in Tables 2A, 2B and 2C:
Table 2A (maximum is two)
<td>ID beaten</td><td>Code Meanings</td><td>Total number of Coupled</td>
<td> 0</td><td>This order is not linked to the next order</td><td>lack</td>
<td> 1</td><td>This order is coupled with one subsequent order</td><td>two</td>
Table 2B (maximum is four)
<td>ID beaten</td><td>Code Meanings</td><td>Total number of Coupled</td>
<td> 00</td><td>This order is not linked to the next order</td><td>lack</td>
<td> 01</td><td>This order is coupled with one subsequent order</td><td>two</td>
<td> 10</td><td>This order is coupled with the next two orders</td><td>three</td>
<td> 11</td><td>This order is linked to the next three orders</td><td>four</td>
Table 2C (maximum is eight)
<td>ID beaten</td><td>Code Meanings</td><td>Total number of Coupled</td>
<td> 000</td><td>This order is not linked to the next order</td><td>lack</td>
<td> 001</td><td>This order is coupled with one subsequent order</td><td>two</td>
<td> 010</td><td>This order is coupled with the next two orders</td><td>three</td>
<td> 011</td><td>This order is linked to the next three orders</td><td>four</td>
<td> 100</td><td>This order is linked to the next four orders</td><td>bake</td>
<td> 101</td><td>This order is coupled with the next five orders</td><td>six</td>
<td> 110</td><td>This order is coupled with the next six orders</td><td>seven</td>
<td> 111</td><td>This order is coupled with the next seven orders</td><td>eight</td>
165 524
It is obvious that each halfword requires a tag, but with this preferred coding scheme, the CPU skips everything except the first instruction tag in the instruction stream. In other words, the byte is checked to determine if it is a compound instruction by checking its identification bits. If it is not the beginning of a compound instruction, its identification bits are zeros. If the byte is the beginning of a compound instruction containing two scalar instructions, the identification bit for the first instruction is "1" and "0" for the second instruction. If the byte is the beginning of a compound instruction containing three scalar instructions, the identification bits are "2" for the first instruction, "1" for the second instruction and "0" for the third instruction. In other words, the identification bits for each halfword determine whether this particular byte is or is not the beginning of a compounded order, while determining the number of instructions that make up the compounded group.
With these exemplary methods for coding compound instructions, it is assumed that if three instructions are combined to form a triple group, the second and third instructions are also combined to form a pair. In other words, if there is a jump to the second order in a triple group, then the identification bit "1" for the second order indicates that the second and third orders will be executed in parallel as a pair, even if the first order in the three was not executed.
It is obvious that, according to the invention, only the instruction stream only needs to be coupled once for a particular configuration of the computer system, and after that any download of the combined instructions will also result in the download of the identification bits assigned to them. This frees from the necessity of inefficient ongoing determination and selection of scalar orders for parallel execution repeated whenever the same or different commands are taken for execution in a so-called superscalar machine.
Despite all the advantages of joining a binary command stream, it turns out to be difficult to do in some computer structures if no technique for determining boundaries in the byte stream is developed. This determination is difficult if different instruction lengths are used, all the more so if data and instructions are moved and if modifications can be made directly on the instruction stream. Of course, at the time of execution, the boundaries of the orders must be known to enable correct execution, but because it is advisable to do the coupling well before the execution of the order, it was possible to develop a specialized technique for combining the orders without information about which data they constitute. This technique is generally described below and can be used to create compound orders formed from larger groups of scalar orders. This technique can apply to all instruction sets of different conventional types of structures, including RISC structures (computers with a reduced instruction list) in which instructions are usually of a fixed length and are not moved with the data.
The coupling technique enables the coupling of two or more scalar instructions from an instruction stream without the need to know the starting point and the length of each separate instruction. Typical orders already contain machine codes in advance at specific positions specifying the order and its length. Adjacent commands suitable for parallel execution in a specific computer configuration are provided with appropriate markers indicating them as candidates for coupling. In System / 370 architecture, where the commands are in terms of length, either two, four or six bytes, the field positions for the machine code are predetermined based on the code of the expected length of the order. The value of each tag determined based on the pre-determined machine code is used to arrange the full sequence of possible commands. Once the actual command boundary has been determined, the corresponding valid tag values are used to identify the start of the compounded command and the other tags generated incorrectly are skipped.
This coupling technique is illustrated, for example, in the drawings of Figures 8-9 and 14-15, in which the rules stipulate that all 2-byte or 4-byte orders are interconnected, i.e. a double-byte order is suitable for parallel execution in this particular computer configuration along with another double-byte order or with another four-byte order. Exemplary coupling rules also mean that all 6-byte orders are not cluttered at all, which means that a six-byte order in this
165 524 a specific computer configuration is only suitable for its single performance. Of course, the invention is not limited to these exemplary coupling rules, but applies to any set of coupling rules specifying criteria for the parallel execution of existing instructions in the particular configuration of a given computer structure.
The instruction set used in this example was taken from the System / 370 architecture. By analyzing the machine code of each command, you can specify the type and length of each command and then generate for that command a control tag containing the identification bits, as described in detail below. Of course, the invention is not limited to any specific structure or set of instructions, and the coupling rules mentioned above are only examples.
The coding scheme for the compounded commands in these preferred embodiments is shown in Tables 2A-2C above: In the first case, with fixed-length commands, not moved with the data, and with a known cell reference point location for the machine code, the coupling can be carried out in accordance with the transmitting rules for use in this particular computer configuration. Since the field reserved for machine code also contains the length of the order, the sequence of scalar orders is predetermined, and each order in the sequence can be considered as a potential candidate for parallel execution with the next order. The first coded value in the control tag indicates that the instruction is not suitable for coupling to the next instruction, while the second coded value in the tag indicates that the instruction can be coupled for parallel execution with the next instruction.
In the second case, with variable-length orders, not moved with data and with a known position of the cell reference point for the machine code and the cell for the order length code, which in the System / 370 is part of the machine code, the coupling can take place in the usual way. As shown in fig. 8, machine codes indicate the sequence 70 as follows: the first instruction is 6 bytes long, the second and third are 2 bytes long, the fourth is 4 bytes long, the fifth is 2 bytes, sixth bytes 6, and the seventh and eighth are 2 bytes long . Vector C 72 in Figure 8 shows the values of the identification bits (indicated in the drawing as coupling bits) for this particular sequence of 70 instructions, the point indicating the beginning of the first instruction is known. Based on the value of such identification bits, the second and third commands form a conjugate pair, which is indicated by "1" in the identification bit of the second order. The fourth and fifth orders form a different conjugate pair, which is indicated by "1: in the identifying bit of the fourth order. The seventh and eighth orders also form a conjugate pair, which is indicated by "1" in the identification bit of the seventh order. Vector C 72 from Fig. 8 it is relatively easy to generate if the instruction bytes are not mixed with the data bytes, and if the instructions have the same length and known boundaries.
Another situation was presented in the third case, where orders are mixed with non-orders, with a reference point given in real time to indicate the beginning of the order. The diagram of Fig. 11 illustrates one way of indicating an instruction reference point, each half word being marked with a reference marker to indicate whether it contains or does not contain the first byte of the instruction. This can occur with both fixed and variable length commands. When using a reference point, it is necessary to specify the position of the data in the byte stream to enable coupling. Accordingly, the coupling unit may skip and skip all non-instruction bytes.
A more complex situation arises if the byte stream contains variable-length orders (without data), but you know where the first order begins. Since the maximum length order is six bytes, and the orders are aligned to double-byte boundaries, three start points are possible for the first order in the stream. Accordingly, the method allows all possible starting points of the first instruction to be included in the text of byte stream 79, as shown in Fig. 9.
Sequence 1 assumes that the first instruction begins at the first byte, and performs the join on this assumption. In this exemplary embodiment, the length field is also the determining factor for the value of the C vector for each possible instruction. Therefore, the C74 vector for
165 524 of sequence 1 is "1" only for the first instruction of a possible conjugate pair formed by a combination of double-byte and four-byte orders.
Sequence 2 assumes that the first instruction begins at the third byte (beginning of the second halfword) and performs the coupling on this assumption. The length field value for the third byte is 2, indicating that the next instruction starts on the fifth byte. Passing through each possible order based on the length field value in the preceding order, all potential sequence 2 commands are generated according to possible identification bits, as shown in vector C 76.
Sequence 3 assumes that the first instruction begins at the fifth byte (beginning of the third halfword) and performs the coupling on this assumption. The length field value for the fifth byte is 4, indicating that the next instruction starts on the ninth byte. Passing through each possible order based on the length field value in the preceding order, all potential sequence 3 commands are generated according to possible identification bits as shown in vector C78.
In some cases, these three different Sequences of potential commands will coincide with one sequence. In Fig. 9 it is noted that these three Sequences converge at the instruction boundaries at the end of the 80th byte. Sequences 2 and 3, at the convergence at the instruction boundary at the end of the 82nd byte, are not in phase relative to the coupling until the end of the sixteenth byte. In other words, these two sequences take into account different pairs of instructions based on the same order sequence. Because the sixteenth byte starts the unconvertible order in 84, the nonphase convergence ends.
If no valid convergence is achieved, then all three possible command sequences must be continued until the end of the window. If, however, important convergence occurs and is detected, then the number of sequences is reduced from three to two (one of the identical sequences becomes inoperative), and in some cases from two to one.
Thus, before convergence, the temporary boundaries and identification bits assigned to each of these sequences are determined for each possible sequence of indications indicating the location of potential compound instructions. From Fig. 9 it can be seen that three separate identification bits are generated in this way for each two bytes of text. In order to ensure compatibility with the pre-processing carried out in the first, second and third cases mentioned, it is desirable to reduce these three possible sequences to a single sequence of identification bits, in which each halfword is assigned only one bit. Because only information is needed about whether the current instruction is coupled to the next, the three bits can be logically added to produce a single sequence in the CC 86 vector.
For the purposes of parallel execution, the juxtaposed identification bits of the juxtaposed CC 86 vector are balanced by a separate C vector of the three sequences 1-3. In other words, the juxtaposed identification bits in the CC vector enable the correct execution of any of the three possible sequences in parallel for compound instructions or single for non-compound instructions. The matching identification bits also work correctly for jumps. For example, if there is a jump to the beginning of the ninth byte 88, then the ninth byte must be the start of the command. Otherwise there is an error in the program. The "1" identification bit assigned to byte 9 is used and the corresponding instruction is executed in parallel with the next instruction.
The various steps of the coupling method illustrated in Figure 9 and above are illustrated in the operation map of Figure 16.
The best moment to provide information about reference points for command boundaries is when you compile. To identify the beginning of each command during compilation, identification tags 101 may be inserted, as shown in Fig. 11. This allows the coupling device to use a simplified procedure for the above-mentioned cases, First, Second and Third. Of course, in order to simplify the operation of the coupling unit and to avoid compilation of the method shown in Fig. 9, the compiler can identify order boundaries and distinguish orders from data by any other method.
165 524
Figure 10 is a flowchart of a possible implementation of a coupler for processing an instruction stream, as in Figure 9. Multiple coupling units 104, 106, 108 are shown and in order to improve performance their number can be as large as the number of halfwords that can be stored in the text buffer. In this version, these three coupling units would start their processing sequences in the first, third and fifth bytes respectively. At the end of the possible sequence of instructions, each coupling unit begins searching for the next possible sequence by offset by six bytes from the previous sequence. Each coupling unit produces coupling identification bits (C vector values) for each half-word in the text. The three sequences originating from the three coupling units are logically summed 110 and the resulting identification bits are memorized in connection with the corresponding text bits.
One of the beneficial properties of complex identification bits in the CC vector is the ability to create many valid coupling bit sequences on the basis of which the instruction is addressed by the jump output. As best seen in Figs. 14-15, differently shaped compound instructions can be obtained from the same stream of bytes.
Figure 14 shows possible combinations of compound instructions, when the computer configuration allows for the issuing and execution of no more than two commands in parallel. If the instruction stream 90 containing the compound instructions is processed in the normal sequence, then a coupled instruction I will be issued for the parallel execution of the identification bit for the first byte in the CC 92 vector. However, if a jump to the fifth byte occurs, then Coupled Order II will be issued for the parallel execution of the identification bit for the fifth byte in the CC 92 vector. Similarly, normal sequential processing of another coupled byte stream 94 will result in the sequential execution of compound instructions IV, V, and VIII (the component instructions in each compound instruction are executed in parallel). In contrast, the jump to the third byte in the coupled instruction stream based on the identification bits in the CC 96 vector will result in the sequential execution of the Combined Orders V and VII, and the instruction beginning in the fifteenth byte (forming the second part of Combined Order VII) will be issued and executed individually . A jump to the seventh byte will result in the sequential execution of Combined Orders VII and VIII, and a jump to the eleventh byte will cause the Combined Order VIII to be executed. In contrast, a jump to the ninth byte in the coupled instruction stream will execute Combined Order VII (it is formed by the second part of Combined Order VI and the first Combined Order VIII). Thus, bits "1" in the CC 96 vector for
Combined Orders IV, VI and VIII are skipped when one of the Combined Orders V or VII is executed. In contrast, the identifier "1" bits in the CC 96 vector for Combined Instructions V and VII are skipped when one of the Combined Orders IV, VI or VIII is executed.
Figure 15 shows possible combinations of compound instructions when the computer configuration allows for the simultaneous issuing and execution of up to three instructions. If the instruction stream 98 containing the compound instructions is processed in the normal sequence, then the Combined Orders X (triple group) and XIII will be executed. In contrast, a jump to the eleventh byte will result in Compound Order XI (triple group), and a jump to the thirteenth byte will result in Compound Order XII (another triple group). Thus, the identifier "2" bits in CC 99 vector for Combined Instructions XI and XII are skipped when Combined Orders X and XIII are executed. On the other hand, when the Combined Order XI is executed, the identification bits for the other three Combined Orders X, XII, XIII are skipped. Similarly, when the Combined Order XII is executed, the identification bits for the remaining three Combined Orders X, XI, XIII are skipped. Several solutions are available for command combination depending on its location and information on the content of the text. In the simplest situation, it would be desirable for the compiler to mark with the help of reference marks which bytes contain the first byte of the instruction and which contain the data. This additional information increases the performance of the coupling device because the exact locations of the orders are known. This means that the coupling could always be performed as in the case of the First, Second and Third cases to generate the vector C identification bits for each compound instruction. The compiler could also include another informa16
165 524, for example, with predicted static jumps, or even insert directives for a coupling unit. .
Other methods could be used to distinguish data from instructions when storing the instruction stream in memory. For example, if data fragments are infrequent, then less than reference marks would normally take for a list of data addresses. These coupler system and software combinations create many options for efficiently creating compound instructions.
1,2,3,4 5,6
<td></td><td>CYCLE</td><td>CYCLE</td><td>CYCLE</td><td>CYCLE</td><td>CYCLE</td><td>CYCLE</td>
<td>Functional unit A</td><td>s</td><td>CS</td><td>CS</td><td> —</td><td>CS</td><td>s</td>
<td>Functional unit B</td><td> —</td><td> —</td><td>CS</td><td> —</td><td>CS</td><td> —</td>
<td>Functional unit C</td><td> —</td><td>CS</td><td> —</td><td>s</td><td>CS</td><td> —</td>
<td>Functional unit D</td><td> —</td><td>CS</td><td> —</td><td></td><td>CS</td><td> —</td>
Combined instruction stream
<td>CS CS CS CS T4</td><td>S</td><td>CS CS T2</td><td>CS CS CS T3</td><td>s</td>
'38 '36
Combined instruction = C Scalar instruction = S Combined scalar command = CS
Coupled order marker = T
One-processor implementation!
FIG. 2
CYCLE 1
CYCLE 2 Command cycles
<img file="PL165524B1_D0001.tif" />
3processor implementation
FIG.3
165 524
<td>Number category</td><td>Category name</td>
<td> 1.</td><td>RR-Format for loading logical, arithmetic, comparisons. • LCR - complete load order • LPR-charge (if) positive • LNR-charge (if) negative • LR - load registry • LTR - load and test • NO-I • OR-OR • XR-exclusive OR • AR-sum • SR-subtract • ALR-add logically • SLR-subtract logically • CLR-compare logically • CR-compare</td>
<td> 2.</td><td>RS-format of commands: offsets (no memory access) • SLR-move logically to the right • SLL-move logically to the left • SRA-move arithmetically to the right • SLA-move arithmetically to the left • SRDL-move logically to the right • SLDL-move logically to the left • SRDA-move arithmetically to the right • SLDA - move arithmetically to the left</td>
<td> 3.</td><td>JUMPING-when reaching the number and indicator • BCT-jump when reaching the number (RX format) • BCTR-jump when the number is reached (RR format) • ΒΧΗ-jump when the high indicator (form F • BXLE-jump when the low indicator (form RS is reached)</td>
<td> 4.</td><td>CONDITIONAL JUMPS<sup>e</sup>BC conditional jump (RX format) • BCR-conditional jump (RR format)</td>
FIG.
4A
FIG.4
FIG.
4B
Figure 4
165 524
<td>Number category</td><td>Category name</td>
<td> 5.</td><td>JUMPS-1 CONNECTION • BAL-jump and Rainbow (RX format) • BALR-jump and connect (RR format) • BAS-jump and save (RX format) • BASR-jump and save (RR format)</td>
<td> 6.</td><td>memorizing • STCM-remember characters under mask • (remember bytes 0-4, RS format) • MVI - transfer immediately (one byte, SI format) • ST-remember (4 bytes) • STC-remember character (one byte) • STM - remember half (2 bytes)</td>
<td> 7.</td><td>LOADING • LM-Tadowe Half (2 bytes) • L-Taduj (4 bytes)</td>
<td>and.</td><td>LA-Load address</td>
<td> 9.</td><td>RX / SI / RS command formats: arithmetic, logic, input, comparison • A-add • AH-add half<sup>0</sup>AL-add logically • N - I • 0 - OR • S-subtract • SM-subtract one half • SL-subtract logically • X-EXCLUSIVE OR • IC-enter the character • ICM - enter the character under the mask (0-4 byte FGTCM command) • C- compare • CH-compared Half • CL-logically compare • CLI - compare logically immediately • CLM - logically compare the character under the mask</td>
<td> 10.</td><td>TM-Test under the hood</td>
F1G.4B
<img file="PL165524B1_D0002.tif" />
FIG.5
165 524
Program in assembler language
<img file="PL165524B1_D0003.tif" />
<img file="PL165524B1_D0004.tif" />
FIG.7
165 524
Byte ^ -Sequence ^ -Vector C
<img file="PL165524B1_D0005.tif" />
Byte = text numbering
Sequence = sequence with program instructions length Vector C = coupling bits for every two bytes
FIG. 8
<img file="PL165524B1_D0006.tif" />
Figure 11
165 524
<img file="PL165524B1_D0007.tif" />
Length field | ę | 2 | 4 | z | 2 | 4 | 4 | 2 | 6 | 4 | 2 | 2 [2 |
Byte position 0 2 4 6 8 1012 14 16 18 20 22 24
<img file="PL165524B1_D0008.tif" />
Length length = possible order length code for every two bytes Byte = text numbering
Sequence = potential sequence of program instructions Vector C — possible join bits for every two bytes Vector CC = compound join bits for every two bytes
FIG.9
165 524
2 4 6 ....... Ν-1 Ν
Battle position IIII IΊ Ί Ί IIIII l
<img file="PL165524B1_D0009.tif" />
Figure 10 <sup>1</sup>0 <sup>Ι</sup>1 <sup>1</sup>2 <sup>1</sup>3 <sup>1</sup>5 <sup>1</sup>6 <sup>1</sup>7
<td>ΒΙΤ</td><td>Function</td>
<td><sup>ι</sup>ο</td><td>If 1, then this command means the beginning of the coupling order</td>
<td></td><td>If l, execute two orders in parallel</td>
<td> *2</td><td>If 1, this compound instruction has more than one execution cycle '</td>
<td> *3</td><td>If 1, suspend piping</td>
<td> *4</td><td>If the instruction is a jump, then if this bit is one, it is envisaged to make that jump</td>
<td> *5</td><td>If 1, this order is memory dependent on the previous compounded order</td>
<td><sup>1</sup>6</td><td>If 1, enable dynamic issuing of orders</td>
<td><sup>1</sup>7</td><td>If 1, this command uses an arithmetic logical unit</td>
F1G.12
165 524
<img file="PL165524B1_D0010.tif" />
figh
Position byte u
Length field
2 4 6 8 10 12 14
Combined order — 90!<sup>4</sup> and<sup>2</sup>l<sup>2</sup>J <sup>6</sup> AND
Combined order lii!
<sup>1</sup> J
and Bit ID for
92'
CC vector
<td> 1</td><td> 0</td><td> 1</td><td>o | o | o</td>
^ -Bit identifier for II
O 2 4 6 8 10 12 14 16 18 20 22 Byte position
<img file="PL165524B1_D0011.tif" />
2 4 6 8 10 12 14 16 18 20 22 24 26 Byte position TUT
31-98
<td colspan="4"> 2 6</td><td> 2</td><td colspan="8"> 2|2| 4 |2| 6</td>
<td></td><td colspan="3"></td><td colspan="3"> '—</td><td colspan="3"></td><td colspan="3"></td>
<td> 0</td><td> 0</td><td> 0</td><td> 0</td><td> 2</td><td> 2</td><td> 2</td><td> 1</td><td> 0</td><td> 0</td><td> 0</td><td> 0</td><td> 0</td>
<td> -</td><td> -</td><td> -</td><td> -</td><td></td><td> 31</td><td>2Γ</td><td>2H</td><td></td><td></td><td></td><td></td><td></td>
Length field
Combined order
Figure 15
<img file="PL165524B1_D0012.tif" />
165 524
<img file="PL165524B1_D0013.tif" />
<img file="PL165524B1_D0014.tif" />
Key:
A-always coupled F1G.17
S- sometimes coupled N-not coupled
FIG.1
<img file="PL165524B1_D0015.tif" />
UP Department of Publications. Circulation of 90 copies
Price: PLN 10,000
Contents14
97 members in 14 offices
Priority claims3
| Document | Office | Kind | Date |
|---|---|---|---|
| 51938490 | United States of America | A | |
| 90519384 | – | – | – |
| US19900519384 | – | – | – |
Members97
| Document | Office | Kind | |
|---|---|---|---|
| HU911099D0 | Hungary | D0 | |
| HU911102D0 | Hungary | D0 | |
| HU911103D0 | Hungary | D0 | |
| CA2037708A1 | Canada | A1 | |
| CA2039640A1 | Canada | A1 | |
| CA2040637A1 | Canada | A1 | |
| EP0454984A2 | European Patent Office (EPO) | A2 | |
| EP0454985A2 | European Patent Office (EPO) | A2 | |
| EP0455966A2 | European Patent Office (EPO) | A2 | |
| WO9117495A1 | World Intellectual Property Organization (WIPO) | A1 | |
| WO9117496A1 | World Intellectual Property Organization (WIPO) | A1 | |
| HUT57454A | Hungary | A | |
| HUT57455A | Hungary | A | |
| HUT57456A | Hungary | A | |
| BR9101913A | Brazil | A | |
| CS93591A2 | Czechoslovakia (until 1993) | A2 | |
| CS93691A2 | Czechoslovakia (until 1993) | A2 | |
| EP0463299A2 | European Patent Office (EPO) | A2 | |
| HU9200024D0 | Hungary | D0 | |
| PL289722A1 | Poland | A1 | |
| EP0481031A1 | European Patent Office (EPO) | A1 | |
| BR9101791A | Brazil | A | |
| PL289723A1 | Poland | A1 | |
| CS93391A3 | Czechoslovakia (until 1993) | A3 | |
| HUT60048A | Hungary | A | |
| EP0496928A2 | European Patent Office (EPO) | A2 | |
| JPH04229326A | Japan | A | |
| JPH04230528A | Japan | A | |
| JPH04233034A | Japan | A | |
| CA2053941A1 | Canada | A1 | |
| JPH04505823A | Japan | A | |
| PL293182A1 | Poland | A1 | |
| JPH04506878A | Japan | A | |
| EP0463299A3 | European Patent Office (EPO) | A3 | |
| EP0481031A4 | European Patent Office (EPO) | A4 | |
| EP0496928A3 | European Patent Office (EPO) | A3 | |
| US5197135A | United States of America | A | |
| EP0545927A4 | European Patent Office (EPO) | A4 | |
| US5214763A | United States of America | A | |
| EP0545927A1 | European Patent Office (EPO) | A1 | |
| EP0455966A3 | European Patent Office (EPO) | A3 | |
| US5295249A | United States of America | A | |
| JPH0683623A | Japan | A | |
| EP0454985A3 | European Patent Office (EPO) | A3 | |
| US5303356A | United States of America | A | |
| EP0454984A3 | European Patent Office (EPO) | A3 | |
| JPH0679273B2 | Japan | B2 | |
| JPH0680489B2 | Japan | B2 | |
| PL165491B1 | Poland | B1 | |
| PL165524B1This record | Poland | B1 | |
| JPH0773036A | Japan | A | |
| CA2040304C | Canada | C | |
| PL166513B1 | Poland | B1 | |
| CZ279899B6 | Czechia | B6 | |
| JPH0776924B2 | Japan | B2 | |
| JPH0778737B2 | Japan | B2 | |
| US5446850A | United States of America | A | |
| US5448746A | United States of America | A | |
| JPH0782438B2 | Japan | B2 | |
| US5465377A | United States of America | A | |
| US5475853A | United States of America | A | |
| EP0455966B1 | European Patent Office (EPO) | B1 | |
| AT131637T | Austria | T | |
| DE69115344D1 | Germany | D1 | |
| JPH087681B2 | Japan | B2 | |
| US5500942A | United States of America | A | |
| US5502826A | United States of America | A | |
| US5504932A | United States of America | A | |
| JP2500082B2 | Japan | B2 | |
| DE69115344T2 | Germany | T2 | |
| EP0454984B1 | European Patent Office (EPO) | B1 | |
| DE69122294D1 | Germany | D1 | |
| EP0454985B1 | European Patent Office (EPO) | B1 | |
| AT146611T | Austria | T | |
| DE69123629D1 | Germany | D1 | |
| DE69122294T2 | Germany | T2 | |
| DE69123629T2 | Germany | T2 | |
| US5701430A | United States of America | A | |
| CA2037708C | Canada | C | |
| CA2040637C | Canada | C | |
| EP0825529A2 | European Patent Office (EPO) | A2 | |
| US5732234A | United States of America | A | |
| HU214423B | Hungary | B | |
| EP0825529A3 | European Patent Office (EPO) | A3 | |
| RU2111531C1 | Russian Federation | C1 | |
| HU216990B | Hungary | B | |
| CA2039640C | Canada | C | |
| EP0496928B1 | European Patent Office (EPO) | B1 | |
| AT189540T | Austria | T | |
| US6029240A | United States of America | A | |
| DE69131956D1 | Germany | D1 | |
| ES2142304T3 | Spain | T3 | |
| EP0545927B1 | European Patent Office (EPO) | B1 | |
| AT194236T | Austria | T | |
| DE69131956T2 | Germany | T2 | |
| DE69132271D1 | Germany | D1 | |
| DE69132271T2 | Germany | T2 |
Numbers
- Publication, DOCDB
- 165524
- Publication, EPODOC
- PL165524B
- Application
- 91289722
- Application, DOCDB
- 28972291
- Application, EPODOC
- PL19910289722
Titles
- English
- METHOD OF PARALLEL PROCESSING OF COMMANDS AND ARRANGEMENT OF A MACHINE OPERATING WITH A SCALED SET OF COMMANDS
Classification
- CPC, 6
- G06F9/382
- G06F9/3017
- G06F9/3808
- G06F9/3842
- G06F9/3853
- G06F9/3885
- IPC, 2
- G06F9 318
- G06F9 38