Method of associativelly parallel data processing and data processing system therefor
Abstract
Multiprocessor parallel computing systems and a byte serial SIMD processor parallel architecture is used for parallel array processing with a simplified architecture adaptable to chip implementation in an air cooled environment. The array provided is an N dimensional array of byte wide processing units each coupled with an adequate segment of byte wide memory and control logic. A partitionable section of the array containing several processing units are contained on a silicon chip arranged with "Picket"s, an element of the processing array preferably consisting of combined processing element with a local memory for processing bit parallel bytes of information in a clock cycle. A Picket Processor system (or Subsystem) comprises an array of pickets, a communication network, an I/O system, and a SIMD controller consisting of a microprocessor, a canned routine processor, and a microcontroller that runs the array. The Picket Architecture for SIMD includes set associative processing, parallel numerically intensive processing, with physical array processing similar to image processing. a military picket line analogy fits quite well. Pickets, having a bit parallel processing element, with local memory coupled to the processing element for the parallel processing of information in an associative way where each picket is adapted to perform one element of the associative process. We have provided a way for horizontal association with each picket. The memory of the picket units is arranged in an array. The array of pickets thus arranged comprises a set associative memory. The set associative parallel processing system on a single chip permits a smaller set of `data' out of a larger set to be brought out of memory where an associative operation can be performed on it. This associative operation, typically an exact compare, is performed on the whole set of data in parallel, utilizing the Picket's memory and execution unit.

Term
Term ended
Expired 13 November 2006, 19.9 years ago.
- Priority
- Filed
- Granted
- Expired
- Today
7 claims: 1 independent, 6 dependent
- 1Zastrzeżenia patentowe 1. Układ przetwarzania równoległego zawierający wiele bloków przetwarzania, znamienny tym, że każdy blok przetwarzania (100) zawiera pamięć lokalną (102), której wyjście jest dołączone do wejścia rejestru przesuwającego (104), zespół przetwarzania bitowo-równoległego (101), do którego jednego wejścia (B) jest dołączone wyjście rejestru przesuwającego (104), oraz połączone szeregowo rejestry maskujące (105, 106), przy czym do wejścia pierwszego rejestru maskującego (105) jest dołączone wyjście zespołu przetwarzania bitowo-równoległego (101) zaś wyjście drugiego rejestru maskującego (106) jest dołączone do drugiego wejścia (A) zespołu przetwarzania bitowo-równoległego (101), wejścia rejestru przesuwającego (104) i wejścia pamięci lokalnej (102).
- 2Układ według zastrz. 1, znamienny tym, że każdy blok przetwarzania (100) zawiera następnie magistralę radiofonicznej transmisji danych - adresów (103), która jest dołączona do wejścia pierwszego rejestru maskującego (105) i do wejścia pamięci lokalnej (102).
- 3Układ według zastrz. 1 albo 2, znamienny tym, że każdy blok przetwarzania (100) zawiera następnie magistralę przenoszenia lewy-prawy (108), która jest dołączona do wejścia rejestru przesuwającego (104).
- 4Układ według zastrz. 1, znamienny tym, że każdy blok przetwarzania (100) zawiera następnie rejestr sterowania - stanu (107), który jest dołączony do zespołu przetwarzania bitowo-równoległego (101).
- 5Układ według zastrz. 1, znamienny tym, że zespół przetwarzania bitowo-równoległego (101) jest zespołem o co najmniej 8-bitowej długości słowa.
- 6Układ według zastrz. 1, znamienny tym, że pamięć lokalna (102) jest pamięcią o pojemności co najmniej 32 x 8 kilobitów.
- 7Układ według zastrz. 1, znamienny tym, że zawiera co najmniej 16 bloków przetwarzania (100).
Independent claims7
188 paragraphs in 4 sections, as filed
The subject of the invention is a parallel processing system, especially designed for use in multiprocess computing systems.
U.S. Patent No. 3,537,074 describes a solution of a table computer with parallel processors and one programmable control block, multiple registers for memorizing complementary vectors, mask registers and elements that respond to a sequence of instructions from one or more control blocks intended for for simultaneous operation on data in vector registers. In the seventies, the parallel processors described by their creator RA Stokes in this patent became known as SIMD machines (Single Instructions - Multiple Data = Multiple Data). Such machines can also be described as comprising a programmable control block for running an array of n in parallel processors, each of which has a memory portion, arithmetic block, program decoding portion, and input / output portion. These systems were usually large-scale systems coupled to the main host computer. An important difference between SIMD processors and simpler processors was that inside such systems, all SIMD processors can have a different set of data in the associated processor, but all processors are controlled by a common control block. SIMD computers also differ from the known simpler Von Neumann processors in that each command uses data vectors rather than individual operands.
167 329
The simplest types of multiprocessor systems are systems such as Multiple Instruction - Multiple Data (MIMD), in which each processor can execute a separate program running on a separate data set. Processors in the MIMD system can perform separate tasks or each of them can perform a different subtask of a common main task.
When thinking about the progress made in the field of parallel SIMD processors, it was judged as described in US Patent No. 4,435,758 entitled A method of performing a conditional jump in SIMD vector processors that they can be used if the tasks to be performed are highly independent and uncompetitive, but if the tasks have competition from the point of view of resource access, then it can be determined that the network of processors synchronous works in SIMD mode. In fact, U.S. Patent No. 4,435,758 discusses the problem and describes the improvements that can be applied to the solution disclosed in U.S. Patent No. 4,101,960 to make a conditional jump.
It has become the norm to describe the most improved SIMD machines as bit-by-bit synchronous processors incorporated in an N x N matrix. Methods for multiplying matrix vectors for such massive parallel architectures, processors attached in an N x N matrix system, physically connected through a mesh topological system and a mesh network overlaid with another switching network intended for reorganization purposes were described in detail in the IBM Technical Discovery Bulletin Vol. 32 No. 3A, August 1989, and used to increase the speed of multiplication of the diluted matrix into a vector.
There are many publications that show that although the task was to design SIMD and SIMD / MIMD machines that would work in a multi-series processor system in which all processors in a given series perform exactly the same orders, each series is programmed differently. This is illustrated, for example, in the IBM Technical Discovery Bulletin, Vol. 32, 8B, January 1990, which describes an architecture with such a configuration that has been referred to as the Parallel Local Operator Engine, abbreviated PLOE, for the processing of specific functions of repeatable memory control.
After a retrospective review of prior art publications for the present invention, it has been found that some of them describe the use of a non-volatile memory processor for static instructions and registers for storing and retrieving data implemented as a single module on one semiconductor wafer. An example is US Patent No. 4,942,516 regarding computer architecture implemented on one integrated circuit board. The type of work presented in it does not apply to comprehensive applications such as SIMD.
Other publications describe different resources for different tasks. For example, there is described a system known to be able to perform parallel matrix multiplication. These are publications that relate to the field referred to as ARTIFICIAL INTELLIGENCE. Addressable content or the content of associative (associative) memories should be addressed when using different processing modules. For applications such as artificial intelligence, in some cases it may be beneficial to discuss their uses based on row selection of the results of previous search operations, i.e. on line logic. This is described in the publications: VLSI for Artificial Intelligence (Integrated circuits with a very large scale of integration for applications such as artificial intelligence), Jose Delgado-Frias and Will R. Moore, Kluwer Academic Publishers, 1989, pages 85-108, VLSI and Rule-Based Systems (Systems on Integrated Circuits with Very Large Integration Scales and Declarative Systems), Peter Kogge, Mark Brule and Charles Stormon. However, other authors give other solutions. Another form of solution represented by intelligent memory modules for massively parallel processing was described in the publication VLSI Systems Design, December 1988, pages 18-28 of the article published by Bob Cushman under the title Matrix Crunching with Masive Parallelism. Others tried to implement parallel
167 329 processing with the use of associative memory on integrated circuits with a very large integration scale to describe the modules of associative memory implemented in integrated circuit technology with very large integration scale, suitable for the use of a completely convertible parallel, associative processing circuit. This approach assumed that the use of classic associative memory architecture would require too many pins for data transfer. See, for example, Parallel Processing on VLSI Associative Memory, S. Hengen and I. Scherson, report awarded with NSF ECS-h404627 and delivered by authors in the Department of Electrical and Computer Engineering University of California Santa Barbara, CA 93106.
The problem that has been highlighted is that it is necessary to develop compact processors for complex (complex) applications. It was found that known studies, which were limited to bit-serial applications with memories up to several thousand bits per elementary processor and with several elementary processors per one chip were not appropriate, because they were characterized by extremely high, even dramatic, packing density, but still suitable for placing in an air-cooled environment and made in the form of lightweight units with small dimensions.
As a result of the review of publications on known solutions related to the subject of the application, it was found that the closest solution, constituting the known state of the art for the subject of the invention, was described in the European Patent Application No. EP-A-208 457, concerning a matrix of processors in which each processor elemental matrix has the ability to choose the element from which it receives its input signal. This application describes a multidimensional matrix of elementary processors that has an increased degree of flexibility to ensure the potential for parallel processing and which could be better utilized without the need for increased manufacturing costs and the complexity of the MIMD processor. The general instruction for local bit-serial execution is sent via a link connecting a logical control block with various parallel elementary processors, while it programmically modifies selected bits of the general instruction for use in local bit lines, where the modified bits are decoded.
A review of publications on this topic also provided information on SIMD elementary processors, contained, for example, in the work Design ofSIMDMicroprocessor Array, CR Josshope, R.OOorman et al., Published in IEEE Proceedings Vol. 136, May 1989. The article provides scientific considerations about the SIMD architecture. A description of the processor with SIMD byte serial architecture is also provided. The article presents solutions for elementary processors with an 8-bit architecture of elementary processor batteries, limited to a printed circuit board, with 1 kbyte random access, and with several elementary processors per board (n), and solutions for local autonomy. However, regardless of the use of such a proposed architecture, associative (associative) processing is not possible. It can be stated that the proposed structure does not provide byte-serial communication with the environment.
To facilitate use of the rest of the description, the following general new terminology is used.
PIKIETA (English PICKET) - the parent of processors, preferably consisting of a combination of elementary processors with local memory for bit-parallel processing of information bytes in a clock cycle.
PICKET MODULE - contains several pickets on one silicon wafer.
PICTURE PROCESSOR SYSTEM (or subsystem) - a general system, consisting of a picket matrix, communication network, input / output system and SIMD control system, consisting of a microprocessor, a standard programmable processor and a microcontroller system that launches the matrix.
PICKET ARCHITECTURE - favorable execution of SIMD architecture with features that combine various conflicting issues, namely:
167 329
- established associative (association) processing,
- parallel digital intensive processing,
- physical matrix processing similar to image processing.
PICTURE MATRIX - a set of geometrically ordered pickets.
The essence of the parallel processing system according to the invention, having a multi-block processing architecture, is that each processing block contains local memory whose output is attached to the shift register input, a bit-parallel processing unit to which one input is connected to the register output shifting, and serially connected masking registers, wherein the input of the first masking register is connected to the bit-parallel processing unit output and the output of the second masking register is connected to the second input of the bit-parallel processing unit, the shift register input and the local memory input.
Preferably, according to the invention, each processing block then includes a radio data transmission bus - addresses, which is connected to the first masking register input and to the local memory input, a left-right transfer bus, which is connected to the shift register input, and a state control register that is attached to the bit-parallel processing assembly.
Further advantages of the invention are obtained when the bit-parallel processing assembly is an assembly with at least 8-bit word length, the local memory is a memory with a capacity of at least 32 x 8 kilobits, and contains at least 16 processing blocks.
The advantage of the solution according to the invention is its ability to perform calculations similar to fast-acting data-processing machines in one-command-many-data (SIMD) mode, but its operation was further improved by the use of more parallel elementary processors. Data dependency issues have been eliminated. When a machine is running one command-many data, no processor or function can have any data dependency among those data that could cause one unit processor to require different numbers of cycles.
Improvement of architecture that was introduced to the system with a large number of picket units generally called pickets, containing a bit-parallel processing unit, with local memory connected to this unit for parallel processing of information in an associative manner, where each station is adapted for transformation one element in the association process. The memory of the picket unit is matrixed. The matrix of such ordered stations contains a set of associative (associative) memories.
The solution according to the invention implemented on one semiconductor wafer provides the possibility that a smaller set of data from a larger set is to be transferred from memory, where associative operations can be performed on this data. These associative actions, which are usually just a comparison, are performed in parallel across the entire data set, using the picket memory and the executing unit.
In the picket matrix, each picket has some data from a larger set. In addition, each picket selects one part of the data from that part. In this way, one part of the data in each set of pickets contains a set of data on which an association is performed by all the pickets in parallel.
The picket technology is extensible, and in the case of 126 kilobytes dynamic random access memory in each individual picket (dynamic random access memory module with a capacity of 16 magabytes) the picket architecture will perform all 24-bit color graphics in the same way as text and 8 -bit color graphics or black graphics that are made using the solution according to our current, favorable example. Experimental manufacturing technology shows that this packing density is predictable and that it can be implemented in the near future as the next generation of manufactured devices capable of operating in an air-cooled environment. For color graphics, our beneficial picket architecture would mean increasing the dynamic capacity of the random access memory on a semiconductor wafer to 128 kilobytes per one picket at 16 pickets per wafer, or 24
167 329 picket units per picket plate with 96 kilobyte memory for each picket could be used for full-color GPUs.
The subject of the invention is shown and explained in more detail in the embodiment based on the attached drawing, in which Fig. 1 is a diagram of a processor representing the known state of the art, Fig. 2 - a pair of basic picket units that are made on a silicon wafer, according to the invention, Fig. 3 illustrates the work of the associative (associative) memory, Fig. 4 shows the basic layout with 16 (N) pickets for the one-command-many-data (SIMD) subsystem, Fig. 5 - the construction of the multi-station processing system, which consists of many picket processors from Fig. 4, Fig. 6 - functional block diagram of the subsystem, and Fig. 7 - the connection circuit of the subsystem control systems with the printed circuit boards from Fig. 5.
Figure 1 shows a typical SIMD system known in the art. In such known systems, the SIMD computer is a single-order computer with a data array of parallel processors comprising a plurality of parallel bit-serial processors, each of which has one SIMD memory associated with it. The input / output (I / O) system acts on the SIMD unit as a system that moves data to fast levels and contains an intermediate memory for bi-directional two-dimensional data transfer between the host computer, which can be a central processor or microprocessor, and the SIMD computer. The input / output (I / O) system consists of input and output processing elements designed to control the flow of data between the host computer and the intermediate memory and to control the flow of data between the intermediate memory and multiple SIMD memories, which are usually organized buffer sections or larger memory areas. In this way, the data entry operations carried out by the input / output (I / O) system include data transmission from the host computer's memory to the intermediate, working memory, and from the indirect memory to the SIMD memory in the second step, with the data being entered also in two steps and involves data transmission via a two-dimensional link between the main (home) computer and the SIMD computer. The input / output (I / O) system can be a separate unit, a sub-unit in the central computer or, often, a unit located inside the SIMD computer, where the SIMD control system acts as a control system for the intermediate input / output buffer.
The SIMD computer is a processor matrix with many elementary processors and a network that connects individual elementary processors with each other and with many conventional individual SIMD memories. The SIMD computer is a parallel matrix processor consisting of a large number of individual elementary processors connected and working in parallel. The SIMD computer contains a control block that generates a sequence of commands for elementary processors and provides timing signals to the computer. The network that connects various elementary processors together creates some form of interconnection of individual elementary processors, which interconnections can be superimposed on possible topologies such as mesh, toroidal polymorphic and others. The memory set is intended for immediate storage of data bits intended for individual elementary processors and here one-on-one responsibility is maintained between the number of elementary processors and the number of memories which may constitute the above-mentioned buffer areas of larger memory.
For example, as shown in Fig. 1, the system includes a parent processor 28. This processor is used to load microcoded programs into the matrix control block 14, which contains intermediate buffer memory, and to exchange data with it, used to control its state over the link data 30, central processor, control block as well as address and control link 31. The parent processor 28 in this example may be any suitable general purpose computer, such as a personal computer or a large computer. In this prior art example, the processor matrix is represented as two-dimensional, however, the matrix can be organized differently, for example, as a three-dimensional or four-dimensional cluster. The matrix processor SIMD consists of
167 329 from the matrix 12 of elementary processors P (i, j) and the matrix control block 14, intended to generate a sequence of general instructions directed to the elementary processors P (i, j). Although this is not shown in Fig. 1, examples from the prior art use elementary processors that operate in one bit mode at a time, and memory areas assigned to elementary processors. Elementary processors are connected via the so-called NEWS network (English abbreviation of the words: North-East-West-South = north - east - west - south) with their respective neighbors using bidirectional bit hicc. The P (i, j) elementary processor is connected to the P (il, j), P (i, j + 1), P (i, jl), P (i + 1, j) elementary processors located in the north direction, eastern, western and southern relative to the considered elementary processor P (i, j). In this typical example, the NEWS network is connected on the perimeter so that the north and south edges are connected in two directions, and the east and west edges are connected in a similar way. In order for data to be entered and output from the processor matrix, the data link 26 connecting the control block to the matrix is connected to the NEWS network. As shown, this is the link attached to the East-West edge of the matrix. Instead of or in addition to the North-South edge, the links can be connected using bidirectional three-state control modules that are attached to the East-West toroidal (envelope) connection. Just like in the preferred embodiment of the invention, which will be discussed later, you can connect 1024 elementary processors in this way, if you use a 32x32 matrix instead of 16x16. The drawing single lines mean single-bit links, while double lines connecting functional elements are a set of connecting, then means bus lines.
In this known solution, the matrix control block generates commands in parallel sent to the elementary processors via the instruction link 18 and generates row and column selection signals transmitted respectively by the line selection links 20 and the column selection link 22. These instructions force the elementary processors to retrieve data from memory, process this data and then to save the data again in memory. For these reasons, each unit processor has access to a bit segment (section or buffer) of the main memory. Here comes the logical conclusion that the main memory of the matrix processor is divided into 1024 separate segments, each of which is allocated to one of 1024 elementary processors in the matrix. This means that in one step of data transmission to the memory or from the memory you can send thirty-two words each with a length of 32 bits. To perform read or write operations, memory is addressed by address members, which are fed to memory address lines via address link 24, and read or write orders are provided in parallel to each unit processor. During the read operation, the row and column selection signals on the row and column selection lines identify which of the elementary processors to perform the operation. As it follows, in the described example you can read a 32-bit word from the memory of up to 32 elementary processors in the selected row, if the matrix is a 32 x 32 matrix. The elementary processor is connected to a segment or block of memory with the index (i, j), which is one bit wide. Since a segment or block of memory is logically connected, according to the principle of one with one, with a corresponding one elementary processor, it can be, and usually is, physically separated on the other board.
Elementary processors P (i, j), from the known solution, are themselves arithmetic-logic units that contain data shift elements and with which each unit is able to remember one bit of information. A multiplexer is provided, connected at the input and output of the arithmetic logic unit and connected to the bidirectional intermediate data register of the memory segment (i, j) assigned and connected to the specific elementary processor P (i, j).
There are separate links for transferring orders and data, and the matrix control block has a micro-instruction memory in which the micro-instruction, specifying the processing operation to be performed by the matrix, is loaded by the main computer 28 using data link 30 and the address and control link 31 . When the operation of the matrix control block is initiated by the main computer 28, a string of micro instructions, a string generated by the micro instructions block that is attached to the memory of micro instructions
167 329 matrix control block 14. The arithmetic logical unit and the block of registers of the matrix control block are used to generate matrix memory addresses, momentum counting, calculation of jump addresses and to perform operations performed by general purpose registers that transmit their output signals to address links matrix control block. The matrix control block also has mask registers for decoding mask and row instructions column, and specific operation orders are sent to the elementary processors via a link of instructions. In this example, the matrix control block has a data buffer inside the control block operably enabled between the master control block data link and the matrix control block data link. From this buffer, data is loaded under the supervision of a micro-order to the control memory in the processor matrix and vice versa. For this purpose, the buffer is implemented as a two-way FIFO buffer (English abbreviation: First-In First-Qut = first in first in out) under the control of the matrix control block.
The solution known from the prior art and discussed above can be compared with the preferred solution according to the invention. Figure 2 illustrates the picket processing block 100, which is a combination of an elementary processor in the form of a bit-parallel processing unit 101 with a local memory 102 attached to this unit to provide the possibility of processing one bit of information in one clock cycle.
In each processing block 100 of the parallel processing system according to the invention, the local memory output 102 is connected to the shift register input 104. The shift register output 104 is then connected to one input B of the bit-parallel processing unit 101 whose output is connected to the first input 105 of two masking registers connected in series 105, 106. The output of the second masking register 106 is connected to the second input A of the bit-parallel processing unit 101, the shift register input 104 and the local memory input 102. The processing block 100 also includes a radio data transmission bus - addresses 103, which is connected to the input of the first masking register 105 and to the local memory input 102, a left-right transfer bus 108, which is connected to the shift register input 104, and a control-state register 107, which is attached to the bit-parallel processing unit 101.
The picket processing block 100 is shaped on a silicon wafer, hereinafter referred to as a picket wafer, with a picket linear matrix in such a way that each picket has an adjacent picket on the side (left or right as shown). In this way, a picket matrix with many local memories is created on the silicon plate, one for each data stream of one byte width, ordered in logical lines or in the form of a linear matrix with links, ensuring communication between neighboring, intended for data transfer to the left and in right. The picket set on the picket plate is geometrically ordered, preferably horizontally, on the plate. Figure 2 shows a typical combination of two pickets from a picket matrix implemented on a multi-station picket plate and a data flow occupying communication paths between each picket processing unit, i.e., elemental processor and memory. Communication paths, which are used to transfer data between memory and elementary processors of the matrix and between neighbors on the right or left are byte paths, and information between neighbors is sent in a slippery manner. The slider can be defined as an element for sending information in one cycle to a place not adjacent to an addressed picket cell that would normally be able to receive information if it was opaque to the message sent until it is reached and received by the nearest active neighbor who picks him up. In this way, the slider functions by sending information to remote cells through off pickets. Suppose station A wants to send a message to a remote station G. Before this cycle, intermediate stations are set to transparent by switching off stations B to F. Then, in the next cycle, station A sends its message to the right, which passes through station B to F, which are transparent because they are off, and station G receives a message because it is still on. In the normal use of the slider, information is transmitted linearly η 329 through the network, but the sliding method can also work in a two-dimensional or even multi-dimensional network.
The cooperation of elementary processors in the preferred example is not bit-serial operation but rather byte-serial operation. Each processor has access to its own attached memory rather than to the local memory block and to the cooperating segment or page of this memory. Instead of a single-bit link, a link with a character width or a width corresponding to a multiple of characters is provided. Instead of bit information, byte information is processed in one clock cycle (or multi-byte information in future systems). In this way, 8, 16 or 32 bits can flow between each picket elementary processor and the allocated memory, corresponding to the width of that memory. In a preferred embodiment of the invention, each picket board has a 32 kB memory with a width of 8 (9) bits and preferably 16 pickets each with such 32 kilobyte memory for each station of the linear matrix. In the preferred embodiment, each associative memory is dynamic random access memory implemented in CMOS integrated circuit technology, and the character byte corresponds to 9 bits (functioning as an 8-bit character with self-control).
Data flow in parallel paths on a byte-wide link between pickets and between elementary processors and their memories is an improvement over the bit-serial structure of prior art systems. However, as will be pointed out later, increasing the degree of parallelism creates new problems that need to be addressed. Some of these solutions are listed below.
The feature that will be highly rated is that in addition to sending a message to the left and right neighbor and to the slider mechanism, which was described earlier, a network connection was also introduced, which is a double byte-wide link so that all pickets can receive the same data at the same time. The signal controlling the pickets and address propagation is also sent via this network link. This is the link that provides comparative data if you are performing a set of associative operations and other comparison or synchronization operations.
Tasks that have a high degree of parallelism in data structure that themselves involve in-station processing for elementary processors to process data under the control of a single instruction stream are applications such as artificial intelligence, context searching and image processing. However, many of these applications, currently possible, could not be used in SIMD processors due to bit-serial processing in a single clock time. For example, the traditional serial elementary processor of the SIMD machine performs one bit of the addition operation in each processor cycle, while the 32-bit parallel machine can perform 32 addition bits in one cycle. The architecture organization, in which 32 kB falls on an elementary processor, uses much larger logical memories, achievable for each elementary processor, than it is ensured by a traditional SIMD machine.
The number of pins on a printed circuit board should be as small as possible, because the data that comes in and out of the board should be minimized. Dynamic random access memory is a simple memory implemented in CMOS integrated circuit technology. It is a memory that provides row-column access by erasing demultiplexed columns within a memory matrix, and providing a row address that reads a row of matrix memory to provide parallel data flow.
The memory, apart from data, contains triple bits, which makes it possible to recognize three logical states, unlike traditional binary (binary) logic, namely: logical one, logical zero and neutral. A neutral state is one that is opposed to a logical one or logical zero. The triple number is placed sequentially in successive memory cells in the memory matrix. Masks are another form of data stored in memory that is directed to the mask register of the picket elementary processor.
167 329
When the memory matrix contains commands, this allows one station to perform an operation other than the other station. The control circuits of individual pickets when performing operations involving most of the pickets, but which does not require the use of all pickets, allow the solution of the invention to be used in areas where solutions based on SIMD operations could not be applied. One simple control function ensures that operations freeze at any picket whose output status corresponds to a special condition. Such a non-zero condition can mean resting. Resting condition is a condition that suspends the operation and puts the picket not in an active state but<sup>in</sup> standby. The second incoming command suspends or allows storage in the memory - depending on the conditions in which the station is located, or on the conditions of the command supplied to the link previously.
When 16 pickets are used on a picket plate, each picket is associated with 32-kilobyte memory, using only 64 plates gives 1024 processors and a 32768 kilobyte memory capacity. The picket matrix is the associative (associative) memory. The solution according to the invention is also useful for numerical image analysis carried out on the basis of intensive data processing in the same way as in the vector calculus. This well-functioning matrix picket processor can now be packed into only two small-sized printed circuit boards. Thousands of such pickets can therefore be placed in a small housing, giving them the form of a unit powered from a low power source, which makes it possible to use the solution according to the invention for image processing with the minimum introduced delays or in the time interval corresponding to the video frame, for example, during flight plane. Features specific to pickets provide the ability to implement large associative memory systems packed in enclosed enclosures.
Figure 3 shows how the associative memory can be fully defined when the compared value is presented for all memory cells and all cells simultaneously respond to the conditions in the line. Associative memories are generally known in the art. In the described solution, using parallel memory pickups of elementary processors that have byte data in order to perform searches, data inputs and masks for searching are provided to place the word K among N words in memory. All matching pickups increase the status of the line, and then a separate operation reads or selects the first K match. This operation, generally called the associative set, can be repeated for subsequent words by the picket memory. Similarly, saving is accomplished through a network activity in which the reported selection line defines a segment and the network data is transferred to all selected stations.
In another embodiment of the invention, although not the most favorable, the random access memory capacity of each station is reduced, enabling the section of the entire associative memory of the type shown in Fig. 3 to be turned on. If it is said that associative memory contains 512 bytes, that each station can contain a set of search indexes and for a single operation it is 512 times 1024 pixels, which means almost 512 x 10<sup>3</sup> comparisons per operation or 512 x 10<sup>9</sup> comparisons per second if one operation requires one microsecond. Going further, you can enter the scope of several terra-comparisons per second. This embodiment provides authoritative attitudes for achieving the feasibility of implementing rules, including extending the scope of information searches at a rate far exceeding the calculation rates currently achieved.
If such operation uses memory and its associated byte-width elementary processors, as shown in Fig. 2, then it becomes possible, in addition to the use of clearly defined algorithms or operational applications, the use of artificial intelligence and parallel programming achievable in situations of SIMD, many other additional applications for a machine implemented in an integrated circuit, which may include:
- parallel arithmetic execution, including matrix multiplication and other tasks that can be performed in special memory machines;
167 329
- performing tasks involving the comparison and processing of images, which tasks can be performed on Von Neumann machines, but whose execution can be significantly accelerated, including applications suitable for extreme parallelism, for example, comparison with a three-dimensional image template;
- data processing based on the contested functions;
- comparing with standards in artificial intelligence applications;
- network control in means ensuring communication between remote local networks (in locks) for quick identification of messages that go to the user to the other side of the network lock;
- modeling of gate levels;
- and in control blocks for the production of integrated circuits with a very high degree of integration to determine the basic conditions of violation.
Processing tasks that leverage the advantages of a memory bank and associative elementary processors will be used by programmers when they take advantage of the benefits and performance of the new system architecture.
By describing the solution according to the invention, one can emphasize the advantages of using a matrix with one gate or one logical element and one processing block 100, i.e. one station. In such a system solution, the process is initiated by assigning each gateway a description as a signal bar that accesses the gate when the system generates input signals and an identification signal. Each time the signal changes, its name is required to be sent over the radio address data bus 103 to all stations and compared in parallel with the names of the expected input signals. If compliance is found, the new signal bit value is recorded in the data flow register in the pickets. When all signal changes are saved, all stations are set to read their contents in parallel with a control word that indicates how to use the current set of input signals to calculate the output. Then, these calculations are performed in parallel, and their result is compared with a fixed value, and the state of all those picket gates whose output signals have changed are recorded in the bit position of the data flow. Next, the external control block queries all pickets and asks for the next gate whose status changes. Then the appropriate name and value of the signal are transferred from the picket to all other pickets remaining in the original state, and the cycle repeats until it is determined that there are no more signal changes, after which the processing is stopped.
Another type of system operation is searching the dictionary entries. Passwords are stored in local station memory 102 so that the first letter of all passwords can be compared to the letter of that password, which requires a password on the data bus address / address 103. All pickets from which the information does not match the selected password are disabled by password are disabled by the specified control word (sign). Then the second letter is compared, with the comparison and off procedure repeated for the next letters (characters) until the active picket blocks are exhausted or the end of the word is reached. At this point, the remaining station blocks are polled and the index of the required data is read by the sequencer.
Figure 4 shows the connection arrangement of several parallel processors and memories, picket blocks ordered in a row arrangement on one silicon wafer, as part of a parallel matrix that can be ordered as a SIMD subsystem, showing the control structure of such a system. The control processor and supervisory (master) processor are also shown here. The figure shows the memory and parallel logic of the elementary processor located on the same board, which in Figure 4 is shown as the part called the picket matrix. Each memory has a width corresponding to n bits, preferably corresponding to a character width of 8 (9) bits, but nothing prevents the use of a memory with a word width of several bytes. It follows that a portion of the memory of the parallel CPU unit will preferably have a width of 8 (9) bits, or, alternatively, 16 or 32 bits. When using modern CMOS technology, it is preferable to use associative memory with a character width (byte o
167 329 9-bit width with self-monitoring) for each picket elementary processor. The memory is directly associated with the elemental processor connected to it according to the one-on-one principle. Each unit processor includes an arithmetic logic unit, masking registers, shift register SR latch 104, state registers and data flow registers DF 105 and 106, which are shown in more detail in the diagram in Fig. 2. Dynamic random access memory and logic of each processor of the picket do not contain any elements of the network of interconnection, because there are direct associations in a one-on-one arrangement between dynamic random access memory with a width of several bits and its elementary processor on the same board.
It can be seen that in Fig. 4, the shift register SR latch 104 is logically engaged between memory and the associated logic of the elementary processor arithmetic and logic, and is in fact the connecting port for each elementary processor along the picket matrix. Each picket board contains a lot of parallel elementary picket processors ordered linearly (which is represented as a straight-mapped link) to provide communication with the picket circuits. The vector address link is a shared memory link, and the data vector address register controls the transition of data to each memory.
Figure 4 also shows the interconnections between the main or microprocessor card, which in a preferred example of the solution according to the invention is a 386 microprocessor implemented as a PS / 2 system after subsystem control, through which general commands go to a standard programmed processor that provides instructions for a 402 sequential order and to the control block 403 executing commands that performs specific micro instructions, triggered by 402 sequential commands. This arrangement can be the same as control block 403 in terms of the functions performed. However, also inside the standard programmable processor, local 405 registers were used, which together with the logical registers of the arithmetic and logic unit (not shown in the figure) form the basis for all addresses that are distributed via the radio bus for all pickets within the 406 picket matrix. In this way, address calculation is performed for all pickets in one arithmetic technology unit without using picket resources or without using picket execution cycles. This important addition provides the flexibility to control picket arrays, allowing you to enable, pause and perform other control functions when performing special tasks and allowing pickets to choose orders or functions.
The 402 instruction sequence controlled from the 407 micro-instruction block feeds instructions to the picket matrix for execution according to the SIMD principle and in the order determined by the main programmable microprocessor, and standard programs from the standard programmable processor being the 408 program library to process data in SIMD mode contained in the picket matrix.
The commands fed to the microprocessor through the subsystem interface are put on a high level of processing, which may include such processes as STARTING, WRITE OBSERVATION, READING RESULT, and they are fed to the microprocessor via the microprocessor CONTROL BLOCK. The microprocessor can be considered as the main system or control processor in the subsystem layout shown in figures 4,5,6 and 7. It is understood that this unit can be an independent device added to a peripheral device (not shown), such as a keyboard and screen monitor. In the case of such a system with an independent unit, the microprocessor system can be considered as a PS / 2 system, in which the files that contain the sequential file directory, which is implemented in the standard programmable processor system, and the processor matrix files are entered in the order shown in Figure 7. The 411 program library may contain program sequences for general program processor control, such as the CALL program and programs with their own names: KALMAN, CONVOLVE and NAV.UPDATE. The selection of these programs is carried out through the user program and in this way general processing can be carried out under conditions when the control functions are carried out by the main external computer or under conditions
167 329 of the control implemented by the user program block 412 located in the MP microprocessor. The data buffer 413 is inserted into the memory by a microprocessor to transfer data to and from the parallel picket processor system. The command sequence 402 is implemented so that it can operate a control stream with a microprocessor and standard programs that are entered into the memory of the 408 work program library. Some of these programs, such as CALLING, CHARGING, LOCKING, SIN, COS, FINDING, MIN, NORMS RANGE and MATRIX MULTIPLICATION are introduced from the set of standard programs through the standard 408 programmed program library.
Inside the standard programmable processor there is also a 407 micro-instruction block, with the help of which it performs functions of controlling the execution of functions at a lower level similar to such as CHARGING, READING, ADDING, MULTIPLY AND MATCHING.
External control such as PREVIOUS / NEXT for each processing unit is preferred and introduced. We also introduce a deterministic implementation of a normalizing floating point byte.
The use of a deterministic approach enables station grouping and group control. The local function is introduced to adapt the system to individual changes of the picket processors.
If the user program requires that it be executed by the processor matrix, then the primary instructions, addresses and data distributed in the network are fed to the picket processor matrix.
The special function that every part of the system uses is determined by the task that should be performed when the user program is compiled.
The flexibility of the subsystem can be illustrated by an example of a solution to the general problem. For example, the matrix multiplication task:
[x] x [y] = [z] can be written as follows:
M xl xR + l
R .
xR x2R yl yM + 1 ...
XM. .
xR xM. .
yM y2M ... yMxC
C zl zR + 1 ...
R .
zR z2R ... zR + C
This problem can be solved by following the steps listed below, clock cycles for operations for each operation, namely:
<td> 01</td><td>Calling matrix multiplication Fx (R, M, C, Xaddr, Yadrr, Zadrr)</td><td> 1</td><td>C</td>
<td> 02</td><td>xSUB = ySUB = zSUB = 1</td><td> 1</td><td> 3</td>
<td> 03</td><td>DO I = 1 is C</td><td> 1</td><td> 3</td>
<td> 04</td><td>DO J = 1 is R</td><td>C</td><td> 3</td>
<td> 05</td><td>z = 0</td><td>CXR</td><td> 5/6*</td>
167 329
<td> 06</td><td>DO K = 1 is M</td><td>C x R</td><td> 3</td>
<td> 07</td><td>* * * Assignment to the associative parallel processor</td><td></td><td></td>
<td> 08</td><td>Zz = Xx x Yy + Zz</td><td>CxRxM</td><td> 204/345*</td>
<td> 09</td><td>* * * Result of return * * *</td><td></td><td></td>
<td> 10</td><td>xSUB = xSUB + R</td><td>CxRxM</td><td> 2</td>
<td> 11</td><td>ySUB = ySUB + 1</td><td>CxRxM</td><td> 2</td>
<td> 12</td><td>NEXT K</td><td>CxRxM</td><td> 3</td>
<td> 13</td><td>xSUB = xSUB - MxR + 1</td><td>CXR</td><td> 2</td>
<td> 14</td><td>ySUB = ySUB - M</td><td>CXR</td><td> 2</td>
<td> 15</td><td>zSUB = zSUB + 1</td><td>CXR</td><td> 2</td>
<td> 16</td><td>NEXT J</td><td>CXR</td><td> 3</td>
<td> 17</td><td>xSUB = 1</td><td>C</td><td> 2</td>
<td> 18</td><td>NEXT with</td><td>C</td><td> 3</td>
<td> 19</td><td>END OF CALL</td><td> 1</td><td> 1</td>
<td></td><td></td><td>transitions</td><td>Scraper</td>
transition
ATTENTION. * means FIXED COMMA (4 bytes) VARIABLE COMMA (1 + 4 bytes).
A closer look at this example suggests that the task specified in instruction 08 requires almost 98% of the cycle time. This is due to the adaptation of the organization of operation in SIMD mode to a parallel picket processor. Other processes (instructions) take only 2% of the cycle time and are assigned to the microprocessor.
Accordingly, an overview of this example of matrix multiplication can be associated with a microprocessor, a standard programmable processor, a local register, or a picket matrix.
In the above example of matrix multiplication, operation according to instruction 01 would be assigned to the main processor, while execution of instructions 02, 05, 10, 11, 13, 14, 15 and 17 would be assigned to the local register, and operation according to instructions 02, 04, 06, 12 , 16, 18 and 19 would be assigned to a standard programmable processor, with the remaining time needed to perform the multiplication operation being used by the picket matrix in response to one command and in one action 08.
Figure 5 shows the construction of a parallel multi-station processor 510 that consists of a large number of parallel picket processors. For applications such as multiple hitting the target, controlled thermonuclear reaction, signal processing, artificial intelligence, image processing obtained from satellite images, recognition of target images, as well as for other applications, a system has been developed that can be implemented in a favorable way, implemented as a SIMD system, consisting of 1024 parallel processors on two to four 511 circuit boards (shown here as consisting of four circuit boards per system) for each processor with 1024 processors. Individual printed circuit boards are inserted into the rack assembly 513, equipped with wedge latches 514. Printed circuit boards are inserted using levers 516, which are intended for inserting and removing printed circuit boards so that when the cover 517 is closed, the mounting system effectively closes and locks the boards in the compartment in which the memory with a capacity of 32 to 62 megabytes is placed. almost 2 billion operations per second. The system is compact, handy, takes up little space, with the matrix consisting of a large number of pickets located in the rear part 518 of the housing, in which part there is also a logical part enabling the interconnection of circuits on printed circuit boards. The processor with 32-megabyte memory is made on 4 printed circuit boards, and the total weight of this device is less than 14 kg. To power the air-cooled processor system, the power supply needed is only 280 W. Each SIMD system has two 520 I / O ports that provide the ability to connect to a main computer or other external devices. Of the many parallel picket processors depicted, each consists of 4 pages
167 329 logical and uses the standard modular packaging used for electronic systems in aviation and a link whose structure enables the connection of external memory, i.e. connection with links such as PI, TM, IEEE 488, whereby the processor can be connected with the input / output port the processor's memory connection, which can be considered an extension of the processor's memory area.
Of the many parallel picket processors shown, each consisting of 1024 parallel elementary processors, each has a 32-byte local memory, and the path assigned to the parallel picket processor has an 8-bit width or a width corresponding to one character (9 bits).
The processors inside each station exchange data with other adjacent processors and with the page via a backed network of connections, preferably crosswise or other.
The system's individual picket processors are housed in a housing containing 2 circuit boards from the four circuit boards that make up the system, a PS / 2 microprocessor on one board and a standard programmable sequential processor chip that is located on the fourth board. This system is illustrated in figures 6 and 7. The individual pickets are processing blocks 100, or the 512 printed circuit boards with pickets can be made from a standard programmable processor so that the processors can independently perform alignment and normalization operations associated with floating point operations.
Processors are controlled in parallel from a common sequential system as discussed herein. The 703 sequential chip board includes a picket processor controller and can be used to picket execute a single instruction chain set up to execute on an array of SIMD processors in a single-byte sequential mode similar to classical serial bit processing. The controller has three layers. Micro-control of pickets is usually done using microcodes, as in modern processors, and is transmitted in parallel to all pickets. Microcontrolling and pickets are synchronized with the same system clock, so that functions controlled by the sequential system can be performed at the same system time. The instructions for introducing into the sequential microcontroller are functions of the standard procedures processor. This 703 sequencing board is a system controller that performs control functions in a loop and performs recursive recursive sequences in a micro-control sequence. Proper picking codes are controlled not by instruction boundaries, but directly by a controller with a library of 408 standard procedures and cycle functions. The standard procedure processor driver contains a large set of macroinstructions that are called by the master system in the subsystem as the master station supervisor. This is the main control system of the picket matrix. Operation of the picket array is checked using type 386 microprocessors. At any given time, all array pickets can execute the same instruction, but it is also possible for processor subgroups to respond individually to the control stream. There are several variants of individual reactions, so that it is possible to achieve local autonomy based on byte control functions for each station (secretion, prohibition, etc.), whose advantages can be used in programming and which can be under the control of the system during program compilation.
In addition, as described, there is local autonomy for addressing memory. The sequential layout of the SIMD controller at work gives a common address to all stations. Each station can expand this address locally to increase data-driven memory access capabilities. In addition, the picket may or may not participate in the operation of the matrix, depending on local conditions.
It is characteristic that it is now possible to introduce the concept of groups into SIMD processing, by providing each picket with the means to assign to one or more of several groups, the processing may be based on this division into groups, although the configuration changes will take place on the run. In one embodiment of the invention, only one group or combination of groups may be active at a time, and then each of them performs the same stream of SIMD processing instructions. Some operations require work only
167 329 subset or group pickets. It is possible to use this feature in the program. It is a way to use this local autonomy of participation in the task at work. Of course, the more pickets are involved in the processing, the better.
One way to increase the number of pickets participating in a task is to enable each of them to execute its own instruction stream. It is basically a MIMD (Multiple Instructions - Multiple Data: many commands - lots of data) work inside SIMD. In the proposed solution it is basically possible to configure the same SIMD machine as a MIMD system or a machine with yet another configuration. This is real thanks to the option of programming pickets to work with their own instruction sequences.
Because each picket can be adapted to work in its own sequence, it is possible to decode the picket level instruction set very simply, allowing for more extensive local processing. The area of the most likely application of this function is the area of making complex decisions, although another area of application interesting for programmers may be simple fixed-point processing.
A simple program of this kind will load the picket program blocks, not exceeding, for example, 2 KB in volume into the local picket memory 102, and this program will be able to be used after the sequential circuit board 703 initiates the SIMD controller local processing under the supervision of executive control starting under specified address xyz. Processing can continue until the controller counts the appropriate number of clock pulses, or the end job signal is detected while monitoring the aggregate SF state stack registers in Figure 4.
In the aggregate state stack (SF - Fig. 4), one SR latch is used for each picket that can be activated to signal the state of the picket. The SIMD controller can check the cumulative value of these latches (one per station) by monitoring the matrix status lines. This matrix state line is a logical combination of the values of each of the picket state latches.
In the following example, we assume that we want to bring the value greater than 250 to the range 500> x> 250. The following procedure uses a cumulative state stack to determine if the task has been completed.
If value <500 then turn off the station
Stat. <condition to disable station
If the stat channel = switched on, then terminated
Value <value - 250
Repeat
Thus, the wet parallel configuration of the picket processors can be combined in various ways, also as a SIMD processor. Such a SIMD machine in a preferred embodiment is programmed to execute a single instruction chain in a classic manner, and coded to work on a matrix of SIMD processors in sequential mode, similar to classic processors with general control of a SIMD controller or sequential system. At the application level, this is done using vector and similar vector instructions, with the vectors being processed inside the processors and shared between them. Vector instructions can be added to macroinstructions, typically 6-10 such instructions.
In such a preferred embodiment, the system is schematically as shown in the functional block diagram of the parallel processor subsystem shown in Fig. 6. Through the I / O ports, the system is controlled from the main interface controller by performing the functions of the subsystem sequencer, similar to the SIMD program with high macro functions level control functions of processing levels. Memory addressing enables the flow of single-byte, 8-bit data, and when performing functions (logical, addition, multiplication, division), logic and arithmetic modulo 8 are used. Floating-point format data blocking is used, autonomous picket operations are implemented, and inactive mode , in partially inactive fashion and with individual addressing.
Figure 7 shows the structure of the subsystem controller. Each of the 512 processor matrix boards (in the drawing illustrating the structure of the subsystem presented in the number of 4 pieces, but which can also be present in the number of 2 SEM E boards) is connected to the 703 sequential system, connected to the 702 subsystem controller, which in turn is opened in either the main system memory side, or towards another subsystem in the set via an integrated circuit 705 attached to the associated microchannel rail 706. In the preferred embodiment of the subsystem, the controller is a universal PS / 2 microprocessor block and works on an integrated Intel 386 microprocessor and 4 MB memory. The personal computer controller 702 is connected to the sequential board via microcannel bus 706 inside the subsystem.
Of course, various modifications and variations of the invention are possible without departing from the idea, and it is therefore understood that the appended claims allow the invention to be implemented other than in the specific example described.
In particular, the following features of the invention may alternatively or together with the characterizing features of the appended claims:
1) Information shifting elements allow it to be transferred in a single cycle to a non-adjacent station block through a station address location that is normally suitable for receiving this information and is not transparent to the message sent before the information was transmitted and received by the nearest active neighbor;
2) The elements of shifting information enable it to be shifted by sending information to a non-adjacent position via an off station that allows the first packet block to transfer information to a remote station, if prior to the transfer cycle the participating picket operations become transparent by turning off the intermediate picket blocks and then in the transfer cycle the controls will trigger the first picket block to send its information to the remote picket block as destination;
3) Elements of message transmission are used linearly across the grille, in two directions of the network, or in three dimensions of the matrix;
4) The picket block set has local autonomy and is also configurable as a matrix of conjugated picket blocks;
5) The system in question should be configured as both a SIMD and MIMD system, and the groups of this set of processing blocks are assigned in accordance with a programmed configuration in which individual system blocks have a certain degree of local autonomy;
6) The system includes a main processor system, the main processor system being connected to an external sequential control system via a bus and equipped with elements for forcing the execution of global instructions in a parallel data processing system;
7) The external control sequencing system is connected to the microcode memory, the microcode memory being programmed to perform functions using standard procedures;
8) The sequential circuit contains high-level macroinstructions for controlling the functions of processing blocks that are connected to this sequential circuit via the bus, addressing the system's local memory allows single-byte data flow and implementation of logic and modulo 8 logic used for logic functions, summation, multiplication, sharing, wherein a floating point operation block is used in these parallel blocks and inactive or partially inactive state with direct addressing of separate picket processing blocks;
9) It is possible to programmatically change the place of implementation of functions inside the main processor system, an external sequential management system equipped with standard procedures, local registers or in a set of picket processing blocks, whereby individual instructions require extensive processing of numerous
167 329 data are assigned to this set of processing blocks configured for SIMD processing;
10) Processing blocks of the system are connected in a matrix of processing blocks with programmable local autonomy, where the picket processing blocks can work with their own instruction sets and can, based on data conditioning, enter or skip operations related to other processing blocks, and the system processing blocks can independently align and normalize operations that are associated with floating point operations;
11) The configuration of the external sequential control system allows performing such operations as CALL, CHARGE, LOCK, SIN, COS, FINDING, MIN, RANGE and MATRIX MULTIPLICATION using standard procedures in the working library of standard procedures;
12) There are procedures to perform the control functions and LOADING, READING, ADDING, MULTIPLY and MATCHING;
13) The processing element exchanges data within an array of adjacent processing blocks and between pages within the system via an internal connection network;
14) An external control processor is used for the picket matrix in which the control microcode is transmitted in parallel to all the pickets in the picket processor group in the matrix, with the control processor and picket processing blocks being synchronized by the same clock, so that functions controlled via an external processor control can be performed at the same system time;
15) The structure, above the external control processor of the picket matrix, uses a microprocessor of the main control system connected to the external control processor via the microcannel bus, where the main microprocessor manages the activity of the picket matrix, and the system is coupled in such a way that all matrix processing blocks can execute the same instruction, although some processor subsets may respond individually to the control information stream;
16) A set of picket parallel processing blocks is arranged along the address bus for communication with an external picket controller, where there is a vector address common for local picket system memories and a register of data address vectors is used to direct the relevant data to each local memory of the picket system. The station is controlled in such a way that the vector address bus is shared by the whole memory, and the vector address data register controls the transmission of data to individual memories, appropriately assigned to the pickets;
17) A set of picket blocks arranged in a matrix and channels for data transfer between picket blocks is used, with communication channels provided with means and elements for bit parallel communication with all other matrix blocks, with coupling channels for transmitting messages from one picket to each other;
18) Messages can be passed from one to another N-dimensional matrix using pickets with galvanic sum;
19) A small part of local memory is implemented in such a way that each of its locations is also used for comparison with a given pattern.
167 329
167 329
100 100
- A_. - ----- X-
<img file="PL167329B1_D0001.tif" />
Fig.2
N
WORDS
<img file="PL167329B1_D0002.tif" />
Figure 3
167 329
<img file="PL167329B1_D0003.tif" />
<img file="PL167329B1_D0004.tif" />
167 329
<img file="PL167329B1_D0005.tif" />
517
516
167 329
External interface
<img file="PL167329B1_D0006.tif" />
FIG. 6
<td colspan="3"></td><td> 1 1</td><td colspan="3"> 703</td><td> 1 1</td><td colspan="2"></td>
<td colspan="3"> 702</td><td> 1 1</td><td colspan="3"></td><td> 1 1 «</td><td colspan="2"> 512</td>
<td></td><td></td><td colspan="3"> 1</td><td> ^-705</td><td colspan="3"> 1</td><td></td>
<td></td><td> 705<sup>from</sup></td><td colspan="3"> _1_ 1</td><td> -706</td><td colspan="3">_L_ 1</td><td> | |</td>
Interconnection System micro-channel connection (inside the system) fTGjZ
167 329
<img file="PL167329B1_D0007.tif" />
FIG.1
UP Department of Publications. Circulation of 90 copies
Price 1.50 PLN
Contents4
12 sheets
Sheet 1 Sheet 2 Sheet 3 Sheet 4 Sheet 5 Sheet 6 Sheet 7 Sheet 8 Sheet 9 Sheet 10 Sheet 11 Sheet 12
119 members in 16 offices
Priority claims4
| Document | Office | Kind | Date |
|---|---|---|---|
| 61159490 | United States of America | A | |
| 61159490 | United States of America | A | |
| 90611594 | – | – | – |
| US19900611594 | – | – | – |
Members119
| Document | Office | Kind | |
|---|---|---|---|
| HU913542D0 | Hungary | D0 | |
| CA2050166A1 | Canada | A1 | |
| EP0485690A2 | European Patent Office (EPO) | A2 | |
| CN1061482A | China | A | |
| HUT59496A | Hungary | A | |
| BR9104603A | Brazil | A | |
| KR920010473A | Republic of Korea | A | |
| PL292368A1 | Poland | A1 | |
| JPH04267466A | Japan | A | |
| CA2064164A1 | Canada | A1 | |
| EP0514043A2 | European Patent Office (EPO) | A2 | |
| MX9206864A | Mexico | A | |
| CA2073516A1 | Canada | A1 | |
| CN1072788A | China | A | |
| EP0544127A2 | European Patent Office (EPO) | A2 | |
| KR930010758A | Republic of Korea | A | |
| JPH05181821A | Japan | A | |
| JPH05233569A | Japan | A | |
| EP0570729A2 | European Patent Office (EPO) | A2 | |
| EP0570741A2 | European Patent Office (EPO) | A2 | |
| EP0570950A2 | European Patent Office (EPO) | A2 | |
| EP0570951A2 | European Patent Office (EPO) | A2 | |
| EP0570952A2 | European Patent Office (EPO) | A2 | |
| JPH0619864A | Japan | A | |
| JPH0628325A | Japan | A | |
| JPH0635872A | Japan | A | |
| JPH0635873A | Japan | A | |
| JPH0635875A | Japan | A | |
| JPH0635876A | Japan | A | |
| JPH0635878A | Japan | A | |
| JPH0636060A | Japan | A | |
| JPH0652124A | Japan | A | |
| JPH0652125A | Japan | A | |
| JPH0668282A | Japan | A | |
| JPH0675931A | Japan | A | |
| EP0544127A3 | European Patent Office (EPO) | A3 | |
| EP0570952A3 | European Patent Office (EPO) | A3 | |
| US5313645A | United States of America | A | |
| EP0514043A3 | European Patent Office (EPO) | A3 | |
| JPH06139200A | Japan | A | |
| EP0570729A3 | European Patent Office (EPO) | A3 | |
| EP0570950A3 | European Patent Office (EPO) | A3 | |
| JPH06214964A | Japan | A | |
| JPH06231092A | Japan | A | |
| TW229289B | Taiwan Province of China | B | |
| EP0485690A3 | European Patent Office (EPO) | A3 | |
| EP0570741A3 | European Patent Office (EPO) | A3 | |
| EP0570951A3 | European Patent Office (EPO) | A3 | |
| SK344091A3 | Slovakia | A3 | |
| CZ344091A3 | Czechia | A3 | |
| PL167329B1This record | Poland | B1 | |
| JPH07287700A | Japan | A | |
| US5475856A | United States of America | A | |
| CZ280210B6 | Czechia | B6 | |
| US5517642A | United States of America | A | |
| JP2512661B2 | Japan | B2 | |
| JP2521401B2 | Japan | B2 | |
| JP2525117B2 | Japan | B2 | |
| JP2533282B2 | Japan | B2 | |
| JP2543306B2 | Japan | B2 | |
| JP2549240B2 | Japan | B2 | |
| JP2549241B2 | Japan | B2 | |
| JP2552075B2 | Japan | B2 | |
| JP2552076B2 | Japan | B2 | |
| JP2557175B2 | Japan | B2 | |
| JP2561800B2 | Japan | B2 | |
| US5588152A | United States of America | A | |
| KR960016880B1 | Republic of Korea | B1 | |
| US5590345A | United States of America | A | |
| US5594918A | United States of America | A | |
| JP2579419B2 | Japan | B2 | |
| US5615309A | United States of America | A | |
| US5615360A | United States of America | A | |
| US5617577A | United States of America | A | |
| US5625836A | United States of America | A | |
| US5630162A | United States of America | A | |
| KR970008529B1 | Republic of Korea | B1 | |
| JP2620487B2 | Japan | B2 | |
| JP2625628B2 | Japan | B2 | |
| RU2084953C1 | Russian Federation | C1 | |
| JP2647315B2 | Japan | B2 | |
| US5708836A | United States of America | A | |
| US5710935A | United States of America | A | |
| US5713037A | United States of America | A | |
| JP2710536B2 | Japan | B2 | |
| US5717943A | United States of America | A | |
| US5717944A | United States of America | A | |
| US5734921A | United States of America | A | |
| US5752067A | United States of America | A | |
| US5754871A | United States of America | A | |
| US5761523A | United States of America | A | |
| US5765011A | United States of America | A | |
| US5765012A | United States of America | A | |
| US5765015A | United States of America | A | |
| US5794059A | United States of America | A | |
| US5809292A | United States of America | A | |
| HU215139B | Hungary | B | |
| US5815723A | United States of America | A | |
| US5822608A | United States of America | A | |
| US5828894A | United States of America | A |
Numbers
- Publication, DOCDB
- 167329
- Publication, EPODOC
- PL167329B
- Application
- 91292368
- Application, DOCDB
- 29236891
- Application, EPODOC
- PL19910292368
Titles
- English
- METHOD OF ASSOCIATIVELLY PARALLEL DATA PROCESSING AND DATA PROCESSING SYSTEM THEREFOR
Classification
- CPC, 9
- G06F15/17343
- G06F15/8015
- F02B2075/027
- G06F7/483
- G06F15/17337
- G06F15/17368
- G06F15/17381
- G06F15/8007
- G06F15/803
- IPC, 9
- G06F15 16
- F02B75 02
- G06F7 57
- G06F9 30
- G06F9 318
- G06F9 38
- G06F13 14
- G06F15 173
- G06F15 80