Shared memory system, parallel processor, and memory lsi
Abstract
This record has no abstract on file.
Term
Term ended
Expired 28 August 2015, 11.1 years ago.
- Priority and filed
- Granted
- Expired
- Today
6 claims: 6 independent, 0 dependent
- 1It is provided between the plurality of processing devices and the shared bus system corresponding to the plurality of processing devices, stores the processing results of the corresponding processing devices, and of another processing device obtained via the shared bus system. In a shared memory system that includes a memory unit that stores the processing results and allows the processing device to obtain the processing results of another processing device from this corresponding memory unit, it is sent from the corresponding processing device and the shared bus system. During the data input means for selecting one of the incoming data, specifying the address, and writing to the memory cell in the memory unit, and the data writing operation by this data input means,Specify the memory cell by the address from the processing device and refer to the processing device.A shared memory system including a data output means for reading data and a write information output means for outputting data to be written to a memory cell in a memory unit corresponding to the processing device to the shared bus system. 複数の処理装置と共有バスシステムとの間に前記複数の処理装置に対応して設けられ、対応する処理装置の処理結果を記憶するとともに、前記共有バスシステムを介して得られる他の処理装置の処理結果を記憶するメモリユニットを備え、処理装置が他の処理装置の処理結果をこの対応するメモリユニットから得られるようにした共有メモリシステムにおいて、 対応する処理装置と共有バスシステムとから送られてくるデータのいずれかを選択し、アドレスを指定して、メモリユニット内のメモリセルに書き込むデータ入力手段と、このデータ入力手段によるデータの書き込み動作中に、メモリセルを前記処理装置からのアドレスで指定して、前記処理装置に対してデータを読み出すデータ出力手段と、処理装置が対応するメモリユニット内のメモリセルに書き込むデータを前記共有バスシステムに出力するライト情報出力手段と、を備えたことを特徴とする共有メモリシステム。
- 2A plurality of processing devices are provided between the plurality of processing devices and the shared bus system corresponding to the plurality of processing devices, the processing results of the corresponding processing devices are stored, and the other processing devices obtained via the shared bus system are stored. In a shared memory system that includes a memory unit that stores the processing results and allows the processing device to obtain the processing results of another processing device from this corresponding memory unit, it is sent from the corresponding processing device and the shared bus system. A write that outputs to the shared bus system a selection means for selecting one of the incoming data and addresses, and data sent from the processing device to the corresponding memory unit and written to the memory cells in the selection means. Information output means, multiple memory cells that can be specified by address in the memory unit,From either the shared bus system or the processing equipmentDuring the write address specifying means for designating the address to write data, the writing means for writing data to the memory cell at the specified address, and the data writing operation by each of the above means,From the processing deviceSpecify the memory cell by addressFor the processing deviceA shared memory system including a read addressing means capable of reading data and a data reading means. 複数の処理装置と共有バスシステムとの間に前記複数の処理装置に対応して設けられ、対応する処理装置の処理結果を記憶するとともに、前記共有バスシステムを介して得られる他の処理装置の処理結果を記憶するメモリユニットを備え、処理装置が他の処理装置の処理結果をこの対応するメモリユニットから得られるようにした共有メモリシステムにおいて、 対応する処理装置と共有バスシステムとから送られてくるデータ及びアドレスのうちいずれか一方のデータ及びアドレスを選択する選択手段と、処理装置から対応するメモリユニットに送られてその中のメモリセルに書き込まれるデータを、前記共有バスシステムに出力するライト情報出力手段と、メモリユニットに、アドレスによって指定できる複数のメモリセルと、前記共有バスシステムまたは処理装置のいずれかからのデータを書き込むアドレスを指定するライトアドレス指定手段及び指定されたアドレスのメモリセルにデータを書き込む書き込み手段と、前記各手段によるデータの書き込み動作中に、前記処理装置からのアドレスでメモリセルを指定して前記処理装置に対してデータを読み出すことができるリードアドレス指定手段及びデータの読み出し手段と、を備えたことを特徴とする共有メモリシステム。
- 3Claim1In the shared memory system described in the above, a synchronization control means is provided to synchronize the timing of reading data from the corresponding memory unit by the processing device and the timing of writing data from the shared bus system to the memory unit to one reference clock. A featured shared memory system. 請求項1に記載の共有メモリシステムにおいて、処理装置が対応するメモリユニットからデータを読み出すタイミングと共有バスシステムからのデータをメモリユニットに書き込むタイミングとを一つの基準クロックに同期させる同期制御手段を設けたことを特徴とする共有メモリシステム。
- 4Claim1In the shared memory system described in the above, each processing device is provided with a means for outputting a synchronization request signal that changes from inactive to active upon completion of processing, and all synchronization request signals from each processing device that performs cooperative processing are active. A shared memory system characterized by providing a synchronization means that enables the processing device that outputs the synchronization request signal to read data from the memory unit by actively switching the synchronization processing completion signal in response to the change to. .. 請求項1に記載の共有メモリシステムにおいて、処理の終了によって非アクティブからアクティブに転じる同期要求信号を出力する手段を各処理装置に設けると共に、協調して処理を行う各処理装置からの同期要求信号が全てアクティブに転じたことを受けて同期処理完了信号をアクティブに転じ、同期要求信号を出力した処理装置によるメモリユニットからのデータの読み出しを可能にする同期化手段を設けたことを特徴とする共有メモリシステム。
- 5Claim4In the shared memory system described in the above, when a write operation to the memory unit occurs at the time when the synchronization processing completion signal turns active, the write operation from the memory unit is continuously generated for a period of time. A shared memory system characterized by having a means for prohibiting read processing. 請求項4に記載の共有メモリシステムにおいて、同期処理完了信号がアクティブに転じた時点でメモリユニットへの書き込み動作が発生している場合、その書き込み動作が連続して発生している期間、前記メモリユニットからの読み出し処理を禁止する手段を備えたことを特徴とする共有メモリシステム。
- 6Claim2A shared memory system according to the above, wherein the data reading means of the memory unit includes a read data latch that latches the data to be read. 請求項2に記載の共有メモリシステムにおいて、前記メモリユニットのデータの読み出し手段に、読み出されるデ-タをラッチするリードデ-タラッチを備えたことを特徴とする共有メモリシステム。
Independent claims6
256 paragraphs, as filed
[Industrial Application Field] The present invention relates to a parallel processing device for exchanging information between a plurality of processing devices, and a memory LSI that can be used in such a device.
[0002] Conventional technology As a shared memory system of a conventional parallel processing device, a method is adopted in which one shared memory is provided on a shared bus system and the shared memory is commonly used by a plurality of processors. There is something. The shared bus system is an arbiter circuit that arbitrates and grants permission to access a shared bus issued from a shared bus or a device connected to the shared bus, a device that inputs / outputs data to the shared memory, and the like. Is appropriately included.
[0003] As a more advanced shared memory system, as in the system shown in Japanese Patent Application Laid-Open No. 5-290000, shared memory (memory unit) distributed to each processor in order to reduce access contention in the shared bus system. ) Is provided. Such shared memory may be referred to by a name such as local shared memory or distributed shared memory.
[0004] As a shared memory system of a parallel processing unit provided with this local shared memory, when the content of the local shared memory of one processor is changed, the content of the local shared memory of one processor is broadcast to the other processor. A broadband parallel processor that also changes the contents of the locally shared memory is known. The system shown in JP-A-5-290000 also belongs to this broadcast type.
[0005] In a parallel processing device of a type having one shared memory in a shared bus system, a large number of read cycles and write cycles from a plurality of processors are complicated on the shared bus system. There is a possibility of conflict. Wasted time is generated on the shared bus system side in order to mediate this access conflict, and the flow is reduced accordingly. Further, in conjunction with this, the standby time on the processor side becomes long, and problems such as an increase in the overhead of the entire processing system also occur.
[0006] If a broadcast-type shared memory system is used as in the system shown in Japanese Patent Application Laid-Open No. 5-290000, only a data write cycle for the local shared memory is generated on the shared bus system. Will be done. Then, the data read cycle from the shared memory is performed independently and in parallel for each processor with respect to the locally shared memory distributed and arranged in each processor. Therefore, access contention does not occur between read cycles for the shared memory of each processor, and the flow is improved.
[0007] However, even if a broadcast-type shared memory system is used, the read cycle from the processor side and the write cycle from the shared bus system side for the local shared memory are on this local shared memory. Conflict (access conflict between read cycle and write cycle). Therefore, even in a parallel processing apparatus having a broadcast-type shared memory system, the effect of removing overheads, wasted time, and the like is not sufficiently obtained. It should be noted that this conflict also occurs in shared memory systems other than the broadcast type.
[0008] In a parallel processing device having a shared memory system, when a plurality of processors coordinating processing need to reliably pass a task processing result to a subsequent task processing in each other's task processing. There is. At this time, it is necessary to consider the delay in data transfer time (communication delay) between processing devices such as processors. Even in the case of a shared memory system using the broadcast method, it is desirable to provide a synchronization means suitable for a parallel processing device that reduces the above-mentioned access contention on the local shared memory.
[0009] Therefore, an object of the present invention is to reduce access conflict on the local shared memory between the read cycle from the processing device side such as the processor and the write cycle from the shared bus system side for the local shared memory. It is an object of the present invention to provide a shared memory system, a parallel processing device, or a memory LSI that can be used in such a device.
[0010] Another object of the present invention is to provide a synchronization means for ensuring that data is reliably passed between tasks while reducing access contention on the local shared memory. It is in.
[0011] [Means for Solving Problems]<u style="single">The above object is provided between the plurality of processing devices and the shared bus system corresponding to the plurality of processing devices, stores the processing results of the corresponding processing devices, and is obtained through the shared bus system. In a shared memory system that includes a memory unit that stores the processing results of another processing device and allows the processing device to obtain the processing results of the other processing device from this corresponding memory unit, the corresponding processing device and the shared bus system. Select one of the data sent from and, specify the address, and specify the data input means to write to the memory cell in the memory unit and the memory cell by the address during the data writing operation by the means. This is achieved by providing a data output means for reading the data and a write information output means for outputting the data to be written to the memory cell in the corresponding memory unit by the processing device to the shared bus system.</u>【0012】<u style="single">The above object is provided between the plurality of processing devices and the shared bus system corresponding to the plurality of processing devices, stores the processing results of the corresponding processing devices, and is obtained through the shared bus system. In a shared memory system that includes a memory unit that stores the processing results of another processing device and allows the processing device to obtain the processing results of the other processing device from this corresponding memory unit, the corresponding processing device and the shared bus system. The shared bus uses a selection means for selecting one of the data and addresses sent from and the data sent from the processing device to the corresponding memory unit and written to the memory cell in the selection means. Write information output means to output to the system, multiple memory cells that can be specified by address in the memory unit, write address specification means to specify the address to write data, and write means to write data to the memory cell of the specified address. It can also be achieved by providing a read address specifying means and a data reading means capable of designating a memory cell by an address and reading data during the data writing operation by each of the above means.</u>[Action] The data input means selects one of the data and the address sent from the processing device or the shared bus system and writes the data and the address to the memory cell in the corresponding memory unit. .. Further, the write information output means outputs the data sent from the processing device to the memory unit to the shared bus system. This data is written to the memory unit of another processing device. On the other hand, the data output means reads the data by designating the memory cell by the address during the writing operation of the data input means.
[0022] With this configuration, it is possible to perform the data writing process and the reading process to the memory cell in the memory unit in parallel. Therefore, the read cycle from the processing device side and the write cycle from the shared bus system side for the memory unit (local shared memory) reduce the access contention caused on this memory unit (local shared memory). Can be done.
At this time, the memory cell designated for writing the data and the memory cell designated for reading the data may be the same. That is, data can be read from the memory cell during the writing process.
Further, in the above-mentioned memory LSI, a read address specifying means for designating a memory cell for reading data, a read means for reading data from the designated memory cell, and a write address for designating a memory cell for writing data are specified. Since the writing means for writing data to the memory cell is provided independently, the data reading process and the data writing process can be performed in parallel. As a result, access contention between the read cycle and the write cycle on the memory LSI can be reduced.
[0025] Further, the synchronization means monitors the synchronization request signal from each processing device that performs processing in cooperation. This synchronization request signal is output to the synchronization means when the processing device finishes processing, and the synchronization means outputs this synchronization request signal after all the synchronization request signals are prepared (after turning active). Allows the output processing device to read data from the corresponding memory unit. As a result, each processing device is stored in the local shared memory in a state where the necessary information from other processing devices that are cooperating in processing is not written in the memory unit (local shared memory) corresponding to each processing device. It is possible to prevent access, obtaining incorrect data, and generating incorrect processing results.
[0026] At this time, by providing the interlock circuit for local synchronization, the data read processing by the processing device is made to wait until the synchronization processing completion signal is generated and the data in the local shared memory is actually rewritten. As a result, it is possible to prevent the processing device from obtaining old data and performing erroneous processing.
[0027] Since the synchronization control means synchronizes the entire shared memory system with one clock, it is possible to eliminate the overhead for synchronizing each means or processing operating asynchronously, and improve the communication latency (delay). can do.
[0028] By latching the data read from the memory unit or the memory cell by the processing device by the read data latch, the writing process can be executed regardless of the reading process.
[0029] Further, the memory LSI of the present invention reads data, a port used only for writing data, a port used only for reading data, a write address designation port for designating an address for writing data, and a port for designating a write address. By providing a read address designation port for designating an address, it is possible to perform data read processing and data write processing in parallel.
[0030] Further, the memory LSI described above makes it possible to execute a write process and a read process, which have conventionally required at least two process cycles, in one process cycle. Even the shortest processing cycle is about 10 ns for ICs or LSIs that use the CMOS process, and about 5 ns for ICs or LSIs that use the bipolar CMOS process. Therefore, in order to execute the write process and the read process, it has conventionally required about 20 ns for an IC or LSI using a CMOS process and about 10 ns for an IC or LSI using a bipolar CMOS process. On the other hand, in the memory LSI of the present invention, by enabling the write process and the read process to be executed in parallel, these processes are executed in one processing cycle, and the write process and the read process are performed in a time of 5 ns or less. Made it possible to run with.
By using the above-described memory LSI of the present invention as a memory unit (local shared memory) of a shared memory system, data reading processing from this memory unit (local shared memory) and data writing to the local shared memory can be performed. A parallel processing device capable of performing processing in parallel can be configured.
[0032] Further, even if the shared memory system or the shared memory system using the memory LSI as described above is not distributed to each processing device and is provided for a plurality of processing devices, the above-mentioned write cycle and read cycle can be combined. It will be possible to reduce access contention.
[0033] In the following description, the shared memory system may be simply referred to as a shared memory.
[Example] In a multiprocessor system composed of a plurality of processors, a waiting process between a shared system (shared memory, shared I / O, etc., which can be freely accessed between processors) and a processor. That is, a method of improving parallel processing efficiency by combining a synchronous processing circuit that executes inter-processor synchronous processing and combining control-flow parallel processing control and data flow-like parallel processing control is available. , There is an example already used in the conventional system as shown in Japanese Patent Application Laid-Open No. 5-2568.
[0035] This Japanese Patent Application Laid-Open No. 5-2568 describes the overall architecture and method. The present invention discloses an optimal synchronization processing method in a highly efficient access method and configuration for the shared memory disclosed in the present specification.
[0036] First, the configuration of FIG. 1 and the interprocessor synchronization processing method during parallel processing will be briefly described.
The system shown in FIG. 1 is one of the shared systems in a multiprocessor system composed of a plurality of processors 0 to n and a shared system which is a source that can be freely accessed from any of them. The shared system controller and shared memory are integrated for each processor so that the shared memory systems 1010 and 1011 to 101n can be regarded as equivalent to the local memory of each processor when viewed from each processor. When one processor changes the contents of the shared memory in its own shared memory system, the shared memory in the shared system of another processor is also changed accordingly. It is said.
[0038] Further, a synchronization processing circuit 1000 is provided to perform synchronization processing between each processor and control parallel processing of each task executed between each processor, as shown in Japanese Patent Application Laid-Open No. 3-234535. In addition, parallel processing control is performed by combining the controller flow and the data flow.
[0039] That is, when a certain processor finishes a certain task, a synchronization request (SREQ) notifying the end of the task is issued to the synchronization processing circuit 1000, and a task for which wait processing (synchronization processing) must be performed is executed. The synchronization processing circuit 1000 operates so that the synchronization completion information (SYNCOK) is kept inactive until the task processing of the other processor is completed and the processor issues a synchronization request (SREQ) to the synchronization processing circuit 1000. To do. And in fact, the process of waiting for a processor is executed when that processor accesses shared memory, at which time it suspends access to the processor's shared memory until it becomes active if SYNCOK is not active, and SYNCOK is active. If there is, it works to allow access to shared memory unconditionally.
[0040] The SYNCOK signal of this example may be considered to have a function substantially equivalent to that of the TEST signal in JP-A-5-2568.
[0041] In FIG. 1, the processors 0 to n are connected to the corresponding shared memory systems 1010 to 101n by a data bus (D), an address bus (A), and a control bus (C), respectively. In this example, when a processor accesses the corresponding shared memory system, the shared system enable (CSEN) that indicates it is activated, signaling the start of the access cycle to the shared memory system.
[0042] The signal corresponding to CSEN can be internally generated by decorating the address signal A or the like in each shared memory system 1010 to 101n, but it precedes on the processor 0 to n side. Since there is a high possibility that the delay time can be reduced by decorating and generating, in this example, CSEN is directly given from the processor side as a signal independent of each signal group of D, A, C. I have to.
[0043] Further, in FIG. 1, each shared memory system 1010 to 101n is connected to a shared bus system (consisting of signal lines REQ, Data, Address, Control, and ACK signal) 1900.
[0044] As described above, in this shared bus system 1900, when a processor changes data (performs write access) to the shared memory, the corresponding address on the shared memory of another processor is used. The shared memory system of the processor that has write-accessed the shared memory with the information to be changed together with the data existing in is provided to be bundled to all other shared memory systems.
That is, if a write cycle for a shared memory system occurs in some processor, that information is transmitted to the shared memory system of another processor via the shared bus system 1900, and each shared memory associated with each processor. The necessary data changes on the corresponding address above are made.
[0046] In the shared bus system 1900, the REQ signal group is a set of bus request signals (REQ) generated from the shared memory controllers in each shared memory system 1010 to 101n at the time of write access to the shared memory. These are input to the bus arbiter circuit 1020. The arbiter circuit 1020 selects one of them, activates the permission signal ACKm corresponding to REQm (m is the request signal corresponding to the processor m), and activates the ACK of the corresponding shared memory system via the ACK signal group. Tell the input.
[0047] When the ACK input turns active, the shared memory controller generates the data and address that are the targets of the write cycle on the shared bus system 1900, and also the arbiter. Circuit 1020 activates a control signal (BUSY) that indicates that the information on those shared buses is active or that the bus is in use.
[0048] The information of the BUSY signal is transmitted to the visit signal (BUSY) input of each of the shared memory systems 1010 to 101n via the control signal (Control) in the shared bus system 1900, and each shared memory controller has its own. By examining the information, it is determined whether or not there is data to be written to the shared memory on the shared bus. If there is data to be written to the shared memory (if the visit signal is active), the valid data is written to the specified address of each shared memory all at once and changed, and each processor It works so that the contents of the shared memory corresponding to are always the same.
[0049] Depending on the system, a method in which the shared memory controller that receives the permission signal (ACK) from the arbiter circuit outputs a visitor signal and transmits it to another shared memory controller is also considered. However, compared to this example, the output of the visit signal requires a longer time (the signal delay is large), so the method of this example will be more effective in a system that requires high-speed operation.
[0050] In addition, depending on the system, various control signals (Control) from bus commands, status information, bus clocks, data transfer protocol control signals, bus status and bus cycle control signals, and sources. Response signals, interrupt vectors, message information signals, etc. may be assigned.
FIG. 2 shows the structure in each shared memory system 1010 to 101n in the present invention. The biggest feature is that the shared memory 2006 has a read address (RA) and its corresponding output data (DO), and a write address (WA) and its corresponding input data (DI). It has a two-port memory structure that is provided as separate ports.
[0052] In the shared memory system, the two-port shared memory 2006, the shared memory control unit 2010, the processor interface 2003, the machine stage controller MSC2002, and various input / output buffer units (2001, 2012 ~ It is composed of a latch unit and a buffer memory unit (2004, 2008, 2009, 2011), a multiplexer unit (2005, 2007), a clock generation circuit 2013, and the like.
[0053] Each shared memory system operates in synchronization with a basic clock such as a processor clock (PCLK) and a system clock (SCLK). PCLK is a clock synchronized with the bus cycle of the processor, and it can be considered that the bus cycle on the processor side operates based on this clock. SCLK is the basic clock of the entire system, and the system can be considered to be synchronized with this clock. In the most ideal condition, if PCLK is generated with reference to SCLK, the entire system including the processor will eventually be operated in synchronization with one basic clock (SCLK in this case). It is considered that efficient timing control becomes possible.
[0054] The features and basic operations of the shared memory system of the present invention are as shown in a) to f) below.
[0055] a) When the shared memory system access enable signal (CSEN) is activated, the shared memory system processor 2010 and PIF2003, MSC2002 obtain the information via the signal input circuit 2001 to the shared memory system. Know that access from the processor has occurred.
[0056] Then, the processor interface PIF2003 and the machine site controller (MSC) 2002 are the address information and the data to be accessed by the processor at an appropriate timing matching the bus cycle and bus protocol of the processor. Exchange information with the processor.
[0057] In this example, the physical address area of the shared memory is decorated on the processor side so that the CSEN signal becomes active when the processor accesses the area.
[0058] Further, MSC2002 operates each processor system including the shared memory system with a single reference clock by adopting the usage method based on Japanese Patent Application Laid-Open No. 2-168340, and the entire system is synchronized. The effect of being able to be constructed as a scale digital circuit system and the effect of ensuring a longer access time (especially during a read cycle) when the processor accesses the shared memory system can be obtained.
[0059] b) When the access bus cycle to the shared memory of the processor is a lead cycle, the data is directly shared using the read ports (DON and RA) of the shared memory 2006, except in special cases. Read from memory 2006. In this case, the address multiplexer MX2007 that multiplexes various address information and gives it to the shared memory 2006 obtains the address information from the processor to the input C via PIF2003 and gives it from the output O1 to the lead address RA of the shared memory 2006. , Read the data corresponding to the RA value from the DO of the shared memory.
The read data value is sent to the processor via PIF2003.
[0061] In AMX2007, the operation of controlling the selection input signal S1 that determines which of A, B, and C on the input side is selected and output to O1 is the RDSEL signal from the shared memory control unit 2010. Do by. At that time, a read latch circuit is provided in the processor interface 2003, the data from the shared memory 2006 is latched there once, and at least the period before and after the timing when the processor reads the data, It is also possible to keep valid data for the processor in a form that secures sufficient setup time and hold time.
[0062] Further, as will be described later, when the clock (bus clock or the like) that defines the bus cycle on the processor side and the clock that defines the timing for writing data to the shared memory side are synchronized, the shared memory If the valid period of the data read from 2006 originally satisfies the setup time and the holding time, the data may be directly given to the processor.
[0063] c) When the access bus cycle to the shared memory of the processor is a write cycle, in this example, the write address value sent from the processor via PIF2003 first responds to the AWBUFCTL signal of the shared memory control unit 2010. Then, it is written to the address write buffer AWBUF2008 at an appropriate timing.
[0064] AWBUF2008 may be configured to store a plurality of write address information in chronological order, output the write address information obtained in the past to O, and give it to the A input of AMX2007. AMX2007 outputs the selected write address value from the address inputs A, B, and C to O2, and gives it to the write address WA input of the shared memory 2006. The selection signal input S2 for performing the selection operation is performed by the write data signal WDSEL of the shared memory control unit 2010.
[0065] Also, the data to be written to the target write address WA is also selected from the processor via PIF2003, once via the data write buffer DWBUF2004, and then by the data multiplexer DMX2005 (input to A and set to O). (Output) and given to the data input DI of shared memory 2006.
[0066] The function of DMX2005 is almost the same as the function of the O2 output side of AMX. However, the selection signal input S for selecting one from the respective write address information input to the inputs A, B, and C and outputting it to O is controlled by the WDSEL signal of the shared memory control unit 2010.
[0067] The function of DWBUF2004 is almost the same as that of AWBUF2008, but the DWTBUFCTL signal of the shared memory control unit 2010 is used as the control signal for latching and storing the data from the processor in DWBUF2004. To. If the control timing of DWBUF2004 and DMX2005 is the same as that of AWBUF2008 and AMX2007 (for example, if the output timing of the address value and data value from the processor is almost the same), the same control signal is used for the selection signal and latch. The signal may be controlled.
The operation of writing the data to the shared memory 2006 is performed by the write enable WE signal from the shared memory control unit 2010. In this example, when the WE signal is activated, the data input to the DI of the shared memory 2006 is reflected in the contents of the memory cell corresponding to the address value input to the RA, and the timing to return the WE signal to inactivity. The data is latched in the memory cell. If the contents of RA and WA of shared memory 2006 show the same address value, the same data as the contents of the data input to DI is output to DO when the WE signal is active. To.
Therefore, in the case of this example, it can be said that the timing of changing the information in the shared memory is determined by the timing of activating the WE signal.
[0070] During the write cycle, not only the contents of the shared memory itself are changed, but also the same data and address information are broadcast-cast to each shared memory corresponding to other processors to display the contents on the shared memory. Need to change. Therefore, it has a function to output the write data output from O of DWBUF2004 and the write address output from O of AWBUF2008 to the shared bus system via the data buffer 2015 and the address buffer 2016, respectively. ..
The ON-OFF operation for the shared bus system between the data buffer 2015 and the address buffer 2016 is performed by the DEN signal and the AEN signal, respectively.
[0072] d) When another processor changes the contents of the shared memory, the address information sent via the shared bus system is obtained from the address buffer 2016, and the data information is obtained from the data buffer 2015, respectively. Write the data information corresponding to the address information to the shared memory 2006.
[0073] In this example, the information obtained via the data buffer 2015 is held once in the data latch 2009, and the information obtained via the address buffer 2016 is held once in the address latch 2011, and then the data information is B of DMX2005. At the input, the address information is input to the B input of AMX2007, and the data to be written from the O output of DMX2005 to the DI input of the shared memory 2006 is targeted to the WA input of the shared memory 2006 from the O2 output of DMX2007. Address is entered.
[0074] The latch timing to the data latch DL2009 and the address latch AL2011 is performed by the CSADL signal of the shared memory control unit 2010. The CSADL signal is operated so that the data and address information are confirmed on the shared bus system, and the latch processing is performed at the timing when they secure sufficient setup time and hold time for DL2009 and AL2011. There is.
[0075] In this example, when the CSADL signal becomes active, the information on the D side of DL2009 and AL2011 is output to the O side to set up the latch circuit, and the information is sent to the timing DL and AL when the CSADL turns inactive. Latched. In this example, since each shared memory system of each processor operates in perfect synchronization in response to the basic clocks (PCLK and SCLK) having the same phase, the data buffer 2015 and the address buffer 2016 are operated during write operation. The timing of inputting / outputting the information required for the shared bus system and the timing of latching the information required for DL2009 and AL2011 generated in synchronization with the timing are the shared memory controller units 2010 in each shared memory system. It can be considered that it is clarified inside.
[0076] By this synchronization, the shared memory control unit 2010 is efficient with little overhead and delay time for generating control signals such as CSADL, DEN, AEN, AWTBUFCTL, DWTBUFCTL that specify these timings. Timing control is possible. In addition, the generation of the WE signal that writes data to the WDSEL that controls DMX2005 and AMX2007 and the shared memory 2006 also responds to the control timing of the CSADL signal that determines the data to DL2009 and AL2011, and the shared memory control unit 2010 You can do it inside.
[0077] e) Conflict control (arbitration control) that reliably allocates the right to use the shared bus system to one processor between the shared memory systems of each processor during the write operation to the shared memory. You will need it.
[0078] When the shared memory control unit 2010 considers that a write cycle (write cycle) to the shared memory system by the processor has occurred from the control signal C from the processor and the shared memory system access enable CSEN. The CSREQ signal is activated and the request signal REQ to the aviator circuit 1020 is generated via the output buffer 2012.
[0079] Then, when the permission signal ACK from the corresponding arbiter circuit 1020 is activated and obtained for the CSACK input of the shared memory control unit 2010 via the input buffer 2014, the own processor uses the shared bus system. With the right, the shared memory controller 2010 will generate a write cycle to the shared memory and shared bus system according to the procedure shown in c).
At this time, if the data buffer DWBUF2004 and the address buffer AWBUF2008 are full, the end of the bus cycle of the processor is pending and waited. The shared memory controller 2010 generates an RDY signal as a signal that determines whether to wait for the end of the bus cycle on the processor side to wait or to end the bus cycle as scheduled and proceed to the next processing.
[0081] When the processor is generating a bus cycle for accessing the shared memory system, if the RDY signal from the shared memory control unit 2010 becomes active, the bus cycle is not put into a waiting state as scheduled. If the bus cycle is terminated and the processor is advanced to the next process and kept inactive, the bus cycle is extended without ending the bus cycle, and as a result, the processor is made to wait.
[0082] Basically, if the buffers 2004 and 2008 are not full and there is free space, the processor latches the data and address information required for DWBUF2004 and AWBUF2008 and proceeds to the next process without waiting. move on. That is, when the processor executes the write operation to the shared memory, if the permission of the shared bus system usage right from the above-mentioned arbiter circuit 1020 continues to be obtained, it is pending in the buffers 2004 and 2008. The data and address information for the written write cycle is stored in chronological order, and the write cycle issued when the buffer is full is extended until the buffer becomes full, and as a result. You will have to wait on the processor side.
[0083] In order to reduce the latency of the writing process to the shared memory, the buffers 2004 and 2008 are completely free during the write operation to the shared memory, and the arbiter circuit 1020 responds to the write operation. If the authorization signal (CSACK) from is immediately activated and the shared bus system is allowed to be used, the write address (WA) and write data (DI) from PIF2003 are sent directly through the C input of the multiplexer 2005,2007. It may be given to the shared memory 2006 and write processing may be executed.
[0084] The control is performed by the shared memory controller 2010 using WDSEL and WE signals. Note that CSEN remains active in the shared memory control unit 2010 as long as valid information exists in these buffers 2004 and 2008.
On the other hand, although the REQ signal is active, the corresponding ACK signal is inactive, and the BUSY signal obtained from the arbiter circuit 1020 via the input buffer 2017 is (shared memory controller-). If active (connected to the CSBUSY input of the unit 2010), it is considered that the write cycle to the shared memory and shared bus system by other processors is permitted and executed, and the shared memory control unit 2010 , D) Generates a write cycle to the shared memory 2006 based on the information that is broadcast from another processor via the shared bus system by the method shown in d).
[0086] f) When the synchronous processing circuit 1000 and the shared memory systems 1010 to 101n operate in conjunction with each other, the processor activates the synchronization request signal SREQ when the necessary task processing is completed and synchronizes the synchronous processing circuit 1000 with the synchronous processing circuit 1000. When accessing the shared memory system when data on the shared memory (data from another processor etc. exists) is required at a later timing (especially -At the time of access), perform local synchronization processing with the processor in the shared memory system so that there is no inconsistency in the data exchange with other processors.
[0087] The synchronization processing circuit 1000 is a processor belonging to a group of processors to be synchronized in advance among the synchronization request signals SREQ from each processor 1110 to 111n, that is, a group in which processing is proceeding in cooperation. If there is at least one inactive state in the SREQ from, the necessary synchronization processing is completed by keeping the synchronization processing completion signal SYNCOK inactive until all the SREQs are activated for the processor to respond to. Tell them that you haven't.
The shared memory system for the processor receives the SYNCOK signal in the signal input circuit 2018 and monitors the synchronization information from the synchronization processing circuit 1000, and the processor owns the shared memory at least when SYNCOK is inactive. When an access cycle to the system (especially a lead cycle) is generated, the shared memory control unit 2010 keeps the RDY signal inactive so that it waits for the end of the processor's bus cycle to wait with the processor. Perform local synchronization operations with the shared memory system.
[0089] As a result, the shared memory is accessed in a state where the necessary information from another processor cooperating on the shared memory 2006 is not written, and as a result, incorrect information is obtained. It is managed so that erroneous processing results are not generated.
[0090] The characteristic feature of this example is that the buffer systems DWBUF2004 and AWBUF2008 are provided, and even if the operation of the processor precedes the access cycle processing of the shared memory system, the access information is transmitted to these buffer systems in chronological order. The point is that it can be stored and post-processed independently and in parallel with the operation of the processor in the shared memory system .
As a result, it is possible to proceed without waiting for the processing of the processor more than necessary. At this time, even if SYNCOK is active, if valid data exists in the buffer systems 2004 and 2008 in the shared memory system of each processor, that is, if the synchronization processing is originally completed. There may be situations where the data that needs to be in shared memory does not yet exist in shared memory.
In this state, when a free read operation from the read port (RA, DI of shared memory 2006) that can be executed in parallel with the write cycle to the shared memory as shown in b) is executed, It may lead to erroneous processing because the required data cannot be obtained.
Therefore, while valid data exists in the buffer system 2004,2008 of the shared memory system corresponding to any processor, write cycles occur continuously on the shared bus system and the arbiter. Utilizing the fact that the BUSY signal from circuit 1020 remains active, if the BUSY signal is active even if the SYNCOK signal turns active, it goes to the shared memory of the processor until the BUSY signal turns inactive. The shared memory processor unit 2010 has a function to prohibit the lead cycle.
That is, when a lead cycle from the processor occurs in this state, the shared memory control unit 2010 keeps the RDY signal inactive until the BUSY signal becomes inactive, and terminates the bus cycle. Delay and make the processor wait.
[0095] In the basic functions of the shared memory system of the present invention shown in a) to f), the following two points are remarkably different from the conventional system.
1) Shared memory A memory unit of 2 ports (consisting of a read port and a write port) that can be operated independently and in parallel is used for the shared memory 2006 in the system. As a result, the read cycle to the shared memory and the write cycle can be executed in parallel, the latency required for the data match processing between the shared memories and the data transfer processing between the processors can be shortened, and the access between the processors can be shortened. Since the loss due to competition can be significantly reduced, the total output to the shared memory system can also be improved.
2) When the synchronous processing circuit 1000 between processors and the shared memory system are operated in conjunction with each other, the information generated by the target task between the tasks managed by the synchronous processing is transmitted via the shared memory. To ensure reliable communication, the processor read cycle is the period from when the synchronization processing circuit notifies the completion of synchronization until the information in the shared memory is actually rewritten to a valid state for the purpose. It is equipped with an interlock circuit for local synchronization that keeps you waiting. As a result, synchronous processing between processors can be performed reliably and consistently in a form that guarantees the validity of passing data between tasks, and automatic processing is automatically performed so that the processor does not obtain old information and perform erroneous processing. Can be managed.
Next, using the simplified embodiment shown in FIG. 3, the two-port shared memory 2006 in the shared memory systems 1010, 1011, ..., 101n of each processor and the peripherals related to the control thereof. The function of the circuit will be described in more detail. In particular, here we will describe the functions and effects of the two-port system.
[0099] The 3004,3005,3006,3007,3008 shown in FIG. 3 correspond to the functions of 2004, 2005, 2006, 2007, and 2008 in FIG. 2, respectively. There may be a lead data latch 3110, but as already mentioned, Figure 2 states that this feature exists within PIF2003. Although the inside of the memory unit 3006 is disclosed in detail in Fig. 3, its peripheral functions are simplified for the sake of simplicity.
[0100] First, during the read process from the processor, the lead address 13001 from the processor side is directly input to the lead address recorder 3103 of the memory unit 3006, and in response to the output from the lead address recorder 3102, the multiplexer 3102 By switching the selection input S, the output corresponding to the specified address is selected by the multiplexer 3102 from the memory cell group 3101 and output to the processor side as RDATA 13003. The multiplexer 3102 may be configured by combining a tri-stage buffer as shown in FIG.
[0101] During the write process from the processor to the shared memory, the write address 13002 is output to the shared bus side via the buffer 3008 to be generated as the address information of the shared bus system, and the processor is directly passed through the multiplexer 3007. It is input to the WADDR decorator 3104 of the memory unit 3006 of. The buffer 3008 may be configured as a queue system that stores address data in chronological order, and may wait for the same function as in 2008 in FIG.
[0102] The write address directly input to the memory unit 3006 is coded by the WADDR recorder 3104 to determine which memory cell in the memory cell group 3101 to write the data to, and the write signal is used. In response to WE, latch the contents of the data WDATA to be written to the selected memory cell. The WE signal is generated by the control unit 3010 in response to signals such as lead / write control signals W / R13005 and shared system select CSEN13006 from the processor. The write data from the processor 13004 is input into the memory unit 3006 as WDATA after passing through the multiplexer 3005.
[0103] Like the write address 13002, the write data 13004 is also output to the shared bus side via the buffer 3004 having the same function as the buffer 3008. The write address and write data output to the shared bus side are broadcast to the shared memory system of another processor via the shared bus, and the write data is latched to the memory cell of the corresponding memory unit.
[0104] In this example, the path of data and address information when the processor writes data to its own memory unit 3006 is directly input to the multiplexer 3005,3007 from the front of the buffer 3004,3008 (A input of the multiplexer). ), And when using only that path as the write path to the memory unit 3006, the signal after passing through the buffers 3004 and 3008 as described in Fig. 2 was used. The control method and conditions are slightly different from the path.
[0105] However, it has already been described that even in FIG. 2, the path is designed to be the same as that in FIG. 3 if the C input of DMX2005 and AMX2007 is selected.
[0106] The advantage of this direct input method is that when the processor rewrites the contents of the shared memory, the data of its own memory unit (shared memory) is changed at an earlier timing than the change of the memory unit of another processor. If there is a possibility that it can be done and the contents of the shared memory that you changed are read again immediately after the change (flag management, semaphore management, holding of shared data that you also use, etc.), it accompanies the change of the shared memory. The point is that it becomes easier to prevent past data from being read due to latency (delay time).
[0107] However, it is easy to control writing using only this direct path in the buffers 3004 and 3008 in about one stage, and the overhead required for the writing process to the shared memory of another processor is recovered and the overhead is recovered first. This is the case when it is used for temporary storage to advance the bus cycle of the processor. When a buffer with such a function is provided, it is necessary to control the control unit 3010 to write data to the memory unit 3006 immediately after confirming the permission signal from the arbiter circuit.
[0108] Further, it is preferable to provide a latch function in PIF2003 of FIG. 2 for holding the write address 13002 and the write data 13004 until the data is written to the own memory unit 3006. It may be clearer to consider such a buffer function as a part of the function of PIF2003 in Fig. 2 and to provide it in the PIF. This is because the buffer latch function can be turned on and off only by the acknowledge CSACK of the shared bus system, and the CSACK signal can be input to PIF2003 without having to control it by the control unit 3010. ..
[0109] When a full-scale buffer is provided, it would be more rational to provide a path (C input) via the buffer in the multiplexers 3005 and 3007 and switch the control as in FIG.
[0110] When each processor reads the contents of the data changed by itself again on its own processing program and uses it, the hardware automatically processes the contents consistently and consistently. Guaranteeing consistency is important. This is because most processors process programs that are supposed to be written and executed sequentially by themselves, and have meaning in the context of data rewriting and data reading operations for resources. This is because they often have them.
[0111] On the other hand, when one processor reads the data changed by another processor, the difference between the time when the other processor changes the data and the time when the processor actually reads the data (information). Information that does not matter (delay time) or how to handle it, for example, assuming that the state quantity (position, speed, acceleration, etc.) depending on the time t with continuity is managed with the sampling time as the minimum time unit, the above-mentioned information delay If the state quantity can be handled by assuming that the time is relatively sufficiently small with respect to the sampling time, or if the sampling time side can be set sufficiently large with respect to the information delay time, the shared memory seen from each processor. It can be considered that there is no problem even if the change time of the above information is slightly different or delayed.
[0112] However, if the sampling time is set so small that the delay time (latency) of the information cannot be ignored, the error of the processing parameter becomes large and becomes a problem. Therefore, when executing an application with a small sampling time that requires real-time performance in this way, a hardware architecture that improves information delay (latency) is required.
[0113] A processor system having sufficient real-time performance (real-time processing capacity) is associated with latency (communication delay between processors and between an external system and a processor) and arithmetic processing time that occur in various places in the system. This performance characteristic is most important for the control processor system, which is a system in which the sampling time is kept small enough in principle with respect to the target sampling time (delay etc.).
[0114] In a real-time processor system having such characteristics, most of the state quantity information does not need to be synchronized between processors associated with data communication by handshake processing or the like, and synchronization is not particularly managed. Sufficient accuracy of processing results can be ensured by information transmission. If reliable data transfer between processors is required, information management may be performed on the shared memory system combined with the synchronization processing circuit 1000 as described above.
[0115] As described above, as a circuit that guarantees that the processor can read and write the contents of its own memory unit 3006 without any contradiction in the flow of its own program processing, in this example, an address comparison circuit 3020 is provided and a write address is sent from the processor. The write address stored in 13002 and the buffer 3008 is taken into the W0 and W1 inputs, respectively, and the lead address 13001 from the processor is taken into the R input, and the write address is taken during the lead cycle from the processor to the memory unit 3006. Comparing the contents of W0 and W1 with the lead address R, if there is even one that matches, the processor unit 3010 outputs to 13007 until all the write cycles to the shared memory for those write addresses are completed. It keeps the RDY-N signal inactive and puts the lead cycle on the processor side to wait.
[0116] In FIG. 3, the operation during the write cycle from the other processor to the shared memory is the same as in the case of FIG. 2, and the selection signal input is performed so that the B input side of the multiplexers 3005 and 3007 is selected, respectively. The controller unit 3010 controls S using the WDSEL signal output and the WASEL signal output, respectively. The buffers 3004 and 3008 are controlled by the control unit 3010 using the CSDTL signal output and the CSADL output, respectively.
[0117] The free state of the buffers 3004 and 3008 is managed by providing a circuit in the control unit 3010 for holding the number of data of the buffer free by counting the increase or decrease. Of course, this function may be provided on the buffer side, and the information from the function may be taken into the control unit 3010 side. The functions of the other input / output signals of the control unit 3010 can be considered to be equivalent to the corresponding input / output signals of the shared memory control unit 2010 shown in FIG.
[0118] Next, FIG. 4 shows an example of the recorders 3103, 3104 and the memory cell 3101 in the memory unit 3006 of FIG. Here, one of the memory cell groups in the memory unit 3006 is shown.
[0119] In order to configure a memory unit, the memory cells 3101 are prepared for the number of data bits, and a plurality of sets thereof are prepared so that they can be specified by an address value (WA, RA). The multiplexer section 3102 may prepare as many tri-stage buffers as the required number of data bits corresponding to each of the plurality of memory cells, and may provide a plurality of the sets as many as the number that can be expressed by the address value. Note that the outputs (ZN) corresponding to the same data in each set are connected. Any one of the plurality of data sets can be selected by specifying the lead address value RA.
[0120] The lead decorator 3103 obtains a lead address RA and a lead enable RE if necessary, and uses the lead address value of the enable signals (EN0, EN1 ---) as the lead address value. Activate one of the corresponding ones.
[0121] The lead data input (RD) of each tri-stage buffer unit 3102 receives the enable signal (active at one level), and if it is active, the contents of the memory cell 3101 are input. Output to ZN (OUTPUT) and keep Z in the float state if inactive. When the RD input is 1, the contents of the memory cell (the value of the D input stored in WR as a trigger signal) are inverted and output to Z, and when the RD input is 0, the tristate type multiplexer section 3102 is described above. ZN is in a floating state.
[0122] The lead enable RE (active at the 1st level) is normally used as an enable signal by making the enable output active immediately after the lead address is fixed and the decoration itself is completed. It has the role of preventing the hazard from riding. If the hazard is large, it may be entangled with the skewer on the wiring, etc., and the tri-stage output Z that is connected may temporarily be short-circuited between the two, but the hazard If is small, there is no particular problem even if the RE signal is eliminated.
[0123] In the examples shown in FIGS. 2 and 3, the leadable RE is not particularly provided. As shown in FIG. 3, if the multiplexer 3102 part adopts a complete multiplexer structure, the lead address RA value itself or a signal equivalent thereto can be directly used for the selection input S. It is possible.
The write decorator 3104, like the lead decorator 3103, is an active write signal WR (active at one level) of a set of memory cells corresponding to one enable signal EN indicated by the write address WA. To generate. The enable signal consists of EN0, EN1, ---, and each enable signal corresponds to each set of memory cells and is connected to the write signal WR input of each set.
[0125] The light decorator 3104 uses a light enable WE signal and is an enable signal (active at one level) that is output when the specified light address WA is determined and the decoration is completed. The proper pulse width to write the WR signal to the memory cell by masking the enable signal except during the period when the WE (active at 1st level) is active so that no hazard occurs. Only surely give to the desired set of memory cells.
[0126] The structure of the memory cell 3101 shown in FIG. 4 is based on a CMOS process, and uses a transfer gate (also referred to as a transparent gate) type 2-input 1-output multiplexer. , The forward output OUT is fed back to one input IA of the transfer gate, data (D) is given to the other input IB, and the selection signal (transistor base input signal) S is used. A gate latch is constructed by giving a WR signal.
That is, when the WR is at the 1st level, the value of the D input is transmitted, and the value is latched at the falling edge of the WR. The value of the D input stored with WR as a trigger signal is inverted and output to the ZN output of the tri-stage buffer 3102 at the time of reading.
FIG. 5 shows the timing of the shared memory access of the present invention. It is assumed that the processor clock (PCLK), which is the standard for the bus protocol on the processor side, and the shared bus clock (BCLK), which is the standard for the bus protocol on the shared bus side, have the same frequency and phase, and both are system standards. It is generated in synchronization with the system clock (SCLK). The frequency of SCLK is twice the frequency of PCLK and BCLK. As described above, a broadcast-type shared memory system is premised, and the shared memory unit in each shared memory system corresponding to each processor is a lead port, which is a feature of the present invention. It is said that it uses a 2-port memory unit that has an independent write port.
[0129] In FIG. 5, access from the shared bus side (one of the processors generated a write cycle to the shared memory system, and the information was blowcasted through the shared bus system). It shows the conflict situation with the lead cycle to the shared memory from the processor side. The status of the bus signal on the write port side on the shared memory is shown in the memory write data (MWD), memory write address (MWA), and memory write enable (MWE). The status of the bus signal on the to-side is shown in the memory read data (MRD), memory lead address (MRA), and lead enable (MRE).
[0130] The bus cycle on the processor side is a 2-processor clock (2 × PCLK), and the access cycle required on the shared memory is also substantially required for 2 × PCLK cycles. However, the actual access time to the shared memory is about 1.5 x PCLK cycle, and the remaining 0.5 x PCLK cycle is required for data hold time, timing adjustment time, processor setup time, etc. It is assumed that it is a good time.
[0131] In FIG. 5, the processor starts a read cycle (lead cycle) to the shared memory in synchronization with PCLK at the beginning of the state S1 to generate the processor address (PA), and the state S2. Generates a processor shared command signal (PRD-N) that commands the read processing of the target data in. On the other hand, on the shared bus side, a visit signal (CSBUSY-N) indicating that the shared bus address (BA) and the write cycle from the shared bus side have become active in synchronization with BCLK at the beginning of state S0. Alternatively, a shared bus write signal (BWT-N) and a shared bus data BD to be written to the shared memory are generated.
[0132] In the present invention, BD and BA are output from the shared memory system 101n of another processor to the shared bus system at almost the same timing, and CSBUSY-N is a BUSY signal from the abiter circuit 1020 (active at 0 level). It is generated and output at a timing slightly earlier than BD and BA. CSBUSY-N is a signal indicating that a shared bus is used, and is one of the control signal information of the shared bus system as explained in FIG.
[0133] By providing the 2-port shared memory, the lead cycle to the shared memory on the processor side is completed at the end of the state S2 without waiting, and at this time, with the BA from the shared bus side. If the processor specifies the address on the same shared memory by PA, the value of BD should be read to the processor as it is, and the processor at the last point of S2 in response to the timing of the RDY-N signal. Read to the side (the read period from shared memory is while the RE is active).
[0134] The write cycle from the shared bus system side is started one PCLK cycle ahead of the lead cycle on the processor side even on the shared memory, that is, without waiting in the state S0, and is started at the state S1. A write command (WE) on the shared memory has been generated, and the BD corresponding to BA has already been confirmed on the shared memory in S1. In this example, when the WE is at the 1st level, the BD value is transmitted to the address corresponding to the BA on the shared memory, and the BD value is latched at that address when the WE falls.
Therefore, when the processor reads the data of the address corresponding to BA after S1, the BD can be read as described above. As shown by the dotted line in Fig. 5, if the processor starts access at the beginning of stage S0, the write cycle from the shared bus side and the lead cycle from the processor side are completely at the same timing. That is, RE and WE are output at the same timing, but the cycles on either side are processed in parallel without waiting, and are completed in the shortest time.
[0136] As described above, according to the present invention, the bus cycle on the shared bus side and the bus cycle on the processor side on the shared memory can be completely processed in parallel, and the data communication between processors having a very short latency is shared. It can be realized on the memory.
[0137] As can be seen from the above, the present invention using the 2-port shared memory has an effect of improving access flow on the processor side and an effect of significantly shortening the data communication latency between processors via the shared memory. And can be obtained. In order to avoid some inconsistencies regarding data transfer between processors, interlocking the bus cycle on the processor side and the bus cycle on the shared bus side as already explained in Fig. 2 and Fig. 3 should be used. Local synchronous processing and parallel processing management between processors linked with the inter-processor synchronous processing circuit as shown in FIG. 1 may be performed.
Next, both PCLK and BCLK are synchronized with SCLK, which is the same as the embodiment shown in FIG. 5, but the timing of the embodiment when the 2-port shared memory is not used as in the present invention. The chart is shown in Figure 6. The access conditions on the processor side and the shared bus side are exactly the same as in FIG. The broadcast-type shared memory system disclosed in Japanese Patent Application Laid-Open No. 5-2568 is the type of this example.
As can be clearly seen from FIG. 6, since there is only one set of the shared memory address (MA) and data (MD), first, the address BA on the shared bus side near the center of the state S0. The data BD becomes active for the shared memory and occupies the shared memory up to the center of the state S2. The write enable WE is output at the same timing as in Fig. 5, and the BD value is latched on the shared memory at the beginning of the state S2.
[0140] Basically, in this embodiment, competition control is performed in which the write operation from the shared bus side is prioritized over the lead operation on the processor side. Therefore, when the write bus cycle on the shared bus side and the lead bus cycle on the processor side compete with each other, the bus cycle on the processor side waits until the bus cycle on the shared bus side ends. In the example shown in Fig. 6, the bus cycle on the processor side that came to read access to the shared memory at stage S1 is waited for one stage (PCLK cycle), and is shared at the end of stage S3. The bus cycle is terminated after obtaining the data PD in the memory.
[0141] When viewed on the shared memory, the write cycle (BA, BD, MWE active) on the shared bus side is executed for two cycles from the vicinity of the center of the state S0 to the vicinity of the center of the state S2, and immediately after that. A read cycle (PA, PD, MRE active) on the processor side is being executed.
[0142] The shared memory controller unit responds to the period when CSBUSY-N is active by allocating the bus cycle on the shared bus side to the shared memory, and responds to the timing when CSBUSY-N turns to inactive (Hi level). Then switch to the bus cycle on the processor side. As shown by the dotted line, when the bus cycle on the processor side is started at stage S0, the waiting time of the bus cycle on the processor side increases to 2 stages (2 x PCLK cycle). The access overhead on the processor side increases.
[0143] From the above, as compared with the example of FIG. 5, the latency of the access overhead on the processor side and the latency from the shared bus side to the processor side via the shared memory are increased by 1 to 2 PCLK cycles. I understand.
[0144] Fig. 7 shows a system in which PCLK and BCLK, which are generally used conventionally in a broadcast-type shared memory system, are in an asynchronous state. Generally, the phases of the respective PCLKs corresponding to each processor are not synchronized with each other, and when the types of processors are different, their periods are often different. Other conditions are the same as in FIGS. 5 and 6.
[0145] In such a system controlled by using an asynchronous reference clock between processors and between a processor and a shared bus system, asynchronous synchronization processing is performed in various places to generate meta states at various levels. Will need to be avoided. In this example, the bus cycle on the shared bus side is synchronized with PCLK, and the data write timing to the shared memory is controlled so that the correct relationship with the access timing on the processor side can be maintained. That is, as a result, the process of synchronizing the bus cycle of the shared bus side, which is originally asynchronous with the bus cycle of the processor side, is performed.
[0146] Actually, the CSBUSY-N and BWT-N signals are synchronized with PCLK by passing them through two or more flip-flops with PCLK as the trigger clock, and the signals are synchronized with PCLK. -With a bar head. With this overhead, the write cycle from the shared bus side on the shared memory starts from the beginning of stage S2, and ends after 2 stages. The end of S3. When the synchronization is completed, the information is notified to the shared bus side in some way, and the original shared memory system that issues the shared bus cycle uses the information to perform the termination processing of the shared bus cycle. In this example, the BSYNC-N signal is returned to the shared bus system as synchronization information of the shared bus cycle.
The shared memory system of the processor generating the shared bus cycle responds to the timing when BSYNC-N becomes active (assuming that the 0 level is the active level and the signal change timing is synchronized with PCLK). Ends the bus cycle being output to the shared bus system. Here, a synchronization signal is internally generated from BSYNC-N, which is asynchronously synchronized using BCLK to synchronize with BCLK, and then in response to the change timing, the bus cycle, that is, BD / Float the output of BA and deactivate BWT-N or CSBUSY-N.
[0148] From this, it can be seen that the shared bus system side also generates overheads for 1 to 2 BCLK synchronization.
[0149] As a result, in the example of FIG. 7, the lead cycle on the processor side is terminated by the BD, BA, MWE signals of the state S3 in which the bus cycle on the shared bus side is terminated on the shared memory. It is the last point of S5 two more stages after the end (the MRE signal corresponding to that lead cycle is active at the beginning of S5 and inactive at the beginning of S6). That is, there is a waiting time of 3 states (3 × PCLK cycle) on the processor side.
[0150] Compared with the access timing of the present invention shown in FIG. 5, the present invention also has an overhead on both the processor side and the shared bus side, and also in terms of communication latency via the shared memory between the processors. It turns out that is far superior to the conventional system.
Next, FIG. 8 shows an embodiment of a ready signal generation circuit linked with a synchronization signal (SYNCOK), which is a major feature of the present invention. As described in FIG. 2, this circuit is used to exchange data on the shared memory consistently between processors that execute related processing when operating in conjunction with the interprocessor synchronization processing circuit 1000. Realize the local synchronization function.
[0152] As already described for the detailed function, even if the synchronization processing in the interprocessor synchronization processing circuit 1000 is completed and the SYNCOK signal becomes active after the task processing by the processor is completed, due to the communication delay. There is a possibility that the information required for the next task processing does not exist in the shared memory. This circuit avoids the situation and further local synchronization processing (interlocking) between the shared memory control unit 2010 and the synchronization processing circuit 1000 to ensure that the required information is obtained from the shared memory. Process) is executed and managed so that there is no contradiction in the context of data transfer on the shared memory.
[0153] Basically, as described above, if CSBUSY is active when SYNCOK is activated, the processor is prohibited from reading the contents of the shared memory until CSBUSY is once deactivated. Specifically, when the above conditions are satisfied, the lead cycle to the shared memory on the processor side is extended by keeping the RDY-N signal inactive to make the processor side wait, and the interlock process is performed.
[0154] When each processor executing the related processing completes the synchronization processing and SYNCOK (active at the 1st level) becomes active, the processor shares at least the necessary processing result in the task already processed in the shared memory. It should have finished issuing all write cycles on the processor side to store to.
Therefore, all the actual write cycles to the shared memory corresponding to the already issued write cycle on the processor side held in the shared memory system of each processor (in the write buffer 2004, 2008, etc.) are all. Until it is completed, that is, until each processor can obtain the information of the matching contents appearing on the shared memory 2006 of all the processors, the information for matching the contents of each shared memory is shared. A write cycle for broadcasting to all shared memory systems corresponding to each processor via the bus system continues to be generated on the shared bus system.
[0156] In response to this, CSBUSY-N, which indicates that the shared bus system is active, is also kept active, so that interlocking is possible by the above-mentioned logic. The logic of this interlock function will be described in detail below using the example of FIG.
It has already been mentioned that CSBUSY-N is an active signal at 0 level and is generated in response to a BUSY signal from the arbiter circuit 1020, which is one of the control signals on the shared bus. .. This is connected to one input RN of the RS flip-flop 8000 via the inverter 8001, and when the CSBUSY signal is inactive (initial state is inactive), 1 is unconditionally output to ZN. , 0 is output to Z relatively. The above state is the initial state.
The output of the NAND gate 8006 is connected to the other input SN, and when the CSBUSY signal is in the active state, the NAND gate 8006 shows the rising edge at which the SYNCOK signal turns active. Detected from the signal of circuit 8005 and the state of the SYNCOK signal, a pulse (Lo pulse) is generated at the SN input of the RS flip-flop 8000.
[0159] When the pulse is generated, the RS flip-flop 8000 is set and 1 is latched on the Z output. However, if the SYNCOK signal is set to 0 level in the initial state, the NAND gate 8006 will be in the state of outputting 1 level, and the CSBUSY signal will have 0 level as the initial value, so the Z output of the RS flip-flop 8000 will be The state is reset to 0, which is consistent with the above initial state.
The NAND gate 8002 has an active CSEN signal indicating that the shared memory has been accessed, the RS flip-flop 8000 is set by a pulse from the NAND gate 8006, and one level is output to Z. And when the CSBUSY signal is active, 0 level is output. This unconditionally deactivates the NAND gate 8003 in the latter stage to one level, that is, deactivates RDY-N, so that access to the shared memory from the processor side is temporary when the interlock conditions are met. Acts to prohibit.
[0161] In this example, the interlock function is designed to operate in both the lead cycle and the write cycle when the processor accesses the shared memory, but only during the lead cycle. If you want to activate the operation, you can decorate the NAND gate 8003 with the condition that the leadable signal from the processor is active (the leadable signal RE with an active level of 1). To the input of the NAND gate 8002).
Note that the NAND gate 8003 should only return the active RDY-N signal to the processor when the CSEN signal is activated, that is, when the processor accesses the shared memory system. It has become.
[0163] Further, the NAND gate 8009 can be used if the shared memory system is accessed, CSEN is at the active level (1 level), and SYNCOK is at the inactive level (0 level) (SYNCOK). Outputs 0 level (connecting the signal-inverted signal to the input of NAND gate 8009), which drives the input of NAND gate 8003 and unconditionally outputs the RDY-N signal. Is set to the inactive level (1 level).
[0164] That is, when the processor accesses the shared memory system, if the synchronization processing for the processor is not completed in the synchronization processing circuit 1000, the operation of accessing the shared memory on the processor side is made to wait. This is equivalent to the local synchronization function disclosed in Japanese Patent Application Laid-Open No. 5-2568. Of course, the NAND gate 8009 may be designed to detect the active state of the readable signal (RE) so that this function works only when a read operation to the shared memory occurs. ..
[0165] As described above, the interlock function of the present invention is linked to the conventional local synchronization function disclosed in Japanese Patent Application Laid-Open No. 5-2568, and is synchronized between processors linked with the shared memory system of the present invention. It can be seen that the processing is supported.
[0166] The condition that the interlock is released and the local synchronization process of the present invention is completed is that the CSBUSY signal turns inactive. When the CSBUSY signal becomes inactive (0 level), the output of the NAND gate 8002 becomes 1 unconditionally, the Z output of the RS flip-flop 8000 is also reset to the 0 level and returned to the initial state, and the interlock is released. It will be released. The SYNCSEL (active at 1st level) input to the NAND gate 8002 is a selection signal that determines whether or not to enable (active) the local synchronization processing function by this interlock circuit.
[0167] As shown by the dotted line in FIG. 8, the CSBUSY signal may be used after passing through several stages of flip-flops using PCLK as a trigger clock. FIG. 8 discloses an example in which the flip-flop 8004 is used in one stage. This makes it possible to remove any hazards on the CSBUSY signal. Also, by delaying the CSBUSY signal for an appropriate amount of time in this way, the interlocking time is until all the necessary data has been written to the shared memory, that is, the lead port. -It is possible to set enough time to cover the time until they can be read reliably from the interlock side.
[0168] In this example, by delaying the CSBUSY signal by 1PCLK cycle, the interlock period is shifted by 1PCLK cycle from the time when the original CSBUSY signal becomes inactive. The control unit is designed so that reading of all necessary data is enabled on the shared memory immediately after the interlock is released. That is, in the case of this embodiment, the last write cycle on the required shared memory is generated before the time when the interlock is released, the RDY-N signal becomes active, and the bus cycle of the processor ends. It suffices if the signals such as light enable (WE), light address, and light data (WD) are active.
[0169] Considering the timing at which the original CSBUSY signal becomes inactive, in this embodiment, one of the two states (1 stage = 1 PCLK period) immediately following that timing. It suffices if the last write cycle on the shared memory is generated at this rate.
As shown in the figure, if the OR logic of the output of the flip-flop 8004 and the CSBUSY signal is taken using the OR gate 8007 and the output is used instead of the output of the flip-flop 8004, the interlock The release time can be set to be almost the same as when the flip-flop 8004 is not used, while keeping the release time almost the same as when the flip-flop 8004 is used. This allows the signal 8008 (to the CSBUSY signal) for the input of the gate 8006 and the RN input of the RS flip-flop 8000, even though there is a write cycle to the shared memory pending when SYNCOK turns active. It is easy to design the state of the signal obtained in response) so that it is not yet active (1 level) at that time due to signal delay.
[0171] FIG. 9 shows another embodiment regarding the access timing to the shared memory of the present invention. In the example of Fig. 5, the access status on the shared memory between the processor side and the shared bus side is shown, but as a condition of the bus cycle, the minimum time of 2 processor cycles (2 × PCLK cycle) per bus cycle is shown. Was assumed to be necessary. In the example of FIG. 9, it is assumed that each bus cycle on the processor side, the shared bus, and the shared memory is configured with a minimum of 1 processor cycle (1 PCLK cycle).
[0172] However, a pipeline bus cycle (the address bus and the data bus are each independently driven in one processor cycle) in which the address is output in advance and the data is input / output in the subsequent states. , And they are offset by one processor cycle from each other) so that the address access time can be secured for a relatively long time.
[0173] The characteristic of the pipeline bus cycle of this embodiment is that after the address (ADDR) for the bus cycle to be output next is output for at least one rate (1 PCLK cycle), the RDY for the previous bus cycle is output. If the -N signal is returned and the bus cycle is completed, the next address value (if the processor is already ready) is output.
[0174] Regarding the exchange of data corresponding to a certain address, input / output to / from the processor is executed at the end of the next stage of the stage at which the address is output. That is, the operation of the data bus is delayed by one status with respect to the operation of the address bus, and the last point of the status at which the data is exchanged with the processor is the above-mentioned data. It is the end of the bus cycle for the address.
[0175] If the processor is in a state where it can output the next address, it is already possible to output to the address bus of the next address at the output state of the data, and the data Since the address can be output in advance in a pipeline one after another in parallel with the input / output, it is called a pipeline addressing or a pipeline bus cycle.
[0176] For example, in the processor (A) of FIG. 9, the address (ADDR) A1 has already completed two or more previous bus cycles at the state S0, so that the processor can prepare the address value A1. It is output immediately at status S0, and since it is a write cycle, data D1 that should be written to the outside from the processor is subsequently output at status S1.
[0177] Since the previous bus cycle has already ended in the state S1, the processor outputs one state of A1 and immediately parallels the output D1 of the data corresponding to A1 to the next. The address A3 (reading cycle) is output. The RDY-N signal for data D1 is taken into the processor (A) at the last point of state S1, and in response to this, the next address information A5 (write cycle) for state S2 is data for A3. -Input operation not Not D3 is being executed in parallel.
[0178] As described above, even if the lead cycle and the write cycle are mixed, the address bus and the data bus are driven in parallel and pipeline for each processor cycle unit, and the address bus and the data bus are substantially driven at one rate / bus. It is possible to realize a cycle.
[0179] In the embodiment of FIG. 9, instead of reliably giving and receiving task-based data in conjunction with the inter-processor synchronization processing circuit, information can be freely transferred between processors without managing synchronization. The access status on the shared memory system when exchanging is shown by taking the case of two processors (processors A and B) as an example. In order to express the access to the shared memory in detail with a small space, all the cycles are read or write cycles to the shared memory, and to other general resources, interprocessor synchronization processing circuit 1000, etc. Access is assumed to be running in parallel with this.
Such a processor system includes a dedicated processor in addition to the main processor, a bus system for accessing a shared memory system, and a bus system for accessing other general resources. This can be achieved by constructing an advanced processing system that has the above separately.
[0181] Data D1, D2, D5 are information written to the shared memory from the processor (A) side, and data D3, D4 are information written to the shared memory from the processor (B) side, and are read. Is also one of those data. If the required data has not yet been written to shared memory and the previous data (previously written data) can be read, prefix it with "not", such as not Dn. It is an expression.
[0182] Whether it is a lead cycle or a write cycle is indicated by an RD / WT signal (Hi-RD, Lo-WT) output at almost the same timing as the address (ADDR). When a write cycle is generated, information (address, data, BUSY signal, etc.) for rewriting shared memory via the shared bus is broadcast to all shared memory systems and on the shared memory of each processor. Will generate a light cycle and change its contents.
[0183] In the bus cycle of the processor in the figure, (W) is attached to the write cycle to the shared memory, and (R) is attached to the lead cycle. The bus cycle on the shared bus is inevitably only the write cycle, and the CSBUSY-N signal (Lo active) is active during the period when the bus cycle is generated.
[0184] On the shared bus, the address information and the data information are output with a deviation of about 1/2 rate from the address and data output timing on the processor side, and the CSBUSY-N signal is output on the shared bus. It is output at almost the same timing as the address. The state of the bus protocol on the shared bus is managed completely in synchronization with the processor clock (PCLK). Specifically, for example, the address A1 output at the state S0 on the processor (A) side is Data D1 output from the beginning of the status S1 on the processor (A) side is output from the beginning of the status S1 on the processor (A) side, and the data D1 is output from the beginning of the status S1 on the shared bus. One rate is output from near the center of.
[0185] The bus protocol state on the shared memory is also managed in synchronization with the processor clock (PCLK), and in FIG. 9, the write data / report corresponding to each of the processors (A) and (B) is mainly used. It shows the state of the processor and the timing at which the write address and light enable (WE) are generated.
Next, the bus cycle on the shared memory will be described in detail. When the CSBUSY-N signal on the shared bus turns active, the address information on the shared bus is gated so that it is valid for the shared memory near the center of the stage, and the write port of the shared memory. The light address (WA) is given to the computer for about one status period. To obtain this timing, use it as a PCLK inversion (PCLK-N) clock and use a gate latch to gate the contents of ADDR on the shared bus at the center of the state, and keep the information transparent. -Latch at the end of the clock, keep it for about 1/2 rate, and then give it to the shared memory as a write address (WA).
On the other hand, for the write data information from the shared bus side, the contents of DATA on the shared bus are gated at the beginning of the status by a gate latch using PCLK as a clock, and the information is transmitted. Then, latch it in the center of the state, keep it for about 1/2 state, and then give it to the shared memory as write data (WD).
[0188] The write enable signal responds to the CSBUSY-N signal so that the write data (WD) is activated in the shared memory by about 1/2 rate from the beginning of the state, that is, when the write data (WD) is given to the shared memory. Generate. While WE is active, the timing is adjusted so that the WA value is valid for shared memory.
[0189] In the present invention, it is assumed that the write cycle to the shared memory is generated in common and substantially the same for the shared memory system of all processors.
The lead cycle on the shared memory is executed in parallel with the write cycle using the lead port as described above, and depends on the bus cycle of the corresponding processor on each shared memory. A completely separate cycle is generated.
[0191] In FIG. 9, WT (A1) / RD (A3), WT (A1) / RD (A1), WT (A5) / RD (A4), WT (A3) / RD (A3), WT (A4). ) / RD (A4), WT (A4) / RD (A1) The lead cycle and the write cycle occur in parallel in the cycle. In addition, (R) is attached when only the lead cycle occurs, and (W) is attached when only the light cycle occurs.
[0192] As more detailed information, RD (Ax) in the case of a lead cycle, WT (Ax) in the case of a write cycle, and a write cycle and a lead cycle occur in parallel in the upper or lower stage of each cycle. In case, it was displayed as WT (Ax) / RD (Ay). Ax and Ay are the address information corresponding to the data information Dx and Dy to be written sent from the processor.
[0193] The input / output directions of the data are opposite to those of the write cycle when the address (RA) and data (RD) are generated on the shared memory in the lead cycle (read is). It can be considered to be almost the same as the case of the write cycle except that the write receives and receives data from the processor to the processor, but when the readable (RE) is present, it turns active. The timing is from the beginning or near the center of the stage where the processor RD becomes active to the last point of the stage.
[0194] Here, let us consider the timing until the data content on the shared memory is changed in response to the write cycle from the processor and the data can be actually read from another processor. In this example, the processor (A) side reads the data D3 corresponding to the address A3 twice (the address output is started at S1 and S4, respectively), and reads the data D4 corresponding to A4. The operation is performed twice (address output is started at S3 and S5, respectively).
However, it is at stage S3 that the processor (B) generates a bus cycle to change the contents of address A3 in shared memory, and data D3 is at the beginning of stage S4. It is output from the processor (B) at, D3 is actually enabled on the shared memory, and it is possible to read from the processor (A) side at the beginning of the state S5 with write enable (WE). From the time it became active.
Similarly, it is at stage S4 that the processor (B) generates a bus cycle to change the contents of address A4 in shared memory, and data D4 is at stage S4. It is output from the processor (B) at the beginning, D4 is actually enabled on the shared memory, and it is possible to read from the processor (A) side when the write enable (WE) is activated at the beginning of S6. Because.
[0197] The processor (A) reads the data D3 corresponding to the address A3 at each of the S2 and S5 states, and reads the data D4 corresponding to A4 at each of the S4 and S6 states. In the state S2, the value of D3 that the processor (B) is trying to rewrite cannot be read, and the contents of the previously set address A3 on the shared memory can be read. Regarding the value of D4, the processor (A) is reading from the shared memory at each stage of S4 and S6, but in S4, the value of D4 that the processor (B) is trying to rewrite cannot be read and is set before. The contents of address A4 on the shared memory can be read.
[0198] Then, in the state S5 in which the D3 from the processor (B) is reflected on the shared memory, the actual value of the D3 can be read by the processor (A), and similarly, the processor (B) can read it. The set actual D4 value can be read by the processor (A) at the state S6.
[0199] From this example, the latency from when the processor (B) outputs data to when the processor (A) takes in the data via the shared memory is 2 states. Understand. Of these, the latency on the shared memory system side is 1 rate (1 PCLK cycle), and it can be seen that a very efficient data sharing mechanism using shared memory has been realized.
As can be seen from FIG. 9, it can be seen that the lead cycle and the write cycle on the shared memory are operating in a state where access contention does not occur completely. In addition, the information exchanged on the shared bus by the broadcast method is only the write cycle, and the bus cycle on the shared memory loses access time due to access conflicts and overheads. Since it has not occurred, both the shared bus side and the processor side can end the bus cycle in one stage.
[0201] This indicates that the most efficient shared memory system can be provided theoretically. For example, the processor system having the bus efficiency of this example has an application that brings about the most severe state for a shared memory system in which all data is input and output via the shared memory system. It is considered that the shared memory system of the present invention has a level of capability that can process up to three processors without any overhead even when executed on the above.
[0202] It is known that the average ratio of write cycles in all bus cycles is about 30%, and if the number of write cycles for three processors, they all go to shared memory. Even if the access is, the shared bus system of the present invention can sufficiently absorb (since the shared bus system supports only the light cycle, the frequency of occurrence of the light cycle on the system determines its performance). Because it is on the level.
[0203] Next, when the shared memory system is used, how many processors can be effectively connected in an actual multiprocessor system will be examined.
[0204] By using this shared memory system, it was shown in the above-mentioned study that up to three processors can be connected with almost no effect on the performance of the system. However, this study assumes that each processor in the system randomly accesses the shared memory system under near-worst conditions, and the processor always uses shared memory in every processor cycle. It is not realistic because it is assumed that you are accessing it.
[0205] Actually, the internal instruction processing of the processor (for example, inter-register operation) is processed by an average of 1 processor clock, and the processing involving access to the outside such as memory is the best 2 processor clocks (of which 1 processor clock is external). Assuming that 50% of the instructions are processed with external access, the average processing time per instruction is 1.5 processor clocks.
That is, the average number of processor clocks required to access external data per instruction is 0.5 clocks, and the ratio to the total bus band (which is assumed to be accessible with one processor clock per data) is It is 33% (0.5 / 1.5 x 100%). In a tightly coupled multiprocessor system, 10% to 30% of the instructions that accompany external access are access to the shared memory system in a general application.
In loosely coupled systems, access to shared memory systems is often less than 1%, but traditional systems still lose system performance due to communication overheads and access conflicts. Is at a level that cannot be ignored. In a conventional tightly coupled multiprocessor system, even if the access frequency to the shared memory system is about 10%, the system performance will be significantly reduced if 3 to 4 processors are connected.
[0208] When the shared memory system is used under the condition that the above processor performance is premised, the random access frequency to the shared memory system from the processor side is about 10% of the previous external access, and 30% of the random access frequency. Assuming a write cycle, due to the characteristics of this shared memory system, the actual frequency of access to shared memory (equal to the occupancy rate of the shared bus system) is only 1% per processor (0.5 x 0.1 x 0.3). /1.5 × 100%). Even assuming that the random access frequency to the shared memory system is about 30% of all external access, the actual access frequency to the shared memory is about 3% (0.5 × 0.3 × 0.3 / 1.5 × 100%).
[0209] This indicates that if one set of the shared memory system is provided, a tightly coupled multiprocessor system composed of about 30 to 100 processors can be effectively operated. Comparing the conventional technology with this technology, it is considered that the superiority of this system is remarkably shown in the system having the number of processors of about 3 to 4 to 100.
[0210] The above study evaluates the performance when only one set of the shared memory system is provided. However, the shared data is well distributed to each shared memory system by providing a plurality of sets of the shared memory system. If it is arranged in this way, it will be possible to support the number of processors that is twice as many as the number of sets of the shared memory system.
[0211] If highly random shared data is distributed evenly among the plurality of sets of shared memory systems, the remainder when the address value is divided by the number of sets of the shared memory systems is the remainder of the shared memory system. Each set is numbered, and a group of addresses having the same remainder value is assigned to a shared memory system having a number corresponding to the remainder value, so that the shared memory is interleaved corresponding to the plurality of sets. The method of sharing is effective.
[0212] If the functions, uses, usages, etc. of the shared data can be classified, the shared data can be distributed by optimizing each classification unit and providing separate shared memory systems to distribute the functions and access the entire data. Can also be designed to be evenly distributed for each shared memory system.
Next, using FIG. 10, in conjunction with the interprocessor synchronization processing circuit 1000, the interlock circuit shown in FIG. 8 enables the synchronization processing completion time and the data to be properly enabled on the shared memory. An example of shared memory access timing when the local synchronization processing function that takes correctness with time is enabled is shown. The conditions are exactly the same as those in FIG. 9, except that the local synchronization function is working.
[0214] The bus state on the processor (A) side operates until the end of the state S3, and the bus state on the processor (B) side operates until the end at the same timing as in FIG. The difference is the operation on the processor (A) side after the start S3 where the SYNCOK signal, which is the synchronization completion information from the synchronization processing circuit 1000, turns active, and the shared bus system after the start S6 that changes accordingly. The above write cycle and the lead and write cycle on the shared memory after the start S4.
[0215] The processor (A) generates a write cycle for the address A5 at the state S2, and then, in parallel with the bus cycle, also requests the synchronization processing circuit 1000 to synchronize with the processor (B). (SREQ) is generated at the beginning of state S3. This point can be regarded as the completion time of the task of processor (A), and the write cycle to the shared memory (write cycle for addresses A1 and A5) generated before the start S3 is processed by the task. It is considered to be the result data that may be used by other processors.
[0216] When the synchronization processing circuit receives the synchronization request (SREQ), the SYNCOK signal for the processor (A) is once set to the inactive level (0 level), and at this point, the processor (B) still processes a predetermined task. Since it has not finished, it keeps the inactive state as it is as shown in Fig. 10.
[0217] The processor (B) has a write cycle to address A4 of the shared memory initiated at stage S4 that is the last to the shared memory in the processor (B) side task that should be synchronized with the processor (A). It is a bus cycle, and the processor (B) generates a synchronization request in the synchronization processing circuit 1000 in parallel with this bus cycle at the beginning of the state S5.
[0218] In response to this, the synchronization processing circuit 1000 considers that the synchronization processing between the processor (A) and the processor (B) has been completed (synchronized), and immediately receives the SYNCOK signal to the processor (A). Returns to the active level (1 level). At this timing, the synchronization processing circuit 1000 receives a synchronization request (SREQ) from the processor (B) and tries to set the SYNCOK signal to the processor (B) once to the inactive level, but it is already on the processor (A) side. Since the task has been completed and a synchronization request (SREQ) has been generated along with it, the synchronization process is completed immediately and the SYNCOK signal is immediately returned to the active level.
[0219] Therefore, since the SYNCOK signal of the processor (B) remains substantially active, the processor (B) is not affected by the change of the SYNCOK signal. That is, the operating efficiency of the processor (B) is substantially the same level as in the example of FIG. 9, and is not affected by the synchronization processing (the processor waits for synchronization, etc., overhead for synchronization, etc.). Is operating at maximum efficiency (without any).
[0220] As described above, in the embodiment of the present invention, among the processors that should synchronize with each other, the processor that outputs the synchronization request at the end can operate without lowering the processing efficiency. Even under such conditions, a synchronization request (synchronization request) is required to ensure that the SYNCOK signal once reliably shifts to the inactive level (0 level) and that the inactive level pulse is reliably generated. The timing to output SREQ) is almost the same as the timing to output the address corresponding to the last write cycle to the shared memory in the task that the processor is about to end, etc. Just do it.
[0221] In the example of FIG. 10, the processor (B) may be designed to output a synchronization request (SREQ) at the timing of outputting the address A4 at the state S4. In this way, guaranteeing the edge at which the SYNCOK signal changes to the active level is an important condition for determining the time at which the interlock is started in the interlock circuit example disclosed in FIG.
[0222] Originally, the processor (A) transfers the synchronization request (SREQ) to the shared memory after the start S4 after generating the synchronization request (SREQ) at the beginning of the start S3, even if the synchronization process is not actually completed. Other processes that do not involve access, the next task process, and the like can be executed in advance. This is because this system supports basically the same function as the local synchronization function disclosed in Japanese Patent Application Laid-Open No. 3-234535.
However, in the example of FIG. 10, due to the space, the shared memory address A4 is accessed immediately after the synchronization request (SREQ) is output, and the SYNCOK signal is not displayed at that time. Since it is active (0 level), the bus cycle on the processor (A) side enters the wait state after start S5 in response to this, and the value of D4 is not read from the shared memory until the local synchronization is completed. You can see that.
[0224] This time, when an interlock function as shown in FIG. 8 is added to the local synchronization function and linked with the interprocessor synchronization processing function using the shared memory system of the present invention, the processor-to-processor data is displayed. It has already been described that the feature of the present invention lies in the fact that there is no contradiction in the context of the transfer.
[0225] Regarding this local synchronization processing method, in the embodiment of Japanese Patent Application Laid-Open No. 3-234535, which is a conventional system, the SYNCOK signal is set to the active level assuming the same access status to the shared memory as in FIG. The active RDY-N signal was returned to the processor (A) as the synchronization processing was completed in the return stage S5, and the processor (A) was advanced to the next processing after the stage S6.
[0226] Therefore, in the conventional method, only the period of 1 PCLK cycle of the state S5 is in the waiting state on the processor (A) side, but by using the shared memory system of the present invention, the shared memory can be released. The drcycle and write cycle can now be executed in parallel, and if the same method as before is used, the processor (A) side shared memory will be performed at a very early timing (at state S5) regardless of the write cycle timing. Dcycle has been executed, and at the time of state S5, the value of D4 for the address A4 from the processor (B) that should actually be received is not yet valid in the shared memory, and the processor (A) Will not be able to obtain the desired data.
[0227] The value of D4 corresponding to the address A4 from the processor (B) actually becomes valid on the shared memory at the time of stage S6, and the processor (A) side is operated until at least that time. You need to wait. In the present invention, such a contradiction in the context of data transfer is normalized by an interlock circuit as shown in FIG.
Next, the operating state of the interlock circuit in the timing diagram of FIG. 10 will be described. When the lead cycle for the address A4 on the shared memory on the processor (A) side is generated in the state S3, the SYNCOK signal is already inactive in S3, so the RDY-N signal in Fig. 8 is received. Fixed to inactive at state S4. Therefore, the bus cycle on the processor (A) side enters the wait state (WAIT CYCLE) until the RDY-N signal turns active after stage S5, and the bus cycle for address A4 is extended (the active RDY-N is). Do not end the bus cycle until it is returned).
However, since the processor (A) is executing the pipeline bus cycle, the state S4 has already output the next address value (A3). This mechanism has already been described.
Next, in the state S5, a synchronization request (SREQ) on the processor (B) side is generated for the synchronization processing circuit 1000, and the synchronization processing circuit 1000 receives the synchronization between the two. Assuming that the synchronization process is completed, the SYNCOK signals for the processors (A) and (B) are returned to active. Since the CSBUSY-N signal is at the active level (0 level) at this point, the function of the interlock circuit of the present invention works in response to the change in the rising edge of the SYNCOK signal.
The interlock circuit keeps the RDY-N signal to the processor (A) side at the inactive level (1 level) until the CSBUSY-N signal returns to the inactive level (1 level) once. In this example, the CSBUSY-N signal is shifted by 1 PCLK cycle by PCLK using a flip-flop as shown in 8004 in Fig. 8, so the internally valid CSBUSY-N signal is a dashed line. It can be considered equivalent to having the indicated timing of change.
Therefore, it is at the beginning of stage S7 that the CSBUSY-N signal internally returns to the inactive level (the original CSBUSY-N signal returns to inactive at stage S6) and responds to it. The RDY-N signal to processor (A) is then activated extensively in S7, and processor (A) terminates the bus cycle to address A4 at the last point of state S7.
[0233] In this example, the lead enable signal RE (active level) is displayed on the NAND gates 8002 and 8006 in FIG. 8 so that the local synchronization function works only when the lead cycle to the shared memory is generated. Since it is assumed that 1) is input, the states in which RDY-N is fixed in the inactive state by the local synchronization function of the present invention are 3 states of S4, S5, and S6. Of these, the control of the RDY-N signal in the state S4 was also supported by the conventional local synchronization function (Japanese Patent Laid-Open No. 5-2568) as described above, and the RDY- in S5 and S6. The control of the N signal is a local synchronization function added by the interlock circuit newly added in this invention.
[0234] As a result, the value of D4 corresponding to the shared memory address A4 generated by the state S3 is taken into the processor (A) at the end of the state S7, and the processor (A) is taken into the state S8. After that, it returns to the normal operation (next process or next task). The bus cycle for the address A3 generated and waited by the processor (A) in the state S4 is the status S8 after the local synchronization processing is completed, and the data D3 supported by the processor (A). It ends with (data rewritten by the processor (B)).
[0235] In the embodiment shown in FIG. 10, the local synchronization process is correctly executed even if the CSBUSY-N signal is used at the same timing (shown by the solid line) instead of being used by shifting it internally. Can be done. That is, if the original CSBUSY-N signal change timing is used, the interlock is released in response to the timing when the CSBUSY-N signal turns inactive in the state S6, and the data bus on the processor (A) side is released. The target data D4 and the active RDY-N signal are generated at the timing indicated by the dotted line in the (DATA) and RDY-N signals.
[0236] In response to this, the processor (A) ends the bus cycle corresponding to the address A4 at the last point of the state S6 and proceeds to the next processing.
[0237] In this embodiment, after the target data is output to the shared memory system from the processor on the side of transmitting the data, until it becomes valid on the shared memory and can be read. We have already mentioned that it requires only one rate of latency. Therefore, in the state S6 following the state S5 where the processor (B) outputs the data D4 corresponding to the address A4 on the shared memory, the data D4 is already valid on the shared memory. Therefore, the processor (A) side reads it at the last point of the state S6, and after the state S7, it operates without any problem even if it shifts to the next cycle.
[0238] Further, if the CSBUSY-N signal is designed to be used internally via the OR gate 8007 shown in FIG. 8, the processor (B) side performs a synchronization request (SREQ) at the state S3. Is output, and the synchronization processing circuit determines that synchronization is completed immediately at that stage, and even if the SYNCOK signal is immediately set to the active level, the CSBUSY-N signal is already set to the active level at that point. Therefore, there is no mistake in the timing for interlocking with the interlock circuit.
[0239] That is, in FIG. 10, it is designed to internally use the CSBUSY-N signal in which the alternate long and short dash line and the solid line are superimposed (the CSBUSY-N signal having the logic to generate 0 level if any of them is 0 level). Just do it. The interlock termination condition is the same as when the CSBUSY-N signal is simply shifted and used (in the case of the alternate long and short dash line).
[0240] As described above, even when the original CSBUSY-N signal is directly used, the interlock start condition is almost the same as when the OR gate 8007 is used. Taking all the conditions into consideration, it can be judged that the method of directly using the CSBUSY-N signal as the signal 8008 in FIG. 8 can maximize the efficiency of this embodiment.
[0241] As described above, in the examples of JP-A-5-2568, the 1-port shared memory unit shown in FIG. 6 is used, and the lead cycle and write to the shared memory are used. When there is a conflict with the cycle, the lead cycle is waited and the write cycle is prioritized so that there is no contradiction in data transfer between processors even when linked with the interprocessor synchronization processing mechanism. It has become.
[0242] Because, by the time the synchronization processing circuit generates the synchronization completion information (SYNCOK), the target processor should have already generated all the necessary write cycles for the shared memory system, and from the lead cycle. This is because if the write cycle is prioritized, the processor side will wait unconditionally until the contents of all the shared memories of each processor are changed by those write cycles.
[0243] That is, the interlock function described here suppresses the side effects associated with the speeding up of data transfer between processors by improving the access efficiency by using the 2-port shared memory unit. It can be said that it is for.
[0244] Finally, regarding the case where each logic circuit portion of the shared memory system is made into an LSI (integrated circuit), a measure that can be taken as a method of dividing the function will be described below.
[0245] i) The memory units 2006 and 3006 are integrated into a one-chip or multiple-chip LSI (integrated circuit). That is, a plurality of memory cells (for example, cells as disclosed in FIG. 4) are provided, and at least a read data (DO) output pin, a corresponding read address (RA) input pin, and a write data (DI) input pin are provided. A write address (WA) input pin corresponding to the input pin and a write enable (WE) signal input pin for instructing the designated memory cell corresponding to the write address (WA) to write the data DI are provided. The memory cells are arranged so as to correspond to the write address WA and the read address RA, and the write data obtained from the write data input pin is placed on the input (D) side of at least one memory cell corresponding to the designation of the WA. A means for setting at least 1 bit of the input DI and latching it on the memory cell to be written by the write signal (WR) generated corresponding to the WE, and at least one memory cell corresponding to the RA designation. Output (Z) is selected, and a memory LSI having a means for outputting to the data output pin as at least 1 bit of the read data DO is manufactured.
Further, in the memory LSI, when the write enable (WE) signal input pin is at the active level, the address values given to the write address (WA) input pin and the read address (RA) input pin are the same. In this case, a function is provided to transmit the value specified for the write data (DI) input pin to the read data (DO) output pin side and output it.
[0247] ii) Each of the shared memory systems 1010 to 101n provided corresponding to each processor is integrated into one LSI. In the 2-port shared memory 2006 in the shared memory system, the one having the same function as that shown in i) is not integrated in this shared memory system LSI, or is not integrated in the shared memory system LSI. The one-chip memory LSI or multiple-chip memory LSI shown in i) is configured as a separate system, and is used by connecting to the shared memory system LSI via DI, DO, RA, WA, and WE pins. In this case, it is necessary to provide the connection signals DI, DO, RA, WA, and WE with the memory LSI as the shared memory system LSI input / output pins.
[0248] Further, the input / output buffer systems 2015 to 2017 and the like that switch the access to the shared bus system may be configured as another one-chip LSI or a plurality of chip LSIs.
[0249] iii) The shared memory control unit 2010 is configured as a 1-chip LSI, and the parts of the shared memory systems 1010 to 101n excluding the shared memory control unit 2010 are divided into the 2-port shared memory 2006 and the shared bus system. Including the input / output buffer system 2015 to 2017, etc., it is configured as yet another 1-chip LSI. Then, the two LSIs are connected and used at the level of the access control signal to the shared memory or the shared bus system.
[0250] According to the above-described embodiment, the following effects can be obtained.
(1) Shared Memory A memory unit of 2 ports (consisting of a read port and a write port) that can be operated independently and in parallel is used for the shared memory in the system. As a result, the read cycle to the shared memory and the write cycle can be executed in parallel, the latency required for the data match processing between the shared memories and the data transfer processing between the processors can be shortened, and the access between the processors can be shortened. Since the loss due to competition can be significantly reduced, it has the effect of improving the total flow to the shared memory system.
(2) By synchronizing the entire shared system with one clock, it is possible to eliminate the overhead for synchronizing the asynchronous circuit, and there is an effect that the communication latency can be improved.
(3) When the synchronous processing circuit between processors and the shared memory system are operated in conjunction with each other, the information generated by the target task between the tasks managed by the synchronous processing is transmitted via the shared memory. To ensure reliable communication, the processor read cycle is the period from when the synchronization processing circuit notifies the completion of synchronization until the information in the shared memory is actually rewritten to a valid state for the purpose. It is equipped with an interlock circuit for local synchronization that keeps you waiting. As a result, synchronous processing between processors can be performed reliably and consistently in a form that guarantees the validity of passing data between tasks, and the processor automatically prevents erroneous processing from getting old information. Has the effect of being manageable.
[Effect of the Invention] According to the present invention, a means for performing a data writing process and a reading process for a memory cell of a memory unit or a memory LSI in parallel is provided on the memory unit or the memory LSI. It is possible to reduce the access conflict between the read cycle and the write cycle of the data.
[0255] Thereby, it is possible to provide a parallel processing apparatus with reduced access contention. Alternatively, it is possible to provide a memory LSI that can be used in such a device. Further, by providing means for monitoring each processing device that performs processing in cooperation and reading data from the memory unit by the processing device in response to the completion of all the processing in these processing devices, each of them is provided. It is possible to prevent the processing device from obtaining erroneous data and generating erroneous processing results.
BRIEF DESCRIPTION OF THE DRAWINGS [Fig. 1] Fig. 1 is a diagram showing a shared system architecture of the present invention.
FIG. 2 is a diagram showing textures in a shared memory system corresponding to each processor.
FIG. 3 is a diagram showing a memory unit and a control unit of the present invention.
FIG. 4 is a diagram showing details of a memory cell and a recorder in the memory unit of the present invention.
FIG. 5 is a diagram showing shared memory access control when PCLK and BCLK of the present invention are synchronized and a 2-port shared memory is used.
FIG. 6 is a diagram showing shared memory access control when PCLK and BCLK of the present invention are synchronized.
FIG. 7 is a diagram showing shared memory access control by a conventional method.
FIG. 8 is a diagram showing a ready signal generation circuit linked with the synchronization signal of the present invention.
FIG. 9 is a diagram showing access timing of the shared memory system of the present invention.
FIG. 10 is a diagram showing access timing when the shared memory system of the present invention is interlocked with an interprocessor synchronization mechanism.
[Description of Code] 1110 ~ 111n ... Processor, 1010 ~ 101n ... Each shared memory corresponding to each processor, 1020 ... Abita circuit, 1000 ... Synchronous processing circuit, 2006 ... 2 Port shared memory, 2010 ... shared memory controller unit, 3006 ... memory unit, 3101 ... memory cell, 3010 ... control unit.
Every citation, both ways
| Document | Relation | Office |
|---|---|---|
| JP01109474A | Cites | Japan |
| JP01285088A | Cites | Japan |
| JP01296486A | Cites | Japan |
| JP02168340A | Cites | Japan |
| JP02211571A | Cites | Japan |
| JP05002568A | Cites | Japan |
| JP05290000A | Cites | Japan |
| JP63228365A | Cites | Japan |
| JP63249239A | Cites | Japan |
4 members in 2 offices
Priority claims2
| Document | Office | Kind | Date |
|---|---|---|---|
| 21844695 | Japan | A | |
| JP19950218446 | – | – | – |
Members4
| Document | Office | Kind | |
|---|---|---|---|
| JPH0962563A | Japan | A | |
| US5960458A | United States of America | A | |
| US6161168A | United States of America | A | |
| JP3661235B2This record | Japan | B2 |
14 legal events, as the office reported them to INPADOC
Over the term
Point at a mark for the eventEvents
| Event | Code | |
|---|---|---|
| Cancellation because of completion of termEXPY | EXPY | |
| Renewal fee payment (event date is renewal date of database)FPAY | FPAY | |
| Renewal fee payment (event date is renewal date of database)FPAY | FPAY | |
| Renewal fee payment (event date is renewal date of database)FPAY | FPAY | |
| Renewal fee payment (event date is renewal date of database)FPAY | FPAY | |
| Renewal fee payment (event date is renewal date of database)FPAY | FPAY | |
| Renewal fee payment (event date is renewal date of database)FPAY | FPAY | |
| Renewal fee payment (event date is renewal date of database)FPAY | FPAY | |
| First payment of annual fees (during grant procedure)A61 | A61 | |
| Written decision to grant a patent or to grant a registration (utility model)A01 | A01 | |
| Decision of grant or rejection writtenTRDD | TRDD | |
| Request for written amendment filedA521 | A521 | |
| Notification of reasons for refusalA131 | A131 | |
| Report on retrievalA977 | A977 |
Numbers
- Publication
- 3661235
- Publication, DOCDB
- 3661235
- Publication, EPODOC
- JP3661235B
- Application
- 21844695
- Application, DOCDB
- 21844695
- Application, EPODOC
- JP19950218446
Titles2
- Japanese
- 共有メモリシステム、並列型処理装置並びにメモリLSI
- English
- Shared memory system, parallel processing device and memory LSI
Classification
- CPC, 2
- G06F15/17
- G06F15/167
- IPC, 4
- G06F15 17
- G06F12 00
- G06F12 06
- G06F15 167