Distributed memory type parallel computer and write data transfer end confirming method thereof
Abstract
This record has no abstract on file.
Term
Term ended
Expired 23 February 2020, 6.6 years ago.
- Priority and filed
- Granted
- Expired
- Today
6 claims: 2 independent, 4 dependent
- 1中央演算部と主記憶部と遠隔制御部とを備えるノードがネットワークを介して複数接続され、任意のノードから他のノードに対してデータを転送する分散メモリ型並列計算機であって、 前記任意のノードにおいて、 前記中央演算部は、 前記遠隔制御部に対して転送データの転送終了を示すEOT(EOT:End Of Transfer) コマンドを発行する発行手段を有し、 前記遠隔制御部は、 前記発行手段による前記EOTコマンドを受け付けると、前記主記憶部に記憶されているデータを取得する取得手段と、 前記取得手段により取得された前記データに前記EOTコマンドを付加する付加手段と、 前記付加手段により前記EOTコマンドを付加した前記転送データを前記他のノードに対して転送する転送手段と、を有し、 前記他のノードにおいて、 前記遠隔制御部は、 前記転送手段により転送される前記転送データのコマンド及びアドレスを受け付けて保持する第1の受け付け手段と、 前記転送手段により転送される前記転送データのデータを受け付けて保持する第2の受け付け手段と、 前記第1の受け付け手段により受け付けられた前記転送デ-タのコマンド及びアドレスの競合調停を行う競合調停手段と、 前記競合調停手段により競合調停された前記アドレスを変換するアドレス変換手段と、 前記競合調停手段により競合調停された前記コマンドがEOTコマンドであるか否かを判定する判定手段と、 前記判定手段により前記コマンドがEOTコマンドであると判定された場合に、前記第2の受け付け手段により受け付けた前記データの一部を、前記EOTコマンドが付加されて転送されたことを示す所定の値に差し替える差し替え手段と、 前記差し替え手段により差し替えられた前記データと前記アドレス変換手段により変換されたアドレスとを前記主記憶部に記憶させる記憶手段と、 前記記憶手段により前記主記憶部への記憶が完了した場合に、その旨を前記任意のノードに通知する完了通知手段と、を有し、 前記中央演算部は、 前記記憶手段により前記主記憶部に前記データの一部が差し替えられて記憶されているか否かを検知する検知手段を有することを特徴とする分散メモリ型並列計算機。
- 2前記差し替え手段は、 前記第2の受け付け手段により受け付けられた前記データの先頭部分を差し替えることを特徴とする請求項1記載の分散メモリ型並列計算機。
- 3前記差し替え手段は、 前記第2の受け付け手段により受け付けられた前記データの最終部分を差し替えることを特徴とする請求項1記載の分散メモリ型並列計算機。
- 4中央演算部と主記憶部と遠隔制御部とを備えるノードがネットワークを介して複数接続され、任意のノードから他のノードに対してデータを転送する分散メモリ型並列計算機のデータ転送終了確認方法であって、 前記任意のノードにおいて、 前記中央演算部から前記遠隔制御部に対して転送データの転送終了を示すEOT(EOT:End Of Transfer) コマンドを発行する発行工程と、 前記発行工程による前記EOTコマンドを受け付けると、前記主記憶部に記憶されているデータを取得する取得工程と、 前記取得工程により取得された前記転送データに前記EOTコマンドを付加する付加工程と、 前記付加工程により前記EOTコマンドを付加した前記転送データを前記他のノードに対して転送する転送工程と、 前記他のノードの遠隔制御部において、 前記遠隔制御部にて前記転送データのコマンド及びアドレスを受け付けて保持する第1の受け付け工程と、 前記転送データのデータを受け付けて保持する第2の受け付け工程と、 前記第1の受け付け工程により保持される前記コマンド及びアドレスの競合調停を行う競合調停工程と、 前記競合調停工程により競合調停された前記アドレスを変換するアドレス変換工程と、 前記競合調停工程により競合調停された前記コマンドがEOTコマンドであるか否かを判定する判定工程と、 前記判定工程により前記コマンドがEOTコマンドであると判定された場合に、前記第2の受け付け工程により受け付けた前記データの一部を、前記EOTコマンドが付加されて転送されたことを示す所定の値に差し替える差し替え工程と、 前記差し替え工程により差し替えられた前記データと前記アドレス変換工程により変換されたアドレスとを前記主記憶部に記憶させる記憶工程と、 前記記憶工程により前記主記憶部への記憶が完了した場合に、その旨を前記任意のノードに対して通知する完了通知工程と、 前記他のノードの中央演算部において、 前記記憶工程により前記主記憶部に前記差し替え工程により前記データの一部が差し替えられて記憶されているか否かを検知する検知工程と、を有することを特徴とする分散メモリ型並列計算機のデータ転送終了確認方法。
- 5前記差し替え工程は、 前記第2の受け付け工程により受け付けられた前記データの先頭部分を差し替えることを特徴とする請求項4記載の分散メモリ型並列計算機のデータ転送終了確認方法。
- 6前記差し替え工程は、 前記第2の受け付け工程により受け付けられた前記データの最終部分を差し替えることを特徴とする請求項4記載の分散メモリ型並列計算機のデータ転送終了確認方法。
Independent claims6
102 paragraphs, as filed
[0001] The present invention relates to a distributed memory type parallel computer and a method for confirming the end of data transfer, in particular, mainly using a CPU of a local node, performing write transfer from a local node to a remote node, and performing a write transfer from the local node to the remote node. The CPU relates to a distributed memory type parallel computer for confirming the completion of the write transfer and a method for confirming the completion of data transfer.
[0002] Conventionally, in a computer configuration method, a distributed memory type parallel computer that realizes high processing capacity by operating a plurality of processing devices (computing nodes) at the same time has been attracting attention.
[0003] When performing calculations using such a distributed memory type parallel computer, it is preferable to reduce the frequency of data transfer between each calculation node, that is, to perform closed programming within the calculation node. ing. The reason is that since each calculation node is connected by a network, the distance between each calculation node is larger than the distance within the calculation node.
[0004] However, in a large-scale science and technology problem, it is almost impossible to perform closed programming in each computing node, and in reality, programming is such that each computing node operates in cooperation. For example, this corresponds to the case where a large-scale problem array is mapped across a plurality of nodes (mapped as a global memory space).
[0005] FIG. 5 is a block diagram showing a configuration of a conventional distributed memory type parallel computer. In FIG. 5, in the conventional distributed memory parallel computer, the first computing node 1 and the second computing node are connected via the network 3. Here, the first calculation node 1 that issues an instruction activation will be referred to as a local node 1, and the second calculation node 2 equivalent thereto will be referred to as a remote node 2.
[0006] The local node 1 has a CPU (Central Processing Unit) 11, an MMU (Main Memory Unit) 12, and an RCU (Remote Control Unit) 13. It is composed of.
[0007] The remote node 2 includes a CPU 21, an MMU 22, and an RCU 23, similarly to the local node 1.
[0008] The network 3 connects the local node 1 and the remote node 2, and is generally configured to have a communication register 31 for synchronization processing between the nodes.
[0009] At the local node 1, the CPU 11 can load and store data to the MMU 12, performs an operation using the loaded data, and stores the result in the MMU 12.
[0010] The RCU 13 accepts a data transfer instruction between the MMU 22s straddling the calculation nodes from the CPU 11, and realizes data transfer between the nodes in cooperation with the RCU 23.
[0011] For example, when the CPU 11 instructs the RCU 13 to transfer the data of the MMU 12 to the MMU 22 of the remote node 2, the RCU 13 loads the data of the MMU 12 and transfers the data to the RCU 23 via the network 3. .. RCU23 stores this transferred data in MMU22. This is called inter-node write transfer.
[0012] Further, when the CPU 11 instructs the RCU 13 to transfer the data of the MMU 22 of the remote node 2 to the MMU 12, the RCU 13 sends a data transfer request to the RCU 23. RCU23 loads the data stored in MMU22 and transfers it to RCU13 via network 3. RCU13 stores this transferred data in MMU12. This is called inter-node read transfer.
FIG. 6 is a block diagram showing a detailed configuration of an RCU in a conventional distributed memory type parallel computer. In FIG. 6, the RCU 13 (23) in the conventional distributed memory type parallel computer has a request reception unit 131 (231), a data reception unit 132 (232), a competition arbitration unit 133 (233), and an address translation unit 134 ( It is composed of 234) and a request / data transmission unit 135 (235).
[0014] The request receiving unit 231 receives and holds an instruction from the CPU 21 or an instruction (command address) from the RCU 13 via the network 3.
[0015] The data receiving unit 232 receives and holds a data portion transferred from the RCU 13 via the network 3.
[0016] The competition arbitration unit 233 selects the requests in the request reception unit 231 one by one (competition arbitration).
[0017] The address conversion unit 234 converts the logical node number into the physical node number, the local JOB number into the remote JOB number, and the in-node logical address into the in-node physical address. In particular, physical node number conversion and remote JOB number conversion are required for instructions that access other nodes (local node 1) via network 3, and in-node physical address conversion is in-node memory (here, in-node memory). Required for instructions to access MMU22).
[0018] The request / data transmission unit 235 is a portion in which the address conversion unit 234 sends an instruction (command address) after address translation and load data from the MMU 22 to another node (RCU 13 via network 3). .. In addition, the data held by the data reception unit 232 is required when storing from another node (RCU13 via network 3) to MMU22.
[0019] Generally, a program in which a local node activates a write transfer instruction and confirms its termination on the remote node is shown as follows.
(Local node program) FLAG = 1 DO I = 1, MNODE2 (I + J) = NODE (I) END DO FLAG = 0 (remote node program) IF FLAG .eq.0 THENCALL NEXT_PROGRAM_SUBEND IF [0021] Distributed memory type Programming in a parallel computer is usually expressed by one program using MPI (Message Passing Interface) / HPF (High Performance Fortran), but here, for the sake of clarity, each computing node is used. I will express it separately.
In a conventional distributed memory array computer, the array NODE1 (I) is mapped to the local node 1, the array NODE2 (I + J) is mapped to the remote node 2, and the parent process is set to CPU 11 of the local node 1. There is.
Further, the synchronization flag FLAG that displays the data transfer status is mapped to the communication register 31 in the network 3. Here, when FLAG = 1 is defined as data transfer, and when FLAG = 0 is defined as data transfer completion.
[0024] Here, let's verify the case where the array copy straddles between the nodes and the copy destination process can jump to the next subroutine after the copy is completed.
The above-mentioned program copies the sequence NODE1 (I) to the sequence NODE2 (I + J), and confirms that the FLAG value has changed from 1 to 0 after the copy is completed. After that), it is a program that can jump to the next subroutine.
Here, the array NODE1 (I) is mapped to the local node 1, the array NODE2 (I + J) is mapped to the remote node 2, and the global flag FLAG is equidistant from both the local node 1 and the remote node 2. It is mapped to the communication register 31 in the network (Inter-node-Network) 3 at a distance. Also assume that the parent process is CPU 11 on local node 1.
[0027] In this case, the CPU 11 of the local node 1 issues the FLAG "1" set instruction, the inter-node write transfer instruction, and the FLAG "0" clear instruction to the RCU 13.
[0028] FIG. 7 is a timing chart showing an operation example of a conventional distributed memory type parallel computer. The operation of the local node 1 is shown below with reference to FIG. 7.
[0029] At the local node 1, the CPU 11 issues a write command to the communication register 31 in the network 3 to the RCU 13 in order to notify the CPU 21 of the remote node 2 that data transfer is in progress.
In the RCU 13, a write instruction from the CPU 11 to the communication register 31 is received by the request receiving unit 131, and after competing arbitration by the contention arbitration unit 133, the request / data transmission unit 135 transmits the instruction to the network 3.
[0031] In network 3, the FLAG of the communication register 31 is set to "1", and a "end reply" indicating that the writing process of the FLAG has been performed normally is returned to the RCU 13, and the RCU 13 sends this end reply. Notifies CPU 11 and ends the process of writing to the communication register 31.
Next, the CPU 11 issues a write transfer instruction to the MMU 22 of the remote node 2 to the RCU 13.
In the RCU 13, a write transfer instruction from the CPU 11 is received by the request receiving unit 131, and after the conflict arbitration unit 133, the address conversion unit 134 performs physical node number conversion / remote JOB number conversion / physical address conversion. To access MMU12. After that, the RCU 13 sends the load data from the MMU 12 and the physical node number / remote JOB number together from the request / data transmission unit 135 to the RCU 23 via the network 3.
Next, in the remote node 2, the RCU 23 receives the write transfer instruction sent from the RCU 13 via the network 3 at the request receiving unit 231 and after competing arbitration with the conflict arbitration unit 233, the address translation unit 234. Translate to a physical address. Further, the transferred data is received by the data receiving unit 232 and written to the MMU22 together with the physical address.
[0035] The RCU 23 sends a write transfer command from the RCU 13 and a "end reply" indicating that the write operation to the MMU 22 is normally completed from the request / data transmission unit 235 to the RCU 13 via the network 3, and the RCU 13 sends the write transfer command to the RCU 13. , This end reply is returned to CPU11 and the write transfer instruction is completed.
Next, the CPU 11 issues a write command to the communication register 31 in the network 3 to the RCU 13 in order to notify the CPU 21 of the remote node 2 that the data transfer is completed.
[0037] In the RCU 13, a request reception unit 131 receives a write instruction to the communication register 31, and after competition arbitration by the competition arbitration unit 133, the request / data transmission unit 135 transmits the instruction to the network 3.
[0038] In network 3, the FLAG value of the communication register 31 is cleared to "0", and a "end reply" indicating that the FLAG can be written normally is returned to the RCU 13, and the RCU 13 sends this end reply to the CPU 11. Is notified to end the writing process to the communication register 31.
Next, the operation of the remote node 2 is shown below. At the remote node 2, the CPU 21 issues a communication register read instruction to the RCU 23. In the RCU 23, a read instruction to the communication register 31 is received by the request receiving unit 231, and after competing arbitration by the contention arbitration unit 233, the request / data transmission unit 235 transmits the instruction to the network 3.
[0040] In network 3, the FLAG value of the communication register 31 is read, and the FLAG value and the "end reply" indicating that the reading of the FLAG has been performed normally are returned to the RCU 23, and the RCU 23 notifies the CPU 21 of this. Ends the reading process of the communication register 31 in the network 3.
[0041] This FLAG reading process is repeated until the FLAG value becomes "0", that is, until the completion of the write transfer is confirmed. Then, when it is confirmed that the FLAG value becomes "0", the CPU21 moves to the next sequence (NEXT_PROGRAM_SUB) and a series of operations is completed.
At this time, in the conventional distributed memory type parallel computer, the latency between each unit / within the unit is defined as follows. It is assumed that T used here corresponds to one machine clock of the distributed memory type parallel computer system.
1. Latency between CPU (11,21) / MMU (12,22): 1T 2. Latency between MMU (12,22) / RCU (13,23): 1T3. Network 3 / RCU (13, 23) Latency between 23): 3T 4. Latency through each unit: 0T [0044] Therefore, in the conventional distributed memory type parallel computer, as shown in Fig. 7, until the completion confirmation of data transfer, that is, CPU21 is FLAG = 62T is required to confirm "0".
However, in the distributed memory type parallel computer shown in the above-mentioned conventional example, the FLAG set processing and the clear processing located before and after the originally intended inter-node write transfer instruction are performed. Since the access is to a network that is far away, there is a problem that the overhead before and after the write transfer becomes large.
[0046] On the other hand, since the CPU of the remote node also continues to issue FLAG read instructions to the RCU, access to a network far away also occurs, and the CPU of the remote node completes the write transfer from the completion of the write transfer. There is a problem that the overhead until confirming is large.
[0047] In order to solve this problem, the following two methods can be considered. First, as the first method, if the FLAG set processing and reset processing can be eliminated in the CPU of the local node, the local node can execute and complete the write data transfer at an early timing. ..
Next, as a second method, if the data transfer confirmation flag FLAG is not in the network but in the vicinity of the remote node in the remote node, the turnaround time is shortened, so that the FLAG confirmation process is speeded up. Then, after the data transfer is completed, it is possible to jump to the next subroutine at an early timing.
If the operation by the above two methods becomes possible, it can be expected that the confirmation of the completion of data transfer at the remote node at the time of write data transfer in the distributed memory type parallel computer will be accelerated.
[0050] In the present invention, the CPU of the local node plays a central role in performing write transfer from the local node to the remote node, and the distribution speeds up a series of operations until the CPU of the remote node confirms the completion of the data transfer. It is an object of the present invention to provide a memory type parallel computer and a method for confirming the end of data transfer thereof.
[Means for Solving the Problem] In order to solve the problem, in the invention according to claim 1, a plurality of nodes including a central calculation unit, a main storage unit, and a remote control unit are connected via a network. It is a distributed memory type parallel computer that transfers data from any node to another node.<u style="single">At the arbitrary node, the central arithmetic unit indicates to the remote control unit the end of transfer of transfer data, EOT (EOT::</u><u style="single">End Of Transfer</u><u style="single">) The remote control unit has an issuing means for issuing a command, and when the remote control unit receives the EOT command by the issuing means, the acquisition means for acquiring the data stored in the main storage unit and the acquisition means for acquiring the data. It has an additional means for adding the EOT command to the data, and a transfer means for transferring the transfer data to which the EOT command is added by the additional means to the other node, and the other node. In the remote control unit, the remote control unit receives and holds the first receiving means that receives and holds the command and address of the transferred data transferred by the transfer means, and the data of the transfer data that is transferred by the transfer means. Converts the second receiving means, the competitive arbitration means for performing competitive arbitration of the command and the address of the transfer data received by the first receiving means, and the address arbitrated by the competitive arbitration means. When the address conversion means to be used, the determination means for determining whether or not the command arbitrated by the conflict arbitration means is an EOT command, and the determination means determine that the command is an EOT command. A replacement means for replacing a part of the data received by the second receiving means with a predetermined value indicating that the EOT command has been added and transferred, and the data and the address replaced by the replacement means. A storage means for storing the address converted by the conversion means in the main storage unit, and a completion notification means for notifying the arbitrary node when the storage in the main storage unit is completed by the storage means. The central calculation unit has a detection means for detecting whether or not a part of the data is replaced and stored in the main storage unit by the storage means. It is a type parallel computer.</u>[0052] The invention according to claim 2 is the invention according to claim 1.<u style="single">Distributed memory type parallel computer</u>In<u style="single">The replacement means is characterized in that the head portion of the data received by the second receiving means is replaced.</u>[0053] The invention according to claim 3 is claimed.<u style="single">1</u>Described<u style="single">For distributed memory type parallel computer</u>Put<u style="single">The replacement means is characterized in that the final portion of the data received by the second receiving means is replaced.</u>[0054] The invention according to claim 4 is the invention.<u style="single">A method for confirming the end of data transfer of a distributed memory type parallel computer in which a plurality of nodes having a central arithmetic unit, a main storage unit, and a remote control unit are connected via a network and data is transferred from an arbitrary node to another node. EOT (EOT:) indicating the end of transfer of transfer data from the central calculation unit to the remote control unit at the arbitrary node.</u><u style="single">End Of Transfer</u><u style="single">) The issuing step of issuing a command, the acquisition step of acquiring the data stored in the main memory when the EOT command by the issuing process is received, and the EOT to the transfer data acquired by the acquisition step. An addition step of adding a command, a transfer step of transferring the transfer data to which the EOT command is added by the addition step to the other node, and a remote control unit of the other node to the remote control unit. The first receiving step of receiving and holding the transfer data command and address, the second receiving step of receiving and holding the transferred data data, and the command and holding by the first receiving step. A competitive arbitration process for performing competitive arbitration of addresses, an address conversion step for converting the address arbitrated by the competitive arbitration process, and whether or not the command mediated by the competitive arbitration process is an EOT command. When the determination step and the determination step determine that the command is an EOT command, a part of the data received by the second reception step is transferred with the EOT command added. A replacement step of replacing the data with a predetermined value indicating that, a storage step of storing the data replaced by the replacement step and the address converted by the address conversion step in the main storage unit, and the main storage step of the storage step. When the storage in the storage unit is completed, the completion notification step of notifying the arbitrary node to that effect, and the replacement step of the central calculation unit of the other node to the main storage unit by the storage process. This is a method for confirming the end of data transfer of a distributed memory type parallel computer, which comprises a detection step of detecting whether or not a part of the data is replaced and stored.</u>[0055] The invention according to claim 5 is claimed.<u style="single">3</u>Described<u style="single">How to confirm the end of data transfer of a distributed memory type parallel computer</u>In<u style="single">The replacement step is characterized in that the head portion of the data received by the second reception step is replaced.</u>[0056] The invention according to claim 6 is claimed.<u style="single">3</u>Described<u style="single">How to confirm the end of data transfer of a distributed memory type parallel computer</u>In <u style="single">The replacement step is characterized in that the final portion of the data received by the second reception step is replaced.</u>[Embodiment of the Invention] Next, a distributed memory type parallel computer according to an embodiment of the present invention and a method for confirming the end of data transfer thereof will be described in detail with reference to the accompanying drawings. With reference to FIGS. 1 to 4, an embodiment of a distributed memory type parallel computer according to the present invention and a method for confirming the end of data transfer thereof is shown. The same components as those of the conventional example shown in FIG. 5 will be described with the same reference numerals.
FIG. 1 is a block diagram showing a schematic configuration of a distributed memory type parallel computer according to an embodiment of the present invention. In FIG. 1, the distributed memory type parallel computer according to the embodiment of the present invention has a first calculation node (hereinafter referred to as a local node) 1 and a second calculation node (hereinafter referred to as a remote node) as a plurality of calculation nodes. ) 2 and the network 3 connecting them. In the embodiment of the present invention, the first calculation node 1 will be described as a local node, but the present invention is not limited to this, and the number of the first calculation node 1 is not limited to this.
[0062] The local node 1 has a CPU (Central Processing Unit) 11, an MMU (Main Memory Unit) 12, and an RCU (Remote Control Unit) 13. It is composed of.
[0063] The remote node 2 is composed of the CPU 21, the MMU 22, and the RCU 23, similarly to the local node 1 described above.
[0064] Further, the network 3 connects to the local node 1 and the remote node 2 described above.
[0065] FIG. 2 is a block diagram showing a schematic configuration of a calculation node in the distributed memory type parallel computer according to the embodiment of the present invention. In FIG. 2, the distributed memory type parallel computer according to the embodiment of the present invention includes a request receiving unit 131 (231), a data receiving unit 132 (232), a competitive arbitration unit 133 (233), and an address conversion unit 134 ( 234), a request / data transmission unit 135 (235), an EOT determination unit 136 (236), and a selector 137 (237).
[0066] In the calculation node according to the embodiment of the present invention, in order to speed up the confirmation of the completion of write transfer at the remote node, a mechanism for issuing a write transfer instruction with an EOT (End Of Transfer) mark is provided. The instruction can be issued from CPU 11 of the local node 1.
[0067] Here, the write transfer command with the EOT mark is a fixed value (for example, "All 1": "FFFFFFFF" in hexadecimal for 4B data and 16 for 8B data) for the value of the first element of the data to be transferred. It replaces "FFFFFFFFFFFFFFFF") with RCU23 of remote node 2 and writes it to MMU22.
[0068] However, it is necessary to perform mapping so that there is no problem even if the first element is replaced with another value. In order to realize this method, the RCU23 of the remote node 2 has an EOT determination unit 136 (236) for recognizing that it is a transfer instruction with EOT, and a selector 137 (237) for replacing the EOT mark with the first element. ) Is newly provided, which is different from the conventional example.
[0069] The request receiving unit 231 receives and holds an instruction from the CPU 21 or an instruction (command address) from the RCU 13 via the network 3.
[0070] The data receiving unit 232 receives and holds a data portion transferred from the RCU 13 via the network 3.
[0071] The competition arbitration unit 233 selects the requests in the request reception unit 231 one by one (competition arbitration).
[0072] The address conversion unit 234 converts the logical node number into the physical node number, the local JOB number into the remote JOB number, and the in-node logical address into the in-node physical address. In particular, physical node number conversion and remote JOB number conversion are required for instructions that access other nodes (local node 1) via network 3, and in-node physical address conversion is in-node memory (here, in-node memory). Required for instructions to access MMU22).
[0073] The request / data transmission unit 235 is a part in which the address translation unit 234 sends the instruction (command address) after address translation and the load data from the MMU22 to another node (RCU13 via the network 3). .. In addition, the data held by the data reception unit 232 is required when storing from another node (RCU13 via network 3) to MMU22.
Next, the EOT determination unit 236 and the selector 237, which are the features of the present invention, will be described.
[0075] The EOT determination unit 236 is a circuit for recognizing that the received command is a write transfer command with an EOT mark, and when the received command is a write transfer command with an EOT mark, the transfer data Instruct selector 237 to replace the first element with the EOT mark.
[0076] The selector 237 normally faces the data receiving unit 232, but when the EOT determination unit 236 instructs to replace the EOT mark, the transferred data is changed to the EOT mark (in hexadecimal for 4B data). Replace with "FFFFFFFF", or "FFFFFFFFFFFFFFFF") in hexadecimal for 8B data.
FIG. 3 is a timing chart showing an operation example of the distributed memory type parallel computer according to the embodiment of the present invention. The operation of the local node 1 is shown below with reference to FIG.
At the local node 1, the CPU 11 issues a write transfer instruction with an EOT mark to the RCU 13, and the RCU 13 loads the transfer data (array NODE1 (I)) from the MMU 12 to the RCU 23 via the network 3. Transfer data.
In the RCU 23, the EOT determination unit 236 recognizes the write transfer command transferred via the network 3 as a write transfer command with an EOT mark, and the selector 237 sets the first element of the data as the EOT mark (4B). If it is data, replace it with "FFFFFFFF" in hexadecimal, and if it is 8B data, replace it with "FFFFFFFFFFFFFFFF" in hexadecimal) and store the data transferred to MMU22.
[0080] The RCU 23 sends an end reply to the RCU 13 notifying that the write transfer instruction with the EOT mark has ended normally, and the RCU 13 returns this end reply to the CPU 11 to complete the processing of the local node 1.
On the other hand, the CPU 21 reads the start address of the write transfer in the MMU22 and repeats the EOT mark confirmation process, and when it is confirmed that the read value is the EOT mark, it recognizes that the write transfer is completed. , The next sequence is started and a series of operations is completed.
[0082] With the above configuration, the FLAG write process between the CPU 11 of the local node 1 and the network 3, that is, the FLAG set and the FLAG clear are eliminated, and the CPU 21 and the network 3 of the remote node 2 are eliminated. Instead of the FLAG read processing (confirming the FLAG value) with and, the EOT mark fixed address is read by the MMU22 in the remote node 2, so that the series of operations is speeded up.
[0083] Based on FIGS. 1 to 3, the operation of confirming the end of transfer at the remote node at the time of write transfer in the distributed memory type parallel computer in the present invention will be described.
First, on the local node 1, the CPU 11 issues a write transfer instruction with an EOT mark to the remote memory (MMU22) to the RCU13.
[0085] In the RCU 13, the request reception unit 131 accepts the write transfer instruction with the EOT mark, the conflict arbitration unit 133 performs the competition arbitration, and then the address conversion unit 134 performs physical node number conversion / remote JOB number conversion / physical address conversion. Access MMU12. Then, the load data from the MMU12 and the physical node number / remote JOB number are combined and sent from the request / data sending unit 135 to the RCU23 (network 3).
Next, in the remote node 2, the RCU 23 receives the write transfer instruction with the EOT mark from the RCU 13 (network 3) at the request receiving unit 231, the conflict arbitration is performed by the conflict arbitration unit 233, and the physical is physically performed by the address conversion unit 234. At the same time as the address is converted, the EOT determination unit 236 recognizes that it is a write transfer instruction with an EOT mark, and the selector 237 instructs the selector 237 to replace the first element of the data with the EOT mark.
[0087] Further, the data is received by the data receiving unit 232, only the first element is replaced with the EOT mark by the selector 237, and the data is written to the MMU22 together with the physical address.
[0088] The RCU 23 sends a notification (end reply) that the write transfer with the EOT mark from the RCU 13 and the write operation to the MMU 22 have been completed normally from the request / data transmission unit 235 to the RCU 13 via the network 3, and the RCU 13 Returns this termination reply to CPU11 and completes the write transfer instruction with the EOT mark.
On the other hand, the CPU 21 of the remote node reads out the address to which the EOT mark of the MMU22 should be written and checks it. This operation continues until the EOT mark is written at the address where the EOT mark of MMU22 should be written. Eventually, the EOT mark is written to the MMU22 by the RCU23, and when the CPU21 confirms it, it recognizes that the write transfer has been completed, and moves to the next sequence (NEXT_PROGRAM_SUB) to complete the series of operations.
[0090] Here, for convenience, the latency between each unit / within a unit is defined as follows. However, T is equivalent to one machine clock of this distributed memory type parallel computer system.
1. Latency between CPU (11,21) / MMU (12,22): 1T 2. Latency between MMU (12,22) / RCU (13,23): 1T 3. Network 3 / RCU (13, 23) Latency between: 3T 4. Passing latency within each unit: 0T [0092] At this time, in the configuration example of the present invention, from the issuance of the write transfer command from the local node to the confirmation of the end of transfer by the remote node. The time (until CPU21 confirms the EOT mark of MMU22) is 12T.
<Other Embodiments> FIG. 4 is a timing chart showing an operation example of a distributed memory type parallel computer according to another embodiment of the present invention. The difference from the embodiment of the present invention is that the replacement position of the EOT mark is changed. In the above-described embodiment, the first element of the transfer data is replaced with the EOT mark, whereas in the other embodiment of the present invention, the last element of the transfer data is replaced with the EOT mark.
First, on the local node 1, the CPU 11 issues a write transfer instruction with an EOT mark to the memory (MMU22) on the remote node side to the RCU 13.
[0095] In the RCU 13, the request reception unit 131 accepts the write transfer instruction with the EOT mark, the conflict arbitration unit 133 arbitrates the competition, and then the address conversion unit 134 performs physical node number conversion / remote JOB number conversion / physical address conversion. To access MMU12. Then, the load data from the MMU 12 and the physical node number / remote JOB number are sent together from the request / data transmission unit 135 to the RCU 23 via the network 3.
Next, in the remote node 2, the RCU 23 receives an EOT-marked write transfer command from the RCU 13 via the network 3 at the request reception unit 231, the competition arbitration at the competition arbitration unit 233, and the address conversion unit 234. At the same time as the physical address is converted, the EOT determination unit 236 recognizes that it is a write transfer instruction with an EOT mark, and the selector 237 instructs the selector 237 to replace the final element of the data with the EOT mark. In addition, the data is received by the data receiving unit 232, only the final element is replaced with the EOT mark by the selector 237, and the data is written to the MMU22 together with the physical address.
[0097] The RCU 23 sends an "end reply" indicating that the write transfer with the EOT mark from the RCU 13 and the write operation to the MMU 22 have been completed normally from the request / data transmission unit 235 to the RCU 13 via the network 3, and the RCU 13 Returns this termination reply to CPU11 and completes the write transfer instruction with the EOT mark.
[0098] On the other hand, the CPU 21 of the remote node reads out the address to which the EOT mark of the MMU22 should be written and checks it. This operation continues until the EOT mark is written at the address where the EOT mark of MMU22 should be written. Eventually, the EOT mark is written to the MMU22 by the RCU23, and when the CPU21 confirms it, it recognizes that the write transfer has been completed, and moves to the next sequence to complete the series of operations.
At this time, in the configuration according to another embodiment of the present invention, the time from the issuance of the write transfer command from the local node to the confirmation of the end of transfer by the remote node (CPU21 confirms the EOT mark of the MMU22). Up to) is 26T.
[0100] Each of the above-described embodiments is a preferred embodiment of the present invention, and can be variously modified and implemented without departing from the gist of the present invention.
[0101] For example, the latency between each unit and the latency within the unit are described using fixed values, but the latency is not limited to the above-mentioned values.
[0102] Further, in the above-described embodiment, one element of the transferred data is described as 4B width or 8B width, but the data width is not limited to the data width.
[Effect of the Invention] As is clear from the above description, according to the distributed memory type parallel computer of the present invention and the data transfer end confirmation method thereof, preprocessing of write transfer by the CPU of the local node (during data transfer). It is possible to eliminate FLAG "1" set) and post-processing (FLAG "0" set indicating that data transfer is completed) to the CPU of the remote node. The reason is that the CPU of the remote node only needs to confirm the transfer data itself by providing the EOT determination unit and the selector in the RCU of the remote node and marking the transfer end mark on the transfer data itself.
[0104] Further, according to the distributed memory type parallel computer of the present invention and the data transfer end confirmation method thereof, the write transfer end confirmation process by the CPU of the remote node can be performed at high speed. The reason is that by providing an EOT judgment unit and a selector in the CPU of the remote node, it is possible to put a transfer end mark on the transfer data itself, and the CPU of the remote node is at the EOT mark write address of the MMU in the same node. This is because the turnaround time at the time of confirmation is greatly shortened because it is only necessary to access.
Further, according to the distributed memory type parallel computer of the present invention and the data transfer end confirmation method thereof, the write transfer confirmation process can be performed at high speed, so that the next sequence of data transfer at the remote node can be executed at an earlier timing. It is possible to improve the performance of the entire system. The reason is that in a distributed memory type program, write transfer processing to the MMU of a remote node frequently appears, and by shortening the confirmation processing time, the execution time of the entire program is shortened.
[0106] Further, according to the distributed memory type parallel computer of the present invention and the data transfer end confirmation method thereof, it can be applied to the general logic of exclusive control. In other words, EOT-marked transfer is an operation in which someone wants to read a certain location but cannot read it until the latest data arrives. For example, this can be applied to speed up the latest data read operation for a disk. Is possible.
BRIEF DESCRIPTION OF THE DRAWINGS [Fig. 1] Fig. 1 is a block diagram showing a schematic configuration of a distributed memory computer according to an embodiment of the present invention.
FIG. 2 is a block diagram showing a schematic configuration of an RCU according to an embodiment of the present invention.
FIG. 3 is a timing chart showing an operation example of a distributed memory computer according to an embodiment of the present invention.
FIG. 4 is a timing chart showing an operation example of a distributed memory computer according to another embodiment of the present invention.
FIG. 5 is a block diagram showing a schematic configuration of a conventional distributed memory computer.
FIG. 6 is a block diagram showing a schematic configuration of an RCU unit in a conventional distributed memory computer.
FIG. 7 is a timing chart showing an operation example of a conventional distributed memory computer.
[Description of code] 1 1st computing node (local node) 2 2nd computing node (remote node) 3 Network 11 CPU (central processing unit) 12 MMU (main memory) 13 RCU (remote control unit) 21 CPU (Central processing unit) 22 MMU (main memory) 23 RCU (remote control unit) 31 Communication register
Every citation, both ways
| Document | Relation | Office |
|---|---|---|
| JP04291660A | Cites | Japan |
| JP2000112912A | Cites | Japan |
9 members in 5 offices
Priority claims2
| Document | Office | Kind | Date |
|---|---|---|---|
| 2000052124 | Japan | A | |
| JP20000052124 | – | – | – |
Members9
| Document | Office | Kind | |
|---|---|---|---|
| CA2337677A1 | Canada | A1 | |
| EP1128274A2 | European Patent Office (EPO) | A2 | |
| AU2321101A | Australia | A | |
| JP2001236335A | Japan | A | |
| US2001021944A1 | United States of America | A1 | |
| AU780501B2 | Australia | B2 | |
| JP3667585B2This record | Japan | B2 | |
| EP1128274A3 | European Patent Office (EPO) | A3 | |
| US6970911B2 | United States of America | B2 |
10 legal events, as the office reported them to INPADOC
Over the term
Point at a mark for the eventEvents
| Event | Code | |
|---|---|---|
| Cancellation because of no payment of annual feesLAPS | LAPS | |
| Renewal fee payment (event date is renewal date of database)FPAY | FPAY | |
| Renewal fee payment (event date is renewal date of database)FPAY | FPAY | |
| Renewal fee payment (event date is renewal date of database)FPAY | FPAY | |
| Certificate of patent or registration of utility modelJAPANESE INTERMEDIATE CODE: R150R150 | R150 | |
| First payment of annual fees (during grant procedure)JAPANESE INTERMEDIATE CODE: A61A61 | A61 | |
| Written decision to grant a patent or to grant a registration (utility model)JAPANESE INTERMEDIATE CODE: A01A01 | A01 | |
| Decision of grant or rejection writtenTRDD | TRDD | |
| Written amendmentJAPANESE INTERMEDIATE CODE: A523A521 | A521 | |
| Notification of reasons for refusalJAPANESE INTERMEDIATE CODE: A131A131 | A131 |
Numbers
- Publication
- 3667585
- Publication, DOCDB
- 3667585
- Publication, EPODOC
- JP3667585B
- Application
- 52124
- Application, DOCDB
- 2000052124
- Application, EPODOC
- JP20000052124
Titles2
- Japanese
- 分散メモリ型並列計算機及びそのデータ転送終了確認方法
- English
- Distributed memory type parallel computer and its data transfer end confirmation method
Classification
- CPC, 1
- G06F15/17
- IPC, 2
- G06F13 00
- G06F15 17