Generating device for data computational node
Abstract
The present application discloses a device for generating a data computing node. The device includes a computing manager and a plurality of computing single boards, each of the computing single boards is connected through a switching network; the computing manager is connected through the switching network Connected to each of the computing boards, used to receive a data calculation request including the calculation demand value of the task to be calculated, calculate the target number value of the calculation single board corresponding to the calculation demand value, and determine the number and the For computing boards with the same target quantity value, the determined computing boards are connected through a reconfigurable network to form a strong computing node for computing the data in the task to be calculated. Through the embodiments of the present application, under the premise of solving the scalability of calculation, not only the data transmission efficiency and data calculation performance are improved, but at the same time, the data calculation performance of the target task is substantially improved by using the strong computing node obtained by tight coupling. Fundamentally solve the problem of local strong communication demand.

Term
6.8 yearsto projected expiry
Projected expiry 19 July 2033, counted from filing; an application has no term until it is granted.
- Priority and filed
- Published
- Today
- Projected expiry
6 claims: 1 independent, 5 dependent
- 1一种数据计算节点的生成装置,其特征在于,包括计算管理器和多个计算单板,每个 所述计算单板通过交换网络相连接; 所述计算管理器通过所述交换网络与每个所述计算单板相连接,用于接收包括有待计 算任务的计算需求值的数据计算请求,计算与所述计算需求值相对应的计算单板的目标数 量值,确定数量与所述目标数量值等同的计算单板,将确定的计算单板通过可重构网络进 行连接,组成用于对所述待计算任务中的数据进行计算的计算强节点。
- 2根据权利要求1所述的装置,其特征在于,每个所述计算单板采用全网状full mesh 互联结构通过所述交换网络相连接。
- 3根据权利要求1所述的装置,其特征在于,所述计算单板包括可重构互联模块和至 少一个计算部件; 每个所述计算部件通过所述可重构互联模块与所述交换网络相连接。
- 4根据权利要求3所述的装置,其特征在于,所述计算强节点的计算单板中每个计算 部件通过所述可重构互联模块与所述可重构网络相连接。
- 5根据权利要求4所述的装置,其特征在于,所述可重构互联模块包括数据分配器。
- 6根据权利要求3所述的装置,其特征在于,所述计算部件包括中央处理器CPU、图形 处理器GPU或专用可重构计算阵列HRCAo
Independent claims6
70 paragraphs, as filed
Technical field of generating device for data computing node
[0001] This application relates to the technical field of high-performance computing, and in particular to a device for generating data computing nodes.
Background technique
[0002] Supercomputers are the embodiment of a country's scientific research strength, and it is of vital importance to national security, economic and social development.
[0003] The current supercomputer architecture is mainly divided into two categories: the homogeneous architecture represented by Jaguar and BlueGene/L, and the heterogeneous architecture represented by Roadrunner.
[0004] Among the above two architectures, the former adopts MPP architecture or cluster architecture to achieve high-performance computing of trillions or even petascales per second, but this structure has higher energy consumption. As the number of computing nodes increases, The energy consumption value has increased significantly, making the scalability of this structure affected by the energy consumption limit. When the scale of the computing node is expanded to tens of millions of computing performance levels, the number of CPU cores of this structure will reach hundreds of thousands As a result, the energy consumption of the entire computing system has increased rapidly.
[0005] In order to solve the scalability problem in the above-mentioned architecture, the heterogeneous architecture mentioned in the latter performs conventional calculations on a general-purpose CPU, while data-intensive calculations are implemented through application accelerators with a configurable structure (such as Cell, GPU, FPGA.ASIC chips, etc.). Due to the higher energy efficiency of the accelerator, the overall energy consumption of the entire system is reduced, making the heterogeneous architecture an important development direction for high-performance computing.
[0006] In the above heterogeneous architecture, although the configurable structure of the accelerator can solve the scalability problem of computing, the accelerator can reduce energy consumption while speeding up data transmission or computing performance, but due to the different acceleration efficiency, even if it is certain To a certain extent, data transmission or computing performance can be improved. Limited by the performance of each computing node, it is still unable to effectively improve the computing performance of the overall system and fundamentally solve the problem of local strong communication requirements.
Summary of the invention
[0007] The technical problem to be solved by this application is to provide a device for generating data computing nodes to solve the problem that the existing system structure cannot substantially effectively improve the computing performance of the overall system and fundamentally solve the local strong communication requirements. technical problem.
[0008] The present application provides a device for generating a data computing node, including a computing manager and a plurality of computing single boards, each of the computing single boards is connected through a switching network;
[0009] The calculation manager is connected to each of the calculation boards through the switching network, and is configured to receive a data calculation request including a calculation demand value of a task to be calculated, and calculate the calculation request value corresponding to the calculation demand value. Calculate the target number value of the single board, determine the computing single board whose number is equivalent to the target number value, and connect the determined computing single board through a reconfigurable network to form a calculation for the data in the task to be calculated Strong computing node.
[0010] In the above device, preferably, each of the computing single boards adopts a fully meshed ful 1 mesh interconnection structure to be connected through the switching network.
[0011] In the above device, preferably, the computing single board includes a reconfigurable interconnection module and at least one computing component;
[0012] Each of the computing components is connected to the switching network through the reconfigurable interconnection module.
[0013] In the above-mentioned device, preferably, each computing component in the computing single board of the computing strong node passes through the reconfigurable
The interconnection module is connected to the reconfigurable network.
[0014] In the above device, preferably, the reconfigurable interconnection module includes a data distributor.
[0015] In the above device, preferably, the computing component includes a central processing unit CPU, a graphics processing unit GPU or a dedicated reconfigurable computing array HRCAo
[0016] It can be seen from the above solution that a device for generating data computing nodes provided by the present application adopts a network interconnection structure that supports the coexistence of a large-scale global switching network and a reconfigurable real-time network to achieve high-bandwidth data transmission in asymmetric configuration. , And the computing single board that performs data computing tasks can be used independently as a single node, or it can be tightly coupled with other computing single boards through a reconfigurable real-time network to form a strong computing node. The embodiment of the application solves the premise of computing scalability This not only improves the data transmission efficiency and data calculation performance, but at the same time, the strong computing node obtained through tight coupling substantially improves the data calculation performance of the target task, which can fundamentally solve the problem of local strong communication requirements.
Description of the drawings
[0017] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings needed in the description of the embodiments. Obviously, the drawings in the following description are only some of the present application. Embodiments, for those of ordinary skill in the art, without creative labor, other drawings can be obtained based on these drawings.
[0018] FIG. 1 is a schematic structural diagram of Embodiment 1 of a device for generating a data computing node provided by this application;
[0019] FIG. 2 is another schematic structural diagram of Embodiment 1 of the present application;
[0020] FIG. 3 is a partial structural diagram of Embodiment 1 of the application;
[0021] FIG. 4 is a schematic diagram of another part of the structure of Embodiment 1 of the application;
[0022] FIG. 5 is a schematic partial structural diagram of Embodiment 2 of a device for generating a data computing node provided by this application;
[0023] FIG. 6 is a schematic diagram of another part of the second embodiment of the application;
[0024] FIG. 7 is a schematic structural diagram of Embodiment 1 of the application;
[0025] FIG. 8 is a schematic diagram of a multi-link aggregation data communication process in Embodiment 2 of the application;
[0026] FIG. 9 is a schematic diagram of another part of the second embodiment of the application;
[0027] FIG. 10 is an application example diagram of the second embodiment of the application;
[0028] FIG. 11 is a schematic diagram of another part of the second embodiment of the application.
Detailed ways
[0029] The following will clearly and completely describe the technical solutions in the embodiments of the present application in conjunction with the drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all implementations. example. Based on the embodiments in this application, all other embodiments obtained by a person of ordinary skill in the art without creative work shall fall within the protection scope of this application.
[0030] Referring to FIG. 1, it shows a schematic structural diagram of Embodiment 1 of an apparatus for generating a data computing node provided by the present application. The apparatus includes a computing manager 101 and a plurality of computing boards 102, each of which The computing single board 102 is connected through the switching network 103.
[0031] Wherein, the computing manager 101 is connected to each computing single board 102 through the switching network 103, and is used to receive data computing requests.
[0032] It should be noted that the data calculation request includes the calculation demand value of the task to be calculated.
[0033] Wherein, after receiving the data calculation request, the calculation manager 101 calculates the target number value of the calculation veneer 102 corresponding to the calculation demand value, and determines that the number is equal to the target number value The computing single board 102 connects the determined computing single board 102 through a reconfigurable network 104, as shown in FIG. 2, to form a strong computing node 105 that performs calculation on the data in the task to be calculated.
[0034] It should be noted that when the target number value is 1, the strong computing node 105 includes only one computing veneer 102. When the target number value is greater than or equal to 2, as shown in FIG. 2, The strong computing node 105 includes at least two computing boards 102.
[0035] It should be noted that, in actual implementation, the computing manager 101 is implemented by a management service device.
[0036] Wherein, the reconfigurable network 104 is a reconfigurable real-time network, and the reconfigurable real-time network provides a reconfigurable tightly coupled communication link for the strong computing node 105, and the strong computing node 105 The computing veneer 102 in the network performs high-bandwidth and low-latency data transmission through the reconfigurable bandwidth reconfigurable network 104, forming a logically tightly coupled strong node for high-speed data calculation.
[0037] Wherein, the calculation manager 101 receives the data calculation request to implement the function of starting data calculation; the calculation manager 101 calculates the target number value of the calculation veneer 102 corresponding to the calculation demand value , Determine the number of computing boards 102 equal to the target number value, and realize the functions of data calculation configuration and task allocation; the computing manager 101 connects the determined computing boards 102 through a reconfigurable network to form The strong computing node 105 calculates the data in the task to be calculated by the strong computing node, and completes the function of task scheduling for data calculation.
[0038] It can be seen from the above solution that the first embodiment of a device for generating a data computing node provided by this application adopts a network interconnection structure that supports the coexistence of a large-scale global switching network and a reconfigurable real-time network to achieve asymmetric configuration. High-bandwidth data transmission, and the computing single board that performs data computing tasks can be independently used as a single computing strong node, or it can be tightly coupled with other computing single boards through a reconfigurable real-time network to form a computing strong node containing multiple computing single boards. On the premise of solving the scalability of calculation, the embodiments of the application not only improve the data transmission efficiency and data calculation performance, but at the same time, the strong calculation node obtained through tight coupling substantially improves the data calculation performance of the target task, which can be obtained from Fundamentally solve the problem of local strong communication demand.
[0039] In practical applications, the switching network 103 includes a large-scale global switching network for high-bandwidth data transmission between the computing boards 102. When each computing single board 102 is connected to each other through the switching network, it adopts a full mesh full mesh interconnection structure to be connected through the switching network. As shown in FIG. 3, it is a schematic structural diagram of the computing single board 102 using a full mesh interconnection structure for connection. In Figure 3, each computing veneer 102 adopts a full mesh interconnection mode. Therefore, in the schematic diagram of the structure of the computing strong node 105 as shown in Figure 4, the computing veneer in each computing strong node 105 adopts The full mesh interconnection structure is connected, and high-bandwidth, low-latency data transmission is carried out through a bandwidth-reconfigurable real-time network, thereby forming a logically tightly coupled strong node.
[0040] Wherein, the tight coupling relationship of the computing strong node 105 can be dynamically assigned according to application requirements, that is, the computing manager 101 calculates the target quantity value, and determines the computing veneer 102 equivalent to the target quantity value. With the support of the full mesh interconnection structure, the strong computing node 105 can logically form different tight coupling relationships through dynamic or static reconstruction; on the physical hardware, there are at least one block such as 2, 3, 4 Up to all n computing boards constitute strong computing nodes of different scales, and n is the number of computing boards 102 connected in the full mesh interconnection structure. As shown in FIG. 4, the first strong computing node 105 is composed of three computing single boards 102, and the second strong computing node 105 is composed of two computing single boards 102.
[0041] Referring to FIG. 5, which shows a partial structural schematic diagram of Embodiment 2 of a device for generating a data computing node provided by the present application, the computing single board 102 includes a reconfigurable interconnect module 121 and at least one computing component 122 ;
[0042] Wherein, each of the computing components 122 is connected to the switching network 103 through the reconfigurable interconnection module 121.
[0043] Wherein, the computing component 122 includes a central processing unit CPU, a graphics processing unit GPU or a dedicated reconfigurable computing array HRCA. The HRCA is an FPGA with an application-oriented structure. In the embodiment of the present application, in addition to reconfigurable logic resources, a hard core for application-oriented customization is added. These hard cores can improve the performance of running applications on this chip and reduce power consumption.
[0044] In the device shown in FIG. 5, each of the computing components 122 exchanges and transmits data with all computing boards 102 connected to the switching network 103 through the reconfigurable interconnection module 121. At the same time, the reconfigurable interconnection module 121 provides communication links between all the computing components 122 in the computing single board 102 where it is located.
[0045] Referring to FIG. 6, which shows another partial structural diagram of the second embodiment of the present application, each computing component 122 in the computing single board 102 of the computing strong node 105 communicates with all computing components 122 through the reconfigurable interconnection module 121. The reconfigurable network 104 is connected.
[0046] In the device shown in FIG. 6, the reconfigurable interconnection module 121 provides reconfigurable tightly coupled communication in the strong computing node to which the computing veneer 102 to which it belongs through the reconfigurable network 104 link. That is, in the device shown in FIG. 7, in the strong computing node 105, the reconfigurable interconnection module 121 in each computing single board 102 communicates to the strong computing node 105 through the reconfigurable network 104. Internally, the data bandwidth between the computing boards 102 can be reconstructed.
[0047] Wherein, a data distributor is provided in the reconfigurable interconnection module 121, and the data distributor calculates the number of links in the full mesh interconnection structure of the strong node 104 according to the number of links in the full mesh interconnection structure of the strong node 104 to complete the bandwidth along different chains. The distribution (or aggregation) function of the path, the data distributor supports both unicast and multicast. The data distributor can carry out data aggregation while being able to carry out data distribution.
[0048] For example, assuming that the single link bandwidth of the full mesh interconnection structure in the computing strong node 105 is M, if the actual communication demand between the two computing single boards 102 is less than or equal to M, the single link is used for direct transmission. ; If the communication requirement between the two computing single boards 102 is 5M, then multiple links can be used for data transmission. As shown in Figure 8, it is a schematic diagram of a five-link aggregation data communication process. In Figure 8, the strong computing node includes 8 computing boards, each circle represents a computing board, and each computing board is set with A reconfigurable interconnection module containing a data distributor. When data is transmitted from the source computing single board to the destination computing single board, the data distributor of the reconfigurable interconnection module in the source computing single board divides the 5M data into 5 chains Data transmission is performed on the target computing single board, and the data distributor of the reconfigurable interconnection module in the target computing single board is used for aggregation to realize data transmission.
[0049] In the embodiments of the present application, when the transmission bandwidth exceeds the single link bandwidth of the full mesh interconnection structure, either a reconfigurable circuit direct connection method or a packet forwarding method may be used for multi-link aggregation communication, where :
[0050] Reconfigurable circuit direct connection mode: through the senders reconfigurable interconnection module, using multiple reconstructed links in the full mesh interconnection structure, the data is directly transmitted to the receiver by means of a circuit; the same switching network In contrast, this method can support the use of multiple links of the computing single board and borrowing the link resources of the reconfigurable high-speed interconnection module in other computing single boards, and the direct data transmission can be carried out using the circuit direct connection method;
[0051] Packet forwarding mode: one-time forwarding by a reconfigurable interconnection module in a distributed configuration in other computing components, and then bandwidth aggregation is completed on the target component;
[0052] Hybrid mode: circuit direct connection combined with packet forwarding mode, in the case of multi-source to multi-destination transmission, through custom
To optimize the debugging of the mixed transmission mode of circuit direct connection and packet forwarding.
[0053] In practical applications, the device shown in FIG. 6 is provided with a memory storage connected to each of the computing components 122, as shown in FIG. 9, the memory storage is used to store all The data transmitted or processed by the calculation component 122 in the data calculation process.
[0054] It should be noted that since the computing single boards in the computing strong node are connected in a full mesh interconnection structure through a reconfigurable interconnection module, the physical connection relationship supports static (or dynamic) configuration to change the reconfigurable interconnection module. Reconstruct it into a tightly coupled computing granule for application requirements. If the current application does not have a large-scale, large-volume communication and transmission requirement in a strong computing node, you can change the computing chip located on each computing board. The reconfigurable interconnection module used for bandwidth aggregation communication is reconstituted into the computing arithmetic operation unit required by the application; conversely, when the algorithm structure mapped to the computing strong node requires more strong communication capabilities, the reconfigurable interconnection module still remains The original communication function setting.
[0055] In the practical application of this application, in order to improve the efficiency of data transmission in the network, management functions such as system monitoring, startup, configuration, task allocation, task scheduling and other information are transmitted from computing networks such as switching networks and reconfigurable networks. After separation, the data transmission is performed by the management network alone, wherein the management network can adopt an Ethernet structure. As shown in FIG. 10, the computing manager 101 is set on a management service device, and the computing manager 101 is connected to each of the computing boards through the global switching network, and at the same time is connected to each of the computing boards through the management network. The computing board is connected, the computing data between the computing manager 101 and the computing board is transmitted through the global switching network, and the functional data between the computing manager 101 and the computing board is Information such as task distribution and scheduling is transmitted through the management network, and when data input and output are realized, it is realized through 10 service devices set on the global switching network and the management network.
[0056] It can be seen from the above that in the actual implementation of this application, each of the computing boards has multiple external network interconnections: a global switching network, a reconfigurable real-time network, and a management network. Among them, the global switching network is the backbone network, which is used to complete the global data exchange between the computing components in the computing single board and the system server, and between the computing components on each computing single board; the reconfigurable real-time network uses high-speed switching or The full mesh interconnection method performs high-bandwidth, low-latency fast data exchange between various computing boards or computing components in a strong computing node, usually in the form of real-time data exchange in the form of intermediate results. The management network (also called configuration and monitoring) network is used for the dynamic configuration of computing components and the monitoring of the operating status of the entire computing strong node, as well as the dynamic management of power and power consumption.
[0057] As shown in FIG. 11, the computing single board further includes a management module, and the computing single board is connected to the management network through the management module. The management module is used to complete the communication of the configuration and monitoring network, that is, complete the loading of its own system; receive the reconfigurable array configuration file of each computing component in the computing single board, and complete the reconfiguration of multiple computing components And management; receive relevant system command information, complete the reconfiguration of the network topology of the reconfigurable high-speed interconnection module; collect and report the operating status of the computing board according to requirements; complete the temperature monitoring of each module on the computing board and each Level voltage management.
[0058] Wherein, in FIG. 11, the computing single board further includes an electronic disk connected to the management module, and the electronic disk is used to store each computing single board, the reconfigurable interconnection module, and the management module. Electrically initialize the configuration data, the configuration data when each module needs to be reconstructed, and record the relevant information and log files in the working state of the computing board.
[0059] It can be seen from the above solution that the second embodiment of the present application cooperates with a reconfigurable interconnection module and a global switching network, a reconfigurable real-time network, and a management network, and configures and reconstructs according to the calculation intensity and communication intensity of different applications. Powerful computing nodes with different computing capabilities, and through the communication between different computing power nodes and the coupling relationship of computing in the application algorithm, establish application-driven unbalanced and asymmetrical configuration computing component configuration schemes and reconfigurable information interaction relationships Interact with
ability.
[0060] It should be noted that the various embodiments in this invention are described in a progressive manner, and each embodiment focuses on the differences from other embodiments, and the same and similar parts between the various embodiments are mutually exclusive. Just see.
[0061] Finally, it should be noted that in this article, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or Imply that there is any such actual relationship or order between these entities or operations. Moreover, the terms "including", "including" or any other variations thereof are intended to cover non-exclusive inclusion, so that an article or device including a series of elements includes not only those elements, but also other elements that are not explicitly listed. Or it also includes elements inherent to such items or equipment. If there are no more restrictions, the element defined by the sentence "including a..." does not exclude the existence of another same element in the article or equipment that includes the element.
[0062] The device for generating a data computing node provided by the present invention is described in detail above, and specific examples are used in this article to illustrate the principle and implementation of the present invention. The description of the above embodiments is only used to help understanding The core idea of the present invention; at the same time, for those of ordinary skill in the art, according to the idea of the present invention, there will be changes in the specific implementation and the scope of application. In summary, this content should not be construed as a reference to this application. limits.
9 sheets
Sheet 1 Sheet 2 Sheet 3 Sheet 4 Sheet 5 Sheet 6 Sheet 7 Sheet 8 Sheet 9
Every citation, both ways
| Document | Relation | Office | Category | Cited during |
|---|---|---|---|---|
| WO2025031329A1 | Cited by | World Intellectual Property Organization (WIPO) | – | International search |
| CN110495144A | Cited by | China | – | Search report |
| CN109445752A | Cited by | China | – | Search report |
| CN108845970A | Cited by | China | – | Search report |
| CN104750659A | Cited by | China | – | Search report |
| CN105786757A | Cited by | China | – | Search report |
| WO2021109698A1 | Cited by | World Intellectual Property Organization (WIPO) | – | International search |
| CN110083449A | Cited by | China | – | Search report |
| CN101441615A | Cites | China | A | Search report |
| CN101620587A | Cites | China | A | Search report |
| CN101620588A | Cites | China | A | Search report |
| CN101630305A | Cites | China | A | Search report |
| CN101655828A | Cites | China | Y | Search report |
| CN101710292A | Cites | China | Y | Search report |
| CN102012838A | Cites | China | A | Search report |
| CN102209041A | Cites | China | A | Search report |
| CN102394903A | Cites | China | A | Search report |
| CN102799563A | Cites | China | A | Search report |
| CN102801750A | Cites | China | A | Search report |
| CN102831011A | Cites | China | A | Search report |
| CN103020002A | Cites | China | A | Search report |
| CN103020002A | Cites | China | A | Search report |
| CN103197976A | Cites | China | A | Search report |
| US2007130446A1 | Cites | United States of America | A | Search report |
| US2010169446A1 | Cites | United States of America | A | Search report |
| US2012079501A1 | Cites | United States of America | A | Search report |
4 legal events, as the office reported them to INPADOC
Over the term
Point at a mark for the eventEvents
| Event | Code | |
|---|---|---|
| Termination of patent right due to non-payment of annual feeCF01 | CF01 | |
| Grant of patent or utility modelGrantedC14 | C14 | |
| Entry into substantive examinationC10 | C10 | |
| PublicationC06 | C06 |
Numbers
- Publication
- 103336756
- Application
- 103071786
Titles2
- Chinese
- 一种数据计算节点的生成装置
- English
- Generating device of data computing node
Classification
- IPC, 1
- G06F15 173