A protocol for maintaining cache coherency in a CMP
Abstract
The present application is a protocol for maintaining cache coherency in a CMP. The CMP design contains multiple processor cores with each core having it own private cache. In addition, the CMP has a single on-ship shared cache. The processor cores and the shared cache may be connected together with a synchronous, unbuffered bidirectional ring interconnect. In the present protocol, a single INVALIDATEANDACKNOWLEDGE message is sent on the ring to invalidate a particular core and acknowledge a particular core.
Term
No projected expiry on record.
- Priority
- Filed
- Published
- Today
13 claims: 13 independent, 0 dependent
- 1200529001 (1) 十、申請專利範圍 1 . 一種在一蜂巢多重處理器中保持快取連貫性的系統 ,包含: 一或多個處理器核心; 一共享快取;以及 一環,其中該環連接該一或多個處理器核心以及該共 享快取。 2 ·如申請專利範圍第1項所述之在一蜂巢多重處理器 中保持快取連貫性的系統,其中該一或多個處理器核心之 每一處理器核心包含一私有快取。 ^ ·如申請專利範圍第1項所述之在一蜂巢多重處理器 中保持快取連貫性的系統,其中該共享快取包含一或多個 快取庫。 4 ·如申請專利範圍第3項所述之在一蜂巢多重處理器 中保持快取連貫性的系統,其中該一或多個快取庫係負責 該系統的一實體位址空間之子集。 5 ·如申請專利範圍第1項所述之在一蜂巢多重處理器 中保持快取連貫性的系統,其中該一或多個處理器核心係 爲完全寫入。 6 ·如申請專利範圍第5項所述之在一蜂巢多重處理器 中保持快取連貫性的系統’其中該一或多個處理器核心將 資料完全寫入至該共享快取。 7 .如申請專利範圍第1項所述之在一蜂巢多重處理器 中保持快取連貫性的系統,其中該一或多個處理器核心包 -13 - 200529001 (2) 含一整合緩衝器。 8·如申請專利範圍第7項所述之在一蜂巢多重處理器 中保持快取連貫性的系統,其中資料係儲存在該整合緩衝 器中。 9 ·如申g靑專利範圍第8項所述之在一蜂巢多重處理器 ** 中保持快取連貫性的系統,其中該整合緩衝器將資料淸除 〃 至該共享快取。 1 0 ·如申請專利範圍第1項所述之在一蜂巢多重處理 | 器中保持快取連貫性的系統,其中該一或多個處理器核心 從該共享快取來存取資料。 1 1 ·如申請專利範圍第8項所述之在一蜂巢多重處理 器中保持快取連貫性的系統,,其中該整合緩衝器將多個儲 存組合至一相同的區塊。 1 2 .如申請專利範圍第1項所述之在一蜂巢多重處理 器中保持快取連貫性的系統,其中該環係爲一同步的且非 緩衝的雙向互連環。 g 1 3 .如申請專利範圍第1 2項所述之在一蜂巢多重處理 器中保持快取連貫性的系統,其中一訊息具有一固定潛伏 環繞互連環。 - 14 - A system for maintaining cache coherency in a cellular multiprocessor comprising:one or more processor cores;a shared cache;and a ring, wherein the ring connects the one or more processor cores and the sharing is fast take. 一種在一蜂巢多重處理器中保持快取連貫性的系統,包含:一或多個處理器核心;一共享快取;以及一環,其中該環連接該一或多個處理器核心以及該共享快取。
- 2A system for maintaining cache coherency in a cellular multiprocessor as described in claim 1 wherein each processor core of the one or more processor cores comprises a private cache. 如申請專利範圍第1項所述之在一蜂巢多重處理器中保持快取連貫性的系統,其中該一或多個處理器核心之每一處理器核心包含一私有快取。
- 3A system for maintaining cache coherency in a cellular multiprocessor as described in claim 1 wherein the shared cache comprises one or more cache libraries. 如申請專利範圍第1項所述之在一蜂巢多重處理器中保持快取連貫性的系統,其中該共享快取包含一或多個快取庫。
- 4A system for maintaining cache coherency in a cellular multiprocessor as described in claim 3, wherein the one or more cache banks are responsible for a subset of a physical address space of the system. 如申請專利範圍第3項所述之在一蜂巢多重處理器中保持快取連貫性的系統,其中該一或多個快取庫係負責該系統的一實體位址空間之子集。
- 5A system for maintaining cache coherency in a cellular multiprocessor as described in claim 1 wherein the one or more processor cores are fully write. 如申請專利範圍第1項所述之在一蜂巢多重處理器中保持快取連貫性的系統,其中該一或多個處理器核心係為完全寫入。
- 6A system for maintaining cache coherency in a cellular multiprocessor as described in claim 5, wherein the one or more processor cores write data completely to the shared cache. 如申請專利範圍第5項所述之在一蜂巢多重處理器中保持快取連貫性的系統,其中該一或多個處理器核心將資料完全寫入至該共享快取。
- 7A system for maintaining cache coherency in a cellular multiprocessor as described in claim 1 wherein the one or more processor cores comprise an integrated buffer. 如申請專利範圍第1項所述之在一蜂巢多重處理器中保持快取連貫性的系統,其中該一或多個處理器核心包含一整合緩衝器。
- 8A system for maintaining cache coherency in a cellular multiprocessor as described in claim 7 wherein the data is stored in the integrated buffer. 如申請專利範圍第7項所述之在一蜂巢多重處理器中保持快取連貫性的系統,其中資料係儲存在該整合緩衝器中。
- 9A system for maintaining cache coherency in a cellular multiprocessor as described in claim 8 wherein the integration buffer deletes data to the shared cache. 如申請專利範圍第8項所述之在一蜂巢多重處理器中保持快取連貫性的系統,其中該整合緩衝器將資料淸除至該共享快取。
- 10A system for maintaining cache coherency in a cellular multiprocessor as described in claim 1 wherein the one or more processor cores access data from the shared cache. 如申請專利範圍第1項所述之在一蜂巢多重處理器中保持快取連貫性的系統,其中該一或多個處理器核心從該共享快取來存取資料。
- 11A system for maintaining cache coherency in a cellular multiprocessor as described in claim 8 wherein the integration buffer combines multiple stores into a same block. 如申請專利範圍第8項所述之在一蜂巢多重處理器中保持快取連貫性的系統,其中該整合緩衝器將多個儲存組合至一相同的區塊。
- 12A system for maintaining cache coherency in a cellular multiprocessor as described in claim 1 wherein the loop is a synchronous and unbuffered bidirectional interconnect loop. 如申請專利範圍第1項所述之在一蜂巢多重處理器中保持快取連貫性的系統,其中該環係為一同步的且非緩衝的雙向互連環。
- 13A system for maintaining cache coherency in a cellular multiprocessor as described in claim 12, wherein a message has a fixed latency surround interconnect. 如申請專利範圍第12項所述之在一蜂巢多重處理器中保持快取連貫性的系統,其中一訊息具有一固定潛伏環繞互連環。
Independent claims13
27 paragraphs, as filed
Agreement to maintain cache coherency in a hive multiprocessor
The present invention relates to an agreement to maintain cache coherency in a CMP (comb multiprocessor).
A cached coherent multiprocessor system with two or more independent processor cores. These cores contain multiple caches to replicate memory data close to where it will be lost. One of the functions of this cache coherent agreement is to keep these caches coherent, that is, to ensure memory consistency.
A cached coherent multi-processor is a special case of a cached coherent multiprocessor system. In a cellular multiprocessor, the individual processor cores are integrated onto a single chip wafer. There is currently no agreement that can be used to ensure cache coherency in a single multi-processor. Therefore, an on-chip cache coherency protocol is required to maintain processor cache coherency on the wafers.
The present invention is based on the above-mentioned problems of the prior art, and has a purpose to provide a system for maintaining cache coherency in a cellular multiprocessor, comprising: one or more processor cores; And a ring, wherein the ring connects the one or more processor cores and the shared cache.
Some preferred embodiments of the invention are described in detail below. However, the present invention may be widely practiced in other embodiments than the following description, and the scope of the present invention is not limited by the examples, which are subject to the scope of the following patents. Furthermore, in order to provide a more succinct description and a better understanding of the invention, the various parts of the drawings are not drawn according to their relative dimensions, and some dimensions have been exaggerated compared to other related dimensions; the irrelevant details are not fully Draw, in order to make the schema simple.
Figure 1 illustrates the design of a cellular multiprocessor that includes multiple processor cores P0, P2, P6, etc., each of which has a private cache, and the processor cores have a single on-chip share Cache 10. The shared cache 10 is composed of a plurality of independent caches (not shown). Each repository of the shared cache 10 is responsible for a subset of the physical address space of the system 5. That is, each shared cache repository is the starting location for a non-overlapping portion of the physical address space. The processor cores P0, P2, etc. and the shared cache 10 can be connected together by a synchronous and unbuffered bidirectional interconnect ring 15.
The hive multiprocessor design in Figure 1 contains a plurality of write-thu core caches, which are relative to write-back. This means that when the core writes data, the core simply writes the data completely to the shared cache 10 instead of placing the written data on a cache of the core. When a core writes a material, the system 5 enables the action of writing the data completely to the shared cache 10 instead of storing the data in its own private cache. This is because the bandwidth of the interconnect ring 15 can be supported.
In a cached coherent multiprocessor system, a data stream is typically required to retrieve the changed data and then completely write the changed data; in addition, it is also used for sacrificed data. However, the system 5 in Figure 1 does not need to be. With the system 5, since the data is no longer changed, an error confirmation code is not required in the private caches. In the shared cache 10, there is another backup of the material. If the data has been tampered with, since there is a backup in the shared cache 10, the system does not need to be modified.
Furthermore, among each of the processor cores P0, P2, P6 and the like, there is a combined write buffer or an integrated buffer (not shown). The modified material is placed in the integration buffer (which is a relatively small structure) rather than placing the modified material into a core of the private cache. The integrated buffer deletes the write back continuously or empties to the shared cache 10. Thus, the system 5 does not need to be completely written immediately; instead, the system 5 places the modified material into the integration buffer. The modified data is placed in the integration buffer and the integration buffer is immediately filled, such that the write to the shared cache 10 is performed in a timely manner. Since this is a cellular multiprocessor shared cache design, other processors on the ring 15 can request data written and extrapolated by one of the processor cores. The data is extrapolated to the shared cache 10 in a timely manner, so that the data is being placed in a common location, i.e., where other processors in the cellular multiprocessor can access the data.
In the system of Figure 1, the caches are blocky, and the storage of the cores is sub-blocks. When a store occurs, it stores 1, 2, 4, 8 or 16 bytes of data. The integration buffer combines the storage of a plurality of different bytes into the same block before performing a full write, which may be beneficial in saving bandwidth.
As previously mentioned, Figure 1 illustrates an unbuffered interconnect ring 15. There is no buffering process in the ring 15, so multiple processes are always around the ring 15. For example, if the system 5 is located in the processor P0 in the ring 15, and a message is being transmitted to the processor P6 in the ring 15, there are five processors between P0 and P6. The system 5 knows that if the message was transmitted during the X period, the system 5 will get the message in X+5 cycles. Since the packet (message) is never blocked in the ring 15, this means that there is no fixed deterministic latency in the system 5.
This coherent agreement is designed to maintain unique coherence in a single hive multiprocessor. Assume a standard "invalid agreement" with the common four-state MESI (modified, shared, invalid) design to maintain cache coherency between chips. The block in a shared cache can be one of four states.
1. The non-existent b. Block X does not exist in the shared cache (or one of the core caches) and is owned by the core C. The block X is located in the shared cache, and the core C The exclusive write privilege with block X exists, is not owned, and the custodian = C d. block X is located in the shared cache, and a single core has a backup of block X. e. exists, is not owned There is no custodian f. Block X is located in the shared cache, but multiple cores have a backup of block X. where: X = the physical address of the request block R = the READ of the requested block X Core W = the core of the WRITE of the start block X = the start shared cache of the block X = O = the core of the block X temporarily
The above agreement describes the main operations enabled by the cores during the READS and WRITES. First, a READ data stream will be described, followed by a WRITE data stream.
Initially, if the state of X is non-existent, in this example, there is no backup of X in the shared cache. Therefore, H transmits a request to the memory buffer for extracting block X from the memory. When H receives block X from the memory, it transfers the block X to R, and since R is the only core with block X backup, H will record R as the custodian.
Next, assuming that the state of X is present, not owned, and custodian = R, in this example, it is assumed that H receives another READ request, but this time comes from R1. H does not contain the requested block, and no private cache has an exclusive backup of the block. H reads the block from the cache and passes it to R1. Since multiple cores (R and R1) now have a backup of block X, H also marks the block as having no custodian.
Finally, assuming that the state of X is present and owned by core O, then in this example, H does not contain the requested block, but core O contains an exclusive backup of the block. Therefore, H transmits an EVICT message to core O. Next, H incorporates the request into block X and waits for a response from core O. Once core O receives the EVICT message, it transmits the updated data of block X to H. Once H receives block X, it transmits block X to R.
In a WRITE data stream, when the core W performs a store to the address X, and the address X does not exist in the combined write buffer, a WRITE message is transmitted to the address X in the ring. Start the shared cache repository H. The initial shared cache repository H can take four possible actions depending on the state of the block X.
Initially, when the state of block X is non-existent, in this example, there is no backup of block X in the shared cache. Therefore, H transmits a request to the memory buffer for extracting block X from the memory. When H receives the block X from the memory, it will transmit a WRITEACKNOWLEDGEMENT signal to the core W, and the record core W is the custodian of the block X.
Next, assuming that the state of X is present, not owned, and the custodian = R, in this example, H does not contain the requested block. H transmits an integrated INVALIDATEANDACKNOWLEDGE signal around the ring, which uses the characteristics of the ring. The portion of the INVALIDATE is only transmitted to the core R (i.e., the custodian) to invalidate the cached backup, and the portion of the WRITEACKNOWLEDGE is only transmitted to the core W. This is advantageous because the shared cache only transmits a message to invalidate the core R and confirm the core W. With this interconnected loop, it is no longer necessary to transmit two separate messages. However, if the custodian is the same as the core that enabled the WRITE, then no "invalidate" is transmitted. All other steps will remain the same. Then, since no other core can now have a cache backup of block X, H records W as the owner of block X and converts the custodian to W.
Next, assuming that the state of X is present, not owned, and there is no custodian, in this example, block X exists but there is no custodian. This means that H does not know which core has a cache backup of block X. Therefore, H transmits a single INVALIDATEANDACKNOWLEDGE message around the ring. The portion of the INVALIDATE is transmitted to all cores to invalidate the cached backup, and the portion of the WRITEACKNOWLEDGE is only transferred to the core W (i.e., the processor requesting the write). Since no other core can now have a cache backup of block X, H converts the custodian to W.
Finally, assuming that the state of X is present and owned by core O, then in this example, H does not contain the requested block, but core O contains an exclusive backup of the block. Therefore, H transmits an EVICT message to core O. Next, H locks the request to block X and waits for a response from core O. When the core O receives the EVICT message, it transmits the updated data of the block X to H. When H receives block X, it transmits a WRITEACKNOWLEDGE message to core W, and records core W as the custodian of block X.
Although the present invention has been described above in terms of several preferred embodiments, it is not intended to limit the invention, and it is to be understood that those skilled in the art can make some modifications and refinements without departing from the spirit and scope of the invention. The scope of the invention is defined by the scope of the appended claims.
<p>P0 processor core</p><p>P2 processor core</p><p>P6 processor core</p><p>5 system</p><p>10Shared cache</p><p>15Interconnecting ring</p>
Many of the ideas of the present invention can be more clearly understood by reference to the following drawings. The related drawings are not drawn to scale, and their functions are merely to show the relevant theorems of the present invention. In addition, numbers are used to indicate corresponding parts of the drawings.
Figure 1 is a block diagram illustrating a cellular multiprocessor in an interconnected loop.
4 priority claims, no other members on record
Priority claims4
| Document | Office | Kind | Date |
|---|---|---|---|
| 10749752 | United States of America | – | |
| 74975203 | United States of America | A | |
| 20030749752 | – | – | – |
| US20030749752 | – | – | – |
Numbers
- Publication
- 200529001
- Publication, DOCDB
- 200529001
- Publication, EPODOC
- TW200529001
- Application
- 93140999
- Application, DOCDB
- 93140999
- Application, EPODOC
- TW20040140999
Titles2
- English
- A protocol for maintaining cache coherency in a CMP
- Chinese
- ???????????????????
Classification
- CPC, 4
- G06F12/0813
- G06F12/0811
- G06F12/0831
- G06F12/084
- IPC, 2
- G06F15 16
- G06F12 08