Barrier transactions in interconnects
Abstract
An interconnection circuit system for a data processing device is disclosed herein. The interconnection circuit system is used to provide a plurality of data paths for at least one initiator to access at least one receiving device through the data path. The interconnection circuit system includes: at least one input terminal for receiving data from the at least one initiator. A transaction request of a device; at least one output terminal for outputting a transaction request to the at least one receiving device; at least one path for transmitting the transaction request between the at least one input terminal and the at least one output terminal; the control circuit system is used for The received transaction request is sent from the at least one input terminal to the at least one output terminal; wherein the control circuit system is used for responding to a blocking transaction request, so as to keep at least some transaction requests in the process relative to the blocking transaction request. A sequence in a transaction request message flow transmitted along one of the above-mentioned at least one path is achieved by rejecting at least some of the above-mentioned transaction requests that will be earlier than the above-mentioned blocking transaction request in the above-mentioned transaction request message flow, as opposed to the above-mentioned At least some of the transaction requests in the transaction request message flow that are later than the blocking transaction request are reordered; wherein the blocking transaction request includes an indicator that indicates which of the transaction requests in the message flow of the transaction request Include at least some of the above transaction requests whose order needs to be maintained.

Term
No projected expiry on record.
- Priority
- Filed
- Granted
- Today
26 claims: 6 independent, 20 dependent
- 1An interconnection circuit system for a data processing device. The interconnection circuit system is used to provide a plurality of data paths for at least one initiating device to access at least one receiving device through the data paths. The interconnection circuit system It includes:at least one input terminal for receiving a transaction request from the at least one initiating device;at least one output terminal for outputting a transaction request to the at least one receiving device;at least one path for connecting the at least one input terminal and The transaction request is transmitted between the at least one output terminal;the control circuit system is used to send the received transaction request from the at least one input terminal to the at least one output terminal;wherein the control circuit system is used to respond to a blocked transaction Request to maintain an order of at least some transaction requests with respect to the blocking transaction request in a transaction request message flow passing along one of the above-mentioned at least one path, which is relative to the transaction request message flow At least some of the transaction requests that are later than the above-mentioned blocking transaction request, refuse to reorder at least some of the above-mentioned transaction requests that are earlier than the above-mentioned blocking transaction request in the above-mentioned transaction request message flow;wherein the above-mentioned blocking transaction request includes an indicator, which indicates Among the above-mentioned transaction requests in the above-mentioned transaction request message flow, which one contains at least some of the above-mentioned transaction requests whose order is to be maintained. 一種用於一資料處理設備的互連電路系統,上述互連電路系統用以提供複數個資料路線,以供至少一起始裝置透過該些資料路線來存取至少一接收裝置,上述互連電路系統包含:至少一輸入端,用以接收來自上述至少一起始裝置的交易請求;至少一輸出端,用以輸出交易請求至上述至少一接收裝置;至少一路徑,用以於上述至少一輸入端與上述至少一輸出端之間傳輸上述交易請求;控制電路系統,用以將上述所接收之交易請求自上述至少一輸入端發送至上述至少一輸出端;其中上述控制電路系統用以回應一阻隔交易請求,以相對於沿著上述至少一路徑其中之一傳遞的一交易請求訊息流中的上述阻隔交易請求而保持至少某些交易請求的一順序,其係藉由相對於上述交易請求訊息流中晚於上述阻隔交易請求的至少某些交易請求,拒絕將在上述交易請求訊息流中早於上述阻隔交易請求的至少某些上述交易請求重新排序;其中上述阻隔交易請求包含一指示元,其指明在上述交易請求訊息流內的上述交易請求中,何者包含其順序將被保持的上述至少某些交易請求。
- 19An initiating device for sending a transaction request to a receiving device through an interconnection, the initiating device comprising:a blocking transaction request generator for generating a blocking transaction request to indicate to the interconnection that it should be maintained through the interconnection A sequence of at least some of the transaction requests in a transaction request message flow, by not allowing at least some of the transaction requests that occurred before the blocking transaction request in the transaction request message flow to occur in relation to the transaction request. Reordering at least some of the above transaction requests after the blocking transaction request;wherein the blocking transaction request generator is used to provide an indicator to the generated blocking transaction request to indicate the transaction request in the transaction request message flow , Which contains at least some of the above transaction requests whose order needs to be maintained. 一種起始裝置,用以透過一互連向一接收裝置發出交易請求,該起始裝置包含:一阻隔交易請求產生器,用以產生阻隔交易請求以向上述互連指明應該保持在通過上述互連之一交易請求訊息流內之至少某些交易請求的一順序,其係藉由不允許將在上述交易請求訊息流發生於上述阻隔交易請求之前的至少某些上述交易請求相對於發生於上述阻隔交易請求之後的至少某些上述交易請求而重新排序;其中上述阻隔交易請求產生器用以提供一指示元予上述所產生的阻隔交易請求,以指明在上述交易請求訊息流內的上述交易請求中,何者包含上述其順序需被保持的至少某些交易請求。
- 21For example, the initiating device described in claim 19, the initiating device further includes:a hazard unit for storing a prominent transaction request that may cause a hazard;and an output terminal for outputting the transaction request to the interconnection;wherein The initiation device is configured to respond to the detection that the hazardous unit is full, so that the output of further transaction requests can be suspended, so as to generate and output a blocking transaction related to at least one address, as well as empty and The above-mentioned hazard unit of any transaction request related to the above-mentioned at least one address. 如請求項19所述的起始裝置,上述起始裝置更包含:危害單元,用以儲存可能產生一危害的突出交易請求;以及一輸出端,用以輸出上述交易請求至上述互連;其中上述起始裝置經設置用以針對偵測到上述危害單元已存滿做出回應,而使得可暫停進一步交易請求的輸出,以產生並輸出與至少一位址相關的一阻隔交易,以及清空與上述至少一位址相關的任何交易請求的上述危害單元。
- 24A data processing equipment comprising:at least one initiating device according to request item 19 for issuing a transaction request;at least one receiving device for receiving the aforementioned transaction request;and an interconnection according to request item 1 for connecting the aforementioned At least one initiating device to the above-mentioned at least one receiving device. 一種資料處理設備,包含:根據請求項19的至少一起始裝置,用以發出交易請求;至少一接收裝置,用以接收上述交易請求;以及根據請求項1的的一互連,用以連接上述至少一起始裝置至上述至少一接收裝置。
- 25A receiving device for receiving a transaction request from an interconnection, the transaction request including a blocking transaction request, the blocking transaction request is used to maintain a sequence of at least some transaction requests in a transaction request message flow, by rejecting At least some of the transaction requests that are earlier than the blocking transaction request in the transaction request message flow are reordered relative to at least some of the transaction requests that are later than the blocking transaction request in the transaction request message flow. The transaction request includes an indicator element to indicate which of the above-mentioned at least some transaction requests should be maintained. The above-mentioned receiving device includes:a response signal generator, and the above-mentioned response signal generator is used to respond to the blocking transaction request to generate and Sending a response to the blocking transaction request and responding to a predetermined indicator in the blocking transaction request to delay the transmission of the response until the receiving device has at least partially processed the transaction request received before the blocking transaction request At least one transaction request. 一種接收裝置,用以接收來自一互連的交易請求,上述交易請求包含阻隔交易請求,上述阻隔交易請求用以保持一交易請求訊息流內至少某些交易請求的一順序,其係藉由拒絕將在上述交易請求訊息流中早於上述阻隔交易請求的至少某些上述交易請求,相對於上述交易請求訊息流中晚於上述阻隔交易請求的至少某些上述交易請求,而重新排序,上述阻隔交易請求包含一指示元,以指明上述至少某些交易請求中何者的順需應予保持,上述接收裝置包含:一回應信號產生器,上述回應信號產生器用以回應上述阻隔交易請求,以產生並傳送對上述阻隔交易請求之一回應,以及回應上述阻隔交易請求中的一預定指示元,以延遲傳送上述回應,直到上述接收裝置已經至少部分處理了於上述阻隔交易請求之前接收之上述交易請求的至少一交易請求。
- 26A method for transmitting data from at least one initiating device to at least one receiving device through an interconnection circuit system. The method includes the following steps:receiving a transaction request from the at least one initiating device at at least one input terminal;Send along at least one of the multiple paths to at least one output end;respond to the receipt of a blocking transaction request: maintain at least some of the transaction request message flow along one of the above-mentioned at least one path A sequence of transaction requests relative to the aforementioned blocking transaction requests by rejecting at least some of the aforementioned transaction requests that are earlier than the aforementioned blocking transaction requests in the aforementioned transaction request message flow, as opposed to later than the aforementioned transaction requests in the aforementioned transaction request message flow. Block at least some of the above transaction requests of the transaction request, and reorder;wherein the above blocking transaction includes an indicator that indicates which of the above transaction requests in the above transaction request message flow contains at least one of the above mentioned transaction requests whose order needs to be maintained. Some transaction requests. 一種用以將資料由至少一起始裝置透過互連電路系統而發送至至少一接收裝置的方法,上述方法包含以下步驟:由位於至少一輸入端的上述至少一起始裝置接收交易請求;將上述交易請求沿著複數個路徑其中至少一路徑向至少一輸出端發送;針對接收到一阻隔交易請求做出回應:在沿著上述至少一路徑其中之一傳遞的一交易請求訊息流內,保持至少某些交易請求相對於上述阻隔交易請求的一順序,其係藉由拒絕將在上述交易請求訊息流中早於上述阻隔交易請求的至少某些上述交易請求,相對於上述交易請求訊息流中晚於上述阻隔交易請求的至少某些上述交易請求,而重新排序;其中上述阻隔交易包含一指示元,其指明在上述交易請求訊息流內的上述交易請求中,何者包含其順序需被保持的上述至少某些交易請求。
Independent claims6
222 paragraphs in 1 section, as filed
Barrier transactions in interconnection
BARRIER TRANSACTIONS IN INTERCONNECTS
The present invention relates to the field of data processing systems. Specifically, the present invention relates to an interconnection circuit system for data processing equipment. The interconnection circuit system provides a plurality of data routes, and one or more start devices (such as a master station) can be stored via this data route. Take one or more receiving devices (e.g., slave stations).
Interconnections can be used to provide connections between different components in a data processing system. These interconnections provide multiple data routes, and one or more initiator devices can access one or more receiving devices via these data routes. The initiating device is a device that generates a transaction request, and thus can be a master station (eg, a processor) or it can be another interconnect. The receiving device is a device that can receive the transaction, and thus can be a slave station (eg, a peripheral component) or another interconnection.
As the system becomes more and more complex and uses multiple processors to communicate with each other or multiple devices, the designer who writes software for a multi-processor system must have more details on the layout of the circuit system components and the potential of the architecture. In order to write software that can ensure the consistent behavior of interactive processing over a long period of time. Even with these detailed understandings, achieving this consistency still requires a lot of effort and sacrifices the performance of the system.
At present, there is a great need to propose a new mechanism so that programmers can use a universal method for any architecture, that is, to ensure the consistency of interactive processing over a long period of time.
A first aspect of the present invention provides an interconnection circuit system for a data processing device. The interconnection circuit system is used to provide a data path for at least one initiating device to access at least one receiving device through the data path, The interconnection circuit system includes: at least one input terminal for receiving a transaction request from the at least one initiating device; at least one output terminal for outputting a transaction request to the at least one receiving device; at least one path for receiving the at least one input The transaction request is transmitted between the terminal and the at least one output terminal; the control circuit system is used to send the received transaction request from the at least one input terminal to the at least one output terminal; wherein the control circuit system is used to respond to a blocking Transaction request, in order to block transaction requests in a transaction request message flow transmitted along one of the above-mentioned at least one path, and maintain a sequence of at least some transaction requests. This function is achieved by relative to the transaction request At least some transaction requests in the message flow that are later than the blocking transaction request, refuse to reorder at least some of the transaction requests that are earlier than the blocking transaction request in the transaction request message flow; wherein the blocking transaction request includes an indicator , Which indicates which of the above transaction requests in the message flow of the above transaction request contains at least some of the above transaction requests whose order will be maintained.
When the system becomes more complicated because it is equipped with multiple processors and multiple peripheral components, it is difficult for the programmer to maintain the required relative transaction sequence without grasping the details of the system architecture used to execute the program. Providing an interconnected circuit system that can respond to blocked transactions enables software designers to ensure consistency of behavior without having to consider the system's architecture and circuit system component layout.
In particular, it provides an interconnection circuit system with a control circuit system, where the control circuit system can be used to respond to blocking transaction requests to maintain the order of at least some of the requests relative to the blocking, which means that the software designer can only understand the data developer In the case of the logical relationship with the consumer, and without understanding the circuit system component layout and delay of the system used to operate the software, it is possible to write software that can be operated on the system. In this way, the interconnection circuit system allows programmers to maintain the relative order of transactions without having to consider the structure.
A blocked transaction means that the transaction has a property such that a plurality of transactions under its control cannot change the order relative to the blocked transaction. Therefore, blocking transactions can be inserted into a transaction request message stream to maintain the order of transactions under its control, and thus prevent certain transactions from proceeding earlier than others. By providing an indicator in the barrier transaction to specify which transaction request the barrier can control the transaction request message flow, more precise control can be provided to the system, and the potential time due to the barrier can be limited to only those transaction requests. A subset of transactions (which include transactions that do not need to be delayed), while allowing other transactions to proceed as usual.
If there is no barrier, the designer must have a detailed understanding of the architectural relationship of multiple agents in the system; when barriers are used, the designer only needs to know the logical relationship between the data developer and the consumer. These logical relationships will not change due to the execution of the software on different architectures, and thus enable designers to create software that can run consistently on all platforms, and make the software system easier to use across platforms.
In fact, the barrier allows the design of hardware and software to be decoupled, so that third-party software can be easily deployed.
In some embodiments, the control circuit system can maintain the order by delaying the transmission of at least some transaction requests (which occur in the transaction request message flow after the blocking transaction request) until it is received Clear the above-mentioned response signal that blocks the transaction.
Generally speaking, the control circuit system does not allow the blocking transaction request to surpass the transaction requests that must maintain the order among the at least some transaction requests and are located in front of the blocking in the transaction message flow. In some "non-blocking" barriers, the barrier also does not allow transaction requests located behind the barrier to pass the barrier, and the order can be maintained. In a system, when the at least some transaction requests that should remain behind the barrier have been delayed (maybe delayed by the blocking circuit system), the transaction requests located behind the barrier may be allowed to pass the barrier because The barrier does not need to maintain the order of these transaction requests, because these transaction requests have already been delayed upstream. However, all transaction requests in front of the barrier must still remain before the barrier, otherwise when the delayed transaction request is allowed to proceed because the response indicates that the barrier has reached a response signal from a response signal generator, the transaction request shall be allowed to proceed. Indicate that all transaction requests before the barrier in the transaction message flow have also reached this point.
In some embodiments, the blocking transaction described above has the same properties as other transaction requests processed by the interconnection, and therefore, the interconnection can provide this additional function, and at the same time be in other aspects with a more traditional interconnection. very similar.
It should be noted that an initiating device is any device that is located upstream of the interconnect and can supply transaction requests. So, for example, it can be another interconnection or it can be a master station. Similarly, a receiving device is a device located downstream of the interconnection and capable of receiving the transaction request, so, for example, it may be a slave station or another interconnection.
In some embodiments, the indicator can be used to indicate a property of the transaction request, and at least some of the transaction requests include transaction requests with the foregoing properties.
One way to specify which transaction requests should be subject to the blocking control is to specify the nature of these requests. In this way, any transaction request of this nature can be identified and processed accordingly. In many cases, it can be transaction requests with certain properties that must be controlled, because only these transaction requests may cause the system to operate incorrectly.
In some embodiments, the aforementioned property includes a source of the aforementioned transaction request.
The barrier can also be exclusive to a specific initial device. This property is useful when the order of transactions generated by the master station relative to each other transaction is important, and there is no need to maintain the order relative to other transactions generated elsewhere. . Allowing a block to delay only transactions from a specific initiating device helps to reduce the latency introduced by the block, because it can reduce the number of blocked transactions that are delayed.
In some embodiments, the indicator indicates a function of the transaction request.
Certain interconnections can be set up so that barriers can be used to delay transaction requests with a predetermined function without delaying other transaction requests. So, for example, they can delay writing to memory but not reading.
In some embodiments, the indicator indicates one or more addresses, and at least some of the transaction requests include transaction requests to the one or more addresses.
Since the format of the barrier transaction is designed to be similar to other transactions in the interconnection, the barrier transaction has an address field, and therefore, a convenient and effective way to limit the impact of the barrier is to restrict these barriers to only one or more destinations. Transactions with addresses. Therefore, only transactions to these addresses will be controlled by the barrier. The above method is very useful when it is important for two or more transactions to the same address or address range so that they cannot be reordered relative to each other. The base address in the block transaction address column can be used to indicate a range of one address in the block, and further fields can be provided to provide an indication of the size of the range.
A further advantage of addressing blocking transactions is that there is no need to replicate these transactions at each diverging node, because they only control transaction requests to certain destinations, and therefore, they only need to be transmitted along the path to these addressed destinations. . Therefore, this method is suitable when the layout of the interconnected circuit system components has many nodes, and if there are many replicated barriers that may cause a storm to be blocked, this method is suitable. Using address blocking can avoid the aforementioned problems.
Address blocking transactions can be used to block all subsequent transactions, or in some embodiments, transactions to a specified address or address range can be blocked, and no further transmission of these transactions is allowed until one of the blocked transactions is received So far. In other embodiments, it can be used in a non-blocking manner. As long as the blocked transaction is to be transmitted to each destination that may have this address, in some implementations, subsequent transaction requests will not be delayed before receiving a response to one of the blocked, and the blocked only stays at The required transaction information flow, and does not allow any subsequent transactions to pass it. In this way, it can provide the required control without excessively increasing the system latency.
Address blocking can also be used in systems with hazardous units. The hazard unit is used to track pending transactions that may cause hazards, because these transactions may affect later transactions. For example, a storage transaction stores data to a specific address, which will affect a reading of that address, and therefore, in the transaction message flow, any reading of the address after the storage It should be kept at the rear of storage. A hazard unit tracks possible hazard transactions and any later transactions that may be a hazard until it has detected that the hazard transaction has been completed. In a more ideal situation, the hazard unit should be small. However, if it is full, the initiating device cannot issue any transaction until a hazard transaction has been completed and the related transaction request has been deleted by the hazard unit. Suspending the system in this way will increase latency, and one way to deal with this problem is to use blocking. A barrier ensures that the sequence of certain transactions can remain unchanged. However, they themselves may cause latency. A certain location barrier is a good solution, because if a barrier is generated to the address of one of the most recently stored hazardous transactions, no subsequent transactions can surpass this hazardous transaction, and therefore there will be no more damage from the hazardous transaction. A hazard of unit removal.
In some embodiments, the interconnection circuit system includes a plurality of regions, each region includes at least one input from an initiating device, and the blocking transaction request includes a region indicator to indicate whether the blocking should be delayed by all The transaction request received by the initiating device may only delay the transaction request received from the initiating device in one of the above-mentioned areas, or the transaction request shall not be delayed.
The method to further limit the functionality of the barrier is to restrict it to a specific area. The area that receives the signal from a specific originating device and has a barrier to the relevant indicator can indicate whether it should delay transactions from all originating devices or only delay transactions from the originating device related to a specific area or not delay transaction requests . If the area is properly set, this can provide a convenient way to limit blocking transactions and reduce latency while maintaining the required functionality.
In some embodiments, the above-mentioned interconnection circuit system further includes a barrier management circuit system.
In addition to the control circuit system used to control the transaction request relative to the blocked transaction, there may also be a blocking management circuit system, which can manage the blocked transaction and delete or merge it when appropriate.
The transmission of blocking transaction requests and the processing of their copying and subsequent responses do not require management costs. Therefore, if there are adjacent blocking transaction requests, it is more advantageous to merge them to form a single blocking transaction request, which can be based on demand And control the order of surrounding transaction requests. In this way, it is only necessary to respond, copy, etc. to this merged request.
In some embodiments, the above-mentioned control circuit system is used to copy a barrier transaction at a divergence point that is located at an entry point of a convergence area, so as to provide a convergence indicator to the above-mentioned copied barrier transaction; and The barrier management circuit system can respond to the detection that the copied barrier transaction leaves the convergence area to remove the convergence indicator and merge at least some of the copied barrier transactions.
When there is a divergent node and a single entry path becomes several exit paths, in order for the barrier transaction to function normally, these barrier transactions must be replicated on these exit paths. This makes it possible to maintain the relevant transaction sequence with respect to the barrier transaction on each path. If the divergent path enters a convergence area, then at least some of these paths will merge. This may cause all replicated barriers to be passed down these merged paths. Therefore, in some embodiments of the present invention, not only the blocking transaction is copied, but also an aggregation indicator is provided for the blocking transaction. This enables the barrier management circuit system to detect the copied barrier transaction that leaves the convergence area, understand that it is related to the other copied barrier transaction, and merge at least some of these copied barrier transactions when appropriate Obstruct. The convergence indicator should also be removed at the edge of the convergence area, and from then on, the barrier can be handled independently.
It should be noted that if a transaction request to a specific address always uses the same path to pass through a convergence area, then the area is not functionally convergent at that location. If this situation applies to all addresses, then The area is not convergent for any location, and it can be treated as a cross-coupling area. There are some advantages to not having any convergence area, and therefore, in some embodiments, the interconnection is designed such that the area is not convergent for the addressed transaction, and therefore, although from the layout of the circuit system components It seems that it is convergent, but it functions as a cross-coupling area.
In some embodiments, the control circuit system may be further used to provide an instruction for the copied barrier to specify the number of copied barrier transactions, so as to indicate the number of barrier transactions that can be combined to the barrier management circuit system. .
The control circuit system can provide an instruction to the copied barrier transaction to specify the number of the copied barrier transaction. This allows the barrier management circuit to know how many potential barrier transactions are available for merging, and collect these barrier transactions at the exit of the convergence area and merge them together when appropriate.
In some embodiments, the aforementioned blocking management circuit system is used to respond to the detection of neighboring blocked transaction requests having the same indicator in the transaction request message stream, so as to merge the aforementioned neighboring blocked transaction requests and provide the aforementioned merged transaction request. The blocked transaction request with the same indicator as above.
In other cases, the blocking transaction request can be removed from the system by merging it together. For example, neighboring blocking transaction requests with the same indicator will have the same effect on the same transaction, and therefore they can be combined into a single blocking transaction request.
In some embodiments, the barrier management circuit system is used to combine barrier transaction requests with different indicators, and provide an indicator to the combined barrier transaction request to specify each of the combined barrier transaction requests. Indicates the transaction referred to by the dollar.
In addition, blocking transaction requests with different indicators can also be combined, as long as a new indicator can be assigned to the combined blocking transaction request to specify the transaction request controlled by each blocking doctrine of the combined blocking transaction Can be controlled by the merged barrier transaction.
For example, when the barrier transaction request has different region indicators, a combined barrier transaction may have a region indicator to indicate the limitation of the two combined regions. This indicator should include the two areas specified by each independent barrier exchange. Therefore, if the two areas overlap, the indicator can be appropriately used to indicate the largest area of the two areas.
In some embodiments, the above-mentioned blocking management circuit system is configured to detect a blocking transaction request followed by a previous blocking transaction request and there is no order between the two to be controlled by the subsequent blocking transaction request. A response is made to: reorder the transaction requests so that the subsequent barrier transaction is moved to the location adjacent to the previous barrier transaction request; and merge the adjacent barrier transaction requests.
Similarly, in other embodiments, the above-mentioned blocking management circuit system is configured to detect an interfering transaction in which a blocking transaction request is followed by a previous blocking transaction request and there is no order between the two being controlled by the preceding blocking transaction request. A response is requested to: reorder the transaction requests so that the previous blocked transaction is moved to a place adjacent to the subsequent blocked transaction request; and merge the adjacent blocked transaction requests.
When there are multiple blocking transaction requests with interfering transaction requests, and none of these interfering transaction requests are controlled by the first blocking transaction request or none of them are controlled by the second blocking transaction request, they can move relative to these requests. The blocked transaction request that has an interference request under its control, and therefore, can be moved to be adjacent to the other blocked request, and can be further merged. What's more, if all the interfering transaction requests are not controlled by any barrier, the two barriers can be moved, and the transactions in between can be reordered, and then the adjacent barriers can be merged. Combining barriers in this way can reduce the management costs associated with transmission blocking transaction requests and transmission responses.
In some embodiments, the aforementioned barrier management circuit system is configured to respond to the detection of a barrier transaction between two or more mergeable transactions, to copy the barrier transaction and to copy the copied barrier transaction One of them is placed on either side of the above two or more transactions and the above two or more transactions are combined.
Generally, transactions that can be merged but separated by a barrier transaction can also be merged, as long as the barrier transaction is copied to either side of the merged transaction. It should be noted that in addition to the merged transaction, there can also be other transactions between the above two barriers.
In some embodiments, the blocking management circuit system is used to detect the blocking transaction request located in the reordering buffer or on a node in the interconnection circuit system, and the node includes a path leading to a bisection In the above-mentioned interconnection circuit system, the halved path is the only communication path between the ingress node and an exit node of the halved path.
The management of blocking such transaction requests can be performed at an entry node leading to a bisecting path or in a reordering buffer. At these points, it can be confirmed that these blocked transaction requests have not been copied elsewhere, and therefore they can be safely merged.
In some embodiments, the above-mentioned control circuit system is used to mark transaction requests that are not controlled by any blocking transaction, and at least some of the above-mentioned transaction requests do not include the above-mentioned marked transaction request.
It is very convenient to be able to mark the transaction request so that the blocking transaction cannot be applied to it. This can be used to adapt to legacy systems that do not have barriers, and therefore use this method to mark transactions with legacy components so that they can function as expected by systems that support barriers. This can also be used to mark certain commands in the above manner when it is obvious that certain commands should not be delayed (for example, those with high priority).
A second aspect of the present invention provides an initiating device that can be used to send a transaction request to a receiving device through an interconnection. The initiating device includes: a blocking transaction request generator for generating blocking transaction requests. The transaction request may indicate to the above interconnection that the order of at least some transaction requests in the transaction request message flow through one of the above interconnections should be maintained, and the method of maintaining the order is to refuse to be earlier than the above blocking in the transaction request message flow. At least some of the above-mentioned transaction requests of the transaction request are reordered with respect to at least some of the above-mentioned transaction requests that are later than the above-mentioned blocking transaction request in the above-mentioned transaction request message flow; wherein the above-mentioned blocking transaction request generator is used to provide an indicator to the above The generated blocking transaction request is used to specify which of the transaction requests in the message flow of the transaction request includes at least some transaction requests whose order needs to be maintained.
In some embodiments, the initiating device includes a barrier transaction generator, and the barrier transaction generator is configured to respond to detection of the output of a highly ranked transaction request to generate and output a barrier transaction .
A highly-ranked transaction generated by an initial device has the effect of suspending the output of subsequent transactions by the master, because a highly-ranked transaction may need to be carried out in the order presented in the initial transaction message stream, and must not be compared to other transactions. rearrange. For example, this is the case in the AXI system, because AXI allows transactions to be reordered and also provides independent address writing and address reading channels. Therefore, when outputting to an interconnect, the initiating device will usually be paused to avoid any potential reordering hazards. The embodiment of the present invention provides a way to ensure the safe operation of the initiating device without having to suspend the initiating device by generating a barrier and the barrier is transmitted after the highly ranked transaction. At this time, since it is known that subsequent transactions cannot be reordered relative to other transactions, subsequent transactions can be safely transmitted. Obviously, a barrier has its own latency. However, due to the early response generation and other barrier management tools, this latency is significantly lower than the latency caused by the suspension of the initial device.
In some embodiments, the initiating device further includes: a hazard unit for storing prominent transaction requests that may cause a hazard; and an output terminal for outputting the transaction request to the interconnection; wherein the initiating device is It is configured to respond to the detection that the above-mentioned harmful unit is full, so that the output of further transaction requests can be suspended, so as to generate and output the above-mentioned harmful unit that blocks the transaction and clears at least one transaction request.
As mentioned above, the hazard unit in the data processing equipment is designed to be small in size, and therefore, it can only store a limited number of transaction requests. When they are full, the device must be suspended until at least one of the stored transaction requests is removed, at which point it can continue to be processed safely. This obviously increases the latency of the system. The embodiments of the present invention deal with this problem by providing a blocking transaction request generator in the device. This method can respond to detecting that the hazardous unit is full and issue a barrier and this allows the initiator to remove any stored transactions related to the barrier from the hazardous unit. Therefore, this allows the initiating device to continue processing and solves the need to suspend it.
In some embodiments, the blocking transaction request is a global blocking transaction request related to all transaction requests, and the initiating device is used to clear the harmful units of all transaction requests.
In some embodiments, the barrier transaction request is a global barrier transaction request, and therefore, issuing this barrier means that all transactions in the hazard unit cannot cause a hazard, and the entire hazard unit can be cleared.
In other embodiments, the blocking transaction is related to the address of one of the transaction requests stored in the harm unit, and the harm unit is used to clear the stored transaction request from the harm unit.
Globally blocking transaction requests has potential problems related to it. Therefore, for the sake of system performance, in a more favorable situation, a certain address blocking transaction request can be output, so that the blocking can stop the transaction request to the specific address, but Other transaction requests can continue. The latent time associated with this barrier is much smaller than the latent time associated with global barriers. The disadvantage is that in the hazard unit, only one or more transaction requests related to this address can be deleted. However, this can free up at least one space and allow the initiating device to continue processing and stop the above-mentioned pause operation until the hazard unit is full again.
In some embodiments, the blocking transaction is related to the address of one of the transaction requests newly stored in the harm unit, and the harm unit is used to clear the newly stored transaction request from the harm unit.
A third aspect of the present invention provides a data processing device, which includes: at least one initiating device according to the second aspect of the present invention for issuing a transaction request; at least one receiving device for receiving the foregoing transaction request; And the interconnection according to the first aspect of the present invention is used to connect the at least one initiating device to the at least one receiving device.
A fourth aspect of the present invention provides a receiving device for receiving a transaction request through an interconnection, the transaction request includes a blocking transaction request, and the blocking transaction request is used to maintain the order of at least some transaction requests in a transaction request message flow , Which is by rejecting at least some of the above-mentioned transaction requests that are earlier than the above-mentioned blocking transaction request in the above-mentioned transaction request message flow, as opposed to at least some of the above-mentioned transaction requests that are later than the above-mentioned blocking transaction request in the above-mentioned transaction request message flow, And reordering; the blocking transaction request includes an indicator to indicate which of the above at least some transaction requests is maintained; the receiving device includes: a response signal generator, the response signal generator for responding to the blocking A transaction request to generate and transmit a response to the blocking transaction request, and to respond to a predetermined indicator in the blocking transaction request to delay the transmission of the response until the receiving device has at least partially processed the blocking transaction request At least one transaction request of the aforementioned transaction request previously received.
A receiving device can respond to receiving the blocking transaction request to generate and transmit a response. However, in some cases, the nature of the blocking prevents the receiving device from responding immediately, but waits until at least some processing of the previously received transaction has been completed or at least partially completed. This method is very useful in certain situations. For example, in one situation, a data synchronization blocking transaction requires not only that all previous transactions have reached its final destination, but also that its processing has been completed. In this case, a response is sent only when the above conditions are all established, which enables the initiating device and interconnection to safely process further transactions once the response is received.
A fifth aspect of the present invention provides a method for sending data from at least one initiating device to at least one receiving device through an interconnection circuit system. The method includes: receiving a transaction from the at least one initiating device at at least one input terminal Request; send the above-mentioned transaction request along at least one of the plurality of paths at least one of the radial at least one output; respond to the receipt of a blocking transaction request to: a transaction request transmitted along one of the above-mentioned at least one path In the message flow, the order of at least some transaction requests relative to the above-mentioned blocking transaction request is maintained by rejecting at least some of the above-mentioned transaction requests that will be earlier than the above-mentioned blocking transaction request in the above-mentioned transaction request message flow, as opposed to the above-mentioned transaction At least some of the above-mentioned transaction requests in the request message flow that are later than the above-mentioned blocking transaction request are reordered; wherein the above-mentioned blocking transaction request includes an indicator that indicates which of the above-mentioned transaction requests in the above-mentioned transaction request message flow includes At least some of the above transaction requests whose order needs to be maintained.
Figure 1 shows an interconnection 10 according to an embodiment of the present invention. The interconnection 10 connects a plurality of master stations 20, 22, 24, and 26 to a plurality of slave stations 30, 32, 34, and 36 via a plurality of paths. These paths may have cross-coupling parts (for example, 40 illustrated in the figure), where the two paths are divided into two paths at individual branch points 41 and 42 respectively, and merged at merging points 44 and 45 respectively. There may also be halved paths (such as 50 illustrated in the figure). These paths are the only connection between two nodes in the interconnection, so that cutting off the path will essentially divide the interconnection. For two.
When transactions move along these different paths, the nature of the paths (that is, whether these paths are cross-coupled or bisected) will affect the order of these transactions. For example, the starting point of a cross-coupling path is a branch point, which can split the transaction information flow into multiple transaction information streams, and a transaction that is located before the branch point after another transaction may be larger than the original one. The transaction arrives at its own destination earlier before it reaches the destination. Transactions moving along the bisecting path must maintain their order, unless some functional units allow reordering, such as a reordering buffer (for example, 60 as illustrated in the figure). The reordering buffer can be used to reorder transactions to allow higher priority transactions to be transmitted from the station before a lower priority transaction.
There are also some so-called re-convergence paths, in which previously separated paths converge, and this may also cause reordering within the transaction message flow. The interconnect 10 does not have any re-convergence paths.
The fact that the order in which multiple transactions arrive at their individual destinations may be different in their delivery order, which may cause some problems for subsequent transactions that depend on the previous transaction (and therefore the previous transaction must be completed first). For example, if there is a save instruction to the same address in the transaction message stream before a load instruction, it is very important to complete the save before loading, otherwise the load action will read incorrect Numerical value. In order to allow the programmer to ensure that the required transactions arrive at the destination in the required order, the interconnection 10 can be configured to respond to the blocked transaction in the transaction message flow, so as to keep the transaction in the interconnection relative to the blocked transaction. order. Therefore, a barrier transaction can be inserted between transactions that should not pass each other, and this can ensure that there is no reordering problem.
The interconnection can respond to the barrier transactions by delaying the transaction after the barrier transaction in the transaction message flow transmitted in the interconnection until a response signal to the barrier transaction has been received. The response signal indicates that it is safe to send a follow-up command. It should be noted that a response signal that can clear a path may be a signal indicating that the earlier transaction has been fully completed, or it may only be a signal indicating that the blocked transaction has been transmitted along a path (if, for example, , The path is a bisected path) signal, or the blocking has reached a node that transmits an early clear response message and can be blocked again.
The interconnection can also only transmit the barrier transaction after the previous transaction along various paths, so that when it is detected that the barrier reaches a specific point, the interconnection can be sure that all previous transactions have passed this point. It does not matter whether it is only transmitting the barrier in the message stream, or delaying the transaction based on the nature of the barrier and whether it is a blocking barrier based on the nature of the barrier.
Blocking means that the transaction behind it in the transaction information flow under its control is blocked somewhere upstream, and therefore, other transactions can surpass the blocking block, because they must not be transactions that need to be kept behind; However, the barrier itself cannot overcome any transaction requests that are located in front of it and under its control. The early response unit can be used to turn on the blocking block. See below for details.
Non-blocking blocking means that no transaction request is blocked, and therefore, the transaction request under its control must be on the correct side of the blocking transaction request. Because there is no upstream blocking, it cannot be switched on with the early response unit. As will be understood from the following discussion, blocking or non-blocking indicators can be used to indicate the different nature of barriers; or, a system can only support one type of barrier, in which case no indicator is needed. Or, a barrier can be a preset barrier type, and in this case, only another type of barrier has an indicator. Or, in the interconnection, all barriers in some areas may behave as blocking barriers, and all barriers in some areas may behave as non-blocking barriers. In this case, the interconnection can be set so that the barrier does not carry any indicator, but the interconnection can handle the barrier in a specific way based on its position in the interconnection.
The process of blocking transactions is controlled by the control circuit system 70. In this figure, this situation is schematically presented with a single block. However, in practical applications, the control circuit system can be dispersed throughout the interconnection and adjacent to the circuit system controlled by it. Therefore, at each branch point, for example, there will be some control circuit system to ensure that at least in some implementations, when a blocked transaction is received, it will be copied, and that the point will be copied Block the transaction down to each exit path. In other embodiments, the replicated barrier will be sent down to all exit paths except one, which will be described in detail below. The control circuit system can be aware that the barrier transaction has been copied, and therefore, can request from each copied barrier transaction before it empties the path used to transmit the transaction (which is located after the original barrier transaction and must remain behind) The response signal.
In the simplest form, a master station (e.g., master station 20) issues a blocking transaction, and then master station 20 blocks all subsequent transactions until it has a response signal from the interconnection (which specifies the master station Subsequent transactions can be transmitted). Or it can be blocked by a control circuit system immediately adjacent to the entry point of the interconnection. The transaction before the barrier transaction and the barrier transaction are transmitted to the interconnection, and the control circuit system 70 controls the sending of these transactions. Therefore, the blocking transaction is copied at the branch point 42 and moved to the merge point 44 and 45. At this point, the transaction enters the bisecting paths 50 and 52, and because these transactions cannot change their position on these paths relative to a barrier, when the barrier transaction reaches the starting point of one of these paths , You can know that all the transactions before it are in front, and will still be in front of it along the path. Therefore, the empty unit 80 can transmit an early response signal, and can respond to the reception of all these signals. The control circuit system at the branch point transmits the early response signal to the master station 20, and then the master station 20 can connect to the receiver. The control blocks the transaction after the transaction and transmits it to the interconnection.
By providing the early response unit 80, the master station 20 can be blocked for a much shorter period of time (compared to having it wait for responses from multiple slave stations to indicate that the blocked transaction has reached the slave station), and so In the future, the potential for blocking the introduction of transactions can be reduced.
The barrier transaction transmitted along the path 50 leaves the interconnection and reaches the slave station 30 without passing through any other path except the bisecting path 50, and therefore, there is no need to respond to the barrier transaction and perform the barrier again, because When the barrier has passed the empty cell, the transaction in front of it must still be in front of it. However, the barrier transaction transmitted along the path 52 will reach the further cross-coupling part 48, and can respond to the barrier transaction being received at the branch point 49, and the control circuit system 70 associated with this branch point will copy the barrier transaction and transfer It passed down to two exit routes and blocked the subsequent entry routes for the transaction requests located behind and under its control. Therefore, in some embodiments, by keeping these subsequent transactions in a buffer in the blocking circuit system 90, these subsequent transactions are delayed until a response signal is received for all the copied blocked transactions . Therefore, the replicated barrier transaction passes through the cross-coupling circuit system 40 and leaves the cross-coupling circuit system to join further bisecting links 52 and 50. As described above, the bisecting path can maintain the order of transactions relative to the barrier, and therefore, the clearing unit 80 can be used to transmit an early response from the beginning of the bisecting path. The blocking circuit system 90 will wait to receive a response from the branch point 49 to the blocking transaction. The branch point 49 replicates the barrier transaction, and further transmits the second barrier transaction down to each path. The branch point 49 will not send a response back to the blocking circuit system 90 until it receives a response to the two blocked transactions transmitted by it respectively. In response to this response, the blocking circuit system 90 allows any subsequent transactions stored in its buffer to be transmitted. When the empty circuit system is located on the last bisecting path before leaving the interconnection, there is no need to issue further barriers for certain barrier types.
As described above, there is a reordering buffer 60 on the bisecting path 52, and this buffer is configured to respond to the block and not allow transactions controlled by the block to be reordered relative to the block.
In the above description, it is assumed that there is a barrier to keep all subsequent transactions behind. However, in some implementations, as detailed below, the barrier may only need to prevent a subset of the subsequent transactions from overtaking it. The subset may come from a specific master station or have a specific function (e.g., Write transaction). In this case, the control circuit system 70 and the blocking circuit system 90 will only delay this subset of transactions, and will allow other transactions to proceed. What's more, at a branch point, if the transaction is controlled by the barrier, the transaction will never be passed down to one of the paths, and there is no need to pass a copied barrier down to the path.
In this regard, the barriers marked as being related to write transactions can essentially block further writes so that no further writes will be issued until a response is received. Because of this nature, the barrier does not need to block reads until a response to the write is received. No further writes can be sent before the response; and therefore, reads can be sent safely, once received A response to the write barrier can send out further writes, and at this point the barrier should also block reading.
Figure 2 shows a second interconnection according to an embodiment of the present invention, which interconnects the master stations 20, 22, 24, and 26 to the slave stations 30, 32, and 34. In this case, the interconnection 12 is already connected to the interconnection 10, and this creates a re-convergence path for some of the transactions (transactions that leave the interconnection 10 and enter the interconnection 12). If the interconnect 10 is directly connected to the slave station, the empty unit 80 can be used to transmit an early response, because the subsequent path is a bisected path and is connected to the outlet of the interconnect, and no further blocking is usually required.
However, if the transaction does not go to the destination device but is transmitted to a further interconnection, it may not be appropriate to block it without responding to blocking transactions traveling along these paths; and if there is A further cross-coupling path or a re-convergence path as described herein can be appropriately connected to a further interconnection to introduce this path to change the way the interconnection responds and blocks early. Therefore, in order to make the interconnection suitable for different purposes, it can include a programmable blocking circuit system 92, which can be programmed to be turned on or off, depending on whether the interconnection is directly connected to a slave station or has Another interconnection of a cross-coupling path or a re-convergence path.
Therefore, the programmable clearing and blocking circuits 82 and 92 are controlled by the control circuit system 76, which can respond to the connection of the interconnection 12 to the interconnection 10 and block any blocking transactions passing therethrough. In this way, the subsequent transaction is not allowed to pass until the response signal is received and its order is ensured.
As mentioned above, blocking transactions can block subsequent transactions under its control, and the above-mentioned transactions can be all transactions or a subset of transactions. The transaction controlled by the specific blocking exchange can be determined by the blocking transaction itself or by the subsequent transaction.
In the embodiment of the present invention, the barrier transaction is designed to look the same as the other interconnected transmissions, and in this way, the interconnection can be used to process these barrier transactions without major redesign of its components. Figure 3 shows a summary of the blocked transaction. In this embodiment, the address field contained in the blocking transaction request is similar to a general transaction request, but these address fields will not be used except for the address blocking described below. It also contains a size column to specify the size of the address range covered by a certain address blocking the exchange; an indicator column that contains indicator bits to specify the various properties of the blocked transaction (for example, what kind of trade). There are two types of blocked transactions, and each bit of the indicator bit has two values to indicate these two different blocked transactions. The above two values can indicate to the interconnection that the transaction is a barrier transaction, and also indicate what type of barrier transaction it belongs to. Therefore, they can specify a System Data Synchronization Block (DSB) transaction, which can be used to separate the transaction, which occurs when the previous transaction has been completed and reached its final destination, and is allowed to send out this data Before any transaction after synchronization block (program sequence). Therefore, the master station can block subsequent transactions by responding to the indication bit indicating such blocking, and there is no chance of an early response to this transaction, and therefore, these blocking transactions will cause significant latency.
It should be noted that since the master station will block subsequent transactions from entering the interconnection until it receives a response from the system DSB, and since the response must come from the destination, other transactions can overtake the DSB because they do not Need to stay after any transaction. However, once they get past the DSB, they can interact with other transaction requests controlled by the barrier. At this point, the barrier becomes relevant to these transactions, and therefore, although they can get past the barrier, if they do , After which they must remain in front of them.
Another type of barrier transaction is the data memory barrier transaction DMB, and transactions controlled by these barriers are not allowed to be reordered relative to this barrier. These barriers are only related to the order of the transaction, not to the nature of its advancement in the system. Therefore, these transactions can take advantage of early responses from clearing units, and these technologies can be used to reduce the potential for blocking transactions.
There are other fields in the blocking transaction, one is the identification or ID field, which is used to identify the main site that generated the transaction; the other is the sharability area column, which can indicate which sharability area the block belongs to (detailed See below); and one is a blocking indicator, which can be set to indicate whether the transaction should be considered blocked. If the indicator is set, it means that the transaction has been blocked, and therefore, clearing the unit and the blocking unit (such as those shown in Figures 1 and 2) will allow the blocking transaction to pass, and no transaction will be made. Respond because they know that subsequent transactions have been blocked upstream. The purpose of this approach will be described below. However, it should also be noted that if there is a re-convergence path (e.g., in the device shown in Figure 2), a blocking unit (e.g., the blocking unit 92 in Figure 2) will The blocking indicator reacts to the blocking transaction, and will block in response.
In some embodiments, the blocking indicator does not exist on the blocking transaction, and where it is determined that the blocking transactions are not blocking, all blocking transactions are treated as blocking blocking transactions.
There may be other control fields for the blocking transaction to indicate whether it only controls transactions with a specific function. Therefore, only one barrier transaction related to the write transaction can be obtained.
The address column can be used to indicate to the block that only transactions to a specific address or range of addresses are controlled. In the latter case, the address field stores a base address, and the size field stores the size of the address range. The advantage of address blocking is that it can control a very specific subset of transactions without slowing down other transactions. What's more, when a transaction is copied at a branch point, if it is known that the address or address range cannot be accessed in one or more paths of the one-port path, the address barrier does not need to be copied, which can reduce mutual The latency of the company and the management cost in the barrier. Since the blocking transaction is very similar to other transactions, there will be an address column, and therefore, the address information can be directly provided in a certain address blocking transaction.
Figure 4 shows that the interconnection transmits a transaction that is not a blocking transaction, and as shown in the figure, its form is very similar to a blocking transaction. There is an indicator column in the locked bit, which can specify whether the transaction should ignore the blocking order transaction. Therefore, marking a field like this in a transaction will allow it to automatically pass any block. This is very useful when the legacy system is used with an interconnect that supports barriers.
As described above with reference to Figure 3, there are areas of sharability, which is a way of segmenting the effect of barriers. This is a method that can further improve the latency of the system, and it will be described in detail with reference to the following figures.
There are many ways to group the main station and the interconnected parts into different areas and control the barriers and transactions in this area to ensure that the correct transaction sequence can be maintained without excessively increasing the latency. In this regard, it is known that if certain rules are applied to the layout of the interconnected circuit system components related to the area, the barrier will have certain available properties to reduce the potential of the barrier. Arranging the area in a special way can limit the allowable circuit system component layout of the system, but it may also increase the potential for blocking. Therefore, the implementation of the various possible areas described below has its own advantages and disadvantages. .
In all regional arrangements, if a barrier transaction is marked as related to a specific region, when it is located outside the region, the transaction can always be connected (except in the repeated convergence region). In its relevant area, a DMB can be connected (except in a cross-coupling area), and in its relevant area, a DSB is always blocked. A system DSB is marked as relevant to the entire interconnection, and therefore, it will never be outside of its relevant area and can always be blocked until a response is received by its destination.
In the first "zero point" implementation method, no part of these areas is selected. Treat all barriers as applicable to all transactions in all parts of the system. Obviously, the effectiveness of this method is low, because the latency caused by the barrier will be very high. However, this approach allows unrestricted and arbitrary regional membership (even if the membership is not effective) and the layout of circuit system components, and thus can always be established. This is logically equivalent to all regions, including all master stations in all regions.
In an alternative "near zero" implementation, there are non-shared areas related to each master station, and outside this area, different methods are used to deal with the barriers related to these areas. When a non-shared barrier is located outside of its shareability area and beyond any position from the input of the master station, it can be treated as being located in the entire interconnection, and therefore in all re-convergence parts of the interconnection Both are non-blocking. Treat other sharable area barrier transactions as being in a zero-point implementation method. This is logically equivalent to making this non-shared area boundary as the sender or master input, and all other areas contain all other masters.
An alternative "simple" implementation has certain restricted circuit system component layouts and better performance. This approach will result in two different solutions, depending on the degree of acceptable line restriction.
All of the above methods can use the following three restrictions on shareability zone membership.
1. The non-shared area of an observer is itself separate.
2. The system shareability area of an observer includes all other observers that can communicate directly or indirectly with it.
3. All members of the internal common area of an observer are also members of the external common area.
The first two of the above restrictions are imposed by the third restriction. In addition, the above two solutions have specific component layout restrictions and possible additional shareability area membership restrictions.
The first of the above two implementations requires a restriction to require each location to be located in a single area, and therefore, it depends on that each location in the interconnection is only one type of regional area (internal, external or system). In order to achieve this goal, an additional shareability zone membership restriction must be implemented:
All members of any sharability area of any observer must regard all other members of the sharability area as members of the same level of sharability area. That is, if observer B is a member of the internal shared area of observer A, relatively, A must be a member of B's internal shareability area.
The component layout restrictions that must be met are as follows:
1. The boundary of the area must include all members of the area
2. Nothing outside of an area can be merged into the area, that is, the boundary of the area must not contain anything located under anything that is not in the boundary of the area
3. All regional boundaries must be located on the regional bisecting link
To explain in a simpler way, in this case, the area boundary can be regarded as the contour line of the component layout that represents the height (where vertical surfaces are allowed but no protruding parts are allowed). Each master station is located at the same height, and the outline of each shareability area is also at the same height as other types of the same type. Vertical screwing may be allowed so that different types of shareability areas are the same, but protrusions that may cause shareability areas to intersect are not allowed.
These component layout restrictions require that nothing can be merged into this area-neither members of the area (which would violate restriction 1) nor non-members (which would violate restriction 2). If a branch downstream of a member leaves the area and later merges into it without merging something outside the area at the same time, the no tomb between leaving and re-entering is still effectively located in the area .
The component layout and regional membership restrictions ensure that in its shareability area, a barrier will not encounter transactions from an observer located outside the area, and when it leaves the area, it has merged All transaction information flows of all members in the area merged with it. They can also ensure that any location outside of any internal shared area is outside of all internal shared areas, and if located outside of any external shared area, it is also outside of all external shared areas.
Because of this, it is only by comparing the shareability area of the barrier with the area type of the area where the branch point is located to determine whether a barrier is a requirement for a branch point, because the barrier Being outside of the area of this restricted system implicitly meets the requirement that members who do not have the shareability area at that location can merge.
This mechanism can be implemented in any of the following ways: clearly indicating that the barrier is outside its shareability area, which would require an explicit detection component at the exit point of the area; or at each relevant branch point Determine the status.
The second of the above two implementations allows multiple locations in multiple regions. This method of implementation depends on the modification of a barrier transaction when it passes through the boundaries of the specified shareability zone, so that once it leaves its shareability zone, the doctrine can be carried out. And make it non-blocking. When it passes through an internal or external shared area of the designated area and faces the non-shared area and marks it as unshared, it can be known that it is located outside of its area and thus can be non-blocking.
In this case, the additional restrictions on the membership of the shareability zone are more relaxed:
For any two sharable areas (A and B), all members of either must also be members of B, or all members of B must also be members of A, or both of the above conditions are true (in this case In, A and B are the same). In other words, the boundary of the area does not cross back.
The layout restrictions of the same circuit system components are required:
1. The boundary of the region must include all members of the region
In order for the layout of circuit system components to have the greatest flexibility, it must just be able to deconstruct the layout components (branch and merge) of the circuit system components so that the regional boundaries can be outlined, so that
2. Nothing located outside an area can be merged into the area-that is, the area boundary must not include anything downstream of anything that is no longer within the area boundary
3. The boundary of the region will cross the region bisecting link
Finally, an additional component layout restriction will be implemented to compensate for problems caused by looser regional membership restrictions:
4. There are no border locations available for different master stations with different numbers of zones (excluded here to master stations that are already outside their external shareability zone).
Limit 4 ensures that when a barrier needs to be modified because it crosses the boundary of an area, it crosses the boundaries of all areas in which it is located. This ensures that the modification operation does not differ depending on the starting device of the barrier.
If a barrier is modified and obtains a non-blocking state, it can of course be turned on (if it is located on a bisecting link), but in some cases, even if it is located on a cross-coupling link, it can still be connected. Turn it on. If the link that crosses the boundary of the region is a bisected link of the region (that is, it is bisected in terms of the region, that is, they will not merge with the path from their own region, but only with other Regional path merger), the blocking transaction can be modified there, and the above connection can also be performed at this point.
In the following cases, restriction 2 can be omitted. If you change the pointed area toward the non-shared area at the exit of the area, you can change the pointed area away from the non-shared area located at the area entry point. This requires a non-saturated area indicator; or a restriction on the number of accessible areas, so that saturation does not occur. In addition, this can cause barriers that have entered an area to be able to block transactions from non-members of that area (because of their increased range).
Figure 5A schematically illustrates the implementation of the above-mentioned areas within an interconnection. The master stations depicted in Figure 5A are located within the interconnection, but in reality, they can of course be located outside the interconnection. Each master station 20, 22, 24, 26, 28 has a message stream or non-shared area 120, 122, 124, 126, 127 directly around it. These message streams or non-functional areas may only be connected to the master station. The resulting transaction is related. Then, there are some lower-level areas, which may contain multiple master stations or the same master station. Therefore, similarly, the master stations 20 and 22 have their non-shared areas, and then have an internal area 121 located around them. The master station 24 has an internal area 125, the master station 26 has a non-shared area 126 and an internal area 127, and the master station 28 has a non-shared area 128 and an internal area 129. Next, there is an outer area set around it. In this embodiment, the aforementioned outer areas are areas 31 and 33. Then there is the system area, which is the entire interconnection. As shown in the figure, these areas are completely within each other and do not intersect in any way. Another limitation is that all exit paths from these areas are bisected paths. By restricting these areas in this way, it can be determined that the trades leaving these areas will leave in a specific way, and when they leave the bisecting path, if the barrier is functioning normally in the area, they will be in the correct Leave in order. This makes it possible to control and block transactions with respect to these areas in a specific way.
FIG. 5B schematically shows an exit node 135 heading to an area including the main stations p0 and p1. This exit node 135 is under the control of the control circuit system 70, and at this point, it is known that any blocking transaction and the transaction under its control are in the correct sequence. Now, as mentioned above, blocking transactions does not necessarily control all transactions, but can control transactions generated by specific master stations or transactions that belong to specific functions.
In the case of shareable areas, blocking transactions are marked as controllable transactions from specific areas. Therefore, a transaction can be marked as a system blocking transaction because it controls all transactions; it can be marked as controlling transactions from a message flow or non-shared area, from an internal area, or from an external area. In any case, when a barrier transaction leaves an area, in this implementation, the barrier transaction can reduce this level, so that if it is an external area barrier, when it leaves the internal area, it can be removed Demotion is a barrier transaction that controls transactions from an internal area, and when it leaves the external area, it will demote the class under its control to a non-shared area where no transactions will be delayed. The above goal can be achieved because at this point, the order of all transactions is scheduled relative to this barrier, and provided that there is no re-convergence path, the interconnection can determine that the order is correct. It should be noted that when the system barriers isolate the area, they will not change, because they will always apply to anything located anywhere.
It should be noted that if there is a re-convergence path in an area, any non-blocking barrier must become blocking when crossing the re-convergence area. If a further interconnection introduces a re-convergence path, and connects this interconnection to an interconnection with these areas, the system with the controlled blocking area will not work. If an interconnection that affects these areas and their hierarchy is added, the system should be controlled so that when the barrier transaction leaves the area, the sharability area indicator in the barrier transaction will not be downgraded.
It should be noted that for the re-aggregation area, certain transactions to a specific address can be restricted to pass through the re-aggregation area along a specific path, and in this case, the re-aggregation area is for the address Rather than converge, an interconnection can be restricted so that for all addresses, transactions will travel along a specific path to a specific address. In this case, any re-convergence area can be treated as a cross-coupling area. , The advantage of this comes from the fact that the repeated convergence area will cause considerable restrictions on the system.
Limited by the way of interconnection configuration, any combination of barrier transactions in a region that is not marked as a non-shared barrier can actually control transactions in any region it encounters, because it will not encounter transactions from another region. trade. A barrier transaction marked as a non-shared barrier will not delay any subsequent transactions, however, no other transactions can be rearranged relative to this transaction. In this way, by arranging the interconnected areas in this way, and by lowering the level of the indicator at the exit of the area, a simple method is provided to determine whether the barrier transaction must delay its encounter All transactions may not delay transactions, and the control component does not need to know exactly which area of the interconnection these transactions are located.
Another possible area implementation method is the "complex" implementation method. It is used in situations where the above component layout restrictions or area membership restrictions are considered too strict. Assuming that restrictions on non-shared and system area membership should be maintained, the information required is a list to clearly list what combination of blocking emitting devices and shareability areas can be considered non-blocking at that location . Therefore, here it is not the barrier itself that determines the blocking properties of the barrier (for example, refer to the implementation method described in Figures 5A and 5B), but the location and the area information stored in the location to determine the barrier of the barrier. Cut the essence.
This can be achieved by two lists at each relevant location, one list is used for the internal shared area, and the other list is used for the external shared area. Each list specifies the group or blocking source located outside the area. Alternatively, a list can be stored in which a two-digit value is used to indicate which shareability area of the source the location is outside of.
Regardless of how the above information is expressed, it is obviously more complicated and difficult to allow the design to be reused, because when the system is reused, the requirements for expressing the area information are different.
Figure 6 shows an example of this interconnection. In this interconnection, four master stations S0, S1, S2 or S3 receive transaction requests. S0 and S1 are located in an inner area 200, and S2 or S3 are located in an inner area 201, and the four master stations are all located in an outer area 202. There are other master stations that are not shown and have other areas.
At location 210, it is located in the inner area for transactions from S2 and in the outer area for transactions from S0 or S1. Therefore, this position can be marked accordingly, and when a barrier is received, it can be determined which area the barrier is related to, and therefore, it can be determined whether the barrier is located outside of its area. Therefore, a barrier suitable for the inner region of S0 and S1 is located outside of the region, and depending on the implementation, the barrier can be marked as such or an early response can be transmitted. Obviously, this is very complicated.
An alternative to this is the conservative and complex implementation. This implementation mode is used when a complicated implementation method is required for component layout and area membership freedom, but it is necessary to avoid the problem of the implementation method and reuse. In this case, each component that must exhibit a specific area-location-specific behavior can be deemed to belong to a specific area level, and the correct behavior can be achieved. If the component believes that it belongs to the smallest area in any area where it is actually located, it will be conservative (but correct) for the blocking behavior that is actually located outside of its own area, and it will be conservative (but correct) when it is actually located The behavior of blocking within its own area is shown correctly. In this regard, it should be noted that the nature of barriers, areas, or transactions can be changed where appropriate, which has enabled them to be handled in a more efficient manner, provided that their changes become more restrictive. Therefore, a barrier marked as internal can be treated as an external barrier, and a transaction marked as suitable for an external area can be marked as suitable for the internal area.
Using this method, these components that must be aware of the zone can be easily programmed, or set to have a single zone (which has a safe default internal zone membership, which can be applied when the power is turned on).
Therefore, in this implementation method, a location in the area is marked as having the properties of the area with the most restrictive behavior in the area where it is a member, and the above-mentioned area is excluding the non-shared area. The lowest level of this area. After that, blocking transactions located at this location will be treated as if they are located in this area. In this arrangement, the area is allowed to become a subset of other areas. In this arrangement, instead of changing the mark on a barrier to adjust the blocking behavior of the barrier when the barrier separates the area without knowing where it is located in the interconnection, it is The location is marked as being in a specific area, which depends on the lowest level system or the smallest common area where it is located.
In the embodiment shown in FIG. 6, for example, it is not necessary to mark the position 210 with three different marks, and only mark it on the smallest common area that belongs to the interior. Therefore, in this case, any barrier marked as internal or external will be regarded as being located in this area, and a barrier from the internal area of S0 and S1 will be regarded as being within its area, even if it is not. Therefore, the early response cannot be transmitted, and the potential for interconnection will increase, which is the disadvantage of this method. However, the labeling of the area is relatively simple, and the determination of whether a barrier is located in the area is also relatively simple.
Figure 7 shows a further interconnection 10, which has a plurality of master stations 20, 22, and 24 and a plurality of slave stations 30, 32. In the meantime, there is a cross-coupling circuit system 40 which is not shown in detail.
This diagram clarifies that the barrier indicator element on the barrier transaction is used to indicate to the interconnection that the barrier has already been interrupted upstream, and that the interconnection does not need to be further blocked. Therefore, in this case, the transaction issued by the GPU 22 is a transaction that is of no importance to the latency, and therefore, a blocking unit 90 will react to the blocking transaction issued and block the subsequent transaction to the blocking , And mark the blocked transaction as a blocked transaction. This means that there is no early response unit in Interconnection 10 that will respond early to this blocking transaction, and will not respond to it and block all subsequent blocking units. Therefore, the blocking transaction will remain blocked until the blocking transaction itself reaches its final destination, at which time a response signal will be transmitted to the GPU 22, and the blocking circuit system 90 will remove the blocking. It should be noted that although the blocking circuit system 90 shown here is located within the interconnect 10, it can also be located within the GPU 22 itself.
The advantage of this is that, as mentioned above, when the potential of transactions from a master site is not important, blocking these transactions at a nearby source means that the blocking transaction will not be blocked by other masters in the cross-coupling area. The transaction sent by the station caused a latent loss to these other master stations. It should be noted that if there is a re-convergence path in the interconnection, the barrier unit in the interconnection may need to be blocked in response to a barrier transaction to ensure the correct order.
In an embodiment that does not use such a barrier indicator, all barriers located in its area are regarded as blocked, unless they are marked as non-blocking in some way.
Picture 8 shows commonly used in Cambridge United Kingdom AXI<sup>TM</sup>ARM of the bus<sup>TM</sup>A connection10. One interconnection10. These buses have parallel read and write channels, and therefore, can be regarded as not having bisecting paths. However, transactions transmitted along these paths are usually linked together, and one of the paths is connected to a transaction along the other path (to a barrier transaction being transmitted). These paths can be regarded as halved paths. Processing, and can be used and early response and subsequent blocking to reduce the potential for blocking transactions. Therefore, in this diagram, a path is shown in which an early response unit 80 generates an early response to a blocked transaction. It also provides a linked transaction that can be sent down to other paths parallel to the path of the blocking transaction. The subsequent blocking unit 90 will not further transmit the blocking transaction until the linked transaction is received. It should be noted that when the barrier may be merged (so it is located at the merge point (e.g. 91) or within a slave station), this kind of transmission of the linked transaction will be required and only when both transactions reach their destination The way to respond. Therefore, when the blocking transaction is transmitted from the blocking unit 90 to the slave station S0, a connected transaction needs to be transmitted, and a response is transmitted only when the blocking transaction and the transaction connected to it have been received. The path used to pass down the blocked transaction is the only path that is actually blocked.
It should be noted that if the read and write message streams are to be merged, the order of the message streams must be sorted at the merge point. Therefore, there must be some control mechanism to ensure that the blocking and the connected transaction reach the merge at the same time point.
Figure 9 shows an interconnection 10 according to an embodiment of the present invention, which includes a plurality of input terminals to receive signals from the master station 20, 22, and 24, and a plurality of output terminals to transmit the transaction request to a plurality of slaves Station (including memory controller 112).
The path used to transmit the transaction includes a bisecting path and a cross-coupling path. The interconnection is used to respond to blocking transaction requests to ensure the order of transactions relative to these blockings. There is also a combined circuit system 100, which is used to combine blocking transactions where appropriate to increase the efficiency of the interconnection. Therefore, if it detects that two barrier transactions are adjacent to each other, it will merge these barrier transactions into a single barrier transaction, and a response from the combined barrier transaction will cause the response signal to be sent to the two merged above Block transactions.
It should be noted that if the barrier transactions have different properties, they may not be suitable for consolidation, or these properties may need to be changed so that the combined barrier transaction has properties that enable both barrier transactions to function. So, for example, if the sharability area of one of these blocking transactions indicates that it can control the transaction from the inner area 1 and the adjacent blocking from the outer area 1 (including the inner area 1 and a further initial device) Transaction, the shareability area of the merged blocking transaction will be the outer area 1. In other words, a sharability area (including the sharability area of the two barrier transactions to be merged) can provide a merged barrier transaction that can function correctly. It should be noted that if there are three or more adjacent blocking transactions, they can also be merged.
The merging circuit system 100 can also merge a barrier transaction that is not adjacent but only has an interference transaction that is not restricted by the barrier. This can be done if the merging circuit system 100 is adjacent to a reordering buffer 110. A reordering buffer is usually used to reorder transactions, so that high-priority transactions can be placed before lower-priority transactions, and can leave to their individual slave stations earlier than the lower-priority transactions. It can also be used together with the merged circuit system 100 to reorder non-adjacent barrier transactions (there are only interfering transactions in which the barrier does not apply). In this case, the blocking transactions can be moved to be adjacent to each other, and they can be merged later. It should be noted that the combined circuit system must be located on the one-two-half connection, otherwise the combined barriers may lead to an incorrect sequence, which is due to the copied barriers moving in other paths.
An alternative to merge barriers is to eliminate them. FIG. 10 shows a circuit similar to that of FIG. 9 but with the blocking elimination circuit system 102. It should be noted that in some embodiments, the merging circuit system 100 and the elimination circuit system 102 are a single unit, and can be used to merge or eliminate barriers depending on the actual situation. The barrier elimination circuit system 102 can be used on any path, including cross-coupling and bisecting paths, and if it detects a barrier followed by another barrier, the previous barrier can be applied to the next barrier. For all non-blocking transactions, and there is no intervening non-blocking transaction applicable to the subsequent blocking, the blocking elimination unit 102 may delay the subsequent blocking until a response to all such previous blockings is received. Once these responses have been received, they can be sent upstream, and a response to the new block can also be sent upstream, and the block can be eliminated. In this way, a barrier can be removed by the interconnection circuit system, and this can increase performance.
This ability to do so has an additional advantage in that it can be used to manage peripheral components that are not very commonly used. FIG. 11 shows a data processing device 2 which has an interconnection circuit system 10, a power management unit 120, and a number of peripheral components 32 and 34 that enter inactive mode for a long time. The control circuitry 70 is used to control different parts of the interconnection circuitry 10.
The transaction request transmitted to these components during the non-active period of the peripheral components 32 and 34 can only be a blocking transaction. It is necessary to respond to these transactions, but if the peripheral component is in a low power consumption mode, it is very unfavorable to wake it up just to respond to a blocking transaction. In order to deal with this problem, a barrier elimination unit 102 can be used. A barrier elimination unit 102 is provided on the path leading to the peripheral elements, which makes it possible to use the barrier elimination unit 102 together with the control circuit system 70 to generate a barrier and transmit it to the peripheral elements. Once the response has been received, the peripheral components can enter the low power consumption mode, and then when a further barrier transaction is received and a response to a previous barrier transaction has been received, only a response can be sent to This further blocks the transaction without awakening the peripheral components, and can eliminate the block transaction.
The control circuit system 70 can cause the barrier elimination unit to generate a barrier transaction in response to the detected non-functioning situation of the peripheral component, and then can contact the power management circuit system to suggest that once a response has been received, let These peripheral components enter a low power consumption mode. Or, once a blocking transaction has been transmitted to the peripheral device, the control circuit system 70 can indicate this situation to the power management circuit system 120, and then the power management circuit system 120 can transmit a low power consumption mode signal to the power management circuit system 120. Peripheral components, and make them enter a low power consumption mode. Alternatively, if the power management circuit system determines that it is time to let the peripheral component enter the low power consumption mode, it can send a signal to notify the peripheral component and the control circuit system 70 at the same time. The entry into the low power consumption mode can be delayed until a blocking transaction has been transmitted to the peripheral component and a response has been received. At this time, the peripheral component can enter a low power consumption mode, and can respond to any subsequent blocking transactions without waking up the peripheral component.
Figure 12 is a flowchart showing the method used to actually move a blocking barrier through an interconnection by further transmitting an early response clear signal and barrier along the interconnection. As a response to a blocked transaction request, an early response signal (a blocked upstream can receive and then clear this signal) is transmitted, so that subsequent transaction requests controlled by the boundary transaction request can be further transmitted. After that, the circuit system itself further transmits the blocking transaction request, but blocks subsequent transaction requests controlled by the blocking transaction request to transmit the blocking transaction request. If there are several paths, it can replicate the blocking transaction request so that it can be transmitted along each of the several paths. The circuit system then waits for a response signal from each blocked transaction request that has been further transmitted. When it receives these response signals, it can connect to the blocked subsequent transaction request and allow it to continue to transmit. In this way, when the blocking signal passes through the interconnection, the potential time of the interconnection can be reduced, and the subsequent transaction request will not be kept at the master station until the blocking transaction request has been completed at the slave station. It should be noted that for a data synchronization blocking request, early response signals are not allowed, and these blockings will cause the interconnection to block subsequent transaction requests controlled by the blocking until each of the peripheral components sent by the blocking The person receives an empty response.
It should also be noted that if the blocking transaction request is a blocking request, an early response signal will not be sent in response to the transaction request. A blocking request is one that is blocked upstream and is also marked to indicate that the blocking should remain upstream. Since there is a block in the upstream, it can be ensured that the transaction that can be used as a block will not be transmitted through the upstream block until a response signal to the block transaction request is received, and therefore, there is no need to block the block transaction ask.
It should be noted that there may be some points in the interconnection circuit system that can transmit an early response to a blocking request, and the blocking will move downstream, and the reissued blocking itself is marked as blocking or non-blocking. Off. If this is non-blocking, it can indicate that there will be transactions that need to be blocked afterwards. This is not the case here, because they have been blocked, but this is acceptable; but if in fact there may be some When a transaction needs to be blocked, it cannot be accepted that it indicates that there is no transaction that needs to be blocked. This is very beneficial for the following situations: when a non-blocking barrier crosses a cross-coupling area and therefore must delay subsequent transactions, but does not want it to block all paths to the exit point, it must always be indicated as non-blocking , Which makes it possible to send an early response from the first position where the response can be made.
Therefore, if the request is not a data synchronization blocking transaction request or a blocking request, an early response is sent, and the copied blocking transaction request must be blocked for subsequent transaction requests related to the blocking transaction request before being transmitted. The exit path. In this way, the upstream path can be connected for subsequent requests, and these subsequent requests can be transmitted, as long as the new block that can reduce the system's latency can still maintain the required sequence. Later, these subsequent requests can be delayed at this point until the node is connected. It should be noted that in some cases, the layout of the circuit system components is set to provide an early response and blocking may actually have undesirable consequences. For example, when each upstream is divided into two, it may There will be no advantage due to the connection, and blocking at this point may be the first and unnecessary blocking.
Perform the following steps to connect the node. It is determined whether a response signal has been received for any of the copied blocking transaction requests that have been further transmitted. When a response signal is received, it is determined whether there is a further response signal waiting for the further blocking request. After all the copied blocking requests have been responded to, the exit path can be connected, and the subsequent transaction request can be further transmitted.
Figure 13 is a flowchart showing the steps of the method to eliminate blocking transaction requests.
A blocking transaction request is received at the start of a bisecting path and further transmitted along the bisecting path.
Then, a subsequent blocking transaction request is received, and it is determined whether any transaction request under its control has been received after the original blocking request. If there is no such transaction, it is determined whether the subsequent blocking transaction request has the same nature as the earlier blocking transaction request. If it is affirmative, the subsequent blocking transaction request is deleted without further transmission, and when a response is received from the earlier block transaction request, a response is transmitted to the blocking transaction request. This is feasible because the subsequent blocking transaction request has the same properties as the first blocking transaction request, and the first blocking transaction request can be used as a boundary of the transaction after the second transaction, and therefore, there is no need to further transmit the transaction request. The second transaction. This can reduce the amount of processing that the control circuit system must perform to control these blocking transaction requests, and has other advantages for peripheral components in low power consumption mode (refer to Figure 14 below for more clarity). It should also be noted that, although not shown in this drawing, if the second blocking request is further transmitted, the interconnection will operate for the second blocking request in the usual manner.
If the subsequent blocking transaction request does not have the same properties as the earlier blocking transaction request, it needs to be further transmitted and aligned in a normal way to respond. For example, if the subsequent blocking transaction request affects a subset of the transactions leading to the earlier blocking transaction request, it has a narrower shareability area or only affects writing (relative to all the transactions that it can respond to) Function).
It should be noted that since this is located on a halved path, when the above-mentioned two blocking transaction requests are received, an early response related to it can be sent separately, and the second blocking transaction request can be easily deleted, and the second blocking transaction request can be easily deleted. The interruption required at the end of the sub-path only applies to the subsequent transaction.
It should also be noted that although this flow chart is drawn for a bisecting path, when the barrier is of the same nature and there is no interfering transaction subject to the barrier, the elimination of a subsequent barrier transaction can be completed on any path, as long as it is for the first When a blocker sends a response, a response is sent to the second blocker.
Figure 14 is a flow chart showing the steps of the method for controlling peripheral components to enter the low power consumption mode. A power off signal is received, which indicates that a peripheral component is about to enter a low power consumption mode. In response to the reception of this signal, a blocking transaction request is transmitted to the peripheral component along a bisecting path and a response is received. Then, the power of the peripheral component is turned off. When a subsequent block transaction request is received, it will determine whether any transaction request under its control is received after the earlier block. If no such request is received, the later blocking transaction request can be deleted (if it has the same nature as the earlier blocking transaction request) and a response to it can be sent. In this way, if the peripheral component is in an inactive state, there is no need to let a blocking transaction request disturb and wake up the peripheral component.
If there is an intermediate transaction controlled by the earlier barrier, or if the second barrier does not have the same properties as the earlier barrier, the barrier cannot be deleted, but it must be further transmitted in a normal manner.
Figure 15 is a flowchart showing the steps of a method for reducing the power consumption of a peripheral component. A blocking transaction request is received and transmitted to the peripheral device, and a response is received as a response. In response to the receipt of this blocking transaction request, it is known that the subsequent blocking transaction request can now be deleted, and therefore, it is an appropriate time for the peripheral component to enter the low power consumption mode. Therefore, a request is sent to the power controller to turn off the power of the peripheral component, and the power of the peripheral component is turned off. At this time, you can delete the received subsequent blocking transaction (provided that the intermediate controlled by the earlier boundary request is not received, and provided that they have the same nature), and send a response to it without interruption Peripheral components during sleep.
If an intermediate transaction controlled by the earlier block has been received, or if the second block does not have the same properties as the earlier block, the block cannot be deleted and must be further transmitted in the normal way.
Figure 16 shows the method steps used to reduce the management costs associated with blocking transaction requests by combining them where feasible.
Therefore, a blocking transaction request is received at the start of a bisecting path and an early response is transmitted, and the blocking transaction request is further transmitted. Since this point is the starting point of a bisecting path, there is no need to block the path for subsequent transaction requests at this point.
Then a subsequent blocking transaction request is received and an early response is transmitted, and the blocking is further transmitted. There is a reordering buffer on the bisection path, and both barriers are stored in it. Then it is determined whether any transaction request controlled by the earlier block is received after the earlier block and before the subsequent block is received. If it is negative, move the two blocked transaction requests to be adjacent to each other at the subsequent blocked position. If there is no interfering transaction controlled by the subsequent block, the subsequent block can be moved to be adjacent to the earlier block. It is then decided whether the nature of these blocking requests is the same. If it is affirmative, these blocks can be combined, and a single block request can be further sent at the correct position in the message stream. If they do not have the same properties, it is determined whether their properties are the same except for the sharable area, and it is also determined whether one of the sharable areas is a subset of the other sharable area. If this is the case, the blocking requests can be merged, and the merged request has the larger area of the two sharability areas. After that, the combined blocking request is further transmitted. If the nature of these blocking requests makes it impossible to merge these blocking requests, then these separate requests are further transmitted.
It should be noted that although the above describes the merger of barriers with reference to a reordering buffer zone, adjacent barrier transactions with appropriate properties can be merged on a bisected path without using a reordering buffer zone.
It should also be noted that in a reordering buffer, if there are transactions that can be reordered (except for a blocking instruction in between to prevent them from reordering), in some implementations, the transaction may be allowed to be merged and generated Two barriers, these two barriers are located on both sides of the resulting merged transaction.
Figure 17 is a flowchart illustrating the method steps for a barrier unit located in the interconnection (such as the interconnection shown in Figure 1) to handle barriers, and illustrates the different properties that a barrier transaction can have.
Therefore, after receiving a blocking transaction, if a subsequent transaction is designated to be controlled by the blocking, it will be blocked until a response is received. If it is not so designated, other properties of the transaction can be considered. If it is designated as not subject to the blocking control, it can be further transmitted without being blocked. If it is not so designated, it is determined whether it has a function designated by the barrier. It should be noted that the barrier may not specify a specific function. In this case, the transaction will be blocked regardless of its function. However, if it does specify a function, the transaction that does not perform this function will not be Blocked, and will be further transmitted. If the subsequent transaction specifies a region, determine whether the region indicator of the transaction is a message flow indicator, if it is affirmative, the transaction will not be blocked, but if it is any other region, it will be blocked . It also determines whether the barrier is a synchronous barrier or whether it has a blocking indicator. In either case, these barriers will not block the subsequent transactions, because these subsequent transactions have been blocked previously, and therefore, any subsequent transactions received will not be affected by these barriers.
It should be noted that although the steps shown here have a specific order, it is of course possible to perform these steps in any order. In addition, when the barrier does not have a specific function indicator, initial device or area, these situations do not need to be considered.
Figure 18 shows an initiating device and a receiving device according to the present technology. The initial device is a processor P0, which can be connected to an interconnect 10 and then to the receiving device 30. The interconnection 10 can be connected to more receiving devices and initiating devices not shown in the figure.
In this embodiment, the barrier generator 130 is used by the initiating device P0 to generate the barrier. The barrier generator generates a barrier transaction request with at least one indicator. This indicator can be an indicator indicating the nature of the blocking transaction request (that is, which transaction request it controls) or it can be an indicator indicating whether the transaction request is blocking, non-blocking, or both . If the indicator indicates that the blocking transaction request is a blocking request, the initiating device P0 will not send any further transaction requests to the interconnection after the blocking transaction request until it receives a response signal for the blocking.
The initiating device P0 also has a processor 150 and a harm unit 140.
When it is expected that transactions to the same address (which are issued by a proxy device) should occur in a specific order, but they can be reordered relative to each other, memory hazards can occur. When they detect a hazard, the master station shall not issue the later transaction until it sees that the earlier transaction has been completed (for a read) or there is a buffered response (for a write).
In order to be able to make decisions about harm, a master station must have a transaction tracking mechanism that records transactions that may lead to a harm. This mechanism can be provided in the form of a hazard unit 140, which stores pending transactions that may cause a hazard to occur until they are no longer likely to cause a hazard. The hazard unit 140 has a limited size, and this size is very small, which is very advantageous. Therefore, the hazard unit may also be full. If this happens, the processor 150 must stop issuing transactions until a transaction of the transactions stored in the harm unit has been completed and can be deleted by the harm unit. If the pending transactions that should be stored in the hazard unit are not stored, then hazards that cannot be corrected may occur. This is not allowed. Obviously, such a stop will increase the latency.
When the transaction has been completed without harm, the harm can be removed. However, if a barrier must be issued to prevent a hazard from occurring, the hazard unit 150 can remove the earlier transaction, because it is unlikely to be harmed by any subsequent transactions that may cause hazards. Therefore, the embodiment of the present invention utilizes the blocking generator 130 to handle such latent time. The blocking generator 130 can detect when the hazard unit 140 is full, and can send a blocking transaction in response to this detection. This blocking transaction prevents subsequent transactions from being reordered relative to it, and therefore, the potential hazards of the transaction stored in the hazard unit can be removed, and then these transactions can be removed by the hazard unit. Therefore, if the barrier generator generates a global barrier (all transactions cannot be reordered relative to it), the hazard unit 140 can be cleared. However, a global barrier itself will cause potential for interconnection, and may not be very ideal.
Therefore, in some embodiments, it is more advantageous to create a certain address barrier, which corresponds to a single address of a transaction located in the hazard unit. Generally speaking, the most recent transaction in the hazard unit will be selected, because in normal operations, this transaction may be the last transaction to be removed. Therefore, the barrier generator 130 detects the address of the most recent transaction in the hazard unit, and issues a barrier related to this address, so that any transactions to the address are not allowed to be reordered with respect to the barrier. This can ensure that the processed transaction is no longer a possible hazard and can be removed by the hazard unit. This frees up space and allows the processor 150 to continue issuing transactions.
In other cases, the barrier generator 130 can also be used to generate barriers. For example, it can detect that the highly ranked transactions issued by the processor must be completed in a specific order. Generally speaking, when outputting a highly ranked transaction, P0 will not output any further transactions until it receives a response signal from the highly ranked transaction indicating that it has been completed. This will of course affect the latency of processor P0. In some embodiments, the barrier generator 130 detects that the processor 150 issues a highly ranked transaction and issues a barrier itself. Once the barrier has been output to the interconnect 10, the response unit 80 in the interconnect will send a response signal to the processor P0, which will clear the barrier and allow the processor P0 to output further transactions. In order to avoid problems caused by the reordering of transactions with respect to the highly ranked, the blocking unit 90 blocks subsequent transactions. When the barrier passes through the interconnection, the interconnection 10 handles the barrier, blocking and clearing when appropriate, and the processor P0 can continue to issue transactions. When the interconnection is assigned a handleable barrier to reduce the latency where feasible, the latency of the system can be reduced (as opposed to a block that occurs at the processor P0 until the highly ranked transaction has been completed).
The receiving device has a connection port 33 for receiving a transaction request from the initiator P through the interconnection 10 and responding to the reception of a blocking transaction request. The response signal generator 34 sends a response and transmits it through the connection port 33 To interconnect 10. In order to respond to the blocking transaction request as a non-blocking transaction request, the receiving device 30 may not send any response. In addition, the receiving device can respond to blocking certain indicators on the transaction request to delay the generation and/or transmission of the response signal until the processing of the previously received transaction request has been at least partially completed. It is possible that certain blocked transaction requests not only require that the earlier transaction request has reached its final destination, but also require that they have been processed. For example, a data synchronization block is the case. Therefore, in response to the receiving device recognizing the blocking (perhaps by the indicator value in the blocking transaction request), the receiving device can delay the transmission of the response signal until the required processing has been completed.
Figure 19 outlines the blocking and the response to the transaction in the cross-coupling and re-convergence area. A transaction message flow with barriers reaches the divergence point 160. The divergence point 160 has a control circuit system 170 and a barrier management circuit system 180 associated therewith. It also has a buffer 162 for storing transactions.
The barrier management circuit system 180 replicates the barrier, and transmits the replicated barrier down to each of the exit paths. The control circuit system 170 is used to block the subsequent transaction from proceeding further, and can monitor the response signal to the copied blocking transaction. In this embodiment, in response to the reception of these two responses, the control circuit system connects the path on which the response has not been received, and allows the subsequent transaction to proceed on this path. Since they are only allowed to advance on this path, they cannot surpass the barrier on this path and the previous transaction, and have already responded to the barrier on other paths, so they have not reordered relative to these transactions. danger. It is particularly advantageous to connect this path early because it is very likely to be the path with the most traffic, because it is the slowest path to respond, and therefore, passing down this path to subsequent transactions as early as possible will help reduce latency.
Before transmitting the subsequent transaction, a blocking indication can also be passed down this path. This may be necessary when there are further points in the path, as is the case here.
At the next divergence point 190, receiving the block will cause a response signal to be sent and clear other paths, and receiving the block representation will cause the block management circuit system 180 to copy the block and pass it down to each exit path, and The control circuit system 170 can act to stop the subsequent transaction. In response to receiving a response signal on one of these paths, other paths can be connected, because this is the only path on which no response is received, and subsequent transactions are further transmitted together with a blocking indication. When a response is received on both paths, the two paths can be connected.
It should be noted that the blocking representation has never been copied and no response is required. It only allows the control circuit system to understand that the subsequent transaction was transmitted before receiving a response, and therefore, if there is a subsequent divergence point, for example, it may need to be blocked.
This is a convenient way to improve the efficiency of processing barriers. In addition, in some embodiments, this method may be particularly advantageous, for example, if the path 200 shown in Figure 19 happens to be a path with low traffic, and it is possible that the previous transaction is a block and a response to it has been received , Then on this copy of the barrier located at the divergence point 190, the barrier management circuit system will perceive that the previous transaction passed down the path 200 is a barrier, and has received a response signal for this barrier, and can delete the After copying the block and responding, other paths can be opened immediately and the subsequent transactions can be transmitted to these paths, thus further reducing the latency of the system.
Figure 20 summarizes the different types of barrier transactions and how to convert them from one type to another when these barrier transactions enter different areas of the interconnection with different requirements. In this way, the blocking nature of a barrier can be removed where feasible and the blocking nature can be reintroduced when necessary to reduce the latency caused by the barrier.
The memory blocking transaction can be blocked or non-blocked, depending on its position in the interconnection. In fact, these barriers have three different behaviors, and they can be regarded as three different types of memory barriers. There is a sequence barrier that does not block but stays in the transaction message flow to separate subsequent transactions from earlier transactions. In an interconnection with multiple regions, it is located outside its region A memory barrier can be used as a sequential barrier; a system clear barrier will be deliberately blocked corresponding to a DSB; and a region clear barrier usually does not block but will be based on the necessary component layout reasons (a cross-coupling area Medium) or optional performance reasons for regional blocking.
In many cases, conversion between these types can be performed. Switching from a sequence barrier to a system or area clear barrier requires a conversion point to block subsequent transactions. The conversion from system or area clearing barriers to sequential barriers does not necessarily require an early response (so that subsequent transactions can be sent), but it can still provide such early responses-this conversion without sending early responses is meaningless, because This will not reduce the latency of the blocking launcher, and will cause more main battles to be blocked in the next cross-coupling area. If an early response is required (and this is allowed-that is, in one or two division areas), if the response to a system or area clears the barrier, because the early response may cause the transaction from behind the barrier to be transmitted, then provide The location of the response must block these late transactions or it must change the block to make it a sequential block.
Generally speaking, it is conceivable that optional conversion is not common. The nature of transaction is very useful for reasons of efficiency management and service quality. Exchange of interests between the launching device and other master stations during the dive time.
Figure 20 shows the allowable barrier conversion, which depends on the context of the conversion and the barrier required for the conversion.
Various aspects and features of the present invention are defined in the scope of the accompanying patent application. Various modifications can be made to the above-described embodiments without departing from the scope of the present invention.
<p>10, 12. . . interconnection</p><p>20, 22, 24, 26, 28, P0, P1, P2, S0, S1, S2, S3. . . Main site</p><p>30, 32, 34, 36. . . Slaves</p><p>31, 33, 210. . . Outer area</p><p>40, 48. . . Cross-coupling part</p><p>41, 42, 49. . . Branch point</p><p>44, 45, 91. . . Merge point</p><p>50, 52. . . Bisection path</p><p>60, 110. . . Reorder buffer</p><p>70, 76, 170. . . Control circuit system</p><p>80. . . Empty cell</p><p>82. . . Empty circuit system</p><p>90, 92. . . Blocking circuit system</p><p>100. . . Combined circuit system</p><p>102. . . Barrier elimination circuit system</p><p>112. . . Memory controller</p><p>120, 122, 124, 126, 127, 128. . . Non-shared area</p><p>125, 129, 200, 201. . . Internal area</p><p>130. . . Barrier generator</p><p>135. . . Exit node</p><p>140. . . Hazard unit</p><p>150. . . processor</p><p>160, 190. . . Divergence point</p><p>162. . . Buffer for storing transactions</p><p>180. . . Barrier management circuit system</p><p>210. . . Location</p>
The present invention is further described by way of example with reference to the embodiments of the present invention and the accompanying drawings, wherein the drawings are as follows:
Figure 1 shows an interconnection according to an embodiment of the present invention;
Figure 2 shows two interconnections connected together according to an embodiment of the present invention;
Figure 3 schematically shows a blocking transaction request according to an embodiment of the present invention;
Figure 4 outlines another transaction request;
Figure 5A schematically illustrates the arrangement of multiple regions in an interconnection;
Figure 5B shows an interconnection area and its egress node according to an embodiment of the present invention;
Figure 6 shows a further arrangement of regions in an interconnect according to a further embodiment of the present invention;
Figure 7 shows a further interconnection according to an embodiment of the present invention;
Figure 8 shows an interconnection with parallel read and write paths according to an embodiment of the present invention;
Figure 9 shows an interconnection according to an embodiment of the present invention, which can be combined to block transactions;
Figure 10 shows an interconnection according to an embodiment of the present invention, which can eliminate blocking transactions;
Figure 11 shows a data processing device with an interconnect according to an embodiment of the present invention;
Figure 12 shows a flow chart of the steps of a method for removing the blockage caused by blocking transactions through an interconnection according to an embodiment of the present invention;
FIG. 13 shows a flowchart of steps of a method for removing a blocking transaction request according to an embodiment of the present invention;
Figure 14 shows a flowchart of a method for controlling peripheral components to enter a low power consumption mode according to an embodiment of the present invention;
15 is a flowchart of steps of a method for reducing the power consumption of a peripheral device according to an embodiment of the present invention;
Figure 16 shows a flow chart of a method for reducing the management costs associated with a blocked transaction by using a merged barrier transaction in a suitable situation according to an embodiment of the present invention;
FIG. 17 is a flowchart of steps of a method for processing a blocked transaction at a blocking unit located in an interconnection according to an embodiment of the present invention;
Figure 18 shows the initiating device and the receiving device according to the present technology;
Figure 19 outlines the transmission and blocking of transactions; and
Figure 20 summarizes the different types of barrier transactions and how to convert them from one type to another when these barrier transactions enter different areas of the interconnection with different requirements.
21 sheets
Sheet 1 Sheet 2 Sheet 3 Sheet 4 Sheet 5 Sheet 6 Sheet 7 Sheet 8 Sheet 9 Sheet 10 Sheet 11 Sheet 12 Sheet 13 Sheet 14 Sheet 15 Sheet 16 Sheet 17 Sheet 18 Sheet 19 Sheet 20 Sheet 21
Every citation, both ways
| Document | Relation | Office | Cited during |
|---|---|---|---|
| TW200502805A | Cites | Taiwan Province of China | Examiner |
| US2008301342A1 | Cites | United States of America | Examiner |
| TW200849051A | Cites | Taiwan Province of China | Examiner |
| US6038646A | Cites | United States of America | Examiner |
| US6038646 | Cites | United States of America | – |
| US20080301342A1 | Cites | United States of America | – |
67 members in 11 offices
Priority claims10
| Document | Office | Kind | Date |
|---|---|---|---|
| 0917946 | United Kingdom | A | |
| 0917946 | United Kingdom | A | |
| 09179466 | United Kingdom | – | |
| 10073427 | United Kingdom | – | |
| 201007342 | United Kingdom | A | |
| 201007342 | United Kingdom | A | |
| 09179466 | – | – | – |
| 10073427 | – | – | – |
| GB20090017946 | – | – | – |
| GB20100007342 | – | – | – |
Members67
| Document | Office | Kind | |
|---|---|---|---|
| GB201007342D0 | United Kingdom | D0 | |
| GB201007363D0 | United Kingdom | D0 | |
| GB201016482D0 | United Kingdom | D0 | |
| US2011087809A1 | United States of America | A1 | |
| US2011087819A1 | United States of America | A1 | |
| GB2474446A | United Kingdom | A | |
| GB2474532A | United Kingdom | A | |
| GB2474533A | United Kingdom | A | |
| GB2474552A | United Kingdom | A | |
| US2011093557A1 | United States of America | A1 | |
| WO2011045555A1 | World Intellectual Property Organization (WIPO) | A1 | |
| WO2011045556A2 | World Intellectual Property Organization (WIPO) | A2 | |
| WO2011045595A1 | World Intellectual Property Organization (WIPO) | A1 | |
| CN102063391A | China | A | |
| US2011119448A1 | United States of America | A1 | |
| US2011125944A1 | United States of America | A1 | |
| TW201120638A | Taiwan Province of China | A | |
| TW201120656A | Taiwan Province of China | A | |
| WO2011045556A3 | World Intellectual Property Organization (WIPO) | A3 | |
| JP2011138481A | Japan | A | |
| TW201128398A | Taiwan Province of China | A | |
| IL218703A0 | Israel | A0 | |
| IL218704A0 | Israel | A0 | |
| IL218887A0 | Israel | A0 | |
| CN102576341A | China | A | |
| EP2488951A1 | European Patent Office (EPO) | A1 | |
| EP2488952A2 | European Patent Office (EPO) | A2 | |
| EP2488954A1 | European Patent Office (EPO) | A1 | |
| KR20120093276A | Republic of Korea | A | |
| KR20120095872A | Republic of Korea | A | |
| KR20120095876A | Republic of Korea | A | |
| CN102713874A | China | A | |
| CN102792290A | China | A | |
| JP2013507708A | Japan | A | |
| JP2013507709A | Japan | A | |
| JP2013507710A | Japan | A | |
| US8463966B2 | United States of America | B2 | |
| US8601167B2 | United States of America | B2 | |
| US8607006B2 | United States of America | B2 | |
| US2014040516A1 | United States of America | A1 | |
| US8732400B2 | United States of America | B2 | |
| GB2474532B | United Kingdom | B | |
| GB2474532A8 | United Kingdom | A8 | |
| GB2474532B8 | United Kingdom | B8 | |
| US8856408B2 | United States of America | B2 | |
| EP2488951B1 | European Patent Office (EPO) | B1 | |
| JP5650749B2 | Japan | B2 | |
| TWI474193BThis record | Taiwan Province of China | B | |
| JP2015057701A | Japan | A | |
| EP2488952B1 | European Patent Office (EPO) | B1 | |
| CN102576341B | China | B | |
| MY154614A | Malaysia | A | |
| IN2792DEN2012A | India | A | |
| IN2853DEN2012A | India | A | |
| IN3050DEN2012A | India | A | |
| CN102792290B | China | B | |
| JP2015167036A | Japan | A | |
| IL218887A | Israel | A | |
| MY155614A | Malaysia | A | |
| EP2488954B1 | European Patent Office (EPO) | B1 | |
| JP5865976B2 | Japan | B2 | |
| IL218704A | Israel | A | |
| TWI533133B | Taiwan Province of China | B | |
| US9477623B2 | United States of America | B2 | |
| KR101734044B1 | Republic of Korea | B1 | |
| KR101734045B1 | Republic of Korea | B1 | |
| JP6141905B2 | Japan | B2 |
Numbers
- Publication
- I474193
- Publication, DOCDB
- I474193
- Publication, EPODOC
- TWI474193B
- Application
- 99133596
- Application, DOCDB
- 99133596
- Application, EPODOC
- TW20100133596
Titles3
- English
- Barrier transactions in interconnection
- English
- BARRIER TRANSACTIONS IN INTERCONNECTS
- Chinese
- 互連中的阻隔交易
Classification
- CPC, 6
- G06F13/362
- G06F13/1621
- H04L49/90
- G06F13/1689
- G06F13/364
- H04L49/9084
- IPC, 1
- G06F17 00