Retention priority based cache replacement policy
18 claims: 10 independent, 8 dependent
- 1データを処理するための装置であって、 メモリアクセス要求の複数のソース であって、実行のためのプログラム命令をフェッチするように構成された、命令フェッチ回路と、前記プログラム命令の制御下で処理操作の対象となるデータ値にアクセスするように構成された、データアクセス回路とを含む、複数のソース と、 前記複数のソースに連結されるキャッシュメモリと、 前記キャッシュメモリに連結され、前記キャッシュメモリへのキャッシュラインの挿入、および前記キャッシュメモリからのキャッシュラインの追い出しを制御するように構成された、キャッシュ制御回路と、を備え、 前記キャッシュ制御回路は、前記キャッシュメモリに挿入される各キャッシュラインと関連付けられるそれぞれの保持優先度値を記憶するように構成され、 前記キャッシュ制御回路は、前記保持優先度値に依存して前記キャッシュメモリから追い出すキャッシュラインを選択するように構成され、また、 前記キャッシュ制御回路は、 (i)前記 命令フェッチ回路と前記データアクセス回路のいずれか1つ が、前記キャッシュメモリへの前記キャッシュラインの挿入をもたらす、メモリアクセス要求を発行したのか、および (ii)前記メモリアクセス要求の特権レベル、のうちの少なくとも1つに依存して前記キャッシュメモリに挿入されるキャッシュラインと関連付けられる、保持優先度値を設定するように構成 され、 前記キャッシュ制御回路は、 高い保持優先度に対応する保持優先度値を有するキャッシュラインに優先して、低い保持優先度に対応する保持優先度値を有するキャッシュラインを追い出し、 最低保持優先度に対応する関連する保持優先度値を有するキャッシュラインの中から追い出すためのキャッシュラインを選択し、 前記最低保持優先度に対応する関連する保持優先度値を伴う、いかなるキャッシュラインもない場合、少なくとも1つのキャッシュラインが、前記最低保持優先度に対応する保持優先度値を有するまで、前記キャッシュメモリ内の全ての前記キャッシュラインの保持優先度を降格させる ように構成された、 装置。
- 2前 記命令フェッチ回路によって発行されるメモリアクセス要求の結果として前記キャッシュメモリに挿入されるキャッシュラインは、命令保持優先度値と関連付けられ、また、前記データアクセス回路によって発行されるメモリアクセス要求の結果として前記キャッシュメモリに挿入されるキャッシュラインは、前記命令保持優先度値と異なるデータ保持優先度値と関連付けられる、請求項1に記載の装置。
- 3前記キャッシュ制御回路は、 (i)前記命令保持優先度値が、前記データ保持優先度値よりも高い保持優先度に対応すること、および (ii)前記命令保持優先度値が、前記データ保持優先度値よりも低い保持優先度に対応すること、のうちの1つを設定するように、フラグ値に応答する、請求項 2 に記載の装置。
- 4前記フラグ値は、ソフトウェアプログラマブルなフラグ値である、請求項 3 に記載の装置。
- 5前記命令フェッチ回路および前記データアクセス回路は、インオーダープロセッサの一部であり、前記命令保持優先度値は、前記データ保持優先度値よりも低い保持優先度に対応する、請求項 2 に記載の装置。
- 6前記保持優先度値は、それらの関連するキャッシュラインのTAG値とともに前記キャッシュメモリ内に記憶される、請求項1に記載の装置。
- 7前記キャッシュ制御回路は、最低保持優先度に対応する関連する保持優先度値を有するキャッシュラインの中から追い出すための前記キャッシュラインをランダムに選択するように構成された、請求項 1 に記載の装置。
- 8前記キャッシュ制御回路は、前記キャッシュメモリの中に既に存在するキャッシュラインへのアクセスを検出し、前記キャッシュラインの保持優先度を昇格させるために前記キャッシュラインの保持優先度値を変更するように構成された、請求項 1 に記載の装置。
- 9前記保持優先度値は、 (i)各アクセスに応じた、最高保持優先度へ向けての前記キャッシュラインの保持優先度の漸増的な昇格、 (ii)最高保持優先度への前記キャッシュラインの直接的な昇格、のうちの1つを行うように変更される、請求項 8 に記載の装置。
- 10前記複数のソースは、汎用プロセッサと、グラフィックス処理ユニットとを含む、請求項1に記載の装置。
- 11前記複数のソースは、複数の汎用プロセッサを含む、請求項1に記載の装置。
- 12前記キャッシュメモリは、少なくとも1つの1次キャッシュメモリおよび2次キャッシュメモリを含むキャッシュメモリの階層内の、前記2次キャッシュメモリである、請求項1に記載の装置。
- 13前記キャッシュメモリは、少なくとも1つの1次キャッシュメモリ、少なくとも1つの2次キャッシュメモリ、および3次キャッシュメモリを含むキャッシュメモリの階層内の、前記3次キャッシュメモリである、請求項1に記載の装置。
- 14前記保持優先度値は、nビットの保持優先度値であり、キャッシュラインの挿入に応じて前記キャッシュ制御回路によって設定することができる、異なる保持優先度値の合計が、2 n -1である、請求項1に記載の装置。
- 15カーネルプログラム特権レベルを伴うメモリアクセス要求の結果として前記キャッシュメモリに挿入されるキャッシュラインは、カーネル保持優先度値と関連付けられ、前記キャッシュレベルに挿入されるキャッシュラインは、ユーザ保持優先度値と関連付けられる、請求項1に記載の装置。
- 16前記カーネル保持優先度値が、前記ユーザ保持優先度値よりも高い保持優先度に対応する、 前記ユーザ保持優先度値が、前記カーネル保持優先度値よりも高い保持優先度に対応する、のうちの1つである、請求項 15 に記載の装置。
- 17データを処理するための装置であって、 メモリアクセス要求を生成するための複数のソース手段 であって、実行のためのプログラム命令をフェッチするように構成された、命令フェッチ手段と、前記プログラム命令の制御下で処理操作の対象となるデータ値にアクセスするように構成された、データアクセス手段とを含む、複数のソース手段 と、 データを記憶するためのキャッシュメモリ手段と、 前記キャッシュメモリ手段へのキャッシュラインの制御挿入、および前記キャッシュメモリ手段からのキャッシュラインの追い出しを制御するためのキャッシュ制御手段と、を備え、 前記キャッシュ制御手段は、前記キャッシュメモリ手段に挿入される各キャッシュラインと関連付けられるそれぞれの保持優先度値を記憶するように構成され、 前記キャッシュ制御手段は、前記保持優先度値に依存して前記キャッシュメモリ手段から追い出すキャッシュラインを選択するように構成され、 前記キャッシュ制御手段は、 (i)前記 命令フェッチ手段と前記データアクセス手段のいずれか1つが 、前記キャッシュメモリ手段への前記キャッシュラインの挿入をもたらす、メモリアクセス要求を発行したのか、および (ii)前記メモリアクセス要求の特権レベル、のうちの少なくとも1つに依存して前記キャッシュメモリ手段に挿入されるキャッシュラインと関連付けられる、保持優先度値を設定するように構成 され、 前記キャッシュ制御手段は、 高い保持優先度に対応する保持優先度値を有するキャッシュラインに優先して、低い保持優先度に対応する保持優先度値を有するキャッシュラインを追い出し、 最低保持優先度に対応する関連する保持優先度値を有するキャッシュラインの中から追い出すためのキャッシュラインを選択し、 前記最低保持優先度に対応する関連する保持優先度値を伴う、いかなるキャッシュラインもない場合、少なくとも1つのキャッシュラインが、前記最低保持優先度に対応する保持優先度値を有するまで、前記キャッシュメモリ内の全ての前記キャッシュラインの保持優先度を降格させる ように構成された、 装置。
- 18データを処理する方法であって、 実行のためのプログラム命令をフェッチするように構成された命令フェッチ回路と、前記プログラム命令の制御下で処理操作の対象となるデータ値にアクセスするように構成されたデータアクセス回路とを含む、 複数のソースでメモリアクセス要求を生成するステップと、 キャッシュメモリ内にデータを記憶するステップと、 前記キャッシュメモリ手段へのキャッシュラインの制御挿入、および前記キャッシュメモリ手段からのキャッシュラインの追い出しを制御するステップと、を含み、前記方法はさらに、 前記キャッシュメモリに挿入される各キャッシュラインと関連付けられるそれぞれの保持優先度値を記憶するステップと、 前記保持優先度値に依存して前記キャッシュメモリ手段から追い出すキャッシュラインを選択するステップと、 保持優先度値を設定するステップであって、 (i)前記 命令フェッチ回路と前記データアクセス回路のいずれか1つ が、前記キャッシュメモリへの前記キャッシュラインの挿入をもたらす、メモリアクセス要求を発行したのか、および (ii)前記メモリアクセス要求の特権レベル、のうちの少なくとも1つに依存して前記キャッシュメモリに挿入されるキャッシュラインと関連付けて、設定するステップと、 高い保持優先度に対応する保持優先度値を有するキャッシュラインに優先して、低い保持優先度に対応する保持優先度値を有するキャッシュラインを追い出すステップと、 最低保持優先度に対応する関連する保持優先度値を有するキャッシュラインの中から追い出すためのキャッシュラインを選択するステップと、 前記最低保持優先度に対応する関連する保持優先度値を伴う、いかなるキャッシュラインもない場合、少なくとも1つのキャッシュラインが、前記最低保持優先度に対応する保持優先度値を有するまで、前記キャッシュメモリ内の全ての前記キャッシュラインの保持優先度を降格させるステップ、 を含む、方法。
Independent claims18
28 paragraphs, as filed
The present invention relates to the field of data processing systems. More specifically, the present invention relates to cache replacement policies for use within a data processing system.
It is known to provide cache memory for data processing systems. Cache memory provides faster and more effective access to frequently used data or instructions. The cache memory generally has a limited size compared to the main memory, and therefore the cache memory contains instructions / data held in the main memory at any given time. Only a subset can be retained. The cache memory has a replacement policy, which is the cache line that should be removed from the cache to make room for a new cache line that should be fetched from main memory and stored in the cache memory. Determine (which may contain data and / or instructions). There are many known examples of cache replacement policies such as LRU (Least Recently Used), round robin, and random.
<p num="0003"> Overviewd from one aspect, the present invention provides an apparatus for processing data, wherein the apparatus is: With multiple sources of memory access requests, A cache memory combined with the plurality of sources and A cache control circuit that is coupled to the cache memory and is configured to control the insertion of a cache line into the cache memory and the evacuation of the cache line from the cache memory. The cache control circuit is configured to store each retention priority value associated with each cache line inserted into the cache memory. The cache control circuit is configured to select a cache line to be expelled from the cache memory depending on the retention priority value. The cache control circuit is (i) Which of the plurality of sources issued a memory access request resulting in the insertion of the cache line into the cache memory, and (ii) It is configured to set a retention priority value associated with a cache line inserted into the cache memory depending on at least one of the privilege levels of the memory access request.</p><p num="0004"> The approach recognizes that an improved replacement policy can be achieved by associating each cache line with a retention priority value that depends on the source of the memory access request and / or the privilege level of the memory access request. The cache line to be expelled from the cache memory is then selected depending on these assigned retention priority values (which can be modified while the cache line is in the cache memory).</p><p num="0005"> The cache control circuit may be configured to expel a cache line having a holding priority value corresponding to a lower holding priority in preference to a cache line having a holding priority value corresponding to a higher holding priority. Therefore, the retention priority value can be used to represent the expected desirability of retaining a given cache line in cache memory. The desire to keep a given cache line in cache memory is due to the fact that it is frequently accessed, or a significant penalty if access is delayed due to the cache line not being in cache memory. It can be the result of the fact that there is. The retention priority value set depending on the source of the original memory access request that triggered the fetch and / or the privilege level of the memory access request that triggered the fetch is achieved by the cache line held in the cache memory. You can increase the benefits (or, from another point of view, reduce the penalty incurred by cache lines that are not kept in cache memory).</p><p num="0006"> In some exemplary embodiments, the plurality of sources is an instruction fetch circuit configured to fetch program instructions for execution and data access configured to access data values to be processed. It has a circuit. As an example, both the instruction fetch circuit and the data access circuit can be part of a general purpose processor, graphics processing unit, or DSP. Either the instruction retention priority value or the data retention priority value is associated with the cache line, depending on whether the cache line was inserted into the cache memory as a result of memory access by the instruction fetch circuit or data access circuit. obtain. The instruction retention priority value and the data retention priority value are different to give different chances that the cache line will then be expelled from the cache memory.</p><p num="0007"> Depending on the situation, it may be desirable for the instruction retention priority value to correspond to a retention priority higher than the data retention priority value. In other situations, it may be desirable to maintain the opposite retention priority relationship. Therefore, in some embodiments, the cache control circuit responds to a flag to set one of the above relationships.</p><p num="0008"> This flag may be controlled by a physical hardware signal, but in some embodiments the flag value is a software programmable flag program value and this value is an instruction or data. Controls whether or not they are preferentially retained in the cache memory by having them assigned different retention priority values when inserted into the cache memory.</p><p num="0009"> In some embodiments, the instruction fetch circuit and the data access circuit are part of an in-order processor, in which the instruction retention priority value is lower than the data retention priority value. It is desirable to have a priority. That is, it holds data in memory in preference to instructions (at least to the extent that it is biased to prioritize holding data over instructions rather than preventing them from being held). Is desirable.</p><p num="0010"> Retention priority values can be stored somewhere in the system (eg, in a cache control circuit), which is, in some embodiments, the retention priority values, along with their associated cache lines. It is advantageous in that it is stored in the cache memory (most likely, if the cache memory contains separate TAG and data memory, along with the TAG value). Cache lines may be extended by one or more bits to easily accommodate the retention priority values associated with them.</p><p num="0011"> When selecting a cache line to evict, the cache control circuit responds to the retention priority value associated with the cache line. In some embodiments, the cache control circuit may be configured to select a cache line to be expelled from among cache lines having an associated retention priority value corresponding to the lowest retention priority value that can be represented. Within this pool of cache lines corresponding to the lowest hold priority, the cache control circuit can be selected according to another policy such as round robin or least recently used, but the associated hold priority corresponding to the lowest hold priority. It is easy to make a random selection from cache lines with degree values.</p><p num="0012"> In some embodiments, a cache controller has at least one cache line, if there is no cache line that corresponds to the lowest hold priority value and therefore has an associated hold priority value that is eligible for eviction. It has a retention priority value that corresponds to the lowest retention priority and is therefore configured to demote the retention priority of all cache lines in the cache memory until it is eligible for eviction. Such demotion can be achieved by changing the stored retention priority values or by changing the mapping between the stored retention priority values and the priorities they represent.</p><p num="0013"> As discussed above, the retention priority value associated with a cache line when it is inserted varies depending on the source or privilege level of the corresponding memory access request. The retention priority value can then be changed individually depending on the activity associated with that cache line, or totally dependent on the activity of the cache memory. In some embodiments, the cache control circuit detects an access (hit) to a cache line that already exists in the cache memory and raises the cache line retention priority in response to such a hit. Is configured to change its retention priority. Therefore, regularly accessed cache lines elevate their retention priority so that they are preferentially retained in cache memory.</p><p num="0014"> The promotion of the cache line retention priority value according to the cache hit can be performed by various different methods. In some embodiments, the hit cache line is assigned its retention priority value towards the highest maximum retention priority value (when this is reached, it is not further promoted in response to the hit) for each access. It may be progressively promoted or, as an alternative, modified to shift directly to the highest retention priority when a hit occurs.</p><p num="0015"> As discussed above, multiple sources with different retention priority values can take the form of an instruction fetch circuit vs. a data access circuit. In other embodiments, the plurality of sources can include a general purpose processor and a graphics processing unit, which are in cache memory where the cache line is shared between the general purpose processor and the graphics processing unit. It has a different retention priority value associated with the cache line that triggers it to be inserted into. As an example, the graphics processing unit can tolerate longer latency associated with missed memory access in cache memory. Therefore, for a general purpose processor, it is desirable to bias the cache resources to the general purpose processor by assigning a retention priority value corresponding to a higher retention priority to the cache line inserted in the cache memory.</p><p num="0016"> Another example of multiple sources could be multiple generic processors that are potentially identical. However, some of these processors can be assigned more speed-sensitive tasks, thus allowing the retention of cache lines in cache memory shared by their multiple general-purpose processors to be rapid to instructions or data. It is appropriate to bias towards the processors that need the most access (for example, some processors perform real-time tasks where performance is important, while others easily have longer latency. Do an acceptable background maintenance task).</p><p num="0017"> The cache memory to which this method is applied can take various different forms at a location in the memory system. In some embodiments, the cache memory is a secondary cache memory, but in other embodiments, the cache memory can be a tertiary cache memory.</p><p num="0018"> It will be appreciated that the retention priority value can have a variety of different bit sizes, depending on the fineness desired to be specified. There is a balance between the degree of fineness supported and the storage resources required to store the retention priority values. In some embodiments, a compromise is achieved by setting the retention priority value to a 2-bit retention priority value.</p><p num="0019"> As mentioned above, the retention priority assigned to a cache line when inserted into cache memory can vary depending on the privilege level associated with the memory access request that fetches the cache line into cache memory. Thus, cache line retention can be biased to prioritize cache lines fetched by memory access requests with higher privilege levels. As an example, the device may execute both the kernel program and one or more user programs, and the privilege level of the kernel program is higher than the privilege level of one or more user programs. In this case, for example, a cache line inserted into cache memory for a kernel program is given a higher retention priority for retention than a cache line inserted for one or more user programs. Can be. In other embodiments, the reverse is also true. That is, the cache line inserted for the user program may be given a higher priority than the cache line inserted for the kernel program.</p><p num="0020"> Overviewd from another aspect of the invention, the invention provides an apparatus for processing data, wherein the apparatus. Multiple source means for generating memory access requests, A cache memory means for storing data, A cache control means for controlling the control insertion of the cache line into the cache memory means and the evacuation of the cache line from the cache memory means is provided. The cache control means is configured to store each retention priority value associated with each cache line inserted into the cache memory means. The cache control means is configured to select a cache line to be expelled from the cache memory means depending on the retention priority value. The cache control means is (i) Which of the plurality of sources issued a memory access request that resulted in the insertion of the cache line into the cache memory means, and (ii) It is configured to set a retention priority value associated with a cache line inserted into the cache memory means depending on at least one of the privilege levels of the memory access request.</p><p num="0021"> Overviewd from a further aspect, the present invention provides a method of processing data, the method of which is: Steps to generate memory access requests from multiple sources, Steps to store data in cache memory and The method further comprises a step of controlling the control insertion of the cache line into the cache memory means and the evacuation of the cache line from the cache memory means. A step of storing each retention priority value associated with each cache line inserted into the cache memory, and A step of selecting a cache line to be expelled from the cache memory means depending on the retention priority value, and It is a step to set the retention priority value. (i) Which of the plurality of sources issued a memory access request that resulted in the insertion of the cache line into the cache memory, and (ii) includes setting a retention priority value associated with a cache line inserted into the cache memory depending on at least one of the privilege levels of the memory access request.</p><p num="0022"> The above and other objectives, features, and advantages of the present invention will become apparent from the following detailed description of the exemplary embodiments, which should be read with the accompanying drawings.</p><p num="0023"> Hereinafter, embodiments of the present invention will be described as merely examples with reference to the accompanying drawings.</p>
<figref num="1">It is a diagram schematically illustrating a data processing system including a plurality of access request sources, a cache hierarchy, and a main memory.</figref><figref num="2">It is a diagram schematically illustrating a general-purpose processor including an instruction fetch circuit and a data access circuit.</figref><figref num="3">FIG. 5 is a schematic diagram of a cache control circuit for inserting cache lines with different associated retention priority values and then expelling the cache lines depending on their retention priority values.</figref><figref num="4">It is a figure which shows schematic the contents of one cache line.</figref><figref num="5">It is a flow diagram which schematically illustrates the operation of the cache control circuit corresponding to the received memory access request.</figref><figref num="6">It is a flow diagram which schematically illustrates the processing operation performed in relation to a cache miss when the retention priority value is set according to the source of the memory access request.</figref><figref num="7">The flow diagram is similar to the flow of FIG. 6 except that the holding priority value is set depending on whether the memory access that caused the mistake has the kernel privilege level associated with the kernel program.</figref>
Figure 1 shows a cache hierarchy with 4 main memories, 1 tertiary cache 6, 2 secondary caches 8, 10, 4 primary caches 12, 14, 16, 18 and two general purpose processors 20, 22. A data processor 2 comprising, one graphics processing unit 24, and a plurality of sources of memory access requests, including a DSP unit 26, is schematically illustrated.
The various cache memories 6 to 18 are arranged in a hierarchical structure. All of the access request sources 20, 22, 24, 26 share a single tertiary cache 6. The two general-purpose processors 20 and 22 share a secondary cache 8. The graphics processing unit 24 and the DSP 26 share a secondary cache 10. Each of the access request sources 20, 22, 24, 26 has its own primary caches 12, 14, 16, 18. The method can be applied to any of caches 6-18, but is found to be particularly useful in secondary caches 8, 10 and tertiary cache 6.
The main memory 4 is also shown in the figure. This main memory 4 supports a memory address space, within which a different range of memory addresses, primarily by different programs or tasks running on 20, 22, 24, 26 to different access request sources. Can be used. A cache line fetched into any of cache memories 6-18 has a corresponding address (actually, a small range of addresses) in main memory 4. As illustrated in FIG. 1, different ranges of memory addresses are used by different ones of multiple user programs 28, 30, 32, and 34, and different ranges of memory addresses 36 are used by the kernel program (specific). Used for (part of the operating system) consisting of code that can be executed frequently in mode of operation. A memory management unit (not shown) may be provided to restrict access to different areas of main memory 4 depending on the source of the current privilege level of the memory access access request. Privilege levels can also be used to control retention priorities. As discussed later, the memory address (s) associated with a cache line determines the retention priority value assigned to that cache line when it is inserted into one of cache memories 6-18. Can be used for. The memory address space is divided into a plurality of ranges of memory addresses so that different retention priority values can be associated with different ranges of memory addresses.
FIG. 2 schematically illustrates a general purpose processor 38 including an instruction fetch circuit 40 and a data access circuit 42. The instruction fetch circuit 40 fetches a series of program instructions passed to the in-order instruction pipeline 44, and the program instructions proceed along the pipeline as they undergo various processing steps. At least one of these processing steps decodes the instruction with an instruction decoder 46 to generate and decode control signals that control processing circuits such as register bank 48, multiplier 50, shifter 52, and adder 54. Performs the processing operation specified by the converted program instruction. The data access circuit 42 is in the form of a load / store unit (LSU), which performs data access to read and write data values, which are executed by processing circuits 48, 50, 52, 54. It is the target of processing operations under the control of the program instructions. The general purpose processor 38 outputs a signal that identifies whether the memory access request is related to instructions or data, and the associated privilege level.
FIG. 2 shows a stylized representative example of the general purpose processor 38, and in fact, it will be recognized that the general purpose processor 38 generally contains a number of additional circuit elements. These circuit elements have been omitted from Figure 2 for clarity.
Data and instructions accessed via the instruction fetch circuit 40 and the data access circuit 42 can be stored at any of the cache layers 6-18 levels and are also stored in the main memory 4. The memory access request is passed through the cache hierarchy until it hits, or to main memory 4 if no hits occur. When no hit occurs, this is a mistake and the cache line containing the data (or instruction) to be accessed is fetched from main memory 4 to cache tiers 6-18. If a mistake occurs in one of the higher-order caches (for example, a mistake occurs in the primary caches 12-18, but a hit occurs in the lower-order cache, such as the secondary or tertiary cache 6 If this happens), the cache line is taken from its lower cache to a cache closer to the source of the higher memory access request. The technique can be used to assign retention priority values at the time of insertion when a cache line is copied from one level in the cache hierarchy to another.
When a memory access request is made by one of the access request sources 20, 22, 24, 26, a signal identifying the source of the memory access request as well as a signal identifying the currently set privilege level of program execution Accompany. As an example, the accompanying signal can distinguish whether the memory access request originates from the instruction fetch circuit 40 or the data access circuit 42. If the result of a memory access request is a mistake, this signal indicating the source of the memory access request is retained when the cache line is inserted into the cache memory where the mistake occurred and is associated with the cache line corresponding to the mistake. Can be used to set the priority value. These different retention priority values, which are then associated with cache lines inserted as a result of instruction fetch or data access, can later be used to control the selection of cache lines to be expelled from the cache memory.
The above describes the setting of different retention priority values based on whether the source is the instruction fetch circuit 40 or the data access circuit 42. Different retention priority values can also be set depending on other factors, such as whether the source was one of the general purpose processors 20, 22 or the graphics processing unit 24. In other embodiments, the two general purpose processors 20, 22 are related, capable of performing different types of tasks and also being used for cache lines that fill the shared cache memory for the tasks. Can have different retention priority values (eg, one processor 20 can perform real-time tasks to which a higher cache retention priority is assigned, and the other processor 22 can have a lower retention priority back. Can perform ground tasks).
Another way to control the retention priority value associated with a cache line inserted into the cache memory (which can be used separately from or in combination with the source identifier) is based on the privilege level associated with that cache line. .. Kernel programs generally run at a higher privilege level than user programs, and cache line retention prioritizes cache lines fetched by the kernel program over cache lines fetched by one or more user programs. Can be biased to (using retention priority values on insertion).
FIG. 3 schematically illustrates a cache control circuit 56 connected to the cache memory 58. The cache control circuit 56 receives a memory access request that specifies a memory address, a signal that identifies whether it is an instruction fetch or a data access from one of the sources of the memory access request, and a privilege level signal. The cache control circuit 56 responds to a software-programmable flag value of 60, which either gave the instruction a higher retention priority than the data, or the data was given a higher retention priority than the instruction. Indicates. In an in-order general purpose processor as illustrated in FIG. 2, it has been found that giving the data cache line a higher retention priority than the instruction cache line provides improved performance. A software programmable flag value of 60 allows this priority to be selected under software control to suit the individual situation. A cache line associated with a memory access request with a kernel privilege level is given a higher retention priority than a cache line associated with a user privilege level and is therefore preferentially retained within cache memory 58. Whether the cache control circuit 56 can provide a memory access request (corresponds to a hit) by the cache line 66 stored in the cache memory 58, or the cache memory 58 does not include the corresponding cache line. Therefore, it controls determining whether a higher hierarchy in the cache memory hierarchy must be referenced (corresponding to a mistake). If a mistake occurs, the data is then returned from a higher hierarchy in the memory hierarchy and stored in cache memory 58 (unless memory access is marked as non-cacheable).
If the cache memory 58 is already full, the existing cache line must be removed from the cache memory 58 to make room for the newly fetched cache line. As described below, the replacement policy may depend on the retention priority value PV stored in the cache memory 58 by the cache control circuit 56 when each cache line is inserted. The retention priority value PV is set to the initial value when the cache line is inserted, but depending on the use of the cache line while the cache line is in the cache memory 58, as well as the possible cache memory 58. It can be modified according to other aspects of the operation (eg, overall demotion to give a pool of potential victim) cache lines.
FIG. 4 schematically illustrates cache line 66. In this embodiment, the cache line contains a TAG value of 68 corresponding to the memory address, from which N words of instructions or data constituting the cache line are taken from the memory address space of the main memory 4. .. The 2-bit retention priority value 70 is also included in the cache line. This retention priority value is coded as indicated by a retention priority value of "00" corresponding to the highest retention priority and a retention priority value of "11" corresponding to the minimum retention priority. In an embodiment further discussed below, when a cache line is inserted into cache memory 58 by the cache controller 56, it is inserted with an initial hold priority value (PV) 70, which hold priority value is the lowest hold priority. It is either a retention priority value of "11", which corresponds to the degree, or a retention priority value of "10", which corresponds to a level one level higher than the lowest retention priority. Other initial values, such as "01" for higher retention priority and "10" for lower retention priority, can be used. When a hit occurs in a cache line already stored in the cache memory, then in some embodiments, if the limit value "00" corresponding to the highest retention priority is reached, the retention priority value is changed. Can be reduced by one. In another embodiment, if a hit occurs for a cache line already stored in the cache memory 58, then its retention priority value may be immediately changed to "00", which corresponds to the highest retention priority.
It will be appreciated that Figure 4 illustrates a cache line with both stored TAGs and data. In practice, the TAG can be stored separately from the data. In this case, the retention priority value can be stored with the TAG along with other bits such as the significant bit and the dirty bit.
The retention priority value is more generally an n-bit retention priority value. This is 2<sup>n</sup>Possible retention priority values K = 0,1, ... 2<sup>n</sup>Give -1. The retention priority values assigned to different classes of cache line inserts (eg data vs. instruction or kernel vs. user) total 2<sup>n</sup>Selected to be -1. This evenly distributes the retention priority values in the available bit space. Different start insert values are, for example, strong priority data (data = 00, instruction = 11), weak priority data (data = 01, instruction = 10), weak priority instruction (data = 10, instruction = 01), or strong priority. It can be selected to be an instruction (data = 11, instruction = 00). These different starting values may be utilized in response to different policies that may be selected between hardware or software configurations.
FIG. 5 is a flow diagram schematically illustrating the operation of the cache control circuit 56 and the cache memory 58 at a given level in the cache layers 6 to 18 when providing a memory access request. At step 72, the process waits for the access request to be received. At step 74, a determination is made as to whether the access request corresponds to data or instructions that already exist in one of the cache lines in cache memory 58. If the memory access request corresponds to a cache line that already exists, this is a hit and the process proceeds to step 76. In step 76, if the lowest possible hold priority value "00" is reached, the hold priority value of the hit cache line is decremented by one (to raise the hold priority level by one). Corresponding). After step 76, step 78 serves to provide a memory access request from cache memory 58, and the process returns to step 72.
If no hits are detected at step 74, the process proceeds to step 80, where a cache miss operation is triggered, as described further below. The process then waits in step 82 for the data or instruction cache line to be returned to cache memory 58. When the data / instruction is returned and inserted into cache memory 58 (or in parallel with such an insert), processing proceeds to step 78, where a memory access request is provided.
FIG. 6 is a flow diagram schematically illustrating a miss operation performed by the cache control circuit 56 and the cache memory 58 in response to a miss, as triggered in step 80 of FIG. At step 84, processing waits until a cache miss is detected. At step 86, it is determined whether the cache memory is currently full. If the cache memory is not currently full, the process proceeds to step 88, where one of the currently empty cache lines is randomly selected as the "victim" and the newly fetched cache line is filled into it. The cache.
If it is determined in step 86 that the cache is already full, the victim must be selected from the currently occupied cache lines. In step 90, it is determined if there is any currently occupied cache line with the lowest retention priority. In the example discussed above, this corresponds to a cache line with a retention priority value of "11". If there is no such cache line with the lowest hold priority, the process proceeds to step 92, where the hold priority level of all cache lines currently held in cache memory 58 is one. Demote. This can be achieved by changing all the stored retention priority values PV. However, a more efficient embodiment may change the mapping between the stored retention priority value and the retention priority level in order to achieve the same result. After the retention priority level of all cache lines has been demoted by one, the process returns to step 90, where it repeats checking any cache line for the lowest retention priority level. The process circulates through steps 90-92 until at least one cache line with the lowest hold priority level is detected in step 90.
When at least one cache line with the lowest retention priority level is detected according to step 90, step 94 serves to randomly select a victim from the cache lines with the lowest retention priority (single cache). If the line has this lowest retention priority, the selection will be this single line). Random selection of Victim from the lowest retention priority cache line is relatively easy to achieve. Other selection policies, such as the least recently used policy or the round-robin policy, may be adopted when selecting the lowest retention priority cache line from this pool. Such different policies generally require additional resources to remember the state associated with their operation.
If a victim cache line is selected in step 94, step 96 determines if the cache line is dirty (including updated values that are different from the cache line currently held in main memory 4). To do. If the cache line is dirty, step 98 serves to write the cache line back to main memory 4 before it is overwritten by the newly inserted cache line.
In step 100, in this embodiment, it is determined whether the memory access that triggered the error originated from the instruction fetch unit 40. If the memory access originated from the fetch unit, the cache line contains at least the instructions that should have been fetched and is marked with a hold priority value of "10" in step 102. If it is determined in step 100 that the memory access is not from the instruction fetch unit, it is from the load / store unit 42 (data access circuit) and in step 104 the retention priority value PV "01" is set. To. Then, in step 106, it waits until it receives a new cache line from a lower level in the memory hierarchy (eg, one of the lower level cache memory or main memory). Upon receiving the cache line in step 106, step 108 then serves to write a new cache line with the initial hold priority value PV to the cache line storing the victim selected in step 94.
FIG. 7 is a flow chart similar to the flow chart of FIG. The difference in Figure 7 is that the retention priority value is determined based on whether the privilege level associated with the cache line where the cache miss occurred is the kernel privilege level that corresponds to the operating system's kernel program. .. Once the victim cache line is selected, then in step 110 it is determined whether the memory access where the cache line miss occurred has the privilege level corresponding to the kernel program. If it is determined in step 110 that the memory access has the kernel privilege level corresponding to the kernel program, step 112 serves to set the retention priority value PV to "01". Conversely, if the memory access has a user privilege level and does not correspond to the kernel program, step 114 serves to set the initial hold priority value PV to "10". Thus, the cache line associated with the kernel program is given a higher initial hold priority and a cache line associated with other programs (eg, user programs).
Although exemplary embodiments of the invention have been described in detail herein with reference to the accompanying drawings, the present invention is not limited to those exact embodiments, and various modifications and modifications and variations and embodiments are made. It should be understood that modifications may be made by one of ordinary skill in the art without departing from the scope and gist of the invention as set forth in the appended claims.
2 Data processing device 4 Main memory 6 Tertiary cache memory 8,10 Secondary cache memory 12,14,16,18 Primary cache memory 20,22,24,26 Access request source 38 General purpose processor 40 instruction fetch circuit 42 Data access circuit 44 In-order instruction pipeline 46 Instruction decoder 48 register bank 50 multiplier 52 shifter 54 adder
7 sheets
Sheet 1 Sheet 2 Sheet 3 Sheet 4 Sheet 5 Sheet 6 Sheet 7
Every citation, both ways
| Document | Relation | Office |
|---|---|---|
| US20080229052A1 | Cites | United States of America |
| US06532520B1 | Cites | United States of America |
| JP10307756A | Cites | Japan |
| JP05314120A | Cites | Japan |
| JP2004110240A | Cites | Japan |
| US20100235579A1 | Cites | United States of America |
| US20100125677A1 | Cites | United States of America |
9 members in 4 offices
Priority claims5
| Document | Office | Kind | Date |
|---|---|---|---|
| 13713999 | United States of America | – | |
| 201213713999 | United States of America | A | |
| 201213713999 | United States of America | A | |
| 13713999 | – | – | – |
| US201213713999 | – | – | – |
Members9
| Document | Office | Kind | |
|---|---|---|---|
| GB201318092D0 | United Kingdom | D0 | |
| CN103870394A | China | A | |
| GB2508962A | United Kingdom | A | |
| US2014173214A1 | United States of America | A1 | |
| JP2014123357A | Japan | A | |
| US9372811B2 | United States of America | B2 | |
| JP6325243B2This record | Japan | B2 | |
| CN103870394B | China | B | |
| GB2508962B | United Kingdom | B |
13 legal events, as the office reported them to INPADOC
Over the term
Point at a mark for the eventEvents
| Event | Code | |
|---|---|---|
| Receipt of annual feesJAPANESE INTERMEDIATE CODE: R250R250 | R250 | |
| Receipt of annual feesJAPANESE INTERMEDIATE CODE: R250R250 | R250 | |
| Receipt of annual feesJAPANESE INTERMEDIATE CODE: R250R250 | R250 | |
| Receipt of annual feesJAPANESE INTERMEDIATE CODE: R250R250 | R250 | |
| Receipt of annual feesJAPANESE INTERMEDIATE CODE: R250R250 | R250 | |
| Certificate of patent or registration of utility modelJAPANESE INTERMEDIATE CODE: R150R150 | R150 | |
| First payment of annual fees (during grant procedure)JAPANESE INTERMEDIATE CODE: A61A61 | A61 | |
| Written decision to grant a patent or to grant a registration (utility model)JAPANESE INTERMEDIATE CODE: A01A01 | A01 | |
| Decision of grant or rejection writtenTRDD | TRDD | |
| Request for written amendment filedJAPANESE INTERMEDIATE CODE: A523A521 | A521 | |
| Notification of reasons for refusalJAPANESE INTERMEDIATE CODE: A131A131 | A131 | |
| Report on retrievalJAPANESE INTERMEDIATE CODE: A971007A977 | A977 | |
| Written request for application examinationJAPANESE INTERMEDIATE CODE: A621A621 | A621 |
Numbers
- Publication
- 6325243
- Publication, DOCDB
- 6325243
- Publication, EPODOC
- JP6325243B
- Application
- 256657
- Application, DOCDB
- 2013256657
- Application, EPODOC
- JP20130256657
Titles2
- Japanese
- 保持優先度に基づくキャッシュ置換ポリシー
- English
- Cache replacement policy based on retention priority
Classification
- CPC, 3
- G06F12/126
- G06F12/0808
- G06F12/08
- IPC, 2
- G06F12 12
- G06F12 08
