EP0905628A2

Reducing cache misses by snarfing writebacks in non-inclusive memory systems

Abstract

The present invention is a method and apparatus for optimizing a non-inclusive multi-level cache memory system. Cache space is provided in a first cache by removing a first cache content, such as a first cache line or block, from the first cache. The removed first cache content is stored in a second cache. This is done in response to a cache miss in the first and second caches. All cache contents that are stored in the second cache are limited to have read-only attributes so that if any copies of the cache contents in the second cache exist in the cache memory system, a processor or equivalent device must seek permission to access the location in which that copy exists, ensuring cache coherency. If the first cache content is required by a processor such as when a cache hit occurs in the second cache for the first cache content, room is again made available, if required, in the first cache by selecting a second cache content from the first cache and moving it to the second cache. Once room is available in the first cache, the first cache content is moved from the second cache to the first cache, rendering the first cache available for write access. Limiting the second cache to read-only access reduces the number of status bits per tag that are required to maintain cache coherency. In a cache memory system using a MOESI protocoal, the number of status bits per tag is reduced to a single bit for the second cache, reducing the tag overhead. An optimization advantage in that the tags may be placed on-chip so as to improve cache bandwidth.

EP0905628A2, drawing sheet 1
Sheet 1 of 5

Term

Term ended

Projected expiry passed 18 September 2018, 8 years ago.

  1. Priority
  2. Filed
  3. Published
  4. Projected expiry
  5. Today

22 claims: 22 independent, 0 dependent

  1. 1
    A method for optimizing a non-inclusive hierarchical cache memory system, comprising the steps of:providing cache space in a first cache by removing a first cache content from said first cache and storing said first cache content in a second cache, said step of providing in response to a cache miss in said first cache and said second cache;obtaining a second cache content by fetching a line of information from a memory store in response to said cache miss in said first cache and said second cache;storing said second cache content within said cache space provided in said first cache;precluding a store operation directed to said first cache content stored in said second cache when a copy of said first cache content exists in the non-inclusive hierarchical cache memory system;andtransferring said second cache content stored within said first cache to said second cache and said first cache content within said second cache to said first cache in response to a cache miss in said first cache and a cache hit in said second cache that arises in response to a cache fetch request to the cache memory system for said first cache content.
  2. 2
    The method in claim 1, further including the step of creating and maintaining tags and states which correspond to said second cache, said tags and states located on the same silicon real estate as a corresponding processor.
  3. 3
    The method in claim 1, wherein said second cache is an external cache.
  4. 4
    The method in claim 1, further including the step of defining said first cache as a level two cache.
  5. 5
    The method in claim 1, further including the step of defining said second cache as a level three cache.
  6. 6
    A method for improving memory latency in a non-inclusive hierarchical cache memory system, comprising the steps of:providing cache space in a first cache by removing a first cache content from said first cache and storing said first cache content in a second cache, said step of providing in response to a cache miss in said first cache and said second cache;obtaining a second cache content by fetching a line of information from a memory store in response to said cache miss in said first cache and said second cache;storing said second cache content within said cache space provided in said first cache;precluding a store operation directed to said first cache content stored in said second cache when a copy of said first cache content exists in the non-inclusive hierarchical cache memory system;transferring said second cache content stored within said first cache to said second cache and said first cache content within said second cache to said first cache in response to a cache miss in said first cache and a cache hit in said second cache that arises in response to a cache fetch request to the cache memory system for said first cache content;andcreating and maintaining tags and states which correspond to said second cache, said tags and states located on the same silicon real estate as a corresponding processor.
  7. 7
    The method in claim 6, further including the step of defining said first cache as a level two cache.
  8. 8
    The method in claim 6, further including the step of defining said second cache as a level three cache.
  9. 9
    A method for improving the rate of cache hits while minimizing cache coherency overhead in a cache memory system, the method comprising the steps of:providing cache space in a first cache by removing a first cache content from said first cache and storing said first cache content in a second cache, said step of providing in response to a cache miss in said first cache and said second cache;obtaining a second cache content by fetching a line of information from a memory store in response to said cache miss in said first cache and said second cache;storing said second cache content within said cache space provided in said first cache;precluding a store operation directed to said first cache content stored in said second cache when a copy of said first cache content exists in the non-inclusive hierarchical cache memory system;transferring said second cache content stored within said first cache to said second cache and said first cache content within said second cache to said first cache in response to a cache miss in said first cache and a cache hit in said second cache that arises in response to a cache fetch request to the cache memory system for said first cache content;creating and maintaining tags and states which correspond to said second cache, said tags and states located on the same silicon real estate as a corresponding processor;defining said first cache as a level two cache;anddefining said second cache as a level three cache.
  10. 10
    An apparatus for optimizing a non-inclusive hierarchical cache memory, comprising:a first cache for storing cache content;a second cache responsive to receiving a first cache content from said first cache when a cache miss in said first cache and said second cache for a second cache content occurs, said second cache precluded from servicing a store operation corresponding to said first cache content when a copy of information corresponding to said first cache content exists in the cache memory, anda second cache content obtained from a memory store in response to said cache miss in said first cache and said second cache, said second cache content stored in said first cache, said second content transferable from said first cache to said second cache and said first content transferable from said second cache to said first cache in response to a cache miss in said first cache and a cache hit in said second cache that arises in response to a fetch request to said cache memory system.
  11. 11
    The apparatus in claim 10, further including a directory for holding tags and states which correspond to said second cache, said directory located on the same silicon real estate as a corresponding processor.
  12. 12
    The apparatus in claim 10, wherein said second cache is an external cache.
  13. 13
    The apparatus in claim 10, wherein said first cache is defined as a level two cache.
  14. 14
    The apparatus in claim 10, wherein said second cache is defined as a level three cache.
  15. 15
    The apparatus in claim 10, further including a processor coupled to said first cache and said second cache;a system request bus coupled to said first cache and said second cache;and a main memory store coupled to said system request bus.
  16. 16
    An apparatus for improving memory latency in a non-inclusive hierarchical cache memory, comprising:a cache line in a first cache, said cache line obtained by removing a first cache line from said first cache in response to a cache miss in said first cache and a second cache, said second cache for storing said first cache line from said first cache, said second cache precluded from servicing a store operation corresponding to said first cache line when a copy of information corresponding to said first cache line exists in the cache memory;a second cache line in said first cache to said second cache and said first content in said second cache to said first cache in response to a cache miss in said first cache and a cache hit in said second cache, said cache miss and said cache hit in said first cache and said second cache, respectively, occurring in response to a fetch request to said cache memory system;andtags and states which correspond to said second cache, said tags and states located on the same silicon real estate as a corresponding processor.
  17. 17
    The apparatus in claim 16, wherein said first cache is defined as a level two cache.
  18. 18
    The apparatus in claim 16, wherein said second cache is defined as a level three cache.
  19. 19
    The apparatus in claim 16, further including a processor coupled to said first cache and said second cache;a system request bus coupled to said first cache and said second cache;and a main memory store coupled to said system request bus.
  20. 20
    An apparatus for minimizing main memory fetches due to cache misses in a cache memory system, comprising:a cache line in a first cache, said cache line obtained by removing a first cache line from said first cache in response to a cache miss in said first cache and a second cache, said second cache for storing said first cache line from said first cache, said second cache precluded from servicing a store operation corresponding to said first cache line when a copy of information corresponding to said first cache line exists in the cache memory;a second cache line in said first cache to said second cache and said first content in said second cache to said first cache in response to a cache miss in said first cache and a cache hit in said second cache, said cache miss and said cache hit in said first cache and said second cache, respectively, occurring in response to a fetch request to said cache memory system;tags and states which correspond to said second cache, said tags and states located on the same silicon real estate as a corresponding processor;defining said first cache as a level two cache;anddefining said second cache as a level three cache.
  21. 21
    The apparatus in claim 20, further including a processor coupled to said first cache and said second cache;a system request bus coupled to said first cache and said second cache;and a main memory store coupled to said system request bus.
  22. 22
    A method for providing a computer system, comprising the steps of:providing a non-inclusive hierarchical cache memory system including:a first cache for storing cache content;a second cache responsive to receiving a first cache content from said first cache when a cache miss in said first cache and said second cache for a second cache content occurs, said second cache precluded from servicing a store operation corresponding to said first cache content when a copy of information corresponding to said first cache content exists in the cache memory;anda second cache content obtained from a memory store in response to said cache miss in said first cache and said second cache, said second cache content stored in said first cache, said second content transferable from said first cache to said second cache and said first content transferable from said second cache to said first cache in response to a cache miss in said first cache and a cache hit in said second cache that arises in response to a fetch request to said cache memory system.
Independent claims22