EP0806726B1

On-line memory monitoring system and methods

Abstract

This record has no abstract on file.

EP0806726B1, drawing sheet 1
Sheet 1 of 5

Term

Term ended

Expired 28 April 2017, 9.4 years ago.

  1. Priority
  2. Filed
  3. Granted
  4. Expired
  5. Today

11 claims: 3 independent, 8 dependent

  1. 1
    A method of improving memory reliability in a computer comprising:(a) providing an error correction code with data stored in memory;(b) detecting and correcting single-bit memory errors as they occur using the error correction code;(c) determining the rate at which single-bit errors occur;and (d) providing a warning if the operating conditions are abnormal whereby to indicate the potential occurrence of multiple-bit errors.     characterised by (e) employing a statistical inference method to determine from said rate, and average failure rates for the memory, whether the memory is running under normal or abnormal conditions;
  2. 2
    On-line memory monitoring apparatus comprising:a processor;a read/write random access memory coupled to the processor, the random access memory being configured to store data words and associated error correction codes;error detection and correction circuitry coupled to the read/write random access memory, the error detection and correction circuitry being configured to determine the specific error correction code to be written to the memory with each data word to be written to memory, and to detect and correct single-bit errors in data and associated error correction codes as read from memory;and an error monitor configured to respond to the error detection and correction circuitry to provide a warning indicating the potential occurrence of multiple-bit errors, if the memory is running under abnormal conditions, characterised in that the monitor generates a log of detected single-bit errors and in that the provision of said warning is computed by statistical inference from the rate at which single-bit memory errors have occurred and average failure rates from the memory.
  3. 3
    The apparatus of Claim 2, wherein the processor is configured to write the corrected data and associated error correction code back into the memory at the memory location from which it was read upon detection and correction of an error in data and associated error correction code read from the memory.
  4. 4
    The apparatus of Claim 3, wherein the error monitor is configured to respond to the error detection and correction circuitry to provide a said warning if the rate at which single-bit memory errors have occurred over either a first or a second time period exceeds first and second predetermined memory error rate limits, respectively.
  5. 5
    The apparatus of Claim 4 wherein the error monitor is configured to provide a said warning indicative of a memory failure if the rate at which single-bit memory errors have occurred over the first time period exceeds the first predetermined memory error rate limit, and of providing a said warning indicative of an unusually high error rate if the rate at which single-bit memory errors have occurred over the second time period exceeds the second predetermined error rate limit.
  6. 6
    A system for on-line memory monitoring responsive to the detection and correction of a single-bit memory error, the system including code configured for storage on a computer-readable apparatus and executable by a computer, the code including a plurality of modules, the system including:a first module configured to maintain a single-bit memory error log;a second module configured to respond to the detection and correction of a single-bit memory error to determine using the memory error log if the rate at which single-bit memory errors have occurred exceeds a predetermined limit;a third module logically coupled to the second module and configured to provide a warning if the second module determines that the rate at which single-bit memory errors have occurred exceeds the predetermined limit;and a fourth module logically coupled to the first module and configured to update the error log upon the detection and correction of a single-bit memory error;wherein the system employs a statistical inference method to determine from said rate, and average failure rates for the memory, whether the memory is running under normal or abnormal operating conditions, abnormal operating conditions corresponding to the single-bit memory rate exceeding the predetermined limit;and wherein the warning is indicative of abnormal operating conditions and the potential occurrence of multiple-bit errors.
  7. 7
    The system of Claim 6, further comprising a fifth module configured to overwrite the memory after detection and correction of a memory error.
  8. 8
    A system of Claim 6, wherein the second module is configured to respond to the detection and correction of a single-bit memory error to determine using the memory error log if the rate at which single-bit memory errors have occurred exceeds a first or a second predetermined limit and the third module is configured to provide a warning of a first character if the second module determines that the rate at which memory errors have occurred exceeds the first predetermined limit, and further comprising a fifth module configured to provide a warning of a second character if the second module determines that the rate at which memory errors have occurred exceeds the second predetermined limit.
  9. 9
    A method as claimed in Claim 1, including the steps of determining the rate at which single-bit memory errors have occurred over a first elapsed time and determining the rate at which single-bit memory errors have occurred over a second elapsed time longer than the first elapsed time, and wherein a said warning is also provided when the rate at which single-bit memory errors have occurred over either the first or the second time period exceeds first and second predetermined memory error rate limits, respectively.
  10. 10
    A method as claimed in Claim 1, further comprising:using the corrected data as error-free data;writing the corrected data and error correction code back into the same memory location from which it was read;and providing a respective warning if the rate at which single-bit memory errors have occurred exceeds a predetermined limit.
  11. 11
    The apparatus as claimed in Claim 2, further comprising:a CPU/memory board having at least one bus connector for connecting to a system bus, wherein the processor is coupled to the bus connector, the processor being configured to write the corrected data and associated error correction code back into the memory at the memory location from which it was read upon detection and correction of an error in data and associated error correction code read from the memory.