Memory error identification based on corrupted symbol patterns
Summary by NHIP
Memory Error Discrimination
The system identifies memory error types by analyzing corrupted symbol patterns within transmitted codewords. It distinguishes chip failures from first or second channel pin failures based on whether corrupted symbols appear adjacent to each other or spaced according to channel width ratios.
Claim Score by NHIP
Abstract
A system includes a memory controller, a buffer, a first channel to couple the memory controller to the buffer, and a second channel to couple the buffer to a memory. The first channel and second channel are to transmit a codeword including a plurality of symbols. A symbol is formed from a plurality of bursts based on data access of the memory. The memory controller is to identify a memory error based on a corrupted symbol pattern of the codeword. The memory controller is to discriminate between a chip failure, a first pin failure of the first channel, and a second pin failure of the second channel, as being a type of the memory error, according to the corrupted symbol pattern.

Term
6.4 yearsleft in the term
Expires 11 February 2033, including 73 days of term adjustment.
- Priority and filed
- Granted
- Today
- Expires
15 claims: 3 independent, 12 dependent
- 1A system comprising:a memory controller;a buffer;a first channel having a first channel width, to couple the memory controller to the buffer;and a second channel having a second channel width, to couple the buffer to a memory;wherein a first channel width differs from a second channel width, and the first channel and second channel are to transmit a codeword including a plurality of symbols, wherein a symbol is formed from a plurality of bursts based on data access of the memory;and wherein the memory controller is to identify a type of memory error based on a corrupted symbol pattern of the codeword, wherein the memory controller is to discriminate between a chip failure, a first pin failure of the first channel, and a second pin failure of the second channel, as being the type of memory error, according to the corrupted symbol pattern.
- 10Broadest claimClaim Score 48, average(NHIP)A method, comprising:receiving a codeword across a first channel coupling a memory controller to a buffer, and a second channel coupling the buffer to a memory, wherein a first channel width differs from a second channel width, and the codeword includes a plurality of data symbols and at least one check symbol, wherein each of the plurality of symbols is formed from a plurality of bursts based on data access of the memory;correcting, using a memory controller and the at least one check symbol, a memory error based on a corrupted symbol pattern of the codeword;and discriminating between a chip failure, a first pin failure of the first channel, and a second pin failure of the second channel, as being a type of the memory error, according to the corrupted symbol pattern.
- 14A non-transitory computer readable medium having instructions stored thereon executable by a processor to cause a memory controller to:receive a codeword across a first channel, coupling a memory controller to a buffer, and a second channel coupling the buffer to a memory, wherein a first channel width differs from a second channel width, and the codeword includes a plurality of data symbols and at least one check symbol, wherein each of the plurality of symbols is formed from a plurality of bursts based on data access of the memory;correct, using the memory controller and the at least one check symbol, a memory error based on a corrupted symbol pattern of the codeword;and discriminate between a chip failure, a first pin failure of the first channel, and a second pin failure of the second channel, as being a type of the memory error, according to the corrupted symbol pattern.
Independent claims3
45 paragraphs in 3 sections, as filed
BACKGROUND
System reliability in computer systems can be affected by system memory, which can be a common source of system failures. Memory modules, such as dual in-line memory modules (DIMMs), may use error-correcting code (ECC) to detect and correct some memory errors. However, ECC may be applied inefficiently and without discriminating between different types of memory failures. This may lead to unnecessary replacement of a memory module, even though the error may be related to a memory channel failure and not the memory module itself.
BRIEF DESCRIPTION OF THE DRAWINGS/FIGURES
<figref idref="DRAWINGS">FIG. 1</figref> is a block diagram of a system including a memory controller according to an example.
<figref idref="DRAWINGS">FIG. 2</figref> is a block diagram of a system including a memory controller according to an example.
<figref idref="DRAWINGS">FIG. 3A</figref> is a block diagram of a data block including a data symbol according to an example.
<figref idref="DRAWINGS">FIG. 3B</figref> is a block diagram of a data block including a data symbol according to an example.
<figref idref="DRAWINGS">FIG. 4</figref> is a flow chart based on discriminating a type of memory error according to an example.
<figref idref="DRAWINGS">FIG. 5</figref> is a flow chart based on discriminating a type of memory error according to an example.
DETAILED DESCRIPTION
Example systems described herein are capable of discriminating between different types of pin failures and chip failures, and may enable stronger protection for errors without a need for increased ECC overhead. These benefits are compatible with buffered memory systems. ECC codeword symbols may be reorganized, to leverage burst access of the memory. A memory controller can analyze a corrupted symbol pattern in codeword symbols, and take different actions for different types of memory errors, thereby improving memory system reliability.
An example system may include a memory controller; a buffer; a first channel, and a second channel. The first channel has a first channel width to couple the memory controller to the buffer, and the second channel has a second channel width to couple the buffer to a memory. The first channel width may differ from the second channel width. The first channel and second channel are to transmit a codeword including a plurality of symbols. A symbol is formed from a plurality of bursts based on data access of the memory. The memory controller is to identify a memory error based on a corrupted symbol pattern of the codeword, and discriminate between a chip failure, a pin failure of the first channel, and a pin failure of the second channel, as being the type of memory error, according to the corrupted symbol pattern.
<figref idref="DRAWINGS">FIG. 1</figref> is a block diagram of a system <b>100</b> including a memory controller <b>102</b> according to an example. System <b>100</b> also includes a first channel <b>110</b>, buffer <b>104</b>, and second channel <b>120</b>. The first channel <b>110</b> is to couple the memory controller <b>102</b> to the buffer <b>104</b>, and the second channel <b>120</b> is to couple the buffer <b>104</b> to a memory <b>106</b>. The system <b>100</b> is to interact with a codeword <b>130</b>, which may be transmitted via the first channel <b>110</b> and/or the second channel <b>120</b>. The codeword <b>130</b> includes a plurality of symbols <b>132</b>. A symbol <b>132</b> includes a plurality of bursts <b>134</b>. The memory controller <b>102</b> is to identify a corrupted symbol pattern <b>136</b> based on the codeword <b>130</b>. The memory controller <b>102</b> also is to discriminate a type of memory error <b>140</b>, based on the corrupted symbol pattern <b>136</b>. Thus, the memory controller <b>102</b> may determine whether the memory error <b>140</b> is based on a chip failure <b>142</b> in the memory <b>106</b>, a first pin failure <b>144</b> of the first channel <b>110</b>, and a second pin failure <b>146</b> of the second channel <b>120</b>.
The first channel <b>110</b>, buffer <b>104</b>, and second channel <b>120</b> enable the memory controller <b>102</b> to interact with the memory <b>106</b> based on buffering. Buffer <b>104</b>, and memory <b>106</b>, are each shown as a single block for convenience. However, buffer <b>104</b> may represent multiple buffers, and memory <b>106</b> may represent multiple separate memories (e.g., memory modules). For example, multiple buffers <b>104</b> may be associated with the first channel <b>110</b>, and the second channel <b>120</b> may be connected to multiple memory ranks. A buffered memory system may use channels of different widths. For example, the memory <b>106</b> may interface with the second channel <b>120</b> based on a much wider data path than the first channel <b>110</b> (the data path between the buffer <b>104</b> and the memory controller <b>102</b>). In an example, the second channel <b>120</b> on one side of the buffer <b>104</b> may have a 144-bit wide channel, and first channel <b>110</b> on the other side of the buffer <b>104</b> may have a 72-bit wide channel. Thus, the first channel <b>110</b> and second channel <b>120</b> have different widths. A high-end system may have a buffer <b>104</b> on-board, with a narrow bus, e.g., a first channel <b>110</b> having a 20-bit width running at much higher frequency compared to the second channel <b>120</b> (e.g., a 144-bit wide channel or similar). Thus, in a buffered memory system, one side of the buffer is narrower and faster, and the other side of the buffer is wider and slower. The widths may reflect a ratio, such as a 1:2 ratio, in terms of channel width. Thus, the narrower channel may run proportionally faster than the wider channel.
Accordingly, the codeword <b>130</b> may be transmitted across the second channel <b>120</b> based on a wide interface, but a shape of the codeword <b>130</b> may be reorganized when transmitted over the narrower interface of the first channel <b>110</b> such that the symbols <b>132</b> are aligned differently relative to pins in the channel. Accordingly, it may be inappropriate to treat all pin failures generically as a subset of chip failures; because the reorganized codeword <b>130</b> may be affected differently from a pin failure in the first channel <b>110</b> in view of the reorganized codeword <b>130</b> and different channel widths. Example systems provided herein may organize the codeword <b>130</b> to efficiently tolerate pin failures on either channel in the buffered system having different channel widths, discriminating between different types of pin failures and chip failures. For example, the corrupted symbol pattern <b>136</b>, arising in view of codeword organization, may enable example systems to correct and identify the different types of errors.
<figref idref="DRAWINGS">FIG. 2</figref> is a block diagram of a system <b>200</b> including a memory controller <b>202</b> according to an example. System <b>200</b> also includes a first channel <b>210</b>, buffer <b>204</b>, and second channel <b>220</b> to interface with memory <b>206</b>. The first channel <b>210</b> and second channel <b>220</b> may transmit data block <b>238</b>. The data block <b>238</b> includes a plurality of data symbols <b>233</b> and check symbols <b>235</b>. The memory controller <b>202</b> is to identify a corrupted symbol pattern <b>236</b> (one type of pattern is shown, others are possible) of a codeword <b>230</b>, which may be based on logic <b>208</b>. The memory controller <b>202</b> is to determine a type of memory error, such as a first pin failure <b>244</b> of the first channel <b>210</b> causing the corrupted symbol pattern <b>236</b>. The memory <b>206</b> includes a plurality of chips <b>237</b> that interface with the second channel <b>220</b> via chip outputs <b>239</b>.
To protect memory <b>206</b> from various sources of errors, error checking and correcting (ECC) codes may be applied to memory <b>206</b>. An ECC dual in-line memory module (DIMM) is a memory module that may include dynamic random access memory (DRAM) chips for data as well as ECC. In an example, the memory <b>206</b> includes two memory modules, each being a DIMM having a 72-bit wide data interface, including 64-bit data and 8-bit ECC. Details of the ECC mechanism may vary based on system designers for different examples. An example may use single bit-error correcting and double bit-error detecting (SEC-DED) code. An 8-bit SEC-DED code can correct 1-bit errors and detect 2-bit errors. Being able to correct 1-bit errors also means that a pin failure may be tolerated (depending on the arrangement of the codeword <b>230</b>), if a pin failure appears to be a 1-bit failure per access (e.g., in the second channel <b>220</b>). <figref idref="DRAWINGS">FIG. 2</figref> illustrates two ECC DIMMs, each having 9×8 DRAM chips <b>237</b>. Data chips are shown in light gray, and ECC chips are show in darker gray. A x8 DRAM has 8 data pins, although other types of chips may be used.
System <b>200</b> may combine the two ECC DIMMs together to form a 144-bit wide channel, with 128 bits of data and 16 bits of ECC. An example approach would be to apply an 8-bit symbol-based Reed-Solomon (RS) code horizontally across the entire 144-bit wide channel. If using burst 4 access, because the total ECC is applied horizontally, the actual burst number (e.g., burst 4, burst 8, and so on) may be varied. Burst 4 is used for clarification and other burst values may be used. DRAMs based on increasing data rates may involve burst lengths longer than 4 or 8, so using, e.g., 2 burst 4 bit or 4 burst 2 bit organization would not affect DRAM behavior regarding the burst transfer. DRAM behavior in modern DRAMs may not be affected by burst length and symbol organization, allowing for flexible combinations of memory types and performance.
The data block <b>238</b> includes a plurality of symbols, represented by a rectangular region that represents an 8-bit symbol. Because the chips <b>237</b> are x8, each burst provides 8 bits. Thus, we see the data block <b>238</b> includes, along a horizontal axis as illustrated, a symbol corresponding to each chip <b>237</b>. Because the memory <b>206</b> is burst 4, the data block <b>238</b> includes, along a vertical axis as illustrated, 4 rows of symbols. Note that the data block <b>238</b> is shown in a “stacked” or “folded” manner, where half of the data block <b>238</b> is positioned above the other half, corresponding to transmission over the narrower first channel <b>210</b>. However, the symbols in each half are stacked 4 high corresponding to burst 4.
Other symbol organizations are possible. The data block <b>238</b> is shown as a folded organization example, with a first set of four bursts (e.g., corresponding to the left DIMM of the memory <b>206</b>) being transmitted, followed by the second set of four bursts (e.g., corresponding to the right DIMM of memory <b>206</b>). However, the contents may be interleaved, or arranged in other ways. For example, the first burst of the left DIMM may be sent, followed by the first burst of the right DIMM, then the second burst of the left DIMM, followed by the second burst of the right DIMM, then the third, and so on. Thus, references to the data block <b>238</b> being folded also may include other variations of symbol organizations, such as interleaved and so on, to arrange the symbols for a narrower symbol arrangement corresponding to the narrower channel width.
The symbols that form the data block <b>238</b> (and codeword <b>230</b>) may include a data symbol <b>233</b> and a check symbol <b>235</b>. Each codeword <b>230</b> may include 16 data symbols <b>233</b> and two check symbols <b>235</b>. The block <b>238</b> is shown including a burst of 4 codewords <b>230</b> stacked vertically. A two check symbol Reed Solomon code may correct one symbol error, which may correspond to a chip failure that affects one symbol per region of ECC code. Therefore, with this 2 check symbol ECC code and 144-bit wide channel organization, it is possible to correct a chip failure. Thus, the system <b>200</b> supports chipkill protection.
Example systems may experience various errors, including a single-bit error, a multi-bit error, complete row failure, or any internal logic failure. There may be input/output (I/O) pin failures, permanent errors, intermittent errors, latent errors, and other errors. All of these types of errors may be detected and/or corrected. Regardless of how the errors manifest, if those errors are confined to a single chip, chipkill may be used to protect against them. However, due to the folded organization of the data block <b>238</b> in view of the different widths of the first channel <b>210</b> and second channel <b>220</b>, errors associated with a pin may manifest differently, and other techniques also may be applied for detection and/or correction.
In example systems (e.g., high-end servers), chipkill-correct may be used to tolerate chip failures and pin failures. Symbol-based Reed-Solomon (RS) codes may be used to implement chipkill-correct, and a wide channel configuration (128-bit data and 16-bit ECC, by tying two ECC DIMMs in lock-step mode) may be used to limit ECC overhead to 12.5%. <figref idref="DRAWINGS">FIG. 2</figref> illustrates a memory channel configuration for chipkill-correct, with a data and ECC block <b>238</b> with burst 4 access (burst 4 is used as an example, but this approach is applicable to other types of access, e.g., burst 8 in DDR3). Each access is composed of 16 data symbols and 2 check symbols (using 8-bit symbols).
Using the example burst 4 access, at each read in this 144-bit wide channel, the system <b>200</b> is to access 64 bytes of data from the DRAM memory <b>206</b>, in view of the 144-bit wide channel. The 64-byte data will be transferred as shown in the block <b>238</b>. For example, a width of the first channel <b>210</b> (between the memory controller <b>202</b> and the buffer <b>204</b>) may be narrower (e.g., 72-bits wide) than the 144-bit wide read from the memory. This difference in channel widths may result in organization where the data originates as a wide shape for the second channel <b>220</b>, and is folded to a narrower shape for the first channel <b>210</b>.
However, if there is a pin failure in the first channel <b>210</b> (between the memory controller <b>202</b> and the buffer <b>204</b>), a larger number of symbols will be affected, because each pin in the first channel <b>210</b> transfers twice as much information (in this particular example demonstrating a 1:2 channel width ratio; other ratios are possible, including non-integer ratios). For example, the second channel <b>220</b> may be twice as wide as the first channel <b>210</b>, so data from the wide second channel <b>220</b> will need to be transferred over a narrower data bus. Thus, symbols corrupted by a pin failure on the first channel <b>210</b> may affect two symbols in the codeword <b>230</b>. The affected symbols may form a corrupted symbol pattern <b>236</b>. Thus, a single pin in the first channel <b>210</b> may affect an entire column of the folded block <b>238</b>, corrupting those symbols corresponding to that column (i.e., two symbols per each of the four codewords <b>230</b>).
The 16 data symbol, 2 check symbol ECC codeword <b>230</b> is shown formed in a folded shape because the data layout is changed and transferred over a narrower channel. Thus, a chip failure is correctable because it causes merely a single-symbol error in codeword <b>230</b>. However, the symbols corrupted by a pin failure in the first channel <b>210</b> appear as two symbol errors per ECC codeword <b>230</b>, which may not be correctable using chipkill described above for this arrangement of 8-bit symbols.
A pin failure in the second channel <b>220</b>, in cases where the second channel <b>220</b> is as wide as the memory <b>206</b>, can be considered to be a subset of a chip failure, because one pin failure would affect one of the multiple pins of a chip <b>237</b>. Such a pin failure may be corrected as described above. However, a pin failure on the first channel <b>210</b>, between the buffer <b>204</b> and the memory controller <b>202</b>, is not a subset of a chip failure because the buffered system using different channel widths and folded arrangement of the codeword may cause more than one symbol to be corrupted. In this arrangement, because each chip produces a symbol, corruption of more than one symbol is comparable to corruption of more than one chip. Thus, a single pin failure in the first channel <b>210</b> may cause corruption to symbols as though multiple chips have failed.
However, by changing the organization of the symbols in view of the different channel widths, it is possible to identify whether a memory error is due to a pin failure in the first channel <b>210</b>, a pin failure in the second channel <b>220</b>, or a chip failure of the memory <b>206</b>. Further, error correction may be able to handle more errors compared to other organization schemes, by changing a number of symbols across the width of the block <b>238</b>. Thus, the system <b>200</b> may inform the operating system and system administrator, providing different guidance to take different actions for different types of failures. Examples provided herein are usable in buffered memory systems where a pin failure is not a subset of a chip failure, and may tolerate pin failures in the first channel <b>210</b>. By discriminating between types of chip failures and/or pin failures; an administrator may be advised to avoid unnecessarily replacing a good DIMM in an attempt to address a pin failure (e.g., a stuck-at fault), increasing efficiency and preventing waste. Depending on memory configurations, examples herein may potentially tolerate a larger number of pin failures (DRAM pin failures) than other chipkill schemes.
Memory controller <b>202</b> may interact with a separate processor, and/or operate as a processor, to perform various functions. Such functionality may be based on logic <b>208</b>. Additional interaction may be such that the memory controller <b>202</b> detects and corrects memory errors based on logic <b>208</b>, and reports some information to (and collects statistics for) a processor, so that even at runtime, a report can be generated regarding memory status. A processor may be used separately from the memory controller <b>202</b>, and/or processor functionality (e.g., logic <b>208</b>) may be integrated with the memory controller <b>202</b>.
The memory controller <b>202</b> may report to the processor such that error information may be available to an operating system (OS) and provided for use by a system operator. For example, if there is a chip failure, even though the system <b>200</b> can tolerate that failure, another failure may exceed the error correction capabilities of a given error correction scheme. Thus, the faulty memory module should be replaced soon, before another error compounds the problem. The memory controller <b>202</b> may report location information to a processor regarding faulty memory, the processor may inform a software layer, and then a system administrator may be notified to physically replace the faulty memory.
In response to a pin failure <b>244</b> in the first channel <b>210</b>, the system <b>200</b> may disable the first channel <b>210</b>. Disabling the first channel <b>210</b> may involve a hardware operation such that the memory controller <b>202</b> reports the error situation to the OS, and that memory on the affected channel might become unavailable (e.g., if there is an additional failure on the channel, such as on a memory associated with that channel). The memory controller <b>202</b> may provide information to the runtime OS to copy all the data, from memory on the affected channel, to a different channel. The memory controller <b>202</b> then may physically (e.g., at a memory controller hardware layer) disable this channel.
In response to a pin failure in the second channel <b>220</b>, a similar OS runtime software operation may be initiated by the memory controller <b>202</b>. All data in the memory associated with the channel (i.e., the memory <b>206</b> connected to the second channel) may be relocated to another unaffected location, and the particular memory <b>206</b> affected by the second channel <b>220</b> may be disabled (thereby disabling the second channel <b>220</b>). Because a failure in the second channel <b>220</b> is likely localized in the particular DIMM associated with the second channel, the loss of memory functionality may involve less data than the first case above (because just the particular DIMMs are affected).
More specifically, there can be multiple DIMMs per channel, so in the first case, access to all the numerous DIMMs associated with the first channel <b>210</b> may be disabled. Thus, in an example system having four memory channels, one quarter of its total memory capacity may be lost by disabling the first channel <b>210</b>. However, in the second case, if there are <b>4</b> DIMMs per channel, then only one memory DIMM may be lost, so the penalty would be one sixteenth of the total memory capacity. Accordingly, discriminating which of the channels has an error is very beneficial to system operation and efficient memory management, to avoid unnecessarily disabling and/or replacing memory and/or channels.
<figref idref="DRAWINGS">FIG. 3A</figref> is a block diagram of a data block <b>338</b>A including a data symbol <b>333</b>A according to an example. The data block <b>338</b>A also includes a check symbol <b>335</b>A, codeword <b>330</b>A, and chip output <b>339</b>A. A symbol (data symbol <b>333</b>A, check symbol <b>335</b>A) is narrower (fewer bits wide) and taller (greater number of bits tall).
Thus, the arrangement of block <b>338</b>A takes advantage of a narrower symbol organization compared to earlier examples. Instead of constructing an 8-bit symbol using 8 bits out of a DRAM chip, a subset of the bits from a chip may be used, in multiple bursts, to provide the full set of bits for that symbol. In an example, a 2-burst of 4 bits may be used to construct an 8-bit symbol, taller than it is wide. Other combinations may be used (e.g., a 4 burst of 2 bits), though not specifically illustrated. Such narrow and tall arrangement is able to take advantage of burst access provided by DDRx DRAM systems. Such DRAM systems may provide n-bit prefetch and burst n access to meet the gap between the slow DRAM core speed and fast bus speed: n is 1 for single-data-rate SDRAM (synchronous DRAM), 2 in DDR, 4 in DDR2, 8 in DDR3, and so on. Constructing an 8-bit symbol from 2-burst of 4 bits does not affect DRAM scheduling nor DRAM access behavior.
<figref idref="DRAWINGS">FIG. 3A</figref> illustrates how data symbols <b>333</b>A and ECC (check symbols <b>335</b>A) may be organized using 2-burst 4-bit symbols. For each 2-burst access, there are 32 data symbols <b>333</b>A and 4 check symbols <b>335</b>A. An ECC codeword <b>330</b>A may be composed of 32 data symbols <b>333</b>A and 4 check symbols <b>335</b>A. Thus, the block <b>338</b>A includes two codewords <b>330</b>A. The 4 check symbol <b>335</b>A error codes can correct 2 symbol errors. A 64 byte data block <b>338</b>A is composed of two sets of codewords <b>330</b>A. When this block <b>338</b>A is transferred over a narrow channel, its shape may be folded as in earlier examples. However, corrupted symbol patterns would be different due to the taller/narrower type of symbols. Chip failures, pin failures at the DRAM side (second channel), and pin failures at the memory controller side (first channel), manifest differently based on different corrupted symbol patterns. A pin failure at the memory controller side (first channel) would affect up to 2 symbols per code word. However, with narrow symbols, and 4 check symbol <b>335</b>A error codes, it is possible to correct the 2 corrupt symbols in the codeword <b>330</b>A caused by the pin failure in the first channel. Thus, not only does the symbol arrangement in block <b>338</b>A enable correction of chip and pin failures, but also enables identification of the type of problem, depending on the corrupted symbol pattern that appears at the memory controller by analyzing the ECC encoding.
<figref idref="DRAWINGS">FIG. 3B</figref> is a block diagram of a data block <b>338</b>B including a data symbol <b>333</b>B according to an example. The data block <b>338</b>B also includes a check symbol <b>335</b>B. <figref idref="DRAWINGS">FIG. 3B</figref> illustrates a first corrupted symbol pattern <b>336</b>B<b>1</b>, a second corrupted symbol pattern <b>336</b>B<b>2</b>, and a third corrupted symbol pattern <b>336</b>B<b>3</b>.
The first corrupted symbol pattern <b>336</b>B<b>1</b> includes two adjacent symbols in a first codeword, along with another two adjacent symbols in the second codeword. Thus, a memory controller may recognize that corruption is causing the same adjacent symbols in successive codewords to become corrupted. Because a chip may provide output for multiple narrow symbols over multiple bursts, the memory controller can conclude that the first corrupted symbol pattern <b>336</b>B<b>1</b> corresponds to a chip failure.
The second corrupted symbol pattern <b>336</b>B<b>2</b> shows a single corrupted symbol per codeword, without an adjacent corrupted symbol, and without a non-adjacent corrupted symbol spaced away as a function of the ratio of the first and second memory channels. Thus, the memory controller may recognize such a symbol error (e.g., one symbol per codeword), spread across both codewords, as a memory pin failure (on the second channel between the DRAMs and the memory buffer).
The third corrupted symbol pattern <b>336</b>B<b>3</b> shows a two non-adjacent corrupted symbols per codeword. The non-adjacent corrupted symbols may be spaced from each other as a function of the ratio of the first and second memory channels, because the errors may arise due to one pin's affect distributed to 2 symbols of the codeword by the folding of the codeword. Thus, the memory controller may recognize such symbol errors (e.g., two non-adjacent symbols per codeword), spread across both codewords, as a buffer pin failure (on the first channel between the buffer and the memory controller).
The example patterns are demonstrated with a 4-check-symbol RS code, which can correct up to 2 symbol errors. Hence, all the described failures can be tolerated, unlike other ECC schemes that are unable to correct errors equivalent to two chip failures. Furthermore, because 2 symbol errors may be tolerated/corrected, this technique is robust enough to handle 2 pin failures in the second channel, because a pin failure in the second channel affects one symbol, and the ECC scheme here provides additional check symbols per codeword due to the narrower nature of each symbol. The schemes/patterns may be modified in view of using different organizations/chips, e.g., different burst/data widths, wherein a symbol is constructed from a portion of a chip's output that is multiplied over a plurality of bursts.
Thus, if an error is corrected, by analyzing the corrupted symbol pattern, the memory controller may identify which type of failure the error stems from. This can be used to improve pin failure tolerance capability. In decoding an ECC in an example, the memory controller may identify which symbol is faulty, and once the faulty symbol location is determined, then the error may be corrected. Thus, by performing error correction, the memory controller (e.g., correction logic) may determine which symbols are corrupted, and may thereby collect corruption location information to analyze how and/or which symbols are corrupted in the entire data block (e.g., 64 byte or 128 byte). The example schemes can provide much higher reliability/availability due to the increased capacity for handling errors.
Not only can examples handle and correct errors, but also notify to the software or system administrator in terms of very specific failure information, such as a pin failure at a specific DIMM chip, a pin failure at the memory controller, and so on, empowering the administrator to take different actions as appropriate and avoiding wasteful memory replacements when the problem is instead caused by a pin failure at the memory controller (because even if the DIMM is replaced, the faulty pin will still cause errors).
Examples herein enable enhancements to detection and correction, especially beneficial to buffered memory systems having different channel widths between the two data buses. However, examples described herein are applicable to systems having first and second channels of the same width, because examples enable differentiation of pin failure types and toleration of additional pin failures.
<figref idref="DRAWINGS">FIG. 4</figref> is a flow chart <b>400</b> based on discriminating a type of memory error according to an example. In block <b>410</b>, a codeword is transmitted across a first channel coupling a memory controller to a buffer, and a second channel coupling the buffer to a memory, wherein a first channel width differs from a second channel width, and the codeword includes a plurality of data symbols and at least one check symbol, wherein each of the plurality of symbols is formed from a plurality of bursts based on data access of the memory. For example, an 8-pin memory chip may provide data for two symbols at a time, and over a four burst access provide 4 8-bit symbols, each symbol fed by four of the pins. In block <b>420</b>, a memory error is corrected, using a memory controller and the at least one check symbol, based on a corrupted symbol pattern of the codeword. For example, the corrupted symbol pattern may include two non-adjacent corrupted symbols that are corrected. In block <b>430</b>, a chip failure, a pin failure of the first channel, and a pin failure of the second channel, are discriminated between as being the type of memory error, according to the corrupted symbol pattern. For example, the corrupted symbol pattern of two non-adjacent corrupted symbols may be detected as a pin failure of the first channel.
<figref idref="DRAWINGS">FIG. 5</figref> is a flow chart <b>500</b> based on discriminating a type of memory error according to an example. Flow starts in block <b>510</b>. In block <b>520</b>, error detection is applied. In block <b>530</b>, it is determined whether there is an error. If there is not an error, flow proceeds to end at block <b>599</b>. If there is an error, flow proceeds to block <b>540</b>. In block <b>540</b>, error correction is applied. In block <b>550</b>, it is determined whether there is a correctable error. If there is not a correctable error, flow proceeds to end at block <b>599</b>. If there is a correctable error, flow proceeds to block <b>560</b>. In block <b>560</b>, a corrupted symbol pattern is identified. In block <b>570</b>, it is determined whether there is a chip failure. If there is a chip failure, flow proceeds to block <b>575</b>. In block <b>575</b>, it is suggested to replace a memory, and flow proceeds to end at block <b>599</b>. If, in block <b>570</b>, there is not a chip failure, flow proceeds to block <b>580</b>. In block <b>580</b>, it is determined whether there is a first channel pin failure. If there is a first channel pin failure, flow proceeds to block <b>585</b>. In block <b>585</b>, it is suggested to disable the first channel, and flow proceeds to end at block <b>599</b>. If, in block <b>580</b>, there is not a first channel pin failure, flow proceeds to block <b>590</b>. In block <b>590</b>, it is determined whether there is a second channel pin failure. If there is a second channel pin failure, flow proceeds to block <b>595</b>. In block <b>595</b>, it is suggested to disable the second channel, and flow proceeds to end at block <b>599</b>. If, in block <b>590</b>, there is not a second channel pin failure, flow proceeds to end at block <b>599</b>.
Those of skill in the art would appreciate that the various illustrative components, modules, and blocks described in connection with the examples disclosed herein may be implemented as electronic hardware, computer software, or combinations of both. Thus, the example blocks of <figref idref="DRAWINGS">FIGS. 1-5</figref> may be implemented using software modules, hardware modules or components, or a combination of software and hardware modules or components. In another example, one or more of the blocks of <figref idref="DRAWINGS">FIGS. 1-5</figref> may comprise software code stored on a computer readable storage medium, which is executable by a processor. As used herein, the indefinite articles “a” and/or “an” can indicate one or more than one of the named object. Thus, for example, “a processor” can include one or more than one processor, such as in a multi-core processor, cluster, or parallel processing arrangement. The processor may be any combination of hardware and software that executes or interprets instructions, data transactions, codes, or signals. For example, the processor may be a microprocessor, an Application-Specific Integrated Circuit (“ASIC”), a distributed processor such as a cluster or network of processors or computing device, or a virtual machine. The processor may be coupled to memory resources, such as, for example, volatile and/or non-volatile memory for executing instructions stored in a tangible non-transitory medium. The non-transitory machine-readable storage medium can include volatile and/or non-volatile memory such as a random access memory (“RAM”), magnetic memory such as a hard disk, floppy disk, and/or tape memory, a solid state drive (“SSD”), flash memory, phase change memory, and so on. The computer-readable medium may have computer-readable instructions stored thereon that are executed by the processor to cause a system (e.g., a rate limit manager to direct hardware rate limiters) to implement the various examples according to the present disclosure.
It is appreciated that the previous description of the disclosed examples is provided to enable any person skilled in the art to make or use the present disclosure. Various modifications to these examples will be readily apparent to those skilled in the art, and the generic principles defined herein may be applied to other examples without departing from the spirit or scope of the disclosure. Thus, the present disclosure is not intended to be limited to the examples shown herein but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.
Contents3
7 sheets
Sheet 1 Sheet 2 Sheet 3 Sheet 4 Sheet 5 Sheet 6 Sheet 7
Every citation, both waysCites: the store holds 24 of 25
| Document | Relation | Office | Cited during |
|---|---|---|---|
| US11625296B2 | Cited by | United States of America | Search report |
| US9870167B2 | Cited by | United States of America | Search report |
| US10977118B2 | Cited by | United States of America | Applicant |
| US2017102896A1 | Cited by | United States of America | Pre-grant |
| US12242344B2 | Cited by | United States of America | Applicant |
| US10268541B2 | Cited by | United States of America | Search report |
| US11010242B2 | Cited by | United States of America | Applicant |
| US2007058410A1 | Cites | United States of America | Search report |
| US2009006900A1 | Cites | United States of America | Search report |
| US2009327596A1 | Cites | United States of America | Search report |
| US2010287445A1 | Cites | United States of America | Search report |
| US2011219197A1 | Cites | United States of America | Search report |
| US2013326293A1 | Cites | United States of America | Search report |
| US2014143633A1 | Cites | United States of America | Search report |
| US5553231A | Cites | United States of America | Search report |
| US6516436B1 | Cites | United States of America | Search report |
| US6697921B1 | Cites | United States of America | Search report |
| US7379316B2 | Cites | United States of America | Applicant |
| US7620875B1 | Cites | United States of America | Applicant |
| US7721140B2 | Cites | United States of America | Applicant |
| US7949931B2 | Cites | United States of America | Applicant |
| US8041990B2 | Cites | United States of America | Applicant |
| US8171377B2 | Cites | United States of America | Applicant |
| US8185800B2 | Cites | United States of America | Applicant |
| US20070058410A1 | Cites | United States of America | Search report |
| US20090006900A1 | Cites | United States of America | Search report |
| US20090327596A1 | Cites | United States of America | Search report |
| US20100287445A1 | Cites | United States of America | Search report |
| US20110219197A1 | Cites | United States of America | Search report |
| US20130326293A1 | Cites | United States of America | Search report |
| US20140143633A1 | Cites | United States of America | Search report |
| Intel E7500 Chipset MCH Inelx4 Single Device Data Correction (x4 SDDC) Implementation and Validation, Aug. 2002. | Non-patent | – | Applicant |
| Intel E7500 Chipset MCH Inelx4 Single Device Data Correction (x4 SDDC) Implementation and Validation, Aug. 2002. | Non-patent | – | Applicant |
2 members in 1 office
Priority claims2
| Document | Office | Kind | Date |
|---|---|---|---|
| 201213689814 | United States of America | A | |
| US201213689814 | – | – | – |
Members2
| Document | Office | Kind | |
|---|---|---|---|
| US2014157054A1 | United States of America | A1 | |
| US8966348B2This record | United States of America | B2 |
42 transactions on the USPTO file
Allowed after 1 non-final rejection.
- Non-final rejections
- 1
- Final rejections
- 0
- RCEs
- 0
- Appeals
- 0
Over time
Point at a mark for the transactionTransactions
| Event | Code | |
|---|---|---|
| Payment of Maintenance Fee, 8th Year, Large EntityM1552 | M1552 | |
| Payment of Maintenance Fee, 4th Year, Large EntityM1551 | M1551 | |
| Correspondence Address ChangeC.ADB | C.ADB | |
| Recordation of Patent Grant MailedPGM/ | PGM/ | |
| Patent Issue Date Used in PTA CalculationAllowedPTAC | PTAC | |
| Email NotificationEML_NTR | EML_NTR | |
| Issue Notification MailedAllowedWPIR | WPIR | |
| Dispatch to FDCD1935 | D1935 | |
| Application Is Considered Ready for IssuePILS | PILS | |
| Issue Fee Payment VerifiedN084 | N084 | |
| Issue Fee Payment ReceivedIFEE | IFEE | |
| Electronic ReviewELC_RVW | ELC_RVW | |
| Email NotificationEML_NTF | EML_NTF | |
| Mail Notice of AllowanceAllowedMN/=. | MN/=. | |
| Notice of Allowance Data Verification CompletedAllowedN/=. | N/=. | |
| Reasons for AllowanceEX.R | EX.R | |
| Date Forwarded to ExaminerFWDX | FWDX | |
| Response after Non-Final ActionA... | A... | |
| Email NotificationEML_NTR | EML_NTR | |
| PG-Pub Issue NotificationPG-ISSUE | PG-ISSUE | |
| Electronic ReviewELC_RVW | ELC_RVW | |
| Email NotificationEML_NTF | EML_NTF | |
| Mail Non-Final RejectionNon-final rejectionMCTNF | MCTNF | |
| Non-Final RejectionNon-final rejectionCTNF | CTNF | |
| Information Disclosure Statement consideredIDSC | IDSC | |
| Case Docketed to Examiner in GAUDOCK | DOCK | |
| Case Docketed to Examiner in GAUDOCK | DOCK | |
| Transfer Inquiry to GAUTI1050 | TI1050 | |
| Application Dispatched from OIPEOIPE | OIPE | |
| Application Is Now CompleteCOMP | COMP | |
| Email NotificationEML_NTR | EML_NTR | |
| Email NotificationEML_NTR | EML_NTR | |
| Change in Power of Attorney (May Include Associate POA)PA.. | PA.. | |
| Filing ReceiptFLRCPT.O | FLRCPT.O | |
| Sent to Classification ContractorPGPC | PGPC | |
| Cleared by OIPE CSRL194 | L194 | |
| Reference capture on IDSRCAP | RCAP | |
| Information Disclosure Statement (IDS) FiledM844 | M844 | |
| Applicants have given acceptable permission for participating foreignAPPERMS | APPERMS | |
| Information Disclosure Statement (IDS) FiledWIDS | WIDS | |
| IFW Scan & PACR Auto Security ReviewSCAN | SCAN | |
| Initial Exam Team nnIEXX | IEXX |
5 legal events, as the office reported them to INPADOC
Over the term
Point at a mark for the eventEvents
| Event | Code | |
|---|---|---|
| Maintenance fee paymentMAFP | MAFP | |
| Maintenance fee paymentMAFP | MAFP | |
| AssignmentAS | AS | |
| Information on status: patent grantGrantedPATENTED CASESTCF | STCF | |
| AssignmentAS | AS |
Numbers
- Publication
- 08966348
- Publication, DOCDB
- 8966348
- Publication, EPODOC
- US8966348
- Application
- 13689814
- Application, DOCDB
- 201213689814
- Application, EPODOC
- US201213689814
Titles
- English
- Memory error identification based on corrupted symbol patterns
Patent term adjustment
- A delay
- +73 daysthe office missed an examination deadline
- Net adjustment
- 73 days
Classification
- CPC, 4
- G06F11/0751
- G06F11/006
- G06F11/10
- G06F11/073
- IPC, 1
- G06F11 00
- USPC, 3
- 714776000
- 714718000
- 714E11159