Detection of transient errors by means of new selective implementation
Abstract
In one embodiment, the present invention includes a method for determining a degree or level of vulnerability for an instruction implemented by a processor, and re-implementing the instruction if the level of vulnerability is greater than a predetermined threshold. The level of vulnerability may correspond to the likelihood of a logic error for an instruction while it resides within the processor. Other embodiments are described and claimed.Vulnerability level, instruction, soft error, threshold, re-execution

Term
Term ended
Projected expiry passed 31 March 2026, 0.5 years ago.
- Priority
- Filed
- Published
- Projected expiry
- Today
29 claims: 4 independent, 25 dependent
- 1프로세서에서 실행된 인스트럭션(instruction)에 대한 취약성 레벨(vulnerability level)을 결정하는 단계 - 상기 취약성 레벨은 상기 인스트럭션에 대한 소프트 에러 가능성에 대응함 - ;및 상기 취약성 레벨이 임계값보다 큰 경우 상기 인스트럭션을 재실행하는 단계 를 포함하는 방법.
- 2제1항에 있어서, 상기 인스트럭션이 실행되었던 프로세서의 제1 실행 유닛이 아닌, 상기 프로세서의 다른 실행 유닛에서 상기 인스트럭션을 재실행하는 단계를 더 포함하는 방법.
- 3제1항에 있어서, 상기 취약성 레벨이 상기 임계값보다 작은 경우 상기 인스트럭션을 재실행하지 않는 단계를 더 포함하는 방법.
- 4제1항에 있어서, 상기 인스트럭션에 의해 점유된 상기 프로세서의 영역 및 상기 인스트럭션의 수명과 관련된 시간 기간 중 적어도 하나에 근거하여 상기 취약성 레벨을 결정하는 단계를 더 포함하는 방법.
- 5제4항에 있어서, 상기 인스트럭션과 관련된 시간 스탬프를 이용하여 상기 시간 기간을 결정하는 단계를 더 포함하는 방법.
- 6제1항에 있어서, 상기 재실행된 인스트럭션의 결과가 상기 인스트럭션의 결과와 매칭되는지를 결정하는 단계;및 매칭되지 않는 경우 상기 프로세서를 플러싱(flushing)하는 단계를 더 포함하는 방법.
- 7적어도 하나의 레지스터 파일 및 상기 적어도 하나의 레지스터 파일에 연결된 적어도 하나의 실행 유닛;및 상기 실행 유닛에 의해 실행된 인스트럭션들의 소프트 에러에 대한 취약성을 결정하는 인스트럭션 검증 유닛 을 포함하는 장치.
- 8제7항에 있어서, 상기 인스트럭션 검증 유닛은, 인스트럭션들 및 관련된 소스 데이터 태그들을 저장하는 버퍼;및 상기 버퍼에 연결되어 취약한 인스트럭션들을 재실행하는 적어도 하나의 실행 유닛을 포함하는 장치.
- 9제8항에 있어서, 상기 인스트럭션 검증 유닛은 상기 인스트럭션에 대한 영역 값 및 상기 인스트럭션에 대한 시간 값 중 적어도 하나에 근거하여 인스트럭션에 대한 취약성 레벨을 결정하는 로직(logic)을 더 포함하는 장치.
- 10제9항에 있어서, 상기 로직은 상기 인스트럭션과 관련된 시간 스탬프에 근거하여 상기 시간 값을 결정하는 장치.
- 11제9항에 있어서, 상기 로직은 상기 취약성 레벨이 임계값보다 큰 경우 재실행을 위해 상기 인스트럭션을 상기 적어도 하나의 실행 유닛에 제공하는 장치.
- 12제11항에 있어서, 상기 임계값은 선택된 성능 매트릭(metric)에 근거하여 조절가능하고, 상기 임계값은 보다 높은 레벨의 성능에 대해 보다 높게 설정되는 장치.
- 13제11항에 있어서, 상기 인스트럭션의 재실행이 상기 인스트럭션의 원래의 실행과 매칭되는 경우 상기 인스트럭션을 회수(retire)하는 인스트럭션 회수 유닛을 더 포함하는 장치.
- 14인스트럭션들을 포함하는 머신-판독가능한 저장 매체를 포함하는 물품에 있어서, 상기 인스트럭션들은 머신에 의해 실행되는 경우 상기 머신으로 하여금, 원래의 결과를 얻기 위해 프로세서에서 인스트럭션을 실행하는 단계;상기 인스트럭션이 소프트 에러에 대해 취약한 경우 재실행된 결과를 얻기 위해 상기 프로세서에서의 상기 인스트럭션을 재실행하는 단계;및 상기 원래의 결과를 상기 재실행된 결과와 비교하는 단계 를 포함하는 방법을 수행하도록 하는, 머신-판독가능한 저장 매체를 포함하는 물품.
- 15제14항에 있어서, 상기 방법은 인스트럭션 회수 단계에서 상기 인스트럭션을 회수하고, 상기 인스트럭션이 상기 소프트 에러에 대해 취약하지 않은 경우 상기 인스트럭션을 재실행하지 않는 단계를 더 포함하는, 머신-판독가능한 저장 매체를 포함하는 물품.
- 16제14항에 있어서, 상기 방법은 상기 비교가 미스매칭을 나타내는 경우 상기 프로세서를 플러싱하는 단계를 더 포함하는, 머신-판독가능한 저장 매체를 포함하는 물품.
- 17제14항에 있어서, 상기 방법은 상기 인스트럭션에 의해 소모된 영역 및 상기 프로세서에서의 상기 인스트럭션의 펜딩 시간(pending time) 중 적어도 하나에 근거하여, 상기 인스트럭션이 상기 소프트 에러에 대해 취약한지 여부를 결정하는 단계를 더 포함하는, 머신-판독가능한 저장 매체를 포함하는 물품.
- 18제17항에 있어서, 상기 방법은, 상기 영역 및 상기 펜딩 시간을 이용하여 취약성 값을 계산하는 단계;및 상기 취약성 값과 임계값을 비교하고, 상기 비교에 근거하여 상기 인스트럭션을 재실행할지 여부를 결정하는 단계를 더 포함하는, 머신-판독가능한 저장 매체를 포함하는 물품.
- 19인스트럭션을 실행하는 적어도 하나의 실행 유닛 및 상기 인스트럭션이 소프트 에러에 대해 취약한 경우 상기 인스트럭션을 재실행하는 중복 실행 유닛을 포함하는 프로세서;및 상기 프로세서에 연결된 DRAM(dynamic random access memory) 을 포함하는 시스템.
- 20제19항에 있어서, 상기 프로세서는 상기 소프트 에러에 대한 상기 인스트럭션의 취약성을 정량화하는 인스트럭션 검증기를 더 포함하는 시스템.
- 21제20항에 있어서, 상기 인스트럭션 검증기는 상기 프로세서에서의 상기 인스트럭션의 크기 및 상기 인스트럭션의 수명 중 적어도 하나에 근거하여 상기 인스트럭션을 재실행할지의 여부를 결정하는 시스템.
- 22제21항에 있어서, 상기 인스트럭션 검증기는 상기 인스트럭션의 상기 크기 및 수명에 근거하여 상기 인스트럭션의 취약성 값을 임계값과 비교하고, 상기 취약성 값이 상기 임계값보다 큰 경우 상기 재실행을 트리거링하는 시스템.
- 23제22항에 있어서, 상기 임계값은 적응적 임계값이며, 상기 시스템은 선택된 성능 레벨에 근거하여 상기 임계값을 조절하는 시스템.
- 24제22항에 있어서, 상기 시스템은 상기 프로세서가 인스트럭션 취약성 감소 기술을 구현하는 경우에 상기 임계값을 감소시키는 시스템.
- 25제20항에 있어서, 상기 인스트럭션 검증기는 상기 인스트럭션의 회수 단계에서 상기 인스트럭션을 재실행할지의 여부를 결정하는 시스템.
- 26제20항에 있어서, 상기 인스트럭션 검증기는 상기 프로세서에서 실행된 인스트럭션들에 대응하는 엔트리들을 저장하는 버퍼를 포함하는 시스템.
- 27제26항에 있어서, 상기 엔트리들은 대응하는 상기 인스트럭션을 상기 프로세서내로 삽입하는 시간에 대응하는 시간 스탬프 필드를 포함하는 시스템.
- 28제27항에 있어서, 상기 인스트럭션 검증기는 상기 시간 스탬프 필드를 이용하여 상기 인스트럭션을 재실행할지의 여부를 결정하는 시스템.
- 29제19항에 있어서, 상기 시스템은 상기 재실행된 인스트럭션의 결과가 상기 인스트럭션의 결과와 매칭되지 않는 경우 상기 프로세서의 적어도 일부분을 플러싱하는 시스템.
Independent claims29
53 paragraphs in 1 section, as filed
DETECTION OF TRANSIENT ERRORS BY MEANS OF NEW SELECTIVE IMPLEMENTATION
FIELD OF THE INVENTION Embodiments of the present invention relate to error detection in a semiconductor device, and more particularly, to error detection in a processor.
Transient errors, often referred to as soft errors, are an increasing source of error in the processor. Due to the reduced size of the devices and the reduced voltage at which they operate, these devices are more susceptible to cosmic particle strikes and parameter changes. Such events, which occur randomly, may cause transient errors that may affect the proper execution of the processor. With each generation of semiconductor manufacturing technology, susceptibility to soft errors is expected to increase.
Certain mechanisms have been used to attempt soft error correction. Typically, these approaches include providing redundant paths for redundant operations on data. However, such redundant paths greatly increase the size and power consumption of the processor, resulting in performance degradation. Moreover, some approaches use simultaneous multithreading (SMT) to detect errors. In such schemes, processing is scheduled on two separate execution paths (eg, two threads in the SMT core). The resulting data is then compared for identity verification. If the results are different, this is an indication of a soft error, and such an error is detected. However, the performance degradation is significant because some hardware focuses on error detection instead of executing other processes, and there is complexity in supporting result comparison and thread coordination.
1 is a flowchart of a method according to an embodiment of the present invention.
2 is a block diagram of a general processor architecture according to an embodiment of the present invention.
3 is a block diagram of a processor according to an embodiment of the present invention.
4 is a block diagram of a multiprocessor system according to an embodiment of the present invention.
5 is a block diagram of a portion of a processor in accordance with an embodiment of the present invention for selective re-execution of instructions.
In various embodiments, a soft error in the processor may be detected and appropriate action may be taken to correct the error. Such soft error detection can be performed with minimal complexity or additional power consumption. Moreover, embodiments may perform error detection using existing processor architectures. Alternatively, a minimal amount of additional hardware may be implemented to perform soft erotica detection.
To perform soft error detection, instructions in the processor pipeline may be selectively duplicated or re-executed based on different parameters. For example, only instructions that are particularly likely to be dependent on soft errors (eg, based on size and/or length of time in the processor) may be selectively duplicated. In this way, a large amount of soft errors can be detected with minimal impact on performance.
Soft error detection according to an embodiment of the present invention may be implemented in different ways. In some embodiments, existing processor architectures may be used to perform soft error detection through one or more algorithms for error detection. In another embodiment, additional controllers, logics and/or functional units may be provided in the processor to handle soft error detection. At a high level, soft error detection can be implemented by identifying instructions in the processor pipeline that are particularly vulnerable to soft errors and re-executing those instructions. If the results of the original and duplicated instructions match, no soft error is indicated. Otherwise, if the results are different, a soft error is indicated, and a recovery mechanism can be applied to correct the error.
Instruction vulnerabilities depend, in large part, on the area the instruction uses in the processor and the amount of time the instruction consumes within the processor. For example, many instructions consume a large number of cycles in the processor before committing, while other instructions move through the pipeline without delay for a single cycle. Moreover, not all instructions use the same hardware resources. For soft error detection coverage, the weakest instructions can be duplicated to provide maximum possible error coverage with minimal impact on performance.
In one embodiment, vulnerable instructions may be replicated upon instruction execution (ie, upon instruction retirement) before the instruction leaves the pipeline. In such an embodiment, a set of arithmetic logic units (ALUs) are included in the processor to validate the outputs of the vulnerable instruction as the instructions reach the top of the reorder buffer (ROB), for example by re-execution of the instructions. )can do.
Because different instructions occupy different amounts of storage and consume different amounts of time in the processor during their lifetime, each instruction's vulnerability to spout errors is different. By identifying such more vulnerable instructions, high coverage soft error detection can be achieved with minimal level of instruction duplication. In this way, a maximum amount of error detection coverage can be achieved with minimal impact on performance, using minimal hardware resources and power consumption. In various embodiments, an instruction's soft error vulnerability is determined by the instruction type (eg, load, store, branch, arithmetic), the time consumed by the instruction within the processor (or a particular processor component), and other factors of the instruction. properties (eg, the source data is being prepared, and has a reduced number of bits because the immediate field is narrow).
Not all instructions occupy the same space within a processor component. For example, load and store instructions may store a memory order buffer index in their entity, but other instructions do not use this field. Likewise, the store and branch instructions do not use the issue queue field to store the destination register index, since they do not produce any result and, therefore, have no destination register assigned to them. These unused bits are not vulnerable to particle attack, thus reducing instruction vulnerability. Also, the vulnerability status of some bits in the issue queue entry may be dependent on input changes or dynamics of a superscalar pipeline. If the source of the instruction is ready when the instruction is dispatched to the issue queue, the source tag fields in the issue queue are vulnerable. Similarly, using a narrow operand identification technique for an immediate field leaves a large portion of an instruction's immediate field not vulnerable, reducing the overall vulnerability of the instruction.
Referring now to Figure 1, there is shown a flow diagram of another method in one embodiment of the present invention. As shown in FIG. 1 , method 10 may begin by associating a time stamp with an instruction (block 15). For example, the front end of the processor may associate a time stamp with the input instruction as the input instruction is placed in a buffer such as a ROB. In various embodiments, instructions may correspond to microoperations (μops), although error detection may be implemented at different instruction granularity levels in other embodiments. Although described herein with respect to a particular processor architecture, it will be understood that the scope of the invention is not so limited, and that soft error detection may be implemented in other locations. Moreover, as the instruction enters the front end of the processor, although described as associating a time stamp with the instruction, the time stamp may instead be associated with the instruction at other points in the processor pipeline.
Still referring to FIG. 1 , an instruction is then injected into the processor pipeline (block 20). Thus, when the instruction is scheduled for execution, the instruction may be performed in the processor pipeline (block 25). After execution, the vulnerability of the instruction may be calculated upon retrieval of the instruction (block 30). For example, the state of an instruction in a ROB is set as ready to execute, when it has finished executing, and waits until it reaches the top of the ROB to be retrieved (i.e., the oldest in time in the instruction). . Then, the vulnerability of the instruction to soft errors can be calculated. Such a calculation may take many different forms in different embodiments. Calculations can take into account both the length of time the instruction resides in the processor, as well as the area of the processor consumed by the instruction. Accordingly, information based on the time stamp as well as the instruction width (including various fields such as instruction, source information, etc.) may be considered.
Next, it may be determined whether the determined instruction vulnerability level is greater than a threshold value (diamond 35). This threshold may be user set and, in some embodiments, may be an adaptive threshold based on a desired level of performance for the processor. If the instruction vulnerability level is below the threshold, control passes to block 40, where the instruction may be retrieved. Thus, the method 10 ends.
Instead, if, at diamond 35 , it is determined that the instruction vulnerability level is greater than the threshold, control passes to block 45 . Here, the instruction may be re-executed (block 45). In some embodiments, the instructions may be duplicated and re-executed in the same processor pipeline, but in various embodiments, one or more additional functional units may be provided for performing re-execution.
After re-execution, the original result may be compared with the re-executed result to determine whether the results match or not (diamond 50). If there is a match, this is an indication that a soft error does not exist, so control passes to block 40 (described above), where the instruction is retrieved. If, instead, at diamond 50 , it is determined that the results are different, control passes to block 55 . Accordingly, at block 55, a soft error is indicated. Such an indication may take many different forms, including a signal to a processor or other such location, some control logic. Based on the indicated error, an appropriate recovery mechanism may be applied (block 60). The recovery mechanism may vary in different embodiments and may include re-executing an instruction, flushing some or all of the various processor resources, or other such recovery mechanism. Although described with this particular implementation in the embodiment of Figure 1, it will be understood that the scope of the invention is not so limited.
As noted above, embodiments may be implemented in many different processor architectures. For purposes of illustration, Figure 2 shows a block diagram of a general processor architecture in accordance with one embodiment of the present invention. 2 , the processor 100 may be an out-of-order processor. However, the scope of the present invention is not so limited, and other embodiments may be implemented in an in-order machine. Processor 100 includes a front end 110 that can receive instruction information and decode such information into one or more microoperations for execution. The front end 110 is connected to an instruction scheduling unit 120 capable of scheduling instructions for execution for a selected one of a plurality of execution units 130 . Such units may vary, but in certain embodiments, integer, floating-point, single instruction multiple data (SIMD), address generation units, and other such execution units may be provided. Moreover, in some embodiments, one or more additional redundant execution units may be provided to perform soft error detection according to embodiments of the present invention.
Still referring to FIG. 2 , when the instructions have been executed in the selected execution unit, the instructions may be provided to the instruction retrieval/verification unit 140 . The unit 140 may be used to retrieve instructions performed in different orders in the processor pipeline, back to in-order retrieval according to the program order. The unit 140 may be further adapted to undergo soft error detection according to an embodiment of the present invention. In particular, upon retrieval, each instruction deemed susceptible to soft errors at a level greater than a given threshold may be re-executed, for example, in additional execution units within execution units 130 . Based on the result of such re-execution, the result of the original instruction is verified, the instruction is recalled, or a soft error is indicated, and an appropriate recovery mechanism is implemented. Although described with the high level architecture in FIG. 2 , it will be understood that the scope of the present invention is not so limited and certain variations are contemplated.
Referring now to FIG. 3, there is shown a more detailed block diagram of a processor in accordance with one embodiment of the present invention. As shown in FIG. 3 , the processor 200 may include various resources for performing instructions. The processor 200 shown in FIG. 3 may correspond to a single-core processor, or, in another embodiment, may alternatively be one of a multi-core or a multi-core processor.
As shown in FIG. 3 , the processor 200 may include a front end 210 including various resources. 3 , which is coupled to a renamer unit 220 that takes instructions and renames logical registers in the instructions into a larger number of physical registers in the processor's register files. A ROB 215 may be provided. From the renamer 220, instructions may be coupled to a trace cache 225, which is coupled to a branch predictor 230 to aid in predicting branches of execution. A microinstruction translation engine (MITE) 238 is coupled to provide translated instructions to the micro-sequencer 235 , which is a unified cache memory 240 (eg, a level 1 or level 2 cache). is connected to The data cache 268 may also be coupled to the unified cache memory 240 , and may be coupled to a load queue 267a and a storage queue 267b .
Still referring to FIG. 3 , renamed instructions may be provided to an execution unit 250 , which receives input instructions, places it in a queue for storage, and places it in a queue for storage, among a number of functional units. It includes an issue queue 252 for scheduling as one. The issue queue 252 is further coupled to a pair of register files, a floating-point (FP) register file 254 and an integer register file 256 . When the source data required for an instruction is provided in a selected register file, the instruction may be executed as scheduled for one of a number of functional units, two of which are shown in Figure 3 for ease of illustration. do. In particular, a first execution unit 260 and a second execution unit 265 may be provided, which may respectively correspond to a floating point logic unit and an integer logic unit. The results of execution of the instructions may be transmitted to various locations via the interconnection network 270 . For example, in a given architecture, the resulting data may be provided back to register files 254 , 256 , load queue 267a and store queue 267b and/or back to ROB 215 .
Moreover, the result data may be provided to the instruction verification unit 280 . Unit 280 may perform soft error detection according to an embodiment of the present invention. As shown, the instruction verification unit 280 may be coupled to the interconnection network 270 , and may further be coupled to the ROB 215 . In various embodiments, the instruction verification unit 280 may include various resources, including a buffer arranged similar to the buffer of the ROB 215 . In one embodiment, the instruction verification unit 280 may include a recheck source buffer (RSB), which may be an extension of the ROB 215 . That is, the RSB may be a first-in-first-out (FIFO) buffer containing the same number of entries as the ROB 215 . In another embodiment, the buffer for instruction verification purposes may simply use ROB 215 . In one embodiment, each entry of ROB 215 (and buffer in instruction verification unit 280, if present) may take the form shown in Table 1 below.
<tables id="1"><table><tgroup xmlns="http://www.oasis-open.org/tables/exchange/1.0" cols="4"><colspec colnum="1" align="justify" colname="col1" colwidth="2670" /><colspec colnum="2" align="justify" colname="col2" colwidth="2670" /><colspec colnum="3" align="justify" colname="col3" colwidth="2670" /><colspec colnum="4" align="justify" colname="col4" colwidth="2670" /><tbody><row><entry align="justify" colname="col1"> instruction</entry><entry align="justify" colname="col2"> source tags</entry><entry align="justify" colname="col3"> time stamp</entry><entry align="justify" colname="col4"> result</entry></row></tbody></tgroup></table></tables>
As shown in Table 1, each entry of ROB 215 may include various fields. In particular, as shown in Table 1, each entry may include an instruction field that may correspond to a microoperation. Moreover, each entry may include source tags to identify the location of the required data. Additionally, each entry may include a timestamp indicating the time the entry was created in the ROB 215 . Finally, as shown in Table 1, each entry may include a result field for storing the result of the instruction. Although described with this particular implementation in the embodiment of Table 1, the scope of the invention is not so limited, and entries in the ROB or RSB may include additional or different fields.
The instruction verification unit 280 may further include its own dedicated functional units 285 for re-executing the selected instruction (or alternatively, it may use already available functional units). RSB maintains the result value generated by all instructions, instruction opcodes, and source tags, so that if an instruction is identified as vulnerable, the result of the instruction (or if it is a memory or branch operation) at commit time. , address) can be verified.
The instruction verification unit 280 may further include a microcontroller, logic, or other resources to perform soft error detection. In particular, the resource may, for example, receive the instruction at execution time, and determine a vulnerability measure for the instruction. This vulnerability measure may take many different forms, and in one embodiment, the instruction verification unit 280 may perform a calculation to determine the instruction vulnerability level according to the following equation.
<maths num="1"><df><img file="KR20080098683A_D0001.tif" /></df></maths>
The instruction vulnerability level is measured by the time stamp from initial instruction insertion into ROB 215 to instruction execution in the occupied bit region of the instruction, including its various fields such as source identifiers, destination identifiers, etc. As such, the instruction may be based on multiplying the time consumed in the processor by the time consumed in the processor, which may correspond to the instruction. In other embodiments, the vulnerability measure may be based on only one of occupied bit area and time spent, or different combinations of these values.
This instruction vulnerability level can be compared to a threshold. As an example of operation, assume a threshold value of 1000. On average, if an instruction covers 50 bits of weak bit space, any instruction that consumes more than 20 cycles in the issue queue is re-executed with selective replication. Such delays may be due to long dependency chains or long latency operations, such as floating point partitioning. In various embodiments, the threshold may vary widely. For example, a threshold of 0 will result in re-execution of all instructions, and a large threshold (eg, 5000) will result in only a small percentage (eg, less than 10%) of instructions being re-entered. In one embodiment, a threshold of 1000 may represent a good tradeoff.
If the vulnerability level is less than the threshold, the instruction is normally performed, and the instruction verification unit 280 takes no further action. Instead, an instruction may be more susceptible to soft errors if it is determined that the instruction vulnerability level is greater than a selected threshold. If the instruction vulnerability value is high, it means that the instruction occupies a large number of bits and/or spends a long time in the processor component, making it more vulnerable. In various embodiments, only instructions higher than the selected vulnerability threshold are replicated. In this way, the maximum amount of error coverage is affected by duplicating the minimum number of instructions. By performing an instruction vulnerability analysis in instruction execution, instructions that are not architecturally correct execution (ACE) that are removed from the pipeline before reaching the execution stage are filtered out. In addition, since the amount of time consumed by the instruction inside the processor is accurately known along with the amount of space occupied by the instruction, the instruction vulnerability information can be more accurately collected at the execution time.
In various embodiments, the instruction verification unit 280 takes information of the instruction, obtains the source data of the instruction, and provides it to one or more additional functional units 285 associated with the instruction verification unit 280, thereby providing the instruction. can be redone. In various embodiments, the vulnerable instruction is re-executed by using information stored within the RSB. In one implementation, source register tags are stored in the RSB and used to access register files 254 and 256 to gather source data for verification. These register files may contain two additional read ports (one port for each source operand) for verification purposes. Thus, accessing register files 254 and 256 for selective re-execution can cover errors that occur while the instruction is in issue queue 252 and source tags are susceptible to particle collision. Alternatively, the source values may be stored in the RSB when they are first read, thereby avoiding having additional ports on the register files.
The result of the re-executed instruction is passed back to the instruction verification unit 280, where the re-executed result is compared with the original result. If the two results match, no further action is taken by the instruction verification unit 280 and the instruction is retrieved. However, if the results do not match, a soft error is indicated, and the instruction verification unit 280 signals the soft error to one or more locations within the processor 200 . In this case, the processor 200 may perform an error recovery mechanism. For example, if the two results do not match, the instruction verification unit 280 may initiate a flush of the processor, resuming execution starting from the faulty instruction. Although described as a specific implementation in the embodiment of Figure 3, it will be understood that the scope of the invention is not so limited.
In certain embodiments, some optimizations are possible. Instead of only verifying instructions at the head or other retrieval location of the ROB, more than one instruction verification may be performed per cycle. Moreover, a number of thresholds and performance metrics may be implemented to identify a period of time when a processor loses performance due to verification. During such periods, only very vulnerable instructions can be copied. For example, during times of low performance, the threshold may be set higher to reduce the number of re-executed instructions. Moreover, an adaptive threshold level that changes according to processor state (eg, error rate, performance, power, etc.) may be used. Thus, depending on a given processor state, one of a number of different vulnerability thresholds may be selected for compensating with the calculated instruction vulnerability values.
In some embodiments, error detection in accordance with embodiments of the present invention may be used in conjunction with vulnerability reduction techniques such as flush and restart or narrow value identification. For example, flushing and restarting the pipeline following an off-chip cache miss will reduce the soft error vulnerability of many instructions by reducing the time they spend in the issue queue, Accordingly, it is possible to reduce the number of re-executed instructions through soft error detection. Moreover, lower threshold levels may be set to increase error coverage where flush and restart mechanisms or other vulnerability reduction techniques are appropriate.
Embodiments may be implemented in several different system types. Referring now to FIG. 4, there is shown a block diagram of a multiprocessor system in accordance with an embodiment of the present invention. As shown in Figure 4, the multiprocessor system is a point-to-point interconnect system, with a first processor 470, although other types of interconnections may be used in other embodiments. may be, but includes a second processor 480 coupled via a point-to-point interconnect 450 . As shown in Figure 4, each of the processors 470 and 480 includes first and second processor cores (ie, processor cores 474a, 474b and processor cores 484a, 484b). may be multicore processors. Although not shown for ease of illustration, the first processor 470 and the second processor 480 (and, more specifically, the cores therein), include weak instruction identification and verification logic to implement the present invention. Soft errors may be detected according to an example. The first processor 470 further includes a memory controller hub (MCH) 472 and point-to-point (PP) interfaces 476 , 478 . Likewise, the second processor 480 includes an MCH 482 and PP interfaces 486 , 488 . As shown in Figure 4, MCHs 472 and 482 couple the processors to respective memories, i.e., memory 432 and memory 434, which are portions of main memory attached locally to each processor. do.
First processor 470 and second processor 480 may be coupled to chipset 490 via PP interconnects 452 and 454 respectively. As shown in FIG. 4 , chipset 490 includes PP interfaces 494 and 498 . Moreover, the chipset 490 includes an interface 492 that couples the chipset 490 with the high performance graphics engine 438 . In one embodiment, an Advanced Graphics Port (AGP) bus 439 may be used to connect the graphics engine 438 to the chipset 490 . The AGP bus 439 may conform to the Accelerated Graphics Port Interface Specification, Revision 2.0, published May 4, 1998, by Intel Corporation of Santa Clara, CA. Alternatively, a point-to-point interconnect 439 may connect these components.
The chipset 490 may be coupled to the first bus 416 via an interface 496 . In one embodiment, first bus 416 is a Peripheral Component Interconnect (PCI) bus, such as PCI Local Bus Specification, Production Version, Revision 2.1 (June 1995), or a PCI Express bus or a third generation input/output ( I/O) bus, such as an interconnect bus, but the scope of the present invention is not so limited.
As shown in FIG. 4 , various I/O devices 414 are connected to the first bus 416 , with a bus bridge 418 connecting the first bus 416 with the second bus 420 . can be connected In one embodiment, the second bus 420 may be a low pin count (LPC) bus. The various devices include, for example, a keyboard/mouse 422 , a communication device 426 , and a second bus 420 including a data storage unit 428 , which in one embodiment may contain code 430 . can be connected to Moreover, the audio I/O 424 may be coupled to the second bus 420 . It should be noted that other architectures are possible. For example, instead of the point-to-point architecture of FIG. 4, the system may implement a multi-drop bus or other such architecture.
Moreover, it will be appreciated that different processor structures and different methods of performing selective instruction duplication may be realized in various embodiments. For example, in some embodiments, instructions may be selectively re-executed during a window of time between when they are reclaimable and when they are actually reclaimed. Different methods of performing such selective reissuance of instructions may be realized. Referring now to FIG. 5, shown is a block diagram of a portion of a processor in accordance with an embodiment of the present invention for selective instruction re-execution.
As shown in FIG. 5 , the processor 500 includes an issue queue 510 coupled to receive input instructions via an internal bus 505 . As further shown in FIG. 5 , an optional queue 520 may be further coupled to receive input instructions. Also, instructions may be provided to the ROB 515 .
5 , issue queue 510 and optional queue 520 are coupled to a selector 525 that can select an instruction from one of these queues for delivery to a register file 530 , a register File 530 may be coupled to one or more execution units 540 . 5 , execution unit(s) 540 may be further coupled to ROB 515 . If a port is available for execution, the instruction in optional queue 520 (its counterpart in issue queue 510 has already been issued) is issued and executed. When the instruction has finished executing its execution, the results are stored in ROB 515 . When the duplicate instruction execution is finished, the result is compared with the stored original result for verification purposes. If the head of ROB 515 is ready to be executed but not verified, execution is stopped.
In some embodiments, ROB 515 may add certain fields (eg, compare bit, verify bit, error detection bit, and bits for storing the result) to each entry. However, in other embodiments, other arrays may be used for such storage. When a ROB entry is assigned to a new instruction, these auxiliary fields in the entry are reset. When instructions have finished their execution, their results are written to ROB 515 . When the original instruction is finished, the result is written and the compare bit is set. The second instruction (i.e., duplicate), if executed, then finds the set compare bit, causes the stored value to be compared with the re-executed result, and sets the verified bit in the ROB entry. The comparison result is stored in the error detection bit to indicate whether the results match or not. One alternative to the compare bit may be a bit associated with the instruction to identify whether it is an original instruction or a duplicate instruction.
As noted above, when an instruction is placed in the issue queue 510, it is also stored in the optional queue 520. Each entry in optional queue 520 may include an opcode, source tags (for reading sources from register file 530), a ROB entry identifier, and a ready bit indicating whether it is ready for reissuance. have. It should be noted that an entry in issue queue 510 may store its corresponding entry in ROB 515 as well as its corresponding entry in optional queue 520 . When the original instruction is issued from the issue queue 510, it signals the corresponding entry in the optional queue 520, setting the ready bit. The total number of entries in the optional queue 520 may be equal to the number of entries in the ROB 515 . It should be noted that the latency of the optional queue 520 is not critical to performance, so it can be implemented with slower and more power efficient designs and even with low power transistors.
Selector 525 may be a multiplexer that selects between instructions from issue queue 510 and optional queue 520 . In various embodiments, selector 525 may prioritize instructions from issue queue 510 . The optional queue 520 may select an instruction to pass to the selector 525 from those with the ready bit set. This can be done with a chain of gates (since it is not in the critical path and does not affect the cycle time), or by multibanking, i.e. only the oldest instruction in the bank completes the port. This can be done by forming a selective queue 520 with as many banks as ports. If a free port exists, an instruction from optional queue 520 is issued, otherwise it waits.
As mentioned above, the soft error vulnerability of an instruction may depend on the area occupied by the instruction in the processor and the time consumed. If an entry in selective queue 520 has a ready bit set, its vulnerability value may be compared to a threshold. It should be noted that since instructions are assigned to issue queue 510 and optional queue 520 at the same time, it is known how many cycles an instruction consumes in issue queue 510 . The elapsed time between placement in selective queue 520 and receipt of a signal to set the ready bit may be used as a measure of time spent. If the vulnerability value is less than the threshold, the entry in the optional queue 520 may be freed, the verified bit in the corresponding ROB entry is set, and the error detected bit in the ROB entry is reset. In some embodiments, time (eg, via time stamps) may be used as the vulnerability value instead of the product of area and time.
For verification purposes, each entry in ROB 515 may include a verified bit, as described above. If the verified bit is not set, execution may stop waiting for verification. If the verified bit is set, the instruction is ready to execute only if its error detected bit is reset. If the error detected bit is set, different actions may be taken. For example, the pipeline may be flushed to re-execute the faulty instruction, or an exception may be raised. Upon branch misprediction, entries are no longer valid in issue queue 510 and optional queue 520 may be removed. The same mechanism used in issue queue 510 may be used to remove entries in optional queue 520 .
It should be noted that redundant hardware itself can be susceptible to particle collisions. However, its vulnerability is zero. That is, if there is a conflict on the optional queue 520, the wrong instruction will execute what is likely to result in a false positive. Since the head of ROB 515 will wait to be verified, a collision may hit the ROB entry identifier or a verified bit in ROB 515, resulting in a deadlock. This can be addressed by parity protection of the ROB entry identifier in the optional queue 520 or by adding a watchdog timer, and if execution is delayed by more than a given number of cycles, the instruction is squashing and can be restarted.
Embodiments may be implemented in code and stored on a storage medium having stored thereon instructions that may be used to program a system to perform the instructions. Storage media include, but are not limited to, any floppy disk, optical disk, compact disk read-only memory (CD-ROM), compact disk rewritable (CD-RW) and magneto-optical disk. tangible disks and RAM such as read-only memory (ROM), dynamic random access memory (DRAM), static random access memory (SRAM), erasable programmable read-only memory (EPROM), flash memory, electrically erasable programmable memory (EEPROM) read-only memory), a semiconductor device such as a magnetic or optical card, or any other tangible medium suitable for storing electronic instructions.
While the present invention has been described with respect to a limited number of embodiments, it will be understood by those skilled in the art that various changes and modifications therefrom are possible. The appended claims cover all such changes and modifications, and are intended to fall within the true spirit and scope of the present invention.
7 sheets
Sheet 1 Sheet 2 Sheet 3 Sheet 4 Sheet 5 Sheet 6 Sheet 7
Every citation, both ways
| Document | Relation | Office | Cited during |
|---|---|---|---|
| KR101287266B1 | Cited by | Republic of Korea | Search report |
| KR101351183B1 | Cited by | Republic of Korea | Examiner |
| US8769504B2 | Cited by | United States of America | Applicant |
9 members in 4 offices
Priority claims3
| Document | Office | Kind | Date |
|---|---|---|---|
| 2006070041 | Spain | W | |
| 2006070041 | Spain | W | |
| WO2006ES70041 | – | – | – |
Members9
| Document | Office | Kind | |
|---|---|---|---|
| WO2007113346A1 | World Intellectual Property Organization (WIPO) | A1 | |
| KR20080098683AThis record | Republic of Korea | A | |
| CN101416163A | China | A | |
| US2009113240A1 | United States of America | A1 | |
| KR100990591B1 | Republic of Korea | B1 | |
| US8090996B2 | United States of America | B2 | |
| US2012047398A1 | United States of America | A1 | |
| US8402310B2 | United States of America | B2 | |
| CN101416163B | China | B |
5 legal events, as the office reported them to INPADOC
Over the term
Point at a mark for the eventEvents
| Event | Code | |
|---|---|---|
| Lapse due to unpaid annual feeLapsedLAPS | LAPS | |
| Written decision to grantGRNT | GRNT | |
| Decision to grant or registration of patent rightE701 | E701 | |
| Notification of reason for refusalE902 | E902 | |
| Request for examinationA201 | A201 |
Numbers
- Publication
- 10-2008-0098683
- Publication, DOCDB
- 20080098683
- Publication, EPODOC
- KR20080098683
- Application
- 107023877
- Application, DOCDB
- 20087023877
- Application, EPODOC
- KR20087023877
Titles2
- Korean
- 새로운 선택적 구현에 의한 일시적 에러의 검출
- English
- Detection of transient errors by a new optional implementation
Classification
- CPC, 5
- G06F11/008
- G06F11/14
- G06F11/1497
- G06F9/30
- G06F11/07
- IPC, 3
- G06F9 30
- G06F11 14
- G06F11 07