Apparatus and method for densely packing a branch instruction predicted by a branch target address cache and associated target instructions into a byte-wide instruction buffer
Abstract
A branch control apparatus in a microprocessor. Aregister receives a first cache line containing a branchinstruction from an instruction cache in response to afetch address. The fetch address hits in a BTAC thatprovides a target address of the branch instruction. TheBTAC also provides an offset of the instruction followingthe branch instruction. The instructions following thebranch instruction are invalidated based on the offset.Muxing logic packs only the valid instructions into a byte-wide instruction buffer that is directly coupled toinstruction format logic. The instruction cache provides asecond cache line containing the target instructions to theregister in response to the target address. Theinstructions preceding the target instructions areinvalidated based on the lower bits of the target address.The muxing logic packs only the valid target instructionsinto the instruction buffer immediately adjacent to thebranch instruction bytes.
Term
No projected expiry on record.
- Priority
- Filed
- Granted
- Today
46 claims: 13 independent, 33 dependent
- 1經濟部智慧財產局員工消費合作社印製 526451 A8 B8 B414twf.doc/006 Qg 六、申請專利範圍 1. 一種位於微處理器中之分支控制裝置,包括: 一指令快取區,用以輸出藉由一擷取位址所選擇之指 令位元組群中之一線; 一指令緩衝區,耦接該指令快取區且用以緩衝指令位 元組群中之該線; 一分支目標位址快取區(BTAC),耦接該擷取位址且 用以提供與位於指令位元組中之該線中之一分支指令之一 位置相關之一偏移資訊;以及 一選擇邏輯,耦接該分支目標位址快取區且用以根據 該偏移資訊使得一部份指令位元組不被提供至該指令緩衝 區。 2. 如申請專利範圍第1項所述之位於微處理器中之分 支控制裝置,其中該偏移資訊明確說明緊接著指令位元組 中之該線中之該分支指令之一指令之位置。 3. 如申請專利範圍第2項所述之位於微處理器中之分 支控制裝置,其中不被提供至該指令緩衝區之該部份指令 位元組,包括在該偏移資訊被特定時,緊接於指令位元組 群中之該線之分支指令後的指令位元組群。 4. 如申請專利範圍第1項所述之位於微處理器中之分 支控制裝置,其中該選擇邏輯包括: 一暫存器,耦接於該指令快取區與該指令緩衝區間且 用以儲存指令位元組群中之該線。 5. 如申請專利範圍第4項所述之位於微處理器中之分 支控制裝置,其中該選擇邏輯更包括: -----------裝--------訂-----— I — (請先閱讀背面之注意事項再填寫本頁) 本紙張尺度適用中國國家標準(CNS)A4規格(210 X 297公釐) 經濟部智慧財產局員工消費合作社印製 526451 A8 B8 C8 8414twf.doc/006 D8 六、申請專利範圍 複數個有效位元,耦接於該暫存器,其中該些有效位 元中之每一個爲與該暫存器中之指令位元組群中之一個相 關連。 6. 如申請專利範圍第5項所述之位於微處理器中之分 支控制裝置,其中該選擇邏輯根據由該分支目標位址快取 區所接收之該偏移資訊植入該些有效位元。 7. 如申請專利範圍第6項所述之位於微處理器中之分 支控制裝置,其中該選擇邏輯使得該暫存器中之指令位元 組群中之每一個具有一對應有效位元以指示該暫存器中之 指令位元組群中之一個爲無效而不被提供至該指令緩衝 區。 8. 如申請專利範圍第7項所述之位於微處理器中之分 支控制裝置,其中該分支目標位址快取區提供一命中信號 給該選擇邏輯以指示該擷取位址是否命中於該分支目標位 址快取區。 9. 如申請專利範圍第8項所述之位於微處理器中之分 支控制裝置,其中如果該命中信號指示該擷取位址命中於 該分支目標位址快取區時,該選擇邏輯根據由該分支目標 位址快取區所接收之該偏移資訊植入該些有效位元。 10. 如申請專利範圍第5項所述之位於微處理器中之分 支控制裝置,其中該選擇邏輯包括: 多工邏輯,耦接於該指令快取區與指令緩衝區之間’ 且用以藉由該些有效位元中之一個指示相關連之該暫存器 中之指令位元組群中之一個爲有效以被提供至該指令緩衝 本紙張尺度適用中國國家標準(CNS)A4規格(210 X 297公釐) ------1 I I I I ^ in —--I — — — — — — (請先閱讀背面之注意事項再填寫本頁) 526451 A8 B8 C8 D8 8414twf.doc/006 六、申請專利範圍 區。 Π.如申請專利範圍第10項所述之位於微處理器中之 分支控制裝置,其中該多工邏輯包括一組多工器以捨棄被 相關連之該些有效位元指示爲無效之指令位元組群。 12. 如申請專利範圍第10項所述之位於微處理器中之 分支控制裝置,其中該多工邏輯包括一組多工器將被相關 連之該有效位元指示爲有效之指令位元組排整於該指令緩 衝區中之一第一空位置。 13. 如申請專利範圍第10項所述之位於微處理器中之 分支控制裝置,其中該多工邏輯包括一組多工器以藉由移 出該指令緩衝區中之一些位元組之方式,將被相關連之該 有效位元指示爲有效之指令位元組移位被相關連之。 14. 如申請專利範圍第13項所述之位於微處理器中之 分支控制裝置,其中該選擇邏輯被配置去接收由一指令格 式化邏輯而來之一移位計數,以指示該些指令位元組由該 指令緩衝區被移出之一數量。 15. 如申請專利範圍第14項所述之位於微處理器中之 分支控制裝置,其中該多工邏輯藉由該移位計數將被相關 連之該有效位元指示爲有效之指令位元組移位。 16. 如申請專利範圍第1項所述之位於微處理器中之分 支控制裝置,其中該指令緩衝區包括一移位暫存器。 17. 如申請專利範圍第16項所述之位於微處理器中之 分支控制裝置,其中該移位暫存器爲一位元組寬度。 18. 如申請專利範圍第1項所述之位於微處理器中之分 -----------裝--------訂--------- (請先閱讀背面之注意事項再填寫本頁) 經濟部智慧財產局員工消費合作社印製 本紙張尺度適用中國國家標準(CNS)A4規格(210 X 297公釐〉 經濟部智慧財產局員工消費合作社印製 526451 A8 B8 pQ 8414twf. doc/ 006 D8 六、申請專利範圍 支控制裝置,其中該指令緩衝區直接耦接用以格式化指令 位元組群之一指令格式化邏輯。 19. 如申請專利範圍第18項所述之位於微處理器中之 分支控制裝置,其中該指令緩衝區中之一底部位元組直接 被提供至被配置爲用以格式化之一指令之第一位元組之一 部份該指令格式化邏輯。 20. 如申請專利範圍第1項所述之位於微處理器中之分 支控制裝置,其中該分支指令包括一 x86分支指令。 21. 如申請專利範圍第1項所述之位於微處理器中之分 支控制裝置,其中該分支目標位址快取區被配置爲回應於 該擷取位址以提供該分支指令之一目摞位址。 22·如申請專利範圍第1項所述之位於微處理器中之分 支控制裝置,其中該目標位址作爲一接下來擷取位址並選 擇性地被提供至該指令快取區,以選擇指令位元組群中之 一第二線,且該第二線包含該指令快取區中之該分支指令 之一目標指令。 23·如申請專利範圍第22項之位於微處理器中之分支 控制裝置,其中該選擇邏輯使得該目標指令被提供至該指 令緩衝區中,且與該指令緩衝區中之該分支指令相鄰。 24·如申請專利範圍第23項之位於微處理器中之分支 控制裝置,其中該選擇邏輯使得在該第二線中之該目標指 令之前之指令位元組群被捨棄且不被提供至該指令緩衝 區° 25·如申請專利範圍第1項之位於微處理器中之分支控 本紙張尺度適用中國國家標準(CNS)A4規格(210 X 297公釐) -----------^^裝--------訂-丨丨II丨丨丨-*^1^ (請先閱讀背面之注意事項再填寫本頁) 經濟部智慧財產局員工消費合作社印製 526451 A8 B8 Q4l4twf.doc/006 t、申請專利範圍 制裝置,其中該指令快取區儲存由該微處理器所執行之可 變長度之指令群。 26. —種位於微處理器中之預先解碼階層,包括: 一指令緩衝區,用以緩衝指令資料以供應給一指令格 式化邏輯; 一選擇邏輯,耦接該指令緩衝區且用以接收藉由來自 於一指令快取區中之一擷取位址所選擇之一第一指令資 料,其中該第一指令資料包括一分支指令;以及 一分支目摞位址快取區(BTAC),耦接該選擇邏輯且用 以提供該分支指令之一目標位址作爲一下一個擷取位址給 該指令快取區; 其中該選擇邏輯被配置爲接收藉由來自於該指令快取 區之該目標位址所選擇之一第二指令資料,且該第二指令 資料包括該分支指令之一目標指令;以及 其中該選擇邏輯被配置爲將該分支指令以及該目標指 令以相互緊鄰之方式寫入該指令緩衝區中。 27. 如申請專利範圍第26項所述之位於微處理器中之 預先解碼階層,其中該分支目摞位址快取區被配置爲回應 於該擷取位址以提供該目標位址。 28·如申請專利範圍第26項所述之位於微處理器中之 預先解碼階層,其中該分支目摞位址快取區被配置爲提供 緊隨於該分支指令之該第一指令資料中之一位置之一指示 給該選擇邏輯。 29·如申請專利範圍第28項所述之位於微處理器中之 39 本紙張尺度適用中國國家標準(CNS)A4規格(210 X 297公釐) -----------裝--------訂---------AW (請先閱讀背面之注意事項再填寫本頁) 526451 A8 B8 C8 D8 8 4 14 twf . do c/Ο Ο 6 六、申請專利範圍 預先解碼階層,其中該選擇邏輯根據該位置之該指示將該 分支指令以及該目標指令以相互緊鄰之方式寫入該指令緩 衝區中。 30. 如申請專利範圍第29項所述之位於微處理器中之 預先解碼階層,其中該選擇邏輯被配置爲接收該目標位 址。 31. 如申請專利範圍第30項所述之位於微處理器中之 預先解碼階層,其中該選擇邏輯根據該目標位址以及該位 置之該指示將該分支指令以及該目標指令以相互緊鄰之方 式寫入該指令緩衝區中。 32. 如申請專利範圍第28項所述之位於微處理器中之 預先解碼階層,其中該分支目標位址快取區被配置爲回應 該擷取位址以提供該指示。 33. 如申請專利範圍第26項所述之位於微處理器中之 預先解碼階層,其中該指令緩衝區包括一移位暫存器。 34. 如申請專利範圍第33項所述之位於微處理器中之 預先解碼階層,其中該移位暫存器爲一位元組寬度。 35. 如申請專利範圍第33項所述之位於微處理器中之 預先解碼階層,其中該選擇邏輯將在該指令緩衝區中之一 最後有效資料位元組之該分支指令以及該目標指令以相互 緊鄰之方式寫入該指令緩衝區中。 36·如申請專利範圍第33項所述之位於微處理器中之 預先解碼階層,其中該選擇邏輯將該分支指令以及該目標 指令寫入該指令緩衝區中之一下一個空位置。 40 -----------·裝--------訂--------- (請先閱讀背面之注意事項再填寫本頁) 經濟部智慧財產局員工消費合作社印製 本’、'氏張尺度適用中國國家標準(CNS)A4規格(210 X 297公釐 526451 A8 B8 C8 D8 84l4twf.doc/006 六、申請專利範圍 37.如申請專利範圍第33項所述之位於微處理器中之 預先解碼階層,其中該指令緩衝區直接耦接於該指令格式 化邏輯。 3 8. —^種提供分支指令與分支指令之目標指令至指令緩 衝區之方法,其方法包括·· 接收來自於一指令快取區之包括該分支指令之一第一 快取線; 緊隨於該分支指令後,接收來自於一分支目標位址快 取區(BTAC)之該第一^快取線中之一^指令之一^偏移資, 快取線,且該第二快取線爲藉由該分支目標位址快取區所 提供之該分支指令之一目標位址所選擇; 捨棄在該第一快取線中之該分支指令之後之指令群; 捨棄在該第二快取線中之該目標指令之前之指令群; 以及 維持在每個捨棄步驟後提供該第一以及該第二快取線 之一部分給該指令緩衝區。 39. 如申請專利範圍第38項所述之提供分支指令與分 支指令之目標指令至指令緩衝區之方法,其中根據該偏移 資訊捨棄在該第一快取線中之該分支指令之後之指令群。 40. 如申請專利範圍第38項所述之提供分支指令與分 支指令之目摞指令至指令緩衝區之方法,其中根據該目標 位址捨棄在該第二快取線中之該目標指令之前之指令群。 41. 如申請專利範圍第38項所述之提供分支指令與分 -----------裝--------訂--------- (請先閱讀背面之注意事項再填寫本頁) 經濟部智慧財產局員Η消費合作社印製 本紙張尺度適用中國國家標準(CNS)A4規格(210 X 297公釐〉 526451 A8 B8 C8 D8 B414twf·doc/006 六、申請專利範圍 支指令之目標指令至指令緩衝區之方法,其方法更包括: 在接收來自於該指令快取區之該第一快取線之前,提 供一擷取位址給該指令快取區; 其中該指令快取區回應該擷取位址以提供該第一快取 線。 42. 如申請專利範圍第41項所述之提供分支指令與分 支指令之目標指令至指令緩衝區之方法,更包括: 在接收來自於該分支目標位址快取區之該偏移資訊之 前,提供該擷取位址給該分支目標位址快取區; 其中該分支目標位址快取區回應該擷取位址以提供該 偏移資訊。 43. 如申請專利範圍第38項所述之提供分支指令與分 支指令之目標指令至指令緩衝區之方法,其方法更包括: 在捨棄位於該第一快取線中之該分支指令之後之指令群之 前,儲存該第一快取線至一暫存器中。 44. 如申請專利範圍第43項所述之提供分支指令與分 支指令之目摞指令至指令緩衝區之方法,其中捨棄在該第 一快取線中之分支指令後面之指令群,包括將在該暫存器 中之分支指令之後之指令群標記爲無效且不提供在該暫存 器中被標記爲無效之指令群給該指令緩衝區。 45. 如申請專利範圍第38項所述之提供分支指令與分 支指令之目標指令至指令緩衝區之方法’其方法更包括: 在捨棄位於該第二快取線中之該目標指令之前之指令 群之前,儲存該第二快取線至一暫存器中。 -------------裝·-------訂·--I I (請先閱讀背面之注意事項再填寫本頁) 經濟部智慧財產局員工消費合作社印製 本紙張尺度適用中國國家標準(CNS)A4規格(210 X 297公釐) 526451 A8 B8 8414twf.doc/006 六、申請專利範圍 46.如申請專利範圍第45項所述之提供分支指令與分 支指令之目標指令至指令緩衝區之方法,其中捨棄在該第 二快取線中之該目標指令之前之指令群,包括將在該暫存 器中之該目標指令之前指令群標記爲無效且不提供在該暫 存區中被標記爲無效之指令群給該指令緩衝區。 (請先閱讀背面之注意事項再填寫本頁) .裂--------訂---I 經濟部智慧財產局員工消費合作社印製 本紙張尺度適用中國國家標準(CNS)A4規格(210 X 297公釐)
122 paragraphs in 1 section, as filed
Apparatus and method for intensively squeezing a branch instruction and a related target instruction predicted by a branch target address cache area into a byte width instruction buffer
<p>100. . . microprocessor</p><p>102. . . Instruction cache area</p><p>142,144,148,166. . . Data bus</p><p>152,162. . . Select address</p><p>116. . . Branch target address cache area</p><p>BEG value. . . 138</p><p>Hit signal. . . 134</p><p>SBI. . . 136</p><p>118. . . Multiplexer</p><p>156,168,172,174. . . control signal</p><p>122. . . Control logic</p><p>124. . . Adder</p><p>142. . . Cache line</p><p>106. . . Valid bit</p><p>108. . . Multiplex logic</p><p>104, 104A, 104B. . . Byte selection register</p><p>112, 112A, 112B, 112C. . . Instruction buffer</p><p>114. . . Instruction formatting logic</p><p>302. . . Cover multiplexer</p><p>304. . . Row multiplexer</p><p>306. . . Maintain/load multiplexer</p><p>402~424. . . Operational steps of the branch control device</p>
1 is a block diagram of a pipelined processor including a branch control device in accordance with a preferred embodiment of the present invention;
Figure 2 is a diagram showing the coupling of the instruction buffer to the instruction formatting logic according to Fig. 1;
Figure 3 is a block diagram of the multiplex logic according to Figure 1;
Figure 4 is a flow chart showing the operation of the branch control device according to Figure 1;
Figure 5 is a block diagram of the branch controller according to Figure 1.
Information about the application
The present application is related to the following U.S. Patent Application, which is incorporated herein by reference.
<tables><img file="TW526451B_D0001.tif" /></tables>
The present invention relates to a branch target address cache area in a pipelined microprocessor, and more particularly to a microprocessor branch when a hit is caused by a branch target address cache area. A method and apparatus for providing a correct instruction stream to an instruction buffer.
Background of the invention
The pipelined microprocessor includes multiple pipelined classes, each of which performs the different functions necessary for execution of the program instructions. Typical pipelined hierarchy functions are instruction fetch, instruction decode, instruction fetch, memory access, and result write back.
The instruction fetching class retrieves the next instruction in the currently executed equation. This next instruction is typically an instruction with the next consecutive memory address. However, in the case of a branch instruction, this next instruction is an instruction on the memory address specified by the branch instruction. This memory address is often referred to as the branch target address. This instruction fetches the fetch from the instruction cache. If these instructions are not provided by the instruction cache area, they will be fetched from the higher memory of the machine's memory hierarchy, such as higher-order cache memory or system memory, into the instruction cache area. . This captured instruction is provided to the instruction decode level.
The instruction decode hierarchy includes instruction decode logic that is used to decode instruction byte received from the instruction fetch hierarchy. In the case of a microprocessor supporting a variable length instruction, such as a microprocessor of the X86 architecture, one function of the instruction decode hierarchy is to format the instruction byte stream into separate instructions. Formatting the instruction byte stream includes determining the length of each instruction. That is, the instruction formatting logic receives the undifferentiated instruction byte stream from the instruction cache area and formats or parses the instruction byte stream into separate byte group groups. Each group of bytes is an instruction, and these groups of instructions form the program executed by the processor. The instruction decode hierarchy also includes a macro-instruction, such as an X86 instruction, which is a microinstruction that can be executed by the remaining pipeline.
This execution level includes execution logic to execute the formatted and decoded instruction set received at the instruction decode level. This execution logic operates on the basis of data retrieved from a register of processors and/or memory. The write back level stores the results of the execution logic into the processor scratchpad group.
In the operation of the pipelined processor, it is important to maintain that each level of the processor is busy with the operations that were originally designed to perform. Especially when the instruction decoding level is ready to decode the next instruction, and the instruction fetching level cannot provide the instruction byte group, the operation of the processor will be affected. In order to prevent the instruction decoding level from being depleted, the instruction buffer is usually placed between the instruction cache area and the instruction formatting logic. And the instruction fetching level tries to maintain several instructions consisting of the bytes to provide decoding to the instruction decoding level without being deficient.
In general, the instruction cache area provides a queue of instruction bytes for a cache line at a time, typically 16 or 32 byte groups. The instruction fetching layer retrieves one or more cache lines in the instruction byte group in the instruction cache area and stores the cache lines into the instruction buffer. When the instruction decode hierarchy is ready to decode an instruction, the instruction decode hierarchy accesses the instruction byte in the instruction buffer instead of waiting in the instruction cache.
The instruction cache area provides a cache line in the instruction byte group, and a capture address is provided to the instruction cache area by the instruction capture layer to select the cache line. During normal operation, since the program instruction group is continuously executed, it is expected that the capture address will simply increase the floor according to the capacity of a cache line. The added capture address is treated as the next consecutive capture address. However, when the branch instruction is fetched by the instruction decode logic (or was previously fetched), the fetch address is updated to the target address of the branch instruction (by the cache line size) instead of being updated. The next consecutive address is retrieved.
However, as the capture address is updated to the branch target address, the address buffer may be implanted in the instruction byte group in the next consecutive instruction group following the branch instruction. Since a branch has occurred, the instruction group following the branch instruction must not be decoded and must not be executed. That is, execution of the appropriate program would require execution of the instruction group on the branch target address, rather than the next consecutive instruction group after the branch instruction. In the program, it can be foreseen that the continuous instruction stream is more common: the instruction byte group in the instruction buffer is erroneously pre-fetched. To remedy this error, the processor must clear all instruction byte groups after the branch instruction, which includes the instruction byte group in the instruction buffer.
Since the instruction decoding level is now in a depletion condition before the instruction buffer is re-implanted by the instruction cache area, clearing the instruction buffer according to the branch instruction is inevitably costly. One solution to this problem is to branch before decoding the branch instruction. The method can use a branch target address cash (BTAC) to cache the retrieved address of the instruction cache line including the pre-executed branch instruction group and its associated target. The address group is implemented.
When the instruction cache address is applied to the BTAC, it is substantially similar to the application of the capture address to the instruction cache area. In the case of an instruction cache fetch address containing a cache line of a branch instruction, the cache line is provided to the instruction buffer. In addition, if the BTAC hits the BTAC, the BTAC will provide a relevant branch destination address. If the branch instruction hits BTAC as expected, the instruction cache fetch address is updated to the target address provided by the BTAC.
Therefore, the instruction cache area will provide one of the instruction queues to the instruction buffer in one time, and there should be a group of instruction bytes after the branch instruction in the cache line. This instruction byte after the branch instruction should not be executed. More specifically, the instruction byte group in the instruction buffer and after the branch instruction should be discarded from the instruction byte stream being supplied to the instruction decode hierarchy. However, since there are still instructions that are ignored and not yet decoded, they are stored in the instruction buffer, so the instruction buffer will not be cleared all at once (as described above, it will be cleared in processors without BTAC). In particular, the branch instruction itself (and any other instruction byte groups in the cache line and before the branch instruction) must be decoded and executed.
However, when a branch instruction is still stored in the instruction buffer and has not been formatted, the instruction group address following the branch instruction in the instruction buffer is not known. This is because the length and position of the branch instruction in the cache line is not known until the branch instruction is formatted. As a result, the position of the branch instruction located in the instruction buffer is not known. Accordingly, the address of the instruction following the branch instruction is not known.
Finally, the instruction decode hierarchy in a processor of variable instruction length includes several portions to format a portion of the associated instruction group. For example, a portion of the formatting logic is used to format the opcode tuples of instructions, typically the first byte, while other portions of the formatting logic are used to format other portions of the instruction. The correct portion of the instruction byte stream must be supplied to the correct portion of the instruction formatting logic.
In general, it is a difficult task to design instruction-formatted logic in a pipelined processor with the ability to perform instruction formatting within the processor cycle time. This has the advantage of providing as much cycle time as possible to format the instruction byte, rather than spending some time controlling the instruction byte to the appropriate portion of the formatting logic. This has the advantage of using an instruction buffer that is directly coupled to the instruction formatting logic within the instruction decode hierarchy. That is, this advantage is that the exclusion logic is required to manipulate the instruction byte from the instruction buffer to the appropriate portion of the instruction formatting logic.
Therefore, what is needed in the pipelined processor is to have the ability to directly couple the instruction buffer to the instruction formatting logic in the processor by using the BTAC based on the instruction cache fetch address group. A branch control device to increase the processor cycle time available for instruction formatting.
Summary of invention
In view of the above, the present invention provides a pipelined capability of having the instruction buffer directly coupled to the instruction formatting logic in the processor by using the associated BTAC of an instruction cache address group. Branch control device in the processor. Accordingly, in order to achieve the above object, the present invention is characterized by providing a branch control device located in a microprocessor. The device includes an instruction cache area, an instruction buffer, a branch target address cache area, and selection logic. The instruction cache area is configured to output one line of the instruction byte group selected by a capture address. The instruction buffer is coupled to the instruction cache area and is used to buffer the line in the instruction byte group. A branch target address cache area (BTAC) is coupled to the capture address group and is configured to provide compensation information regarding a location of a branch instruction in the line in the instruction byte. The selection logic couples the branch target address cache area and is used to cause a portion of the instruction byte group to be not provided to the instruction buffer according to the compensation information.
In other aspects, the invention features a pre-decoding hierarchy located in a microprocessor. The pre-decoding hierarchy includes an instruction buffer, a selection logic, and a branch target address cache area. The instruction buffer is used to buffer the instruction data to provide instruction formatting logic. The selection logic is coupled to the instruction buffer and is configured to receive the first instruction data selected from the instruction cache area and retrieve the address, wherein the first instruction material includes a branch instruction. The branch target address cache area (BTAC) is coupled to the selection logic and is used to provide the target address of the branch instruction as the next capture address to the instruction cache area. Additionally, the selection logic is configured to receive the second instruction material selected by the target cache address from the instruction cache area, and the second instruction material includes one of the branch instruction target instructions. The selection logic is configured to write the branch instruction and the target instruction to the instruction buffer in close proximity to each other.
In still another aspect, the present invention is further characterized by providing a method for providing a branch instruction and a branch instruction target instruction to an instruction buffer, the method comprising receiving a first cache line including a branch instruction from an instruction cache area And receiving offset information from the instruction byte group immediately following the branch instruction in the first cache line of the branch target address cache area (BTAC). The method also includes receiving a second cache line including the target instruction from the instruction cache area. And the second cache line is selected by the target address of the branch instruction provided by the branch target address cache area. The method also includes discarding the instruction following the branch instruction in the first cache line and discarding the instruction prior to the target instruction in the first cache line. And maintaining a portion of the first and second cache lines to the instruction buffer after each discarding step.
In summary, the present invention has the use of a pre-instruction buffer associated with it, and an instruction buffer using a byte width can be directly coupled to the instruction formatting logic. In other words, no multiplexer is needed between the instruction buffer and the instruction formatting logic to pick the correct instruction byte to the correctly formatted portion. Although the instruction buffer of the byte width is usually smaller than the instruction buffer of the non-byte width, the direct coupling between the instruction buffer and the instruction formatting logic makes the instruction formatting logic essentially reduce the time to limit the formatting logic. pulse. The time limit is reduced by the processor clock cycle time that is increased by overwriting the formatting instructions, rather than spending time controlling the appropriate portion of the instruction byte group from the instruction buffer to the instruction formatting logic. This direct coupling is possible because the present invention only places valid byte groups into the instruction buffer. The present invention treats a group of instruction bytes that are considered invalid due to a branch instruction, and discards it before it is supplied to the instruction buffer.
The above and other objects, features and advantages of the present invention will become more apparent and understood.
Simple illustration
1 is a block diagram of a pipelined processor including a branch control device in accordance with a preferred embodiment of the present invention;
Figure 2 is a diagram showing the coupling of the instruction buffer to the instruction formatting logic according to Fig. 1;
Figure 3 is a block diagram of the multiplex logic according to Figure 1;
Figure 4 is a flow chart showing the operation of the branch control device according to Figure 1;
Figure 5 is a block diagram of the branch controller according to Figure 1.
Main component symbol description
100. . . microprocessor
102. . . Instruction cache area
142,144,148,166. . . Data bus
152,162. . . Select address
116. . . Branch target address cache area
BEG value. . . 138
Hit signal. . . 134
SBI. . . 136
118. . . Multiplexer
156,168,172,174. . . control signal
122. . . Control logic
124. . . Adder
142. . . Cache line
106. . . Valid bit
108. . . Multiplex logic
104, 104A, 104B. . . Byte selection register
112, 112A, 112B, 112C. . . Instruction buffer
114. . . Instruction formatting logic
302. . . Cover multiplexer
304. . . Row multiplexer
306. . . Maintain/load multiplexer
402~424. . . Operational steps of the branch control device
Preferred embodiment
Referring to Figure 1, a block diagram of a pipelined processor 100 including a branch control device in accordance with the present invention is shown. In one embodiment, microprocessor 100 includes a processor of the x86 architecture. In one embodiment, the microprocessor 100 has a 13-level pipeline including: an instruction fetch hierarchy, a plurality of instruction cache access levels, an instruction format hierarchy, an instruction decode or a conversion hierarchy, and a The temporary access level, an address calculation level, multiple data cache access levels, multiple execution levels, one storage level, and one write back level.
Microprocessor 100 includes an instruction cache area 102 for a cache instruction byte group. This instruction byte group is received from memory via data bus 166. Instruction cache area 102 includes a cache line that stores an array of instruction byte groups. The cache line of this array is indexed by a capture address 152. Therefore, the capture address 152 selects one of the cache lines in the array. The instruction cache area 102 outputs the selected cache line in the instruction byte group via the data bus 142.
In one embodiment, the instruction cache area 102 includes a 64KB, 4-channel group of associated cache areas, and each channel has a 32-bit tuple line group. In one embodiment, half of the cache lines selected in the instruction byte group are provided by the instruction cache area 102 for a time. In other words, 16 bytes are provided in every two separate periods. In one embodiment, the instruction cache area 102 is similar to the US Patent Specification Serial No. 09/849,736, entitled "SPECULATIVE BRANCH TARGET ADDRESS CACHE" (label number CNTR: 2021). Described.
The microprocessor 100 also includes a branch target address cache (hereinafter referred to as BTAC) 116. The BTAC 116 also receives the retrieved address 152 of the instruction cache area 102. The BTAC 116 stores the retrieved address group of the branch instruction group that was previously executed. The BTAC 116 includes a storage element that stores the branch instruction target address after the branch instruction executed by the microprocessor 100. The storage element also stores information about other uncertain branches when the target address of the branch instruction is cached. In particular, the storage element stores location information of instructions following the branch instruction in the cache line. The capture address 152 indexes the storage elements of the array in the BTAC 116 to select one of the storage elements.
The BTAC 116 selects a target address 132 and a speculative branch information (hereinafter referred to as SBI) 136 from the storage element by the capture address 152. In one embodiment, the SBI 136 includes a branch instruction length, whether the branch instruction is deformed by a plurality of instruction cache lines, whether the branch is a call or a return instruction or a direction for predicting the branch instruction. News. The above may refer to the previously mentioned U.S. patent specification entitled "Uncertain branch target address cache area".
The BTAC 116 also outputs a BEG value 138 in which the offset information for the instruction group following the associated branch instruction in the instruction cache line 142 selected by the capture address 152 is specified. This BEG value 138 is cached in the BTAC 116 alone with the instruction target address after the branch instruction is executed. This BEG value 138 is calculated by the index of the branch instruction added to the length of the branch instruction. If the microprocessor 100 branches the target address 132 provided by the BTAC 116, all of the instruction sets in the cache line after the branch instruction are considered invalid. That is, the instruction byte group after the branch instruction will not be executed, and the reason is that the branch instruction has been obtained. Therefore, for the execution of a unique program, the instruction byte group following the branch instruction must be discarded before execution.
The BTAC 116 also outputs a hit signal 134 to indicate whether the capture address 152 hits in the BTAC 116. In one embodiment, BTAC 116 is similar to that described in the previously mentioned U.S. patent specification entitled "SPECULATIVE BRANCH TARGETADDRESS CACHE". In particular, the BTAC 116 is an indeterminate BTAC because the instruction cache line provided by the microprocessor 100 in the instruction cache area 102 is decoded so that it exists in the cache line selected by the capture address. The target address 132 provided by the BTAC 116 is branched before the known or unknown branch instruction. That is to say, even if there is a possibility that the BTAC 116 is hit by the capture address to select the cache line without any branch instruction, the microprocessor 100 may branch indefinitely.
Microprocessor 100 also includes control logic 122, hit signal 134, SBI 136, and BEG value 138 to all be provided as inputs to control logic 122. The operation of control logic 122 is described in more detail below.
Microprocessor 100 also includes a multiplexer 118. The multiplexer 118 receives at least three types of addresses as inputs and corresponding to the control signal 168 output by the control logic to select one of the inputs as the capture address l52 for output to the instruction cache area 102. The multiplexer 118 receives the target address 132 output by the BTAC 116. The multiplexer 118 also receives the next consecutive capture address 162 and provides the next consecutive capture address 162 to the multiplexer 118. The next consecutive capture address 162 is a pre-fetched address that is incremented according to the cache line size in the instruction cache area 102 and received by the adder 124. This decision target address is provided by the execution logic in microprocessor 100.
The multiplexer 118 also receives a parsed target address 164. The parsed target address 164 is provided by execution logic in the microprocessor 100. This execution logic decodes the parsed target address 164 based on a full decoding of a branch instruction. If, after branching the target address 132 provided by the BTAC 116, and the branch of the microprocessor 100 later determines to be the wrong branch, the microprocessor 100 will clear the pipeline and branch other profiling target addresses 164 or The capture address of the cache line of the instruction group after the branch instruction is included to correct the error. In one embodiment, assuming that the microprocessor 100 determines that no branch instruction exists in the cache line l42, the microprocessor 100 will remove the pipeline and the capture address group for the cache line including the branch instruction itself. Make a branch to correct the error. The correction of this error is described in U.S. Patent Specification Serial No. 09/849,658, entitled Detecting and Correcting Error Uncertain Branch Target Address Cache Branch and Tag Number CNTR2022.
In one embodiment, multiplexer 118 also receives other target addresses predicted by other branch prediction elements, such as a call/back stack and a branch target buffer (BTB), depending on the branch. The instruction indicator is to cache the target address of the indirect branch instruction. The multiplexer 118 selectively overwrites the target address 132 provided by the BTAC 116 by calling/back stacking or BTB. It is described in U.S. Patent Specification Serial No. 09/849,799, entitled Uncertain Branch Target Address Cache Area, which is based on the branch and selectively overwritten by the second prediction device, and the label number is CNTR2052.
The microprocessor 100 also includes a byte select register 104 to receive a selected one of the command bit groups output by the instruction cache area 102 via the data bus 142. In one embodiment, the byte select register 104 has a 16 byte width to receive 16 instruction bytes output by the instruction cache area 102 for a time. Microprocessor 100 also includes 16 valid bits 106 corresponding to each of the 16 of the byte select registers 104. The valid bit 106 will indicate the corresponding byte in the select scratchpad, regardless of whether the corresponding byte is a valid byte when the microprocessor 100 is temporarily buffered, formatted, and executed. Control logic 122 outputs a control signal based on HIT 134, SBI 136, and BEG 138 output from BTAC 116 to select to reset or clear valid bit 106. In the byte select register 104, the instruction byte group following a fetched branch instruction is marked as invalid by the control logic 122. In addition to this, the instruction byte group before the target instruction of a branch instruction is also marked as invalid. In an embodiment, the valid bit 106 is included in the control logic 122.
The microprocessor 100 also includes multiplex logic 108 and receives the byte output by the byte select register 104 via the data bus 144. The multiplex logic 108 discards the invalid byte set output by the byte select register 104. For the description of the multiplex logic 108 and its operation, please refer to the related figures 3 and 4 below.
Microprocessor 100 also includes an instruction buffer 112 for receiving valid instruction bytes output by multiplex logic circuit 108 via data bus 146. The advantage is that the multiplex logic 108 selectively selects only the valid byte from the byte output from the byte select register based on the control signal 156 output from the control logic 122 to provide the instruction buffer. 112. The instruction byte group corresponding to the valid bit 106 in the byte select register 104 and indicated as invalid is discarded by the multiplexer 108 and is not provided to the instruction buffer 112. .
In one embodiment, the instruction buffer 112 stores 13 instruction bytes. The instruction buffer 112 includes a byte width shift register to store the instruction byte group. The advantage is that the instruction buffer 112 moves out of the already formatted instruction byte on a byte granular substrate, thus retaining the first byte of the next instruction at the bottom of the instruction buffer. And the following is a more detailed description.
Microprocessor 100 also includes instruction formatting logic 114 to receive the instruction byte group output by instruction buffer 112 via data bus 148. Instruction formatting logic 114 examines the contents of instruction buffer 112 and formats or parses the instruction bytes it contains into separate instructions. In particular, the instruction formatting logic 114 determines the byte size of the instruction at the bottom of the instruction buffer 112. Instruction formatting logic 114 provides formatted instructions to the remaining pipelines of microprocessor 100 for more decoding and execution. This has the advantage that the instruction buffer 112 buffers the instruction byte to reduce the chance of the instruction formatting logic 114 being depleted.
Instruction formatting logic 114 provides the instruction size at the bottom of instruction buffer 112 to control logic 122 via control signal 172. Control logic 122 uses the instruction size of this control signal 172 to control the shifting of the command value of the instruction byte by being removed from instruction buffer 112 by control signal 174. That is, control signal 172 provides a shift count service for multiplex logic 108 and instruction buffer 112. Control logic 122 also uses control signal 172 to control the loading of instruction bytes into instruction buffer 112. Control logic 122 also uses control signal 172 to control the operation of multiplex logic 108 through control signal 156. In one embodiment, the instruction formatting logic 114 is capable of formatting a plurality of instructions at a clock cycle of 100 cycles per processor.
Referring now to Figure 2, there is shown the coupling of instruction buffer 112 to instruction formatting logic 114 in Figure 1. The instruction formatting logic 114 includes several separate portions that are different instruction byte sets for formatting one of the instructions in the microprocessor 100. In the embodiment illustrated in the second figure, the instruction formatting logic 114 includes a portion for formatting byte 0, a portion for formatting byte 1, and a formatting byte. Part of 2... A part to format the byte N. Correspondingly, the instruction buffer 112 includes the bytes 1, 2, ..., N.
The formatting logic 114 is configured to read a byte 0 of the formatting instruction directly from the instruction buffer 112 via the data bus 148 of FIG. Similarly, formatting logic 114 is configured to read a portion of byte 1 of the formatted instruction directly from instruction buffer 112 via byte 148. The formatting logic 114 is configured to read a portion of the byte 2 of the formatted instruction directly from the instruction buffer 112 via the data bus 148. And so on to the byte N.
Thus, from FIG. 2, it can be observed that the present invention has the advantage of providing a direct coupling between instruction buffer 112 and instruction formatting logic 114 such that no manipulation logic is required for manipulation of the instruction byte. Therefore, the clock characteristics of the microprocessor 100 are likely to be improved. This advantage is in part due to the byte width instruction buffer 112 shifting an entire instruction out of a time so that the byte at the bottom of the instruction buffer is always the byte 0 of the next instruction to be formatted by the instruction logic. format.
An advantage of the present invention is that there is no need to provide a manipulation logic between instruction buffer 112 and instruction formatting logic 114 to manipulate the instruction byte to the correct portion of formatting logic 114. In general, the steering logic has a decisive influence on the timing and thus increases the clock cycle of the microprocessor 100. However, with the advantages of the present invention, the periodic clock of microprocessor 100 may be shortened when employing a BTAC 116 as described herein.
Referring now to Figure 3, there is shown a block diagram of multiplex logic 108 in accordance with Figure 1 of the present invention. Multiplex logic 108 includes a device having 13 16:1 overlay multiplexers 302, and each overlay multiplexer 302 receives BSR[15:0], which is a data stream. Row 144 is from the byte select register 104 in FIG. Each overlay multiplexer 302 selects a byte. The byte is provided on the output M[12:0] by the overlay multiplexer 302.
The overlay multiplexer 302 is controlled by the control logic 122 via the control signal 156 of FIG. The multiplexer 302 performs an unmasked or discarded operation on the invalid instruction byte (referred to by the effective bit 106) stored in the byte select register 104 in an interconnected manner and provides only The valid instruction byte is on the bottom of the output multiplexer 302 output M[12:0]. That is, the first significant byte in the byte select register 104 will be provided on the output byte M[0], the second valid byte on the output byte M[1]... So that when the 13th significant byte exists in the byte select register 104, the 13th effective byte will be provided on the output byte M[12]. So, for example, assume that the byte 2 (eg, BSR[2]) output by the byte select register 104 is the first valid byte, and the first valid byte will It is output on M[0]; assuming BSR[3] is valid, it will be output on M[1], and so on.
The multiplex logic 108 includes a device having 24 13:1 packed multiplexers 304. Each of the finishing multiplexers 304 receives the 13 bytes M[12:0] output by the overlay multiplexer 304. Each of the finishing multiplexers 304 selects a byte. The selected byte is provided on output A[23:0]. The multiplexer 304 is controlled by the control logic 122 via a control signal 156. Each of the row multiplexers 304 interconnects the received valid bit groups from M[12:0] of the multiplexer 302 to the first in the instruction buffer 112 in an interconnected manner. Pre-shift location. That is, the multiplexer 304 aligns the received significant bit groups on M[12:0] to the vacant bits in the instruction byte before the instruction buffer 112 moves out of the already formatted instruction. Group location.
The top multiplexer 304 is not necessarily all 13:1 multiplexers. In one embodiment, for byte 23, the multiplexer 304 is a 1:1 multiplexer that receives only M[12] as input, and for the tier 22, the row is more The worker 304 is a 2:1 multiplexer that receives only M[12:11] as an input, and for the byte 21, the multiplexer 304 is only receiving M[12:10] as an input. A 3:1 multiplexer and analogy down to byte 12, the multiplexer 304 is a 12:1 multiplexer that receives only M[12:1] as input.
In one embodiment, the hood multiplexer 302 and the multiplexer 304 may be combined into a single device having a plurality of multiplexers to operate the combined functions described above in FIG.
Multiplex logic 108 is a device 306 that includes a set of 13 2:1 hold/loading multiplexers, and each of the sustain/load multiplexers 306 receives two input bits. group. One of them is the output A[n] corresponding to the input multiplexer 304. For byte 0, the sustain/load multiplexer 306 receives one A[0] as input by the multiplexer 304, and the multiplexer 306 for the byte 1 by the multiplexer 306 The multiplexer 304 receives an input A[1], and so on. For the byte 12, the sustain/load multiplexer 306 receives an input A from the multiplexer 304. 12]. For the sustain/load multiplexer 306, the second input is one of the corresponding outputs IB[12:0] provided on the bus 148 for the instruction buffer 112 in FIG. That is, for byte 0, the sustain/load multiplexer 306 receives the second IB[0] as input from the instruction buffer 112, and maintains/loads the multiplexer for the byte 1. 306 receives the second IB[1] as input by instruction buffer 112, and so on. For byte 12, the sustain/load multiplexer 306 receives the second as input from instruction buffer 112. IB [12].
The sustain/load multiplexer 306 selects one of the two inputs depending on whether the corresponding byte is valid in the instruction buffer 112. Control logic 122 controls maintenance/load multiplexer 306 via control signal 156. Thus, for example, assume that the byte 5 in the instruction buffer 112 is valid, and for the byte 5, the corresponding sustain/load multiplexer 306 selects the IB[5] input to maintain the instruction. The value in buffer 112 is instead of an instruction byte received by byte select register 112. Conversely, assuming that the byte 5 in the instruction buffer 112 is an invalid instance (eg, the location of the byte 5 in the instruction buffer 112 is empty), the corresponding sustain/load multiplexer 306 is For byte 5, the A[5] input is selected to receive an instruction byte by the byte select register 104 instead of maintaining the value of IB[5]. The sustain/load multiplexer 306 provides the selected input on the output X[12:0].
The multiplex logic 108 includes a set of devices having 13 shift multiplexers 308, and each shift multiplexer 308 receives the output (A[23:0]) included by the multiplexer 304. The 12 input bytes consisting of different combinations between the outputs (X[12:0]) of the multiplexer 306 are maintained/loaded. Shift multiplexer 308 receives X[11:0] for byte 0, and X[12:1] for byte 1, for byte 2 Said shift multiplexer 308 receives A[13] and X[12:2], for byte 3, shift multiplexer 308 receives A[14:13] and X[12:3], By analogy to bit tuple 12, shift multiplexer 308 receives A[23:13] and X[12].
Shift multiplexer 308 shifts the byte selected by mask over multiplexer 302, packed multiplexer 304, and sustain/load multiplexer 306 based on the number of instruction byte transferred out of instruction buffer 112. . That is, when the instruction formatting logic 114 formats an instruction, the instruction buffer 112 reads and determines the byte size of the formatted instruction, and this size determines the number of bytes that will be shifted out of the instruction buffer 112. . The size of the formatted instruction is shift count 172 in FIG. 1, and control logic 122 uses this shift count 172 and controls shift multiplexer 308 via control signal 156 to select the appropriate input to produce the shift.
The output of shift multiplexer 308 is provided to a corresponding one of the instruction buffers 112 to load the masked, aligned, maintained/loaded, and shifted instruction byte groups. The loading of instruction buffer 112 is controlled by control logic 122 via control signal 174.
Thus, it can be observed that the multiplex logic 108 of FIG. 3 has the advantage that only the active instruction byte output from the byte select register 104 is densely squeezed into the instruction buffer 112, and in the following illustration There is a more detailed description.
Referring now to Figure 4, there is shown a flow chart of the operation of the branch control device in accordance with Figure 1 of the present invention. The flow begins at block 402.
In block 402, the retrieved address 152 in FIG. 1 is provided to the instruction cache area 102 in FIG. 1 to select one of the instruction byte groups from the instruction cache area 102. In addition, when the capture address 152 is provided to the BTAC 116 in FIG. 1 and the target address 4 of the capture address 152 is stored in the BTAC 116, the capture address 152 is generated in the BTAC 116. A hit. The next step is from block 402 to block 404.
In block 404, the cache line in the instruction group (also including the branch instruction) that has been selected in step 402 is loaded into the byte select register 104 in FIG. In addition, because the fetch address 152 in the cache line containing the branch instruction is cached in the BTAC 116, the control logic 122 detects the hit address 152 hit in the BTAC 116. The processes below step 406 and step 414 are essentially simultaneous operations. The flow first described includes blocks 406, 408, and 412.
In block 406, control logic 122 derives BEG value 138 from BTAC 116, which specifies the offset information for the instruction byte group following the branch instruction selected in the cache line in step 402. Because the byte select register 104 is the same size as a cache line, and the cache line containing the branch instruction is loaded into the byte select register 104 in step 404, therefore, the BEG The value 138 is the offset information of the instruction group after the branch instruction in the byte select register 104. For example, suppose the branch instruction starts at byte 3 and is 2 bytes long. Before the branch instruction, the instruction group will start before the byte 5 in the byte selection register, by BTAC116. The BEG value 138 provided will be 5. The next step is from block 406 to block 408.
In block 408, control logic 122 overwrites the valid bit group for all of the instruction groups in the byte select register 104 following the branch instruction based on the BEG value 138 obtained from BTAC 116 in step 406. Therefore, all of the instruction byte groups marked as invalid in the byte select register 104 will then be discarded and not provided to the instruction buffer 112 of FIG. For example, assuming the BEG value 138 is 5, the bytes 5 through 15 in the byte select register 104 will be marked as invalid. The next step is from block 408 to block 412.
In block 412, the instruction byte set in the byte select register 104 and marked as valid in step 408 is loaded into the instruction buffer 122 by the control logic 122. Conversely, the instruction byte in the byte select register 104 and marked as invalid in step 408 is not loaded into the instruction buffer 112. Assuming that there are already programmed instruction byte groups in the instruction buffer 112, these formatted instruction byte groups are before the valid byte is loaded by the byte select register 104. The instruction buffer 112 is shifted out. The number of valid tuples in the instruction buffer 112 is limited by the number of valid tuple locations in the instruction buffer 112 after the number of significant bytes from the byte select register 104 entering the instruction buffer 112 is shifted out of the formatted instruction group.
According to the above description of FIG. 3, the valid byte from the byte select register 104 is directly loaded after the last valid and non-formatted instruction byte in the instruction buffer. To the instruction buffer 112. For example, assume that there are 15 bytes in the byte select register 104 that are valid, and after the formatted byte is removed, there are 9 bytes in the instruction buffer 112 that are empty. In this case, only the first nine valid bytes from the byte select register 104 are loaded into the instruction buffer 112.
Described below is another concurrent flow of steps, including blocks 414, 416, and 418.
In block 414, the target address of the branch instruction is obtained by BTAC 116 and provided to multiplexer 118 in FIG. The target address 132 obtained by the BTAC 116 corresponds to the BEG value 138 obtained by the BTAC 116 in step 406. The next step is block 414 to block 416.
In block 416, multiplexer 118 selects the target address provided by BTAC 116 as the next capture address 152 of instruction cache line 102. Selection of the target address 132 will cause the microprocessor 100 to indefinitely branch the cached target address 132 of the branch instruction contained in the selected cache line in step 402. That is, the target address 132 will be used as the next capture address 152 to select the cache line containing the branch target instruction from the instruction cache area 102. The next step is block 416 to block 418.
In block 418, a cache line in the instruction byte group of the target instruction containing the branch instruction is loaded into the byte select register 104. In step 418, the cache line including the target instruction byte is loaded to the byte select register 104 and in step 412, the valid byte including the branch instruction is selected by the byte select register 104. The operations loaded into the instruction buffer 102 are essentially simultaneous. Next, the process of combining blocks 412 and 418 and performing 412, 418 to block 422.
In block 422, the control logic 122 overwrites the valid bit group, and the valid bit group is selected for temporary storage of all the bytes prior to the target instruction obtained by the BTAC 116 in accordance with the target address 132 in step 414. The instruction byte group in the device 104. Therefore, all instruction byte groups that are marked as invalid in the byte select register 104 will then be discarded and not provided to the instruction buffer 112. For example, assuming that the least significant bits of the target address 132 are 0x7, the byte 0 to the byte 6 in the byte select register 104 will be marked as invalid. The next step is block 422 to block 424.
In block 424, the active instruction byte of the target instruction is now included after the last byte of the branch instruction, and is directly loaded into the instruction buffer 112 by the byte select register 104. An example of the operation of the branch control device described above can be described in accordance with FIG.
Referring now to Figure 5, there is shown a block diagram of one embodiment of a branch controller in Figure 1 in accordance with the present invention. Figure 5 shows the contents of two different levels of the byte select register 104 in Figure 4, labeled as 104A and 104B. Figure 5 also shows the contents of the three different levels of instruction buffer 112 in Figure 4, which are labeled 112A, 112B, and 112C.
Instruction buffer 112A shows the starting content of instruction buffer 112 in this embodiment. Instruction buffer 112A contains 7 valid bytes and 6 empty (or invalid) byte locations. Thus, in this embodiment, the byte locations 0 through 6 in each of the instruction buffers 112A contain an instruction byte, labeled A through G.
The byte select register 104A displays the byte select register 104 after being loaded into the branch instruction contained in step 404 of FIG. 4, and in step 408 of FIG. The 406BTAC 116 receives the offset information such that the byte after the valid bit is overwritten selects the contents of the register 104.
In byte select register 104A, byte 0 contains an instruction byte labeled Q. Byte 1 contains an instruction byte that is labeled R. Byte 2 contains the first byte of a two-tuple branch instruction, which is labeled JCC and is the operational byte of the x86 traditional jumper instruction. Byte 3 contains the second byte of this JCC instruction and is represented as an alternative, so it is labeled disp.
In this embodiment, before the branch instruction group has a byte offset of 4 after the branch instruction, in step 406 of FIG. 4, the BEG value 138 obtained by the control logic 122 from the BTAC 116 is 4. Thus, in step 408 of FIG. 4, control logic 122 causes bytes 0 through 3 to be active and 4 through 15 to be inactive.
Instruction buffer 112B shows that instruction buffer 112 is loaded by the byte select register 104A after shifting out five formatted instruction bytes (e.g., instruction bytes A through E) in this embodiment. The contents of a valid instruction byte (such as Q, R, JCC, and disp). That is, the instruction buffer 112B displays the content after execution of step 412 in FIG. The invalid instruction byte after the branch instruction and occupying the positions of bytes 4 through 15 in the byte select register 104A is discarded by the multiplexer 302 in FIG. 3 and is not loaded into the instruction buffer. In area 112.
Byte 0 of instruction buffer 112B contains instruction byte F that is offset by byte 5 of instruction buffer 112A. Byte 1 of instruction buffer 112B contains instruction byte G offset by byte 6 of instruction buffer 112A. That is, byte 0 and byte 1 of instruction buffer 112B are maintained by sustain/load multiplexer 306.
In byte 2 of instruction buffer 112B, byte 0, which is included by byte select register 104A, is loaded into instruction buffer 112B by multiplex logic 108 and according to step 412 in FIG. The instruction byte Q of the most significant byte in the write-shift instruction buffer 112B is placed next to it. Similarly, the byte 3 of the instruction buffer 112B includes the byte 1 of the byte select register 104A through the multiplex logic 108 and is loaded into the byte 3 of the instruction buffer 112B. Instruction byte R. Similarly, the byte 4 of the instruction buffer 112B includes the byte 2 of the byte select register 104A through the multiplex logic 108 and is loaded into the byte 4 of the instruction buffer 112B. JCC opcode instruction byte. Similarly, byte 5 of instruction buffer 112B contains the JCC instruction by byte 3 of byte select register 104A through multiplex logic 108 and loaded into byte 5 of instruction buffer 112B. The alternative byte. The byte select register 104B displays the contents of the cache line containing the branch target instruction loaded by the byte select register 104 in step 4 of FIG. In this embodiment, the branch target instructions placed in the bytes 13 through 15 of the byte select register 104B are labeled X, Y, and Z. Bytes X, Y, and Z construct one or two or three instructions based on their instruction length. Because in this embodiment, the first byte of the target instruction is placed in the byte 13 of the byte select register 104B, the control logic 122 must bit the bit by clearing the associated valid bit 106. Tuples 0 through 12 are marked as invalid. The bytes 0 through 12 of the byte select register 104B will be discarded by the overlay multiplexer 302 and will not be provided to the 112 instruction buffer.
The instruction buffer 112C shows three valid instruction bytes (e.g., X,) that the instruction buffer 112C does not remove any formatted instruction bytes and load the target instruction from the byte select register 104B. After Y and Z). In other words, the instruction buffer 112C displays the content after execution of step 424 in FIG. The invalid instruction byte that occupies bytes 0 through 12 of the byte select register 104B and precedes the target instruction is discarded by the overlay multiplexer 302 and is not provided to the instruction buffer 112.
Bytes 0 through 5 of instruction buffer 112C contain the same values as bytes 0 through 5 of instruction buffer 112B before any occurrences of the shift occur. The reason for this is that the instruction formatting logic 114 in Figure 1 cannot format the instructions in the previous periodic clock. An example of the reason is that the instruction formatting logic 114 in the microprocessor 100 pipeline cannot format an instruction that is a field (eg, by a floating point instruction or the execution of a previous branch instruction, which requires expense). Many clocks are used to complete).
The byte 6 of the instruction buffer 112C contains instructions for the byte 13 of the byte select register 104B to be loaded into the byte 6 of the instruction buffer 112C by the multiplex logic 108 and according to step 424. Byte X. The instruction byte X is loaded into the instruction buffer 112C and is immediately adjacent to the last byte of the branch instruction (e.g., adjacent to the disp byte in the instruction buffer 112C). Similarly, the byte 7 of the instruction buffer 112C contains the byte 7 of the instruction buffer 112C that was loaded by the byte 14 of the byte select register 104B through the multiplex logic 108 and according to step 424. The instruction byte in the Y. Similarly, byte 8 of instruction buffer 112C contains the byte 15 of the byte select register 104B through the multiplexer logic 108 and is loaded into the instruction buffer 112C according to step 424. Command byte Z in 8.
It is worth noting that branch instructions may be included in two different instruction cache lines. In one embodiment, BTAC 116 assigns a hit signal 134 associated with the retrieved address 152 of the first cache line (eg, with a cache line containing the first portion of the branch instruction). In an alternate embodiment, the BTAC 116 will assign a hit signal 134 associated with the second cache line (e.g., with the capture address 152 of the cache line 3 containing the second portion of the branch instruction. In other words, the BTAC 116 cache The second part of the branch instruction fetches the address 152 instead of the first part.
From the discarding example, the operation of the branch control device in FIG. 1 is to intensively squeeze only the valid instruction byte into the byte width instruction buffer 112, remove the formatted instruction byte, and will not The formatted instruction byte is stored at the bottom of the instruction buffer 112. As shown in FIG. 2, the instruction buffer is directly coupled to the instruction formatting logic 114 such that the active instruction byte at the bottom of the instruction buffer 112 is provided directly to the instruction formatting logic 114 without the need to add manipulation logic. . The instruction byte is directly supplied from the instruction buffer 112 to the instruction formatting logic 114. The present invention has the advantage of increasing the availability of instruction formatting time and higher than required between the instruction buffer 112 and the instruction formatting logic 114. The availability of instruction formatting time under the control logic architecture.
Another advantage of the present invention is that the branch control device operates with a byte width and is directly coupled to the instruction buffer to be operated according to the branch prediction control predicted by the BTAC before the branch instruction is decoded by the instruction decoding logic. Branch.
Although the present invention has been disclosed in the above preferred embodiments, it is not intended to limit the present invention. For example, the instruction buffer may vary in size, and more or fewer instruction bits may be stored than in the above embodiment. group. In addition, the size of the byte select register and the instruction cache line are also changed. Therefore, any person skilled in the art can make various changes and modifications without departing from the spirit and scope of the invention, and the scope of the invention is defined by the scope of the appended claims.
18 members in 3 offices
Priority claims4
| Document | Office | Kind | Date |
|---|---|---|---|
| 09898583 | United States of America | – | |
| 89858301 | United States of America | A | |
| 20010898583 | – | – | – |
| US20010898583 | – | – | – |
Members18
| Document | Office | Kind | |
|---|---|---|---|
| CN1369780A | China | A | |
| CN1375767A | China | A | |
| CN1376977A | China | A | |
| TW526451BThis record | Taiwan Province of China | B | |
| TW530205B | Taiwan Province of China | B | |
| TW564369B | Taiwan Province of China | B | |
| US6823444B1 | United States of America | B1 | |
| US2005044343A1 | United States of America | A1 | |
| US2005198479A1 | United States of America | A1 | |
| US2005198481A1 | United States of America | A1 | |
| US2006010310A1 | United States of America | A1 | |
| CN1249575C | China | C | |
| CN1270234C | China | C | |
| CN1279442C | China | C | |
| US7159098B2 | United States of America | B2 | |
| US7162619B2 | United States of America | B2 | |
| US7203824B2 | United States of America | B2 | |
| US7234045B2 | United States of America | B2 |
2 legal events, as the office reported them to INPADOC
Over the term
Point at a mark for the eventEvents
| Event | Code | |
|---|---|---|
| Expiration of patent term of an invention patentMK4A | MK4A | |
| Issue of patent certificate for granted invention patentGrantedGD4A | GD4A |
Numbers
- Publication
- 526451
- Publication, DOCDB
- 526451
- Publication, EPODOC
- TW526451B
- Application
- 90127266
- Application, DOCDB
- 90127266
- Application, EPODOC
- TW20010127266
Titles5
- English
- Apparatus and method for densely packing a branch instruction predicted by a branch target address cache and associated target instructions into a byte-wide instruction buffer
- Chinese
- 將藉由分支目標位址快取區所預測之分支指令與相關目標指令密集擠入位元組寬度指令緩衝區之裝置及方法
- English
- APPARATUS AND METHOD FOR DENSELY PACKING A BRANCH INSTRUCTIONPREDICTED BY A BRANCH TARGET ADDRESS CACHE AND ASSOCIATEDTARGET INSTRUCTIONS INTO A BYTE-WIDE INSTRUCTION BUFFER
- Unlabeled
- 將藉由分支目標位址快取區所預測之分支指令與相關目標指令密集擠入位元組寬度指令緩衝區之裝置及方法
- Unlabeled
- Apparatus and method for intensively squeezing a branch instruction and a related target instruction predicted by a branch target address cache area into a byte width instruction buffer
Classification
- CPC, 1
- G06F9/3806
- IPC, 4
- G06F9 00
- G06F9 30
- G06F9 42
- G06F12 02