Apparatus and method for data processing and computer program product
Abstract
A data processing system is provided with an instruction (ADD8TO16) thatunpacks non-adjacent portions of a data word using sign or zero extension andcombines this with a single-instruction-multiple-data type arithmetic operation, suchas an add, performed in response to the same instruction. The instruction is wellsuited to use within systems having a data path (2) including a shifting circuit (6)upstream of an arithmetic circuit (8).

Term
No projected expiry on record.
- Priority
- Filed
- Granted
- Today
15 claims: 15 independent, 0 dependent
- 1一種資料處理裝置,該裝置至少包含:(i)一平移電路;(ii)一算術電路;和(iii)一指令解碼器,在一制令下達之後動作,用以控制上述平移電路和上述算術電路的一指令,且對一資料字組Rn及一資料字組Rm執行一運算產生一特定的值,該動作包含下列步驟:(iv)選擇多個上述資料字組Rm的非相鄰多位元部分組成一位元長度A的多位元部分;(v)隨意地以一般的的平移量平移上述多個多位元部分至平移後的位元位置;(vi)將上述多組多位元部份每個都從上述之A位元長度提升(promote)至B位元長度用以構成多組已提升的多位元部份,如此一來上述已提升的多位元部份可以連接構成一已提升的資料字組P;及(vii)執行多組獨立的算術運算當作輸入運算元,其為B位元長度之個別的位元位置部份由上述已提升之資料字組P和上述的字組Rn構成一結果字組Rd。
- 2如申請專利範圍第1項所述之裝置,其中B=2*A。
- 3如申請專利範圍第1項所述之裝置,其中上述多個為位元部分會平移至以平移的位元位置使得一最低的位元位置部分從第零位元位置往上延伸。
- 4如申請專利範圍第1項所述之裝置,其中將上述多位元部分從位元長度A提升至位元長度B的步驟至少包含下列之一:(i)有號延伸(sign extending)上述多位元部分至位元長度為B;且(ii)零延伸(zero extending)上述多位元部分至位元長度為B。
- 5如申請專利範圍第1項所述之裝置,其中上述獨立的算術運算為獨立的加法運算。
- 6如申請專利範圍第1項所述之裝置,其中上述的資料字組位元長度為C且C=N*B,其中N為大於一的整數。
- 7如申請專利範圍第2項所述之裝置,其中C=B*2。
- 8如申請專利範圍第1項所述之裝置,其中B=16,且A=8。
- 9如申請專利範圍第1項所述之裝置,其中上述一般位移量是B-A。
- 10如申請專利範圍第1項所述之裝置,其中上述指令是一單一-指令-多重-資料指令。
- 11如申請專利範圍第1項所述之裝置,其中上述指令結合一資料值卸載運算和一算術運算。
- 12如申請專利範圍第1項所述之裝置,其中上述平移電路是上述裝置的資料路徑中上述算術電路的逆運作。
- 13如申請專利範圍第1項所述之裝置,其中該提升電路可用以將上述多位元部分由位元長度A提升為位元長度B,且該提升電路用以和上述平移電路並聯,而當執行上述指令時上述平移電路將有限制範圍的一般平移量當作資料值提供予上述平移電路傳遞,與其相比較的是當執行其他指令時上述平移電路提供的是一般平移量範圍。
- 14一種處理資料的方法,該方法至少包含解碼和執行一產生特定值的指令,該指令執行下列步驟:(i)選擇多個上述資料字組Rm的非相鄰多位元部分用以組成多個位元長度為A的多位元部分;(ii)隨意地以一般平移量平移上述多個多位元部分至平移後的位元位置;(iii)將上述多組多位元部份每個都從上述之A位元長度提升(promote)至B位元長度用以構成多組已提升的多位元部份,如此一來上述已提升的多位元部份可以連接構成一已提升的資料字組P;且(iv)執行多組獨立的算術運算當作輸入運算元,其為B位元長度之個別的位元位置部份由上述已提升之資料字組P和上述的字組Rn構成一結果字組Rd。
- 15一種計算機程式產品,至少包含一計算機程式,用以控制計算機執行如申請專利範圍第14項中所述的方法。
Independent claims15
54 paragraphs, as filed
Data processing device and method and computer program product
<p>2. . . Data path</p><p>4. . . Register group</p><p>6. . . Translation circuit</p><p>8. . . Adder</p><p>10. . . Signed/Zero Extend Mask</p><p>12. . . Multiplexer</p><p>14. . . Data path</p><p>l6. . . Register group</p><p>18. . . Translation circuit</p><p>20. . . Adder</p><p>twenty two. . . Select and combine logic circuits</p>
Figure 1 is a diagram illustrating the actions of the first single instruction multiple data (SIMD) type data processing instruction;
Figure 2 is a diagram illustrating the data path of a processing device that appropriately executes the data processing instructions in Figure 1;
Figures 3 and 4 illustrate two further types of SIMD data processing instructions; and
Figure 5 is a diagram illustrating the data path in a data processing system that appropriately executes the data processing commands in Figures 3 and 4.
Field of invention:
The present invention is related to the field of data processing systems. In more detail, the present invention is related to processing a single instruction multiple data type operations.
Background of the invention:
Single instruction multiple data operation is a known technology. By this technology, the data word group operated with the single finger actually represents multiple data values, and the data word group uses individual data values in the specified operation Do this. In a working data processing system, this type of instruction can increase efficiency and is useful in reducing code size and accelerating processing. This technology is very common, but it is not limited to applications that process data values representing physical signals, such as digital signal processing applications.
When extending the data processing capabilities of a data processing system, an important consideration is any degree of size, complexity, cost, and regular power consumption, which can be used to support additional processing capabilities. When reducing the additional recurring consumption can increase the estimate of processing power is very beneficial.
The purpose and summary of the invention:
From one aspect, the current invention provides a data processing device. The above-mentioned device at least includes a translation circuit, an arithmetic circuit, and an instruction interpreter. The latter is used to control the translation circuit and arithmetic after an instruction is issued. The circuit performs an operation on the data word group Rn and the data word group Rm, where the above operation generates a bit, which is given when selecting multiple sets of non-adjacent multi-bit parts of the data word group Rm , Used to form multiple sets of multi-bit parts of the A-bit length; selectively shift the upper multiple sets of multi-bit parts by the total amount of general translation element positions; combine the above-mentioned multiple sets of multi-bit parts Each part is promoted from the above-mentioned A-bit length to B-bit length to form multiple groups of promoted multi-bit parts, so that the above-mentioned promoted multi-bit parts can be connected to form one The promoted data word group P; and multiple independent arithmetic operations are performed as input operands, which are the individual bit positions of the B-bit length from the promoted data word group P and the above word group Rn constitutes a result word group Rd.
The present invention provides a new data processing command in a data processing system for fetching data values in a data block and also for performing a single command multiple data type arithmetic operation on the fetched data values. The present invention confirms that fetching non-adjacent data values in a data word group can be implemented, and compared with traditional value fetching commands for fetching adjacent data values, there is a relatively small additional recurring consumption. In particular, the need for additional data paths to separate the bit positions of previous adjacent data values can be avoided. Instead, for example, existing masks and block shifting circuits can be used. Furthermore, the simplification of the value function also enables a single instruction to provide arithmetic operations built on operands without causing the processing loop limitation problem.
In general form, the present invention can be applied to the selection of non-adjacent multi-bit parts of any length compared with the increased length. A particularly effective and convenient implementation is the selection of multi-bit parts in actual work. The length of the element part is half of the promoted multi-bit part, and it is matched with the promoted adjacent multi-bit part in the promoted data word to obtain an operation length equal to the length of the input data word .
There are several ways to increase the length of the selected multi-bit part. Two particularly useful methods are to use sign extensions or leading zero extensions.
Arithmetic operations combined with unpacking can take several forms. However, in a particularly preferred embodiment, the arithmetic operations are independently performed odd operations on the individual promoted multi-bit parts. This command is particularly useful in many actual data processing situations, such as the calculation of the sum of absolute differences as part of the MPEG mobile complement calculation.
As mentioned before, the present invention can take many effective ways to utilize the existing processing sources in the data processing system. A special case is that in a system, a translation circuit in the data path provides the inverse operation of the arithmetic circuit. This arrangement allows the unloading to be performed with arbitrary translation before the arithmetic operation.
In a preferred embodiment, in order to provide the required functions without adding additional processing cycle time consumption, a lifting circuit responsible for increasing the length of the selected multi-bit part (for example, using signed extension or leading zero extension) Both and part of the translation circuit are provided, and the general specified translation sum range is limited, so that the first part of the translation circuit and the boost circuit can be combined to perform the required calculations without extending the time data value..... .
Viewed from another aspect, the present invention provides a data processing method. The above method includes at least the steps of decoding and executing an instruction, by selecting a plurality of non-adjacent multi-bit elements of the data word group Rm Part, the instruction will produce a result, which constitutes a multi-bit part of a plurality of bit lengths A, freely translate the above-mentioned multi-bit parts by a general translation amount to the bit position after translation; and the above-mentioned multi-bit part Each of the multi-bit parts is increased from the bit length A to the bit length B to form a plurality of increased multi-bit parts, so the above-mentioned multi-bit parts can be adjacent to form an increased data word group P; and From the above-mentioned promoted data word group P and the above-mentioned data word group Rn, the bit positions of individual length B are used as operands to perform multiple independent arithmetic operations to form a result data word group Rd.
Based on the above-mentioned technology, the present invention also provides a computer program product storing and controlling general-purpose computer programs, for example, including data processing instructions with the above-mentioned arithmetic format.
The above-mentioned features and advantages of the present invention and other objects will be clearly described in conjunction with the following embodiments and related drawings.
Schematic description
Figure 1 is a diagram illustrating the actions of the first single instruction multiple data (SIMD) type data processing instruction;
Figure 2 is a diagram illustrating the data path of a processing device that appropriately executes the data processing instructions in Figure 1;
Figures 3 and 4 illustrate two further types of SIMD data processing instructions; and
Figure 5 is a diagram illustrating the data path in a data processing system that appropriately executes the data processing commands in Figures 3 and 4.
Symbol description of main components
2. . . Data path
4. . . Register group
6. . . Translation circuit
8. . . Adder
10. . . Signed/Zero Extend Mask
12. . . Multiplexer
14. . . Data path
l6. . . Register group
18. . . Translation circuit
20. . . Adder
twenty two. . . Select and combine logic circuits
Detailed description of the invention:
Figure 1 illustrates the operation of the first single instruction multiple data instruction (SIMD) type data processing instruction, called ADD8TO16. This command includes signed and unsigned types, which correspond to the natural extension added to the front end of the input operand data word group selection part, as if it were extended to the length of the part to be processed. The first input operand data word group is stored in the register Rm of the data processing device. The data word group is composed of four octet parts p0, p1, p2, and p3. Depending on whether a rotateright operation is a specific instruction, the multi-bit parts p0 and p2 or the multi-bit parts p1 and p3 will be selected as the input data block in the register. This arbitrary right-turn translation operation can also be calculated on 16 and 24 bits if necessary. This will effectively partially swap the high and low bit order. In the example description in Figure 1, the non-adjacent parts p2 and p2 are selected in the non-rotation (translation) type, and the dashed line refers to other types.
When the multi-bit part has been selected, each is increased from an 8-bit length to a 16-bit length using zero extension or signed extension. The shaded parts of the data block P in the figure that have been promoted indicate these extended parts.
The second input data word group is stored in the register Rn and consists of two 16-bit data values. This example illustrates the execution of a single instruction multi-data addition operation. By this operation, the extended p0 value will be added to the 16-byte group a0 in the low position of Rn, and the extended p2 value will be added to the high of Rn. 16 bytes of position a2. Such addition can be regarded as one of full-width additions, and the full-width addition has a discontinuous carry chain between the 15th and 16th bits in the result value. In this way, other SIMD-type arithmetic operations can also be performed, for example, SIMD subtraction.
The output result data word group produced by the instruction in Figure 1 produces the 16-bit sum of the low position of p0 and a0, and the 16-bit group of the high position contains the sum of p2 and a2. This command is particularly useful when determining the sum of absolute differences between individual data values. By this operation a0 and a2 represent cumulative values and p0 to p3 represent the absolute values of individual significant differences, such as pixel difference values. This type of operation is generally used in MPEG animation estimation processing, and the ability to perform this operation at high speed is very advantageous.
The example in Figure 2 illustrates data path 2 in the data processing system used to implement the instructions in Figure 1. A register group 4 contains 32-bit data words for use. The input operation data word group stored in Rm and Rn is read by the register group and the result data word group is written back to the register Rd in the register group 4. The data path 2 includes a translation circuit 6 and an adder circuit 8. Many other data processing instructions provided by the system tools utilize the translation circuit 6 and the adder circuit 8 in many different ways. The data path 2 is carefully designed so that the time required for the data value to pass through the translation circuit 6 and the addition circuit 8 can be appropriately matched with the data processing cycle time. In a system, the hardware resources are effective for each data block transmitted through data path 2, and the system can effectively use the hardware resources of path 2. A sign/zero extension is connected in parallel with the lower bit part of the masking circuit 10 and the translation circuit. A multiplexer 12 can select the output result of the full translation circuit 6 or the output result of the signed/zero extension and mask circuit 10 as the input of the adder 8. The input of the other adder 8 is the input operand data block of Rn.
When the fingering shown in Figure 1 is executed, the input operand data word group of Rm will be provided to the translation circuit 6. In the translation circuit, the right translation of any 8-bit element will be applied to the data word group, which is the same as in the device Whether there are specific indicators in the relevant. The right rotation of any 16-bit and 24-bit positions can also be applied to this data block. In a translator based on a multi-level multiplexer, such limited possible translation can be relatively simply provided by the first part of the translation circuit 6. (For example, in a 32-bit system, the multiplexer on the first layer can provide 16-bit translation while the multiplexer on the second layer provides eight-bit translation.) Therefore, through the translation circuit 6 in a certain amount The value of the arbitrary translation can be provided to the sign/zero extension and mask circuit 10. The circuit 10 masks the unselected multi-bit part of the input operand data word group that may be translated in Rm, and extends the masked part with the zero or sign of the multi-bit part selected by them. replace. The output result of the sign/zero extension and mask circuit 10 passes through the multiplexer 12 to the first input of the adder 8. The second input of the adder 8 is the input operand data block of Rn. The adder 8 performs a SIMD addition on the input (in other words, two parallel 16 bits are added with a carry circuit. The 15th and 16th bits are actually discontinuous). The output result of the adder 8 will be written to the temporary Register Rd in register group 4.
Another option is that the signed/zero extension and mask circuit 10 can treat Rm (unrotated) as its input and it will execute the four possible signed bits of 0, 8, 16, 24. Rotate and then produce the mask. The translation circuit 6 translates all 32 bits of Rm in parallel.
Figures 3 and 4 illustrate two SIMD instructions packed in half-words. The PKHTB instruction in Figure 3 changes the first half of a fixed-length input operand data block of the register Rn and the variable position half-word length of the second input operand data block of the register Rm The degree part is combined into an output data block, and it is separately stored as the first half and the second half of the register Rd. The instruction PKHBT then combines the fixed-length second half of the input operand data block of the register Rn and the variable position length bit part of the second input operand data block of the register Rm into one output The data word group is stored separately as the second half and the first half of the register Rd. It can be seen that the selected part of the input operand data word group of Rn has not shifted in the output data word group. This allows this part to be provided by a simple mask or selection circuit and represents a very small additional hardware load. The half-word part of the variable position in Figure 3 is selected from bit positions 15 to 0 after Rm has been shifted to the right by k bit positions. Similarly, according to the instruction, the variable-length half-word portion of Rm in Figure 4 is selected from bit positions 31 to 16 after the word has been shifted by k bit positions.
The instructions in Figures 3 and 4 combined with variable-length translation and packing functions are particularly useful for adjusting the Q value change during the operation of the fixed point calculation value.
Figure 5 illustrates the data path 14, which is particularly suitable for executing the commands of Figures 3 and 4. A register set 16 then provides input operand data words, which in this example is 32-bit data, and stores output data words. The data path includes a translation circuit 18, an addition circuit 20, and a selection and combination circuit 22.
In operation, the non-translational input operation element data word group of Rn is transferred from the register group 16 to the selection and combination logic 22. In the command in the example in Figure 3, the most significant 16-bit group of the value in Rn is selected and composes the output data word group in Rd. In the command in the example in Figure 4, the least significant 16-bit group of the input operand data word group in Rn is selected and the least significant bit that constitutes the output data word group in Rd is transferred. The input operation element data word group in Rm is transmitted through the full translation circuit 18. In the instruction in Figure 3, an arithmetic right shift of k-bit position and the lowest-meaning 16-bit group output by the translation circuit 18 is selected by the selection and combination circuit 22, forming the lowest-meaning 16-bit group of the output data word group of Rd Tuple. In the instruction of the example in FIG. 4, the translation circuit 18 provides a left logic translation of k-bit positions and supports the selection and combination of the result of the circuit 22. The selection and combination circuit 22 selects the highest significance 16-byte group of the translation circuit 18 and uses it to form the highest significance 16-byte group in the Rd output data word group.
It can be seen that the selection and combination circuit 22 and the addition circuit 20 are connected in parallel. Therefore, assuming that the data path 14 is carefully designed to perform full translation and addition operations in one processing cycle, the relatively straight forward selection and combination operations can be performed at the normal time of the adder 20 without any processing cycle restrictions. Provided in.
It can be understood that the data processing instructions in the above description and definition in the application have been defined in terms of the achieved result value. It can be seen that the same result value can be obtained in many different processing steps and sequences. The present invention includes all these types that use a single instruction to produce the same result value.
Although the illustrative embodiments of the present invention have been described in detail herein with reference to the corresponding drawings, it should be understood that the correct embodiments are not used to limit the scope of the present invention, and those skilled in the art can comment on the present invention. The above-mentioned embodiments are modified, but these modifications do not deviate from the spirit and scope of the patent applied for by the present invention.
3 sheets
Sheet 1 Sheet 2 Sheet 3
Every citation, both ways
| Document | Relation | Office | Cited during |
|---|---|---|---|
| US9436447B2 | Cited by | United States of America | Applicant |
| TWI498817B | Cited by | Taiwan Province of China | Examiner |
| US10228919B2 | Cited by | United States of America | Applicant |
30 members in 11 offices
Priority claims5
| Document | Office | Kind | Date |
|---|---|---|---|
| 0024311 | United Kingdom | A | |
| 0024311 | United Kingdom | A | |
| 00243113 | United Kingdom | – | |
| 20000024311 | – | – | – |
| GB20000024311 | – | – | – |
Members30
| Document | Office | Kind | |
|---|---|---|---|
| GB0030533D0 | United Kingdom | D0 | |
| GB0122802D0 | United Kingdom | D0 | |
| US2002040378A1 | United States of America | A1 | |
| US2002040427A1 | United States of America | A1 | |
| GB2367650A | United Kingdom | A | |
| GB2367659A | United Kingdom | A | |
| WO0229552A1 | World Intellectual Property Organization (WIPO) | A1 | |
| WO0229553A1 | World Intellectual Property Organization (WIPO) | A1 | |
| JP2002132497A | Japan | A | |
| US2002065860A1 | United States of America | A1 | |
| GB2370893A | United Kingdom | A | |
| IL151395A0 | Israel | A0 | |
| EP1323031A1 | European Patent Office (EPO) | A1 | |
| CN1432151A | China | A | |
| KR20030066631A | Republic of Korea | A | |
| TW548587BThis record | Taiwan Province of China | B | |
| RU2002124769A | Russian Federation | A | |
| JP2004511039A | Japan | A | |
| GB2367650B | United Kingdom | B | |
| GB2370893B | United Kingdom | B | |
| CN1196998C | China | C | |
| JP3723115B2 | Japan | B2 | |
| US6999985B2 | United States of America | B2 | |
| RU2279706C2 | Russian Federation | C2 | |
| MY129332A | Malaysia | A | |
| US7260711B2 | United States of America | B2 | |
| KR100880614B1 | Republic of Korea | B1 | |
| IL151395A | Israel | A | |
| JP5133491B2 | Japan | B2 | |
| EP1323031B1 | European Patent Office (EPO) | B1 |
2 legal events, as the office reported them to INPADOC
Over the term
Point at a mark for the eventEvents
| Event | Code | |
|---|---|---|
| Expiration of patent term of an invention patentMK4A | MK4A | |
| Issue of patent certificate for granted invention patentGrantedGD4A | GD4A |
Numbers
- Publication
- 548587
- Publication, DOCDB
- 548587
- Publication, EPODOC
- TW548587B
- Application
- 90121381
- Application, DOCDB
- 90121381
- Application, EPODOC
- TW200190121381
Titles4
- Chinese
- 資料處理裝置及方法與電腦程式產品
- English
- APPARATUS AND METHOD FOR DATA PROCESSING AND COMPUTERPROGRAM PRODUCT
- Unlabeled
- 資料處理裝置及方法與電腦程式產品
- Unlabeled
- Data processing device and method and computer program product
Classification
- CPC, 9
- G06F7/505
- G06F7/38
- G06F7/49931
- G06F9/30014
- G06F9/30025
- G06F9/30032
- G06F9/30036
- G06F2207/3828
- G06F9/30038
- IPC, 10
- G06F7 00
- G06F7 50
- G06F7 505
- G06F7 74
- G06F7 76
- G06F9 30
- G06F9 302
- G06F9 305
- G06F9 315
- G06F9 38