Simd integer multiply high with round and shift
Abstract
This record has no abstract on file.
Term
Term ended
Expired 22 December 2023, 2.8 years ago.
- Priority
- Filed
- Granted
- Expired
- Today
4 claims: 4 independent, 0 dependent
- 1乗算上位丸めシフト処理を実行するためのコンピュータにより実現される方法であって、 当該方法は、L個のデータ要素の第1セットを有する第1レジスタにおける第1オペランドと、L個のデータ要素の第2セットを有する第2レジスタにおける第2オペランドとを特定する単一命令に応答して、 マイクロプロセッサが、 各ペアが、前記L個のデータ要素の第1セットからの第1データ要素と、前記L個のデータ要素の第2セットの対応するデータ要素位置からの第2データ要素とを有するL個のデータ要素ペアを掛け合わせ、L個の積のセットを生成するステップと、 前記L個の積のそれぞれを右に14ビットシフトし、L個のシフトされた値を18ビット長となるように生成するステップと、前記L個のシフトされた値のそれぞれの最下位ビット位置に“1”を付加することによって、前記L個のシフトされた値のそれぞれを丸め処理し、L個の丸められた値を生成するステップと、 前記L個の丸められた値のそれぞれを右に1ビットだけスケーリングし、L個のスケーリングされた値のセットを生成するステップと、L個の切り捨てられた値を取得するため、前記L個のスケーリングされた値から最下位の16ビットを選択することによって、前記L個のスケーリングされた値のそれぞれを切り捨て処理し、L個の切り捨てられた値を生成するステップと、 前記単一命令の最終結果として、前記L個の切り捨てられた値を前記単一命令により示される宛先レジスタに格納するステップと、 を実行することによって前記単一命令を実行することからなり、 各切り捨て処理された値は、それのデータ要素のペアに対応するデータ要素位置に格納されることを特徴とする方法。
- 2単一命令を受け付け、該単一命令に応答して、マイクロプロセッサのハードウェア実行ユニットに2つのオペランドに対してPacked乗算上位丸めシフト処理を実行させるステップと、前記マイクロプロセッサのハードウェア実行ユニットにおいて前記単一命令を実行し、切り捨て処理された結果のセットを生成するステップと、 Packedデータ要素として宛先レジスタに前記切り捨て処理された結果のセットを格納するステップと、 から構成される方法であって、 前記Packed乗算上位丸めシフト処理は、 Packedデータ要素の第1セットの各データ要素と、Packedデータ要素の第2セットの対応するデータ要素とを乗算し、積のセットを生成し、 前記積のセットのそれぞれを右に14ビットシフトし、その後に丸め処理して、18ビット長となるように結果のセットを生成し、 前記結果のそれぞれから複数のビットを選択し、切り捨て処理された結果のセットを生成することから構成され、 前記単一命令は、 前記Packed乗算上位丸めシフト処理に関する情報を提供するため、前記Packed乗算上位丸めシフト処理に対する前記切り捨てられた結果のセットが、前記結果のセットの上位ビット又は下位ビットから構成されるか示すオペコードを指定する第1フィールドと、 前記Packedデータ要素の第1セットを有する第1オペランドに対して、第1ソースアドレスを指定する第2フィールドと、 前記Packedデータ要素の第2セットを有する第2オペランドに対して、第2ソースアドレスを指定する第3フィールドと、 から構成されるフォーマットを有することを特徴とする方法。
- 3単一命令に応答してPacked乗算丸めシフト処理を実行するマイクロプロセッサのハードウェア実行ユニットから構成される装置であって、 前記ハードウェア実行ユニットは、前記単一命令に応答して、 Packedデータ要素の第1セットの各データ要素と、Packedデータ要素の第2セットの対応するデータ要素とを乗算し、積のセットを生成し、シフトされた値のそれぞれの最下位ビット位置に“1”を付加することによって、前記積のセットのそれぞれを丸め処理し、結果のセットを生成し、 前記結果のセットのそれぞれを右に14ビットシフトし、18ビット長となるように結果の中間セットを生成し、 前記結果の中間セットのそれぞれから複数のビットを選択し、切り捨てられた結果のセットを生成し、 最終結果として前記切り捨てられた結果のセットを格納し、 前記単一命令は、 前記Packed乗算丸めシフト処理に関する情報を提供するため、前記Packed乗算上位丸めシフト処理に対する前記切り捨てられた結果のセットが、前記結果のセットの上位ビット又は下位ビットから構成されるか示すオペコードを指定する第1フィールドと、 前記Packedデータ要素の第1セットを有する第1オペランドに対して、第1ソースアドレスを指定する第2フィールドと、 前記Packedデータ要素の第2セットを有する第2オペランドに対して、第2ソースアドレスを指定する第3フィールドと、 から構成されるフォーマットを有することを特徴とする装置。
- 4第1命令を格納するメモリと、前記メモリから前記第1命令をフェッチするプロセッサと、 から構成されるシステムであって、 前記プロセッサは、前記第1命令の実行に応答して、 Packedデータ要素の第1セットの各データ要素と、Packedデータ要素の第2セットの対応するデータ要素とを乗算し、積のセットを生成し、シフトされた値のそれぞれの最下位ビット位置に“1”を付加することによって、前記積のセットのそれぞれを丸め処理し、一時的結果のセットを生成し、 前記一時的結果のセットのそれぞれをスケーリングし、スケーリングされた一時的結果のセットを生成し、 前記スケーリングされた一時的結果のそれぞれから複数のビットを選択し、切り捨て処理された結果のセットを生成し、 最終結果として前記切り捨て処理された結果のセットを格納し、 前記第1命令は、 前記Packed乗算丸めシフト処理に関する情報であって、符号付き整数のPacked乗算丸めシフト処理を示す情報を提供するオペコードであって、前記切り捨てられた結果のセットのそれぞれの上位ビットを選択するためのオペコードを指定する第1フィールドと、 前記Packedデータ要素の第1セットを有する第1オペランドに対して、第1ソースアドレスを指定する第2フィールドと、 前記Packedデータ要素の第2セットを有する第2オペランドに対して、第2ソースアドレスを指定する第3フィールドと、 から構成されるフォーマットを有することを特徴とするシステム。
Independent claims4
90 paragraphs, as filed
The present disclosure relates to the technical fields of processing equipment, related software and software sequences that perform mathematical operations.
Computer systems are becoming more and more widespread in today's society. Computer processing power has increased the efficiency and productivity of workers in a wide range of fields. As the cost of purchasing and owning computers diminished, more consumers were able to take advantage of the latest, faster machines. In addition, many people enjoy the use of notebook computers due to their flexibility in their use. Mobile computers make it easy for users to carry and work with their data when they are out of the office or on the go. Such situations are common among marketing staff, corporate executives, and even students.
As processor technology advances, new software code is generated to run on machines with advanced processors. In general, users expect and demand higher performance from their computers, regardless of the type of software they are using . Here, there is one problem that can arise from the instructions and processing types being executed inside the processor. That is, some types of processing may take a lot of time to complete, depending on the complexity of the processing and / or the type of circuit required. This motivates us to optimize the way we perform complex processing inside the processor.
Media applications have driven the development of microprocessors for decades. In fact, many of the computer performance improvements in recent years have been facilitated by media applications. Significant advances have been made in the corporate sector for more entertaining educational and communication purposes, but the performance improvements described above have occurred primarily in the consumer sector. Nevertheless, future media applications will require even higher computing power. As a result, future personal computers (PCs) will not only be easy to use, but will also have more audiovisual features. More importantly, it will be the fusion of computer and communication.
Therefore, in the current computer, not only the reproduction of audio and video data collectively as contents but also the display of images is becoming an increasingly common application. Filtering and convolution processing are the most commonly performed processing on content data such as image, audio and video data. Since these processes require a large amount of calculation, high-level data for efficient execution by utilizing various data storage devices such as single-instruction multiplex data (SIMD) registers. Parallel processing is provided.
<p> Many existing architectures require unnecessary data type changes, which reduces instruction throughput and significantly increases the number of clock cycles required for data ordering for arithmetic operations.</p><p> In view of these problems, it is an object of the present invention to provide a readable medium by a method, an apparatus, a system and a machine for performing a packed multiplication upper rounding shift process.</p>
<p> In order to solve the above problems, the method according to the present invention receives a step of receiving a first operand having a first set of L data elements and a second operand having a second set of L data elements. L pieces each having a first data element from the first set of L data elements and a second data element from the corresponding data element position of the second set of L data elements. A step of multiplying a pair of data elements of the above to generate a set of L products, a step of rounding each of the L products to generate L rounded values, and a step of generating the L rounded values. Each of the rounded values is scaled to generate L scaled values, and each of the L scaled values corresponds to a pair of its data elements for storage at the destination. It is characterized by consisting of steps of truncation processing so that it is stored at the position of the data element to be processed.</p><p> In order to solve the above problems, the method according to the present invention includes a step of multiplying each data element of the first set of packed data elements and the corresponding data element of the second set of packed data elements to generate a set of products. Packed Multiply Rounding, which consists of a step of rounding and shifting each of the sets of products to generate a set of results, and a step of selecting multiple bits from each of the sets of results and generating a set of truncated results. A method consisting of a step of receiving an instruction to execute shift processing on two operands and a step of executing the instruction and generating the truncated set of results to be stored in the destination register as a packed data element. , The instruction specifies a first field that identifies an operating code that provides information about the Packed Multiply Round Shift process, a second that identifies a first source address for a first operand that has a first set of the Packed data elements. It is characterized by having a format consisting of a field, a third field that specifies a second source address for a second operand having a second set of the Packed data elements.</p><p> In order to solve the above problems, the device according to the present invention is a device consisting of an execution unit that executes one or more instructions of an instruction set including at least one instruction for executing a packed multiplication rounding shift process. The execution unit multiplies each data element in the first set of packed data elements with the corresponding data element in the second set of packed data elements in response to the at least one instruction that executes the packed multiplication rounding shift process. Combine, generate a set of products, round and shift each of the sets of products, generate a set of results, select multiple bits from each of the results, generate a set of truncated results, The at least one instruction is a first field for specifying an opcode for providing information about the Packed Multiply Round Shift process and a first source address for a first operand having a first set of the Packed data elements. It is characterized by having a format consisting of a second field for specifying a second field for specifying a second source address for a second operand having a second set of the Packed data elements.</p><p> In order to solve the above problems, the system according to the present invention comprises a memory for storing data and instructions and a processor connected to the memory via a bus and executing a multiplication round shift process in response to a multiplication round shift instruction. The processor is composed of a bus unit that receives the multiplication rounding shift instruction from the memory and an execution unit that is connected to the bus unit and executes the multiplication rounding shift instruction. The instruction causes the execution unit to generate a set of products by multiplying each data element in the first set of packed data elements with the corresponding data element in the second set of packed data elements, respectively. Is rounded and shifted to generate a set of results, and by selecting a plurality of bits from each of the results, a set of truncated results is generated.</p><p> In order to solve the above problems, the machine-readable medium according to the present invention is a machine-readable medium that stores a program, and the program that can be executed by the machine contains a first set of L data elements. The step of receiving the first operand having, the step of receiving the second operand having the second set of L data elements, and the first data element and L from the first set of L data elements, respectively. A step of multiplying L pairs having a second data element from the corresponding data element position of the second set of data elements to generate a set of L products, and a product of the L products. A step of rounding each to generate L rounded values, a step of scaling each of the L rounded values to generate L scaled values, and storage at the destination. Therefore, it is characterized in that a method including a step of truncating each of the L scaled values so as to be stored at the data element position corresponding to the pair of the data elements is executed.</p>
<p> As described above, according to the present invention, it is possible to obtain a readable medium by a method, an apparatus, a system and a machine for executing the Packed multiplication upper rounding shift process.</p>
Hereinafter, embodiments of the present invention will be described with reference to the drawings. Here, the present invention is not limited to the accompanying drawings. Also, in the drawings, the same reference symbols indicate the same elements.
The following description describes an example of SIMD integer multiplication rounding shift processing. In the following description, specific details such as processor type, microarchitecture state, events, feasible mechanisms, etc. are given to provide a more complete understanding of the invention. However, those skilled in the art will recognize that the present invention can be practiced beyond such specific details. Moreover, well-known configurations, circuits, etc. are not shown in detail so as not to unnecessarily obscure the present invention.
The following examples are described with respect to processors, but other embodiments can also be applied to other types of integrated circuits and logic devices. Similar techniques and teachings of the present invention can be readily applied to other types of circuits or semiconductor devices that can enjoy higher pipeline throughput and performance. The teachings of the present invention are applicable to any processor or machine that performs data manipulation. However, the invention is not limited to processors or machines that perform 256-bit, 128-bit, 64-bit, 32-bit or 16-bit data processing, but any processor or machine that requires manipulation of packed data. Can be applied to.
In the following description, various specific details are given for illustration purposes to provide a complete understanding of the present invention. Those skilled in the art will recognize that these specific details are not always necessary for the practice of the present invention. Also, well-known electrical structures and circuits are not given in detail so as not to unnecessarily obscure the present invention. Furthermore, the following description gives examples, and the accompanying drawings show various examples for illustration purposes. However, these examples should not be construed as limiting. These examples do not comprehensively list all possible realizations of the invention, but are merely intended to provide an example of the invention.
The following examples describe instruction processing and arrangement of execution units and logic circuits, but other examples of the present invention can be achieved by software. In one embodiment, the method according to the invention is implemented by machine-executable instructions. These instructions cause a programmable general purpose or application processor to perform each step of the invention. The present invention is provided as computer program products or software that include a machine or computer-readable medium that contains instructions used to program a computer (or other electronic device) that performs processing according to the invention. Alternatively, each step of the invention may be performed by specific hardware elements, including wiring logic that performs these steps, or by any combination of programmed computer and custom hardware components. It may be executed. Such software can be stored in memory within the system. Similarly, such code can be distributed over the network or via other computer-readable media.
Therefore, the medium that can be read by the machine is not limited to the following, but is a floppy disk (registered trademark), an optical disk, a CD (Compact Disc), a CD-ROM (CD Read-Only Memory), an optical magnetic disk, and a ROM. (Read-Only Memory), RAM (Random Access Memory), EPROM (Erasable Programmable Read-Only Memory), EEPROM (Electrically Erasable Programmable Read-Only) Readable by machines (eg computers) such as Memory), magnetic or optical cards, flash memory, transmission over the Internet, electronic, optical, acoustic or other carrier signals (eg carrier waves, infrared signals, digital signals, etc.) Any mechanism for storing and transmitting information in the form is included. Thus, computer readable media include any type of media / machine readable medium suitable for storing or transmitting electronic instructions or information in a form readable by a machine (eg, a computer). In addition, the invention may also be available for download as computer program products. In that case, the program is transferred from the remote computer (eg, the server) to the requesting computer (eg, the client). The transfer of the program may be carried out by other forms of data signals realized by electronic, optical, acoustic or carrier waves, or by other propagation media over communication links (eg, modems, network connections, etc.).
Design may go through various stages, from production to simulation to manufacturing. Data representing a design may represent the design in various ways. First, the hardware is represented using a hardware description language and other functional description languages, which is useful in simulations. In addition, logic and / or transistor gate circuit level models are generated at some stage in the design process. Moreover, at some stage, most designs reach data levels that represent the physical arrangement of various devices in the hardware model. When conventional semiconductor manufacturing techniques are used, the data representing the hardware model may be data that identifies the presence or absence of various features of the mask layer of the mask used to generate the integrated circuit. In any design representation, this data can be stored in any form of machine readable medium. Optical or electrical waves, memories, magnetic or optical storage devices such as disks that are generated or modulated for the transmission of such information are machine readable media. Any of these media can "carry" or "express" design and software information. When an electrical carrier that indicates and carries a code or design is transmitted, a new copy is made to the extent that the electrical signal is copied, buffered, or retransmitted. Therefore, a communication provider or a network provider can copy an object (carrier wave) that realizes the technique of the present invention.
Today's processors utilize many execution units to process and execute various codes and instructions. While some instructions complete immediately, some instructions require enormous clock cycles, so not all instructions are generated equally. The faster the instruction throughput, the better the overall performance of the processor. Therefore, it is desirable to execute as many instructions as possible at high speed. However, some instructions have greater complexity and require more execution time and processor resources. For example, floating point instructions, load / store processing, data transfer, etc.
As more and more computer systems are used in the Internet and multimedia applications, additional processor support has been introduced in the past. For example, a single instruction multiple data (SIMD) integer / floating point instruction or a streaming SIMD extension (SSE) is an instruction that reduces the total number of instructions required to perform a particular program task. These instructions enable high-speed software performance by performing parallel processing on multiple data elements. This allows performance improvements to be achieved in a wide range of applications, including video, audio, and image / photo processing. Realization of SIMD instructions in microprocessors and similar logic circuits usually involves many issuances. Moreover, the complexity of SIMD processing often creates the need for additional circuitry for accurate data processing and manipulation.
Two's-complement notation is an effective way to represent signed numbers. The most significant bit of 2's complement represents its sign and the remaining bits represent its magnitude. Fixed-point number calculations allow multiplication in integer processors without causing overflow. Decimal calculations are of great benefit to digital signal processing programming in the absence of multiplication overflow problems. Multiplying two 16-bit numbers requires 32 bits for the result, and the 32-bit result produced by multiplying two 16-bit fixed-point numbers is rounded to 16 bits by introducing a minimum error. The conversion of a 16-bit integer is to divide the decimal value of the integer by 32768. In one embodiment, attention is paid to the upper 16 bits of the product produced by multiplying two decimals. However, the upper 16 bits of the result are half of the expected decimal result. This product needs to be shifted to the left to multiply the result by 2. This gives the final correct product. Decimal operations also require sign extensions for multipliers and multiplicands.
The need for a shift to the left can also be explained as an arrangement of decimal positions. For example, when multiplying decimals, the decimal point is ignored and placed at the end. The decimal point is arranged so that the total number of digits to the right of the decimal point of the multiplier and multiplicand is equal to the number of digits to the right of the decimal point of their product. Similarly, the "decimal point" here for decimal arithmetic is located to the right of the leftmost (sign) bit, and to the right of this point is 15 bits (digits). However, there are a total of 30 bits to the right of the decimal point in the source. Without the shift, the 32-bit result would have 31 bits to the right of the decimal point. By shifting the number to the left by one bit, the number of bits to the right of the decimal point can be effectively reduced to 30.
The embodiments of the present invention can improve the accuracy of fixed point integer SIMD instructions. The fixed point integer format is similar to that of fixed point decimal arithmetic. The fixed point format of "1.15" in one embodiment represents a number whose binary point has a signed value located between the 14th and 15th bits. Here, the bit position is counted from 0 from the rightmost bit. Therefore, the rightmost or least significant bit is in the 0th position. The bit position immediately to the left is the first bit, and so on. This 1.N numeric format is often used in digital signal processing (DSP) applications. The embodiments according to the invention also provide further improvements in accuracy through rounding and shift processing techniques. Further improvements in accuracy obtained from the embodiments of the present invention contribute to easier programming of many applications. In addition, this further improvement in accuracy allows for faster execution of algorithms such as the Discrete Cosine Transform (DCT), which are often used in video and image processing applications.
An example application for SIMD integer multiply high with round and shift instruction is in high quality video. 16x16-bit multiplication with 16-bit results is very popular in video encoders and decoders, especially in inverse DCT, DCT, quantization (Q) and inverse Q blocks. The accuracy of multiplication has a great effect on the overall image quality. The performance improvement and speedup according to the embodiment of the present invention have a greater influence on the inverse DCT calculation. In addition to DCT calculations, it is also useful for Q and inverse Q calculations, which are basically 16-bit multiplications.
Generally, in the computer industry, the IEEE standard 1180-1990 is often used to realize the 8 × 8 inverse discrete cosine transform. Although this standard relates to video conferencing, some of the standards apply to encoders and decoders in various MPEG formats. However, it is difficult to comply with the IEEE 1180-1990 standard while maintaining high performance. This trade-off often results in non-compliant high performance or compliant low performance. Moreover, coding into a standard is a time-consuming and iterative process, especially if an inadequate algorithm is selected.
Compliance with the IEEE 1180-1990 standard is facilitated by the implementation of the multiplier high rounding shift instruction. An embodiment of the SIMD integer multiplication upper rounding shift instruction according to the present invention can provide the same 1.15 data format for input / output data elements in a packed data environment. This simplifies code writing and programming with an instruction set that includes an example of multiplication upper rounding shift processing. Similarly, the accessibility of compilers associated with high-level languages is also possible. To improve the performance and accuracy of video, audio and image encoders / decoders (codecs), developers may utilize languages and compilers made possible by examples of fixed point SIMD instructions such as integer multiplication rounding shifts. it can. An instruction set with SIMD capabilities helps avoid the previously required redundant algorithms in iterating similar data.
Each input to multiplication in one implementation follows the 1.15 format. In one embodiment of multiplication-upward rounding shift processing with memory, a tentative 18-bit value with 2.16 format is generated from the high-order bits of a 32-bit product by multiplying two 16-bit data values. This tentative 18-bit value is rounded for precision by adding "1" to the least significant bit. Although some techniques simply discard all the low-order bits, the rounding process according to the embodiments of the present invention allows the error to fall within some acceptable threshold for inverse DCT coding. This rounded value is shifted left by one bit for further precision and to obtain the desired output format. A 16-bit result with the 1.15 format is extracted from the rounded, shifted 18-bit value. The rounding and shifting operations performed on the tentative values can provide 2 bits with more precision by simply taking the upper 16 bits of the 32-bit product. For example, in the general embodiment described herein, the rounding process provides 1 bit with additional precision by extracting the upper 16 bits from the product of 32 bits. Similarly, the shift process provides one bit with additional precision for the rounded product. These descriptions describe examples for 16-bit length integer values, but other examples can be applied to data values of any bit length.
FIG. 1A is a block diagram of a computer system as an example composed of a processor including an execution unit that executes an instruction for multiplication upper rounding shift processing according to an embodiment of the present invention. System 100 includes components such as processor 102 that utilize an execution unit that includes logic to execute the data processing algorithms according to the invention as in the embodiments described herein. System 100 is based on PENTIUM III, PENTIUM 4, Xeon, Itanium and / or XScale microprocessors available from Intel Corporation in Santa Clara, California. It is represented by a processing system. However, other systems (including other microprocessors, engineering workstations, set-top boxes, etc.) may be utilized. In one embodiment, sample system 100 may run a version of the WINDOWS® operating system available from Microsoft Corporation in Redmond, Washington. However, other operating systems (eg, UNIX® and Linux), embedded software and / or graphical user interfaces may be used. The present invention is not limited to a particular combination of hardware circuits and software.
In other embodiments of the invention, it is available in other devices such as portable devices and embedded applications. Examples of portable devices include mobile phones, Internet protocol devices, digital cameras, PDAs (Personal Digital Assistants), portable personal computers, and the like. Embedded applications include microcontrollers, digital signal processors (DSPs), system-on-chips, network computers (NetPCs), set-top boxes, network hubs, wide area network (WAN) switches, or misaligned memory copies or moves. Other systems etc. are included. Furthermore, in order to improve the efficiency of multimedia applications, there is an architecture realized to enable instructions to process multiple data at the same time. As data types and volumes increase, the performance of computers and their processors must be improved so that they can be manipulated in more efficient ways.
FIG. 1A is a block diagram of a computer system 100 composed of a processor 102 with one or more execution units 108 processing algorithms including SIMD integer multiplication superordinates with rounding and shift instructions according to the present invention. For example, processor 102 can receive a program instruction requesting SIMD multiplication superposition processing for the Packed data operand. This embodiment is described for a single processor desktop or server system, but other embodiments may be included in a multiprocessor system. System 100 is an example of a hub architecture. The computer system 100 includes a processor 102 that processes a data signal. The processor 102 includes a complex instruction set computer (CISC) microprocessor, a reduced instruction set computer (RISC) microprocessor, and VLIW (Very Long Instruction). Word) It can be a macro processor, a processor that implements a combination of instruction sets, or other processor equipment such as a digital signal processor. The processor 102 is connected to a processor bus 110 capable of transmitting data signals between the processor 102 and other components in the system 100. The components of System 100 perform their existing functions well known to those skilled in the art.
In one embodiment, processor 102 includes Level 1 (L1) internal cache memory 104. Depending on the architecture, processor 102 has a single internal cache or multiple levels of internal cache. In another embodiment, the cache memory may be provided outside the processor 102. Other embodiments may also include a combination of both an internal cache and an external cache, depending on the required embodiment. The register file 106 can store various types of data in various registers, including integer registers, floating point registers, status registers and instruction pointer registers.
Execution unit 108 includes logic for performing integer and floating point processing and is provided in processor 102. Processor 102 may also include a microcode ROM that stores microcode for certain macro instructions. In this embodiment, the execution unit 108 includes logic for handling the packed instruction set 109. In one embodiment, the Packed instruction set 109 includes a Packed multiplication super-instruction to acquire the relevant superordinate portion of the resulting product. By including the Packed Instruction Set 109 of the General Purpose Processor 102 Instruction Set along with the associated circuits that execute the instructions, even if the processing utilized by many multimedia applications is performed by using the General Purpose Processor 102 Packed Data. Good. This allows many multimedia applications to run more efficiently by using the full width of the processor's data bus to perform processing on the packed data. This eliminates the need to send smaller data units to the processor's data bus to perform one or more operations on a data element at the same time. Other embodiments of Execution Unit 108 are also available in microcontrollers, embedded processors, graphics devices, DSPs and other types of logic circuits. System 100 includes memory 120. The memory 120 may be a DRAM (Dynamic Random Access Memory) device, a SRAM (Static Random Access Memory) device, a flash memory device, or another memory device. The memory 120 can store instructions and / or data represented by a data signal that can be executed by the processor 102.
The system logic chip 116 is connected to the processor bus 110 and the memory 120. The system logic chip 116 of the illustrated embodiment is a memory controller hub (MCH). The processor 102 can communicate with the MCH 116 via the processor bus 110. MCH116 provides a high bandwidth memory path 118 to memory 120 for storing instructions and data and for storing graphics commands, data and textures. The MCH 116 guides the data signal between the processor 102, the memory 120 and the other components of the system 100 and bridges the data signal between the processor bus 110, the memory 120 and the system I / O 122. In some embodiments, the system logic chip 116 may include a graphics port for connection to the graphics controller 112. The MCH 116 is connected to the memory 120 via the memory interface 118. The graphics card 112 is connected to the MCH 116 via the AGP (Accelerated Graphics Port) interconnect 114.
System 100 uses a dedicated hub interface bus 122 to connect the MCH 116 to the I / O controller hub (ICH) 130. The ICH130 provides a direct connection to some I / O devices via the local I / O bus. The local I / O bus is a high-speed I / O bus that connects peripherals to memory 120, a chipset, and a processor 102. Some examples include voice controllers, firmware hubs (flash BIOS) 128, wireless transmitters 126, data storage devices 124, existing I / O controllers including user input and keyboard interfaces, and USB (Universal Serial Bus). The serial expansion port and the network controller 134. The data storage device 124 may consist of a hard disk drive, a floppy disk® drive, a CD-ROM device, a flash memory device, or other mass storage device.
In other embodiments of the system, the execution unit that executes the Packed Multiplication Superinstruction can be used with the system on chip. An embodiment of a system-on-chip comprises a processor and memory. The memory for such a system is flash memory. The flash memory is provided on the same chip as the processor and other system components. In addition, other logical blocks such as memory controllers and graphics controllers can also be placed on the system on chip.
FIG. 1B shows another embodiment of the data processing system 140 that implements the principles of the present invention. An embodiment of the data processing system 140 is an Intel® Personal Internet Client Architecture (PCA) application processor with Intel XScale® technology (as described at "www.intel.com"). It will be appreciated by those skilled in the art that the embodiments described herein can be used with other processing systems without departing from the scope of the invention.
The computer system 140 includes a processing core 159 capable of performing SIMD processing including multiplication upper rounding shift. In one embodiment, processing core 159 represents a processing unit of any type of architecture, not limited to CISC, RISC, VLIW type architectures. The processing core 159 may also be suitable for manufacture in one or more processing techniques, and by being represented in sufficient detail on machine readable media, the processing core 159 is suitable for facilitating this manufacture. It may be a product.
The processing core 159 is composed of an execution unit 142, a register file set 145, and a decoder 144. The processing core 159 may also include additional circuitry (not shown) that is not necessary for the understanding of the present invention. The execution unit 142 is used to execute the instruction received by the processing core 159. In addition to recognizing typical processor instructions, Execution Unit 142 can recognize the instructions in the Packed Instruction Set 143 for performing processing on the Packed data format. Packed instruction set 143 includes instructions that support data merging, and may further include other packed instructions. Execution unit 142 is connected to register file 145 by the internal bus. The register file 145 represents a storage area in the processing core 159 for storing information including data. As mentioned above, it will be understood that it does not matter which storage area is used to store the Packed data. The execution unit 142 is connected to the decoder 144. The decoder 144 is used to decode the instructions received by the processing core 159 to the control signal and / or the microcode input point. In response to these control signals and / or microcode input points, execution unit 142 performs appropriate processing.
The processing core 159 is not limited to the following, but is, for example, an SDRAM (Synchronous Dynamic Random Access Memory) control 146, a SRAM (Static Random Access Memory) control 147, a burst flash memory interface 148, and a PCMCIA (Personal Computer Memory Card). On bus 141 for communicating with various other system units including International Association) / CF (Compact Flash) card control 149, LCD (LCD) control 150, DMA (Direct Memory Access) controller 151, and alternative bus master interface 152. Be connected. In one embodiment, the data processing system 140 also includes an I / O bridge 154 for communicating with various I / O devices via the I / O bus 153. Such an I / O device is not limited to the following, but is, for example, UART (Universal Asynchronous). It may consist of Receiver / Transmitter) 155, USB156, Bluetooth wireless UART157, and I / O expansion interface 158.
One embodiment of the data processing system 140 comprises a processing core 159 capable of performing SIMD processing, including shift merging processing, for mobile, network and / or wireless communication. The processing core 159 includes discrete transforms such as Walshuadamard transform, fast Fourier transform (FFT), discrete cosine transform (DCT) and their respective inverse transforms, as well as color space transform, video coding motion prediction or video decoding motion prediction. It may be programmed by various audio, video, image forming and communication algorithms, including compression / decompression techniques and modulation / demodulation (MODEM) functions such as pulse code modulation (PCM).
FIG. 1C shows another embodiment of a data processing system capable of performing SIMD multiplication superposition processing. According to another embodiment, the data processing system 160 includes a main processor 166, a SIMD coprocessor 161, a cache memory 167 and an input / output system 168. The input / output system 168 may be selectively connected to the wireless interface 169. The SIMD coprocessor 161 can perform SIMD processing including multiplication superposition. The processing core 170 may be suitable for manufacturing in one or more processing techniques, and by representing in sufficient detail on machine readable media, the processing core 170 is all of the data processing system 160 including it. Alternatively, it may be suitable for facilitating a part of production.
In one embodiment, the SIMD coprocessor 161 is composed of an execution unit 162 and a register file set 164. One embodiment of the main processor 165 includes a decoder 165 that recognizes instructions in instruction set 163, including SIMD Packed multiplication superordinate instructions, for execution by execution unit 162. In another embodiment, the SIMD coprocessor 161 also comprises at least a portion of a decoder 165B that decodes the instructions in instruction set 163. The processing core 170 also includes additional circuits (not shown) that are not necessary for the understanding of the present invention.
During operation, the main processor 166 executes a stream of data processing instructions that control common types of data processing operations, including the interaction of the cache memory 167 with the I / O system 168. SIMD coprocessor instructions are embedded in a stream of data processing instructions. The decoder 165 of the main processor 166 issues these SIMD coprocessor instructions as the type to be executed by the mounted SIMD coprocessor 161. Therefore, the main processor 166 issues these SIMD coprocessor instructions (or control signals representing SIMD coprocessor instructions) on the coprocessor bus 166, and is received by the SIMD coprocessor mounted from the coprocessor bus 166. In this case, the SIMD coprocessor 161 receives and executes the received SIMD coprocessor instruction.
Data may be received via wireless interface 169 for processing by SIMD coprocessor instructions. As an example, voice communication may be received in the form of a digital signal and processed by SIMD coprocessor instructions to regenerate a digital voice sample representing the voice signal. As another example, compressed audio and / or video may be received in digital bitstream format and processed by SIMD coprocessor instructions to regenerate digital audio samples and / or motion video frames. In one embodiment, the processing core 170, the main processor 166 and the SIMD coprocessor 161 are integrated into a single processing core 170 consisting of an execution unit 162, a register file set 164 and a decoder 165, and instructions including SIMD multiplication superordinate instructions. Recognize the instructions in set 163.
FIG. 2 is a block diagram of a microarchitecture for a processor 200 of an embodiment having a logic circuit that performs a Packed Integer Multiply Upper Rounding Shift process according to the present invention. SIMD integer multiplication upper processing by rounding and shifting processing may also be referred to as Packed multiplication upper rounding shift processing (PMUL upper processing), or multiplication upper processing. In one embodiment of the Packed Multiplier instruction, the instruction extracts data from two memory blocks, multiplies the corresponding data elements from each block to obtain a tentative result, and rounds and shifts this tentative result. , Truncate this intermediate result to the desired higher part of each product for storage in the resulting merged data block. SIMD multiplication super-instructions are also called PMULHRSW or Packed multiplication super-rounding shifts. In this embodiment, the merge process is also performed to perform processing on data elements having sizes such as bytes, words, double words, and quad words. Although the description here relates to integers and integer processing, other embodiments of the present invention may be used in floating point numbers and floating point processing.
The in-order front end 201 is part of a processor 200 that fetches a macro instruction to be executed and prepares the instruction for later use in the processor pipeline. The front end 201 of this embodiment includes a plurality of units. The instruction prefetcher 226 fetches macroinstructions from memory, supplies them to the instruction decoder 228, and decodes them into elements called microinstructions or microprocessing (or microops or uops) that can be executed by the machine. The trace cache 230 receives the decrypted uops and decomposes them into ordered program sequences or traces in the uop queue 234 for execution. When the trace cache 230 faces a complex macro instruction, the microcode ROM 232 provides the uop needed to complete the process.
Many macro instructions are converted into one micro op, while other macro instructions require multiple micro ops to complete the process. In one embodiment, if more than 4 micro ops are needed to complete the macro instruction, the decoder 228 accesses the microcode ROM 232 and executes the macro instruction. In one embodiment, the multiplication upper rounding shift instruction is decoded into a small number of micro ops for processing by the instruction decoder 228. In another embodiment, the instructions for the Packed Multiply Top Rounding Shift Algorithm are stored in the Micro ROM 232 if a large number of micro ops are required to complete the process. The trace cache 230 refers to the input point PLA (Programmable Logic Array) to determine the correct microinstruction pointer to read the microcode sequence for the merge algorithm in the microcode ROM232. When the microcode ROM232 completes the ordering of the microops for the current macroinstruction, the machine frontend 201 resumes fetching the microops from the trace cache 230.
Some SIMD and other multimedia type instructions are considered complex instructions. Most floating-point instructions are also complex. Further, when the instruction decoder 228 faces a complex macro instruction, the microcode ROM 232 is accessed at an appropriate position to extract the microcode sequence for the macro instruction. The various micro ops required to execute this macro instruction are communicated to the out-of-order execution engine 203 for execution in the appropriate integer and floating point execution units.
The out-of-order execution engine 203 provides microinstructions for execution. Out-of-order execution logic has multiple buffers for smoothing and ordering the flow of microinstructions to optimize performance as they enter the pipeline and are scheduled for execution. .. Allocation or allocator logic allocates machine buffers and resources that each uop needs to run. Register rename logic renames the logical register at the input of the register file. Allocation logic also goes to one of two uop queues for memory and non-memory processing before the instruction schedulers of memory scheduler, fast scheduler 202, slow / normal floating point scheduler 204, and simple floating point scheduler 206. Assign an input for each uop in. The uop schedulers 202, 204, and 206 determine when uop is ready to run, based on the scheduler's dependent input register operand source readiness and the availability of execution resources that uop needs to perform its processing. .. The high-speed scheduler 202 of this embodiment schedules every half cycle of the main clock cycle, while the other schedulers can schedule only once per main processor clock cycle. The scheduler arbitrates the dispatch port and schedules uop for execution.
The register files 208 and 210 are arranged between the schedulers 202, 204 and 206 and the execution units 212, 214, 216, 218, 220, 222 and 224 of execution block 211. There are register files 208 and 210 for integer and floating point operations, respectively. Each of the register files 208 and 210 of this embodiment also includes a bypass network that bypasses or transfers termination results that have not yet been written to the register file to a new dependent uop. The integer register file 208 and the floating point register file 210 can also communicate data with each other. In one embodiment, the integer register file 208 is divided into two register files, one for lower 32-bit data and the other for upper 32-bit data. In one example floating point register file, 210 has a 128-bit wide input. This is because floating point instructions typically have operands that are 64 to 128 bits wide.
Execution block 211 includes execution units 212, 214, 216, 218, 220, 222 and 224 that actually execute the instruction. This part contains register files 208 and 210 that store the integer and floating point data operand values that the microinstruction needs to execute. The processor 200 of this embodiment is composed of a plurality of execution units including an address generation unit (AGU) 212, AGU214, high-speed ALU216, high-speed ALU218, low-speed ALU220, floating-point ALU222, and floating-point movement unit 224. In this embodiment, the floating-point execution blocks 222 and 224 execute floating-point processing, MMX processing, SIMD processing, and SSE processing. The floating point ALU222 of this embodiment has a 64-bit unit floating point divider for performing micro ops on division, square root and remainder. In the embodiments of the present invention, any processing relating to floating point numbers is performed on floating point hardware. For example, the conversion between integer and floating point formats involves a floating point register file. Similarly, the floating point division process is performed in the floating point divider.
Non-floating point numbers and integer types, on the other hand, are handled by integer hardware resources. Simple and frequently used ALU operations are processed by the fast ALU execution units 216 and 218. The high-speed ALU 216 and 218 of this embodiment can execute high-speed processing with an effective waiting time of half a clock cycle. In one embodiment, most complex integer operations are passed to the slow ALU220. The slow ALU220 includes integer execution hardware for long latency types of processing such as multiplication, shift, flag logic and branching. Memory load / store processing is performed by AGU212 and 214. In this embodiment, the integers ALU216, 218 and 220 are described for performing integer processing on 64-bit data operands. In other embodiments, ALU216, 218 and 220 can be implemented to support various data bits such as 16, 32, 128, 256. Similarly, floating point units 222 and 224 can be implemented to support operands with different bit widths. In one embodiment, floating point units 222 and 224 are executed in 128-bit wide Packed data operands for SIMD and multimedia instructions.
The word "register" is used here to refer to the onboard processor storage used as part of a macro instruction that identifies an operand. In other words, the registers referred to here are those that can be seen from outside the processor (from the programmer's point of view). However, the registers of one embodiment are not limited to a particular type of circuit. Rather, the registers of one embodiment may be capable of storing and providing data and performing the functions described herein. The registers described here are, for example, dedicated physical registers, dynamically allocated physical registers by using register renaming, combinations of dedicated physical registers and dynamically allocated physical registers, and the like. It can be realized by the circuit inside the processor using various techniques. In one embodiment, the integer register stores 32-bit integer data. The register file of one embodiment also contains eight multimedia SIMD registers for packed data. For the following explanation, the register is a data register that can hold packed data, such as the 64-bit wide MMX register (mm register) in a microprocessor capable of MMX technology from Intel Corporation in Santa Clara, California. Is interpreted as. Such MMX registers are available in both integer and floating point formats and can be operated by packed data elements with SIMD and SSE instructions. Similarly, 128-bit wide XMM registers for SSE2 technology are also available to hold such Packed data operands. In this embodiment, the registers do not need to distinguish between the two data types when storing packed data and integer data.
FIG. 3A shows various signed and unsigned Packed data type representations in 128-bit wide multimedia registers according to an embodiment of the present invention. The Packed Byte format of this example contains 6 Packed Byte data elements. Bytes are defined as 8-bit data. The unsigned packed byte representation 302 indicates the storage of unsigned packed bytes in the SIMD register. The information of each byte data element is from the 7th bit to the 0th bit for the 0th byte, from the 15th bit to the 8th bit for the 1st byte, and the 23rd bit for the 2nd byte. Is stored in the 16th bit, and finally in the 128th to 120th bits for the 15th byte. Therefore, all available bits are used in the register. This storage arrangement brings about an improvement in the storage efficiency of the processor. When 16 data elements are accessed, one process is executed in parallel for the 16 data elements.
Signed Packed Byte Representation 304 indicates the storage of signed Packed Bytes. The 8th bit of all byte data elements is a sign indicator. The Packed Word format of this example contains eight Packed Word data elements. Each Packed word contains 16 bits of information. The unsigned Packed word representation 306 shows how the 7th to 9th words are stored in the SIMD register. The signed packed word representation 308 is similar to the unsigned packed word-in register representation 306. Here, the 16th bit of each word data element is a code indicator. The Packed doubleword format is 128 bits long and contains four Packed doubleword data elements. Each Packed doubleword element contains 30 bits of information. The unsigned Packed doubleword representation 310 shows how doubleword elements are stored. The signed packed doubleword representation 312 is similar to the unsigned packed doubleword-in register representation 310. Here, the required sign bit is the 32nd bit of each doubleword data element. Packed quadwords are 128 bits long and contain two Packed quadword data elements.
In general, a data element is a piece of data that is stored in one register or memory area along with other data elements of the same length. In the Packed data sequence for SSE2 technology, the number of data elements stored in the XMM register is 128 bits divided by the bit length of each data element. Similarly, in a packed data sequence for MMX and SSE technology, the number of data elements stored in the MMX register is 64 bits divided by the bit length of each data element. Although the data type shown in Figure 3A is 128 bits long, the embodiments of the present invention are also operational with operands of 64 bits wide or other sizes.
Figure 3B shows other in-register data storage formats. Each Packed data can contain multiple independent data elements. Three packed data formats are shown: Packed Half 341, Packed Single 342 and Packed Double 343. One embodiment of Packed Half 341, Packed Single 342 and Packed Double 343 includes fixed point data elements. In other embodiments, one or more of Packed Half 341, Packed Single 342 and Packed Double 343 may contain floating point data elements. Another embodiment of Packed Half 341 is 128-bit long, including eight 16-bit data elements. One embodiment of Packed Single 342 is 128 bits long and contains four 32-bit data elements. One embodiment of Packed Double 343 is 128 bits long and contains two 64-bit data elements. It will be appreciated that such a packed data format can be further extended to other register lengths of, for example, 96 bits, 160 bits, 192 bits, 224 bits, 256 bits or more. ..
Figure 3C shows a type of processing encoding format described in "IA-32 Manual for Intel Architecture Software Developers 2" available from Intel Corporation via "www.intel.com/design/litcentr". An example of a processing coding format having 32 or more bits corresponding to (opcode) and a register / memory operand addressing mode is shown. The type of rounding shift multiplication is encoded by one or more fields 361 and 362. Up to two operands are located per instruction, including up to two source operand identifiers 364 and 365. In one embodiment of the shift merge instruction, the destination operand identifier 366 is the same as the source operand identifier 364. In other embodiments, the destination operand identifier 366 is identical to the source operand identifier 365. Therefore, in the example of the shift merge process, one of the source operands specified by the source operand identifiers 364 and 365 is overwritten by the result of the multiplication upper rounding shift process. In one embodiment of the shift merge instruction, operand identifiers 364 and 365 can be used to identify 64-bit source and destination operands.
Figure 3D shows another processing code (opcode) format 370 with 40 bits or more. Opcode format 370 supports opcode format 360 and is a selective prefix byte. It consists of byte) 378. The type of multiplication high rounding shift processing is encoded by one or more fields 378, 371 and 372. Up to two operand positions per instruction are specified by the source operand identifiers 374 and 375 and the prefix byte 378. In one embodiment of the Packed Multiplier Rounding Shift, the prefix byte 378 is used to identify the 128-bit source and destination operands. In one embodiment of the multiplication higher instruction, the destination operand identifier 376 is the same as the source operand identifier 374. In another embodiment, the destination operand identifier 376 is the same as the source operand identifier 375. Therefore, in the example of the multiplication upper processing, one of the source operands specified by the source operand identifiers 374 and 375 is overwritten by the result of the multiplication upper processing. Opcode formats 360 and 370 are register-to-register, partially identified by MOD fields 363 and 373 and selective scale-index-base and displacement bytes. to register, memory to register, register by memory, register by register, register by immediate, register by register Enables register to memory addressing.
In another embodiment, as shown in FIG. 3E, 64-bit single instruction multiplex data (SIMD) arithmetic processing is performed through coprocessor data processing (CDP) instructions. Processing coding (opcode) format 380 indicates a CDP instruction having CDP opcode fields 382 and 389. In another embodiment of the multiplication upper rounding shift process, the type of CDP instruction is encoded by one or more fields 383, 384, 387 and 388. Up to three operand positions are specified per instruction, including up to two source operand identifiers 385 and 390 and one destination operand identifier 386. One embodiment of the coprocessor can operate on 8, 16, 32 and 64-bit values. In one embodiment, multiplication higher processing is performed on fixed point or integer data elements. In some embodiments, the merge instruction may be conditionally executed utilizing condition field 381. For some multiplication superordinate instructions, the size of the source data is encoded by field 383. In some embodiments of the shift merge instruction, zero (Z), negative (N), carry (C) and overflow (V) detections are performed in the SIMD field. For some instructions, the type of saturation is encoded by field 384.
In one embodiment of the invention, the packed multiplication upper rounding shift is represented by the instruction formats PMULHRSW mm1, mm2 / m64. The PMULHRSW in this example helps to remember the packed multiplication high rounding shift word. In this case, it is an instruction with two source operands mm1 and mm2 / m64. The instructions of this embodiment are executed by a 64-bit Packed data block composed of a plurality of smaller data elements. In this case, each data element has a length of 16 bits or words. Each Packed data block can contain four words that form a total of 64 bits. The first source operand "mm1" is a 64-bit MMX register. In this embodiment, the 64-bit MMX register "mm1" from the first source operand is also the destination of the result of the Packed multiplication high rounding shift process. The second source operand "mm2 / m64" in this example can be a 64-bit MMX register (mm2) or a 64-bit memory location (m64).
While the examples described below generally relate to 64-bit length operands and data blocks, examples of multiplication high round shift instructions can also be processed by 128-bit packed data blocks. For example, the instruction format of one embodiment can be represented as PMULHRSWxmm1, xmm2 / m128. Each of the two source operands in this case is 128-bit long and consists of eight 16-bit word-sized data elements. The first source operand "xmm1" is a 128-bit XMM register. In this embodiment, the XMM register "xmm1" is also the destination of the result. The second source operand "xmm2 / m128" in this embodiment is a 128-bit XMM register (xmm2) or a 128-bit memory unit (m128). In this embodiment, each data block can include a signed integer. In one embodiment, the signed integer is in two's complement format.
Further, although the examples described herein relate to packed data blocks composed of word-sized data elements, other various sized data elements are also considered. For example, another embodiment of the Packed Multiply Rounding Shift Instruction may be performed on a data element having a length of bytes, doublewords or quadwords. Similarly, the length of data operands is not limited to 64 and 128. For example, instructions from other embodiments may be executed for 256-bit long Packed operands.
FIG. 4B is a block diagram of an embodiment of logic for executing SIMD integer multiplication upper rounding shift processing on a data operand according to the present invention. The PMULHRSW for the multiplication upper rounding shift process (or multiplication upper for simplification) according to this embodiment starts with two pieces of information, the first data operand DATAA410 and the second data operand DATAB420. In one embodiment, the PMULHRSW multiplication super-instruction is decrypted into one microprocess. In another embodiment, the instruction is decoded into a variable number of micro ops in order to perform a multiplication higher processing on the data operand.
Here, DATAA410, DATAB420 and RESULTANT440 are generally referred to as operands or data blocks, but are not limited to the following, and include registers, register files and memory areas. In one embodiment, DATAA410 and DATAB420 are 64-bit wide MMX registers (or, in some examples, referred to as "mm"). Depending on the particular embodiment, the data operands can be of other width, such as 128 or 256 bits. The first operand 410 and the second operand 420 are data blocks containing x data elements, and when each data block is 1 byte (8 bits), each has a total 8 x bit width. Therefore, each data segment is 8x bit wide. Where x is 8, each operand is 8 bytes or 64 bits wide. In other embodiments, the data element may be a nibble (4 bits), a word (16 bits), a double word (32 bits), a quad word (64 bits), and the like. In other embodiments, x may be a data element width such as 16, 32, 64, etc.
The first Packed operand 410 in this embodiment is composed of four data elements A3, A2, A1 and A0. The second Packed operand 420 also consists of four data elements B3, B2, B1 and B0. The data elements here have the same length and are each composed of one word (16 bits) of data. However, in other embodiments of the invention, each data segment is processed with a longer 128-bit operand consisting of 1 byte (8 bits), and the 128-bit wide operand has a 16-byte wide data segment. Similarly, if each data segment is doubleword (32-bit) or quadword (64-bit), the 128-bit operand has a 4-doubleword-width or 2-quadword-width data segment, respectively. The embodiments of the present invention are not limited to data operands or data segments of a specific length, and can be sized to suit each embodiment.
Operands 410 and 420 are located in registers, memory areas, register files, or a combination thereof. The data operands 410 and 420 are sent to the multiplication upper rounding shift calculation logic 430 of the processor's execution unit along with the multiplication upper rounding shift instruction. The instruction should be decrypted in advance in the processor pipeline until the PMULHRSW instruction reaches the execution unit. Therefore, the multiplication super-instruction can follow the format of microprocessing (uop) or other decoding formats. In this embodiment, the two data operands 410 and 420 are received in the multiplication upper rounding shift calculation logic 430. Since this example is executed for a 64-bit width operand, the tentative space 431 must hold the product of 128-bit wide intermediate results. A temporary space with a width of 256 bits is required for a data operand with a width of 128 bits.
Logic 430 of this embodiment first multiplies the corresponding data values at each element position in order to obtain the product A × B. Each intermediate 32-bit value of A × B for the four positions is truncated to 18 bits. In this embodiment, truncation is performed by shifting each 32-bit value to the right by 14 bits and removing these bits. This leaves 18 bits for each tentative value. One "1" is added to the least significant bit of this embodiment for rounding. The 16 bits to the immediate right of the most significant bit of each rounded value are output at each data element position in the result 440. At the leftmost data element position in this embodiment, the result is equal to the bit [16: 1] of "((A3 × B3) >> 14) +1". The selection of rounded result bits [16: 1] scales this value appropriately, similar to decimal operations.
Other embodiments of the invention include, for example, 128/256/512 bit wide operands, bit / byte / word / doubleword / quadword size data segments, 8/16/32 bit wide shift counts, etc. It is feasible for operands and data segments of other lengths. Therefore, the examples of the present invention are not limited to operands, data segments and shift counts of a specific length, and can be sized suitable for each embodiment.
At run time, the Packed Integer Multiply Up-Round Shift Instruction in one example performs a SIMD signed 16-bit x 16-bit multiplication of the Packed Signed Integer Words of the 1st and 2nd Source operands for an exact 32 bits. Generate an intermediate product. The intermediate product in one embodiment is first truncated to the upper 18 bits. This 18-bit selection gives 18-bit intermediate precision. By adding "1" to the least significant bit of the 18-bit value, the rounded value is rounded. In other words, the rounding process adds "1" to the bit value in the 14th bit of the original 32-bit intermediate product. The final result is obtained by selecting 16 bits to the immediate right of the most significant bit of the 18-bit value. In this example, each value in the result contains one sign bit. Each data element resulting from this example can have a fixed point integer format of "1.15". The multiplication upper rounding shift instruction of this embodiment stores 16 bits of each rounding-shifted intermediate 32-bit value at an appropriate position of the destination operand.
In this embodiment, the results of this and other data element positions are packed into a data block result of the same size as the source data operand. For example, if the source Packed data operand is 64 or 128 bits wide, the resulting Packed data block will also be 64 or 128 bits wide, respectively. In addition, source data operands for coding are obtained from registers or memory areas. In this embodiment, the resulting Packed data block overwrites the data in the SIMD register for one of the source data operands.
FIG. 4B is a block diagram of the operation of the integer multiplier upper rounding shift process for the selected data element position. DATA ELEMENT A450 is from the first source operand. DATA ELEMENT B452 is from the second source operand. The multiplication upper rounding shift process 454 of this embodiment is started by multiplying the data elements in order to generate a product of the intermediate value TEMP456. For two 16-bit wide source data elements, the product is a 32-bit median. In this embodiment, the most significant 18 bits of TEMP456 are used for rounding scaling processing. Further accuracy in the calculation can be achieved by maintaining 18 bits. The multiplication upper rounding shift process 454 is continued by performing rounding and scaling processing on the intermediate value 456 in order to obtain the latest intermediate value 458. In this embodiment, the rounding process is performed by adding "1" to the 14th bit of the 32-bit intermediate value TEMP456. By the way, the 14th bit of the 32-bit value is also the least significant bit of the 18-bit wide portion of interest. A shift process is performed on the 32-bit rounded value to scale the intermediate value. A 1-bit left shift is performed on the rounded value to reach the latest median value 458. The latest median 458 is truncated to give RESULT 460. In this embodiment, the bits of interest are the upper 16 bits of the latest 32-bit median 458, which are stored as RESULT460. The lower 16 bits are truncated in the truncation process.
FIG. 5 is a block diagram of an embodiment of the circuit 500 that executes the multiplication upper rounding shift process according to the present invention. The circuit 500 of this embodiment is provided in a vector complex integer unit. Since this integer unit is an embodiment of the PMULHRSW instruction with a 128-bit operand, it is divided into eight parts, each of which performs a 16-bit × 16-bit multiplication. The 64-bit operand embodiment requires four parts. In Figure 5, the SRC Y ELEMENT 502 is sent to the radix-4 booth recode block 504 of radix-4. The SRC X ELEMENT 502 will be received at the booth mux 508. Boothmax produces nine cross product vectors 509.
In the multiplication process by manual calculation, the process is started by extracting the least significant bit of a certain operand (A) and multiplying this bit with each digit of the other operand (B) in bit units. A one-line result is generated for each bit of multiplication target A. Each of these lines is known as a partial product. For example<tables num="1"><img file="JP4480997B2_D0001.tif" /></tables> Since a large number of multiplications requires a lot of hardware to handle all the partial products, a booth recoding technique is performed in one embodiment to simplify the calculations. In the booth recoding process, slightly more than half of the partial products (N bits / 2 + 1) are generated in the same way as the manual calculation. For example, instead of obtaining the above four partial products, the booth recoding process produces three partial products. Therefore, for a multiplier of 16 × 16, the partial product to be added is 16/2 + 1, that is, 9. This method is called radix 4 here. Each 16-bit multiplication array is a booth-coded array of radix 4. The booth coding process generated nine partial products, which were reduced by the carry sum adder (CSA) tree structure and adders. In one embodiment, the entire 16-bit array structure of the CSA tree looks like this:
<tables num="2"><img file="JP4480997B2_D0002.tif" /></tables> This embodiment is configured to handle negative multiplication. The "S" represents the sign and the "P" is used to describe the lower 2 bits of the previous partial product. For example, "pp" of partial product 1 is the least significant 2 bits of partial product 0. The essence of the leading sign extension is to roll off the sign bit. This is similar to bit inversion of 2's complement, which makes negative numbers positive before multiplication. Similarly, the essence of the "P" bit is to give +1 for the 2's complement inversion of the negative-to-positive conversion.
Bit [31:16] can be considered as the high-order result bit of multiplication. However, in the multiplication upper rounding shift, the rounding process and the shift process are dealt with before the final result. In one embodiment, the rounding process relates to adding "1" to bit position 14 of the array. However, there is no free space in the 14th bit of the partial product tree to easily add "1". In the eighth line, bits 13, 12 and 11 are empty positions. Similarly, there is a free space in bit 11 on line 7. Adding a "1" to all four positions, as shown in the R bits below, extends the "1" to bit position 14. In the rounding technique of this embodiment, the CSA compression tree 510 is as follows.
<tables num="3"><img file="JP4480997B2_D0003.tif" /></tables> The embodiments of the present invention utilize CSA to help reduce the number of partial product terms from 9 to 2 before the 32-bit adder 514. In one embodiment, the CSA compression tree reduces the number of partial product terms (using 4: 2 CSA) first from 9 to 6, then from 6 to 4, and finally from 4 to 2. This technique avoids the need for nine 32-bit adders. The output of the CSA tree 510 in this example is the partial product term reduced to 2. One is the final CSA total term and the other is the carry out term. In order to logically add these two terms for complete results, the carry term must be shifted one bit to the left to properly match the sum term. For example, bit 0, the least significant bit of the carry term, needs to be aligned with bit 1 of the sum term.
The 32-bit adder 514 generates FULL RESULT 515 by adding SUM512 and CARRY511. The SUM512 of this embodiment is SUM [31: 0]. Carry511 is Carry [30: 0] shifted 1 bit to the left. The bit associated with this embodiment is bit [30:15]. These 16 bits are shifted by 1 bit from the product of the above multiplications. In this embodiment of circuit 500, this shift process is implemented by code 516 with result max 518 and result max. Therefore, the RESULTANT520 of the signed integer multiplication upper rounding shift process is the 16 bits immediately to the right of the most significant bit of the FULL RESULT 515, that is, FULL RESULT [30:16]. In this embodiment, the results from each of the eight array structures for each pair of data elements are concatenated to obtain the final 128-bit result.
FIG. 6A shows the operation of the Packed multiplication upper rounding shift instruction according to the first embodiment of the present invention. The 64-bit wide source operand DATA A601 is a hexadecimal number 479C, respectively.<sub>16</sub>, 1AF7<sub>16</sub>, C000<sub>16</sub>And 0200<sub>16</sub>Consists of four data elements 602, 603, 604 and 605 that store. Similarly, the 64-bit wide source operand DATA B611 is each in hexadecimal D76E.<sub>16</sub>, 2BC5<sub>16</sub>, C0FF<sub>16</sub>And 0220<sub>16</sub>It is composed of four data elements 612, 613, "614 and 615 having. The Packed Multiplier Rounding Scaling Instruction according to an embodiment of the present invention, along with DATA A601 and DATA B611 as source operands, produces RESULTANT operand 621. The Packed Multiplier Rounding Scaling Process 620 of this example produces a result for each corresponding pair of source data elements. In this example, the four data elements of RESULTANT621 are hexadecimal E94E.<sub>16</sub>622、0938<sub>16</sub>623, 1F81<sub>16</sub>624 and 0009<sub>16</sub>Has 625.
FIG. 6B shows the more detailed behavior of the Packed Multiplier instruction at the particular data element position of FIG. 6A. Continuing from the example in Figure 6A, the position of the second data element from the left is described in more detail here. The value of the second leftmost data element 603 of DATA A601 is 1AF7<sub>16</sub>(Or, in binary, 001 1010 1111 0111). The value of the second leftmost data element 613 of DATA B611 is 2BC5<sub>16</sub>(Or, in binary, 0010 1011 1100 0101). Packed Multiplication In the upper rounding scaling process, these two values are first multiplied and 049C3D13.<sub>16</sub>(0000 0100 1001 1100 0011 1101 0001 0011<sub>2</sub>) Product 631 is obtained. This product 631 is treated as the first tentative intermediate value TEMP630.
The rounded portion 633 of this process is performed on the product 631. In this embodiment, the rounding process is to add "1" to the 14th bit 632 of the product 631. The result of rounding 633, 634, creates a new TEMP630. The result of rounding 634 is 049C7D13<sub>16</sub>(0000 0100 1001 1100 0111 1101 0001 0011<sub>2</sub>). The rounding result 634 is scaled to give the desired result in this example. The scaling process 636 here is executed as a 1-bit left shift of 634 as a result of the rounding process of TEMP630. Therefore, bits 30 to 15 are shifted up from bit positions 31 to 16. TEMP630 is truncated to 16-bit values, and the most significant 16 bits (upper part) of the rounded-shifted value are output as RESULTANT623. RESULTANT623 is the second data element position from the left of Packed RESULTANT612. In this example, RESULTANT 623 is 0938.<sub>16</sub>(0000 1001 0011 1000<sub>2</sub>).
An example of Packed Multiplier Rounding Scaling (PMULHRSW) processing at the second data element position of a pair of 64-bit operands is shown below.
<tables num="4"><img file="JP4480997B2_D0004.tif" /></tables> In the above example, one or both of the source data operands can be 64-bit data registers in the processor enabled by MMX / SSE technology or 128-bit data registers by SSE2 technology. Depending on the embodiment, these registers can be 64/128/256 bits wide. Similarly, one or both of the source operands can be memory areas other than registers. In one embodiment, the resulting destination is an MMX or XMM data register. In addition, the resulting destination may be the same register as one of the source operands. For example, in one architecture, the multiplication upper rounding shift instruction has a first source operand MM1 and a second source operand MM2. The predetermined destination for the result can in this case be the register of the first source operand MM1.
FIG. 7A is a flowchart 700 showing an embodiment of a method of performing an integer multiplication rounding shift process on a Packed data operand to obtain an upper part of the product. The length L is used to represent the width of the operand and data block. Depending on the particular embodiment, L is used to indicate the length in terms of the number of data segments, the number of bits, the number of bytes, the number of words, and so on. In block 710, the first data operand A of length L is received to perform the Packed Integer Multiply Upper Rounding Shift process. At block 720, the second data operand B of length L for PMULHRSW processing is received. In block 730, the execution instruction of the multiplication upper rounding shift is processed.
Further explanation is given regarding what occurs for each data element position in the details of the multiplication upper rounding shift process in block 730 of this embodiment. In one embodiment, the multiplication upper rounding shift processing for all of the resulting Packed data element positions is processed in parallel. In another embodiment, some parts of the data element are processed at the same time. In block 731, the tentative value TEMP is calculated by multiplying the value of the element from operand A by the value of the element from operand B. At block 732, this tentative value is rounded. In one embodiment, the upper 18 bits of the tentative value are used in the calculation for higher accuracy. In other embodiments, other numbers of bits may be of interest. After the rounding of block 732, the tentative value is scaled in block 733. In this embodiment, the scaling process is to shift the tentative value to the left by one bit. At block 734, the tentative value is truncated to the required number of bits and stored at the destination as the resulting value. The resulting value for each of the different pairs of source data elements is placed at the appropriate data element position corresponding to the source element pair of the resulting Packed operand.
FIG. 7B is a flowchart showing another embodiment of the method of acquiring the relevant high-order portion of the product obtained as a result of the Packed Integer Multiply Rounding Shift process. In this embodiment, the operand is composed of word-sized data elements. However, other embodiments may be implemented with data elements of other sizes, such as bytes, doublewords or quadwords. In block 742, the control signal of the multiplication upper rounding shift process is decoded. At block 744, the determination of the operand size in the process is checked. In one embodiment, the operand size can be determined by the control signal decoded in block 742. For example, the operand size can be coded by instruction. If it is determined that the operand size is 64 bits long, the register file and / or memory is accessed in block 746 and the operand data is acquired according to the location of the data. In one embodiment, the source operands may be in SIMD registers and / or memory areas. In the 64-bit length operands of this embodiment, each operand has a 4-word size data element.
The calculation of these four pairs of source data elements is shown as a set of four expressions in block 747. The first equation "TEMP [31: 0] = A [15: 0] x B [15: 0]" represents the multiplication of the source data elements. The second equation "INT (TEMP [31: 0] >> 14) +1" represents the rounding process of the intermediate result. In this embodiment, the tentative value is shifted right by 14 bits and "1" is added to the least significant bit. In other words, the upper 18 bits of the intermediate result are retained and a "1" is added to the original 14th bit. The third equation "DEST [15: 0] = TEMP [16: 1]" represents the shift and truncation of the rounded result. In this case, each resulting day or element is a word, which requires 16 bits. The remaining 18 bits [16: 1] are extracted here. In this embodiment, the shift to the left is performed by taking the 16 bits immediately to the left of the least significant bit. The truncated value is stored as a result of the data element position. In block 747, these three equations, unless the bit range is filled with the correct value for that position (ie [15: 0], [31:15], [47:32] and [63:48]). Is repeated for each data element position.
If it is determined in block 744 that the operand size is 128 bits long, the register file and / or memory is accessed in block 748 to obtain the required operand data. Each operand has an 8-word size data element, as opposed to the 128-bit length operand of this embodiment. Similar to the 64-bit path, in block 749 each of the eight pairs of source data elements is processed by the set of three expressions above. The correct bit range for this set of eight expressions in a 128-bit path is [15: 0], [31:15], [47:32], [63:48], [79:64], [95:80]. , [111: 96] and [127: 112]. The two paths described in this example relate to 64-bit and 128-bit operands, but the six 16-bit values for which other various length operands can be used in other examples are each. Corresponds to each data element position and is stored in each data element position of DEST.
SIMD Integer Multiply Round Shift Techniques have been disclosed. Although specific examples have been described and shown with the accompanying drawings, such examples are merely exemplary and do not limit the scope of the invention. Also, the invention is not limited to the particular configurations and arrangements exemplified and described, and one of ordinary skill in the art will be able to make various other modifications under the present disclosure. In such technical fields where progress is rapid and further progress is not readily predictable, the disclosed examples are facilitated by technological progress without departing from the principles of the present disclosure or the scope of the appended claims. It is possible to make such modifications.
The present invention is not limited to the above specific embodiment, and various modifications and changes can be made within the gist of the present invention.
<figref num="1A">FIG. 1A is a block diagram of a computer system composed of a processor having an execution unit that executes SIMD instructions for integer multiplication upper rounding shift processing according to an embodiment of the present invention.</figref><figref num="1B">FIG. 1B is a block diagram of a computer system as another example according to another embodiment of the present invention.</figref><figref num="1C">FIG. 1C is a block diagram of a computer system as yet another example according to another embodiment of the present invention.</figref><figref num="2">FIG. 2 is a block diagram of a microarchitecture of an example processor having a logic circuit that executes a packed integer multiplication upper rounding shift process according to the present invention.</figref><figref num="3A">FIG. 3A shows various Packed data type representations of multimedia registers according to an embodiment of the present invention.</figref><figref num="3B">FIG. 3B shows the Packed data types according to other embodiments.</figref><figref num="3C">FIG. 3C shows an embodiment of the processing coding (opcode) format of the Packed multiplication upper rounding shift instruction.</figref><figref num="3D">Figure 3D shows other processing coding formats.</figref><figref num="3E">Figure 3E shows yet another processing coding format.</figref><figref num="4A">FIG. 4A is a block diagram of an embodiment of the logic for executing the SIMD integer multiplication upper rounding shift process for the data operand according to the present invention.</figref><figref num="4B">FIG. 4B is a block diagram of the operation of the integer multiplication upper rounding shift process for the selected data element position.</figref><figref num="5">FIG. 5 is a block diagram of an embodiment of a circuit that executes the multiplication upper rounding shift process according to the present invention.</figref><figref num="6A">FIG. 6A shows the operation of the Packed multiplication upper rounding shift instruction according to the first embodiment of the present invention.</figref><figref num="6B">FIG. 6B shows the more detailed behavior of the Packed Multiplier instruction at the particular data element position of FIG. 6A.</figref><figref num="7A">FIG. 7A is a flowchart showing an embodiment of a method of executing integer multiplication, rounding, and shifting processing on the Packed data operand for acquiring the upper part of the product.</figref><figref num="7B">FIG. 7B is a flowchart showing another embodiment of the method of obtaining the relevant high-order part of the product resulting from the Packed Integer Multiply Rounding Shift process.</figref>
Code description
100, 140, 160 computer system 102, 166, 200 processors 104, 167 cache 106, 208, 210 register files 108 Execution unit 109 Packed instruction set 110 processor bus 112 graphics / video card 114 AGP interconnect 116 Memory Controller Hub (MCH) 118 Memory interface 120 memory 122 Dedicated hub interface bus 124 Data storage device 126 wireless transmitter 128 flash BIOS 130 I / O Controller Hub (ICH) 134 Network controller 141 Bus 142, 162 Execution unit 143 Packed instruction set 144, 165 decoder 145, 164 register file 146 SDRAM control 147 SRAM control 148 Burst flash memory interface 149 PCMCIA / CF card control 150 LCD control 151 DMA control 152 Alternate bus master interface 153 I / O Bus 154 I / O bridge 155 UART 156 USB 157 Bluetooth UART 158 I / O extended interface 159, 170 processing core 161 SIMD coprocessor 163 instruction set 168 I / O system 169 Wireless interface 201 front end 202 High speed scheduler 203 Out of Order Engine 204 Slow / Normal Floating Point Scheduler 206 Simple Floating Point Scheduler 211 Execution block 212, 214 Address Generation Unit (AGU) 216, 218 High speed ALU 220 low speed ALU 222 Floating point ALU 224 Floating point move unit 226 Instruction prefetcher 228 Instruction decoder 230 trace cache 232 Microcode ROM 234 uop queue 430 Multiplication Upper Rounding Shift Calculation Logic
Every citation, both ways
| Document | Relation | Office |
|---|---|---|
| JP11500547A | Cites | Japan |
| JP2002527808A | Cites | Japan |
| JP2001516916A | Cites | Japan |
| JP03268024A | Cites | Japan |
14 members in 7 offices
Priority claims5
| Document | Office | Kind | Date |
|---|---|---|---|
| 610833 | United States of America | – | |
| 61083303 | United States of America | A | |
| 61083303 | United States of America | A | |
| 2003610833 | – | – | – |
| US20030610833 | – | – | – |
Members14
| Document | Office | Kind | |
|---|---|---|---|
| US2004267857A1 | United States of America | A1 | |
| TW200500940A | Taiwan Province of China | A | |
| NL1025106A1 | Netherlands (Kingdom of the) | A1 | |
| KR20050005730A | Republic of Korea | A | |
| JP2005025718A | Japan | A | |
| CN1577257A | China | A | |
| RU2003137661A | Russian Federation | A | |
| RU2263947C2 | Russian Federation | C2 | |
| TWI245219B | Taiwan Province of China | B | |
| KR100597930B1 | Republic of Korea | B1 | |
| NL1025106C2 | Netherlands (Kingdom of the) | C2 | |
| CN100541422C | China | C | |
| US7689641B2 | United States of America | B2 | |
| JP4480997B2This record | Japan | B2 |
28 legal events, as the office reported them to INPADOC
Over the term
Point at a mark for the eventEvents
| Event | Code | |
|---|---|---|
| Cancellation because of no payment of annual feesLAPS | LAPS | |
| Receipt of annual feesJAPANESE INTERMEDIATE CODE: R250R250 | R250 | |
| Receipt of annual feesJAPANESE INTERMEDIATE CODE: R250R250 | R250 | |
| Receipt of annual feesJAPANESE INTERMEDIATE CODE: R250R250 | R250 | |
| Receipt of annual feesJAPANESE INTERMEDIATE CODE: R250R250 | R250 | |
| Receipt of annual feesJAPANESE INTERMEDIATE CODE: R250R250 | R250 | |
| Receipt of annual feesJAPANESE INTERMEDIATE CODE: R250R250 | R250 | |
| Receipt of annual feesJAPANESE INTERMEDIATE CODE: R250R250 | R250 | |
| Receipt of annual feesJAPANESE INTERMEDIATE CODE: R250R250 | R250 | |
| Receipt of annual feesJAPANESE INTERMEDIATE CODE: R250R250 | R250 | |
| Renewal fee payment (event date is renewal date of database)FPAY | FPAY | |
| Receipt of annual feesJAPANESE INTERMEDIATE CODE: R250R250 | R250 | |
| Renewal fee payment (event date is renewal date of database)FPAY | FPAY | |
| Renewal fee payment (event date is renewal date of database)FPAY | FPAY | |
| Certificate of patent or registration of utility modelJAPANESE INTERMEDIATE CODE: R150R150 | R150 | |
| Certificate of patent or registration of utility modelJAPANESE INTERMEDIATE CODE: R150R150 | R150 | |
| First payment of annual fees (during grant procedure)JAPANESE INTERMEDIATE CODE: A61A61 | A61 | |
| Written decision to grant a patent or to grant a registration (utility model)JAPANESE INTERMEDIATE CODE: A01A01 | A01 | |
| Written decision to grant a patent or to grant a registration (utility model)JAPANESE INTERMEDIATE CODE: A01A01 | A01 | |
| Decision of grant or rejection writtenTRDD | TRDD | |
| Transfer to examiner for re-examination before appeal (zenchi)AppealJAPANESE INTERMEDIATE CODE: A911A911 | A911 | |
| Request for written amendment filedJAPANESE INTERMEDIATE CODE: A523A521 | A521 | |
| Decision of refusalJAPANESE INTERMEDIATE CODE: A02A02 | A02 | |
| Request for written amendment filedJAPANESE INTERMEDIATE CODE: A523A521 | A521 | |
| Notification of reasons for refusalJAPANESE INTERMEDIATE CODE: A131A131 | A131 | |
| Report on retrievalJAPANESE INTERMEDIATE CODE: A971007A977 | A977 | |
| Request for written amendment filedJAPANESE INTERMEDIATE CODE: A523A521 | A521 | |
| Written request for application examinationJAPANESE INTERMEDIATE CODE: A621A621 | A621 |
Numbers
- Publication
- 4480997
- Publication, DOCDB
- 4480997
- Publication, EPODOC
- JP4480997B
- Application
- 425711
- Application, DOCDB
- 2003425711
- Application, EPODOC
- JP20030425711
Titles2
- Japanese
- SIMD整数乗算上位丸めシフト
- English
- SIMD Integer Multiply Upward Rounding Shift
Classification
- CPC, 9
- G06F9/30014
- G06F7/52
- G06F5/01
- G06F7/49947
- G06F7/5338
- G06F9/30036
- G06F9/3885
- G06F2207/382
- G06F2207/3828
- IPC, 11
- G06F7 38
- G06F7 496
- G06F9 305
- G06F9 315
- G06F5 01
- G06F7 52
- G06F7 523
- G06F7 53
- G06F9 30
- G06F9 302
- G06F9 38