Processing modules for computer architecture for broadband networks
29 claims: 2 independent, 27 dependent
- 1コンピュータ・プロセッサにおいて、 複数の第1処理ユニットを有し、各々の前記第1処理ユニットが、前記第1処理ユニットに関連づけられたローカル・メモリを含み、 前記第1処理ユニットによるデータの処理を制御する第2処理ユニットを有し、 前記第2処理ユニットが、前記データを前記第1処理ユニットの前記ローカル・メモリへ転送指示し、前記第1処理ユニットで、前記第1処理ユニットの前記ローカル・メモリの前記データを処理するように作動可能であり、 前記データはプログラムに関連づけられたものであり、 前記第1処理ユニットは、前記プログラムとプログラムに関連づけられたデータを処理するものであり、 前記第2処理ユニットは、前記第1処理ユニットによる前記プログラムとプログラムに関連づけられたデータの処理を制御し、 更に、前記第2処理ユニットは、前記プログラムと前記プログラムに関連付けられたデータを前記第1処理ユニットの前記ローカル・メモリへ転送指示し、前記第1処理ユニットで、前記第1処理ユニットの前記ローカル・メモリの前記プログラムと前記プログラムに関連付けられたデータを処理するように作動可能であり、 前記第1処理ユニットは、前記プログラムと前記プログラムと関連づけられたデータとが格納されるメイン・メモリにアクセスが可能であり、 前記第2処理ユニットは、前記メイン・メモリに格納されたプログラムと前記プログラムと関連付けられたデータを前記第1処理ユニットの前記ローカル・メモリへ転送指示し、 前記メイン・メモリが、複数のメモリ・ロケーションを含み、各々の前記メモリ・ロケーションが当該メモリ・ロケーションに関連付けられたメモリ・セグメントを含み、 前記各々のメモリ・セグメントには、前記メモリ・セグメントと関連付けられたメモリ・ロケーションに格納されたデータが最新のものであるか否かを示す状態情報が記録され、 前記状態情報が、前記メモリ・セグメントと関連付けられたメモリ・ロケーションに格納された前記データが最新のものではないことを示す場合に、第1処理ユニットから前記データの読み出し要求があった場合には、当該データの読み出しを阻止するとともに、当該読み出し要求を行った第1処理ユニットの識別子と当該第1処理ユニットと関連付けられたローカル・メモリの内の記憶位置とが前記メモリ・セグメントに記録され、 前記データが最新のものに更新された場合には、当該更新されたデータが、前記識別子が記録された前記第1処理ユニットの前記記録された記憶位置へと転送されるよう構成されている、プロセッサ。
- 2請求項1に記載のプロセッサにおいて、前記メイン・メモリがダイナミック・ランダム・アクセス・メモリであることを特徴とするプロセッサ。
- 3請求項1に記載のプロセッサにおいて、前記第1処理ユニットの各々が単一命令、複数データ・プロセッサであることを特徴とするプロセッサ。
- 4請求項1に記載のプロセッサにおいて、前記第1処理ユニットの各々が、1組のレジスタと、複数の浮動小数点演算ユニットと、前記1組のレジスタと前記複数の浮動小数点演算ユニットとを接続する1以上のバスとを含むことを特徴とするプロセッサ。
- 5請求項4に記載のプロセッサにおいて、前記第1処理ユニットの各々が、複数の整数演算ユニットと、前記複数の整数演算ユニットと前記1組のレジスタとを接続する1以上のバスとをさらに含むことを特徴とするプロセッサ。
- 6請求項1に記載のプロセッサにおいて、光インターフェースと、光導波路とをさらに有し、前記光インターフェースが、前記プロセッサによって生成された電気信号を、前記プロセッサから伝送するための光信号に変換するとともに、前記プロセッサまで伝送された光信号を電気信号に変換することが可能で、前記光導波路が前記光信号を伝送するために前記光インターフェースと接続されていることを特徴とするプロセッサ。
- 7請求項1に記載のプロセッサにおいて、前記ローカル・メモリがスタティック・ランダム・アクセス・メモリであることを特徴とするプロセッサ。
- 8請求項1に記載のプロセッサにおいて、画素データを生成するレンダリング・エンジンと、前記画素データを一時的に格納するフレーム・バッファと、前記画素データをビデオ信号に変換する表示制御装置と、をさらに有することを特徴とするプロセッサ。
- 9請求項1に記載のプロセッサにおいて、前記プログラムと関連付けられた前記データがスタック・フレームを含むことを特徴とするプロセッサ。
- 10請求項1に記載のプロセッサにおいて、前記各々の第1処理ユニットが、制御装置を有し、前記プログラムと前記関連付けられたデータの前記処理時、前記メイン・メモリから前記第1処理ユニットに関連付けられたローカル・メモリへ前記データの転送を指示することを特徴とするプロセッサ。
- 11請求項1に記載のプロセッサにおいて、前記メイン・メモリが、複数のメモリ・バンク・コントローラと、前記第1処理ユニットの各々と前記メイン・メモリとの間で接続を行うためのクロス・バー・スイッチとを有することを特徴とするプロセッサ。
- 12請求項1に記載のプロセッサにおいて、前記各々の第1処理ユニットが、前記メモリ・ロケーションからのデータの読み出し、あるいは、前記メモリ・ロケーションへのデータの書き込みを禁止する手段をさらに有することを特徴とするプロセッサ。
- 13請求項1に記載のプロセッサにおいて、ダイレクト・メモリ・アクセス・コントローラをさらに有することを特徴とするプロセッサ。
- 14請求項13に記載のプロセッサにおいて、前記第2処理ユニットが、前記ダイレクト・メモリ・アクセス・コントローラにコマンドを出すことによって、前記プログラムと、前記プログラムと関連付けられた前記データとを、前記第1処理ユニットに関連付けられたローカル・メモリへ転送を命令することと、前記コマンドに応答して、前記ダイレクト・メモリ・アクセス・コントローラが、前記プログラムを前記第1処理ユニットに関連付けられたローカル・メモリへ転送することを特徴とするプロセッサ。
- 15請求項14に記載のプロセッサにおいて、前記第1処理ユニットが、前記ダイレクト・メモリ・アクセス・コントローラにコマンドを出すことによって、前記プログラムを処理するために、前記メイン・メモリから、前記第1処理ユニットに関連付けられたローカル・メモリへ前記データを転送を命令することと、前記コマンドに応答して、前記ダイレクト・メモリ・アクセス・コントローラが、前記第1処理ユニットに関連付けられたローカル・メモリへ前記データを転送することを特徴とするプロセッサ。
- 16請求項15に記載のプロセッサにおいて、前記第1処理ユニットが、前記ダイレクト・メモリ・アクセス・コントローラにコマンドを出すことによって、前記第1処理ユニットに関連付けられたローカル・メモリから、前記プログラムの前記処理の結果データを、前記メイン・メモリへ転送する命令を出すことと、前記コマンドに応答して、前記ダイレクト・メモリ・アクセス・コントローラが、前記第1処理ユニットに関連付けられたローカル・メモリから、前記メイン・メモリへ前記結果データを転送することを特徴とするプロセッサ。
- 17処理装置において、 複数の第1処理ユニットを有して成る1つ以上のプロセッサ・モジュールを有し、前記第1処理ユニットのそれぞれは当該第1処理ユニットに関連づけられたローカル・メモリを備え、 前記第1処理ユニットによるデータの処理を制御する第2処理ユニットを有し、 前記第2処理ユニットが、前記データを前記第1処理ユニットのローカル・メモリへ転送指示し、その後、前記第1処理ユニットが、前記ローカル・メモリから前記データを処理するように作動可能であり、 前記データは、プログラムに関連づけられたものであり、 前記第1処理ユニットは、前記プログラムとプログラムに関連づけられたデータを処理するものであり、 前記第2処理ユニットは、前記第1処理ユニットによる前記プログラムと前記プログラムに関連付けられたデータの処理を制御するものであり、 前記第2処理ユニットが、前記プログラムと前記プログラムに関連付けられたデータを前記第1処理ユニットのローカル・メモリへ転送指示し、その後、前記第1処理ユニットが、前記ローカル・メモリから前記プログラムと前記プログラムと関連付けられたデータを処理するように作動可能であり、 前記プログラムと前記プログラムに関連付けられたデータを格納するためのメイン・メモリを更に有し、 前記第2処理ユニットは、前記プログラムと前記プログラムに関連付けられたデータを前記メイン・メモリから前記ローカル・メモリへ転送指示することによって、前記第1処理ユニットに前記プログラムを処理する指示を出し、 前記メイン・メモリが、複数のメモリ・ロケーションを含み、各々の前記メモリ・ロケーションが当該メモリ・ロケーションに関連付けられたメモリ・セグメントを含み、 前記各々のメモリ・セグメントには、前記メモリ・セグメントと関連付けられたメモリ・ロケーションに格納されたデータが最新のものであるか否かを示す状態情報が記録され、 前記状態情報が、前記メモリ・セグメントと関連付けられたメモリ・ロケーションに格納された前記データが最新のものではないことを示す場合に、第1処理ユニットから前記データの読み出し要求があった場合には、当該データの読み出しを阻止するとともに、当該読み出し要求を行った第1処理ユニットの識別子と当該第1処理ユニットと関連付けられたローカル・メモリの内の記憶位置とが前記メモリ・セグメントに記録され、 前記データが最新のものに更新された場合には、当該更新されたデータが、前記識別子が記録された前記第1処理ユニットの前記記録された記憶位置へと転送されるよう構成されている、処理装置。
- 18請求項17に記載の処理装置において、少なくとも1つの前記プロセッサ・モジュール用の前記複数の第1処理ユニットの数が8つであることを特徴とする処理装置。
- 19請求項17に記載の処理装置において、最低1つの前記プロセッサ・モジュール用の前記第1処理ユニットの数が4つであることを特徴とする処理装置。
- 20請求項17に記載の処理装置において、前記プロセッサ・モジュールの各々が、ただ1つの前記第2処理ユニットを有することを特徴とする処理装置。
- 21請求項18に記載の処理装置において、前記プロセッサ・モジュールの各々が、ダイレクト・メモリ・アクセス・コントローラをさらに有し、前記ダイレクト・メモリ・アクセス・コントローラが、前記第1処理ユニットおよび前記第2処理ユニットから出力されるコマンドに応答して、前記プログラムと前記関連付けられたデータを前記メイン・メモリと前記ローカル・メモリとの間で転送することを特徴とする処理装置。
- 22請求項17に記載の処理装置において、前記プロセッサ・モジュールの各々が、前記第1処理ユニットと前記第2処理ユニットとの通信のため1つのローカル・バスを、さらに有することを特徴とする処理装置。
- 23請求項17に記載の処理装置において、前記プロセッサ・モジュール間の通信を行うためのモジュール・バスをさらに有することを特徴とする処理装置。
- 24請求項17に記載の処理装置において、前記各々の第1処理ユニットが、複数の浮動演算ユニットと複数の整数演算ユニットを有することを特徴とする処理装置。
- 25請求項17に記載の処理装置において、1つ以上の光インターフェースをさらに有し、前記各々の光インターフェースが、前記処理装置からの伝送のために前記プロセッサ・モジュールの電気信号を光信号に変換し、また、前記処理装置へ伝送された光信号を電気信号に変換することが可能であることを特徴とする処理装置。
- 26請求項17に記載の処理装置において、前記の少なくとも1つのプロセッサ・モジュールが、画素データを生成するレンダリング・エンジンと、前記画素データを一時的に格納するフレーム・バッファと、前記画素データをビデオ信号に変換する表示制御装置と、をさらに有することを特徴とする処理装置。
- 27請求項17に記載の処理装置において、前記プロセッサ・モジュールの数が1であることを特徴とする処理装置。
- 28請求項17に記載の処理装置において、前記プロセッサ・モジュールの数が2であることを特徴とする処理装置。
- 29請求項17に記載の処理装置において、前記プロセッサ・モジュールの数が4であることを特徴とする処理装置。
Independent claims29
102 paragraphs, as filed
The present invention relates to an architecture for a computer processor and a computer network, and relates to an architecture for a computer processor and a computer network in a broadband environment. Furthermore, the present invention relates to a programming model for such an architecture.
Computers and computing devices in modern computer networks (such as local area networks (LANs) used in office networks and global networks such as the Internet) have been designed primarily for stand-alone computing. It was. Sharing data and application programs (applications) over computer networks was not a major design goal for these computers and computing devices. These computers and computing devices are also commonly designed with a wide range of different types of processors manufactured by a variety of different manufacturers (Motorola, Intel, Texas Instruments, Sony, etc.). .. Each of these processors has its own specific instruction set and instruction set architecture (ISA). That is, it has its own particular set of assembly language instructions and a structure for the main arithmetic unit and storage device that executes these instructions. Programmers are required to understand the instruction set and ISA of each processor and write applications for these processors. The mix of different types of computers and computing devices on today's computer networks complicates the sharing and processing of data and applications. Furthermore, in order to make adjustments for this mixed environment, it is often necessary to prepare multiple versions of the same application.
There is a wide range of computers and computing devices connected to global networks, especially the Internet. In addition to personal computers (PCs) and servers, these computing devices include cellular phones, mobile computers, personal digital assistants (PDAs), set-top boxes, digital televisions and other equipment. included. A major problem has arisen due to the sharing of data and applications in a mixture of different products in computers and computing devices.
Several methods have been tried to solve these problems. Among these techniques are, among other things, good interfaces and complex programming techniques. These solutions often require the realization of a substantial increase in processing power. In addition, these solutions often result in a substantial increase in the time required to process the application and the time required to transmit data over the network.
In general, data is transmitted over the Internet separately from the corresponding application. This approach eliminates the need to send the application itself to each set of transmission data for the application. Therefore, while this approach minimizes the amount of bandwidth required, it often causes user dissatisfaction. In other words, the client-side computer may not be able to obtain a proper application or the latest application for using this transmission data. In addition, this approach requires that a plurality of applications having different versions be prepared for each application corresponding to a plurality of different types of ISAs and instruction sets used by processors on the network.
The Java® model attempts to solve this problem. This model uses a small application (applet) that adheres to strict security protocols. Applets are sent from a server-side computer over a network and run by a client-side computer (client). All Java applets run on the client-side Java virtual machine because it is necessary to avoid sending different versions of the same applet to clients using different ISAs. A Java virtual machine is software that emulates a computer with Java ISA and an instruction set. However, this software is executed by a client-side ISA and a client-side instruction set. On the client side, the ISA and instruction set are different, but the Java virtual machine version given is one. Therefore, it is not necessary to prepare a different version for each of a plurality of applets. Each client can execute a java applet by downloading only the appropriate Java virtual machine corresponding to the ISA and instruction set of the client.
<p> Although the challenge of having to write different versions of applications for each different ISA and instruction set has been solved, the Java processing model requires an additional layer of software on the client-side computer. This additional layer of software slows down the processor significantly. This slowdown is especially noticeable for real-time multimedia applications. In addition, the downloaded Java applet may contain viruses and processing malfunctions. These viruses and malfunctions can cause client database corruption and other damage. The security protocol used in the Java model solves this problem by providing software called a "sandbox" (a space in memory on the client side where Java applets cannot write any more data). Although attempts have been made to resolve this, this software-driven security model is frequently unstable at its run and requires more processing.</p><p> Real-time, multimedia network applications are becoming more and more important. These network applications require very high speed processing. In the future, thousands of megabits of data per second may be needed for such applications. It is very difficult to reach such a processing speed with the current architecture of the network, especially the architecture of the Internet, and the programming model currently implemented in the Java model and the like.</p><p> Therefore, there is a need for new computer architectures, new architectures for computer networks, and new programming models. It is hoped that this new architecture and programming model will solve the problem of sharing data and applications between various members of the network without any computational burden. It is also desirable that this new computer architecture and programming model solve the inherent security problems that arise when sharing applications and data between members of the network.</p>
<p> In one embodiment of the invention, there is a category or form of a computer, a computing device, and a computer network (or, instead of a computer network, a computer network system or a computer system including a plurality of computers). A new architecture is provided for (and can also be). In other embodiments, the present invention provides a new programming model for these computers, computing devices and computer networks.</p><p> According to the present invention, all members of a computer network (all computers and computing devices on the network) consist of a common computing module. This common computing module has a uniform structure and preferably the same ISA is used. Members of the network include clients, servers, PCs, mobile computers, gaming machines, PDAs, set-top boxes, electrical equipment, digital televisions, and other devices that use computer processors. The uniform modular structure enables efficient high-speed processing of applications and data by network members and high-speed transmission of applications and data over the network. This structure also simplifies the configuration of network members of various sizes and processing powers, and simplifies the creation of processing applications by these members.</p><p> Also according to the present invention, in a computer network, a plurality of first processing units having a plurality of processors connected to the network, each of which has the same instruction set architecture, and the first processing unit. It has a second processing unit for controlling one processing unit, and the first processing unit is operable to process software cells transmitted over the network, said software. Each of the cells has a program compatible with the instruction set architecture, data associated with the program, and the software cell among all of the software cells transmitted over the network. Also provided is a computer network characterized by having an identifier (eg, a cell identification number) for unique identification.</p><p> According to the present invention, a computer system having a plurality of processors connected to a computer network, wherein each of the processors has a plurality of first processing units having the same instruction set architecture, and the above. It has a second processing unit for controlling a first processing unit, and the first processing unit can operate to process software cells transmitted over the network, said. Each of the software cells has a program compatible with the instruction set architecture, data associated with the program, and the software cell among all of the software cells transmitted over the network. Also provided is a computer system characterized by having an identifier (eg, a cell identification number) for uniquely identifying the.</p><p> In addition, according to the present invention, in a data stream of software cells for transmission over a computer network, the computer network has a plurality of processors, each of which is in the processor. A program for processing by one or more of the programs, data associated with the program, and a global identifier that uniquely identifies the software cell among all software cells transmitted over the network. A data stream characterized by having, is also provided. In the above configuration, the present invention can be provided in the form of "data structure" or "data having the above-mentioned structure" instead of the form of "data stream".</p><p> In other embodiments, the present invention provides a new programming model for transmitting data and applications over a network and for processing data and applications between members of the network. This programming model uses software cells that are transmitted over the network and can be processed by any member of the network. Each software cell has the same structure and can contain both applications and data. The high-speed processing and transmission speeds provided by the modular computer architecture allow for high-speed processing of these cells. The application code is preferably based on the same common instruction set and ISA. Each software cell preferably contains a global identifier (global ID) and information that describes the amount of computational resources required to process the cell. Since all computational resources have the same basic structure and use the same ISA, the specific resource that performs the processing of this cell can be placed anywhere on the network and can be dynamically allocated. it can.</p><p> The basic processing module is the processor element (PE). The PE preferably comprises a processing unit (PU), a direct memory access controller (DMAC) and a plurality of additional processing units (APU). In a preferred embodiment, one PE comprises eight APUs. The PU and APU communicate in real time using a shared dynamic random access memory (DRAM) that preferably has a crossbar architecture. The PU manages the schedule and general management of data and application processing by the APU. APU performs this process in parallel and independently. DMAC controls access to data and applications stored in shared DRAM by PU and APU.</p><p> According to this modular structure, the number of PEs used by a member of the network is based on the processing power required by that member. For example, one server can use four PEs, one workstation can use two PEs, and one PDA can use one PE. The number of PE APUs assigned to process a particular software cell depends on the complexity and size of the programs and data in that cell.</p><p> In a preferred embodiment, multiple PEs are associated with one shared DRAM. Preferably, the DRAM is divided into multiple sections, each of which is divided into multiple memory banks. In a particularly preferred embodiment, the DRAM has 64 memory banks, each bank having a storage capacity of 1 megabyte. Each section of DRAM is preferably controlled by a bank controller, and each DRAM of PE is preferably accessible to each bank controller. Therefore, the DMAC of each PE of this embodiment can access any part of the shared DRAM.</p><p> In another aspect, the invention provides a synchronization system and method for reading data from an APU from a shared DRAM and writing data to the shared DRAM. This system prevents conflicts between multiple APUs and multiple PEs that share DRAM. According to this system and method, a DRAM area is specified and multiple full-empty bits are stored. Each of these full-empty bits corresponds to a designated area of DRAM. Since this synchronization system is integrated into the DRAM hardware, it prevents the computational overhead of data synchronization schemes performed in the software.</p><p> Further, according to the present invention, a sandbox is provided in the DRAM to provide security against corruption of the program processing data of another APU caused by the program processing data of one APU. Each sandbox defines a shared DRAM area where data cannot be read or written.</p><p> In another aspect, the invention provides a system and method for a PU to issue commands to an APU to initiate processing of an application and data by the APU. These commands are called APU Remote Processing Instructions (ARPCs), which allow the PU to manage and coordinate the parallel processing of applications and data by the APU without the APU acting as a coprocessor. Become. In one embodiment of the invention, the computer processor has a plurality of first processing units, each said first processing unit comprising a local memory associated with the first processing unit, said first. The second processing unit has a second processing unit that controls the processing of the program by the processing unit and the data associated with the program, and the second processing unit transfers the data associated with the program and the program to the first processing unit. It is characterized in that the transfer instruction to the local memory is instructed, and the first processing unit can operate to process the program of the local memory of the first processing unit and the data associated with the program. A computer processor is also provided. Also, in a computer processor, there are a plurality of first processing units, each of which includes a local memory associated with the first processing unit, the program by the first processing unit and the said. It has a second processing unit that controls the processing of the software cell containing the data associated with the program, and the second processing unit reads and reads the information recorded in the software cell itself. According to the information, the processing unit that performs the processing of the software cell is designated from the plurality of first processing units, and the program and the data associated with the program are input to the first processing unit. It is characterized in that the transfer instruction to the local memory is instructed, and the first processing unit can operate to process the program of the local memory of the first processing unit and the data associated with the program. A computer processor is also provided. In another aspect of the invention, the processing apparatus has one or more processor modules having a plurality of first processing units, each of which is associated with the first memory. The program has a second processing unit that controls the processing of the program and the data associated with the program by the first processing unit, and the second processing unit associates the program with the program. The data is instructed to be transferred to the local memory of the first processing unit, and then the first processing unit can be operated to process the program and the data associated with the program from the local memory. A processing device characterized by being present is also provided. Also, in a processing apparatus, the processing apparatus has one or more processor modules having a plurality of first processing units, and each of the first processing units has a local memory associated with the first memory. It has a second processing unit that controls the processing of the software cell containing the program and the data associated with the program by the first processing unit, and the second processing unit records in the software cell itself. The processed information is read, and the processing unit that performs the processing of the software cell is designated from the plurality of first processing units according to the read information, and the program is associated with the program. The data is instructed to be transferred to the local memory of the first processing unit, and then the first processing unit can operate to process the program and the data associated with the program from the local memory. Processing equipment is also provided.</p><p> In other embodiments, the present invention provides a system and method for setting up a dedicated pipeline structure for streaming data processing. According to this system and method, an APU coordination group and a memory sandbox coordination group associated with these APUs are set up by the PU to process these streaming data. The dedicated pipeline APU and memory sandbox remain dedicated to the pipeline during periods of no data processing. In other words, the dedicated APUs and the sandboxes associated with these dedicated APUs are reserved during this period.</p><p> In another embodiment, the invention provides an absolute timer for task processing. This absolute timer is independent of the clock frequency used by the APU for processing applications and data. The application is written based on the time for the task, as defined by the absolute timer. Even if the clock frequency used by the APU increases due to improvements in the functions of the APU, the time for a predetermined task defined by the absolute timer remains the same. According to this method, it is not necessary to prevent the processing of old applications written on the assumption of slow processing time in the old APU from being performed by these new APUs, and the processing time is improved by the new version of APU. Will be possible.</p><p> The present invention also provides another method that enables a new APU with a higher processing speed to be used for processing an old application written on the premise of a slower processing speed in the old APU. In this method, the instructions (microcode) used by the APU when processing these older applications are analyzed during the handling of problems in coordinating the parallelism of the APU caused by the speed improvement. No operation (NOOP) instructions are inserted into the instructions executed by some of these APUs so that the order of processing by the APUs is maintained in the order that the program expects. By inserting these "NOOPs" into these instructions, the correct timing for executing all the instructions by the APU is maintained.</p><p> In another embodiment, the invention provides a chip package that includes an integrated circuit with an integrated optical waveguide.</p>
FIG. 1 shows the entire architecture of the computer system 101 according to the present invention.
As illustrated in this figure, the system 101 includes a network 104, to which multiple computers and computing devices are connected to the network. Examples of networks 104 include LANs, global networks such as the Internet, or other computer networks.
Among the computers and computing devices (members of the network) connected to the network 104 are client-side computers 106, server computers 108, personal information appliances (PDA) 110, digital television (DTV) 112 and others. Includes wired or wireless computers and computing devices. The processors used by the members of network 104 consist of the same common computing modules. Also, these processors preferably perform processing according to all the same ISAs and preferably the same instruction set. The number of modules contained within an individual processor is determined by the processing power required by that processor.
For example, server 108 in system 101 performs more data processing and application processing than client 106, so it will contain more computing modules than client 106. On the other hand, the PDA110 performs only the minimum amount of processing. Therefore, the PDA110 contains only a minimum number of computing modules. DTV112 performs a processing level between client 106 and server 108. Therefore, the DTV112 contains several computing modules between the client 106 and the server 108. As described below, each computing module includes a processing controller and a plurality of identical processing units that perform parallel processing of data and applications transmitted over the network 104.
This homogeneous configuration of system 101 improves adaptability, processing speed and processing efficiency. Each member of System 101 is one or more of the same compute modules (or part of a compute module) It doesn't matter on which computer or computing device the actual processing of the data and applications is performed, because the processing is performed using. In addition, individual application and data processing can be shared among members of the network. Throughout the system, by uniquely identifying the cells containing the data and applications processed by the system 101, the processing result is sent to the computer or computing device that requested the processing, regardless of where this processing was performed. It becomes possible to transmit. Since the modules that perform this process have a common structure and a common ISA, the computational burden of additional layers of software to achieve compatibility between processors is avoided. This architecture and programming model improves the processing speed required to run real-time multimedia applications and the like.
To take advantage of the additional benefits of processing speed and efficiency improved by system 101, the data and applications processed by this system are uniquely identified into software cells 102, each of the same format. It will be packaged. Each software cell 102 may or may contain both application and data. Each software cell also includes a cell identifier for identifying the cell within the network 104 and the entire system 101, for example, an ID that globally identifies the software cell. This structural uniformity of software cells and the unique identification of software cells within the network improves the processing of applications and data on any computer or computing device in the network. For example, the client 106 can create the software cell 102, but since the processing capacity on the client 106 side is limited, the software cell can be transmitted to the server 108 for processing. it can. Therefore, the software cell can move the entire network 104 and perform processing based on the availability of processing resources on the network.
Also, the homogeneous structure of the system 101 processor and software cells can prevent many of today's heterogeneous network problems. For example, inefficient programming models (such as virtual machines such as Java's virtual machines) that try to allow application processing on any ISA with any set of instructions are avoided. Therefore, the system 101 can realize wideband processing much more efficiently and much more effectively than today's networks.
The underlying processing module for all members of network 104 is the processor element (PE). Figure 2 illustrates the structure of PE. As shown in this figure, PE201 comprises a processing unit (PU) 203, DMAC205, and a plurality of additional processing units (APUs), namely APU207, APU209, APU211, APU213, APU215, APU217, APU219, APU221. The local PE bus 223 transmits data and applications between the APU, the DMAC205, and the PU203. The local PE bus 223 may have a conventional architecture or the like, or may be realized as a packet-switched network. When implemented as a packet-switched network, it requires more hardware, while increasing the available bandwidth.
PE201 can be configured using various methods to realize digital logic circuits. However, PE201 is preferably configured as a single integrated circuit on a silicon substrate. Alternative materials for substrates include gallium arsenide, gallium-aluminum arsenic, arsenic and other so-called III-B compounds using a wide variety of dopants. PE201 can also be realized using superconducting materials (such as fast single flux quantum (RSFQ) logic processing).
PE201 is closely associated with Dynamic Random Access Memory (DRAM) 225 via high bandwidth memory connectivity 227. DRAM225 functions as the main memory for PE201. Although the DRAM 225 is preferably a dynamic random access memory, other means, such as static random access memory (SRAM), include magnetic random access memory (MRAM), optical memory. Alternatively, DRAM 225 can be realized by using holography memory or the like. DMAC205 improves data transfer between DRAM225 and PE201's APU and PU. As will be further described below, DMAC205 specifies an exclusive area in DRAM225 for each APU, but only the APU can write data into this exclusive area, and only the APU can write this exclusive area. Data cannot be read from the target area. This exclusive area is called the "sandbox".
The PU203 may be a standard processor capable of stand-alone processing of data and applications. At the time of operation, the PU manages the schedule and general management of data and application processing by the APU. The APU is preferably a single instruction, multiple data (SIMD) processor. Under the control of PU203, APU processes these data and applications in parallel and independently. The DMAC205 controls access to the data and applications stored in the shared DRAM 225 by the PU203 and APU. Although it is desirable that the PE201 preferably contains eight APUs, a number of APUs slightly higher or lower than this number may be used in the PE depending on the required processing power. It is also possible to combine (package together) several PEs such as PE201 to improve processing power.
For example, as shown in FIG. 3, four PEs may be packaged (combined together) in one or more chip packages to form a single processor for members of network 104. This configuration is called a wideband engine (BE). As shown in FIG. 3, BE301 includes four PEs (PE303, PE305, PE307, PE309). Communication between these PEs takes place via the BE bus 311. Communication between the shared DRAM 315 and these PEs is performed by the broadband memory connection unit 313. Instead of the BE bus 311 the communication between the PEs of the BE301 can be done via the DRAM 315 and this memory connection.
The input / output (I / O) interface 317 and the external bus 319 communicate between the broadband engine 301 and the other members of the network 104. Each PE of BE301 executes data and application processing in a parallel and independent manner similar to the parallel and independent processing of application and data performed by the PE's APU.
FIG. 4 is a diagram illustrating the structure of the APU. The APU402 includes local memory 406, registers 410, four floating point arithmetic units 412 and four integer arithmetic units 414. However, again, depending on the processing power required, a number of floating-point arithmetic units 412 and integer arithmetic units 414 that are slightly higher or lower than four may be used. In one preferred embodiment, the local memory 406 includes a storage capacity of 128 kilobytes and the capacity of the register 410 is 128 x 128 bits. The floating point unit 412 operates suitably at a speed of 32 billion floating point operations (32 GLOPS) per second, and the integer operation unit 414 operates preferably at a speed of 32 billion operations (32 GOP) per second.
Local memory 406 is not cache memory. The local memory 406 is preferably configured as SRAM. No support for cache coherency, or cache integrity, for APU is required. PUs may also require cache integrity to support direct memory access (DMA) initiated by the PU. However, there is no need for cache integrity support for DMA initiated by the APU, or for access from and to external devices.
The APU 402 also includes a bus 404 for transmitting applications and data to and from the APU. In one preferred embodiment, the bus has a width of 1024 bits. The APU 402 also includes internal buses 408, 420 and 418. In one preferred embodiment, bus 408 has a width of 256 bits and communicates between local memory 406 and register 410. The buses 420 and 418 communicate between the register 410 and the floating-point arithmetic unit 412, and between the register 410 and the integer arithmetic unit 414, respectively. In one preferred embodiment, the width of the buses 418 and 420 from register 410 to floating-point unit 412 or integer arithmetic unit 414 is 384 bits, and the bus from floating-point arithmetic unit 412 or integer arithmetic unit 414 to register 410. The width of 418 and 420 is 128 bits. Wider data flow from register 410 is processed by the wider width of the above bus from register 410 to floating-point unit or integer arithmetic unit, wider than the width from floating-point arithmetic unit 412 or integer arithmetic unit 414 to register 410. Allowed during. A maximum of 3 words are required for each calculation. However, the result of each calculation is generally only one word.
Figures 5-10 further illustrate the modular structure of the processors of the members of network 104. For example, as shown in Figure 5, one processor can contain a single PE502. As mentioned above, this PE generally includes PU, DMAC and 8 APUs. Each APU contains local storage (LS). On the other hand, the processor may have the structure of the visualizer (VS) 505. As shown in FIG. 5, VS505 has PU512, DMAC514 and four APUs (APU516, APU518, APU520, APU522). The space in the chip package normally occupied by the other four APUs of the PE is in this case occupied by the pixel engine 508, the image cache 510 and the cathode ray tube controller (CRTC) 504. The optical interface 506 may be included in the chip package, depending on the communication speed required for the PE502 or VS505.
This standardized modular structure makes it possible to easily and efficiently configure modifications for many other processors. For example, the processor shown in FIG. 6 has two chip packages (chip package 602 with BE and chip package 604 with four VSs). Input / output (I / O) 606 provides an interface between BE in chip package 602 and network 104. Bus 608 communicates between chip package 602 and chip package 604. Data flow is controlled by the ingress and egress processor (IOP) 610 for input and output to and from I / O 606. I / O 606 can be manufactured as an ASIC (Application Specific Integrated Circit). The output from VS is the video signal 612.
FIG. 7 illustrates a chip package (or other locally connected chip package) for the BE702 with two optical interfaces 704 and 706 for ultrafast communication to other members of the network 104. The BE702 can function as a server or the like on the network 104.
The chip package of FIG. 8 has two PE802 and 804 and two VS806 and 808. The I / O 810 provides an interface between the chip package and the network 104. The output from the chip package is a video signal. This configuration can function as an image processing workstation or the like.
FIG. 9 illustrates yet another configuration. This configuration includes 1/2 of the processing power of the configuration illustrated in FIG. One PE902 will be provided in place of the two PEs, and one VS904 will be provided in place of the two VSs. The I / O 906 has half the bandwidth of the I / O illustrated in FIG. Such a processor can also function as an image processing workstation.
The final configuration is illustrated in FIG. This processor consists of only a single VS1002 and I / O 1004. This configuration can function as a PDA or the like.
FIG. 11 illustrates the integration of optical interfaces into the processor chip package of network 104. These optical interfaces convert the optical signal into an electrical signal and the electrical signal into an optical signal. In addition, these optical interfaces can be composed of various materials including gallium arsenide, aluminum / gallium arsenide, germanium and other elements and compounds. As shown in this figure, the optical interfaces 1104 and 1106 are assembled on top of the BE1102 chip package. The BE bus 1108 communicates with the PE of BE1102, namely PE1110, PE1112, PE1114, PE1116 and their optical interfaces. The optical interface 1104 contains two ports (port 1118 and port 1120), and the optical interface 1106 contains two ports (port 1122 and port 1124). Ports 1118, 1120, 1122, and 1124 are connected to optical waveguides 1126, 1128, 1130, and 1132, respectively. Optical signals are transmitted through these optical waveguides through the ports of optical interfaces 1104 and 1106 to and from BE1102.
A plurality of BEs may be collectively connected in various configurations by using such an optical waveguide and four optical ports of each BE. For example, as shown in FIG. 12, two or more BEs (BE1152, BE1154, BE1156, etc.) can be connected in series through such an optical port. In this example, the BE1152 optical interface 1166 is connected to the BE1154 optical interface 1160 optical port through its optical port. Similarly, the optical port of BE1154 optical interface 1162 is connected to the optical port of BE1156 optical interface 1164.
FIG. 13 illustrates a matrix configuration. In this configuration, the optical interface of each BE is connected to two other BEs. As shown in this figure, one of the optical ports of BE1172's optical interface 1188 is connected to the optical port of BE1176's optical interface 1182. The other optical port of the optical interface 1188 is connected to the optical port of the BE1178 optical interface 1184. Similarly, one optical port on the BE1174 optical interface 1190 is connected to the other optical port on the BE1178 optical interface 1184. The other optical port of the optical interface 1190 is connected to the optical port of the BE1180 optical interface 1186. This matrix structure can be extended to other BEs as well.
Either a serial configuration or a matrix configuration can be used to configure a processor for network 104 of any desired size and power. Needless to say, additional ports may be added to the BE's optical interface or to processors with fewer PEs than the BE to form other configurations.
FIG. 14 is a diagram illustrating the control system and structure of BE for DRAM. Similar control systems and structures are used in processors of different sizes and containing slightly different numbers of PEs. As shown in this figure, a crossbar switch connects each DMAC1210, which consists of four PEs with BE1201, to eight bank controls 1206. Each bank control 1206 controls eight banks 1208 (only four are shown) of DRAM 1204. Therefore, the DRAM 1204 will have a total of 64 banks. In a preferred embodiment, the DRAM 1204 has a capacity of 64 megabytes and each bank has a capacity of 1 megabyte. The smallest addressable unit in each bank is a 1024-bit block in this preferred embodiment.
The BE1201 also includes the switch unit 1212. Switch unit 1212 provides access to DRAM 1204 on other APUs in BE that are closely connected to BE1201. Therefore, it is possible to connect the second BE closely with the first BE, and each APU of each BE should address twice as many memory locations as the APU normally has access to. Is possible. Direct reading of data from the DRAM of the first BE to the DRAM of the second BE, or from the DRAM of the second BE to the DRAM of the first BE via a switch unit such as switch unit 1212. It is possible to directly write the data of.
For example, as shown in FIG. 15, in order to perform such a write, the APU of the first BE (such as the APU1220 of BE1222) causes the DRAM of the second BE (not the DRAM1224 of BE1222 as usual). , BE1226 DRAM1228, etc.) is issued with a write command to the memory location. The BE1222 DMAC1230 sends a write command to the bank control 1234 via the crossbar switch 1221, which in turn transmits the command to the external port 1232 connected to the bank control 1234. The BE1226 DMAC1238 receives a write command and forwards this command to the BE1226 switch unit 1240. The switch unit 1240 identifies the DRAM address contained in the write command and sends the data stored in the DRAM address to the bank 1244 of the DRAM 1228 via the bank control 1242 of BE1226. Therefore, the switch unit 1240 allows both DRAM 1224 and DRAM 1228 to function as a single memory space for the BE1222 APU.
FIG. 16 illustrates the configuration of 64 DRAM banks. These banks consist of eight rows (1302, 1304, 1306, 1308, 1310, 1312, 1314, 1316) and eight columns (1320, 1322, 1324, 1326, 1328, 1330, 1332, 1334). ing. Each row is controlled by one bank controller. Therefore, each bank controller controls 8 megabytes of memory.
Figures 17 and 18 illustrate different configurations for storing and accessing DRAM in the smallest addressable storage unit (such as a 1024-bit block). In Figure 17, the DMAC1402 stores eight 1024-bit blocks 1406 in a single bank 1404. In FIG. 18, DMAC1412 reads and writes data blocks containing 1024 bits, but these blocks are distributed between the two banks (bank 1414 and bank 1416). Therefore, each of these banks contains 16 blocks of data, and each block of data contains 512 bits. This distribution makes it possible to improve DRAM access even faster, which is useful for processing certain applications.
Figure 19 illustrates the architecture of DMAC1506 in PE. As illustrated in this figure, structural hardware, including the DMAC1506, is deployed through all PEs so that each APU1502 has direct access to the structural node 1504 of the DMAC1506. Each node executes logical processing suitable for memory access by the target APU that the node directly accesses.
FIG. 20 illustrates another embodiment of DMAC, namely a non-distributed architecture. In this case, the structural hardware of the DMAC1606 is centralized. The APU1602 and PU1604 communicate using the DMAC1606 via the local PE bus 1607. DMAC1606 is connected to bus 1608 via a crossbar switch. Bus 1608 is connected to DRAM 1610.
As mentioned above, all of the multiple APUs in a PE can independently access the data in the shared DRAM. As a result, the second APU may request these data while the first APU is processing some data in its local storage. If the data is output from the shared DRAM to the second APU at that time, the data may become invalid due to the ongoing processing of the first APU, which can change the value of the data. Therefore, if the second processor receives data from the shared DRAM at that time, an error result may occur in the second processor. For example, such data may include specific values for global variables. If the first processor changes its value during its processing, the second processor will receive a value that is no longer in use. Therefore, some method is needed to synchronize the reading and writing of data from and to the memory location by the APU within the range of the shared DRAM. In this method, reading from the memory location of the target data that another APU is currently working on in its local storage, and therefore not up-to-date, and in the memory location that stores the latest data. It is necessary not to write data to and to.
To solve these problems, for each addressable memory location in the DRAM, in the DRAM to store the state information related to the data stored in that memory location. Allocates additional memory segments. This state information includes the full / empty (F / E) bits, the APU identifier (APU ID) that requests the data from the memory location, and the local APU that is the read destination to read the requested data. Contains the storage address (LS address). The memory location where the DRAM can be addressed can be of any size. In one preferred embodiment this size is 1024 bits.
Setting the F / E bit to 1 indicates that the data stored in the memory location is up to date. On the other hand, setting the F / E bit to 0 indicates that the data stored in the associated memory location is not up to date. When this bit is set to 0, even if the APU requests the data, the APU prevents the data from being read immediately. In this case, the APU ID that identifies the APU requesting the data and, when the data is up-to-date, the memory location in the local storage of this APU to read the data from. The LS address to be used is entered into the additional memory segment.
Additional memory segments are also allocated for each memory location in the APU's local storage. This additional memory segment stores one bit called a "busy bit". This busy bit is used to reserve the associated LS memory location for storing unique data retrieved from DRAM. If the busy bit is set to 1 for a particular memory location in local storage, the APU can use this memory location only for writing these unique data. On the other hand, if the busy bit is set to 0 for a particular memory location in local storage, the APU can use this memory location for writing arbitrary data.
An example showing how the F / E bit, APU ID, LS address and busy bit are used to synchronize the reading and writing of data from the PE shared DRAM and to the PE shared DRAM is illustrated in the figure. Illustrated in 21-35.
As shown in Figure 21, one or more PEs (such as PE1720) use DRAM 1702. PE1720 includes APU1722 and APU1740. The APU1722 includes a control logic circuit 1724, and the APU1740 contains a control logic circuit 1742. The APU1722 also includes local storage 1726. This local storage contains multiple addressable memory locations 1728. The APU1740 contains local storage 1744, which also contains multiple addressable memory locations 1746. All of these addressable memory locations are preferably preferably 1024 bits in size.
An additional segment of memory is associated with an addressable memory location for each LS. For example, memory segments 1729 and 1734 are associated with local memory locations 1731 and 1732, respectively, and memory segment 1752 is associated with local memory location 1750. Busy bits as described above are stored within each of these additional memory segments. The local memory location 1732 is indicated with some crosses indicating that this memory location contains data.
The DRAM 1702 includes multiple addressable memory locations 1704, including memory locations 1706 and 1708. These memory locations are preferably 1024 bits in size. Additional segments of memory are also associated with each of these memory locations. For example, additional memory segment 1760 is associated with memory location 1706 and additional memory segment 1762 is associated with memory location 1708. The state information associated with the data stored in each memory location is stored in the memory segment associated with the memory location. This state information includes the F / E bit, APU ID and LS address as described above. For example, for memory location 1708, this state information includes the F / E bit 1712, APU ID 1714, and LS address 1716.
Using this state information and busy bits, reading from the shared DRAM and between synchronized shared DRAMs and writing data to the synchronized shared DRAM between APUs of PEs or one group of PEs. It can be performed.
FIG. 22 is a diagram illustrating the start of synchronous writing of data from the LS memory location 1732 of the APU1722 to the memory location 1708 of the DRAM 1702. Synchronous writing of these data is initiated by the control logic circuit 1724 of the APU1722. The F / E bit 1712 is set to 0 because memory location 1708 is empty. As a result, it is possible to write the data in LS memory location 1732 into memory location 1708. On the other hand, if this bit is set to 1 and memory location 1708 is in the full state, indicating that it contains the latest valid data, control circuit 1722 will receive an error message, which will result in an error message. Writing data to the memory location is prohibited.
The result of a successful synchronous write of data to memory location 1708 is shown in Figure 23. This written data is stored in memory location 1708 and the F / E bit 1712 is set to 1. This setting indicates that the memory location 1708 is full and that the data in this memory location is the most up-to-date valid data.
FIG. 24 is a diagram illustrating the start of a synchronous read of data from memory location 1708 of DRAM 1702 to LS memory location 1750 of local storage 1744. To initiate this read, the busy bit in memory segment 1752 of LS memory location 1750 is set to 1 and this memory location is reserved for the data. Setting this busy bit to 1 prevents the APU1740 from storing any other data in this memory location.
As shown in Figure 25, control logic 1742 then issues a synchronous read command to memory location 1708 on DRAM 1702. The F / E bit 1712 associated with this memory location is set to 1, so the data stored in memory location 1708 is considered to be up-to-date and valid data. As a result, the F / E bit 1712 is set to 0 when preparing to transfer data from memory location 1708 to LS memory location 1750. This setting is shown in Figure 26. Setting this bit to 0 indicates that the data in memory location 1708 will be invalid after reading these data.
As shown in FIG. 27, the data in memory location 1708 is then read from memory location 1708 to LS memory location 1750. FIG. 28 is a diagram showing the final state. A copy of the data in memory location 1708 is stored in LS memory location 1750. The F / E bit 1712 is set to 0 to indicate that the data in memory location 1708 is invalid. This invalidation is the result of the above data changes made by APU1740. The busy bit in memory segment 1752 is also set to 0. This setting indicates that the APU1740 can use the LS memory location 1750 for any purpose, that is, the LS memory location is no longer in a reserved state waiting to receive unique data. Therefore, the APU1740 can access the LS memory location 1750 for any purpose.
Figures 29-35 show that the F / E bit for the DRAM 1702 memory location is set to 0 and that the data at this memory location is neither up-to-date nor valid. Illustrates a synchronous read of data from a memory location in DRAM 1702 (such as memory location 1708) to an LS memory location in APU's local storage (such as LS memory location 1752 in local storage 1744), if any. ing. To initiate this transfer, the busy bit in memory segment 1752 of LS memory location 1750 is set to 1 and this LS memory location is reserved for this data transfer, as shown in Figure 29. .. As shown in FIG. 30, control logic 1742 then issues a synchronous read command to memory location 1708 in DRAM 1702. The F / E bit (F / E bit 1712) associated with this memory location is set to 0, so the data stored in memory location 1708 is invalid. As a result, the signal is transmitted to control logic 1742, which prevents immediate reading of data from this memory location.
As shown in Figure 31, the APU ID 1714 and the LS address 1716 for this read command are written into memory segment 1762. In this case, the APU ID for APU1740 and the LS memory location for LS memory location 1750 are written into memory segment 1762. Therefore, when the data in the range of memory location 1708 is up-to-date, this APU ID and LS memory location are used to determine the destination memory location to carry the latest data. To.
The data in memory location 1708 will be valid and up-to-date when the APU writes data into this memory location. Figure 29 illustrates the synchronous writing of data from memory location 1732 of the APU1722 into memory location 1708. This synchronous write of these data is allowed because the F / E bit 1712 for this memory location is set to 0.
After this write, the data in memory location 1708 will be the latest valid data, as shown in Figure 33. Therefore, the APUID 1714 and LS address 1716 obtained from memory segment 1762 are immediately read from memory segment 1762, and then this information is removed from this segment. The F / E bit 1712 is also set to 0 in anticipation of an immediate read of the data in memory location 1708. APU, as shown in Figure 34 Upon reading ID 1714 and LS address 1716, this information is immediately used to read valid data in memory location 1708 to LS memory location 1750 on APU1740. The final state is illustrated in Figure 35. This figure shows the valid data copied from memory location 1708 to memory location 1750, the busy bits in memory segment 1752 set to 0, and the F / in memory segment 1762 set to 0. The E-bit 1712 is illustrated. Setting this busy bit to 0 allows the APU1740 to access the LS memory location 1750 for any purpose. Setting this F / E bit to 0 indicates that the data in memory location 1708 is no longer up-to-date or valid.
FIG. 36 summarizes the above operations and the various states of the DRAM memory location, which are the states of the F / E bits, the APU ID, and the memory corresponding to the memory location. Based on the LS address stored in the segment. This memory location can have three states. These three states are the empty state 1880, where the F / E bit is set to 0 and no information is provided for the APU ID or LS address, and the F / E bit is set to 1 for the APU ID or LS address. On the other hand, there is a full state 1882 in which no information is provided, and a blocking state 1884 in which the F / E bit is set to 0 and information is provided for the APU ID and LS address.
As shown in this figure, in empty state 1880, synchronous write operations are allowed, resulting in a transition to full state 1882. However, when the memory location is empty, the data in the memory location is not up-to-date, resulting in a transition to blocking state 1884 for synchronous read operations.
In full state 1882, synchronous read operations are allowed, resulting in a transition to empty state 1880. On the other hand, synchronous write operations in full state 1882 are prohibited to avoid overwriting valid data. If such a write operation is attempted in this state, the state does not change and the error message is transmitted to the corresponding control logic of the APU.
Blocking state 1884 allows synchronous writing of data into memory locations, resulting in a transition to empty state 1880. On the other hand, synchronous read operations in blocking state 1884 are prohibited. This is to prevent a conflict with the previous synchronous read operation that caused this blocking state. If a synchronous read operation is attempted in blocking state 1884, the error message is transmitted to the corresponding control logic of the APU without any state change.
The above-mentioned method of synchronously reading data from the shared DRAM and synchronously writing data to the shared DRAM removes the calculation resource normally dedicated as a processor for reading data from the external device and writing data to the external device. It can also be used for. This input / output (I / O) function can also be performed by the PU. However, this change in synchronization scheme may be used by an APU running the appropriate program to perform this function. For example, using this method, a PU that receives an interrupt request for data transmission from an I / O interface initiated by an external device may delegate the processing of this request to this APU. The APU then issues a synchronous write command to the I / O interface. This interface now sends a signal to the external device that data can now be written into the DRAM. The APU then issues a synchronous read command to the DRAM, setting the DRAM's associated memory space to the blocking state. The APU also sets the busy bit to 1 for the memory location of the APU's local storage where it needs to receive data. In the blocking state, the ID of the APU and the address of the associated memory location of the APU's local storage are contained within the additional memory segment associated with the DRAM's associated memory space. The external device then issues a synchronous write command to write the data directly to the associated memory space of the DRAM. Since this memory space is in a blocking state, data is immediately read from this space into the memory location of the APU's local storage identified within the additional memory segment. The busy bits for these memory locations are then set to 0. When the external device completes the writing of data, the APU outputs a signal to the PU indicating that the transmission is completed.
Therefore, using this method, data transfer processing from an external device can be performed with the minimum computational load on the PU. However, it is desirable that the APU delegated this function be able to issue an interrupt request to the PU, and it is desirable that the external device directly access the DRAM.
Each PE's DRAM contains multiple "sandboxes". A sandbox defines a shared DRAM area, beyond which no particular APU or pair of APUs can read or write data. These sandboxes provide security against data corruption processed by another APU due to data processed by one APU. These sandboxes also allow software cells to download software cells from network 104 into a particular sandbox without the potential for data corruption in the entire DRAM. In the present invention, the sandbox is provided in hardware consisting of DRAM and DMAC. By providing these sandboxes within this hardware instead of software, you get the benefits of speed and security.
The PE PU controls the sandbox assigned to the APU. Since the PU normally runs only reliable programs such as the operating system, this method does not compromise security. According to this method, the PU builds and maintains the key management table. Figure 37 illustrates this key management table. As shown in this figure, each entry in the key management table 1902 contains an identifier (ID) 1904 for the APU, an APU key 1906 for that APU, and a key mask 1908. The use of this key mask will be described below. Key management table 1902 is a static random access memory (SRA)<u style="single">M</u>) Is preferably stored in relatively fast memory and associated with DMAC. Entry into key management table 1902 is controlled by the PU. When the APU requests to write data to a specific storage location (storage location) in DRAM or read data from a specific storage location in DRAM, the DMAC sends the memory access key associated with that storage location to the memory access key. On the other hand, the APU key 1906 assigned to the APU in the key management table 1902 is evaluated.
As shown in FIG. 38, a dedicated memory segment 2010 is allocated for each addressable storage location 2006 in DRAM 2002. The memory access key 2012 for this storage location is stored in this dedicated memory segment. As mentioned above, an additional dedicated memory segment 2008, also associated with each addressable storage location 2006, stores synchronization information for writing data to and reading data from the storage location. To.
Upon activation, the APU issues a DMA command to the DMAC. This command contains the address of DRAM2002 storage location 2006. Before executing this command, DMAC examines the requesting APU key 1906 using the APU ID 1904 in key management table 1902. Next, the DMAC is the memory access key 2012 stored in the dedicated memory segment 2010 associated with the storage location of the DRAM to which the APU requests access, and the APU key of the requesting APU. Compare with 1906. If the two keys do not match, the DMA command will not be executed. On the other hand, if the two keys match, the DMA command proceeds and the requested memory access is performed.
FIG. 39 shows an example of another embodiment. In this example, the PU also maintains memory access control table 2102. The memory access control table 2102 contains an entry for each sandbox in the DRAM. In the particular example of Figure 39, the DRAM contains 64 sandboxes. Each entry in the memory access control table 2102 contains a sandbox identifier (ID) 2104, a base memory address 2106, a sandbox size 2108, a memory access key 2110, and an access keymask. 2110 and are included. Base memory address 2106 provides an address in the DRAM, which indicates the first part of a particular memory sandbox. Sandbox size 2108 gives the size of the sandbox, and therefore this size gives the endpoint of a particular sandbox.
FIG. 40 is a flow chart showing the steps for executing a DMA command using the key management table 1902 and the memory access management table 2102. In step 2202, the APU issues a DMA command to the DMAC for access to one or more specific memory locations in the sandbox. This command contains sandbox ID 2104, which identifies the particular sandbox to which the access request is made. In step 2204, the DMAC utilizes the APU ID 1904 to look up the requesting APU key 1906 in the key management table 1902. At step 2206, the DMAC utilizes sandbox ID 2104 in a command that looks up the memory access key 2110 associated with the sandbox in memory access management table 2102. At step 2208, the DMAC compares the APU key 1906 assigned to the requesting APU with the access key 2110 associated with the sandbox. At step 2210, a decision is made as to whether the two keys match. If the two keys do not match, processing proceeds to step 2212, where the DMA command does not proceed and an error message is sent to the requesting APU and / or PU. On the other hand, if the two keys match in step 2210, the process proceeds to step 2214, where DMAC executes the DMA command.
Key masks for APU keys and memory access keys give the system great flexibility. The key mask for the key converts the masked bits to wildcards. For example, if the key mask 1908 associated with the APU key 1906 has its last two bits set to "mask", for example by setting these bits in the key mask 1908 to 1, then the APU The key can be either 1 or 0 and will just match the memory access key. For example, suppose the APU key is 1010. Normally, this APU key only allows access to sandboxes with 1010 access keys. However, if the APU key mask for this APU key is set to 0001, then this APU key can be used to access sandboxes with either a 1010 or 1011 access key. .. Similarly, an APU with an APU key of either 1010 or 1011 can access access key 1010 with a mask set to 0001. Since both the APU key mask and the memory key mask can be used at the same time, it is possible to set accessibility by APU for a large number of variations of sandboxes.
The present invention also provides a new programming model for the processor of System 101. Software cell 102 is used in this programming model. It is possible to transmit these cells for processing to any processor on the network 104. The new programming model also takes advantage of System 101's unique modular architecture and System 101's processors.
Software cells are processed by APU directly from APU's local storage. APU does not work directly on any data or program in DRAM. The data and programs in the DRAM are loaded into the APU's local storage before the APU processes these data and programs. Therefore, APU's local storage will contain program counters, a stack, and other software elements for running these programs. The PU controls the APU by issuing a DMA command to the DMAC.
The structure of software cell 102 is illustrated in FIG. As shown in this figure, software cells such as software cell 2302 include route selection information section 2304 and body portion 2306. The information contained in Route Selection Information Section 2304 is determined by the protocol of network 104. The route selection information section 2304 contains header 2308, destination ID 2310, source ID 2312 and response ID 2314. The destination ID includes the network address. Under the TCP / IP protocol, for example, a network address is an Internet Protocol (IP) address. Further, the destination ID 2310 contains the identifiers of the PE and APU of the transmission destination to which the cell should be transmitted for processing. Source ID 2314 contains the network address, which identifies the PE and APU, launches the cell from this PE and APU, and if necessary, the destination PE and APU provide additional information about the cell. It becomes possible to obtain. Response ID 2314 contains the network address, which identifies the PE and APU to which the query for the cell and the result of cell processing are sent.
The body part 2306 of the cell contains information unrelated to the network protocol. The disassembled portion of FIG. 41 illustrates the details of the cell body portion 2306. The header 2320 of the cell body portion 2306 identifies the start of the cell body. The cell interface 2322 contains the information necessary to use the cell. This information includes the globally unique ID2324, the required APU2326, the sandbox size 2328, and the ID2330 of the previous cell.
The globally unique ID 2324 uniquely identifies software cell 2302 throughout the network 104. A global unique ID 2324 is created based on the source ID 2312 (such as the unique identifier of the PE or APU in the source ID 2312) and the time and date of creation or transmission of software cell 2302. The required APU2326 gives the minimum number of APUs required to execute a cell. Sandbox size 2328 provides a protected amount of memory within the required APU associated with the DRAM required to execute the cell. The previous cell ID 2330 provides the identifier of the previous cell in a group of cells (such as streaming data) that requires sequential execution.
Execution section 2332 contains the core information of the cell. This information includes DMA command list 2334, program 2336, and data 2338. Program 2336 contains programs (called applets) executed by APU, such as APU programs 2360 and 2362 , and data 2338 contains data processed using these programs. The DMA command list 2334 contains a set of DMA commands needed to start the program. These DMA commands include the DMA commands 2340, 2350, 2355, 2358. The PU issues these DMA commands to the DMAC.
DMA command 2340 contains VID 2342. VID2342 is the virtual ID of the APU that is associated with the physical ID when the DMA command is issued. DMA command 2340 also includes load command 2344 and address 2346. Load command 2344 commands the APU to read certain information from DRAM and put it into local storage. Address 2346 gives a virtual address in the DRAM that contains this particular information. This particular information may be the program from program section 2336, data from data section 2338, or other data. Finally, the DMA command 2340 contains the address 2348 of the local storage. This address identifies the address of the local storage where the information is likely to be loaded. The DMA command 2350 contains similar information. Other DMA commands can also be used.
The DMA command list 2334 also contains a series of kick commands, such as kick commands 2355 and 2358. A kick command is a command that initiates processing of cells sent by the PU to the APU. The DMA kick command 2355 includes a virtual APU ID 2352, a kick command 2354, and a program counter 2356. Virtual APU ID 2352 identifies the target APU to kick, kick command 2354 gives the associated kick command, and program counter 2356 gives the address for the program counter to execute the program. The DMA kick command 2358 gives similar information to the same APU or another APU.
As mentioned above, the PU treats the APU as an independent processor, not as a coprocessor. Therefore, to control the processing by the APU, the PU uses a command similar to a remote procedure call. These commands are called "APU Remote Procedure Calls (ARPC)". The PU executes ARPC by issuing a series of DMA commands to DMAC. DMAC loads the APU program and its associated stack frames into APU's local storage. The PU then issues the first kick to the APU and executes the APU program.
Figure 42 illustrates the ARPC steps for running an applet. These steps performed by the PU at the start of processing the applet by the designated APU are shown in the first part 2402 of Figure 42, and the steps performed by the designated APU are shown in the second part 2404 of Figure 42. There is.
At step 2410, the PU evaluates the applet and then specifies the APU for processing the applet. At step 2412, the PU allocates space for the applet to run in the DRAM by issuing a DMA command to the DMAC that sets the required memory access keys for the sandbox. At step 2414, the PU enables the transmission of the applet completion signal in response to an interrupt request to the designated APU. At step 2418, the PU issues a DMA command to the DMAC that loads the applet from the DRAM into the APU's local storage. At step 2420, a DMA command is executed and the applet is read from DRAM to local storage. At step 2422, the PU issues a DMA command to the DMAC that loads the stack frame associated with the applet from the DRAM into the APU's local storage. At step 2423, the DMA command is executed and the stack frame is read from the DRAM to the APU's local storage. At step 2424, the PU assigns a key to the APU by DMAC to read data from one or more hardware sandboxes specified in step 2412 and to that one or more hardware sandboxes. Issue a DMA command that allows the APU to write data. At step 2426, the DMAC updates the key management table (KTAB) with the key assigned to the APU. At step 2428, the PU issues a DMA command kick to the APU to start processing the program. Depending on the particular applet, the PU may issue other DMA commands when running a particular ARPC.
As mentioned above, the second part 2404 of FIG. 42 illustrates the steps taken by the APU when the applet is executed. At step 2430, the APU begins executing the applet in response to the kick command issued at step 2428. At step 2432, at the direction of the applet, the APU evaluates the applet's associated stack frame. At step 2434, the APU issues multiple DMA commands to the DMAC to load the data that the stack frame specifies from the DRAM to the APU's local storage as needed. At step 2436, these DMA commands are executed and the data is read from the DRAM to the APU's local storage. At step 2438, APU runs the applet and prints a result. At step 2440, the APU issues a DMA command to the DMAC and stores the result in the DRAM. At step 2442, a DMA command is executed and the applet results are written from the APU's local storage to the DRAM. At step 2444, the APU issues an interrupt request to the PU to transmit a signal indicating that ARPC is complete.
The ability of APUs to perform tasks independently under the direction of the PU allows one group of APUs and the memory resources associated with one group of APUs to be dedicated to the execution of extended tasks. For example, one PU is dedicated to receiving data transmitted over one or more APUs and a group of memory sandboxes associated with these one or more APUs over network 104 over an extended period of time. And can also be dedicated to transmission to one or more other APUs and their associated memory sandboxes for further processing of the data received during this time. This capability is particularly suitable for processing streaming data (such as streaming MPEG or streaming ATRAC audio or video data) transmitted over network 104. The PU dedicates one or more APUs and their associated memory sandboxes to receive these data, and one or more other APUs and their associated memory sandboxes to decompress and process these data. Can be dedicated. In other words, the PU can establish a dedicated pipeline relationship between a group of APUs and their associated memory sandbox to perform such data processing.
However, in order to perform such processing efficiently, the pipeline's dedicated APU and memory sandbox remain dedicated to the pipeline during times when the applet containing the data stream is not processed. Is desirable. In other words, it is desirable that the dedicated APUs and their associated sandboxes be left in the reserved state during these hours. Reserving, or reserve, the APU and its associated memory sandbox at the end of applet processing is called "end of residence." Resident termination is performed in response to a command from the PU.
Figures 43, 44, and 45 illustrate the configuration of a dedicated pipeline structure for processing streaming data (such as streaming MPEG data), including a group of APUs and their associated sandboxes. As shown in FIG. 43, the components of this pipeline structure include PE2502 and DRAM2518. PE2502 includes multiple APUs including PU2504, DMAC2506 and APU2508, APU2510, APU2512. Communication between PU2504, DMAC2506 and these APUs is via PE bus 2514. The DMAC2506 is connected to the DRAM2518 by a broadband width bus 2516. The DRAM 2518 contains multiple sandboxes (sandbox 2520, sandbox 2522, sandbox 2524, sandbox 2526, etc.).
FIG. 44 illustrates the steps for setting up a dedicated pipeline. At step 2610, PU2504 assigns APU2508 to handle network applets. The network applet has a program for processing the network protocol of network 104. In this case, this protocol Transmission control protocol / Internet protocol (TCP / IP). TCP / IP data packets that follow this protocol are transmitted over network 104. Upon receipt, the APU2508 processes these packets, assembles the data in the packets and puts them into software cell 102. At step 2612, PU2504 instructs APU2508 to perform a resident termination when the processing of the network applet is complete. At step 2614, PU2504 assigns APU2510 and 2512 to process the MPEG applet. At step 2615, PU2504 instructs APU2510 and 2512 to perform resident termination when the processing of the MPEG applet is complete. At step 2616, PU2504 designates sandbox 2520 as the source sandbox for access by APU2510 and the destination sandbox for access by APU2508. At step 2618, PU2504 designates sandbox 2522 as the destination sandbox for access by APU2510 and the source sandbox for access by APU2512. At step 2620, PU2504 designates sandbox 2524 as the destination sandbox for access by APU2512 and the source sandbox for access by APU later in the pipeline. In step 2622, PU2504 designates sandbox 2526 as the destination sandbox for access by the APU later in the pipeline and the source sandbox for access. At step 2624, APU2510 and APU2512 send synchronous read commands to memory blocks within the range of source sandbox 2520 and source sandbox 2522, respectively, to set these memory blocks to the blocking state. Finally, the process moves to step 2628, where the dedicated pipeline setup is complete and the pipeline dedicated resources are reserved. Will be done. In this way, the APU2508, 2510, 2512, etc. and their associated sandboxes 2520, 2522, 2524 and 2526 enter the reserved state.
FIG. 45 illustrates the processing steps of streaming MPEG data by this dedicated pipeline. At step 2630, APU2508 processes a network applet and, within its local storage, receives TCP / IP data packets from network 104. At step 2632, APU2508 processes these TCP / IP data packets, assembles the data in these packets, and puts them into software cell 102. At step 2634, the APU2508 checks the software cell header 2320 (Figure 23) to determine if the cell contains MPEG data. If the cell does not contain MPEG data, in step 2636, APU2508 puts the cell into the generic sandbox specified in DRAM2518 to process other data by other APUs not contained in the dedicated pipeline. To transmit. The APU2508 also notifies the PU2504 about this transmission.
On the other hand, if the software cell contains MPEG data, in step 2638 the APU2508 will have the ID 2330 of the cell before that cell (Figure<u style="single">41</u>) To identify the MPEG data stream to which the cell belongs. At step 2640, APU2508 selects APU in a dedicated pipeline for processing cells. In this case, the APU2508 chooses the APU2510 to process these data. This selection is based on the previous cell ID 2330 and the load balancing factor. For example, if previous cell ID 2330 indicates that the previous software cell of the MPEG data stream to which the software cell belongs was sent to APU2510 for processing, then the current software cell is also for normal processing. Is sent to APU2510. At step 2642, APU2508 issues a synchronous write command to write MPEG data to sandbox 2520. Since this sandbox is preset to the blocking state, in step 2644 the MPEG data is automatically read from the sandbox 2520 to the local storage of the APU2510. At step 2646, the APU2510 processes MPEG data in its local storage to generate video data. At step 2648, the APU2510 writes video data to the sandbox 2522. At step 2650, the APU2510 issues a synchronous read command to the sandbox 2520, which prepares the sandbox for receiving additional MPEG data. At step 2652, the APU2510 performs the resident termination process. This process puts the APU into a reserved state, during which the APU waits in the MPEG data stream to process additional MPEG data.
Other dedicated structures can be configured between a group of APUs and their associated sandboxes for other types of data processing. For example, as shown in Figure 46, you can set up a dedicated group of APUs (APU2702, 2708, 2714, etc.) and perform geometric transformations on 3D objects to generate a 2D display list. It will be possible. Further processing these 2D display lists by other APUs (render) It is possible to generate pixel data. To perform this process, sandboxes are dedicated to the APU2702, 2708, 2414 for storing 3D objects and the resulting display list. For example, the source sandboxes 2704, 2710, and 2716 are dedicated to storing 3D objects processed by APU2702, APU2708, and APU2714, respectively. Similarly, the destination sandboxes 2706, 2712, and 2718 are dedicated to storing the display list that results from the processing of these 3D objects by APU2702, APU2708, and APU2714, respectively.
The tuning APU2720 is dedicated to receiving display lists from destination sandboxes 2706, 2712, 2718 in its local storage. The APU2720 makes adjustments between these display lists and sends these display lists to other APUs for rendering pixel data.
The system 101 processor also uses an absolute timer. This absolute timer outputs a clock signal to the other elements of the APU and PE. This clock signal does not depend on the clock signal that drives these elements, and is faster than this clock signal. The use of this absolute timer is shown in the figure<u style="single">47</u>Illustrated in.
As shown in this figure, this absolute timer determines the time budget for task performance by the APU. This time budget sets the completion time for these tasks, which is longer than the time required for the APU to process the task. As a result, for each task, there will be a busy time and a standby time within the time budget. All applets are written to work on this time budget, regardless of the actual processing time of the APU.
For example, a specific task can be performed during the busy time 2802 of Time Budget 2804 for a specific APU of PE. Standby time 2806 occurs during the time budget because busy time 2802 is less than time budget 2804. During this standby time, the APU goes into sleep mode, which consumes less power.
Until the time budget 2804 expires, other elements of other APUs or PEs do not anticipate the outcome of task processing. Therefore, regardless of the actual processing speed of the APU, the processing result of the APU is constantly adjusted using the time budget determined by the absolute timer.
In the future, the processing speed by APU will be even faster. However, the time budget set by the absolute timer remains the same. For example, figure<u style="single">47</u>As shown in, future APUs will perform tasks in a shorter amount of time, and therefore standby times will be even longer. Therefore, the busy time 2808 is shorter than the busy time 2802 and the standby time 2810 is longer than the standby time 2806. However, since the program is written to perform processing based on the same time budget set by the absolute timer, the coordination of processing results between APUs is maintained. As a result, the faster APU can process the program written for the slower APU without causing a conflict when the result of the processing is expected.
For the adjustment problem of parallel processing of APU caused by the improvement of operation speed or the difference in operation speed, instead of the absolute timer that determines the adjustment between APUs, the APU is used in the PU or one or more specified APUs. It is also possible to have the analysis of the specific instruction (microcode) being executed performed during the processing of the applet. It is possible to insert a no operation (NOOP) instruction into the instruction and execute this instruction by some of the APUs to properly perform the APU processing expected by the applet step by step. .. By inserting these NOOPs into the instructions, it is possible to maintain the correct timing for the APU to execute all the instructions.
Although the present invention has been described above with respect to specific embodiments, it should be understood that these embodiments are merely exemplary to illustrate the principles and applications of the invention. Therefore, it is possible to make numerous modifications to the above illustrated embodiments without departing from the spirit and scope of the invention as defined by the appended claims, as well as other configurations. It is possible to devise.
<figref num="1">The entire architecture of the computer network according to the present invention is illustrated.</figref><figref num="2">It is a figure which illustrates the structure of the processor element (PE) by this invention.</figref><figref num="3">It is a figure which illustrates the structure of the wide band engine (BE) by this invention.</figref><figref num="4">It is a figure which illustrates the structure of the additional processing unit (APU) by this invention.</figref><figref num="5">It is a figure which illustrates the structure of the processor element by this invention, a visualizer (VS), and an optical interface.</figref><figref num="6">It is a figure which illustrates one combination of the processor elements by this invention.</figref><figref num="7">FIG. 5 illustrates another combination of processor elements according to the present invention.</figref><figref num="8">FIG. 5 illustrates yet another combination of processor elements according to the present invention.</figref><figref num="9">FIG. 5 illustrates yet another combination of processor elements according to the present invention.</figref><figref num="10">FIG. 5 illustrates yet another combination of processor elements according to the present invention.</figref><figref num="11">It is a figure which illustrates the integration of the optical interface in a chip package by this invention.</figref><figref num="12">It is a figure which shows one configuration of the processor which uses the optical interface of FIG.</figref><figref num="13">It is a figure which shows another configuration of the processor which uses the optical interface of FIG.</figref><figref num="14">It is a figure which illustrates the structure of the memory system by this invention.</figref><figref num="15">It is a figure which illustrates the writing of the data from the 1st wideband engine to the 2nd wideband engine by this invention.</figref><figref num="16">It is a figure which shows the structure of the shared memory for the processor element by this invention.</figref><figref num="17">It is a figure which illustrates one structure for the memory bank shown in FIG.</figref><figref num="18">FIG. 5 illustrates another structure for the memory bank shown in FIG.</figref><figref num="19">It is a figure which illustrates the structure for DMAC by this invention.</figref><figref num="20">It is a figure which illustrates the alternative structure for DMAC by this invention.</figref><figref num="21">It is a figure which illustrates the data synchronization operation by this invention.</figref><figref num="22">It is a figure which illustrates the data synchronization operation by this invention.</figref><figref num="23">It is a figure which illustrates the data synchronization operation by this invention.</figref><figref num="24">It is a figure which illustrates the data synchronization operation by this invention.</figref><figref num="25">It is a figure which illustrates the data synchronization operation by this invention.</figref><figref num="26">It is a figure which illustrates the data synchronization operation by this invention.</figref><figref num="27">It is a figure which illustrates the data synchronization operation by this invention.</figref><figref num="28">It is a figure which illustrates the data synchronization operation by this invention.</figref><figref num="29">It is a figure which illustrates the data synchronization operation by this invention.</figref><figref num="30">It is a figure which illustrates the data synchronization operation by this invention.</figref><figref num="31">It is a figure which illustrates the data synchronization operation by this invention.</figref><figref num="32">It is a figure which illustrates the data synchronization operation by this invention.</figref><figref num="33">It is a figure which illustrates the data synchronization operation by this invention.</figref><figref num="34">It is a figure which illustrates the data synchronization operation by this invention.</figref><figref num="35">It is a figure which illustrates the data synchronization operation by this invention.</figref><figref num="36">It is a memory diagram of three states exemplifying various states of a memory location by the data synchronization method of this invention.</figref><figref num="37">It is a figure which illustrates the structure of the key management table for a hardware sandbox by this invention.</figref><figref num="38">It is a figure which illustrates the storage method of the memory access key for a hardware sandbox by this invention.</figref><figref num="39">It is a figure which illustrates the structure of the memory access management table for a hardware sandbox by this invention.</figref><figref num="40">It is a flow chart which shows the step of accessing a memory sandbox using the key management table of FIG. 37 and the memory access management table of FIG. 39.</figref><figref num="41">It is a figure which illustrates the structure of the software cell by this invention.</figref><figref num="42">It is a flow chart which shows the step of issuing a remote processing instruction to APU by this invention.</figref><figref num="43">It is a figure which illustrates the structure of the pipeline for exclusive use for streaming data processing by this invention.</figref><figref num="44">FIG. 5 is a flow chart showing the steps performed by the dedicated pipeline of FIG. 43 when processing streaming data according to the present invention.</figref><figref num="45">FIG. 5 is a flow chart showing the steps performed by the dedicated pipeline of FIG. 43 when processing streaming data according to the present invention.</figref><figref num="46">It is a figure which illustrates other structure of the dedicated pipeline for streaming data processing by this invention.</figref><figref num="47">It is a figure which illustrates the absolute timer system for adjusting the parallel processing of an application and data by APU by this invention.</figref>
Code description
101 system 1010 key 102 cells 104 network 106 client 108 server computer 1104 optical interface 1108 bus 1118 port 1122 port 1126 Optical Waveguide 1160 optical interface 1162 optical interface 1164 optical interface 1166 Optical interface 1182 optical interface 1184 optical interface 1186 optical interface 1188 optical interface 1188 optical interface 1190 optical interface 1190 optical interface 1206 control 1212 units 1221 Crossbar exchange 1232 External port 1234 control 1240 unit 1242 control 1244 bank 1406 block 1414 bank 1416 bank 1504 node 1607 bus 1608 bus 1722 control circuit 1724 Control logic circuit 1726 storage 1728 Location 1729 segment 1731 location 1732 location 1742 Control logic circuit 1746 location 1750 location 1752 segment 1760 segment 1762 segment 1880 Empty state 1882 Full state 1884 Blocking state 1902 Key management table 1906 key 1908 mask 2006 Storage location 2008 segment 2010 segment 2012 key 2102 Access control table 2106 address 2110 key 2110 key mask 223 bus 227 High bandwidth memory connection 2302 cell 2308 header 2320 header 2322 interface 2332 Execution section 2334 list 2520 sandbox 2522 sandbox 2524 sandbox 2526 sandbox 2704 sandbox 2706 Destination sandbox 301 wideband engine 311 bus 313 Broadband memory connection 317 interface 319 External bus 406 memory 408 internal bus 410 register 412 Floating point unit 414 Integer arithmetic unit 420 bus Optical interface in 506 package 508 engine 510 image cache
47 sheets
Sheet 1 Sheet 2 Sheet 3 Sheet 4 Sheet 5 Sheet 6 Sheet 7 Sheet 8 Sheet 9 Sheet 10 Sheet 11 Sheet 12 Sheet 13 Sheet 14 Sheet 15 Sheet 16 Sheet 17 Sheet 18 Sheet 19 Sheet 20 Sheet 21 Sheet 22 Sheet 23 Sheet 24 Sheet 25 Sheet 26 Sheet 27 Sheet 28 Sheet 29 Sheet 30 Sheet 31 Sheet 32 Sheet 33 Sheet 34 Sheet 35 Sheet 36 Sheet 37 Sheet 38 Sheet 39 Sheet 40 Sheet 41 Sheet 42 Sheet 43 Sheet 44 Sheet 45 Sheet 46 Sheet 47
Every citation, both ways
| Document | Relation | Office |
|---|---|---|
| JP05242057A | Cites | Japan |
| JP57176456A | Cites | Japan |
| JP63019058A | Cites | Japan |
| JP05151183A | Cites | Japan |
| JP10269165A | Cites | Japan |
| JP08161283A | Cites | Japan |
| JP08235143A | Cites | Japan |
| JP56123051A | Cites | Japan |
| JP08180018A | Cites | Japan |
136 members in 10 offices
Priority claims5
| Document | Office | Kind | Date |
|---|---|---|---|
| 09816752 | United States of America | – | |
| 81675201 | United States of America | A | |
| 81675201 | United States of America | A | |
| 2001816752 | – | – | – |
| US20010816752 | – | – | – |
Members136
| Document | Office | Kind | |
|---|---|---|---|
| US2002135582A1 | United States of America | A1 | |
| US2002138637A1 | United States of America | A1 | |
| US2002138701A1 | United States of America | A1 | |
| US2002138707A1 | United States of America | A1 | |
| WO02077826A1 | World Intellectual Property Organization (WIPO) | A1 | |
| WO02077838A1 | World Intellectual Property Organization (WIPO) | A1 | |
| WO02077845A1 | World Intellectual Property Organization (WIPO) | A1 | |
| WO02077846A1 | World Intellectual Property Organization (WIPO) | A1 | |
| WO02077848A1 | World Intellectual Property Organization (WIPO) | A1 | |
| US2002156993A1 | United States of America | A1 | |
| JP2002342165A | Japan | A | |
| JP2002351850A | Japan | A | |
| JP2002358289A | Japan | A | |
| JP2002366533A | Japan | A | |
| JP2002366534A | Japan | A | |
| US6526491B2 | United States of America | B2 | |
| JP3411273B2 | Japan | B2 | |
| JP2003271570A | Japan | A | |
| TW556087B | Taiwan Province of China | B | |
| JP2003281107A | Japan | A | |
| JP3454808B2 | Japan | B2 | |
| KR20030081532A | Republic of Korea | A | |
| JP2003303134A | Japan | A | |
| KR20030085037A | Republic of Korea | A | |
| KR20030085038A | Republic of Korea | A | |
| KR20030086319A | Republic of Korea | A | |
| KR20030086320A | Republic of Korea | A | |
| US2003229765A1 | United States of America | A1 | |
| EP1370948A1 | European Patent Office (EPO) | A1 | |
| EP1370961A1 | European Patent Office (EPO) | A1 | |
| EP1370968A1 | European Patent Office (EPO) | A1 | |
| EP1370969A1 | European Patent Office (EPO) | A1 | |
| EP1370971A1 | European Patent Office (EPO) | A1 | |
| JP3483877B2 | Japan | B2 | |
| TW574653B | Taiwan Province of China | B | |
| JP2004046861A | Japan | A | |
| JP2004078979A | Japan | A | |
| JP3515985B2 | Japan | B2 | |
| CN1494690A | China | A | |
| CN1496511A | China | A | |
| CN1496516A | China | A | |
| CN1496517A | China | A | |
| CN1496518A | China | A | |
| TW594492B | Taiwan Province of China | B | |
| JP2004252990A | Japan | A | |
| US6809734B2 | United States of America | B2 | |
| US6826662B2 | United States of America | B2 | |
| TWI227401B | Taiwan Province of China | B | |
| US2005078117A1 | United States of America | A1 | |
| US2005081181A1 | United States of America | A1 | |
| US2005081209A1 | United States of America | A1 | |
| US2005081213A1 | United States of America | A1 | |
| US2005097302A1 | United States of America | A1 | |
| US2005120187A1 | United States of America | A1 | |
| US2005120254A1 | United States of America | A1 | |
| US2005138325A1 | United States of America | A1 | |
| US2005160097A1 | United States of America | A1 | |
| JP3696563B2 | Japan | B2 | |
| US2005268048A1 | United States of America | A1 | |
| WO2006038714A2 | World Intellectual Property Organization (WIPO) | A2 | |
| WO2006038717A2 | World Intellectual Property Organization (WIPO) | A2 | |
| JP2006107513A | Japan | A | |
| JP2006107514A | Japan | A | |
| WO2006038714A3 | World Intellectual Property Organization (WIPO) | A3 | |
| US7093104B2 | United States of America | B2 | |
| TW200629053A | Taiwan Province of China | A | |
| US2006190614A1 | United States of America | A1 | |
| TW200634553A | Taiwan Province of China | A | |
| CN1279469C | China | C | |
| CN1279470C | China | C | |
| TWI266200B | Taiwan Province of China | B | |
| US7139882B2 | United States of America | B2 | |
| CN1291327C | China | C | |
| WO2006038717A3 | World Intellectual Property Organization (WIPO) | A3 | |
| WO2006038717B1 | World Intellectual Property Organization (WIPO) | B1 | |
| US7231500B2 | United States of America | B2 | |
| US7233998B2 | United States of America | B2 | |
| KR20070064432A | Republic of Korea | A | |
| EP1805626A2 | European Patent Office (EPO) | A2 | |
| US2007168538A1 | United States of America | A1 | |
| US2007186077A1 | United States of America | A1 | |
| EP1370948A4 | European Patent Office (EPO) | A4 | |
| CN101040268A | China | A | |
| US2007288701A1 | United States of America | A1 | |
| KR100840113B1 | Republic of Korea | B1 | |
| US7392511B2 | United States of America | B2 | |
| US2008162877A1 | United States of America | A1 | |
| KR100847982B1 | Republic of Korea | B1 | |
| CN100412848C | China | C | |
| US2008250414A1 | United States of America | A1 | |
| US2008256275A1 | United States of America | A1 | |
| KR100866739B1 | Republic of Korea | B1 | |
| US7457939B2 | United States of America | B2 | |
| KR20080108588A | Republic of Korea | A | |
| US7496673B2 | United States of America | B2 | |
| EP1370969A4 | European Patent Office (EPO) | A4 | |
| EP1370971A4 | European Patent Office (EPO) | A4 | |
| EP1370961A4 | European Patent Office (EPO) | A4 | |
| KR100890134B1 | Republic of Korea | B1 | |
| US7509457B2 | United States of America | B2 |
34 legal events, as the office reported them to INPADOC
Over the term
Point at a mark for the eventEvents
| Event | Code | |
|---|---|---|
| Cancellation because of completion of termEXPY | EXPY | |
| Receipt of annual feesJAPANESE INTERMEDIATE CODE: R250R250 | R250 | |
| Receipt of annual feesJAPANESE INTERMEDIATE CODE: R250R250 | R250 | |
| Receipt of annual feesJAPANESE INTERMEDIATE CODE: R250R250 | R250 | |
| Receipt of annual feesJAPANESE INTERMEDIATE CODE: R250R250 | R250 | |
| Receipt of annual feesJAPANESE INTERMEDIATE CODE: R250R250 | R250 | |
| Receipt of annual feesJAPANESE INTERMEDIATE CODE: R250R250 | R250 | |
| Receipt of annual feesJAPANESE INTERMEDIATE CODE: R250R250 | R250 | |
| Receipt of annual feesJAPANESE INTERMEDIATE CODE: R250R250 | R250 | |
| Receipt of annual feesJAPANESE INTERMEDIATE CODE: R250R250 | R250 | |
| Renewal fee payment (event date is renewal date of database)FPAY | FPAY | |
| Certificate of patent or registration of utility modelJAPANESE INTERMEDIATE CODE: R150R150 | R150 | |
| Certificate of patent or registration of utility modelJAPANESE INTERMEDIATE CODE: R150R150 | R150 | |
| First payment of annual fees (during grant procedure)JAPANESE INTERMEDIATE CODE: A61A61 | A61 | |
| Written decision to grant a patent or to grant a registration (utility model)JAPANESE INTERMEDIATE CODE: A01A01 | A01 | |
| Written decision to grant a patent or to grant a registration (utility model)JAPANESE INTERMEDIATE CODE: A01A01 | A01 | |
| Decision of grant or rejection writtenTRDD | TRDD | |
| Notification of resignation of power of attorneyJAPANESE INTERMEDIATE CODE: A7424RD04 | RD04 | |
| Request for written amendment filedJAPANESE INTERMEDIATE CODE: A523A521 | A521 | |
| Notification of resignation of power of attorneyJAPANESE INTERMEDIATE CODE: A7424RD04 | RD04 | |
| Notification of reasons for refusalJAPANESE INTERMEDIATE CODE: A131A131 | A131 | |
| Request for written amendment filedJAPANESE INTERMEDIATE CODE: A523A521 | A521 | |
| Notification of reasons for refusalJAPANESE INTERMEDIATE CODE: A131A131 | A131 | |
| Request for written amendment filedJAPANESE INTERMEDIATE CODE: A523A521 | A521 | |
| Written permission of extension of timeJAPANESE INTERMEDIATE CODE: A602A602 | A602 | |
| Written request for extension of timeJAPANESE INTERMEDIATE CODE: A601A601 | A601 | |
| Notification of reasons for refusalJAPANESE INTERMEDIATE CODE: A131A131 | A131 | |
| Report on retrievalJAPANESE INTERMEDIATE CODE: A971007A977 | A977 | |
| Request for written amendment filedJAPANESE INTERMEDIATE CODE: A523A521 | A521 | |
| Written request for application examinationJAPANESE INTERMEDIATE CODE: A621A621 | A621 | |
| Request for written amendment filedJAPANESE INTERMEDIATE CODE: A821A521 | A521 | |
| Request for written amendment filedJAPANESE INTERMEDIATE CODE: A523A521 | A521 | |
| Notification of change in applicantJAPANESE INTERMEDIATE CODE: A711A711 | A711 | |
| Notification of appointment of power of attorneyJAPANESE INTERMEDIATE CODE: A7423RD03 | RD03 |
Numbers
- Publication
- 4597553
- Publication, DOCDB
- 4597553
- Publication, EPODOC
- JP4597553B
- Application
- 63697
- Application, DOCDB
- 2004063697
- Application, EPODOC
- JP20040063697
Titles2
- Japanese
- コンピュータ・プロセッサ及び処理装置
- English
- Computer processor and processing equipment
Classification
- CPC, 2
- G06F9/4843
- G06F15/177
- IPC, 7
- G06F15 177
- G06F15 80
- G06F9 54
- G06F9 48
- G06F12 06
- G06F15 16
- G06F15 173
