On-demand multi-thread multimedia processor
Abstract
Devices include multimedia processors that can simultaneously support multiple applications for various types of multimedia such as graphics, audio, video, cameras, games, and so on. A multimedia processor includes a storage resource that can be configured to store information about an application, data and state information, and an assignable processor for performing various types of processing about the application. Configurable storage resources include an instruction cache for accumulating instructions about the application, a register bank for accumulating data about the application, a context register for accumulating state information about threads of the application, and the like. The processing device includes an arithmetic logic unit (ALU) core, an elementary function core, a logic core, a texture sampler, a load control device, a flow controller, and the like. The multimedia processor allocates a configurable portion of storage resources to each application and dynamically allocates a processor to the application when requested by these applications. [Selection diagram] Fig. 1

Term
1.4 yearsto projected expiry
Projected expiry 21 February 2028, counted from filing; an application has no term until it is granted.
- Priority
- Filed
- Published
- Today
- Projected expiry
27 claims: 7 independent, 20 dependent
- 1複数のアプリケーションを同時にサポートするためのマルチメディア・プロセッサを含むデバイスであって、前記マルチメディア・プロセッサは、 前記複数のアプリケーションに関する命令、データ、状態情報を蓄積するための構成可能な記憶リソースと、 前記複数のアプリケーションに関する処理を実行するための割り当て可能な処理装置と、 を具備し、 前記マルチメディア・プロセッサは、各アプリケーションに前記記憶リソースの構成可能な一部分を割り当て、前記アプリケーションによって要求された際、前記複数のアプリケーションに前記処理装置を動的に割り当てることを特徴とするデバイス。
- 2前記構成可能な記憶リソースは、前記複数のアプリケーションに関するデータを格納するための命令キャッシュを有し、 各アプリケーションは、前記命令キャッシュの構成可能な一部分を割り当てられることを特徴とする請求項1に記載のデバイス。
- 3前記構成可能な記憶リソースは、前記複数のアプリケーションに関するデータを格納するためのレジスタ・バンクを有し、 各アプリケーションは、前記レジスタ・バンクの構成可能な一部分に割り当てられることを特徴とする請求項1に記載のデバイス。
- 4前記構成可能な記憶リソースは、 前記複数のアプリケーションに関する命令、又はデータを格納するための記憶装置であって、仮想メモリ及び物理メモリと関連付けられる記憶装置と、なお、各アプリケーションは、前記仮想メモリの構成可能な一部分を割り当てられる、 各アプリケーションに対して割り当てられた前記仮想メモリの一部分を前記物理メモリの対応した一部分にマッピングするための少なくとも1つのテーブルと、を有することを特徴とする請求項1に記載のデバイス。
- 5前記マルチメディア・プロセッサは、前記アプリケーションに割り当てられた前記記憶リソースの一部分内の記憶に関するキャッシュ・メモリ、又はメイン・メモリから、要求された際に、各アプリケーションに関する命令及びデータを取り出すためのロード制御装置を更に具備することを特徴とする請求項1に記載のデバイス。
- 6前記処理装置は、独立に動作し、各処理装置は、与えられたタイム・スロット内で前記複数のアプリケーションの任意の一つに対して割り当て可能であることを特徴とする請求項1に記載のデバイス。
- 7前記マルチメディア・プロセッサは、前記マルチメディア・アプリケーションの任意の一つに割り当てられなかった処理装置の電力を落とすことを特徴とする請求項1に記載のデバイス。
- 8前記処理装置は、演算論理装置(ALU)コア、初等関数コア、論理コア、テクスチャ・サンプラ、ロード制御装置及びフロー・コントローラの少なくとも1つを有することを特徴とする請求項1に記載のデバイス。
- 9前記マルチメディア・プロセッサは、前記複数のアプリケーションからスレッドを非同期に受信するためのインプット・インターフェース装置と、 前記複数のアプリケーションに結果を非同期に供給するためのアウトプット・インターフェース装置と、 を更に具備することを特徴とする請求項1に記載のデバイス。
- 10前記マルチメディア・プロセッサは、電力消費を削減するために前記マルチメディア・プロセッサの読み込みに基づいてクロックの速度を調節することを特徴とする請求項1に記載のデバイス。
- 11前記マルチメディア・プロセッサは、前記処理装置がある特定の時間間隔に割り当てられた時間の割合に基づいて読み込みを決定することを特徴とする請求項10に記載のデバイス。
- 12前記マルチメディア・プロセッサは、前記複数のアプリケーションに関する複数のスレッドを同時にサポートすることを特徴とする請求項12に記載のデバイス。
- 13前記構成可能な記憶リソースは、前記複数のスレッドに関する複数のコンテキスト・レジスタを具備し、 各コンテキスト・レジスタは、関連付けられたスレッドに関する状態情報を格納することを特徴とする請求項12に記載のデバイス。
- 14各スレッドに関する前記状態情報は、プログラム・カウンタ、ダイナミック・フロー・コントロールに関するポインタを格納するためのスタック、関連したアドレス指定に関するアドレス・レジスタ、条件の計算結果を格納するための述語レジスタ及びロード要求及びデータバック条件を追跡するためのロード参照カウンタの少なくとも1つを有することを特徴とする請求項13に記載のデバイス。
- 15。 各スレッドは、前もって決められたサイズまでのデータ単位で処理することを特徴とする請求項12に記載のデバイス。
- 16各スレッドはイメージ内の4つのピクセルまで、又は4つの頂点までで処理することを特徴とする請求項12に記載のデバイス。
- 17前記マルチメディア・プロセッサは、前記複数のアプリケーションに適用可能な単一の命令セットをサポートすることを特徴とする請求項1に記載のデバイス。
- 18前記複数のアプリケーションは、グラフィックス・アプリケーション、オーディオ・アプリケーション、ビデオ・アプリケーション、カメラ・アプリケーション、ゲーム・アプリケーションの少なくとも一つを有することを特徴とする請求項1に記載のデバイス。
- 19複数のアプリケーションを同時にサポートすることと、 前記アプリケーションに関する命令、データ及び状態情報を蓄積するため各アプリケーションに記憶リソースの構成可能な一部を割り当てることと、 前記アプリケーションによって要求された際、前記複数のアプリケーションに処理装置を動的に割り当てることと、 を具備することを特徴とする方法。
- 20前記割り当てることは、前記アプリケーションに関する命令を蓄積するために各アプリケーションに命令キャッシュの構成可能な一部分を割り当てることを具備することを特徴とする請求項19に記載の方法。
- 21前記割り当てることは、前記アプリケーションに関するデータを蓄積するために各アプリケーションにレジスタ・バンクの構成可能な一部分を割り当てることを具備することを特徴とする請求項19に記載の方法。
- 22前記複数のアプリケーションからスレッドを非同期に受信することと、 前記複数のアプリケーションに非同期に結果を与えることと、 を更に具備することを特徴とする請求項19に記載の方法。
- 23複数のアプリケーションを同時にサポートする手段と、 前記アプリケーションに関する情報、データ、状態情報を格納するために各アプリケーションに記憶リソースの構成可能な一部分を割り当てる手段と、 前記アプリケーションによって要求された際、前記複数のアプリケーションに処理装置を動的に割り当てる手段と、 を具備することを特徴とする装置。
- 24割り当てるための前記手段は、前記アプリケーションに関する命令を蓄積するために各アプリケーションに命令キャッシュの構成可能な一部分を割り当てるための手段を具備することを特徴とする請求項23に記載の装置。
- 25前記割り当てるための前記手段は、前記アプリケーションに関するデータを蓄積するために各アプリケーションにレジスタ・バンクの構成可能な一部分を割り当てるための手段を具備することを特徴とする請求項23に記載の装置。
- 26前記複数のアプリケーションから非同期にスレッドを受信する手段と、 前記複数のアプリケーションに非同期にスレッドを送信する手段と、 を更に具備することを特徴とする請求項23に記載の装置。
- 27複数のアプリケーションを同時にサポートするためのマルチメディア・プロセッサを具備する無線装置であって、 前記マルチメディア・プロセッサは、 前記複数のアプリケーションに関する命令、データ及び状態情報を格納するための構成可能な記憶リソースと、 前記複数のアプリケーションに関する処理を実行するための割り当て可能な処理装置と、なお、前記マルチメディア・プロセッサは、各アプリケーションに前記記憶リソースの構成可能な一部分を割り当て、前記アプリケーションによって要求された際、前記複数のアプリケーションに前記処理装置を動的に割り当てる、 前記記憶リソースに読み込むための命令とデータを格納するためのキャッシュ・メモリと、 を具備することを特徴とする無線装置。
Independent claims27
63 paragraphs, as filed
The present disclosure relates to electronics in general, and specifically to processors.
Processors are widely used for various purposes such as computer computing, communication, and networking. The processor is a general purpose processor such as a central processing unit (CPU), a specialized processor such as a digital signal processor (DSP), or a graphics processing unit (GPU). General-purpose processors support a comprehensive set of functions and instructions used in various types of applications. General-purpose processors are inefficient for certain applications that require unique processing. In contrast, specialized processors support a limited set of specialized features and instructions customized for a unique application. This is because they are given a specialized processor to efficiently support the application for what was designed. However, the range of applications supported by specialized processors is limited.
Devices such as mobile phones, personal digital assistants (PDAs), or laptop computers support various types of applications. It is desirable to run these applications as efficiently as possible and with small hardware in order to reduce costs, power, etc.
Devices that include multimedia processors that can support multiple applications at the same time are listed here. These applications relate to various types of multimedia such as graphics, audio, video, cameras and games. A multimedia processor includes configurable storage resources for storing state information, data and instructions about an application, and assignable processors for performing various types of processing about the application. .. Configurable storage resources include an instruction cache for storing instructions about the application, register banks for storing data about the application, and context registers for storing state information about threads for the application. registers) etc. are included. Processing devices include arithmetic logic unit (ALU) cores, elementary function cores, logic cores, texture samplers, load controllers, flow controllers, and the like. In addition, they operate as described below. The multimedia processor allocates some of the configurable storage resources to each application and dynamically allocates the processor to the application when requested by these applications. Therefore, each application does not need to monitor an independent virtual processor and recognize other applications running at the same time. A multimedia processor is an input interface device for receiving threads asynchronously from an application, an output interface device for asynchronously supplying results to an application, and cache memory as needed, and / or from main memory. It also includes a load controller for retrieving instructions and data about the application.
The multimedia processor determines the read based on the percentage of time the processor is allocated to the application. Multimedia processors adjust clock speeds based on reads to reduce power consumption.
The various aspects and features of the disclosure are described in more detail below.
<figref num="1">Block diagram of a multimedia system.</figref><figref num="2">Block diagram of the multimedia processor.</figref><figref num="3">Diagram showing the allocation of storage resources to N applications.</figref><figref num="4">The figure which shows the allocation of the processing equipment for N applications.</figref><figref num="5">The figure which shows the virtual processor for each of N applications.</figref><figref num="6">Block diagram of the thread scheduler.</figref><figref num="7">The figure which shows the design of the storage device with a virtual memory architecture.</figref><figref num="8">The figure which shows the logical and physical reference table for the storage device of FIG.</figref><figref num="9">The figure which shows the process for supporting a multimedia application.</figref><figref num="10">Block diagram of wireless communication device.</figref>
Detailed description of the invention
FIG. 1 shows a block diagram of the multimedia system 100. System 100 is a stand-alone system or part of a larger system such as a computing system (eg laptop computer), wireless communication device (eg mobile phone), game system (eg game machine). Is. System 100 supports N multimedia applications referenced as 1 to N applications. In general, N is any integer. Applications are also called programs, software programs, and so on. Multimedia applications relate to any type of multimedia such as graphics, audio, video, cameras, games, etc. Applications start and end at different times, and any number of applications run in parallel at a given moment.
System 100 supports two-dimensional (2D) and / or three-dimensional (3D) graphics. 2D or 3D images are represented by polygons (generally triangles). Each triangle is composed of pixels. Each pixel has various properties such as spatial coordinates, color values, and texture coordinates. Each property has up to 4 components. For example, spatial coordinates may be given by either three components, x, y and z, or four components, x, y, z and w. Where x and y are horizontal and vertical coordinates, z is depth and w is homogeneous coordinates. Color values may be given by either three components, r, g and b, or four components, r, g, b and a. Where r is red, g is green, b is blue, and a is the transparency factor that determines the transparency of one pixel. Homogeneous coordinates are generally given by u and v, which are horizontal and vertical coordinates. One pixel may be associated with other properties.
System 100 includes a multimedia processor 120, a texture engine 180 and a configurable cache memory 190. The multimedia processor 120 performs various types of processing related to multimedia applications, as described below. The texture engine 180 performs graphics operations such as texture mapping. Here, texture mapping is a complex graphics operation that associates pixel color correction with the color of a texture image. The cache memory 190 is a high-speed memory that can store data and instructions related to the multimedia processor 120 and the texture engine 180. The system 100 may also include other devices.
The multimedia processor 120 performs processing for N applications. The multimedia processor 120 divides each application into a series of threads, eg, transparent to the application and automatically. A thread (or thread of execution) refers to a particular operation that executes a set of one or more instructions. Threads allow an application to perform multiple processes simultaneously by different devices, and allow different applications to share processing and storage resources.
In the design shown in FIG. 1, the multimedia processor 120 includes an input interface device 122, an output interface device 124, a thread scheduler 130, a flow controller 132, a master controller 134, and an assignable processor 140. Includes configurable storage resources 150 and load controller 170. The input interface device 122 receives threads from N applications and supplies these threads to the thread scheduler 130. The thread scheduler 130 schedules and manages various functions that schedule and manage threads as described below. The flow controller 132 assists in application / program flow control. The master controller 134 receives information such as processing mode and data format, and sets operations of various devices in the multimedia processor 120 accordingly. For example, master controller 134 decodes commands, sets up state for an application, and controls a state update sequence.
In the design shown in FIG. 1, the assignable processor 140 includes an ALU core 142, an elementary function core 144, a logic core 146 and a texture sampler 148. A core generally refers to a processing device in an integrated circuit. The terms "core", "engine", "machine", "processor", "processing device", "hardware device", etc. can be used interchangeably. In general, the specifiable processing device 140 may include any number of processing devices or may include any type of processing device. Each processing device may operate independently of other processing devices.
The ALU core 142 performs arithmetic operations such as addition, subtraction, multiplication, integration and addition, inner product, absolute value, reciprocity, comparison, and saturation. The ALU core 142 is composed of one or more scalar ALUs and / or one or more vector ALUs. The scalar ALU can run on one component at a time. The vector ALU can operate on multiple (eg, 4) components at once. Elementary function core 144 computes transcendental elementary functions such as sine, cosine, reciprocal, logarithm, exponential, square root, and reciprocal of square root, which are widely used in graphic applications. The elementary function core 144 improves performance by computing the elementary function in much less time than it takes to perform a polynomial approximation of the elementary function using simple instructions. The elementary function core 144 includes one or more elementary function devices. Each elementary function device computes an elementary function for one component at a time.
The logical core 146 includes logical operations (eg AND, OR, XOR, etc.), bit operations (eg left shift and right shift), integer operations, comparisons, data buffer management operations (eg, push), and fetch (pop). ) Etc.) and / or perform other operations. In addition, logical core 146 also performs format conversions, such as converting integers to floating point numbers and vice versa. The texture sampler 148 performs preprocessing for the texture engine 180. For example, the texture sampler 148 reads the texture coordinates, attaches the code and / or other information, and sends its output to the texture engine 180. In addition, the texture sampler 148 commands the texture engine 180 to receive the results from the texture engine.
In the design shown in FIG. 1, the configurable storage resource 150 includes a context register 152, an instruction cache 154, a register bank 156, and a constant buffer 158. In general, the configurable storage resource 150 includes any number of storage devices and any type of storage device. The context register 152 stores state information or context about threads from N applications. The instruction cache 154 stores instructions related to threads. These instructions point to specific operations to be performed in each thread. Each operation is an arithmetic operation, an elementary function, a logical operation, a memory access operation, and the like. The instruction cache 154 loads instructions from cache memory 190 and / or main memory (not shown in FIG. 1) via a load controller as needed. Register bank 156 stores intermediate and final results as well as data about the application from the processor. The constant buffer 158 stores constants used by the logical unit 140 (eg, ALU core 142 and logical core 146), such as scale factors and filter weights.
The load controller 170 controls reading instructions, data and constants for N applications. The load controller 170 interfaces with the cache memory 190 to load constants, data, and instructions from the cache memory into the instruction cache 154, register bank 156, and constant buffer 158. Further, the load control device 170 writes the data and the result of the register bank 156 to the cache memory 190. The output interface device 124 receives the final results for the executed thread from register bank 156 and supplies these results to the application. The input interface device 122 and the output interface device 124 provide an asynchronous interface to external devices (for example, a camera, a display device, etc.) associated with N applications.
FIG. 1 shows an example of the design of the multimedia processor 120. In general, the multimedia processor 120 includes any set of assignable processors and any set of configurable storage resources. A configurable storage resource stores instructions, data, state information, etc. about the application. The processing device performs any type of processing related to the application. Even if the controller 132 and the load control device 170 are not included in the device 140, the flow controller 132 and the load control device 170 may also be considered as specifiable processing devices. In addition, the multimedia processor 120 may include other processing devices, storage devices and control devices not shown in FIG. The multimedia processor 120 allocates some of the configurable storage resources to each application and dynamically allocates processing units to the application when requested by these applications.
The multimedia processor 120 is one or more graphics application program interfaces (APIs) such as Open Graphics Library (OpenGL), Direct3D (Direct3D), and Open Vector Graphics (OpenVG). To execute. These various graphics APIs are well known in the art. In addition, the multimedia processor 120 supports 2D graphics, 3D graphics, or both.
FIG. 2 shows a block diagram of the design of the multimedia processor 120 of FIG. In this design, the thread scheduler 130 is interfaced with the ALU core 142, the elementary function core 144, the logical core 146 and the texture sampler 148 in the assignable processor 140. Further, the thread scheduler 130 is interfaced with the input interface device 122, the flow controller 132, the master controller 134, the context register 152, the instruction cache 154, and the load control device 170. The register bank 156 is interfaced with the load controller 170, the output interface device 124, and the ALU core 142, the elementary function core 144, the logical core 146, and the texture sampler 148 in the assignable processing device 140. Further, the load control device 170 is connected to the instruction cache 154, the constant buffer 158, and the cache memory 190 by an interface. Also, the various devices in the multimedia processor 120 may be interfaced with another device in other ways.
The main memory 192 may be part of system 100 or may be outside system 100. The main memory 192 is a larger, slower memory located farther (eg, off-chip) from the multimedia processor 120. The main memory 192 stores all instructions and data for N applications executed by the multimedia processor 120. Instructions and data in main memory 192 are loaded into cache memory 190 when and as needed.
The multimedia processor 120 is designed and processed to appear as an independent virtual processor to each executed application. Each application is allocated sufficient storage resources for instructions, data, constants and state information. Each application has an individual state (eg, program counter, data format, etc.) maintained by the multimedia processor 120. In addition, each application is assigned a processor based on the instructions executed for that application. N applications may run simultaneously without interfering with another application and unaware of the other application. The multimedia processor 120 adjusts performance targets for each application based on application requirements and / or other factors (eg, priorities).
Figure 3 shows an example of allocating storage resources to N applications. Each application is allocated part of context register 152, part of instruction cache 154, part of register bank 156, and part of constant buffer 158. For each storage device, the portion assigned to a given application can be 0 or non-zero depending on the storage requirements of that application.
Context register 152 dynamically allocates threads from N applications and stores various types of information about threads, as described below. Context register 152 is updated when the thread is accepted, executed, and terminated. The instruction cache 154 and the register bank 156 are assigned to each application at the start of execution, for example, based on the request of the application. For each application, the portion allocated to instruction cache 154 and / or register bank 156 may change during application execution based on its requirements and other factors. The constant buffer 158 stores constants used for any application. The constants for a given application are loaded into the constant buffer 158 when needed and are then available to all applications.
The storage device may be designed to support flexible allocation of storage resources to the application, as described below. In addition, the storage device may be designed for easy memory access by the application, as described below.
Figure 4 shows an example of assigning processing equipment to N applications. A separate timeline is maintained in each processing unit such as ALU core 142, elementary function core 144, logical core 146, texture sampler 148, flow controller 132 and load controller 170. The schedule for each processing device is partitioned into time slots. A time slot is the smallest unit of time assigned to an application that matches one or more clock cycles. The processing device has time slots of the same or different durations.
Time slots for the ALU core 142 are assigned to any application. In the example shown in FIG. 4, the ALU core 142 is assigned to application 1 (app 1) at time slots t and t + 1, and is assigned to application 3 at time slot t + 2, time slot t +. It is assigned to application N in 3. Similarly, time slots for elementary function core 144, logical core 146, texture sampler 148, flow controller 132 and load controller 170 are assigned to any application. The multimedia processor 120 dynamically allocates processing units to applications on demand based on the processing requests of these applications.
Figure 5 shows a virtual processor for each of the N applications. Each application observes a virtual processor that has all of the processing equipment utilized by that application. Each application is assigned a processing device based on the processing request of that application. In addition, the designated processor is indicated in the schedule for that application. In the example shown in Figure 5, application 1 is assigned an ALU core 142 in time slots t and t + 1, then a logical core 144 in time slot t + 2, and then time slot t +. 3 for load controller 170, then time slot t + 4 for ALU core 142, then time slot t + 5 for load controller 170, then time slot t + 6 for logical core 144 It is assigned and so on. Application 1 does not use the elementary function core 144, the texture sampler 148 and the flow controller 132 during the time slot period shown in FIG. 5 and is not assigned. Applications 2 through N are assigned processing devices in a different order.
As shown in FIG. 5, each application is assigned the appropriate processing unit within the multimedia processor 120. The specific processing device assigned to each application may change over time depending on the processing request. Each application does not need to be aware of other applications and does not need to know the allocation of processing devices to other applications.
The multimedia processor 120 supports multi-threading to achieve parallel execution of instructions, improving overall efficiency. Multi-threading refers to running a large number of threads in parallel by different processors. Thread scheduler 130 accepts threads from N applications, determines which threads are ready for execution, and sends these threads to different processors. The thread scheduler 130 manages the execution of threads and utilization of the processing device.
FIG. 6 shows a block diagram of the design of the thread scheduler 130 of FIGS. 1 and 2. In this design, the thread scheduler 130 includes a central thread scheduler 610, a high-level decoder 612, a resource usage monitor device 614, an active queue 620 and a sleep queue 622. Context register 152 includes T context registers 630a to 630t for T threads. Here, T is an arbitrary value.
The central thread scheduler 610 communicates with processors 132, 142, 144, 146, 148 and 170, and context registers 630a through 630t via request (Req) and grant interfaces. The scheduler 610 issues a request to the instruction cache 154 and receives a hit / miss instruction in the response. In general, communication between these devices is accomplished by various mechanisms such as control signals, pauses, messages and registers.
The central thread scheduler 610 executes various functions to schedule threads. The central thread scheduler 610 decides whether to accept new threads from N applications, send threads for execution, and release / remove terminated threads. For each thread, the central thread scheduler 610 determines if the resources required by that thread (eg, instructions, processors, register banks, etc.) are available. If the requested resource is available, spawn a thread and place it on the active queue 620, and if none of the threads are available, place the thread on sleep queue 622. The active queue 620 stores threads that are ready to run, and the sleep queue 622 stores threads that are not ready to run.
In addition, the central thread scheduler 610 also manages thread execution. At each scheduling interval (eg, each time slot), the central thread scheduler 610 selects the number of candidate threads in the active queue 620 for evaluation and possible processing. The central thread scheduler 610 determines which processing device to use for the candidate thread, checks for storage device read / write conflicts, and sends different threads to different processing devices for execution. The multimedia processor 120 supports the execution of M threads at the same time. Where M is an appropriate value (eg, M = 12). In general, M is the size of a storage resource (eg, instruction cache 154 and register bank 156), call or delay time for load processing, processing equipment pipeline, so that the processing equipment can use it as much as possible. And / or selected based on the size of other factors.
The central thread scheduler 610 updates the status to the appropriate thread state. If (a) the next instruction for the thread is not found in the instruction cache 154, (b) the next instruction is waiting for a result from the previous instruction, or (c) some other sleep conditions are met. If so, the central thread scheduler 610 places threads on sleep queue 622. If the sleep conditions are incorrect, the central thread scheduler 610 moves threads from sleep queue 622 to active queue 620.
The central thread scheduler 610 maintains a program counter for each thread and updates the program counter when an instruction is executed or the program flow changes. The scheduler 610 seeks assistance from the flow controller 132 to control the program flow for the thread.
The flow controller 132 handles if / else statements, loops, calling subroutines, branches, switch instructions, missing pixels, and / or other flows that change the instruction. The flow controller 132 evaluates one or more conditions for each such instruction, and if the conditions (s) are met, it indicates a one-way change in the program counter to the conditions (s). If they do not match, the change in the program counter in the other direction is shown. In addition, the flow controller 132 also executes other functions related to dynamic program flow. The central thread scheduler 610 updates the program counter based on the result of the flow controller 132.
In addition, the central thread scheduler 610 manages context registers 152 and updates these registers when a thread is accepted, executed, and completed. Context register 152 stores various types of information for threads. For example, the context register 630 for a thread is (1) the application / program identifier (ID) for the application to which the thread belongs, (2) a program counter that indicates the current instruction for the thread, and (3) valid and invalid pixels for the thread. A coverage mask that points to, (4) an active flag that points to the pixel to process in the case of a flow that modifies the instruction, (5) a resume instruction pointer that points to the case where the pixel is inactive and reactivated, ( 6) Stack for storing return instruction pointers for dynamic flow control, (7) Address register for relative addressing, (8) Predicate register for storing state calculation results stores), (9) load reference counters that track load requests and databack conditions, and / or (10) store other information. Context register 152 can also store less, more, or different information.
The multimedia processor 120 supports a comprehensive set of instructions for a variety of multimedia applications. This set of instructions includes operations, elementary functions, logic, bitwise, flow control and other instructions.
Decoding at two levels of the instruction is performed to improve performance. The high-level decoder 612 displays the instruction type, operand type, source and destination identifiers (IDs), and / or other information used for scheduling. Perform a high level decryption of the instruction to determine. Each processor includes a separate instruction decoder that performs low-level decoding of instructions for that processor. For example, the instruction decoder for the ALU core 142 handles only ALU-related instructions, the instruction decoder for the elementary function core 144 handles only the instructions for the elementary function, and so on. Decoding at two levels simplifies the design of the central thread scheduler 610 as well as the instruction decoder for the processor.
The resource usage monitoring device 614 monitors the usage of the processing device, for example, by continuously tracking the percentage of time allocated to each processing device. The monitoring device 614 adjusts the processing of the processing device to save battery power while supplying the appropriate performance. For example, monitor device 614 adjusts the clock speed of the multimedia processor 120 based on the multimedia processor's reads to reduce power consumption. Further, the monitoring device 614 also adjusts the clock speed of each processing device based on the rate of reading or utilization of the processing device. The monitor device 614 selects the fastest clock speed for full loading and the slower clock speed for less loading. In addition, monitor device 614 suspends / powers down any processing device that is not assigned to any application and enables / powers up that processing device when assigned to any application.
Each thread is packet-based and processes data units up to a predetermined size. The size of the unit of data is selected based on the design of the processing device and the storage device, the characteristics of the data being processed, and the like. In one design, each thread processes up to 4 pixels, or up to 4 vertices, in the image. Register bank 156 includes four register banks. The register bank stores up to four components of each characteristic for (a) pixels, that is, one component for each register bank, or (b) a component of characteristics for pixels, that is, a register. -Accumulate one pixel for each bank. The ALU core 142 includes four scalar ALUs or one vector ALU that can operate on up to four components at a time.
The storage device (eg, instruction cache 154 or register bank 156) is implemented in a virtual memory architecture that allows the application to efficiently allocate storage resources and easily recalls the storage resources allocated by the application. The virtual memory architecture utilizes virtual memory and physical memory.
The application allocates a partition of virtual memory and makes memory access via the virtual address space. Different partitions of virtual memory are allocated to different partitions of physical memory that store instructions and / or data.
Figure 7 shows the virtual memory architecture and the design of the storage device 700. The storage device 700 is used for the instruction cache 154, the register bank 156, and the like. In this design, the storage device 700 is seen by the application as virtual memory 710. The virtual memory 710 is divided into a plurality of (for example, S) logical tiles or divisions referred to as 1 to S logical tiles. In general, S is any integer greater than or equal to N. S have the same size or different sizes. Each application can be assigned any number of contiguous logical tiles based on the application's memory usage and available tiles. In the example shown in FIG. 7, application 1 is assigned logical tiles 1 and 2, application 2 is assigned logical tiles 3-6, and so on.
The storage device 700 implements a physical memory 720 that stores instructions and / or data related to the application. The physical memory 720 contains S physical tiles from 1 to S. Each logical tile in virtual memory 710 is mapped to one physical tile in physical memory 720. An example of mapping for some logical tiles is shown in Figure 7. In this example, the physical tile 1 stores instructions and / or data for the logical tile 2, and the physical tile 2 stores commands and / or data for the logical tile S-1.
The use of logical and physical tiles simplifies tile assignment and tile management for applications. The application requires a certain amount of storage resources for instructions and data. The multimedia processor 120 allocates one or more tiles in the instruction cache 154 and one or more tiles in the register bank 156 to the application. The application may be assigned additional logical tiles, fewer logical tiles, or different logical tiles as needed.
Figure 8 shows the logical tile reference table in Figure 7 (look up). table) (LUT) 810 design and storage The physical address reference table 820 for the storage device 700 is shown. In this design, the logical tile reference table 810 contains N entries for N applications, i.e. one entry for each application. N entries are indexed by application ID. The entry for each application contains a field for the first logical tile assigned to the application and another field for the number of logical tiles assigned to the application. In the example shown in FIG. 8, application 1 is assigned two logical tiles starting from logical tile 1, application 2 is assigned four logical tiles starting from logical tile 3, and application 3 is assigned two logical tiles. It is shown that 8 logical tiles are assigned starting from logical tile 7. Each application may be assigned contiguous logical tiles to simplify the generation of addresses for memory access. However, the application may be assigned logical tiles in any order, for example, so that logical tile 1 is assigned to any application.
In the design shown in FIG. 8, the physical address reference table 820 contains S entries for S logical tiles, that is, one entry for each logical tile. The S entries in table 820 are indexed by logical tile address. The entry for each logical tile points to the physical tile to which that logical tile is mapped. In the examples shown in FIGS. 7 and 8, logical tile 1 is mapped to physical tile 4, logical tile 2 is mapped to physical tile 1, logical tile 3 is mapped to physical tile i, and logical tile 4 is physical. For example, it is mapped to tile S-2. Reference tables 810 and 820 are updated each time an application is assigned additional logical tiles, fewer logical tiles, and / or different logical tiles. The application is allocated different amounts of storage resources simply by updating the lookup table without having to actually transfer instructions or data between physical tiles.
Therefore, the storage device is associated with virtual memory and physical memory. Each application is assigned a portion of the virtual memory that can be installed. At least one table is used to map the portion of virtual memory allocated to each application to the corresponding portion of physical memory.
Each application is generally allocated a finite amount of storage resources in the instruction cache 154 and register bank 156 to store instructions and data for that application, respectively. Cache memory 190 stores additional instructions and data for the application. A cache miss is returned to the thread scheduler 130 whenever an instruction to the application is not available in the instruction cache 154. At that time, the thread scheduler 130 issues an instruction request to load the load controller 170. Similarly, whenever data for an application is not available in register bank 156 or the storage device overflows with data, a data request is issued to load controller 170.
The load control device 170 receives an instruction request from the thread scheduler 130 and a data request from another device. The load controller 170 may (a) load the instructions and / or data requested by the cache memory 190 or the main memory 192 and / or (b) the cache memory 190 or the main. Mediate these various requests and generate memory requests to write data to memory 192.
The storage device in the multimedia processor 120 stores a small portion of the instructions and data currently used by the application. Cache memory 190 stores a larger portion of instructions and data that will be used by the application. The multimedia processor 120 supports unlimited instruction and data access via cache memory 190. This capability allows the multimedia processor 120 to support applications of any size. In addition, the multimedia processor 120 supports generic memory and texture loads between cache memory 190 and main memory 192.
Figure 9 shows process 900 supporting multimedia applications. Multiple applications are supported simultaneously, for example by a multimedia processor (block 912). A portion of the available storage resources is allocated to each application to store instructions, data and state information about the application (block 914). In block 914, each application allocates part of an installable instruction cache to store instructions to the application, part of an installable register bank to store data to the application, and state information to the application. For example, assign one or more context registers to store. Processing devices are dynamically assigned to applications when requested by these applications (block 916). Threads are received asynchronously from the application and scheduled to run (block 918). The thread execution result is provided asynchronously to the application (block 920).
The multimedia processors described herein are used in wireless communication devices, portable devices, game devices, computing devices, consumer electronic devices, computers and the like. Typical uses of multimedia processors for wireless communication devices are described below.
FIG. 10 shows a block diagram of the design of the wireless communication device 1000 in the wireless communication system. The wireless device 1000 is a mobile phone, a terminal, a handset, a personal digital assistant (PDA), or other device. A wireless communication system is a code division multiple access (CDMA) system, a global mobile communication system (GSM) system, or other system.
The wireless device 1000 can supply bidirectional communication via a reception path and a transmission path. On the receiving path, the signal transmitted by the base station is received by the antenna 1012 and fed to the receiver (RCVR) 1014. Receiver 1014 supplies samples to digital section 1020 for conditioning, digitizing, and further processing the received signal. In the transmission path, the transmitter (TMTR) 1016 receives the data transmitted from the digital section 1020, processes and coordinates the data, and produces a modulated signal that is transmitted to the base station via the antenna 1012. ..
Digital section 1020 includes, for example, modem processor 1022, digital signal processor (DSP) 1024, video / audio processor 1026, controller / processor 1028, display processor 1030, central processing unit (CPU) / reduction instruction set computer. Includes various processing devices, interface devices and storage elements such as (RISC) 1032, multimedia processor 1034, camera processor 1036, internal memory / cache memory 1038 and external bus interface (EBI) 1040. Modem processor 1022 performs processes related to data transmission and reception (eg, coding, modulation, demodulation and decoding). The DSP1024 performs specialized processing on the wireless device 1000. The video / audio processor 1026 is a camcorder, video assist (video) Video content for video applications such as playback) and video conferencing (eg, still images, video and video text. In addition, the video / audio processor 1026 performs audio content for audio applications. It also performs processing related to (for example, synthesized audio). The controller / processor 1028 directs the operations of various devices within the digital section 1020. The display processor 1030 directs the video, graphics on the display device 1050. And the processing to prepare the text display. The CPU / RISC1032 performs the general-purpose processing related to the wireless device 1000. The multimedia processor 1034 executes the processing related to the multimedia application and described above. The camera processor 1036 performs processing related to the camera (omitted in FIG. 10). The internal memory / cache memory 1038 is in the digital section 1020. Stores data and / instructions for various devices and implements the cache memory 190 of FIGS. 1 and 2. The EBI 1040 has a digital section 1020 (eg, internal memory / cache memory 1038) and main memory 1060. Facilitates the transfer of data between. Multimedia applications can be any processor in digital section 1020. It is executed against the service.
Digital section 1020 runs on one or more processors, microprocessors, DSPs, RISCs, and so on. In addition, the digital section 1020 is also made of one or more application specific integrated circuits (ASICs) and / or some other types of integrated circuits (ICs).
The multimedia processors described herein run within a variety of hardware devices. For example, microprocessors include ASICs, DSPs, digital signal processing devices (DSPDs), programmable logic devices (PLDs), and field programmable gate arrys (FPGAs). , Processors, controllers, microcontrollers, microprocessors, electronic devices and multimedia processors running within other electronic devices may or may not include integrated memory / embedded memory.
The device running the multimedia processor described herein may be a stand-alone device or part of a device. The device is an ASIC such as (i) a stand-alone IC, (ii) stored data, and / or a set of one or more ICs containing memory ICs for instructions, (iii) a mobile station modem (MSM). , (Iv) Modules embedded within other devices, (v) Mobile phones, wireless devices, handsets or mobile devices (vi) and others.
The above statements of the present disclosure are provided to allow any person skilled in the art to manufacture and use the present disclosure. Various amendments to this disclosure will be readily apparent to those skilled in the art, even if the generic principles defined herein apply to a variety of other things without departing from the spirit or scope of the disclosure. good. Therefore, this disclosure is not intended to be limited to the examples described herein and should be given the broadest scope consistent with the novel features and principles disclosed herein.
11 sheets
Sheet 1 Sheet 2 Sheet 3 Sheet 4 Sheet 5 Sheet 6 Sheet 7 Sheet 8 Sheet 9 Sheet 10 Sheet 11
Every citation, both ways
| Document | Relation | Office | Cited during |
|---|---|---|---|
| KR101359717B1 | Cited by | Republic of Korea | Search report |
| JP2013539892A | Cited by | Japan | Examiner |
| JP2000509528A | Cites | Japan | Examiner |
| JP2004157636A | Cites | Japan | Examiner |
| JP2004259274A | Cites | Japan | Examiner |
| JP2005182791A | Cites | Japan | Examiner |
| JPH10187533A | Cites | Japan | Examiner |
19 members in 10 offices
Priority claims9
| Document | Office | Kind | Date |
|---|---|---|---|
| 11677362 | United States of America | – | |
| 67736207 | United States of America | A | |
| 67736207 | United States of America | A | |
| 2008054620 | United States of America | W | |
| 2008054620 | United States of America | W | |
| 2007677362 | – | – | – |
| 2008054620 | – | – | – |
| US20070677362 | – | – | – |
| WO2008US54620 | – | – | – |
Members19
| Document | Office | Kind | |
|---|---|---|---|
| US2008201716A1 | United States of America | A1 | |
| CA2676184A1 | Canada | A1 | |
| WO2008103854A1 | World Intellectual Property Organization (WIPO) | A1 | |
| TW200842757A | Taiwan Province of China | A | |
| KR20090115211A | Republic of Korea | A | |
| EP2126690A1 | European Patent Office (EPO) | A1 | |
| CN101627367A | China | A | |
| US7685409B2 | United States of America | B2 | |
| JP2010519652AThis record | Japan | A | |
| RU2009135022A | Russian Federation | A | |
| RU2425412C2 | Russian Federation | C2 | |
| KR101118486B1 | Republic of Korea | B1 | |
| TWI367453B | Taiwan Province of China | B | |
| JP5149311B2 | Japan | B2 | |
| EP2126690B1 | European Patent Office (EPO) | B1 | |
| BRPI0807951A2 | Brazil | A2 | |
| CA2676184C | Canada | C | |
| CN101627367B | China | B | |
| BRPI0807951B1 | Brazil | B1 |
25 legal events, as the office reported them to INPADOC
Over the term
Point at a mark for the eventEvents
| Event | Code | |
|---|---|---|
| Receipt of annual feesJAPANESE INTERMEDIATE CODE: R250R250 | R250 | |
| Receipt of annual feesJAPANESE INTERMEDIATE CODE: R250R250 | R250 | |
| Receipt of annual feesJAPANESE INTERMEDIATE CODE: R250R250 | R250 | |
| Receipt of annual feesJAPANESE INTERMEDIATE CODE: R250R250 | R250 | |
| Receipt of annual feesJAPANESE INTERMEDIATE CODE: R250R250 | R250 | |
| Receipt of annual feesJAPANESE INTERMEDIATE CODE: R250R250 | R250 | |
| Receipt of annual feesJAPANESE INTERMEDIATE CODE: R250R250 | R250 | |
| Receipt of annual feesJAPANESE INTERMEDIATE CODE: R250R250 | R250 | |
| Receipt of annual feesJAPANESE INTERMEDIATE CODE: R250R250 | R250 | |
| Receipt of annual feesJAPANESE INTERMEDIATE CODE: R250R250 | R250 | |
| Receipt of annual feesJAPANESE INTERMEDIATE CODE: R250R250 | R250 | |
| Renewal fee payment (event date is renewal date of database)FPAY | FPAY | |
| Certificate of patent or registration of utility modelJAPANESE INTERMEDIATE CODE: R150R150 | R150 | |
| Certificate of patent or registration of utility modelJAPANESE INTERMEDIATE CODE: R150R150 | R150 | |
| First payment of annual fees (during grant procedure)JAPANESE INTERMEDIATE CODE: A61A61 | A61 | |
| Written decision to grant a patent or to grant a registration (utility model)JAPANESE INTERMEDIATE CODE: A01A01 | A01 | |
| Decision of grant or rejection writtenTRDD | TRDD | |
| Request for written amendment filedJAPANESE INTERMEDIATE CODE: A523A521 | A521 | |
| Written submission of copy of amendment under article 19 pctJAPANESE INTERMEDIATE CODE: A524A524 | A524 | |
| Notification of reasons for refusalJAPANESE INTERMEDIATE CODE: A131A131 | A131 | |
| Request for written amendment filedJAPANESE INTERMEDIATE CODE: A523A521 | A521 | |
| Written permission of extension of timeJAPANESE INTERMEDIATE CODE: A602A602 | A602 | |
| Written request for extension of timeJAPANESE INTERMEDIATE CODE: A601A601 | A601 | |
| Notification of reasons for refusalJAPANESE INTERMEDIATE CODE: A131A131 | A131 | |
| Report on retrievalJAPANESE INTERMEDIATE CODE: A971007A977 | A977 |
Numbers
- Publication
- 2010519652
- Publication, DOCDB
- 2010519652
- Publication, EPODOC
- JP2010519652
- Application
- 2009551017
- Application, DOCDB
- 2009551017
- Application, EPODOC
- JP20090551017
Titles2
- Japanese
- オン-デマンド・マルチ-スレッド・マルチメディア・プロセッサ
- English
- On-Demand Multi-Thread Multimedia Processor
Classification
- CPC, 14
- G06F9/5016
- G06F12/0842
- G06F9/30145
- G06F9/30167
- G06F9/382
- G06F9/383
- G06F9/3851
- G06F9/3885
- G06F12/10
- G06F9/45558
- G06F2009/45579
- G06F2009/45583
- Y02D10/00
- G06F9/38
- IPC, 2
- G06F9 50
- G06F9 46
Designated states4
- Regional, 4
- Zimbabwe
- Turkmenistan
- Türkiye
- Togo