On-demand multi-thread multimedia processor
27 claims: 5 independent, 22 dependent
- 1REIVINDICAÇÕES Dispositivo que inclui um processador de para suportar simultaneamente múltiplas o processador de multimídia compreendendo:recursos de armazenagem configuráveis para armazenar instruções, dados e informações de estado para as múltiplas aplicações;e unidades de processamento atribuíveis para executar processamento para as múltiplas aplicações, onde o processador de multimídia aloca uma porção configurável dos recursos de armazenagem para cada aplicação e dinamicamente atribui as unidades de processamento às múltiplas aplicações como solicitado pelas aplicações.
- 2Dispositivo, de acordo com a reivindicação 1, em que os recursos de armazenagem configuráveis compreendem um cache de instrução para armazenar instruções para as múltiplas aplicações, onde cada aplicação é alocada uma porção configurável do cache de instrução.
- 3Dispositivo, de acordo com a reivindicação 1, em que os recursos de armazenagem configuráveis compreendem bancos de registrador para armazenar dados para as múltiplas aplicações, em que cada aplicação é alocada uma porção configurável dos bancos de registrador.
- 4Dispositivo, de acordo com a reivindicação 1, em que os recursos de armazenagem configuráveis compreendem uma unidade de armazenagem para armazenar instruções ou dados para as múltiplas aplicações, a unidade de armazenagem sendo associada a uma memória virtual e uma memória física, em que cada aplicação é alocada uma porção configurável da memória virtual;e pelo menos uma tabela para mapear a porção da memória virtual alocada a cada aplicação para uma porção correspondente da memória física. 2/6
- 5Dispositivo, de acordo com a reivindicação 1, em que o processador de multimídia compreende ainda:uma unidade de controle de carga para buscar instruções para cada aplicação, como solicitado, a partir de uma memória cache ou uma memória principal para armazenagem na porção dos recursos de armazenagem alocados para a aplicação.
- 6Dispositivo, de acordo com a reivindicação 1, em que as unidades de processamento operam independentemente e cada unidade de processamento é atribuível a qualquer uma das múltiplas aplicações em uma dada partição de tempo.
- 7Dispositivo, de acordo com a reivindicação 1, em que o processador de multimídia desliga unidades de processamento não atribuíveis a qualquer uma das aplicações múltiplas.
- 8Dispositivo, de acordo com a reivindicação 1, em que as unidades de processamento compreendem pelo menos um entre um núcleo de unidade de lógica aritmética (ALU), um núcleo de função elementar, um núcleo de lógica, um amostrador de textura, uma unidade de controle de carga, e um controlador de fluxo.
- 9Dispositivo, de acordo com a reivindicação 1, em que o processador de multimídia compreende ainda:uma unidade de interface de entrada para receber de forma assíncrona, encadeamentos a partir das múltiplas aplicações;e uma unidade de interface de saída para fornecer, de forma assíncrona, resultados para as múltiplas aplicações.
- 10Dispositivo, de acordo com a reivindicação 1, em que o processador de multimídia ajusta velocidade de 3/6 10, em que carregamento relógio com base em carregamento do processador de multimídia para reduzir consumo de energia.
- 11Dispositivo, de acordo com a reivindicação o processador de multimídia determina o com base em percentagem de tempo que as unidades de processamento são atribuídas em um período de tempo específico.
- 12Dispositivo, de acordo com a reivindicação 1, em que o processador de multimídia suporta simultaneamente uma pluralidade de encadeamentos para as múltiplas aplicações.
- 13Dispositivo, de acordo com a reivindicação 12, em que os recursos de armazenagem configuráveis compreendem uma pluralidade de registradores de contexto para a pluralidade de encadeamentos, cada registrador de contexto armazenando informações de estado para um encadeamento associado.
- 14Dispositivo, de acordo com a reivindicação 13, em que as informações de estado para cada encadeamento compreendem pelo menos um entre um contador de programa, uma pilha para armazenar ponteiros para controle de fluxo de endereço para endereçamento de atributo para armazenar de cálculo de condição, e um contador de referência de carga para rastrear solicitações de carga e condições de retorno de dados.
- 15Dispositivo, de acordo com a reivindicação 12, em que cada encadeamento opera em uma unidade de dados de até um tamanho predeterminado.
- 16Dispositivo, de acordo com a reivindicação 12, em que cada encadeamento opera em até quatro pixels ou até quatro vértices em uma imagem. registradores registradores dinâmico, relativo, resultados 4/6
- 17Dispositivo, de acordo com a reivindicação 1, em que o processador de multimídia suporta um único conjunto de instruções aplicáveis para as múltiplas aplicações.
- 18Dispositivo, de acordo com a reivindicação 1, em que as múltiplas aplicações compreendem pelo menos uma entre uma aplicação de gráfico, uma aplicação de áudio, uma aplicação de vídeo, uma aplicação de câmera, e uma aplicação de jogos.
- 19Método compreendendo:suportar simultaneamente múltiplas aplicações;alocar uma porção configurável de recursos de armazenagem para cada aplicação para armazenar instruções, dados, e informações de estado para a aplicação;e atribuir dinamicamente unidades de processamento para as múltiplas aplicações como solicitado pelas aplicações.
- 20Método, de acordo com a reivindicação 19, em que a alocação compreende alocar uma porção configurável de um cache de instrução para cada aplicação para armazenar instruções para a aplicação.
- 21Método, de acordo com a reivindicação 19, em que a alocação compreende alocar uma porção configurável de bancos de registrador para cada aplicação para armazenar dados para a aplicação.
- 22Método, de acordo com a reivindicação 19, compreendendo ainda:receber de forma assíncrona encadeamentos a partir das múltiplas aplicações;e fornecer de forma assíncrona resultados para as múltiplas aplicações. 5/6
- 23Equipamento compreendendo:meio para suportar simultaneamente múltiplas aplicações;meio para alocar uma porção configurável de recursos de armazenagem para cada aplicação para armazenar instruções, dados, e informações de estado para a aplicação;e meio para atribuir dinamicamente unidades de processamento para as múltiplas aplicações como solicitado pelas aplicações.
- 24Equipamento, de acordo com a reivindicação 23, em que o meio para alocar compreende meio para alocar uma porção configurável de um cache de instrução para cada aplicação para armazenar instruções para a aplicação.
- 25Equipamento, de acordo com a reivindicação 23, em que o meio para alocar compreende meio para alocar uma porção configurável de bancos de registrador para cada aplicação para armazenar dados para a aplicação.
- 26Equipamento, de acordo com a reivindicação 23, compreendendo ainda:meio para receber, de forma assíncrona, encadeamentos a partir das múltiplas aplicações;e meio para fornecer, de forma assíncrona, resultados para as múltiplas aplicações.
- 27Dispositivo sem fio compreendendo:um processador de multimídia para suportar simultaneamente múltiplas aplicações, o processador de multimídia compreendendo recursos de armazenagem configuráveis para armazenar instruções, dados e informações de estado para as múltiplas aplicações, e 6/6 unidades de processamento atribuíveis para executar processamento para as múltiplas aplicações, em que o processador de multimídia aloca uma porção configurável dos recursos de armazenagem a cada aplicação e 5 dinamicamente atribui as unidades de processamento às múltiplas aplicações como solicitado pelas aplicações;e uma memória cache para armazenar instruções e dados para carregamento para os recursos de armazenagem. 1/10 Ο CC ΧΖ Ο 1^. φ φ σ σ Φ Φ J0 σ ο <5 £= φ 32 c Ο c ο Ο ο οο XJο Osl ο <ο Φ Õ Ε ο Τ3 Ο σ (0 w « ν ο ο CM <Μ Φ Ό Φ ι_ -1—* C LU σ> 'Φ > ο 2 2 $? ω α> ο ΙΕ Φ < Ό <ο· ÍFRJ <Λ ο Έ φ Φ Ό Ο φ φ 8 Φ 5? Ω. Ο «5 <f LL1 ω ο -·—» c φ ...§ φ φ Ο Φ Ο C LU Φ *σ ο ι_ ιφ 4= ο Έ ™ Φ Ο CL ~ •*3CM ο «ο η Φ ,π φ as Ό Ό t Ό Τ3 Έ φ .22 φ φ ζ> Ό C ο CO V6 ”· Γ 7 Ιφ ο φ ο φ CL 2/10 Processador de Multimídia £Z — C LU LU 3/10 <ο ΙΟ te ο Ό (β 4/10 m cl CL < xr CL CL < m cl CL < m cl CL < NQ. CL < CN CL CL < CO Q. Q. < CL CL CL CL < co CL Q. < CL CL < m CL CL < co CL CL CN CL CL CO co + CN Partição de Tempo
Independent claims27
126 paragraphs in 7 sections, as filed
(54) Title: MULTIMEDIA PROCESSOR MULTI- (57) Summary:
CHAINED ON DEMAND (30) Unionist Priority: 2/21/2007 us 11 / 677,362 (73) Owner (s): Qualcomm Incorporated (72) Inventor (s): CHUN YU, Guofang Jiao, Yun Du (74) Attorney (s) ): Montaury Pimenta, Machado &
Lioce (86) International Order: pct uS2008054620 de
21/02/2008 (87) International Publication: wo 2008 / i03854de 28/08/2008
Threads from | Application 1
Threads from Application N <sup>1</sup>
<img file="BRPI0807951A2_D0001.tif" />
MULTI-MEDIA CHAINED PROCESSOR UNDER DEMAND
FUNDAMENTALS
FIELD
The present disclosure refers generally to electronics, and more specifically to a processor.
FUNDAMENTALS
Processors are widely used for various purposes such as computing, communication, networking, etc. A processor can be a general purpose processor like a central processing unit (CPU) or a specialized processor like a digital signal processor (DSP) or a graphics processing unit (GPU). A general purpose processor can support a generic set of instructions and generic functions, which can be used by applications of various types. A general purpose processor can be inefficient for certain applications with specific processing requirements. In contrast, a specialized processor can support a limited set of instructions and specialized functions, which can be customized for specific applications. This allows the specialized processor to efficiently support the applications for which it is designed. However, the range of applications supported by the specialized processor may be limited.
A device like a cell phone, a personal digital assistant (PDA), or a laptop computer can support applications of various types. It is desirable to run these applications as efficiently as possible and with as little hardware as possible to reduce cost, energy, etc.
SUMMARY
2/26
A device including a multimedia processor that can simultaneously support multiple applications is described here. These applications can be for various types of multimedia such as graphics, audio, video, camera, games, etc. The multimedia processor comprises configurable storage resources to store instructions, data, and status information for applications and assignable processing units to perform various types of processing for applications. Configurable storage resources can include an instruction cache to store instructions for applications, register banks to store data for applications, context recorders to store state information for application threads, etc. Processing units may include an arithmetic logic unit (ALU) core, an elementary function core, a logic core, a texture sampler, a load control unit, a flow controller, etc. that can operate as described below. The multimedia processor allocates a configurable portion of the storage resources for each application and dynamically assigns the processing units to the applications as requested by those applications. Each application thus observes an independent virtual processor and does not need to be aware of the other applications running simultaneously. The multimedia processor can also include an input interface unit for asynchronously receiving application threads, an output interface unit for asynchronously providing results for the applications, and a load control unit for searching instructions and data for the applications. applications, as needed, from cache memory and / or main memory.
3/26
The multimedia processor can determine loading based on the percentage of time that processing units are assigned to applications. The multimedia processor can adjust clock speed based on charging to reduce power consumption.
Various aspects and characteristics of the disclosure are described in further detail below.
BRIEF DESCRIPTION OF THE DRAWINGS
Figure 1 shows a block diagram of a multimedia system.
Figure 2 shows a block diagram of a multimedia processor.
Figure 3 shows the allocation of storage resources to N applications.
Figure 4 shows the allocation of processing units to N applications.
Figure 5 shows a virtual processor for each of the N applications.
Figure 6 shows a block diagram of a chaining programmer.
Figure 7 shows a drawing of a storage unit with a virtual memory architecture.
Figure 8 shows physical and logical query tables for the storage unit in Figure 7.
Figure 9 shows a process for supporting multimedia applications.
Figure 10 shows a block diagram of a wireless communication device.
DETAILED DESCRIPTION
Figure 1 shows a block diagram of a multimedia system 100. System 100 can be a stand-alone system or part of a larger system such as a
4/26 texture coordinates, etc.
computer system (for example, a laptop computer), a wireless communication device (for example, a cell phone), a game system (for example, a game console), etc. System 100 supports N multimedia applications, which are referred to as applications 1 through N. In general, N can be any integer value. An application can also be referred to as a program, a software program, etc. A multimedia application can be for any type of multimedia such as graphics, audio, video, camera, games, etc. Applications can start and end at different times, and any number of applications can be run in parallel at any given time.
System 100 can support two-dimensional (2-D) and / or three-dimensional (3-D) graphics. A 2-D or 3-D image can be represented with polygons (typically triangles). Each triangle can be composed of image elements (pixels). Each pixel can have several attributes such as space, color values, coordinates. Each attribute can have up to four components. For example, space coordinates can be given by three components x, y and z or four components x,
w.
where y and y are horizontal and vertical coordinates, z is depth, and w is a homogeneous coordinate. Color values can be given by three components rg, eb or four components r, g, bea, where r is red, g is green, b is blue and a is a transparency factor that determines the transparency of a pixel. Texture coordinates are typically given by horizontal and vertical coordinates, u and v. A pixel can also be associated with other attributes.
System 100 includes a multimedia processor
120, a 180 texture engine, and a configurable cache memory 190. The multimedia processor 120 can
5/26 perform various types of processing for multimedia applications, as described below. The texture engine 180 can perform graph operations like texture mapping, which is a complex graph operation involving modifying the color of pixels with the color of a texture image. Cache memory 190 is fast memory that can store instructions and data for multimedia processor 120 and texture engine 180. System 100 can include other units.
The multimedia processor 120 performs processing for N applications. The multimedia processor 120 can divide the processing of each application into a series of threads, for example, automatically and transparently to the application. A thread (or thread) can indicate a specific task that can be performed with a set of one or more instructions. Chaining allows an application to have multiple tasks executed simultaneously by allowing different processing and flow resources 132, processing different units and applications to share storage.
In the drawing shown in figure 1, the multimedia processor 120 includes an input interface unit 122, an output interface unit 124, a thread programmer 130, a master controller controller 134, assignable units 140, storage resources configurable 150, and a load control unit 17 0. The input interface unit 122 receives threads from the N applications and supplies those threads to the thread programmer 130. The thread scheduler 130 performs several functions to schedule and manage thread execution, as described below. The controller
6/26 flow 132 assists with program / application flow control. Master controller 134 receives information such as processing mode, data format, etc. and configure the operation of multiple units on the multimedia processor 120 accordingly. For example, master controller 134 can decode command, set status for applications, and control a status update sequence.
In the drawing shown in Figure 1, assignable processing units 140 include an ALU core 142, an elementary function core 144, a logic core 146, and a texture sampler 148. A core generally refers to a processing unit in a core, motor, integrated circuit. The terms processor, hardware unit, etc., can be interchangeable. In general, processing, assignable machine units, processing unit units 140 may include any number of units of arithmetic unit multiply processing and any type processing. Each processing unit can operate independently from the other processing units.
The ALU 142 core can perform operations such as addition, subtraction, multiplication, and accumulate, point product, absolute, negation, comparison, saturation, etc. The ALU core 142 may comprise one or more scalar ALUs and / or one or more vector ALUs. A scalar ALU can operate on a component at a time. A vector ALU can operate on multiple components (for example, four) at once. The elementary function core 144 can compute transcendental elementary functions such as sine, cosine, reciprocal, logarithm, exponential, square root, reciprocal square root, etc., which can be widely used by applications of
7/26 chart. The elementary function kernel 144 can improve performance by computing elementary functions in much less time than the time required to perform polynomial approximations of elementary functions using simple instructions. The elementary function core 144 may comprise one or more elementary function units. Each elementary function unit can compute an elementary function for one component at a time.
The logic core 146 can perform logical operations (eg, AND, OR, XOR, etc.), bitwise operations (eg, left and right shifts), integer operations, comparison, management operations data buffering (eg, pushing, jumping, etc.), and / or other operations. The logic kernel 146 can also perform format conversion, for example, from integers to floating point numbers, and vice versa. The texture sampler 148 can perform preprocessing for texture engine 180. For example, the texture sampler 148 can read texture coordinates, attach code and / or other information, and send its output
<td>for the engine</td><td colspan="2">180 texture.</td><td>sampler</td><td>texture</td><td> 148</td>
<td>can provide</td><td>also</td><td>instructions</td><td>for the engine</td><td>texture</td><td> 180</td>
<td colspan="2">and receive results</td><td>engine</td><td>: texture.</td><td></td><td></td>
<td>At the</td><td>drawing</td><td>shown</td><td>in figure 1,</td><td>resources</td><td>in</td>
<td>storage</td><td colspan="2">configurable 150</td><td colspan="2">include recorders</td><td>in</td>
<td>context 152</td><td>, one</td><td>cache of</td><td>instruction 154,</td><td>banks</td><td>in</td>
<td>recorder</td><td>156, and</td><td>a buffer</td><td>constant 158.</td><td colspan="2">Generally,</td>
configurable storage features 150 can include any number of storage units and any type of storage unit. Context registers 152 store state or context information for threads from N applications. Instruction cache 154 stores instructions for threads.
8/26
These instructions indicate specific operations to perform for each thread. Each operation can be an arithmetic operation, an elementary function, a logic operation, a memory access operation, etc. Instruction cache 154 can be loaded with instructions from cache memory 190 and / or a main memory (not shown in figure 1), as needed, via load control unit 170. The register banks 156 store data for the applications as well as intermediate and final results from processing units 150. The constant buffer 158 stores constant values (for example, scale factors, filter weights, etc.) used by the processing units. processing 140 (e.g., ALU core 142 and logic core 146).
The load control unit 170 can control the loading of instructions, data and constants for N applications. The load control unit 170 interfaces with cache memory 190 and loads instruction cache 154, register banks 156, and constant buffer 158 with instructions, data, and constants from cache memory 190. Load control unit 170 also writes data and results in register banks 156 to cache memory 190. Output interface unit 124 receives the final results for the threads executed from register banks 156 and provides these results for the applications. The input interface unit 122 and output interface unit 124 can provide asynchronous interfaces to external units (e.g., camera, display unit, etc.) associated with N applications.
Figure 1 shows an example drawing of a multimedia processor 120. In general, the
Multimedia 9/26 120 can include any set of assignable processing units and any set of configurable storage resources. Configurable storage resources can store instructions, data, status information, etc., for applications. Processing units can perform any type of processing for applications. Flow controller 132 and load control unit 170 can also be considered as assignable processing units although they are not included in units 140. Multimedia processor 120 may also include other processing, storage and control units not shown in the figure 1. The multimedia processor 120 can allocate a configurable portion of the storage resources to each application and dynamically assign the processing units to the applications as requested by those applications.
The multimedia processor 120 can implement one or more graphics application programming interfaces (APIs), such as Open Graphics Library (OpenGL), Direct3D, Open Vector Graphics (OpenVG), etc. These various charting APIs are known in the art. The multimedia processor 120 can also support 2-D graphics, or 3-D graphics, or both.
Figure 2 shows a block diagram of a multimedia processor drawing 120 in figure 1. In this drawing, the thread programmer 130 interfaces with the ALU core 142, elementary function core 144, logic core 146, and sampler texture 148 in assignable processing units 140. Thread programmer 130 also interfaces with input interface unit 122, flow controller 132, master controller 134, context recorders 152,
10/26 instruction cache 154, and load control unit 170. Register banks 156 interface with the ALU core 142, elementary function core 144, logic core 146, and texture sampler 148 in the processing units assignable 140, load control unit 170, and output interface unit 124. Load control unit 170 interfaces with instruction cache 154, constant buffer 158, and cache memory 190. The various units in the multimedia processor 120 can interface with each other in other modes as well.
A main memory 192 may be part of system 100 or it may be external to system 100. Main memory 192 is a slower, larger memory located further away (for example, off-chip) from multimedia processor 120. Main memory 192 can store all instructions and data for N applications being
<td>performed</td><td colspan="2">by the processor</td><td>in</td><td>multimedia</td><td> 120 .</td><td>At</td>
<td>instructions</td><td>and data in</td><td colspan="3">main memory 192</td><td>can</td><td>to be</td>
<td>loaded</td><td>in the memory</td><td>cache</td><td> 190</td><td>When is it</td><td colspan="2">according</td>
<td>required.</td><td></td><td></td><td></td><td></td><td></td><td></td>
<td>O</td><td>processor</td><td>in</td><td colspan="2">multimedia 120</td><td>can</td><td>to be</td>
<td>projected</td><td>and operated from</td><td>such</td><td>mode</td><td>that seems</td><td>as</td><td>one</td>
independent virtual processor for each application being run. Each application can be assigned sufficient storage resources for instructions, data, constant and status information. Each application can have its own individual status (for example, program counter, data format, etc.), which can be maintained by the multimedia processor 120. Each application can also be assigned processing units assigned based on the instructions to be executed for that application. N applications can run simultaneously without interfering with each other and without having to be aware of other applications. O
11/26 multimedia processor 120 can adjust the performance target for each application based on application demand and / or other factors, for example, priority.
Figure 3 shows an example allocation of storage resources for applications N. Each application can be allocated a portion of context registers 152, a portion of instruction cache 154, a portion of register banks 156, and a portion of constant buffer 158. For each storage unit, the portion allocated to a given application can be zero or non-zero depending on the storage requirements of that application.
can be based on factors. The constants
Context recorders 152 dynamically assigned to threads to N applications and can store various types of information for threads, as described below. Context recorders 152 can be updated as threads are accepted, executed and completed. Instruction cache 154 and register banks 156 can be allocated to each application at the beginning of execution, for example, based on the requirements of the application. For each application, the portions allocated to instruction cache 154 and / or register banks 156 may change during application execution based on their demand and other constant buffer 158 may be used for application. Constants for a given application can be loaded into the constant buffer 158 when needed and can then be available for use by all applications.
store any
Storage units can be designed to support flexible allocation of storage resources to applications, as described below. The units of
12/26 storage can also be designed to simplify memory access by applications, as also described below.
Figure 4 shows an example assignment of processing units for N applications. A separate timeline can be maintained for each processing unit as an ALU core 142, elementary function core 144, logic core 146, texture sampler 148, flow controller 132, and load control unit 17 0. The time line for each processing unit can be divided into time partitions. A time partition is the smallest unit of time that can be allocated to an application and can correspond to one or more clock cycles. Processing units can have time partitions of the same duration or different durations.
The time partitions for the ALU 142 core can be assigned to any of the applications. In the example shown in figure 4, the core of ALU 142 is assigned to application 1 (App 1) in time partitions tet + 1, application 3 in time partition t + 2, application N in time partition t + 3, etc. Similarly, the time partitions for elementary function core 144, logic core 146, texture sampler 148, flow controller 132, and load control unit 170 can be assigned to any of the applications. The multimedia processor 120 can dynamically assign the processing units to the applications on demand based on the processing requirements of those applications.
Figure 5 shows a virtual processor for each of the N applications. Each application observes a virtual processor having all the processing units used by that application. Each application can be assigned processing units with
13/26 based on the processing demand for that application, and the assigned processing units can be shown on a timeline for that application. In the example shown in figure 5, application 1 is assigned the core of ALU 142 in time partitions tet + 1, then the logic core 144 in time partition t + 2, then the load control unit 170 in the partition of time t + 3, then the core of ALU 142 in the time partition t + 4, then the load control unit 170 in the time partition t + 5, then the logic core 144 in the time partition t + 6 , etc. Application 1 does not use and is not assigned elementary function core 144, texture sampler 148, and flow controller 132 during the time partitions shown in figure 5. Applications 2 through N are assigned processing units in different sequences.
As shown in figure 5, each application can be assigned relevant processing units in the multimedia processor 120. The specific processing units to be assigned to each application can change over time depending on processing requirements. Each application does not need to be aware of the other application or the processing unit assignments for the other applications.
The multimedia processor 120 can support multiple threads to achieve parallel instruction execution and improve overall efficiency. Multiple threads refers to the execution of multiple streams in parallel by different processing units. Thread programmer 130 can accept threads from N applications, determine which threads are ready to run, and dispatch those threads to different processing units. Thread programmer 130 can manage execution
14/26 of threads and use of processing units.
Figure 6 shows a block diagram of a thread programmer drawing 130 in figures 1 and 2. In that drawing, thread programmer 130 includes a central thread programmer 610, a high level decoder 612, a video monitor unit. resource usage 614, an active queue 620 and a queue 622. Context registers 152 include T context registers 630a through 630t for T threads, where T can be any value.
central chaining programmer 610 can communicate with processing units 132, 142, 144, 146, 148 and 170 and context recorders 630a through 630t via grant and request interfaces (Req). Programmer 610 can issue requests for instruction cache 154 and receive hit / miss indications in response. In general, communication between these units can be achieved with various mechanisms such as control signals, interruptions, messages, recorders, etc.
The central thread programmer 610 can perform various functions to program threads. The central thread programmer 610 can determine whether to accept new threads from N applications, dispatch threads that are to be run, and release / remove threads that are completed. For each thread, the central thread programmer 610 can determine if resources (for example, instructions, processing units, recorder banks, etc.) required by that thread are available, activate the thread, and queue it up if 620 required resources are available, and queue the thread 622 if any resources are not
15/26 is available. Active queue 620 stores threads that are ready to run, and queue 622 stores threads that are not ready to run.
The 610 central thread programmer can also manage thread execution. At each scheduling interval (for example, each time partition), central thread programmer 610 can select a number of candidate threads in active queue 620 for evaluation and possible dispatch. The central thread programmer 610 can determine the processing units to use for candidate threads, check read / write conflicts for storage units, and dispatch different threads to different processing units for execution. The multimedia processor 120 can support execution of M threads simultaneously, where M can be any appropriate value (for example, M = 12). In general, M can be selected based on the size of the storage resources (for example, instruction cache 154 and register banks 156), the latency or delay for loading operations, the pipelines of the processing units, and / or others factors so that processing units are used as fully as possible.
The 610 central thread scheduler can update thread status and status as appropriate. Central thread scheduler 610 can queue thread 622 if (a) the next instruction for the thread is not found in instruction cache 154, (b) the next instruction is waiting for results from a previous instruction , or (c) some other waiting conditions are met. The central thread scheduler 610 can move a thread from queue 622 to the
16/26 active queue 620 when waiting conditions are no longer true.
The central thread programmer 610 can maintain a program counter for each thread and can update the program counter as instructions are executed or the program flow is changed. Programmer 610 can request assistance from flow controller 132 to control program flow for threads.
The flow controller 132 can handle if / else instructions, loops, subroutine calls, branches, switching instructions, pixel kill and / or other flow change instructions. The flow controller 132 can evaluate one or more conditions for each such instruction, indicate a change in the program counter in a way if the condition (s) is (are) met, and indicate a change in the counter program otherwise if the condition (s) is not met. Flow controller 132 can also perform other functions related to dynamic program flow. The central thread programmer 610 can update the program counter based on results from flow controller 132.
The central thread scheduler 610 can also manage context recorders 152 and update these registers as threads are accepted, executed and completed. Context recorders 152 can store various types of information for threads. For example, a context register 630 for a thread can store (1) a program / application identifier (ID) for the application to which the thread belongs, (2) a program counter that points to the current instruction · for the chaining, (3) a
17/26 cover mask indicating valid and invalid pixels for the chain, (4) an active indicator indicating which pixels to operate in case of a flow change instruction, (5) a restart instruction hand indicating when a pixel will reactivate if inactive, (6) a stack that stores return instruction pointers for dynamic flow control, (7) address registers for relative addressing, (8) attribute recorders that store storage condition calculation results, (9) a load reference counter that tracks load requests and data return conditions, and / or (10) other information. Context recorders 152 can also store less information, more information or different information.
The multimedia processor 120 can support a comprehensive set of instructions for various multimedia applications. This set of instructions can include arithmetic instructions, elementary function, logic, bit sense, flow control and other instructions.
Two-level decoding of instructions can be performed to improve performance. The high-level decoder 612 can perform high-level decoding of instructions to determine type of instructions, type of operand, source and destination identifiers (IDs), and / or other information used for programming. Each processing unit can include a separate instruction decoder that performs low-level instruction decoding for that processing unit. For example, an instruction decoder for the core of ALU 142 can handle only instructions related to ALU, an instruction decoder for elementary function core 144 can only handle instructions for elementary functions, etc. Two-level decoding
18/26 can simplify the design of central thread programmer 610 as well as instruction decoders for processing units.
The resource usage monitor unit 614 monitors the use of processing units, for example, by keeping an eye on the percentage of time that each processing unit is assigned. The 614 monitor unit can dynamically adjust the operation of the processing units to conserve battery power while providing the desired performance. For example, monitor unit 614 can adjust the clock speed for multimedia processor 120 based on loading the multimedia processor to reduce power consumption. Monitor unit 614 can also adjust the clock speed for each individual processing unit based on load or percentage utilization of that processing unit. The monitor unit 614 can select the highest clock speed for full charging and can select progressively slower clock speed for progressively lower charging. Monitor unit 614 can also disable / turn off any processing unit that is not assigned to any application and can enable / turn on the processing unit when it is assigned.
Each thread can be packet based and can operate on a data unit of up to a predetermined size. The size of the data unit can be selected based on the design of the storage and processing units, the characteristics of the data being processed, etc. In a drawing, each thread operates at up to four pixels or up to four vertices in an image. Registrar banks 156 may include four banks
19/26 register that can store (a) up to four components of each attribute for pixels, one component per register bank, or (b) attribute components for pixels, one pixel per register bank. The ALU 142 core can include four scalar ALUs or a vector ALU that can operate on up to four components at a time.
A storage unit (for example, instruction cache 154 or register banks 156) can be implemented with a virtual memory architecture that allows efficient allocation of storage resources for applications and easy access to the storage resources allocated by applications. The virtual memory architecture can use both virtual memory and physical memory. Applications can be allocated sections of virtual memory and can perform memory access through the virtual address space. Different sections of virtual memory can be mapped to different sections of physical memory, which store instructions and / or data.
Figure 7 shows a drawing of a storage unit 700 with a virtual memory architecture. The storage unit 700 can be used for instruction cache 154, register banks 156, etc. In this drawing, the storage unit 700 appears as a virtual memory 710 for the applications. Virtual memory 710 can be divided into multiple (S) logic tiles or sections, which are referred to as logic tiles 1 through S. In general, S can be any integer value equal to or greater than N. S tiles can be the same size or different sizes. Each application can be allocated any number of consecutive logical tiles based on that application's memory usage and available tiles. In the example shown in figure 7, application 1 is allocated logic tiles 1 and 2, application 2 is allocated logic tiles 3 through 6, etc.
20/26
The storage unit 700 implements a physical memory 720 that stores instructions and / or data for the applications. Physical memory 720 includes physical tiles S 1 through S. Each logical tile of virtual memory 710 is mapped to a physical tile of physical memory 720. An example mapping for some logical tiles is shown in figure 7. In this example, the physical tile 1 stores instructions and / or data for logical tile 2, physical tile 2 stores instructions and / or data for logical tile Sl, etc.
The use of logical tiles and physical tiles can simplify the allocation of tiles to applications and the management of tiles. An application can request certain amounts of storage resources for data instructions. The multimedia processor 120 can allocate one or more tiles in instruction cache 154 and one or more tiles in register banks 156 for the application application, additional, smaller or different logic tiles can be allocated as needed.
Figure 8 shows a drawing of a logical tile lookup table (LUT) 810 and a physical address lookup table 820 for storage unit 700 in figure 7. In this drawing, the logical tile lookup table 810 includes entries N for applications N, one entry for each application. N entries can be indexed by the application ID. The entry for each application includes a field for the first logical tile allocated to the application and another field for the number of logical tiles allocated to the application. In the example shown in figure 8, application 1 is allocated two logic tiles that start with logic tile 1, application 2 is allocated four logic tiles that start with logic tile 3, application 3 is allocated eight logic tiles that start with logical tile 7, etc. Each application can be allocated
A in
21/26 consecutive logic tiles to simplify the generation of addresses for memory access. However, applications can be allocated logical tiles in any order, for example, logical tile 1 can be allocated to any application.
In the drawing shown in figure 8, the physical address lookup table 820 includes S entries for logic tiles S, one entry for each logical tile. The S entries in table 820 can be indexed by logical tile address. The entry for each logical tile indicates the physical tile to which that logical tile is mapped. In the example shown in figures 7 and 8, logical tile 1 is mapped to physical tile 4, logical tile 2 is mapped to physical tile 1, logical tile 3 is mapped to physical tile i, logical tile 4 is mapped for the S-2 physical tile, etc. Query tables 810 and 820 can be updated whenever an application is allocated additional logic tiles, in smaller and / or different numbers. Applications can be allocated different amounts of storage resources by simply updating the lookup tables, without actually having to transfer instructions or data between physical tiles.
A storage unit can thus be associated with a virtual memory and a physical memory. Each application can be allocated a configurable portion of the virtual memory. At least one table can be used to map the portion of virtual memory allocated to each application to a corresponding portion of physical memory.
Each application is typically allocated limited amounts of instruction caching resources
154 and register banks 156 to store instructions and data, respectively, for that application. Cache memory 190 can store additional instructions and data for
22/26 the applications. Whenever an instruction for an application is not available in instruction cache 154, a cache error can be returned to the thread programmer 130, who can then issue an instruction request to the load control unit 170. Similarly, whenever data for an application is not available in register banks 156 or whenever a storage unit overflows with data, a data request can be issued to the load control unit 170.
Load control unit 170 can receive instruction requests from thread programmer 130 and data requests from other units. The load control unit 170 can arbitrate these various requests and generate memory requests to (a) load the desired instructions and / or data from cache memory 190 or main memory 192 and / or (b) write data to memory 190 cache or 192 main memory.
The storage units in the multimedia processor 120 can store small portions of instructions and data that are currently used for applications. Cache memory 190 can store larger portions of instructions and data that could be used for applications. Multimedia processor 120 can support unlimited instruction and data access through cache memory 190. This capability allows the 120 multimedia processor to support applications of any size. Multimedia processor 120 can also support generic memory load and texture load between cache memory 190 and main memory 192.
Figure 9 shows a 900 process for supporting multimedia applications. Multimedia applications are
23/26 supported simultaneously, for example, by a multimedia processor (block 912). A configurable portion of storage resources is allocated to each application to store instructions, data and status information for the application (block 914). For block 914, each application can be allocated a configurable portion of an instruction cache to store instructions for the application, a configurable portion of register banks to store data for the application, one or more context registers to store state information for application, etc. The processing units are dynamically assigned to the applications as requested by those applications (block 916). Threads can be received asynchronously from applications and programmed for execution (block 918). The results of the execution of the threads can be provided asynchronously for the applications (block 920).
The multimedia processor described here can be used for wireless communication devices, portable devices, gaming devices, computing devices, consumer electronic devices, computers, etc. An exemplary use of the multimedia processor for a wireless communication device is described below.
Figure 10 shows a block diagram of a drawing of a wireless communication device 1000 in a wireless communication system. The wireless device 1000 can be a cell phone, a terminal, a telephone device, a personal digital assistant (PDA), or some other device. The wireless communication system can be a code division multiple access system (CDMA), a Global System for mobile communication (GSM), some other system.
or
24/26
The wireless device 1000 is capable of providing bidirectional communication via a receive path and a transmit path. In the reception path, signals transmitted by the base stations are received by an antenna 1012 and supplied to a receiver (RCVR) 1014. The receiver 1014 conditions and digitizes the received signal and provides samples for a digital section 1020 for further processing. In the transmission path, a transmitter (TMTR) 1016 receives data to be transmitted from the digital section 1020, processes and conditions the data, and generates a modulated signal, which is transmitted through the antenna 1012 to the base stations.
The digital section 1020 includes several processing units, interface and memory such as a 1022 modem processor, a 1024 digital signal processor (DSP), a 1026 video / audio processor, a 1028 controller / processor, a processor display unit 1030, a central processing unit (CPU) / reduced instruction set computer (RISC) 1032, a multimedia processor 1034, a camera processor 1036, a cache / internal memory 1038, and an external bus interface (EBI) 1040. The modem processor 1022 performs processing for transmitting and receiving data (for example, encoding, modulation, demodulation and decoding). The DSP 1024 can perform specialized processing for the wireless device 1000. The 1026 video / audio processor performs processing on video content (for example, still images, moving videos, and moving text) for video applications such as camcorder, video playback and video conferencing. The 1026 video / audio processor also performs processing for audio content (for example, synthesized audio) for audio applications. O
25/26 controller / processor 1028 can guide the operation of several units in the digital section 1020. The display processor 1030 performs processing to facilitate the display of video, graphics, and text on a display unit 1050. CPU / RISC 1032 can perform general purpose processing for processors, microprocessors, wireless device 1000. The 1034 multimedia processor performs processing for multimedia applications and can be implemented as described above for figures 1 through 8. The 1036 camera processor performs processing for a camera (not shown in figure 10). Cache / internal memory 1038 stores data and / or instructions for various units in digital section 1020 and can implement cache memory 190 in figures 1 and 2. EBI 1040 facilitates data transfer between digital section 1020 (eg cache / internal memory 1038) and main memory 1060. Multimedia applications can be run by any of the processors in digital section 1020.
The digital section 1020 can be implemented with one or more processors, microprocessors, DSPs, RISCs, etc. The digital section 1020 can also be manufactured in one or more integrated circuits of specific application (ASICs) and / or some other type of integrated circuits (ICs).
The multimedia processor described here can be implemented on several hardware devices. For example, the multimedia processor can be implemented in ASICs, DSPs, digital signal processing devices (DSPDs), programmable logic devices (PLDs), field programmable port arrangements (FPGAs), controllers, microcontrollers, other electronic devices
26/26 electronic units. The multimedia processor may or may not include built-in / integrated memory.
A device that implements the multimedia processor described here can be an independent unit or it can be part of a device. The device may be (i) an independent IC, (ii) a set of one or more ICs that can include memory ICs to store data and / or instructions, (iii) an ASIC such as a mobile station modem (MSM), (iv) a module that can be incorporated into other devices, (v) a cell phone, wireless device, telephone, or mobile unit, (vi) etc.
The previous description of the disclosure is provided to allow anyone skilled in the art to make or use the disclosure. Various changes to the disclosure will be readily apparent to those skilled in the art, and the generic principles defined here can be applied to other variations without departing from the spirit or scope of the disclosure. Thus, the disclosure is not intended to be limited to the examples described here, but the broadest scope compatible with the new principles and characteristics disclosed here must be agreed.
multimedia applications,
Contents7
11 sheets
Sheet 1 Sheet 2 Sheet 3 Sheet 4 Sheet 5 Sheet 6 Sheet 7 Sheet 8 Sheet 9 Sheet 10 Sheet 11
19 members in 10 offices
Priority claims9
| Document | Office | Kind | Date |
|---|---|---|---|
| 11677362 | United States of America | – | |
| 67736207 | United States of America | A | |
| 67736207 | United States of America | A | |
| 2008054620 | United States of America | W | |
| 2008054620 | United States of America | W | |
| 11677362 | – | – | – |
| 2008054620 | – | – | – |
| US20070677362 | – | – | – |
| WO2008US54620 | – | – | – |
Members19
| Document | Office | Kind | |
|---|---|---|---|
| US2008201716A1 | United States of America | A1 | |
| CA2676184A1 | Canada | A1 | |
| WO2008103854A1 | World Intellectual Property Organization (WIPO) | A1 | |
| TW200842757A | Taiwan Province of China | A | |
| KR20090115211A | Republic of Korea | A | |
| EP2126690A1 | European Patent Office (EPO) | A1 | |
| CN101627367A | China | A | |
| US7685409B2 | United States of America | B2 | |
| JP2010519652A | Japan | A | |
| RU2009135022A | Russian Federation | A | |
| RU2425412C2 | Russian Federation | C2 | |
| KR101118486B1 | Republic of Korea | B1 | |
| TWI367453B | Taiwan Province of China | B | |
| JP5149311B2 | Japan | B2 | |
| EP2126690B1 | European Patent Office (EPO) | B1 | |
| BRPI0807951A2This record | Brazil | A2 | |
| CA2676184C | Canada | C | |
| CN101627367B | China | B | |
| BRPI0807951B1 | Brazil | B1 |
4 legal events, as the office reported them to INPADOC
Over the term
Point at a mark for the eventEvents
| Event | Code | |
|---|---|---|
| Patent or certificate of addition of invention granted [chapter 16.1 patent gazette]GrantedPRAZO DE VALIDADE: 10 (DEZ) ANOS CONTADOS A PARTIR DE 14/05/2019, OBSERVADAS AS CONDICOES LEGAIS. (CO) 10 (DEZ) ANOS CONTADOS A PARTIR DE 14/05/2019, OBSERVADAS AS CONDICOES LEGAISB16A | B16A | |
| Decision: intention to grant [chapter 9.1 patent gazette]B09A | B09A | |
| Patent application procedure suspended [chapter 6.1 patent gazette]B06A | B06A | |
| Formal requirements before examination [chapter 6.20 patent gazette]PARECER 6.20B06T | B06T |
Numbers
- Publication
- PI0807951
- Publication, DOCDB
- PI0807951
- Publication, EPODOC
- BRPI0807951
- Application
- 7951
- Application, DOCDB
- PI0807951
- Application, EPODOC
- BR2008PI07951
Titles2
- Portuguese
- PROCESSADOR MULTIMÍDIA MULTI-ENCADEADO SOB DEMANDA
- English
- MULTIMEDIA MULTI-CHAINED PROCESSOR ON DEMAND
Classification
- CPC, 14
- G06F9/5016
- G06F12/0842
- G06F9/30145
- G06F9/30167
- G06F9/382
- G06F9/383
- G06F9/3851
- G06F9/3885
- G06F12/10
- G06F9/45558
- G06F2009/45579
- G06F2009/45583
- Y02D10/00
- G06F9/38
- IPC, 4
- G06F9 38
- G06F9 455
- G06F12 08
- G06F12 10
