Apparatus, system and method for manufacturing an extensible configurable tape for processing images
Abstract
The invention relates to an electronic device and a method for manufacturing an extensible configurable tape for processing images. According to the invention, the electronic device consists of a device () for parallel processing consisting of: a multitude of processing elements, each of them being configurated to execute instructions, a memory subsystem comprising a multitude of memory areas and an interconnection subsystem configured so as to couple the multitude of processing elements and the memory subsystem, the interconnection system including a local interconnection and a global interconnection and a processor () communicating with the device () for parallel processing and which is configured to run a module stored in the memory which is configured to receive a data flow graph associated to a data processing process, the graph comprising a multitude of nodes and a multitude of arcs which connect two or more nodes, each node identifying an operation and each arc identifying a relation between the nodes connected and to allocate a first node of the multitudine of nodes to a first processing element of the device for parallel processing, and a second node from the multitudine of nodes to a second processing element of the device for parallel processing, thus rendering parallel the operations associated with the first node and with the second node.

Term
7.1 yearsto projected expiry
Projected expiry 6 November 2033, counted from filing; an application has no term until it is granted.
- Priority and filed
- Published
- Today
- Projected expiry
30 claims: 3 independent, 27 dependent
- 1CLAIMS REVENDICĂRI 1. A parallel processing device, the processing device comprising:1. Un dispozitiv de procesare paralelă, dispozitivul de procesare cuprinzând: a multitude of processing elements, each configured to execute instructions;o multitudine de elemente de procesare, fiecare configurat să execute instrucțiuni;a memory subsystem comprising a plurality of memory areas, including a first memory area associated with one of the multiple processing elements, this first memory area comprising a plurality of random access memory (RAM) blocks, each block having ports individual reading and writing;and an interconnection system configured to couple the plurality of processing elements and the memory subsystem, this interconnection system including: un subsistem de memorie care cuprinde o multitudine de zone de memorie, inclusiv o primă zonă de memorie asociată unuia dintre multiplele elemente de procesare, această primă zonă de memorie cuprinzând o multitudine de blocuri de memorie cu acces aleatoriu (RAM), fiecare bloc având porturi individuale de citire și de scriere;și un sistem de interconectare configurat să cupleze multitudinea de elemente de procesare și subsistemul de memorie, acest sistem de interconectare incluzând: a local interconnect configured to couple the first memory area and one of the multiple processing elements;and a global interconnect configured to couple the first memory area and the rest of the multiple processing elements. o interconectare locală configurată să cupleze prima zonă de memorie și unul dintre multiplele elemente de procesare;și o interconectare globală configurată să cupleze prima zonă de memorie și restul multiplelor elemente de procesare.
- 12A method of operating a parallel processing system, the method comprising:providing a plurality of processing elements, including a first processing element and a second processing element, each of the multiple process ing elements being configured to execute instructions;12. O metodă de operare a unui sistem de procesare paralelă, metoda cuprinzând: asigurarea unei multitudini de elemente de procesare, inclusiv un prim element de procesare și un al doilea element de procesare, fiecare dintre multiplele elemente de procesare fiind configurat să execute instrucțiuni;asigurarea unui subsistem de memorie care cuprinde o multitudine de zone de memorie, inclusiv o primă zonă de memorie asociată primului element de procesare, această primă zonă de memorie cuprinzând o multitudine de blocuri de memorie cu acces aleatoriu (RAM), fiecare bloc având porturi individuale de citire și de scriere;providing a memory subsystem comprising a plurality of memory areas, including a first memory area associated with the first processing element, this first memory area comprising a plurality of random access memory (RAM) blocks, each block having individual ports reading and writing;receiving - by an arbitration block associated with one of the multiple blocks of RAM memory through a local interconnection of an interconnection system - a first request to access the memory from the first processing element;and sending - by the arbitration block through global interconnection - a first authorization message to the first processing element, to authorize the access of the first processing element to one of the multiple blocks of RAM. recepționarea - de către un bloc de arbitrare asociat unuia dintre multiplele blocuri de memorie RAM printr-o interconectare locală a unui sistem de interconectare - unei prime solicitări de accesare a memoriei din partea primului element de procesare;și trimiterea - de către blocul de arbitrare prin interconectarea globală - unui prim mesaj de autorizare către primul element de procesare, pentru a autoriza accesul primului element de procesare la unul dintre multiplele blocuri de memorie RAM.
- 2021. An electronic device comprising:21. Un dispozitiv electronic cuprinzând: a parallel processing device comprising: un dispozitiv de procesare paralelă cuprinzând: . a multitude of processing elements, each configured to execute instructions;. o multitudine de elemente de procesare, fiecare configurat să execute instrucțiuni;. a memory subsystem comprising a plurality of memory areas, including a first memory area associated with one of the multiple elements of (λ-2 0 1 3 - 0 0 8 1 2 - 0 6 '11 - 2013 processing, this first area memory comprising a plurality of random access memory (RAM) blocks, each block having individual read and write ports;and _ an interconnection system configured to couple the plurality of processing elements and the memory subsystem, this interconnection system including: . un subsistem de memorie care cuprinde o multitudine de zone de memorie, inclusiv o primă zonă de memorie asociată unuia dintre multiplele elemente de (λ-2 0 1 3 - 0 0 8 1 2 - 0 6 ‘11- 2013 procesare, această primă zonă de memorie cuprinzând o multitudine de blocuri de memorie cu acces aleatoriu (RAM), fiecare bloc având porturi individuale de citire și de scriere;și _un sistem de interconectare configurat să cupleze multitudinea de elemente de procesare și subsistemul de memorie, acest sistem de interconectare incluzând: a local interconnect configured to couple the first memory area and one of the multiple processing elements;and a global interconnect configured to couple the first memory area and the rest of the multiple processing elements;o interconectare locală configurată să cupleze prima zonă de memorie și unul dintre multiplele elemente de procesare;și o interconectare globală configurată să cupleze prima zonă de memorie și restul multiplelor elemente de procesare;a processor, communicating with the parallel processing device, configured to run a module stored in memory that is configured: un procesor, comunicând cu dispozitivul de procesare paralelă, configurat să ruleze un modul stocat în memorie care este configurat: receiving a data flow graph associated with a data processing process, the data flow graph comprising a plurality of nodes and a plurality of arcs connecting two or more of the plurality of nodes, each node identifying an operation and each arc identifying a relationship between the connected nodes;and assigning a first node of the plurality of nodes to a first processing element of the parallel processing device, and a second node of the plurality of nodes to a second processing element of the parallel processing device, thus parallelizing the operations associated with the first knot and with the second knot. să recepționeze un grafic de flux de date asociat unui proces de procesare de date, graficul de flux de date cuprinzând o multitudine de noduri și o multitudine de arce care conectează două sau mai multe dintre multitudinea de noduri, fiecare nod identificând o operație și fiecare arc identificând o relație dintre nodurile conectate;și să aloce un prim nod din multitudinea de noduri la un prim element de procesare al dispozitivului de procesare paralelă, iar un al doilea nod din multitudinea de noduri la un al doilea element de procesare din dispozitivul de procesare paralelă, paralelizând astfel operațiile asociate cu primul nod și cu al doilea nod.
Independent claims3
226 paragraphs in 2 sections, as filed
DEVICE, SYSTEMS AND METHODS FOR CARRYING OUT
A CONFIGURABLE AND EXTENSIBLE PICTURE PROCESSING / SEQUENCE TAPE (SEQUENCE IN SEQUENCE) Reference to Related Application [0001] This application claims the benefit of the prior priority date of the UK patent application no. GB1314263.3, titled "Configurable and Extensible Image Processing Tape," which was filed on August 8, 2013 by Linear Algebra Technologies Limited and is fully explicitly included in this document by reference.
Scope This application is generally related to processing devices suitable for image and video processing.
Background Computational image and video processing is very demanding in terms of memory bandwidth, as image resolutions and frame rates are high, with aggregate values of the order of many hundreds of megapixels per second. Moreover, since this domain is relatively in the beginning, new algorithms always appear. That is why it is difficult to fully implement them in hardware, as the hardware components may be unable to adapt to algorithm changes. at the same time, a software approach, based on implementation only at the processor level, is unrealistic. As a result, a flexible architecture / infrastructure is generally desirable, which may include processors and hardware accelerators.
At the same time, the demand for such video and image processing comes largely from portable electronic devices, such as tablets and mobile devices, where power consumption is a key consideration. Therefore, there is a general need for a flexible infrastructure that combines multi-core processors and hardware accelerators with a high bandwidth memory subsystem, allowing them to ensure a sustained data transfer rate under reduced power consumption. , according to the requirements for portable electronic devices.
Summary According to the object of the present application, there is presented an apparatus, systems and methods for realizing a configurable and extensible image processing band.
<img file="RO129804A0_D0001.tif" />
^ - 2 0 1 3 - 0 0 8 1 2 - IL
6 -11- 2013 The object presented includes a parallel processing device. The processing device includes a plurality of processing elements, each configured to execute instructions, and a memory subsystem comprising a plurality of memory areas, including a first memory area associated with one of the multiple processing elements. The first memory area comprises a plurality of random access memory (RAM) blocks, each having individual read and write ports. The parallel processing device may include an interconnection system configured to connect the plurality of processing elements and the memory subsystem. The interconnect system may include a local interconnect configured to couple the first memory area and one of the multiple processing elements, as well as a global interconnect, configured to couple the first memory area and the rest of the plurality of processing elements.
In any of the implementations presented here, one of the multiple RAM blocks is associated with an arbitrage block, this arbitrage block being configured to receive requests for access to memory from one of the multiple processing elements and to allow access to one. from multiple processing elements to one of multiple RAM blocks.
In any implementation presented here, the arbitration block is configured to grant access to one of the multiple blocks of RAM according to a round robin algorithm.
In any of the implementations described herein, the arbitration block comprises a conflict detector configured to monitor memory access requests targeting one of the multiple RAM blocks and to determine if there are two or more of the multiple processing elements that are trying to. to access at the same time the same block between multiple blocks of RAM.
In any of the implementations presented here, the conflict detector is coupled to a plurality of address decoders, each of these multiple address decoders being coupled to one of the multiple processing elements and being configured to determine whether one of the multiple address elements processing tries to access one of the many RAM blocks associated with the arbitrage block.
In any of the implementations presented herein, the plurality of processing elements comprises at least one vector processor and at least one hardware accelerator.
0-2 ^ 3- 0 0 8 1 2 -0 6 -11-2013 In any of the implementations presented here, the parallel processing device includes a plurality of controllers of the memory areas, each configured to provide access to one of the multiple memory areas.
In any of the implementations presented here, the interconnection system comprises a first bus configured to ensure communication between at least one vector processor and the memory subsystem.
In any of the implementations presented here, the interconnection system comprises a second bus system, configured to ensure communication between at least one hardware accelerator and the memory subsystem.
In any of the implementations presented here, the second bus system comprises a memory area address filter configured to mediate communication between at least one hardware accelerator and the memory subsystem, receiving a request to access memory from at least one hardware accelerator and allowing at least one hardware accelerator to access the memory subsystem.
In any of the implementations presented here, one of the multiple processing devices comprises a buffer to increase the flow of the memory system, the number of elements within the buffer being greater than the number of cycles required to retrieve data from the subsystem. of memory.
The subject matter of the present application includes a method of operating a parallel processing system. The method includes providing a plurality of processing elements, including a first processing element and a second processing element, each of the multiple processing elements being configured to execute instructions. The method also includes providing a memory subsystem comprising a plurality of memory areas, including a first memory area associated with the first processing element, this first memory area comprising a plurality of random access memory (RAM) blocks, each block having individual ports for reading and writing. In addition, the method includes receiving - by an arbitration block associated with one of the multiple blocks of RAM by a local interconnection of an interconnection system - a first request to access memory from the first processing element. In addition, the method includes sending - by the arbitration block via global interconnection - a first authorization message to the first processing element, to authorize access of the first processing element to one of the multiple blocks of RAM.
C ~ 2 0 1 3 - oo 8 1 2 - O 6; 11-2013 In any of the implementations presented here, the method also includes receiving
- by the arbitration block through a global interconnection of the system of interconnection to a second request to access the memory from a second processing element, as well as sending - by the arbitration block through global interconnection - to a second message of authorization to the second processing element to authorize the access of the second processing element to one of the multiple blocks of RAM.
In any of the implementations presented here, the method also includes sending by the arbitration block - a plurality of authorization messages to the plurality of processing elements, to authorize access to one of multiple RAM blocks according to a round robin algorithm.
In any of the implementations presented here, the method also includes monitoring
- by a conflict detector in the composition of the arbitration block - the requests for accessing the memory addressed to one of the multiple blocks of RAM, as well as identifying the situation in which two or more of the multiple processing elements try to access the same block between the multiple RAM blocks.
In any of the implementations presented herein, the plurality of processing elements comprises at least one vector processor and at least one hardware accelerator.
In any of the implementations presented herein, the method further includes providing a plurality of controllers of the memory areas, each controller being configured to provide access to one of the multiple memory areas.
In any of the implementations presented here, the method further includes ensuring communication between at least one vector processor and the memory subsystem through a first bus system of the interconnection system.
In any of the implementations presented herein, the method further includes ensuring communication between at least one hardware accelerator and the memory subsystem through a second bus system of the interconnection system.
In any of the implementations presented here, the second bus system comprises a memory area address filter configured to mediate communication between at least one hardware accelerator and the memory subsystem, receiving a request to access memory from at least one hardware accelerator and allowing at least one hardware accelerator to access the memory subsystem.
CV 2 0 1 3 - 0 0 8 1 2 0 6 -11- 2013 The subject matter of the application includes an electronic device. The electronic device includes a parallel processing device. The processing device includes a plurality of processing elements, each configured to execute instructions, and a memory subsystem comprising a plurality of memory areas, including a first memory area associated with one of the multiple processing elements. The first memory area comprises a plurality of random access memory (RAM) blocks, each having individual read and write ports. The parallel processing device may include an interconnection system configured to connect the plurality of processing elements and the memory subsystem. The interconnect system may include a local interconnect configured to couple the first memory area and one of the multiple processing elements, as well as a global interconnect, configured to couple the first memory area and the rest of the plurality of processing elements. Also, the electronic device includes a processor that communicates with the parallel processing device, configured to run a module stored in memory. The module is configured to receive a data flow graph associated with a data processing process, the data flow graph comprising multiple nodes and multiple arcs connecting two or more of the multiple nodes, each node identifying an operation and each arc identifying a relationship between the connected nodes; and assigning a first node of the multiple nodes to a first processing element of the parallel processing device, and a second node of the multiple nodes to a second processing element of the parallel processing device, thus parallelizing the operations associated with the first node and with the second node.
In any of the implementations presented here, the data flow graph is presented in an extensible markup language (XML) format.
In any of the implementations presented herein, the module is configured to allocate the first node in the plurality of nodes to the first processing element based on a previous performance of a memory subsystem in the parallel processing device. In any of the implementations presented here, the memory subsystem of the parallel processing device comprises a counter configured to count a number of memory conflicts within a predetermined period of time, and the previous performance of the memory subsystem comprises the number of conflict conflicts. memory measured by the meter. In any of the implementations presented here, the module is configured to allocate the first node in the plurality of nodes to the first processing element while the parallel processing device operates at least a portion of the data flow graph.
Ρ- 2 ^ 1 3 - 0 0 8 1 2 -d 6 Ή- 2013 \ S In any of the implementations presented here, the module is configured to receive a plurality of data flow graphs and to allocate all operations associated with the multitude. of flow graphs to a single processing element in the parallel processing device.
In any of the implementations presented herein, the module is configured to interleave memory accesses by the processing elements, to reduce memory conflicts.
In any of the embodiments presented herein, the electronic device comprises a mobile device.
In any of the implementations presented here, the data flow graph is specified using an application programming interface (API) associated with the parallel processing device.
In any of the implementations presented here, the module is configured to provide input image data to the plurality of processing elements by dividing the image input data into strips and providing an image input data strip to one of the multiple processing elements.
In any of the implementations presented here, the number of image input data strips is identical to the number of elements between the multiple processing elements. Schema description The present application will be described below with reference to schematics.
FIG. 1 describes a computational image processing platform from Chimera.
FIG. 2 describes a multicore architecture of a Cell processor.
FIG. 3 describes an architecture with an efficient low power microprocessor (ELM).
FIG. 4 illustrates an improved memory subsystem according to certain implementations.
FIG. 5 illustrates a section of the parallel processing device according to certain implementations.
FIG. 6 illustrates a centralized collision detection system in a block control logic according to certain implementations.
α'2013-00812-0 δ -îl- ^ 3 FIG. 7 illustrates a distributed collision detection system in a block control logic according to certain implementations.
FIG. 8 presents an arbitration block for reporting a collision signal to an applicant according to certain implementations.
FIG. 9 illustrates a cycle oriented arbitration block according to certain implementations.
FIG. 10 illustrates a mechanism to reduce the latency of memory access due to arbitrage access to memory according to certain implementations.
FIG. 11 illustrates a planning software application according to certain implementations.
FIG. 12 illustrates a hierarchical structure of a system that contains a parallel processing device according to certain implementations.
FIG. 13 illustrates how the description of the directed acyclic graph (DAG) or the data flow graph can be used to control the operations of a parallel processing device according to certain implementations.
FIG. 14A-14B illustrates the planning and issuing of activities by the compiler and the planner according to certain implementations.
FIG. 15 illustrates the operation of a real-time DAG compiler according to certain implementations.
FIG. 16 compares a schedule generated by an OpenCL planner with a schedule generated by the proposed online DAG planner according to certain implementations.
FIG. 17 illustrates a barrier mechanism for synchronizing an operation of processors and / or filter accelerators according to certain implementations.
FIG. 18 illustrates the parallel processing device with different types of processing elements according to certain implementations.
FIG. 19 illustrates the proposed multi-core memory subsystem according to certain implementations.
FIG. 20 illustrates a single area of the connection matrix-CMX infrastructure according to certain implementations.
FIG. 21 illustrates a beam architecture (switch architecture) for accelerator memory controller (AMC) according to certain implementations.
<sup>c</sup>'' θ 1 3 - 0 0 8 Î 2 - 0 6; 11-2013 FIG. 22 illustrates an AMC controller with beam ports according to certain implementations.
FIG. 23 illustrates a read operation using a CMA according to certain implementations.
FIG. 24 illustrates a write operation using an AMC according to certain implementations.
FIG. 25 illustrates the parallel processing device according to certain implementations.
FIG. 26 illustrates an electronic device that includes a parallel processing device according to certain implementations.
Detailed Description One of the possible ways to interconnect such different processing resources (for example, processors and hardware accelerators) is by using a bus like the one presented in the computational engine for photographs from Chimera, designed by NVidia. FIG. 1 illustrates the computational engine for photos from Chimera. Chimera Photo Computing Engine 100 includes multiple graphics processing unit (GPU) cores 102 connected to a multinucleus ARM processor subsystem 104 and image signal processing hardware (HW) accelerators (ISP) 106 via an unarchived bus infrastructure 108 (for example, a single-hierarchy bus system that connects all processing elements). The computational engine for photos from Chimera is generally presented as a software framework that abstractes the details of the base 102 GPUs, the CPUs 104 and the ISP blocks 106 from the programmer. In addition, the computational engine for photographs 100 from Chimera describes the data flow through the computational photographic engine as having been accomplished by two information buses 108-0, 108-1, the first bus 108-0 carrying image or frame data, and the second bus. 108-1 carrying status information associated with each frame.
The non-hierarchical bus infrastructure, as in the case of Chimera, can be cheap and convenient to implement. However, non-hierarchical bus infrastructure can have several considerable disadvantages if used as a means of interconnecting heterogeneous processing elements (for example, processing elements of different types), such as GPU cores 102, CPUs 104 and ISP 106 blocks. first, the use of a bus to interconnect computational resources implies the possibility to distribute 8 e <sup>?</sup> 3 - 0 0 8 1 2 - ο 6 -H- 2013 memory in the entire local system at each central processing unit (CPU) 104, at a graphics processing unit (GPU) 102 and / or at a signal processor block image type (ISP) 106. Therefore, memory cannot be allocated flexibly within the processing band according to the computational band requirements for photographs that the programmer wants to implement. This lack of flexibility may either make it difficult to implement certain aspects of image and video processing, or limit implementation in terms of frame rate, image quality or other points of view.
Secondly, the use of unarchived bus infrastructure can lead to the situation where different computational resources (CPUs 104, GPUs 102 and ISP blocks 106) have to compete for the bandwidth of the bus. This competition requires arbitrage, which reduces the bandwidth available on the bus. Therefore, the theoretical bandwidth available for the actual activity progressively decreases. As a result of reducing bandwidth, a processing band may no longer meet the performance requirements of the application in terms of frame rate, image quality and / or power.
Third, the lack of sufficient memory near a certain computational resource may require the transfer of data back and forth between the memory associated with a particular GPU 102, CPU 104, or hardware ISP block 106, and another computational resource. This unavailability of memory can lead to additional bus bandwidth consumption and arbitrary overhead. Moreover, the unavailability of memory also increases power consumption. Therefore it can be difficult or even impossible to support a certain algorithm at a certain target frequency of the frames.
Fourth, using a non-hierarchical bus infrastructure may create difficulties in constructing a processing band from heterogeneous processing elements, which may each have different latency characteristics. For example, GPU cores 102 are designed to tolerate latency by running multiple superimposed process threads to support multiple delayed memory accesses (generally external DRAM memory) so as to cover latency, while ordinary CPUs 104 and ISP hardware 106 blocks are not designed to tolerate latency.
Another way of interconnecting different processing resources is found in an IBM-designed Cell processor architecture, illustrated in FIG. 2. Cell 200 processor architecture includes local storage (LS) 202 available to each processor 204, also known as synergistic execution unit 9 (V 2 0 1 3 - 0 0 8 1 2 -0 6 -11- 2013
V
SXU). The Cell 200 processor relies on a time-shared infrastructure and direct memory access (DMA) transfers 206 to programmatically plan data transfers between LS 202 of one processor and LS 202 of another processor. The difficulty of the Cell 200 architecture lies in the complexity faced by the programmer in trying to explicitly plan background data transfers with hundreds of cycles ahead (due to the high latency levels in the Cell 200 architecture), so as to ensure the availability of shared data for each 204 processor at the required time. If the programmer does not explicitly schedule background data transfers, the 204 processors may be blocked, which would affect performance.
Another way to interconnect different processing resources is by using a shared multi-core memory subsystem, for efficient data sharing between processors of a multi-core processing system. This multi-core shared memory subsystem is used in the efficient lowpower microprocessor (ELM) system. FIG. 3 illustrates the ELM system. The ELM 300 system includes an assembly 302, the primary physical unit for computing resources in an ELM system. Assembly 302 includes a cluster of four poorly coupled processors 304. The four-processor cluster 304 shares local resources, such as an assembly memory 306 and an interface with the interconnection network. The assembly memory 306 captures instructions and data in working sets close to the processors 304, and the memory banks are configured so that the memory can be accessed simultaneously by the local processors 304 and the network interface controller. Each processor 304 within an assembly 302 is allocated a preferred memory bank from the memory of the assembly 306. The accesses of a processor 304 to its preferred bank benefit from priority over the accesses of other processors and the network interface (which they will block). Instructions and data belonging to only one of the processors 304 can be stored in its preferred memory bank, to ensure deterministic access times. Referees who control access to read and write ports tend to establish affinity between processors 304 and memory banks 306. In this way, the software can more confidently estimate the available bandwidth and latency when accessing data that can be shared by multiple processors.
However, the ELM 300 architecture may have high power consumption due to physically large random access memory (RAM) blocks. Furthermore, the ELM 300 architecture may suffer from low throughput if many data is shared between processors 304. In addition, there is no provision for data sharing between (V 2 0 1 3 - 0 0 8 1 2 -fl 6 J1- 2013
<img file="RO129804A0_D0002.tif" />
304 processors and hardware accelerators, which in some cases can be advantageous in terms of power and performance.
Those presented here relate to a device, systems and methods designed to enable multiple processors and hardware accelerators and access data shared with other processors and hardware accelerators. This document presents a device, systems, and methods for simultaneously accessing shared data, without being blocked by a local processor that has a high affinity (for example, a higher priority) for accessing local storage.
The apparatus, systems and methods presented provide substantial advantages over existing multinucleus memory architectures. Existing multicore memory architectures use a single monolithic block of RAM for each processor, which can limit the bandwidth at which data can be accessed. The architecture presented may provide a mechanism for accessing memory at a considerably higher bandwidth compared to existing multicore memory architectures that use a single monolithic RAM block. The architecture shown achieves this higher bandwidth by instantiating multiple physical RAM blocks per processor, rather than instantiating a single large RAM block per processor. Each RAM block can include a dedicated access arbitrage block and surrounding infrastructure. Therefore, each block of RAM in the memory subsystem can be accessed independently by others by multiple processing elements in the system, such as vector processors, reduced instruction processor (RISC), hardware accelerators or DMA engines.
It is somewhat counterintuitive that the use of multiple small RAM instances is advantageous compared to the use of a single large RAM instance, because a memory bank based on a single large RAM instance has a higher spatial efficiency compared to one based on multiple RAM instances. smaller. However, the power dissipation for smaller RAM instances is generally significantly reduced compared to a single large RAM instance. Moreover, if a single large physical RAM instance would reach the same bandwidth as the multiple instance RAM blocks, the large physical RAM instance would have a substantially higher power consumption than the cumulative consumption of multiple physical RAM instances. Therefore, at least from the perspective of power dissipation, the memory subsystem may have more to gain if multiple physical RAM instances are used than if a single large RAM instance is used.
The memory subsystem with multiple physical RAM instances can have an added advantage in that, in general, the cost per RAM access - for example, the access time of (£ -2 0 1 3 - 0 0 8 1 2 -0 6 -.11- 2013 memory or power consumption - is much lower for smaller RAM blocks than larger RAM blocks, this is due to the shorter bit lines used to read / write data from RAM blocks. In addition, the access time for read and write operations in the case of smaller RAM blocks is also reduced (due to the reduced time constants of the resistor-capacitor (RC) circuits associated with the shorter bit lines). Therefore, the processing elements coupled to the memory subsystem with multiple RAM blocks can operate at a higher frequency, which reduces the static power due to the resting current losses. This can be especially useful when processors and memory are isolated in areas of power. For example, when a particular processor or filter accelerator has completed its activity, the power range associated with the respective processor or filter accelerator may be advantageously isolated. Therefore, the memory subsystem in the presented architecture has superior characteristics in terms of available bandwidth and power dissipation.
In addition, a memory subsystem with multiple RAM instances, each with arbitrary accesses, can provide numerous data sharing modes between processors and hardware accelerators, without dedicating a RAM block to a particular processor by blocking the RAM block. In principle, if a larger RAM is subdivided into N sub-blocks, the bandwidth available for data increases approximately by the N factor. This idea assumes that data can be partitioned in a timely manner to reduce concomitant sharing (for example, an access conflict) by multiple processing elements. For example, when a consumer processor or consumer accelerator reads data from a data buffer fed by a manufacturing processor or a production accelerator, a concurrent sharing of a data buffer occurs, resulting in a access conflict.
In certain implementations, the presented architecture may provide mechanisms for reducing concomitant data sharing. In particular, the presented architecture can be adapted to reduce the concomitant sharing through a static memory allocation mechanism and / or a dynamic memory allocation mechanism. For example, in the static memory allocation mechanism, the data is mapped to different portions of memory before the program is launched - for example, in the compilation phase of the program -, to reduce the concomitant sharing of data. On the other hand, in the dynamic memory allocation model, the data is mapped to different portions of memory during program execution. The static memory allocation mechanism provides a predictable memory allocation mechanism for data <2 0 1 3 - 0 0 8 1 2 - Ο 6 -11- 2013
<img file="RO129804A0_D0003.tif" />
and is not characterized by substantial overhead in terms of power or performance, As an example, the presented architecture may be used in conjunction with a scheduler running on a controller (eg, a RISC supervisory processor) or one or multiple processors that mediate access to partitioned data structures on multiple RAM blocks. The scheduler can be configured to interleave the start times of different processing elements working on fragments (for example, lines or blocks) of data (for example, an image frame) so as to reduce simultaneous access to shared data.
In certain implementations, a hardware arbitrage block may be added to the planner. For example, the hardware arbitrage block can be configured to mediate processor accesses (such as vector processors) to shared memory through a shared deterministic interconnect, designed to reduce processor blockage. In some cases, the hardware arbitrage block can be configured to perform cycle-oriented planning. Cycle-oriented planning may include planning for use of a resource at processor cycle granularity, not activity level granularity, which may require multiple processor cycles. Planning resource allocation to processor cycle granularity can ensure enhanced performance.
In other implementations, a plurality of hardware accelerators may be added to the scheduler, each of which may include an input buffer and an output buffer for data storage. The input buffer and the output buffer can be configured to absorb (or hide) the variation of delays that occur in accessing external resources, such as external memory. Input and output buffer may include a first-in, first-out (FIFO) buffer, and FIFO buffer may include a sufficient number of slots to store a sufficient amount of data and / or instructions so as to absorb the variation of delays that occur in accessing external resources.
In certain embodiments, the apparatus, systems and methods presented provide a parallel processing device. The parallel processing device may include a plurality of processors, such as a parallel processor, each of which may execute instructions. The parallel processing device may also include a plurality of memory areas, each memory area being associated with one of the parallel processing devices and giving the respective processor preferential access in d - 2 0 1 3 0 0 8 1 2 - 0 6 -11- 2013 related to other processing devices in the parallel processing device. Each memory area may include a plurality of RAM blocks, each RAM block may include a read port and a write port. In some cases, each memory area may be provided with a memory zone controller to provide access to a connected memory area. Processors and RAM blocks can be connected to each other via a bus. In some cases, the bus can connect any of the processors to any of the memory areas. Where appropriate, each RAM block may include a control logic on the memory block. The logic of control on the memory block is sometimes called the logic of control on the block or block of arbitration.
In some embodiments, the parallel processing device may also include at least one hardware accelerator configured to perform a predefined processing function, such as image processing. In some cases, the predefined processing function may include a filtering operation.
In some embodiments, at least one hardware accelerator may be coupled to the memory areas through a separate bus. The separate bus may include an Accelerator Memory Controller (AMC), configured to receive requests from at least one hardware accelerator and allow the hardware accelerator access to a memory area through the corresponding memory area controller. It is thus appreciated that the path of memory access applied by hardware accelerators may be different from that applied by vector processors. In some implementations, at least one hardware accelerator may include an internal memory buffer (for example, a FIFO memory) to compensate for delays in accessing memory areas.
In some embodiments, the parallel processing device may include a host processor. The host processor can be configured to communicate with the AMC via a host bus. The parallel processing device can also be provided with an application programming interface (API). The API provides a high-level interface for vector processors and / or hardware accelerators. In some embodiments, the parallel processing device may operate in conjunction with a compiler providing instructions for the parallel processing device, in some cases, the compiler is configured to run on a host processor, which is distinct from the processing elements, such as be it a vector processor or a graphics accelerator. In some cases, the compiler is configured to receive a data flow graph through the Image / Video API 1206 (FIG. 12), specifying an image processing process. Compiler cr 2 0 1 5 - OO 8 1 2 - O 6 Jul 2013
<img file="RO129804A0_D0004.tif" />
it can also be configured to map one or more aspects of the data flow graph to one or more processing elements, such as a vector processor or a hardware accelerator. In some implementations, a data flow graph may include nodes and arcs, each node identifying an operation and each arc identifying a relationship between nodes (for example, operations), such as the order in which the operations are performed. The compiler can be configured to allocate a node (for example, an operation) to one of the processing elements to parallelize the calculation of the data flow graph. In some implementations, the data flow graph may be provided in an extensible markup language (XML) format. In some implementations, the compiler can be configured to allocate multiple data flow graphs to a single processing element.
In some implementations, the parallel processing device may be configured to measure its performance and communicate the information to the compiler. Therefore, the compiler can use the previous performance information received from the parallel processing device to determine the allocation of current activities to the processing elements in the parallel processing device. In some implementations, performance information may indicate a number of access conflicts that occur on one or more processing elements in the processing device.
In some cases, the parallel processing device can be used in video applications, which can be computationally expensive. In order to meet the computational demand of video applications, the parallel processing device can configure its memory subsystem so as to reduce the access conflicts between the processing units while accessing the memory. For this purpose, as discussed above, the parallel processing device may subdivide the monolithic memory banks into multiple physical RAM instances, instead of using the monolithic memory banks as a single physical memory block. Through this subdivision, each physical RAM instance can be arbitrated for write and read operations, thus increasing the bandwidth available so many times, how many physical RAM instances exist in the memory bank.
In some implementations, cycle-oriented hardware arbitrage may also provide for more programmable traffic classes and scheduling masks. Multiple traffic classes and programmable scheduling masks can be controlled using the scheduler. The cycle-oriented hardware arbitrage block can include a port arbitrage block, which can be configured to allocate a unique shared resource to multiple applicants according to a round robin algorithm. In the round robin algorithm, applicants (e.g.
<img file="RO129804A0_D0005.tif" />
<img file="RO129804A0_D0006.tif" />
6
20 »the processing elements) receive access to a resource (for example, memory) in order to receive the requests from the applicants. In some cases, the port arbitration block can intensify the round robin algorithm to take into account multiple classes of traffic. The single shared resource may include a RAM block, shared registers, or other resources that vector processors, filter accelerators, and RISC processors can access to share data. In addition, the arbitration block may allow the disruption of resource allocation according to the round robin algorithm with a priority vector or maximum priority vector. The priority vector or the highest priority vector can be provided by a planner to prioritize certain traffic classes (for example, video traffic classes), depending on the needs of the particular application of interest.
In some embodiments, a processing element may include one or more processors, such as a vector processor or a vector processing unit with hybrid streaming architecture, a hardware accelerator and a hardware filter operator.
FIG. 4 illustrates a parallel processing device with a memory subsystem, which allows multiple processors (for example, hybrid processing units for streaming - SHAVE) to share a multi-port memory subsystem according to implementations. In particular, FIG. 4 shows a parallel processing device 400, which is suitable for processing image and video data. The processing device 400 comprises a plurality of processing elements 402, such as a processor. In the configuration exemplified in FIG. 4, processing device 400 includes 8 processors (SHAVE 0 402-0 - SHAVE 7 402-7). Each processor 402 may include two read-write units 404, 406 (LSU0, LSU1), through which data can be read from memory and written to memory 412. Each processor 402 may also include an instruction unit 408 in which instructions may be loaded. A particular implementation in which the processor includes a SHAVE, this SHAVE may include one or more reduced instruction set (RISC) processors, digital signal processors (DSP), a very long instruction word processor instruction word -VLIW) and / or a graphics processing unit (GPU). Memory 412 comprises a plurality of memory areas 412-0 ... 412-7 here referred to as connection matrix (CMX) areas. Each memory area 412 is associated with an appropriate processor 402-7.
The parallel processing device 400 also includes an interconnection system 410 which couples the processors 402 and the memory areas 412. The de \ -2 0 1 3 - 0 0 8 1 2 -0 6 61-2013 interconnection system 410 is called here inter-SHAVE interconnection (ISI). ISI may include a bus through which the processors 402 can read or write data to any component of any memory area 412.
FIG. 5 illustrates a section of the parallel processing device according to certain implementations. Section 500 includes a single 402-N processor, a 412-N memory area associated with the 402-N single processor, ISI 410 that couples the 402-N single processor, and other memory areas (not shown in the figure) and a block control logic. 506 for arbitrage communication between a block in memory area 412-N and processors 402. According to the illustration in section 500, the 402-N processor can be configured to directly access the 412-N memory area associated with the 402-N processor; the 402-N processor can also access other memory areas (not shown in the figure) via ISI.
In certain embodiments, each 412-N memory area may include a plurality of RAM blocks or physical RAM blocks 502-0 ... 502-N. For example, a 412-N memory area with a capacity of 128 kB can include four 32-kB memory blocks with a single port (for example, physical RAM elements) organized in the form of 4,000 32-bit words. In some implementations, a block 502 may also be called a logical RAM block, in some implementations, a block 502 may include a single-port metal-oxide-semiconductor (CMOS) complementary RAM. The advantage of a single-port RAM is that it is generally available in most semiconductor processes. In other implementations, a block 502 may include a multi-port CMOS RAM. In some implementations, each block 502 may be associated with a control logic on block 506. The control logic on block 506 is configured to receive requests from processors 402 and to allow access to the individual read and write ports of the block. associated 502. For example, when a 402-N processing element wants to access data in a 502-0 RAM block, before the 402-N processing element sends the request for memory data directly to the 502-0 RAM block, the 402- processing element N cannot send a memory access request to the control logic on block 506-0 associated with RAM block 502-0. The request to access the memory may include a memory address of the data requested by the 402-N processing element. Subsequently, the block control logic 506-0 can analyze the memory access request and determine whether the 402-N processing element can access the requested memory. If the 402-N processing element can access the memory required, the control logic on the 506-0 block can send a
<img file="RO129804A0_D0007.tif" />
<2 0 1 3 “0 0 8 1 2 - 0 6-.11- 2013 message to allow access to the 402-N processing element, and subsequently, the 402-N processing element can send a memory data request to 502-0 RAM block. Since there is the possibility of simultaneous access by multiple processing elements, in some implementations, the control logic on block 506 may include a conflict detector, which is configured to detect a situation where two or more processing elements, such as be it a processor or an accelerator, it tries to access any of the blocks in a memory area. The conflict detector can monitor access to each block 502 to detect a simultaneous access attempt. The conflict detector can be configured to report to the planner that an access conflict has to be resolved.
FIG. 6 illustrates a centralized collision detection system in a block control logic according to certain implementations. The conflict detection system may include a centralized arbitration block 608, which includes a plurality of conflict detectors 604 and a plurality of one-hot (one-bit active) 602 encoders . In some implementations, the one-hot address encoder 602 is configured to receive a memory access request from one of the processing elements 402 and to determine whether the memory access request is for data stored in the RAM block 502 associated with the address encoder. one-hot 602. Each conflict detector 604 may be coupled to one or more one-hot address encoders 602, which are also coupled to one of the processing elements 402 that can access the block 502 associated with the conflict detector 602. In some embodiments, a conflicts 604 can be coupled to all one-hot address encoders 602 associated with a particular RAM block 502.
If the request for accessing the memory is for data stored in the RAM block 502 associated with the one-hot address encoder 602, then the one-hot address encoder 602 can communicate a bit value "l" to the respective conflict detector 604. RAM block if the memory access request is not for data stored in the RAM block 502 associated with the one-hot address encoder 602, and then the one-hot address encoder 602 can communicate a bit value "0" to the conflict detector 604 of the respective RAM block.
In some embodiments, the one-hot address encoder 602 is configured to determine whether the access request is for data stored in the RAM block 502 associated with the one-hot address encoder 602 by analyzing the target address of the memory access request. For example, when the RAM block 502 associated with the one-hot address encoder 602 is α- 2 0 1 3 - 0 0 8 1 2 - 0 6 -J1- 2013 designated with a memory address range from 0x0000 to OxOOff, then the one-hot 602 address encoder can determine if the target address of the memory access request is in the range 0x0000 to OxOOff. If yes, the memory access request is for data stored in RAM block 502 associated with the one-hot address encoder 602; if not, the memory access request is not for data stored in the RAM block 502 associated with the one-hot address encoder 602. In some cases, the one-hot address encoder 602 may use a range comparison block to determine whether the target address of the memory access request is within the address range associated with a 502 RAM block.
Once the conflict detector 604 receives bit values from all one-hot address encoders 602, the conflict detector 604 can count how many "1" exist in the received bit values (for example, it can collect the bit values). to determine if there is more than one processing element 402 that at that time requests access to the same RAM block 502. If there is more than one processing element that requests access to the same RAM block 502, the conflict detector 604 may report a conflict. FIG. 7 illustrates a distributed collision detection system in a block control logic according to certain implementations. The distributed conflict detection system may include a distributed arbitrator 702, which includes a plurality of conflict detectors 704. The functioning of the distributed conflict detection system is substantially similar to the functioning of the centralized conflict detection system. In this case, the conflict detectors 704 are arranged distributed. In particular, the distributed arbitrator 702 may include conflict detectors 704 which are arranged in series, each conflict detector 704 being coupled to a single subset of one-hot address encoders 602 associated with a particular RAM block 502. This arrangement is different from the centralized conflict detection system, in which a conflict detector 704 is coupled to all one-hot address encoders 602 associated with a particular RAM block 502.
For example, when a particular RAM block 502 can be accessed by 64 processing elements 402, a first conflict detector 704-0 may receive a memory access request from 32 processing elements, and a the second conflict detector 704-1 can receive a memory access request from the other 32 processing elements. The first conflict detector 704-0 can be configured to analyze one or more memory access requests received from the 32 processing elements coupled to it and determine a first number of elements, out of the 32 coupled processing elements. to it, which requests access to a particular 502-0 RAM block. In parallel, the second detector of ^ 2013-00872--.
ο 6-11-11 2013 W?
conflicts 704-1 can be configured to analyze one or more memory access requests received from the 32 processing elements coupled to it and determine a first number of elements, out of the 32 processing elements coupled to it, which requests access to the respective 502-0 RAM block. Then, the second conflict detector 704 can gather the first and second numbers to determine how many of the 64 processing elements require access to the respective 502-0 RAM block.
Once a conflict detection system detects a conflict, the conflict detection system can send a stop signal to an applicant 402. FIG. 8 presents an arbitration block for reporting a collision signal to an applicant according to certain implementations. Further particularizing, the outputs of the interval comparison blocks from the conflict detection systems are combined using an OR gate to generate a stop signal for the applicant. The signal half indicates that there is more than one processing element trying to access the same physical RAM subblock within the memory area associated with the requestor. Upon receiving the stop signal, the applicant may interrupt the operation of accessing the memory until the conflict disappears. In some implementations, the conflict may be removed from hardware independent of the program code. In some implementations, the arbitration block may operate at cycle granularity. In such implementations, the arbitration block allocates resources to processor cycle granularity, not activity level granularity, the latter being able to include multiple processor cycles. Such cycle-oriented planning can enhance system performance. The arbitration block can be implemented in hardware, so that the arbitration block can perform cycle oriented planning in real time. For example, in any particular instance, the arbitration block implemented in the hardware can be configured to allocate resources for the next processor cycle.
FIG. 9 illustrates a cycle oriented arbitration block according to certain implementations. The cycle-oriented arbitration block may include an arbitration block for ports 900. The arbitration block for ports 900 may include a first port selection block 930 and a second port selection block 932. The first port selection block 930 is configured to determine which of the memory access requests (identified as a bit position in the client request vector) is allocated to the memory area port [0] to access a coupled memory area at the port [0] of the zone, and the second selection block 932 is configured to determine which of the request vectors you have
Α-2 0 1 3 - 0 0 8 1 2 - β 6 -11- 2013 Kh * / the client is assigned to the port [1] of the area for accessing a memory area coupled to the port [1] of the area.
The first port selection block 930 includes a leading "1" detector (LOD) 902-0 and a second LOD 902-1. The first LOD 902-0 is configured to receive a request vector from the client, which can include a plurality of bits. Each bit in the request vector from the client indicates whether or not a message access request has been received from a requestor associated with the position of the respective bit. In some cases, the request vector from the client operates in the "logical level 1" mode. Once the first LOD 902-0 receives the request vector from the client, the first LOD 902-0 is configured to detect a bit position, counting from left to right, where the request becomes for the first time non-zero, thus identifying the first request to access the memory, counting from left to right, to the first block for selecting ports 930. In parallel, the client request vector can be masked by a logic operator AND 912 to generate a mask client request vector generated by a mask register 906 and a mask shifter to the left 904. The mask register 906 can be set by a processor communicating with the mask register 906, and the mask shifter to the left 904 may be configured to move to the left the mask represented by the mask register 906. The second LOD 902-1 can receive the masked vector of the client request from the logic operator AND 912, then detecting the first 1 of the masked vector of the client request.
The output from the first LOD 902-0 and the second LOD 902-1 are then forwarded to port winner selection block 908 [0]. Port Winner Selection Block 908 [0] receives two more additional entries: a priority vector and a maximum priority vector. Port Winner Selection Block 908 [0] is configured to determine which of the received memory access requests should be allocated to port [0] of the memory area, based on input priorities. In some implementations, the priorities of the entries can be sorted as follows: starting from the highest priority vector, continuing with the priority vector that divides the masked LOD vector into priority and non-priority requests, followed by the unmasked LOD vector, which has the lowest priority. in other implementations other priorities may be specified.
While the first port select block 930 can be configured to determine whether the client request vector can be allocated to the port [0] of the memory area, the second port select block 932 can be configured to determine if ¢ -2 0 1 3 - 0 0 9 ^ 2-0 6 -11- 2U13
<img file="RO129804A0_D0008.tif" />
the request vector from the client can be assigned to the port [1] of the memory area. The second port selection block 932 includes a first trailing one detector (TOD) 912-0, a second TOD 912-1, a mask register 914, a mask shifter to the right 916, a block 918 of the winner of the port [1] and a masking block with logic ȘI 920. TOD 912 is configured to receive a request vector from the client, which can include a plurality of bits, and detect a bit position, counting from right to left, for which the vector first becomes non-zero. The operation of the second port selection block 932 is substantially similar to the first port selection block 930, except that it operates from the right to the left of the input vector, selecting the last 1 of the input request vector using a detector. of the last 1 912-0.
Outputs of port winners selection blocks 908, 918 are also communicated to the same winners detection block 910, which is configured to determine whether the same memory access request has gained access to both the port [0] of the area. , as well as the port [1] of the area. If the same request vector from the client has gained access to both the port [0] of the area and to the port [1] of the area, the same block of detection 910 of the winners selects one of the zone ports to route the request and allocates the other immediately lower request port in the input vector. This avoids the over-allocation of resources for a specific request, thus improving the allocation of resources to competing applicants.
Port arbitration block 900 works as follows: starting from the left side of the 32-bit client request vector and the hidden LOD 902-1, it communicates the position of the first masked request vector, if this masked request vector is not exceeded by a higher priority entry coming from the priority vectors or with the highest priority; the applicant corresponding to the LOD position wins and receives access to the port [0]. The LOD position is also used to advance the position of the mask using the left-hand shifter 904 on 32 bits and is also used for comparison with the LOD allocation at port 1, to verify if the same applicant has received access to both ports, in case wherein only one port is granted, switching a bistable to grant alternate access between ports 0 and 1 in the case of successive detections of the same winner. If the LOD output from the masked detector 902-1 was given priority by an appropriate bit 1 of the priority vector, the requesting client receives access to port 0 for 2 consecutive cycles. if there is no first 1 in the masked vector of the client request and there is no request with ^ -2 0 1 3 - 0 0 8 1 2 -0 6 -11- 2013 higher priority, the unmasked LOD wins and receives access to port 0. In any of the above, bit 1 of the vector with the highest priority will take precedence over any previous request and grant the applicant unrestricted access to port 0.
The logic at the bottom of the diagram starts from the right of the request vector, otherwise it functions in the same way as the upper part, which starts from the left of the request vector. In this case, the operation of the arbitration block of port 1 is identical to the portion of port 0 of the logic from the point of view of priorities, etc.
In some embodiments, a processing element 402 may include a memory buffer to reduce a latency of memory access due to arbitrary access to memory. FIG. 10 illustrates a mechanism to reduce the latency of memory access due to arbitrage access to memory according to certain implementations. In a typical memory access arbitrage model, the memory access arbitrage block consists of a band, which results in a fixed overhead arbitration penalty when allocating a shared resource, such as a RAM block. 502, to one of the multiple processing elements (for example, applicants). For example, when an applicant 402 sends a memory access request to an arbitration block 608/702, it takes at least four cycles until the applicant 402 receives an access message, because each of the next steps takes at least one cycle: (1) analyzing the request to access the memory to the one-hot address encoder 602, (2) analyzing the output of the one-hot address encoder 602 to the arbitration block 608/702, (3) sending a message granting access to the encoder. of one-hot address 602 from the arbitration block 608/702 and (4) sending the message of granting access to the applicant 402 from the one-hot address encoder 602. Subsequently, the requester 402 must send a memory data request to the RAM block 502 and receive data from the memory block 502, each of these steps lasting at least one cycle. Therefore, a memory access operation has a latency of at least five cycles. This fixed penalty would reduce the bandwidth of the memory subsystem.
This latency problem can be solved using an access request buffer 1002, which is stored in the processing element 402. For example, the memory access request buffer 1002 may receive access requests to memory from the processing element to each processor cycle and can store them until they are ready to be sent to the 608/702 memory arbitrage block. The active buffer 1002 synchronizes the frequency with which the requests for £ 2013-00812-0 6 -J1- 2013 memory are sent to the memory arbitration block 608/702 and the frequency at which data is received from the memory subsystem. In some embodiments, the buffer may include a queue. The number of items in buffer 1002 (for example, buffer depth) may be greater than the number of cycles for retrieving data from the memory subsystem. For example, when the access latency to RAM is 6 cycles, the number of items in the buffer 1002 may be 10. The buffer 1002 may reduce the latency penalty resulting from arbitrariness, improving the flow of the memory subsystem, in principle, using memory -Buffer access memory requests, applicants can be allocated up to 100% of the total memory bandwidth.
It will be understood that a possible problem related to the use of several RAM instances is that, allowing simultaneous access of several processing elements to substances within a bank, a memory dispute may result.
The present document presents at least two ways to address memory disputes. First of all, care will be taken when designing the software, as we will describe later, to avoid a memory dispute and / or a memory conflict by carefully arranging the data in the memory subsystem, so that reduces dispute and / or memory conflicts. In addition, software development tools associated with the parallel processing device may allow reporting of memory dispute or memory conflict in the software design phase. Therefore, memory dispute or memory conflict issues can be corrected by improving the data arrangement in response to the memory dispute or the reported memory conflict in the software design stage.
Second, as we will describe below, the ISI block within the architecture is configured to detect port conflicts (disputes) in the hardware and to interrupt the low-priority processing elements. For example, the ISI block is configured to analyze memory access requests from processing elements, serve the sequence of memory access requests, and route memory access requests in order of priority, so that all read or write operations data from all processing elements to be finalized in order of priority.
The order of priority between the processing elements can be set in several ways. In some implementations, the priority order can be statically defined at system design. For example, the priority order can be coded as a reset state for system registers, so that when a system is started, it starts with ^ - 2 0 1 3 - 0 0 8 1 2 -0 6 -11-2013
<img file="RO129804A0_D0009.tif" />
a set of pre-allocated priorities. In other implementations, the order of priority can be determined dynamically using user-programmable registers.
In certain implementations, programmers may plan the arrangement of data for their software applications, so as to reduce the disputes of shared memory sub-blocks within a memory area. In some cases, the planning of the data arrangement can be assisted by an arbitration panel. For example, the arbitration block can detect a memory dispute, allow, on the basis of priority, access of memory by a processing element associated with the activity with the highest degree of priority, may interrupt other processing elements that dispute its memory and may run the dispute process by process, until it is resolved.
FIG. 11 illustrates a planning software application according to certain implementations. In this application, the planning software can coordinate an implementation of a 3x3 blur filter within a processing strip. The planning software can, in the execution phase, determine an order of operations and coordinate the operations of the processing elements. An 1100 data flow graph for the processing band includes element 1 - element 5 1102-1110. The element 1 1102 may include an input buffer 1112, a processing block 1144 and an output buffer 1114. The input buffer 1112 and the output buffer 1114 may be implemented using a bistable. In some embodiments, each of the other elements 1104-1110 may have a structure substantially similar to that of element 1 1102.
In some implementations, the element 2 1104 may include a processing element (for example, a vector processor or a hardware accelerator) that can filter an input with a 3x3 blur filter. Element 2 1104 can be configured to receive input data from shared buffer 1118, which temporarily stores output data of element 1 1102. To apply a 3x3 blur filter to an input, element 2 1104 can receive at least 3 lines of data from the shared buffer 1118 before it can begin operation. Thus, the software scheduler 1120, which can run on a RISC processor 1122, can detect that the shared buffer 1118 contains the correct number of data lines before signaling to element 2 1104 that it can begin the filtering operation.
After the initial signal that there are 3 data lines, the software scheduler 1120 can be configured to signal to element 2 1104 each time a new additional line is added to the 3-line rolling memory 1118. In addition to line synchronization with line, the arbitration and synchronization for each element is performed cycle by cycle
6 (- 2 0 1 3 - 0 0 8 1 2 - 0 6 -11- 2013 pQÎ of the processing band. For example, element 1 1102 may include a hardware accelerator that produces a complete output pixel at each cycle. at this output, the hardware accelerator can maintain buffer 1112 so that processing block 1114 has sufficient data to continue operations. In this way, the processing block 1114 can produce a sufficient output to maintain the flow of the element 1102 as high as possible.
In some implementations, a software tool chain can predict memory conflicts by analyzing the software program that uses the memory subsystem. The software tool chain may include an integrated development environment (IDE) based on a graphical user interface (GUI) (for example, an Eclipse-based IDE), from which the programmer can edit code, call the compiler, assembler, and to source errors when necessary. The software tool chain can be configured to predict memory conflicts through a dynamic analysis of programs running on multiple processors using a system simulator that emulates the entire processing, bus, memory elements and peripherals. Also, the software tool chain can be configured to record, in a log file or on a display device, if different programs running on different processors or hardware resources attempt concurrent access to a particular block of a memory area. The software tool chain can be configured to perform cycle-by-cycle recordings.
In some embodiments, the processing band 1100 may also include one or more hardware counters (for example, one counter for each memory instance) to be incremented each time a memory conflict occurs. These counters can then be read by a debugger (for example, JTAG) and displayed on a screen or recorded in a file. Further analysis of the log files by the system programmer can allow different scheduling of memory access so as to reduce the possibility of conflicts at the memory ports.
A key difficulty for IBM CELL architecture programmers (illustrated in FIG. 2) is scheduling data transfers hundreds of cycles ahead so that data can be controlled by DMA and stored in local storage. (local storage LS) before a vector processor accesses the data. Some implementations of the presented architecture may address this issue by performing hardware access arbitrage and planning and recording conflicts in hardware counters that can be read by ^ -2013-00812-0 6-11-11, 2013.
<img file="RO129804A0_D0010.tif" />
user. In this way, the architecture presented can be used to create a high performance tape for video / image processing.
FIG. 12 illustrates a hierarchical structure of a system that contains a parallel processing device according to certain implementations. System 1200 may include a parallel computing system 1202 with a plurality of processing elements, such as filters, and a software application 1204 running on the parallel computing system 1204, an application programming interface (API) 1206 for the interface between application 1204 and parallel computing system 1202, a compiler 1208 that compiles the software application 1204 for running in the parallel computing system 1202 and a scheduler 1210 for controlling the operations of the processing elements in the parallel computing system 1202.
In some embodiments, the presented parallel processing device may be configured to operate in conjunction with a processing band description tool (for example, a software application) 1204, which allows the description of image processing bands as a graph of data flow. The processing band description tool 1204 is capable of describing the image / video processing bands in a flexible manner, independent of the underlying hardware / software platform. In particular, the data flow graph, used by the tool for describing the processing band, allows the description of activities independent of the processing elements (for example, processor and filter accelerator resources) that can be used to implement the data flow graph. . The output data resulting from the processing band description tool may include a description of the directed acyclic graph (DAG) or the data flow graph. The description of the DAG or data flow graph can be stored in an appropriate format, such as XML.
In some implementations, the description of the DAG or the data flow graph may be accessible to all other tools in the 1200 system and may be used to control the operations of the parallel processing device in accordance with the DAG. FIG. 13 illustrates how the DAG description or data flow graph can be used to control the operations of a parallel processing device according to some implementations.
Prior to the actual operation of the computing device, a compiler 1208 for the parallel processing device 1202 can take (1) a description of the data flow graph 1306 and (2) a description of the available resources 1302 and can generate a list of activities 1304 that indicates how DAG can be performed on several processing elements. For example, when an activity cannot be performed on a single processing element, compiler 1208 can divide the activity into several processing elements;
Α-2 0 1 3 - 0 0 8 1 2 -0 6 -11- 2013
<img file="RO129804A0_D0011.tif" />
when the activity can be performed on a single processing element, compiler 1208 can allocate the activity of a single processing element.
In some cases, when the activity would use only part of the capabilities of a processing element, compiler 1208 may merge and schedule the execution of multiple activities on a single processing element sequentially, to the extent that it can be supported by a processing element. FIG. 14A illustrates the planning and issuing of activities by the compiler and the planner according to some implementations. The advantage of scheduling activities using a compiler and scheduler is that the compiler and scheduler can plan activities automatically based on the operations performed by the activities. This is a great advantage over the previous situation, where a programmer had to manually determine the planning of the code running on a processing element or a group of processing elements that carry out a certain activity, including when planning the data transfers made. from DMA from peripherals to CMX, from CMX to the CMX block and from CMX back to peripherals. This work was difficult and prone to errors, and the use of DFGs allows the automation of this process, saving time and increasing productivity.
During the execution of a computing device, the scheduler 1210 can dynamically plan the activity at the processing elements available based on the activity list 1304 generated by compiler 1208. Planner 1210 can run on RISC host processor 1306 in a multi-core system and can plan activities at the level of processing elements, such as a multitude of vector processors, filter accelerators, and direct memory access (DMA) processing units, using statistics from hardware performance monitors and the 1308 timer. In some implementations, the hardware and timed performance monitors 1308 may include interrupt counters, CMX conflict counters, bus cycle counters (ISI, APB and AXI) and cycle counters, which can be read by scheduler 1210.
In some implementations, planner 1210 may allocate activities to available processing elements based on the statistics received from hardware performance monitors and from timers 1308. Hardware and timed performance monitors 1308 can be used to increase the efficiency of the processing elements or to perform an activity using fewer processing elements, so as to save energy or allow the calculation of other activities in the bulkhead.
Q-2? 13-00 8 12-0 6 -11- 2013 For this purpose, hardware performance monitors 1308 can provide a performance metric. The performance metric can be a number that indicates the level of activity of a processing element. The performance metric can be used to control the number of processing elements instantiated to perform an activity. For example, when the performance metric associated with a particular processing element is greater than a predetermined threshold, the scheduler 1210 can instantiate an additional processing element of the same type as the first processing element, thus distributing the activity to multiple processing elements. Another example: when the performance metric associated with a particular processing element is below a predetermined threshold, scheduler 1210 can remove one of the instantiated processing elements of the same type as the first processing element, thus reducing the number of processing elements that perform a certain activity.
In some implementations, planner 1210 may prioritize the use of processing elements. For example, scheduler 1210 can be configured to determine whether it would be preferable for the task to be allocated to a processor or hardware filter accelerator. In some implementations, the scheduler 1210 may be configured to change the CMX buffer arrangement in the memory subsystem so that the system can meet the execution configuration criteria. Execution configuration criteria may include, for example, image processing rate (frames per second), power consumption, the amount of memory used by the system, the number of processors in operation and / or the number of accelerators in operation.
There are several ways in which an output buffer may be disposed in memory. In some cases, the output buffer may be physically adjacent to the memory. In other cases, the output buffer may be divided into "chunks" or "zones". For example, the output buffer may be divided into N vertical strips, where N represents the number of processors allocated to the image processing application. Each strip is located in another CMX area. This arrangement can favor processors, as each processor can access the input and output buffer locally. However, this arrangement may be unfavorable to filter accelerators, which may cause many conflicts for them. Filter accelerators often process data from left to right. Therefore, all filter accelerators would initiate their processes by accessing the first image strip, which would cause many conflicts from the beginning. In other cases, the output buffer may be interlaced. For example, the output buffer can be divided over all
0 “2013-00812-0 6 -11- 2013 CMX zones, in a predetermined size range. This predetermined size can be 128 bits. The interleaved layout of the output buffer may favor filter accelerators, because the distribution of accesses on CMX zones reduces the risk of conflicts.
In some implementations, a buffer (such as an input or output memory) may be allocated depending on the hardware and / or software nature of its manufacturers and consumers. Consumers are more important because they generally require more bandwidth (filters usually read more lines and produce a single line). Hardware filters are programmed according to the buffer memory arrangement (they allow addressing the adjacent, interlaced and distributed memory).
FIG. 14B illustrates a process for automatically scheduling an activity using the compiler and scheduler according to some implementations. The compiler determines a list of activities to be performed by the parallel processing device based on the DAG. In step 1402, the scheduler is configured to receive the task list and keep the task list in separate queues. For example, when the task list includes (1) activities to be performed by the DMA, (2) activities to be performed by a processor, and (3) activities to be performed by a hardware filter, the scheduler can store activities in three separate queues: for example , a first queue for the DMA, a second queue for the processor and a third queue for the hardware filter. In steps 1404-1408, the compiler is configured to output activities to the associated hardware components as these components become available for new activities. For example, in step 1404, when the DMA becomes available for performing an activity, the execution compiler is configured to unpack the first queue for the DMA and communicate to the DMA the queued activity. Similarly, in step 1406, when the processor becomes available for performing an activity, the execution compiler is configured to undo the second queue for the processor and communicate to the processor the activity removed from the queue. Also, in step 1408, when the hardware filter becomes available for performing an activity, the execution compiler is configured to undo the third queue for the hardware filter and communicate to the hardware filter the activity removed from the queue.
In some implementations, scheduler 1210 may use counter values from hardware performance monitors and chronometers 1308 to adjust the use of processing elements, especially when running multiple processing tapes (for example, an application software 1204) in the matrix of processing elements, because these processing strips were not necessarily designed together. For example, ^ -2 0 1 5 - 0 0 8 1 2 -0 6 -11-2013
<img file="RO129804A0_D0012.tif" />
if the actual bandwidth allocated to each processing band is below the expected value and numerous conflicts occur when accessing CMX memory, planner 1210 can use this information to intersperse the execution of 2 processing bands, changing the order in which activities are taken from the 2 queues of processing bands and thus reducing memory conflicts.
In some implementations, the DAG compiler can run in real time (for example, online). FIG. 15 illustrates the operation of a real-time DAG compiler according to certain implementations. The real-time DAG compiler 1502 can be configured to receive upon entry an XML description of the DAG, a description of the available processing elements and any user-defined constraints, such as number of processors, frame rate, power dissipation envisaged, etc. Then, the real-time DAG compiler 1502 can be configured to schedule DAG components at the processing elements, for example, the DMA processing unit, a processor, a hardware filter, and a memory, to ensure that the DAG as specified can meet the constraints. user-defined when mapped to system resources. In some implementations, the real-time DAG compiler 1502 can determine whether the activities in the DAG can be performed in parallel by a breadth-first algorithm. If the width of the DAG is greater than the number of processing elements available for performing the activity in parallel (for example, the amount of available processing power is smaller than the parallelism of the DAG), the real-time DAG compiler 1502 can "fold" the activities so that they should be performed sequentially on the available processing elements.
FIG. 16 compares a schedule generated by an OpenCL planner with a schedule generated by the proposed online DAG planner according to certain implementations. The program produced by the proposed scheduler 1208/1502 can eliminate the redundant copies and DMA transfers present in a typical OpenCL schedule. These data transfers are present in an OpenCL schedule because the GPU used to perform the processing on a DAG activity is located away from the processor executing the schedule. In a typical application processor used on a mobile device, large blocks of data are transferred back and forth between the processor executing the schedule and the GPU processing. In the proposed model, all the processing elements share the same memory space, thus not being necessary to copy in both directions and thus achieving a considerable savings in time, bandwidth and power dissipation.
^ -2013-00812-ο β -11- 2013 In some implementations, when an activity would use only part of the capabilities of a processing element, scheduler 1210 may be configured to merge and schedule the execution of multiple activities on one single processing element sequentially, to the extent that a processing element can be supported, as in FIG. 14.
In an image processing application, a scheduler can be configured to divide processing activities between processors, dividing an image into strips. For example, the image can be divided into vertical or horizontal jacks of predetermined width.
In some implementations, the scheduler may predetermine the number of processors used for a particular image processing application. This allows the planner to predetermine the number of strips for the image. In some implementations, a filtering operation can be performed by the processors in series. For example, when there are 5 software filters running by the application, the 402 processors can each be configured to run the first software filter simultaneously at a first time, the second software filter simultaneously at a second time. This means that the computational burden is more evenly balanced between the processors allocated to the respective image processing application. This is because the processors are configured to run the same list of filters simultaneously in the same order.
When too many processors are allocated to the image processing application, the processors may spend a long time idle, waiting for the hardware filter accelerators to complete their activities. On the other hand, when too few processors are allocated to the application, hardware filter accelerators can spend a long time at rest. In some implementations, the planner 1210 can be configured to detect these situations and to adapt accordingly. In other implementations, the scheduler 1210 can be configured to override processors for particular image processing applications and allow processors to reduce their power once they have completed their activity before hardware filter accelerators.
In some implementations, the scheduler may use a barrier mechanism to synchronize the processing elements, such as hardware filter accelerators and processors. The scheduler output data may include a command flow. These commands may include (1) start commands for processing elements, such as hardware filter accelerators and processors, and (2) barrier commands. A command-bar indicates that the processing elements must wait and not pass a next set of commands until after all the processing elements in the group have reached the command-barrier, ii-201 3-0 0 0 8 1 2 -0 6 -.11- 2013
<img file="RO129804A0_D0013.tif" />
even if some processing elements have practically completed their activity. In some implementations, the planner may communicate this command-barrier based on the dependencies between the activities performed by the processing elements.
FIG. 17 illustrates a barrier mechanism for synchronizing the processing elements according to certain implementations. The flow of orders includes barrier orders (1702, 1712) and activity orders (1704, 1706, 1708, 1710). Each activity command can be associated with a processing element and, as in the graph below, the activity commands can be completed at different times. Therefore, the planner can include a barrier command 1712, so that the processing elements do not move on to future activities until the barrier command 1712 is cleared. This barrier mechanism can be considered as a temporary organization of parallel activities in a band. processing.
In some implementations, the barrier mechanism is implemented in hardware using interrupt signals 1714. For example, the scheduler can program a bit mask, which specifies which processing elements belong to a group. As the processing elements complete the allocated activities, interrupt signals associated with the respective processing elements are declared. Once all the interrupt signals associated with the processing elements in the group have been declared, the controller of the processing elements can receive a global interrupt signal, which shows that all the processing elements have reached the barrier command.
Interrupting sources may include SHAVE vector processors, RISC processors, hardware filters, or external events. In particular, hardware filters allow many modes, including a non-circular buffer mode, in which the input / output buffer contains a frame, and the filter can be configured to output a single interrupt signal either when processing the entire frame. input, or when writing the entire corresponding output frame. Filters can also be programmed to operate on lines, fixes, or blocks in frames using appropriate settings for image size, address range / baseline memory buffer, etc.
An important challenge in the case of a complex parallel processing device is the way to program the processing elements in the parallel processing device, especially for embedded systems that are very sensitive to power and have few resources (eg resources). computational and memory). Computational imaging, and especially video and image processing, is very demanding for systems
Ct “2 0 1 3 ~ 0 0 0 1 2 -0 6 -11- 2013 incorporated in terms of performance, as the dimensions and frequencies of the staff are very large, increasing more and more from year to year.
The solution to this problem presented here is to provide an application programming interface (API) 1206 that allows high-level application writing by a programmer without requiring in-depth knowledge of the details of the multicore 1202 processor architecture. Using software API 1206, the programmer can quickly create new image or video processing lanes without knowing the implementation details in depth, because the details on the implementation of functions in software, programmable processors, or hardware are abstracted, the programmer not confronting them. For example, an implementation of a blur filter is given as an implementation of reference software running on one or more processors or hardware accelerator filters. The programmer can initially use a software blur filter implementation, and then switch to using a hardware filter without altogether changing the processing band implementation, as not the programmer, but ISI, AMC, and CMX arbitrage block have the role of determining which one. processor and hardware resources gain access to physical memory blocks and in what order.
While the multi-port memory approach described above is suitable for memory sharing under high bandwidth conditions and low latency between identical processors, it is not ideal for sharing bandwidth with other devices. These other devices can be hardware accelerators and other processors with different latency requirements, especially for applications that require very large bandwidth, such as computational video and image processing.
The presented architecture can be used in conjunction with a multi-port memory subsystem, to provide the larger bandwidth required for multiple simultaneous accesses from a multitude of programmable VLIW processors with high latency requirements, a large group of programmable hardware filters for image / video processing, as well as a bus interface to allow control and access of data by a conventional host processor and peripherals. FIG. 18 illustrates the parallel processing device with different types of processing elements according to certain implementations. The parallel processing device includes a plurality of processors 1802 and a plurality of filter accelerators 1804, and the plurality of processors 1802 and the plurality of filter accelerators 1804 may be coupled £ 2013-00812-fl 6-11-11, to memory subsystem 412 by ISI 410, respectively by the accelerator memory controller (AMC) 1806.
The AMC 1806 subsystem and the multicore memory subsystem (CMX) 412 provide chip storage, facilitating processing of low power digital streaming signal at the 1802 processors, as well as the 1804 hardware filter accelerators for certain image processing applications. /video. In some implementations, CMX 412 memory is organized into 16 128 kB zones, organized in 64-bit (2 MB) words in total. Each processor 1802 can have direct access to an area of memory subsystem 412 and indirect access (with higher latency) to all other areas of memory subsystem 412. Processors 1802 can use CMX 412 memory to store instructions or data, and accelerators The 1804 hardware filter uses CMX 412 memory to store data.
To facilitate data sharing between heterogeneous processing elements, allowing latent intolerant processors 1802 to achieve high performance when accessing a shared CMX memory 412 with hardware filter accelerators 1804, hardware filter accelerators 1804 are designed to tolerate latency. This is accomplished by providing each hardware filter accelerator (filter) 1804 with local FIFO memories, which gives more elasticity to the synchronization, as well as with a beam switch to share access to the CMX, ISI being able to support inter-SHAVE communication without disputes with hardware filter accelerators, as in FIG. 10 according to certain implementations. In addition to conflicts at the output ports, the conflict with accessing an ISI-410 port is possible. If more than one external zone attempts to access the same memory area in any cycle, a port conflict may occur. The mapping of the LSU port to the ISI interconnection port is fixed, making it possible for SHAVE 0 1802-0 to access zone 2 through LSU port 0 and for SHAVE 11 (1702-11) to access zone 2 through LSU port 1 without conflicts. The ISI matrix can allow the transfer of 8 x 2 ports x 64 bits of data at each cycle. For example, SHAVE N 1802 can access the N + 1 area through LSU port 0 and 1, and all 8 SHAVE processors can access simultaneously without any interruption.
In some implementations, the memory subsystem 412 can be logically divided into zones (blocks). FIG. 19 illustrates the proposed multi-core memory subsystem according to certain implementations. FIG. 19 illustrates a detailed interconnection by bus between AXI, AHB, SHAVEs, ISI and CMX, as well as filter accelerators, AMC and CMX. The diagram shows two AMC input ports and two AMC output ports, 2 ISI input ports and 2 £ <-2013- 0 0 8 1 2 -0 6 -41- 2013 ISI output ports, L2 cache connections and mutual exclusion (mutex), as well as internal write / read arbitrage and source multiplexing for addressing the 4 memory blocks and FIFO memory, as well as selecting the output destination of the 4 memory block outputs to ISI and AMC.
Each zone can be connected to two of 16 possible ISI input sources, including 12 SHAVEs, DMA, Texture Management Unit (TMU) and AHB bus interface to the board-mounted host processor. . Likewise, each zone has 2 ISI output ports that allow an area to send data to 2 of 16 possible destinations, including 12 SHAVEs, DMA, texture management unit (TMU) and AHB and AXI bus interfaces to the processor - host mounted on the board. In the preferred implementation, the memory area contains 4 physical RAM blocks with input arbitrage block which, in turn, connects to the local SHAVE processor (2 LSUs and 2 64-bit instruction ports), 2 ISI input ports , 2 AMC and FIFO input ports used for inter-SHAVE messaging, as well as a messaging FIFO, L2 cache and mutex exclusion blocks.
On the way out of a CMX area, the input to the destination selection block is connected to the 4 RAM instances, as well as to the L2 cache and the mutex hardware blocks. The exits from the destination selection block, illustrated in FIG. 20 as the 2002 block, connects to the 2 local LSU ports and the instruction ports (SP_1 and SP_0), as well as 2 ISI output ports and 2 AMC output ports. The 2 ISI ports allow connecting a local area to two destinations out of the 12 possible processors, DMA, TMU AXI and AHB host buses. The processors are provided with access to memory through inter-SHAVE (ISI) interconnection, which connects to the 2 64-bit ISI inputs and 2 64-bit ISI output ports contained in an area of the multicore memory subsystem. The deterministic access characterized by high bandwidth and low latency offered by the ISI interconnection reduces the interruptions at the processor level and offers a high computational flow.
FIG. 21 illustrates an AMC beam architecture according to certain implementations. The AMC crossbar 1806 can be configured to connect the 1804 image processing hardware filters to the AMC ports of the CMX 412 multi-core memory areas. The AMC 1806 can include one or more area controllers 2102, preferably one for each CMX 412 block. The controllers of the ports of zone 2102 are, in turn, connected to the slice address request filter - SARF 2104. In turn, SARF 2104 is connected to the AMC clients (in this implementation, the AMC clients are ^ - 2013-00812-0 6 -11- 2013 image processing hardware accelerators). SARF 2104 accepts read / write requests from filter accelerators and offers them request or permission signals, accepting data and addresses from SIPP ports that have received write access and providing read data to those who have received read access. . In addition, SARF offers AXI mastering on the AXI host bus for the host processor in the system, which allows the host to query (read / write) CMX memory via the AMC beam switch.
In some implementations, there is provided a controller of the ports of zone 2102 in AMC 1806, which communicates with the 2 read ports and the 2 write ports in the associated CMX memory area 412, as in FIG. 21. Seen from the filter accelerator portion of the CMX 412 memory subsystem, each hardware filter is connected to a traverse switch port 1806 from the accelerator memory controller (AMC). AMC 1806 has a pair of 64-bit read ports and a pair of 64-bit write ports that connect it to each CMX 412 memory area (there are 16 zones, totaling 2 MB in the preferred implementation). Connecting the image processing hardware accelerators to the AMC 1806 via read or write client interfaces and providing a local buffer within the accelerators allows for latency requirements to be relaxed, leaving more bandwidth available for ISI and processors, with synchronization with high degree of determination that allows to reduce the interruption of the processors.
FIG. 20 illustrates a single area of the CMX infrastructure according to certain implementations. The area contains a referee and source multiplexing, which allows up to 4 out of eight possible 64-bit sources to access the 4 physical SRAM blocks in the CMX area, shared L2 cache, shared mutex hardware blocks for inter-processor negotiation of mutual exclusion process threads, as well as a 64-bit FIFO memory used for small bandwidth inter-SHAVE messaging. The six input sources are: AMCoutl and AMCoutO, which are connected to the corresponding slice_port [l] and slice_port [0] ports of the ACM, as in FIG. 21; to which are added 2 ISI ports (ISIoutl and ISIoutO); 2 LSU ports (LSU_1 and LSU_0); and, finally, 2 instruction ports (SP_1 and SP_0), which, combined, allow 128-bit instruction reading from CMX. Source arbitrator and multiplexing generates the 64-bit read / write address and data for controlling the 4 SRAM blocks in response to the priority access of the 8 input sources, while the 64-bit FIFO inter-SHAVE communication input is connected to the arbitrator and to the source multiplexer, the output can only be read by the local processor in the CMX area. In practice, each processor communicates with the FIFO messaging memory of the other ^ -2013-00812-0 6 -11- 2013 processor through the ISIoutl and ISIoutO ports and through the ISI infrastructure external to the CMX zone, which interconnects the CMX, zones and processors.
In addition to the arbitration between 64 applicants from each of the 2 AMC ports in an area, as shown in FIG. 20, an additional 2: 1 arbitrator is provided to arbitrate between AMC ports 1 and 0. The purpose of this 2: 1 arbitrator is to prevent saturation by one or the other of the 2 AMC ports of the entire bandwidth of the AMC port, in fact which would lead to excessive interruptions at one of the requesting ports. This additional feature ensures a more balanced allocation of resources in the presence of more demanding port bandwidth requesters, thus maintaining a higher throughput throughout the architecture. Similarly, a 2: 1 arbitrator arbitrates between the two processor ports SP1 and SP0, for similar reasons.
The arbitration and multiplexing logic also controls the access of the processors, either directly or through ISI, to a shared L2 cache, through a second level of arbitrage that shares access according to a strict round robin algorithm between 16 sources. possible, a 64-bit port being connected between the level two arbitrator and each of the 16 CMX zones. Likewise, the same logic allows access to the 32 hardware mutexes that are used for negotiating between mutual exclusion processors among the process threads running on the 12 plate-mounted processors and the two 32-bit RISC processors (via AHB and AXI bus from ISI).
The priority in the preferred implementation is that SP_1 and SP_0 have the highest priority degree, followed by LSUl and LSU_0, then ISIoutl and ISIoutO, and AMCoutl and AMCoutO and, finally, FIFO have the highest priority degree. little. The reason for this priority allocation is that SP_1 and SP_0 control the processor's program access to the CMX, and the processor will immediately pause if the following instruction is not available, followed by LSU l and LSU 0, causing the processor to stop again; likewise, ISIoutl and ISIoutO come from other processors and will cause them to stop if data is not immediately available. AMCoutl and AMCoutO ports have the lowest priority, they have integrated FIFO memories and can thus tolerate a high level of latency before interrupting. FIFO from the processor is only required for low bandwidth messaging between processors, thus having the lowest priority of all.
Once the arbitrator has allowed up to 4 sources to access the 4 SRAM, L2 cache, mutex and FIFO blocks, the output data from the six read data sources, including 4 SRAM, L2 cache and mutex blocks, is selected. , these data being directed to up to 4 of 8 <2013-00812-0 6 -π- 2βί3
<img file="RO129804A0_D0014.tif" />
possible 64-bit destination ports; 4 at the processor associated with the memory area (SP_1, SP_0, LSU l and LSU O); 2 ISI associates (ISIoutl and ISIoutO) and, finally, 2 AMC associates (AMCoutl and AMCoutO). No prioritization is required on the output multiplexer, only 4 64-bit sources need to be distributed to 8 ports of destination.
FIG. 22 illustrates an AMC controller with beam ports according to certain implementations. AMC's port controller 2202 includes a round robin referee 2204, which connects port controller 2202 to the 1804 accelerators whose requests have been filtered through the processor. Thereafter, the arbitrator may transmit valid requests from AMC clients to the FIFO memory of the port controller. In case of read requests, a request response (read client ID and line index) is sent to FIFO TX read ID. Data retrieved from the zone port and valid signals are entered into the port controller read logic, which displays the read client ID and line indices in the FIFO Rd TX ID and transmits the read data from the appropriate area port and the valid signals to FIFO Rd data, from which they can be read by the requesting AMC client On the CMX side of FIFO, the port interrupt logic presents the requests from FIFO and ensures the control of the zone ports for the 2 AMC input ports in the associated CMX memory area.
The number of read and write client interfaces to CMX are configurable separately. Any customer can address any (or any) area of the CMX. With 16 memory zones in CMX, 2 ports in each zone, and a processor frequency of 600 MHz in the system, the maximum total width of data memory that can be provided to customers is 143 GB / s: Max. = 600 MHz * (64/8) * 2 * 16 = l, 536el 1 B / s = 143 GB / s.
At the higher 800 MHz frequency of the processor, the frequency band increases to 191 GB / sec. The AMC controller arbitrates the simultaneous accesses of the read / write interfaces from the hardware accelerator blocks connected to it. At most two read / write addresses in each memory area can be allocated to each processor cycle, resulting in a maximum area memory bandwidth of 8.9 GB / s at a 600 MHz frequency of the system processor . Customer access is not restricted to the CMX address space. Any access that comes out of the CMX address space is transmitted to the master program of the AXI bus from AMC.
FIG. 23 illustrates a read operation using a 1806 AMC according to certain implementations. In this illustration, 4 data words are read from address range A0-3. The AMC client (for example, a 1804 filter accelerator) first declares a request at ^ * 2013-00812'0 6 -11- 2013
<img file="RO129804A0_D0015.tif" />
port controller input. Port controller 2202 responds by issuing a permission signal (gnt) which, in turn, causes the client to issue the AO, Al, A2 and, finally, A3 addresses. The corresponding rindex values appear on the ascending arc of the processor cycle (clk) that corresponds to each permission. It can be seen that the synchronization can be very elastic on the client side compared to the index data and addresses, which are transmitted from the CMX area to the port controller. The deterministic synchronization on the CMX side of the port controller allows efficient access to the shared CMX between AMC clients and processors, which have high latency sensitivity, and FIFO memories and local storage from AMC clients allow high synchronization variability on the AMC client side. (for example, the filter accelerator) from the CMX 412 memory subsystem.
FIG. 24 illustrates a write operation using a 1806 AMC according to certain implementations. The synchronization diagram illustrates the transfer of 4 data words to CMX via AMC. The AMC client issues a request, and at the next ascending arc of the processor cycle (clk), the permission signal (gnt) rises, transferring the data word DO associated with the A0 address through the AMC. The gnt signal then goes down for one processor cycle, and on the next ascending arc, the gnt signal rises for 2 processor cycles, providing access for Dl and D2 to the Al or A2 addresses, before the gnt signal goes down again. . At the next ascending clk arc, the gnt signal rises again, allowing the data word D3 to be transferred to address A3, after which the req and gnt signals descend to the next clk arc, pending the next read / write request.
The software framework for streaming image processing tape (SIPP) used in conjunction with the model in FIG. 12 provides a flexible approach to implementing image processing tapes using CMX 412 memory for scan line buffers, frame blocks (subsections of frames) or even entire frames at high resolution using an external DRAM tablet in one module. , connected to a substrate to which the image / video processing matrix is attached. The SIPP framework deals with complexities such as image border management (pixel replication) and cyclic line memory buffer management, making the implementation of ISP (image signal processing) functions in software (from the processor level) easier and easier. generic.
FIG. 25 illustrates the parallel processing device 400 according to certain implementations. The parallel processing device 400 may include a memory subsystem (CMX) 412, a plurality of filter accelerators 1804, and a bus structure 1806 for arbitrating access to the memory subsystem 412. The memory subsystem (CMX) cț-2 0 1 3 - 0 0 8 1 2 -0 6 -11- 2013
412 it is constructed in such a way as to allow a plurality of processing elements 402 to access, in parallel, the data memory and program code without causing interruptions. These processing elements 402 may include, for example, SHAVE (vector processing unit with hybrid streaming architecture) VLIW (very long instruction word) processors as required, parallel access to uninterrupted data and program code or a filter accelerator. In addition, the memory subsystem (CMX) 412 may provide for a host processor (not included in the illustration) to access the CMX 412 memory subsystem through a parallel bus such as AXI (not included in the illustration). In certain implementations, each processing element 402 can read / write up to 128 bits per cycle through its LSU ports and can read program code up to 128 bits per cycle through its instruction port. In addition to the ISI and AMC interfaces for processors and filter accelerators respectively, CMX 412 offers simultaneous access to read / write memory via the AHB and AXI bus interfaces. AHB and AXI are standard ARM parallel interface buses that allow a processor, memory and peripherals to be connected using a 1806 shared bus infrastructure. The CMX 412 memory subsystem can be configured to handle a maximum of 18 128-bit memory accesses per cycle.
Accelerators 1804 include a collection of image processing hardware filters that can be used in SIPP 1200 software. Accelerators 1804 can exempt processing elements 1802 from some of the most computationally demanding functionality. The diagram shows how a multitude of filter accelerators 1804 can be connected to AMC 1804 which performs address filtering, arbitrage and multiplexing. Multiple serial MIPI 2502 camera interfaces can be connected to AMC 1804, with the preferred implementation being a total of 12 MIPI serial colors connected in 6 groups of 2 colors. AMC 1804 is also connected to the AXI and APB interfaces, to allow the 2 RISC processors in the reference implementation to access CMX memory through AMC. The final element of the diagram is CMX 412, whose access is arbitrated by AMC 1804, allowing multiple hardware filter accelerators 1804 simultaneous access to physical RAM instances in CMX 412 memory. Also shown is a reference filter accelerator 1804, in this case a 5x5 2D filter, which contains a fpl6 arithmetic processing band (16-bit floating point format such as IEEE754), an associated controller for interrupts in the bandwidth. processing, a line buffer read client for storing an input line in the processing band ¢ 16, a control input of the beginning of the line and the client writing the line buffer for storing the output data of the q-2 band 0 1 3 - 0 O 8 1 2 - O 6-11-2013 fp processing 16. For for accelerators to fit within the SIPP, they require access to CMX memory at high bandwidth, this access being provided by the Accelerator Memory Controller (AMC).
In some implementations, the CMX 412 memory subsystem may be divided into 128kB blocks or areas associated with their neighboring processing element 402 for high speed access, low power consumption. Within an area, memory is organized in the form of smaller blocks, for example independent SRAM blocks 3x32 kB, 1x16 kB and 2x8 KB. The physical size of RAM can be chosen as a trade-off between the use of space and the flexibility of the configuration. Any processing element 402 can access physical RAM anywhere in the memory subsystem (CMX) 412 with the same latency (3 cycles), but access outside the local area of a processor is limited in bandwidth and will have more power consumption. higher than accessing a local memory area. Generally, to reduce power consumption and increase performance, a processing element 402 can store data locally in a dedicated memory area.
In some implementations, each physical RAM may have a 64-bit width. If more than 402 processing elements attempt to access the same physical RAM, a conflict may occur which results in a CPU shutdown. CMX will automatically arbitrate port conflicts, ensuring that no data is lost. For each port conflict, a processing element 402 is interrupted for one cycle, which reduces flow. By carefully arranging the data (by the programmer) within CMX 412, port conflicts can be avoided and processor cycles can be better utilized.
In some implementations, a plurality of processors are equipped with CMX accelerators and memory.
It can be seen from the image processing hardware architecture of FIG. 25 that each 1804 filter accelerator may include at least one AMC read and / or write client interface for accessing CMX 412 memory. The number of AMC 1806 read / write client interfaces may be configured as required. AMC 1806 can include a 64-bit pair of ports in each CMX 412 memory area. AMC 1806 routes requests from its clients to the appropriate CMX 412 area (by partially decoding the address). Simultaneous requests from multiple clients for the same memory area can be arbitrated by a round robin algorithm. The retrieved read data from CMX 412 is routed back to the requesting AMC read clients.
<^ 2 0 1 3 - 0 0 8 12-0 6 7 »- 2013 [0177] AMC Clients (Accelerators) 1804 presents a complete 32-bit address at AMC 1806. Clients access that is not mapped to memory space The CMX are transmitted to the master program of the AXI bus from AMC. Simultaneous accesses (outside of the CMX memory space) by different clients are arbitrated by a round robin algorithm.
AMC 1806 is not limited to providing access to CMX 412 for 1804 filter accelerators; any hardware accelerator or third party may use the AMC 1806 to access the CMX memory and the larger memory space of the platform if its memory interfaces are adequately adapted to the AMC's read / write client interfaces.
The image processing hardware (SIPP) band may include 1804 filter accelerators, a 1806 arbitrage block, MIPI 2502 control, APB and AXI interfaces and CMX 412 multiport memory links, as well as a 5x5 exemplary hardware filter. This configuration allows a plurality of 1802 processors and hardware filter accelerators 1804 for image processing applications to share a 412 memory subsystem consisting of a plurality of physical single-port RAM (random access memory) blocks.
The use of single-port memories increases the energy and spatial efficiency of the memory subsystem, but limits bandwidth. The proposed configuration allows these RAM blocks to behave as a multi-port virtual memory subsystem, capable of serving multiple simultaneous read and write requests received from multiple sources (processors and hardware blocks), using multiple physical RAM instances and providing arbitrary access to them. to serve multiple sources.
The use of an application programming interface (API) and application-level data partitioning are important to ensure reduced competition among processors or between processors and filter accelerators for accessing physical RAM blocks, thus increasing the data bandwidth. for processors and hardware in the case of a given configuration of the memory subsystem.
In some embodiments, the parallel processing device 400 may be located in an electronic device. FIG. 26 illustrates an electronic device that includes a parallel processing device according to certain implementations, the electronic device 2600 may include a processor 2602, memory 2604, one or more interfaces 2606, and a parallel processing device 400.
V2 0 1 3 - 0 0 8 1 2 -0 6 -11- 2013 The electronic device 2600 may have 2604 memory, such as a computer-readable media, flash memory, a magnetic disk drive, a optical drive, programmable read-only memory (PROM) and / or read-only memory (ROM). The electronic device 2600 can be configured with one or more processors 2602 that process instructions and run software that can be stored in memory 2604. The processor 2602 may also communicate with memory 2604 and interfaces 2606 for communication with other devices. Processor 2602 can be any applicable processor, such as a chip system that combines a central processing unit, an application processor and flash memory, or a small instruction set processor (RISC).
In some implementations, compiler 1208 and scheduler 1210 may be implemented in software stored in memory 2604 and running on processor 2602. Memory 2604 may be a non-readable media readable by a computer, flash memory, a magnetic disk drive. , an optical drive, a read-only programmable memory (PROM), a read-only memory (ROM) or any other memory or combination of memories. The software can run on a processor capable of executing computer instructions or computer code. Also, the processor could be implemented in hardware using an application specific integrated circuit (ASIC), a programmable logic array (PLA), a programmable logic gate array (FPGA) ) or any other integrated circuit.
In some embodiments, compiler 1208 may be implemented in a separate computing device that communicates with electronic device 2600 via interface 2606. For example, compiler 1208 may operate on a server communicating with electronic device 2600.
2606 interfaces can be implemented in hardware or software. The 2606 interfaces can be used to receive control data and information from the network as well as from local sources such as a TV remote control. The electronic device can provide a variety of user interfaces, such as a keyboard, touch screen, trackball, touchpad and / or mouse. Also, the electronic device may include speakers and a display device in certain implementations.
In some embodiments, a processing element in the parallel processing device 400 may include an integrated chip capable of executing computer instructions or computer code. Also, the processor could be implemented in hardware using an application specific integrated circuit (ASIC), a ^ * 2013-00812-0 6-11-11 2013 programmable logic array (PLA), a field programmable gate array (FPGA) or any other integrated circuit.
In some implementations, the parallel processing device 400 may be implemented as a system on chip (SOC). In other implementations, one or more blocks of the parallel processing device may be implemented as a separate chip, and the parallel processing device may be made as a system in a single module (system in package - SIP). In some implementations, the parallel processing device 400 may be used for data processing applications. Data processing applications may include image processing applications and / or video processing applications. Image processing applications may include an image processing process, including an image filtering operation; Video processing applications may include a video decoding operation, a video encoding operation, a video analysis operation to detect motion or objects in video data. Further applications of the present invention include machine learning and classification based on a sequence of images, objects or video data and augmented reality applications, including those where a game application extracts geometric data from multiple camera views, including depth cameras. , and extract data from multiple views from which wireframe geometry can be extracted (for example, through a point cloud) for further shading to the vertex by a GPU.
The 2600 electronic device may include a mobile device, such as a cell phone. The mobile device can communicate with a multitude of radio access networks using a multitude of access technologies, as well as with wired communications networks. The mobile device can be a smartphone that offers advanced capabilities, such as text editing, web browsing, games, electronic book reader capabilities, and a full keyboard. The mobile device can run an operating system such as Symbian OS, iPhone OS, Blackberry from RIM, Windows Mobile, Linux, Palm WebOS and Android. The screen can be a touch screen that can be used to enter data into the mobile phone, and the screen can be used instead of the full keyboard. The mobile device may be capable of running applications or communicating with applications made available by servers in the communications network. The mobile device can receive updates and other information from these applications on the network.
Also, the electronic device 2600 may include numerous other devices, such as TVs, video projectors, TV receivers or TV receivers, digital video recorders (DVRs), computers, netbooks, laptops, tablets and any other <2 0 1 3. - 0 0 8 1 2 - 0 6 -11- 2013 audio-video equipment that can communicate with a network. Also, the electronic device may keep in its stack or memory coordinates for positioning around the globe, profile information or other location information.
It will be appreciated that although several different configurations have been described here, the characteristics of each can be advantageously combined in a variety of forms, to obtain an advantage.
In the above specification, the application has been described by referring to specific examples. It is evident, however, that it can be modified in different ways without departing from the general spirit and scope of the invention defined in the appended claims. For example, connections can be any type of connection suitable for transferring signals from or to the respective nodes, units or devices, for example, through intermediates. Consequently, in the absence of any other details or suggestions, the connections may be, for example, direct or indirect connections.
It should be understood that the architectures depicted here have only the role of example and that, in fact, many other architectures can be implemented that perform the same functionality, in the abstract, but of course, any component configuration meant to achieve the same functionality is effectively "associated" so that the desired functionality is achieved. Therefore, any two components here combined to achieve a certain functionality can be viewed as being "associated" with one another so that the desired functionality is achieved, regardless of architectures or intermediate components. Likewise, any two components thus associated can be regarded as "operably connected" or "operably coupled" to one another in order to achieve the desired functionality.
In addition, those skilled in the art will recognize that the boundaries between the functionality of the operations described above are merely illustrative. The functionality of several operations can be combined into one operation and / or the functionality of a single operation can be distributed into additional operations. Moreover, alternative implementations may include multiple instances of a particular operation, and the order of operations may be changed in different other implementations.
However, other modifications, variations and alternatives are possible. Specifications and drawings should be viewed in the same way as being illustrative, not restrictive.
In the claims, any reference signs placed in parentheses will not be construed as limiting the claim. The words "comprising" do not exclude the presence of elements or stages other than those mentioned in a claim. Moreover, the terms "one" as-2 O 1 3 - OO 8 1 2 - O 6 -11- 2013 and "o" are used herein in the sense of one or more than one. Also, the use of introductory expressions such as "at least one" and "one or more" in claims should not be construed as suggesting that the introduction of another claim element by undeclared articles "one" or "one" limits any particular claim comprising said 5 claim element introduced to inventions containing a single such element, even when the same claim includes the introductory expressions "one / more or more" or "at least one", and undocumented items such as "one" or "one". The same is true for the use of determined articles. In the absence of other details, terms such as "first" and "second" are used to arbitrarily distinguish between the elements described by those terms. Thus, these terms did not necessarily indicate the temporal or other prioritization of the respective elements. The mere fact that certain measures are repeated in claims that differ between them does not indicate that a combination of these methods cannot be used advantageously.
Contents2
27 sheets
Sheet 1 Sheet 2 Sheet 3 Sheet 4 Sheet 5 Sheet 6 Sheet 7 Sheet 8 Sheet 9 Sheet 10 Sheet 11 Sheet 12 Sheet 13 Sheet 14 Sheet 15 Sheet 16 Sheet 17 Sheet 18 Sheet 19 Sheet 20 Sheet 21 Sheet 22 Sheet 23 Sheet 24 Sheet 25 Sheet 26 Sheet 27
94 members in 10 offices
Priority claims2
| Document | Office | Kind | Date |
|---|---|---|---|
| 201300812 | Romania | A | |
| RO20130000812 | – | – | – |
Members94
| Document | Office | Kind | |
|---|---|---|---|
| GB201314263D0 | United Kingdom | D0 | |
| RO129804A0This record | Romania | A0 | |
| US2015046673A1 | United States of America | A1 | |
| US2015046674A1 | United States of America | A1 | |
| US2015046675A1 | United States of America | A1 | |
| US2015046677A1 | United States of America | A1 | |
| US2015046678A1 | United States of America | A1 | |
| WO2015019197A2 | World Intellectual Property Organization (WIPO) | A2 | |
| US9146747B2 | United States of America | B2 | |
| WO2015019197A3 | World Intellectual Property Organization (WIPO) | A3 | |
| WO2016016726A2 | World Intellectual Property Organization (WIPO) | A2 | |
| WO2016016730A1 | World Intellectual Property Organization (WIPO) | A1 | |
| KR20160056881A | Republic of Korea | A | |
| EP3031047A2 | European Patent Office (EPO) | A2 | |
| WO2016016726A3 | World Intellectual Property Organization (WIPO) | A3 | |
| CN105765623A | China | A | |
| JP2016536692A | Japan | A | |
| CN106796504A | China | A | |
| KR20170061661A | Republic of Korea | A | |
| EP3175320A1 | European Patent Office (EPO) | A1 | |
| EP3175355A2 | European Patent Office (EPO) | A2 | |
| KR20170067716A | Republic of Korea | A | |
| US9727113B2 | United States of America | B2 | |
| CN107077186A | China | A | |
| JP2017525046A | Japan | A | |
| JP2017525047A | Japan | A | |
| US2017293346A1 | United States of America | A1 | |
| BR112017001975A2 | Brazil | A2 | |
| BR112017001981A2 | Brazil | A2 | |
| US9910675B2 | United States of America | B2 | |
| US9934043B2 | United States of America | B2 | |
| US10001993B2 | United States of America | B2 | |
| EP3175355B1 | European Patent Office (EPO) | B1 | |
| US2018246725A1 | United States of America | A1 | |
| EP3410294A1 | European Patent Office (EPO) | A1 | |
| US2018349147A1 | United States of America | A1 | |
| EP3175320B1 | European Patent Office (EPO) | B1 | |
| JP6491314B2 | Japan | B2 | |
| EP3506053A1 | European Patent Office (EPO) | A1 | |
| JP2019109926A | Japan | A | |
| US10360040B2 | United States of America | B2 | |
| CN106796504B | China | B | |
| JP6571078B2 | Japan | B2 | |
| CN110515658A | China | A | |
| US2019370005A1 | United States of America | A1 | |
| JP2019220201A | Japan | A | |
| US10521238B2 | United States of America | B2 | |
| KR102063856B1 | Republic of Korea | B1 | |
| KR20200003293A | Republic of Korea | A | |
| US10572252B2 | United States of America | B2 | |
| CN107077186B | China | B | |
| CN105765623B | China | B | |
| JP6695320B2 | Japan | B2 | |
| CN111240460A | China | A | |
| US2020192666A1 | United States of America | A1 | |
| US2020241881A1 | United States of America | A1 | |
| JP2020129386A | Japan | A | |
| CN112037115A | China | A | |
| KR102223840B1 | Republic of Korea | B1 | |
| KR20210027517A | Republic of Korea | A | |
| JP2021061036A | Japan | A | |
| KR102259406B1 | Republic of Korea | B1 | |
| KR20210065208A | Republic of Korea | A | |
| US11042382B2 | United States of America | B2 | |
| US11188343B2 | United States of America | B2 | |
| EP3506053B1 | European Patent Office (EPO) | B1 | |
| KR102340003B1 | Republic of Korea | B1 | |
| KR20210156845A | Republic of Korea | A | |
| JP7025617B2 | Japan | B2 | |
| FI3506053T3 | Finland | T3 | |
| JP2022058622A | Japan | A | |
| JP7053713B2 | Japan | B2 | |
| EP3982234A2 | European Patent Office (EPO) | A2 | |
| EP3982234A3 | European Patent Office (EPO) | A3 | |
| US2022147363A1 | United States of America | A1 | |
| US2022179657A1 | United States of America | A1 | |
| KR102413501B1 | Republic of Korea | B1 | |
| JP2022097484A | Japan | A | |
| KR20220092646A | Republic of Korea | A | |
| KR102459716B1 | Republic of Korea | B1 | |
| KR20220148328A | Republic of Korea | A | |
| JP2022169703A | Japan | A | |
| EP4116819A1 | European Patent Office (EPO) | A1 | |
| US11567780B2 | United States of America | B2 | |
| US11579872B2 | United States of America | B2 | |
| BR112017001975B1 | Brazil | B1 | |
| KR102515720B1 | Republic of Korea | B1 | |
| US2023132254A1 | United States of America | A1 | |
| BR112017001981B1 | Brazil | B1 | |
| KR102553932B1 | Republic of Korea | B1 | |
| KR20230107412A | Republic of Korea | A | |
| US11768689B2 | United States of America | B2 | |
| US2023359464A1 | United States of America | A1 | |
| JP7384534B2 | Japan | B2 |
Numbers
- Publication
- 129804
- Publication, DOCDB
- 129804
- Publication, EPODOC
- RO129804
- Application
- 812
- Application, DOCDB
- 201300812
- Application, EPODOC
- RO20130000812
Titles2
- English
- APPARATUS, SYSTEM AND METHOD FOR MANUFACTURING AN EXTENSIBLE CONFIGURABLE TAPE FOR PROCESSING IMAGES
- Romanian
- APARAT, SISTEM ŞI METODĂ PENTRU A REALIZA O BANDĂ CONFIGURABILĂ ŞI EXTENSIBILĂ DE PROCESARE DE IMAGINI
Classification
- IPC, 2
- G06F9 30
- G06F9 28