Computing machine having improved computing architecture and related system and method
Abstract
Vector transfer machine between peer entities (40), comprising: a central processor (42) that can be operated to execute a program, and, in response to the program, can be operated to generate first major data; a channeling accelerator (44) comprising pipelines of permanent connections and which is coupled to the central processor (42) and which can be operated to receive the first main data and to generate first channeling data from the first main data , the channeling accelerator presenting configurable internal interconnections and a microprogram memory (52) that can be operated to store a configuration microprogram to configure the internal interconnections, characterized in that the peer vector machine (40) further comprises a configuration register (70) of the accelerator coupled to the central processor (42) and that can be operated to store configuration information of the pipeline accelerator that is independent with respect to the Program, and the central processor (42) can be operated to receive the configuration information from the register (70) and to configure the pipeline accelerator (44) with a view to generating the first pipeline data by supplying the information of configuration to the pipeline accelerator before executing the program by downloading the configuration information in the microprogram memory, and because a single bus (50) is provided for communication of the central processor with the pipeline accelerator and the core processor and the pipeline accelerator communicate through messages comprising the first main data and a header containing the pipeline desired permanent connections of data destination.

Term
Term ended
Projected expiry passed 31 October 2023, 2.9 years ago.
- Priority
- Filed
- Published
- Projected expiry
- Today
9 claims: 6 independent, 3 dependent
- 1ES 2 300 633 T3 REIVINDICACIONES 1. Máquina de transferencia de vectores entre entidades pares (40), que comprende:un procesador central (42) que se puede hacer funcionar para ejecutar un programa, y, en respuesta al programa, se puede hacer funcionar para generar unos primeros datos principales;un acelerador de canalización (44) que comprende canalizaciones de conexiones permanentes y que está acoplado al procesador central (42) y que se puede hacer funcionar para recibir los primeros datos principales y para generar unos primeros datos de canalización a partir de los primeros datos principales, presentando el acelerador de canalización unas interconexiones internas configurables y una memoria de microprograma (52) que se puede hacer funcionar para almacenar un microprograma de configuración para configurar las interconexiones internas, caracterizada porque la máquina de vectores entre pares (40) comprende además un registro de configuración (70) del acelerador acoplado al procesador central (42) y que se puede hacer funcionar para almacenar información de configuración del acelerador de canalización que es independiente con respecto al programa, y el procesador central (42) se puede hacer funcionar para recibir la información de configuración a partir del registro (70) y para configurar el acelerador de canalización (44) con vistas a generar los primeros datos de canalización mediante el suministro de la información de configuración al acelerador de canalización antes de ejecutar el programa descargando la información de configuración en la memoria del microprograma, y porque se proporciona un único bus (50) para la comunicación del procesador central con el acelerador de canalización y el procesador central y el acelerador de canalización se comunican a través de mensajes que comprenden los primeros datos principales y un encabezamiento que contiene la canalización de conexiones permanentes deseada de destino de los datos.
- 2Máquina de vectores entre pares (40) según la reivindicación 1, en la que el procesador central (40) se puede hacer funcionar además para:recibir unos segundos datos;y generar los primeros datos principales a partir de los segundos datos.
- 3Máquina de vectores entre pares (40) según cualquiera de las reivindicaciones anteriores, en la que el procesador central (42) se puede hacer funcionar además para:recibir los primeros datos de canalización del acelerador de canalización;y procesar los primeros datos de canalización.
- 4Máquina de vectores entre pares (40) según cualquiera de las reivindicaciones anteriores, que comprende además:una memoria de interfaz (48) acoplada al procesador central (42) y al acelerador de canalización (44) y que presenta una primera sección de memoria;en la que el procesador central (42) se puede hacer funcionar para, almacenar los primeros datos principales en la primera sección de memoria, y proporcionar los primeros datos principales de la primera sección de memoria al acelerador de canalización (44).
- 5Máquina de vectores entre pares (40) según cualquiera de las reivindicaciones anteriores, que comprende además:una memoria de interfaz (48) acoplada al procesador central (42) y al acelerador de canalización (44) y que presenta una primera y una segunda secciones de memoria;en la que el procesador central (42) se puede hacer funcionar para almacenar los primeros datos principales en la primera sección de memoria, proporcionar los primeros datos principales desde la primera sección de memoria al acelerador de canalización (44), ES 2 300 633 T3 recibir los primeros datos de canalización del acelerador de canalización (44), almacenar los primeros datos de canalización en la segunda sección de memoria, recuperar los primeros datos de canalización de la segunda sección de memoria hacia el procesador central (42), y procesar los primeros datos de canalización.
- 6Máquina de vectores entre pares (40) según cualquiera de las reivindicaciones anteriores, en la que:el acelerador de canalización (44) comprende un circuito integrado de lógica programable;y la información de configuración comprende un microprograma que se puede hacer funcionar para configurar el circuito integrado de lógica programable.
- 7Método, que comprende:generar unos primeros datos principales ejecutando un programa con un procesador central (42);y generar unos primeros datos de canalización a partir de los primeros datos principales con un acelerador de canalización (44) que comprende canalizaciones de conexiones permanentes y que está acoplado al procesador central caracterizado porque se recibe desde un registro (70) con el procesador central (42) información de configuración del acelerador de canalización que es independiente con respecto al programa, y se configura el acelerador de canalización (44) para generar los primeros datos de canalización mediante el suministro de la información de configuración al acelerador de canalización con el procesador central (42) antes de ejecutar el programa, en el que dicho suministro de la información de configuración comprende la descarga de la información de configuración en una memoria de microprograma que se puede hacer funcionar para almacenar un microprograma de configuración con vistas a configurar interconexiones internas del acelerador de canalización, y en el que dicha generación de los primeros datos de canalización comprende el envío, hacia el acelerador de canalización, de un mensaje que comprende los primeros datos principales y un encabezamiento que contiene la canalización de conexiones permanentes de destino deseada.
- 8Método según la reivindicación 7, en el que la generación de los primeros datos principales comprende la generación de los primeros datos principales a partir de los segundos datos de canalización generados por el acelerador de canalización (44).
- 9Método según las reivindicaciones 7 u 8, que comprende además la generación de los segundos datos principales a partir de los primeros datos de canalización ejecutando el programa con el procesador central (42).
Independent claims9
121 paragraphs in 8 sections, as filed
IS 2 300 633 T3
DESCRIPTION
Piped coprocessor.
Priority claim
The present application claims priority over US Provisional Application Serial No. 60 / 422,503, filed October 31, 2002, which is incorporated herein by reference.
Cross references with related requests
The present application is related to US Publication No. 2004/0181621 entitled COMPUTING MACHINE HAVING IMPROVED COMPUTING ARCHITECTURE AND RELATED SYSTEM AND METHOD; No. 2004/013 6241 entitled PIPELINE ACCELERATOR FOR IMPROVED COMPUTING ARCHITECTURE AND RELATED SYSTEM AND METHOD; No. 2004/0170070 entitled PROGRAMMABLE CIRCUIT AND RELATED COMPUTING MACHINE AND METHOD; and no. 2004/0130927 entitled PIPELINE ACCELERATOR HAVING MULTIPLE PIPELINE UNITS AND RELATED COMPUTING MACHINE AND METHOD; submitted all of them on October 9, 2003, and which have a common owner.
Background
A typical computing architecture for processing relatively large amounts of data in a relatively short period of time comprises multiple interconnected processors that share the processing load. By sharing the processing load, these multiple processors can typically process data faster than a single processor for a given clock frequency. For example, each of the processors can process a respective part of the data or execute a respective part of a processing algorithm.
FIG. 1 is a schematic block diagram of a conventional computing machine 10 featuring a multiprocessor architecture. Machine 10 comprises a master processor 12 and coprocessors 14i to 14<sub>n</sub>, which communicate with each other and with the master processor through a bus 16, an input port 18 to receive raw data from a remote device (not shown in Fig. 1), and an output port 20 to provide data processed to remote source. The machine 10 further comprises a memory 22 for the master processor 12, respective memories 24<sub>1</sub> to 24<sub>n</sub> for coprocessors 14<sub>1</sub> to 14<sub>n</sub>, and a memory 26 that is shared by the master processor and the coprocessors through bus 16. Memory 22 acts as both program and work memory for the master processor 12, and each of the memories 241 to 24n acts as both program and work memory for a respective coprocessor 14<sub>1</sub> to 14<sub>n</sub>. Shared memory 26 allows master processor 12 and coprocessors 14 to transfer data to each other, and to / from the remote device through ports 18 and 20 respectively. Master processor 12 and coprocessors 14 further receive a A common clock signal that controls the speed at which the machine 10 processes the raw data.
In general, the computing machine 10 efficiently divides the raw data processing between the master processor 12 and the coprocessors 14. The remote source (not shown in Fig. 1), such as a set of sonar elements (Fig. 5 ), loads the raw data through port 18 into a shared memory section 26, which acts as a first-in-first-out (FIFO) buffer (not shown) for the raw data. The master processor 12 retrieves the raw data from memory 26 via bus 16, and then the master processor and coprocessors 14 process the raw data, transferring data between them as needed via bus 16. Master processor 12 loads the processed data in another FIFO buffer (not shown) defined in shared memory 26, and the remote source retrieves the processed data from this FIFO through port 20.
In an example operation, the computing machine 10 processes the raw data by sequentially performing n + 1 respective operations on the raw data, in which these operations together constitute a processing algorithm such as a Fast Fourier Transform (FFT). More specifically, machine 10 forms a data processing pipeline from master processor 12 and coprocessors 14. For a given frequency of the clock signal, such channeling typically allows machine 10 to process raw data faster than a machine having only a single processor.
After retrieving the raw data from the raw data FIFO (not shown) in memory 26, the master processor 12 performs a first operation, such as a trigonometric function, on the raw data. This operation produces a first result, which is stored by processor 12 in a first result FIFO (not shown) defined within memory 26. Typically, processor 12 executes a program stored in memory 22, and performs the actions described above under the control of the program. Processor 12 may also use memory 22 as work memory to temporarily store data that is generated by the processor in intermediate intervals of the first operation.
Next, after retrieving the first result from the first result FIFO (not shown) in memory 26, the coprocessor 141 performs a second operation, such as a logarithmic function, on the first result. This second operation produces a second result, which is stored by coprocessor 14<sub>1</sub> in
ES 2 300 633 T3 a FIFO of second results (not shown) defined within memory 26. Typically, coprocessor 14i executes a program stored in memory 24<sub>1</sub>, and performs the actions described above under the control of the program. The coprocessor 14<sub>1</sub> can also use memory 24<sub>1</sub> as work memory to temporarily store data that is generated by the coprocessor in intermediate intervals of the second operation.
Then coprocessors 242 to 24<sub>n</sub> sequentially perform third operations n<sup>ths</sup> on the second results (n-1)<sup>we are</sup> in a similar way to that described above for coprocessor 24<sub>1</sub>.
Operation n ™<sup>1</sup><sup>1</sup>, which is performed by coprocessor 24<sub>n</sub>, produces the final result, that is, the processed data. Coprocessor 24<sub>n</sub> loads the processed data into a processed data FIFO (not shown) defined within memory 26, and the remote device (not shown in Fig. 1) retrieves the processed data from this FIFO.
Since the master processor 12 and the coprocessors 14 are simultaneously performing different operations of the processing algorithm, typically the computing machine 10 can process the raw data faster than a computing machine that has a single processor that performs the different operations sequentially. Specifically, the single processor cannot retrieve a new set of raw data until it performs all of the n + 1 operations on the previous set of raw data. However, using the pipeline technique described above, the master processor 12 can retrieve a new set of raw data after only performing the first operation. Consequently, for a given clock frequency, this pipeline technique can increase the speed at which the machine 10 processes raw data by a factor of approximately n + 1, compared to a single processor machine (not shown in Fig. 1).
Alternatively, computing machine 10 may process the raw data in parallel by simultaneously executing n + 1 instances of a processing algorithm, such as an FFT, on the raw data. That is, if the algorithm comprises n + 1 sequential operations as described above in the previous example, in that case each of the master processor 12 and the coprocessors 14 sequentially perform all of the n + 1 operations on sets respective of the raw data. Consequently, for a given clock frequency, this parallel processing technique, such as the channeling technique described above, can increase the speed at which the machine 10 processes the raw data by a factor of approximately n + 1 compared to with a single processor machine (not shown in Fig. 1).
Unfortunately, although the computing machine 10 can process data faster than a single processor computing machine (not shown in Fig. 1), the data processing speed of the machine 10 is normally significantly less than the frequency of the processor. processor clock. Specifically, the data processing speed of the computing machine 10 is limited by the time required by the master processor 12 and the coprocessors 14 to process data. For the sake of brevity, an example of this speed limitation is described in combination with the master processor 12, although it is understood that this description also applies to the coprocessors 14. As described above, the master processor 12 runs a program that controls the processor to manipulate data in a desired way. This program comprises a sequence of instructions that are executed by processor 12. Unfortunately, processor 12 typically requires multiple clock cycles to execute a single instruction, and it often must execute multiple instructions to process a single data value. For example, suppose that processor 12 is to multiply a first data value A (not shown) by a second data value B (not shown). During a first clock cycle, processor 12 retrieves a multiplication instruction from memory 22. During the second and third clock cycles, processor 12 respectively retrieves A and B from memory 26. During a fourth clock cycle, processor 12 multiplies A and B, and, during a fifth clock cycle, stores the product resulting in memory 22 or 26 or provides the resulting product to the remote device (not shown). This is the best possible situation, since in many cases the processor 12 requires additional clock cycles for supplementary tasks such as initializing and closing counters. For this reason, at best processor 12 requires five clock cycles, or an average of 2.5 clock cycles per data value, to process A and B.
Accordingly, the speed at which the computing machine 10 processes data is typically significantly less than the frequency of the clock that controls the master processor 12 and the coprocessors 14. For example, if the processor 12 is driven by clock pulses at 1.0 Gigahertz (GHz) although it requires an average of 2.5 clock cycles per data value, in that case the effective data processing speed is equal to (1.0 GHz) / 2.5 = 0.4 GHz. Typically, this effective data processing speed is characterized in units of operations per second. Thus, in this example, for a 1.0 GHz clock speed, processor 12 would be assigned a nominal data processing speed of 0.4 Gigaoperations / second (Gops).
FIG. 2 is a block diagram of a hard-wired data pipeline 30 that can typically process data faster than a processor for a given clock frequency, and typically at substantially the same rate at which it is used. activates the pipeline using clock pulses. The pipeline 30 comprises operator circuits 32<sub>1</sub> to 32<sub>n</sub> each performing a respective operation on respective data without executing program instructions. That is, the desired operation is "burned" in a circuit 32 such that it implements the operation automatically, without the need for program instructions.
IS 2 300 633 T3
By removing the overhead associated with executing program instructions, pipeline 30 can typically perform more operations per second than a processor for a given clock frequency.
For example, pipeline 30 can often solve the following equation faster than a processor for a given clock frequency:
Y (xk) = (5xk + 3) 2<sup>xk</sup> (1) where x<sub>k</sub> represents a sequence of raw data values. In this example, the operator circuit 32<sub>1</sub> is a multiplier that calculates 5x<sub>k</sub>, circuit 32<sub>2</sub> is an adder that calculates 5x<sub>k</sub>+3, and circuit 32<sub>n</sub> (n = 3) is a multiplier that calculates (5x<sub>k</sub>+3)2<sup>xk</sup>.
During a first clock cycle k = 1, circuit 32<sub>1</sub> receives data value x<sub>1</sub> and multiply it by 5 to generate 5x<sub>1</sub>.
During a second clock cycle k = 2, circuit 322 receives 5xi from circuit 32i and adds 3 to generate 5xi + 3. Also, during the second clock cycle, circuit 321 generates 5x2.
During a third clock cycle k = 3, circuit 32<sub>3</sub> get 5x<sub>1</sub>+3 from circuit 32<sub>2</sub> and perform a multiplication by 2<sup>x1</sup> (specifically move 5x1 + 3 to the right x1 positions) to generate the first result (5x1 + 3) 2<sup>x1</sup>. Also during the third clock cycle, circuit 321 generates 5x<sub>3</sub> and circuit 32<sub>2</sub> generates 5x<sub>2</sub>+3.
Pipeline 30 continues with the processing of subsequent raw data values xk in this manner until all raw data values have been processed.
Therefore, after a delay of two clock cycles after receiving a raw data value x<sub>1</sub> -This delay is usually called pipeline latency 30- the pipeline generates the result (5x<sub>1</sub>+3)2<sup>x1</sup>, and after this it generates a result for each clock cycle.
Thus, omitting latency, pipe 30 has a data processing speed equal to the clock speed. In comparison, assuming that the master processor 12 and the coprocessors 14 (Fig. 1) have data processing speeds that are 0.4 times the clock speed as in the previous example, the pipeline 30 can process data 2.5 times faster than computing machine 10 (Fig. 1) for a given clock speed.
Still referring to Fig. 2, a designer may select the implementation of pipeline 30 in a programmable logic IC (PLIC), such as an in-situ programmable gate array (FPGA), since a PLIC allows flexibility of the design and modifications greater than an Application Specific IC (ASIC). To configure permanent connections within a PLIC, the designer simply sets interconnect configuration registers arranged within the PLIC in predetermined binary states. The combination of all these binary states is usually called a "microprogram". Typically, the designer loads this firmware into a non-volatile memory (not shown in Fig. 2) that is coupled to the PLIC. When the PLIC is “turned on”, it downloads the firmware from memory to the interconnect configuration registers. Thus, to modify the operation of the PLIC, the designer simply modifies the firmware and allows the PLIC to download the modified firmware into the interconnect configuration registers. This ability to modify the PLIC simply by modifying the firmware is particularly useful during the prototype testing phase and for updating the pipeline 30 "in situ".
Unfortunately, the hardwired pipeline 30 typically cannot execute all algorithms, particularly those involving meaningful decision making. Typically, a processor can execute a decision-making instruction (for example, conditional instructions such as "if A, then go to B, if not go to C") about as fast as it can execute an instruction from a operation (eg "A + B") of comparable length. However, while pipeline 30 may be capable of making a relatively simple decision (eg, "A> B?"), It typically cannot execute a relatively complex decision (eg, "if A, then go to B, if not go to C ”). Furthermore, while it is possible that pipeline 30 could be designed to execute such a complex decision, the size and complexity of the required circuitry often makes such a design impractical, particularly when an algorithm comprises multiple different complex decisions.
Therefore, processors are typically used in applications that require significant decision making, and hard-wired pipelines are typically limited for "compute intensive" applications that involve little or no decision making.
Furthermore, as described below, it is typically much easier to design / modify a processor-based computing machine, such as the computing machine 10 of Fig. 1, than to design / modify a permanent connection pipeline such as the pipeline. 30 of Fig. 2, particularly when pipeline 30 comprises multiple PLICs.
IS 2 300 633 T3
Typically, computing components, such as processors and their peripherals (eg, memory), comprise industrially standardized communication interfaces that facilitate the interconnection of the components to form a processor-based computing machine.
Typically, a standard communication interface comprises two layers: a physical layer and a service layer.
The physical layer comprises the circuitry and the corresponding circuit interconnections that form the interface and the operating parameters of this circuitry. For example, the physical layer comprises the pins that connect the component to a bus, the buffers that hold data received from the pins, and the controllers that push data over the pins. The operating parameters include the acceptable voltage range of the data signals received by the pins, the timing of the signals for writing and reading data, and the supported operating modes (eg, burst mode, search mode). Conventional physical layers comprise transistor-transistor logic (TTL) and RAMBUS.
The service layer comprises the protocol by which a computing component transfers data. The protocol defines the data format and the manner in which the component sends and receives the formatted data. Conventional communication protocols include File Transfer Protocol (FTP) and TCP / IP (expand).
Consequently, as manufacturers and other vendors typically design computer components that have industrially standardized communication layers, it is typically possible to design the interface of that component and interface it to other computer components with relatively little effort. This allows most of the time to be spent designing the other parts of the computing machine, and the machine can be easily modified by adding or removing components.
Designing a computing component that supports an industry standard communication layer saves design time by using an existing physical layer design from a design library. This option also ensures that the component can easily interface with other commercially available computer components.
Furthermore, designing a computing machine using computing components that support a common industrially standardized communication layer allows the designer to interconnect the components with reduced time and effort. Because the components support a common interface layer, the designer can interconnect them via a system bus with little design effort. Furthermore, since the supported interface layer is industrially standardized, the machine can be easily modified. For example, different components and peripherals can be added to the machine as system design evolves, or next generation components can be easily added / designed as technology evolves. Furthermore, since the components support a common industrially standardized service layer, an existing software module that implements the corresponding protocol can be incorporated into the software of the computing machine. In this way, the components can be communicated via interface with little effort since the interface design is essentially already done, and therefore it is possible to concentrate on the design of the parts (for example, the software) of the machine that make the machine perform the desired function (s).
Unfortunately, however, there are no known, industrially standardized communication layers for components, such as PLICs, used to form permanent connection conduits such as conduit 30 of Fig. 2.
Therefore, to design a pipeline that features multiple PLICs, typically a significant amount of time is consumed and significant effort is required designing and debugging the communication layer between PLICs "from scratch." Typically, such an ad hoc communication layer depends on the parameters of the data being transferred between the PLICs. Similarly, to design a pipeline that interfaces with a processor, it would take a significant amount of time and significant effort to design and debug the communication layer between the pipeline and the processor from scratch.
Similarly, to modify such a pipeline by adding a PLIC therein, typically a significant amount of time is consumed and significant effort is exerted in designing and debugging the communication layer between the added PLIC and existing PLICs. Similarly, to modify a pipeline by adding a processor, or to modify a computing machine by adding a pipeline, it would take a significant amount of time and exert significant effort in designing and debugging the layer. communication between the pipeline and the processor.
Therefore, referring to Figs. 1 and 2, due to the difficulties of interfacing multiple PLICs and interfacing a processor with a pipeline, users are often forced to compromise when designing a computing machine. For example, with a processor-based computing machine, the user is forced to compromise computing-intensive speed for complex decision-making capabilities and design / modification flexibility. Conversely, with a computing machine based on permanent connection pipelines, the user is obliged
ES 2 300 633 T3 to give up complex decision-making capabilities and design / modification flexibility due to speed-intensive computation. Furthermore, due to the difficulties of interfacing multiple PLICs, it is often impractical for a user to design a pipeline based machine whose number of PLIC circuits is not reduced. As a consequence, normally a practical pipe-based machine has limited functionality. Furthermore, due to difficulties in interfacing a processor with one PLIC, it would be impractical to interface a processor with more than one PLIC. As a consequence, the benefits of combining a processor and a pipeline would be minimal.
For this reason, a need has arisen for a new computing architecture that allows the combination of the decision-making power of a processor-based machine with the computationally intensive speed of a machine based on permanent connection pipelines.
Document EP-A-0 945 788 discloses a data processing system comprising a core of a digital signal processor and a coprocessor, which responds to commands from the core of the digital signal processor. The applicant observes that the orders sent by the digital signal processor include operation codes that define the type of calculation, with which the orders are programming instructions for the coprocessor, and said coprocessor must execute these programming instructions to process and generate resulting data. The applicant also observes that, during its operation, the core of the digital signal processor controls the data and coefficients used by the coprocessor by loading the data to be processed in the data memory and the coefficients in the coefficient memory. After the transfer of data to be processed, the core of the digital signal processor signals to the coprocessor the command corresponding to the desired algorithm of signal processing. In this way, the orders sent by the core of the digital signal processor to the coprocessor are part of the program code that is executed by the core of the digital signal processor in the process of generating the data to be fed to the coprocessor. .
S. Bakshi et al., "Partitioning and Pipelining for Performance-Constrained Hardware / Software Systems," IEEE Trans. on VLSI Systems, vol. 7, No. 4, December 1999, pages 419 to 432, disclose (Figure 2) a pipelined architecture comprising one or more processors, one or more ASICs, and one or more memory chips that all communicate through one or more buses. The present Applicant notes that ASICs are permanent patch circuits, which cannot be (re) configured later.
G. Lecurieux-Lafayette, "Un seul FPGA dope le traitement d'images", Electronique, CEP Communication, Paris, n ° 55, 1996, pages 98, 101 to 103, describe an "Image computer" comprising a microprocessor and a coprocessor based on an FPGA, in which the FPGA is reconfigurable during operation.
Summary
In one embodiment of the invention, a peer-to-peer vector transfer machine is provided as set forth in claim 1, and a method as set forth in claim 7.
Since the peer-to-peer vector machine comprises both a processor and a permanent connection pipeline accelerator, it can usually process data more efficiently than a computing machine that includes only processors or only permanent connection pipes. For example, the pairwise vector machine can be designed so that the central processor performs the non-mathematically intensive decision making and operations while the accelerator performs the mathematically intensive operations. By shifting mathematically intensive operations to the throttle, the peer-to-peer vector machine, for a given clock frequency, can typically process data at a rate that exceeds the rate at which a single-processor machine can process the data.
Brief description of the drawings
Fig. 1 is a block diagram of a computing machine having a conventional multiprocessor architecture.
Fig. 2 is a block diagram of a conventional permanent connection pipeline.
Fig. 3 is a schematic block diagram of a computing machine exhibiting a peer-to-peer vector architecture according to one of the embodiments of the invention.
Fig. 4 is a schematic block diagram of an electronic system incorporating the peer-to-peer computing machine of Fig. 3 according to one of the embodiments of the invention.
Detailed description
FIG. 3 is a schematic block diagram of a computing machine 40, featuring a peer-to-peer vector architecture in accordance with one embodiment of the invention. In addition to a central processor 42, the peer-to-peer vector machine 40 comprises a pipeline accelerator 44, which performs at least a part of the data processing, and which therefore effectively replaces the bank of coprocessors 14 of the
In this way, the central processor 42 and the accelerator 44 are "peer entities" that can transfer vectors of data from one side to the other. Since accelerator 44 does not execute program instructions, it typically performs mathematically intensive operations on data significantly faster than a bank of coprocessors for a given clock frequency. Consequently, by combining the decision-making ability of processor 42 and the computation-intensive ability of accelerator 44, machine 40 exhibits the same capabilities as a conventional computing machine such as machine 10, although it can typically process data faster than this last. In addition, as described in the previously cited US publications n ° 2004/0181621 entitled COMPUTING MACHINE HAVING IMPROVED COMPUTING ARCHITECTURE AND RELATED SYSTEM AND METHOD and 2004/0136241 entitled PIPELINE ACCELERATOR FOR IMPROVED COMPUTING ARCHITECTURE AND RELATED SYSTEM AND METHOD, the fact of supplying to the accelerator 44 the same communication layer as the central processor 42 facilitates the design and modification of the machine 40, particularly when the communications layer corresponds to an industrial standard. Furthermore, when the accelerator 44 includes multiple components (for example, PLIC circuits), supplying these components with this same communication layer facilitates the design and modification of the accelerator, particularly when the communication layer corresponds to an industrial standard. . On the other hand, the machine 40 may also provide other advantages as described below and in the patent applications cited above.
In addition to central processor 42 and pipeline accelerator 44, peer vector computing machine 40 comprises processor memory 46, interface memory 48, bus 50, microprogram memory 52, optional raw data input ports 54 and 92 (port 92 shown in Fig. 4), optional processed data output ports 58 and 94 (port 94 shown in Fig. 4), and an optional router 61.
The central processor 42 comprises a processing unit 62 and a message handler 64, and the processor memory 46 comprises a processing unit memory 66 and a handler memory 68, which respectively act as both program and job memories. for the processor unit and the message handler. Processor memory 46 further comprises an accelerator configuration register 70 and a message configuration register 72, which store respective configuration data that allows central processor 42 to configure the operation of accelerator 44 and the structure of messages that generates message handler 64.
Pipeline accelerator 44 is disposed in at least one PLIC (not shown) and comprises permanent link pipes 74 to 74<sub>n</sub>, which process respective data without executing program instructions. The firmware memory 52 stores the configuration firmware corresponding to the accelerator 44. If the accelerator 44 is arranged in multiple PLICs, these PLICs and their respective firmware memories can be arranged on multiple circuit boards, that is, daughter cards (not shown ). The accelerator 44 and daughter cards are further described in the previously cited US Publications Nos. 2004/0136241 titled PIPELINE ACCELERATOR FOR IMPROVED COMPUTING ARCHITECTURE AND RELATED SYSTEM AND METHOD and 2004/0130927 titled PIPELINE ACCELERATOR HAVING MULTIPLE PIPELINE UNITS AND RELATED METHOD MACHINE UNITS AND RELATED METHOD MACHINE UNITS . Alternatively, throttle 44 may be arranged in at least one ASIC, and therefore may have internal interconnections that are not configurable. In this alternative, the machine 40 may bypass the firmware 52. Furthermore, although the accelerator 44 is shown to comprise multiple pipes 74, it may comprise only a single pipeline.
Still referring to FIG. 3, the operation of the peer-to-peer vector machine 40 according to one of the embodiments of the invention will now be described.
Peer-to-peer vector machine setup
When the peer vector machine 40 is first activated, the processing unit 62 configures the message handler 64 and the pipeline accelerator 44 (when the accelerator is configurable) so that the machine executes the desired algorithm. Specifically, the processing unit 62 executes a main application program that is stored in memory 66 and which causes the processing unit to configure the message handler 64 and the accelerator 44 as described below.
To configure message handler 64, processing unit 62 retrieves message format information from register 72 and provides this format information to message handler, which stores this information in memory 68. When machine 40 processes Data As will be described later, the message handler 64 uses this format information to generate and decrypt data messages that have a desired format. In one embodiment, the format information is written in Extensible Markup Language (XML), although it may be written in another language or data format. Since the processing unit 62 configures the message handler 64 each time the peer vector machine 40 is activated, the format of the message can be modified simply by modifying the format information stored in register 72. Alternatively, an external message configuration library (not shown) can store information corresponding to multiple message formats, and the main application can be designed and / or modified so that processing unit 62 updates record 72 from parts selected files from the library, and then download the desired formatting information from the updated registry to the
ES 2 300 633 T3 messages 64. The format of the messages and their generation and decryption are further described below and in the previously cited US publication No. 2004/0181621 entitled COMPUTING MACHINE HAVING IMPROVED COMPUTING ARCHITECTURE AND RELATED SYSTEM AND METHOD.
Similarly, to configure the distribution of pipeline accelerator interconnects 44, processing unit 62 retrieves a configuration firmware from register 70 and downloads this firmware into memory 52 via message handler 64 and the bus. 50. The accelerator 44 then configures itself by downloading the firmware from memory 52 to its interconnect configuration registers (not shown). Since the processing unit 62 configures the accelerator 44 each time the vector machine between pairs 40 is activated, the distribution of the interconnections - and therefore the operation - of the accelerator 44 can be modified simply by modifying the microprogram stored in the register. 70. Alternatively, an external throttle setting library (not shown) can store a firmware corresponding to multiple throttle settings 44, and the main application can be designed and / or modified such that processing unit 62 updates register 70 from of selected parts of the library, and then download the desired firmware from the updated register to memory 52. In addition, the external library or register 70 may store firmware modules that define different parts and / or functions of the accelerator 44. Thus, these modules can be used to facilitate the design and / or modification of the accelerator 44. Additionally, processing unit 62 can use these modules to modify throttle 44 while machine 40 is processing data. The configuration of the interconnects of the accelerator 44 and the firmware modules is further described in the aforementioned US Publication No. 2004/0170070 entitled PROGRAMMABLE CIRCUIT AND RELATED COMPUTING MACHINE AND METHOD.
In addition, the processing unit 62 may "surface-configure" the pipeline accelerator 44 while the peer vector machine 40 is processing data. That is, the processing unit 62 can configure the operation of the accelerator 44 without altering the distribution of accelerator interconnects. Such a surface configuration is further described below and in US Publication No. 2004/0136241 entitled PIPELINE ACCELERATOR FOR IMPROVED COMPUTING ARCHITECTURE AND RELATED SYSTEM AND METHOD.
Data processing with the peer-to-peer vector machine
In general, the peer-to-peer vector machine 40 efficiently divides the raw data processing between the central processor 42 and the pipeline accelerator 44. For example, the central processor 42 can perform most or all of the operations of decision making related to the data, and the accelerator 44 may perform most or all of the mathematically intensive operations on the data. However, the machine 40 can divide the data processing in any desired manner.
Core processor operation
In one embodiment, the central processor 42 receives the raw data from and provides the resulting processed data to a remote device such as an array of sonar elements (Fig. 4).
The central processor 42 first receives the raw data from the remote device through the input port 54 or bus 50. The peer-to-peer vector machine 40 may comprise a FIFO (not shown) to temporarily store the received raw data.
Next, the processing unit 62 prepares the raw data to be processed by the pipeline accelerator 44. For example, the unit 62 can determine, for example, which of the raw data is to be sent to the accelerator 44 or in what sequence it will be sent. they are going to send the raw data. Alternatively, unit 62 may process the raw data to generate intermediate data with a view to sending it to accelerator 44. The preparation of the raw data is further described in the aforementioned US Publication No. 2004/0181621 entitled COMPUTING HAVING IMPROVED COMPUTING ARCHITECTURE AND RELATED SYSTEM AND METHOD.
While the raw data is being prepared, the processing unit 54 can also generate one or more "surface configuration" commands to modify the operation of the accelerator 44. Unlike the firmware that configures the distribution of the interconnects of the accelerator 44 when the accelerator is activated. In machine 40, a surface configuration command controls the operation of the throttle without altering the distribution of its interconnections. For example, a surface configuration command may control the size of the data strings (eg, 32-bit or 64-bit) that are processed by the accelerator 44. The surface configuration of the accelerator 44 is further described in the US publication cited above. n ° 2004/0136241 entitled PIPELINE ACCELERATOR FOR IMPROVED COMPUTING ARCHITECTURE AND RELATED SYSTEM AND METHOD.
The processing unit 62 then loads the prepared data and / or the surface configuration order (s) into a corresponding location in the interface memory 48, which acts as a FIFO buffer between the unit 62 and the throttle 44.
Next, the message handler 64 retrieves the prepared data and / or the software command (s) from the interface memory 48 and generates message objects comprising the data and / or the command (s). ) and related information. Typically, the accelerator 44 needs four identifiers that describe the data / command (s) and the information.
ES 2 300 633 T3 mation related (collectively referred to as "information"): a) the desired destination of the information (eg pipeline 74i), b) priority (eg, whether the throttle processes this data before or after of previously received data), c) the length or end of the message object, and d) the unique instance of the data (for example, sensor signal number nine of a thousand sensor array of elements). To facilitate this determination, message handler 64 generates message objects that have a predetermined format as described above. In addition to the prepared data / the surface configuration command (s), a message object typically comprises a header comprising the four identifiers described above and which may also comprise identifiers that describe the type of information that the object comprises (e.g. example, data, order), and the algorithm by which the data is to be processed. This last identifier is useful when the destination pipeline 74 implements multiple algorithms. The handler 64 may retrieve the header information from interface memory 48, or it may generate the header based on the interface memory location from which it retrieves the prepared data or the order (s) ). By decrypting the message header, the router 61 and / or the accelerator 44 can direct the information within the message object to the desired destination, and have the destination process the information in a desired sequence.
There are alternative embodiments for generating the message objects. For example, although each message object is described as comprising either data or a shallow configuration command, an individual message object may comprise both data and one or more commands. Furthermore, although the message handler 64 is described as receiving the data and commands from the interface memory 48, it can receive the data and commands directly from the processing unit 54.
The generation of message objects is further described in the aforementioned US Publication No. 2004/0181621 entitled COMPUTING MACHINE HAVING IMPROVED COMPUTING ARCHITECTURE AND RELATED SYSTEM AND METHOD.
Pipeline Accelerator
The pipeline accelerator 44 receives and decrypts the message objects from the message handler 64 and efficiently directs the data and / or commands within the objects to the desired destination (s). This technique is particularly useful when the number of algorithms implemented by processing unit 62 and pipes 74 is relatively small, and therefore router 61 can be omitted. Alternatively, when the number of algorithms implemented by the processing unit 62 or the number pipelines 74 is relatively high, the router 61 receives and decrypts the message objects from the message handler 64 and efficiently directs the data and / or commands that are found. inside the objects towards the desired destination (s) within the throttle 44.
In one of the embodiments in which there are a reduced number of algorithms of the processing units and pipelines 74, each pipeline simultaneously receives a message object and analyzes the header to determine whether or not it is a desired recipient of the message. If the message object is destined for a specific pipeline 74, then that pipeline decrypts the message and processes the data / order (s) retrieved. However, if the message object is not destined for a specific pipeline 74, then that pipeline ignores the message object. For example, suppose a message object comprises data to be processed by pipeline 74<sub>1</sub>. Therefore channeling 74<sub>1</sub> parses the message header, determines that it is a desired destination for the data, retrieves the message data, and processes the retrieved data. Conversely, each of the pipes 742 to 74<sub>n</sub> parses the message header, determines that it is not a desired destination for the data, and therefore does not retrieve or process the data. If the data within the message object is destined for multiple pipes 74, then the message handler 64 generates and sends a sequence of respective message objects comprising the same data, one message for each destination pipeline. Alternatively, the message handler 64 may simultaneously send the data to all the destination pipes 74 by sending a single message object that has a header identifying all of the destination pipes. The retrieval of data and surface configuration commands from message objects is further described in the aforementioned US Publication No. 2004/0136241 entitled PIPELINE ACCELERATOR FOR IMPROVED COMPUTING ARCHITECTURE AND RELATED SYSTEM AND METHOD.
In another embodiment where there are a large number of processes from the processing units or pipes 74, each pipeline receives message objects from router 61. Although router 61 should ideally send message objects only towards target pipeline 74, the target pipeline continues to analyze the header to determine whether or not it is a desired recipient of the message. This analysis identifies potential errors in the routing of the message, that is, exceptions. If the message object is destined for a target pipeline 74, then that pipeline decrypts the message and processes the data / order (s) retrieved. However, if the message object is not destined for the target pipeline 74, in that case said pipeline ignores the processing corresponding to that message object, and can also issue a new message to the central processor 42 indicating that an exception of the routing. Routing exception handling is described in the aforementioned US Publication No. 2004/0181621 entitled COMPUTING MACHINE HAVING IMPROVED COMPUTING ARCHITECTURE AND RELATED SYSTEM AND METHOD.
IS 2 300 633 T3
The pipeline accelerator 44 then processes the incoming data and / or commands retrieved from the message objects.
When it comes to data, the destination pipeline (s) 74 perform a respective operation or operations on it. As described in conjunction with Fig. 2, since the pipes 74 do not execute program instructions, they can typically process data at a rate that is substantially equal to the frequency of the pipeline clock.
In a first embodiment, a single pipeline 74 generates resulting data by processing the incoming data.
In a second embodiment, multiple pipelines 74 generate resulting data by serially processing the incoming data. For example, pipeline 74 may generate a first intermediate data by performing a first operation on the incoming data. Next, pipeline 74<sub>2</sub> it can generate a second intermediate data by performing a second operation on the first intermediate data, and so on, until the final pipeline 74 of the chain generates the data that is the result.
In a third embodiment, multiple pipelines 74 generate the resulting data by processing the incoming data in parallel. For example, pipeline 741 may generate a first set of the resulting data by performing a first operation on a first set of the incoming data. At the same time, pipeline 742 can generate a second set of resulting data by performing a second operation on a second set of the incoming data, and so on.
Alternatively, pipelines 74 may generate resulting data from the incoming data according to any combination of the above three embodiments. For example, pipeline 741 may generate a first set of the resulting data by performing a first operation on a first set of the incoming data. At the same time, pipelines 742 and 74n can generate a second set of resulting data by serially performing second and third operations on a second set of the incoming data.
In any of the above embodiments and alternatives, a single pipeline 74 can perform multiple operations. For example, pipeline 741 can receive data, generate a first intermediate data by performing a first operation on the received data, temporarily store the first intermediate data, generate a second intermediate data by performing a second operation on the first intermediate data, and so on. until you generate the result data. A number of techniques are available to move pipe 741 from performing the first operation to performing the second operation, and so on. Such techniques are described in the aforementioned US Patent Application Serial No. 10 / 683,929 entitled PIPELINE ACCELERATOR FOR IMPROVED COMPUTING ARCHITECTURE AND RELATED SYSTEM AND METHOD (File No. 1934-12-3).
For a shaping command, accelerator 44 sets the bits in the corresponding shaping register (s) (not shown) as indicated by the message header. As described above, setting these bits typically changes the operation of the throttle 44 without changing the distribution of their interconnections. This situation is similar to setting bits in a processor control register to, for example, set an external pin as an input pin or output pin, or select an addressing mode. Additionally, a shallow configuration command can split a record or table (an array of records) to hold data. Another surface configuration command or an operation performed by the accelerator 44 may load data into the surface configured table or register. The surface configuration of the accelerator 44 is further described in the previously cited US Publication No. 2004/0136241 entitled PIPELINE ACCELERATOR FOR IMPROVED COMPUTING ARCHITECTURE AND RELATED SYSTEM AND METHOD (File No. 1934-13-3).
Pipeline accelerator 44 then provides the resulting data to central processor 42 via router 61 (or directly in case the router is bypassed) for further processing.
Alternatively, the throttle 44 provides the resulting data to the remote destination (Fig. 4) either directly through the output port 94 (Fig. 4), or indirectly through the router 61 (if present). , bus 50, central processor 42, and output port 58. Accordingly, in this alternate embodiment, the resulting data generated by throttle 44 is the final processed data.
When the accelerator 44 provides the resulting data to the central processor 42 - either for further processing or for transferring it to the remote device (Fig. 4) - it sends this data in a message object that has the same format as the message objects generated by the message handler 64. Like the message objects generated by the message handler 64, the message objects generated by the accelerator 44 comprise headers that specify, for example, the destination and priority of the resulting data. For example, the header may instruct the message handler 64 to pass the resulting data to the remote device through port 58, or it may specify which part of the program executed by the processing unit 62 is to control the processing of the data. By using the same message format, accelerator 44 presents the same interface layer as core processor 42. This situation facilitates the design and modification of the peer-to-peer vector machine 40, particularly where the interface layer is an industry standard.
IS 2 300 633 T3
The structure and operation of the pipeline accelerator 44 and the pipes 66 are further described in the previously cited US publication No. 2004/0136241 entitled PIPELINE ACCELERATOR FOR IMPROVED COMPUTING ARCHITECTURE AND RELATED SYSTEM AND METHOD (docket No. 1934-133).
Receiving and processing from the pipeline accelerator with the core processor
When message handler 64 receives a message object from accelerator 44, it first decrypts the message header and directs the retrieved data to the indicated destination.
If the header indicates that the data is to be moved to the remote device (Fig. 4) through port 58, then the message handler 64 can provide the data directly to port 58, or to a FIFO buffer (not shown) from the port formed in interface memory 48 or other memory and then from the buffer to port 58. Multiple ports 58 and multiple respective remote devices are also contemplated.
However, if the header indicates that the processing unit 62 is to further process the data, then the message handler 62 stores the data in a location in interface memory 48 that corresponds to the program part of the processing unit that will control the data processing. More specifically, at this time the same header indirectly indicates which part (s) of the program executed by the processing unit 54 is (are) going to control the processing of the data. Consequently, the message handler 64 stores the data at the location (eg, a FIFO) of the interface memory 48 corresponding to this part of the program.
As described above, interface memory 48 acts as a buffer between accelerator 44 and processing unit 62, and thus enables data transfer when the processing unit is not synchronized with the accelerator. For example, this lack of synchronization can occur when the accelerator 44 processes data faster than the processing unit 62. By using interface memory 48, throttle 44 is not slowed down by the slower response of processing unit 62. This further avoids the inefficiency penalties associated with indeterminate response time of the processing unit to treatment interruptions. The indeterminate treatment, by the processing unit 62, of the output messages of the accelerator 44 would unnecessarily complicate the design of the accelerator forcing the designer to provide: a) either a storage and processing means for the output messages of those that have been backed up, or b) rest controls throughout the pipeline to avoid overwriting the messages that have been backed up. Therefore, the use of interface memory 48, which acts as a buffer between accelerator 44 and processing unit 62, has several desirable consequences a) accelerators are easier to design, b) accelerators need a smaller infrastructure and may contain larger PLIC applications, c) Accelerators can be streamlined to run faster as the output data is not “locked” by a slower processor.
Next, for data that has been stored by message handler 64 in interface memory 48, processing unit 62 retrieves the data from interface memory. The processing unit 62 can interrogate the interface memory 48 to determine when new data has arrived at a specific location, or the message handler 64 can generate an interrupt or other signal that notifies the processing unit of the arrival of the data. . In one embodiment, before the processing unit 62 retrieves data, the message handler 64 generates a message object that comprises the data. More specifically, the program executed by processing unit 62 can be designed to receive data in message objects. In this way, the message handler 64 could store a message object in interface memory 48 instead of just storing the data. However, a message object typically occupies significantly more memory space than the data it contains. Accordingly, to save memory, message handler 64 decrypts a message object from pipeline accelerator 44, stores the data in memory 48, and then effectively regenerates the message object when processing unit 62 is ready to receive the messages. data. The processing unit 62 then decrypts the message object and processes the data under the control of the program part identified in the message header.
Next, the processor unit 62 processes the retrieved data under the control of the destination part of the program, generates processed data, and stores the processed data in a location in interface memory 48 that corresponds to the desired destination of the processed data. .
The message handler 64 then retrieves the processed data and delivers it to the indicated destination. To retrieve the processed data, the message handler 64 may interrogate the memory 48 to determine when the data has arrived, or the processing unit 62 may notify the message handler of the arrival of the data with an interrupt or other signal. To deliver the processed data to its desired destination, the message handler 64 may generate a message object comprising the data, and send the message object back to the accelerator 44 for further processing of the data. Alternatively, the handler 56 may send the data to port 58, or to another location in memory 48 for further processing by the processing unit 62.
IS 2 300 633 T3
The reception and processing of data, by the central processor, coming from the pipeline accelerator 44 is further described in the previously cited US publication No. 2004/0181621 entitled COMPUTING MACHINE HAVING IMPROVED COMPUTING ARCHITECTURE AND RELATED SYSTEM AND METHOD (file no 1934-12-3).
Alternative data processing techniques using the peer-to-peer vector machine
Still referring to Fig. 3, there are alternatives to the above-described embodiments in which the central processor 42 receives and processes data, and then sends the data to the pipeline accelerator 44 for further processing.
In one alternative, the central processor 42 performs all the processing on at least part of the data, and thus sends this data to the pipeline accelerator 44 for further processing.
In another alternative, pipeline accelerator 44 receives the raw data directly from the remote device (Fig. 4) through port 92 (Fig. 4) and processes the raw data.
The accelerator 44 can then send the processed data directly back to the remote device through port 94, or it can send the processed data to the central processor 42 for further processing. In the latter case, the accelerator 44 can encapsulate the data in message objects as described above.
In still another alternative, the accelerator 44 may comprise, in addition to the permanent connection pipes 74, one or more instruction executing processors, such as a Digital Signal Processor (DSP), to complement the computation-intensive capabilities of the pipes.
Example of implementation of the peer-to-peer vector machine
Still referring to FIG. 3, in one embodiment, channel bus 50 is a standard 133 MHz PCI bus, channel 74 are comprised of one or more standard PMC cards, and memory 52 is one of the flash memories that are each located on a respective PMC card.
Example of application of the peer-to-peer vector machine
FIG. 4 is a block diagram of a sonar system 80 incorporating the peer-to-peer vector machine 40 of FIG. 3 in accordance with one embodiment of the invention. In addition to machine 40, system 80 comprises an array 82 of transducer elements 84 to 84<sub>n</sub> to receive and transmit sonar signals, digital-to-analog converters (DAC converters) 86<sub>1</sub> to 86<sub>n</sub>, analog-to-digital converters (ADC converters) 88<sub>1</sub> to 88<sub>n</sub>, and a data interface 90. As the generation and processing of sonar signals are normally mathematically intensive functions, the machine 40 can normally perform these functions more quickly and efficiently than a conventional computing machine - such as the multiprocessor machine 10 (Fig. 1) - for a given clock frequency as described above in combination with Fig. 3.
During a transmit mode of operation, matrix 82 transmits a sonar signal in a medium such as water (not shown). First, the peer-to-peer vector machine 40 converts raw signal data received at port 92 into n digital signals, one for each of the array elements 84. The magnitudes and phases of these signals dictate the pattern of the signal beam. matrix transmission 82. Next, machine 40 provides these digital signals to interface 90, which provides these signals to respective DACs 86 for conversion to respective analog signals. For example, the interface 90 may act as a buffer that serially receives the digital signals from the machine 40, stores these signals until it receives and temporarily stores the n's in their entirety, and then simultaneously provides these sequential signal samples to the devices. Respective DACs 86. The transducer elements 84 then convert these analog signals into respective sound waves, which interfere with each other to form the beams of a sonar signal.
During a receive mode of operation, matrix 82 receives a sonar signal from the medium (not shown). The received sonar signal is composed of the part of the transmitted sonar signal that is reflected by remote objects and the sound energy emitted by the environment and remote objects. First, the transducer elements 84 receive respective sound waves that make up the sonar signal, convert these sound waves into n analog signals, and provide these analog signals to the ADCs 88 for conversion into n respective digital signals. Interface 90 then provides these digital signals to peer vector machine 40 for processing. For example, interface 90 may act as a buffer that receives digital signals from ADCs 88 in parallel and then provides these signals serially to machine 40. The processing that machine 40 performs on the digital signal dictates the matrix receive beam pattern 82. Additional processing steps such as filtering, band shift, spectral transformation (eg Fourier transform) are applied to digital signals. Machine 40 then provides the data from the processed signals through port 94 to another apparatus such as a display device to view the location of objects.
IS 2 300 633 T3
The peer-to-peer vector machine 40 may also incorporate other non-sonar systems, although it has been described in conjunction with the sonar system 80.
The above description is presented to enable one skilled in the art to embody the invention and make use of it. Various modifications of embodiments will become apparent to those skilled in the art, and the generic principles herein may be applied to other embodiments and applications. Thus, the present invention is not limited to the embodiments just described, but should be granted the broadest scope in accordance with the principles and features disclosed herein.
Contents8
4 sheets
Sheet 1 Sheet 2 Sheet 3 Sheet 4
82 members in 10 offices
Priority claims33
| Document | Office | Kind | Date |
|---|---|---|---|
| 20020422503P | United States of America | – | |
| 42250302 | United States of America | P | |
| 42250302 | United States of America | P | |
| 20030683929 | United States of America | – | |
| 20030683932 | United States of America | – | |
| 20030684053 | United States of America | – | |
| 20030684057 | United States of America | – | |
| 20030684102 | United States of America | – | |
| 68392903 | United States of America | A | |
| 68392903 | United States of America | A | |
| 68393203 | United States of America | A | |
| 68393203 | United States of America | A | |
| 68405303 | United States of America | A | |
| 68405303 | United States of America | A | |
| 68405703 | United States of America | A | |
| 68405703 | United States of America | A | |
| 68410203 | United States of America | A | |
| 68410203 | United States of America | A | |
| 0334557 | United States of America | W | |
| 0334557 | United States of America | W | |
| 2003US34557 | World Intellectual Property Organization (WIPO) | – | |
| 422503P03781552 | – | – | – |
| 683929 | – | – | – |
| 683932 | – | – | – |
| 684053 | – | – | – |
| 684102 | – | – | – |
| US20020422503P | – | – | – |
| US20030683929 | – | – | – |
| US20030683932 | – | – | – |
| US20030684053 | – | – | – |
| US20030684057 | – | – | – |
| US20030684102 | – | – | – |
| WO2003US34557 | – | – | – |
Members82
| Document | Office | Kind | |
|---|---|---|---|
| US2002147892A1 | United States of America | A1 | |
| DE10212642A1 | Germany | A1 | |
| US6633965B2 | United States of America | B2 | |
| CA2503611A1 | Canada | A1 | |
| CA2503613A1 | Canada | A1 | |
| CA2503617A1 | Canada | A1 | |
| CA2503620A1 | Canada | A1 | |
| CA2503622A1 | Canada | A1 | |
| WO2004042560A2 | World Intellectual Property Organization (WIPO) | A2 | |
| WO2004042561A2 | World Intellectual Property Organization (WIPO) | A2 | |
| WO2004042562A2 | World Intellectual Property Organization (WIPO) | A2 | |
| WO2004042569A2 | World Intellectual Property Organization (WIPO) | A2 | |
| WO2004042574A2 | World Intellectual Property Organization (WIPO) | A2 | |
| AU2003287317A1 | Australia | A1 | |
| AU2003287318A1 | Australia | A1 | |
| AU2003287319A1 | Australia | A1 | |
| AU2003287320A1 | Australia | A1 | |
| AU2003287321A1 | Australia | A1 | |
| US2004130927A1 | United States of America | A1 | |
| US2004133757A1 | United States of America | A1 | |
| US2004133763A1 | United States of America | A1 | |
| US2004136241A1 | United States of America | A1 | |
| US2004158688A1 | United States of America | A1 | |
| TW200416594A | Taiwan Province of China | A | |
| US2004170070A1 | United States of America | A1 | |
| US2004181621A1 | United States of America | A1 | |
| US2004189686A1 | United States of America | A1 | |
| WO2004042574A3 | World Intellectual Property Organization (WIPO) | A3 | |
| WO2004042560A3 | World Intellectual Property Organization (WIPO) | A3 | |
| EP1559005A2 | European Patent Office (EPO) | A2 | |
| WO2004042562A3 | World Intellectual Property Organization (WIPO) | A3 | |
| KR20050084628A | Republic of Korea | A | |
| KR20050084629A | Republic of Korea | A | |
| KR20050086423A | Republic of Korea | A | |
| KR20050086424A | Republic of Korea | A | |
| EP1570344A2 | European Patent Office (EPO) | A2 | |
| KR20050088995A | Republic of Korea | A | |
| EP1573514A2 | European Patent Office (EPO) | A2 | |
| EP1573515A2 | European Patent Office (EPO) | A2 | |
| EP1576471A2 | European Patent Office (EPO) | A2 | |
| US6990562B2 | United States of America | B2 | |
| WO2004042561A3 | World Intellectual Property Organization (WIPO) | A3 | |
| WO2004042569A3 | World Intellectual Property Organization (WIPO) | A3 | |
| JP2006515941A | Japan | A | |
| US7061485B2 | United States of America | B2 | |
| JP2006518056A | Japan | A | |
| JP2006518057A | Japan | A | |
| JP2006518058A | Japan | A | |
| JP2006518495A | Japan | A | |
| US7103793B2 | United States of America | B2 | |
| EP1570344B1 | European Patent Office (EPO) | B1 | |
| DE60318105D1 | Germany | D1 | |
| US7373432B2 | United States of America | B2 | |
| US7386704B2 | United States of America | B2 | |
| ES2300633T3This record | Spain | T3 | |
| US7418574B2 | United States of America | B2 | |
| US2008222337A1 | United States of America | A1 | |
| DE60318105T2 | Germany | T2 | |
| DE10212642B4 | Germany | B4 | |
| AU2003287317B2 | Australia | B2 | |
| TWI323855B | Taiwan Province of China | B | |
| AU2003287319B2 | Australia | B2 | |
| AU2003287321B2 | Australia | B2 | |
| AU2003287318B2 | Australia | B2 | |
| KR100996917B1 | Republic of Korea | B1 | |
| AU2003287320B2 | Australia | B2 | |
| KR101012744B1 | Republic of Korea | B1 | |
| KR101012745B1 | Republic of Korea | B1 | |
| KR101035646B1 | Republic of Korea | B1 | |
| US7987341B2 | United States of America | B2 | |
| JP2011154711A | Japan | A | |
| JP2011170868A | Japan | A | |
| KR101062214B1 | Republic of Korea | B1 | |
| JP2011175655A | Japan | A | |
| JP2011181078A | Japan | A | |
| CA2503613C | Canada | C | |
| US8250341B2 | United States of America | B2 | |
| CA2503611C | Canada | C | |
| JP2013236380A | Japan | A | |
| JP5568502B2 | Japan | B2 | |
| JP5688432B2 | Japan | B2 | |
| CA2503622C | Canada | C |
Numbers
- Publication
- 2300633
- Publication, DOCDB
- 2300633
- Publication, EPODOC
- ES2300633T
- Application
- 3781552
- Application, DOCDB
- 03781552
- Application, EPODOC
- ES20030781552T
Titles2
- Spanish
- COPROCESADOR CANALIZADO.
- English
- CHANNEL COCKROW.
Classification
- CPC, 2
- G06F9/3879
- G06F15/7839
- IPC, 5
- G06F9 38
- G06F9 30
- G06F9 445
- G06F9 46
- G06F15 78