Digital signal processor including a programmable network
38 claims: 2 independent, 36 dependent
- 1デジタル信号プロセッサであって、 複数のメモリユニット(0、・・・、n)と、 1つ以上の専用機能を実施するように構成された複数のアクセラレータユニット(0、・・・、m)と、 データ経路フロー制御に関連する命令を実行するように構成された実行ユニットを含むプロセッサコア(146)と、 前記命令の実行に応じて、前記複数のメモリユニットと、複数のアクセラレータユニットと、プロセッサコアとの間に選択的に接続性を提供するように構成されたプログラム可能回路網(250)と、を特徴とするデジタル信号プロセッサ。
- 2請求項1に記載のデジタル信号プロセッサであって、特定の命令の実行に応じて、前記プログラム可能回路網(250)は、前記複数のメモリユニット(0、・・・、n)の内の与えられた1つを前記複数のアクセラレータユニット(0、・・・、m)の内の与えられた1つに結合するように構成されるデジタル信号プロセッサ。
- 3請求項1に記載のデジタル信号プロセッサであって、特定の命令の実行に応じて、前記プログラム可能回路網(250)は、前記複数のメモリユニットの1つ以上のメモリユニット(0、・・・、n)を前記プロセッサコアに結合するように構成されるデジタル信号プロセッサ。
- 4請求項1乃至3のいずれかに記載のデジタル信号プロセッサであって、特定の命令の実行に応じて、前記プログラム可能回路網(250)は、前記複数のアクセラレータユニット(0、・・・、m)の2つ以上のアクセラレータユニットを共にチェーン状に結合し、また、更に、前記チェーンの第1アクセラレータユニットを前記複数のメモリユニット(0、・・・、n)の与えられた1つ及び前記プロセッサコア(146)に結合するように構成されるデジタル信号プロセッサ。
- 5請求項4に記載のデジタル信号プロセッサであって、前記複数のアクセラレータユニット(0、・・・、m)の各アクセラレータユニットは、前記チェーンの他のアクセラレータユニットに接続された場合、前記プロセッサコアによる介在なしで接続される、前記アクセラレータユニットと通信を行うように構成されるデジタル信号プロセッサ。
- 6請求項1乃至5のいずれかに記載のデジタル信号プロセッサであって、前記プログラム可能回路網(250)には、前記プロセッサコア(146)、前記各複数のメモリユニット(0、・・・、n)、及び前記各複数のアクセラレータユニット(0、・・・、m)への接続のための複数のそれぞれのインターフェイスポートが含まれるデジタル信号プロセッサ。
- 7請求項6に記載のデジタル信号プロセッサであって、それぞれの各インターフェイスポートには、読み出し・書き込みポート対が含まれ、前記各読み出し・書き込みポート対には、読み出し要求信号、データ利用可能信号及び複数のデータラインが含まれるデジタル信号プロセッサ。
- 8請求項1乃至7のいずれかに記載のデジタル信号プロセッサであって、前記プロセッサコア(146)には、更に、ベクトル命令を実行するように構成された1つ以上の実行ユニットが含まれ、前記1つ以上の実行ユニットは、データが含まれるベクトルに作用するデジタル信号プロセッサ。
- 9請求項8に記載のデジタル信号プロセッサであって、前記実行ユニットには、1つ以上の命令実行パイプラインが含まれ、各々クロックサイクル当り単一の動作を実行するように構成されたデジタル信号プロセッサ。
- 10請求項9に記載のデジタル信号プロセッサであって、前記実行ユニットは、単一命令複数データ(SIMD)命令を実行するように構成されたデジタル信号プロセッサ。
- 11請求項9又は10に記載のデジタル信号プロセッサであって、前記各1つ以上の実行パイプラインは、前記ベクトルの異なるデータに対して同じ命令を実行するように構成されたデジタル信号プロセッサ。
- 12請求項9に記載のデジタル信号プロセッサであって、1つ以上の前記実行ユニットの1つ以上の前記実行パイプラインは、複素乗算累算器ユニットであるデジタル信号プロセッサ。
- 13請求項8に記載のデジタル信号プロセッサであって、1つ以上の前記実行ユニットは、実数部及び虚数部を有する複素数値化データに対して演算を行う複素ベクトル命令を実行するように構成された複素実行ユニットであるデジタル信号プロセッサ。
- 14請求項13に記載のデジタル信号プロセッサであって、前記複素実行ユニットは、本来あらゆるデータを複素数値化データとして解釈するように構成されたデジタル信号プロセッサ。
- 15請求項13に記載のデジタル信号プロセッサであって、前記複素実行ユニットには、1つ以上の命令実行パイプラインが含まれ、各々クロックサイクル当り単一の複素演算を実行するように構成されたデジタル信号プロセッサ。
- 16請求項15に記載のデジタル信号プロセッサであって、前記1つ以上の命令実行パイプラインの1つ又は複数には、前記複素ベクトル命令を実行するように構成された複素演算論理回路ユニットが含まれるデジタル信号プロセッサ。
- 17請求項1に記載のデジタル信号プロセッサであって、前記プロセッサコア(146)には、更に、実数部及び虚数部を有する複素数値化データに関する演算を実施するように構成された複素乗算累算器ユニットが含まれるデジタル信号プロセッサ。
- 18請求項1乃至17のいずれかに記載のデジタル信号プロセッサであって、前記1つ以上の専用機能のそれぞれ与えられた機能は、異なる無線通信標準規格に関連するデジタル信号プロセッサ。
- 19請求項1乃至18のいずれかに記載のデジタル信号プロセッサであって、前記各複数のメモリユニット(0、・・・、n)には、読み出し又は書き込みトランザクションの受信に応じて、ローカルメモリ位置に対応するアドレスを生成するように構成されたアドレス生成ユニットが含まれるデジタル信号プロセッサ。
- 20請求項1乃至19のいずれかに記載のデジタル信号プロセッサであって、前記各複数のメモリユニット(0、・・・、n)、前記複数のアクセラレータユニット(0、・・・、m)、前記プロセッサコア(146)、及び前記プログラム可能回路網(250)は、単一の集積回路上に製造されるデジタル信号プロセッサ。
- 21請求項1乃至20のいずれかに記載のデジタル信号プロセッサであって、前記複数のアクセラレータユニット(0、・・・、m)の少なくとも幾つかのアクセラレータユニットは、ベースバンド信号処理に関連する前記専用機能の構成可能ハードウェア実装品であるデジタル信号プロセッサ。
- 22請求項1乃至21のいずれかに記載のデジタル信号プロセッサであって、前記デジタル信号プロセッサは、無線用途用のプログラム可能なベースバンドプロセッサとして用いられるデジタル信号プロセッサ。
- 23請求項1乃至21のいずれかに記載のデジタル信号プロセッサであって、前記デジタル信号プロセッサは、媒体プロセッサとして用いられるデジタル信号プロセッサ。
- 24マルチモード無線通信装置(100)であって、 無線周波数信号を送受信するように構成された無線周波数フロントエンドユニット(130)と、 前記無線周波数フロントエンドユニット(130)に結合されたプログラム可能デジタル信号プロセッサと、を特徴とし、 前記プログラム可能デジタル信号プロセッサには、 複数のメモリユニット(0、・・・、n)と、 各々1つ以上の専用機能を実施するように構成された複数のアクセラレータユニット(0、・・・、m)と、 データ経路フロー制御に関連する命令を実行するように構成された実行ユニットを含むプロセッサコア(146)と、 前記命令の実行に応じて、前記複数のメモリユニット(0、・・・、n)と、前記複数のアクセラレータユニット(0、・・・、m)と、前記プロセッサコア(146)との間に選択的に接続性を提供するように構成されたプログラム可能回路網(250)と、 が含まれるマルチモード無線通信装置。
- 25請求項24に記載の無線通信装置であって、特定の命令の実行に応じて、前記プログラム可能回路網は、前記複数のメモリユニットの内の与えられた1つを前記複数のアクセラレータユニットの内の与えられた1つに結合するように構成される無線通信装置。
- 26請求項24に記載の無線通信装置であって、特定の命令の実行に応じて、前記プログラム可能回路網は、前記複数のメモリユニットの1つ以上のメモリユニットを前記プロセッサコアに結合するように構成される無線通信装置。
- 27請求項24に記載の無線通信装置であって、特定の命令の実行に応じて、前記プログラム可能回路網は、前記複数のアクセラレータユニットの2つ以上のアクセラレータユニットを共にチェーン状に結合するように、また、更に、前記チェーンの第1アクセラレータユニットを前記複数のメモリユニットの与えられた1つと前記プロセッサコアとの内の1つに結合するように構成される無線通信装置。
- 28請求項27に記載の無線通信装置であって、前記複数のアクセラレータユニットの各アクセラレータユニットは、前記チェーンの他のアクセラレータユニットに接続された場合、それが前記プロセッサコアによる介在なしで接続される前記アクセラレータユニットと通信を行うように構成される無線通信装置。
- 29請求項24に記載の無線通信装置であって、前記プログラム可能回路網には、前記プロセッサコアへの、前記各複数のメモリユニットへの、及び前記各複数のアクセラレータユニットへの接続のための複数のそれぞれのインターフェイスポートが含まれる無線通信装置。
- 30請求項29に記載の無線通信装置であって、それぞれの各インターフェイスポートには、読み出し・書き込みポート対が含まれ、前記各読み出し・書き込みポート対には、読み出し要求信号、データ利用可能信号及び複数のデータラインが含まれる無線通信装置。
- 31請求項24に記載の無線通信装置であって、前記プロセッサコアには、更に、実数部及び虚数部を有する複素数値化データに対して演算を行う複素ベクトル命令を実行するように構成された複素実行ユニットが含まれる無線通信装置。
- 32請求項31に記載の無線通信装置であって、前記複素実行ユニットには、各々クロックサイクル当り単一の複素演算を実行するように構成された複数の命令実行パイプラインが含まれる無線通信装置。
- 33請求項32に記載の無線通信装置であって、前記各複数の命令実行パイプラインには、前記複素ベクトル命令を実行するように構成された複素演算論理回路ユニットが含まれる無線通信装置。
- 34請求項32に記載の無線通信装置であって、前記複素実行ユニットは、単一命令複数データ(SIMD)命令を実行するように構成された無線通信装置。
- 35請求項24に記載の無線通信装置であって、前記プロセッサコアには、更に、実数部及び虚数部を有する複素数値化データに関する演算を実施するように構成された複素乗算累算器ユニットが含まれる無線通信装置。
- 36請求項35に記載の無線通信装置であって、前記複素乗算累算器ユニットは、本来あらゆるデータを複素数値化データとして解釈するように構成される無線通信装置。
- 37請求項24に記載の無線通信装置であって、前記プログラム可能デジタル信号プロセッサは、複数の無線通信標準規格によって確立されたパラメータ内において動作するように構成される無線通信装置。
- 38請求項24に記載の無線通信装置であって、前記複数のアクセラレータユニットの少なくとも幾つかのアクセラレータユニットは、複数の無線通信標準規格に準拠する信号のベースバンド信号処理に関連する前記専用機能の構成可能ハードウェア実装品である無線通信装置。
Independent claims38
55 paragraphs, as filed
The present invention relates to digital signal processors, especially programmable digital signal processors.
In a relatively short period of time, the use of wireless devices and especially mobile phones has increased dramatically. The worldwide spread of this wireless device has led to the emergence of numerous wireless standards and the convergence of wireless products. And this has increased interest in software defined radio (SDR).
As stated by the SDR Forum, SDR is "a collection of hardware and software technologies that enable reconfigurable system architectures for wireless networks and user terminals." SDR provides an efficient and relatively inexpensive solution to the problem of building multi-mode, multi-band, multi-function radio devices that can be improved and enhanced by software. Thus, SDR can be seen as a technology that grants special powers applicable in a wide range of areas of the wireless industry.
Many wireless communication devices use wireless transmitters and receivers that include one or more digital signal processors (DSPs). One type of DSP used in wireless communication equipment is a baseband processor (BBP), which can handle many of the signal processing functions associated with processing a received radio signal and preparing to transmit the signal. For example, BBP may provide modulation and demodulation, as well as channel coding and synchronization functions. Many traditional BBPs are implemented as application specific integrated circuit (ASIC) devices, which can support a single radio standard. In many cases, ASIC_BBP can provide excellent performance. However, ASIC issues may be limited to operating within the wireless standards for which on-chip hardware is designed.
SDR solutions require increased flexibility in wireless baseband processors to meet requirements for time-to-market, cost, and product life. Large-scale parallel processing is required for baseband processors to meet the requirements of demanding applications such as wireless local area networks (LAN), 3rd / 4th generation mobile telephone communications, and digital video broadcasting. To.
<p> To that end, various programmable BBP (PBBP) solutions based on extremely complex very long instruction word (VLIW) and / or multiple processor core machines have been proposed. These traditional PBBP solutions often have drawbacks such as large chip area and potentially limited performance compared to their ASIC equivalents. Therefore, it is desirable to have a programmable DSP architecture that can support a number of different modulation techniques, bandwidth and mobility requirements, and has acceptable area and power consumption.</p>
<p> Various embodiments of a programmable baseband digital signal processor, including a programmable network, are disclosed. In one embodiment, the digital signal processor includes a plurality of memory units, a plurality of accelerator units, and a processor core. The digital signal processor also includes a memory unit, an accelerator unit, and a processor. Includes a programmable network that can be configured to selectively provide connectivity to and from the saccore. Each accelerator unit may be configured to perform one or more dedicated functions independently of the processor core. The processor core may include an execution unit that may be configured to execute instructions related to data path flow control. The programmable network may be configured to selectively provide connectivity depending on the execution of instructions. According to the present invention, the processing capacity is improved, and this is achieved by maintaining flexibility.</p><p> In one particular embodiment, the programmable network is configured to combine a given one in a memory unit with a given one in an accelerator unit in response to the execution of a particular instruction. Can be done.</p><p> In other particular embodiments, the programmable network may be configured to couple one or more memory units to the processor core, depending on the execution of a particular instruction. In yet another particular embodiment, in response to the execution of a particular instruction, the programmable network is such that two or more accelerator units are chained together, and further, the first accelerator unit of the chain. Is configured to combine with one of the given one of the memory units and the processor core.</p><p> In another embodiment, the radio communication device includes a radio frequency front-end unit configured to transmit and receive radio frequency signals, and a programmable digital signal processor coupled to the radio frequency front-end unit. One such digital signal processor may be a baseband digital signal processor. A programmable digital signal processor includes a plurality of memory units, a plurality of accelerator units, and a processor core. The programmable digital signal processor further includes a programmable network that can be configured to selectively provide connectivity between the memory unit, the accelerator unit, and the processor core.</p><p> Each accelerator unit may be configured to perform one or more related dedicated functions independent of the processor core. The processor core may include an execution unit that may be configured to execute instructions related to data path flow control. The programmable network may be configured to selectively provide connectivity depending on the execution of instructions.</p>
Various modifications and other forms of the present invention may be permitted, the particular embodiment of which is shown in the drawings as an example and will be described in detail herein. However, the drawings and their detailed description are not intended to limit the invention to the particular embodiments disclosed, and conversely, all within the spirit and scope of the invention as defined by the appended claims. It should be understood that it seeks to cover the modifications, equivalents, and other options of. It should be noted that the headings are for structural purposes only and are not intended to be used to limit or interpret the description or claims. Moreover, the term "possible, may" is used throughout the application to mean acceptable (ie, possible, possible) rather than compulsory (ie, must). Please note that. The term "included" and its derivatives mean "include without limitation". The term "connected" means "directly or indirectly connected", and the term "combined" means "directly or indirectly connected".
Next, FIG. 1 shows a block diagram of an embodiment of a multimode wireless communication device including a programmable baseband processor. In the illustrated embodiment, some basic division of the wireless communication system is shown from the viewpoint of both function and hardware. In particular, the multi-mode wireless communication device 100 has a receiving sub. A system 110 and a transmission subsystem 120 are included, each of which is connected to an antenna 125. It should be noted that in various embodiments, the multimode wireless communication device may be a handheld mobile phone communication device or the like. Further, it should be noted that components having a reference identifier that includes both numbers and letters may be referred to by numbers only, as appropriate.
The receiving subsystem 110 includes a portion of an RF front end 130 that is connected to an analog-to-digital converter (ADC) 140. The ADC 140 is connected to a programmable baseband processor (PBBP) 145A, which is further coupled to the application processor (s) 150. Transmission subsystem 120 includes an application processor (s) 160 coupled to PBBP145B, which is connected to a digital-to-analog converter (DAC) 170. The DAC 170 is also coupled to a portion of the RF front end 130. Note that the PBBP145A and 145B can be implemented as one programmable processor and, in some embodiments, can be made on a single integrated circuit. Also note that in some embodiments the ADC140 can be implemented as part of the PBBP145A.
The PBBP145 performs many functions in both the transmit subsystem 120 and the receive subsystem 110. Within the transmit subsystem 120, the PBBP145B may convert data from the application source into a format suitable for the radio channel. For example, the transmission subsystem 120 may perform functions such as channel coding, digital modulation, and coding shaping. Channel coding means using various methods for error correction (eg, convolutional coding) and error detection (eg, using cyclic redundancy code (CRC)). Digital modulation means the process of mapping a bitstream to a stream of complex samples. The first (sometimes only) step in digital modulation is a specific signal assembly, such as two-phase phase shift keying (BPSK), four-phase phase shift keying (QPSK), or quadrature amplitude modulation (QAM). Is to map a group of bits to. There are various ways to map a group of bits to the amplitude and phase of a radio signal. In some cases, the second step domain conversion may be applied. Fast Fourier Transforms (IFFTs) can be used for this step in orthogonal frequency division multiplexing (OFDM) systems (ie, modulation methods in which information is sent at multiple adjacent frequencies at the same time). In a spread spectrum system such as Code Division Multiple Access (CDMA), for example (a "spread spectrum" method that allows a large number of users to share an RF spectrum by assigning a separate "code" to each in-use user). Each code is multiplied by a spread sequence of 1 and -1. The final step is code shaping, which converts a square wave into a band-limited signal using a finite impulse response (FIR) bandpass filter. Channel coding and mapping functions typically operate at the bit level (rather than at the word level) and are generally not suitable for implementation on programmable processors. However, as described in more detail below, in various embodiments of PBBP145, these and other functions are
The PBBP145 can perform functions such as synchronization, channel equalization, demodulation, and forward error correction. For example, the receive subsystem 110 recovers the code from a distorted analog baseband signal and also converts them to a bitstream with an acceptable bit error rate (BER) to turn the application into an application processor (one or one). Multiple) Can be run at 150.
Synchronization can be divided into several steps. The first step may include detecting an incoming signal or frame, sometimes referred to as "energy detection". In connection with this , Antenna selection and gain control, etc. can also be performed. The next step is code synchronization, which aims to find the exact timing of the incoming code. All of the above actions are usually based on compound self or cross-correlation.
In many cases, the receiving subsystem 110 needs to make some kind of correction for defects in the radio channel. This correction is known as channel equalization. In OFDM systems, channel equalization can involve simple scaling and rotation of each subcarrier after performing an FFT. In CDMA systems, "lake" receivers are often used to combine incoming signals from multiple signal paths with different path delays. In some systems, a least squares (LMS) adaptive filter may be used. As with synchronization, convolution-based algorithms can be used for most of the operations associated with channel estimation and equalization. These algorithms are generally not similar enough to share the same fixed hardware in traditional ASIC embodiments. Therefore, the ASIC solution is not fully adaptable. However, they can be efficiently implemented on programmable DSP processors such as the PBBP145.
Demodulation can be seen as the opposite of modulation. Demodulation usually involves performing a correlation or "despreading" of the FFT in an OFDM system with the spreading sequence in a CDMA system. The final step of demodulation may be to convert the composite code into bits according to the signal aggregate. Like channel coding, deinterleaving and channel decoding may not be suitable for firmware implementation. However, as will be described in more detail below, Viterbi or turbo decoding is an extremely excessive function that can be used for convolutional codes but can be realized as one or more hardware accelerators.
Programmable baseband processor architecture FIG. 2 shows a block diagram of an embodiment of the programmable baseband processor of FIG. By providing dynamic reconfigurability, the PBBP145 may support a variety of radio standards at multiple modes of operation (ie, preamble receive, payload receive, and transmit) and at different data rates. To achieve the desired reconfigurability, various embodiments of the PBBP145 use a programmable connectivity network between the processor core, numerous memory units, and various hardware accelerators. It may include a central processor core that manages the DSP flow by controlling the interconnect.
In FIG. 2, the PBBP 145 includes a processor core 146 and a plurality of data memory units indicated by 0 to n, where n may be any number. The PBBP145 further includes a plurality of hardware accelerators indicated by 0 to m, where m may be any number. Further, the PBBP 145 includes a programmable network 250 connected between the processor core 146 and each of the data memory and the accelerator. Further, the PBBP 145 contains the integer and coefficient memory units indicated by 220 and 215, respectively, which are connected to the processor core 146 via the programmable network 250. Finally, the PBBP 145 includes a medium access layer (MAC) interface unit 225 connected between the programmable network 250 and the host / MAC processor (not shown).
Processor core In the illustrated embodiment, the processor core 146 includes a control register CR265 and a control unit 260 connected to a programmable network 250. Processor core 146 further includes a complex multiplication accumulator (CMAC) unit 270 and a complex operation logical unit (CALU) 280, both of which are independently coupled to the programmable network 250. The processor core 146 also has a power connected to the CMAC 270. Includes a tor controller 275A and a vector controller 275B connected to the CALU280.
The control unit 260 includes an ALU261, a separate multiplication accumulator unit 262, and a set of register files (RF) 263. In one embodiment, the control unit 260 may function as a reduced instruction set controller (RISC) configured to execute integer instructions.
Each CALU280 contains a accumulator (not shown) and contains four ALUs, represented by 282A to 282D. CALU280 further includes a vector storage unit 283 and a vector reading unit 284. Note that in one embodiment, the vector storage unit 283 and the vector read unit 284 can be shared among the four ALUs, but the four ALUs can function to operate in parallel. Also note that in one embodiment, the vector controllers 275A and 275B can be realized as a single shared unit that can be shared between the CMAC 270 and the CALU 280.
CMAC270 can be optimized for operations on complex vectors. Therefore, the CMAC 270 contains a number of complex data paths that can be operated together or separately. In one embodiment, the data paths CMAC0 and CMAC1 may include two complex data paths, each containing a multiplier, an adder, and a cumulative register (all not shown). Therefore, CMAC270 can be referred to as a four-way CMAC data path. Also, in addition to multiplication and addition, CMAC0 and CMAC1 may perform rounding and scaling operations, respectively, to support saturation. In one embodiment, the CMAC270 operation can be split into three pipeline steps. In addition, CMAC0 and CMAC1 can each perform N-element vector operations in N / 2 clock cycles. In addition, CMAC0 and CMAC1 may support operations relating to complex values stored in accumulation registers (eg, complex addition, subtraction, conjugate, etc.). For example, CMAC270 calculates complex multiplication values such as (AR + jAI) * (BR + jBI) in one clock cycle, and also calculates complex cumulative values in one clock cycle, and performs complex vector operations (for example, complex convolution). , Conjugate complex convolution, and complex vector inner product).
In one embodiment, processor core 146 can function as a DSP processor with a large number of single instruction multiplex data (SIMD) execution units. In particular, the data paths are both grouped into SIMD clusters, where each cluster may use vector controllers 275A and 275B, vector storage unit 283, and vector read unit 284. The cluster can perform different tasks, while each data path in the cluster can execute a single instruction for multiplex data at each clock cycle. In particular, the four-way CALU280 and four-way CMAC270 can function, for example, as SIMD clusters to perform four parallel operations in parallel, such as four correlations, i.e., despreading four different codes. Similarly, the CMAC270 may implement, for example, two parallel radix 2FFT butterflies or one radix 4FFT butterfly.
Instruction set architecture In one embodiment, the instruction set architecture for processor core 146 may include three classes of compound instructions. The first class instructions are RISC instructions, which act on 16-bit integer operands. The RISC instruction class contains most control instructions and can be executed within control unit 260 of processor core 146. The next class of instructions are DSP instructions, which act on complex-valued data with real and imaginary parts. DSP instructions can be executed against one or more SIMD clusters. The third class instruction is a vector instruction. DSP instructions because vector instructions act on large datasets and can take advantage of advanced address modes and vector loop support. Can be regarded as an extension of. With few exceptions, vector instruction sets affect complex data types.
Many baseband reception algorithms can be broken down into task chains, but there are few backward dependencies between tasks. This property allows different tasks to be performed in parallel on the SIMD execution unit, as well as to take advantage of it using the instruction set architecture described above. Vector operations act on large vectors, so one instruction is issued every clock cycle, which can reduce the complexity of the control path. Moreover, since vector SIMD instructions act on long vectors, many RISC instructions can be executed during vector operations. Thus, in one embodiment, the processor core 146 may be a single instruction issuing machine per clock cycle, and each SIMD cluster and control unit may execute instructions for each clock cycle in a pipeline processing manner. .. Therefore, PBBP145 can be thought of as running two threads in parallel. The first thread contains the program flow and other processing using the control unit 260. The second thread contains complex vector operations performed on the SIMD cluster. FIG. 3 shows an instruction execution pipeline of one embodiment of the processor core of FIG.
With reference to FIGS. 2 and 3 collectively, the left column of FIG. 3 represents time (execution clock cycle unit). The remaining columns represent the execution pipelines of the complex SIMD cluster and control unit 260 (eg, CMAC270 and CALU280) and the issuance of instructions to them. In particular, in the first clock cycle, a complex vector instruction (eg, CVL.256) is issued to the CMAC 270. As shown, the vector instruction completes through many cycles. In the next clock cycle, a vector instruction is issued to CALU280. In the next clock cycle, an integer instruction is issued to the control unit 260. In the next few cycles, any number of integer instructions may be issued to the control unit 260 while the vector instructions are being executed.
More generally, a processor may include one or more execution units configured to execute vector instructions. That is, one or more execution units act on a vector containing data. An example of an execution unit for executing such vector instructions is CMAC, but other types of such execution units can also be used. An architecture based on the present invention may include one or more execution units for executing any known type of vector instruction in any combination. The execution unit may be configured to operate with a complex quantification vector, i.e., a complex vector instruction acting on complex quantification data having real and imaginary parts. Alternatively, the execution unit may be configured to operate on real numbers.
Note that in one embodiment, the control flow can be stopped until any vector operation is completed using the "idle" instruction to provide control flow synchronization and to control the data flow. I want to. For example, if a vector instruction is executed by the corresponding SIMD execution unit, the control unit 260 may execute an "idle" instruction. The "idle" instruction may stop the control unit 260, for example, until a display such as a flag is received by the control unit 260 from the corresponding SIMD execution unit.
Hardware accelerator As mentioned above, many baseband features may be provided by dedicated hardware accelerators used in combination with programmable cores to provide multimode support for a wide range of all wireless standards. The choice of which function to promote should be taken into consideration. As an example, features that are regularly implemented and used by some radio standards are good candidates for acceleration.
For example, in one embodiment, the following functions: decimeter / filter, 4 "finger" lake functions for CDMA and DSSS modulation schemes, 4FFT / modified Walsh transforms for OFDM modulation schemes and IEEE 802.11b, demappers, Convolution / turbo encoder / Viterbi decoder, configurable block interleaver, configurable scrambler, and CRC accelerator can be implemented using accelerators 0-m in FIG. Note that in other embodiments, other numbers and types of functions can be achieved using accelerators 0-m.
In one embodiment, the decimeter / filter accelerator may include configurable filters such as FIR filters that can be used for standards such as ADC and IEEE 802.11a. Similarly, a four-finger lake accelerator may include a accumulator unit and a simple complex multiplier capable of multiplying the sample values from the set {0 +/- 1 and 0 +/- i}. The lake accelerator may further include a local complex memory for delayed path storage, a despread code generator, and a matched filter (all not shown) capable of performing multiple path lookup and channel estimation functions. A radix 4 FFT / Modified Fourier Transform (FFT / MWT) accelerator may include a radix 4 butterfly (not shown) and a flexible address generator (not shown). In one embodiment, the FFT / MWT accelerator may perform a 64-point FFT in 54 clock cycles and also support the IEEE802.11b standard to perform a modified Discrete Fourier transform in 18 clock cycles. Convolution / Turbo Encoder Viterbi Decoder Accelerators may include reconfigurable Viterbi decoders and turbo encoders / decoders that provide support for convolution and turbo error correction codes. In one embodiment, the decoding of the convolution code can be performed by the Viterbi algorithm. On the other hand, the turbo code can be decoded by using the soft output Viterbi algorithm. A configurable block interleaver accelerator can be used to rearrange the data, spread adjacent data bits in time, and in the case of OFDM, spread between different frequencies. In addition, a scrambler accelerator can be used to scramble the data with pseudo-random data to ensure a uniform distribution of 1s and zeros in the transmitted data stream. The CRC accelerator may include a linear feedback shift register (not shown) or other algorithms for generating the CRC.
Memory unit Memory management and allocation can be important considerations for efficient use of the SIMD architecture of processor core 146. Thus, the data memory system architecture includes several relatively small data memory units (eg DM0-DMn). In one embodiment, the data memory DM0-DMn can be used to store complex data during processing. Each of these memories can be implemented to have two alternating memory banks, which allows two consecutive addresses (vector elements) to be accessed in parallel. Further, each data memory DM0-DMn may include an address generation unit (eg, 405A-405n shown in FIG. 4) that can be configured to perform modulo addressing as well as FFT addressing. As described further below, each DM0-DMn may independently and dynamically connect to any accelerator and processor core 146 via a programmable network 250. The coefficient memory 215 can be used to store the FFT and filter coefficients, look-up tables, and other data not processed by the accelerator. Integer memory 220 can be used as a packet buffer to store a bitstream for MAC interface 225. Both the coefficient memory 215 and the integer memory 220 are coupled to the processor core 146 via the programmable network 250.
Programmable network The programmable network 250 includes data paths, memory, accelerators and external inputs. It is configured to interconnect the surfaces. Therefore, the programmable network 250 is a crossbar in which connections are set up from one input (write) port to one output (read) port, and any input port is connected to any output port in an NxM configuration. Can behave in the same way as. However, in some embodiments, the connection between some memory and some arithmetic units may not be necessary. In this way, the programmable network 250 can simplify the programmable network 250 by optimizing so that only certain memory configurations are possible. The interconnection of programmable networks 250 and the like can reduce the complexity of network and accelerator interfaces while still enabling many simultaneous communications by eliminating the need for arbiters and addressing logic circuits. Note that in one embodiment, the programmable network 250 can be implemented using a multiplexing device or, for example, a combined logic circuit configuration such as an AND / OR configuration. The AND / OR configuration, when tested, had slightly smaller network hardware compared to the examples that included the multiplexing device.
In one embodiment, the programmable network 250 is implemented as two sub-networks. The first subnetwork may be used for sample-based transfers and the second subnetwork may be a series network used for bit-based transfers. Since bit-based transfer requires redundant framing and deframing of data chunks that are not equal to the data width of the network, splitting the two networks can improve the processing power of the network. In such an embodiment, each subnetwork can be realized as a separate crossbar switch composed of processor cores 146. The programmable network 250 may also be configured such that accelerators having related functions are directly connected to each other in a chain and data memories are connected to each other. This type of network configuration allows data to flow seamlessly between accelerator units without the intervention of processor core 146, which allows processor core 146 to circuit only during the creation and disappearance of network connections. It is possible to get involved in the network.
As mentioned above, not all memories need to be connected to all arithmetic elements, and the programmable network 250 can be optimized to allow only certain memory configurations. In those embodiments, the programmable network 250 may be referred to as a "partial network". To transfer data between these partial networks, several memory blocks within one or more data memory units (eg DM0) may be allocated to both subnetworks. These memory blocks can be used as ping-pong buffers between tasks. Useless memory movements can be avoided by "swap" memory blocks between arithmetic elements. This strategy can provide efficient and predictable data flow without wasted memory transfer operations.
FIG. 4 shows a further aspect of the programmable network embodiment of FIG. In the illustrated embodiment, each unit connected to the programmable network 250 (eg, processor core 146, DM0-n, accelerator 0-m, etc.) has an interface port with at least one read / write port pair. included. Each read / write port pair includes a "data incoming" and "data outgoing" signal and a "handshake incoming" and "handshake outgoing" signal. In one embodiment, the incoming / outgoing data signal may each be a multiple bit data path, while the incoming handshake signal may be a read request (RR) signal and the outgoing handshake signal may be data available. It may be a (DAV) signal. Similarly, the programmable network 250 includes a plurality of corresponding interface ports (eg, interface ports 0-n), each having the same port signal.
Further, the processor core 146 includes a network configuration port that can be used to send network configuration information to the programmable network 250. In one embodiment, processor core 1 The 46 may configure a network connection by using a dedicated assembly instruction or by writing a configuration vector to a control register such as the control register 265 in FIG.
Note that the processor core 146 can be implemented as a cluster SIMD architecture, so multiple data memories DM0-DMn can be connected to the processor core 146 at the same time. When the programmable network 250 is configured in this way, each data memory may be connected to its own SIMD cluster port.
Further, as shown in FIG. 5, the network interface ports of the programmable network 250 allow the connection of chained accelerators (eg, accelerators 0, 2, 3), which are between themselves. Automatically synchronizes and communicates with, and can operate independently of the processor core 146 (ie, without interaction with it) and without any kind of arbiter or network master unit. As mentioned above, this protocol allows simultaneous operation of processor core 146 and any number of accelerators without synchronization overhead within processor core 146, which frees processor core 146 and is a useful base. Band processing can be performed. In addition, the number of memory accesses can be reduced because intermediate storage may not be required when sending data between accelerators.
The programmable network 250 is configured to allow exclusive memory access to any unit (eg, processor core 146, accelerator 2, etc.) and store the algorithm output, thereby eliminating stall cycles due to access conflicts. There is a possibility. After completing the task, the reconfiguration of the programmable network 250 allows the accelerator or interface to "hang over" the entire memory, including the output data, thereby eliminating data movement between memories. Can be done.
6A and 6B are timing diagrams showing timing between units connected to an embodiment of programmable network 250. FIG. 6A shows typical timing of memory request, while FIG. 6B shows typical timing of data request for a unit (eg, accelerator) slower than the requester. The timing diagram of FIG. 6A includes a clock signal, a read request signal (RR), a data availability signal (DAV) and a data signal, while FIG. 6B contains an additional stall signal.
In one embodiment, the interface port logic circuits in the programmable network 250 and in each of the processor core 146, accelerator 0-m, and data memory 0-n may be configured to automatically synchronize between units. .. Thus, once the programmable network 250 is configured by processor core 146 to connect to two devices (eg, processor core 146 and DM0), the RR signal of the device requesting data is data. Cannot be idle as long as is available. In particular, if the transmitting unit is configured to provide data as fast as the requester can request the data, the RR signal cannot be idle as long as the requester needs the data. As shown in Figure 6A, the RR signal is asserted by the requester for 3 clock cycles. After the next clock cycle in which the RR is asserted, the DAV is asserted by the source for 3 clock cycles while the data is being sent. Therefore, three data "blocks" or units are sent in three clock cycles. Two cycles after the RR is deasserted, the requester reasserts the RR for two cycles. One cycle after the RR is asserted, the source asserts the DAV for two cycles and the data is sent in those two cycles.
However, some sources request as quickly as a requester can request data. Not all sources can provide data to ESTA. Thus, if there are three or more prominent read request cycles, the requester can be configured to stall the RR signal. For example, in Figure 6B, RR is asserted for two cycles and the source is not asserting DAV.
Therefore, the requester stalls the RR signal until the DAV is asserted for one cycle and data is sent. The RR signal is then asserted for only one cycle since only one data cycle was sent, and there are two prominent requests again. Therefore, when the requester requests data from a slow source, the requester alternately asserts and stalls the RR signal, allowing the source to catch up. However, it should be noted that in other embodiments, it is believed that the requester may stall after a prominent requirement of less than or greater than 2.
FIG. 7 shows a typical pipeline operation of an embodiment of the programmable baseband processor of FIGS. 2 and 4. During payload processing operations related to the IEEE802.11a standard, in one embodiment, the processing flow includes receiving and processing even and odd codes in three pipeline stages (ie, at any time). Data from three different OFDM codes (each containing 80 input samples) can be processed in different parts of the PBBP145.
In particular, during the odd code interval stage 1, the odd code is received by the ADC front end / filter and the sample is stored in DM0. During the odd code interval stage 2, the processor core 146 performs an operation on the even sample stored in DM1 and stores the result in DM3. During the odd code interval stage 3, the accelerator chain performs an independent operation on the result stored in DM2 and transfers the result to the MAC layer interface. Similarly, during even code interval stage 1, the even code is received by the ADC front end / filter and the sample is stored in DM1. During the even-numbered stage 2, the processor core 146 performs an operation on the odd-numbered samples stored in DM0 and stores the result in DM2. During the even-numbered stage 3, the accelerator chain performs an independent operation on the result stored in DM3 and transfers the result to the MAC layer interface.
As mentioned above, the programmable network 250 can be dynamically reconfigured during operation to facilitate data flow between accelerators, data memory, and processor cores 146. In addition, flow control can be provided between running processes using idle instructions, interrupts, and flags.
Referring collectively to FIGS. 2 and 7, in the first pipeline stage, processor core 146 constitutes programmable network 250 and decimeter / filter accelerator (eg, accelerator 0) in data memory (eg, eg accelerator 0). Connect to DM0) (block 600). Upon receipt of the odd OFDM code, accelerator 0 performs decimation and frequency offset correction, and the corrected sample may be transmitted via DM0 to programmable network 250. Upon receiving the full code, in one embodiment, the sample counter function in accelerator 0 (not shown) may generate an interrupt for processor core 146 (block 605). In response to an interrupt, processor core 146 may reconfigure programmable network 250 to connect accelerator 0 to a different data memory (eg DM1) so that a sample of the next sign is transferred to DM1. .. At about the same time, DM0 and other data memory (eg DM2) can be connected to processor core 146 (block 610). Accelerator 0 can then receive an even sign, perform an operation on it, and write the corresponding sample to DM1 (block 615).
In the second pipeline stage, the processor core 146 can perform an operation on the sample stored in DM0 in response to the interrupt generated by accelerator 0. In one embodiment, processor core 146 may perform FFT and channel correction for the codes now available in DM0 and also perform some phase and channel tracking tasks (block 635). The corrected frequency domain sample can be transferred from processor core 146 to DM2 via programmable network 250 (block 640).
When processor core 146 has finished sending the results to DM2, processor core 146 has a programmable network by connecting additional chained accelerators (eg, accelerators 1-4) and DM2 to the inputs of the chain. 250 can be reconstructed (block 645). For example, the memory-to-accelerator chain contains a DM2 connection to the demapper, the demapper is connected to the deinterleaver, the deinterleaver is connected to the Viterbi decoder, and the Viterbi decoder connects to the MAC layer interface. obtain. At about the same time, when accelerator 0 finishes sending even samples to DM1, processor core 146 causes programmable network 250 to reconnect accelerator 0 to DM0 (block 620), thus causing accelerator 0 to have the next odd sign. Can be processed and the sample stored in DM0 (block 625).
When accelerator 0 finishes for the next odd sign, an interrupt is generated in block 605, as described above. In response to an interrupt, processor core 146 may reconfigure programmable network 250 to connect accelerator 0 to DM1 and processor core 146 to DM0 and DM2 (block 630). Note that processor core 146 can be idle. For example, processor core 146 can wait for accelerator 0 to finish storing samples in one of its data memories.
The third pipeline stage contains an accelerator chain that operates independently of the operation performed by processor core 146 for the results stored in DM2. For example, the accelerator chain may perform demapping and channel decoding operations and transfer the resulting bitstream to the MAC layer interface (block 660). When the accelerator chain finishes the operation on the DM2 data, the accelerator chain can generate an interrupt to the processor core 146. Once the results are prepared within DM3, processor core 146 may reconfigure the programmable network 250 to connect DM3 to the inputs of the accelerator chain (block 665). The accelerator chain performs operations on the results stored in DM3 and forwards the resulting bitstream to the MAC layer interface (block 670).
Due to the flexible nature of the architectures and microarchitectures described above, the PBBP145 may provide support for many wireless standards and for many modes in those standards.
Furthermore, the present invention is highly suited for implementing programmable baseband processors for wireless applications. However, the present invention can be used as a media communication processor, i.e., a processor specifically designed for the generation and distribution of digital media. Therefore, it can be used to build multimedia subsystems that process any combination of audio, video, graphics, fax and modem operation.
The embodiments have been described in considerable detail, but once the disclosure is fully recognized, a number of changes and amendments will be apparent to those skilled in the art. The following claims are to be construed as including all such changes and amendments.
<figref num="1">FIG. 6 is a block diagram illustrating an embodiment of a multimode wireless communication device including a programmable baseband processor.</figref><figref num="2">The block diagram which shows one Embodiment of the programmable baseband processor of FIG.</figref><figref num="3">The figure which shows the instruction issuing pipeline of one Embodiment of the processor core of FIG.</figref><figref num="4">FIG. 2 illustrates a further aspect of the programmable baseband processor embodiment of FIG.</figref><figref num="5">The figure which shows the typical network connection in one embodiment of the programmable baseband processor of FIG. 2 and FIG.</figref><figref num="6A">A timing diagram showing a typical timing mode between units connected to one embodiment of the programmable network of FIGS. 2 and 4.</figref><figref num="6B">FIG. 4 is a timing diagram showing another typical timing mode between units connected to one embodiment of the programmable network of FIGS. 2 and 4.</figref><figref num="7">The figure which shows the pipeline flow which described the typical operation of the embodiment of the programmable baseband processor of FIG. 1, FIG. 2, and FIG.</figref>
8 sheets
Sheet 1 Sheet 2 Sheet 3 Sheet 4 Sheet 5 Sheet 6 Sheet 7 Sheet 8
Every citation, both waysCites: the store holds 4 of 5
| Document | Relation | Office | Cited during |
|---|---|---|---|
| KR20170067716A | Cited by | Republic of Korea | Search report |
| US11768689B2 | Cited by | United States of America | Applicant |
| JP2005508532A | Cites | Japan | – |
| JP2000284970A | Cites | Japan | – |
| JP2000513523A | Cites | Japan | – |
| US20020186043A1 | Cites | United States of America | – |
| Eric Tell, Anders Nilsson and Dake Liu,“A Low Area and Low Power Programmable Baseband Processor Architecture”,Proceedings of the 9th International Database Engineering & Application Symposium (IDEAS'05),[online],2005年 7月,p.347-351,[retrieved on 2012-01-25]. Retrieved from the Internet,URL,<http://ieeexplore.ieee.org/stamp/stamp.jsp?tp=&arnumber=1530969> | Non-patent | – | – |
| Magnus Karlsson, Mark Vesterbacka, and Wlodek Kulesza,“A METHOD FOR INCREASING THE THROUGHPUT OF FIXED CEFFICIENT DIGIT-SERIAL/PARALLEL MULTIPLIERS”,Proceedings of the 2004 International Symposium on Circuits and Systems (ISCAS'04),[online],2004年 5月,Vol.2,p.425-428,[retrieved on 2012-01-31]. Retrieved from the Internet,URL,<http://ieeexplore.ieee.org/stamp/stamp.jsp?tp=&arnumber=1329299> | Non-patent | – | – |
| BAHMAN BARAZESH, JEAN-CLAUDE MICHALINA, and ANDRE PICCO,“A VLSI Signal Processor with Complex Arithmetic Capability”,IEEE TRANSACTIONS ON CIRCUITS AND SYSTEMS,[online],1988年 5月,VOL.35, NO.5,p.495-505,[retrieved on 2012-01-31]. Retrieved from the Internet,URL,<http://ieeexplore.ieee.org/stamp/stamp.jsp?tp=&arnumber=1776> | Non-patent | – | – |
| Eric Tell, Anders Nilsson, Dake Liu,“A Programmable DSP core for Baseband Processing”,The 3rd International IEEE-NEWCAS Conference,[online],2005年 6月,p.403-406,[retrieved on 2012-01-25]. Retrieved from the Internet,URL,<http://ieeexplore.ieee.org/stamp/stamp.jsp?tp=&arnumber=1496739> | Non-patent | – | – |
23 members in 6 offices
Priority claims9
| Document | Office | Kind | Date |
|---|---|---|---|
| 11135964 | United States of America | – | |
| 13596405 | United States of America | A | |
| 13596405 | United States of America | A | |
| 2006000602 | Sweden | W | |
| 2006000602 | Sweden | W | |
| 2005135964 | – | – | – |
| 2006000602 | – | – | – |
| US20050135964 | – | – | – |
| WO2006SE00602 | – | – | – |
Members23
| Document | Office | Kind | |
|---|---|---|---|
| US2006271764A1 | United States of America | A1 | |
| US2006271765A1 | United States of America | A1 | |
| WO2006126943A1 | World Intellectual Property Organization (WIPO) | A1 | |
| WO2006126943A1 | World Intellectual Property Organization (WIPO) | A1 | |
| WO2007018468A1 | World Intellectual Property Organization (WIPO) | A1 | |
| US7299342B2 | United States of America | B2 | |
| WO2007018468A8 | World Intellectual Property Organization (WIPO) | A8 | |
| KR20080034095A | Republic of Korea | A | |
| EP1913487A1 | European Patent Office (EPO) | A1 | |
| EP1913488A1 | European Patent Office (EPO) | A1 | |
| KR20080042837A | Republic of Korea | A | |
| CN101203846A | China | A | |
| CN101238455A | China | A | |
| US7415595B2 | United States of America | B2 | |
| JP2008546072A | Japan | A | |
| JP2009505215A | Japan | A | |
| EP1913487A4 | European Patent Office (EPO) | A4 | |
| JP5000641B2This record | Japan | B2 | |
| JP5080469B2 | Japan | B2 | |
| CN101203846B | China | B | |
| KR101256851B1 | Republic of Korea | B1 | |
| KR101394573B1 | Republic of Korea | B1 | |
| EP1913487B1 | European Patent Office (EPO) | B1 |
16 legal events, as the office reported them to INPADOC
Over the term
Point at a mark for the eventEvents
| Event | Code | |
|---|---|---|
| Cancellation because of no payment of annual feesLAPS | LAPS | |
| Receipt of annual feesJAPANESE INTERMEDIATE CODE: R250R250 | R250 | |
| Receipt of annual feesJAPANESE INTERMEDIATE CODE: R250R250 | R250 | |
| Receipt of annual feesJAPANESE INTERMEDIATE CODE: R250R250 | R250 | |
| Receipt of annual feesJAPANESE INTERMEDIATE CODE: R250R250 | R250 | |
| Receipt of annual feesJAPANESE INTERMEDIATE CODE: R250R250 | R250 | |
| Renewal fee payment (event date is renewal date of database)FPAY | FPAY | |
| Certificate of patent or registration of utility modelJAPANESE INTERMEDIATE CODE: R150R150 | R150 | |
| Certificate of patent or registration of utility modelJAPANESE INTERMEDIATE CODE: R150R150 | R150 | |
| First payment of annual fees (during grant procedure)JAPANESE INTERMEDIATE CODE: A61A61 | A61 | |
| Written decision to grant a patent or to grant a registration (utility model)JAPANESE INTERMEDIATE CODE: A01A01 | A01 | |
| Written decision to grant a patent or to grant a registration (utility model)JAPANESE INTERMEDIATE CODE: A01A01 | A01 | |
| Decision of grant or rejection writtenTRDD | TRDD | |
| Notification of reasons for refusalJAPANESE INTERMEDIATE CODE: A131A131 | A131 | |
| Notification of resignation of power of attorneyJAPANESE INTERMEDIATE CODE: A7424RD04 | RD04 | |
| Written request for application examinationJAPANESE INTERMEDIATE CODE: A621A621 | A621 |
Numbers
- Publication
- 5000641
- Publication, DOCDB
- 5000641
- Publication, EPODOC
- JP5000641B
- Application
- 2008513415
- Application, DOCDB
- 2008513415
- Application, EPODOC
- JP20080513415
Titles2
- Japanese
- プログラム可能回路網を含むデジタル信号プロセッサ
- English
- Digital signal processor with programmable network
Classification
- CPC, 10
- G06F9/30036
- G06F15/7857
- G06F9/30079
- G06F9/3851
- G06F9/3885
- G06F9/3891
- H04B1/0003
- Y02D10/00
- G06F9/38
- G06F9/3888
- IPC, 4
- G06F9 38
- G06F17 16
- H03K19 173
- H04L29 06
