Integrated circuit
Abstract
This record has no abstract on file.
Term
Term ended
Expired 2 April 2017, 9.5 years ago.
- Priority
- Filed
- Granted
- Expired
- Today
9 claims: 1 independent, 8 dependent
- 1A computing system, a main central processing unit (CPU) microprocessor, a digital signal processing unit (DSP) microprocessor having an instruction set different from that of the main CPU microprocessor, the main CPU microprocessor and the DSP microprocessor. Combined storage device、 A file-based operating system in the storage device in which the DSP performs operations on the main CPU during a time interval in which the main CPU is otherwise occupied, thereby enhancing the performance of the computing system.Arranged toThe computing system can run at least the software portion of a software application, and the file-based operating system determines where in virtual memory such software portion starts and ends.ShowhandleSaid file-based operating system including, Let the DSP microprocessor perform the main CPU microprocessor functionIncludes programDSP kernel softwareWearWithTo identify where the sender and receiver handles areThe main CPU microprocessorRun the file-based operating system、A program that controls the main CPU microprocessor and sends information from the position indicated by the sender handle to the position indicated by the receiver handle.The file-based operating systemIncluding、Where in the storage device to cause the DSP microprocessor to perform a function on behalf of the main CPU microprocessorThe sender handle and the receiver handleWorks based on existenceThe DSP microprocessorThe program that controlsThe DSP kernel softwareIncluding Computational system. 計算システムであって、 主中央処理装置(CPU)マイクロプロセッサ、 前記主CPUマイクロプロセッサとは異なる命令集合を有するディジタル信号処理装置(DSP)マイクロプロセッサ、 前記主CPUマイクロプロセッサと前記DSPマイクロプロセッサとに結合された記憶装置、 前記記憶装置内のファイルベースオペレーティングシステムであって、前記主CPUが他に係わって占有されている時間間隔中前記DSPが前記主CPUの動作を実行し、それによって前記計算システムの性能を高めるように配置され、前記計算システムはソフトウェアアプリケーションの少なくともソフトウェア部分を実行することができ、前記ファイルベースオペレーティングシステムは仮想メモリ内の何処でそのようなソフトウェア部分が開始し終了するかを示すハンドルを含む前記ファイルベースオペレーティングシステム、 前記DSPマイクロプロセッサに主CPUマイクロプロセッサ機能を行わさせるプログラムを含むDSPカーネルソフトウェアを備え、何処に送り手ハンドルと受け手ハンドルとが存在するかを特定するために前記主CPUマイクロプロセッサが前記ファイルベースオペレーティングシステムを実行し、前記主CPUマイクロプロセッサを制御して送り手ハンドルによって示される位置から受け手ハンドルによって示される位置へ情報を送るプログラムを前記ファイルベースオペレーティングシステムが含み、前記DSPマイクロプロセッサに前記主CPUマイクロプロセッサの代わりに機能を実行させるために前記記憶装置内の何処に前記送り手ハンドルと前記受け手ハンドルとが存在するかに基づいて動作する前記DSPマイクロプロセッサを制御するプログラムを前記DSPカーネルソフトウェアが含む 計算システム。
260 paragraphs, as filed
The present invention relates to generally improved personal computer (hereinafter referred to as PC) circuits, computer systems, and operating methods thereof.
[0002] [Problems to be Solved by the Invention] Early computers required a large space and occupied the entire room. From time to time, minicomputers and desktop computers have been on the market.
[0003] Popular desktop computers include "Apple " (Motorola 680x0 microprocessor-based) and "IBM -Compatible Machines" (Intel or other x86 microprocessor-based). ) And so on, these are known as PCs and have become very popular for office and home use. There have also been introduced high-end desktop computers called workstations based on some superscalar and other very high performance microprocessors such as SuperSPARC .
[0004] Further developed, notebook dimensions or palmtop computers are optionally battery-powered for mobile user applications. Such notebooks and small computers are challenging technologies that have the conflicting goals of miniaturization, constantly faster, higher performance, more flexibility and longer battery life. Further, a desktop sealed container called a docking station has a portable computer fitted in the docking station, and improvement of such a portable computer / docking station system is desired. However, in all these systems, the choice of central processing unit (hereinafter referred to as the CPU) generally determines the processing power of the system, and an add-in-card, that is, an embedded card, is added to the CPU. It is CPU-centric in the sense that it adds specific applications or functions such as modems and multimedia. Improvements in circuits, integrated circuit devices, and all types of computer systems, and above all, how to tackle the challenges just mentioned, are desired to be described herein.
[0005] Overall, and in one form of the invention, the PC system is a main CPU microprocessor, a file-based operating system, and a digital signal processor (hereinafter referred to as DSP). ) Includes microprocessors, which are arranged to allow the DSP to perform main CPU operations during the time interval in which the main CPU is otherwise occupied, thereby increasing the bandwidth of this PC system. This PC system can also include a large number of CPUs and / or a large number of DSPs.
[0006] In other forms of the invention, combined video / imagery systems including DSP microprocessors, video / audio control logic circuits, and compression / decompression electronics are both DSP microprocessors, main CPU microprocessors, The memory management circuit coupled to both of these microprocessors, and the memory circuit connecting the memory circuit to the memory management path and the local bus, so that the memory circuit is the DSP, the main CPU, the video / audio control logic circuit, And acts as a unified memory architecture for compression / decompression electronics, and the DSP performs processing functions for both video / audio control logic and compression / decompression electronics. In addition, software that virtualizes parts of the video / image function may be added to the system.
[0007] The present invention provides a system for comprehensively enhancing the software of a PC.
[0008] The present invention provides a system that improves performance by collectively handling add-in features, that is, embedded features, through an add-in card.
[0009] The present invention provides a methodology for integrating features within core logic for the realization of a motherboard.
[0010] The present invention integrates all the functions of the system on the CPU chip.
[0011] The present invention has at least one piece of application hardware reduced to the main CPU, DSP, and substantially only the physical layer, so that DSP relates to the signal mediated by this physical layer. It provides a system that virtualizes and executes the rest of the application. This system may employ a large number of applications and layers. The DSP virtualizes, for example, a local area network (hereinafter referred to as LAN), a video controller, an image compression / decompression electronic circuit, a fax machine, and a modem.
[0012] The present invention includes a DSP core, a master / slave bus interface, a memory circuit including a first-in first-out (hereinafter referred to as FIFO) function coupled to this interface, and an integrated circuit having a RAM function coupled to the DSP core. provide.
[0013] The present invention includes an integrated DSP core, a master / slave bus interface, and a memory circuit including a FIFO function coupled to the interface, and a single instruction / multiplex data control circuit for coupling the memory to the DSP core. Provides a circuit.
The present invention includes a DSP core, an interface circuit including a master / slave bus interface and a translation circuit, and a memory circuit including a FIFO function coupled to the master-slave bus interface and a RAM function coupled to the translation circuit. Provided is an integrated circuit in which the DSP core is coupled to the memory circuit and an interface.
[0015] The present invention includes a first bus interface circuit, a display controller circuit having a display interface, and a bus mastering circuit having a second bus interface circuit in which the display controller circuit is coupled to the second bus interface circuit. Provided is a video controller integrated circuit including.
[0016] The present invention provides the following software system. That is, a software system having an operating system, at least one multimedia driver, an x86 object code application, a non-x86 DSP code application that virtualizes a hardware application, and DSP kernel software, which kernel software is the operating system and ( Or) perform real-time interrupts and / or direct memory access (hereinafter referred to as DMA) virtualization and / or multithreaded multitasking operations within the DSP core in connection with multimedia real-time events. To perform memory transaction functions and / or input / output (hereinafter referred to as I / O) transaction functions that can be operated on the DSP core and / or are otherwise performed by x86. Let me.
[0017] In the present invention, a processing element, an interconnection circuit connected to the processing element, a memory electronic circuit connected to the processing element (or the interconnection circuit), and a multiplexing bus connected to the interconnection circuit. , And a computing system having a virtualization circuit connected to the multiplexing bus to perform a preselected application function.
[0018] The present invention relates to a first-multiplexed bus interface circuit, a second-multiplexed bus interface circuit, and at least one signal processing element connected to the first-multiplexed bus interface circuit and the second-multiplexed bus interface circuit. Provides a virtualization circuit that performs a preselected application function having.
Other improved PC devices, systems, and methods of operation thereof are also claimed within the scope of the claims.
The present invention is understood by reference to the following detailed description in connection with the accompanying drawings.
[0021] In these accompanying drawings, parts having the same function are indicated by the same reference numerals unless otherwise specified.
[Embodiments of the Invention] With reference to FIG. 27 first, a block diagram of the improved PC100 of the present invention is shown. In FIG. 27, the microprocessor unit (hereinafter referred to as MPU) block 2702 is a 486 (or P5) CPU, or any other type of x86 CPU, a dynamic RAM (hereinafter referred to as DRAM) controller circuit. And peripheral device interconnection (hereinafter referred to as PCI) including a bridge circuit. The PC100 has multimedia application capabilities.
The local CPU bus 2706 connects the MPU block 2702 to the DRAM 2714. This DRAM preferably has a capacity of at least 4M. However, it is clear that more or less memory is adopted in this way. Local Bus 2706 is designed to keep CPUs independent so that different CPUs can be used on the same bus.
[0024] The PCI bus 2710 is connected to the CPU in the block 2702 via the PCI bridge in the CPU block 2702. In this way, the PCI2710 bus provides a wide bandwidth bus and also provides connectors for a wide variety of peripheral devices that would otherwise potentially unfavorably load the local bus 2706.
The peripheral processing unit (hereinafter referred to as PPU) 2718 acts as a system bridge for connecting the PCI bus 2710 to the ISA / AT bus 2734. In addition, a PCMCIA card or a PCMCIA bridge 2726 for the PCMCIA card bus connects to the PCI bus. The network bridge 2730 connected to the PCI bus may not be as complex as a LAN bridge, or may be a wide area network (hereinafter referred to as WAN) bridge, a radio frequency (hereinafter referred to as RF) bridge, or a non-simultaneous transmission mode. It may be a bridge (hereinafter referred to as ATM) or a comprehensive service digital network (hereinafter referred to as ISDN) bridge. Each of these types of bridges produces a significant amount of traffic directly on the PCI bus and indirectly on the local bus 2706. The PC100 has a keyboard and mouse (not shown in FIG. 27) connected to it for user input, and a CPU output user observation display or CRT (not shown in FIG. 27).
[0026] In FIG. 27, buses 2706, 2710, and 2734 are not interconnected. Connected to the ISA / AT bus 2734, for example, games 2739, hard disk drive 2742, printer 2746, fax data modem 2750, telephone proxy answerer (also known as DTAD) 2754, professional audio block 2758, and CD for multimedia PCs. (Compact Disk) A collection of optional accessory peripheral devices such as, but not limited to, device 2762 is displayed. Each of these ISA blocks is an accessory peripheral device that performs part-time functions, and many (marked with an asterisk) include a DSP.
[0027] In general, the other blocks connected to the PCI bus 2710 in FIG. 27 cost about the same as all peripheral devices connected to the ISA / AT bus 2734. For example, the total system cost of a professional audio system 2758 should be less than $ 10, or about 0.5% of the total PC system cost.
[0028] Cost reduction roadmaps, future sophistication, and compatibility with existing (heritage) software goals are important goals for any improved PC system. The system embodiments of the present invention confirm and demonstrate that improvements at the system level can achieve these goals. In particular, the system embodiments ensure that none of them included in the PC motherboard preferably occupy only a low cost, software-free, and essentially negligible motherboard area.
It can be said that backward compatibility with DOS and Windows 3.11 is provided by the "virtual" hardware of the present invention. Virtual hardware means that all the features shown in Figure 27 are fixed and dedicated, and even if these features are distributed, they are not updatable. The software can help virtualize them. If one chip achieves all of the functions shown in Figure 27, then this chip would be highly programmable. Therefore, some of the various embodiments of the invention are sufficiently programmable to "virtualize" all of the functionality shown in FIG. 27 by combining all the redundant and incompatible hardware shown in FIG. 27. And finally provide a single chip that performs these functions.
[0030] Windows is a recent and future operating system for PCs (hereinafter referred to as OS) and dynamically links to libraries, so that the software does not have to be compiled and linked before the execution time. For example, dynamic link loading (hereinafter referred to as DLL) is not linked until the run time and provides virtual software. This reduces the software dimensions.
In general, traditional CPU systems use fixed CPU hardware to implement a variety of applications, but the CPU is involved in the AT bus or all I / O ports shown in Figure 27. I can't stand. An embodiment of the invention described later herein allows the hardware to do just that.
For multimedia applications, the DMA2719 and interrupt controller 2720 in the system bridge 2718 have fixed hardware and fixed functionality. In one embodiment of the invention, DMA and interrupt control are virtualized. Two key challenges in servicing multimedia are needed for the ability to serve wideband and real-time interrupts for multimedia data. These challenges can create obstacles and the possibility of interruptions that introduce constraints or limitations in the ability to service real-time events. All of this is overcome in the embodiments of the present invention by virtualizing the hardware, for example, making the hardware programmable to take on a number of personalities. When hardware takes on the personality of an interrupt handler, it is a part of the CPU or an otherwise real-time capable CPU and OS extension, not a coprocessor or ancillary equipment.
[0033] The virtual hardware has mobility in the CPU and the OS at the same time. Since no method is known to physically achieve this, virtualization is realized according to the teachings of the present invention. Programmable DSP solutions and cores serve as the basis for virtual hardware, which is customized and improved as described herein.
The DSP core is a highly advantageous basis for running the various virtualizations required for the embodiments of the present invention. However, DSPs utilize memory peripheral devices and local DSP buses, which introduces additional challenges. Using the teachings of the present invention, the opposite DSPs are made to "look" like a x86 CPU, like a Pentium (R), or vice versa.
Software links are used to tightly couple DSPs and their peripherals to x86 CPUs, and to do this, use popular operating systems, such as, but not limited to, Windows 95. To do. This software link connects the two more tightly than if the DSP and its peripherals were on the same chip. Therefore, everything is made into software, the line between hardware and software is not clear, and it achieves some significant advantage. The end result is a single chip with software that can be remotely advanced by telephone line or by transfer of control information.
[0036] Figure 19a shows the dominant layered architecture within Windows 3.11. Application block 1902 is a multimedia application, such as a voice application. Ordinary blocks under block 1902 are transparent to independent software vendors (ISVs). Windows applications use a client / server model, in which case the application is a client, and the server handles whatever the request is. The issue is often found by the correct server. Block 1906 is a first layer MMSYSTEM (multimedia system) under block 1902 that communicates with multimedia hardware and plays audio .WAV files. Windows.WAV driver block 1910 is the second or first driver layer in this system that handles requests from application 1902 and does not involve compression. Play the signal in .WAV format. The lowest layer in the CPU environment is the DSP driver 1914, which is a virtual device driver that virtualizes the hardware in the hardware conformance layer (hereinafter referred to as HAL). The Audio Compression Manager (AVM) driver 1918 provides compression / decompression functionality within the system and communicates with block 1914. This completes the image mapping from the CPU2702 side.
It provides a way within the Windows architecture to plug into the client / server architecture. Therefore, DSP servers or audio compression servers can be plugged in through Windows, all of which have the same mechanism. Figure 19a shows how to plug any desired architecture into Windows. It is also the key to backwards compatibility and establishes a good (bland) new non-intrusive composition for the future.
[0038] Next, in Figure 19b, the virtual hardware environment 1922 is a TI TMS320C5x DSP with special DSP kernel software that is compatible with up to a third-party installation software base already available (eg, modems, audio equipment, etc.). However, it is realized without limitation. Everything runs on a preemptive basis, and priority is calculated in real time and executed dynamically. However, something must supply the OS so that it can calculate priority in real time. Some of that is real-time kernel software 1922. It complements the Windows OS, and is not incompatible with it, but works to extend it. This provides a multi-threaded and multi-tasking system.
Next, while the audio converter block 1926 performs audio compression / decompression on the DSP side, the ACM driver block 1918 can also perform the same function on the CPU side. However, if the CPU is occupied and does not have time to carry out this mechanism, the DSP may handle it. If the CPU and DSP are free, either or both of these can perform this function due to the ability of the software to obscure the line between the CPU and DSP. Therefore, all, as a favorable result through the OS, the CPU can embody the DSP, and the DSP can embody the CPU. Traditionally, CPUs perform data movements because they can perform movements in 32- or 64-bit chunks. However, in the present invention, an algorithmically intensive block in the application call to the DSP does this. In this way, it can be said that the application is accelerated.
The DSP codec driver 1930 couples to the stereo codec 1934. Preferably, for this system embodiment, one chip is appropriately used with the external stereo codec 1934 for the DSP to take into account codec updates.
FIG. 1 shows a block diagram of the improved PC system 100 of the present invention. In FIG. 1, the PC system 100 includes a CPU 102, which is connected to the cache 104 and the host bridge 108. The host bridge 108 is connected to the main memory bus 106 (sometimes referred to as the CPU bus or local bus), thereby connecting to the main memory 112 and also to the PCI bus 116. The PC system 100 typically has a keypad and / or mouse (not shown) connected to it for user input, and a CPU output user observation display or CRT (shown in Figure 1). Has not been). The virtual DSP circuit 200 is connected to the PCI bus 116. The P1394 / USB (upper wave band) block 120, which provides a low to high speed serial capture port for system 100, is connected to the PCI bus 116. The I / O block 124 is shown in FIG. 1 and is connected to the PCI bus 116. This I / O block 124 may be an ATM / LAN, ISDN or RF link. The DSP block 200 optionally connects to all these various I / O systems.
[0042] The DSP 200 is capable of "virtualizing" the hardware required to perform the above-mentioned functions or applications of the block 124, and thus may include or replace the I / O block 124. it can.
[0043] In FIG. 1, past modems are virtualized by providing suitable modem software in the DSP block 200. Other similar applications that may be virtualized are speakerphones, speech, DSVD (digital simultaneous voice and data modem), T.120 transport layer, VDSVD (video digital simultaneous voice and data). The "oval block" in Figure 1 shows some of these applications.
Other virtualizable applications include vector quantization for video compression / decompression (hereinafter referred to as VQ), MPEG for video compression / decompression, ISDN-based chamber conference H.320, ATN and LAN-based H. There are .321, H.322, and telephone line video conferencing.
Still other virtualizable applications include 3D graphic display applications and 3D audio for transport directional audio.
Many such applications include commercial audio, games, hard disk devices, printers, fax data modems, telephone proxy answerers (DTADs), multimedia CD devices, and data / file compression, the latter of which Performed by DSP to reduce system traffic and CPU load, there are also format conversions, and digital filters and digital conversions.
[0047] Fig. 1 includes the P1394 / USB block 120, because the domestic market supplies PCs as ostensibly household appliances, but it is not always in operation. For example, a PC can be easily switched on and off, and the user does not interface to it. Two cables may be connected to this PC, one for USB and one for P1394, and the RF interface may be connected as well. When a user wants a telephone surrogate answerer, they have a low-cost telephone, that is, the telephone has a USB cable that extends from now on to the USB jack of the PC, and also has a wire to a regular wall-mounted telephone jack. If the PC is in another room in the house, RF is used or the house is wired to P1394.
[0048] To borrow costs from a cellular phone, remove the hardware from the design of the cellular phone and move it to a PC. Therefore, the cellular phone is no longer an isolated system.
Therefore, for each home application, the hardware is removed from the application as is currently known and transferred to a PC that virtualizes this hardware.
[0050] The RF hardware of the I / O block 124 is also virtualized so that it is only its physical layer. Therefore, PC systems are truly multifunctional and multitasking.
FIG. 2 shows a more detailed block diagram of the virtual DSP block 200 of the improved PC system 100 of the present invention. In FIG. 2, block 200 is connected to PCI bus 116 and ISA bus 128. In another embodiment, block 200 is still connected to local bus 106.
[0052] In FIG. 2, the PCI bus 116 is connected to the hardware interface electronic circuit 210, this interface is connected to the first layer, this layer is then connected to the interface electronic circuit 214, and this circuit is then 2 Combined with two Texas Instruments TMS320C5x DSP cores 218 and 222. The electronic circuit 210 includes a PCI slave interface for speech transcoding (True speech Slave), a PCI slave interface for Windows-based modem data pump (WinModem Slave), and a PCI superbus master interface for DMA scatter-gather. Includes sound blasters (R, Creative Labs) similar to I / O ports, and a DMA interface. The electronic circuit 210 and the PCI bus 116 also establish a PCI / ISA bridge for multimedia.
The heart of block 200 is one or more DSP cores 218, 222. The electronic circuit 214 is a FIFO RAM, that is, the electronic circuit 214 is not a hardware FIFO but a main memory 112 that performs a FIFO function, that is, a piece of RAM. The RAM 112 looks like a regular RAM to the DSP cores 218 and 222, while the RAM 112 looks like a FIFO to the PCI bus side of this interface. Many translations are done in electronic circuit 214. Data and operands flow between CPU 102 and block 200. CPU102 does not operate in bytes or bits. Even if CPU 102 operates in units of 32 or 64 bits, it operates through bursts through cache 104 and main memory 112 based on cache lines. This data stream incompatibility consumes many cycles in translation.
Interface hardware 214 is designed to take a 32- or 64-bit wide data path, extract any portion of it into bytes, and then supply 16-bit words to the DSP core. Therefore, a complex instruction set computer (hereinafter referred to as CISC) or a reduced instruction set computer (hereinafter referred to as RISC) for a single instruction multiple data (hereinafter referred to as SIMD) architecture is realized to solve an interface challenge. ..
Therefore, in an advantageous manner, the interface electronic circuit 214 performs the following functions. That is, 1) Inject a 64- or 128-bit wide burst as a FIFO compatible burst. 2) Translate them into a preselected DSP format. 3) Acts as RAM for the DSP core.
Since different clocks run on different chips, the clocking challenge is solved within block 210 by decoupling the DSP core operation from the PCI operation. From the CPU side, this system is well synchronized with PCI operation. From the DSP side, this system is asynchronous with PCI operation. Avoiding a large number of standby states and increasing the DSP operating speed in this way does not lead to a severe grinding halt when interfacing with the CPU.
DMA transactions are preferably based on flow-I / O. The CPU runs in virtual space, while the DMA works in physical space. The translation is preferably performed by the DSP under software control using Windows 95, without relying on old or new hardware in CPU 102. The DSP will be the main DMA engine. In another embodiment, the DMA engine in CPU 102 can be decoupled from DSP 218, and DSP core 218 and DSP core 222 are associated with the DMA engine. The DSP assembles and executes the DMA in the way CPU 102 or Windows desires. For example, a DMA processing capacity of 100 Mbyte / s can be easily achieved using the DMA core of the current technology.
[0058] The block 210 itself is a PCI agent, and the cost is reduced in the sense that it is a single master, multi-slave RCI agent in hardware except when it is in operation. The bus control capability of block 210 causes all slaves to be driven in operation, but has the ability to do so when called to control the PCI bus by an application. The software calls the slave to be the master with the appropriate configuration data. This multiplex interface works when the CPU polls if it is not occupied elsewhere. If bus mastering is required, but the CPU is otherwise occupied, let the DSP become the bus master through the master in block 210 to achieve the same. Therefore, for the same application, the DSP works as a host instead of the CPU 102 itself.
[0059] The software architecture of Windows requires that it look like a file transaction under Windows. File transactions are memory transactions and it does not matter whether the CPU or DSP carries them out. The CPU can only do one thing at a time, so the DSP does the work during system downtime and fills the downtime, so no hardware is needed. This ability to cherry pick and work during pauses is an important advantage of this research study that cannot be accomplished by the CPU. Video chips, DSPs, CPUs, etc., all of these chips are controlled by the same constraint that every transaction must be a memory transaction or a file transaction (or an I / O transaction that is less important for this purpose). Will be done. Everything is treated equally under the OS.
[0060] FIG. 3 is a more detailed electrical circuit diagram (partially schematic, partially block diagram) of a preferred embodiment of the improved computer system portion of FIG. FIG. 3 schematically shows the hardware interface electronic circuit 210 and the interface electronic circuit 214 of FIG. In particular, it can be seen that the PCI master / slave interface circuit 304 is connected to the PCI bus 116. The PCI master / slave interface circuit 304 includes electronic circuits for master operation and slave operation, and is connected to a hardware layer 305, which is a PCI configuration control and state register electronic circuit 306, a PCI I / O spatial register electronic circuit. Includes 308, dual port read / write FIFO electronic circuit 310, and DSP I / O spatial register electronic circuit 312. Hardware layer 305 is then connected to interface / codec DMA control circuit 316, the latter to DSP (not shown in FIG. 3).
FIG. 4a is a more detailed electrical circuit diagram (partially schematic, partially block diagram) of a preferred embodiment of the improved computer system portion of FIG. In Figure 4a, the PCI bus 116 is connected to a single chip 420 on the motherboard or add-in card via a gateway to master / slave interface 424. FIG. 4a also shows a second chip 460 connected to the PCI bus 116. The two chips 420, 460 show two different ways of classifying the various functionality according to the teachings of the present invention.
The chip 420 has an on-chip acceleration bus 343 and is preferably a desktop PC. General-purpose buses GPI401 and GPI402 are provided for either chips 420 or 460 (not shown in Figure 4a). Within chip 420, IDSP logic block 428 contains at least all the logic of block 200 of FIG. Similarly, the IDSP logic block 428 of chip 420 contains at least all the logic of block 200 of FIG. The chip 460 includes two zoom video (hereinafter referred to as ZV) buses, and is preferably used in a portable PC. The graphics / video controller 432 has video capture / compression / decompression capabilities as well as 2D and 3D graphics capabilities. Controller 432 is required to be a bus master, but by convention it is always a slave. The proposed new unified memory architecture (called UMA) favors bus masters.
[0063] Typical or conventional memory cycles are too long, so the memory must also change. Memory cannot be a commercial product in terms of functionality. So the memory needs additional functionality and it has to be interfaced with the new block. The present invention takes advantage of the memory having this additional functionality.
[0064] In FIG. 4b, the PCI bus 116 is connected to the DSP / video chip 420 in FIG. 4a to form a PC system. Chip 420 is connected to memory controller 484, which has other couplings with CPU 102. The memory controllers 484 and CPU 102 are further coupled to and control the data buffer 488, which accesses the UMA memory on the local bus 106. By using the data buffer 488, the CPU must also request memory access rather than controlling it. Advantageously, the functionality of the video / graphics block 432 is integrated within the PC system. Video / graphics capabilities are advantageous because the chip 420 is probably pad bound (ie, the circuit has many pins and does not occupy all of the silicon needed to fit all the required pads). It exists in space. All of the various applications listed here will be brought to the IDSP Acceleration Bus.
[0065] The UMA block 492 includes an extended data output DRAM (Extended Data Output DRAM, hereinafter referred to as EDO) for the low-end market. The EDO has an invariant pin count, supports a 486 level CPU, and its data output remains active longer to avoid precharging. Another suitable UMA is a fast EDO with burst capability to fill the cache line. High level UMA is a multi-bank burst EDO. Skilled workers select the memory type according to the performance and market price points desired for a particular application system.
[0066] From UMA block 492, one embodiment has an uncompressed output with a bandwidth of up to 11 Gbyte / s on bus 106. It can be said that this is achieved by either a broadband bus or a high-speed bus such as a RAM bus operating above 500 MHz. In Figure 4b, controller 432 accesses UMA492 and sends uncompressed video to display 436 (not shown in Figure 4b). Access to memory over wideband memory and PCI bus 116 maintains wideband and short latency (eg 2ms).
The acceleration bus 434 of FIG. 4a is connected to the DSP 428 for external or internal access to take into account the acceleration of controller 432. Similarly, Figure 4b also includes an acceleration bus for external or internal access.
FIG. 5 shows a block diagram of an embodiment of an improved computer system of the present invention for host-dependent asymmetric multiprocessing. In Figure 5, the P5 processor is the CPU. In Figure 5, any CPU / IDSP intelligence in the PC can be centralized through the virtual technology of the invention in the amount needed to support the application and disbanded after the application is stopped. can do. The user only needs to have a given amount of computational resources required for a given application at any given time.
[0069] A large number of IDSPs that meet the needs of any particular application are shown in Figure 5. These chips are designed differently to properly interface to the ISA and PCI buses. Some examples are shown in Figure 5. IDSP blocks are shown attached to these buses, but they may be attached to either the north bridge or the south bridge, and may be integrated within the function of that block. Good.
[0070] FIG. 6a shows a schematic block diagram of an embodiment of a CPU model that uses a superscalar extension for an improved computer system of the present invention. In Figure 6b, the new chip has three pipelines, two conventional superscalar pipelines and a third DSP operation pipeline. That is, the superscalar CISC / RISC and DSP CPU architectures have three operations dispatched on a single chip. There is microcode memory in CISC. The DSP core is a DSP hardware micromemory, which dispatches DSP operations. The DSP core can be "empty" until the user clicks on the Windows icon and then the DSP code is cached from the hard disk to main memory and local memory for execution. In this way, the x86 CPU does not have to perform non-standard operations and dispatch DSP operations. Such a combination architecture is compatible with Windows 95.
FIG. 6b shows one such memory cache hierarchy. Disk storage 2742 is at one pole of the spectrum defined in Figure 6b. Next is the main memory 112. The external single access memory is memory 214 anywhere else in the ISA bus, any other memory on the PCI bus, or in the chip 200 or in the third level PC system of Figure 6b. The second level has single access on-chip memory, such as in configurable DSP core add-on memory. At the first level, a dual access on-chip DSP core memory with B0, B1, B2 memory of Texas Instruments TMS320Cx core is provided, by way of example, but not limited to.
[0072] This is a software cache including a block 210 that operates as a cache, a CPU that operates as a cache, and an OS. Therefore, the DSP builds a cache operation cycle in software, and control goes above the hierarchy.
As an example, one V.34 modem uses 64K of code and data spacing as before, as modem speed negotiation requires this modem to speak to any other modem. .. During the 75ms period, this modem knows what the other modem is talking about, and then switches all the way through only a piece of code to maintain the specific application that corresponds to that modem sought as a result of negotiations. Put in. In this way, less code and data is required compared to the prior art.
With a transfer rate of 132 Mbyte / s, many applications can easily move the required code from memory to the DSP local memory that executes it.
FIG. 7 shows how a shared memory model is used to combine a DSP and a CPU. A shared memory modem based on the Windows architecture tightly couples the DSP and CPU asymmetrically on any symmetric software architecture beneath it.
[0076] In FIG. 7, Windows has an architecture that calls a file drawer-like handle that tells where in virtual memory the software starts and ends. These handles provide a mechanism for indicating the location of memory resources within the virtual memory space. These handles allow the DSP to be done by the host CPU. The host states that the application needs to manipulate certain memory contents and where the sender handle is and where the destination handle is. Content-based Windows sends information to the location defined by the recipient handle. Knowing where 1) the sender handle is and 2) where the receiver handle is, the DSP goes into operation and handles the transfer on behalf of the CPU. According to the application, the OS virtual memory manager tells the CPU where the physical address is. The CPU also has on-chip hardware that assists in determining the physical address. The DSP asks the OS's virtual memory manager as a superbus master (for example, this master can intersect page boundaries unlike a simple bus master). Unlike CPUs that treat handles as virtual addresses, DSPs lock handles to make them real physical addresses while the application is running. A DSP lock is a utility written to the main DSP driver in the HAL layer that activates an existing utility in the OS virtual memory manager to restore the physical address.
FIG. 8 is a schematic diagram of an embodiment of a multimedia extension model for an improved computer system of the present invention. In particular, it can be seen that the real-time service provided by the IDSP of the present invention also provides real-time priority to the multithreaded and multitasking Windows OS interrupt scheduler. Figures 32 and 33, which will be discussed later, show details of how to handle interrupts to give them real-time properties. FIG. 35 shows the details of a real-time service.
FIG. 9 is a schematic diagram of an embodiment of a system / cache / virtual memory model for an improved computer system of the present invention. A DSP core is a cache that moves data and instructions back and forth very efficiently. The CPU moves the data from the hard disk to this cache. The improved cache operation method for a PC of the present invention will be described.
[0079] In FIG. 9, the cord is taken out from the hard disk device (also referred to as HDD) to the main memory 112 during the run time. During the run time, a piece of code is retrieved from main memory 112 into block 214 in FIG. Block 210 further moves the code to the DSP memory 918 and C5x memory of FIG. 9, independent of the DSP and CPU. The cache operation method is shown in Fig. 6b. In FIG. 2, the block 210 performs a part of the cache operation method. In Figure 19b, a piece of code that corresponds to CPU downtime is paged through a virtual space in the CPU that is now extended to include memory on the DSP side in block 300. In this way, the OS is used to translate virtual memory into physical memory. In FIG. 9, DARAM is double access static RAM and SARAM is single access static RAM.
[0080] FIG. 10 is a schematic block diagram of an example of an MPEG reproduction filter graph model that may be used in the improved computer system of the present invention. When the user clicks on a software object in Figure 10, Windows dynamically allocates memory and runs object linking embedding (hereafter referred to as OLE). MPEG works from object to object expression instead of layer to layer expression. The sender may include image capture data for stretching.
FIG. 11 is a schematic block diagram of an embodiment of a virtual I / O hardware-PCI DMA and multimedia real-time interrupt handler model for an improved computer system of the present invention. In Figure 11, the system blocks interrupts in real time. The DSP space is an external static RAM (hereinafter referred to as SRAM) that resides in the main memory 112 for actual purposes. Therefore, the sender space, the receiver space, and the DSP space are all included in the same space. In this way, you dynamically obtain resources that would otherwise not be available to your Windows system. The heritage code runs in favor of Windows under system 100 compatiblely. The application (also known as APP) uses parts of the main memory and effectively decomposes these into the indicated blocks of the PCI bus scatter-collection DMA controller.
FIG. 12 is a schematic block diagram of parallel processing by a CPU and IDSP in a frame for an improved computer system of the present invention. In Figure 12, given a frame, the host CPU must perform all processing and monopolize a large number of slots within time zone 1204. Advantageously, the combination of DSP and CPU provides two time zones 1210 and 1212, where the DSP processes frames within the time zone 1210 and the Discrete Cosine Transform (hereinafter referred to as DCT) is the CPU time. There are two short time slots in band 1212.
[0083] Therefore, a relatively medium-performance DSP has sufficient resources to perform the main image processing, because this DSP is dedicated to its signal processing task. LANs, modems, and other applications are time-divisioned with image processing, but during other time intervals unrelated to the analysis points in FIG.
FIG. 13 is a simplified block diagram of an MPEG encoder 1300 that may be employed by the improved computer system of the present invention. In particular, the incoming video image is reordered within block 1301 and then fed to motion estimation block 1303 to determine which area of the image is changing. The output from the motion estimation block 1303 is supplied to the image / storage prediction block 1305, the output multiplexer 1307, and the adder 1309. The output from the image / storage prediction block 1305 is supplied to the adder 1311 and the adder 1309. The adder 1309 supplies the input to the DCT block 1313, which feeds the output to the quantization block 1315. The quantization block 1315 supplies the output to the variable length encoder 1317 and the inverse quantization block 1319. The inverse quantization block 1319 supplies the output to the inverse DCT block 1321 and the latter supplies the output to the adder 1311. The variable length encoder block 1317 supplies output to the multiplexer 1307, which supplies output to buffer 1323, which outputs encoded video data. Such encoders are implemented in hardware or software.
FIG. 14 is a simplified block diagram of the MPEG Decoder 1400 that may be employed by the improved computer system of the present invention. The incoming encoded video is stored in the input buffer 1401, which outputs to the demultiplexer 1403, which outputs to the image storage / prediction block 1405 and the variable length decoder 1409. The image storage / prediction block 1405 outputs to the adder 1407. The variable length decoder 1409 outputs to the inverse quantization block 1411, which supplies the quantization block steps from the demultiplexer 1403. The inverse quantization block 1411 outputs to the inverse DCT block 1413, and the latter outputs to the adder 1407. The adder 1407 outputs to the image storage / prediction block 1405 and the image reorder block 1415, which outputs the decoded video image.
[0086] FIG. 15 is a simplified block diagram of an embodiment of video resolution for the notebook computer of the present invention, which may employ a P5 processor as the CPU. In Figure 15, the system works with the PCMCIA standard, Amendment 2.1, as shown in Figure 4a.
The ZV fits video input resources contained within the framebuffer without the need for an extra framebuffer. PCI is not isochronous, bursting and interrupting, and is therefore extremely disadvantageous for video. According to the present invention, isochronous capability is added by a dedicated backdoor private highway bus ZV connected to the frame buffer. The acceleration bus links the blocks together within the chip 420 or 460 of FIG.
FIG. 16 is a simplified block diagram of an embodiment of video resolution for the desktop computer of the present invention, which may employ a P5 processor as the CPU. FIG. 16 shows a desktop PC that does not use PCMCIA or a card bus. Therefore, while playing MPEG video, the CPU performs the video decoding function, while the DSP performs the audio decoding function, or vice versa. If the data comes from a CD ROM, system synchronization is performed by the CPU. If the data comes from an external camera 1394 or other external image capture system, the DSP will perform system synchronization because this data first arrives at the DSP. In this way, the system and method advantageously solves the question of which processor should perform system synchronization.
FIG. 17 is a simplified block diagram of an application pipeline for an improved computer system of the present invention. In Figure 17, a single processor, even a superscalar processor, cannot be pipelined at the application level. This is distinctly different from the pipeline processor hardware pipeline. What this means is that you run different parts of the pipeline at the same time to pipeline your application.
The stage is established by the nature of the application. For example, the DSP stretches N frames of video, while the x86 CPU simultaneously outputs processor data for frame N-1 to the screen, and then repeats this cycle in a pipeline-related manner. The DSP also performs filtering, scaling, and color conversion, while the CPU inputs or outputs data.
FIG. 18 is a block diagram showing the system level of another embodiment of the hardware architecture for the improved computer system according to the present invention. FIG. 18 shows an alternative arrangement of the simplified block diagram of FIG. FIG. 18 shows a PDC that uses the PCI bus and has slot 1810 for the plug-in PCI card 1820. The plug-in PCI card 1820 includes a PCI interface (I / F) 1822, DSP 1824, and associated codec 1826, as shown in FIG. 3 or 4a. All others are standard PC hardware. Hard disk drive 1830 is part of standard PC hardware. Codec 1826 handles audio functions.
The PCI card 1820 is both a PCI master and a PCI slave. The host CPU can access the PCI configuration register and the main memory pointer register in slave mode. In master mode, the DSP can retrieve program code and data using scattered-collected DMA and store it in main memory.
FIG. 19b is a more detailed electrical circuit diagram (partially schematic, partially block diagram) of a preferred embodiment of the improved computer system of FIG. FIG. 19b is similar to FIG. 3, but further shows details of DSP1950 and related memories 1952, 1954. In addition, its connection to the codec chip 1956, as well as the chip 300 (shown in Figure 3), the speaker 1962, and the microphone 1960 is shown. 38 to 44 show details of one arrangement in this embodiment.
[0094] The DSP1950 runs independently of the CPU. When the PCI card is a slave, the CPU downloads the initialization code to the DSP after reset, where the DSP becomes independent. The CPU can start the DSP by setting the START bit in the PCI I / O space register 308. The START bit is an interrupt to the DSP, causing the DSP to start code execution. The DSP independently executes as a PCI master until the code execution is completed. When the DSP algorithm is finished, the DONE bit is set in the PCI Status Word 308 and an interrupt to the CPU is generated on the PCI bus.
[0095] The CPU supplies a pointer address to the main memory. The three required addresses are the base address for the DSP's 128K bytes for the program space, the address for the 128K bytes for the read space (sender), and the address for the 128K bytes for the write space (receiver). Depending on the application, they point to different areas of memory, or point to the same area of memory. The CPU controls the DSP by writing to register 308, which has a START bit for the DSP. The DSP uses this bit as an interrupt. This interrupt either initiates the algorithm loaded on the DSP or indicates that the host has sent a command to the DSP.
The other bit in register 308 is the DSP reset bit. This bit is sent to the reset pin of the DSP. This bit aborts the current DSP task and starts the DSP in bootload order. To load the new task into the DSP, you can either reset the DSP (which takes additional time to bootload) or set the MASTER ABORT bit to the host-controlled register 308. When the MASTER ABORT bit is set, an interrupt is generated on the DSP, which causes the originally loaded boot code to be executed. This boot code loads the DSP task from the host main memory currently pointed out by the RCI I / O spatial memory program spatial pointer. The DSP indicates that the task has finished by setting the EVENT bit in the PCI I / O space CMD / STATUS register 308 and raises a PCI interrupt (the bit is set again before the task is completed). , Indicates that the DSP is working).
The PCI application card DSP architecture is initially memory-free. This attracts the attention of system manufacturers who need to keep costs as low as possible. However, this architecture does not prevent the addition of external program space memory or external data space memory, as shown in memory 1952, 1954. If DSP application software can provide the time and overhead of sequential 8-bit access, one 8-bit wide external RAM may be used instead of the 16-bit 2-RAM chip memory system to achieve data memory. In order to achieve a memoryless system, the internal memory of the DSP will be divided into a program space and a data space. If the program is too large to fit in internal memory, the code must be retrieved from main memory onto the PCI bus as needed. The raw data is fetched from the main memory, and the data resulting from the DSP operation is returned to the main memory.
[0098] In the absence of external memory, the only way to make the code executable in the DSP is to reset the DSP and use a ROM-based bootloader to make a small initialization program from main memory. Is to load. The bootloader version of DSP (BDSP) is used for the following description. This DSP reads the reset global memory address FFFF to determine what type of bootloading will occur. The value of 1100 in the four least significant bits indicates a 16-bit parallel I / O load that uses a handshaking XF signal and a basic I / O (BIO) signal. When the BIO is low, the data bus is driven. After the boot load is complete, the BIO must remain high to prevent bus collisions on the DSP data bus.
By loading 16-bit words via I / O boot mode, the CPU can quickly load the basic main memory initialization program into the DSP. The DSP is in a reset state, so it can only respond to PCI bus transactions as a slave. The PCI host will control the initialization program values loaded in the DSP. This allows the initialization program to be modified by software. After this routine is loaded, the DSP is unreset so that the DSP can execute the initialization program (also the BIO selection bit to allow the DSP's BIO signal to be driven by the FIFO state signal. Must be set). The initialization program acts as a PCI bus master and enters the main memory to actually retrieve the DSP application software. By bootloading only the initialization program, the DSP PCI bus master can actually control the loading of DSP application code.
If the DSP application software cannot fit into the DSP internal memory as a whole, more code needs to be swapped in. This code is retrieved from main memory. To retrieve a lot of code, a DSP application may execute a BLDP (Block Load Program to Data) instruction, which retrieves Global spatial data and moves it into program space. If global data (Global If the Data) space is specified to be external memory and the program space is specified to be internal memory, the BDSP will provide the external access needed to move the code. As DSPs are fetching the global data space to the outside world, their application software needs a mechanism to tell the PCI interface that the DSP code is being fetched on the bus and not data. The I / O port register bits will be used to indicate whether the data being retrieved is program space code, sender data, receiver data, or DSP-determined data. The I / O port register bit is set to access the reset program space.
The first boot-loaded code brought into the application code and used is preferably maintained in memory. This aborts DSP operation by forcing the CPU to interrupt to reload this application. If this code is lost, the CPU loses control of the DSP and must reset the DSP to reload the program code. DSP application software can transfer data from external memory to internal memory and restore it with the BLDD (Block Load Data to Data) command.
[0102] The PCI interface is both a PCI master and a PCI slave. This interface needs to be a slave so that the CPU can access the configuration space. The CPU can access either the memory space or the I / O space to perform some control over the DSP and assemble the main memory pointer. The DSP will be the PCI master to access the main memory for code and data. There is also a FIFO 310 included in the PCI interface. This FIFO preferably holds 64 bytes of data. Since the BDSP has a 16-bit external data bus, the FIFO between the PCI bus and the BDSP bus must convert the DSP program and data from 32 bits to 16 bits. By packing two instructions or two data words in each main memory location, the number of PCI buses required for a given application can be reduced by half.
The configuration registers required by the PCI standard are used to specify areas in the PCI I / O space and memory address space with 0s and 1s. As long as the PCI space is required, it is five double words. Since the address space should be 2 and is required, 8 addresses must be secured. Registers above these five required double words are not used and will return 0 if read. These five double words are used for program space pointers, data space pointers, DSP commands, DSP state words, and I / O ports for bootloading BDSPs.
[0104] The PCI host is PCI. Write the I / O spatial register 308. The DSP is controlled by these registers. Base addresses 0 and 1 in the configuration register point to the area of PCI space to which this chip responds. Both base addresses point to the same space, but base address 0 is configured as an I / O space. Base address 1 is configured as a memory space. This provides greater flexibility for system designers designing either memory-based systems or I / O-based systems. 20 bytes are required to communicate with and control the DSP. 32 bytes are reserved in the PCI configuration register. Registers 0x64 to 0x1F will be allocated and only 0 will be returned. Their addresses range from base address + 0x0 to base address + 0x13. These registers not only allow the CPU to determine which portion of main memory should be used for DSP memory, but also allow the state of the DSP to be started, stopped, and monitored. The lower 12 bits of the main memory program (and data) space register are reserved. This relocates the program and data space within the 4K boundaries.
[0105] When the DSP algorithm is finished, the EVENT INTERRUPT bit in the state word is set. When this bit is set, an interrupt is generated to the application software through the PCI bus. Since this interrupt is controlled by DSP software, interrupt enable is not implemented in hardware.
[0106] DSP hardware control terms directly control the operation of DSP hardware. The least significant 4 bits are the retry counter bits. The retry counter bit causes the PCI macro function to retry the PCI transaction if it initially fails. This value is reset to 0000, which will cause the above features to retry indefinitely. This value can be varied to be between 1 and 15 retries. If the retry counter does not run because the transaction is not executed, the Retry Counter Expired bit in the MISC CTRL register is set. This state is available to the host.
The DSP interface ASIC comprises a PCI interface, a PCI FIFO, a DSP (connected device) interface, and a CODEC interface. This research requires that the BDSP be a discrete chip.
The DSP has 5 user interrupts available. One of these interrupts is unmaskable and four are maskable. The CODEC interface requires 3 interrupts and the PCI interface register requires 2 interrupts and a RESET pin. These interrupts are used as follows: That is, NMI- Master abort (from PCI host to DSP), INT1-CDRQ (from codec), INT2-PDQR (from codec), INT3-IRQ / IRQ2 (from codec 1 and codec 2), and INT4-command (from PCI host to DSP) ). The command interrupt is the lowest priority, but this is not a problem as there is sufficient housekeeping between exiting the RESET pin and responding to the DSP algorithm. The master censor interrupt is the highest priority and is used to cause the DSP to reload its initial DSP code. Interrupts 1, 2, and 3 are generated by the codec. If a second codec is required by the user, IRQ interrupts can be shared. The IRQ signal and the IRQ2 signal are gated together to generate one interrupt to the DSP. The IRQ signal, CDRQ signal, and PDRQ signal are inverted by the ASIC. The IRQ2 signal supplied by the user must be low and active.
The DSP accesses a memory-mapped I / O port, which is used to control PCI macros for operation mastering.
The CODEC register is used by the DSP to communicate with the discrete CODREC on the PCI application board. The FIFO state is available to the DSP on the I / O port. Depending on this ASIC or future revisions of the ASIC, the FIFO may have different dimensions. Knowing what size FIFO the DSP must work with to tune the performance will help the DSP. The FIFO dimension register displays the maximum number of words that a FIFO can hold. The number of DSP words to transfer is written to other I / O ports. This value is doubled in hardware to get the number of bytes to transfer. The maximum number of words that can be transferred is 32768 (65536 bytes). This means that the most significant bit of register 0x59 cannot be used. If 1 is written to that bit, it is ignored.
The PCI address offset (0x58) is an address in DSP space that is transferred into PCI space and used to transfer data to or from the host. The DSP accesses the host main memory by writing the DSP pointer to the PCI adder pointer register, which picks up 2 DSP words and creates a 32-bit PCI pointer. This is useful if the DSP needs to calculate the location in main memory that has the database in the scatter-collection table. The PCI address space selection bit in the PCI control register determines whether to use the DSP pointer or PCI pointer written in the PCII / O space when a PCI bus transaction is started.
The codec state / mode register (0x5F) is readable for all 16 bits. Only its lower 8 bits are writable (the upper 8 bits are only read), BIO selection allows flexible selection of FIFO states or user signals connected to the DSP BIO input. All of the state signals, and therefore BIO, are high and active.
[0113] DSP is PCI Performs reads and writes from or to the PCI bus through a 64-byte FIFO on the ASIC. The PCI FIFO is configured as a 2-parallel 16-word FIFO that synchronizes with the PCI clock. The DSP control signal is synchronized with the DSP's CLKOUT1 signal. There are two options for clocking the processor. The first option is to divide the PCI clock by 2 and use the DSP's CLKIN2 input to run the DSP internally at the PCI clock speed (up to 33MHz) (this is to double the CLKIN2 signal by 2). become). As a result, the PCICLK signal has the same frequency as the CLKOUT1 signal. However, the CLKOUT1 signal can be phase-shifted up to about 270 degrees. Delays through clock-divided electronic circuits and phase shifts through DSP phase-locked loops (hereinafter referred to as PLLs) can cause a total delay of up to 25 ns. The second option is to divide the PCICLK signal by 1 and use the DSP's 1 divided input option (DSP does not support this mode). The end result would be that the PCICLK signal is still at the same frequency as the CLKOUT1 signal, but in different phases.
The DSP's RDz (read permission--low and active) and WEz (write permission--low and active) signals are used to generate FIFO read and write pulses. Reads by the DSP must be performed with at least 3 software standbys to ensure that clock delays are taken into account.
[0115] DSP writes must be performed with at least two software standbys to ensure that clock delays are taken into account.
[0116] The DSP manages the data coming from and going to the FIFO. This data stream cannot be transparent to the DSP unless it overruns and underuns the FIFO [the DSP can also outrun the PCI bus and overrun the FIFO]. To get the DSP to use its resources best, it is supposed to use a BIO signal as a flag to tell the DSP when the FIFO has data to read or when the FIFO is empty during writing. There are two ways to perform a write from the DSP. The first method is to wait until the FIFO is empty (before reading or writing) and then read or write a number of words that satisfy the FIFO. This will handle 32 words in the current hardware configuration and guarantees that the FIFO will not be overrun or underrun. After reading or writing 32 words, the DSP can put the BIO signal in a loop (which indicates that the FIFO is ready to receive the other 12 words) and read or write the other 32 words. ..
The second method is to transfer blocks of data larger than 12 words. This method uses a BIO signal, but the DSP puts the BIO signal in a loop before each read or write to the FIFO. Half-empty and half-full flags can be used to display the FIFO status via the BIO signal for these operations. The DSP uses that I / O space to set a bit in the control register to initiate a PCI transaction. The control register parameters are the PCI address space selection bit, the FIFO reset bit, the PCI macro bit, the DSP state bit, the start bit, and the BIO selection bit. A status register is also needed to check the status of the FIFO before the transaction starts. The parameters required to initiate a PCI bus transaction are the DSP address (which translates to the PCI main memory address), the number of words to transfer (the hardware translates the number of words into the number of bytes), and The direction.
[0118] Signal data phase transfer occurs when only two words are transferred through the PCI bus. These two words only satisfy one 32-bit transaction. PCI macros need to be aware of this special situation in order to handle the PCI protocol correctly. The value (number of words) written by the DSP is doubled to get the number of bytes, so the minimum number of bytes is 2. There are actually four bytes to transfer. If less than 4 bytes are transferred, the DSP must change the PCI bus byte allow bits, or two of the 4 bytes of the PCI data phase are not valid. The directional bit is important because it is a determinant within the sender of the FIFO input. The directional bits are set automatically when the DSP writes to the FIFO (this prevents the bits from having to be set, the FIFO written, and then the transition started). When reading from the PCI bus, the directional bits are set by writing to the register that initiates the transaction. The BIO signal is selected by the value of the BIO selection register. This value must be changed each time there is a change between the read and write operations.
After these parameters are written to the interface ASIC, PCI transfer is initiated when the start bit is written by the DSP. An example of DSP order is as follows: That is, [0120] DSP read order 1. Write the address of the transferred data. 2. Write the number of words to transfer ( 2). 3. Write the DONE bit (clear it), the working bit (set it), the transfer direction (0 = read), the BIO selection, and the START bit.
[0121] Once the start bit is set, the hardware waits for the 3.1 PCI bus to be free (the P2CDGNTCUPL signal transitions to inactive). 3.2 Write to a register to set the CD2PREQCULPL signal and the CD2PGNTCUPL signal. 3.3 Wait for the P2CDGNTCUPL signal to transition to active. 3.4 Write to register to set CD2PMSTCUPL signal (to prevent other transactions). 3.5 Strobe the new address, number of bytes, direction, and polyphase state to the macro. 3.6 Start the operation to fill the FIFO. 4. Wait for the macro to indicate that the FIFO is not empty. (Bring this state to the DSP as BIO). Five. Read from the FIFO. Reads can be made from the FIFO with a 2-cycle / 1 wait state instruction. These standby states are software standby states. The FIFO read pulse is generated by the DSP read signal. If the FIFO has no data, the BIO will not be set and the operation will be postponed (for this you may have to use internal time out). By using the BIO signal, the DSP waits for data. Loop back to 3. and repeat until completion. The DONE bit can be optionally set when the PCI interrupt has been generated.
If the PCI bus slave can run without waiting during the last two data phases, the PCI macro will satisfy the FIFO. If there is no wait in the last data phase, the FIFO will not write the last word. The FIFO will have the FIFO Full-1 times word inside.
DSP Write Order 1. Wait for the macro to indicate that the FIFO is empty. (Have this state in the DSP as BIO.) 2. Write to the FIFO. Writes can be made to the FIFO with a 3 cycle / 1 wait state instruction. These standby states are software standby states. The FIFO write pulse is generated by the DSP write signal. 3. Write the address of the data to be transferred. 4. Write the number of words to transfer (32 in this case). 5. Write the DONE bit (reset it), the working bit (set it), the BIO selection, and the start bit.
[0124] Once the start bit is set, the hardware waits for the 5.1 PCI bus to be free (the P2CDGNTCUPL signal transitions to inactive). 5.2 Write to a register to set the CD2PREQCUPL signal and the CD2PGNTCUPL signal. 5.3 Wait for the P2CDGNTCUPL signal to transition to active. 5.4 Write to register to set CD2PMSTCUPL signal (to prevent other transactions). 5.5 Strobe the new address, number of bytes, direction, and polyphase state to the macro. 5.6 Warn the DSP that the FIFO is empty (use the BIO signal). Loop back to 3. and repeat until completion.
[0125] When the DSP accesses the main memory, the 16-bit word-oriented address of the DSP must be translated into a 32-bit byte-oriented address. This part of the translation requires that the 16 bits of the DSP be shifted 1 bit to the left (multiplied by 2). This means that instead of addressing 64K words of memory space, the DSP can actually access 128K bytes of PCI host main memory. The PCI specification requires that a PCI initiator wishing to use linear addressing (the address increments linearly during burst transfer) leave the 2 least significant bits of the address as 00. This imposes a constraint on the DSP software that all fetches from main memory must start with an odd address (least significant bit 0). The DSP data space and instruction space will be relocatable on 4K boundaries for flexibility in applications in main memory.
[0126] The PCI bus is a selectable width bus. Any 4-byte allow signal can be activated to indicate that any or all of the 4 bytes on this bus are valid data. It is assumed that the DSP has a 16-bit bus, which has 16 bits of valid data for each transfer. All byte grant signals are active (all data are 32 bits) while the DSP is the initiator on the PCI bus. There are two 16 bits in parallel to capture PCI data for transfer between the PCI and DSP buses. The DSP muxes the output of several FIFOs to get 16-bit words. Multiplexing is controlled by the least significant bit of the DSP. 0 on address bit 0 selects the least significant FIFO and 1 on address bit 0 selects the most significant FIFO.
The stereo audio codec 1956 is an 8-bit device with an ISA bus interface. This codec is connected to a data bus that is separate from the DSP data bus. The DSP accesses this codec via I / O spatial registers. To prevent the DSP from being constrained to wait for the slow ISA interface, ASIC316 implements a state machine that controls actual reads and writes to and from the DSP.
This state machine that controls the codec is controlled via the CODEC Statue / Mode register in the DSP I / O address space. The state machine starts when 1 is written to the start bit. Depending on whether a read or write is required, and whether PIO or DMA access is required, the state machine follows the appropriate protocol and writes those bytes to the I / O spatial registers or the contents of the codec. To read and put in those registers. The BYTE XFER COUNT bit determines how many bytes are transferred during DMA operation (bits 1 to 4). The ADDR SEL bit determines to which address the least significant byte is written during PIO access (bits 0 to 3). When the codec state machine completes, the codec state bit DONE is set. This bit is cleared when the state machine starts. This bit is provided for information only for the DSP.
[0129] The PCI card user has an option to include external memory in the program space and data space of the DSP. This memory (socketed for easy installation and removal) allows the DSP algorithm to maintain its current form until the code can be rewritten to take advantage of access to host main memory. Will be able to. Data space is available to keep local data separate from host main memory. Global memory is designated as an interface to the host system. This interface will go through a FIFO that resides in the global memory space. This memory is accessed using software standby instead of a READY signal.
The bootload process uses global memory location 0xFFFF and memory map I / O port 50h. This position is available in global memory if the BIO bit (DSP command word --bit 3) in the PCI I / O space register is set high. If the BIO is set low, a read from position 0xFFFF will result in a bus collision between the bootload data and the global memory data. If the user has to use global memory, the BR # signal from the DSP to the ASIC will be gated with a user address decode signal and the ASIC FIFO will be activated during the user's global memory access. To prevent.
[0131] The PC BIO system (hereinafter referred to as BIOS) software executed at power-on is responsible for discovering what kind of PCI card is installed in the system. This software is so-called power-on self-test (hereinafter referred to as POST) code and is responsible for performing generic initialization of PCI applications. After completing the initialization, insert the card into the main memory and take out its own program code.
Once the PC is powered up, the DSP is held in the reset state until the host DSP driver releases the DSP reset. At this point, the driver writes to the PCI I / O spatial command word (COMMAND WORD) register to release the DSP reset. The CPU then bootloads the BDSP. This device driver downloads 16-bit words to the DSP via the PCI I / O Spatial Bootload DATA register by looking for the BDSP to set the XF, and then the valid data is the bootload data. After writing to the register, set the BIO (in the command word register). This is PCI No additional hardware is required in the ASIC or on the application card. This causes some overhead with respect to the host DSP driver. This is also the most flexible research study, as the bootload code can be varied within the software. Once the DSP is initialized, the DSP waits for the CPU to write to the register and set the COMMAND bit. The COMMAND bit causes an interrupt to the DSP and causes the DSP to start execution. One of the first things this initial DSP code should do is to check for the existence of data memory. If data memory is found, it must be determined whether the memory is 8-bit or 16-bit wide. The existence of DSP global memory gives DSP applications the opportunity to store up to 64K words (or bytes) of data in high-speed memory on the PCI card without having to go through the PCI bus to access the PCI card. means.
FIG. 20 is a schematic block diagram of a memory allocation and locking model for an improved computer system of the present invention. The memory allocation and locking shown in FIG. 20 corresponds to that shown in FIG. It is known through DSP software that there is a pointer to DSP space. It is also known where the sender and receiver tables are located for each transaction. Figure 20 provides a directory for the memory architecture. FIG. 11 represents a single sender space and a continuous receiver space. In Windows95, this can be said about virtual space, but physical addresses are scattered throughout the space. Therefore, there is the term DSP-scattering-collection bus mastering.
FIG. 21 is an overview block diagram of a sender and receiver data DMA transfer table model for an improved computer system of the present invention. In FIG. 21, each block 2110, 2120, 2130, 2140 is a (scattered locked) area within the memory 112 or portion of the memory 112 used by the application during the run time. At that moment a directory is assembled, which splits into many pieces within the area. Linked lists allow applications to fly around in virtual space. Ping-pong buffers 2150, 2160 research studies are used. The receiver is the sender for the CPU that succeeds this CPU.
FIG. 22 is a schematic block diagram of the internal structure of the sender data DMA transfer table model for the improved computer system of the present invention. FIG. 22 shows a more detailed structure of the sender DMA transfer table compared to FIG.
[0136] FIG. 23 is a detailed list of regions for the sender DMA transfer table of FIG. FIG. 24 is a schematic block diagram of the internal structure of the receiver data DMA transfer table model for the improved computer system of the present invention. FIG. 25 is a more detailed block diagram of the programmer / dataspace subtable of FIG. FIG. 24 shows the details of the receiver transfer table structure. FIG. 25 shows the details of the programmer / data space block of FIG. FIG. 23 shows the details of the list area in any one of FIGS. 22, 24, and 25. These figures show how DSPs and host CPUs share and co-exist when using the same memory space, such as local memory 112.
FIG. 26 is an electrical circuit block diagram of another embodiment of an improved computer system according to the present invention. In the case of this embodiment, the memory management logic, the memory controller, and the memory, and the cache are combined into one block and connected to the CPU via the local bus. The local bus is then connected to the PCI bridge, which is then connected to the PCI bus. In addition, the video / graphics chip 2614 and the image capture / decompression block 2618 provide 2D and 3D graphics and the various video image functions previously mentioned herein.
[0138] Figure 28 is improved con according to the present invention is an electrical circuit block diagram of another embodiment of a computer system. In the case of this embodiment, the CPU has a memory controller and PCI bridge integrated on its chip and is therefore directly connected to the memory and PCI bus. An integrated multimedia graphics controller is connected to the PCI bus and display. In addition, two optional chips are provided and connected to the PCI bus. One chip is for peripheral I / O COMBO with DSP. The other chip is for PCMCIA controllers with DSP. The BIOS of the PC is connected to the I / O block so that the keyboard, mouse, and parallel port (hereinafter referred to as PP) are connected. Not only a hard disk device (hereinafter referred to as HDD) and a floppy disk device (hereinafter referred to as FDD), but also a series port (hereinafter referred to as SP) is connected to the DSP part.
FIG. 29 is an electrical circuit block diagram of another embodiment of the improved computer system according to the present invention. FIG. 29 is similar to FIG. 28, but the wireless LAN of FIG. 29 is integrated with the PCMCIA controller block and sophisticated to include 1394 data streams.
FIG. 29 shows an example in which one chip provides all functionality and a bare bond physical layer provides I / O. Everything other than the physical layer is the part of the PC. This example avoids multiple peripheral electronics and can focus on the physical layer that requires Federal Communications Commission (FCC) approval. The line protection device (hereinafter referred to as DAA) is software and has the ability to annotate voice with a speech codec. The LAN focuses on the physical layer, and the DSP virtualizes other LAN functions. In the ultimate digital innovation of cost reduction, everything except the physical layer is virtualized, and the DSP200 is at the core of this innovation.
FIG. 30 is an electrical circuit block diagram of an example application / driver model for an improved computer system of the present invention. The VxD environment (hereinafter referred to as VXDE) provides DSP-related low-level services for Windows applications, such as multimedia applications. The Windows 3.1 model of the DSP driver shown in Figure 30 includes Windows VxD and Windows DLL. A DLL interface to or to a Windows application performs application callbacks and file I / O, handles interrupts, and interfaces with the VxD. The VxD is only responsible for locking down and freeing the physical memory that may be used by the DSP.
[0142] Under Windows 95, the above model is no longer suitable because the Windows driver that the DSP driver DLL needs to communicate with is 16-bit or 32-bit. In order for both the 16-bit and 32-bit Windows drivers to be able to communicate with the DSP driver, the above model is modified as shown in Figure 31.
FIG. 31 is an electrical circuit block diagram of another embodiment of an application / driver model for an improved computer system of the present invention. The model shown in FIG. 31 has a Windows VxD, a 16-bit Windows DLL, and a 32-bit Windows DLL. These DLLs are only needed for communication between the Windows driver and the DSP VxD. However, for this model, the DSP VxD is particularly responsible: it interfaces with the DSP, handles application callbacks and file I / O, etc., performs interrupts, and is given the name DSP VxD environment. ing.
VXDE provides DSP-related low-level services for Windows applications, such as multimedia applications. For multimedia applications, these services include: NodeAdvise, NodeAllocate, NodeDestroy, NodeGetAttr, NodeGetData, NodeGetPosition, NodePause, NodePuttData, NodeResetStream, NodeRun, NodeSetAttr, NodeSignalEvent, NodeCovertData, NodeWaitSemaphore, NodeCreatSemaphore, and NodeDestroySemaphore.
Windows multimedia applications communicate with VXDE through a 16-bit DLL or 32-bit DLl, which in the case of 16-bit through a callback function or in the case of 32-bit through an event signaling mechanism [Semaphore]. Communicate with Windows applications. The VXDE communicates with the DSP by writing to the memory map port of the DSP, and the DSP communicates with the VXDE by generating a hardware interrupt. However, this hardware is virtualized by VXDE through VPICD.
For any Windows application that wants to communicate with VXDE, all that needs to be done is a set of APIs listed above, along with a 16-bit or 32-bit DLL (referred to as DSAPI.DLL or DSPAP32.DLL). Is to use. Preferably, the 16-bit DLL puts a parameter in the call structure, marks this structure as a call from the 16-bit side, and then calls VXDE through the entry point. The entry point is located using the Int 2F function 1684h. The 32-bit DLL, on the other hand, calls VXDE through the DEVICEIOCONTROL interface. On the VXDE side, when a call is received from the entry point, the pointer of the call structure is translated into a linear address and the call is dispatched to the corresponding function. This feature translates any type of pointer from a SELECTOR / OFFSET address to a linear address and then processes this call. If the call is received from the DEVICEIOCONTROL interface, the call will be dispatched directly to the corresponding function without any translation required.
A 12-bit Windows application that wants to be notified when the DSP hardware has completed a particular task can do so by putting a valid event handle in the OVERLAPPED structure and sending a pointer to VXDE. .. Therefore, it needs to call WaitForSigleObject at the time of the event. For applications that do not want to be notified, when calling VXDE, the parameter containing the pointer to the OVERLAPPED structure is assumed to be zero.
A 16-bit Windows application that wants to be informed when the DSP hardware has completed a particular task can use the callback feature.
[0149] FIGS. 32 and 33 are schematic block diagrams showing various effects in the application / driver model of FIGS. 31 and 32 on the improved computer system of the present invention. VXDE virtualizes the IRQ used to communicate with the DSP through VPICD. Any interrupt to the specified IRQ will be dispatched to VXDE. When an interrupt occurs, VXDE first clears the interrupt to allow any future interrupt, and then starts servicing the interrupt. If the interrupt is the result of a call from a 16-bit application, the callback function is handled through a scheduled event (Call_Priority_Event), during which the callback function is called through the Vmm service Simulate_Far_Call. This is because when the VxD handles a hardware interrupt, it is constrained by the numbering service it calls. If the interrupt is the result of a call from a 32-bit application, the application's notification is also handled through the scheduling event, where the notification is realized by calling the Win32 service vWin32_DIOCCompletionRoutine using EBX = overlapped.Internal.
FIG. 34 is a schematic block diagram of an embodiment of a virtual memory model for an improved computer system of the present invention. The model in Figure 34 shows how all communications are redirected to DSP hardware through serial ports.
[0151] DSPVxD provides a set of services to port drivers and port virtualization VxDs to redirect communication through serial ports to DSP hardware. At the system boot time when the port virtualization VxD is loaded, this virtualization VxD initializes its execution environment, installs I / O trapping handlers, installs port I / O contention handlers, and so on. Virtualize a computer output microfilm (hereinafter referred to as COM) device.
[0152] When a DOS application in a disk operating system (hereinafter referred to as DOS) attempts to acquire a COM port, ownership of that port is given to real-mode virtual memory (hereinafter, virtual memory is referred to as VM). , The IRQ associated with that COM is virtualized through VPICD, so interrupts from that IRG can be reflected in the correct VM. Any communication through the COM port is trapped through the installed port I / O trapping handler and redirected to the DSP handler using the set of services provided by the DSP VxD.
When a running Windows application on system VMz attempts to acquire a COM port, the port driver calls the contention handler installed by port virtualization VxD to transfer ownership of that COM port to the system VM. Set and turn off I / O trapping. Any communication through the COM port is redirected to the DSP hardware by a port driver that uses the set of services provided by DSP VxD.
FIG. 35 is a schematic block diagram of another embodiment of the virtual memory model for the improved computer system of the present invention. The model in Figure 34 can be changed to Figure 35 so that the interface with the hardware occurs only within the port driver. The advantage of doing this is that when changing hardware, only the port driver needs to be modified. The downside is the additional delay associated with DOS-based applications.
FIG. 36 is an electrical circuit block diagram of an embodiment of an improved computer system portion according to the present invention. FIG. 36 shows the integration of the graphic controller with the memory controller and the sharing of memory through these memory data buses. The CPU supplies an address to the memory controller, controls it, and receives data on its host bus via a data buffer.
FIG. 37 is an electrical circuit block diagram of an embodiment of an improved computer system portion according to the present invention. FIG. 37 illustrates the use of VZ (zoom video) as a point-to-point unidirectional video bus between a PC card socket and a video graphic array (hereinafter referred to as VGA) controller. This figure shows how a TV in a window can be achieved in a portable computer with a low cost PC card. MPEG or video conferencing cards can also be plugged into PCMCIA slots.
[0157] FIGS. 38-44 are electrical circuit block diagrams of an example showing a breadboard of parts of an improved computer system according to the present invention. This embodiment is transformed into a DC / PC question card, which is a multi-layer PC board that meets the PCI short card definition. This card has a digital area, an analog area, and a blank area, and it is possible to connect a child card to the blank area in a short position. The analog ground plane for the analog area and the digital ground plane for the digital area shall be connected close to the codec chip. The end plate of this board houses a C5x embroidery header (closest to the motherboard), three stereo horn jacks (3.5mm), two RCA844 jacks, and an RJ11 horn jack (farthest from the motherboard).
The header for the child card is a double 90 degree pin on the motherboard. The child card has a double header socket made for short boards overall. The address range available via the I / O bus on the child card is decomposed to give flexibility in assigning software wait states to different address ranges.
The DSP C5x series has two input pins that determine which type of input clock operation scheme to use. The CLLKMD1 and CLKD2 signals allow four different locking behaviors (one of which is reserved for testing). This board design has two clock operation options.
[0160] FIG. 45 is a schematic block diagram of an example of an MPEG reproduction filter graph model adopted in the improved computer system of the present invention. In Figure 10, when the user clicks on a software object, the Windows application dynamically allocates memory and runs OLE. The "COM interface" circle in Figure 45 is the interface between the application and the filter graph manager, and the MCI block controls the media interface. MPEG works from object to object expression instead of layer to layer expression. The sender may include image capture data for stretching.
FIG. 46 is an electrical circuit block diagram of another embodiment of an improved computer system according to the present invention. FIG. 46 is similar to FIG. 29, but the PC card bus block has an integrated IDSP, and a ZV bus for video. Southbridge also includes an integrated IDSP function. Graphics and video capabilities are integrated in a single chip connected to the PCI bus.
FIG. 47 is a block diagram of software applications and related applications used in the improved computer system of the present invention. FIG. 47 is similar to the lower portion of FIG. 29.
FIG. 48 is a simplified block diagram of an audio decoder used in the improved computer system of the present invention. Incoming coded audio is demultiplexed, error checked, and any accompanying data is output. Various audio subbands are output to the dequantization block, which is fed with the quantization block step as part of the data. The inverse quantization block 1411 outputs to the inverse filter bank block, the latter outputs the decoded audio signal.
FIG. 49 is a schematic block diagram of a direct DSP component and how this component interfaces with its driver and emulation block. Such components are employed in the improved computer system of the present invention.
FIG. 50 is a schematic block diagram of another embodiment of the virtual memory model for the improved computer system of the present invention. The model shown in FIG. 50 is similar to that of FIG. 12, but has a Windows direct DSP HAL, a 16-bit Windows DLL, and a 32-bit Windows DLL. These DLLs are only needed for direct communication between the Windows driver and the DSP. However, for this model, the direct DSP has all the responsibilities: interface with the DSP, handle interrupts, perform application callbacks and file I / O, etc., and is given the name of the DSP environment directly. ing.
FIG. 51 is a schematic block diagram showing various layers of software between a Windows application that implements an improved computer system of the present invention and the underlying PC hardware.
FIG. 52 is a schematic block diagram showing different effects within the different models of FIG. 50 on the improved computer system of the present invention. Figure 52 is similar to Figure 32, but also contains API level blocks.
[0168] FIG. 53 is a schematic block diagram showing different effects within the various virtual models of FIG. 50 on the improved computer system of the present invention. Figure 52 is similar to Figure 33, but also contains API level blocks.
FIG. 54 is a schematic block diagram of another embodiment of the virtual model for the improved computer system of the present invention. FIG. 54 is similar to FIG. 34, but also contains the HAL block directly.
FIG. 55 is a schematic block diagram of another embodiment of the virtual model for the improved computer system of the present invention. FIG. 55 is similar to FIG. 35, but also contains the HAL block directly.
FIG. 56 is a schematic block diagram of another embodiment of the virtual model for multimedia that can be used on the improved computer system of the present invention. This model uses diretx nomenclature and shows various interfaces between the application and the driver components.
[0172] FIG. 57 is a schematic block diagram showing various layers of software between a Windows application that implements an improved computer system of the present invention and the underlying PC hardware. This model shows different interfaces between applications, driver components, and different hardware blocks.
FIG. 58 is a schematic block diagram showing various ways in which application functionality selected to provide a block containing a DSP according to the techniques of the present invention can be collectively combined. This figure shows various integrations that use the functionality typically attached to the PCI bus.
[0174] FIG. 59 is a schematic block diagram showing an alternative method in which application functionality selected to provide a block containing a DSP according to the techniques of the present invention is collectively combined. This figure shows the various integrations that use the functionality typically associated with the PCI bus and how some of this functionality is included with the northbridge to accelerate CPU operation. In addition, a second PCI bus has been shown for various networking functionality or high speed access.
FIG. 60 is a schematic block diagram of another embodiment used in the improved computer system of the present invention. Figure 60 shows the use of interface 5550, which provides access to the CPU and memory. Connected to interface 5550, there are multiple blocks 5510, 5520, 5530, each containing a DSP and used to virtualize one or more applications according to the teachings of the present invention. These blocks are then connected to interface 5560, the latter appropriately connected to the PCMCIA, PCI, and / or ISA buses. In addition, the high speed ZV bus is also shown to interconnect with a suitable one of blocks 5510, 5520, 5530.
FIG. 61 is a high-level schematic block diagram of other embodiments adopted in the improved computer system of the present invention. FIG. 61 shows how the integrated block 5500 of FIG. 60 is classified as a processing element connected to a memory and an appropriate AD conversion block and DA conversion block. Additional I / O is also shown.
Although the present invention has been described with reference to the illustrated examples, it is not intended to be a limiting interpretation of this description. Other embodiments of the invention, as well as the various variations of the illustrated practices, are feasible and obvious to those skilled in the art by reference to this description. Therefore, the appended claims are considered to include any such modifications and examples that fall within the true scope of the invention.
The following sections are further disclosed with respect to the above description.
(1) A computing system, a main CPU microprocessor, a DSP microprocessor having an instruction set different from that of the main CPU microprocessor, and a storage device coupled to the main CPU microprocessor and the DSP microprocessor. , And a calculation that includes a file-based operating system in the storage device arranged so that the DSP performs the operation of the CPU during a time interval in which the main CPU is otherwise occupied to enhance the performance of the system. system.
(2) The computing system according to paragraph 1, further comprising a video integrated circuit coupled to the DSP microprocessor and the main CPU microprocessor, the storage further including a disk and a DRAM, said A computing system in which a DRAM is coupled to the microprocessor, the main CPU microprocessor, and the video integrated circuit in a unified memory architecture.
(3) Hardware according to paragraph (1), further comprising at least one application device supporting an application of the system, wherein the application device is substantially reduced to only the physical layer. A computing system in which the DSP microprocessor uses signals mediated by the physical layer to virtualize and execute the rest of the application.
(4) The computing system according to paragraph 1, further comprising kernel software in the storage device for execution by the DSP microprocessor.
(5) The computing system according to paragraph 4, further comprising an I / O port coupled to the DSP microprocessor, the DSP micro that the kernel software properly coordinates with the file-based operating system. If the processor operation is defined and the main CPU microprocessor is occupied to perform a given function representing virtual hardware, the DSP microprocessor performs the function and the main CPU and said If both DSPs are free, either the main CPU microprocessor or the DSP microprocessor can perform the function as determined by the file-based operating system, and the virtual hardware is the CPU. A calculation system having a mobility related to the I / O and the I / O.
(6) In the computing system described in paragraph 4, the kernel software uses the file-based operating system to define interrupt-based DSP microprocessor operation, and priorities for a given function are calculated in real time. A computing system in which functions are dynamically performed and the kernel software can supply the file-based operating system, thereby calculating real-time priorities.
(7) In the computing system according to paragraph 1, the file-based operating system reads the program space for the DSP microprocessor and further as a shared memory model between the main CPU microprocessor and the DSP microprocessor. Includes software that supplies the DSP microprocessor with a pointer address to the storage device for data sender space and write data receiver space, whereby the arrangement uses the file-based operating system to make the DSP microprocessor the main. A computing system that is tightly coupled to the CPU microprocessor.
(8) The computing system according to paragraph 1 capable of executing at least a part of a software application, wherein the file-based operating system starts and ends the software part anywhere in the virtual memory. Includes the software that defines the handles that tell, and further includes the DSP kernel that defines the behavior of the system causing the DSP microprocessors to perform main CPU functions, where the main CPU microprocessors have sender and receiver handles. The file-based operating system defines the behavior of sending information to the position defined by the receiver handle, where the DSP kernel has said the sender handle and the receiver handle to the DSP microprocessor. A computing system that defines behavior based on whether a function should be performed on behalf of the main CPU.
(9) In the computing system according to paragraph 1, the main CPU microprocessor uses a virtual address, and the file-based operating system includes a virtual memory manager that provides a physical address corresponding to the virtual address. A computing system that further includes a DSP kernel that defines the behavior of the computing system using physical addresses for DSP microprocessor functions.
(10) In the computing system described in paragraph 1, the file-based operating system provides a virtual device driver (VxD) that communicates with the application through a callback function for 16-bit applications and through a semaphore for 32-bit applications. A computing system that contains the software to define.
(11) In the computing system described in paragraph 1, the operation of virtualizing interrupts used by the file-based operating system to communicate between the DSP microprocessor and the main CPU microprocessor is defined. A computing system.
(12) A data input having a width, a data output having a width different from the data input, an address input, an address output, and a first bus having an address, a data width, and a first bus clock frequency, and the first bus. An integrated circuit device having a second bus with a different data width and a device such as that used as an interface to the second bus having an address, which is synchronized with the first bus clock frequency. At least two parallel multilingual FIFOs with operating clock inputs for the FIFO, multiplexer electronics with control inputs, bytes of data between the data inputs and the different widths of data outputs via the FIFO. The multiplexer electronic circuit coupled so as to multiplex the data, and an address translation circuit that translates an address between the address input and the address output, the address input being the control input of the multiplexer electronic circuit. An integrated circuit device comprising the address translation circuit, which has a coupled least significant bit so that the data bytes are differently multiplexed according to the address input least significant bit.
(13) The integrated circuit device according to paragraph 12, further comprising a processor having an instruction set and integrated on the integrated circuit device, the processor having the address input and the address output. Combined, integrated circuit device.
(14) The integrated circuit device according to paragraph 13, in order to have a byte use permission output, set the byte use permission output, and combine data from the FIFO with the data output. An integrated circuit device that is more responsive to the processor.
(15) The integrated circuit device according to paragraph 13, further comprising a northbridge electronic circuit including a DRAM memory controller and a PCI bus interface.
(16) The integrated circuit device according to paragraph 13, further comprising a card bus controller electronic circuit.
(17) The integrated circuit device according to paragraph 13, further comprising a series bus controller electronic circuit selected from the group including a general purpose series bus (USB) and an IEEE1394-compliant series bus.
(18) The integrated circuit device according to paragraph 12, further comprising a scattering-collecting DMA electronic circuit coupled to the address output.
(19) The integrated circuit device according to paragraph 12, further comprising a bus control block including both a master circuit and a slave circuit coupled to the data output.
[0198] (20) A main CPU microprocessor, a first bus coupled to the main CPU microprocessor, which has an address line, a data width, and a first bus clock frequency, and instructions different from those of the first bus and the main CPU microprocessor. A second microprocessor having a set, a second bus having a data width different from that of the first bus, having an address line and being coupled to the second microprocessor, and a data input having a width. An integrated circuit device having a data output having a width different from that of the data input, an address input, and an address output, and being coupled as an interface to the first bus and the second bus. A multiplexer electronic circuit with at least two parallel multilingual FIFOs, control inputs, having clock inputs for the operation of the FIFO synchronized to the clock frequency, with the data inputs and the data outputs of different widths via the FIFO. The multiplexer electronic circuit coupled so as to multiplex the bytes of data between them, and an address translation circuit that translates an address between the address input and the address output, wherein the address input is the multiplexer electronic circuit. The integrated circuit device having the least significant bit coupled to said control input of, and thus having said address translation circuit in which data bytes are differently multiplexed according to said address input least significant bit. 2 A computing system in which a microprocessor is coupled to the address input and the data input via the second bus.
(21) An improved PC system (100) that includes a main CPU microprocessor (102), a file-based operating system (FIG. 19a), and a DSP microprocessor (200), and includes a main CPU ( The DSP (200) is arranged to perform the operation of the main CPU during the time interval when 102) is busy with others, thereby achieving an increase in the bandwidth of the PC system. This PC system can also include multiple CPUs and / or multiple DSPs.
BRIEF DESCRIPTION OF THE DRAWINGS FIG. 1 is an electrical circuit block diagram of an embodiment of an improved computer system according to the present invention.
FIG. 2 is a detailed electrical circuit block diagram (partially schematic, partially block diagram) of a preferred embodiment of the improved computer system portion of FIG.
FIG. 3 is a more detailed electrical circuit diagram (partially schematic, partially block diagram) of a preferred embodiment of the improved computer system portion of FIG.
FIG. 4 is a detailed electrical circuit block diagram (partially schematic, partially block diagram) of a preferred embodiment of the improved computer system portion of FIG. 1, where a is the configuration within the chip. Figure, b is the whole connection diagram.
FIG. 5 is a block diagram of an embodiment of an improved computer system of the present invention for host-dependent asymmetric multiprocessing.
FIG. 6 is a schematic block diagram of an embodiment using a superscalar extension for an improved computer system of the present invention, where a is a CPU model diagram and b is a memory cache hierarchical structure diagram.
FIG. 7 is a schematic diagram of an embodiment of a shared memory model for an improved computer system of the present invention.
FIG. 8 is a schematic diagram of an embodiment of a multimedia extension model for an improved computer system of the present invention.
FIG. 9 is a schematic diagram of an embodiment of a system / cache / virtual memory model for an improved computer system of the present invention.
FIG. 10 is a schematic block diagram of an embodiment of an MPEG reproduction filter graph model that may be used in the improved computer system of the present invention.
FIG. 11 is a schematic block diagram of an embodiment of a virtual I / O hardware-PCI DMA and multimedia real-time interrupt handler model for an improved computer system of the present invention.
FIG. 12 is a schematic block diagram of parallel processing within a frame for an improved computer system of the present invention.
FIG. 13 is a simplified block diagram of an MPEG encoder that may be employed by the improved computer system of the present invention.
FIG. 14 is a simplified block diagram of an MPEG decoder 1400 that may be employed by the improved computer system of the present invention.
FIG. 15 is a simplified block diagram of an embodiment of a video solution for a notebook computer of the present invention.
FIG. 16 is a simplified block diagram of an embodiment of video resolution for the desktop computer of the present invention.
FIG. 17 is a simplified block diagram of an application pipeline for DSP algorithms for an improved computer system of the present invention.
FIG. 18 is an electrical circuit block diagram of another embodiment of an improved computer system according to the present invention.
FIG. 19 is an improved computer system flow diagram and electrical circuit diagram of the present invention, where a is a system flow diagram and b is a more detailed electrical circuit diagram (partial) of a preferred embodiment of the improved computer system of FIG. Schematic diagram, partially block diagram).
FIG. 20 is a schematic block diagram of a memory allocation and physical locking model for an improved computer system of the present invention.
FIG. 21 is an overview block diagram of a sender and receiver data DMA transfer table model for an improved computer system of the present invention.
FIG. 22 is a schematic block diagram of the internal structure of a sender data DMA transfer table model for an improved computer system of the present invention.
FIG. 23 is a block diagram showing details of a region list for the sender DMA transfer table of FIG.
FIG. 24 is a schematic block diagram of the internal structure of a receiver data DMA transfer table model for an improved computer system of the present invention.
FIG. 25 is a more detailed block diagram of the programmer / data space subtable of FIG.
FIG. 26 is an electrical circuit block diagram of another embodiment of an improved computer system according to the present invention.
FIG. 27 is an electrical circuit block diagram of another embodiment of an improved computer system according to the present invention.
FIG. 28 is an electrical circuit block diagram of yet another embodiment of the improved computer system according to the present invention.
FIG. 29 is an electrical circuit block diagram of yet another embodiment of the improved computer system according to the present invention.
FIG. 30 is an electrical circuit block diagram of an embodiment of a virtual memory model for an improved computer system of the present invention.
FIG. 31 is an electrical circuit block diagram of another embodiment of a virtual memory model for an improved computer system of the present invention.
FIG. 32 is a schematic block diagram showing various effects in virtual memory of FIGS. 31 and 32 on an improved computer system of the present invention.
FIG. 33 is a schematic block diagram showing various effects in the virtual memory of FIGS. 31 and 32 on the improved computer system of the present invention.
FIG. 34 is a schematic block diagram of another embodiment of a virtual memory model for an improved computer system of the present invention.
FIG. 35 is a schematic block diagram of yet another embodiment of a virtual memory model for an improved computer system of the present invention.
FIG. 36 is an electrical circuit block diagram of an embodiment of an improved computer system portion according to the present invention.
FIG. 37 is an electrical circuit block diagram of an embodiment of an improved computer system portion according to the present invention.
FIG. 38 is a schematic block diagram of an embodiment of an improved computer system portion of the present invention.
FIG. 39 is a schematic block diagram of an embodiment of an improved computer system portion of the present invention.
FIG. 40 is a schematic block diagram of an embodiment of an improved computer system portion of the present invention.
FIG. 41 is a schematic block diagram of an embodiment of an improved computer system portion of the present invention.
FIG. 42 is a schematic block diagram of an embodiment of an improved computer system portion of the present invention.
FIG. 43 is a schematic block diagram of an embodiment of an improved computer system portion of the present invention.
FIG. 44 is a schematic block diagram of an embodiment of an improved computer system portion of the present invention.
FIG. 45 is a schematic block diagram of an example of an MPE playback filter graph model for an improved computer system of the present invention.
FIG. 46 is an electrical circuit block diagram of another embodiment of an improved computer system according to the present invention.
FIG. 47 is a block diagram of software applications and related applications used in the improved computer system of the present invention.
FIG. 48 is a simplified block diagram of an audio decoder used in an improved computer system of the present invention.
FIG. 49 is a schematic block diagram of a direct DSP component and how to interface this element with a driver and an emulation block.
FIG. 50 is a schematic block diagram of another embodiment of a virtual memory model for an improved computer system of the present invention.
FIG. 51 is a schematic block diagram showing various layers of software between a Windows application that implements an improved computer system of the present invention and the underlying PC hardware.
FIG. 52 is a schematic block diagram showing different effects within the different models of FIG. 50 on the improved computer system of the present invention.
FIG. 53 is a schematic block diagram showing different effects within the various virtual models of FIG. 50 on the improved computer system of the present invention.
FIG. 54 is a schematic block diagram of another embodiment of a virtual model for an improved computer system of the present invention.
FIG. 55 is a schematic block diagram of another embodiment of a virtual model for an improved computer system of the present invention.
FIG. 56 is a schematic block diagram of another embodiment of a virtual model for multimedia that can be used on an improved computer system of the present invention.
FIG. 57 is a schematic block diagram showing various layers of software between a Windows application that implements an improved computer system of the present invention and the underlying PC hardware.
FIG. 58 is a schematic block diagram showing various ways in which application functionality selected to provide a block containing a DSP according to the techniques of the present invention can be collectively combined.
FIG. 59 is a schematic block diagram illustrating an alternative method in which application functionality selected to provide a block containing a DSP according to the techniques of the present invention is collectively combined.
FIG. 60 is a schematic block diagram of another embodiment adopted in the improved computer system of the present invention.
FIG. 61 is a schematic block diagram of high water levels of other embodiments adopted in the improved computer system of the present invention.
[Code Description] 100 PC System 102 CPU104 Cache 106 Main Memory Bus 108 Host Bridge 112 Main Memory 116 PCI Bus 120 P1394 / USB Block 124 I / O Block 128 ISA Bus 200 DSP Block 210 Hardware Interface Electronic Circuit 214 Interface Electronic Circuit , Memory 218 DSP core 222 DSP core 304 PCI master / slave interface circuit 305 Hardware layer 306 PCI configuration control and status register Electronic circuit 308 PCI I / O Spatial register Electronic circuit 310 Double port read / write FIFO electronic circuit 312 DSP I / O Spatial Register Electronic Circuit 316 Interface / Codex DMA Control Circuit 420 Chip 424 Master / Slave Interface 428 DSP432 Card Bus Controller 460 Chip 488 Data Buffer APP Application FFD Flopy Disk Device HDD Hard Disk Device MUX multiplexer PCI Peripheral Device Interconnect PP Parallel Port PPU Peripheral Processor SP Serial Port UMA Unified Memory Architecture VxD Video Driver WAN Wide Area Network
52 members in 4 offices
Priority claims4
| Document | Office | Kind | Date |
|---|---|---|---|
| 014734 | United States of America | – | |
| 1473496 | United States of America | P | |
| 1996014734 | – | – | – |
| US19960014734P | – | – | – |
Members52
| Document | Office | Kind | |
|---|---|---|---|
| EP0817096A2 | European Patent Office (EPO) | A2 | |
| JPH1083304A | Japan | A | |
| US5987590A | United States of America | A | |
| EP0964540A2 | European Patent Office (EPO) | A2 | |
| US6141744A | United States of America | A | |
| US6148389A | United States of America | A | |
| US6170048B1 | United States of America | B1 | |
| US6170049B1 | United States of America | B1 | |
| US6179489B1 | United States of America | B1 | |
| US6421527B1 | United States of America | B1 | |
| US6574213B1 | United States of America | B1 | |
| US6678267B1 | United States of America | B1 | |
| EP0964540A3 | European Patent Office (EPO) | A3 | |
| EP0817096A3 | European Patent Office (EPO) | A3 | |
| US6744757B1 | United States of America | B1 | |
| US6757256B1 | United States of America | B1 | |
| US6765904B1 | United States of America | B1 | |
| US6801499B1 | United States of America | B1 | |
| US6801532B1 | United States of America | B1 | |
| US6804244B1 | United States of America | B1 | |
| US2004252700A1 | United States of America | A1 | |
| US2004252701A1 | United States of America | A1 | |
| US2006039280A1 | United States of America | A1 | |
| JP3774538B2This record | Japan | B2 | |
| US7574351B2 | United States of America | B2 | |
| US7606164B2 | United States of America | B2 | |
| US2009268724A1 | United States of America | A1 | |
| US2009323679A1 | United States of America | A1 | |
| US7653045B2 | United States of America | B2 | |
| EP0964540B1 | European Patent Office (EPO) | B1 | |
| US2010085986A1 | United States of America | A1 | |
| DE69942077D1 | Germany | D1 | |
| EP0817096B1 | European Patent Office (EPO) | B1 | |
| DE69739934D1 | Germany | D1 | |
| US7822021B2 | United States of America | B2 | |
| US2011004808A1 | United States of America | A1 | |
| US8024182B2 | United States of America | B2 | |
| US8050254B2 | United States of America | B2 | |
| US2011301947A1 | United States of America | A1 | |
| US2012008645A1 | United States of America | A1 | |
| US8224643B2 | United States of America | B2 | |
| US2012259624A1 | United States of America | A1 | |
| US8456991B2 | United States of America | B2 | |
| US8463601B2 | United States of America | B2 | |
| US2013230043A1 | United States of America | A1 | |
| US2013246057A1 | United States of America | A1 | |
| US2013250938A1 | United States of America | A1 | |
| US8666735B2 | United States of America | B2 | |
| US2014135068A1 | United States of America | A1 | |
| US8995430B2 | United States of America | B2 | |
| US9106533B2 | United States of America | B2 | |
| US9153228B2 | United States of America | B2 |
23 legal events, as the office reported them to INPADOC
Over the term
Point at a mark for the eventEvents
| Event | Code | |
|---|---|---|
| Cancellation because of no payment of annual feesLAPS | LAPS | |
| Receipt of annual feesR250 | R250 | |
| Receipt of annual feesR250 | R250 | |
| Renewal fee payment (event date is renewal date of database)FPAY | FPAY | |
| Renewal fee payment (event date is renewal date of database)FPAY | FPAY | |
| Renewal fee payment (event date is renewal date of database)FPAY | FPAY | |
| Renewal fee payment (event date is renewal date of database)FPAY | FPAY | |
| Renewal fee payment (event date is renewal date of database)FPAY | FPAY | |
| Renewal fee payment (event date is renewal date of database)FPAY | FPAY | |
| Renewal fee payment (event date is renewal date of database)FPAY | FPAY | |
| Renewal fee payment (event date is renewal date of database)FPAY | FPAY | |
| Renewal fee payment (event date is renewal date of database)FPAY | FPAY | |
| Certificate of patent or registration of utility modelR150 | R150 | |
| First payment of annual fees (during grant procedure)A61 | A61 | |
| Written decision to grant a patent or to grant a registration (utility model)A01 | A01 | |
| Decision of grant or rejection writtenTRDD | TRDD | |
| Request for written amendment filedA521 | A521 | |
| Written permission of extension of timeA602 | A602 | |
| Written request for extension of timeA601 | A601 | |
| Notification of reasons for refusalA131 | A131 | |
| Report on retrievalA977 | A977 | |
| Request for written amendment filedA521 | A521 | |
| Written request for application examinationA621 | A621 |
Numbers
- Publication
- 3774538
- Publication, DOCDB
- 3774538
- Publication, EPODOC
- JP3774538B
- Application
- 8421097
- Application, DOCDB
- 8421097
- Application, EPODOC
- JP19970084210
Titles2
- Japanese
- パーソナルコンピュータ回路、コンピュータシステム、及びその動作方法
- English
- Personal computer circuits, computer systems, and how they operate
Classification
- CPC, 8
- G06F9/50
- G06F9/3879
- G06F9/544
- G06F13/28
- G06F13/4018
- G06F13/4027
- G06F15/7864
- G06F2209/509
- IPC, 5
- G06F9 38
- G06F9 50
- G06F13 28
- G06F13 40
- G06F15 78