Memory subsystem
Abstract
The embodiment of the present invention relates to a memory subsystem implemented in the parallel pipelined integrated circuit implementation of the calculation engine or connected to the parallel pipelined integrated circuit implementation of the calculation engine and accessed by the parallel pipelined integrated circuit implementation , The calculation engine is designed to solve complex calculation problems. Additional embodiments of the present invention relate to memory subsystems implemented in various different types of electronic devices or connected to and accessed by various different types of electronic devices. One embodiment of the invention includes a memory controller and one or more separate memory devices, the memory controller being implemented in a first integrated circuit or other electronic system. An alternative embodiment of the present invention incorporates a memory controller in one or more memory devices that are connected to a calculation engine implemented by an integrated circuit or another electronic device and are used by the calculation engine. Or another electronic device to access. In an alternative embodiment of the present invention, the memory controller and the memory are integrated together in a computing engine or another electronic device. Alternative embodiments of the present invention include a multiple access memory interfaced with a simpler memory controller to connect to a calculation engine or other electronic device or be integrated in the calculation engine or other electronic device.

Term
3.2 yearsleft in the term
Expires 21 December 2029.
- Priority
- Filed
- Granted
- Today
- Expires
25 claims: 1 independent, 24 dependent
- 1一种存储器子系统,所述存储器子系统包括: 存储器;以及 存储器控制器,所述存储器控制器 提供一个或更多个数据流接口,用于接收数据流输入,所述数据流输入包括根据预定 义的数据结构组织的多个数据, 为所述存储器内的各个存储器单元和二维存储器区域提供随机存取接口,每个各个存 储器单元和每个二维存储器区域与描述所述存储器内的相应位置的(X, y)坐标对相关联, 在数据流接口和随机存取接口之间进行仲裁,以使通过所述数据流接口和所述随机存 取接口接收的同时请求的存储器存取串行化, 通过将所述数据流输入写入到存储器来执行通过所述数据流接口请求的存储器存取, 以及 通过从存储器单元和二维存储器单元区域读取值和将值写入到存储器单元和二维存 储器单元区域来执行通过所述随机存取接口请求的单存储器单元和二维存储器单元区域 的存储器存取。
- 2如权利要求1所述的存储器子系统,在单集成电路计算引擎内实现。
- 3如权利要求1所述的存储器子系统,在计算引擎芯片组内实现。
- 4如权利要求1所述的存储器子系统,在存储器集成电路内实现。
- 5如权利要求1所述的存储器子系统,其中所述存储器子系统由具有下述频率的 fastclk时钟信号控制,所述频率至少是所述存储器控制器通过所述一个或更多个数据流 接口中的任一个接收的最快时钟信号的频率的η倍,其中η大于2,并且其中从数据流接口 接收的所述时钟信号便利于所述存储器控制器通过所述数据流接口接收的所述数据流的 数据单元的存储器控制器接收的同步。
- 6如权利要求5所述的存储器子系统,其中所述存储器子系统包括仲裁器,所述仲裁 器在所述数据流接口和所述随机存取接口之间进行仲裁,以使通过所述数据流接口和所述 随机存取接口接收的同时请求的存储器存取串行化,所述仲裁器包括与每个数据流接口对 应的数据源仲裁器块和与所述随机存取接口对应的存储器定序器块。
- 7如权利要求6所述的存储器子系统,其中所述一个或更多个数据源仲裁器块和所述 存储器定序器仲裁器块在一个序列中被链接在一起,所述数据源仲裁器块和所述存储器定 序器仲裁器块在所述序列内的位置限定所述数据源仲裁器块和所述存储器定序器仲裁器 块中的每个的优先级,其中第一位置中的所述数据源仲裁器块具有最高优先级,其余优先 级随着相关联的数据源仲裁器块在所述序列中的位置增大而降低,并且所述存储器定序器 仲裁器块具有最低优先级。 如权利要求7所述的存储器子系统,其中每个数据源仲裁器块为单个写操作获取存 储器控制,但是然后交出存储器控制,以允许其他数据源仲裁器块或所述存储器定序器仲 裁器块获取存储器控制。
- 89. 如权利要求7所述的存储器子系统,其中所述存储器定序器仲裁器块获取存储器控 制,直到来自通过数据流接口输入的数据流的数据损失的可能性迫使所述仲裁器从所述存 储器定序器仲裁器块撤销存储器控制并且将存储器控制授予所述数据流接口为止。
- 910. 如权利要求7所述的存储器子系统,其中单数据源仲裁器块或所述存储器定序器 CN 102369552 Β 仲裁器块在任何时间被授予存储器控制,其中与当前通过数据流接口请求存储器存取的所 述数据流接口相关联的较高优先级的数据源仲裁器块阻止将存储器控制授予较低优先级 的数据源仲裁器块,并且其中所述最高优先级的存储器存取请求接收数据源仲裁器块或存 储器定序器仲裁器块接着被授予存储器控制。
- 1011. 如权利要求6所述的存储器子系统,还包括一个或更多个数据源块和存储器定序 器块,所述一个或更多个数据源块每个实现数据流接口,并且每个与用于所述数据流接口 的数据源仲裁器块相关联,所述存储器定序器块实现所述随机存取接口,所述随机存取接 口提供对各个存储器单元和二维存储器区域的存取,并且与所述存储器定序器仲裁器块相 关联。
- 1112. 如权利要求11所述的存储器子系统,其中每个数据源块包括FIFO队列,所述FIFO 队列用于临时储存通过所述数据流接口输入的所述数据流的数据单元,所述数据流接口由 所述数据源块实现。
- 1213. 如权利要求12所述的存储器子系统,其中被包括在通过所述数据流接口接收的数 据流中的冗余数据单元被所述数据源块丢弃,而不是被插入到FIFO队列中,所述数据流接 口由所述数据源块实现。
- 1314. 如权利要求5所述的存储器子系统,其中所述fastclk的频率足够高,以允许所述 存储器子系统在储存通过所述一个或更多个数据流接口输入的所述数据流的同时满足通 过所述随机存取接口接收的存储器存取请求。
- 1415. 如权利要求1所述的存储器子系统,其中所述存储器包括一个或更多个存储器集 成电路。
- 1516. 如权利要求1所述的存储器子系统,其中所述数据流接口从视频摄像机接收数据 流,并且所述随机存取接口从视频编解码器接收随机存取存储器请求。
- 1617. 一种多路存取存储器,所述多路存取存储器包括: 存储器单元的直线网格,每个存储器单元储存信息比特; 行解复用器,所述行解复用器通过沿着行解复用器在存储器单元的直线网格内选择的 一行存储器单元的移位操作来将比特从数据流引导到存储器单元中;以及 行解复用器和列解复用器对,所述行解复用器和列解复用器对将在随机存储器存取写 请求中接收的比特引导到由所述行解复用器和列解复用器对选择的存储器单元中,并且从 由所述行解复用器和列解复用器对选择的存储器单元检索通过随机存储器存取读请求所 请求的比特。 1 如权利要求17所述的多路存取存储器,其中所述多路存取存储器包括一个或更多 个行解复用器,所述行解复用器每个通过沿着由所述行解复用器选择的一行存储器单元的 移位操作来将比特从数据流引导到存储器单元中,每个行解复用器与所述存储器单元的直 线网格的不同区域相关联。
- 1719. 如权利要求18所述的多路存取存储器,其中所述多路存取存储器被同时存取,以 通过移位操作并且通过随机存储器存取写请求和随机存储器存取读请求从数据流写入至 少一行。
- 1820. 如权利要求17所述的多路存取存储器,其中存储器单元包括反向数据储存单元, 所述反向数据储存单元静态地没有更新地储存在数据储存单元输入上接收的数据比 CN 102369552 Β 特的补码,并且连续地将储存的所述数据比特的补码输出到存储器单元移位输出。
- 1921. 如权利要求20所述的多路存取存储器,其中所述数据储存单元为包括两个反向器 的触发器。
- 2022. 如权利要求21所述的多路存取存储器,其中所述存储器单元还包括: 第一移位操作晶体管,所述第一移位操作晶体管的源极输入连接至存储器单元数据流 输入,并且所述第一移位操作晶体管的栅极连接至存储器单元Shiftl输入; 第二移位操作晶体管,所述第二移位操作晶体管的漏极输出连接至所述数据储存单元 输入,并且所述第二移位操作晶体管的栅极连接至存储器单元Shift2输入; 连接的反向器,所述反向器的输入连接至所述第一移位操作晶体管的漏极输出,并且 所述反向器的输出连接至所述第二移位操作晶体管的源极输入; 存储器单元随机存取输出晶体管,所述存储器单元随机存取输出晶体管的栅极连接至 存储器单元存储器单元随机存取读输入,所述存储器单元随机存取输出晶体管的源极输入 连接至所述存储器单元移位输出,并且所述存储器单元随机存取输出晶体管的漏极输出连 接至存储器单元随机存取读输出;以及 存储器单元随机存取输入晶体管,所述存储器单元随机存取输入晶体管的栅极连接至 存储器单元存储器单元随机存取写输入,所述存储器单元随机存取输入晶体管的源极输入 连接至存储器单元随机存取输入,并且所述存储器单元随机存取输入晶体管的漏极输出连 接至所述数据储存单元输入。
- 2123. 如权利要求22所述的多路存取存储器,其中通过以下方式将比特从所述存储器单 元数据流输入移位到所述存储器单元中,所述方式即,拉高Shiftl存储器单元输入,拉低 Shiftl存储器单元输入,并且拉高Shift2存储器单元输入。
- 2224. 如权利要求22所述的多路存取存储器,其中通过拉高所述存储器单元随机存取写 输入来将比特从所述存储器单元随机存取输入写入到所述存储器单元中。
- 2325. 如权利要求22所述的多路存取存储器,其中通过拉高所述存储器单元存储器单元 随机存取读输入来在所述存储器单元随机存取读输出上从所述数据储存单元读取比特。
- 2426. 如权利要求22所述的多路存取存储器,其中在所述存储器单元的直线网格中,存 储器单元行用作移位寄存器,其中除了每行中最后的存储器单元之外的存储器单元的存储 器单元移位输出都连接至所述行中的下一个存储器单元的所述存储器单元数据流输入。
- 2527. 如权利要求17所述的多路存取存储器,与存储器控制器耦合,以构成存储器子系 统。 CN 102369552 Β
Independent claims25
287 paragraphs, as filed
Memory subsystem
[0001] Cross-reference of related applications: This application is a partial continuation application of U.S. Application No. 12/322, 571 filed on February 4, 2009, and U.S. Application No. 12/322, 571 is on January 12, 2009 Part of the submitted US application No. 12/319, 750 continues to apply.
Technical field
[0002] The present invention relates to electronic storage, and in particular, to a memory subsystem in a highly parallel pipelined integrated circuit calculation engine or accessed by a highly parallel pipelined integrated circuit calculation engine, or in various types of electronic devices A memory subsystem used in any of or accessed by any of various different types of electronic devices.
Background technique
[0003] Computing machines are undergoing rapid development. Early electronic computers were usually fully sequential processing machines that executed a stream of instructions one by one, which together constituted a computer program. For many years, electronic computers have typically included a single main processor that can quickly execute a relatively small set of simple instructions, including memory fetching, memory storage, arithmetic, and logic instructions. The solution of a computing task is solved by programming a set of instructions and then executing the program on a single-processor computer system.
[0004] In the relatively early days of the development of electronic computers, various ancillary and supporting tasks began to be moved from the main processor to dedicated auxiliary processing components. As an embodiment, a separate I/O controller was developed to offload many repetitive and computational bandwidth-consuming tasks associated with the exchange of information between the main memory and various external devices, including Human capacity storage device, communication device, display device and user input device. This merging of multiple processor elements into a single-main processor computer system is the beginning of a trend to increase computational parallelism.
[0005] Parallel computing is currently the main trend in the design of modern computing machines. At one extreme, each processor core usually provides simultaneous parallel execution of multiple instruction streams and assembly-line simultaneous execution of multiple instructions. Most computers, including personal computers, now incorporate at least two processor cores, and often many processor cores, in each monolithic integrated circuit. Each processor core can execute multiple instruction streams relatively independently. An electronic computer system can contain multiple multi-core processors, and can be gathered together into a large distributed computing network that includes tens to thousands to hundreds of thousands of separate computer systems that communicate with each other , And each computer system performs one or more divisible parts of a large distributed computing task.
[0006] As computers have developed toward parallel and massively parallel computing systems, many of the most difficult and annoying problems associated with parallel computing have been found to be related to the decomposition of large-scale computing tasks into relatively independent subtasks Link, each subtask can be executed by a different processing entity. When the problem is not properly decomposed, or when the problem cannot be decomposed, for parallel execution, the use of parallel computer machines usually provides little or no benefit, and in the worst case, will actually cause Slower execution than can be achieved by traditional software implementations executed on a single-processor computer system. When multiple computing entities compete for shared resources, or based on calculation results jointly generated by other processing entities, it will consume a lot of computing and communication resources to manage the parallel operations of multiple computing entities. Often, the communication overhead and computational overhead may be greater than the parallel calculations performed on multiple processors or other computing entities.
CN 102369552 Β
The benefits of the method are much more important. In addition, parallel computing can involve a lot of financial costs, as well as a lot of power consumption and cooling costs.
[0007] Therefore, although judging from biological systems, parallel computing seems to be a logical method for efficiently computing many computing tasks, and the development trend has appeared within a short period of time in the development of electronic computers, but parallel computing is still associated with many complexities. , Cost and shortcomings are associated. Although many problems can theoretically benefit from parallel computing methods, the currently available technologies and hardware for parallel computing generally cannot provide cost-effective solutions for many computing problems, especially for those who need to be limited by size constraints, Complicated calculations performed in real-time within a device subject to heat dissipation constraints, power consumption constraints, and cost constraints. For this reason, computer scientists, electrical engineers, researchers, and developers in many computing-dominated fields, manufacturers and sellers of electronic devices and electronic computers, and finally users of electronic devices and electronic computers all recognize the need to continue to develop efficiently. Realize a new method of parallel computing engine for solving practical problems. Specifically, computer scientists, electrical engineers, researchers, and developers in many computing-dominated fields, manufacturers and vendors of electronic devices and electronic computers, and others seek those that can be used in or associated with parallel computing engines. An efficient, low-power, and cost-effective subsystem that includes an efficient, low-power, and cost-effective memory subsystem.
Summary of the invention
[0008] The embodiment of the present invention relates to a memory implemented in a parallel pipelined integrated circuit implementation of a calculation engine or connected to a parallel pipelined integrated circuit implementation of a calculation engine and accessed by these parallel pipelined integrated circuit implementations Subsystem, the calculation engine is designed to solve complex calculation problems. Additional embodiments of the present invention relate to memory subsystems implemented in various different types of electronic devices or connected to and accessed by various different types of electronic devices. One embodiment of the invention includes a memory controller and one or more separate memory devices, the memory controller being implemented in a first integrated circuit or other electronic system. An alternative embodiment of the present invention incorporates a memory controller in one or more memory devices that are connected to a calculation engine implemented by an integrated circuit or another electronic device and are used by the calculation engine. Or another electronic device to access. In an alternative embodiment of the present invention, the memory controller and the memory are integrated together in a computing engine or another electronic device. Alternative embodiments of the present invention include a multiple access memory interfaced with a simpler memory controller to connect to a calculation engine or other electronic device or be integrated in the calculation engine or other electronic device.
Description of the drawings
[0009] FIG. 1 illustrates a digitally encoded image.
[0010] FIG. 2 illustrates two different pixel value encoding methods according to two different color and brightness models.
[0011] FIG. 3 illustrates digital coding using the Y'CrCb color model.
[0012] FIG. 4 illustrates the output of a video camera.
[0013] FIG. 5 illustrates the function of a video codec.
[0014] FIG. 6 illustrates various data objects on which video encoding operations are performed during video data stream compression and compressed video data stream decompression.
[0015] FIG. 7 illustrates the division of a video frame into two slice groups.
[0016] FIG. 8 illustrates a second level of video frame segmentation.
CN 102369552 Β
[0017] FIG. 9 illustrates the general concept of intra prediction.
[0018] FIGS. 10A-1 illustrate nine 4X4 luma block intra prediction modes.
[0019] FIGS. 11A-11D illustrate four modes for intra prediction of 16×16 luma blocks using an illustration convention similar to that used in FIGS. 10A-1.
[0020] FIG. 12 illustrates the concept of inter prediction.
[0021] FIGS. 13A-D illustrate an interpolation process for calculating the pixel value of a block in the search space of a reference frame, and the interpolation process can be considered to occur at fractional coordinates.
[0022] FIGS. 14A-C illustrate different types of frames and some of the different types of inter prediction that are possible for these frames.
[0023] FIG. 15 illustrates the generation of difference macroblocks.
[0024] FIG. 16 illustrates motion vector and intra prediction mode prediction.
[0025] FIG. 17 illustrates the decomposition, integer transformation, and quantization of difference macroblocks.
[0026] FIG. 18 provides the derivation of integer transform and inverse integer transform used in H.264 video compression and video decompression, respectively.
[0027] FIG. 19 illustrates the quantization process.
[0028] FIG. 20 provides a digital embodiment of bad coding.
[0029] Figures 21A-B provide an embodiment of arithmetic coding.
[0030] FIGS. 22A-B illustrate a common artifact and a filtering method used to improve the artifact as the final step of decompression.
[0031] Figure 23 summarizes H.264 video data stream encoding.
[0032] FIG. 24 illustrates the H.264 video data stream decoding process in a block diagram manner similar to that used in FIG. 23.
[0033] FIG. 25 is a very high-level diagram of a general-purpose computer.
[0034] FIG. 26 illustrates many aspects of the video compression and decompression process, which, when considered, provide a deep understanding of the much more computationally efficient new method of implementing the video codec according to the present invention.
[0035] FIG. 27 illustrates the basic features of the integrated circuit implementation of the video codec according to the method of the present invention.
[0036] FIG. 28 illustrates an embodiment of the present invention in which the integrated circuit 2802 includes a memory 2804, which is external in the embodiment illustrated in FIG. 27.
[0037] FIG. 29 illustrates an alternative embodiment of the present invention in which a digital video camera is included in an integrated circuit implementation of a combined video camera and video codec.
[0038] FIGS. 30-32 illustrate the overall timing and data flow within an integrated circuit implementation of the video codec according to the present invention.
[0039] FIGS. 33A-B provide block diagram illustrations of a single integrated circuit implementation of a video codec according to the present invention.
[0040] FIG. 34 illustrates the overall system timing and synchronization of a single integrated circuit implementation of the video codec according to the present invention.
[0041] FIG. 35 provides a table of various types of objects that are transferred from the video cache to the processing element along the data object bus in a single integrated circuit implementation of the video codec according to the present invention.
[0042] FIGS. 36A-B illustrate at an abstract level the operation of processing elements within a single integrated circuit implementation of a video codec that characterizes an embodiment of the present invention.
[0043] FIG. 37 illustrates a motion estimation processing element characterizing an embodiment of the present invention.
CN 102369552 Β
[0044] FIG. 38 illustrates an intra-prediction and inter-prediction processing element that characterizes an embodiment of the present invention, the processing element including a pair of processing elements.
[0045] FIG. 39 shows a block diagram of a rotten encoding processing element characterizing an embodiment of the present invention.
[0046] FIG. 40 illustrates an example of storage requirements for a video cache in the video codec implementation illustrated in FIG. 33A.
[0047] FIG. 41 illustrates the operation of the luminance macroblock circular queue (4002 in FIG. 40) during nine advanced processing cycles.
[0048] FIG. 42 illustrates an embodiment of a video cache controller of a video codec that characterizes an embodiment of the present invention.
[0049] FIG. 43 provides a table indicating an example of overall computational processing performed by each of certain processing elements of a video codec characterizing an embodiment of the present invention.
[0050] FIGS. 44A-E provide high-level VHDL definitions of various processing elements in a single integrated circuit implementation of a video codec according to an embodiment of the present invention as shown in FIG. 33A.
[0051] FIG. 45 illustrates the components and functions of the memory subsystem of a video camera that characterize various embodiments of the present invention.
[0052] FIGS. 46A-E illustrate a series of video systems that characterize embodiments of the present invention, and these video systems characterize ways to improve the integration between the subsystems of the video system.
[0053] FIG. 47 illustrates the generalized interfaces provided by the memory controller embodiment of the present invention to the camera, video codec, and memory.
[0054] FIGS. 48A-H illustrate the components and operation of the components of a memory controller that characterizes an embodiment of the present invention.
[0055] FIGS. 49A-C illustrate an implementation of the arbiter discussed with reference to FIGS. 48B-C, which is a component of a memory controller that characterizes an embodiment of the present invention.
[0056] FIG. 50 provides a simple illustration regarding timing considerations for a memory controller arbiter implemented in a memory controller that characterizes an embodiment of the present invention.
[0057] FIGS. 51-54 provide schematic diagrams of a memory controller characterizing an embodiment of the present invention.
[0058] FIG. 55 illustrates the operation of a multiple access memory characterizing an embodiment of the present invention.
[0059] FIG. 56 abstractly illustrates the operation of a multiple access memory characterizing an embodiment of the present invention.
[0060] FIG. 57 illustrates a multi-plane memory system according to an embodiment of the present invention.
[0061] FIG. 58 illustrates the division of a memory partition associated with each camera in a multiple access memory according to an embodiment of the present invention.
[0062] FIG. 59 illustrates writing a frame to a multiple access memory according to various embodiments of the present invention.
[0063] FIG. 60 illustrates a signal inverter.
[0064] FIG. 61 shows a schematic diagram of a memory unit or a memory unit of a multiple access memory used to characterize an embodiment of the present invention and a symbolic representation of the memory unit.
[0065] FIGS. 62A-C illustrate the shifting of data into a memory cell of a multiple access memory that characterizes an embodiment of the invention.
[0066] FIGS. 63A-C illustrate writing the Boolean value "0" into the memory cell currently storing the Boolean value "1" according to an embodiment of the present invention.
CN 102369552 Β
[0067] FIGS. 64A-B illustrate the output of the value currently stored in the memory cell of the multiple access memory to the output signal line according to an embodiment of the present invention.
[0068] FIGS. 65A-B illustrate the writing of values into a memory cell characterizing an embodiment of the present invention through two input signal lines.
[0069] FIGS. 66A-B show an embodiment of a 4X4 memory storage array using 16 memory cells of the type illustrated in FIG. 61 and a symbolic representation of the 4X4 memory storage array.
[0070] FIG. 67 shows a schematic diagram of a larger capacity memory based on a 4X4 memory storage array (such as the 4X4 memory storage array shown in FIG. 66A) according to an embodiment of the present invention.
[0071] FIG. 68 illustrates a schematic diagram of a two-dimensional access decoder block shown in the memory storage array illustrated in FIG. 67 as an embodiment of the present invention.
[0072] FIG. 69 illustrates a memory controller characterizing an embodiment of the present invention connected to the multiple access memory interface discussed above with reference to FIGS. 65-68.
Detailed ways
[0073] Embodiments of the present invention relate to a memory subsystem that can be implemented in a computing engine or can be connected to a computing engine, the computing engine with low power consumption, low heat dissipation, large computing bandwidth and low task execution latency (latency) Perform complex calculation tasks. The calculation engine is implemented as individual integrated circuits or chips, which are characterized by a number of simultaneously operating processing elements in accordance with the present invention to provide highly parallel calculations. The following methods make it possible to effectively use the currently executing processing elements, that is, to appropriately decompose complex computing tasks, to efficiently access shared information and data objects in the integrated circuit, and to efficiently and hierarchically control processing Tasks and subtasks.
[0074] The processing elements access the computing objects they operate on through an object bus, which interconnects the processing elements with the on-board object cache. In many embodiments, the on-board object cache is connected to or coupled to a larger-capacity object memory through an object memory controller. In some embodiments of the present invention, the larger-capacity object memory may be implemented as an external component. In certain embodiments of the present invention, the calculation control implemented by the calculation engine of the present invention is provided by a microprocessor controller based on a relatively low frequency clock, where one or more high frequency clock signals control the processing within the processing element. In some embodiments of the present invention, the processing elements are logically arranged into an assembly line-style pipeline, where computing objects are usually processed sequentially by the processing elements along the pipeline, and are moved between processing elements and/or from objects. The cache moves back and forth. The calculation object, rather than the arbitrary size of data units (such as bytes or words), is used as the center to organize processing element calculation, cache access, memory access and data transmission.
[0075] A large number of different computing tasks can be solved by the design and development of highly parallel integrated circuit implementations of the computing engine according to the embodiments of the present invention. In the following, a parallel pipelined integrated circuit implementation of a video codec as a specific embodiment of the present invention will be discussed. Various alternative implementations of the integrated circuit implementation of the video codec can be utilized in a variety of electronic devices, including mobile phones equipped with video cameras, digital video cameras, personal computers, surveillance equipment , Remote sensors, aircraft and spacecraft, and many other types of equipment. It is emphasized here that throughout the following discussion, video codec implementations are specific examples of many different parallel pipelined integrated circuit computing engines that characterize implementations of the present invention.
[0076] The described parallel integrated circuit implementation of the video codec is designed to perform complex computational tasks. The following discussion is organized into six subsections of 264 compressed video signal decompression standards; (2) Principles of parallel integrated circuit design for solving complex computing tasks according to the present invention; (3) Implemented as a single integrated circuit according to the present invention of
CN 102369552 Β
Η. 264 video codec; (4) according to the present invention is a video system embodiment characterized by increased integration with the memory subsystem; (5) a first family of memory sub-systems characterizing a set of embodiments of the present invention System; and (6) a family of memory subsystems characterizing the second set of embodiments of the present invention. It should be noted that although the examples are mainly provided in the context of the H·264 standard, these are merely examples, and the present invention is by no means limited to the H·264-based implementation. In the first subsection below, a general description of the calculation tasks performed by a specific embodiment of the parallel pipelined integrated circuit calculation engine is given. The described implementation is a video codec that compresses the original video signal and decompresses the compressed video signal according to the H.264 or MPEG-4 AVC, compressed video signal decompression standard. For readers who are already familiar with the decompression standard of H·264 compressed video signals, you can skip the first section. In the second subsection, the principle of parallel integrated circuit design according to the embodiment of the present invention is described. The parallel integrated circuit design can be applied to any of many complex computing tasks. In the third subsection, the pair is implemented as a single integrated circuit Η. 264 video codec is described in detail. In the fourth subsection, various implementations of the single integrated circuit video codec are discussed, which provide ways to improve the integration of the memory subsystem with the video codec and ultimately with the imaging system. In the fifth section, the first family of RAM-based memory subsystems that characterize the embodiments of the present invention are discussed. Finally, in the sixth subsection, the second family of high-efficiency memory systems that characterize the embodiments of the present invention are discussed.
[0077] The first subsection: H.264 compressed video signal decompression standard
[0078] This first subsection provides an overview of the H.264 compressed video signal decompression standard. This subsection provides a description of the computational problems solved by the specific implementation of the parallel pipelined integrated circuit computing engine that characterizes the embodiment of the present invention. Those readers who are familiar with H·264 can skip the first section of this book and continue to the second section below.
[0079] FIG. 1 illustrates a digitally encoded image. Digitally encoded images can be still photos, video frames, or any of various graphic objects. Generally, a digitally encoded image includes a digitally encoded sequence of numbers that together describe the rectangular image 101. A rectangular image has a horizontal dimension 102 and a vertical dimension 104, and the ratio of the horizontal dimension 102 and the vertical dimension 104 is called the "aspect ratio" of the image.
[0080] A digitally encoded image is broken down into tiny display units, which are called "pixels." In FIG. 1, the small part 106 in the upper left corner of the displayed image is shown enlarged twice. Each enlargement step is a 12 times enlargement, generating a final 144 times enlargement of the extremely small part of the upper left corner of the digitally encoded image 108. When magnified by 144 times, it is seen that a small part of the displayed image is divided into small squares by a linear coordinate grid, and each small square (for example, square 110) corresponds to a pixel or represents a pixel. Video images are digitally encoded as a series of data units, each data unit describing the light-emitting characteristics of a pixel in the displayed image. A pixel can be considered as a unit in a matrix, and the position of each pixel is described by a horizontal coordinate and a vertical coordinate. Alternatively, the pixels can be considered as a long linear sequence of pixels generated in a raster scan order or some other predefined order. Generally, logical pixels in a digitally encoded image are converted into light emitted from one or several extremely small display elements of the display device. The number that digitally encodes the value of each pixel is converted into one or more electronic voltage signals to control the display unit to emit light with suitable hue and intensity, so as to be based on the pixel value encoded in the digitally encoded image To control all the display units, the display device faithfully reproduces the encoded images for human viewers to watch. Digitally encoded images can be displayed on The cathode ray tube, LCD or plasma display device and other such light-emitting display devices contained in televisions, computer display monitors can be printed on paper or composite film by computer printers, and can be sent to remote devices through digital communication media, Can be stored on mass storage devices and computer memory, and can be processed by various image processing applications.
CN 102369552 Β
[0081] There are various methods and standards for encoding color and emission intensity information into data units. Figure 2 illustrates two different pixel value encoding methods according to two different color and brightness models. The first color model 202 is represented by a cube. The volume in the cube is indexed by three orthogonal axes, and the three orthogonal axes are R'axis 204, Β<sup>ζ</sup>The shaft 206 and the L shaft 208. In this embodiment, each axis is increased by 256 increments, the 256 increments corresponding to all possible values of 8-bit bytes, among which the replaceable R, G, and B models are used less or more Increase for most purposes. The volume of the cube characterizes all possible color and brightness combinations that can be displayed by the pixels of the display device. The R'axis, G, axis, and B, axis correspond to the red component, blue component, and green component of the colored light emitted by the pixel. The luminous intensity of the display unit is generally a non-linear function of the voltage supplied to the data unit. In the RGB color model, the G component value 127 in the byte-coded G component will guide half of the maximum voltage that can be applied to a display unit to be applied to a specific display unit. However, when half of the maximum voltage is applied to the display unit, the emission brightness may significantly exceed half of the maximum brightness emitted at the full voltage. For this reason, a nonlinear transformation is applied to the increments of the RGB color model to generate R'G, B, increments of the color model, so that the scaling is linear with respect to the perceived brightness. When up to 256 brightness levels can be designated for each of the red component, the blue component, and the green component of the light emitted by the pixel, the encoding for a specific pixel 210 may include three 8-bit bytes, for a total of 24 bits. When a larger number of brightness levels can be specified, a larger number of bits are used to characterize each pixel, and when a smaller number of brightness levels can be specified, a smaller number of bits can be used. Bits to encode each pixel.
[0082] Although R, G, B, color models are relatively easy to understand, especially considering the red-emitting phosphor, green-emitting phosphor, and blue-emitting phosphor structure of the display unit in the CRT screen, it is very important for video signal compression and decompression. For compression, a variety of related but different color models are more useful. One such alternative color model is the Y'CrCb color model. The Y'CrCb color model can be abstractly characterized as a double vertebral body volume 212. The double vertebral body volume 212 has a central horizontal plane 214 containing the orthogonal Cb axis and the Cr axis, and the length of the double vertebral body corresponding to the Y axis The vertical axis 216. In this color model, the Cr axis and the Cb axis are color specifying axes, where the horizontal middle plane 214 represents all possible hues that can be displayed, and the Y'axis represents the brightness or intensity of the displayed hues. Specifying R'G, B, the values of the red, blue, and green components in the color model can be directly transformed into equivalent Y'CrCb values through a simple matrix transformation 220. Therefore, when the 8-bit number is used to encode the Y, Cr, and Cb components emitted by the display unit according to the Y, CrCb color model, the 24-bit data unit 222 can be used to encode the value of a single pixel.
[0083] For image processing, when the Y, CrCb color model is used, a digitally encoded image can be considered as three separate pixated planes superimposed on each other. Figure 3 illustrates digital coding using the Y'CrCb color model. The digitally encoded image as shown in FIG. 3 can be considered as Y, image 302 and two chrominance images 304 and 306. The Y, plane 302 basically encodes the brightness value of the image, and is equivalent to the monochrome representation of the digitally encoded image. The two chromaticity planes 304 and 306 together characterize the hue or color at each point in the digitally encoded image. For many video processing and video image storage purposes, it is convenient to extract the Cr plane and the Cb plane to generate the Cr plane 308 and the Cb plane 310 with half the resolution. In other words, instead of storing the intensity value and two chromaticity values of each pixel, the intensity value is stored for each pixel, and a pair of chromaticity values are stored for each 2X2 square containing four pixels. Therefore, all four pixels in the upper left corner of the image 312 are encoded to have the same Cr value and Cb value. For each 2 X 2 area of the image 320, four intensity values 322 and two chrominance values 324 (48 bits in total, or in other words, 12 bits per pixel) can be used to perform this area. Digital coding.
[0084] FIG. 4 illustrates the output of a video camera. Video camera 402 is characterized as lens 404 and electronic output
CN 102369552 Β
A sensor 406 is produced. The video camera generates a clock signal 408, and the rising edge of each pulse of the clock signal 408 corresponds to the beginning of the next data packet (eg, data packet 410). In the embodiment shown in FIG. 4, each data packet contains an 8-bit intensity value or chrominance value. The digital video camera also generates a line signal or line signal 412, which is high during a time period corresponding to the output of the entire line of the digitally encoded image. The digital video camera additionally outputs a frame signal 414, which is high during the period of time when one digital image or one frame is output. The clock signal, the line signal, and the frame output signal together specify the time for outputting each intensity value or chrominance value, outputting each line of the frame, and outputting each frame of the video signal. The data output 416 of the video camera is shown in more detail as a packet sequence 420 at the bottom of FIG. 4. Referring to the 2X2 pixel area shown in FIG. 3 (320 in FIG. 3) and using the same indexing convention as the intensity value 322 and chrominance value 324 for encoding this area in FIG. 3, the indexing convention in FIG. 4 The content of the data stream 420 can be understood. The two intensity values of the 2X2 square area of pixels 422-426 are sent as part of the pixel values of the first row together with the first set of two chromaticity values 428-429 of the 2X2 square area of pixels, of which two chromaticity values 428-429 is between the first two intensity values 422-423 Is sent from time to time. Subsequently, the chrominance values 430-431 are repeated between the second pair of intensity values 424 and 426 as part of the pixel intensity of the next row. The repetition of chrominance values facilitates certain types of real-time video data stream processing. However, the second pair of chroma values 430-431 are redundant. As discussed with reference to FIG. 3, the chromaticity plane is extracted so that only two chromaticity values are associated with each 2×2 area containing four pixels.
[0085] FIG. 5 illustrates the function of a video codec. As discussed above with reference to Figures 1-4, the video camera 502 generates a stream 504 of digitally encoded video frames. At 30 frames per second, assuming a frame of 1920X1080 pixels, and assuming that each pixel uses 12-bit encoding, the video camera generates 93 megabytes of data per second. One minute of continuous video capture will generate 5.5 gigabytes of data. Small handheld electronic devices manufactured according to currently available designs and technologies cannot process, store, and/or send data at this rate. In order to generate a manageable data transmission rate, the video codec 506 is used to compress the data stream output from the camera. The Η.264 standard provides a video compression ratio of approximately 30:1. Thus, the video codec 506 compresses the 93MB/s data stream entered from the camera to generate a compressed video data stream 508 of approximately 3MB/s. In contrast to the original video data stream generated by the camera, the video codec outputs a compressed video data stream at a data rate that can be processed for storage or transmission by a handheld device. The video codec can also receive the compressed video data stream 510 and decompress the compressed data to generate an output original video data stream 512 for use by the video display device.
[0086] Since the video signal usually contains a relatively large amount of redundant information, the video codec can achieve a compression ratio of 30:1. As an embodiment, the video signal generated by shooting two children throwing a ball back and forth contains a relatively small amount of rapidly changing information and a relatively large number of static or slowly changing objects. The rapidly changing information is the image of the child and the ball. The static or slowly changing objects include background scenery and lawns on which children play. While the image of the child's silhouette and the ball may change significantly from frame to frame during the shooting process, the background object may remain relatively constant throughout the shooting period or at least for a relatively long period of time. In this case, most of the information encoded in the frames following the first frame can be completely redundant. Video compression technology is used to identify redundant information and efficiently encode the redundant information, thus greatly reducing the total amount of information included in the compressed video signal.
[0087] The compressed video stream 508 (520) is shown in more detail in the lower part of FIG. 5. According to the H.264 standard, the compressed video stream includes a network abstraction layer ("NAL") packet sequence, such as NAL packet 522. Each NAL packet includes an 8-bit header, such as the header 524 of the NAL packet 522. The first bit 526 must always be zero, the next two bits 528 indicate whether the data contained in the packet is associated with the reference frame, and the last 5 bits 530 together constitute a type field, which indicates the type of the packet and its data The nature of the payload. Packet types include packets containing encoded pixel data and encoded metadata,
CN 102369552 Β
It also includes packages that characterize various types of separators, the encoded metadata describes how the data part has been encoded, and the separators include sequence end separators and stream end separators. The body 532 of the NAL packet usually contains encoded data.
[0088] FIG. 6 illustrates various data objects on which video encoding operations are performed during video data stream compression and compressed video data stream decompression. From the viewpoint of video processing, the video frame 602 is considered to be composed of a two-dimensional array of macroblocks 604, and each macroblock includes an array of 16×16 data values. As discussed above, video compression and decompression generally operate independently on Y'frames containing intensity values and chrominance frames containing chrominance values. The human eye is usually much more sensitive to changes in brightness than to spatial changes in color. Therefore, as discussed above, the initial useful compression is obtained simply by extracting two chromaticity planes. Assuming the 8-bit representation of the intensity value and the chrominance value, before the extraction, the 2X2 square pixel can be characterized by 12 bytes of coded data. After extraction, the four pixels of the same 2X2 square can be characterized by only 6 bytes of data. Therefore, by reducing the spatial resolution of the color signal, 2: The compression ratio of 1. Although a macro block is the basic unit for performing compression and decompression operations, for some compression and decompression operations, the macro block can be further divided. The intensity macro block or the luminance macro block each contains 256 pixels 606, but can be divided to generate 16×8 partitions 608, 8×16 partitions>8×8 partitions 612, 8×4 partitions 614, 4×8 partitions 616, and 4×4 partitions 618. Similarly, each chroma macroblock contains 64 encoded chroma values 620, but can be further divided to generate 8X4 partitions 622, 4X8 partitions 624, 4X4 partitions 626, 4X2 partitions 628, 2X4 partitions 630, and 2X2 partitions 632. In addition, 1X4, 1X8, and 1X16 pixel vectors can be used in some operations.
[0089] According to the H.264 standard, each video frame can be logically divided into slice groups, where the division is specified by a slice-group map. Many different types of slice group partitions can be specified by suitable slice group mapping. Figure 7 illustrates the division of a video frame into two slice groups. The video frame 702 is divided into a first checkerboard slice group 704 and a supplementary checkerboard slice group 706. Both the first slice group and the second slice group contain an equal number of pixel values, and each contains half of the total number of pixel values in the frame. According to a substantially arbitrary mapping function, a frame can be divided into substantially any number of slice groups, each slice group including a substantially arbitrary portion of all pixels.
[0090] FIG. 8 illustrates a second level of video frame segmentation. Each slice group (for example slice group 802) can be divided into several slices 804-806. Each slice contains several adjacent pixels in raster scan order (adjacent within the slice group, but not necessarily adjacent within a frame). The slice group 802 may be an entire video frame, or may be a partition of a frame according to any slice group division function. Some of the compression and decompression operations can be performed slice by slice.
[0091] In summary, video compression and decompression techniques are performed on video frames and various subsets of video frames, the subsets including slices, macroblocks, and macroblock partitions. Generally, the intensity plane object or the luminance plane object is operated independently of the chrominance plane object. Since the chroma plane is decimated by half in each dimension, and the overall 4:1 compression, the size of the chroma macro block and the macro block partition is usually half of the size of the luma macro block and the luma macro block partition.
[0092] As suggested by the H.264 standard, the first step in video compression is to use one of two different common prediction techniques to, in one case, from adjacent macroblocks or macros in the same frame Block partition predicts the pixel value of the currently considered macro block or macro block partition, and in another case, predicts the pixel value of the currently considered macro block or macro block partition from the spatially adjacent macro block or macro block partition , The spatially adjacent macroblocks or macroblock partitions appear in the frames before or after the frame of the macroblock or macroblock partition that is being predicted. The first type of prediction is spatial prediction, which is called "intra prediction". The second type of prediction is temporal prediction, which is called "inter-frame prediction". Intra prediction is the only type of prediction that can be used for certain frames, which are called "reference frames". Intra prediction is also the default prediction used when encoding macroblocks. For macroblocks of non-reference frames, inter prediction is first tried. When the inter-frame prediction is successful, the intra-frame prediction is not used for the macroblock.
CN 102369552 Β
However, when inter prediction fails, intra prediction can be used as the default prediction method.
[0093] FIG. 9 illustrates the general concept of intra prediction. Consider the macroblock C 902 that occurs during the macroblock-by-macroblock compression of a video frame. As discussed above, the 16×16 luma macroblock 904 can be encoded using 256 bytes. However, if the content of a macroblock can be calculated from adjacent macroblocks in the image, then a considerable amount of compression is theoretically possible. For example, consider the four neighboring macroblocks of macroblock C 902 currently under consideration. The four macro blocks include a left macro block 904, an upper left diagonal macro block 906, an upper macro block 908, and an upper right diagonal macro block 910. If one of a number of different prediction functions 912 can be used to calculate the pixel value in C based on one or more of these neighboring macroblocks, the content of the macroblock can be simply encoded for prediction The number specifier or indicator of a function. If the number of prediction functions is less than or equal to, for example, 256, the specifier or indicator for the selected prediction function may be encoded in single-byte information. Therefore, if the selected one of the 256 possible prediction functions can be used to calculate the content of the macroblock from the neighborhood of the macroblock, a rather amazing 256 can be achieved: The compression ratio of 1. Unfortunately, because there are too many possible macroblocks to accurately predict with only 256 prediction functions, the spatial prediction method for H·264 compression usually does not achieve such a magnitude compression ratio. For example, when each pixel is encoded with 12 bits, there are K = 4096 different possible pixel values and 4096<sup>256</sup>Different possible macroblocks. However, for H·264 video compression, especially for relatively static video signals with large image regions that do not change rapidly and with relatively uniform intensity and color, intra-frame prediction can significantly benefit the overall compression ratio.
[0094] Η can be performed according to nine different modes for 4X4 luminance macroblocks or according to four different modes for 16X16 luminance macroblocks. 264 intra prediction. Figures 10A-I illustrate nine 4 X 4 luma block intra prediction modes. The illustration convention used in all these figures is similar and is described with reference to FIG. 10A. The 4X4 luminance macroblock being predicted is represented in the figure by the 4X4 matrix 1002 at the bottom right of the figure. In this way, the pixel value 1004 on the uppermost left side in the 4×4 matrix being predicted in FIG. 10A contains the value "A". The unit adjacent to the 4X4 brightness block represents the pixel value in the adjacent 4X4 brightness block in the image. For example, in FIG. 10A, the values "A" 1006, "B" 1007, "C" 1008, and "D" 1009 are the data values contained in the 4X4 luminance block directly above the 4X4 luminance block 1002 being predicted. Similarly, the units 1010-1013 represent the pixel values in the last vertical column of the 4X4 luma block to the left of the 4X4 luma block being predicted. In the case of the mode 0 prediction illustrated in FIG. 10A, the value in the last row of the upper adjacent 4X4 luminance block is vertically copied downward to the column of the currently considered 4X4 luminance block 1002. Therefore, in FIG. 10A, the mode 0 prediction constitutes the downward vertical prediction represented by the downward pointing arrow 1020 shown in FIG. 10A. In FIGS. 10B-10I, the same graphical convention as that used in FIG. 10A is used to show the remaining eight types of intraframes used to predict 4X4 luminance blocks. The prediction modes, and therefore, these eight intra prediction modes are completely independent and self-explanatory. Each mode other than mode 2 can be considered as a space vector indicating the direction in which pixel values in adjacent 4×4 blocks are converted to the block being predicted.
[0095] FIGS. 11A-11D illustrate four modes for intra prediction of 16×16 luma blocks using an illustration convention similar to that used in FIGS. 10A-1. In Figures 11A-D, the block being predicted is the 16X16 block in the lower right part of the matrix 1102, the leftmost vertical column 1104 is the rightmost vertical column of the left-adjacent 16X16 luminance block, and the top horizontal row 1106 is the upper-adjacent The bottom row of the 16X16 luma block. The upper leftmost unit 1110 is the lower right corner unit of the upper left diagonal 16×16 luminance block. The 16X16 prediction mode is similar to a subset of the 4X4 intra prediction mode. Except for the mode 4 shown in Figure 11D, the mode 4 is a relatively complex planar prediction mode, which is the next row of the 16X 16 luma block adjacent from above. All pixels in the rightmost vertical column of the 16X16 luminance block adjacent to the left are calculated for each pixel's predicted value. Generally, the mode that generates the closest approximation of the current block being intra-predicted is selected as the intra-frame applied to the currently considered block
CN 102369552 Β
Forecast mode. The predicted pixel value can be compared with the actual pixel value; the pixel value uses any of a variety of comparison metrics, including the average pixel value difference between the predicted block and the considered block, and the mean square error of the pixel value , Variance and other such measures.
[0096] FIG. 12 illustrates the concept of inter prediction. As discussed above, inter prediction is temporal prediction, and can be considered as motion-based prediction. For illustration purposes, consider the current frame 1202 and the reference frame 1204 that occurs before or after the current frame in the video signal. At the current moment of video compression, the current macroblock 1206 needs to be predicted from the content of the reference frame. An embodiment of the process is illustrated in FIG. 12. In the reference frame, for the current frame, the reference point 1210 is selected as the coordinates of the currently considered block 1206 applied to the reference frame. In other words, the process starts when the currently considered block in the current frame is at an equivalent position in the reference frame. Then, in the bounded search space indicated by the thick solid line 1212 square in FIG. 12, each block in the search area is compared with the currently considered block in the current frame to identify the search area 1212 of the reference frame 1204 The most similar block in the currently considered block. If the difference between the content of the block with the closest pixel value in the search area and the currently considered block is lower than the threshold value, the closest block selected from the search area predicts the content of the currently considered block. The block selected from the search area may be an actual block, or may be an estimated block at fractional coordinates relative to a linear pixel grid, wherein the pixel value in the estimated block is interpolated from the actual pixel value in the reference frame. Therefore, make Using inter-frame prediction, instead of encoding the currently considered macroblock 1206 into 256 pixel values, the currently considered macroblock 1206 can be encoded as the identifier of the reference frame and the digital representation of the vector, the vector pointing from the reference point 1210 The macro block selected from the search area 1212. For example, if it is found that the selected interpolation block 1214 most closely matches the currently considered block 1206, the currently considered block may be encoded as an identifier of the reference frame 1204 and a digital representation of the vector 1216, such as in the video signal The offset between the frame and the current frame, the vector 1216 represents the spatial displacement of the selected block 1214 from the reference point 1210.
[0097] A variety of different metrics can be used to compare the content of the actual block or interpolation block in the search area of the reference frame 1212 with the content of the currently considered frame 1206, the metrics including the average absolute value between pixel values. Pixel value difference or mean square error. The C++-like pseudo code 1220 provided in FIG. 12 as an alternative description of the aforementioned inter-frame prediction process. The encoded displacement vector is called a motion vector. The spatial displacement of the selected block from the reference point in the reference frame corresponds to the temporal displacement of the currently considered macroblock in the video stream, which usually corresponds to the actual motion of the object in the video image.
[0098] FIGS. 13A-D illustrate an interpolation process for calculating the pixel value of a block in the search area of a reference frame, and the interpolation process can be considered to occur at fractional coordinates. The Η.264 standard allows 0. 25 Resolution. Consider the 6×6 block of pixels 1302 on the left of FIG. 13A. The interpolation process can be considered as the calculation of the translational expansion of the actual pixels in two dimensions and the interpolation between the expanded pixels. Figures 13A-D illustrate the calculation of higher resolution interpolation values between the middle four pixels 1304-1307 in a 6X6 block of actual pixel values. The right side of Figure 13A illustrates the expansion 1310. In this embodiment, the pixel values 1304-1307 have been spatially expanded in two dimensions, and 21 new units have been added to form a 4X4 matrix with original pixel values 1304-1307. The remaining pixels of the 6X6 matrix of pixels 1302 have also been translated and expanded. FIG. 13B illustrates an interpolation process that generates an interpolation 1312 between actual pixel values 1304 and 1306. As shown by the dashed line 1314 in FIG. 13B, a vertical filter is applied along a column of pixel values, the pixel values including the original pixel values 1304 and 1306. Calculate the interpolation value Y 1312 according to formula 1316. In this embodiment, according to formula 1322, the value Υ, 1320 is interpolated by linear interpolation of two vertically adjacent values. The interpolation 1324 can be calculated similarly by linear interpolation between the values 1312 and 1306. The vertical filter 1314 can be similarly applied to calculate the interpolation in the column containing the original values 1305 and 1307. Figure 13C
CN 102369552 Β
The illustration illustrates the calculation of the interpolation in the horizontal line between the original values 1304 and 1305. In this embodiment, similar to the application of the vertical filter in FIG. 13B, the horizontal filter 1326 is applied to the actual pixel value. The intermediate point interpolation is calculated by formula 1328, and the quarter point value on either side of the intermediate point can be obtained by linear interpolation according to formula 1330 and a similar formula for right interpolation between the intermediate point and the original value 1305 . The same horizontal filter can be applied to the last row containing the original values 1306 and 1307. FIG. 13D illustrates the calculation of the intermediate interpolation point 1340 and the adjacent quarter points between the interpolated intermediate point values 1342 and 1344. All remaining values can be obtained by linear interpolation.
[0099] FIGS. 14A-C illustrate different types of frames and embodiments of different types of inter prediction that are feasible with respect to these different types of frames. As shown in FIG. 14A, the video signal includes a sequence of linear video frames. In FIG. 14A, the sequence starts with frame 1402 and ends with frame 1408. The first type of frame in the video signal is called "I" frame. The pixel value of the macro block of the I frame cannot be predicted by inter-frame prediction. The I frame is a type of reference point within the decompressed video signal. The content of the encoded I frame depends only on the content of the original signal I frame. Therefore, when systematic errors occur in decompression involving problems associated with inter prediction, video signal decompression can be restored by skipping forward to the next I reference frame and restarting decoding from that frame. Such errors do not propagate across the I frame barrier. In FIG. 14A, the first frame 1402 and the last frame 1404 are I frames.
[0100] The next type of frame is illustrated in FIG. 14B. The P frame 1410 may contain blocks that have been inter-predicted from the I frame. In FIG. 14B, block 1412 has been coded as a motion vector and reference frame 1402 identifier. The motion vector represents the temporal movement of the position of the block 1414 in the reference frame 1402 to the block 1412 in the P frame 1410. A P frame represents a type of prediction constraint frame that contains blocks that can already be predicted from a reference frame through inter prediction. The P frame represents another type of fence frame in the encoded video signal. Figure 14C illustrates a third type of frame. B-frames 1416-1419 may contain blocks predicted from one or two other B-frames, P-frames, or I-frames through inter-frame prediction. In FIG. 14C, the B frame 1418 contains the block 1420 that is inter-predicted from the block 1422 in the P frame 1410. The B frame 1416 contains the predicted block 1428 from both the block 1428 in the B frame 1417 and the block 1430 in the reference frame 1402. Block 1426. B-frames can make best use of inter-frame prediction, and therefore, achieve the highest compression due to inter-frame prediction, but also have a higher possibility of causing various errors and abnormalities in the decoding process. When a block (for example, block 1426) is predicted from two other blocks, the block is encoded as two different reference frame identifiers and motion vectors, and the prediction block is generated as two prediction blocks from which it is predicted The possible weighted average of the pixel values in.
[0101] As mentioned above, if intra prediction and/or inter prediction are completely accurate, an extremely high compression ratio can be obtained. Characterizing a block as one or two motion vectors and frame offsets is definitely much more concise than expressing as 256 different pixel values. It is even more efficient to characterize the block as one of 13 different intra prediction modes. However, as can be realized by the large number of different possible macroblock values, in terms of the macroblock value as a 256-byte encoded value, neither intra-frame prediction nor inter-frame prediction can generate the content of a block in a video frame. Precise prediction of, unless the video signal containing the video frame does not contain noise and contains almost no information, such as a uniform, unchanged, and pure-color background video. However, even if intra-frame prediction and inter-frame prediction cannot accurately predict the content of a macroblock, generally speaking, they can often estimate the content of a macroblock relatively close. This estimation can be used to generate a difference macroblock, which characterizes the difference between the actual macroblock and the predicted value obtained by intra prediction or inter prediction for the macroblock. When the prediction is satisfactory, the resulting difference block usually contains only a few or even zero pixel values.
[0102] FIG. 15 illustrates an embodiment of the generation of difference macroblocks. In the embodiment of FIG. 15, the macro block is shown as a three-dimensional chart in which the height of the column above the two-dimensional surface of the macro block represents the magnitude of the pixel value within the macro block. In FIG. 15, the actual macroblocks within the currently considered frame are shown as a three-dimensional chart 1502 at the top. The middle three-dimensional chart representation passes
CN 102369552 Β
Predicted macroblocks obtained by intra-frame prediction or inter-frame prediction. Note that the three-dimensional graph of the predicted macroblock 1504 is completely similar to the actual macroblock 1502. Figure 15 characterizes the situation where intra prediction or inter prediction has produced a very close estimate of the actual macroblock. Subtracting the predicted macroblock from the actual macroblock produces a difference macroblock, which is shown as a lower three-dimensional graph 1506 in FIG. 15. Although Figure 15 is an exaggeration of the best-case prediction, it illustrates that, compared to the actual final predicted macroblock, the difference macroblock usually contains not only smaller amplitude values, but also often contains fewer non-zero values. It should also be noted that the actual macroblock can be completely restored by adding the difference macroblock to the predicted macroblock. Of course, the predicted pixel value can exceed or be lower than the actual pixel value, so the difference macroblock can contain both positive and negative values. However, as an example, the shift of the origin can be used to generate all positive difference macroblocks.
[0103] Just as the pixel values in the macroblock can be predicted from the values in the blocks that are spatially adjacent and/or temporally adjacent to the macroblock, it is also possible to predict the motion vector generated by inter-frame prediction and pass The mode generated by intra prediction. Figure 16 illustrates an embodiment of motion vector and intra prediction mode prediction. In Figure 16, the currently considered block 1602 is shown within a block grid of a portion of the frame. The neighboring blocks 1604-1606 have been compressed through intra prediction or inter prediction. Therefore, there are intra prediction modes or inter prediction motion vectors associated with these adjacent, already compressed blocks, and the intra prediction mode is a type of displacement vector. Therefore, a reasonable assumption is that the space vector or time vector associated with the currently considered block 1602 according to whether intra prediction or inter prediction is used will be associated with the adjacent, already compressed block 1604-1606 The space vector or time vector is similar. In fact, the space vector or time vector associated with the currently considered block 1602 can be predicted as the average of the space vector or time vector of the neighboring blocks as shown in the vector addition 1610 on the right of FIG. 16. Therefore, instead of directly encoding the motion vector or the inter prediction mode, the H·264 standard calculates the difference vector based on vector prediction by subtracting the predicted vector 1622 from the actually calculated vector 1622. Between frames The temporal motion of the block and the spatial consistency within the frame will be expected to be roughly correlated, and therefore, the predicted vector will be expected to closely approximate the actual calculated vector. Therefore, the size of the difference vector is usually smaller than the actually calculated vector, and thus, fewer bits can be used to encode the difference vector. Furthermore, as with the difference macroblock, the actual calculated vector can be accurately reconstructed by adding the difference vector to the prediction vector.
[0104] Once the difference macro block is generated through inter prediction or intra prediction, the difference macro block is decomposed into 4×4 difference blocks according to a predetermined order, and each 4×4 difference block is transformed through integer transformation to generate the corresponding coefficient block , And then quantize the coefficients of the coefficient block to generate a final quantized coefficient sequence. The advantage of intra-frame prediction and inter-frame prediction is that the transformation of the difference block usually generates a large number of trailing zero coefficients. These trailing zero coefficients can be efficiently compressed by the subsequent bad coding step.
[0105] FIG. 17 illustrates one embodiment of decomposition, integer transformation, and quantization of difference macroblocks. In this embodiment, the difference macroblock 1702 is decomposed into 4X4 difference blocks 1704-1706 in the order described by the numerical designation of the unit of the difference macroblock in FIG. 17. The integer transformation 1708 calculation is performed on each 4X4 difference block to generate The corresponding 4X4 coefficient block 1708. The coefficients in the transformed 4×4 block are serialized according to the zigzag serialization mode 1710 to generate a linear coefficient sequence, and then the coefficient sequence is quantized by the quantization calculation 1712 to generate a quantized coefficient sequence 1714. Many of the already discussed steps in video signal compression are lossless. The macro block can be regenerated losslessly from the intra prediction method or the inter prediction method and the corresponding difference macro block. There is also an exact inverse of the integer transformation. However, since once quantized, the approximate value of the original coefficient can be reproduced by the approximate inverse of the quantization method (referred to as "re-scaling"), so the quantization step 1712 is a form of lossy compression. Since high-resolution chroma data cannot be recovered from low-resolution chroma data, chroma plane decimation is another lossy compression step. Quantization and chroma plane extraction are actually two lossy compression steps in H.264 video compression technology.
CN 102369552 Β
[0106] FIG. 18 provides the derivation of integer transform and inverse integer transform used in H.264 video compression and video decompression, respectively. The symbol "X 1802 represents a 4X4 difference block or residual block (for example, 1704-1706 in Figure 17)<sub>ο</sub>The discrete cosine transform is defined by the first set of expressions 1804 in FIG. 18. The discrete cosine transform is a well-known transform similar to the discrete Fourier transform. As shown in Expression 1806, the discrete cosine transform is an operation based on matrix multiplication. The discrete cosine transform can be factored as shown in expression 1808 in FIG. 18. The elements of the matrix C 1810 include the rational number "d" 1812. In order to estimate the discrete cosine transform efficiently, the number can be approximated to 1/2, so that the approximate matrix element 1814 in FIG. 18 is obtained. This estimation of multiplying two rows of matrix C in order to generate all integer elements generates the integer transformation 1818 and the corresponding inverse integer transformation 1820 in FIG. 18.
[0107] FIG. 19 illustrates the quantization process. It can be assumed that any integer value is in the range 0-255. As a simple example, consider that the number 1902 encoded with 8 bits can therefore be between 0 (1904 in Figure 19) and 255 (1906 in Figure 19). Value range. As shown in FIG. 19, the quantization process can be used to encode the 8-bit number 1902 with only three bits 1908 by inverse linear interpolation from an integer in the range 0-255 to an integer in the range 0-7. In this case, the integer values 0-31 represented by 8-bit encoded numbers are all mapped to the value 0 (1912 in Fig. 19) ο 32 integer values in the continuous range are mapped to the values 1-7 ο Therefore, for example , The quantization of the integer 200 (1916 in Figure 19) generates a quantized value of 6 (1918 in Figure 19). The 8-bit value can be regenerated from the 3-bit quantized value by simple multiplication. The 3-bit quantized value can be multiplied by 32 to generate an approximation of the original 8-bit number. However, the approximate number 1920 may only have one of the values 0, 32, 64,... 224. In other words, quantization is a form of numerical extraction or loss of precision. The rescaling process or multiplication can be used to reproduce the number that estimates the original value that was quantized, but the accuracy lost in the quantization process cannot be restored. Generally, quantization is expressed by Formula 1922, and the inverse or rescaling of quantization is expressed by Formula 1924. The value "Qstep" in these formulas controls the accuracy lost in the quantization process. In the embodiment illustrated on the left side of FIG. 19, Qstep has the value "32". A smaller Qstep value provides a smaller loss of accuracy, but also provides less compression, and a larger value provides a greater compression, but also provides a greater loss of accuracy. For example, in the embodiment shown in FIG. 19, if Qstep is 128 instead of 32, an 8-bit number can be encoded with one bit, but the rescaling will only generate two values 0 and 128. It should also be noted that the scaling value can be vertically shifted as indicated by arrows 1926 and 1928 by an additional addition step after rescaling. For example, in the embodiment shown in FIG. 19, instead of generating values 0, 32, 64, ... 224, 16 is added to the scaling value to generate corresponding values 16, 48, ..., 240, so that the The gap at the top of the vertical number axis is not that large.
[0108] After quantizing the residual block or the difference block and collecting the difference vectors and other objects generated as the data stream from the steps upstream of the bad coding, the bad encoder is applied to the partially compressed data stream to generate bad coded data The badly coded data stream includes the payload of the NAL packet described above with reference to FIG. 5. Rotten coding is a lossless coding technique that takes advantage of the statistical non-uniformity in the partially coded data stream. A well-known example of bad coding is Morse code. Morse code uses single pulse codes of frequently occurring letters (such as "E" and "T") and infrequently encountered letters (such as "Q" and "Z"). ) Four-pulse or five-pulse coding.
[0109] FIG. 20 provides a digital embodiment of bad coding. Consider a four-symbol string 2002 including 28 symbols, and each character is selected from one of the letters "A", "B", "C", and "D". As shown in the encoding table 2004, a simple and intuitive encoding of the 28-symbol string would be to assign one of four different 2-bit codes to each of the four letters. Using this 2-bit encoding, a 56-bit encoded symbol string 2006 equivalent to the symbol string 2002 is generated. However, the analysis of the symbol string 2002 revealed the percentage incidence of each symbol shown in the table 2010. "A" is the symbol that appears most frequently so far, and "D" is the symbol that appears least frequently so far. Characterize better coding by coding table 2012,
CN 102369552 Β
The encoding table 2012 uses a variable length representation of each symbol. "Α", which is the most frequently occurring symbol, is assigned the code "0". The least frequently occurring symbols "B" and "D" are assigned codes "110" and "111", respectively. This encoding is used to generate an encoding symbol string 2014 using only 47 bits. Generally, for symbols with the occurrence probability of P, binary bad coding should generate -logzP bits of coded symbols. Although in the embodiment shown in FIG. 20, for a long symbol sequence that clearly has an uneven symbol appearance distribution, the improvement in the code length is not large, but the bad coding generates a relatively high compression ratio. [0110] One type of bad coding is called "arithmetic coding". A simple example is provided in Figures 21A-B. The arithmetic coding illustrated in Figures 21A-B is a version of a context adaptive coding method. In this embodiment, the 8-symbol sequence 2102 is encoded as a decimal value with 5 digits after the decimal point. 04016 (2104 in Figure 21A), which can be converted to small numbers by any of various known binary number encodings. The numeric value 04016 is coded to generate a binary coded symbol string. In this simple embodiment, the symbol occurrence probability table 2106 is continuously updated during the encoding process. Since when the symbol appearance probability is adjusted according to the symbol appearance frequency observed during encoding, the encoding method dynamically changes with time, so this is recommended. For context adaptation. In the beginning, due to the lack of a better set of initial probabilities, the probabilities of all symbols were set to 0.25. At each step, use intervals. The interval of each step is represented by a number line (for example, the number line 2108). At the beginning, the interval changed in the range of 0-1. At each step, the interval is divided into four partitions according to the probability of the current symbol appearance frequency table. Since the initial table contains an equal probability of 0.25, in the first step, the interval is divided into four equal parts. In the first step, the first symbol "A" 2110 in the symbol sequence 2102 is encoded. The interval partition 2112 corresponding to this first symbol is selected as the interval 2114 for the next step. In addition, since the symbol "A" is encountered, the probability of occurrence of the symbol "A" is increased by 0.03 and the probability of the remaining symbols is reduced by 0. 01 to adjust the symbol appearance probability in the next version of Table 2116. The next symbol is still "A" 2118, so the first interval partition 2119 is selected again as the next interval 2120 used in the third step. This process continues until all symbols in the symbol string have been used. The last symbol "A" 2126 selects the first interval 2128 among the last intervals calculated in the process. Note that the size of the interval decreases at each step, and it is usually necessary to specify a larger number of decimal places. The symbol string can be encoded by selecting any value in the last interval 2128. The value .04016 falls within this interval, and therefore, characterizes the encoding of the symbol string. As shown in FIG. 21B, the original symbol string can be regenerated by starting the process again using the initial equivalent symbol appearance frequency probability table 2140 and the initial interval 0-1 2142. The code .04016 is used to select the first partition 2144 corresponding to the symbol "A". However, in steps similar to the steps in the forward process shown in FIG. 21A, the code .04016 is used to select each subsequent partition of each subsequent interval until the final symbol string 2148 is regenerated. [0111] Although this embodiment illustrates the general concept of arithmetic coding, since this embodiment assumes infinite precision arithmetic, and because the symbol appearance frequency probability table adjustment algorithm will quickly lead to inoperable values, it is false. Examples. The actual arithmetic coding does not assume infinite precision arithmetic, but uses technology to adjust the interval to give the interval specification and choice within the precision provided by any particular computer system. The Η.264 standard specifies several different coding schemes. These codes One of the schemes is the context-adaptive arithmetic coding scheme. The table look-up process is used to encode frequently occurring symbol strings generated by upstream encoding techniques to facilitate subsequent decompression. The frequently occurring symbol strings include partially compressed Various metadata and parameters included in the data stream.
[0112] When the video data stream is compressed according to the H.264 technology, the subsequent decompression can obtain certain types of artifacts. As an example, Figures 22A-B illustrate a common artifact and a filtering method that is used as the final step of decompression to improve the artifact. As shown in FIG. 22A, without filtering, the decompressed video image can appear blocky. Since decompression and compression are performed block by block, each block boundary can characterize significant discontinuities in the compression/decompression process, which discontinuities result in visually perceptible blocks of the displayed decompressed video image. Figure 22B
CN 102369552 Β
The illustration illustrates the deblocking filter method used to improve blocking artifacts in H.264 decompression. In this technique, in order to smooth the discontinuity of the pixel value gradient across the block boundary, a vertical filter similar to the filter for pixel value interpolation discussed above with reference to FIGS. 13A-D is moved along all the block boundaries.Device2210 and horizontal filter 2212. The three pixel values on each side of the boundary can be affected by the deblocking filter method. On the right side of Fig. 22B, an embodiment of the application of a deblocking filter is shown. In this embodiment, the filter 2214 is characterized as a vertical column containing four pixel values on either side of the block boundary 2216. The application of the filter generates filtered pixel values for the first three pixel values on either side of the block boundary. As an embodiment, the filter value χ* of the pixel 2218 is calculated from the pre-filter values of the pixels 2218, 2220, 2221, 2222, and 2223. In order to re-establish the continuous gradient across the boundary, the filter tends to average the pixel values or blur the pixel values.
[0113] Figure 23 summarizes the H.264 video data stream encoding. Figure 23 provides a block diagram and, therefore, a high-level description of the encoding process. However, this diagram, together with the previous discussion and the previously referenced figures, provides a basic overview of H·264 encoding. When necessary, additional details are disclosed to describe the specific video codec implementation of the present invention. It should be noted that there are many subtle points, details and special situations in video encoding and video decoding that cannot be addressed in the overview section of this document. For ease of communication and simplification, most of the embodiments here are based on the H·264 standard, however, it should never be understood that the present invention presented here is limited to the H·264 application. The official H. 264 is over 500 pages long. These many details include, for example, special cases caused by various boundary conditions, specific details, and optional alternative methods that can be applied in various context-sensitive situations. Consider, for example, intra prediction. The intra prediction mode depends on the availability of pixel values in specific neighboring blocks. For boundary blocks without neighbors, many of the modes cannot be used. In some cases, in order to make it possible to use a specific intra prediction mode, the unavailable neighboring pixel values can be interpolated or estimated. Many interesting details in the encoding process are related to the following operations: selecting the best prediction method, quantization parameters, and other such parameter selections to optimize the compression of the video data stream. Η. The 264 standard does not specify how to perform compression, but instead, specifies the format and content of the encoded video data stream and how the encoded video data stream will be decompressed. The H·264 standard also provides a variety of different levels of different computational complexity, among which the high-end level supports additional steps and methods that are computationally more expensive but more efficient. The present summary is intended to provide sufficient background for understanding the description of various embodiments of the present invention provided later, but is by no means intended to constitute a complete description of H·264 video encoding and decoding.
[0114] In FIG. 23, a stream of frames 2302-2304 is provided as an input to the encoding method. In this embodiment, as discussed above, the frame is decomposed into macroblocks or macroblock partitions for subsequent processing. In the first processing step, an attempt is made to perform inter-frame prediction on the currently considered macroblock or macroblock partition from one or more reference frames. When, as determined in step 2308, the intra-frame prediction is successful and one or more motion vectors are generated, the actual original macroblock is subtracted from the actual original macroblock in the difference step 2310, and the prediction generated by the motion estimation and compensation step 2306 is subtracted Macro block to generate a corresponding residual macro block, which is output to the data path 2312 ± ο through the step of differencing. However, if the inter prediction fails as determined in step 2308, then the intra prediction is started Step 2314 is to perform intra prediction on the macro block or macro block partition, and then in step 2310, the macro block or macro block partition is subtracted from the actual original macro block or macro block partition to generate the output to the data path 2312 Residual macroblock or residual macroblock partition. Then the residual macroblock or residual macroblock partition is transformed by the transform step 2316, and quantized by the quantization step 2318. It may be reordered in step 2320 to encode more efficiently, and then bad encoding is performed in step 2322 to A stream of output NAL packet 2324 is generated. Usually, compressed While balancing the cost, timeliness, and memory usage of various prediction methods, the embodiment seeks to utilize a prediction method that provides the closest prediction of the considered macroblock. Any of a variety of different ranking and selection criteria for applying prediction methods can be used.
CN 102369552 Β
[0115] Continuing to follow the embodiment of FIG. 23, after quantization in step 2318, the quantized coefficients are input to the reordering stage 2320 and the bad encoding stage 2322, and are also input to the inverse quantizer 2326 and the inverse transform step 2328 to The residual macroblock or residual macroblock partition is regenerated, and the residual macroblock or residual macroblock partition is output to the data path 2330 through an inverse transform step. The residual macroblock or macroblock partition output through the inverse transform step is usually not the same as the residual macroblock or residual macroblock partition output through the difference step 2310 to the data path 2312±. Recall that quantization is a lossy compression technique. Therefore, the inverse quantization step 2326 generates estimates of the original transform coefficients, rather than accurately reproducing the original transform coefficients. Therefore, although the inverse integer transform will generate an exact copy of the residual macroblock or macroblock partition, if it is applied to the original coefficients generated by the integer transform step 2316, since the inverse integer transform step 2328 is applied to the rescaling coefficients, In step 2328, only an estimate of the original residual macroblock or macroblock partition is generated. Then in the addition step 2332, the estimated residual macroblock or macroblock partition is added to the corresponding predicted macroblock or macroblock partition to generate a decompressed version of the macroblock. The decompressed, but unfiltered, macroblock version is input to the intra prediction step 2312 through the data path 2334 for subsequent processing Intra prediction of the block. Perform the steps of the deblocking filter 2336 on the decompressed macroblocks to generate filtered and decompressed macroblocks, and then combine the macroblocks to generate decompressed images 2338-2340, decompressed images 2338 -2340 can then be input to the motion estimation and compensation step 2306. One ingenuity involves the input of decompressed frames to the motion estimation and compensation step 2306, and the decompressed but unfiltered macroblocks and macroblock partitions are input to the intra prediction step 2314. Recall that in order to predict the value in the currently considered macroblock or macroblock partition, both intra-frame prediction and most motion estimation and compensation use neighboring blocks. In the case of spatial prediction, the relative block in the current frame is used. Neighboring blocks, or in the case of temporal inter prediction, use neighboring blocks in the previous frame and/or following frame. However, consider the recipient of the compressed data stream. The recipient will not be able to access the original original video frames 2302 and 2304. Therefore, during decompression, the receiver of the encoded video data stream will use previously decoded or decompressed macroblocks for predicting the content of subsequently decoded macroblocks. If the encoding process uses the original video frame for prediction, the encoder will use data that is different from the data subsequently available to the decoder for prediction. This will cause significant errors and artifacts in the decoding process. In order to prevent this, the encoding process is used to The decompressed macroblocks and macroblock partitions and the decompressed and filtered video frames in the inter prediction step and the intra prediction step, so that intra prediction and inter prediction use the same data for any decompression process. The content of the available macroblocks and macroblock partitions is predicted, and any decompression process can only rely on the coded video data stream for decompression. Therefore, the decompressed but unfiltered macroblock and macroblock partition that are input to the intra prediction step 2314 through the data path 2334 are the neighboring blocks from which the current macroblock or macroblock partition is subsequently predicted, and the motion estimation and The compensation step 2306 uses the decompressed and filtered video frames 2338-2340 as reference frames for processing other frames.
[0116] FIG. 24 illustrates an exemplary H.264 video data stream decoding process in a block diagram manner similar to that used in FIG. 23. Decompression is much simpler than compression. The NAL packet stream 2402 is input to the bad decoding step 2404, the bad decoding step 2404 applies inverse bad coding to generate quantized coefficients, and the reordering step 2406 reorders the quantized coefficients to be the same as those performed by the reordering step 2320 in FIG. 23. Complementary sorting. The information in the bad decoding stream can be used to determine the parameters with which the data was originally encoded, the parameters including whether intra prediction or inter prediction was used during the compression of each block. Through step 2408, the data allows the selection of inter prediction in step 2410 or the selection of intra prediction in step 2412 to generate prediction values for the macroblocks and macroblock partitions provided to the addition step 2416 along the data path 2414. In step 2418, the inverse quantizer rescales the reordered coefficients, and in step 2420, the inverse integer transform is applied to generate the residual or residual macroblock or macroblock partition estimate, and the estimate is added in the addition step 2416. To predicted macroblocks or macroblock partitions based on previously decompressed macroblocks or macroblock partitions. The addition step generates decompressed macroblocks or macroblock partitions to generate decompressed video frames 2424-2426, which are decompressed in step 2422
CN 102369552 Β
A deblocking filter is applied to the reduced macroblock or macroblock partition to generate the final decompressed video frame. The decompression process is basically equivalent to the lower part of the compression process shown in FIG. 23.
[0117] The second subsection: the principle of parallel integrated circuit design for solving complex computing tasks according to the present invention
[0118] The problem of implementing a calculation engine that performs H.264 compression and decompression is to illustrate an exemplary problem domain of the present invention. In this section, the principle used to develop a parallel pipelined integrated circuit calculation engine for performing H.264 compression and decompression is described as an example of the overall method for characterizing the design of the calculation engine of the embodiment of the present invention. The present invention is by no means limited to the H.264 embodiment.
[0119] One way to implement a video codec that implements the H.264 video compression and decompression discussed in the first subsection is to program the encoding and decoding process with software and execute the program on a general-purpose computer. Figure 25 is a very high-level diagram of a general-purpose computer. The computer includes a processor 2502, a memory 2504, a memory/processor bus 2506, and a bridge 2508, and the memory/processor bus 2506 interconnects processors and memories. The bridge interconnects the processor/memory bus 2506 with the high-speed data input bus 2510 and the internal bus 2512, and the internal bus 2512 connects the first bridge 2508 and the second bridge 2514. The second bridge is connected to each device 2516-2518 through a high-speed communication medium 2520. One of these devices is an I/O controller 2516 that controls the mass storage device 2520.
[0120] Consider the execution of a software program that implements a video codec. In this embodiment, the software program is stored on the mass storage device 2520, and is paged into the memory 2504 as needed. The processor 2502 fetches the instructions of the software program from the memory for execution. Therefore, the execution of each instruction involves at least one memory fetch, and may also involve the processor accessing stored data in the memory (and ultimately in the mass storage device 2520). A large part of the actual computing behavior in general-purpose computer systems is devoted to transferring data and program instructions between mass storage devices, memories, and processors. In addition, with regard to video cameras or other data input devices that generate large-capacity data at a high data transmission rate, there may be a lot of competition between the video camera and the processor for both the memory and the large-capacity storage device. This kind of competition can continue to reach the saturation of various buses and bridges in general computer systems. In order to use the software implementation of the video codec to implement real-time video compression and decompression, a very large part of the available computing resources and power consumed by the computer is dedicated to data transmission and instruction transmission, rather than actually performing compression and decompression. The parallel processing method can be expected as a feasible method to improve the computational throughput of a software-implemented video codec. However, in general computing systems, it is far from trivial tasks to properly decompose the problem to make full use of multiple processing components, and It may not solve the competition for memory resources and the exhaustion of data transmission bandwidth in the computer system, or may even worsen the competition for memory resources and the exhaustion of data transmission bandwidth in the computer system.
[0121] The next implementation that can be considered would be to use any of various system-on-a-chip design methods to move the software implementation to hardware. A video codec implemented by a system-on-chip will provide certain advantages over a general-purpose computer system that implements a software implementation of the video codec. Specifically, the program instructions can be stored on the board or in the flash memory, and various calculation steps can be implemented in a logic circuit instead of being implemented as instructions that are sequentially executed by the processor. However, the system-on-chip implementation of the video codec is generally still sequential in nature, and does not provide a high-throughput parallel computing method.
[0122] FIG. 26 illustrates many aspects of the video compression and decompression process, which, when considered, provide a deep understanding of the much more computationally efficient new method of implementing the video codec according to the present invention. First, the H.264 standard already provides advanced problem decomposition that is subject to parallel processing solutions. As discussed above, each video frame 2602 is decomposed into macroblocks 2604-2613, and in order to compress the video frame in the forward direction, macroblock-based or macroblock partition-based operations are performed on the macroblocks and macroblock partitions. , And decompress the macro block in reverse to reconstruct the solution
CN 102369552 Β
Compressed frame. As discussed above, there must be a correlation between frames and between macroblocks during the encoding process and during the decoding process. However, as shown in FIG. 26, the correlation between the macroblock and the macroblock and the correlation between the macroblock partition and the macroblock partition are generally forward correlations. The starting macroblock in the starting frame of the sequence 2613 does not depend on the following macroblocks, and can be compressed completely based on its own content. When the compression continues frame by frame through the raster scanning process of macroblocks, the subsequent macroblocks may depend on the macroblocks in the previously compressed frame, especially for inter prediction, and may depend on the previously compressed ones in the same frame. Macro block, Especially for intra prediction. However, relevance is well constrained. First, the correlation is limited by the maximum distance in sequence, space, and time 2620. In other words, there are only adjacent macroblocks in the current frame and macros in the search area centered on the position of the current frame in a relatively small number of reference frames. Blocks may contribute to compressing any given macroblock. If the correlation is not well constrained in time, space, and sequence, a very large memory capacity will be required to accommodate the intermediate results required to compress consecutive macroblocks. Such memory is expensive, and as the complexity and size of memory management tasks increase, it quickly starts to consume available computing bandwidth. Another type of constraint is that for a given macroblock 2622, there may only be a relatively small and maximum number of correlations. This kind of constraint also helps to limit the necessary size of the memory, and also helps to limit the computational complexity. As the number of correlations grows, computational complexity can grow geometrically or exponentially. In addition, when the necessary communication between processing entities is well restricted, parallel processing solutions to complex computing problems are feasible and manageable. Otherwise, the result is that communication between separate processing entities quickly overwhelms the available computing bandwidth. Another feature of the video codec problem is that the processing of each macroblock in the forward compression direction or the reverse decompression direction is a stepwise process 2624. As discussed above, these sequential steps include inter-frame and intra-frame prediction, residual macroblock generation, Main transformation, quantization, object reordering and bad coding. These steps are separate, and generally speaking, the result of one step is directly fed to the following step. Therefore, just as cars or electrical appliances can be manufactured step by step along the assembly line, macro blocks can be processed in the assembly line.
[0123] In many different problem domains, there may be features implemented by the video codec discussed with reference to FIG. 26 that motivate the massively parallel processing implementation of the video codec according to the present invention. In many cases, computational problems can be broken down in many different ways. In order to apply the method of the present invention to any particular problem, as the first step of the method, it is necessary to select a problem decomposition that produces some or all of the features discussed above with reference to FIG. 26. For example, the problem of video data stream compression can be broken down in an alternative and unfavorable way. For example, an alternative decomposition would be to analyze the entire video data stream or most blocks of the frame to perform motion detection before macroblock processing. In some aspects, this larger granularity method can provide significant advantages for motion detection and motion detection-based compression. However, this alternative problem decomposition requires a large internal memory, and the motion detection step will be too complicated and computationally inefficient to be easily included in the step-by-step processing of easy-to-calculate and manageable data objects. adapt.
[0124] Again, it is emphasized that although the present invention is described in the context of implementing a video codec, the method of the present invention is applicable to a wide range of efficient computing engines designed to solve a variety of different computing problems. For those problems that can be decomposed and formulated to provide the features discussed with reference to FIG. 26, the method of the present invention provides efficiency in calculating bandwidth, cost, power consumption, and other important efficiencies that stimulate and constrain the development of computing engines, devices, and systems. .
[0125] FIG. 27 illustrates the basic features of a single integrated circuit implementation of a video codec according to the method of the present invention. Those components implemented in a single integrated circuit are shown in the large dashed block 2702. The video codec embodiment additionally uses the external memory 2704 and the external optics and electronics of the video camera 2706. Additional external components of the video camera system include power supplies, various additional electromechanical components, housings, interconnecting components for external devices, and other such components.
[0126] As discussed above with reference to FIG. 4, the video camera provides a data stream and an electronic timing signal input 2708 to
CN 102369552 Β
Video codec. The data stream is directed to the memory 2704 and the microprocessor controller part 2710 in the integrated circuit. The microprocessor controller part 2710 can access the timing signal output of the video camera to coordinate the behavior of the video codec. The memory 2704 is dual-ported, so that when the video data stream enters from the digital video camera 2706, the previously stored original video data can be extracted from the external memory to the internal cache memory 2712 to provide to several processing elements 2714-2719 Each of them. In Figure 27, six processing elements 2714-2719 are shown, but in the specific embodiment discussed below, there are in fact a greater number of processing elements. The number of processing elements is a parameter determined by the problem domain and selected by the design. The completely separate processing elements of one embodiment may alternatively be combined together in another embodiment.
[0127] In the embodiment of FIG. 27, the microprocessor controller 2710 executes instructions stored in the flash memory 2720. The microprocessor controller communicates with the memory 2722, the cache memory 2724, the system clock 2726, and the multiple processing elements 2728 through various signal paths. Within the integrated circuit, a large amount of data flow occurs through the object bus 2730. The object bus transfers the objects related to the video data (mainly macroblocks and macroblock partitions) to the processing elements. In addition, the object bus can also transfer shared parameters and metadata, including objects that describe macroblocks and macroblock partition objects, as well as higher-level structures within the current frame and video data stream.
[0128] In this embodiment, each processing element performs a step of stepwise processing of video data objects, which are mainly macroblocks and macroblock partitions. The types of video objects input to the processing element and the types of video and data objects output through the processing element depend on the specific steps in the compression process implemented by the processing element. The processing element performs a large number of calculations that are performed to compress the video data. The processing method is a very pipelined and assembly-line method. In this method, a given raw data macroblock enters the first processing element 2714, and is processed in a stepwise manner along the subsequent processing element sequence in the processing element pipeline. Transform. The overall assembly line processing is controlled by the timing signal of the calculation steps realized by the relatively low-frequency clock signal. The processing steps within each processing element are controlled by a relatively high frequency clock signal. An important aspect of the single integrated circuit implementation of the video codec is that the low-frequency calculation step timing signal provides the timing signal for the microprocessor controller, but does not provide absolute control of the assembly line process. Generally speaking, each step of the step-by-step advanced processing should be performed within the time sequence signal interval of a single low-frequency calculation step. However, there may be cases where the processing element cannot complete its task within a time interval. These conditions are detected by the advanced control logic provided by the microprocessor controller 2710. In this case, the microprocessor controller can delay the start of the subsequent calculation steps, so that even if the processing element The task has exceeded the low-frequency timing interval, and the processing element can also complete its task. Thus, the microprocessor controller control provides an important level of flexibility in the overall control of the video compression and video decompression process. If this flexibility is not provided, the low-frequency interval needs to be set to at least the largest possible time interval required by any processing element in the system to complete the most computationally complex task that the processing element may encounter. In the case that the most complex task only occurs infrequently (for example, once every 1,000 macroblocks), the processing elements will spend most of the time in the low-frequency time interval during the processing of the remaining 999 macroblocks with small computational requirements. Be idle during the time period. By providing more flexible microprocessor controller control of the overall assembly line process, the low-frequency timing signal interval can be set to a reasonable value, which specifies the time interval during which most of the macroblocks can be processed, and can be real-time The low-frequency timing signal interval is adjusted in a context-sensitive manner to adapt to relatively infrequent and computationally intensive macroblocks.
[0129] In this embodiment, the on-board object cache 2712 provides different types of flexibility. The cache memory provides dynamic buffering for data objects that can accommodate different amounts of data required at a specific point in video compression. Like the timing flexibility provided by the microprocessor controller control, the flexible cache memory allows efficient processing of general processing tasks with less memory-intensive access while adapting to specific contexts.
CN 102369552 Β
Related memory requirements. The higher frequency timing interval provided by the clock 2726 allows clocked processing within the processing element that is implemented as a logic circuit rather than as instructions executed by a microprocessor controller. It is this clock-controlled, logic circuit-based implementation that provides the large computational bandwidth realized by the overall single integrated circuit of the video codec. If most of the video compression and video decompression processes are executed by instruction execution on the processor, most of the overall computational overhead will be consumed by the instruction fetch cycle. The object memory controller is responsible for exchanging objects between the on-board object cache memory and the object memory.
[0130] Finally, the object bus 2730 facilitates an object pipeline-based implementation. If macroblocks and macroblock partitions are transmitted within an integrated circuit as bitstreams or byte streams, VHF communication processing will be required for the transfer of macroblocks and macroblock partitions to and from the processing element . By providing a wide high-capacity object bus, the data sent to each processing element can be transferred from the cache memory in a computationally and time-efficient manner.
[0131] In summary, the implementation of complex computing tasks according to the present invention involves the design and fabrication of a single integrated circuit that implements a problem-specific computing engine. The calculation engine includes a microprocessor controller that provides high-level control of processing within the integrated circuit, and also provides a large number of parallel pipelined processing elements that perform a large number of calculation processes. The processing elements operate in parallel to provide a very high computing bandwidth, and are provided with object-based data through the object bus and the intermediate processing element data path. The objects, such as macroblocks and macroblock partitions, are natural for the processing elements to operate on. Object. The high-frequency timing in the processing element is provided by the system clock, and the low-frequency advanced calculation step control is provided by the microprocessor controller, thereby providing the flexibility of the total timing of the assembled processing of the calculation task to improve the efficiency and throughput of the calculation engine .
[0132] An alternative implementation of a video codec is shown in FIGS. 28 and 29. FIG. 28 illustrates an embodiment of the present invention in which the integrated circuit 2802 includes a memory 2804, which is external in the embodiment illustrated in FIG. 27. Figure 29 illustrates an alternative embodiment of the invention in which a digital video camera is included in a single integrated circuit implementation of a combined video camera and video codec.
[0133] FIGS. 30-32 illustrate the overall timing and data flow within a single integrated circuit implementation of the video codec according to the present invention. Upon completion of the previous step, the microprocessor controller checks the processing entity, the cache memory, and if necessary, also checks the memory to ensure that all data objects required to perform the next high-level calculation step are available for transmission to the data needed The processing element of the object. Therefore, the microprocessor controller checks to ensure that the data objects are available, and when needed, facilitates the grouping of data objects 3001-3005 in the cache for the processing elements to access in the next high-level calculation step, and checks every One processing element has generated any data that needs to be provided to another processing element for the next high-level calculation step, and is currently storing this data.
[0134] It is again noted here that the microprocessor controller control provides the flexibility of the overall control of the integrated circuit. In many cases, whether a specific data object or multiple specific data objects need to be prepared for transmission in the next step depends on the position of the step in the sequence of steps in the overall video encoding task. As an embodiment, the start macroblock of the start reference frame of the video stream is first processed by the first processing element, and the first processing element provides no results to the subsequent processing elements. As another embodiment, for inter-frame prediction, the reference frame in the video stream is not processed. Therefore, in any given low-frequency timing interval, the data objects required to process the subsequent advanced calculation steps can be changed in a context-sensitive manner. Moreover, in some advanced calculation steps, one or more of the processing elements may not be in working state. The implementation complexity of context-sensitive and time-varying control within the processing element itself will undesirably require complex processing element implementation. However, by using a microprocessor controller that executes instructions to provide a higher level of control, many levels of decision-making and time-related and context-sensitive control changes can be made using firmware instead of highly complex logic circuits.
CN 102369552 Β
Road to achieve.
[0135] In this embodiment, once all the data objects are available for the next processing step, and all the processing elements are ready to start executing the next step, the microprocessor controller as shown in FIG. The processing element starts processing the start signal of the next step. As shown in Figure 31, at the beginning of the next step, the data objects are transferred to the processing elements that need them. Then, as shown in FIG. 32, the processing elements perform their respective tasks, generate output for subsequent steps, and request the cache memory for data objects that will be needed in the subsequent steps. When the processing of the current high-level calculation step illustrated in FIG. 32 is completed, the state shown in FIG. 30 is reached, and then the processing element is prepared to start the subsequent processing steps. Again, it should be emphasized that during each low-frequency time interval, each processing element performs its calculation tasks on data objects that are different from the data objects being processed by other processing elements. For example, while one residual macroblock is being transformed by integer transformation, another macroblock is analyzed for inter-frame prediction or intra-frame prediction.
[0136] In summary, the high-level conceptual components of the calculation engine that characterizes an embodiment of the present invention include: (1) Problem decomposition, which results in a reasonable size of calculation object (each calculation object has a bound to other data Relevance), step-by-step processing of processed calculation objects and later processed calculation objects, and each calculation object basically includes data structure values, such as the values of elements of one-dimensional, two-dimensional or higher-dimensional arrays or multi-field records Or the value of the field in the structure; (2) Assembly line processing element series, each of the processing elements executes the high-level steps of the step-by-step processing of calculation objects, and the high-level steps are executed on different calculation objects in parallel; (3) An on-board object cache that buffers enough computing objects so that relevant data and objects used to process the computing objects along the series of processing elements can be loaded from the memory to the object cache at the beginning During the step-by-step processing of computing objects, the memory is not repeatedly accessed; (4) Object bus, which allows processing elements to access objects stored in the object cache in object-level access transactions (5) A low-frequency clock cycle and one or more high-frequency clock cycles, the low-frequency clock cycle is used to control step-by-step processing, and the high-frequency clock cycle is used for small-granularity control of the step calculation by the processing element; (6) Microprocessor controller or other control sub-components for coordination The high-level steps of the processing element are executed, and the execution of the high-level steps of the processing element is synchronized; and (7) an object memory controller for loading an object from the memory to the object cache, and for loading the object in the object cache Store in memory.
[0137] For certain problem domains, the single integrated circuit implementation of the calculation engine provides advantages in manufacturing, chip packaging, device footprint, power consumption, calculation delay, and other such advantages. For other problem domains, the overall calculation engine can be implemented as two or more separate calculation engines, where the problem domain is divided into higher-level sub-domains, and each sub-domain is executed by a separate calculation engine. The sub-domain is further divided into tasks, and each task is executed by a processing element in the computing engine. This method can also provide manufacturing advantages and improved modularity. For some other types of problem domains, in order to utilize already developed integrated circuits, for example, a single integrated circuit implementation of the calculation engine can be combined with another integrated circuit to realize a device.
[0138] In a specific embodiment of the H.264 compression and decompression calculation engine that characterizes an embodiment of the present invention, the calculation objects include macroblocks and macroblock partitions as discussed above with reference to FIG. 6, as discussed above with reference to FIG. 12 The motion vector and various data and parameter objects that describe the video stream context of macroblocks and macroblock partitions. Processing elements include inter prediction processing elements, intra prediction processing elements, motion estimation processing elements, direct integer transformation processing elements, inverse integer transformation processing elements, quantization and scaling processing elements, dequantization and scaling processing elements, and bad coding processing elements And bad decoding processing components. The object cache stores the above-mentioned types of objects, including macroblocks and macroblock partitions. Object bus is there
CN 102369552 Β
Transfer macroblocks and macroblock partitions between the processing element and the object cache, reducing the need for processing elements to execute byte-oriented or word-oriented communication protocols to access computing objects. The low-frequency clock period generally controls the stepwise macroblock processing performed by the assembly line processing element series, and the higher frequency clock period controls the computational processing performed by the processing element. The microprocessor controller performs the overall control and synchronization of the stepwise macroblock processing, ensuring that each processing element that executes the next processing step can obtain the desired object before starting the next processing step in the processing element. Finally, the memory controller operates to exchange calculation objects (including macroblocks and macroblock partitions) between the mass random access memory and the object cache memory.
[0139] The third subsection: H.264 video codec implemented as a single integrated circuit according to an embodiment of the present invention
[0140] In this concluding section, a specific example of a calculation engine characterizing an embodiment of the present invention is described. Again, it is emphasized that the embodiments of the present invention can be designed and implemented to perform any of a large number of different computing tasks, including image processing tasks, 3D media compression and decompression, various types of computational filtering, pattern matching, and neural networks. Network implementation. The following discussion of the H·264 video codec calculation engine is intended to provide a detailed description of an embodiment of the present invention, and is not intended to limit the scope of the appended claims to be designed to perform H·264 video compression and / Or decompressed calculation engine, general video application or any other specific problem domain. This particular implementation is a single integrated circuit implementation of the video codec. Alternative embodiments may utilize a multiple calculation engine approach, or may combine a single integrated circuit calculation engine with additional integrated circuits.
[0141] FIGS. 33A-B provide block diagram illustrations of a single integrated circuit implementation of a video codec according to the present invention. In view of the above discussion with reference to Figures 27-32 and Figures 23-24, most of the diagrams provided in Figure 33A are essentially self-descriptive. The single integrated circuit implementation of the video codec includes separate processing element 3302 for motion estimation, processing element 3304 for intra and inter prediction, processing element 3306 for residual block calculation, and direct integer transform The processing element 3308 for quantization and scaling, the processing element 3310 for bad encoding, the processing element 3313 for bad decoding, the processing element 3314 for inverse quantization and inverse scaling, the inverse integer transform The processing element 3316 and the processing element 3318 for the deblocking filter. Processing element 3302 corresponds to block 2306 in FIG. 23, processing element 3304 corresponds to blocks 2306 and 2314 in FIG. 23, processing element 3306 corresponds to operation 2310 in FIG. 23, and processing element 3308 corresponds to block 2316 in FIG. 23, Processing element 3310 corresponds to block 2318 in FIG. 23, processing elements 3312 and 3313 correspond to block 2322 in FIG. 23, processing element 3314 corresponds to block 2326 in FIG. 23, processing element 3316 corresponds to block 2328 in FIG. 23, and processing element 3318 corresponds to block 2336 in FIG. 23. Note that the video codec as described with reference to FIG. 5 may receive raw video data 3320 from a video camera and generate compressed video data 3322 as output, or may receive compressed video data 3324 as input and generate 3326 as output raw video data. The reorderer block 2320 in FIG. 23 can be incorporated in the processing elements 3310 and 3314 or in the processing elements 3312 and 3313 of the video codec embodiment. It should also be noted that the video memory controller 3330 is responsible for directing input video data to the external memory 3332 and exchanging data objects between the video cache, the external memory and the object bus 3340 in the single integrated circuit implementation. Figure 33B provides the key to Figure 33A. Note that the object bus 3340 can be considered to include separate luminance object bus, chrominance object bus, motion vector object bus, parameter/data object bus, and internal microprocessor controller bus.
[0142] FIG. 33A provides details on the input and output of each processing element in an embodiment of the present invention, and thus provides the interaction of each processing element with the object bus, the video cache, and the video memory controller. The video memory controller 3330 routes the video data from the camera to the external memory. Microprocessor controller 3342
CN 102369552 Β
The representative processing element initiates a memory request to the video memory controller, and the request is satisfied by the video memory controller accessing the requested data object from the external memory and storing the requested data object in the video cache memory. Therefore, most of the computational overhead associated with dividing the video data signal into frames, macroblocks, and macroblock partitions is performed in the video memory controller. Another aspect of massively parallel processing is provided by the single integrated circuit of the video codec. The implementation mode provides.
[0143] The multiplexer 3344 provides a path from the quantization processing element 3310 to the dequantization and de-scaling processing element 3314 during video compression, and provides from the bad decoder 3313 to the dequantization and de-scaling processing element 3314 during video decompression. route of. The motion estimation processing element 3302 operates on luma macroblocks and macroblock partitions, while the remaining processing elements operate on both luma and chroma macroblocks and/or macroblock partitions. The SPI port 3350 in FIG. 33A is a serial-parallel interface that allows writing and/or reading of the flash memory through SPI interface signals.
[0144] FIG. 34 illustrates the overall system timing and synchronization of a single integrated circuit implementation of a video codec according to an embodiment of the present invention. As discussed above, the short-interval clock pulse signal 3402 controls the execution steps within the processing element during the processing of each overall step in the assembly line processing of macroblocks and macroblock partitions. As discussed above, the processing element starts to execute the next high-level calculation step when receiving the start signal 3404 from the microprocessor controller, and generates a completion signal pulse 3406 when each high-level calculation step is completed. As discussed above, the long-interval clock pulse signal 3410 is generally controlled by a microprocessor controller along the assembly line of processing elements to control the advanced stepwise pipelined processing of macroblocks and macroblock partitions. Generally speaking, during each low frequency interval 3412, each processing element executes the next overall step in processing. However, as discussed above, in some cases, since the processing element starts to process each advanced calculation step when it receives the start signal from the processor, if the processing element fails to complete its task, the next process The step may not start when the low-frequency clock signal transitions from low to high.
[0145] FIG. 35 provides a table showing embodiments of various types of objects that can be stored from a video cache along a data object bus in a single integrated circuit implementation of the video codec according to the present invention. Transfer to the processing element. The table shows two main categories of objects: (1) video object 3502; and (2) data object 3504. The video object includes macroblocks and macroblock partitions from both the luma plane and the chroma plane as discussed above with reference to FIG. 3 and the motion vector object as discussed above with reference to FIG. 12. Data objects include various types of information about the current context of the currently processed macroblock or macroblock partition, the slice to which the macroblock or macroblock partition belongs, the nature of the frame including the macroblock or macroblock partition, and other information. Such information. The object may also contain the parameter information discussed above with reference to FIG. 19, such as a quantization parameter. The computing bandwidth is significantly increased by using the object bus 3340, instead of requiring processing elements to execute byte-based or word-based protocols to access data objects from memory caches and memory, the object bus 3340 is customized to provide a separate The object required by the processing element of the object. The wide data object bus provides extremely high internal data transfer rates within the integrated circuit.
[0146] FIGS. 36A-B illustrate at an abstract level the operation of processing elements within a video codec integrated circuit implementation that characterizes an embodiment of the present invention. As discussed above, for the overall synchronization of the low-frequency advanced step processing cycle, the processing element receives the start pulse 3602 from the microprocessor controller and outputs the completion pulse 3604 to the microprocessor controller. The processing element receives one or more objects and other data 3606 from the previous processing element and/or object bus in the pipeline, and outputs one or more objects and/or other data 3608 to the next in the processing element pipeline Processing element and/or object bus. Of course, the first processing element in the pipeline does not receive the object from the previous processing element, and the last processing element in the pipeline generates output from the integrated circuit implementation of the video codec, instead of outputting the object or other data to the processing element . As discussed above, the processing element receives for controlling the processing element
CN 102369552 Β
The high-frequency clock pulse signal of the logic circuit in the component to perform complex calculation tasks. Note that the processing element passes data and results along the pipeline through the pipeline memory, which is completely different from the object bus.
[0147] FIG. 36B illustrates synchronization and timing control of processing elements. As discussed above, the processing element performs calculation tasks according to the high-frequency clock signal 3620. When the start signal pulse 3622 is received, the task starts, and the processing element declares the task to be completed through the completion signal pulse 3624.
[0148] FIG. 37 illustrates a motion estimation processing element characterizing an embodiment of the present invention. The motion estimation processing element receives as input a luminance object corresponding to the current macroblock, plus one or more luminance objects that characterize the reference macroblock from the reference frame stored in the memory. The motion estimation processing element generates a motion vector object as an output.
[0149] FIG. 38 illustrates an intra-prediction and inter-prediction processing element that characterizes an embodiment of the present invention and includes a pair of processing elements. The intra prediction processing element 3802 receives the luminance and chrominance horizontal and vertical pixel vectors and the data describing the most adjacent block from adjacent blocks, and generates as output four types for the entire macroblock or 16 4X4 macroblock partitions, respectively 16X16 intra prediction mode or one of nine 4X4 intra prediction modes. For each chroma macroblock, one of four chroma intra prediction modes is generated. As with inter prediction, the intra prediction processing element selects the mode that provides the best estimate of the currently considered macroblock. In order to find the specific partition that provides the most effective prediction, according to the level of compression complexity achieved by the video codec, as discussed with reference to FIG. 6, the macroblock can be partitioned in many different ways. The inter prediction processing element 3804 receives both the reference macro block, the luma macro block, and the chroma macro block and the motion vector, and generates a predicted macro block or a macro block partition as an output.
[0150] The two processing element implementations of the intra and inter prediction processing elements (3304 in FIG. 33A) illustrate one design parameter. The number and complexity of processing elements can vary according to many different design considerations and the complexity of the tasks performed by the processing elements. For example, when a very high bandwidth implementation is required, it may be necessary to implement any specific task as several parallel processing elements within the processing element pipeline. In lower bandwidth implementations, these parallel processing elements can be combined together in a single processing element. Another important point is revealed in the implementation of intra and inter prediction processing elements. As discussed above, the entire H·264 standard includes various levels of compression and decompression. Higher levels provide better compression, but at the cost of greater computational complexity. The specific single integrated circuit implementation of the video codec can implement higher-level standards and intermediate and lower-level standards, and the actual operation can be controlled by parameters input to the single integrated circuit and stored in the flash memory. Therefore, the single integrated circuit implementation can provide flexible operation according to multiple parameters.
[0151] FIG. 39 shows a block diagram of processing elements that characterize bad coding of an embodiment of the present invention. The processing element receives luma objects, chroma objects, and motion vector objects, as well as various types of data objects, and applies various bad coding schemes as discussed above to generate the final coded output that is packed into NAL data units.
[0152] FIG. 40 illustrates an example of storage requirements for a video cache in the video codec implementation illustrated in FIG. 33A. In one embodiment, the video cache contains enough macroblocks, macroblock partitions, and motion vector objects so that a given object does not have to be exchanged between the video cache and external memory during the sequence of operations. The sequence is executed as sequential steps starting from the first processing element and proceeding to the last processing element. Various types of objects are stored in circular queues in the video cache, many of which are divided to contain the currently considered macroblock information in one partition, and adjacent macroblock information in another partition. Therefore, for example, the video cache memory includes a circular queue 4002, which contains 16 luminance macroblocks, which are divided into two partitions, each of which has 8 macros. Piece.
CN 102369552 Β
[0153] FIG. 41 illustrates the operation of the luma macroblock circular queue (4002 in FIG. 40) during the interval of nine advanced calculation steps. At time interval t<sub>o</sub>During 4102, the next original video data luminance macroblock is input to slot 0 4104, and the corresponding adjacent reference macroblock from the reconstructed frame is placed in slot 8 4106. During each consecutive time interval, additional original data macroblocks and reference macroblocks are input into consecutive slots in the circular queue. When the macro block in slot 0 4104 is accessed and modified by subsequent processing elements, the content of the macro block changes gradually during the time interval. Finally, at interval t<sub>7</sub> During 4108, the last processing element encodes and outputs the content of the macro block, so that during the interval 4110, the new original data macro block can be placed in slot 0 4112. Therefore, the circular queue contains the macroblock data video cache for assembled processing by all processing elements, and then after the last processing element has used the macroblock, it is replaced with the new original data macroblock and the reference macroblock The macro block. During each low-frequency timing signal time interval, all encoding or decoding steps are performed, but each processing element performs its tasks on different macroblocks or macroblock partitions during the low-frequency timing signal time interval.
[0154] FIG. 42 illustrates an embodiment of a video cache controller of a video codec that characterizes an embodiment of the present invention. The video cache 4202 is accessed through multiplexers 4204-4206 controlled by the circular buffer read-write address pointer. Therefore, each processing element can store different read and write address pointers from other processing elements at a given moment, so that each processing element can access the appropriate slot in the circular queue. When the block advances through the pipeline of the processing element, the read address pointer and the write address pointer associated with the block increase from the processing element to the processing element to ensure that the processing element accesses the appropriate slot in the circular queue without having to The video cache memory transfers data internally to transfer between the video cache memory and the external memory.
[0155] FIG. 43 provides a table indicating the total computational processing performed by each of certain processing elements of a video codec characterizing an embodiment of the present invention. It can be understood from this table that the size of the computational bandwidth provided by the massively parallel processing in the single integrated circuit implementation of the video codec according to an embodiment of the present invention. In order to implement a computing engine and software that will provide the same computing bandwidth, the processor that executes the software will need to work at an astonishing speed far greater than the clock speed supported by the currently available processors.
[0156] A popular integrated circuit design language is the Very High Speed Integrated Circuit Hardware Description Language ("VHDL"). Figures 44A-E provide high-level VHDL definitions of the various processing elements in the single integrated circuit implementation of the video codec according to one embodiment of the present invention as shown in Figure 33A. In FIG. 44A, definitions 4402 of various objects are provided first. Then, under the bold name of each processing element, provide the VHDL definition of the input and output of the processing element. For example, in the lower part of FIG. 44A, the input and output 4404 of the motion estimation processing element are provided. The motion estimation processing element receives four logical signal inputs 4406, a luminance macro block 4408, and a luminance reference macro block 4410, and generates three logical signals 4412-4414 and a motion vector object 4416 as outputs.
[0157] The fourth subsection: Implementation of a video codec characterized by improved integration with the memory subsystem as a whole according to various embodiments of the present invention
[0158] In this section and the following two sections, various different memory subsystems that characterize embodiments of the present invention are discussed. This subsection focuses on the various implementations of the video codec implemented by a single integrated circuit discussed in the previous subsections. These implementations feature fully integrated cameras, video codecs, and memory with increased integration. Different implementations provide different features that can be found in various market gaps for specific uses and different features for different types of imaging systems. For example, the integrated lens, sensor, memory, and codec implementations discussed below can provide efficiency and cost advantages, while the less integrated implementations of the present invention can provide greater modularity and lower initial design and manufacturing. The cost advantage.
CN 102369552 Β
[0159] FIG. 45 illustrates the components and functions of the memory subsystem of a video camera that characterize various embodiments of the present invention. The video camera system may include one or more video cameras shown in FIG. 45 as lens and sensor pairs 4502-4505, which transmit data in a stream of information encoding units (eg, bytes) to electronic storage 4506 .
[0160] As discussed above, a video camera can generate a large amount of data, and can generate data asynchronously with respect to other video cameras connected to the video camera system and a video camera data processing subsystem, such as The previously described single integrated circuit video codec. Video data is stored in the memory as frames. As discussed above, each frame is described by the position in a linear frame sequence, various video camera parameters, and a large number of data values. The video camera parameters include frame width and frame height. Each of the data values uses a fixed few bits to encode color and brightness information for different pixels in the frame. The single pixels, pixel vectors, pixel blocks, and various other data objects used by the video codec are retrieved from the memory and placed in the memory cache 4508 for the various processing elements of the video codec Access, and can be returned from the memory cache to the memory in a changed form during the video camera data processing.
[0161] Although the data path 4510 from the video cameras 4502-4505 to the memory 4506 is basically unidirectional, the data path 4512 interconnecting the cache memory 4508 and the memory 4506 is bidirectional. In several embodiments of the present invention, the memory controller 4514 provides data formatting, data addressing, and arbitration functions that allow one or more video cameras 4502-4505 to pass through the cache memory 4508 in the video codec. When accessing the memory at the same time, the memory 4506 is accessed simultaneously and asynchronously.
[0162] FIGS. 46A-E illustrate a series of video systems that characterize embodiments of the present invention, and these video systems characterize ways to increase the integration between the subsystems of the video system. The first embodiment of the present invention shown in FIG. 46A features a separate video camera 4602, a memory-bank 4604, and a video codec 4606 component subsystem, where the video codec includes sub-components Both the memory controller 4608 and the cache memory. This embodiment of the invention is discussed in detail in the previous section. The advantages of this embodiment of the present invention include the use of standard off-the-shelf memory chips for the memory bank, and a relatively high degree of modularity. When faster, more capable memory chips become available, the memory bank can be updated through memory chip replacement and relatively simple reparameterization of the video codec.
[0163] FIG. 46B shows the next embodiment of the present invention. In this embodiment, the memory controller function 4610 is directly incorporated into one or more of the memory chips 4612 of the memory bank instead of being incorporated in the video codec. This embodiment of the present invention provides a simple interface between the video codec and the memory subsystem, and in some embodiments, can reduce the communication overhead between the video codec and the memory subsystem.
[0164] FIG. 46C illustrates a third embodiment of the present invention. In this embodiment of the present invention, as in the previously discussed embodiments of the present invention, the video camera data is directly input into the memory subsystem 4616 instead of input into the memory subsystem through the video codec. In this embodiment, the video codec is further simplified, and the data communication overhead is further reduced.
[0165] FIG. 46D illustrates a fourth embodiment of the present invention. In this embodiment of the invention, the video codec and memory are combined into a single integrated circuit 4618. Since the data communication between the memory controller and the video codec is internalized in a single integrated circuit, instead of requiring a high-speed data link or bus between the video codec and the memory subsystem, by doing so, the communication The overhead will be further reduced.
[0166] FIG. 46E illustrates a fully integrated embodiment of the present invention. In this embodiment, as 4620 from above
CN 102369552 Β
As seen from the side 4622, the camera is fully integrated with the memory subsystem and video codec to create a single-chip system. In this embodiment of the present invention, camera data is directly input into the integrated circuit, without the need for a high-speed bus or serial link connecting the camera to the integrated circuit. In addition, a two-dimensional or higher-dimensional memory architecture can allow two-dimensional camera data to be directly input to the corresponding two-dimensional memory without serialization and parallelization, which improves the speed and improves the camera data to the electronic memory Storage efficiency. The final embodiment of the invention can take advantage of many new developments in materials science, nanotechnology, and nanoelectronics that provide much greater two-dimensional and three-dimensional memory densities. Full integration can also improve many potential points of failure and reduce the cost of subsystem integration.
[0167] As mentioned above, the ordering of the various implementations illustrated in FIGS. 46A-E is not meant to imply that one embodiment of the present invention necessarily replaces or is replaced by another embodiment. Each of the embodiments of the invention shown in Figures 46A-E provides a different set of tradeoffs and balances between various different costs, efficiencies, and desired characteristics. The fully integrated system shown in Figure 46E can accommodate multiple cameras, but the underlying integrated circuit will generally be larger to accommodate additional cameras, and although a two-camera implementation in which the cameras are rotated several degrees relative to each other can provide stereoscopic images The capture is to facilitate the determination of the depth in the captured image, but a fully integrated system can be characterized by relatively rigid geometric constraints. The first embodiment of the present invention shown in FIG. 46A allows the use of currently existing off-the-shelf memory chips to perform memory subsystem updates relatively easily. It is expected that the various different families of embodiments of the present invention can be based on the various architectures shown in Figures 46A-E to provide a range of performance, cost, and trade-offs.
[0168] Section 5: Characterizing the first family of memory subsystems of the present invention
[0169] In the first set of embodiments of the present invention, the memory subsystem includes a set of off-the-shelf memory chips and the memory controller subsystem in the video codec integrated circuit discussed above with reference to FIG. 46A, or various integrated A more advanced memory controller subsystem in a video data processing system based on a video camera. In the video data processing system, the memory controller is integrated in one or more memory chips. For example, refer to FIG. 46B- The embodiment of the invention discussed in C. In these embodiments of the present invention, the memory controller is used as an arbiter to coordinate the video camera and video codec as well as the serializer from one-dimensional data to one-dimensional data and the parallelizer from one-dimensional data to one-dimensional data to the memory. Asynchronous simultaneous access. The memory subsystem of the first set of embodiments of the present invention can also be coupled with other types of computing engines, or be included in other types of systems that require high-efficiency memory subsystems.
[0170] FIG. 47 illustrates the generalized interfaces provided to the camera, video codec, and memory by the memory controller embodiment of the present invention. In FIG. 47, the memory controller 4702 is represented by a rectangle. The memory and cache memory interfaces are shown on the right side 4704 of the rectangle, and the video camera and video codec interfaces are shown on the left side 4706 of the rectangle. The memory controller can interface with one or more video cameras. In Figure 47, the memory controller is shown to interface with four video cameras 4708-4711. Each video camera is connected to the input data path, line signal ("Is"), frame signal ("fs") and pixel clock signal ("pixclk") (for example, for the input data path 4716 of the video camera 4708, 1s 4718, fs 4720 And pixclk 4722) interface connection. The input data path 4716 includes a set of parallel input signal lines, each input signal line corresponding to a bit in a larger unit of information (for example, a byte). The Is signal 4718, the fs signal 4720 and the pixclk signal 4722 have been described in the previous chapters. To briefly review, Is indicates a line boundary, and fs indicates a frame boundary within the serialized camera data signal. pixclk is the clock signal supplied to the camera, where the camera usually inputs one pixel for each valid transition, or, in other words, every One pixclk interval or tick signal (tick) is input to one pixel.
[0171] The memory controller that characterizes various embodiments of the present invention provides a slightly more complex interface for the video codec or video processing subsystem of the video system. This interface includes video frames stored in memory
CN 102369552 Β
The input of the X coordinate 4730 and the y coordinate 4732 of the pixel or pixel block, the input of the frame number 4734, the input of the opcode 4738, the "program" 4740, the "write" 4742 and the "select" 4744 signal input, and when the memory controller supports When input from two or more video cameras, the camera's index 4736 is input. It should be noted that the interface described with reference to FIG. 47 is an exemplary interface, and specific embodiments of the present invention may utilize various different specific interfaces including different data paths and signals and/or more numbers. Or a smaller number of data paths and signals. For example, instead of inputting the X coordinate 4730 and the y coordinate 4732 to the memory controller, an alternative interface can provide input of linear pixel addresses within a linear pixel sequence. In this case, when the two implementations support the same maximum frame size, linear pixel address input will require twice the signal line used for X-coordinate input or y-coordinate input in the previously discussed implementation.
[0172] In some embodiments of the present invention, the x-coordinate and the y-coordinate include an (x, y) coordinate pair, which describes the position of the uppermost point on the left in the block or is stored in the memory The pixel vector or the position of a specific pixel within the video frame. Both the X coordinate and the y coordinate are supplied through a set of parallel signal lines with sufficient width to represent the largest possible coordinate within the largest possible frame size supported by the video system as a binary number. Note that, contrary to normal mathematical conventions for coordinates X and y, the y coordinate corresponds to the row index within the frame, and the X coordinate corresponds to the column index. Again, alternatively, when the frame is considered a linear pixel sequence, rather than a two-dimensional pixel array, the pixels within the frame can be addressed by linear index or position.
[0173] The frame number 4734 is also supplied by a set of sufficient number of parallel signal lines to express the largest possible frame index for the frame stored in the memory as a binary number. For example, two signal lines are enough to express the number of each frame in the set {0,1,2,3}. Similarly, the camera index 4736 is supplied by a sufficient number of parallel signal lines to express the index or camera number of each video camera connected to the memory controller. Operation code 4738 is also provided by a sufficient number of parallel signal lines to encode all required operation codes into binary numbers. In one embodiment of the present invention, the "program", "write" and "select" signal lines are single signal lines that provide binary "0" or "1", and "0" or "1" is also called off. And open or low and high. Any possible unary signal encoding can be utilized within the specific memory controller implementation of the present invention. For example, the Boolean value "0" input to the "write" signal line may indicate a write operation in some embodiments of the present invention, but may indicate a read operation in an alternative embodiment of the present invention. In any particular embodiment of the invention, the convention utilized is fixed. In one embodiment of the present invention, the memory controller is initialized by setting the "program" signal line 4740 to high and then passing through the remaining signal lines and data paths Some of them provide initialization data. For example, the maximum frame size may be specified as the product of the maximum possible value of the X coordinate and the y coordinate input through the X coordinate signal path 4730 and the y coordinate signal path 4732. Similarly, the maximum number of frames that can be stored for each camera can be input using the frame number signal line 4734. The maximum number of supported cameras can be input through the camera index signal path 4736 as the maximum possible camera index. The video codec requests memory-to-cache or cache-to-memory operations in the following manner, namely, inputting the operation code through the opcode signal path 4738, inputting the frame number through the frame number signal path, and passing the camera When the index signal path is used to input the camera index and the X coordinate and y coordinate are input through the X signal path and the y signal path, set the "Select" signal line to high. In the single video camera implementation of the present invention, the camera index signal path is not required. In one embodiment of the present invention, the cache-to-memory operation is initiated in the following manner, that is, when the memory-to-cache operation is requested by setting the "write" signal line to low, the " The write signal line 4744 is set to high. The precise timing and sequence of the input signals can vary from one embodiment of the invention to another. In some embodiments of the present invention, each opcode characterizes a different type of memory access, as described in the previous section
CN 102369552 Β
As discussed, the memory access types include access to pixels, pixel vectors of various sizes, pixel blocks of various sizes, and data objects.
[0174] The generalized memory controller interface illustrated in FIG. 47 includes a first interface for the memory, the first interface including a bidirectional data path 4750, an address path 4752, and a control path 4754. The memory controller interface additionally includes a second interface for the cache memory, and the second interface also includes a data path 4756, an address path 4758, and a control path 4760. Data is exchanged between the memory and the memory cache through a set of parallel signal lines 4750 and 4756, respectively. In some embodiments of the present invention, the data path is bidirectional and is used to write data to and from the memory The memory receives both of the data. In other embodiments of the present invention, separate unidirectional data paths can be used to write data to and read data from the memory. Each data path includes a sufficient number of parallel signal lines to carry information Units, such as bytes, words, long words, or longer bit sequences. In many implementations, each signal line in the data path corresponds to a different side within the memory bank or memory architecture. For example, an 8-signal line data path can interconnect the memory controller with eight memory chips, where each byte is received from cameras distributed outside the 8 memory chips. Various other mappings and data path sizes are possible, especially including wider data paths and various types of interleaved mappings of data units and storage planes.
[0175] The memory address is supplied to the memory and the cache memory through address paths 4752 and 4758 as part of each access request. In some embodiments of the present invention, the memory controller converts serialized camera data and related signals or two-dimensional data block specifications into linear memory addresses. The memory is usually also supplied with control signals through the control paths 4754 and 4760, such as the "select" signal and the "write" signal (4742 and 4744) supplied by the video codec to the memory controller.
[0176] FIGS. 48A-H illustrate the components and operation of the components of the memory controller that characterizes one embodiment of the present invention. These components will be described in more detail in the following paragraphs with reference to the following drawings. Figures 48A-H all use similar graphical conventions, which are discussed next with reference to Figure 48A.
[0177] The memory controller 4802 is characterized as a rectangle. As discussed above, the memory controller may be implemented as a module or subsystem within the video codec, alternatively, may be implemented as a subsystem or module within one or more memory chips, or may be implemented as Subsystems or modules within video codecs and memory integrated circuits. It is even possible to implement the memory controller as an independent component in some complex systems. The memory controller includes a camera block corresponding to each video camera connected to the memory controller. In FIGS. 48A-H, three camera blocks 4804-4806 are shown. The camera blocks receive data and signals from the video camera, and output control signals and data to the memory, thereby enabling the video camera to send data to the memory controller. Data stream interface. The memory sequencer 4808 implements the video codec-to-memory controller interface discussed with reference to FIG. 47. Although a single memory sequencer is shown in FIGS. 48A-H, alternative embodiments may utilize multiple memory sequencers to facilitate higher levels of interconnection between the video codec and the memory controller. The degree of parallelism.
[0178] The memory subsystem is illustrated with a rectangle 4810, which is connected to the memory controller through a data input path 4812, a data output path 4814, an address path 4816, and a control path 4818. As discussed above, certain memory subsystems may combine the data input path 4812 and the data output path 4814 into a single bidirectional data path. The memory controller additionally includes an arbiter 4820, a clock input 4822, three multiplexers 4824-4826, and an address conversion unit 4828. Finally, the memory controller includes a data input port 4829 for receiving data from the cache memory and a data output port 4830 for sending data to the cache memory. In some embodiments of the invention, these ports can be combined in a bidirectional data port. The same or separate address conversion components can be used
CN 102369552 Β
To specify cache memory access, or cache memory access can be specified by another system (such as a video codec).
[0179] FIGS. 48B-C illustrate the function of the arbiter (4820 in FIG. 48A). As shown in FIG. 48B, the arbiter receives a request signal from each of the camera block and the memory sequencer. The camera block or the memory sequencer initiates the memory operation by setting the request signal line to "1", which interconnects the camera block or the memory sequencer and the arbiter. The arbiter 4820 arbitrates between asynchronous requests and synchronous requests received from the camera block and the memory sequencer block, and at each moment, selects a single one of the camera block and the memory sequencer to perform the next request service. . In other words, the arbiter continuously receives indications of pending requests from the camera and the memory sequencer, and serializes the incoming request stream by allowing only a single camera or memory sequencer to perform memory transactions at each moment . The arbiter sends the enable signal encoded as the binary value "1" in the presently discussed embodiment of the present invention to each camera block in the memory sequencer block through the enable signal lines 4836-4839. At any given moment, only a single enable signal line is high. The arbiter needs to ensure that only a single camera block or memory sequencer block is performing memory transactions at any given moment, and also to ensure fair processing of requests initiated by the camera block and memory sequencer over time, so that All camera blocks and memory sequencers get enough memory bandwidth, and no input data fails to reach the memory. As shown in Figure 48C, the enable signal line 4836-4839 is also input to three multiplexers 4824-4826 to allow the multiplexer to select data coordinates, pixel coordinates, pixel vector coordinates or pixel block coordinates, and control from the currently enabled The signal of the camera block or memory sequencer.
[0180] FIG. 48D illustrates the function of the first multiplexer 4824. Each of the camera block 4804-4806 and the memory sequencer block 4808 can send data to the memory 4810. The data path output by each camera block and the data input port 4829 controlled by the memory sequencer block are connected to the first multiplexer 4824, and the first multiplexer 4824 selects one of the four input data paths for receiving The data forwarded to the memory 4810 by the first multiplexer. In FIG. 48D, the memory sequencer block is currently enabled by the arbiter, and therefore, the first multiplexer 4824 sends the data received from the cache memory by the first multiplexer 4824 through the data input port 4829 to the memory. As shown in FIG. 48E, the multiplexer 4825 sends the coordinates of the data from one of the camera blocks or the coordinates of the pixels, pixel vectors, pixel blocks or data objects from the memory sequencer block to the address conversion unit 4828 at a given moment. The address conversion unit 4828 converts the coordinates into the corresponding linear memory address, and the address conversion unit then forwards the linear memory address to the memory. For camera input, the linear memory address is calculated as the position within the line of the frame, which is added to the sum of the camera's base offset and the frame offset. Similarly, as shown in FIG. 48F, the multiplexer 4826 receives control from each of the camera block and the memory sequencer block. The control signal is sent from the camera block or the memory sequencer block, which is currently enabled by the arbiter, to the memory at any given moment. As shown in FIG. 48G, when the memory sequencer block is enabled by the arbiter, during the read request initiated by the memory sequencer block to the memory, data flows from the memory through the memory controller via the data output port 4830 to the high-speed memory. Buffer memory.
[0181] As shown in FIG. 48H, the clock input 4822 is input to the memory controller and is routed to each of the camera block and the memory sequencer, arbiter, and is routed to any other memory control that uses a clock synchronization signalDeviceComponents. The clock signal called "fastclk" in the memory controller is usually a multiple of the fastest expected pixclk signal. The multiple η varies according to the number of camera blocks, and in some embodiments of the present invention, the multiple η may be a configurable parameter. In one embodiment of the present invention, when the four camera blocks are interconnected with the memory through the memory controller, a multiple of n = 16 is used. The internal memory controller clock signal called "fastclk" below needs to have a sufficiently high frequency to Allows the memory controller to simultaneously provide all storage initiated by the camera block and the memory sequencer block
CN 102369552 Β
Device access request service. For example, in the case of four video cameras connected to the memory controller, when all four cameras send data at the fastest possible rate of one pixel per pixclk interval, during the span of 16 fastclk intervals, on average, Four pixels are sent through the memory controller. Since the memory controller performs internal operations according to the fastclk frequency, including memory access operations, even under the full data transfer load supplied by all four video cameras, the memory controller will pass enough fastclk cycles through the memory sequencer block ( And from this, the memory bandwidth) is provided to the video codec.
[0182] FIGS. 49A-C illustrate an implementation of the arbiter discussed with reference to FIGS. 48B-C, which is a component of a memory controller that characterizes an embodiment of the present invention. Figure 49A provides a state transition diagram and a symbolic representation of the camera block arbiter block. The camera block arbiter block is a component of the memory controller arbiter and is implemented as a state machine. The symbolic representation 4902 of the camera arbiter block indicates that the camera arbiter block receives two input signals and outputs two output signals. Input signals include Block4n signal 4904 and Req signal 4906. The camera arbiter block outputs Grant signal 4908 and Block-Out signal 4910. The input signal and the output signal are single signal line signals, each of which carries one of two binary values "0" and "1" at any particular moment. Therefore, the state of the camera arbiter block can be expressed as four Boolean terms or quantities. The Boolean terms or quantities are currently input to the camera arbiter block through two input signal lines and the camera arbiter block passes two output signals. Corresponding to the value of the line output. The state is re-evaluated and can be changed every fastclk interval.
[0183] FIG. 49A also provides a state transition diagram 4912, which characterizes the operating characteristics of the camera arbiter block. In the state transition diagram 4912, each circle (eg circle 4914) represents a different stable state of the camera arbiter block. The straight arrow of the interconnected state characterizes the state transition, such as the state transition arrow 4916, and the curved arrow (such as the curved arrow 4918) indicates that there is no state transition, or, in other words, that the camera arbiter block is in two or more fastclk Keep in single state during the interval. Each state (for example, state 4914) is designated by a name, for example, the name is the name "pending" with which the state 4914 is designated, and also includes an indication of the current output of the two output signals. In the state pending 4914, for example, the camera arbiter block outputs "0" to the Grant output, and outputs "1" to the BlockOut output. The state transition occurs for a specific value of one or two input signals. For example, when the Req input signal 4906 is low, or as expressed in Figure 49A, when the Boolean expression ""Req" has the value TRUE, the camera arbiter block transitions from state 4920 to state 4922. When using a single variable Boolean The expression shows that when the state transitions, the result of inputting a specific value through a single input signal line is a state transition. When a two-variable Boolean expression shows the state transition, then two input signals The number line must have a specific value for the state transition to occur.
[0184] The camera arbiter block can be in any of three states: (1) IDLE state 4922; (2) PENDING state 4914; and (3) GRANT state (4920). In the idle state, the associated camera has no memory control that is requesting to perform a memory access operation. In the pending state, the associated camera is sending data written to the memory, but is prevented by another camera or video codec from gaining memory control. In the granted state, the camera associated with the camera arbiter block can access the memory.
[0185] When powered on, the initial state of the camera arbiter block is "idle" 4922, where the camera arbiter block outputs Boolean "0" to the Grant output signal line, and will communicate with the signal received through the Block-In input signal line The opposite value is output to the Block-Out output signal line. When the Req input signal line is low, the camera arbiter block remains in the idle state 4922. When the Req input signal becomes high or the input value "1", when the Block-In input signal line is low, the camera arbiter block transfers 4926 to the grant state 4920, or when the Block-In input signal line is high, transfers 4928 to
CN 102369552 Β
Pending state 4914. In the idle state 4922, the cameras associated with the camera arbiter block are not currently initiating memory requests, and no other cameras are currently initiating memory requests. When a higher priority camera has requested a memory transaction, the currently considered camera arbiter cannot transfer to the grant state, but instead transfers to the pending state to wait until no higher priority camera is requesting a memory transaction. The pending state 4914 represents a state in which the camera associated with the camera arbiter block is requesting a memory transaction, but since a higher priority camera is currently performing a memory transaction with the memory, the aforementioned memory transaction Cannot be activated. In the pending state, the camera arbiter block outputs the boolean value "1" to the block-out signal line to prevent lower priority cameras or memory sequencers from being granted access to the memory until the camera arbiter The camera associated with the block has a chance to execute the requested memory transaction. The grant state 4920 represents a camera arbiter block state in which the associated camera can perform memory transactions. In the grant state 4920, as indicated by the camera arbiter block outputting "1" to the grant signal, the camera associated with the camera arbiter block is Enable. Note that the camera cannot remain in the grant state 4932 for more than a single fastclk interval. Therefore, the camera is only allowed to transfer a single data unit to the memory before handing memory control to another currently requested block.
[0186] FIG. 49B shows a state transition diagram and symbolic representation of the memory sequencer arbiter block in the memory controller arbiter that characterizes an embodiment of the present invention. The symbolic representation 4940 of the memory sequencer arbiter block instructs the memory sequencer arbiter block to receive the same two input signals as the input signal received by the camera arbiter block, and only output the Grant output signal instead of the camera arbiter Both the Grant output signal and the Block-Out output signal of the block output. Therefore, the state transition diagram of the memory sequencer arbiter block 4942 is simpler than the state transition diagram of the camera arbiter block. The feature of the memory sequencer arbiter block is that there is only a single idle state 4944 in the state transition diagram, and importantly, the memory sequencer arbiter block can remain in the grant state 4946 during multiple fastclk intervals. In the multi-memory sequencer implementation, all except the last memory sequencer arbiter block includes a Block-Out output signal and a state transition diagram slightly different from the state transition diagram shown in FIG. 48B.
[018] FIG. 49C illustrates an embodiment of a memory controller arbiter within a memory controller that characterizes an embodiment of the present invention. The arbiter includes a series of camera arbiter blocks 4960-4962 and a single memory sequencer block 4964. According to the order of the associated camera arbiter blocks in the camera arbiter block sequence, the cameras associated with the camera arbiter blocks 4960-4962 are prioritized in descending order. Therefore, the camera associated with the camera arbiter block 4960 has the highest priority. Therefore, the memory sequencer associated with the memory sequencer arbiter block 4964 has the lowest priority. However, as discussed above, the memory sequencer can remain in the granted or enabled state for multiple fastclk intervals, while a higher priority camera can only control the memory within one fastclk interval. The high-priority block is connected to the low-priority block through the Block-Out output signal line of the high-priority block and the Block-In input signal line of the low-priority block. When a particular block is in the pending or granted state, all low-priority blocks are prevented from transitioning from the pending state to the granted state.
[0188] FIG. 50 provides a simple graphical illustration of timing considerations for a memory controller arbiter implemented within a memory controller that characterizes an embodiment of the present invention. In Figure 50, the complete loop of circle 5002 represents one pixclk clock interval. Each increment along the circle (eg increment 5004) characterizes a fastclk interval. When all four cameras are sending data at the maximum data transfer rate, each camera sends one data unit at each Pixclk interval or during the time span represented by a full circle in FIG. 50. Even though the memory controller in Figure 50 performs memory transfers on behalf of the camera in two fastclk intervals, no matter the timing or sequence of the data units received from the camera, at least half of the bandwidth of the memory controller (not shaded in Figure 50) Can still be used for storage
CN 102369552 Β
Sequencer or video codec. In Figure 50, it is shown that the memory transfer occurs in one fastclk interval, so that 3/4 of the fastclk interval can be used for video codec memory operations. When the multiple n of the fastclk interval of each pixclk interval increases, the increased part of the memory controller bandwidth can be used for the video codec. Therefore, the number η of fastclk intervals within the pixclk interval can be adjusted to balance between the service requested by the video codec memory access and the service requested by the camera memory, and to keep the memory controller clock rate as low as possible In order to consume as little power as possible and ensure that the memory access does not exceed the maximum access rate to the memory.
[0189] FIGS. 51-54 provide schematic diagrams of a memory controller characterizing an embodiment of the present invention. Figure 51 provides a schematic diagram of a single camera memory controller characterizing an embodiment of the present invention. In this embodiment, in addition to the memory sequencer block 5104, there is a single camera block 5102. The memory controller is driven by the fastclk signal 5106, and the fastclk signal 5106 oscillates n times the pixclk 5108 signal of the camera. The arbiter includes a camera arbiter block 5110 and a memory sequencer arbiter block 5112. In the embodiment illustrated in FIG. 51, the memory control signal is generated directly from the arbiter output signal and the signal input to the memory controller through a logic gate. , Instead of generating memory control signals through a separate control multiplexer. The first multiplexer 5114 sends the camera data or the memory cache data to the memory according to whether the camera block or the memory sequencer is currently enabled by the Grant signal from the respective arbiter block. The camera block 5012 converts the camera input signals Is, fs, and pixclk into X coordinates and y coordinates. The X coordinates and y coordinates are outputs 5116 and 5118 to the coordinate multiplexer 5120, and the coordinate multiplexer 5120 transmits the input coordinates to On the address conversion unit 5122, the address conversion unit 5122 generates the address conversion unit and sends it to the memory Memory address. The memory sequencer block 5104 converts the input X coordinate 5124 and y coordinate 5126 into the memory X coordinate 5128 and y coordinate 5130, respectively.
[0190] FIG. 52 shows a schematic diagram of a multi-camera memory controller characterizing an embodiment of the present invention. In this embodiment, a plurality of camera blocks 5202-5203 are respectively associated with a plurality of camera arbiter blocks 5204 and 5205.
[0191] FIG. 53 provides a schematic diagram illustrating a camera block included in the memory controller implementation shown in FIGS. 51 and 52. The camera block includes a first-in first-out ("FIFO") buffer 5308, in which the information unit input by the associated video camera is temporarily stored for output to the memory. This FIFO buffer provides sufficient flexibility in data transmission timing to allow the arbiter to balance competing requests from multiple cameras and from the memory sequencer. In the embodiment shown in FIG. 53, the camera block generates the request signal 5310 only when the FIFO buffer is almost full. Otherwise, the camera block defers to other camera blocks and memory sequencer blocks. Specifically, the memory sequencer is allowed to execute memory transactions initiated by the video codec to the fullest possible extent, but to ensure that the All camera data is transferred to the memory to prevent any camera data loss. In one embodiment of the invention, the FIFO for each camera block includes space for storing four data units. The camera block also detects redundant data in the data stream input by the camera, and when the redundant data is detected, sends a redundant UV signal 5312. In the previous chapter, the camera output redundant data was discussed. This redundant UV signal allows the camera to be discarded The redundant data in the data stream sent by the camera is not inserted into the FIFO to prevent unnecessary power consumption and allocate memory controller cycles for sending redundant camera data to the memory. In an alternative embodiment, the FIFO queue may contain entries having a size sufficient to accommodate a plurality of information units, which are sorted relative to the FIFO entry and in the FIFO queue. In some embodiments of the present invention, the camera block can also change the order of the information units in the data stream by reordering the received information units in the FIFO queue. As discussed, the information unit can be a byte, a word, or another fixed-size binary information unit.
[0192] In certain embodiments of the invention, a more complex, content-addressable FIFO ("CAM FIFO") block is utilized.
CN 102369552 Β
CAM FIFO includes multiple slots, each slot contains multiple entries, and each slot is associated with X and y coordinates, Y, U, or V indicators and counters. The CAM FIFO allows camera inputs to be reordered for output to memory and output larger data units, each of which contains multiple camera input data units. As an example, although the camera block receives 8-bit pixel values from the camera, the FIFO can be implemented to output 32-bit words. In this embodiment, as shown in FIG. 53, when the counter associated with the slot has the largest counter value as indicated internally and as indicated by the read-OK signal used, the FIFO is almost full and will When the next slot to be transmitted is full, the request signal is pulled high together with the enable and almost full signals to determine the output of the request signal.
[0193] FIG. 54 provides a schematic diagram of a memory controller-to-memory interface in an embodiment of the present invention. The "write" input 5402 and the "select" input 5404 of the memory 5406 are generated from the arbiter Grant signals 5408-5411 and the "write" signal, and the "write" signal is the input 5412 to the memory sequencer.
[0194] The memory controller implementation of the present invention discussed in this section thus provides several advantages for video processing systems. As discussed above, the memory controller provides fair arbitration for the multiple cameras that input data streams into the memory and for the video codec that processes the data input from the cameras, without starvation or data loss. The memory controller filters the video camera data so that memory controller cycles are not wasted to transfer redundant data to the memory, and memory bandwidth that is otherwise wasted on redundant data can be replaced for video codec initiation Memory access operations. The memory controller that characterizes the embodiment of the present invention provides a simple interface to the video codec through the memory sequencer subsystem, so that the result of a single high-level memory operation is that two dimensions can be exchanged between the memory and the memory cache. Data, the two-dimensional data includes macroblocks and vectors. The high clock frequency required to process the input data from the camera and the memory request from the video codec can be gathered in the memory controller instead of spreading in the entire video codec, reducing energy consumption and video codec design And the cost of production.
[0195] In general, the memory subsystems of the currently discussed memory subsystem family can be used in various applications in which one or more data sources simultaneously connect to the memory controller data stream interface. The data stream is sent to the memory, and at the same time, the calculation engine or other device reads data from the memory and sends the data to the memory in a random access manner, and specifically, exchanges with the memory in a single two-dimensional random access memory operation. Dimensional data object.
[0196] Section 6: Characterizing the second family of memory subsystems of the second set of embodiments of the present invention
[0197] As discussed in the previous subsections, the implementation of the present invention featuring a memory bank implemented by a separate standard memory chip is made possible by a memory controller that operates at a relatively high clock rate. The memory exchange operation is performed between the video camera and the memory and between the video codec and the memory through the memory cache. Due to the need to serialize and parallelize the data between the memory cache and the memory, and to exchange data through several data communication interfaces, a large amount of data communication overhead is incurred. In addition, since the standard memory chip can only service a single memory request at each moment, the memory controller operates at a relatively high frequency to provide enough arbiter cycles to fairly provide a fair share of the simultaneous transmission of cameras and video codecs. Arbitrate between memory accesses.
[0198] The new type of memory disclosed in this section (called "multi-access memory") can be used in video processing systems and many other types of computing systems to provide parallel memory for multiple memory access devices. Without arbitration and/or multiplexing. This multiple access memory can provide a data storage medium that is much more efficient than traditional RAM memory, and can be customized to provide access rates and operations that match the needs of video processing systems or other types of computing systems. Speed and performance.
CN 102369552 Β
[0199] FIG. 55 illustrates the operation of a multiple access memory characterizing an embodiment of the present invention. In a video processing system application (such as the video processing system discussed in the previous section), the memory 5502 is divided into multiple partitions 5504-5507, and each partition is respectively associated with a specific video camera 5508-5511. The memory may include an additional partition 5512 for storing data unique to the video codec. Each camera can write the frame line 5514-5517 into the memory at the same time without interfering with other cameras and not being interfered by other cameras. At the same time, the video codec can access any part of the memory that is not currently written by the camera (for example, macro block 5518) for writing or reading. In FIG. 55, the macro block 5518 is written from the memory to the corresponding macro block 5520 in the cache memory 5522. In this way, the multi-access memory supports simultaneous access to each of the multiple video cameras and the video codec through the memory cache. The only limitation of multiple parallel access is that a single line in a frame stored in the memory cannot be accessed by multiple entities at the same time.
[0200] FIG. 56 abstractly illustrates the operation of a multiple access memory characterizing an embodiment of the present invention. The memory includes a grid of memory storage elements 5602, and each element is also referred to as a "unit" or "unit" for storing a single bit. Multiple grids can be combined into multiple faces in the memory subsystem, where information units (such as bytes) are distributed across the entire face, or distributed according to more complex distribution patterns. In the video data processing system, the lines of the video frame are stored in the horizontal rows of the grid, and the horizontal rows are, for example, the shaded horizontal rows 5604 in FIG. 56. The camera data is input to the memory through the line multiplexer 5608 through the sequential shift operation discussed below. As discussed above, camera data is a stream of information units that can be interpreted as information units within a line of a video frame by using the Is signal and fs signal supplied by the camera. The row demultiplexer 5608 demultiplexes the camera data stream into rows of frames, and fills each row continuously, one row at a time, through repeated sequential shift operations during the memory write operation. The memory can also be accessed in a random manner through the column demultiplexer 5610 and the row demultiplexer 5612. The column demultiplexer 5610 and the row demultiplexer 5612 together provide two-dimensional access to the multiple access memory. Random access. In order to select a specific memory unit 5620 to perform For read or write access, as shown in FIG. 56, the x coordinate 5614 is input to the column multiplexer 5610, and the y coordinate 5615 is input to the row multiplexer 5612. In this way, the memory supports the video codec in the video processing system application. One-dimensional random access and basic one-dimensional write access for writing the frame line from the data stream input from the camera to the memory. Since each camera in the multi-camera video processing system is provided with a different and separate memory partition and line demultiplexer, the cameras do not interfere with each other and can write to the memory at the same time. Assuming that the access part of the memory is not simultaneously written by the camera, the video codec that accesses the memory through the column multiplexer 5614 and the row multiplexer 5612 can access any part of the memory. The simultaneous access of the camera and the video codec to a specific memory unit can be prevented by the conflict detection subsystem, or alternatively, the video codec can be used to monitor the camera input behavior and ensure that there is no access to the current data stream from the camera Input video frames to prevent.
[0201] FIG. 57 illustrates a multi-faceted memory system according to an embodiment of the present invention. In FIG. 57, each face (for example, face 5702) of the storage face group 5704 is a memory grid, such as the memory grid 5602 shown in FIG. 56. Each of the camera and the video codec accesses the memory through the decoder 5706 and the data channel 5708. The common signal 5710 output from the decoder drives the memory access operations in all memory planes. On the contrary, the data input from the data channel is distributed across the entire storage surface. In the embodiment shown in FIG. 57, each bit of one byte of data is sent to and stored in a separate storage plane. Therefore, FIG. 57 shows eight storage planes, the bits of each received byte are written to these eight storage planes, and the bits of each byte sent from the memory are retrieved from these eight storage planes. As mentioned above, information units other than bytes can be received by the multiple access memory and can be distributed according to various distribution schemes, including more complicated interleaving schemes.
CN 102369552 Β
[0202] FIG. 58 illustrates the division of a memory partition associated with each camera in a multiple access memory according to an embodiment of the present invention. In FIG. 58, the memory partition 5802 associated with the first camera is further divided into frame-sized memory areas, such as a memory area 5804. The memory system treats consecutively sorted frames in the memory as a memory FIFO for processing. These frames basically form a circular FIFO queue 5806, and the video camera data is continuously streamed to the circular FIFO queue 5806. The current frame pointer 5808 indicates the current frame in the memory partition to which the camera is currently sending data or will then send data to the current frame. The frame pointer is increased by detecting edges in the fs signal output from the camera to the memory.
[0203] FIG. 59 illustrates writing a frame to a multiple access memory according to various embodiments of the present invention. In Figure 59, the two-dimensional representation 5902 of the frame data is shown as a linear grid indexed with a column index X and a row index y. As shown in FIG. 59, data is written to the multiple access memory 5904 in a mirror reflection manner, and the mirror reflection reverses the column index X. In the lower part of FIG. 59, the first four data transfer shift operations during the transfer of one row of data to the multiple access memory are illustrated. The first data unit 5906 in the row is transferred to the first data storage unit 5908 in the corresponding row of the multiple access memory. The next data unit 5910 is transferred to the first data storage element 5908 of the memory row, and the value existing in the first memory element is simultaneously transferred to the second data storage unit 5912 in the row. In other words, each data unit is input through a shift operation, where the entire multiple access memory row operates as a very long shift register. Each subsequent data value is transferred to the first data storage unit of the memory row, where all data values currently stored in the row are shifted one position to the right. This operation generates a mirror reflection of the video frame in the multiple access memory. The multiple access memory contains lines of sufficient size to hold the video frame of the largest possible width. You can use fewer shift operations to store narrower video frames in a multiple-access memory. The video can be shifted line by line through another shift operation The frame is moved out of the opposite side of the memory to read the video frame from the multiple access memory. This shift-based video frame reading operation can be used, for example, to send decompressed video data to a display device for display.
[0204] Next, an implementation of the multiple access memory discussed above with reference to FIGS. 55-59 is provided. Figure 60 illustrates the signal inverter. The symbol representation of the signal inverter 6002 is provided in the center of FIG. 60. The signal inverter reverses the input "1" digital signal to "0" and reverses the input "0" digital signal to "1". As shown in the schematic diagram 6004 in FIG. 60, the signal inverter can be implemented using two complementary metal oxide semiconductor ("CMOS") transistors. The p-type transistor 6006 and the n-type transistor 6008 are connected in series. The source of the p-type transistor is connected to the voltage corresponding to the Boolean value "1" 6010, and the drain of the n-type transistor is connected to the ground 6012. Inputting the boolean value "0" 6014 activates the p-type transistor 6006 and stops the n-type transistor 6008, thereby obtaining an output 6016 of the boolean value "1". Similarly, as shown in the schematic diagram 6020 in FIG. 60, input the boolean value "1" "Get the output of the boolean value "0".
[0205] FIG. 61 shows a schematic diagram of a memory unit or a memory cell of a multiple access memory used to characterize an embodiment of the present invention and a symbolic representation of the memory cell. The symbolic representation 6102 of the memory cell shows a memory cell that receives five input signals: (1) "Ain, the twos complement of the value that is input through a random access memory access operation to be stored in the memory cell; (2) SIn , The value input to the row demultiplexer 5608 in FIG. 56 to be input to the memory cell based on the shift operation; (3) Shift1 and Shift2 signals 6102-6107, used to shift the value input on the signal line Sin To the memory cell; (4) Awrite, to control the signal line "Ain 6104 ± input value to be written into the memory cell; and (5) OutRd 6109, to control the output of the memory cell content to the output signal line Out 6112. Memory The unit continuously outputs the currently stored value to the output signal line SOut 6110. As shown by the symbolic representation of the memory unit 6102 in FIG. 61, the memory unit outputs the input signals Shiftl, Shift2, Awrite, and OutRdo unchanged in FIG. 61. Schematic diagram of the memory cell 6120ο Set the signal line OutRd to high conduction
CN 102369552 Β
The transistor T4 6122, and causes the data stored in the memory cell in the flip-flop 6131 to be output to the output signal line "Out" 6124. Inputting the Boolean value "1" to the input signal line "Shiftl" 6125 activates the transistor T1 6126 to input the value of the signal line Sin 6127 ± to the inverter II 6128. Inputting the Boolean value "1" to the input signal line Shift2 6129 activates the transistor T2 6130 to input and output the reverse Sin input to the flip-flop 6131, which includes inverters 12 6132 and 13 6133. The trigger positively stores the complement of the input value without incremental refresh<sub>o</sub>The flip-flop thus constitutes a static bit storage unit. In contrast, the inverter II 6128 relies on dynamic charge transfer, which is used to reverse the input signal Sin and transfer the reversed signal 2 to the flip-flop 6131. In the memory write shift operation cycle, the first pass Input "1" to Shiftl to activate the transistor T1 to input the signal Sin to the inverter II, and then stop or turn off the transistor T1, and then activate the transistor T2 to reverse the reverse Sin signal The device II is transmitted to the trigger 6131. The interval of the shift operation is fast enough that the value input to the inverter Π does not disappear before the complement of the value is output to the flip-flop. In order to input the value of the signal "Ain 6136 into the flip-flop 6131, the input signal Awrite 6134 activates or closes the transistor T3 6135. Use Sin, Shiftl and Shift2 inputs to write camera data into the memory unit, and pass Awrite, OutRd and " Ain input writes data to the data unit of the video codec and retrieves data from the data unit of the video codec.
[0206] FIGS. 62A-C illustrate the shifting of data into a memory cell of a multiple access memory that characterizes an embodiment of the invention. Initially, as shown in FIG. 62A, the memory cell stores the Boolean value "0" in the flip-flop 6202. Currently, the boolean value "1" is input through the input signal line Sin 6204. As shown in FIG. 62B, the input signal line Shift1 6206 is pulled high to activate the transistor T1, and the input value "1" is transferred from Sin to the inverter II 6207. Next, as shown in FIG. 62C, pull down the input signal line Shift1 to stop the transistor T1 6208, and to activate the transistor T2 6212 and input the Boolean value "0" output by the inverter Π into the flip-flop 6214, pull The high input signal line Shift2 6210, in the flip-flop 6214, the input value "0" is inverted again and stored as the Boolean value "1". Note that in FIG. 62B, if the second memory cell is located to the right of the illustrated memory cell, and the Sin input of the second memory cell is connected to the SOut output 6216 of the illustrated memory cell, and the illustrated memory cell is The Shift1 and Shift2 of the memory unit are directly connected to the Shift1 and Shift2 of the second memory unit. The previous value will be shifted to the second memory unit for storage during the shift operation. Therefore, interconnecting the memory cells together in a row produces a shift register.
[0207] FIGS. 63A-C illustrate writing the Boolean value "0" to the memory cell currently storing the Boolean value "1" according to an embodiment of the present invention. Figures 63A-C use the same illustration conventions used in Figures 62A-C.
[0208] FIGS. 64A-B illustrate the output of the value currently stored in the memory cell of the multiple access memory to the output signal line according to an embodiment of the present invention. In FIG. 64A, the memory cell currently stores the Boolean value "1" 6402. By pulling up the OutRd signal line 6404 as shown in FIG. 64B, the transistor T4 6406 is activated, and the stored Boolean value "1" is output to the output signal line Out 6408. Figures 65A-B illustrate the writing of values into a memory cell characterizing an embodiment of the present invention via two input signal lines. As shown in Figure 65A, the memory cell currently stores the Boolean value "0" 6502. The input of the Boolean value "1" to the signal line Awrite 6504 and the input of the Boolean value "0" to the signal line "Ain 6506" result in the Boolean value "1" 6508 being stored in the memory cell.
[0209] FIGS. 66A-B show an embodiment of a 4X4 memory storage array using 16 memory cells of the type illustrated in FIG. 61 and a symbolic representation of the 4X4 memory storage array. In FIG. 66A, 16 memory cells (including memory cell 6602) of the type illustrated in FIG. 61 are arranged to provide a two-dimensional grid of a 4×4 memory storage array. Each memory cell stores a single bit of information. As discussed above with reference to Figure 56, each 4X4 array
CN 102369552 Β
One row can be written by a series of four shift operations as discussed with reference to FIG. 59, or it can be written by two-dimensional random access as discussed with reference to FIG. 56. During a random access write, data is input to the Ain input of the memory cell. Therefore, during a random access write, the memory cells are accessed one by one. Similarly, the Out output of a specific memory cell is during a random access read operation. Output value. In Figure 66A, "Ain input is shown connected to a two-dimensional input signal line grid labeled "AJ, where X and y are the coordinates of the memory cells in the 4X 4 array, and the Out output is connected To the two-dimensional input signal line grid marked "", where X and y are the coordinates of the memory cells in the 4X4 array. Figure 66B shows the symbolic representation of the 4X4 memory cell array. The characteristics of the 4X4 memory cell array are transferred Wr and Rd inputs and φΐ and Shen 2 inputs to each memory cell in the row, Wr and Rd inputs are connected to the Awrite and OutRd inputs of the first memory cell in each row, and φΐ and φ2 inputs are connected to each row Shiftl and Shift2 inputs of the first memory cell.
[0210] FIG. 67 shows a schematic diagram of a larger capacity memory based on a 4X4 memory storage array (for example, the 4X4 memory storage array shown in FIG. 66A) according to an embodiment of the present invention. The row decoder 6702 together with the demultiplexers 6730 and 6731 operate as the row demultiplexer 5608 in FIG. 56. The column decoder 6704 and the row decoder 6706 together operate as the column demultiplexer 5614 and the row demultiplexer 5612 in FIG. 56. Collect the signal line YL at the output. 6708 and signal line YL| 6710 ± collect the output value from each row of memory cells. The output of a specific memory cell in the row is read from the common output collection signal line for the row. Two-dimensional access decoder block ("2DDecode (2D decoding)") 2X2 blocks of 6720-6723 are driven by row decoder 6706 and column decoder 6704 to output appropriate Wr signals and Rd signals to a 4X4 array. Note that only a single row of the multiple access memory shown in FIG. 67 is selected at any time for writing by the row decoder 6702 and the demultiplexers 6730 and 6731 based on the shift operation. Therefore, as discussed with reference to FIG. 56, the row demultiplexer including the row decoder 6702 and the demultiplexers 6730 and 6731 appropriately performs the row write operation based on the shift operation in the following manner, namely, X and y coordinates supplied by the camera block of the memory controller Perform decoding and alternately pull up and down the φΐ 6736 and φ2 6737 input signals in each interval of two consecutive Pixclk tick signals. Figure 67 is intended to illustrate that a large number of memory cells can be interconnected to form the arbitrarily large memory surface illustrated in Figure 57 to implement a multiple access memory that allows multiple cameras to perform simultaneously Based on the write access of the shift register and the video codec used in the video processing system for random access.
[0211] FIG. 68 illustrates a schematic diagram of a two-dimensional access decoder block shown in the memory storage array illustrated in FIG. 67 as an embodiment of the present invention. Each two-dimensional access decoder block outputs the Wr signal and the Rd signal to the corresponding 4X4 array. When the column and row output by the column decoder (6704 in Figure 67) and the row decoder (6706 in Figure 67) match the index of the 2DDecode, when the input Read is high, the 2DDecode outputs the Rd signal, And when the input Write (write) is high, the Wr signal is output.
[0212] FIG. 69 illustrates a memory controller interfaced with the multiple access memory discussed above with reference to FIGS. 65-68 that characterizes an embodiment of the present invention. Like the memory controller illustrated in FIGS. 48A-H, the memory controller 6902 interfaced with the multiple access memory 6904 includes a camera block 6906-6908 associated with each camera and a video codec associated with it. The memory sequencer block 6910. Like the previously described memory controller, each camera block in the memory sequencer block outputs data, two-dimensional addresses, and clock signals. However, the memory controller shown in FIG. 69 does not use an arbiter or a multiplexer, but directly passes the output signal to the decoder 6912-6915 of the multiple access memory. The camera data can therefore be sent directly to the memory without arbitration or multiplexing, and data can be exchanged with the memory cache through the memory sequencer without arbitration or multiplexing. The video codec can monitor the current frame pointer used by the camera block to ensure that the video codec does not initiate a connection with the camera.
CN 102369552 Β
Memory access operations that conflict with memory access operations, or, in an alternative embodiment of the present invention, a simple conflict detection circuit can be included in the memory controller to ensure memory access operations initiated by the memory sequencer Does not conflict with the write access operation of the camera.
[0213] Although the invention has been described in terms of specific embodiments, the intention is not to limit the invention to these embodiments. For those skilled in the art, modifications within the spirit of the present invention will be obvious. For example, any one of various integrated circuit design specification languages (including VHDL and Verilog) can be used to encode the memory controller and the multiple access memory that characterize the embodiment of the present invention. The meaning of the binary signal can be arbitrarily assigned, and different embodiments of the present invention can utilize different signal coding conventions. A memory subsystem of any size can be made according to various embodiments of the present invention. The memory subsystem is discussed in the context of video system applications, but the memory subsystem can be used in other applications characterized by simultaneous data stream writing and two-dimensional random access to the memory. Any of many different integrated circuit fabrication techniques can be used to implement various alternative embodiments of the present invention. Many alternative circuit and subsystem component designs can be designed to implement memory subsystems that characterize embodiments of the present invention.
[0214] For illustrative purposes, the foregoing description uses specific terminology to provide a thorough understanding of the present invention. However, those skilled in the art will understand that specific details are not required to implement the present invention. For the purposes of illustration and description, the foregoing description of specific embodiments of the present invention is presented. They are not intended to be exhaustive or to limit the invention to the precise form disclosed. In view of the above teaching, many modifications and changes are possible. In order to best explain the principles of the present invention and its practical application, embodiments are shown and described, so that those skilled in the art can make best use of the present invention and various modifications suitable for the specific use under consideration. Kind of implementation. It is intended that the scope of the present invention is defined by the appended claims and their equivalents.
CN 102369552 Β
99 sheets
Sheet 1 Sheet 2 Sheet 3 Sheet 4 Sheet 5 Sheet 6 Sheet 7 Sheet 8 Sheet 9 Sheet 10 Sheet 11 Sheet 12 Sheet 13 Sheet 14 Sheet 15 Sheet 16 Sheet 17 Sheet 18 Sheet 19 Sheet 20 Sheet 21 Sheet 22 Sheet 23 Sheet 24 Sheet 25 Sheet 26 Sheet 27 Sheet 28 Sheet 29 Sheet 30 Sheet 31 Sheet 32 Sheet 33 Sheet 34 Sheet 35 Sheet 36 Sheet 37 Sheet 38 Sheet 39 Sheet 40 Sheet 41 Sheet 42 Sheet 43 Sheet 44 Sheet 45 Sheet 46 Sheet 47 Sheet 48 Sheet 49 Sheet 50 Sheet 51 Sheet 52 Sheet 53 Sheet 54 Sheet 55 Sheet 56 Sheet 57 Sheet 58 Sheet 59 Sheet 60 Sheet 61 Sheet 62 Sheet 63 Sheet 64 Sheet 65 Sheet 66 Sheet 67 Sheet 68 Sheet 69 Sheet 70 Sheet 71 Sheet 72 Sheet 73 Sheet 74 Sheet 75 Sheet 76 Sheet 77 Sheet 78 Sheet 79 Sheet 80 Sheet 81 Sheet 82 Sheet 83 Sheet 84 Sheet 85 Sheet 86 Sheet 87 Sheet 88 Sheet 89 Sheet 90 Sheet 91 Sheet 92 Sheet 93 Sheet 94 Sheet 95 Sheet 96 Sheet 97 Sheet 98 Sheet 99
Every citation, both ways
| Document | Relation | Office | Category | Cited during | Relevant claims |
|---|---|---|---|---|---|
| CN1942870A | Cites | China | X | Search report | 1-4、16 |
| US5418907A | Cites | United States of America | X | Search report | 17-27 |
| US6993637B1 | Cites | United States of America | Y | Search report | 5-15 |
| CN1589030A | Cites | China | A | Search report | 1-27 |
| US6401176B1 | Cites | United States of America | A | Search report | 1-27 |
| US2008158601A1 | Cites | United States of America | A | Search report | 1-27 |
22 members in 4 offices
Priority claims19
| Document | Office | Kind | Date |
|---|---|---|---|
| 12319750 | United States of America | – | |
| 31975009 | United States of America | A | |
| 31975009 | United States of America | A | |
| 12322571 | United States of America | – | |
| 32257109 | United States of America | A | |
| 32257109 | United States of America | A | |
| 12380410 | United States of America | – | |
| 38041009 | United States of America | A | |
| 38041009 | United States of America | A | |
| 2009068999 | United States of America | W | |
| 2009068999 | United States of America | W | |
| 12319750 | – | – | – |
| 12322571 | – | – | – |
| 12380410 | – | – | – |
| PCTUS2009068999 | – | – | – |
| US20090319750 | – | – | – |
| US20090322571 | – | – | – |
| US20090380410 | – | – | – |
| WO2009US68999 | – | – | – |
Members22
| Document | Office | Kind | |
|---|---|---|---|
| US2010177585A1 | United States of America | A1 | |
| US2010177828A1 | United States of America | A1 | |
| WO2010080644A2 | World Intellectual Property Organization (WIPO) | A2 | |
| WO2010080645A2 | World Intellectual Property Organization (WIPO) | A2 | |
| WO2010080646A2 | World Intellectual Property Organization (WIPO) | A2 | |
| US2010220215A1 | United States of America | A1 | |
| WO2010080646A3 | World Intellectual Property Organization (WIPO) | A3 | |
| WO2010080644A3 | World Intellectual Property Organization (WIPO) | A3 | |
| WO2010080645A3 | World Intellectual Property Organization (WIPO) | A3 | |
| DE112009004320T5 | Germany | T5 | |
| CN102356635A | China | A | |
| CN102369522A | China | A | |
| CN102369552A | China | A | |
| DE112009004344T5 | Germany | T5 | |
| DE112009004344T8 | Germany | T8 | |
| DE112009004408T5 | Germany | T5 | |
| US8566515B2 | United States of America | B2 | |
| US8660193B2 | United States of America | B2 | |
| CN102369552BThis record | China | B | |
| US2015012708A1 | United States of America | A1 | |
| US2015288974A1 | United States of America | A1 | |
| CN102369522B | China | B |
8 legal events, as the office reported them to INPADOC
Over the term
Point at a mark for the eventEvents
| Event | Code | |
|---|---|---|
| Termination of patent right due to non-payment of annual feeCF01 | CF01 | |
| Grant of patent or utility modelGrantedC14 | C14 | |
| Succession or assignment of patent rightASS | ASS | |
| Transfer of patent application or patent right or utility modelC41 | C41 | |
| Correction of patent for invention or patent applicationC53 | C53 | |
| Change of bibliographic dataCORRECT: APPLICANT; FROM: MAXIM INTEGRATED PRODUCTS, INC. TO: MAXIM INTEGRATED PRODUCTS INC.COR | COR | |
| Entry into substantive examinationC10 | C10 | |
| PublicationC06 | C06 |
Numbers
- Publication
- 102369552
- Publication, DOCDB
- 102369552
- Publication, EPODOC
- CN102369552B
- Application
- 801580069
- Application, DOCDB
- 200980158006
- Application, EPODOC
- CN20098158006
Titles2
- Chinese
- 存储器子系统
- English
- Memory subsystem
Classification
- CPC, 11
- G06F13/1663
- G06F12/0875
- G11C19/00
- H04N19/61
- H04N19/11
- H04N19/91
- H04N19/186
- H04N19/82
- H04N19/423
- H04N19/43
- H04N19/436
- IPC, 1
- G06F12 02