Improvements in memory management for systems for generating 3-dimensional computer images
9 claims: 2 independent, 7 dependent
- 1三次元コンピュータ映像を発生するシステムに使用するためのメモリマネージメントシステムにおいて、 a)映像を複数の長方形エリアに細分化する手段と、 b)各長方形エリアに対するオブジェクトデータを記憶する1つ又は複数の第1部分、及び前記オブジェクトデータから導出される深さデータを記憶する1つ又は複数の第2部分をいつでも有しているメモリと、 c)前記メモリの1つ又は複数の第1部分にオブジェクトデータを記憶する手段と、 d)前記オブジェクトデータから各長方形エリアに対する深さデータを導出するための手段と、 e)各長方形エリアに対する深さデータを前記メモリの1つ又は複数の第2部分に記憶する手段と、 f)既存のコンテンツの少なくとも一部分に置き換えるように前記メモリの1つ又は複数の第1部分のうちの1つ以上に更なるオブジェクトデータをロードする手段と、 g)前記記憶された深さデータを検索する手段と、 h)新たなオブジェクトデータ及び記憶された深さデータから各長方形エリアの各画素に対する更新された深さデータを導出し、そしてその更新された深さデータを、以前に記憶された深さデータに置き換えるように記憶する手段と、 i)前記メモリにロードすべき更なるオブジェクトデータがなくなるまで前記特徴e)、f)、g)及びh)で機能を繰り返し遂行させる手段と、 j)表示のために、映像データ及び陰影付けデータを前記深さデータから導出するための手段と、 を備え、 メモリの1つ又は複数の第1部分にオブジェクトデータを記憶する前記手段は、1つの長方形エリアのみに入るオブジェクトデータを、その長方形エリアに割り当てられたメモリのブロックに記憶し、且つ2つ以上の長方形エリアに入るオブジェクトデータを、2つ以上の長方形エリアに入るオブジェクトデータを記憶するためのグローバルリストとして割り当てられたメモリのブロックに記憶するように構成される、メモリマネージメントシステム。
- 2各長方形エリアに割り当てられるメモリの少なくとも一部分、及びグローバルリストとして割り当てられるメモリの少なくとも一部分が、要件に基づいてメモリの未使用部分から割り当てられて、各長方形エリアに割り当てられるメモリの少なくとも一部分のサイズ及び位置と、グローバルリストとして割り当てられるメモリの少なくとも一部分のサイズ及び位置とが動的に変化するようにする、請求項1に記載のメモリマネージメントシステム。
- 3前記グローバルリストは、2つ以上の長方形エリアに入るオブジェクトに関するオブジェクトデータと、単一の長方形エリアに入るが別の長方形エリアとの境界に接近したオブジェクトに関するオブジェクトデータも、記憶するように構成される、請求項1又は2のいずれかに記載のメモリマネージメントシステム。
- 4長方形エリア間の境界の各側で、単一の長方形エリアに入るオブジェクトを含む連続する基本的エリアの数は、境界からの距離が増加するにつれて増加し、そして単一の長方形に入るオブジェクトを含む連続する基本的エリアの数が所定のスレッシュホールドに交差するところの基本的エリアと境界との間に入るオブジェクトのオブジェクトデータは、グローバルリストに記憶される、請求項3に記載のメモリマネージメントシステム。
- 5三次元コンピュータ映像を発生するシステムに使用するためのメモリをマネージする方法において、 コンピュータが、 a)映像を複数の長方形エリアに細分化するステップを実行し、 b)メモリが、各長方形エリアに対するオブジェクトデータを記憶する1つ又は複数の第1部分、及び前記オブジェクトデータから導出される深さデータを記憶する1つ又は複数の第2部分をいつでも有しており、 前記コンピュータが、更に、 c)前記メモリの1つ又は複数の第1部分にオブジェクトデータを記憶するステップと、 d)前記オブジェクトデータから各長方形エリアに対する深さデータを導出するステップと、 e)各長方形エリアに対する深さデータを前記メモリの1つ又は複数の第2部分に記憶するステップと、 f)既存のコンテンツの少なくとも一部分に置き換えるように前記メモリの1つ又は複数の第1部分のうちの1つ以上に更なるオブジェクトデータをロードするステップと、 g)前記記憶された深さデータを検索するステップと、 h)新たなオブジェクトデータ及び記憶された深さデータから各長方形エリアの各画素に対する更新された深さデータを導出し、そしてその更新された深さデータを、以前に記憶された深さデータに置き換えるように記憶するステップと、 i)前記メモリにロードすべき更なるオブジェクトデータがなくなるまで前記特徴e)、f)、g)及びh)で機能を繰り返し遂行させるステップと、 j)表示のために、映像データ及び陰影付けデータを前記深さデータから導出するステップと、を実行し、 メモリの1つ又は複数の第1部分にオブジェクトデータを記憶する前記ステップは、1つの長方形エリアのみに入るオブジェクトデータを、その長方形エリアに割り当てられたメモリのブロックに記憶し、且つ2つ以上の長方形エリアに入るオブジェクトデータを、2つ以上の長方形エリアに入るオブジェクトデータを記憶するためのグローバルリストとして割り当てられたメモリのブロックに記憶するように構成される、方法。
- 6各長方形エリアに割り当てられたメモリの少なくとも一部分、及びグローバルリストとして割り当てられたメモリの少なくとも一部分は、要件に基づいてメモリの未使用部分から割り当てられて、各長方形エリアに割り当てられるメモリの少なくとも一部分のサイズ及び位置と、グローバルリストとして割り当てられるメモリの少なくとも一部分のサイズ及び位置とが動的に変化するようにする、請求項5に記載の方法。
- 7前記映像データ及び陰影付けデータが特定の長方形エリアに対して導出されると、その長方形エリアに割り当てられたメモリの少なくとも一部分が空きとマークされ、そして前記映像データ及び陰影付けデータが全ての長方形エリアに対して導出されると、グローバルリストとして割り当てられたメモリの少なくとも一部分も、空きとしてマークされる、請求項5又は6のいずれかに記載の方法。
- 8前記グローバルリストは、2つ以上の長方形エリアに入るオブジェクトに関するオブジェクトデータと、単一の長方形エリアに入るが別の長方形エリアとの境界に接近したオブジェクトに関するオブジェクトデータとを記憶するように構成される、請求項5から7のいずれかに記載の方法。
- 9長方形エリア間の境界の各側で、単一の長方形エリアに入るオブジェクトを含む連続する基本的エリアの数は、境界からの距離が増加するにつれて増加し、そして単一の長方形に入るオブジェクトを含む連続する基本的エリアの数が所定のスレッシュホールドに交差するところの基本的エリアと境界との間に入るオブジェクトのオブジェクトデータは、グローバルリストに記憶される、請求項8に記載の方法。
Independent claims9
77 paragraphs, as filed
The present invention relates to memory management used in a system for generating a three-dimensional computer-generated image.
Applicant's UK Patent No. 2281682 describes a 3D rendering system for polygons in which each object appears to be defined as a myriad of sets of surfaces. Each basic area of the screen (eg, pixels) on which the image should be displayed is projected from the viewpoint through which rays are projected onto the 3D screen. Then, the position where the projected light beam intersects each surface is determined. From these intersections it is possible to determine if the intersection surface is visible in the basic area. The basic area is then shaded for display based on the result of that decision.
The system can be embodied in a pipeline-type processor with a large number of cells, each capable of performing surface intersection calculations. Therefore, a large number of surface intersections can be calculated at the same time. Each cell is loaded with a set of coefficients that define the surface on which the cross test should be performed.
Improvements to this configuration are described in Applicant's UK Pat. No. 2,298,111. In this document, the video is divided into sub-regions or tiles, and those tiles can be processed sequentially. It has been proposed to use a variable tile size and project a bounding box around a complete object so that only tiles that fall within that bounding box need to be processed. This is done by determining the distribution of objects on the visible screen in order to select the appropriate tile size. The surfaces that define the various objects are then stored in a list known as a display list, thereby avoiding the need to store the same surface for each tile. This is because one object made up of many surfaces appears in many tiles. An object pointer that identifies an object in the display list is also stored. There is one object pointer list for each tile. The tiles can be rendered sequentially using the ray projection technique until all the objects in each tile have been processed. This is a useful method as it does not require any effort to render an object that is known to be invisible on a particular tile.
Further improvements to this system are proposed in Applicant's International Patent Application No. PCT / GP99 / 03707, where tiles in a bounding box that are not needed to display a particular object are rendered. Discarded before.
FIG. 1 shows a processor 101 of the format used in the existing systems described above. In essence, there are three components. The tile acceleration unit (TA) 103 performs a tile processing operation, i.e. selects an appropriate tile size, divides the visible screen into tiles, and displays tile information, i.e., 3D object data for each tile, into display list memory 105. Supply. The video compositing processor (ISP) 107 uses the 3D object data in the display list memory to perform the ray / surface crossover test described above. This produces depth data for each basic area of the visible screen. The video data derived from the ISP 107 is then fed to a texture and shading processor (TSP) 109, which applies the texture and shading data to the surface determined to be visible, and the video and shading. The attached data is output to the frame buffer memory 111. Therefore, the appearance of each basic area of the display is determined to represent a 3D image.
In the systems described above, problems can arise as the complexity of the scene to be rendered increases, complex scenes require more 3D object data for each tile to be stored in display list memory. This means higher storage requirements. When the display list memory runs out of space, parts of the scene simply cannot be rendered, and this form of video collapse becomes increasingly unacceptable.
To solve this problem, Applicant's International Patent Application No. PCT / GB01 / 02536 proposes the idea of partial rendering. The state of the system (ISP and TSP) is stored in memory before the tile rendering is complete, and this state is later reloaded to finish the rendering. This process is referred to as "z / frame buffer load and storage".
The screen is divided into a number of areas called macro tiles, each macro tile consisting of a rectangular area of the screen. The display list memory is then divided into blocks, which are listed in the free storage list. Blocks from the free storage are then assigned to macro tiles as needed. The tile processing operation stores data associated with each macro tile in each block. (The tile processing operation performed by the TA fills the display list memory and is therefore sometimes referred to as memory allocation.) When the display list memory is full or reaches a certain threshold, the system After selecting a macro tile, performing a z / framebuffer load, and rendering the contents of the macrotile, use the z / framebuffer storage operation to save it. Therefore, the depth data for the macro tile is stored based on the data previously loaded in the display list. When such rendering is complete, the system frees the memory blocks associated with the macro tile, thereby making those blocks available for further memory. (The rendering process is known as memory deallocation because it frees up space in the display list memory.) Therefore, the scene for each tile consists of a number of tiles followed by partial rendering. Each partial rendering updates the stored depth data. This means that an upper bound on memory consumption is imposed and the memory bandwidth consumed by the system is minimized.
An embodiment of a processor of the format used in a partial rendering system is shown in FIG. It will be clear that this is a modification of FIG. The z-buffer memory 209 is linked to the ISP 207 via the z-compression / decompression unit 211. This goes into operation when the system renders a complex scene and the display list memory 205 is not large enough to include all the surfaces that need to be processed for a particular tile. The display list is loaded with data by TA203 for all tiles until it is substantially full (or until the specified threshold is reached). However, this represents only a portion of the initial data. This video is rendered one tile at a time by ISP207. The output data for each tile is given to TSP213, which uses the texture data to texture the tile. At the same time, since the video data is incomplete, the result from the ISP 207 (ie, the depth data) is stored in the buffer memory 209 via the compression / decompression unit 211 for temporary storage. Rendering of the remaining tiles is then continued with incomplete video data until all tiles have been rendered and stored in framebuffer memory 215 and z-buffer memory 209.
The first portion of the display list is then discarded and additional video data is loaded into it. Since the processing is sequentially performed for each tile by the ISP 207, the related portion of the data from the z-buffer memory 209 is loaded via the z-compression / decompression unit 18 and combined with the new video data from the display list memory 205. be able to. Next, new depth data for each tile is supplied to the TSP 213, which is combined with the texture data and then supplied to the frame buffer 215.
This process continues until all video data has been rendered for all tiles in the scene. Therefore, the z-buffer memory fills the temporary storage, which allows smaller display list memory than is needed to render particularly complex scenes. The compression / decompression unit 211 optionally allows the use of a small z-buffer memory.
Therefore, as described in International Patent Application No. PCT / GB01 / 02536, when the display list memory is full or reaches a threshold, the system selects the macro tile to render and displays the display. Free the list memory. In the above application, the selection of the macro tile to be rendered depends on a number of factors, and for example, the macro tile that releases most of the memory to the free storage unit can be selected.
<p num="0014"> The inventor of the present invention has found that various improvements can be made to memory management in this system.</p><p num="0015"> An object of the present invention is to provide a memory management system and method that reduces the occupied area of memory and improves performance when compared to the known systems described above. Yet another object of the present invention is to provide a memory management system and method capable of handling a large number of applications running simultaneously.</p>
<p num="0016"> According to the first aspect of the present invention, in a memory management system for use in a system that generates a three-dimensional computer image, a) means for subdividing the image into a plurality of rectangular areas, and b) for each rectangular area. A memory that always has one or more first parts that store object data, and one or more second parts that store depth data derived from the object data, and c) one of the memories. Alternatively, a means for storing the object data in a plurality of first parts, d) a means for deriving the depth data for each rectangular area from the object data, and e) one or a plurality of memory for the depth data for each rectangular area. Means for storing in the second part and f) means for loading additional object data into one or more of one or more first parts of memory to replace at least a portion of existing content, g) Means for retrieving the stored depth data and h) Derivation of updated depth data for each pixel in each rectangular area from new object data and stored depth data, and the updated depth. Means to store the data to replace previously stored depth data and i) function with the features e), f), g) and h) until there is no more object data to load into the memory. A system is provided that includes means for repetitive execution and j) means for deriving video data and shading data from depth data for display.</p><p num="0017"> Therefore, the object data and the depth data are stored in a single memory. Therefore, the memory requirement can be relaxed. Note that the one or more parts allocated to the object data are not fixed as the amount of memory required for the object data changes as the features perform their function. If two or more parts of memory are allocated to object data at a particular time, those parts may be adjacent in memory or interspersed with other parts of memory allocated for different purposes. You may. Similarly, one or more portions assigned to the depth data are not fixed as the amount of memory required for the depth data changes as the features perform their function. If at a particular time two or more parts of memory are allocated to depth data, those parts may be adjacent in memory or with other parts of memory allocated for different purposes. It may be scattered.</p><p num="0018"> The means for subdividing the image into a plurality of rectangular areas can be configured to select the size of the rectangular area based on the particular image to be generated. The rectangular areas may be of equal size and shape or may be different.</p><p num="0019"> Preferably, one or more first parts of memory and one or more second parts of memory are allocated from unused parts of memory according to requirements, said features c), d), e). , F), g), h) and i) perform their functions, the size of one or more first and second parts in memory and one or more first and second parts. Make the position of the part change dynamically.</p><p num="0020"> Preferably, the features d) and e) are always present when one or more second portions are always reserved for the depth data of at least one rectangular area and there is object data stored in memory. To be able to perform those functions.</p><p num="0021"> In another embodiment, the feature e) comprises means for compressing the depth data before it is stored, and one or more for the depth data of at least two rectangular areas. When the second part is always reserved and there is object data stored in the memory, the features d) and e) can always perform their functions.</p><p num="0022"> By reserving an appropriate division of memory in this way, it is always possible to derive further depth data and store it in one or more second parts of memory. However, this means that even if the image is complicated, it can always occur. In one embodiment, the reserved quantity is sufficient for the depth data of each pixel area of one macro tile. In another embodiment, the reserved quantity is sufficient for the depth data of each pixel area of the two macro tiles.</p><p num="0023"> In one embodiment, a means of storing object data in one or more first portions of memory stores object data that fits in only one rectangular area in a block of memory allocated to that rectangular area, and Object data that fits in two or more rectangular areas is configured to be stored in a block of memory allocated as a global list for storing object data that fits in two or more rectangular areas. In this case, the video data and shading data for a particular rectangular area are also derived from the depth data for that macro tile and also from the depth data in the global list. Once the video data and shading data are derived for a particular rectangular area, that block of memory allocated for the depth data of that rectangular area can be marked as free. Once the video data and shading data have been derived for all rectangular areas, the global list can also be marked as empty.</p><p num="0024"> Further, according to the first aspect of the present invention, in the method of managing memory in a system that generates a three-dimensional computer image, a) a step of subdividing the image into a plurality of rectangular areas and b) each rectangular area. A step of preparing a memory having one or more first parts for storing object data for, and one or more second parts for storing depth data derived from the object data, and c. ) A step of storing the object data in one or more first parts of the memory, d) a step of deriving the depth data for each rectangular area from the object data, and e) storing the depth data for each rectangular area in the memory. Steps to store in one or more second parts and f) load additional object data into one or more of one or more first parts of memory to replace at least a portion of existing content. Steps, g) searching for stored depth data, h) deriving updated depth data for each pixel in each rectangular area from new object data and stored depth data, and its The steps of storing the updated depth data to replace the previously stored depth data and i) the steps e), f), g) and i) until there is no more object data to load into the memory. Also provided is a method comprising a step of repeating h) and a step of deriving video data and shading data from depth data for j) display.</p><p num="0025"> The aspects described in relation to the method of the first aspect of the present invention can also be applied to the system of the first aspect of the present invention, and the aspects described in relation to the system of the first aspect of the present invention. Can also be applied to the method of the first aspect of the present invention.</p><p num="0026"> According to the second aspect of the present invention, in a memory management system for use in a system that generates a three-dimensional computer image, a) means for subdividing the image into a plurality of rectangular areas, and b) each rectangular area. A memory that stores object data related to the objects in the video to be entered, i) at least one part allocated to each rectangular area to store object data related to the objects in each rectangular area, and ii) two or more rectangular areas. A memory containing at least one part allocated as a global list for storing object data about the objects to be entered, c) means for storing the object data in the memory, and d) video data and shading data for each rectangular area. Derivation means for deriving from object data and e) object data for each rectangular area from each part of the memory and also from the global list if the rectangular area includes objects that fall into at least one other rectangular area. , A system including means for supplying to the derivation means and f) means for storing the video data and shading data derived by the derivation means for display.</p><p num="0027"> Allocating a portion of memory as a global list means that you only need to write the object data of objects that fit into two or more rectangular areas into memory once. This reduces the amount of memory required and also reduces the time required to store such object data.</p><p num="0028"> Preferably, at least a portion of the memory allocated to each rectangular area and at least a portion of the memory allocated as a global list are allocated from the unused portion of the memory according to the requirements, said features c), d) and e). Allows the size and location of at least a portion of the memory allocated to each rectangular area to dynamically change and the size and location of at least a portion of the memory allocated as a global list as it performs those functions.</p><p num="0029"> Preferably, the derivation means includes means for deriving depth data for each rectangular area from object data, and shading means for deriving video data and shading data from the depth data.</p><p num="0030"> Fortunately, the global list also stores object data for objects that fall into two or more rectangular areas and objects that fall into a single rectangular area but are close to the boundaries of another rectangular area. It is composed of. This improves processing for basic areas that fall near the boundaries between macro tiles.</p><p num="0031"> However, in this case, it is necessary to determine which object data is stored in the global list and which is stored in the portion assigned to the appropriate rectangular area. In one embodiment, on each side of the boundary between rectangular areas, the number of contiguous basic areas containing objects that go into a single rectangular area increases as the distance from the boundary increases, and a single rectangle. The object data of the objects that fall between the boundary and the basic area where the number of consecutive basic areas containing the objects that enter intersect a predetermined threshold is stored in the global list. The threshold can be determined by a number of factors, including the number of objects in the scene.</p><p num="0032"> According to the second aspect of the present invention, in a method of managing memory in a system that generates a three-dimensional computer image, a) a step of subdividing the image into a plurality of rectangular areas and b) entering each rectangular area. It is a step of storing object data related to video objects in memory, i) allocating at least a part of memory to each rectangular area, storing object data related to objects in each rectangular area in that part, and ii) storing object data in memory. A step performed by allocating at least a portion as a global list and storing object data for objects that fall into two or more rectangular areas in the global list, and c) assigning object data for each rectangular area from each part of memory. And if the rectangular area also includes objects that fall into at least one other rectangular area, it is also a step of supplying to the derivation means from the global list, the derivation means providing video data and shading data for each rectangular area. A method is provided that includes a step of deriving, d) a step of deriving video data and shading data in the deriving means, and e) a step of storing the video data and shading data for display. To.</p><p num="0033"> There are two effects of allocating a part of memory as a global list. First, the object data is stored only once, which reduces the amount of memory required. Second, since the object data is only stored once, the object data for objects that fall into two or more rectangular areas need only be written once in memory, i.e. not each macro tile portion of memory, but global. All you have to do is write it to the list.</p><p num="0034"> Preferably, at least a portion of the memory allocated to each rectangular area, and at least a portion of the memory allocated as a global list, is allocated from the unused portion of memory according to requirements, as described in steps c), d) and As e) is performed, the size and location of at least a portion of the memory allocated to each rectangular area and the size and location of at least a portion of the memory allocated as a global list are dynamically changed.</p><p num="0035"> Preferably, in the step d) of deriving the video data and the shading data in the deriving means, the depth data for each rectangular area is derived from the object data, and the video data and the shading data are derived from the depth data. including.</p><p num="0036"> Preferably, when the video data and shading data are derived for a particular rectangular area, at least a portion of the memory allocated to that rectangular area is marked as free, and the video data and shading data are all rectangular. When derived for an area, at least a portion of the memory allocated as a global list is also marked as free. Since the global list is not marked as empty until the video data and shading data are derived for all rectangular areas, the large global list that reduces the amount of data repetition until it can be marked as empty, and takes a long time, A compromise must be made with a small global list that can be marked as free relatively quickly but does not significantly reduce the amount of data iterations.</p><p num="0037"> Preferably, the global list is configured to store object data for objects that fall into two or more rectangular areas and object data that falls within a single rectangular area but is close to the boundary with another rectangular area. Will be done.</p><p num="0038"> In this case, preferably, on each side of the boundary between the rectangular areas, the number of contiguous basic areas containing objects that go into a single rectangular area increases as the distance from the boundary increases, and is single. The object data of the objects that fall between the basic areas and the boundaries where the number of consecutive basic areas, including the objects that fit in the rectangle, intersect a given threshold, is stored in the global list.</p><p num="0039"> The aspects described in relation to the method of the second aspect of the present invention can also be applied to the system of the second aspect of the present invention, and the aspects described in relation to the system of the second aspect of the present invention. Can also be applied to the method of the second aspect of the present invention.</p><p num="0040"> According to a third aspect of the present invention, a system that generates a three-dimensional computer image is such that the generated image is divided into a plurality of rectangular areas and each rectangular area is divided into a plurality of smaller areas. In the memory for use, the image generated is provided with a part for storing the object data for the object entering each rectangular area and a part for the pointer from each small area to the object data for the object entering the small area. The object in is divided into triangles, the object data contains triangle data, vertex data, and pointers between those triangle data and vertex data, and the memory is the image that the system processes a small area and the image in that small area. When generating a portion of, the system uses a pointer from each small area to the triangle data for the object that enters that small area, and a pointer between the triangle data and the vertex data, thereby that small area. A memory is provided that is configured to be able to fetch the vertex data for a single fetch.</p><p num="0041"> According to a fourth aspect of the present invention, in a method of generating a three-dimensional computer image, a) a step of subdividing the image into a plurality of rectangular areas MTn, and b) one or a plurality of items for storing object data. The step of preparing a memory having one part and one or more second parts for storing the depth data derived from the object data at any time, and c) the object data for each rectangular area MTn is stored in 1 of the memory. A step of loading into one or more first parts, where each rectangular area MTn is one or more of the memory until the total size of one or more first parts of the memory used exceeds a predetermined threshold. Steps such as using each block Pn of a plurality of first parts, and d) select one or more blocks Pn of the first part of the memory, and obtain depth data for each rectangular area MTn from the object data of MTn. A step of deriving and storing the derived depth data for MTn in one or more second parts of memory, wherein the block Pn is i) further after the step d) has been performed. Object data can be loaded into blocks Pn and replaced with all or part of existing content, or ii) additional blocks Pm of one or more first parts of memory are selected and each rectangular area MTm And e) a step selected so that the depth data for MTm can be derived from the object data of MTm and the derived depth data for MTm can be stored in one or more second parts of the memory. If i) of step d) is satisfied, the steps c) and d) are repeated until there is no more object data loaded in the memory, or ii) of step d) is satisfied. In this case, the step d) of repeating the step d) is repeated until there is no more object data loaded in the memory, and f) the video data and the shading data for each rectangular area MTn are stored in one or more second memory. For each rectangular area MTn stored in the partA method is provided that includes a step of deriving from the depth data and a step of e) storing the video data and the shading data for display.</p><p num="0042"> When the process is to be serialized, the selection of the rectangular area in step d) is very high because the amount of one or more first parts of memory used exceeds a predetermined threshold after step c). It becomes important. This is because after the rectangular area is selected, more object data must be able to be loaded or more depth data must be derived and stored.</p><p num="0043"> Note that as the steps of this method are performed, the amount of memory required for the object data changes, so one or more parts allocated to the object data are not fixed. If two or more parts of memory are allocated to object data at a particular time, those parts may be adjacent in memory or interspersed with other parts of memory allocated for different purposes. You may. Similarly, one or more portions assigned to the depth data are not fixed as the amount of memory required for the depth data changes as the steps of this method are performed. If at a particular time two or more parts of memory are allocated to depth data, those parts may be adjacent in memory or with other parts of memory allocated for different purposes. It may be scattered.</p><p num="0044"> The predetermined threshold in step c) may be 100% full, or another threshold, eg, 75% of full. Step d) can be started when approaching the threshold to maintain parallel processing as much as possible.</p><p num="0045"> Preferably, one or more first parts of memory and one or more second parts of memory are allocated from unused parts of memory according to requirements when the steps of the method are performed. In addition, the size of one or more first and second parts in memory and the position of one or more first and second parts are made to change dynamically.</p><p num="0046"> Preferably, step d) can always be performed when one or more second parts are always reserved for depth data and there is object data stored in one or more first parts of memory. To do so.</p><p num="0047"> In this case, one or more reserved second portions of memory are at least sufficient to store depth data for each basic area of one rectangular area. Alternatively, one or more reserved second portions of memory are at least sufficient to store depth data for each basic area of the two rectangular areas.</p><p num="0048"> The reserved part is, at a minimum, able to continue processing. Preferably, more of the one or more second parts are available for the new depth data. This improves processing.</p><p num="0049"> In one embodiment, step c) includes loading object data for each rectangular area MTn into one or more first portions of memory, and object data for an object that fits in only one rectangular area is memory. Object data for an object stored in each block Pn of one or more first parts of the memory and falling into two or more rectangular areas is stored in the block PGI of one or more first parts of memory.</p><p num="0050"> In one embodiment, the memory allocated to each block Pn may be equal at a particular point. In this case, after step c), if the memory block Pn is the same size for each rectangular area MTn, before additional object data can be loaded into one or more first parts of memory. Step d) is performed for each rectangular area.</p><p num="0051"> However, memory is usually not evenly allocated to each block Pn. In this case, after step c), if the memory block Pn is not the same size for each rectangular area MTn, the block Pn of one or more first parts of the memory selected in step d) It becomes the maximum block Pn.</p><p num="0052"> According to a fifth aspect of the present invention, in a memory management system for use in a system that generates three-dimensional computer images in a plurality of applications executed at the same time, a) the images of each application are divided into a plurality of rectangular areas. Means for subdivision, b) at least one memory for storing the object data in each rectangular area and depth data derived from the object data for each application, and c) at least one object data for each application. Means for storing in one memory, d) means for deriving depth data of each rectangular area from object data for each application, and e) for each application, at least one memory for depth data of each rectangular area. Means to store in, f) means to load additional object data for each application into at least one memory and replace it with existing content for each application, and g) store depth data for each application. The means of searching and h) the updated depth data for each pixel in each rectangular area for each application is derived from the new object data and stored depth data for each application and updated for each application. Means for storing the depth data and replacing it with the previously stored depth data for each application, and i) the feature e until there is no more object data for the application to be loaded into at least one memory. ), F), g) and h) are means to repeatedly perform the function, and j) means to derive video data and shading data from depth data for each application for display. , K) A memory management system comprising means for storing and updating the progress made by the features c), d), e), f), g), h) and i) for each application. Is provided.</p><p num="0053"> Therefore, one system can run two or more applications at the same time. The feature k) allows the system to track the internal state of each application.</p><p num="0054"> In one embodiment, at least one memory includes one memory per application.</p><p num="0055"> In this embodiment, each memory stores one or more first parts for storing object data for each rectangular area of each application and depth data derived from the object data of each application. It includes one or more second parts.</p><p num="0056"> In this embodiment, preferably one or more first parts of memory and one or more second parts of memory are allocated from unused parts of memory based on requirements, said feature c). The size of one or more first and second parts in memory and one or more when d), e), f), g), h) and i) perform their function. The positions of the first and second parts are made to change dynamically.</p><p num="0057"> In this embodiment, preferably, one or more second parts are always reserved for the depth data of each application, and there is object data stored in one or more first parts of memory. In addition, the features d) and e) always allow them to perform their functions.</p><p num="0058"> In another embodiment, at least one memory comprises a single memory, of which portion is allocated to each application as needed. This is a more efficient way to use memory and can reduce the total memory requirement.</p><p num="0059"> In this embodiment, preferably, when a part of a single memory is always reserved for depth data and there is object data stored in the memory, the features d) and e) always perform those functions. It can be so.</p><p num="0060"> In this case, preferably, the reserved portion is sufficient for the depth data for each basic area of one rectangular area to be stored, regardless of the number of applications.</p><p num="0061"> According to a fifth aspect of the present invention, in a method of generating a three-dimensional computer image by a plurality of applications executed at the same time, for each application, a) a step of subdividing the image into a plurality of rectangular areas. , B) A step of preparing at least one memory for storing the object data for each rectangular area and the depth data derived from the object data, c) a step of storing the object data in at least one memory, and d. ) The step of deriving the depth data of each rectangular area from the object data, e) the step of storing the depth data of each rectangular area in at least one memory, and f) at least one additional object data for each application. A new step to load into memory and replace with existing content, g) to search for stored depth data, and h) to update updated depth data for each pixel in each rectangular area. Steps to derive from object data and stored depth data, and to store updated depth data and replace it with previously stored depth data, and i) additional to load into at least one memory. The steps e), f), g) and h) are repeated until the object data is exhausted, and j) the video data and the shading data are derived from the depth data for display. The steps of this method are performed by one system for all plurality of applications, which system time-divides between the plurality of applications and for each application the steps c), d). , E), f), g), h) and i) are also provided for methods of memorizing and updating the progress.</p><p num="0062"> The aspects described in relation to the method of the fifth aspect of the present invention can also be applied to the system of the fifth aspect of the present invention, and the aspects described in relation to the system of the fifth aspect of the present invention. Can also be applied to the method of the fifth aspect of the present invention.</p><p num="0063"> Aspects described in relation to one aspect of the invention can also be applied to another aspect of the invention.</p>
<figref num="1">It is the schematic of the 1st known rendering and texturing system.</figref><figref num="2">FIG. 5 is a schematic representation of a second known rendering and texturing system that provides improvements over the system of FIG.</figref><figref num="3">It is the schematic of the rendering and texture processing system by embodiment of this invention.</figref><figref num="4">It is the schematic of the display list memory by embodiment of this invention.</figref><figref num="5">FIG. 6 is a schematic diagram showing one possible configuration of a dynamic parameter management (DPM) system.</figref><figref num="6a">The rendering process of the first embodiment, which is composed of FIGS. 6a to 6f and the memory is uniformly allocated to the macro tiles, is shown.</figref><figref num="6b">The rendering process of the first embodiment, which is composed of FIGS. 6a to 6f and the memory is uniformly allocated to the macro tiles, is shown.</figref><figref num="6c">The rendering process of the first embodiment, which is composed of FIGS. 6a to 6f and the memory is uniformly allocated to the macro tiles, is shown.</figref><figref num="6d">The rendering process of the first embodiment, which is composed of FIGS. 6a to 6f and the memory is uniformly allocated to the macro tiles, is shown.</figref><figref num="6e">The rendering process of the first embodiment, which is composed of FIGS. 6a to 6f and the memory is uniformly allocated to the macro tiles, is shown.</figref><figref num="6f">The rendering process of the first embodiment, which is composed of FIGS. 6a to 6f and the memory is uniformly allocated to the macro tiles, is shown.</figref><figref num="7a">FIG. 7 shows the rendering process of the second embodiment, which comprises 7a to 7g and the memory is non-uniformly allocated to the macro tiles.</figref><figref num="7b">FIG. 7 shows the rendering process of the second embodiment, which comprises 7a to 7g and the memory is non-uniformly allocated to the macro tiles.</figref><figref num="7c">FIG. 7 shows the rendering process of the second embodiment, which comprises 7a to 7g and the memory is non-uniformly allocated to the macro tiles.</figref><figref num="7d">FIG. 7 shows the rendering process of the second embodiment, which comprises 7a to 7g and the memory is non-uniformly allocated to the macro tiles.</figref><figref num="7e">FIG. 7 shows the rendering process of the second embodiment, which comprises 7a to 7g and the memory is non-uniformly allocated to the macro tiles.</figref><figref num="7f">FIG. 7 shows the rendering process of the second embodiment, which comprises 7a to 7g and the memory is non-uniformly allocated to the macro tiles.</figref><figref num="7g">FIG. 7 shows the rendering process of the second embodiment, which comprises 7a to 7g and the memory is non-uniformly allocated to the macro tiles.</figref><figref num="8a">FIG. 5 is a schematic diagram of a rendering and texturing system configured to execute two simultaneous applications according to an embodiment of the present invention.</figref><figref num="8b">FIG. 5 is a schematic representation of a rendering and texturing system configured to run two simultaneous applications according to another embodiment of the invention.</figref>
The existing system has already been described with reference to FIGS. 1 and 2. Embodiments of the present invention will be described in detail with reference to FIGS. 3 and later.
FIG. 3 is a schematic diagram of a rendering and texture processing system according to an embodiment of the present invention. System 301 is similar to that of FIG. 2 and includes TA303, ISP311 and TSP313, and a frame buffer 315. However, in this case, the display list memory 305 and the Z-buffer memory 307 are both part of a block of memory called the parameter memory 309. The allocation between the display list and the Z-buffer in the parameter memory is described below. FIG. 3 does not show the z-compression / decompression unit as in FIG. 2, but such a unit can be included.
FIG. 4 is a schematic view of a display list memory according to an embodiment of the present invention. Display list 401 includes control streams and object data described in detail below.
As mentioned above, the tiles on the visible screen are grouped together into macro tiles, each macro tile being an area of the screen consisting of a large number of tiles. Macro tile boundaries are defined by configuration registers. In one embodiment, there are four macro tiles. In another embodiment, there are 16 macro tiles. In fact, there may be any number of macro tiles, and the macro tiles do not necessarily have to be the same shape and size (although if they are not the same, the process is potentially more complicated). The size of the macro tiles is designed to provide sufficient particle size for memory recycling, while not incurring excessively high infrastructure costs. Larger macro tiles result in lower memory usage and less on-chip memory, while smaller macro tiles allow memory to be recycled more quickly and increase memory efficiency.
As is apparent in FIG. 4, macro tiles 403a, 403b include both control streams 405a, 405b and object data 407a, 407b to group the tiles. Each object in the tile is stored as a number of triangles with surfaces and vertices. The object data is stored in a vertex block that stores data about each triangle in the macro tile and each vertex in the macro tile. Object data is written only once for macro tiles. The control stream data is rather similar to the object pointer of PCT / GB01 / 02536. It gives a pointer from the memory partition allocated for a particular tile to the object parameters for the objects that appear on that tile. Therefore, there is control stream data for each tile in the macro tile. This saves copy surface data for objects that appear in many tiles. When a particular tile is rendered, the control stream data allows the appropriate object data (triangle data and vertex data) to be fetched from the vertex blocks. Known configurations require fetching triangle data for a particular tile, determining the vertices of that triangle, and then fetching the appropriate vertex data before the ray / intersection procedure can be performed for rendering. there were. However, in this embodiment of the invention, a link is provided between the triangular data in the vertex block and the appropriate vertex data. Therefore, appropriate vertex data can be fetched directly. This is important. This is because you only need to fetch once, not twice. Also, it is only necessary to fetch the reduced amount of data. This is because one vertex is associated with two or more triangles. In one particular embodiment, the vertex block contains up to 32 vertices, 64 triangles.
In FIG. 4, it is clear that there is a global list 409 as well as a block of display list memory for each macro tile requested by the free storage device 413. It is also possible that the object data traverses two or more macro tiles. In this case, the object data 411 is assigned to a global list that contains only the object data in two or more macro tiles. Therefore, the display list is grouped into a macro tile plus global list. All object and control stream data in a macro tile is only addressed by the tiles that are in a given macro tile, and the object data that is in two or more macro tiles is in the global list and this Data can be addressed from the control stream.
The memory of FIG. 4 is shown as having a macro tile block at one end of the memory, followed by a global list, and then a free storage device, but this is merely an approximation and the memory. Note that it should not be interpreted as an indication of how to divide. In practice, blocks of global lists and blocks of various macro tiles can be scattered in memory based on when they are requested. The unallocated portion of memory is left in the free storage device. Therefore, both the amount of memory allocated to a particular function and its location in the display list change dynamically.
In an improved aspect, the global list contains yet another object data rather than simple object data that fits within two or more macro tiles. Typically, the system uses a pipeline processor that has a large number of cells, with adjacent cells acting on adjacent pixels. As we approach the boundaries between macro tiles, some objects that fit in only one macro tile may be mixed with objects that fit in both macro tiles on each side of the boundary. Therefore, the ISP needs to switch between readings from the macro tile block assigned to the macro tile on one side of the boundary and readings from the global list, and when crossing the boundary based on an array of objects. You need to switch several times. In order to avoid this, it is meaningful to store the object data for the objects that are close to the boundary in the global list even if they are contained in only one macro tile. This prevents the ISP from switching between readings from the macro tile block and readings from the global list several times. Therefore, the global list stores object data for objects that fall into two or more macro tiles, and also stores object data for objects that fall into only one macro tile but are close to the boundary with another macro tile. To do.
At a sufficient distance from the boundary, all objects enter only their macro tiles, so the ISP can simply read from the appropriate macro tile block in memory. As it approaches the boundary, some pixels require the ISP to read from the global list. The determination of whether to store object data in a macro tile block of memory or in a global list is made based on the number of adjacent pixels containing the object in only one macro tile. As you approach the boundary, the number of consecutive pixels containing objects that fit in only one macro tile decreases. If the number of consecutive pixels falls below a certain number, whether the subsequent object data (ie, between that pixel and the boundary) should be stored in the global list, even if these objects only fit in that macro tile. Judgment is made. Similarly, as you leave the boundary, the number of consecutive pixels containing objects that fit in only one macro tile increases. When the number of consecutive pixels rises above the threshold, subsequent object data is stored in the appropriate macro tile block in memory instead of the global list.
The conventional configuration described in PCT / GB01 / 02536 does not include a global list, but simply includes a block of memory allocated to each macro tile. This allows all memory associated with a particular macro tile to be released after that macro tile has been rendered, without affecting the display list of other macro tiles, but this is especially true. This means that there are a large number of copies of the data when there is a large object that covers a large number of macro tiles. Global lists need to render all macro tiles before they can be freed, but the effect of global lists is that such duplication can be largely avoided and memory consumption can be reduced. Also, the object data for objects that go into two or more macro tiles need only be written once to the memory, that is, the global list, not once for each macro tile, so the amount of time required to store such object data. Is shortened.
For example, when the screen is divided into four macro tiles, about 50% of the memory can be allocated to the global list for some scenes. The minimum amount of memory for the global list should be similar to the memory used for macro tiles. However, global list memory should be as limited as possible. This is because, as mentioned above, this memory cannot be recovered from the scene without rendering all the macro tiles in the scene. This will be described below.
As is apparent in FIG. 4, the data in the display list is grouped at that location, memory is allocated to each macro tile as needed, and can be released once this macro tile has been rendered. Once all the macro tiles have been rendered, the global list space can also be removed.
According to embodiments of the present invention, the display list memory is managed by a dynamic parameter management (DPM) system. Memory is allocated to the system during the tiling phase, and this memory is deallocated during the rendering phase. The DPM system manages memory allocation and deallocation during these stages, and if memory is not available, it works to schedule the data to be rendered, which frees up more memory. Therefore, the scene value of vertex data can always be tiled and rendered.
One possible configuration of the DPM system is shown in FIG. The DPM system 501 includes a dynamic parameter manager DPM503, an ISP505, a macro tile engine MTE507, a tile engine TE509, and a memory 511. The MTE507 generates object data for each macro tile and puts it in the appropriate block of memory 511. Therefore, the MTE 507 controls the object data portion of the display list memory. The TE 509 uses the object data from the MTE 507 to generate control stream data for each tile and put it in memory 511. Therefore, the TE509 controls the control stream portion of the display list memory. To track the consumption of control stream memory, the TE 509 includes a tail pointer cache 509a described below. The ISP505 uses the object data to derive depth data and stores it in the z-buffer portion of memory 511.
The process of memory allocation and deallocation will be described below with reference to FIG. These two operations are performed asynchronously and are allocated and deallocated from the internal chunk of parameter memory.
The DPM maintains a page manager 503a, which stores the linked list state and allocates and deallocates blocks of memory. The linked list state allows a particular block of memory to be associated with a macro tile, freeing the correct block of memory when the macro tile is rendered. This is shown in FIG. 5 503a. The DPM also tracks the internal state of TAs, ISPs, and TSPs, namely the current state of object data, depth data, video data, and shading data, which shows how many scenes are loaded into the display list. , And how much depends on how much has been rendered so far. When TA hardware processes object data and control stream data through separate parts of the pipeline, it is assigned separately to these two structures. DPM is only interested in allocating and deallocating pages to MTEs and TEs. Once assigned, the page is no longer of interest and all you need to know is when the assigned page can be terminated and unassigned. For this reason, flags are used to indicate when partial rendering is complete and the memory allocated for that macro tile can be unpaged. Then, when all the macro tiles have been rendered, the flag indicates that the global list page can also be cleared.
In one particular embodiment, memory is allocated to the object data in 4096 byte chunks for a given macro tile or global list, which is required by the macro tile engine that places the object data in the macro tiles. Consumed accordingly. The tile engine takes in the object data of macro tiles and generates control stream data for each tile. First, 4096 bytes of memory are allocated to the macro tile for the control stream. Memory from this page is then allocated to control stream data for a given tile (in the macro tile) as requested by the tile accelerometer. The control stream data is allocated in 16 32-bit word blocks. When the assigned page is fully used, the new page is assigned to the macro tile or global list by DPM. In this embodiment, when the first page of the control stream data is 4096 bytes and there are 16 4-byte words for each tile, the control stream data for 64 tiles is stored and then assigned. The page is full. Since the control stream data is allocated in blocks of 16 words (in this embodiment), there is a separate tail pointer cache 509a that tracks the consumption of this control stream memory. This tail pointer cache contains the address of the next available control stream. When one block is completely consumed (after 64 tiles in the embodiment), a new control stream block is allocated, which is updated to the tail pointer cache 509a.
Memory is deallocated when the rendering process takes place. When the macro tile is complete, the DPM appends the page assigned to that macro tile back to free storage. Due to the linked list structure, this writes the head of the macro tile list value to the tail address of the free storage list, and updates the free list tail address to the macro tile free list address.
Traditional systems could render a scene once it was tiled. Therefore, tile processing (by TA) and rendering (by ISP) were done in series. Now, in partial rendering, ISPs and TAs must operate on the same data at the same time, i.e. in parallel. To achieve this without collapsing the memory allocation list, consider tinkering with TAs and ISPs and handling different data fragments. This is done by double buffering the head and tail information for each macro tile and global list. This means that TA can act on the data at the same time as the ISP. This is because they are each directed to individual head and tail information.
Ideally, the ISP and TA should operate in parallel, and there should always be enough memory to allocate memory to the TA as needed. However, in some cases (eg, the macro tiles are small or the video is very complex), there may not be enough memory to allocate immediately.
If TA assignment is not possible, the system checks to see if rendering is in progress. If so, the system will continue to try to allocate memory as memory is steadily available for rendering. When sufficient memory is freed by rendering, TA allocation proceeds.
Eventually, all pending renderings are complete, and the DPM must perform rendering management, i.e. select macro tiles to render before more allocations can be made.
Therefore, during the process, performance will steadily decline as the available memory space decreases. The process is eventually serialized as the amount of memory space decreases, and at this stage the choice of macro tiles for which partial rendering should be performed becomes very important. In PCT / GB01 / 02536, when the display list memory fills up to a certain threshold, macro tiles to be rendered are selected to free more memory. It is suggested that the macro tiles to be rendered are selected based on the amount of memory released. An improved way to choose which macro tile to render is described below.
As mentioned above, both the z-buffer memory and the display list memory are housed in a block of memory, i.e., parameter memory. How the z-buffer and display list are allocated will be described with reference to FIGS. 6 and 7.
The hardware reserves a portion of the parameter memory as a z-buffer, which is at least as large as the size of the macro tile, i.e. enough memory to store depth data for each pixel of the macro tile. .. That is, even if the memory is used up, some reserved memory is always left in order to perform partial rendering.
The entire parameter memory must be larger than the z-buffer memory so that it can recover from partial rendering caused by insufficient memory. The minimum amount of total parameter memory must be equal to the sum of the total z-buffer memory and the space for one yet another macro tile. Therefore, even when partial rendering is performed on all macro tiles (so that the z-buffer is full), more object data for yet another macro tile can still be inserted into the display list memory. If the memory required for the z-buffer is 1 unit (0.25 units for each macro tile), the minimum parameter memory required is the sum of 1 and 0.25, or 1.25. It is a unit.
For example, if the screen size is 1024x1024 pixels, the z-buffer format is 32 bits per pixel, and the screen is divided into four macro tiles, then a total z-buffer of about 4 MB is required. Therefore, the total parameter memory should be about 5MB (ie, the sum of 4MB plus an additional 1MB for another macro tile). Alternatively, if the screen size is 2048x2048 pixels, the z-buffer format is 32 bits per pixel, and the screen is divided into 16 macro tiles, then a total z-buffer of about 16 MB is required. Therefore, the total parameter memory should be about 17MB (ie, the sum of 16MB plus an additional 1MB for another macro tile).
(In the conventional configuration shown in FIG. 2, compression was performed before storing in the z-buffer and decompression was performed when reading from the z-buffer. Such a compression / decompression unit was used in the present invention. In some cases, it is actually necessary to reserve a z-buffer equal to the two macro tiles. Therefore, when new depth data is stored, that data is stored before the previous depth data is decompressed. There must be space to compress. In this case, two macrotile value z-buffers (rather than one) must always be reserved.)
The choice of macro tile to render is based on the need to be able to insert new object data for the macro tile or perform another partial rendering after that partial rendering. If none of these are satisfied, the system is blocked and no further action can be taken. Therefore, the macro tiles to be rendered are carefully selected.
Two examples of macro tile selection will be described below.
Example 1: FIG. 6A-6F In this first embodiment, memory is evenly allocated to each macro tile. As shown in FIG. 6a, the visible screen has four macro tiles. Memory allocations for these macro tiles are shown in FIGS. 6b-6f.
First, all memory is allocated to each macro tile in 0.25 units. The z-buffer portion (equal to the size of one macro tile, i.e. 0.25) is reserved-ZB0. This is shown in FIG. 6a.
The reserved z-buffer ZB0 can be used for partial rendering of one macro tile. They all use the same amount of memory space, so there is no problem with any of them. Partial rendering is performed on the macro tile MT3. This gives the results shown in FIG. 6b. The macro tile MT3 is freed up to ZB0 by partial rendering, but the freed space must be reserved in more z-buffer-ZB1. Therefore, there is no free space for inserting new object data, but ZB1 can be used to perform another partial rendering.
Therefore, the reserved z-buffer ZB1 can be used for partial rendering of another macro tile, in this case MT2. Partial rendering is performed, giving the results shown in FIG. 6c. The macro tile MT2 is freed up to ZB1 by partial rendering, but this freed space must be reserved for more z-buffer-ZB2. Therefore, there is no free space to insert new object data, but ZB2 can be used to perform another partial rendering.
Therefore, the reserved z-buffer ZB2 can be used for partial rendering of another macro tile, in this case MT1. Partial rendering is performed, giving the results shown in FIG. 6d. The macro tile MT1 is freed up to ZB2 by partial rendering, but this freed space must be reserved for more z-buffer-ZB3. Therefore, although there is still no free space to insert new object data, ZB3 can be used to perform another partial rendering.
Therefore, the reserved z-buffer ZB3 can be used for partial rendering of the final macro tile MT0. Partial rendering is performed, giving the results shown in FIG. 6e. Since the macro tile MT0 has been released and all the macro tiles are now allocated the z-buffer, there is no need to reserve any more z-buffers.
Therefore, free space can be used to insert more object data for one tile. The scene is then completed, with partial rendering, more object data inserted, and so on, at which point the parameter memory and z-buffer memory can be deallocated.
Therefore, memory is initially evenly allocated to the four macro tiles, so all tiles must be rendered before more object data can be inserted. As an optimization, if the system finds that the memory is evenly distributed to all macro tiles at any stage, then all macro tiles can be rendered immediately.
The case shown in FIG. 6 is the worst scenario because the maximum memory that can be released by partial rendering is 0.25 units. In other cases where memory is not evenly allocated to macro tiles, partial rendering can always free up more memory space.
Example 2: FIG. 7A-7G In this second embodiment, memory is non-uniformly allocated to each macro tile. In this case, when the memory is full, the macro tile with the most memory allocated is selected for rendering. This is a more common case and is shown in FIGS. 7a-7g. FIG. 7a shows the macro tiles on the screen, while FIG. 7b-7g shows the memory usage for each macro tile. Macro tiles for video are of equal size, but memory allocation is different based on the amount of object data in each macro tile.
First, all memory is allocated to macro tile MT2 in 0.4375 units, to macro tile MT0 in 0.25 units, to macro tile MT3 in 0.1875 units, and to macro tile MT1 in 0.125 units. The z-buffer portion (equal to the size of one macro tile on the screen, ie 0.25) is reserved-ZB0. This is shown in FIG. 7b.
The reserved z-buffer ZB0 can be used for partial rendering of one macro tile. Macro tile MT2 is selected because the maximum amount of memory is allocated. When partial rendering is performed on macro tile MT2, 0.25 units of free space must be reserved as z-buffer-ZB1, but the remaining space is free for more object data. .. This is shown in FIG. 7c.
More object data in MT2 is then loaded into that free space until the memory is full. This is shown in FIG. 7d.
The reserved z-buffer ZB1 can then be used for partial rendering of macro tile MT1. That's because the macro tile is currently allocated the maximum amount of memory. When partial rendering is performed in MT1, the space freed by it must be reserved in more z-buffer-ZB2. Therefore, there is no free space to insert more object data. This is shown in FIG. 7e.
Since ZB0 has already been assigned to MT2, another partial rendering can be performed in MT2. This frees up 0.1875 free space. ZB2 can also be used for partial rendering of MT3. In total, this frees up 0.375 space, of which 0.25 must be reserved as z-buffer ZB3. This is shown in FIG. 7f.
At this stage, the freed space can be used for more object data for MT0, MT2 or MT3, in which case partial rendering will use the z-buffer assigned to MT0, MT2 and MT3. Is done. Alternatively, the freed space is used for more object data for MT1. In this case, ZB3 is used for partial rendering of macro tile MT1. This is shown in FIG. 7g. The macro tile MT1 has been released and all macro tiles are now allocated z-buffers, eliminating the need to reserve more z-buffers.
The embodiment shown in FIG. 7 is easier to handle than the embodiment shown in FIG. This is because the memory is not evenly allocated to the macro tiles, so more than 0.25 units of memory is freed after the first partial rendering.
Figures 6 and 7 show two examples of rendering management when memory is low, processes are effectively serialized, and the choice of macro tiles to render at each stage becomes very important. Is shown. These two cases show that the process can continue with the minimum amount of z-buffer equivalent to one macro tile always reserved.
As mentioned above, global list memory should be as limited as possible because it cannot recover this memory from the scene without rendering all the macro tiles in the scene. The global list is limited to the extra memory allocated in addition to the minimum requirement of z-buffer size x 1.25 described above. If this is 50% of macrotile memory, the total minimum memory requirement is 1.75 for the size of the z-buffer (ie, 1 unit for display list memory, 0.25 reserved for z-buffer, And about the global list 0.5).
However, in order to be able to render any complex scene, the global object buffer must be recoverable, so this embodiment allows multiple macro tiles to be split with respect to their global object list, and then all. This is made possible by rendering the macro tiles of. Splitting is achieved by creating a new context for all macro tiles and continuing to render into the new context. In parallel with this, the tiled data is processed and released back to the system. This method allows both the assign and deallocation parts of the pipeline to remain active.
In certain embodiments of the invention, the device supports up to 256 MB of parameter memory divided into 4 kB pages or blocks, up to 16 macro tiles per scene, and one global list.
The system described above can be used to run two or more applications at the same time. For example, if you open two windows, an application, on your PC screen, you can use the same hardware to generate video data for both applications. FIG. 8a shows a first configuration for generating video data for two applications using the hardware of FIG. FIG. 8b shows a second configuration for generating video data for two applications using the hardware of FIG. 8a and 8b can be very easily extended to three or more applications.
The system 801 of FIG. 8a includes a TA801, an ISP811, a TSP813, and a frame buffer 815. However, in contrast to the system of FIG. 3, the TA803 and ISP811 can access two separate memories. The first memory 809A is for the first application and includes a display list memory 805A for the first application and a z-buffer memory 807A for the first application. The second memory 809B is for a second application and includes a display list memory 805B for the second application and a z-buffer memory 807B for the second application.
This system operates exactly like the configuration of FIG. 3 described above. The system allocation to each application can be allocated by time division multiplexing (TDM) or by another method.
The system 801'in FIG. 8b includes a TA801', an ISP811', a TSP813', and a frame buffer 815'. However, in contrast to the system of FIG. 8a, the memory for the two applications is housed in a chunk of memory 805. This reduces the memory required. The chunk of memory 805 allocates memory to each application as needed. In FIG. 8b, some of the memory 809A'is allocated to the first application. That portion of memory 809A'includes display list memory and z-buffer memory. Some of the memory 809B'is allocated to the second application. That portion of memory 809B'includes display list memory and z-buffer memory. Memory that has not yet been allocated is left in the free storage device 807.
Further, the system of FIG. 8b operates in exactly the same manner as the configuration of FIG. 3 described above. The system allocation to each application can be allocated by time division multiplexing (TDM) or by another method. Since the memory is allocated to each application as needed, the memory is used more efficiently than in the case of FIG. 8a.
In both of the above embodiments, when actually running two or more applications using the same hardware, the system must include some means of storing the internal state for each application. The internal state is the current state of the TA and ISP for the application, that is, the currently stored object data, depth data, video data, and shading data. This provides a record of the progress that has occurred in storing object data, rendering, etc. for a particular application. Therefore, when hardware is swapped between a large number of applications, it is known where to start when each application is reached.
In the above embodiment, it is always necessary to reserve enough memory for the depth data of one macro tile so that any complex scene can occur while only one application is running. It was noticed that there was. So how would this rule apply when more than one application runs?
First, consider the embodiment of FIG. 8a in which the memory of each application is separate. In this case, the memory of each application needs to reserve the z-buffer for one macro tile. That is, in total, there are n memories for n applications that are executed simultaneously, and each of these n memories has a reserved portion of the z-buffer.
Next, consider the embodiment of FIG. 8b, in which the memory for a large number of applications is housed in a chunk of memory. In this case, it is only necessary to reserve the z-buffer for one macro tile, regardless of the total number of applications. This is because only one partial rendering is done at a particular time, so the same reserved memory space can be used for all applications. This is a further effect of the configuration of FIG. 8b.
303: TA 305: Display list memory 307: Z-buffer memory 311: ISP 313: TSP 315: Frame buffer memory 401: Display list 403a, 403b: Macro tile 405a, 405b: Control stream 407a, 407b: Object data 409: Global list 413: Free storage device 501: DPM system 503: Dynamic Parameter Manager DPM 505: ISP 507: Macro tile engine MTE 509: Tile engine TE 511: Memory
20 sheets
Sheet 1 Sheet 2 Sheet 3 Sheet 4 Sheet 5 Sheet 6 Sheet 7 Sheet 8 Sheet 9 Sheet 10 Sheet 11 Sheet 12 Sheet 13 Sheet 14 Sheet 15 Sheet 16 Sheet 17 Sheet 18 Sheet 19 Sheet 20
Every citation, both ways
| Document | Relation | Office |
|---|---|---|
| JP07134776A | Cites | Japan |
| JP04112386A | Cites | Japan |
| JP2003536153A | Cites | Japan |
| JP2002529865A | Cites | Japan |
| 今村和宏, 外2名,”レンダラ解体新書”,CG WORLD,日本,株式会社ワークスコーポレーション,2002年 3月 1日,第43巻,p.28-57 | Non-patent | – |
20 members in 5 offices
Members20
| Document | Office | Kind | |
|---|---|---|---|
| GB0619327D0 | United Kingdom | D0 | |
| GB2442266A | United Kingdom | A | |
| WO2008037954A2 | World Intellectual Property Organization (WIPO) | A2 | |
| US2008186318A1 | United States of America | A1 | |
| GB0815937D0 | United Kingdom | D0 | |
| GB0815938D0 | United Kingdom | D0 | |
| GB2442266B | United Kingdom | B | |
| GB2449398A | United Kingdom | A | |
| GB2449399A | United Kingdom | A | |
| GB2449398B | United Kingdom | B | |
| WO2008037954A3 | World Intellectual Property Organization (WIPO) | A3 | |
| GB2449399B | United Kingdom | B | |
| EP2070049A2 | European Patent Office (EPO) | A2 | |
| JP2010505172A | Japan | A | |
| JP2012198931A | Japan | A | |
| JP2012198932A | Japan | A | |
| JP5246563B2 | Japan | B2 | |
| US8669987B2 | United States of America | B2 | |
| JP5545554B2This record | Japan | B2 | |
| JP5545555B2 | Japan | B2 |
21 legal events, as the office reported them to INPADOC
Over the term
Point at a mark for the eventEvents
| Event | Code | |
|---|---|---|
| Receipt of annual feesJAPANESE INTERMEDIATE CODE: R250R250 | R250 | |
| Receipt of annual feesJAPANESE INTERMEDIATE CODE: R250R250 | R250 | |
| Receipt of annual feesJAPANESE INTERMEDIATE CODE: R250R250 | R250 | |
| Receipt of annual feesJAPANESE INTERMEDIATE CODE: R250R250 | R250 | |
| Receipt of annual feesJAPANESE INTERMEDIATE CODE: R250R250 | R250 | |
| Receipt of annual feesJAPANESE INTERMEDIATE CODE: R250R250 | R250 | |
| Receipt of annual feesJAPANESE INTERMEDIATE CODE: R250R250 | R250 | |
| Receipt of annual feesJAPANESE INTERMEDIATE CODE: R250R250 | R250 | |
| Receipt of annual feesJAPANESE INTERMEDIATE CODE: R250R250 | R250 | |
| Receipt of annual feesJAPANESE INTERMEDIATE CODE: R250R250 | R250 | |
| Certificate of patent or registration of utility modelJAPANESE INTERMEDIATE CODE: R150R150 | R150 | |
| First payment of annual fees (during grant procedure)JAPANESE INTERMEDIATE CODE: A61A61 | A61 | |
| Written decision to grant a patent or to grant a registration (utility model)JAPANESE INTERMEDIATE CODE: A01A01 | A01 | |
| Decision of grant or rejection writtenTRDD | TRDD | |
| Request for written amendment filedJAPANESE INTERMEDIATE CODE: A523A521 | A521 | |
| Written permission of extension of timeJAPANESE INTERMEDIATE CODE: A602A602 | A602 | |
| Written request for extension of timeJAPANESE INTERMEDIATE CODE: A601A601 | A601 | |
| Notification of reasons for refusalJAPANESE INTERMEDIATE CODE: A131A131 | A131 | |
| Report on retrievalJAPANESE INTERMEDIATE CODE: A971007A977 | A977 | |
| Request for written amendment filedJAPANESE INTERMEDIATE CODE: A523A521 | A521 | |
| Written request for application examinationJAPANESE INTERMEDIATE CODE: A621A621 | A621 |
Numbers
- Publication
- 5545554
- Application
- 137478
Titles2
- Japanese
- 三次元コンピュータ映像を発生するシステムのためのメモリマネージメントの改良
- English
- Improvement of the memory management for the system which generates a three-dimensional computer image
Classification
- CPC, 4
- G06T1/60
- G06T15/005
- G06T2200/08
- G06T15/405
- IPC, 1
- G06T15 00
