Methods and apparatus for navigating an image
15 claims: 4 independent, 11 dependent
- 1演算装置および記憶装置を有するコンピュータが少なくとも1つのオブジェクトを有するイメージをズーミングする方法であって、前記演算装置に、 式p=d’・z a に基づき、前記 記憶 装置に格納された少なくとも1つのオブジェクトの少なくともいくつかの要素の拡大および/または縮小されたサイズを計算するステップであって、 pは前記計算された拡大および/または縮小のズームレベルでの前記オブジェクトの1つまたは複数の要素のピクセル単位での線形サイズであり、d’は物理単位での前記オブジェクトの前記1つまたは複数の要素の帰属線形サイズであり、zは物理線形サイズ/ピクセル単位での前記ズームレベルであり、 前記ズームレベルzが第1の範囲にあるとき、前記帰属線形サイズd’はd1、-1<a<0として計算し、 前記ズームレベルzが前記第1の範囲に連続する第2の範囲にあるとき、前記帰属線形サイズd’はd2、a=-1として計算し、 前記ズームレベルzが前記第1の範囲から前記第2の範囲に変化するとき、前記線形サイズpが連続的に変化するよう前記帰属線形サイズd1、d2が定められている、計算するステップと、 前記計算した拡大および/または縮小されたサイズの前記オブジェクトに基づき前記イメージの拡大および/または縮小されたイメージを生成するステップと を実行させることを特徴とする方法。
- 2d’およびaのうちの少なくとも1つは前記オブジェクトのうちの1つまたは複数の要素について異なりうることを特徴とする請求項1に記載の方法。
- 3前記ズームレベルzの指数aはズームレベルz0からz1の範囲内で-1<a<0であり、z0はz1よりも小さい物理線形サイズ/ピクセルであることを特徴とする請求項1に記載の方法。
- 4z0、z1、d’、およびaのうちの少なくとも1つは前記オブジェクトのうちの1つまたは複数の要素について異なりうることを特徴とする請求項3に記載の方法。
- 5前記ズームレベルzが前記第2の範囲にあるときd’=c・dであり、cは定数であり、dは前記オブジェクトの前記1つまたは複数の要素の物理単位での実線形サイズまたは帰属線形サイズであることを特徴とする請求項1に記載の方法。
- 6演算装置および記憶装置を有するコンピュータの前記演算装置に、 式p=d’・z a に基づき、前記記憶装置に格納された少なくとも1つのオブジェクトの少なくともいくつかの要素の拡大および/または縮小されたサイズを計算するステップであって、 pは前記計算された拡大および/または縮小のズームレベルでの前記オブジェクトの1つまたは複数の要素のピクセル単位での線形サイズであり、d’は物理単位での前記オブジェクトの前記1つまたは複数の要素の帰属線形サイズであり、zは物理線形サイズ/ピクセル単位での前記ズームレベルであり、 前記ズームレベルzが第1の範囲にあるとき、前記帰属線形サイズd’はd1、-1<a<0として計算し、 前記ズームレベルzが前記第1の範囲に連続する第2の範囲にあるとき、前記帰属線形サイズd’はd2、a=-1として計算し、 前記ズームレベルzが前記第1の範囲から前記第2の範囲に変化するとき、前記線形サイズpが連続的に変化するよう前記帰属線形サイズd1、d2が定められている、計算するステップと、 前記計算した拡大および/または縮小されたサイズの前記オブジェクトに基づき前記オブジェクトを有するイメージの拡大および/または縮小されたイメージを生成するステップと を実行させるためのプログラムを記録したコンピュータ読取り可能な記憶媒体。
- 7d’およびaのうちの少なくとも1つは前記オブジェクトのうちの1つまたは複数の要素について異なりうることを特徴とする請求項6に記載のコンピュータ読取り可能な記憶媒体。
- 8前記ズームレベルzの指数aはズームレベルz0からz1の範囲内で-1<a<0であり、z0はz1よりも小さい物理線形サイズ/ピクセルであることを特徴とする請求項6に記載のコンピュータ読取り可能な記憶媒体。
- 9z0、z1、d’、およびaのうちの少なくとも1つは前記オブジェクトのうちの1つまたは複数の要素について異なりうることを特徴とする請求項8に記載のコンピュータ読取り可能な記憶媒体。
- 10前記ズームレベルzが前記第2の範囲にあるときd’=c・dであり、cは定数であり、dは前記オブジェクトの前記1つまたは複数の要素の物理単位での実線形サイズまたは帰属線形サイズであることを特徴とする請求項6に記載のコンピュータ読取り可能な記憶媒体。
- 11処理装置と、 少なくとも1つのオブジェクトを有するイメージ、および、前記処理装置に前記少なくとも1つのオブジェクトを有するイメージをズーミングさせるプログラムを記録した記憶装置であって、前記プログラムは、 式p=d’・z a に基づき、前記少なくとも1つのオブジェクトの少なくともいくつかの要素の拡大および/または縮小されたサイズを計算するステップであって、 pは前記計算された拡大および/または縮小のズームレベルでの前記オブジェクトの1つまたは複数の要素のピクセル単位での線形サイズであり、d’は物理単位での前記オブジェクトの前記1つまたは複数の要素の帰属線形サイズであり、zは物理線形サイズ/ピクセル単位での前記ズームレベルであり、 前記ズームレベルzが第1の範囲にあるとき、前記帰属線形サイズd’はd1、-1<a<0として計算し、 前記ズームレベルzが前記第1の範囲に連続する第2の範囲にあるとき、前記帰属線形サイズd’はd2、a=-1として計算し、 前記ズームレベルzが前記第1の範囲から前記第2の範囲に変化するとき、前記線形サイズpが連続的に変化するよう前記帰属線形サイズd1、d2が定められている、計算するステップと、 前記計算した拡大および/または縮小されたサイズの前記オブジェクトに基づき前記イメージの拡大および/または縮小されたイメージを生成するステップとを 前記処理装置に実行させる 、記憶装置とを備えたことを特徴とするコンピューティング装置。
- 12d’およびaのうちの少なくとも1つは前記オブジェクトのうちの1つまたは複数の要素について異なりうることを特徴とする請求項11に記載のコンピューティング装置。
- 13前記ズームレベルzの指数aはズームレベルz0からz1の範囲内で-1<a<0であり、z0はz1よりも小さい物理線形サイズ/ピクセルであることを特徴とする請求項11に記載のコンピューティング装置。
- 14z0、z1、d’、およびaのうちの少なくとも1つは前記オブジェクトのうちの1つまたは複数の要素について異なりうることを特徴とする請求項13に記載のコンピューティング装置。
- 15前記ズームレベルzが前記第2の範囲にあるときd’=c・dであり、cは定数であり、dは前記オブジェクトの前記1つまたは複数の要素の物理単位での実線形サイズまたは帰属線形サイズであることを特徴とする請求項11に記載のコンピューティング装置。
Independent claims15
261 paragraphs, as filed
The present invention relates to methods and devices for navigating, such as zooming and panning, across an image of an object to provide the appearance of smooth, continuous navigating movements.
Most traditional graphical computer user interfaces (GUIs) are designed to use fixed spacial scale visual components, but for a long time the visual components have been on the display. Although it does not have a fixed spatial scale, it has been recognized that it can actually be represented and manipulated so that it can be panned and / or zoomed in or out. The ability to zoom in and out on an image is relevant, for example, to viewing maps, browsing text layouts such as newspapers, displaying digital photos, displaying blueprints or diagrams, and displaying other large datasets. Is desirable.
Many existing computer applications, such as Microsoft Word, Adobe Photo Shop, and Adobe Acrobat, include zoomable components. Generally, the zoom function provided by these computer applications is not a very important aspect for the user to interact with the software, and the zoom function is used only occasionally. These computer applications allow the user to pan smoothly and continuously across the image (eg, use scroll bars or cursors to move the displayed image left, right, up, or down). It is a thing. However, an important problem with these computer applications is that the user cannot zoom smoothly and continuously. In practice, it provides discrete steps of zooming in, such as 10%, 25%, 50%, 75%, 100%, 150%, 200%, 500%. When the user uses the cursor to select the desired zoom, the image responds with a sharp change to the selected zoom level.
Internet-based computer applications also have undesired quality of discontinuous zoom. The computer application that underlies the www.mapquest.com website illustrates this point. On the MapQuest website, users can enter one or more addresses and receive a road map image in response. Figures 1-4 are examples of images that can be obtained from the MapQuest website in response to a regional map inquiry in Long Island, NY, USA. The MapQuest website allows users to zoom in and out to discrete levels such as 10 levels. Figure 1 is a rendering at zoom level 5 of approximately 100 meters / pixel. Figure 2 is an image of zoom level 6 at approximately 35 meters / pixel. Figure 3 is an image of zoom level 7 at approximately 20 meters / pixel. Figure 4 is an image of zoom level 9 at approximately 10 meters / pixel.
As can be seen by comparing Figures 1 to 4, the steep transition between zoom levels causes a sudden and steep loss of detail when zooming out and a sudden and steep addition of detail when zooming in. For example, Figure 1 (zoom level 5) does not show local, secondary, or connecting roads, but the next zoom level, Figure 2, shows secondary and secondary roads. A connecting road suddenly appears. These steep discontinuities are very unpleasant when using the MapQuest website. However, keep in mind that even if the MapQuest software application was modified to display, for example, open roads at zoom level 5 (Figure 1), the results would still be unsatisfactory. The visual density of the map varies with the zoom level so that the results are satisfactory at a certain zoom level (eg level 7 in Figure 3), but when zoomed in, the roads are not dense and the map is overly sparse. May be visible. When zoomed out, the roads eventually mix with each other, quickly forming nests with no gaps that make individual roads indistinguishable.
The ability to provide smooth, continuous zooming on road map images is often problematic due to the different levels of roughness associated with road categories. U.S. roads include A1: Major highways, A2: Main roads, A3: Highways, secondary roads, and connecting roads, A4: General roads, city roads, and local roads, A5: Unpaved roads. There are approximately five road categories (based on Tiger / Line Data distributed by the US Census). These roads can be considered as elements of the whole object (ie road map). The roughness of the road elements is reflected in the fact that there are significantly more A4 roads than A3 roads, there are significantly more A3 roads than A2 roads, and there are significantly more A2 roads than A1 roads. In addition, the physical dimensions of the roads (eg their width) vary significantly. The A1 road can be about 16 meters wide, the A2 road can be about 12 meters wide, the A3 road can be about 8 meters wide, the A4 road can be about 5 meters wide, and the A5 road can be about 2.5 meters wide.
The MapQuest computer application addresses these different levels of roughness by displaying only the road categories that appear to be appropriate at a particular zoom level. For example, a national view might show only A1 roads, a state view might show A1 and A2 roads, and a group view might show A1, A2, and A3 roads. Even if MapQuest was modified to allow continuous zooming of road maps, this approach would lead to the sudden appearance and disappearance of confusing, visually unpleasant, road categories during zooming.
<p> In view of the above, the art provides images of complex objects that allow smooth and continuous zooming of the image while preserving the visual differences between the elements of the object based on size or importance. There is a need for new methods and devices for navigating.</p>
<p> According to one or more aspects of the invention, at least some elements of at least one object are magnified and / or so as to be non-physically proportional to one or more zoom levels associated with zooming. Methods and devices for performing various actions, including zooming in or out of an image with at least one object that is scaled down, are contemplated.</p><p> Non-physically proportional scaling is the formula p = c d z<sup>a</sup>In this formula, p is the linear size of one or more elements of the object in pixels at the zoom level, c is a constant, and d is one or more of the objects. The linear size of multiple elements in physical units, z is the physical linear size / zoom level in pixels, and a is a scale exponent of a -1.</p><p> Under non-physical scaling, the scale exponent a is not equal to -1 (usually -1 <a <0) within zoom levels z0 to z1, and z0 is a physical linear size / pixel below z1. is there. Preferably, at least one of z0 and z1 may be different for one or more elements of the object. Note that a, c, and d may also differ depending on the element.</p><p> At least some elements of at least one object can also be zoomed in and out so that they are physically proportional to one or more zoom levels associated with zooming. Physically proportional scaling can be expressed by the formula p = c · d / z, where p is the linear size of one or more elements of the object at the zoom level in pixels. Is a constant, d is the linear size of one or more elements of the object in physical units, and z is the physical linear size / zoom level in pixels.</p><p> The methods and devices described above and / or described herein are any known processor, programmable digital device or system, programmable array that can operate to run standard digital circuits, analog circuits, software and / or firmware programs. Note that this can be achieved using any known technique, such as a logical device, or any combination of the above. The present invention can also be embodied in a software program for storage in a suitable storage medium and for execution by a processing unit.</p><p> The elements of an object can have different degrees of roughness. For example, as mentioned above, the roughness of the elements of a road map object is significantly higher on A4 roads than on A3 roads, significantly higher on A3 roads than on A2 roads, and on A2 roads than on A1 roads. It appears because there are quite a lot of. The degree of roughness in the road category is also reflected in characteristics such as average road length, intersection frequency, and maximum curvature. The roughness of the elements of other image objects also appears in other aspects, but there are too many to enumerate all. Therefore, the scaling of elements in a given given image is physical based on (i) the degree of roughness of these elements and (ii) the zoom level of a given given image. May be proportional to or non-physically proportional to. For example, an object can be a road map, an element of an object can be a road, and different degrees of roughness can be a road hierarchy. Therefore, the scaling of a given road in a given given image is physically proportional based on (i) the road hierarchy of the given road and (ii) the zoom level of the given given image. Or it may be non-physically proportional.</p><p> According to one or more other aspects of the invention, one or more actions of receiving multiple pre-rendered images of road maps with different zoom levels on the client terminal and one or more including zooming information on the client terminal. To get an intermediate zoom level intermediate image that corresponds to the zooming information of the navigation command, so that the action of receiving the user navigation command and the display of the intermediate image on the client terminal provide a smooth navigation look. Methods and devices for performing various actions are contemplated, including actions that blend two or more pre-rendered images.</p><p> According to one or more other aspects of the invention, an action of receiving a plurality of pre-rendered images of at least one object with different zoom levels on a client terminal, of the at least one object. At least some elements are scaled up and / or scaled down to create multiple given images, and this scaling is (i) physically proportional to the zoom level, and (ii) non-physically proportional to the zoom level. Between the action to receive, the action to receive one or more user navigation commands containing zooming information on the client terminal, and the intermediate zoom level corresponding to the zooming information of the navigation command, which is at least one of Methods and devices for performing various actions, including the action of blending two or more pre-rendered images to obtain an image and the action of displaying an intermediate image on a client terminal. It is planned.</p><p> According to one or more other aspects of the invention, an action of transmitting a plurality of pre-rendered images of road maps with different zoom levels to a client terminal over a communication channel and a plurality of pre-rendered images on the client terminal. The action of receiving the image, the action of issuing one or more user navigation commands containing zooming information using the client terminal, and the display of the intermediate image on the client terminal provide a smooth navigation appearance. To perform a variety of actions, including blending two or more pre-rendered images to get an intermediate zoom level intermediate image that corresponds to the zooming information of the navigation command. Methods and devices are contemplated.</p><p> According to one or more other aspects of the invention, an action of transmitting a plurality of pre-rendered images of at least one object with different zoom levels to a client terminal over a communication channel, the at least one of which. At least some elements of an object are scaled up and / or scaled down to create multiple given images, and this scaling is (i) physically proportional to the zoom level, and (ii) to the zoom level. At least one of non-physically proportional, an action to send, an action to receive multiple pre-rendered images on the client terminal, and one or one containing zooming information using the client terminal. An action that issues multiple user navigation commands, an action that blends two pre-rendered images to get an intermediate zoom level intermediate image that corresponds to the zooming information of the navigation command, and an intermediate image on the client terminal. Methods and devices for performing various actions are contemplated, including actions that display.</p><p> Other aspects, features, and advantages will become apparent to those skilled in the art as described herein with the accompanying drawings.</p><p> It will be appreciated that the shapes shown in the drawings are intended to illustrate the invention, and the invention is not limited to the precise arrangements and means illustrated.</p>
Then, referring to the drawings, the same numbers represent the same elements, and Figures 5-11 show a series of images representing the road network in Long Island, NY, USA, where each image has a different zoom level (or resolution). ). Before delving into the technical details of how the present invention is practiced, the desired feature that results from the use of the present invention, at least the appearance of smooth and continuous navigation, while maintaining the integrity of the information. , Especially in relation to zooming, these images will be described.
It should be noted that the various aspects of the invention described below are applicable in contexts other than navigation of road map images. In fact, the range of images and implementations for which the present invention can be used cannot be enumerated in their entirety due to the large number. For example, the features of the present invention can be used to navigate images such as human anatomical maps, complex topographic maps, engineering drawings such as wiring diagrams and blueprints, and Gene Ontology. However, the present invention has been found to have specific applicability for navigating images with different levels of detail or roughness of its elements. Therefore, for brevity and clarity, various aspects of the invention are described in relation to a particular example, an image of a road map.
Although it is not possible to demonstrate a smooth, continuous zooming appearance in a patent document, this feature is through experimentation and prototype development by running a suitable software program on a Pentium®-powered computer. Has been proven. The road map image 100A illustrated in Figure 5 is a zoom level that can be characterized in units of physical length / pixel (or physical linear size / pixel). In other words, the zoom level z represents the actual physical linear size represented by a single pixel in image 100A. In Figure 5, the zoom level is about 334 meters / pixel. Those skilled in the art will appreciate that this zoom level can be expressed in other units without departing from the claimed gist and scope of the invention. Figure 6 is the same road map image 100B as Figure 5, but with a zoom level of about 191 meters / pixel.
According to one or more aspects of the invention, a user of a software program that embodies one or more aspects of the invention may zoom in or out between the levels shown in FIGS. 5 and 6. Can be done. Such zooming is to have the appearance of a smooth, continuous transition between the 334m / pixel level (Figure 5) and the 191m / pixel level (Figure 6), and any level between them. It is important to keep in mind. Similarly, users can see z = 109.2 meters / pixel (Figure 7), z = 62.4 meters / pixel (Figure 8), z = 35.7 meters / pixel (Figure 9), z = 20.4 meters / pixel (Figure 10), and You can zoom to other levels such as z = 11.7 meters / pixel (Figure 11). Again, advantageously, transitions through these zoom levels and any level between them have the appearance of smooth, continuous motion.
Another important feature of the invention, shown in FIGS. 5-11, is that there are few or no details that appear or disappear abruptly when zooming from one level to another. The details shown in Figure 8 (zoom level z = 62.4 meters / pixel) can also be found in Figure 5 (zoom level z = 334 meters / pixel). This is true even if the image object, which in this case is a road map, contains elements of varying degrees of roughness (ie, roads). In fact, the road map 100D of FIG. 8 includes at least A1 highways such as 102, A3 secondary roads such as 104, and A4 general roads such as 106. These details are still found in image 100A in Figure 5, even on A4 general road 106, but are significantly zoomed out when compared to image 100D in Figure 8.
In addition, at zoom level z = 334 meters / pixel (Figure 5), A4 roads 106 can be seen, but A1, A2, A3, and A4 roads can be distinguished from each other. Even the distinction between A1's main highway 102 and A2's main road 108 is distinguishable from each other, as opposed to the relative weights given to these roads in the rendered image 100A.
Advantageously, if the user continues to zoom in to, for example, zoom level z = 20.4 meters / pixel as shown in image 100F of FIG. 10, the ability to distinguish between road hierarchies is also maintained. The weight of A1's main highway 102 is significantly increased compared to the zoom level z = 62.4 meters / pixel in Figure 8, but with other details such as A4's open road 106 or even A5's dirt road. It does not increase enough to erase it. Nevertheless, the weight of low-level roads, such as A4 general road 106, is significantly increased compared to the weight of its relative roads at zoom level z = 62.4 meters / pixel in Figure 8.
Therefore, there is considerable dynamic range between the zoom level shown in Figure 5 and the zoom level shown in Figure 11, and the details remain nearly consistent (ie, suddenly appear and disappear while zooming smoothly). Even if there are no roads to do), the information that the user is trying to get at a given zoom level is not hidden by unwanted artifacts. For example, at a zoom level of z = 334 meters / pixel (Figure 5), the user may want to get a general idea of what major highways exist and in which direction they extend. This information can be easily obtained even if the A4 general road 106 is displayed. At a zoom level of z = 62.4 meters / pixel (Figure 8), the user may want to know if a particular A1 main highway 102 or A2 main road 108 is servicing a particular town or area. Again, the user can obtain this information without being disturbed by other, more detailed information, such as the presence and extent of A4's open road 106, or even A5's unpaved road. Ultimately, at zoom level z = 11.7m / pixel, users are interested in finding certain A4 public roads such as 112 and are obstructed by fairly large roads such as the A1 main highway 102. You can do this without.
To achieve one or more of the various aspects of the invention described above, one or more computing devices execute one or more software programs that cause the computing devices to perform appropriate actions. It is intended to do. In this regard, then see FIGS. 12-13, which are flow diagrams showing process steps preferably performed by one or more computing devices and / or related devices.
This process flow is preferably performed by a commercially available computing device (such as a computer with a Pentium®), but this process step is optional without departing from the spirit and scope of the claimed invention. It can be performed using some other technique of. In fact, the hardware used is standard digital circuits, analog circuits, any known processor capable of running software and / or firmware programs, programmable read-only memory (PROM), and programs. Using one or more programmable digital devices or systems, such as Possible Array Logical Devices (PALs), and any other known or deployed techniques, such as any combination of the above, and others. , Can be implemented. In addition, the methods of the invention can be embodied in any known or unfolded medium-storable software program.
In Figure 12, Action 200 creates multiple images (each with a different zoom level or resolution) and blends two or more images to achieve a smooth navigation look, such as zooming (Action). 206), it is a figure which shows the embodiment of this invention. Although not required to practice the present invention, the techniques shown in FIG. 12 are intended to be used in connection with the relationship between the service provider and the client. For example, a service provider creates multiple pre-rendered images (Action 200) and makes resources available to the user's client terminal over a communication channel such as the Internet (Action 202). Expand. Alternatively, the pre-rendered image can be a built-in or related part of an application program that the user loads and runs on his computer.
When using the blending method, if the image object is a road map, 30 meters / pixel, 50 meters / pixel, 75 meters / pixel, 100 meters / pixel, 200 meters / pixel, 300 meters / pixel, 500 meters / pixel. Experiments have shown that image sets work well at pixel, 1000m / pixel, and 3000m / pixel zoom levels. However, it should be noted that any number of images can be used at any number of resolutions without departing from the scope of the present invention. In fact, more or fewer images with a particular zoom level different from the previous example may cause other image objects to work optimally in other contexts.
Regardless of how the client terminal acquires the image, in response to a user-initiated navigation command such as a zooming command (action 204), the client terminal preferably creates an intermediate resolution image that matches the navigation command. It is possible to work to blend two or more images to do (Action 206). This blending has been incorporated herein by reference in its entirety (eg, Lance Williams, Pyramidal Parametrics, Computer Graphics, Proc. SIGGRAPH '83, 17 (3): 1-11 It can be performed by several methods, such as the well-known trilinear interpolation technique described in (1983). Other methods for interpolating images, such as bicubic linear interpolation, are also useful in connection with the present invention, and other methods may be developed in the future. It should be noted that the present invention does not require or depend on any particular one of these blending methods. For example, as shown in Figure 8, the user may want to navigate to a zoom level of 62.4 meters / pixel. This zoom level can be between two pre-rendered images (for example, between a zoom level of 50 meters / pixel and a zoom level of 75 meters / pixel in this example), with a desired zoom level of 62.4 meters. / Pixels can be achieved using trilinear interpolation techniques. In addition, it is possible to obtain any zoom level between 50 meters / pixel and 75 meters / pixel using the blending method described above, which is smooth and continuous if done immediately enough. You can get the appearance of the navigation. The blending technique can be performed up to other zoom levels, such as the 35.7 m / pixel level shown in Figure 9. In these cases, the blending technique can be performed as a pre-rendered image between 30 meters / pixel and 50 meters / pixel in the previous example.
In the blending technique described above, the computing power of the processing unit in which the present invention is executed performs a rendering operation on the first instance in order to achieve a high image frame rate for smooth navigation, and / Or (ii) Can be used if it is not high enough to perform image rendering "just-in-time" or "running" (eg in real time). However, as described below, other embodiments of the invention contemplate the use of known or deployed high performance processing units that can be rendered in client terminals for blending and / or high frame rate applications.
The process flow of FIG. 13 shows the detailed steps and / or actions performed by the present invention, preferably to create one or more images. Action 220 retrieves information about the image object (s) using any known or expanded technique below. These image objects have typically been modeled with appropriate primitives such as polygons, lines, and points. For example, if the image object is a road map, any Universal Transverse You can easily get a model of the road in the Mercator (UTM) zone. The model is usually in the form of a list of line segments (in any coordinate system) that include the roads in the zone. The list is any known or expanded rendering as long as it incorporates certain techniques for determining the weight (eg, apparent or actual thickness) of a given primitive within a pixel (spatial) domain. You can use a process to convert it to an image (pixel image) within a spatial domain. Continuing with the road map example above, the rendering process incorporates certain techniques for determining the weight of the lines that model the roads in the road map within the spatial domain. These techniques are described below.
Action 222 (Figure 13) classifies the elements of the object. For road map objects, this classification can take the form of recognizing existing categories: A1, A2, A3, A4, and A5. In fact, these road elements have different degrees of roughness and can be rendered differently based on this classification, as described below. In action 224, mathematical scaling is applied to different road elements based on the zoom level. Mathematical scaling may vary based on element classification, as described in more detail below.
As a background, there are two traditional techniques for rendering image elements such as map roads: real physical scaling and pixel width presetting. The real-physical scaling technique directs the road map to be rendered as if it were displaying a real-physical image of the road on different scales. For example, the A1 highway is 16 meters wide, the A2 road is 12 meters wide, the A3 road is 8 meters wide, the A4 road is 5 meters wide, and the A5 road is 2.5 meters wide. it can. This is acceptable to the viewer when zooming in on a small area of the map, but when zooming out, all major and non-major roads are too thin to be seen. At some zoom level, such as the state level (eg about 200 meters / pixel), all roads are completely invisible.
Pixel width presetting techniques indicate that every road has a certain pixel width, such as one pixel width on the display. Major roads, such as highways, can be emphasized, for example, by making them 2 pixels wide. Unfortunately, this technique changes the visual density of the map as you zoom in and out. At some zoom level, for example a small group level, this result is satisfactory. However, when zoomed in, the roads are not dark and the entire map looks sparse. Further zooming out causes the roads to mix with each other, quickly forming tight nests that make individual roads indistinguishable.
According to one or more aspects of the invention, in action 224, the image is (i) physically proportional to the zoom level, or (ii) zoom, depending on the parameters described in more detail below. Created so that at least some image elements are scaled up and / or scaled, either non-physically proportional to the level.
The fact that scaling is "physically proportional to the zoom level" means that if the size of the element appears to change with distance from the human eye, the number of pixels representing the road width increases or decreases with the zoom level. Note that this means. Given the perspective formula, the apparent length y of an object of physical size d, y = c d / x Where c is a constant that determines angled perspective and x is the distance of the object from the viewer.
In the present invention, the linear size of an object of physical linear size d'at display pixel p is p = d' z<sup>a</sup>Where z is the zoom level in units of physical linear size / pixel (eg meters / pixel) and a is the power law. For a = -1 and d'= d (the actual physical linear size of the object), this equation is dimensionally correct and equals the perspective formula at p = y and z = x / c. This represents the equality between physical zooming and fluoroscopic transformation, zooming in is equivalent to moving the object closer to the viewer, and zooming out is equivalent to moving the object away.
To perform non-physical scaling, a can be set to an exponent other than -1 and d'can be set to a physical linear size other than the actual physical linear size d. In the context of a road map, p can represent the width of the road displayed in pixels, d'can represent the imputed width of the physical unit, and is "non-physically proportional to the zoom level". Means that the road width in display pixels increases or decreases with the zoom level, i.e. a -1, in a way other than physically proportional to the zoom level. Scaling is distorted in that it achieves certain desired results.
Note that "linear size" means one-dimensional size. For example, assuming an arbitrary 2D object and doubling its "linear" size, the area is 4 = 2<sup>2</sup>Only increase. In two dimensions, the linear size of an object's elements can include length, width, radius, diameter, and / or any other dimension that can be read by the ruler on the Euclidean plane. Line thickness, line length, circle or disk diameter, one side length of a polygon, and the distance between two points are all examples of linear size. In this sense, the two-dimensional "linear size" is the distance between two identified points of an object on the 2D Euclidean plane. For example, the linear size is (dx<sup>2</sup>+ dy<sup>2</sup>) Can be calculated by taking the square root, and in this equation dx = x1-x0, dy = y1-y0, and the two identified points are obtained by Cartesian coordinates (x0, y0) and (x1, y1). ..
The concept of "linear size" is inevitably extended to dimensions above 2 dimensions, for example assuming a volume object, doubling its linear size means that the volume is 8 = 2.<sup>3</sup>Means only increase. Similar linear size surveys can be defined for non-Euclidean spaces such as the surface of a sphere.
Any exponent a <0 reduces the render size of the element on zoom out and increases it on zoom in. If a <-1, the render size of the element will decrease faster than the proportional physical scaling at zoom out. Conversely, if -1 <a <0, the size of the rendered element decreases slower than the proportional physical scaling when zooming out.
According to at least one aspect of the invention, for a given length of a given object, p, so that the user does not experience sudden jumps or discontinuities in the element size of the image during navigation. It is possible to make (z) nearly continuous (as opposed to the most extreme discontinuities, that is, the sudden appearance or disappearance of elements during navigation). In addition, p (z) accompanies zooming out so that the element of the object becomes smaller as it zooms out (for example, the road becomes narrower) and the element of the object becomes larger as it zooms in. It is preferable that the amount decreases monotonically. This gives the user a sense of physicality about the objects in the image.
The aforementioned scaling features can be more fully understood by looking at Figure 14, which is a log-log graph of metric / pixel zoom levels relative to the pixel-by-pixel rendered line width for the A1 highway. .. (Plotting logarithm (z) on the x-axis and logarithm (p) on the y-axis plots log (x)<sup>a</sup>) = A It is convenient because the plot becomes a straight line due to the relationship of log (x). ) The basic features of the zoom level plot for line (road) width are: (i) When zoomed in (eg up to about 0.5 meters / pixel), road width scaling can be physically proportional to the zoom level. (ii) When zoomed out (eg above about 0.5 meters / pixel), road width scaling can be non-physically proportional to zoom level, and (iii) When further zoomed out (eg, above about 50 meters / pixel, depending on the parameters described in more detail below), road width scaling can be physically proportional to the zoom level. That there, Is.
For zones where road width scaling is physically proportional to zoom level, p = d' z<sup>a</sup>The scaling formula of is adopted, and a = -1 in this formula. In this example, a reasonable value for the physical width of the actual A1 highway is about d'= 16 meters. Therefore, when zoomed out to at least a certain zoom level z0, for example z0 = 0.5 m / pixel, the rendering width of the line representing the A1 highway decreases monotonically with physical scaling.
The zoom level of z0 = 0.5 is chosen to be the internal scale below which the physical scaling is applied. This avoids the non-physical appearance of road maps when combined with other detail-scale GIS content of actual physical dimensions. In this example, z0 = 0.5 meters / pixel, or 2 pixels / meter, corresponding to a scale of approximately 1: 2600 when represented as a map scale on a 15-inch display (1600 x 1200 pixel resolution). At d = 16 meters, which is a reasonable physical width for A1 roads, the rendered road will appear to be its actual size when zoomed in (0.5 meters / pixel or less). At a zoom level of 0.1 meters / pixel, the rendered line is about 160 pixels wide. At a zoom level of 0.5 meters / pixel, the rendered line is 32 pixels wide.
For zones where road width scaling is non-physically proportional to zoom level, p = d' z<sup>a</sup>The scaling formula of is adopted, in which -1 <a <0 (between zoom levels z0 and z1). In this example, non-physical scaling is performed between approximately z0 = 0.5 m / pixel and z1 = 3300 m / pixel. Again, if -1 <a <0, the width of the rendered road decreases more slowly during zoom-out than proportional physical scaling. This advantageously keeps the A1 road visible (and distinguishable from other smaller roads) when zooming out. For example, as shown in Figure 5, A1 road 102 is kept visible at zoom level z = 334 meters / pixel and can be distinguished from other roads. Assuming the physical width of the A1 road is d'= d = 16 meters, the width of the line rendered using physical scaling is about 0.005 pixels at a zoom level of about 3300 meters / pixel, which is virtually invisible. Render like this. However, using non-physical scaling of -1 <a <0 (in this example, a is about -0.473), the width of the rendered line is about 0.8 pixels at a zoom level of 3300 meters / pixel, which is clearly visible. Render like this.
Note that the value of z1 is chosen so that the given road is still on the most zoomed out scale with "greater than reality" importance. For example, if the entire United States is rendered on a 1600 x 1200 pixel display, the resolution would be approximately 3300 meters / pixel, or 3.3 kilometers / pixel. Looking at the world as a whole, there is no reason to assume that US highways will be more important than national highways.
Therefore, at zoom levels above z1, which is about 3300 meters / pixel in the above example, road width scaling is physically proportional to zoom level, but preferably has a large d'on p (z) continuity. (Much larger than the actual width d). In this zone, p = d' z<sup>a</sup>The scaling formula of is used, and a = -1 in this equation. A new imputed physical width of the A1 highway, eg d'= 1.65 kilometers, is selected to make the road width rendered at z1 = 3300 meters / pixel continuous. The new values for z1 and d'are preferably chosen so that the line rendering width is a reasonable number of pixels on the outer scale z1. In this case, at a zoom level (3300 meters / pixel) where the entire country can be seen on the display, the A1 road can be about 1/2 pixel wide, which is narrow but still clearly visible, which is 1650 meters. That is, it corresponds to the imputed physical road width of 1.65 kilometers.
Based on the above, we propose the following specific set of equations for the line width rendered as a function of zoom level. p (z) = d0 z<sup>-1</sup>, Z z0 p (z) = d1 z<sup>a</sup>, Z0 <z <z1 p (z) = d2 z<sup>-1</sup>, Z z1 The form of p (z) has six parameters, z0, z1, d0, d1, d2, and a. z0 and z1 mark the scale on which p (z) changes. In the zoom-in zone (z z0), zooming is physical (ie, the exponent of z is -1) and the physical width is d0, preferably corresponding to the actual physical width d. Zooming is physical even in the zoom-out zone (z z1), but the physical width is d1 and generally does not correspond to d. Between z0 and z1, the rendered line width scales with an exponent of a that can be a value other than -1. Given consecutive preferences for p (z), it is sufficient to specify z0, z1, d0, and d2 to uniquely determine d1 and a, which is clearly shown in Figure 14. Has been done.
The above method for A1 roads is also applicable to other road elements of road map objects. An example of applying these scaling techniques to roads A1, A2, A3, A4, and A5 is shown in the log-log graph in Figure 15. In this example, z0 = 0.5m / pixel for all roads, but it can vary between elements depending on the context. The A2 road is generally slightly smaller than the A1 road, with d0 = 12 meters. In addition, A2 roads are "important" at the US state level, for example, so z1 = 312 meters / pixel, which is approximately the rendering resolution of a single state (about 1/10 of the country on a linear scale). At this scale, we know that a line width of 1 pixel is desirable, so d2 = 312 meters is a reasonable setting.
The general techniques for roads A1 and A2 outlined above can be used to establish the parameters of the remaining elements of the road map object. For A3 roads, d0 = 8 meters, z0 = 0.5 meters / pixel, z1 = 50 meters / pixel, and d2 = 100 meters. For A4 streets, d0 = 5 meters, z0 = 0.5 meters / pixel, z1 = 20 meters / pixel, and d2 = 20 meters. Furthermore, for A5 dirt roads, d0 = 2.5 meters, z0 = 0.5 meters / pixel, z1 = 20 meters / pixel, and d2 = 20 meters. Note that using this parameter setting, the A5 dirt road becomes more like a street as you zoom out of the zoom level, but when zoomed in, it is half the width on the physical scale. I want to.
The logarithmic-logarithmic plot in Figure 15 is an overview of scaling behavior by road type. Note that on all scales the apparent width is A1> A2> A3> A4> = A5. Also note that all indices occur near a = -0.41 except for dirt roads. All dashed lines have a gradient of -1 and indicate physical scaling of different physical widths. From top to bottom, the corresponding physical widths of these dashed lines are 1.65 kilometers, 312 meters, 100 meters, 20 meters, 16 meters, 12 meters, 8 meters, 5 meters, and 2.5 meters.
When interpolation between multiple pre-rendered images is used, in many cases all lines or others with the correct pixel width determined by the resulting interpolation and physical and non-physical scaling formulas. It can be guaranteed that the ideal rendering of the basic geometric elements of is indistinguishable or almost indistinguishable to the human eye. To understand this alternative embodiment of the invention, some background on drawing antialiased lines is presented below.
The description of drawing the dealiased line is presented according to the road map example described in detail above, where all the basic elements are lines and the line width follows the scaling formula described above. See Figure 16A, where a 1-pixel wide vertical line is on a white background so that the horizontal position of the line is exactly aligned with the pixel grid and simply consists of a 1-pixel wide black pixel column on a white background. It is drawn in black. According to various aspects of the present invention, it is desirable to consider and deal with the case where the line width is a non-integer pixel. As shown in Figure 16B, the endpoints of the line remain fixed, but when the line weight is increased to 1.5 pixels wide, the dealiased graphics display has 25% of the pixel columns to the left and right of the center column. It is drawn in gray. Referring to Figure 16C, these side columns are drawn in 50% gray with a width of 2 pixels. Referring to FIG. 16D, at 3 pixels wide, the side columns are 100% black, resulting in three solid black columns, as expected.
This technique for drawing non-integer-width lines on a pixelated display results in a sense of visual continuity (or illusion) with changes in line width, even with a small number of pixels. Even if the widths are different, the lines with different widths can be clearly distinguished. This technique is commonly known as dealiased line drawing, where the line integral of the intensity function over the perpendicular to the drawn line (or the "1 intensity" function for a black line on a white background) is the line width. Designed to guarantee equality to. This method is easily generalized for lines whose endpoints are not exactly centered on the pixel, lines in non-vertical orientations, and curves.
Drawing the dealiased vertical lines in Figures 16A-D is alpha blending two images, one (image A) with a line 1 pixel wide and the other (image B) with a line 3 pixels wide. Note that it can also be done by (-blending). Alpha blending is (1-alpha) for each pixel on the display.<sup>*</sup>(Corresponding pixel in image A) + alpha<sup>*</sup>Assign (corresponding pixel in image B). As the alpha changes between zero and one, the effective width of the rendered line changes smoothly between one and three pixels. This alpha blending technique, in the most common case, produces good visual results only if the difference between the two rendered line widths in images A and B is less than 1 pixel, otherwise it does not. , Lines may have haloed in the middle width. This same technique can be applied to render points, polygons, and any other basic graphical element of different linear sizes.
Returning to FIGS. 16A-D, the 1.5 pixel wide line (Fig. 16B) and the 2 pixel wide line (Fig. 16C) are the 1 pixel wide line (Fig. 16A) and the 3 pixel wide line (Fig. 16D). It can be constructed by alpha blending between. With reference to FIGS. 17A-C, a 1-pixel wide line (FIG. 17A), a 2-pixel wide line (FIG. 17B), and a 3-pixel wide line (FIG. 17C) are shown in any orientation. The same principle that lines are exactly aligned with the pixel grid applies to this arbitrary orientation in Figures 17A-C, but for good results, the spacing between the alpha-blended line widths is 2. Must be smaller than a pixel.
In the context of this map example, you can select different sets of images for pre-rendering by referring to the log-log plots in Figures 14-15. For example, we then refer to FIG. 18, which is similar to FIG. 14, except that FIG. 18 contains a set of horizontal and vertical lines. The horizontal line indicates the line width between pixels 1 to 10 that is incremented by 1 pixel. The vertical lines are spaced so that the line width over the spacing between two adjacent vertical lines varies by 2 pixels or less. Therefore, vertical lines represent a set of zoom values suitable for pre-rendering, and alpha blending between two adjacent thus pre-rendered images is almost always the case when rendering a line representing a road with continuously varying widths. It will produce equivalent features.
The interpolation between the six resolutions represented by the vertical lines shown in Figure 18 is sufficient to accurately render the A1 highway using scaling curves shown at approximately 9 meters / pixel and beyond. Is. Rendering less than about 9 meters / pixel renders these displays in a vector rather than interpolating between pre-rendered images, as these displays are so zoomed in and the number of roads displayed is very small. , Computationally more efficient (and more efficient with respect to data storage requirements) and does not require pre-rendering. At resolutions above about 1000 meters / pixel (such displays include large parts of the Earth's surface), rendering uses lines that are 1 pixel wide, so only the final pre-rendered image is used. It is possible. Lines thinner than a single pixel render the same pixel more faintly. Therefore, to create an image with an A1 line 0.5 pixel wide, you can multiply the 1 pixel wide line image by 0.5 alpha.
In fact, a slightly larger set of resolutions is pre-rendered so that there is no scaling curve in Figure 15 that varies beyond one pixel over each interval between resolutions. By reducing the permissible changes to one pixel, the rendering quality can be improved as a result. In particular, the "SYSTEM AND METHOD FOR EXACT RENDERING IN A ZOOMING USER" filed on March 1, 2004, which is intended and described (eg, the entire disclosure of which is incorporated herein by reference). US Patent Application No. 10/790253, entitled "INTERFACE" (see Reference No. 489/2)) Tiling techniques can be considered in the context of the present invention. This tiling technique can be used to resolve an image at a particular zoom level, even if its level does not match the pre-rendered image. If each image in a slightly larger resolution set was pre-rendered and tiled at the appropriate resolution, the result would appear to be a continuous variation in the width of all lines according to the scaling formulas disclosed herein. , A complete system for zooming and panning through road maps of any complexity.
Additional details regarding other techniques for blending images that can be used in connection with the practice of the present invention (eg, the entire disclosure of which is incorporated herein by reference, filed June 5, 2003. See US Provisional Patent Application No. 60/475897 entitled "SYSTEM AND METHOD FOR THE EFFICIENT, DYNAMIC AND CONTINUOUS DISPLAY OF MULTI RESOLUTIONAL VISUAL DATA"). Other details regarding blending techniques that can be used in connection with the practice of the present invention (eg, "SYSTEM AND METHOD FOR FOVEATED," filed March 12, 2003, the entire disclosure of which is incorporated herein by reference. See US Provisional Patent Application No. 60/453897) entitled "SEAMLESS, PROGRESSIVE RENDERING IN A ZOOMING USER INTERFACE").
Advantageously, using the aforementioned aspects of the present invention, the user can enjoy the appearance of smooth and continuous navigation at various zoom levels. Moreover, when zooming from one level to another, the details rarely appear or disappear abruptly or at all. This represents a significant advance in cutting-edge technology.
It is contemplated that various aspects of the invention are applicable to a number of products, including interactive software applications over the Internet, automotive software applications, and the like. For example, the present invention can be used on an internet website that provides a map and travel direction to a client terminal in response to a user request. Alternatively, various aspects of the invention can also be used in in-vehicle GPS navigation systems. The present invention can also be incorporated into medical imaging devices, which allows detailed information about, for example, the patient's circulatory system, nervous system, etc. to be rendered and navigated as described above. The scope of the invention is numerous and cannot be enumerated in its entirety, but those skilled in the art will appreciate that they fall within the scope of the invention as intended and claimed herein. ..
The present invention can also be used in connection with other scopes where the rendered image provides a means for advertising and other advanced transactions. Additional details relating to these aspects and uses of the invention (named "METHOD AND APPARATUS FOR EMPLOYING MAPPING TECHNIQUES TO ADVANCE COMMERCE" filed herein and with additional details incorporated herein by reference in its entirety. See also US Provisional Patent Application No. 60/553803 (see reference number 489/7).
An appendix is attached to this document. This appendix is part of the disclosure herein.
Although the present invention has been described herein with reference to specific embodiments, it will be appreciated that these embodiments merely exemplify the principles and scope of the invention. Accordingly, in the exemplary embodiments, a number of modifications can be made without departing from the spirit and scope of the invention as defined by the appended claims, and other arrangements can be devised. Let's understand that there is.
(appendix) By BLAISE HILARY AGUERA Y ARCAS Patent application for systems and methods for accurate rendering in zooming user interfaces Kaplan & Gilman, LLP Agent reference number 489/2
(Document name) Statement (Title of the Invention) Systems and methods for accurate rendering in zooming user interfaces
(Technical field) This application is filed in US Provisional Application No. 60/452075 filed on March 5, 2003, US Provisional Application No. 60/453897 filed on March 12, 2003, and US Provisional Application No. 60 filed on June 5, 2003. It claims the priority of / 475897 and US Provisional Application No. 60/474313 filed May 30, 2003.
The present invention generally relates to a graphical zooming user interface (ZUI) for a computer. More specifically, the present invention is a computationally efficient method that results in good user responsiveness and interactive frame rates, as well as usually degradation of image quality. Zoomable, both in an accurate way in the sense that vector plots, text, and other non-graphic content are eventually rendered, without leading resampling, and without interpolation of other images that also lead to degradation. A system and method for progressively rendering various visual contents.
(Background technology) Most modern graphical computer user interfaces (GUIs) are designed to use fixed space scale visual components. However, with the advent of the computer graphics field, it has been found that visual components can be represented and manipulated in such a way that they do not have a fixed spatial scale on the display but can be zoomed in or out. Zoomable components are desirable, to name a few, for viewing maps, browsing large heterogeneous text layouts such as newspapers, displaying digital photo albums, and visualizing large datasets. It is clear in many application examples. Even when viewing regular documents such as spreadsheets and reports, it is often useful to have a quick overview of the document and then be able to zoom in on the area of interest. Many current computer applications include Microsoft® and other Office® products (Zoom under the View menu), Adobe® Photoshop®, Adobe®. Includes zoomable components such as Acrobat®, QuarkXPress®, etc. In most cases, these applications can zoom in and out on the document, but the visual components of the application itself do not necessarily zoom in and out. Continuous panning across a document is standard (ie, use scrollbars or cursors to convert the displayed document to left, right, up, or down), but continuously in a user-friendly way. The ability to zoom and pan is not available in prior art systems.
First, some definitions will be explained. A display is a device used to output a rendered image to a user. Framebuffers are used to dynamically represent the content of at least part of the display. The display refresh rate is the rate at which the physical display or part of it is refreshed using the contents of the framebuffer. The frame rate of the framebuffer is the rate at which the framebuffer is updated.
For example, in a normal personal computer, the display refresh speed is 60 to 90 Hz. For example, most digital video has a frame rate of 24-30Hz. Therefore, each frame of digital video will actually be displayed at least twice while the display is refreshed. Since multiple framebuffers can be used at different frame rates, they can be displayed almost simultaneously on the same display. For example, this would happen if two digital videos with different frame rates were playing in different windows on the same display.
One of the problems with the Zooming User Interface (ZUI) is that the visual content must be displayed in different resolutions when the user zooms. The ideal solution to this problem would be to display an accurate, newly calculated image based on the underlying visual content in every successive frame. The problem with these techniques is that it is computationally infeasible to accurately recalculate each resolution of the visual content in real time when the user zooms, if the underlying visual content is complex.
As a result of the above, many prior art ZUI systems use multiple precomputed images, each representing the same visual content but with different resolutions. We refer to each of these different precomputed images as a level of detail (LOD). The complete set of LODs, conceptually organized as a stack of images with decreasing resolution, shown in Figure 1 is called the LOD Pyramid. In these traditional systems, when zooming occurs, the system interpolates between LODs and displays the resulting image at the desired resolution. Although this technique solves computational problems, it often involves a loss of information due to the fact that blurry, unrealistic final eclectic images are often displayed, representing interpolation of different LODs. These interpolation errors are especially noticeable when the user has the opportunity to stop zooming and display a still image at a selected resolution that does not exactly match the resolution of any LOD.
Another problem with interpolation between precomputed LODs is that this technique typically treats vector data in the same way as photographic or image data. Vector data, such as blueprints or line art, is displayed by processing a set of abstract instructions using a rendering algorithm that can render lines, curves, and other basic shapes at any desired resolution. Text rendered using scalable fonts is an important special case of vector data. Image or photo data (including text rendered using bitmap fonts) is not generated that way, but is displayed either by interpolation between precomputed LODs or by resampling the original image. It must be. The present inventors refer to the latter as non-vector data in the present specification.
Conventional systems that use rendering algorithms to redisplay vector data at the new resolution of each frame during a zoom sequence must limit the system itself to simple vector drawing only to achieve interactive frame rates. It doesn't become. On the other hand, prior art systems that pre-calculate LODs for vector data and interpolate between them are visually prone to interpolation errors, especially at the sharp edges that are typical of most vector data renderings. The quality deteriorates significantly. This degradation is usually unacceptable for text content, which is a special case of vector data.
(Disclosure of Invention) (Problems to be solved by the invention) An object of the present invention is to create a ZUI that replicates the zooming effect that a user would feel if they actually saw a physical object and moved it closer to them.
An object of the present invention is to create a ZUI that displays an image at an appropriate resolution to avoid or eliminate interpolation errors in the final displayed image. Another object of the present invention is to allow the user to optionally zoom in considerably within the vector content while maintaining a clear, unblurred view and maintaining the interactive frame rate.
Another object of the present invention is to optionally zoom out until the user gets the appearance of complex vector content while performing both maintaining an overall overview of the content and maintaining the interactive frame rate. To be able to do it.
Another object of the present invention is to prevent the user from perceiving transitions between LODs or rendering qualities during interaction.
Another object of the present invention is to allow a safe degradation of image quality due to blurring if the information normally required to render a portion of the image is still incomplete.
Another object of the present invention is to gradually improve image quality by making the focus clearer as more complete information needed to render a portion of the image becomes available. is there.
An object of the present invention is to render both vector and non-vector data optimally and separately.
These and other objects of the invention will become apparent to those skilled in the art by reviewing the following specification.
(Means to solve the problem) The aforementioned and other issues with prior art allow users to view images at dynamically changing resolutions when zooming in or out, panning, or otherwise changing their image view. It is overcome according to the present invention with respect to the hybrid strategy for implementing ZUI. These changes in the view are called navigation. Zooming the image to a resolution different from any predefined LOD is accomplished by displaying the image at a new resolution interpolated from the predefined LOD that "surrounds" the desired resolution. "Surrounding LOD" means the lowest resolution LOD higher than the desired resolution and the highest resolution LOD lower than the desired resolution. If the desired resolution is higher than the resolution of the LOD with the highest available resolution or lower than the resolution of the LOD with the lowest available resolution, then there will be only a single "surrounding LOD". In this document, dynamic interpolation of an image at a desired resolution based on a precomputed set of LODs is referred to as mip mapping or trilinear interpolation. The latter term indicates that bilinear sampling is used to resample the surrounding LODs, followed by linear interpolation (ie, trilinear) between these resampled LODs. (See, for example, Lance Williams, Pyramidal Parametrics, Computer Graphics (Proc. SIGGRAPH '83) 17 (3): 1-11 (1983), the entire disclosure of which is incorporated herein by reference.) By Williams. An obvious modification or extension to the introduced mip mapping technique uses resampling and / or interpolation of the enclosing LOD. In the present invention, it is not important whether the resampling and interpolation operations are zero-order (closest-order), linear, higher-order, or more generally non-linear.
According to the invention described herein, if the user defines an exact desired resolution that is rarely at the resolution of one of the predefined LODs, the final image is preferably first intermediate final. It is displayed by displaying the image. The intermediate final image is a first image that is displayed at the desired resolution before the image is improved, as will be described later. The intermediate final image may correspond to an image that will be displayed at the desired resolution using prior art.
In a preferred embodiment, the transition from the intermediate final image to the final image can be gradual, as described in more detail below.
In an enhanced embodiment, the invention includes an unreasonable increment (ie, an expansion or contraction factor between consecutive LODs that cannot be expressed as a ratio of two integers), as described in more detail below. LODs can be placed at intervals in any resolution increment.
In other enhanced embodiments, parts of the image in each different LOD are shown as tiles, which are rendered to minimize any imperfections perceived by the user. In other embodiments, the displayed visual content consists of multiple LODs (potentially a superset of surrounding LODs), each of which gradually turns the display into a final image to hide imperfections. Appears in the appropriate proportions and location for fading in.
Rendering of various tiles in multiple LODs can tolerate computational complexity so that the system can run on standard computers at the typical clock speeds available on most laptop and desktop personal computers. It is implemented to optimize the appearance of the visual content while keeping it within a reasonable level.
The present invention uses a predefined LOD for fast zooming and panning, but a mixed strategy that renders and displays the correct LOD if the view is stable enough. including. Accurate LODs are rendered and displayed at a precise resolution chosen by the user, which is usually different from predefined LODs. This mixed strategy creates a continuous "perfect rendering" illusion with much less computation, as the human visual system is not sensitive to the details of the visual content while the movement is stationary. be able to.
(Best mode for carrying out the invention) FIG. 2 is a flow chart showing a basic technique for carrying out the present invention. The flow diagram of FIG. 2 represents an exemplary embodiment of the invention, which begins execution when the image is displayed at an initial resolution. Although the present invention can be used in the client / server model, it should be noted that the client and server may be on the same machine or on different machines. So, for example, a set of discrete LODs may be stored remotely on the host computer side, and the user can connect to this host via a local PC. The actual hardware platform and system used is not important to the present invention.
The flow diagram begins at block 201 with an initial view of the image at a particular resolution. In this example, the image is considered static. The image is displayed in block 202. The user can navigate the image, for example by moving the computer mouse. The initial view displayed in block 202 will change as the user navigates the image. It should be noted that the underlying image can be dynamic in itself, as in the case of motion video, but for illustration purposes here the image itself is treated as static. As mentioned above, any image to be displayed can also have text or other vector data and / or non-vector data such as photographs and other images. The present invention and the entire description below are applicable regardless of whether the image contains vector data, non-vector data, or both.
Regardless of the type of visual content displayed in block 202, this method transfers control to decision point 203 where navigation input may be detected. If no such input is detected, the method loops back to block 202 and continues to display static visual content. If a navigation input is detected, control will be transferred to block 204, as shown in the figure.
Decision point 203 can be implemented by a continuous loop in software looking for a particular signal to detect motion, an interrupt system in hardware, or any other desired method. The particular technique used to detect and analyze navigation requirements is not important to the present invention. Regardless of the method used, the system can detect a request indicating a desire to navigate the image. Although most of the discussion herein relates to zooming, it should be noted that the techniques are also applicable for zooming, panning, or otherwise navigating. In fact, the techniques described herein are applicable to any type of dynamic transformation or modification in the way the image is viewed. Such deformations may indicate, for example, 3D transformation and rotation, image filtering, local stretching, dynamic spatial distortion applied to selected areas of the image, or more information. Any other type of strain can be included. Another example is a virtual magnifying glass that can be moved over an image, magnifying a portion of the image underneath this virtual magnifying glass. If it is detected at decision point 203 that the user has started navigating, block 204 may render and display a new view of the image, which may have a different resolution than the previously displayed view, for example. it can.
One of the simple prior art techniques for displaying a new view is based on LOD interpolation when the user zooms in or out. The LODs selected can be two LODs that "surround" the desired resolution, the resolution of the new view. In traditional systems, interpolation is performed continuously while the user zooms, often directly in hardware to achieve speed. The combination of motion detection at decision point 205 and the near-immediate display of a properly interpolated image at block 204 results in a continuous zoom of the image as the user navigates. .. The interpolated image looks realistic and clear enough because the image moves while zooming in or out. Interpolation errors are minimally detected by the human visual system because they are hidden by a constantly changing view of the image.
At decision point 205, the system tests whether movement has almost stopped. This can be done using a variety of techniques, including measuring the rate of change of one or more parameters of the view, for example. That is, this method checks whether or not the user has reached the point where the zooming ends. When such stabilization is confirmed at decision point 205, control is transferred to block 206, and control is returned to block 203 after the exact image has been rendered. In this way, the system will eventually display an accurate LOD at any desired resolution.
In particular, the display is not simply rendered and displayed by interpolating two predefined LODs, but the original used to render text or other vector data when the initial view was displayed in block 202. It can be rendered and displayed by re-rendering the vector data using an algorithm. Non-vector data can also be resampled for rendering and display with the exact requested LOD. The requested re-rendering or resampling can be performed at the exact resolution required for display at the desired resolution, as well as precisely at the correct position of the display pixels for the underlying content calculated based on the desired view. It can also be run on the corresponding sampling grid. As an example, converting images every 1/2 pixel on the display surface on the display does not change the required resolution, but changes the sampling grid, which requires accurate LOD re-rendering or resampling.
The aforementioned system in Figure 2 uses predefined LOD-based interpolation while the view is changing (eg, navigation is occurring), but the exact view is when the view is nearly stationary. Represents a mixed method of being rendered and displayed in.
As used herein, the term rendering refers to the computerized generation of tiles in a particular LOD based on vector or non-vector data. For non-vector data, it is possible to render at any resolution by resampling the original image at a higher or lower resolution.
Next, we proceed to how to render and display different parts of the visual content needed to achieve the exact final image represented by block 206 in Figure 2. Referring to FIG. 3, if it is determined that navigation has stopped, control is transferred to block 303 and the interpolated image is displayed immediately, just as it was during zooming. This interpolated image, which can be temporarily displayed after the navigation is stopped, is referred to by the present inventors as an intermediate final image, or simply an intermediate image. This image is generated from the interpolation of the surrounding LOD. In some cases, the intermediate image can be interpolated from more than two discrete LODs, or from two non-discrete LODs surrounding the desired resolution, as described in more detail below.
When the intermediate image is displayed, it proceeds to block 304, causing the image to begin a gradual fade in the exact rendering direction of the image, which we call the final image. The final image differs from the intermediate image in that the final image may not contain any predefined LOD interpolation. Alternatively, the final image or part thereof can include newly rendered tiles. For photographic data, the newly rendered tiles can result from resampling of the original data, and for vector data, the newly rendered tiles can result from rasterization at the desired resolution.
Also note that you can skip directly from blocks 303 to 305 and immediately replace the interpolated image with the final exact image. However, in a preferred embodiment, step 304 is performed so that the transition from the intermediate final image to the final image is carried out progressively and smoothly. This gradual fading, sometimes referred to as blending, progressively focuses the image when navigation is stopped, producing the same effect as autofocus in a camera or other optical instrument. The illusion physically created by this effect is an important aspect of the present invention.
Next, we describe how this fading or blending can be performed to minimize the perceived irregularities, sudden changes, seams, and other imperfections in the image. However, it will be appreciated that the particular technique of fading is not important to the present invention and that many variants will be apparent to those skilled in the art.
Different LODs have different numbers of samples per physical area of the underlying visual content. Therefore, the first LOD can utilize a 1 inch x 1 inch area displayable object to generate a single 32 x 32 sample tile. However, the information utilizes the same 1 "x 1" area, which can also be represented as 64 x 64 samples, and thus higher resolution tiles.
The present inventors define a concept called irrational tiling. The tiling subdivision that we write as the variable g is defined as the ratio of the linear tiling grid size in the high resolution LOD to the linear tiling grid size in the next low resolution LOD. In Williams's treatise, which introduces trilinear interpolation, g = 2. In the prior art, this same value of g has been used. LODs can be subdivided into tiles in any manner, but in an exemplary embodiment, each LOD is subdivided into a grid of square or rectangular tiles containing a fixed number of samples (excluding edges of visual content as needed). To be made. Conceptually, for g = 2, each tile at a given LOD is then "split" into 2x2 = 4 tiles at a higher resolution LOD, as shown in Figure 4 (again). Potentially excluding edges).
Granulality 2 tiling has fundamental drawbacks. Typically, when the user zooms in at a random point in the tile, the increase in zoom by g times is a single additional tile corresponding to the next high resolution LOD near the point where the user's zoom is heading. Will need to be rendered. However, if the user is zooming in on a grid line in the tiling grid, two new tiles, one on each side of the line, will need to be rendered. Finally, if the user is zooming in on the intersection of the two grid lines, four new tiles will need to be rendered. If these events, which require 1, 2, or 4 new tiles per gx zoom, are randomly scattered throughout the extended zooming sequence, the overall performance remains the same. However, grid lines in any integer subdivision tiling (ie, g is an integer) retain the grid lines in any high resolution LOD.
For example, consider zooming in on the center of a very large image tiling with subdivision 2. The present inventors set the (x, y) coordinates of this point to (1/2, 1/2), and the visual content has corners (0,0), (0,1), (1,0). , And adopt the rule that it fits within the square of (1,1). Since this center is the intersection of two grid lines, each time the user reaches each high resolution LOD, four new tiles must be rendered, so zooming at this particular point is faster. Is slow and inefficient. On the other hand, suppose the user zooms in on an unreasonable point, a grid point (x, y) where x and y cannot be represented as a ratio of two integers. Examples of such numbers are pi (= 3.14159 ...) and the square root of 2 (= 1.414213 ...). Therefore, it is easy to demonstrate that the sequence of 1, 2, and 4 obtained by the number of tiles required to be rendered for each gx zoom is quasi-random, that is, does not follow a periodic pattern. it can. It is clear that this kind of pseudo-random sequence is more desirable from a performance standpoint, and from a performance standpoint there are no significant points regarding zooming.
Forced tiling solves this problem and is interpreted as g itself being an irrational number, usually the square root of 3, 5, or 12. This means that an average of 3, 5, or 12 tiles in a given LOD (correspondingly) will be contained within a single tile in the next low resolution LOD, but in consecutive LODs. Note that the tiling grid of is no longer "matching" on any grid line in this scheme (potentially at the tip of the visual content, x = 0 and y = 0, or something along each axis. Except for other preselected single grid lines). If g is chosen so that it is not the nth root of any integer (pi is such a number), the LOD does not share any grid lines (again, potentially x = 0 and y = 0). except). Therefore, it can be seen that each tile can randomly overlap with 1, 2, or 4 tiles in the next lower LOD, but for g = 2, this number is always 1.
For forced tiling subdivision, zooming in at any point will generate a pseudo-random stream requesting 1, 2, or 4 tiles, and performance will be uniform on average no matter where you zoom in. Perhaps the greatest advantage of forced tiling is related to panning after a large zoom. If the user pans the image after zooming in too much, at some point the grid lines will move onto the display. This is usually the case when the area opposite this grid line corresponds to a lower resolution LOD than the rest of the area on the display, but it is desirable that these resolution differences be as small as possible. However, for the integer g, this difference is often quite large, as grid lines can overlap over many consecutive LODs. This creates a "deep crack" in the resolution over the entire node area, as shown in FIG. 6 (a).
On the other hand, grid lines in forced tiling never overlap with adjacent LOD grid lines (again, one grid line in each direction that may be in one corner of the image may be excluded. ), There is no discontinuity in the resolution of multiple LODs. This increase in smoothness at relative resolution makes the illusion of spatial continuity more compelling.
Figure 6 (b) shows the advantages gained by forced tiling subdivision. FIG. 6 shows a cross section through several LODs of visual content, with each bar representing a cross section of a rectangular tile. Therefore, the second level from the top, which has two bars, can be an LOD of 2x2 = 4 tiles. The top-to-bottom curve 601 represents the boundaries of the visible area of the visual content in the relevant LOD during the zooming operation, and as the resolution increases (zooms in for more detail), the area under investigation. Decreases. Dark bars (eg 602) represent tiles that have already been rendered during zooming. Light-colored bars haven't been rendered yet and can't be displayed. When the tiling is an integer, as in Figure 6 (a), sudden changes in resolution across the space are common, and when the user pans to the right after zooming, there are four spatial boundaries indicated by the arrows. Note that the LOD suddenly "quits". The resulting image will appear sharp on the left side of this border and will be extremely blurry on the right side. The same visual content, represented using forced tiling subdivision, does not have these resolution "cracks" and adjacent LODs do not share tile boundaries, except as shown on the far left. Computationally, this shared boundary can occur at most one on the x-axis and one on the y-axis. In the illustrated embodiment, these shared boundaries are located at y = 0 and x = 0, but can be located at any other location if present.
Another advantage of force tiling subdivision is that finer control of g is possible, especially in the whole useful range where g is not so large, because there are significantly more irrational numbers than integers. This additional freedom helps to tune the zooming performance of certain applications. If g is set to the unreasonable square root of an integer (sqrt (2), sqrt (5), sqrt (8)), in the above embodiment, every other LOD grid line is accurately aligned and g If is an unreasonable cube root, every other LOD is accurately aligned, and so on. This gives an additional benefit in limiting the complexity of compound tiling as defined below.
An important aspect of the invention is the order in which the tiles are rendered. More specifically, different tiles in different LODs are rendered optimally so that all visible tiles are rendered first. Invisible tiles may not render at all. Within the visible tile set, rendering proceeds in increasing resolution order, so the tiles in the low resolution LOD are rendered first. Within any particular LOD, tiles are rendered in increasing distance from the center of the display, which we refer to as fossa rendering. Heapsort, quicksort, and many other sorting algorithms can be used to sort these tiles in the order described above. To implement this ordering, it is possible to use the lexigraphic key to sort the "requests" to render the tile, the outer subkey is visibility, the middle subkey. Is the resolution of the sample per physical unit, and the inner subkey is the distance to the center of the display. Other methods for ordering tile rendering requests can also be used. The actual rendering of tiles is optimally performed as a process parallel to the navigation and display described herein. When rendering and navigation / display proceed as a parallel process, the user's responsiveness can remain high, even if the tiles render slowly.
Next, the tile rendering process in the exemplary embodiment will be described. If the tile represents vector data, such as alphabetic typography for stroke-based fonts, rendering the tile involves running an algorithm to rasterize the alphabetic data and, in some cases, transmitting that data from the server to the client. Is done. Alternatively, the data supplied to the rasterization algorithm is sent to the client, which can execute the algorithm for rasterizing the tiles. In another example, rendering a tile containing relevant digitally sampled photographic data involves resampling that data to generate the tile with the appropriate LOD. With pre-stored discrete LODs, rendering may simply involve transmitting tiles to the client computer for subsequent display. Tiles between discrete LODs, such as tiles in the final image, may require some additional calculations as described above.
Whenever tiles are rendered and the image begins to fade to the correct image, different combinations of different tiles from different LODs can be included in the actual display. Therefore, any part of the display can contain, for example, LOD 1 to 20%, LOD 2 to 40%, and LOD 3 to 40%. Regardless of the tiles displayed, the algorithm attempts to render tiles from various LODs with the most favorable priority to supply the rendered tiles for the most needed display. The actual display of rendered tiles will be described in more detail below with reference to Figure 5.
The following describes a method for drawing multiple LODs using an algorithm that can guarantee the spatial and temporal continuity of the image details. This algorithm uses high resolution tiles in preference to low resolution tiles that cover the same display area, while spatially blending to avoid sharp borders between LODs, as well as when enabled ( That is, when high resolution tiles are rendered), all rendered tiles are designed to be optimally used, using temporally gradual weight blending for more detailed blending. Unlike prior art, this algorithm and its variants result in more than two LODs blending together at a given point on the display, resulting in a blending factor that changes smoothly over the entire display area. It is possible to generate a blending coefficient that evolves even after the user has stopped navigating. Nevertheless, this exemplary embodiment is computationally efficient and renders a partially transparent image, or an image in which the overall transparency varies over the entire image area, as will be apparent below. Can be used if.
This specification defines a composite tile area, or simply a composite tile. To define composite tiles, we consider all LODs stacked on top of each other. Each LOD has its own tile grid. The composite grid is then formed by projecting all the grids from all the LODs onto a single plane. The composite grid is then composed of various composite tiles of different sizes, defined by the boundaries of the tiles from all the different LODs. This is conceptually shown in Figure 7. Figure 7 shows tiles from three different LODs, 701 to 703, all representing the same image. LOD Imagine 701-703 stacked on top of each other. In these cases, if the corners 750 of each of these LODs were aligned in a row and stacked on top of each other, the area indicated by 740 would be inside the area indicated by 730, at 730 and 740. The area indicated will be inside the area indicated by 720. Region 710 in FIG. 7 indicates the existence of a single composite tile 710. Each composite tile is inspected frame by frame, and the frame rate can typically be greater than 10 frames per second. Note that, as mentioned earlier, this frame rate is not necessarily the refresh rate of the display.
FIG. 5 is a flow diagram showing an algorithm for updating the framebuffer as the tiles are rendered. The placement configuration in Figure 5 is intended to work on any composite tile in the displayed image each time the framebuffer is updated. So, for example, frame duration (frame) If the duration) is 1/20 second, then each composite tile on the entire screen will preferably be inspected and updated every 1/20 second. If a composite tile is manipulated based on the process in Figure 5, the composite tile may lack associated tiles in one or more LODs. The process in Figure 5 attempts to display each composite tile as a weighted average of all available overlapping tiles in which the composite tile is present. Note that the weighted average can be expressed as a relative ratio of each LOD, as composite tiles are defined in such a way that they fit exactly within one tile at any given LOD. This process determines the appropriate weights for each LOD in the composite tile and gradually changes those weights over space and time to progressively fade the image towards the final image described above. Try to get it done.
A composite grid contains multiple vertices defined to be any intersections or corners of grid lines within the composite grid. These are called composite grid vertices. We define the opacity of each LOD at each composite grid vertex. This opacity can be expressed as a weight between 0.0 and 1.0, so for an image where the desired result is completely opaque, the sum of all LOD weights at each vertex should be 1.0. Is. The current weight at any particular point in time for each LOD at each vertex is maintained in memory.
The algorithm for updating the weights of the vertices proceeds as follows.
The following variables, which take a number between 0.0 and 1.0, centerOpacity, cornerOpacity at each corner (4 if the tiling is a rectangular grid), and edgeOpacity at each edge (4 if the tiling is a rectangular grid) are: It is kept in the memory of each tile. When a tile is rendered first, all its opacity, as listed, is usually set to 1.0.
During the drawing pass, the algorithm starts with the highest resolution LOD and goes through compound tiling once for each associated LOD. In addition to this per-tile variable, the algorithm also maintains the variables levelOpacityGrid and opacityGrid. Again, both of these variables are numbers between 0.0 and 1.0 and are maintained for each vertex in the composite tiling.
The algorithm passes through each LOD in order from highest resolution to lowest resolution to perform the following operations: First, 0.0 is assigned to the levelOpacityGrid at all vertices. Then, for each rendered tile in that LOD (which may be a subset of the tile set for that LOD if some have not yet been rendered), the algorithm is set to the centerOpacity, cornerOpacity, and edgeOpacity values of the tile. Based on this, update a part of the levelOpacityGrid that touches the tile as follows.
If the entire vertex is inside the tile, it will be updated using centerOpacity.
If the vertices are, for example, on the left edge of the tile, they will be updated with the left edgeOpacity.
If the vertex is in the upper right corner, for example, it will be updated with the upper right cornerOpacity.
"Update" means to set the new value to the lowest current value, or the value used for the update, if the existing levelOpacityGrid value is greater than 0.0. If the existing value is zero (ie this vertex has not been touched yet), just set the levelOpacityGrid value to the value used for the update. The end result is that the levelOpacityGrid at each vertex position will be set to the lowest nonzero value used for the update.
The algorithm then passes through the levelOpacityGrid and sets any vertex that touches an unrendered tile called a hole to 0.0. This guarantees the spatial continuity of the blending, and in the current LOD, whenever a composite tile fits inside a hole, the opacity drawing should fade to zero at all vertices bordering the hole. is there.
In the extended embodiment, the algorithm relaxes all levelOpacityGrid values to further improve the spatial continuity of LOD blending. The above situation can be visualized so that every vertex is like a tent pole and the levelOpacityGrid value at that point is the height of the tent pole. This algorithm makes it possible that the height of the tent pole at all adjacent points on the hole is zero and that inside the rendered tile the tent pole is set to some (possibly) non-zero value. Guarantee. In the extreme case, probably all values inside the rendered tile are set to 1.0. For illustrative purposes, the border value is 0.0, assuming that the neighbors of the rendered tile have not yet been rendered. It has not yet been determined how narrow the "margin" is between the 0.0 border tent pole and one of the 1.0 internal tent poles. If this margin is too small, the transition can be too rapid when measured as an opacity derivative over the entire space, even if the blending is technically continuous. The relaxation operation smoothes the tent and always holds a value of 0.0, but in some cases lowers the other tent poles to make the function defined by the tent surface smoother, i.e. its maximum spatial derivative (maximum). spacial derivative) is restricted. This is not important to the present invention in which the various methods are used to perform this operation, one approach is to use, for example, selective low band filtering, leaving zeros and any nonzero values. Replace locally with the weighted average of its neighbors. Other methods will be apparent to those skilled in the art.
The algorithm then passes through all the composite grid vertices and considers the corresponding values of levelOpacityGrid and opacityGrid at each vertex, and if levelOpacityGrid is greater than 1.0-opacityGrid, levelOpacityGrid is set to 1.0-opacityGrid. Then, for each vertex again, the corresponding value of levelOpacityGrid is added to the opacityGrid. Due to the steps above, this will not cause the opacityGrid to be above 1.0. These steps in the algorithm ensure that the higher resolution LOD contributes to as much opacity as possible if available, so that the lower resolution LOD "sees through" only when there are holes. can do.
The final step in scanning the current LOD is to actually draw the composite tile in the current LOD, using the levelOpacityGrid as the per-vertex opacity value. In the extended embodiment, it is possible to multiply the levelOpacityGrid with a scalar overallOpacity variable in the range 0.0 to 1.0 just before drawing, which allows the entire image to be drawn with the partial transparency given by overallOpacity. it can. Note that drawing an image-containing polygon, such as a rectangle, with different opacity at each vertex is a standard procedure. For example, it can be done using industry standard texture mapping features that use OpenGL or Direct3D graphics libraries. In reality, the opacity drawn inside each of these polygons is spatially interpolated, resulting in a smooth change in opacity across the polygon.
In other extended embodiments of the algorithm described above, tiles not only maintain their current values (called current values) for centerOpacity, cornerOpacity, and edgeOpacity, but also parallel values called targetCenterOpacity, targetCornerOpacity, and targetEdgeOpacity. Also maintain the set (called the target value). In this extended embodiment, when the tiles are rendered first, the current values are all set to 0.0, but the target values are all set to 1.0. Then, after each frame, the current value is adjusted to a new value closer to the target value. This can be done using some formulas, but as an example, newValue = oldValue<sup>*</sup>(1-b) + targetValue<sup>*</sup>It is feasible with b, in which b is a rate greater than 0.0 and less than 1.0. A value of b close to 0.0 will result in a very slow transition to the target value, and a value of b close to 1.0 will result in a very fast transition to the target value. This method of opacity update results in exponential convergence towards the target, giving a visually pleasing impression of temporal continuity. The same result can be achieved with other equations.
The above description describes a preferred embodiment of the present invention. The present invention is not limited to these preferred embodiments, and various modifications that are consistent with the appended claims are also included in the present invention.
(A brief description of the drawing) (Figure 1) A diagram showing the LOD pyramid (in this case, the bottom of the pyramid representing the highest resolution representation is a 512 × 512 sample image, and the continuous reduction of this image is indicated by factor 2). (Fig. 2) is a flow chart for use in an exemplary embodiment of the present invention. (Figure 3) Another flow diagram showing how the system displays the final image after zooming. (Fig. 4) It is a figure which shows the LOD pyramid of FIG. 1 to which the grid line which shows the subdivision of each LOD is added to the rectangular tile which is the same size in the sample unit. (Fig. 5) Another flow diagram for use in connection with the present invention, showing the process for displaying rendered tiles on a display. (Fig. 6) Fig. 6 is a diagram showing a concept called unreasonable tiling, which is explained in more detail in the present specification. (Fig. 7) Fig. 7 is a diagram showing composite tiles and tiles constituting composite tiles, which are described in more detail in the present specification.
<img file="JP4861978B2_D0001.tif" />
<img file="JP4861978B2_D0002.tif" />
<img file="JP4861978B2_D0003.tif" />
<img file="JP4861978B2_D0004.tif" />
<img file="JP4861978B2_D0005.tif" />
<img file="JP4861978B2_D0006.tif" />
<img file="JP4861978B2_D0007.tif" />
(Document name) Statement (Title of the Invention) Systems and methods for pitted, seamless, progressive rendering in zooming user interfaces
(Technical field) The present invention generally relates to a zooming user interface (ZUI) for a computer. More specifically, the present invention is a system and method for progressively rendering arbitrary large or complex visual content in a zooming environment while maintaining excellent user responsiveness and high frame rate. In some situations, rendering quality needs to be temporarily degraded to meet these goals, but the present invention leverages the well-known properties of the human visual system to reduce this degradation. Mask on a large scale.
(Background technology) Most modern graphical computer user interfaces (GUIs) are designed to use fixed space scale visual components. However, with the advent of the computer graphics field, it has been found that visual components can be represented and manipulated in such a way that they do not have a fixed spatial scale on the display but can be zoomed in or out. Zoomable components are desirable, to name a few, for viewing maps, browsing large heterogeneous text layouts such as newspapers, displaying digital photo albums, and visualizing large datasets. It is clear in many application examples. Even when viewing regular documents such as spreadsheets and reports, it is often useful to have a quick overview of the document and then be able to zoom in on the area of interest. Many current computer applications include Microsoft® and other Office® products (Zoom under the View menu), Adobe® Photoshop®, Adobe®. Includes zoomable components such as Acrobat®, QuarkXPress®, etc. In most cases, these applications can zoom in and out on the document, but the visual components of the application itself do not necessarily zoom in and out. In addition, zooming is usually not a very important aspect of user interaction with the software, and zoom settings are only modified from time to time. Continuous panning across a document is standard (ie, use scrollbars or cursors to convert the displayed document to left, right, up, or down), but almost without exception, zooms continuously. There is no function. A more generalized zooming framework allows you to zoom any kind of visual content, zooming with panning. It will be part of an equally important user experience. Ideas along these lines appeared in many films as early as the 1960s as an ultra-modern computer user interface (for example, Stanley Kubrick's 2001: A Space Odyssey, Turner Entertainment, Time Warner. (1968)), the trend continues to recent films (eg Steven Spielberg's "Minority Report", 20th Century Fox and DreamWorks Pictures (2002)). From the 1970s to the present, several continuous zooming interfaces have been devised and / or developed (early appearances were WC Donelson, Spatial Management of Information, Proceedings of Computer Graphics SIGGRAPH (1978), ACM Press, p.203-9. A recent example is Zanvas.com), which was launched in the summer of 2002. In 1991, some of these ideas were formally approved (eg, Ken Perlin and Jacob Schwartz At New York University (Fractal Computer User Centerface with Zooming Capability), see US Pat. No. 5,341,466. ). Zooming developed by Perlin and its collaborators com). In 1991, some of these ideas were formally approved (eg, Ken Perlin and Jacob Schwartz At New York University (Fractal Computer User Centerface with Zooming Capability), see US Pat. No. 5,341,466. ). Zooming developed by Perlin and its collaborators com). In 1991, some of these ideas were formally approved (eg, Ken Perlin and Jacob Schwartz At New York University (Fractal Computer User Centerface with Zooming Capability), see US Pat. No. 5,341,466. ). Zooming developed by Perlin and its collaborators com). In 1991, some of these ideas were formally approved (eg, Ken Perlin and Jacob Schwartz At New York University (Fractal Computer User Centerface with Zooming Capability), see US Pat. No. 5,341,466. ). Zooming developed by Perlin and its collaborators com). In 1991, some of these ideas were formally approved (eg, Ken Perlin and Jacob Schwartz At New York University (Fractal Computer User Centerface with Zooming Capability), see US Pat. No. 5,341,466. ). Zooming developed by Perlin and its collaborators 203-9. A recent example is Zanvas.com), which started in the summer of 2002. In 1991, some of these ideas were formally approved (eg, Ken Perlin and Jacob Schwartz At New York University (Fractal Computer User Centerface with Zooming Capability), see US Pat. No. 5,341,466. ). Zooming developed by Perlin and its collaborators 203-9. A recent example is Zanvas.com), which started in the summer of 2002. In 1991, some of these ideas were formally approved (eg, Ken Perlin and Jacob Schwartz At New York University (Fractal Computer User Centerface with Zooming Capability), see US Pat. No. 5,341,466. ). Zooming developed by Perlin and its collaboratorsPad, the prototype of the user interface, and its successor, Pad ++, have seen some development since then (Perlin explains the subsequent development at http://mrl.nyu.edu/projects/zuni/). doing.).
(Disclosure of Invention) (Problems to be solved by the invention) However, as far as we know, major applications based on all ZUIs (Zooming User Interfaces) are not yet in large numbers, partly due to some technology deficiencies, and the present invention addresses one of those technology deficiencies. ..
(Means to solve the problem) The present invention embodies a novel idea on which the newly developed zooming user interface framework (hereinafter referred to as Voss from its practical name) is based. Voss is more powerful, more responsive, more visually compelling, and more versatile than its predecessors, due to some innovations in its software architecture. The patent relates to Voss techniques specifically for object tiling, level-of-detail blending, and rendering queuing.
Multi-resolution visual objects are typically rendered from a discrete set of images sampled at different resolutions or levels of detail (image pyramids). In some technical contexts where continuous zooming is used, such as 3D gaming, two adjacent levels of detail surrounding the desired level of detail are blended to render each frame, which is usually desired. This is not the case where the level of detail is exactly one of the discrete sets. These techniques are sometimes referred to as trilinear filtering or mipmapping. In most cases, the mip-mapped image pyramid is pre-created and continuously maintained in short-term memory (ie RAM) during the zooming operation, so that any required level of detail is always available. In some enhanced 3D rendering scenarios, the image pyramid must itself be rendered within an animation loop, but in these cases the complexity of this first rendering pass is such that it does not affect the overall frame rate. You have to control it well.
In this context, it is desirable to be able to continuously navigate by zooming and panning through any visual complexity of unlimited amounts of content. This content cannot be rendered quickly and is not immediately available, but may require downloading from a remote location over a low bandwidth connection. Therefore, it is not always possible to render the level of detail at a frame rate (first pass) that is comparable to the desired display frame rate (second pass). In addition, it is not possible to keep a pre-created image pyramid for all content in memory, the image pyramid must be rendered or re-rendered as needed, and this rendering will have the desired frame rate. May be slower than.
The present invention is based solely on strategies for prioritizing the rendering (potentially slow) of the image pyramid portion associated with the current display and partial information, i.e., the currently available subset of the image pyramid. It includes both strategies for presenting the user with a smooth and continuous perception of the rendered content. These strategies are combined to take advantage of available computing power or bandwidth in a near-optimal manner, masking any image degradation caused by an imperfect image pyramid to the extent possible. Spatial and temporal blending is utilized to avoid discontinuities or sudden changes in image sharpness.
An object of the present invention is to allow sampled (i.e., pixelated) visual content to be rendered within a zooming user interface without the final image quality degradation of traditional trilinear interpolation.
Another object of the present invention is to allow arbitrary large or complex visual content to be displayed within a zooming user interface.
Another object of the invention is any complex visual content, even if this content is ultimately represented using a large amount of data, and these data are stored in remote locations. It is to be able to display almost immediately even if it is shared via a low bandwidth network.
Another object of the present invention is to allow a user to optionally zoom in significantly on visual content while maintaining an interactive frame rate.
Another object of the present invention is to allow the user to optionally zoom out significantly to get an overview of complex visual content, both in the process of preserving the appearance of the entire content and in the process of maintaining the interactive frame rate. Is to do.
Another object of the present invention is to minimize the transition between levels of detail or rendering quality that the user perceives during the interaction.
Another object of the present invention is continuous blurring if detailed visual content is so far unavailable because the information required for rendering is unavailable or because rendering is still in progress. By (blurring), only the image quality is allowed to be degraded (graceful).
Another object of the present invention is to safely improve image quality by gradually sharpening when rendering of a portion of visual content is first available.
These and other objects of the present invention will become apparent to those skilled in the art from the review described herein below.
Prior art: multi-resolution image and zooming user interface From a technical point of view, the zooming user interface is a generalization of the usual concepts based on visual computing, allowing it to overcome some limitations inherent in traditional user / computer / document interaction models. It is a thing. One of these limitations is the size of a document that can be "opened" from a computer application, and traditionally had to "load" the entire document before it could be viewed or edited. Transfer all document information from some repository (eg from a hard disk or over a network) to short-term memory at the time of opening, even if a particular computer has a large amount of short-term memory (usually RAM) available. Because of this, limited bandwidth makes the delay from issuing an "open" command to the start of viewing or editing unbearably long, and is aware of this limitation.
Still digital images provide both a prime example of this problem and an example of how the computer science community has surpassed the standard model for visual computing in overcoming this problem. Table 1 below shows different bandwidths for typical compression sizes for a variety of different image types, from the smallest useful images (thumbnails sometimes used as icons) to the largest images commonly used today. Indicates the download time. The shaded frame indicates an image size that makes interactive browsing difficult or impossible at a particular connection speed.
<tables num="1"><img file="JP4861978B2_D0008.tif" /></tables>
Almost every image currently on the web is less than 100K (0.1MB), as most users connect to the web with a DSL or lower bandwidth and larger images will take longer to download. is there. Even with local settings, it is rare to encounter images larger than 500K (0.5MB) on a typical user's hard drive. So many illustrated books, road map books, maps, newspapers, and works of art in the average home can quickly grow to tens of megabytes when digitized at maximum resolution. The fact that it contains an image of is often demonstrated to be useful for that large (ie, more detailed) image.
Until a few years ago, the lack of large-scale images was due to the lack of storage space in the repository, but due to improvements in hard drive technology, simplification of CDROM burning, and increasing popularity of large-scale network servers, repository space has become available. It is no longer a limiting factor. The main bottleneck today is bandwidth, followed by short-term memory (ie RAM) space. This problem is actually much worse than what was suggested by the table above, and in most contexts users are interested in displaying an entire set of images, not just a single image. Therefore, if the image is larger than the medium size, it is impractical to wait while downloading the images one after another.
Current image compression standards, such as JPEG2000 (http://www.jpeg.org/JPEG2000.html), are designed to address this issue exactly. Rather than storing image content in a linear fashion (ie, typically from top to bottom and left to right, the entire pixel in a single path), it is based on multi-resolution decomposition. The image is first resized to a resolution scale hierarchy, usually half, for example a 512x512 pixel image is 256x256 pixels, 128x128, 64x64, 32x32, 16x16, Resized to 8x8, 4x4, 2x2, and 1x1. Obviously, the details are captured only in high resolution and broad, which uses a fairly small amount of information. strokes) are captured in low resolution. This is why images of different sizes are called level of detail, or LOD for short. At first glance, this set of storage requirements for images of different sizes may appear to be higher than the requirements for high-resolution images alone, but in reality it is not, and low-resolution images are the next higher-resolution image. Acts as a "predictor" of the image. This makes it possible to encode the entire image hierarchy much more efficiently than is normally possible with a non-hierarchical representation of only high resolution images.
Imagine that a series of multiple resolution versions of an image are stored in the repository in ascending order of size, usually when the image is transferred across a data link into the cache, the user is at a lower resolution for the entire image. You will be able to get an overview, after which the more detailed parts will be "filled" as the transmission progresses. This is called incremental or progressive transmission. When done properly, any image, no matter how large, even if the bandwidth of the connection to the repository is medium, almost instantly across the space (though not all details). It has the characteristic of being displayable. The final time required to download all detailed images is the same, but the order in which this information is sent has been changed so that the large features of the image are transmitted first, which is It is much more useful to the user than transmitting pixel information in "reading order", from top to bottom and from left to right, as well as all details.
Hidden in this improvement is a new concept of what it means to "open" an image that does not fit into the traditional application model described above. Now imagine the notion that the user can view the image while downloading it, and its usefulness stems from the fact that a rough portion of the image is available immediately after the download begins, perhaps well before the download is complete. View. Therefore, it makes no sense for the application to keep the user waiting for the download to complete, instead the application should immediately display the possible part of the document and download the details "in the background". There are no delays or unnecessary interruptions in user interaction while continuing. This requires the application to perform multiple tasks at once, called multithreading. Modern web browsers use multithreading in a slightly different position to display the text layout of a web page, while maintaining user interaction while simultaneously downloading images on the web page. Please note. In this case, you can think of the embedded image itself as an additional level of detail that enhances the basic level of detail that consists of the minimum required text layout of the web page. This analogy will later prove its importance.
Obvious hierarchical image representation and progressive transmission of image documents are advances beyond linear representation and transmission. However, further progress is important if the image, at its highest level of detail, has more information (ie, more pixels) than the user's display can display at one time. With current display technology, this always applies to the bottom four images in Table 1, but smaller displays (such as PDA screens) may not be able to display even the bottom eight. This makes zooming a requirement for large images, and displaying images larger than the display is useless if it is not possible to zoom in to find additional details.
When the download of a large image is started, the user is probably looking at it as a whole. The first level of detail is often too coarse and the displayed image looks uneven or blurry, depending on the type of interpolation used to spread the small amount of available information over a large display area. It will be. The image will then be progressively refined, but at some point the display will be "saturated" with information and downloading any additional details will have no visual effect. Therefore, there is no point in continuing to download beyond that point. However, assuming the user decides to zoom in to see a particular area in more detail, the effective projection size of the image is significantly larger than the physical screen. So, in the download model mentioned above, higher levels of detail are increasing. It is necessary to download to order). The difficulty is that every level of detail contains about four times as much information as the previous level of detail, and as the user zooms in, the download process inevitably cannot keep up. To make matters worse, most of the downloaded information is wasted because it consists of high-resolution details outside the display area. It is clear that we need the ability to download only selected parts of a level of detail, that is, we should download only visible details. With this change, it is possible to create an image browsing system that can not only display images of any large size, but also efficiently navigate (ie, zoom and pan, etc.) these images at any level of detail.
The previous model of document access is sequential in nature, that is, the entire information object is transmitted in linear order. This model, on the other hand, is random access, that is, only selected parts of the information object are requested, and these requests can be made in any order, over an extended period of time, i.e. during a display session. Here, the computer and repository engage in extended dialogs that parallel the user's "dialogs" that appear on the display along with the document.
For efficient random access, it is convenient (although not absolutely necessary) to subdivide each level of detail into a grid so that grid squares or tiles are the basic unit of transmission. The pixel size of each tile can be maintained at or below a certain size so that each increasing level of detail contains about four times as many tiles as the previous level of detail. If that dimension may not be an exact multiple of the nominal tile detail, small tiles will occur at the edges of the image, and at the lowest level of detail, the entire image will be smaller than a single nominal tile. The resulting tiling image pyramid is shown in Figure 2. Note that the "vertices" of the pyramid, where the downscaled image is smaller than a single tile, look like the untiled image pyramid in Figure 1. The JPEG2000 image format contains all the same features described when representing tiled multi-resolution and random access images.
So far, we've only discussed static images, but with application-specific modifications, this same technique can be applied to almost any type of visual document. This includes, but is not limited to, large texts, maps, or mixed documents such as other vector graphics, spreadsheets, videos, and web pages. So far, we have implicitly considered display-only applications, that is, applications that only need to define actions or methods that correspond to open and draw. It is clear that other methods may be desirable, such as editing commands performed by a paint program for static images and editing commands performed by a word processor for text. Further consider the problem of text editing, where normal actions, such as inserting typed input, are relevant only over a range of spatial scales with respect to the underlying document. Interactive editing is no longer possible if the text is zoomed out to the point where it can no longer be read. You can also see that interactive editing is no longer possible when a single character zooms in enough to cover the entire screen. Therefore, the zooming user interface can also limit certain methods of actions to the level of detail associated with them.
If the visual document is represented internally as more abstract data, such as text, spreadsheet entries, or vector graphics, rather than as an image, generalize the tiling concept introduced in the previous section. It is necessary to. For still images, the process of rendering the tiles once acquired is self-evident, as the information (once stretched) is the exact pixel x pixel content of the tile. Further speed bottlenecks are usually the transfer (eg download) of compressed data to a computer. However, in some cases, the speed bottleneck lies in the rendering of the tiles, and the information used to perform the rendering is already stored locally or is so compact that there is no download delay. there is a possibility. Therefore, we then refer to the creation of a complete, fully drawn tile in response to a tile drawing request as a tile rendering, with the understanding that this may be a slow process. Whether this is slow is because the requested data is large and must be downloaded over a slow connection, or because the rendering process itself is computationally intensive. It is irrelevant.
A complete zooming user interface combines these ideas in such a way that users can view large, sometimes dynamic composite documents whose sub-documents are usually spatially non-overlapping. .. These subdocuments can further include that subdocument (usually non-overlapping), and so on. Documents therefore form a tree, where each document has a collection of subdocuments, a pointer to a child, each of which is contained within the spatial boundaries of the parent document. Each of these documents is called a node, borrowed from the programming terms for trees. Drawing methods are defined for all nodes at all levels of detail, but other methods for application-specific features can only be defined for a particular node, and their actions are at a particular level of detail. Can only be limited. Therefore, some nodes can be editable static images using commands such as painting, others are editable text, others are designed for viewing and clicking. It can be a web page that has been edited. All of these can co-exist within a "supernode", a common large spatial environment that can be navigated by zooming and panning.
Proper implementation of the zooming user interface has some direct consequences, including:
--Very large documents can be browsed as a whole without downloading them from the repository, limiting even documents that are larger than available short-term memory or otherwise exorbitant in size. Can be displayed without.
--Content is only downloaded during navigation as needed, resulting in optimal and efficient use of available bandwidth.
--Zoom and pan are spatially intuitive operations and can be organized in a way that makes it easy to understand large amounts of information.
--Since "screen space" is inherently unlimited, you don't have to minimize windows, use multiple desktops, or hide windows behind each other to work with multiple documents or views at the same time. Alternatively, the documents can be arranged and configured as desired, and the user can zoom out to see an overview of all of them, or zoom in to see a particular one. This does not preclude the possibility of rearranging and configuring the positions (or even scales) of these documents so that any combination of them can be seen at a useful scale on the screen at the same time. Also, it does not necessarily exclude combinations of zooming that use older methods.
-Since zooming is the original aspect of navigation, any kind of content can be displayed at an appropriate spatial scale.
--High resolution display no longer suggests shrinking text and images to smaller sizes (sometimes unreadable), allowing more content to be displayed at once, depending on the zoom level. Or allow the content to be displayed in normal size and higher fidelity.
--Visually impaired people can easily navigate the same content as people with normal eyesight simply by zooming in further.
These benefits are especially valuable now that the information available to regular computers connected to the web is skyrocketing. Ten years ago, very large documents of this kind that could be viewed in a ZUI were rare, and moreover, these documents took up a lot of space, so most computer-enabled repositories (eg 40MB). It rarely fits on the hard disk). However, we are now facing a very different situation where the server can easily store large amounts of documents and document hierarchies, and make this information available to any client connected to the web. Is. Nonetheless, the bandwidth of the connection between these potentially large repositories and regular users is much lower than the bandwidth of the connection to the local hard disk. This is exactly the scenario where ZUI offers its greatest advantage over traditional graphical user interfaces.
(Best mode for carrying out the invention) For a view of a particular node at a particular desired resolution, there are several tile sets that need to be rendered to include at least one sample per screen pixel in a particular LOD. Note that the view is usually not precisely displayed at the resolution of one of the node's LODs, but at an intermediate resolution between two of them. So ideally, in a zooming environment, the client will generate a set of visible tiles both just above and just below the actual resolution and use some interpolation to render the pixels on the display based on this information. To do. The most common scenario is linear interpolation, both spatially and between levels of detail, which is commonly referred to in the graphics literature as trilinear interpolation. Closely related techniques are commonly used in 3D graphics architectures for texturing (SL Tanimoto and T. Pavlidis, A hierarchical data structure for picture processing, Computer Graphics and Image Processing, Vol.4, p.104-119 (1975); Lance Willliams, Pyramidal Parametrics, ACM SIGGRAPH Conference Proceedings (1982)).
Unfortunately, tile downloads (or programmatic rendering) are often slow, and not all tiles you need are always available, especially during fast navigation. Therefore, the innovation in this patent is to present the viewer with a spatially and temporally continuous coherent image that approximates this ideal image in an environment where tile downloads or creations occur slowly and asynchronously. Focus on the combination of strategies.
In the following, we will use two variable names, f and g. f refers to the tile sampling density for the display defined in # 1. The tiling subdivision described as the variable g is defined as the ratio of the linear tiling grid size at one LOD to the linear tiling grid size at the next lower LOD. None of the innovations presented herein rely on the constant g, but it is generally presumed to be constant over different levels of detail for a given node. In the JPEG2000 example described in the previous section, g = 2, and conceptually each tile is "split" into 2x2 = 4 tiles with the next higher LOD. Subdivision 2 is by far the most common in similar applications, but in this context g can take other values.
1. Detail level tile request queuing. First, we introduce a system and method for queuing tile requests that allows clients to gradually "focus" composite images by analogy using optics.
When faced with irregular, and sometimes low-bandwidth connectivity issues to information repositories containing hierarchically tiled nodes, the question of how the zooming user interface requests tiles during navigation. Must be dealt with. In many situations, all of these requirements will be met in a timely manner, or throughout the time the information is relevant (ie, before the user zooms or pans elsewhere). It is unrealistic to assume that it will be. Therefore, it is desirable to intelligently prioritize tile requests.
The "outermost" rule for tile request queuing is an increasing level of detail for displays. This zoom-dependent "relevance level of detail" is obtained by the number f = (linear tile size in tile pixels) / (tile length projected onto the screen measured in screen pixels). When f = 1, the tile pixel to screen pixel is 1: 1 and when f = 10, the information in the tile is much more detailed than what the display can display (10).<sup>*</sup>If 10 = 100 tile pixels fit within a single screen pixel and f = 0.1, the tiles are coarse with respect to the display (every tile pixel is 10)<sup>*</sup>Must be "stretched" or interpolated to cover 10 = 100). This rule states that if the area of the display is undersampled (ie, only loosely defined) with respect to the rest of the display, the client's first priority is to fill this "resolution hole". Guarantee. If there are no multiple level of detail in the hole, then add the request for all level of detail in f <1 plus the request for the next higher level of detail (to allow blending of LODs, see # 5). Queued in ascending order. Strictly speaking, only the finest level of detail of these needs to be rendered in the current view, and the coarser level of detail is verbose in that it defines a low resolution image on the display, so at first glance , It is presumed that this leads to unnecessary overhead. However, these coarser levels cover a larger area, which is generally much larger than the display. In fact, the most crude level of detail for any node contains only a single tile due to its structure, so clients rendering any view of a node will always queue this "outermost" tile first. It will be put in.
This is an important point for the robustness of the display. Robustness means that the client never "losses" what is displayed in response to the user's pan and zoom, even if there is a large backlog of tile requests waiting to be fulfilled. The client simply displays the best (ie, highest resolution) image available in any area of the display. In the worst case, this will be the outermost tile, the first tile ever requested in relation to the node. Therefore, based solely on the first tile request, any spatial portion of the node can always be rendered, and all subsequent tile requests can be considered as incremental subdivision.
The overall effect is that the display may appear blurry after resizing or zooming, as the fallback on low resolution tiles gives the image a blurry impression. The image is then sharpened when the tile requirements are met.
Simple calculations show that the overhead generated by requesting "redundant" low-resolution tiles is actually small, especially at the cost of the robustness of properly defining node images everywhere from the beginning. Is small.
2. Cueing for foveated tile requests. Within the relevant level of detail, tile requests are queued by increasing the distance to the center of the screen, as shown in Figure 3. This technique is inspired by the human eye, which has a central region, a fossa, specialized for high resolution. Since zooming is usually associated with an interest in the central area of the display, queuing a fossa tile request usually reflects an implicit prioritization of visual information as the user zooms in. In addition, the blur remaining on the edges of the display is less noticeable than near the center, as the user's eyes generally spend more time looking at the area near the center than the edges of the display.
The temporary and relative increase in sharpness near the center of the display created by zooming in using the pitted tile request also reflects the natural consequences of zooming out, as shown in Figure 4. This figure shows two alternative "navigation paths", a single document that occupies approximately two-thirds of the display, assuming that the user can see it in very high resolution on the first line. Stay stationary while looking at (or node). Initially, the node content is represented by a single low resolution tile, then the next LOD tile is available so that the node content can be seen at twice the resolution using 4 (= 2x2) tiles. And followed by 4x4 = 16 and 8x8 = 64 tile versions. The second line looks at what happens when the user zooms in on the shaded square until the image displayed on the first line is fully refined. Tiles with a high level of detail are requeued, but in this case only tiles that are partially or fully visible. Refinement progresses to a point comparable to the first line (in terms of the number of tiles visible on the display). The third line shows what is available if the user zooms out again and how to fill in the lost details. Note that all levels of detail are displayed, but in practice very detailed levels are probably omitted from the display on the third line, as they represent more detail than the display can convey.
Note that zooming out usually leaves the center of the display filled with more detailed tiles than the edges. Therefore, this tile request order consistently prioritizes the sharpness of the central area of the display during all navigation.
Blending of LOD in time. Without further refinement, when the tiles needed for the current display are downloaded or built and drawn first, some of the underlying coarser tiles that are likely to represent the same content are immediately invisible. As such, the user experiences this transition as a sudden change in blurring in some areas of the display. These sudden transitions are annoying and unnecessarily draw the user's attention to the details of software implementation. A common technique for ZUI design by our inventors is to create a seamless visual experience for the user, thereby noting the existence of tiles or other aspects of the software that should remain "covered". You won't be attracted. Therefore tiles are not visible as soon as they are first available, but are blended into several frames, usually about 1 second. This blending function is linear (ie, the opacity of the new tile is a linear function of the time since the tile became available, such that the new tile is 50% opaque at half the fixed blend interval). , Exponential function, or any other interpolation function. In an exponential blend, every short constant time interval corresponds to a constant percentage change in opacity. For example, new tiles are 20% more opaque per frame, resulting in 20%, 36%, 49%, 59%, 67%, 74%, 79%, 83%, 87%, 89%, 91. A series of opacity over consecutive frames such as%, 93%, etc. occurs. Mathematically, the exponent never reaches 100%, but in reality the opacity is indistinguishable from 100% in a short period of time. Exponential blending has the advantage that the largest increase in opacity occurs near the start of the blend, so that new information is immediately invisible to the user while still maintaining acceptable temporal continuity. In the reference embodiments of the present invention, the illusion is that the area of the display is smoothly focused as the required information becomes available.
4. Continuous LOD. In situations where the download or creation of tiles lags behind the user's navigation, adjacent areas of the display may have different levels of detail. The previous innovation (# 3) addressed the issue of temporal discontinuity at the level of detail, but another innovation is needed to address the issue of spatial discontinuity at the level of detail. If these spatial discontinuities are not corrected, they will appear to the user as seams in the image such that the visual content is drawn more clearly on one side of the seams. The problem is that the opacity of each tile can be varied across the tile area, especially if an area on the display with a relatively low level of detail touches the tile edge. The solution is to have zero transparency. In some situations, it is also important that the opacity at each corner of the tile is zero if the corners touch areas with a relatively low level of detail.
FIG. 5 is the simplest reference embodiment of the present inventors on a method of decomposing each tile into rectangles and triangles called tile fragments so that the opacity changes continuously over each tile fragment. Is shown. Tiles X bounded by a square aceg have adjacent tiles L, R, T, and B on the left, right, top, and bottom, each sharing an edge. In addition, it also has adjacent parts TL, TR, BL, and BR that share a single corner. Suppose tile X exists. Its "inner square" iiii is completely opaque. (Note that repeated lowercase letters indicate the same vertex opacity value.) However, the opacity of the surrounding rectangular frame is determined by the presence (and complete opacity) of adjacent tiles. To. Therefore, in the absence of tile TL, point g is completely transparent, and in the absence of L, point h is completely transparent, and so on. The inventors of the present invention refer to the boundary region of the tile (iiii outside X) as a blending flap.
FIG. 6 shows a reference method used to interpolate the opacity of the entire debris. Part (a) shows a rectangle with constant opacity. Part (b) is a rectangle with two opposite edges with different opacity, and the overall opacity of the interior is simply linear interpolation based on the shortest distance from the two edges of each interior point. Part (c) shows a bilinear method for interpolating the opacity of the entire triangle when the opacity of all three corners abc is different. Conceptually every internal point p subdivides the triangle into three sub-triangles with regions A, B, and C, as shown in the figure. The opacity of p is simply the weighted sum of the corner opacity, and the weight is the fractional region of the three subtriangles (ie, A, B, and C divided by the total triangle region A + B + C). This is because this formula gives equal opacity for vertices when p moves to a vertex, and when p is on the edge of a triangle, the opacity is linear interpolation between two connected vertices. Therefore, it is easily verified.
This method is non-existent over the entire tiling surface, as the opacity within the debris is determined by the opacity at its vertices as a whole, and adjacent debris always shares the vertices (ie there is no T-joint). Guarantee that the transparency changes smoothly. This strategy, combined with # 3 temporal LOD blending, ensures that the relative level of detail visible to the user is a continuous function over the display area and in time. It avoids both spatial seams and temporal discontinuities and presents the user with a visual experience reminiscent of optics that results in a continuously focused scene. When navigating a large document, the speed at which the scene is focused is either the bandwidth function of the connection to the repository or the speed function of the tile rendering, whichever is slower. Finally, in combination with the fossa prioritization of innovation point # 2, successive levels of detail are biased in such a way that the central area of the display is first focused.
5. Generalized Linear-Mipmap-Linear LOD Blending. So far, we have discussed strategies and reference embodiments to ensure spatial and temporal smoothness in apparent LODs on nodes. However, we have not yet addressed how to blend levels of detail during successive zooming operations. The method used is a generalization of trilinear interpolation in which adjacent levels of detail are linearly blended over the entire intermediate region of the scale. At each level of detail, each tile debris is spatially averaged with adjacent tile debris of the same level of detail for spatial smoothness, and temporally averaged for smoothness over time, opacity at draw time. Have. If the level of detail is insufficiently sampling the display, i.e. f <1 (see # 1), the target opacity is 100%. However, if the display is oversampled, the target opacity is linearly (or using any other monotonic function) reduced to zero if the oversampling is g times. Similar to trilinear interpolation, this results in continuous blending through the zoom operation so that the perceived level of detail never changes suddenly. However, unlike traditional trilinear interpolation, which always involves blending two levels of detail, the number of levels of detail blended in this scheme can be one, two, or more. Numbers greater than two are temporary and are caused by tiles at multiple levels of detail that have not yet been fully blended in time. Single levels are also common in that lower than ideal LODs generally occur when they are "substitute" for 100% opacity of higher LODs that have not yet been downloaded or built and are not blended. Is temporary.
The simplest reference embodiment for rendering a set of tile debris for a node is to use the so-called "painter algorithm", where all tile debris are in order from back to front, i.e. the coarsest (minimum LOD). Rendered from to the most detailed (the highest LOD that oversamples the display below g times). All target opacity except the highest LOD is 100%, but if temporal blending is incomplete, it may be temporarily rendered with lower opacity. The highest LOD has variable opacity, which depends on how much the display is sampled, as mentioned above. It is clear that this reference embodiment is suboptimal in that it may render debris that is completely hidden by debris that is rendered later. A more optimal embodiment is possible by using data structures and algorithms that are similar to the data structures and algorithms used to remove surfaces hidden in 3D graphics.
6. Exercise prediction. It is especially difficult to keep asking for tile requirements during fast zoom or pan. Even during these high-speed navigation patterns, zoom or pan motion is properly predicted locally by linear extrapolation (ie, it is difficult to perform sudden inversions or reorientations). Tend. Therefore, this temporal motion coherence is utilized to generate slightly advanced tile requests and improve visual quality. This is done by making tile requests using a virtual viewport that stretches, expands, or contracts the direction of motion when panning or zooming, and anticipating requests for additional tiles. After navigation, the virtual viewport relaxes over a short time interval back to the actual viewport.
None of the above innovations are constrained by rectangular tiling, but on a grid such as triangular or hexagonal tiling, or heterogeneous tiling consisting of a mixture of these shapes, or completely arbitrary tiling. Note that it is explicitly generalized to any tiling pattern that can be defined. The only explicit change that needs to be made to address these alternative tilings is to define a tile-shaped triangulation similar to Figure 5, where the opacity at the edges and inside can all be controlled separately. Is.
<img file="JP4861978B2_D0009.tif" />
<img file="JP4861978B2_D0010.tif" />
<img file="JP4861978B2_D0011.tif" />
<img file="JP4861978B2_D0012.tif" />
<img file="JP4861978B2_D0013.tif" />
<img file="JP4861978B2_D0014.tif" />
(Document name) Statement (Title of the Invention) Systems and methods for efficient, dynamic, and continuous display of multi-resolution visual data
(Technical field) The present invention generally relates to multi-resolution images. More specifically, the present invention relates to systems and methods for efficiently blending together visual representations of content at different resolutions or levels of detail in real time. This method is perceptual continuity, even in highly dynamic contexts where the data to be visualized may be changing and only a portion of the data may be available at a given point in time. Guarantee. The present invention has applications in several areas, including, but not limited to, a zooming user interface (ZUI) for computers.
(Background technology) In many situations, including the display of complex visual data, these data are stored or calculated hierarchically as a set of representations at different levels of detail (LOD). Many multi-resolution methods and representations have been devised for different types of data, including (but not limited to) wavelets for digital images and progressive meshes for 3D models. .. Multi-resolution methods are also used in mathematical and physical simulations, where in some cases very long calculations can be performed more "coarsely" or "in more detail", and the present invention relates to such simulations and multi-resolution visuals. It also applies to other situations where the data can be generated interactively. Further, the present invention is applied in situations where visual data can be obtained "on the run" at different levels of detail, for example from cameras with machine-controllable pan and zoom. The present invention is a general purpose method for dynamically displaying such multi-resolution visual data on one or more 2D displays (such as a CRT or LCD screen).
In describing the present invention, as a main example, we will use wavelet decomposition of large-scale digital images (such as those used in the JPEG2000 image format). As its starting point, this decomposition employs original pixel data, which is usually an array of samples on a regular rectangular grid. Each sample typically represents the color or brightness measured at a point in the space corresponding to its grid coordinates. In some applications, the grid can be very large, for example tens of thousands of samples (pixels) on each side. Especially when the server (where the image is stored) is connected to the client (where the image will be displayed) by a low bandwidth connection and you want to browse these images remotely, you can interact with these large sizes. There are quite a few problems with displaying by type. If the image data is sent from the server to the client in a simple raster order, all the data must be transmitted before the client produces an overview of the entire image. This can take a long time. Generating such an overview is computationally expensive and may require downsampling, for example, a 20,000 x 20,000 pixel image to 500 x 500 pixels. Not only are these operations very slow to perform interactively, but they also require the client to have enough memory to store all the image data, in this example 1.2 gigabytes per 8-bit RGB color image. (GB) (= 3)<sup>*</sup>20,000^2)。
Almost every image currently on the web is less than 100K (0.1MB), as most users connect to the web with a DSL or lower bandwidth and larger images will take longer to download. is there. Even with local settings, it is rare to encounter images larger than 500K (0.5MB) on a typical user's hard drive. So many illustrated books, road map books, maps, newspapers, and works of art in the average home can quickly grow to tens of megabytes when digitized at maximum resolution. The fact that it contains an image of is often demonstrated to be useful for that large (ie, more detailed) image.
Until a few years ago, the lack of large-scale images was due to the lack of non-volatile storage space (repository space), but due to improvements in hard drive technology, simplification of CDROM burning, and increasing popularity of large-scale network servers. Repository space is no longer a limiting factor. The main bottleneck today is bandwidth, followed by short-term memory (ie RAM) space.
Current image compression standards, such as JPEG2000 (http://www.jpeg.org/JPEG2000.html), are designed to address this issue exactly. Rather than storing image content in a linear fashion (ie, typically from top to bottom and left to right, the entire pixel in a single path), it is based on multi-resolution decomposition. The image is first resized to a resolution scale hierarchy, usually half, for example a 512x512 pixel image is 256x256 pixels, 128x128, 64x64, 32x32, 16x16, Resized to 8x8, 4x4, 2x2, and 1x1. The present inventors refer to the coefficient of how much the size of each resolution differs from the next higher size as subdivision, which is 2 here, but this subdivision is represented by the variable g. Subdivision may vary with scale, for example, g is assumed here to be constant throughout the "image pyramid", but is not limited to this. Obviously, the details are captured only in high resolution, and the rough parts that use a fairly small amount of information are captured in low resolution. This is why images of different sizes are called level of detail, or LOD for short. At first glance, this set of storage requirements for images of different sizes may appear to be higher than the requirements for high-resolution images alone, but in reality it is not, and low-resolution images are the next higher-resolution image. Acts as a "predictor" of the image. This makes it possible to encode the entire image hierarchy much more efficiently than is normally possible with a non-hierarchical representation of only high resolution images.
Imagine that a series of multi-resolution versions of an image are stored in the server's repository in ascending order of size, usually when the image is transferred from the server to the client, the user has a low resolution overview of the entire image. Will be available, after which the more detailed parts will be "filled" as the transmission progresses. This is called incremental or progressive transmission and is one of the major advantages of multi-resolution representation. If progressive transmission is done properly, the client can (almost instantly) fill the entire space of any image, no matter how large, even if the bandwidth of the connection to the server is medium. It has the property of being displayable (although not all details). The final time required to download all detailed images is the same, but the order in which this information is sent has been changed so that the large features of the image are transmitted first, which is It is much more useful to the client than transmitting pixel information in full detail as well as top-to-bottom and left-to-right "reading order".
For efficient random access in dynamic and interactive contexts, it is (not absolutely necessary) to subdivide each level of detail into a grid so that grid squares or tiles are the basic unit of transmission. However, it is convenient. The pixel size of each tile can be maintained at or below a certain size so that each increasing level of detail contains about four times as many tiles as the previous level of detail. If that dimension may not be an exact multiple of the nominal tile detail, small tiles will occur at the edges of the image, and at the lowest level of detail, the entire image will be smaller than a single nominal tile. Therefore, assuming 64x64 pixel tiles, the 512x512 pixel image described above has 8x8 tiles at the highest level of detail, 4x4 at 256x256 level, 128x128 level. It has 2x2, and a single tile at the remaining level of detail. The JPEG2000 image format includes the aforementioned features for displaying tiled multi-resolution and random access images.
When the details of a large tiled JPEG 2000 image are displayed interactively by the client on a 2D display with limited size and resolution, some particular set of adjacent tiles may have certain details. It is necessary to generate accurate renderings at the level. However, in this dynamic context, all of these may not be available. However, coarser level of detail tiles are often available, especially if the user starts with an extensive overview of the image. Since tiles with a coarser level of detail spatially spread over a larger area, it is likely that the entire area will be covered by some combination of available tiles. This suggests that the available image resolution is not constant over the entire display area.
In a previously filed provisional patent application, the inventor proposed a method for "fading out" the edges of tiles that touch blank spaces at the same level of detail, thereby "subjecting" at the level of detail. Sudden visual discontinuities in sharpness are avoided, which would otherwise result if the "range" is incomplete. The edge area of the tile reserved for blending is called the blending flap. A simple reference embodiment for displaying the completed composite image is the "painter algorithm", where all relevant tiles at the crudest level of detail (ie, tiles with overlapping display areas) are drawn first, followed by All relevant tiles at a progressively more detailed level of detail are drawn. As mentioned above, blending was applied to the edges of the incomplete area at each level of detail. The result, as desired, is that the coarser level of detail is "show-through" only where it is not hidden by the higher level of detail.
While this simple algorithm is useful, it has some drawbacks, firstly it wastes processing time because the tiles are drawn even if they end up being partially or completely hidden. is there. Especially in simple calculations, each display pixel is often log<sub>2</sub>It turns out that it is drawn (f) times (re) times, where f is the magnification factor of the display for the lowest level of detail. Second, this technique relies on compounding within the framebuffer, which means that at the midpoint of the drawing operation, the area to be drawn has no final appearance and is double buffered or related methods. Use to create the need to perform offscreen compounding to avoid resolution flicker display. Third, unless additional compound operations are applied, this technique can only be used for opaque rendering, for example, the final rendering will have 50% opacity everywhere and other content will be "transparent". It is impossible to guarantee that it can be "made". This is because the painter algorithm relies strictly on the effect of one "paint layer" (ie the level of detail) that completely hides the underlying layers, where the level of detail is hidden and where it is not. I don't know in advance.
(Disclosure of Invention) (Problems to be solved by the invention) The present invention solves these problems while retaining all the advantages of the painter algorithm.
(Means to solve the problem) One of these advantages is the ability to deal with any type of LOD tiling, including non-rectangular or irregular tiling, as well as irrational grid tiling. I have applied for another provisional patent application. Tiling generally consists of polygonal subdivision or tessellation of the area containing the visual content. For tiling, which is useful in multi-resolution contexts, it is generally desirable that the area of tiles at low level of detail is larger than the area of tiles at high level of detail, and the multiplication factor for their size difference is subdivision g. Therefore, the present inventors assume that this is constant (but is not limited to this). The following describes an improved algorithm using an irrational but rectangular tiling grid. Those skilled in the art will appreciate the generalization to other tiling schemes.
(Best mode for carrying out the invention) The improved algorithm consists of four stages. In the first stage, a composite grid is constructed within the reference frame of the image from the superposition of visible parts of all tile grids within all levels of detail drawn. When the tessellation innovations (detailed in another provisional patent application) are used, the result is an irregular composite grid outlined in Figure 1. In addition, this grid is augmented by grid lines that correspond to the x and y values needed to draw the tile's "blending flap" at each level of detail (the resulting grid is too dense and visual). It is not shown in Fig. 1 because it is crowded with tiles). This composite grid can be defined by a sorted list of x and y values for the grid lines, and all rectangles and triangles that need to draw all visible tiles (including their blending flaps). Has the property that the vertices of are at the intersection of the x and y grid lines. Suppose there are n grid lines parallel to the x-axis and m grid lines parallel to the y-axis. Then a two-dimensional n with entries corresponding to the squares in the grid<sup>*</sup>Build a table of m. Each grid entry has two fields: opacity, which is initially set to zero, and a reference list of specific tiles, which is initially blank.
The second step is to go through tiles sorted by decreasing level of detail (as opposed to simple implementation). Each tile covers an integer composite grid square. For each of these squares, check if the table entry has less than 100% opacity, and if so, add the current tile to the list and increase the opacity accordingly. The per-tile opacity used in this step is stored in the tile data structure. When this second step is complete, the composite grid will contain the entries corresponding to the correct tile pieces to draw within each grid square, along with the opacity when drawing these "tile debris". Normally, these opacity add up to 1. Overall hidden low resolution tiles are not referenced anywhere in this table, but partially hidden tiles are only referenced by partially visible tile debris.
The third stage of the algorithm is traversing the composite grid, where the vertices are averaged with adjacent vertices of the same level of detail and then the vertices retain the total opacity (usually 100%) of each vertex. By readjusting the opacity, the debris opacity of the vertices of the composite grid is adjusted. This implements a subdivided version of the spatial smoothing of the scale described in another provisional patent application. Subdivision results from the fact that composite grids are generally denser than the 3x3 grid per tile defined in innovation # 4, especially for low resolution tiles. (In the best LOD, the structure makes the composite gridting at least as detailed as necessary.) This allows the averaging technique by actually creating a smoother blending flap with more tile debris. You will be able to achieve considerable smoothness at a remarkable level of detail.
Finally, in the fourth stage, the composite grid is traversed again and the tile debris is actually drawn. This algorithm involves multiple passes through the data and a certain amount of bookkeeping, which results in much less drawing to be done in the end, much more than a simple algorithm. It improves performance and makes every tile fragment rendered visible to the user, sometimes with low opacity. Some tiles may not be drawn at all. This is in contrast to a simple algorithm that draws any tile that intersects the displayed area as a whole.
An additional advantage of this algorithm is that you can draw partially transparent nodes simply by changing the total opacity target from 100% to some lower value. This is a simple algorithm because every level of detail except the most detailed level must be drawn with the highest opacity in order to fully "overpaint" any underlying still low resolution tile. It is impossible.
If the view is rotated in the xy plane with respect to the node, some minor changes are needed for efficiency. Composite grids can be constructed in the usual way, and if a larger coordinate area is visible along the diagonal, it can be larger than the grid if it is not rotated. However, when passing through tiles, only visible tiles (by simple cross-polygon criteria) need to be considered. Further, the composite grid square outside the display area does not need to be updated when traversing in the second or third stage or drawing in the fourth stage. It should be noted that some other implementation details can be modified to optimize performance, and the algorithm is presented herein in the most understandable form of its operation and original features. Any graphics programmer in the art can easily add the details of the optimized implementation. For example, you don't have to maintain a tile list per tile debris, instead, when each level of detail is complete, you can instantly draw with the correct opacity and always have a single tile identification per debris. It only needs to be stored. Another exemplary optimization is that the total opacity rendering that has not yet been performed, represented by (area) x (remaining transparency), can be tracked, and as a result, if everything has already been drawn. The algorithm can be terminated early and then there is no need to "visit" a low level of detail if it is not needed.
This algorithm can be generalized to any polygonal tiling pattern by using constrained Delaunay triangulation instead of a grid for storing vertex opacity and tile debris identifiers. This data structure efficiently creates a triangular segment whose edges include every edge in all original LOD grids, and access to a particular triangle or vertex is n * log (n) times in sequence. (n is the number of vertices or triangles added) A viable and efficient operation. In addition, the resulting triangle is the basic primitive used for graphics rendering on most graphics platforms.
<img file="JP4861978B2_D0015.tif" />
(Document name) Statement (Title of the Invention) Methods and Devices for Using Image Navigation Techniques to Improve Commercial Transactions
(Technical field) The present invention relates to methods and devices for applying image navigation techniques in improving commerce, such as providing a new environment for advertising and purchasing of products and / or services.
(Background technology) Digital mapping and geographic information applications are a fast-growing industry. These are seeing a sharp increase in investment from companies in many different markets, namely candidates such as Federal Express, clothing stores, and fast food chains. Over the last few years, mapping has become one of the few software applications on the web of great interest (so-called "killer applications"), alongside search engines, webmail, and dating. ing.
Mapping should be an advanced visual representation in principle, but at the moment almost all of its end-user utilities are in drive guidance. The map image that always accompanies the drive guide is usually not very well rendered, conveys little information, is not conveniently navigable, and is just a good look. Clicking on the pan or zoom control causes a long delay, during which the web browser becomes unresponsive, after which a new map image appears that has little visual relevance to the previous image. In principle, it should be more effective for a computer to navigate a digital map than for us to navigate a paper road map, but in reality, a computer-based visual navigation of a map is more effective. It is still inferior.
(Disclosure of Invention) (Problems to be solved by the invention) The present invention is intended to be used in combination with novel techniques to enable continuous and fast visual navigation of maps (or any other image), even over low bandwidth connections. doing.
(Means to solve the problem) This technique relates to a new technique for continuously rendering maps in pan and zoom environments. This is an application of fractal geometry to the rendering of lines and points, allowing road networks (1D curves) and point points (0D points) to be drawn at all scales, creating the illusion of continuous physical zooming. While producing, it still maintains the "visual density" of the demarcated map. Related techniques apply to text labels and icon content. This new rendering technique avoids the effects of sudden appearances or disappearances of small roads during zooming, the negative effects inherent in digital map drawing. Details of this navigation technique can be found (eg, a US patent application (surrogate) named "METHOD AND APPARATUS FOR NAVIGATING AN IMAGE" of the same date as this application, the entire disclosure of which is incorporated herein by reference. See person reference number 489/9)). This navigation technique is sometimes referred to herein as "Voss".
With Voss technology, several new commercial value business models for Internet mapping are feasible. These models employ proven and successful businesses such as Yahoo! Maps and MapQuest as their starting point, both of which are profitable from geographic advertising. However, the methods of the present inventors are not limited to advertising, and utilize the functions of new technologies to add considerable value to both companies and end users. The essential idea is to allow businesses and people to rent "real estate" on a map, usually at their physical address, and embed zoomable content within it. This content can be displayed in the form of an icon when displayed on a wide area map, that is, in the case of McDonalds, such as a golden arch, but if it is displayed finely thereafter, it is optional smoothly and continuously. It is decomposed into various types of web-like content. By integrating dynamic user content and corporate / residential address data into a mapping application such as the present inventor, a kind of "geographical worldwide web" becomes possible.
Existing geographic information services (GIS) providers such as car navigation systems, mobile phones and PDAs, real estate rentals or sales, three-line advertising, and many more, in addition to a large horizontal market for both consumers and retailers. And synergistic with related industries. Possible business relationships in these areas include technology licensing, strategic alliances, and direct vertical integration sales.
The functionality of the new navigation technique of the present invention is described in detail in the aforementioned US patent application. In the case of this application, the most relevant aspects of the basic technology are:
· Smooth zoom and pan through the 2D world with perceptual continuity and extended bandwidth management -Infinite precision coordinate system that allows unlimited nesting of visual content Ability to nest content stored on many different servers, where spatial containment is equivalent to hyperlinks.
Additional details can be found with respect to the latter two elements (eg, "SYSTEM AND METHOD FOR INFINITE PRECISION COORDINATES IN A ZOOMING", filed May 30, 2003, the entire disclosure of which is incorporated herein by reference. See US Provisional Patent Application No. 60/474313, entitled "USER INTERFACE").
Maps consist of many layers of information, and ultimately the Voss map application allows the user to turn most of these layers on and off, allowing the map to be highly customizable. The layers include:
1. Road 2. Waterway 3. Administrative division 4. Aerial photo-based orthoimagery (aerial photo "unwarp" in digital format to fully tile the map) 5. Topographic map 6. Public infrastructure locations such as schools, churches, payphones, toilets, etc. 7. Labels for each of the above 8. Cloud cover, precipitation, and other weather conditions 9. Traffic conditions 10. Advertising 11. Personal and commercial user content, etc.
The most important layers from a typical user's point of view are 1-4 and 7. Advertising / user content layers 10-11, which are of particular interest in this patent application, are also of considerable interest. Many of the map layers, including 1-7, are already available from the US federal government in high quality and for a small amount of money. Value-added layers like 8-9 (and others) can be made available at any time, even during development or even after deployment.
Some companies provide annotated extended road data that is important for generating drive guidance but not for visual geography, showing one-way streets, highway entrance ramps, and other features. ing. The most relevant commercial geographic information service (GIS) provided with respect to this application is geocoding, which allows location addresses to be converted into precise latitude and longitude coordinates. It has been determined that the acquisition of geocoding services is not exorbitantly expensive.
In addition to map data, national occupational phonebook / personal phonebook data may also be useful in implementing the present invention. This information may also be licensed. Use national occupational / personal phonebook data in combination with geocoding to enable geographic user search or filtering (eg, "Highlight all restaurants in Manhattan") To. Perhaps most importantly, a directory list combined with geocoding greatly simplifies the association of corporate and individual users with geographic locations, allowing "real estate" to be rented or outsourced via online transactions. , Does not require large-scale sales force.
National telephone and address databases designed for telemarketing are cheaply available on CD, but they are not always of high quality and their scope is usually only a part and often old. Several companies offer robust directory servers with APIs designed for software-oriented companies like the present invention. The most suitable is W3Data (www.w3data.com), which offers a national phone list called "neartime" using an XML-based API for a minimum of $ 500 / month, starting at $ 0.10 / hit and in volume. It goes down to $ 0.05 / hit for 250,000 / month, or $ 0.03 / hit for more than 1,000,000 / month. It covers the entire United States and Canada. Reverse query is also possible to search for a name by phone number. "Near time" data is updated at least every 90 days. Combined with the 90-day cache of entries already obtained by the inventors, this is a very economical way to obtain a high quality national list. Nightly updated "real-time" data is also available, but more expensive ($ 0.20 / hit). The real-time data is the same as that used by the 411 operator.
There are also service providers similar to W3Data that can generate a national list of companies by category with equivalent business and pricing models.
Traditional ad-based business models, as well as "media player" models that include intellectually owned data formats and downloadable plug-ins (such as Flash, Adobe Acrobat, and Real Player), are usually later or earlier. I have a problem I don't understand. Ad spots become valuable ads only if people have already seen them, and plugins (even if they're free) are worth downloading only if they already have useful content for display. Yes, content will only be invested if you already have a user base that can easily display it.
Voss mapping applications require downloadable client software and can be profitable by advertising it, but without the drawbacks of traditional ad-based business models. Even before "renting" a significant amount of commercial space, the present invention is a useful and visually appealing way to display a map and retrieve addresses, i.e. similar to existing mapping applications. It will provide features with a significantly improved visual interface. Further, the method of the present invention provides a limited but useful service for non-profit users whose user base is free. Restricted services consist of hosting a small amount (5-15MB) of server space per user in the user's geographic location, usually at home. The client software should include a simple authoring feature that allows the user to drag and drop images and text onto their "physical address" and then view it by any other authorized user with the client software. Can be done. (Password protection is available). Photo albums that share only potential attract a significant number of users, especially because zooming user interface techniques are clearly useful for navigating digital photo collections over limited bandwidth. be able to. With a reasonable annual fee, you can use additional server space. This is a very horizontal market that can be a major source of income.
Providing useful services from the beginning, a technique that works well for search engines and other useful (and now useful) web services, creates the usual uncertain which one comes later. It is avoided to do. This is in stark contrast to online dating services, which are not useful until, for example, a user base is built.
According to various aspects of the invention, sources of income include: 1. Commercial "rental" of space on the map corresponding to the physical address 2. "Plus Services" (defined below) fees for commercial users 3. "Plus service" fees for non-profit users 4. Professional zoomable content authoring software 5. License or partnership with vendors and service providers for PDAs, mobile phones, car navigation systems, etc. 6. Information
Basic commercial rentals of space on the map can be priced using a combination of the following variables: 1. Number of sites on the map 2. Map area per site in square meters (land occupied area) 3. Desirability of real estate based on total display statistics 4. Server space required to host content in MB
Plus services for commercial users are aimed at companies that want to implement franchises, e-commerce or other more sophisticated web usage, and companies that want to raise awareness of their advertising.
1. Increased Visible Height--Visible but unobtrusive icons or "flags" can indicate the location of a business site from a greater distance (more zoomed out) than would otherwise be visible. Some levels are offered, such as.
2. Focusing Priority--Voss focuses the image area when data becomes available during navigation. By default, all visual content is treated equally and the focus is from the center of the screen to the outside. Focusing priorities focus product content faster than others, increasing the user's "peripheral vision" prominence. This feature is tuned to deliver commercial value without compromising the user's navigation experience.
3. Incorporation of traditional web hyperlinks into zoomable content--these are clearly marked (for example, with traditional underlined blue text) and the web browser opens when the user clicks. We can charge to include these hyperlinks, or we can charge for each click, like Google.
4. Geographic area rental with reference to an external commercial server that itself hosts zoomable content of any type and size--this is the # 3 enthusiast version of any kind of e-business. It can be executed via a map. We can either charge a flat rate again or charge each user who connects to an external server (which no longer requires a click, just zoom in).
5. Bulletin Board--As in real life, many highly visible areas of the map have considerable free space. The company purchases this space with hyperlinks and content containing hyperlinks and "hyperjumps" that, when clicked, would cause the user to jump through the space to a commercial site somewhere else on the map. , Can be inserted. In contrast to regular commercial space, board space does not need to be rented in a fixed location, which location can be generated while user navigation is in progress.
This last service raises issues of "ecology" or visual aesthetics and map usability. It is desirable that the map be attractive and leave a sincere service to the user that suggests restrictions and sophisticated "parcel restrictions" in advertising. If the map becomes too cluttered with bulletin boards or other content that does not reflect real-world geography, users will be turned off and the map will be less valuable as an advertisement and e-commerce.
When the statistics of the present inventors are collected, the value of many of these commercial "plus services" can be quantitatively demonstrated. Quantitative evidence of competitive advantage should increase sales of these accessories.
Plus services for non-profit users consist of several of the same products that differ in scale, price, and market entry.
Limited authoring of zoomable Voss content is possible from within the free client. This includes inserting text, dragging and dropping digital photos, and setting passwords. Professional authoring software is a modified version of the client designed to allow for more flexible zoomable content creation, as well as mechanisms for creating hyperlinks and hyperjumps and inserting custom applets. be able to.
The present invention can be used to generate large amounts of total and individual information about spatial attention density, navigation routes, and other patterns. These data have commercial value.
<figref num="1">It is a figure which shows the image of the zoom level 5 acquired from the MapQuest website.</figref><figref num="2">It is a figure which shows the image of the zoom level 6 acquired from the MapQuest website.</figref><figref num="3">It is a figure which shows the image of the zoom level 7 acquired from the MapQuest website.</figref><figref num="4">It is a figure which shows the image of the zoom level 9 acquired from the MapQuest website.</figref><figref num="5">FIG. 5 illustrates an image of a long island created at a zoom level of about 334 meters / pixel according to one or more aspects of the invention.</figref><figref num="6">FIG. 5 illustrates an image of a long island created at a zoom level of about 191 meters / pixel according to one or more other aspects of the invention.</figref><figref num="7">FIG. 6 illustrates an image of a long island created at a zoom level of about 109.2 meters / pixel according to one or more other aspects of the invention.</figref><figref num="8">FIG. 5 illustrates an image of a long island created at a zoom level of about 62.4 meters / pixel according to one or more other aspects of the invention.</figref><figref num="9">FIG. 5 illustrates an image of a long island created at a zoom level of about 35.7 meters / pixel according to one or more other aspects of the invention.</figref><figref num="10">FIG. 5 illustrates an image of a long island created at a zoom level of about 20.4 meters / pixel according to one or more other aspects of the invention.</figref><figref num="11">FIG. 6 illustrates an image of a long island created at a zoom level of about 11.7 meters / pixel according to one or more other aspects of the invention.</figref><figref num="12">FIG. 5 is a flow diagram illustrating process steps that can be performed to provide smooth, continuous navigation of an image, according to one or more aspects of the invention.</figref><figref num="13">FIG. 5 is a flow diagram illustrating other process steps that can be performed to smoothly navigate an image according to various aspects of the invention.</figref><figref num="14">In a diagram showing a log-logarithmic graph of metric / pixel zoom levels relative to pixel line widths showing physical and non-physical scaling according to one or more other aspects of the invention. is there.</figref><figref num="15">FIG. 6 shows a log-logarithmic graph showing changes in physical and non-physical scaling in FIG.</figref><figref num="16">A to D are diagrams showing alias-removed vertical lines with endpoints precisely centered in pixel coordinates.</figref><figref num="17A">A to C are diagrams showing slanted alias-removed lines whose endpoints are not positioned exactly in pixel coordinates.</figref><figref num="18">Zoom level for the line width in Figure 14, including horizontal lines that indicate incremental line widths and vertical lines that are spaced apart so that the line width over the distance between two adjacent vertical lines varies by 2 pixels or less. It is a figure which shows the graph represented by the logarithm-logarithm of.</figref>
33 sheets
Sheet 1 Sheet 2 Sheet 3 Sheet 4 Sheet 5 Sheet 6 Sheet 7 Sheet 8 Sheet 9 Sheet 10 Sheet 11 Sheet 12 Sheet 13 Sheet 14 Sheet 15 Sheet 16 Sheet 17 Sheet 18 Sheet 19 Sheet 20 Sheet 21 Sheet 22 Sheet 23 Sheet 24 Sheet 25 Sheet 26 Sheet 27 Sheet 28 Sheet 29 Sheet 30 Sheet 31 Sheet 32 Sheet 33
Every citation, both ways
| Document | Relation | Office | Cited during |
|---|---|---|---|
| WO03107312A2 | Cites | World Intellectual Property Organization (WIPO) | Examiner |
| JP2000046566A | Cites | Japan | Examiner |
| JP2002245473A | Cites | Japan | Examiner |
| WO03107312A1 | Cites | World Intellectual Property Organization (WIPO) | – |
| JP2000046566A | Cites | Japan | – |
| JP2002245473A | Cites | Japan | – |
69 members in 10 offices
Priority claims14
| Document | Office | Kind | Date |
|---|---|---|---|
| 10803010 | United States of America | – | |
| 55380304 | United States of America | P | |
| 55380304 | United States of America | P | |
| 60553803 | United States of America | – | |
| 80301004 | United States of America | A | |
| 80301004 | United States of America | A | |
| 2005008812 | United States of America | W | |
| 2005008812 | United States of America | W | |
| 2004553803 | – | – | – |
| 2004803010 | – | – | – |
| 2005008812 | – | – | – |
| US20040553803P | – | – | – |
| US20040803010 | – | – | – |
| WO2005US08812 | – | – | – |
Members69
| Document | Office | Kind | |
|---|---|---|---|
| US827935A | United States of America | A | |
| US907259A | United States of America | A | |
| WO2004081869A2 | World Intellectual Property Organization (WIPO) | A2 | |
| US2004233219A1 | United States of America | A1 | |
| WO2004109446A2 | World Intellectual Property Organization (WIPO) | A2 | |
| US2005001849A1 | United States of America | A1 | |
| WO2004109446A3 | World Intellectual Property Organization (WIPO) | A3 | |
| US2005184989A1 | United States of America | A1 | |
| US2005206657A1 | United States of America | A1 | |
| CA2558833A1 | Canada | A1 | |
| CA2559678A1 | Canada | A1 | |
| CA2812008A1 | Canada | A1 | |
| WO2005089403A2 | World Intellectual Property Organization (WIPO) | A2 | |
| WO2005089434A2 | World Intellectual Property Organization (WIPO) | A2 | |
| WO2004081869A3 | World Intellectual Property Organization (WIPO) | A3 | |
| US2005268044A1 | United States of America | A1 | |
| US2005270288A1 | United States of America | A1 | |
| US7042455B2 | United States of America | B2 | |
| WO2006052390A2 | World Intellectual Property Organization (WIPO) | A2 | |
| US7075535B2 | United States of America | B2 | |
| US2006176305A1 | United States of America | A1 | |
| AU2006230233A1 | Australia | A1 | |
| CA2599357A1 | Canada | A1 | |
| WO2006105158A2 | World Intellectual Property Organization (WIPO) | A2 | |
| US2006235941A1 | United States of America | A1 | |
| US7133054B2 | United States of America | B2 | |
| US2006267982A1 | United States of America | A1 | |
| WO2005089434A3 | World Intellectual Property Organization (WIPO) | A3 | |
| EP1756521A2 | European Patent Office (EPO) | A2 | |
| US2007047101A1 | United States of America | A1 | |
| US2007047102A1 | United States of America | A1 | |
| WO2006052390A3 | World Intellectual Property Organization (WIPO) | A3 | |
| EP1759354A2 | European Patent Office (EPO) | A2 | |
| US2007104378A1 | United States of America | A1 | |
| US7224361B2 | United States of America | B2 | |
| WO2006052390A9 | World Intellectual Property Organization (WIPO) | A9 | |
| EP1810249A2 | European Patent Office (EPO) | A2 | |
| US7254271B2 | United States of America | B2 | |
| US2007182743A1 | United States of America | A1 | |
| US7286708B2 | United States of America | B2 | |
| JP2007529786A | Japan | A | |
| KR20070116925A | Republic of Korea | A | |
| EP1864222A2 | European Patent Office (EPO) | A2 | |
| JP2008501160A | Japan | A | |
| US2008031527A1 | United States of America | A1 | |
| US2008050024A1 | United States of America | A1 | |
| CN101147174A | China | A | |
| US7375732B2 | United States of America | B2 | |
| JP2008517540A | Japan | A | |
| JP2008535098A | Japan | A | |
| WO2006105158A3 | World Intellectual Property Organization (WIPO) | A3 | |
| RU2007136099A | Russian Federation | A | |
| WO2005089403A3 | World Intellectual Property Organization (WIPO) | A3 | |
| US7546419B2 | United States of America | B2 | |
| US7554543B2 | United States of America | B2 | |
| CN101501664A | China | A | |
| BRPI0607611A2 | Brazil | A2 | |
| US7724965B2 | United States of America | B2 | |
| AU2006230233B2 | Australia | B2 | |
| US7912299B2 | United States of America | B2 | |
| US7930434B2 | United States of America | B2 | |
| CN101147174B | China | B | |
| JP4831071B2 | Japan | B2 | |
| JP4861978B2This record | Japan | B2 | |
| EP1864222A4 | European Patent Office (EPO) | A4 | |
| CA2559678C | Canada | C | |
| EP1810249A4 | European Patent Office (EPO) | A4 | |
| CA2558833C | Canada | C | |
| CA2812008C | Canada | C |
28 legal events, as the office reported them to INPADOC
Over the term
Point at a mark for the eventEvents
| Event | Code | |
|---|---|---|
| Cancellation because of no payment of annual feesLAPS | LAPS | |
| Receipt of annual feesJAPANESE INTERMEDIATE CODE: R250R250 | R250 | |
| Receipt of annual feesJAPANESE INTERMEDIATE CODE: R250R250 | R250 | |
| Receipt of annual feesJAPANESE INTERMEDIATE CODE: R250R250 | R250 | |
| Receipt of annual feesJAPANESE INTERMEDIATE CODE: R250R250 | R250 | |
| Receipt of annual feesJAPANESE INTERMEDIATE CODE: R250R250 | R250 | |
| Receipt of annual feesJAPANESE INTERMEDIATE CODE: R250R250 | R250 | |
| Receipt of annual feesJAPANESE INTERMEDIATE CODE: R250R250 | R250 | |
| Receipt of annual feesJAPANESE INTERMEDIATE CODE: R250R250 | R250 | |
| Written notification of registration of transferJAPANESE INTERMEDIATE CODE: R350R350 | R350 | |
| Request for change of ownership or part of ownershipJAPANESE INTERMEDIATE CODE: R313111S111 | S111 | |
| Request for change of ownership or part of ownershipJAPANESE INTERMEDIATE CODE: R313113S111 | S111 | |
| Receipt of annual feesJAPANESE INTERMEDIATE CODE: R250R250 | R250 | |
| Receipt of annual feesJAPANESE INTERMEDIATE CODE: R250R250 | R250 | |
| Renewal fee payment (event date is renewal date of database)FPAY | FPAY | |
| Certificate of patent or registration of utility modelJAPANESE INTERMEDIATE CODE: R150R150 | R150 | |
| Certificate of patent or registration of utility modelJAPANESE INTERMEDIATE CODE: R150R150 | R150 | |
| First payment of annual fees (during grant procedure)JAPANESE INTERMEDIATE CODE: A61A61 | A61 | |
| Written decision to grant a patent or to grant a registration (utility model)JAPANESE INTERMEDIATE CODE: A01A01 | A01 | |
| Written decision to grant a patent or to grant a registration (utility model)JAPANESE INTERMEDIATE CODE: A01A01 | A01 | |
| Decision of grant or rejection writtenTRDD | TRDD | |
| Request for written amendment filedJAPANESE INTERMEDIATE CODE: A523A521 | A521 | |
| Notification of reasons for refusalJAPANESE INTERMEDIATE CODE: A131A131 | A131 | |
| Request for written amendment filedJAPANESE INTERMEDIATE CODE: A523A521 | A521 | |
| Notification of reasons for refusalJAPANESE INTERMEDIATE CODE: A131A131 | A131 | |
| Request for written amendment filedJAPANESE INTERMEDIATE CODE: A523A521 | A521 | |
| Notification of reasons for refusalJAPANESE INTERMEDIATE CODE: A131A131 | A131 | |
| Written request for application examinationJAPANESE INTERMEDIATE CODE: A621A621 | A621 |
Numbers
- Publication
- 4861978
- Publication, DOCDB
- 4861978
- Publication, EPODOC
- JP4861978B
- Application
- 2007504079
- Application, DOCDB
- 2007504079
- Application, EPODOC
- JP20070504079
Titles2
- Japanese
- イメージをナビゲートするための方法および装置
- English
- Methods and devices for navigating images
Classification
- CPC, 5
- G06F3/04845
- G06F2203/04806
- G06T3/40
- G09B29/007
- G09B29/10
- IPC, 3
- G06T11 60
- G06F3 048
- G06F3 0484
