A method and a device for collaborative computing
26 claims: 23 independent, 3 dependent
- 1多次元立方体データ構造におけるデータ分析のための方法であって、 データセットを処理して第1の多次元立方体データ構造をもたらすことを含み、前記データセットは、1つ以上のテーブルを含むテーブル構造を有し、さらに、 1つ以上の次元制限を前記第1の多次元立方体データ構造に適用することによって、第2の多次元立方体データ構造を生成することを含 み、 前記1つ以上の次元制限を前記第1の多次元立方体データ構造に適用することは、次元制限を適用して、表示中の前記第2の多次元立方体データ構造の第1の特定の部分をもたらすことを含み、 前記第1の特定の部分は、特定の条件を満たさない複数の値を集約することに起因する集約された値を含み、 前記第1の多次元立方体データ構造は、前記データセットの1つ以上の計算変数について特定の数学関数を算定した結果を含み、前記第1の多次元立方体データ構造は、前記データセットの1つ以上の次元変数のすべての固有値について区分化され、 前記データセットを処理して前記第1の多次元立方体データ構造をもたらすことは、前記テーブル構造の1つ以上のテーブルからデータ項目を順次読出すことと、1つ以上のデータ記録を含む中間データ構造を投入することとを含み、前記1つ以上のデータ記録の各々1つは、各次元変数についてのフィールドと、前記特定の数学関数によって暗示される1つ以上の数式についての集約フィールドとを含み、 前記第2の多次元立方体データ構造は、前記中間データ構造を全面的に検討することによって生成され、それにより特定の条件を満たさない複数の値を集約することに起因する集約された値を生成する、 方法。
- 2前記1つ以上の次元制限を前記第1の多次元立方体データ構造に適用することは、1つ以上のユーザインターフェイス要素を構成することを含む、請求項1に記載の方法。
- 3前記1つ以上の次元制限を前記第1の多次元立方体データ構造に適用することは、特定の表示オプションの選択に応答して、次元制限を前記第2の多次元立方体データ構造の選択された次元変数に適用することを含む、請求項1または2に記載の方法。
- 4前記1つ以上の次元制限を前記第1の多次元立方体データ構造に適用することは、表示オプションを適用して、表示中の第2の多次元立方体データ構造の第2の特定の部分をもたらすことをさらに含む、請求項 1 に記載の方法。
- 5前記第1の特定の部分は、前記第1の多次元立方体データ構造に含まれるテーブルの特定の複数の行を含む、請求項 1~4のいずれか1項 に記載の方法。
- 6前記第1の特定の部分は、前記第1の多次元立方体データ構造に含まれるテーブルの特定の複数の行を含み、前記特定の複数の行のそれぞれの値は累算され、所定の値と比較される、請求項 1~5 のうちいずれか1項に記載の方法。
- 7前記第1の特定の部分は、前記第1の多次元立方体データ構造に含まれるテーブルの第1の特定の複数の行を含み、前記 第1の 特定の複数の行は大きい順に並べられる、請求項 1~6 のうちいずれか1項に記載の方法。
- 8前記第1の特定の部分は、前記第1の多次元立方体データ構造に含まれるテーブルの第1の特定の複数の行を含み、前記 第1の 特定の複数の行は小さい順に並べられる、請求項 1~6 のうちいずれか1項に記載の方法。
- 9前記第1の特定の部分は、特定の条件を満たす1つ以上の値を含む、請求項 1~6 のうちいずれか1項に記載の方法。
- 10データセットを処理して前記第1の多次元立方体データ構造をもたらすことは、前記テーブル構造の1つ以上の次元変数について数学関数を算定することを含む、請求項1~ 9 のうちいずれか1項に記載の方法。
- 11投入するアクションは、前記データ項目について、各次元変数の現在の値を識別することと、データ項目に基づいて前記1つ以上の数式の各々1つを算定することと、各次元変数の現在の値に基づいて、適切な集約フィールドにおける前記算定の結果を集約することとを含む、請求項 1~10のいずれか1項 に記載の方法。
- 12前記第1の多次元立方体データ構造は、各次元変数のすべての固有値についての集約フィールドの内容に基づいて、前記特定の数学関数を算定することによって生成される、請求項 11 に記載の方法。
- 13前記第1の多次元立方体データ構造に基づいて、前記特定の条件を満たさない1つ以上の次元変数の値を識別することをさらに含む、請求項 1~12のいずれか1項 に記載の方法。
- 14前記中間データ構造を全面的に検討することは、前記特定の条件を満たさない前記1つ以上の次元変数の値に関連付けられた前記集約フィールドの内容を集約し、それにより前記特定の条件を満たさない複数の値を集約するために前記特定の数学関数を算定することを含む、請求項 1~13のいずれか1項 に記載の方法。
- 15多次元立方体データ構造におけるデータ分析のための装置であって、 コンピュータ実行可能な命令を有するメモリと、 前記メモリに機能的に結合されたプロセッサとを含み、前記プロセッサは、前記コンピュータ実行可能な命令により、 データセットを処理して第1の多次元立方体データ構造をもたらすように構成され、前記データセットは、1つ以上のテーブルを含むテーブル構造を有し、さらに、 1つ以上の次元制限を前記第1の多次元立方体データ構造に適用することによって、第2の多次元立方体データ構造を生成するように構成されて おり、 前記1つ以上の次元制限を前記第1の多次元立方体データ構造に適用することは、次元制限を適用して、表示中の前記第2の多次元立方体データ構造の第1の特定の部分をもたらすことを含み、 前記第1の特定の部分は、特定の条件を満たさない複数の値を集約することに起因する集約された値を含み、 前記第1の多次元立方体データ構造は、前記データセットの1つ以上の計算変数について特定の数学関数を算定した結果を含み、前記第1の多次元立方体データ構造は、前記データセットの1つ以上の次元変数のすべての固有値について区分化され、 前記データセットを処理して前記第1の多次元立方体データ構造をもたらすことは、前記テーブル構造の1つ以上のテーブルからデータ項目を順次読出すことと、1つ以上のデータ記録を含む中間データ構造を投入することとを含み、前記1つ以上のデータ記録の各々1つは、各次元変数についてのフィールドと、前記特定の数学関数によって暗示される1つ以上の数式についての集約フィールドとを含み、 前記第2の多次元立方体データ構造は、前記中間データ構造を全面的に検討することによって生成され、それにより特定の条件を満たさない複数の値を集約することに起因する集約された値を生成する、 装置。
- 16前記プロセッサはさらに、1つ以上のユーザインターフェイス要素を構成するように構成されている、請求項 15 に記載の装置。
- 17前記1つ以上の次元制限の各々は、前記第2の多次元立方体データ構造における1つ以上の次元変数の値の表示数を制約する、請求項 15 または 16 に記載の装置。
- 18前記プロセッサはさらに、特定の表示オプションの選択に応答して、次元制限を前記第2の多次元立方体データ構造の選択された次元変数に適用するように構成されている、請求項 15 または 16 に記載の装置。
- 19前記プロセッサはさらに、表示オプションを適用して、表示中の前記第2の多次元立方体データ構造の第2の特定の部分をもたらすように構成されている、請求項 15~18のいずれか1項 に記載の装置。
- 20前記第1の特定の部分は、前記第1の多次元立方体データ構造に含まれるテーブルの特定の複数の行を含む、請求項 15~18のいずれか1項 に記載の装置。
- 21前記第1の特定の部分は、前記第1の多次元立方体データ構造に含まれるテーブルの特定の複数の行を含み、前記特定の複数の行のそれぞれの値は累算され、所定の値と比較される、請求項 15~20 のうちいずれか1項に記載の装置。
- 22前記第1の特定の部分は、前記第1の多次元立方体データ構造に含まれるテーブルの第1の特定の複数の行を含み、前記 第1の 特定の複数の行は大きい順に並べられる、請求項 15~21 のうちいずれか1項に記載の装置。
- 23前記第1の特定の部分は、前記第1の多次元立方体データ構造に含まれるテーブルの第1の特定の複数の行を含み、前記 第1の 特定の複数の行は小さい順に並べられる、請求項 15~21 のうちいずれか1項に記載の装置。
- 24前記第1の特定の部分は、特定の条件を満たす1つ以上の値を含む、請求項 15~23 のうちいずれか1項に記載の装置。
- 25前記プロセッサはさらに、前記データセットの1つ以上の次元変数について数学関数を算定するように構成されている、請求項 15~23 のうちいずれか1項に記載の装置。
- 26請求項1~14のいずれか1項に記載の方法をコンピュータに実行させるためのプログラム。
Independent claims26
161 paragraphs, as filed
0001Technical field The present invention relates to methods and devices for data analysis.
0002Prior art US7058621 discloses how to act on a database to extract information and present it to the user. The database contains a data table that contains the values of multiple variables. The information is to be extracted by computing at least one mathematical function that acts on one or more selected computational variables. The information presented is to be partitioned for one or more selected taxonomies. In this method, the steps of identifying all boundary tables, the steps of identifying all connection tables, the step of selecting a start table from the boundary table and connection tables, and the value of each selected variable in the boundary table are set. As the calculation produces a final data structure that contains the results of a mathematical function for all the eigenvalues of each class variable, with the steps to build a transformation structure that links to the corresponding values of one or more connection variables in the start table. Includes steps to calculate mathematical functions for each data record in the starting table by using a transformation structure.
0003US20100017436 discloses methods and equipment for extracting information. The information is extracted from the database using a computer-implemented method with a chain of main calculations, in which the first main calculation (P1) is in the dataset (R0) that represents the database. The first selection item (S1) is applied to generate the first result (R1), and the second main calculation (P2) puts the second selection item (S2) on the first result (R1). It acts to produce a second result (R2). The first and second results (R1, R2) are cached in computer memory (10) for reuse in subsequent iterations of this method, thereby extracting the first and / or second results. Reduce the need to perform the main calculation of 2 (P1, P2). Caching is at least the first selection identifier value (ID1) as a function of the first selection (S1) and the second as a function of at least the second selection (S2) and the first result (R1). And the selection identifier value (ID3) of, and the first selection identifier value (ID1) and the first result (R1), and the second selection identifier value (ID3) and the second result (R2). ) And each are stored as associated objects in the data structure (12). Each of the identifier values is generated as a statistically unique electronic fingerprint by the hash function (f). Reuse calculates the first and second select identifier values (ID1, ID3) during the next iteration, and in some cases looks up the first and / or second results (R1, R2). Accompanied by accessing the data structure (12).
<p num="0004"> Outline of the invention In one aspect, a method and system for the user's database method and interaction with the system is provided. In one aspect, a user interface may be generated to facilitate dynamic display generation for viewing the data. The system may include visualization components to dynamically generate one or more visual representations of the data to present in the state space.</p><p num="0005"> In one aspect, the disclosure relates to methods for data analysis. The method may include processing the dataset to result in a first multidimensional cubic data structure, the dataset having a table structure containing one or more tables. The method may further include generating a second multidimensional cube data structure by applying one or more dimension restrictions to the first multidimensional cube data structure.</p><p num="0006"> In another aspect, the disclosure relates to computing devices. The computing device includes a memory with computer-executable instructions and a processor functionally coupled to the memory, which processes the data set with computer-executable instructions to form a first multidimensional cube. Configured to provide a data structure, the dataset has a table structure containing one or more tables, and by applying one or more dimension limits to the first multidimensional cube data structure. It may be configured to generate two multidimensional cubic data structures.</p><p num="0007"> Some of the additional benefits are mentioned in the discussion below. Alternatively, it may be learned by practice. Benefits will be realized and obtained by the elements and combinations specifically pointed out in the attached claims. It should be understood that, as described, both the above overview and the detailed description below are merely exemplary and explanatory, and are not limiting.</p><p num="0008"> A brief description of the drawing The accompanying drawings incorporated herein by reference and forming part of this specification exemplify examples and serve to explain the principles of methods and systems along with explanations.</p>
0009<figref num="1">It is a figure which shows an exemplary table 1-5.</figref><figref num="2">It is a block flowchart of an exemplary method for extracting information from a database.</figref><figref num="3">It is a figure which shows an exemplary table 6-12.</figref><figref num="4">It is a figure which shows an exemplary table 13-16.</figref><figref num="5">It is a figure which shows an exemplary table 17, 18, and 20-23.</figref><figref num="6">It is a figure which shows an exemplary table 24-29.</figref><figref num="7">It is a figure of an exemplary operating environment.</figref><figref num="8">It is a figure which shows how selection acts on a range to generate a data subset.</figref><figref num="9">It is a figure which shows an exemplary user interface.</figref><figref num="10a">It is a figure which shows another exemplary user interface.</figref><figref num="10b">It is a figure which shows another exemplary user interface.</figref><figref num="11a">It is a block flowchart of an exemplary method.</figref><figref num="11b">It is a figure of an exemplary operating environment.</figref><figref num="12a">It is a figure which shows an exemplary user interface.</figref><figref num="12b">It is a figure which shows another exemplary user interface.</figref><figref num="13a">It is a block flowchart of an exemplary method.</figref><figref num="13b">It is another block flowchart of an exemplary method.</figref><figref num="13c">It is another block flowchart of an exemplary method.</figref><figref num="14">It is a figure which shows an exemplary user interface.</figref><figref num="15">It is a block flowchart of an exemplary method.</figref><figref num="16">It is a figure which shows an exemplary user interface.</figref><figref num="17a">It is a figure which shows an exemplary table.</figref><figref num="17b">It is a figure which shows an exemplary table.</figref><figref num="17c">It is a figure which shows an exemplary table.</figref><figref num="17d">It is a figure which shows an exemplary table.</figref><figref num="17e">It is a figure which shows an exemplary table.</figref><figref num="17f">It is a figure which shows an exemplary table.</figref><figref num="18a">It is a figure which shows an additional exemplary table.</figref><figref num="18b">It is a figure which shows an additional exemplary table.</figref><figref num="18c">It is a figure which shows an additional exemplary table.</figref><figref num="19">It is a block flowchart of an exemplary method.</figref><figref num="20">It is a figure which shows an exemplary method for data analysis according to one or more aspects of this disclosure.</figref><figref num="21">FIG. 5 illustrates an exemplary computing device for data analysis according to one or more aspects of the present disclosure.</figref>
0010Detailed explanation Before disclosing and describing the methods and systems of the invention, it should be understood that the methods and systems are not limited to any particular method, particular component, or particular configuration. It should also be understood that the terminology used herein is intended only to describe a particular embodiment and is not intended to be limiting.
0011As used in the specification and the accompanying claims, the singular includes multiple objects unless the context clearly dictates otherwise. The range may be expressed here as "approximately" from one particular value to "approximately" another particular value. When such a range is represented, another embodiment includes from that particular value to that other particular value. Similarly, if the value is expressed as an approximation by the use of the antecedent "approximately", it will be understood that that particular value forms another embodiment. Furthermore, it will be understood that the end points of each range are significant in relation to the other end point and independently of the other end point.
0012"Optional" or "optionally" means that the event or situation described below may or may not occur, and that description causes the event or situation described below. It means to include cases and cases that do not occur.
0013Throughout the description and claims herein, the word "contains" and variations of this wording, such as "contains" and "contains," mean "including, but not limited to,". , For example, are not intended to exclude other adjuncts, components, integers or steps. "Exemplary" means "an example of ..." and is not intended to convey a representation of a preferred or ideal example. "... etc." Is not used in a limited sense, but is used for explanatory purposes.
0014The components that can be used to perform the disclosed methods and systems are disclosed. These and other components are disclosed herein. Combinations, subsets, interactions, groups, etc. of these components, with specific reference to each, may not explicitly disclose various individual and collective combinations, as well as their substitutions. It is understood that if any is disclosed, each is hereby specifically considered and described for all methods and systems. This applies to all aspects of the application, including but not limited to steps in the disclosed methods. Thus, if there are various additional steps or actions that can be performed, each of these additional steps can be performed in any particular embodiment or combination of embodiments of the disclosed method. Is understood.
0015Here, the present invention and the system will be described as an example with reference to FIGS. 1 to 6 of the drawings. FIG. 1 shows the contents of the database after identification of the relevant data tables according to the disclosed method, and FIG. 2 shows a series of steps in one embodiment of the method according to one or more aspects of the present disclosure. 3 to 6 show an exemplary data table.
0016The database contains a plurality of data tables (tables 1 to 5) as shown in FIG. Each data table contains the data values of multiple data variables. For example, in Table 1, each data record contains data values for the data variables "Product", "Price", and "Part". If a field in the data recording does not have a specific value, this field is considered to hold a null-value. Similarly, in Table 2, each data record contains the values of the variables "Date", "Client", "Product", and "Number". Data values are typically stored in the form of ASCII-encoded strings.
0017A method according to one or more aspects of the present disclosure can be implemented, for example, by a computer program in response to execution by a processor. In the first step (step 101), the program reads all the data records in the database, for example, using the SELECT statement to select all the tables in the database, in this case tables 1-5. put out. Typically, the database is read into the computer's primary memory.
0018In order to increase the calculation speed, it is preferable that each eigenvalue of each data variable in the database is assigned a different binary code, and that the data record is stored in a binary coded format (step 101). ). This is typically done when the program first reads the data record from the database. For each input table, the following steps are performed. First, the column names in the table, such as variables, are read continuously. Each time a new data variable appears, the data structure is instantiated for it. The internal table structure is then instantiated to include all data records in binary format, and the data records are continuously read and binary coded. For each data value, the data structure of the corresponding data variable is checked to see if the value was previously assigned a binary code. If assigned, the binary code is inserted in place in the table structure described above. If not assigned, the data value is added to the data structure, the new binary code, preferably the next binary code in ascending order, is assigned and then inserted into the table structure. In other words, for each data variable, a unique binary code is assigned to each unique data value.
0019Tables 6-12 of FIG. 3 show the binary codes assigned to different data values of some data variables contained in the database of FIG.
0020After reading all the data records in the database, the program analyzes the database to identify all connections between the data tables (step 102). A connection between two data tables means that these data tables share a variable. Different algorithms for performing such analyzes are known in the art. After analysis, all data tables are virtually connected. In Figure 1, such virtual connections are indicated by double-ended arrows (a). Virtually connected data tables should form at least one so-called snowflake structure, eg, a branched data structure with only one connection path between any two data tables in the database. For this reason, the snowflake structure does not contain any loops. If a loop occurs in a virtually connected data table, for example if two tables share more than one variable, in some cases special features known in the art to break such a loop. Depending on the algorithm, snowflake structures may still be formed.
0021After this initial analysis, the user can start exploring the database. In doing so, the user defines a mathematical function, which may be a combination of mathematical formulas (step 103). Suppose the user wants to extract the total sales by year and by client from the database in Figure 1. The user can use the corresponding mathematical function "SUM (x)<sup>*</sup>"y)" is defined and the computational variables "price" and "quantity" to be included in this function are selected. The user also selects the classification variables "Client" and "Year".
0022The computer program then identifies all relevant data tables, such as all data tables that contain any one of the selected compute and classification variables (such data tables are called boundary tables). (Step 104), as well as identify all intermediate data tables (such data tables are called connection tables) in the connection path between these boundary tables in the snowflake structure. For clarity, the first frame (A) of FIG. 1 contains a group of related data tables (tables 1-3). Obviously, in this particular case, there is no connection table.
0023In this case, all values of the selected computational variable, for example all occurrences of frequency data, must be included for the calculation of the mathematical function. In Figure 1, selected variables that require such frequency data ("price", "quantity") are indicated by thick arrows (b), while the remaining selected variables are dotted (b ). It is indicated by. Here, a subset (B) can be defined that includes all boundary tables (Tables 1-2) that contain such computational variables and a connection table between such boundary tables in the snowflake structure. The frequency requirement for a particular variable is determined by the mathematical formula that includes it. Finding the mean or median requires frequency information. In general, this also applies to finding sums, while finding maximum or minimum values does not require frequency data for computational variables. Also, classifiers generally do not require frequency data.
0024The starting table is then selected, preferably from the data tables in subset (B), most preferably from the data table with the most data records in this subset (step 105). In Figure 1, table 2 is selected as the starting table. Therefore, the start table contains selected variables ("client", "quantity") and connection variables ("date", "product"). These connection variables link the starting table (Table 2) to the bounding tables (Tables 1 and 3).
0025The transformation structure is then constructed, as shown in Tables 13 and 14 of FIG. 4 (step 106). This transformation structure puts each value of each connection variable ("date", "product") in the start table (table 2) into the corresponding selected variable ("year", respectively) in the boundary table (tables 3 and 1). Used to convert to a "price") value. Table 13 continuously reads the data records of Table 3 and creates a link between each eigenvalue of the connection variable ("date") and the corresponding value of the selected variable ("year"). Will be built. There is no link from value 4 ("Date: 1999-01-12"). This is because this value is not included in the bounding table. Similarly, Table 14 continuously reads the data records in Table 1 and creates a link between each eigenvalue of the connection variable ("Product") and the corresponding value of the selected variable ("Price"). It is built by doing. In this case, the value 2 ("Product: Toothpaste") is linked to the two values of the selected variable ("Price: 6.5"). This is because this connection has occurred twice in the bounding table. Thus, the frequency data is included in the transformation structure. There is also no link from value 3 ("Product: Shampoo").
0026Once the transformation structure is built, a virtual data record is created. Such virtual data records, as shown in Table 15, apply to all selected variables in the database ("client", "year", "price", "quantity"). When constructing a virtual data record (steps 107 to 108), first, a certain data record is read from the start table (table 2). Next, the value of each selected variable ("client", "number") in the current data recording of the start table is incorporated into the virtual data recording. Also, by using the transformation structure (Tables 13-14), each value of each connection variable ("date", "product") in the current data record of the start table is the corresponding selected variable ("year"). , "Price"), and this value is also incorporated into the virtual data recording.
0027At this stage (step 109), the virtual data recording is used to build the intermediate data structure (Table 16). Each data record in the intermediate data structure adapts to each selected classifier (dimension) and the aggregate field for each formula implied by the mathematical function. The intermediate data structure (Table 16) is constructed based on the values of the selected variables in the virtual data recording. Therefore, each formula is calculated based on one or more values of one or more related compute variables in the virtual data recording, and the result is the current value of the classification variable ("client", "year"). Aggregates in the appropriate aggregation fields based on the combination.
0028The above procedure is repeated for all data records in the start table (step 110). Thus, by continuously reading the data records in the start table, by incorporating the current values of the selected variables in the virtual data records, and by calculating each formula based on the contents of the virtual data records. An intermediate data structure is constructed. If the current combination of classification variable values in the virtual data record is new, a new data record is created in the intermediate data structure to hold the result of the calculation. In other cases, suitable data records are quickly found and the results of the calculations are aggregated in the aggregate field. Thus, as the start table is thoroughly reviewed, data records are added to the intermediate data structure. Preferably, the intermediate data structure is a data table associated with an efficient index system such as an AVL or hash structure. In most cases, the aggregate field is implemented as an add register, where the results of the calculated formulas are accumulated. In some cases, for example, when calculating the median, aggregate fields are instead implemented to hold all the individual results for the unique combination of values of the identified class variable. Note that the procedure for constructing an intermediate data structure from the start table requires only one virtual data record. Therefore, the contents of the virtual data record are updated for each data record in the start table. This will minimize the memory requirements when running computer programs.
0029The procedure for constructing the intermediate data structure will be further described with reference to Tables 15-16. When creating the first virtual data record R1 as shown in Table 15, the values of the selected variables "client" and "number" are taken as is from the first data record in the start table (table 2). Next, the conversion structure (Table 13) transfers the value "1999-01-02" of the connection variable "date" to the value "1999" of the selected variable "year". Similarly, the transformation structure (Table 14) transfers the value "toothpaste" of the connection variable "product" to the value "6.5" of the selected variable "price", thereby forming the virtual data record R1. .. Next, as shown in FIG. 16, a data record is created in the intermediate data structure. In this case, the intermediate data structure has three columns, two of which hold the selected classifiers ("client", "year"). The third column holds the aggregate field, where the formula ("x") that acts on the selected computational variables ("quantity", "price").<sup>*</sup>The calculation results of y ") are aggregated. When calculating the virtual data record R1, the current value of the classifier (binary code: 0,0) is read first and incorporated into this data record of the intermediate data structure. Next, the current value of the calculated variable (binary code: 2,0) is read. The formula is calculated for these values and added to the associated aggregate field.
0030The virtual data record is then updated based on the start table. The updated virtual data record R2 remains unchanged because the transformation structure (Table 14) shows a duplicate of the selected variable "Price" value "6.5" for the connection variable "Product" value "Toothpaste". It is the same as R1. Next, the virtual data record R2 is calculated as described above. In this case, the intermediate data structure contains the data record corresponding to the current value of the classification variable (binary code: 0,0). Therefore, the calculation result of the formula is accumulated in the associated aggregate field.
0031The virtual data record is then updated based on the second data record in the start table. When calculating this updated virtual data record R3, a new data record is created in the intermediate data structure, and so on.
0032In this example, the null value is represented by the binary code -2. Further, in the illustrated example, the virtual data recording holding a null value (-2) of any one of the calculated variables can be directly excluded. Because null values are formulas ("x"<sup>*</sup>This is because it cannot be calculated in y "). Also, all null values (-2) of the classification variable are treated as any other valid value and placed in the intermediate data structure.
0033After a full review of the starting table, the intermediate data structure contains four data records, each with a unique combination of classifier values (0,0; 1,0; 2,0; 3,-2). Includes the corresponding cumulative results (41; 37.5; 60; 75) of the calculated formula.
0034Preferably, the intermediate data structure is also processed to eliminate one or more classification variables (or dimensional variables). Preferably, this is done during the process of constructing the intermediate data structure as described above. Each time a virtual data record is calculated, additional data records are created or discovered in the intermediate data structure if they already exist. Each of these additional data records is defined to hold a summary of the calculation results of the formula for all values of one or more classification variables. Therefore, when the starting table is thoroughly examined, the intermediate data structure is the aggregated result for all unique combinations of classifier values and the aggregated result after exclusion of each associated classifier. Will include both.
0035This procedure of eliminating dimensions in intermediate data structures is further described with reference to Tables 15 and 16. When the virtual data record R1 is calculated (Table 15) and the first data record (0,0) is created in the intermediate data structure, additional data records are created in this structure. Such additional data records are defined to retain the corresponding results when one or more dimensions are eliminated. In Table 16, a taxonomy variable is assigned a binary code of -1 in the intermediate data structure to show that all values for this variable are calculated. In this case, three additional data records are created, each holding a new combination of classifier values (-1,0; 0, -1; -1.-1). Calculation results are aggregated in the associated aggregation fields of these additional data records. The first (-1,0) of these additional data records stipulates that the aggregated results for all values of the classification variable "client" when the classification variable "year" has the value "1999" are retained. Has been done. The second additional data record (0, -1) is defined to hold the aggregated results for all values of the classification variable "year" when the classification variable "client" is "Nisse". A third additional data record (-1, -1) is defined to hold the aggregated results for all values of both classification variables ("client" and "year").
0036When the virtual data record R2 is calculated, the result is in the aggregate field associated with the current combination of class variable values (binary code: 0,0), and the associated additional data record (binary code). Aggregate in the aggregate field associated with: -1,0; 0, -1; -1, -1). When the virtual data record R3 is calculated, the results are aggregated in the aggregate field associated with the current combination of classification variable values (binary code: 1,0). The result is also in the aggregate field of the newly created additional data record (binary code: 1, -1) in the intermediate data structure, and the associated existing data record (binary code: -1,0; It is also aggregated in the aggregate field associated with -1, -1).
0037After a full review of the starting table, the intermediate data structure contains 11 types of data records as shown in Table 16. Preferably, if the intermediate data structure applies to more than one classifier, then the intermediate data structure is for each classifier that is excluded, for each unique combination of the values of the remaining classifiers, for all of this classifier. It will include the calculated results aggregated over the values.
0038Once the intermediate data structure is constructed, the formulas contained in the intermediate data structure ("x"<sup>*</sup>Mathematical function ("SUM (x") based on the result of (y ")<sup>*</sup>By calculating y) "), a final data structure, such as a multidimensional cube, is created, as shown in the non-binary notation in Table 17 of FIG. 5 (step 111). In doing so, the results in the aggregate field for each unique combination of classifier values are combined. In the illustrated case, the creation of the final data structure is straightforward due to the trivial nature of this mathematical function. The contents of the final data structure may then be presented to the user in a two-dimensional table as shown in Table 18 of FIG. 5 (step 112). Alternatively, if the final data structure contains many dimensions, the data may be presented in a pivot table that allows the user to interactively move up and down the dimensions, as is well known in the art.
0039A second example of the disclosed method can be described below with reference to Tables 20-29 of FIGS. 5-6. This description will simply elaborate on one aspect of this example: the construction of transformation structures containing data from connection tables, and the construction of intermediate data structures for more complex mathematical functions. In this example, the user wants to extract sales data for each client from a database that includes the data tables shown in tables 20-23 of FIG. For ease of explanation, binary coding is omitted in this example.
0040The user should classify the result for each client Mathematical function a) "IF (Only (Environment index) ='I') THEN Sum (Number)<sup>*</sup>Price)<sup>*</sup>2, ELSE Sum (Number<sup>*</sup>Price)) ", and b)" Avg (Number)<sup>*</sup>Price) "has already been specified.
0041Mathematical function (a) should double sales for products belonging to the product family with the Environmental index'I', while using actual sales for other products. Is specified. The mathematical function (b) is included for reference.
0042In this case, the selected taxonomy variables are "environmental index" and "client", and the selected computational variables are "quantity" and "price". Tables 20, 22 and 23 are identified as bounding tables, while table 21 is identified as a connection table. Table 20 is selected as the starting table. Therefore, the start table contains selected variables ("quantity", "client") and connection variables ("product"). The connection variable links the start table (table 20) to the boundary table (tables 22-23) via the connection table (table 21).
0043Next, the formation of the conversion structure will be described with reference to Tables 24 to 26 in FIG. The first part of the transformation structure (Table 24) continuously reads the data records of the first boundary table (Table 23), with each eigenvalue of the connection variable ("Products") and the selected variable ("Products"). It is constructed by creating a link to the corresponding value of the "environmental indicator"). Similarly, the second part of the transformation structure (Table 25) was selected with each eigenvalue of the connection variable ("price group") by continuously reading the data records of the second boundary table (Table 22). It is constructed by creating a link to the corresponding value of a variable ("price"). Next, the data record of the connection table (table 21) is continuously read. Instead of the corresponding values of the connection variables ("Products") in Table 21, the values of the connection variables ("Products" and "Prices", respectively) in Tables 24 and 25 are substituted. The results are merged in one final transformation structure as shown in Table 26.
0044Then, by continuously reading the data records in the start table (Table 20), the current values of the selected variables ("Environment Index", "Client", "Quantity", "Price") are virtualized. Intermediate data structures are constructed by using transformation structures (Table 26) to incorporate into the record and by calculating each formula based on the current contents of the virtual data record.
0045For clarity, Table 27 shows the corresponding contents of the virtual data record for each data record in the start table. As mentioned in connection with the first example, only one virtual data recording is required. The contents of this virtual data record are updated, eg, replaced, for each data record in the start table.
0046Each data record in an intermediate data structure, as shown in Table 28, adapts to the value of each selected classifier ("client", "environmental indicator") and the aggregate field for each formula implied by the mathematical function. To do. In this case, the intermediate data structure contains two aggregate fields. One aggregate field is a formula ("x") that processes the selected computational variables ("quantity", "price").<sup>*</sup>It includes the aggregation result of y ") and the counter for the number of such processes. The layout of this aggregate field should calculate the average quantity ("Avg (x)"<sup>*</sup>y) ") given by the fact. The other aggregate field is designed to hold the lowest and highest values of the classification variable "Environmental Indicators" for each combination of classification variable values.
0047As in the first example, by calculating a mathematical formula for the current content of the virtual data record (each row in Table 27), and based on a combination of the current values of the classification variables ("Client", "Environmental Indicators"). The intermediate data structure (Table 28) is constructed by aggregating the results in the appropriate aggregate fields. The intermediate data structure also includes a data record in which the value "<ALL>" (all) is assigned to one or both of the classification variables. The corresponding aggregate field contains the aggregate result if one or more classification variables (dimensions) are excluded.
0048When an intermediate data structure is constructed, a final data structure, such as a multidimensional cube, is created by calculating mathematical functions based on the calculation results of mathematical formulas contained in the intermediate data structure. Each data record in the final data structure, as shown in Table 29, adapts to the value of each selected class variable ("client", "environmental indicator") and the aggregate field for each mathematical function selected by the user. To do.
0049The final data structure is constructed based on the results in the aggregated fields of the intermediate data structure for each unique combination of classification variable values. If function (a) is calculated by sequentially reading the data records in table 28, the program first checks if both values in the last column of table 28 are equal to'I'. If they are equal, the relevant result contained in the first aggregate field in table 28 is multiplied by 2 and stored in table 29. If they are not equal, the relevant results contained in the first aggregate field in Table 28 are stored in Table 29 as is. When function (b) is calculated, a formula ("x") that processes the selected computational variables ("quantity", "price")<sup>*</sup>The aggregation result of y ") is divided by the number of such processes. Both of them are stored in the first aggregate field in Table 28. The result is stored in the second aggregate field in Table 29.
0050With the present disclosure, it is easy to see that the user is free to choose mathematical functions to incorporate computational variables into these functions, and to choose classification variables to present the results. is there.
0051Instead of or in addition to the illustrated procedure of building an intermediate data structure based on a series of data records from the start table, it is conceivable to first build a so-called join table, albeit with reduced memory efficiency. This join table thoroughly reviews all data records in the start table and uses a transformation structure to convert each value of each connection variable in the start table to the value of at least one corresponding selected variable in the bounding table. It is built by doing. Therefore, the data recording of the join table will include all combinations in which the values of the selected variables occur. Next, an intermediate data structure is constructed based on the contents of the join table. For each record in the join table, each formula is calculated and the results are aggregated in the appropriate aggregate field based on the current value of each selected class variable. However, this alternative procedure requires more computer memory to extract the requested information.
0052It should be recognized that mathematical functions may contain mathematical formulas with different and conflicting requirements for frequency data. In this case, for each such formula, steps 104-110 (FIG. 2) are repeated and the results are stored in one common intermediate data structure. Alternatively, one final data structure, eg, a multidimensional cube, may be constructed for each formula, and the contents of these cubes are fused during presentation to the user.
0053As will be appreciated by those skilled in the art, the methods and systems may take the form of fully hardware embodiments, fully software embodiments, or a combination of software and hardware aspects. The method and system may also take the form of a computer program product that resides on a computer-readable storage medium and has computer-readable program instructions (eg, computer software) embodied in the storage medium. More specifically, this method and system may take the form of computer software implemented by the web. Any suitable computer-readable storage medium may be utilized, including hard disks, CD-ROMs, optical storage devices, or magnetic storage devices.
0054Examples of this method and system will be described with reference to block diagrams and flowcharts of methods, systems, equipment, and computer program products. It will be appreciated that each block in the block diagram and flowchart, as well as a combination of blocks in the block diagram and flowchart diagram, can be realized by computer program instructions. These computer program instructions are loaded into a general purpose computer, a dedicated computer, or other programmable data processing device to generate a machine, and instructions that are executed on the computer or other programmable data processing device. , A means for realizing a function specified in a block of a flowchart or a plurality of blocks may be created.
0055These computer program instructions are also stored in a computer-readable memory that can be instructed to function in a particular manner on a computer or other programmable data processing device, and the instructions stored in the computer-readable memory. May produce a product that includes computer-readable instructions to implement the functions specified in the block of the flowchart or in the blocks. Computer program instructions are also loaded onto a computer or other programmable data processing device so that a series of operating steps is performed on the computer or other programmable device to spawn the processes implemented by the computer. In that way, instructions executed on a computer or other programmable device may provide a step for achieving a function identified in a block of flowcharts or blocks.
0056Therefore, the blocks in the block diagram and flowchart diagram are a combination of means for performing the specified function, a combination of steps for performing the specified function, and a program instruction means for performing the specified function. To support. In addition, each block of the block diagram and the flowchart diagram, and the combination of the blocks in the block diagram and the flowchart diagram are determined by a dedicated hardware-based computer system that performs a specified function or step, or a combination of dedicated hardware and computer instructions. It is feasible.
0057Those skilled in the art will appreciate that functional descriptions are provided and that each function can be performed by software, hardware, or a combination of software and hardware. In one aspect, the method and system may include data analysis software 706 as shown in FIG. 7 and described below. In one exemplary aspect, the method and system may include a computer 701 as shown in FIG. 7 and described below.
0058FIG. 7 is a block diagram showing an exemplary operating environment for performing the disclosed method. This exemplary operating environment is merely an example of an operating environment and is not intended to suggest limitations on the use or scope of functionality of the operating environment architecture. Also, the operating environment should not be construed as having a dependency or requirement for any one or a combination of the components illustrated in the exemplary operating environment.
0059This method and system can operate in many other general purpose or dedicated computing system environments or configurations. Examples of well-known computing systems, environments, and / or configurations that may be suitable for use in this system and method include, but are not limited to, personal computers, server computers, laptop devices, and multiprocessor systems. Additional examples include set-top boxes, programmable consumer electronics, networked PCs, minicomputers, mainframe computers, distributed computing environments that include any of the systems or devices described above, and the like.
0060The disclosed methods and processing of the system can be performed by software components. The disclosed systems and methods can be explained in the general context where computer-executable instructions, such as program modules, are executed by one or more computers or other devices. In general, a program module includes computer code, routines, programs, objects, components, data structures, etc. that perform a particular task or achieve a particular abstract data type. The disclosed method can also be practiced in a grit-based distributed computing environment where tasks are performed by remote processors linked through a communication network. In a distributed computing environment, program modules may be located on both local and remote computer storage media, including memory storage.
0061Those skilled in the art will also appreciate that the systems and methods disclosed herein are feasible via general purpose computing equipment in the form of computer 701. The components of computer 701 may include one or more processors or processors 703, system memory 712, and system bus 713, which connects various system components, including processor 703, to system memory 712. Not limited to them. When there are a plurality of processing units 703, the system may utilize parallel computing.
0062The system bus 713 is a bus structure of several possible types, including a memory bus or memory controller, a peripheral bus, an accelerated graphics port, and a processor or local bus using any of the various bus architectures. Represents one or more of. As an example, such architectures include Industry Standard Architecture (ISA) buses, Micro Channel Architecture (MCA) buses, Enhanced ISA (EISA) buses, and Video Electronics Standards Association (Video). Electronics Standards Association (VESA) Local Bus, Accelerated Graphics Port (AGP) Bus, Peripheral Component It may include Interconnects (PCI), PCI Express Bus, Personal Computer Memory Card Industry Association (PCMCIA), Universal Serial Bus (USB), and the like. Bus 713, and all buses identified in this description, may also be implemented through wired or wireless network connections, including processor 703, mass storage 704, operating system 705, data analysis software 706, data 707, network. Each subsystem, including adapter 708, system memory 712, I / O interface 710, display adapter 709, display device 711, and man-machine interface 702, is one or more remote computing devices 714a in physically separate locations. , 714b, 714c may be included and connected through this form of bus, effectively providing a completely decentralized system.
0063The computer 701 typically includes a variety of computer readable media. An exemplary readable medium may be any available medium accessible by computer 701, eg, in a non-limiting sense, both volatile and non-volatile media, removable and non-removable. Includes both media. System memory 712 includes computer-readable media in the form of volatile memory such as random access memory (RAM) and / or non-volatile memory such as read only memory (ROM). System memory 712 typically includes data such as data 707 and / or program modules such as operating system 705 and data analysis software 706, which have immediate access to processing unit 703 and / or processing. Currently being processed by part 703.
0064In another aspect, the computer 701 may also include other removable / non-removable, volatile / non-volatile computer storage media. As an example, FIG. 7 shows a mass storage device 704 capable of providing non-volatile storage of computer code, computer readable instructions, data structures, program modules, and other data for the computer 701. For example, in a non-limiting sense, the mass storage device 704 is a hard disk, a removable magnetic disk, a removable optical disk, a magnetic cassette or other magnetic storage device, a flash memory card, a CD-ROM, a digital versatile disk (digital). versatile disk (DVD) or other optical storage device, random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), etc. May be good.
0065Optionally, any number of program modules, including operating system 705 and data analysis software 706 as an example, may be stored on mass storage 704. Each of the operating system 705 and the data analysis software 706 (or any combination thereof) may include elements of the programming and data analysis software 706. Data 707 may also be stored on mass storage 704. Data 707 may be stored in any of one or more databases known in the art. Examples of such databases include DB2®, Microsoft® Access, Microsoft® SQL Server, Oracle®, mySQL, PostgreSQL and the like. The database may be centralized or distributed across multiple systems.
0066In another aspect, the user may enter commands and information into the computer 701 via an input device (not shown). Examples of such input devices include, but are not limited to, keyboards, pointing devices (eg, "mouse"), microphones, joysticks, scanners, tactile input devices such as gloves, and other body covers. These and other input devices may be connected to processing unit 703 via a man-machine interface 702 coupled to system bus 713, but are also known as parallel ports, game ports, and IEEE1394 ports (also known as firewire ports). , Serial port, or other interface and bus structure such as Universal Serial Bus (USB).
0067In yet another aspect, the display device 711 may also be connected to the system bus 713 via an interface such as the display adapter 709. It is conceivable that the computer 701 may have two or more display adapters 709 and the computer 701 may have two or more display devices 711. For example, the display device may be a monitor, an LCD (Liquid Crystal Display), or a projector. In addition to the display device 711, other output peripherals may include components such as speakers (not shown) and printers (not shown) that can be connected to the computer 701 via the input / output interface 710. Any step and / or result of the method may be output to the output device in any form. Such output may be any form of visual representation including, but not limited to, text, graphical, animated, audio, tactile, and the like.
0068Computer 701 can operate in a networked environment using logical connections to one or more remote computing devices 714a, 714b, 714c. As an example, the remote computing device may be a personal computer, a portable computer, a server, a router, a network computer, a peer device, or another common network node. The logical connection between the computer 701 and the remote computing devices 714a, 714b, 714c may be made via a local area network (LAN) and a general wide area network (WAN). .. Such network connections may be through network adapter 708. The network adapter 708 is feasible in both wired and wireless environments. Such network environments have traditionally been common in offices, enterprise-scale computer networks, intranets, and the Internet 715.
0069For illustration purposes, application programs such as operating system 705 and other executable program components are shown here as separate blocks, but such programs and components are repeatedly stored in different storage components of computing device 701. It is recognized that it exists and is executed by the computer's data processor. An example implementation of data analysis software 706 (eg, a compiled instance of such software) is one of the methods of the present disclosure, such as the exemplary methods presented in FIGS. 19-20 and the related description. One or more may be embodied or included, stored on or transmitted through some form of computer-readable medium. Both of the disclosed methods are embodied in and performed by computer-readable and / or computer-executable instructions embodied on computer-readable media such as system memory 712 or mass storage 704. It may be. For example, in response to the execution of data analysis software 706, processor 703 uses the methods described herein (eg, exemplary methods of FIGS. 19-20) and at least one or more of the disclosed systems. May be realized. The computer-readable medium may be any available medium accessible by the computer or computing device. As an example, in a non-limiting sense, computer-readable media may include "computer storage media" and "communication media". A "computer storage medium" is volatile and non-volatile, removable and realized by any method or technique for storing information such as computer-readable instructions, data structures, program modules, or other data. Includes non-removable media. Illustrative computer storage media include RAM, ROM, EEPROM, flash memory or other memory technologies, CD-ROMs, digital versatile devices.
0070The method and system may employ artificial intelligence techniques such as machine learning and iterative learning. Examples of such techniques are expert systems, case-based reasoning, Bayesian networks, behavior-based AI, neural networks, fuzzy systems, evolutionary computation (eg, genetic algorithms), swarm intelligence (eg, ant algorithms). ), And hybrid intelligent systems (eg, expert inference rules generated through neural networks, or generation rules from statistical learning), but not limited to them.
0071The methods and systems described above enable real-time relevant data mining and visualization. In one aspect, the method and system may manage the associations between datasets, where any data point in the analytical dataset is associated with any other data point in the dataset. The dataset may be hundreds of tables with thousands of fields.
0072In one aspect, a method and system for interacting with a user's disclosed database method and system is provided. In one aspect, a user interface may be generated to facilitate dynamic display generation for viewing the data. As an example, a particular view of a particular dataset or data subset generated for a user may be referred to as a state space or session. The system may include visualization components to dynamically generate one or more visual representations of the data to present in the state space.
0073Figure 8 shows how selection affects the range to generate a subset of the data. The data subset can form a state space, which is based on the selection state given by the selection. In one aspect, the selected state (or "user state") may be defined by the user clicking a list box or graph in the application's user interface. Applications should now host multiple graphical objects (charts, tables, etc.) that calculate one or more mathematical functions (also known as "expressions") for one or more dimensions (classification variables) for a subset of data. It can be designed. The result of this calculation creates a chart result that is a multidimensional cube that can be visualized in one or more graphical objects.
0074The application allows the user to explore the range by making different selections by clicking on the graphical object to select the variable, which results in changes in the chart results. At any point during the investigation, there is a current state space, which is associated with the current selected state acting on the range (which always remains the same).
0075As shown in Figure 8, when the user makes a new selection, the inference engine computes a subset of the data. Also, identifier ID1 for selection and range can be generated based on the filter in selection and range. The identifier ID2 for the data subset is then generated based on the data subset definition, typically the bit sequence, which defines the contents of the data subset. Finally, ID1 may be used as the lookup identifier and ID2 may be cached. Similarly, use ID2 as the lookup identifier to cache the data subset definition.
0076In FIG. 8, the chart calculation is performed in the same manner. Here we have two sets of information: a data subset and related chart properties. The latter is typically a mathematical function that has both computational and classification variables (dimensions), but is not limited to it. Both of these information sets are used to calculate the chart results, and both of these information sets are also used to generate the identifier ID3 for input to the chart calculation. ID2 has already been generated in the previous step, and ID3 is generated as the first step in the chart calculation procedure.
0077The identifier ID3 is formed from ID2 and related chart properties. ID3 is seen as an identifier for a particular chart generation instance that contains all the information needed to calculate a particular chart result. In addition, the chart result identifier ID4 is created from the chart result definition, typically the bit sequence, which defines the chart result. Finally, use ID3 as the lookup identifier and cache ID4. Similarly, use ID4 as the lookup identifier to cache the chart result definition.
0078Multiple graphical objects (or visual representations) have graphs, charts, trees, multidimensional depictions, images (computer-generated or digitally captured), video / audio displays that describe the data, and different data analysis in each area. It may be of virtually any display or output type, including a hybrid representation in which the output is segmented into the display area of. The user may select one or more default visual representations, but the next visual representation may be generated based on further analysis of the form most suitable for the data and the next dynamic selection. As shown in Figure 9, there are several list boxes on the left side of the interface, and on the right side of the user interface there is a graphical object that reflects the selection (or was not selected) in the list box. Has been done. The placement of list boxes and graphical objects is a design choice. In one aspect, the user may select data points, or the visualization component may instantaneously filter and reaggregate other fields and corresponding visual representations based on the user's selection. In one aspect, filtering and reaggregation may be completed without querying the database. In one aspect, the visual representation may be presented to the user using a meaningfully applied color scheme. For example, user selections may be highlighted in green, datasets related to selections may be highlighted in white, and unrelated data may be highlighted in gray. Coloring in a meaningful way provides an intuitive navigation interface in the state space.
0079As shown in Figure 10a, a layout containing several graphical objects is provided to the user. This dataset reflects movie data. For example, movie director, movie title, movie actor, movie length, movie rating, movie release date, and so on. As shown in Figure 10b, once the user selects a director, the graphical objects are dynamically adjusted in real time. In this example, the user has selected the "Emeric Pressburger" director. In response to this selection, all of the graphical objects are adjusted to reflect the data associated with "Emeric Pressburger".
0080In this way, the methods and systems provided allow the user to instantiate a session that can transform raw data into ready-to-use analysis. While one user can manipulate the interface to generate meaningful visual representations, methods and systems are also provided to facilitate collaborative sessions where multiple users can interact with the interface simultaneously or substantially simultaneously. The interface.
0081In one aspect, a user may share his session with one or more other users. As a result, users can discover and deploy new analytics in a real-time collaborative environment. Each user may make choices that are visible to all users. In some cases, constraints may be enforced so that only a few users can make a choice. In yet another example, a temporary list of one user (eg, search, dropdown, etc.) may be hidden from others.
0082In one aspect, two or more users may share a common session. The first time a session is created is called the primary session, while the next users who join are called the secondary session. In one aspect, only the primary session can invite others to join, while in another, any user can invite others to join. The system can be configured such that all aspects of the secondary session mirror all aspects of the primary session. If the primary session has section access reduction, they are mirrored in the secondary session. Section access reduction may be a mechanism that provides data security. For example, when a user clicks on a list box, that user may be restricted from viewing the reduced amount of data to another user who has superior section access. For example, one user may be able to see all the movie directors, while another user can only see one movie director. In one aspect, access rights or data security checks do not apply to the secondary session.
0083All users, both primary and secondary, can share interactions with the user interface that interacts with the system (eg, mouse clicks). Any user who clicks can send the change to one or more of the other clients if the click changes the selection state. Any clicks that affect only the local client and do not involve a message / response from the server are not shared. If two or more clients click "simultaneously", the server clicks 2 each click, much like one client clicks once, then clicks a second time to cancel the first click. It may be treated as one or more asynchronous clicks.
0084In one aspect, the primary user may invite the secondary users to join their session using a panel that drops down from the collaborative toolbar icon. The email invitation may allow the primary user to identify the email address and some additional text that can be placed within the body of the email. When the "Invite" button is pressed, an email may be sent to the recipient to join the session, along with a standard message, additional messages included by the primary user, and a URL.
0085Invitation to join a session can be performed using a specially formatted URL. This URL can link to the system and specific interface workspaces. In addition, this URL can provide additional parameters that are single-use keys for identifying and joining the appropriate session. Once this URL is clicked (for example, sent to the server), it can be disabled, so it can only be used once and cannot be forwarded.
0086When a secondary user joins a session, the primary user may be notified. This notification may be a message attached to the toolbar icon that indicates a change in the state of the collaborative toolbar icon (eg, a color change) and who joined the session. Once a secondary user joins a session, one or more other users may browse the list of users who are currently sharing the session, or delete them in some situations.
0087In another aspect, the primary user may invite the secondary users to join their session using a panel that drops down from the collaborative toolbar icon. An additional option for inviting secondary users is by searching the user directories that have access to the system. Primary users may use directory search results to invite users directly.
0088In one aspect shown in FIG. 11a, 1101 is the step of initiating a primary session for the first user, 1102 is the step of requesting cooperation from the second user, and 1103 is the step of the second user. Interaction by either user, including the step of initiating a secondary session to provide a single state space for collaborative real-time data analysis to the first and second users in 1104. Is reflected in a single state space, providing a method for co-computing.
0089In one aspect shown in Figure 11b, a collaborative session may include a single low-level shared session that can connect to two or more higher-level XML transformers. XML transformers can be connected to each other via synchronous logic. Each XML transformer can connect to the end point of a web session and the other end point can connect to a web browser. As such, commands and selections made by either of the XML transformers may affect the shared low-level session, and state changes may be sent back to both XML transformers. The XML transformer that issued the command may return the state change to the client. The other XML transformer may return the changed state through the client async mechanism.
0090In yet another aspect, methods and systems for time-shift coordination are provided. Within a single state space, users can create and share notes about the various objects contained within that state space. These notes can be shared with one or more other users, who can respond by leaving their own note comments. Each user can save state-space "snapshots" (bookmarks) and data using each memo. The notes may be searchable by the users for efficient access to state space notes and associated snapshots.
0091FIG. 12a shows a graphical object with an attached memo and a memo thread that can be viewed after selecting the memo. Figure 12b shows the change in state space after selection of the saved selection state associated with the memo.
0092As an example, the user adds a new note by right-clicking on an object displayed in the state space and selecting Memo from the context menu, providing the user with a menu option for viewing existing notes. You may. Optionally, all objects in the state space with existing notes may be identified (eg, by icon, color change, etc.). Similarly, the number of attached notes for each object may be displayed. Therefore, the resulting memo may be linked to both the object and the selected state. An object may have one or more memos and one or more memo threads (a series of comments based on the memos). The user may make notes after the user analyzes the dataset and arranges the state space accordingly. The user may choose to attach a snapshot of the current state space to the memo. The system may then create a hidden bookmark and attach it to a note. In one aspect, multiple snapshots of a state space may be associated with a note, for example reflecting a comparison of two different analyzes.
0093To browse the memo and the associated state space, the user may select the desired memo and the memo text will be presented to the user. The user may then choose to add additional information to the memo thread and apply bookmarks to modify the current state space to reflect the state space associated with the memo. In another example, the state space may be automatically updated to reflect the state space associated with the memo when the memo is selected.
0094For notes, permissions may be adjusted to control access to the notes by different classes of users. For example, users in one class may be able to view notes but not create notes, while users in another class may create, edit, and delete notes.
0095The method for time-shift coordination is feasible in various ways. For example, a memo (single memo or memo thread) may be linked to a particular selection state and stored in a single "bookmark". Therefore, one bookmark may contain several notes for each object. By applying this bookmark, these notes will be visible. In yet another example, notes may be linked to several selected states, each memo can correspond to one particular selected state, and all subsequent replies in the memo thread will have the same selected state. May belong to. Selections belonging to a particular note may be stored in a temporary hidden bookmark. In yet another example, the memo may be linked to raw data or data in the input field. For this reason, the memo is viewed as a text entry field.
0096In one aspect shown in FIG. 13a, 1301a is a step of creating a state space that reflects the selected state, 1302a is a step of creating a memo, and 1303a is a step of attaching a memo to an object in the state space, 1304a. Provides a method and system for time-shift collaborative analysis, including a step of saving the selected state and a step of associating the saved selection state with a memo in 1305a.
0097In yet another aspect shown in Figure 13b, 1301b is the step of creating a state space that reflects the selected state, 1302b is the step of creating a memo, and 1303b is the step of attaching a memo to an object in the state space. Methods and systems for time-shift coordinated analysis are provided, including.
0098In yet another aspect shown in Figure 13c, 1301c presents an object in a state space with an attached memo, 1302c receives a selection of memos, and 1303c presents a memo into a memo. Methods and systems for time-shift coordinated analysis are provided, including steps to adjust the state space to reflect the associated conserved selection state.
0099In one aspect, the methods and systems provided allow the user to create multiple states in a single space and apply these states to specific objects in space. The user can make copies of these objects and then put them in different states. Objects in a given state are unaffected by user selection in other states. The methods and systems provided allow users to generate graphical objects that represent different state spaces (and thus different selection states) in a single view.
0100The use of alternative states allows the simultaneous use of multiple selections in space, allowing comparison of those selections in a single visual representation or separate visual representations. The user may select data items for comparative analysis and then make override selections that affect comparative analysis in real time. FIG. 14 shows an exemplary realization of an alternative state.
0101The list box on the left is logically associated with state space X and is located in the state space X container, and the list box on the right is logically associated with state space Y and is located in the state space Y container. There is. In this example, the result graph (chart) shows the results of calculating mathematical functions (mathematical expressions) in both state space X and state space Y. Therefore, the user can define the state space X by clicking in the list box on the left, and display the corresponding calculation result on the result graph. Similarly, the user can define the state space Y by clicking in the list box on the right and display the corresponding calculation result in the result graph.
0102A state identifier for system processing may be assigned to each state. In one aspect, at least two states may be enabled: the default state and the inherited state. The default state may be a state in which many uses occur. Objects may inherit state from higher level objects such as sheets and containers. This means that states are inherited, such as document-sheet-sheet object. Sheets and sheet objects are always in inherited state unless overridden. As an example, a document may be an application document, a sheet may be a tab in such a document, and a container may be an area on a tab that may contain one or more objects. The object may be any text or graphical object, such as a list box, pie chart, bar chart, and so on. Sheets and sheet objects (for example, containers and graphical objects) are always in inherited state, but you can override the inherited state for that sheet or sheet object by associating the sheet or sheet object with a clear state space. It is possible for the user.
0103In one aspect, the lower level may automatically inherit the higher level state space. As shown in Figure 14, if a sheet (eg browse) is assigned a default state space X, all containers and individual objects on this sheet are also associated with this state space unless otherwise specified. Will. Therefore, the user only has to associate the container / object with the state space Y as desired.
0104Charts and other object expressions inherit the state of the object that contains the expression. Charts and object expressions may refer to alternative states. This means that no matter where the expression occurs, the expression may refer to a different state than the object that contains the expression.
0105The method and system use the default state to calculate charts and aggregates by adopting state definitions in terms of per-field selections and defining sets in terms of a subset of rows per table. May drive a subset of the data. This default behavior 1. defines a set of data independent of the current selection to allow for alternative states; 2. multiple through the use of mathematical operators such as merge, intersect, and exclude. It may be modified in two distinct points: combining a set of.
0106The alternative state plays a role in the first part of defining the selection state in which the set can be generated. For processing purposes, the default state may be represented by a "$", while all data may be represented by a "1", regardless of state and selection. The alternative state introduces two additional syntax elements.
01071. An expression may be based on an alternative state. Example: sum ({[Group 1]} Sales) calculates sales based on the selection in state'group 1'.
0108sum ({$} Sales) calculates sales based on the selection in the default state. Both of these formulas may be present in a single chart. This allows users to compare multiple states within a single object. A state reference in an expression overrides the state of the object. FIG. 14 may be seen as an example of such a realization. The state space X may be the default state space (represented by $) and the state space Y may be the state space "group 1". Therefore, the left-hand bar in the result graph may be given by the mathematical function Sum ({$} Sales), while the right-hand bar in the result graph is given by the mathematical function Sum ({[Group 1]} Sales). May be given. This is an example of the fact that wherever an expression occurs, the expression may refer to a different state than the object that contains the expression.
0109Instead of displaying the calculation results for state spaces X and Y in one and the same result graph, they may be displayed in separate graphs. In such an example, one of the graphs would be associated with the expression Sum ({[Group 1]} Sales) and the other graph would be associated with the expression Sum ({$} Sales).
01102. Field selections in one state may be used as modifiers in another state.
0111Example: sum ({[Group 1] <Region = $ :: Region>} Sales) This syntax uses the selections in the Region field from the default states and uses them to modify the state'group 1'. The effect is to keep the region fields "synchronized" between the default state and the'group 1'for this expression. Therefore, the selection in the object associated with the first state space (for example, by the user clicking a value in the list box associated with state space X) is in addition to (or its) in the first state space. It can be used to modify a second state space (for example, state space Y) (instead). In FIG. 14, this means that when the user makes a selection in a particular list box on the left hand side to modify the state space X, the corresponding modification (selection) is automatically made to the state space Y. Can be used to ensure.
0112Set operator with state (+,<sup>*</sup>,-, /) Can be used. The following formula is valid and will count the separate invoice numbers in either the default state or state 1.
0113Example: count ({$ + State 1} DISTINCT [Invoice Number]) counts the separate invoice numbers for the <default> state and the merger of state 1.
0114count ({1-State 1} DISTINCT [Invoice Number]) counts separate invoice numbers that are not in state 1.
0115count ({State 1<sup>*</sup>State 2} DISTINCT [Invoice Number]) counts separate invoice numbers for both <default> state and state 1.
0116For this reason, the method and system provide a way to logically combine data in different state spaces by using logical operators known from the following Boolean algebra: + = Merger (A + B includes all elements of both A and B)<sup>*</sup>= Fellowship (A<sup>*</sup>B contains all the elements of A that also belong to B) -= Difference (AB includes all elements of A that do not belong to B) / = XOR (A / B contains all elements found in only one of A and B).
0117The use of set operators makes it possible to combine data from two or more state spaces into a single equation for display, for example, in a graph.
0118In one aspect shown in FIG. 15, in 1501, a step of presenting a first user interface element associated with a first state space and a second user interface element associated with a second state space, In 1502, it includes a step of receiving a selection in the first and second user interface elements, and in 1503, a step of presenting a result graph showing the selected state of the first state space and the selected state of the second state space. , A method for data analysis is provided. In one aspect, the first state space and the second state space may contain the same data set or may contain different data sets.
0119In one aspect, methods and systems for utilizing dimensional restrictions are provided. Dimensional limits can be set for various chart types, or more generally for most graphical objects described herein. An option called "dimension limit" may be presented to the user to control the number of dimensional values displayed in a given chart or graphical object. The user can select multiple values, for example one of "first N values", "maximum M values" and "minimum ... values". N and M are natural numbers that indicate the concentration of a set of values intended to be provided (or returned). In one aspect, N and M may be provided as options for the dimensional controls "first N values" and "maximum M values". These values, or dimensions, are computers that are encoded (or programmed or configured) in the data analysis software 706 according to the aspects described here, for example, the values that the system can return to the visualization component. You can control how the 701) can sort (for example, the display 711 operates or is configured to operate in response to the execution of the data analysis software 706 by the processor 703). In one aspect, sorting only occurs for the first expression (except for the pivot table if the primary sort can override the one-dimensional sort). In one aspect shown in FIG. 16, one or more user interface elements can be presented to apply one or more dimensional restrictions. For example, a sliding selection tool can be presented to allow the user to apply the "show only ..." dimension limit. The example in FIG. 16 shows that the dimension restriction application only displays the top 6 sales reps.
0120Dimensional restrictions may be applied to generate the data to be displayed in a chart (graph, table, etc.). These dimensional restrictions may include one or more of the following:
0121"Show only ..." This option is selectable if the user wants to see the first, largest, or smallest x values. If this option is set to 5, you will see 5 values. If the dimension enabled Show Other, the segment called Other would occupy one of the five display slots.
0122The "First ... Values" option will return rows based on the option selected on the Sort tab of the Properties dialog. If the chart is a straight table, the rows will be returned based on the primary sort at that time. In other words, the user can change the display of the values by double-clicking on any of the column headers and making that column a primary sort.
0123The "Maximum ... Values" option returns the rows in descending order based on the first formula in the chart. When used in a straight table, the illustrated dimension values will remain constant while the expressions are sorted interactively. If the order of the expressions is changed, the dimensional values will (and may change).
0124The "Minimum ... Values" option returns the rows in ascending order, based on the first expression in the chart. When used in a straight table, the illustrated dimension values will remain constant while the expressions are sorted interactively. If the order of the expressions is changed, the dimensional values will (and may change).
0125"Display only values that are ..." This option is selectable if the user wants to display all dimension values that meet the specified criteria for this option. Choose to display the value based on the overall percentage or the exact quantity. The "Relative to whole" option allows for a relative value mode similar to the "Relative" option on the Expression tab of the Properties dialog. The value may be entered as a calculated expression.
0126"Display only the values that become ... when accumulated" If this option is selected, all rows up to the current row are accumulated and the result is compared to the value set for this option. The "Relative to Whole" option allows for a relative value mode similar to the "Relative Value" option on the Expression tab of the Properties dialog (based on the first value, maximum or minimum value). Compare the cumulative value with the grand total. The value may be entered as a calculated expression.
0127Different display options are also offered, including one or more of the following: "Show other" Enabling this option will generate a segment called "Other" in the chart. All dimension values that do not meet the comparison criteria for display constraints will be grouped into a segment called "Other". If there is a dimension after the selected dimension, "Collapse Inner Dimension" will control whether individual values for the next / inner dimension are displayed on the chart.
0128"Global grouping mode" This option applies only to the internal dimensions. When this option is enabled, constraints will be calculated only for the selected dimensions. All previous dimensions will be ignored. When this is disabled, constraints are calculated based on all preceding dimensions.
0129Using dimension limits with the selected option Show Other is shown in Figure 17a, which contains the variables "Customer," "Product," and "Sales," given for customers A to F and products X and Y. A simplified example based on the dataset shown is described here.
0130Example 1 Suppose a user wants to visualize sales for each "customer". This corresponds to calculating the mathematical function Sum (Sales) for the dimensional variable "customer". This results in the following multidimensional cube (which can be visualized as a graph or table as shown in Figure 17b):
0131Example 2 Here, the user applies the dimension limit "Show only the first three values" to the dimension "Customers" for cube generation, while also checking the "Show others" box. Suppose. This results in the cube shown in Figure 17c. As shown, the sales are shown for customers A and B, but the sales for the remaining customers (C to F) are aggregated to the "Other" value.
0132Example 3 Instead, the user applies the "Show only up to 3 values" dimension limit to the "Customer" dimension for cube generation, while also checking the "Show others" box. Suppose. This results in the cube shown in Figure 17d. As shown, the sales are shown for customers A and C, but the sales for the remaining customers (B and D ~ F) are aggregated to the "Other" value.
0133Example 4 Instead, assume that the user applies the "Show only values greater than or equal to 50" dimension limit to the "Customer" dimension for cube generation, while also checking the "Show others" box. To do. This results in the cube shown in Figure 17e. As shown, the sales are shown for customers A, B and C, but the sales for the remaining customers (D to F) are aggregated to the "Other" value.
0134Example 5 Instead, the user applies the dimension limit "display only the maximum value, which is 80% of the total when accumulated" to the dimension "customer" for the generation of the cube, while the "display other" box. Also assume that you have checked. This results in the cube shown in Figure 17f. As shown, sales are shown for customers A, B, C and F, but sales for the remaining customers (D and E) are aggregated to the "Other" value.
0135All of these examples make use of the calculations described here. It should be understood that the above example has been simplified to facilitate the understanding of dimensional limits. However, in practice, one or more complex mathematical functions may be calculated for large amounts of data connected across many different tables.
0136To sequentially calculate mathematical functions for one or more dimensions (classification variables), the data may be processed in binary coding format by using transformation structures and based on the start table. This is illustrated with reference to Tables 15 and 16 of FIG.
0137Here, Table 15 shows the use of virtual data records that are sequentially updated for each record in the start table, and Table 16 shows how intermediate data structures are populated based on the sequentially updated contents of the virtual data records. Indicates whether it will be done. The intermediate data structure contains aggregate fields used to aggregate the calculation results of the formula for each existing unique combination of classification variable values. In Table 16, the intermediate data structures aggregate the calculation results for the combination of "client" and "year" (0,0), (1,0), (2,0), (3, -2). doing. A value of -2 indicates a null value.
0138Table 16 also shows how dimensions are "excluded" or "folded" in intermediate data structures, which is that formulas are aggregated for all values of one or more classification variables. Means. In this process, additional data records are added to the intermediate data structure to maintain the aggregation of calculation results for the folded dimensions. In Table 16, the intermediate data structures are the data recording (-1,0), (-1, -2) when the "client" is collapsed and the data recording when the "year" is collapsed (-1,2). One data record (-) when both 0, -1), (1, -1), (2, -1), (3, -1) and "client" and "year" are collapsed Includes 1, -1) and. A value of -1 for a variable therefore indicates that the calculation results for all the values of the variable have been aggregated.
0139The data in the intermediate data structure is then used to build a multidimensional cube as shown in Table 17 of FIG. Tables 28 and 29 of FIG. 6 show slightly more advanced examples of the intermediate data structure and the resulting multidimensional cube, respectively. Here, more complex mathematical functions are calculated in the multidimensional cube (Table 29), and intermediate data structures (Table 28) are needed for the correct calculation of the mathematical functions in the multidimensional cubes shown in Tables 28 and 29. Includes an aggregate field that aggregates the calculation results of a certain formula.
0140Returning to Examples 1-5 above, we generate an overall multidimensional cube (see the full table in Example 1 above), in which, for example, the first two customers and their sales data (Example 2). ), Or simply select data such as the two customers with the highest sales and their sales data (eg 3), and it should be recognized that certain dimensional limits are applicable.
0141Difficulties arise when the "other" value should be calculated. This is because this value cannot be defined when a multidimensional cube is created (because its contents are only known once the multidimensional cube is created). The "Other" value corresponds to the aggregation of calculation results for a particular value of one or more classification variables ("Customer" in the example above). In the above example, the math function is a simple addition, and the calculation result of the math function for the "other" value is a simple sum (in the cube) for the "customer" that should be included in the "other" value. It may be obtained by adding to. However, if the mathematical function is more complex, for example if it contains an average quantity (see Tables 28-29 above), the "other" value can be obtained by combining the data in the cube. Can not.
0142One solution is to start the calculation of a new multidimensional cube that contains an aggregate field for a particular value of the classifier that defines the "other" value. In the situation of Example 2, this new cube would be calculated to contain a new "customer" designated as "other" containing the aggregated results for customers C through F.
0143To minimize data processing, the method and system use intermediate data structures (eg, existing or previously populated intermediate data structures) to populate a multidimensional cube with "other" values. Available. As mentioned earlier, aggregate fields in intermediate data structures are defined to allow dimensions to be collapsed (excluded). In some respects, the calculation of "other" values may be considered as a partial exclusion of dimensions in intermediate data structures.
0144Therefore, in Examples 2-4, the dimension limit identifies the value of the variable "customer" to be contained in the cube, along with the corresponding sales. The "other" value of the cube is populated by aggregating the sales for the remaining values of the variable "customer" by thoroughly examining the intermediate data structure.
0145In Example 5, the dimension limit requires that the total sales be known. Total sales data is known only once the intermediate data structure has been generated (corresponding to the exclusion of the dimension "customer"). To populate the "Other" value, the intermediate data structure identifies the maximum value (sales) in the aggregate field for different "customers" until it reaches at least 80% of total sales, and the remaining "customer" sales. To calculate the content of the "other" value by aggregating, it will be thoroughly reviewed once again.
0146There are situations where the "other" value may not be calculated correctly based on the intermediate data structure. For example, when the calculation (stated in US7058621) requires special attention to frequency data. In one embodiment, the method and system include components that detect the potential need for special attention to frequency data. If such a potential need is detected, the method and system may refuse to enter an "other" value. In one variant, the method and system instead use "other" (eg, using a process-intensive alternative that is usually avoided by calculating the value of "other" based on intermediate data structures). You may start calculating a new multidimensional cube containing the value of. In one example, during the generation of a multidimensional cube, whenever the software detects that two or more data records in the intermediate data structure have been updated based on the contents of one virtual data record, it goes to frequency data. The potential need for special attention may be flagged.
0147Example of global grouping mode Assume the multidimensional cube shown in Figure 18a. Here, this cube is generated to calculate sales for the two dimensions (classification variables) "product" and "customer". Now suppose the user applies the dimension limit "Show only the maximum 3 values" to the variable "Customer" while also checking the "Show others" box. This will result in the multidimensional cube shown in Figure 18b.
0148As shown, the process identifies the two customers with the highest sales of product X and the two customers with the highest sales of product Y, and the "other" value for product X. And generate "other" values for product Y. The "Other" value for product X is the cumulative sales for customers C to F, and the "Other" value for product Y is the cumulative sales for customers B and D to F. The "Other" value is generated as described above (eg, by a full review of the intermediate data structure).
0149Instead, suppose the user applies the same dimension limit to the variable "Customer" and also checks the "Global Grouping Mode" box (while checking the "Show Other" box). .. This will result in the multidimensional cube shown in Figure 18c.
0150Global grouping mode allows the process to identify the two customers with the highest sales of all products (eg, combined product X and product Y). The cube is the "other" value, which is the cumulative sales data for product X for these two customers, and the sales for the remaining "customers" (for example, customers B and D to F) for product X. It also includes sales data for product Y for these two customers, and an "other" value that accumulates sales for the remaining "customers" (for example, customers B and D to F) for product Y. Is generated.
0151Thus, the global grouping mode allows dimension limits to be applied only to selected dimensions (customers).
0152In one aspect shown in FIG. 19, and in view of the various features described herein in relation to dimension limits, at 2301, a data processing event is performed on the dataset to execute the first multidimensional cubic data structure. A method for data analysis is provided in 2302 that includes a step that results in a second multidimensional cube data structure by applying one or more dimension limits to the multidimensional cube data structure. .. The first data processing event may include computing mathematical functions for one or more dimensional variables in the dataset. One or more dimension limits may include "display only ...", "display only values that are ...", "display only values that become ... when accumulated", and so on. .. In one aspect, the second multidimensional cube data structure may be displayed according to one or more of "Show Others", "Global Grouping", and so on.
0153An option called "dimension limit" may be presented to the user to control the number of dimension values displayed in a given chart. The user can select multiple values, for example one of "first ... values", "maximum ... values" and "minimum ... values". .. These values control how the system sorts the values that the system returns to the visualization component. In one aspect, sorting only occurs for the first expression (except for the pivot table if the primary sort can override the one-dimensional sort).
0154FIG. 20 shows a flowchart of an exemplary method 2000 for data analysis according to one or more aspects of the present disclosure. A computing device such as the computer 701, or a processor integrated with or functionally coupled to it (such as the processor 703), can implement at least some of the exemplary methods 2000. In 2010, the dataset will be processed to bring the first multidimensional cube data structure. The dataset has a table structure that includes one or more tables. The implementation of 2010 (eg, executed by a processor such as processor 2120 or processor 703) can be referred to as a processing action. In one aspect, processing a dataset to result in a first multidimensional cube data structure involves computing mathematical functions for one or more dimensional variables in a table structure.
0155In 2020, a second multidimensional cube data structure will be generated by applying one or more dimension limits to the first multidimensional cube data structure. In one aspect, applying one or more dimension limits to a first multidimensional cube data structure involves constructing one or more user interface elements. In another aspect, applying one or more dimensional limits to a first multidimensional cube data structure sets the dimensional limits to the selection of a second multidimensional cube data structure in response to a selection of specific display options. Includes applying to the given dimensional variables. In yet another aspect, applying one or more dimensional limits to the first multidimensional cube data structure applies the dimensional limits to the first identification of the second multidimensional cube data structure being displayed. Including bringing in the part of. In yet another aspect, applying one or more dimensional limits to the first multidimensional cube data structure applies display options to the second identification of the second multidimensional cube data structure being displayed. Including further bringing in the part of.
0156In one aspect, the first particular part contains a number of specific rows of the table contained in the first multidimensional cube data structure. In another aspect, the first particular part contains certain rows of a table contained in the first multidimensional cube data structure, and the values of each of the particular rows are cumulative and predetermined. Compared to the value. In yet another aspect, the first particular part contains the first particular rows of the table contained in the first multidimensional cube data structure, the particular rows being arranged in descending order. In yet another aspect, the first particular part contains the first particular rows of the table contained in the first multidimensional cube data structure, the particular rows being arranged in ascending order. In an additional or alternative aspect, the first particular part contains one or more values that satisfy a particular condition.
0157In one embodiment, the first particular part includes aggregated values resulting from aggregating a plurality of values that do not meet a particular condition. In one aspect of such an embodiment, the first multidimensional cube data structure contains the result of calculating a particular mathematical function for one or more computational variables in the dataset, and the first multidimensional cube data structure , All eigenvalues of one or more dimensional variables in the dataset are partitioned.
0158In an additional or alternative aspect of such an embodiment, processing a dataset to result in a first multidimensional cubic data structure reads data items sequentially from one or more tables in the table structure. Each one of one or more data records is implied by a field for each dimensional variable and a particular mathematical function, including that and submitting an intermediate data structure containing one or more data records. Contains aggregate fields for one or more formulas. In one aspect, populating an intermediate data structure containing one or more data records identifies the current value of each dimensional variable for a data item and each of one or more formulas based on the data item. It involves calculating one and aggregating the results of the calculation in the appropriate aggregation field based on the current value of each dimensional variable. In another aspect, the first multidimensional cube data structure is generated by calculating a particular mathematical function based on the contents of the aggregated fields for all the eigenvalues of each dimensional variable. In yet another aspect, the second multidimensional cubic data structure is generated by a full review of the intermediate data structure, thereby being aggregated due to the aggregation of multiple values that do not meet certain conditions. Generate a value.
0159In one embodiment, exemplary method 2000 may include identifying the values of one or more dimensional variables that do not meet certain conditions, based on a first multidimensional cubic data structure. In one aspect, a full review of the intermediate data structure aggregates the contents of the aggregate fields associated with the values of one or more dimensional variables that do not meet certain conditions, and thus does not meet certain conditions. It may include calculating a particular mathematical function to aggregate multiple values.
0160FIG. 21 shows an exemplary computing device 2100 capable of implementing (eg, performing) at least one or more of the methods of the present disclosure. As shown, the computing device 2100 includes a processor 2110 that is functionally coupled to memory 2120 via bus 2115. Processor 703 may embody or embody processor 2110, system memory 712 may embody or embody memory 2120, and bus 713 may embody bus 2115. , Or may be embodied. In one embodiment, the computer-executable instructions contained in the data analysis software 2124 can configure processor 2110 to process the dataset to result in a first multidimensional cubic data structure, with one dataset. It has a table structure including the above tables. Also, in one aspect, such an instruction can configure processor 2110 to apply one or more dimensional limits to a first multidimensional cube data structure to result in a second multidimensional cube data structure. is there. In one aspect, each of the one or more dimensional limits constrains the display number of values of one or more dimensional variables in the second multidimensional cube data structure.
0161In another aspect, processor 2110 can be configured to apply dimensional limits to a second multidimensional cubic data structure in response to a selection of specific display options. In yet another aspect, the processor 2110 can be further configured to apply dimensional restrictions to provide a first particular portion of the second multidimensional cubic data structure being displayed. The first specific part may include a plurality of specific rows of the table contained in the first multidimensional cube data structure. In addition to or instead, the first particular part contains certain rows of the table contained in the first multidimensional cube data structure, and the values of each of the particular rows are accumulated. , Compared to a given value. Further or instead, the first specific part contains the first specific rows of the table contained in the first multidimensional cube data structure, and the specific rows are arranged in descending order. In one scenario, the first particular part contains the first particular rows of a table contained in the first multidimensional cube data structure, and the particular rows are arranged in ascending order. In other scenarios, the first particular part may contain one or more values that satisfy a particular condition.
0162In one aspect, processor 2110 can also be configured to constitute one or more user interface elements. In another aspect, the processor can be configured to apply display options to provide a second particular part of the second multidimensional cube data structure being displayed. In yet another aspect, the processor can also be configured to compute mathematical functions for one or more dimensional variables in the dataset.
0163The methods and systems of the present disclosure have been described in the context of preferred embodiments and specific examples. The scope is not intended to be confined to the particular embodiment described, as the examples here are intended to be exemplary rather than restrictive in all respects.
0164Unless otherwise stated, none of the methods described herein is intended to be construed as requiring the steps to be performed in a particular order. Therefore, it is specifically stated in the claims or description that a method claim does not actually describe the order in which it should be followed by that step, or that the steps should be limited to a particular order. If not, in no respect is it intended that the order should be inferred. This is any possible non-expression for interpretation, including logical matters relating to the construction of steps or action flows, obvious meanings derived from grammatical construction or punctuation, the number or type of examples described herein. This applies to the foundation.
0165It will be apparent to those skilled in the art that various modifications and changes can be made without departing from scope or spirit. Other embodiments will be apparent to those skilled in the art by reviewing the specification and practices disclosed herein. The specification and examples are considered merely by way of illustration and the true scope and spirit is intended to be indicated by the claims.
33 sheets
Sheet 1 Sheet 2 Sheet 3 Sheet 4 Sheet 5 Sheet 6 Sheet 7 Sheet 8 Sheet 9 Sheet 10 Sheet 11 Sheet 12 Sheet 13 Sheet 14 Sheet 15 Sheet 16 Sheet 17 Sheet 18 Sheet 19 Sheet 20 Sheet 21 Sheet 22 Sheet 23 Sheet 24 Sheet 25 Sheet 26 Sheet 27 Sheet 28 Sheet 29 Sheet 30 Sheet 31 Sheet 32 Sheet 33
Every citation, both ways
| Document | Relation | Office |
|---|---|---|
| JP2002539563A | Cites | Japan |
| JP2008059433A | Cites | Japan |
36 members in 5 offices
Members36
| Document | Office | Kind | |
|---|---|---|---|
| CA2851350A1 | Canada | A1 | |
| CA2851351A1 | Canada | A1 | |
| CA2851352A1 | Canada | A1 | |
| WO2013070139A2 | World Intellectual Property Organization (WIPO) | A2 | |
| WO2013070140A1 | World Intellectual Property Organization (WIPO) | A1 | |
| WO2013070141A2 | World Intellectual Property Organization (WIPO) | A2 | |
| US2013159307A1 | United States of America | A1 | |
| US2013159882A1 | United States of America | A1 | |
| US2013159901A1 | United States of America | A1 | |
| WO2013070141A3 | World Intellectual Property Organization (WIPO) | A3 | |
| WO2013070139A3 | World Intellectual Property Organization (WIPO) | A3 | |
| US8745099B2 | United States of America | B2 | |
| EP2776948A2 | European Patent Office (EPO) | A2 | |
| EP2776949A1 | European Patent Office (EPO) | A1 | |
| EP2776950A2 | European Patent Office (EPO) | A2 | |
| JP2014533402A | Japan | A | |
| JP2015501967A | Japan | A | |
| JP2015504548A | Japan | A | |
| US2016026664A1 | United States of America | A1 | |
| JP5947910B2 | Japan | B2 | |
| JP6139546B2This record | Japan | B2 | |
| US9727597B2 | United States of America | B2 | |
| JP6285362B2 | Japan | B2 | |
| US2018150493A1 | United States of America | A1 | |
| US10262017B2 | United States of America | B2 | |
| US10366066B2 | United States of America | B2 | |
| US2020019540A1 | United States of America | A1 | |
| US10685005B2 | United States of America | B2 | |
| US2020349140A1 | United States of America | A1 | |
| CA2851350C | Canada | C | |
| CA2851351C | Canada | C | |
| US11106647B2 | United States of America | B2 | |
| CA2851352C | Canada | C | |
| US11151107B2 | United States of America | B2 | |
| US2022114154A1 | United States of America | A1 | |
| US11580085B2 | United States of America | B2 |
16 legal events, as the office reported them to INPADOC
Over the term
Point at a mark for the eventEvents
| Event | Code | |
|---|---|---|
| Receipt of annual feesJAPANESE INTERMEDIATE CODE: R250R250 | R250 | |
| Receipt of annual feesJAPANESE INTERMEDIATE CODE: R250R250 | R250 | |
| Receipt of annual feesJAPANESE INTERMEDIATE CODE: R250R250 | R250 | |
| Receipt of annual feesJAPANESE INTERMEDIATE CODE: R250R250 | R250 | |
| Receipt of annual feesJAPANESE INTERMEDIATE CODE: R250R250 | R250 | |
| Receipt of annual feesJAPANESE INTERMEDIATE CODE: R250R250 | R250 | |
| Receipt of annual feesJAPANESE INTERMEDIATE CODE: R250R250 | R250 | |
| Certificate of patent or registration of utility modelJAPANESE INTERMEDIATE CODE: R150R150 | R150 | |
| First payment of annual fees (during grant procedure)JAPANESE INTERMEDIATE CODE: A61A61 | A61 | |
| Written decision to grant a patent or to grant a registration (utility model)JAPANESE INTERMEDIATE CODE: A01A01 | A01 | |
| Decision of grant or rejection writtenTRDD | TRDD | |
| Request for written amendment filedJAPANESE INTERMEDIATE CODE: A523A521 | A521 | |
| Written request for extension of timeJAPANESE INTERMEDIATE CODE: A601A601 | A601 | |
| Notification of reasons for refusalJAPANESE INTERMEDIATE CODE: A131A131 | A131 | |
| Report on retrievalJAPANESE INTERMEDIATE CODE: A971007A977 | A977 | |
| Written request for application examinationJAPANESE INTERMEDIATE CODE: A621A621 | A621 |
Numbers
- Publication
- 6139546
- Application
- 2014541001
Titles2
- Japanese
- 多次元立方体データ構造におけるデータ分析のための方法および装置
- English
- Methods and equipment for data analysis in multidimensional cubic data structures
Classification
- CPC, 7
- G06F16/2264
- G06F16/26
- G06F16/22
- G06F16/283
- G06F16/2282
- G06F3/048
- G06F3/04842
- IPC, 1
- G06F17 30
