Similar instance retrieving device
Abstract
[Task] Even when there is a set of attributes that are correlated with each other, a similar case search device that outputs an appropriate similar case based on an operation principle that is easy for the user to understand is provided in consideration of the relationship.
Solution.The case holding means 101 holds a pair of cases described by a predetermined set of attribute values and the value of the class to which each case belongs, and the correlation rule generating means 110 outputs the case holding means 101. Correlation rule is generated and output based on the above case and the value of the corresponding class. The correlation rule holding means 109 holds the correlation rule from the correlation rule creating means 110. Further, the similarity calculation means 103 utilizes the correlation rule output by the correlation rule holding means 109 to output the total similarity between the query which is a new case and the case output by the case holding means 101. After calculation, the similarity case extraction means 111 extracts a case having a high overall similarity output by the similarity calculation means 103 and outputs the case together with the value of the corresponding class.

Term
Term ended
Projected expiry passed 7 November 2020, 5.9 years ago.
- Priority and filed
- Published
- Projected expiry
- Today
10 claims: 1 independent, 9 dependent
- 1【特許請求の範囲】 【請求項1】 予め定められた属性の値の組で記述される事例および各事例が所属するクラスの値を対にして保持する事例保持手段と、前記属性やクラスとこの属性やクラスに対応する値とを対にして構成したアイテムの組で構成される条件部および結論部からなり、条件部と結論部が同時に生起することが多いことを意味する相関ルールを保持する相関ルール保持手段と、新たな事例であるクエリーと前記事例保持手段が出力する前記事例との間の総合類似度を前記相関ルール保持手段が出力する前記相関ルールを利用して算出する類似度算出手段と、前記類似度算出手段が出力する前記総合類似度が高い事例を抽出して対応するクラスの値とともに出力する類似事例抽出手段とを備えたことを特徴とする類似事例検索装置。
- 2【請求項2】 前記類似事例検索装置はさらに、前記事例保持手段が出力する前記事例および対応するクラスの値を入力として、相関ルールを生成して出力する相関ルール生成手段を含み、前記相関ルール保持手段は前記相関ルール生成手段が出力する前記相関ルールを入力として保持することを特徴とする請求項1記載の類似事例検索装置。
- 3【請求項3】 前記類似度算出手段は、前記事例保持手段が出力する前記事例と前記クエリーとの間の事例間類似度を算出する事例間類似度算出手段と、前記相関ルール保持手段が出力する相関ルールの中で、結論部がクラスに対応するアイテムである相関ルールを抽出して出力するクラス含有相関ルール抽出手段と、前記クラス含有相関ルール抽出手段が出力する前記相関ルールと、前記クエリーと、前記事例保持手段が出力する前記事例を入力として、前記クエリーと前記事例が前記相関ルールの条件部を満たすか否かを判定し、前記クエリーと前記事例とが両方とも前記相関ルールの条件部を満たす場合には肯定を意味する条件部判定結果を、そうでない場合には否定を意味する条件部判定結果を出力する条件部判定手段と、前記条件部判定手段が出力する前記条件部判定結果を入力として、前記条件部判定結果が肯定を意味する場合には類似度を増加することを内容とする類似度制御命令を、否定を意味する場合には類似度を変更しないことを内容とする類似度制御命令を生成して出力する類似度制御手段と、前記事例間類似度算出手段が出力する前記事例間類似度と、前記類似度制御手段が出力する前記類似度制御命令を入力として、前記類似度制御命令の内容に従って前記事例間類似度を変更して総合類似度として出力する類似度統合手段とを備えたことを特徴とする請求項1記載の類似事例検索装置。
- 4【請求項4】 前記類似度算出手段はさらに、前記クラス含有相関ルール抽出手段が出力する前記相関ルールと、前記事例保持手段が出力する前記事例に対応するクラスの値を入力として、前記事例に対応するクラスの値が前記相関ルールの結論部であるアイテムに対応するクラスの値と同一であるか否かを判定し、同一である場合は肯定を意味するクラス判定結果を、そうでない場合には否定を意味するクラス判定結果を出力するクラス判定手段を含み、前記類似度制御手段は、前記条件部判定手段が出力する前記条件部判定結果に加えて前記クラス判定手段が出力する前記クラス判定結果を入力として、前記条件部判定結果および前記クラス判定結果がともに肯定を意味する場合には類似度を増加することを内容とする類似度制御命令を、それ以外の場合には類似度を変更しないことを内容とする類似度制御命令を生成して出力することを特徴とする請求項3記載の類似事例検索装置。
- 5【請求項5】 前記類似度算出手段は、前記事例保持手段が出力する前記事例と前記クエリーとの間の事例間類似度を算出する事例間類似度算出手段と、前記相関ルール保持手段が出力する相関ルールの中で、クラスに対応するアイテムを含まない相関ルールを抽出して出力するクラス非含有相関ルール抽出手段と、前記クラス非含有相関ルール抽出手段が出力する前記相関ルールと、前記クエリーと、前記事例保持手段が出力する前記事例を入力として、前記クエリーと前記事例が前記相関ルールの条件部を満たすか否かを判定し、前記クエリーと前記事例とが両方とも前記相関ルールの条件部を満たす場合には肯定を意味する条件部判定結果を、そうでない場合には否定を意味する条件部判定結果を出力する条件部判定手段と、前記クラス非含有相関ルール抽出手段が出力する前記相関ルールと、前記クエリーと、前記事例保持手段が出力する前記事例を入力として、前記クエリーと前記事例が前記相関ルールの結論部を満たすか否かを判定し、前記クエリーと前記事例とが両方とも前記相関ルールの結論部を満たす場合には両方肯定を意味する結論部判定結果を、前記クエリーと前記事例とが両方とも前記相関ルールの結論部を満たさない場合には両方否定を意味する結論部判定結果を、それ以外の場合には不一致を意味する結論部判定結果を出力する結論部判定手段と、前記条件部判定手段が出力する前記条件部判定結果と前記結論部判定手段が出力する前記結論部判定結果を入力として、前記条件部判定結果および前記結論部判定結果に応じて、類似度を増加する,類似度を減少する,類似度を変更しない、のいずれかを決定し、決定した内容の類似度制御命令を生成して出力する類似度制御手段と、前記事例間類似度算出手段が出力する前記事例間類似度と、前記類似度制御手段が出力する前記類似度制御命令を入力として、前記類似度制御命令の内容に従って前記事例間類似度を変更して総合類似度として出力する類似度統合手段とを備えたことを特徴とする請求項1記載の類似事例検索装置。
- 6【請求項6】 前記類似度制御手段は、前記条件部判定手段が出力する前記条件部判定結果と前記結論部判定手段が出力する前記結論部判定結果を入力として、前記条件部判定結果が肯定を意味し、かつ前記結論部判定結果が両方肯定を意味する場合には類似度を増加することを内容とする類似度制御命令を生成して出力することを特徴とする請求項5記載の類似事例検索装置。
- 7【請求項7】 前記類似度制御手段は、前記条件部判定手段が出力する前記条件部判定結果と前記結論部判定手段が出力する前記結論部判定結果を入力として、前記条件部判定結果が肯定を意味し、かつ前記結論部判定結果が両方否定を意味する場合には類似度を増加することを内容とする類似度制御命令を生成して出力することを特徴とする請求項5記載の類似事例検索装置。
- 8【請求項8】 前記類似度制御手段は、前記条件部判定手段が出力する前記条件部判定結果と前記結論部判定手段が出力する前記結論部判定結果を入力として、前記条件部判定結果が肯定を意味し、かつ前記結論部判定結果が不一致を意味する場合には類似度を減少することを内容とする類似度制御命令を生成して出力することを特徴とする請求項5記載の類似事例検索装置。
- 9【請求項9】 前記類似度制御手段は、前記条件部判定手段が出力する前記条件部判定結果と前記結論部判定手段が出力する前記結論部判定結果を入力として、前記条件部判定結果が否定を意味する場合には類似度を変更しないことを内容とする類似度制御命令を生成して出力することを特徴とする請求項5記載の類似事例検索装置。
- 10【請求項10】 前記相関ルール保持手段は相関ルールとともに対応する相関の強さを保持し、前記類似度制御手段は、対応する相関ルールの相関の強さが大きいほど類似度変更の程度が大きくなるように類似度制御命令を生成することを特徴とする請求項3または請求項5に記載の類似事例検索装置。
Independent claims10
253 paragraphs in 1 section, as filed
Description: TECHNICAL FIELD [Detailed description of the invention]
【0001】
[Technical field to which the invention belongs]
In the present invention, when past cases are accumulated as a pair with the value of the class to which the query belongs, the class to which the query belongs is searched for a case similar to a new case (hereinafter referred to as "query"). It relates to a similar case search device that supports value prediction.
【0002】
[Conventional technology]
Regarding the system that predicts the value of the class to which it belongs by searching for similar cases of a given case (query), for example, the document "Nearest Neighbor Method and Memory-Based Reasoning" by Takao Mouri (Journal of the Japanese Society for Artificial Intelligence, It is described in Vol.12, No.2, pp.188-195, March 1997).
【0003】
In this system, it is premised that a case where the value of the class to which it belongs is known is given in advance. When a case (query) that does not know which value class belongs to is given, the work performed by using the system uses the value of the class to which the query belongs, the case given in advance, and the value of the corresponding class. To predict. Here, it is assumed that each case is described by the values of some predetermined attributes. In this system, in order to support the prediction of the value of the class to which the query belongs, the one with high similarity to the query is extracted from the case and output together with the value of the corresponding class. Based on the idea that "the value of the class of similar cases is often the same", the value of the class of the query can be predicted from the value of the output class.
【0004】
In the system described above, since the extracted cases and the values of the corresponding classes are determined by the similarity, the definition of similarity has a great significance for the performance of the system. Since the attributes that make up a case often differ in the degree to which they are useful for classification, research is being actively conducted on how to handle each attribute with a weight according to its importance.
【0005】
However, the method of individually weighting each attribute has a problem that it does not consider the case where there is a set of attributes that are correlated with each other. Therefore, a method of appropriately converting a set of attributes constituting a case into a new set of attributes and then weighting the new attribute is also being studied. For example, the method of using quantification type 2 is described in the document "An Optimal Weighting Criterion of Case Indexing for Both Numeric and Symbolic Attributes" (AAAI-94 Case-Based Reasoning Workshop, pp.123-127, 1994) by Mohri et al. is described.
【0006】
[Problems to be Solved by the Invention]
According to the method of converting the attribute set in advance, it is possible to deal with the case where there is an attribute set having a correlation with each other. However, since the new attribute after conversion is calculated by combining a plurality of original attributes, when the calculation result desired by the user is not sufficiently obtained, the user can track the calculation result in the system. There is a problem that it is often difficult to understand the operation.
【0007】
The present invention has been made to solve the above-mentioned problems, and even when there are a set of attributes that are correlated with each other, the operating principle is easy for the user to understand in consideration of the relationship. It is an object of the present invention to provide a similar case search device that outputs an appropriate similar case based on the above.
【0008】
[Means for solving problems]
The similar case search device according to the first invention includes a case holding means for holding a pair of cases described by a predetermined set of attribute values and the value of the class to which each case belongs, and the attribute or class. It consists of a condition part and a conclusion part consisting of a set of items composed of a pair of values corresponding to this attribute and class, and a correlation rule that means that the condition part and the conclusion part often occur at the same time. Similarity that calculates the total similarity between the holding correlation rule holding means and the query that is a new case and the case output by the case holding means using the correlation rule output by the correlation rule holding means. It is provided with a degree calculation means and a similar case extraction means that extracts a case having a high overall similarity output by the similarity calculation means and outputs it together with a corresponding class value.
【0009】
Further, the similar case search device according to the second invention generates and outputs a correlation rule by inputting the case and the corresponding class value output by the case holding means into the configuration according to the first invention. The correlation rule holding means holds the correlation rule output by the correlation rule generating means as an input.
【0010】
Further, in the similar case search device according to the third invention, the similarity calculation means is a case-to-case similarity calculation means for calculating the case-to-case similarity between the case and the query output by the case holding means. , The class-containing correlation rule extracting means that extracts and outputs the correlation rule whose conclusion part is the item corresponding to the class among the correlation rules output by the correlation rule holding means, and the class-containing correlation rule extracting means that outputs. By inputting the correlation rule, the query, and the case output by the case holding means, it is determined whether or not the query and the case satisfy the condition part of the correlation rule, and the query and the case The condition unit determination means that outputs the condition unit determination result meaning affirmative if both satisfy the condition unit of the correlation rule, and the condition unit determination result meaning negative if both of them satisfy the condition unit, and the condition unit determination The condition unit determination result output by the means is input, and when the condition unit determination result means affirmative, a similarity control command containing an increase in similarity is input, and when the condition unit determination result means negative, the similarity is similar. The similarity control means that generates and outputs a similarity control instruction that does not change the degree, the inter-case similarity that is output by the inter-case similarity calculation means, and the similarity control means that outputs the inter-case similarity control means. It is provided with a similarity integration means that takes the similarity control command as an input, changes the similarity between cases according to the content of the similarity control command, and outputs it as a total similarity.
【0011】
Further, in the similar case search device according to the fourth invention, the similarity calculation means further receives the correlation rule output by the class-containing correlation rule extracting means and the case output by the case holding means as inputs. It is determined whether or not the value of the class corresponding to the case is the same as the value of the class corresponding to the item which is the conclusion part of the correlation rule. In some cases, the similarity control means includes a class determination means that outputs a class determination result meaning denial, and the similarity control means outputs the condition unit determination result output by the condition unit determination means and the class determination means outputs. When the class judgment result is input and the condition part judgment result and the class judgment result both mean affirmative, the similarity control command is used to increase the similarity. In other cases, the similarity is increased. It generates and outputs a similarity control instruction whose content is not to change.
【0012】
Further, in the similar case search device according to the fifth invention, the similarity calculation means is a case-to-case similarity calculation means for calculating the case-to-case similarity between the case and the query output by the case holding means. , The class-free correlation rule extraction means that extracts and outputs the correlation rule that does not include the item corresponding to the class among the correlation rules output by the correlation rule holding means, and the class-free correlation rule extraction means that outputs. By inputting the correlation rule, the query, and the case output by the case holding means, it is determined whether or not the query and the case satisfy the condition part of the correlation rule, and the query and the case If both satisfy the condition part of the correlation rule, the condition part determination result meaning affirmation is output, and if not, the condition part determination result meaning negative is output, and the class is not included. By inputting the correlation rule output by the correlation rule extraction means, the query, and the case output by the case holding means, it is determined whether or not the query and the case satisfy the conclusion part of the correlation rule. When both the query and the case satisfy the conclusion part of the correlation rule, a conclusion part determination result meaning affirmation is obtained, and when both the query and the case do not satisfy the conclusion part of the correlation rule. The conclusion part judgment result which means the denial of both is output, and the conclusion part judgment result which means the disagreement in other cases is output, and the condition part judgment result which is output by the condition part judgment means. Using the conclusion unit determination result output by the conclusion unit determination means as an input, the similarity is increased, the similarity is decreased, or the similarity is not changed according to the condition unit determination result and the conclusion unit determination result. The similarity control means that determines one of the above and generates and outputs a similarity control command of the determined content, the case-to-case similarity that is output by the case-to-case similarity calculation means, and the similarity control means. It is provided with a similarity integration means that receives the output similarity control command as an input, changes the similarity between cases according to the content of the similarity control command, and outputs the total similarity.
【0013】
Further, in the similar case search device according to the sixth invention, the similarity control means inputs the condition unit determination result output by the condition unit determination unit and the conclusion unit determination result output by the conclusion unit determination unit. , When the condition part determination result means affirmation and the conclusion part determination result means both affirmative, a similarity control command is generated and output, which includes increasing the similarity. ..
【0014】
Further, in the similar case search device according to the seventh invention, the similarity control means inputs the condition unit determination result output by the condition unit determination unit and the conclusion unit determination result output by the conclusion unit determination unit. , When the condition part judgment result means affirmative and the conclusion part judgment result means both negative, a similarity control command is generated and output, which includes increasing the similarity. ..
【0015】
Further, in the similar case search device according to the eighth invention, the similarity control means inputs the condition part determination result output by the condition part determination means and the conclusion part determination result output by the conclusion part determination means. , out to generate a similarity control command to the content to reduce the degree of similarity in the case where the condition part determination result indicates a positive, and the conclusion part determination result means mismatch is to force ..
【0016】
Further, in the similar case search device according to the ninth invention, the similarity control means inputs the condition unit determination result output by the condition unit determination unit and the conclusion unit determination result output by the conclusion unit determination unit. When the condition unit determination result means negative, a similarity control instruction is generated and output, which includes not changing the similarity.
【0017】
Further, in the similarity case search device according to the tenth invention, the correlation rule holding means holds the corresponding correlation strength together with the correlation rule, and the similarity control means has a large correlation strength of the corresponding correlation rule. The similarity control instruction is generated so that the degree of similarity change becomes larger.
【0018】
BEST MODE FOR CARRYING OUT THE INVENTION
Hereinafter, embodiments of the present invention will be described with reference to the drawings. As a specific example, the explanation will be given using a case where the weather data at a certain point is an example, and a problem that the weather 3 hours after the measurement of the weather data is a class. Embodiment 1. Figure 1 shows a block diagram of the similar case search device in Embodiment 1. In FIG. 1, the case holding means 101 holds a plurality of cases and the values of the corresponding classes. Here, each case is described as a set of predetermined attribute values. Also, the class takes one of a finite number of predetermined values.
【0019】
Figure 2 shows examples of meteorological data and corresponding class values. Here, the example is described by a set of values of each attribute such as temperature, atmospheric pressure, wind direction, wind speed, and visibility. For example, Case 1 represents a state in which the temperature is 9.5 [degrees], the atmospheric pressure is 996 [hPa], the wind direction is north, the wind speed is 2.3 [m / s], and the visibility is 2 [km]. In addition, the class shows the weather after 3 hours, and the first line of Fig. 2 expresses that the weather was rainy 3 hours after the condition described in Case 1. .. What we want to do in the present invention is to predict the value of the class to which the new case belongs based on the stored past case and the value of the corresponding class when a new case appears. The example of FIG. 2 is equivalent to predicting the weather three hours later based on the past data when new meteorological data is measured. In addition, although the case where the purpose is prediction is described here, when the case expresses the problem and the class expresses the solution to the problem, the appropriate solution is decided when a new problem appears. It can also be used when you want to.
【0020】
In the following, examples will be expressed in the form of a comma-separated list of attribute values. For example, Case 1 in Figure 2 is represented as (9.5, 996, North, 2.3, 2). In addition, when expressing including the class, the value of the class is added to the end of the list by separating it with a semicolon. For example, when Case 1 in Fig. 2 is expressed including the class, it is expressed as (9.5, 996, North, 2.3, 2; Rain).
【0021】
Now, in FIG. 1, the correlation rule holding means 109 holds one or more correlation rules. Here, the correlation rule is a rule that consists of a condition part and a conclusion part composed of a set of items corresponding to an attribute value or a class, and means that the condition part and the conclusion part often occur at the same time. That is. Correlation rules may be given by experts based on their experience, but may be derived from cases as described below.
【0022】
An example of correlation rules for meteorological data is shown below. * Correlation rule 1 "Temperature: 15-17.5, Atmospheric pressure: 990-1000, Wind direction: East Weather after 3 hours: Rain" * Correlation rule 2 "Temperature: 15-17.5, Wind direction: South Visibility: Evil" In Correlation Rule 1, the part "Temperature: 15-17.5, Atmospheric pressure: 990-1000, Wind direction: East" represents the condition part, and the part "Weather after 3 hours: Rain" represents the conclusion part. In Correlation Rule 2, the part "Temperature: 15-17.5, Wind direction: South" represents the condition part, and the part "Visibility: Evil" represents the conclusion part. In other words, Correlation Rule 1 states that "the temperature is in the range of 15-17.5 [degrees], the atmospheric pressure is in the range of 990-1000 [hPa], the wind direction is east, and the weather three hours later is rainy. The situation that the temperature is in the range of 15-17.5 [degrees] and the wind direction is south and the visibility is bad at the same time. It often happens. "
【0023】
In FIG. 1, the similarity calculation means 103 calculates the total similarity between the query 102 and each case in the case holding means 101. The query is a new case input by the user, and is data in the same format as the case held in the case holding means 101. When calculating the total similarity, the similarity calculation means 103 uses not only the values of the query and the case but also the correlation rule in the correlation rule holding means 109. The details of the operation of the similarity calculation means 103 will be described later.
【0024】
The similar case extracting means 111 extracts only the cases having a high overall similarity among the cases output by the similarity calculating means 103, and outputs them together with the values of the corresponding classes. After all, the case with high overall similarity with the query and the value of the corresponding class will be output. Since it is highly likely that similar cases belong to a class with the same value, it is possible to predict the value of the class corresponding to a given query by using the result of the similar case extraction means 111. it can.
【0025】
When calculating the similarity between a query and a case, we use a correlation rule that shows the relationship between multiple attribute values and class values, so not only the individual effects of the attribute values that make up the case, but also multiple specific values. There is an effect that it is possible to output an appropriate similar case for a given query, considering the effect that occurs when the conditions related to the attribute value overlap.
【0026】
By the way, the correlation rule generating means 110 takes the values of a plurality of cases and the corresponding classes output by the case holding means 101 as inputs, and generates a correlation rule that holds for many cases and the values of the corresponding classes. .. The correlation rule holding means 109 holds the correlation rule output by the correlation rule generating means 110 as an input.
【0027】
Here, the operation of the correlation rule generating means 110 will be described. Correlation rule generation is carried out in three stages: pre-processing, mining processing, and post-processing.
【0028】
First, in the preprocessing, a procedure called discretization is performed with respect to the attributes constituting the case according to the instruction separately given by the user, if it is a continuous value attribute. Discretization is to divide the range that a continuous value attribute can take with an appropriate number of boundary values, give a new name to each area, and map the original attribute value to the corresponding area name to make it a discrete value attribute. It is a procedure to do.
【0029】
For example, temperature is a continuous value attribute, so discretization is necessary. If the boundary values are set to "7.5", "10", "12.5", "15", and "17.5" when the temperature is discretized, the values after discretization are, for example, "less than 7.5", "7.5-10", and "10-12.5". "12.5-15" "15-17.5" "17.5 or more" will be 6 pieces. Here, for example, "7.5-10" represents the range of "7.5 or more and less than 10." In the case of the meteorological data shown in Fig. 2, the temperature, atmospheric pressure, wind speed, and visibility are continuous value attributes, so they are discretized. Figure 3 shows the data after appropriate discretization. In FIG. 3, for example, the temperature attribute value 9.5 is discretized to "7.5-10", and the visibility attribute value 0.5 is discretized to "bad" (meaning poor visibility).
【0030】
In the preprocessing, the attribute name is further concatenated with respect to the attribute value name, and the class name is concatenated with respect to the class value. Here, the concatenation of the attribute name and the attribute value name, and the concatenation of the class name and the class value are referred to as "items". For example, the temperature attribute value "7.5-10" is converted to the item "Temperature: 7.5-10". Figure 4 shows an example of the data after converting all the attribute values and class values of the meteorological data in Fig. 3 into items.
【0031】
Although discretization and itemization have been described above, in the preprocessing, other types of data processing, such as deletion of unnecessary data, may be performed if necessary.
【0032】
After applying the pretreatment, the values of each case and class are converted into a sequence of items, for example, as shown in FIG. Such an arrangement of items will be referred to as a "record" here. One record corresponds to one case and class value. The mining process applies an algorithm for generating association rules to this set of records. As a method of generating association rules, for example, there is a method called April by R. Agrawal et al., Reference "Fast Algorithms for Mining Association Rules" (Proc. Of the 20th Int'l Conference on Very Large Databases, 1994), and Japanese Patent Laid-Open No. It is detailed in 8-287106. In these documents, correlation rule generation is based on two indicators called "support" and "confidence".
【0033】
For example A<sub>1</sub>, A<sub>2</sub>, ......, A<sub>k</sub> B Correlation rule expressed in the form of, that is, a set of items (hereinafter referred to as "item set") A<sub>1</sub>, A<sub>2</sub>, ......, A<sub>k</sub>Records that include, and records that also include item B (item A)<sub>1</sub>, A<sub>2</sub>, ......, A<sub>k</sub>In the case of the correlation rule that there is a correlation with (records containing, B), the item set A<sub>1</sub>, A<sub>2</sub>, ......, A<sub>k</sub>The ratio of the number of records including, B to the total number of records is called the support of the correlation rule, and the item set A<sub>1</sub>, A<sub>2</sub>, ......, A<sub>k</sub>The ratio of records that also include item B among the records that include item B is called the certainty of the correlation rule. Here A<sub>1</sub>, A<sub>2</sub>, ......, A<sub>k</sub>The part of is called the condition part, and the part of B is called the conclusion part. The above-mentioned document describes a method of extracting a correlation rule at high speed so that the degree of support and the degree of certainty are equal to or higher than each predetermined lower limit. By performing mining processing on the data shown in FIG. 4, for example, the following correlation rule is generated. * Correlation rule 1 "Temperature: 15-17.5, Atmospheric pressure: 990-1000, Wind direction: East Weather after 3 hours: Rain" * Correlation rule 2 "Temperature: 15-17.5, Wind direction: South Visibility: Evil" [0034]
Post-processing is the process of removing unnecessary correlation rules generated by the mining process. This may be done by an expert based on experience by looking at the correlation rules, but a method has also been devised to automatically detect and delete unnecessary correlation rules. Here, a method of detecting an unnecessary correlation rule will be described by focusing on the strength of the correlation. As the strength of the correlation, for example, the degree of certainty may be taken. It is assumed that the following correlation rule is generated. "A<sub>1</sub>, A<sub>2</sub>, ......, A<sub>n</sub> B " Now consider determining whether this correlation rule is unnecessary. First, the strength of the correlation of this correlation rule is represented by f (a). Furthermore, consider the following form of correlation rule. A<sub>j1</sub>, A<sub>j2</sub>, ......, A<sub>jk</sub> B (where j<sub>1</sub>, j<sub>2</sub>, ..., j<sub>k</sub>Is 1 j<sub>1</sub><j<sub>2</sub><... <j<sub>k</sub>An integer that satisfies <n, 1 k <n) Assuming that m such correlation rules are extracted, the strength of the correlation corresponding to each is determined by f (b).<sub>i</sub>) (Here, i = 1, ..., m). At this time, f (a) f (b)<sub>l</sub>) Satisfies the original correlation rule "A<sub>1</sub>, A<sub>2</sub>, ......, A<sub>n</sub> B " Is judged to be unnecessary and will be deleted.
【0035】
In the above method, "For a certain correlation rule, consider adding an item to the condition part to generate a correlation rule. Even though the item is added to increase the restriction, the strength of the correlation is high. If it did not increase, the correlation strength of the generated correlation rule only reflects the correlation strength of the original correlation rule, so the generated correlation rule can be considered unnecessary. Is based on the idea.
【0036】
Since the appropriate correlation rule is semi-automatically extracted from the set of cases by the above correlation rule generation method, it is not necessary to know in advance the relationship that holds between the attribute values and class values that make up the case. , Has the effect.
【0037】
Next, the similarity calculation means 103 will be described in detail. The similarity calculation means 103 is composed of a case-to-case similarity calculation means 104, a class-containing correlation rule extraction means 105, a condition unit determination means 106, a similarity control means 107, and a similarity integration means 108.
【0038】
The case-to-case similarity calculation means 104 calculates the case-to-case similarity between the case output by the case holding means 101 and the query 102. The degree of similarity between cases is a numerical expression of how similar a case and a query are. If the similarity between cases between case c and query q is expressed as sim (c, q), sim (c, q) can be calculated by, for example, the following formula.
【0039】
[Number 1]
<img file="JP2002149697A_D0001.tif" />【0040】
Where n is the total number of attributes, w<sub>i</sub>Is the weight of the i-th attribute, c<sub>i</sub>Is the value of the i-th attribute in case c, q<sub>i</sub>Is the value of the i-th attribute of query q, and s<sub>i</sub>(c<sub>i</sub>, q<sub>i</sub>) Is the attribute value c for the i-th attribute<sub>i</sub>And q<sub>i</sub>Represents the similarity between attribute values between. w<sub>i</sub>, Also s<sub>i</sub>Must be set appropriately in advance. This s<sub>i</sub>Is, for example, the attribute value c for the i-th attribute<sub>i</sub>And q<sub>i</sub>When is a continuous value, it is obtained based on the difference between the two.
【0041】
Hereinafter, the calculation of the similarity between cases will be shown by a specific example. The similarity between attribute values for each attribute of the meteorological data shown in Fig. 2 shall be determined as follows. First, regarding the temperature, the similarity between the attribute values with respect to the attribute values x and y is [0042]
[Number 2]
<img file="JP2002149697A_D0002.tif" />【0043】
So, regarding atmospheric pressure, the similarity between the attribute values for the attribute values x and y is [0044]
[Number 3]
<img file="JP2002149697A_D0003.tif" />【0045】
And, regarding the wind speed, the similarity between the attribute values for the attribute values x and y is [0046]
[Number 4]
<img file="JP2002149697A_D0004.tif" />【0047】
It shall be calculated in. However, here, it is assumed that the range in which the temperature can change is 20 [degrees], the range in which the atmospheric pressure can change is 50 [hPa], and the range in which the wind speed can change is 5 [m / s]. In addition, it is assumed that the similarity between attribute values related to wind direction and visibility is defined as shown in Fig. 5 and Fig. 6, respectively. In FIG. 5, for example, it is shown that the similarity between the attribute values of "wind direction: north" and "wind direction: east" is 0.5. In addition, FIG. 6 shows that, for example, the similarity between the attribute values of "visibility: 0 [km] to 1 [km] range" and "visibility: 1 [km] to 5 [km] range" is 0.3. Has been done. It is also assumed that the weights of temperature, atmospheric pressure, wind direction, wind speed, and visibility are 1,1,1,1,1 respectively.
【0048】
Here, the calculation of similarity between cases between Case 1 (9.5, 996, North, 2.3, 2) and Query (8.5, 986, East, 1.3, 10) will be described. At this time, based on the above-mentioned definition of similarity between attribute values for each attribute, the similarity between attribute values of temperature, atmospheric pressure, wind direction, wind speed, and visibility is 0.95, 0.8, respectively from the following equations and Fig. 5 and Fig. 6, respectively. It becomes 0.5, 0.8, 0.2.
【0049】
[Number 5]
<img file="JP2002149697A_D0005.tif" />【0050】
By adding the similarity between attribute values in consideration of the weight and dividing by the sum of the weights, the similarity between cases 1 and the query is calculated to be 0.65.
【0051】
By the way, the class-containing correlation rule extracting means 105 of FIG. 1 extracts and outputs a correlation rule such that the conclusion part is an item corresponding to the class in the correlation rule output by the correlation rule holding means 109.
【0052】
The condition part determination means 106 inputs the correlation rule output by the class-containing correlation rule extraction means 105, the query 102, and the case output by the case holding means 101, and whether the query 102 and the case satisfy the condition part of the correlation rule. Judge whether or not. Then, when both the query 102 and the case satisfy the condition part of the correlation rule, the condition part judgment result meaning affirmation is output, and when not, the condition part judgment result meaning negation is output.
【0053】
The similarity control means 107 receives the condition unit determination result output by the condition unit determination unit 106 as an input, and issues a similarity control command having the content of increasing the similarity when the condition unit determination result means affirmative. , When it means negation, it generates and outputs a similarity control instruction that does not change the similarity. Here, as the similarity control instruction whose content is to increase the similarity, for example, "add a numerical value larger than 0" or "multiply a numerical value larger than 1" can be considered. However, the similarity between cases is assumed to be 0 or more.
【0054】
Here, the concept regarding the generation of similarity control instructions will be described. Correlation rules such that the conclusion part is an item corresponding to the class are considered to express the condition of the attribute that has a strong influence on the class. Therefore, if both the query and the case satisfy the condition of the attribute, it is assumed that the query and the case have the same kind of influence on the class, and the similarity is increased so that the case can be easily extracted. To do.
【0055】
The similarity integration means 108 receives the case-to-case similarity output by the case-to-case similarity calculation means 104 and the similarity control command output from the similarity control means 107 as inputs, and the case-to-case similarity according to the content of the similarity control command. Is changed and output as the total similarity.
【0056】
Here, the case where the following correlation rules exist in the correlation rule holding means 109 will be described as an example. * Correlation rule 1 "Temperature: 15-17.5, Atmospheric pressure: 990-1000, Wind direction: East Weather after 3 hours: Rain" * Correlation rule 2 "Temperature: 15-17.5, Wind direction: South Visibility: Evil" Then, the conclusion of Correlation Rule 1 is "Weather after 3 hours: Rain", which is an item related to the class. On the other hand, the conclusion of Correlation Rule 2 is "Visibility: Evil", which is not a class item. Therefore, the class-containing correlation rule extraction means 105 outputs only the correlation rule 1.
【0057】
Now consider the overall similarity calculation between the query (16.2, 996, east, 0.7, 0.5) and the case (17.2, 991, east, 2.3, 2; rain) recorded in the case holding means 101. The inter-case similarity calculation means 104 calculates the inter-case similarity by weighting and adding the similarity between attribute values and dividing by the sum of the weights as described above.
【0058】
On the other hand, the condition unit determining means 106 determines whether or not the query and the case satisfy the condition unit of the correlation rule. The condition part of the above correlation rule 1 is "temperature: 15-17.5, atmospheric pressure: 990-1000, wind direction: east". The query is (16.2, 996, east, 0.7, 0.5) and the case is (17.2, 991, east, 2.3, 2; rain), both of which are the temperatures described in the Conditional Section of Correlation Rule 1. It is included in the range of atmospheric pressure and wind direction. Therefore, since both the query and the case satisfy the condition part of the correlation rule 1, the condition part determination means 106 outputs the condition part determination result meaning affirmation.
【0059】
At this time, the similarity control means 107 outputs the condition unit determination result meaning affirmation from the condition unit determination unit 106, and therefore outputs a similarity control command whose content is to increase the similarity. In response to this, the similarity integration means 108 increases the similarity between cases and outputs it as the total similarity.
【0060】
Next, consider the total similarity calculation between the query (16.2, 996, west, 0.7, 0.5) and the case (17.2, 991, east, 2.3, 2; rain) recorded in the case holding means 101. In this case, since the wind direction is west in the query and the wind direction is east in the condition part of correlation rule 1, this query does not satisfy the condition part of correlation rule 1. Therefore, the condition unit determination means 105 outputs the condition unit determination result meaning denial.
【0061】
At this time, the similarity control means 107 outputs the condition unit determination result meaning denial from the condition unit determination means 106, and therefore outputs a similarity control command containing the content that the similarity is not changed. In response to this, the similarity integration means 108 outputs the inter-case similarity output by the inter-case similarity calculation means 104 as the total similarity as it is.
【0062】
In this way, when there is a correlation rule that means "when multiple conditions overlap, it is often a class of a specific value", and both the query and the case satisfy the multiple conditions, Increase the similarity between the query and its case. As a result, the case is likely to be output as a similar case.
【0063】
By performing the above processing, when it is meaningful that multiple conditions of attributes overlap, the total similarity is calculated in consideration of the overlap, and as a result, appropriate similar cases are likely to be output. It has the effect of becoming.
【0064】
By the way, when the condition part determination result meaning affirmation is output from the condition part determination means 106, the similarity control means 107 generates a similarity control instruction having the content of increasing the similarity. Sometimes it is necessary to determine the degree of increase in similarity. At this time, it is determined that the degree of increase in similarity is large for the correlation rule having a large correlation strength, and the degree of increase in similarity is small for the correlation rule having a small correlation strength. .. As the strength of the correlation, for example, the degree of certainty may be taken. By calculating the degree of similarity change using the strength of the correlation in this way, there is an effect that the total similarity can be calculated appropriately.
【0065】
In the above, we have described the case where there is only one correlation rule to be considered for the pair of query and case for which the total similarity should be calculated. When there are a plurality of target correlation rules, all the similarity control instructions by each correlation rule may be applied. Alternatively, the correlation rule having the strongest correlation may be selected from the target correlation rules, and only the similarity control instruction according to the selected correlation rule may be applied.
【0066】
Embodiment 2. Figure 7 shows a block diagram of the similar case search device in Embodiment 2. The configuration and operation of the parts other than the similarity calculation means 103 are the same as those in the first embodiment. The second embodiment is different from the first embodiment in that the similarity calculation means 103 is further configured to include the class determination means 704. In the similarity calculation means 103, the operations of the case-to-case similarity calculation means 701, the class-containing correlation rule extraction means 702, the condition part determination means 703, and the similarity integration means 706 are the same as those described in the first embodiment.
【0067】
In the class determination means 704 of FIG. 7, the correlation rule output by the class-containing correlation rule extraction means 702 and the value of the class corresponding to the case output by the case holding means 101 are input, and the value of the class corresponding to the case is correlated. Determine if it is the same as the value of the class corresponding to the item that is the conclusion part of the rule. Then, if they are the same, the class determination result meaning affirmative is output, and if not, the class determination result meaning negative is output.
【0068】
The similarity control means 705 uses the class judgment result output by the class judgment means 704 as an input in addition to the condition part judgment result output by the condition part judgment means 703, and both the condition part judgment result and the class judgment result mean affirmative. In some cases, a similarity control instruction whose content is to increase the similarity is generated, and in other cases, a similarity control instruction whose content is not to change the similarity is generated and output.
【0069】
The concept of generating the similarity control instruction is the same as that described in the first embodiment, but in the present embodiment, the case where the similarity is increased is not only the condition part of the correlation rule but also the conclusion part. It is limited to those that also meet.
【0070】
Here, the case where the following correlation rules exist in the correlation rule holding means 109 will be described as an example. * Correlation rule 1 "Temperature: 15-17.5, Atmospheric pressure: 990-1000, Wind direction: East Weather after 3 hours: Rain" * Correlation rule 2 "Temperature: 15-17.5, Wind direction: South Visibility: Evil" As described in the first embodiment, the class-containing correlation rule extraction means 702 outputs only the correlation rule 1.
【0071】
Now consider the overall similarity calculation between the query (16.2, 996, east, 0.7, 0.5) and the case (17.2, 991, east, 2.3, 2; rain) recorded in the case holding means 101. The inter-case similarity calculation means 701 calculates the inter-case similarity by weighting and adding the similarity between attribute values and dividing by the sum of the weights as described above.
【0072】
The condition unit determination means 703 determines whether or not the query and the case satisfy the condition unit of the correlation rule 1. In this case, the condition unit determination means 703 outputs the condition unit determination result meaning affirmation.
【0073】
The class determination means 704 determines whether or not the value of the class corresponding to the case is the same as the value of the class corresponding to the item which is the conclusion part of the correlation rule. Since the class value corresponding to the case is "rain" and the conclusion part of the above correlation rule 1 is "weather after 3 hours: rain", it is determined that the class values are the same. In this case, the class determination means 704 outputs a class determination result meaning affirmation.
【0074】
At this time, the similarity control means 705 increases the similarity because the condition unit determination means 703 outputs the condition unit determination result meaning affirmation and the class determination means 704 outputs the class determination result meaning affirmation. Outputs a similarity control command containing the above. In response to this, the similarity integration means 706 increases the similarity between cases and outputs it as the total similarity.
【0075】
Next, consider the total similarity calculation between the query (16.2, 996, east, 0.7, 0.5) and the case (17.2, 991, east, 2.3, 2; sunny) recorded in the case holding means 101. In this case, unlike the above example, the class value corresponding to the case is "sunny", and the conclusion part of the above correlation rule 1 is "weather after 3 hours: rain", so the class values are the same. It is judged that it is not. In this case, the class determination means 704 outputs a class determination result meaning denial.
【0076】
At this time, since the similarity control means 705 outputs a class determination result meaning negative from the class determination means 704, the similarity is that the similarity is not changed regardless of the output of the condition unit determination means 703. Outputs a degree control command. In response to this, the similarity integration means 706 outputs the inter-case similarity output by the inter-case similarity calculation means 701 as the total similarity as it is.
【0077】
Next, consider the total similarity calculation between the query (16.2, 996, west, 0.7, 0.5) and the case (17.2, 991, east, 2.3, 2; sunny) recorded in the case holding means 101. In this case, since the wind direction is west in the query and the wind direction is east in the condition part of correlation rule 1, this query does not satisfy the condition part of correlation rule 1. Therefore, the condition unit determination means 703 outputs the condition unit determination result meaning denial.
【0078】
At this time, since the similarity control means 705 outputs the condition part determination result meaning negative from the condition part determination means 703, the content is that the similarity is not changed regardless of the output of the class determination means 704. Output the similarity control command. In response to this, the similarity integration means 706 outputs the inter-case similarity output by the inter-case similarity calculation means 701 as the total similarity as it is.
【0079】
In this way, when there is a correlation rule that means "when multiple conditions overlap, it is often a class of a specific value", the query and the case satisfy the multiple conditions, and the case is the case. If the correlation rule is satisfied, increase the overall similarity between the query and the case. As a result, the case is likely to be output as a similar case.
【0080】
By performing the above processing, when it is meaningful that multiple conditions of attributes overlap, the total similarity is calculated in consideration of the overlap, and as a result, appropriate similar cases are likely to be output. It has the effect of becoming.
【0081】
Embodiment 3. Figure 8 shows a block diagram of the similar case search device in Embodiment 3. The configuration and operation of the parts other than the similarity calculation means 103 are the same as those in the first embodiment. In the third embodiment, the class-containing correlation rule extracting means 105 in the first embodiment is replaced with the class-free correlation rule extracting means 802, and the conclusion part determination means 804 is included in the embodiment. It is different from Form 1. In the similarity calculation means 103, the operations of the case-to-case similarity calculation means 801, the condition unit determination means 803, and the similarity integration means 806 are the same as those described in the first embodiment.
【0082】
The class-free correlation rule extracting means 802 of FIG. 8 extracts and outputs a correlation rule that does not include an item corresponding to the class in the correlation rule output by the correlation rule holding means 109.
【0083】
The conclusion part determination means 804 inputs the correlation rule output by the class-free correlation rule extraction means 802, the query 102, and the case output by the case holding means 101, and whether the query and the case satisfy the conclusion part of the correlation rule. If both the query and the case satisfy the conclusion part of the correlation rule, it means that both are affirmative. If both the query and the case do not satisfy the conclusion part of the correlation rule. The conclusion part judgment result which means denial is output to, and the conclusion part judgment result which means disagreement is output in other cases.
【0084】
The similarity control means 805 receives the condition part judgment result output by the condition part judgment means 803 and the conclusion part judgment result output by the conclusion part judgment means 804 as inputs, and has contents according to the condition part judgment result and the conclusion part judgment result. Generates and outputs a similarity control instruction.
【0085】
Here, the concept regarding the generation of similarity control instructions will be described. In general, the correlation rule shows the relationship that holds frequently between the values of attributes, and it is considered that the case where this relationship holds corresponds to the "normal state" and the case where it does not hold corresponds to the "unusual state". It is possible.
【0086】
Therefore, check the establishment status of the correlation rule for each query and case, that is, whether it is a "normal state" or an "unusual state", and increase the similarity so that the case can be easily extracted when the states match. Can be considered. It is also conceivable to reduce the similarity so that the cases are less likely to be extracted when the states do not match. Based on the above idea, the following three methods can be used as a method of generating the similarity control instruction of the contents according to the condition part judgment result and the conclusion part judgment result.
【0087】
In the first method, the similarity control means 805 increases the similarity when the condition part judgment result means affirmative and the conclusion part judgment result means both affirmative. Generates and outputs an instruction.
【0088】
In the second method, the similarity control means 805 increases the similarity when the condition part judgment result means affirmative and the conclusion part judgment result means both negative. Generates and outputs an instruction.
【0089】
In the third method, the similarity control means 805 is a similarity control instruction in which the similarity is reduced when the condition part judgment result means affirmation and the conclusion part judgment result means inconsistency. Is generated and output.
【0090】
Here, as the similarity control instruction whose content is to increase the similarity, for example, "add a numerical value larger than 0" or "multiply a numerical value larger than 1" can be considered. Further, as the similarity control instruction whose content is to reduce the similarity, specifically, for example, "add a numerical value smaller than 0" or "multiply a numerical value smaller than 1" can be considered. However, the similarity between cases is assumed to be 0 or more.
【0091】
By the way, since the above three methods do not compete with each other, they may be used together. At this time, if the condition unit determination result means negation, a similarity control instruction is generated that does not change the similarity. When the condition part judgment result means affirmative, the similarity is increased when the conclusion part judgment result means both affirmative or both negative, and when the conclusion part judgment result means both affirmative, the similarity is decreased. Generate a similarity control command containing the above. In the present embodiment, a specific example will be shown of the case where the above three methods are used in combination.
【0092】
The case where the following correlation rule exists in the correlation rule holding means 109 will be described as an example. * Correlation rule 1 "Temperature: 15-17.5, Atmospheric pressure: 990-1000, Wind direction: East Weather after 3 hours: Rain" * Correlation rule 2 "Temperature: 15-17.5, Wind direction: South Visibility: Evil" Correlation rule 1 then contains the item "Weather after 3 hours: Rain" for the class. Correlation rule 2, on the other hand, does not include items related to the class. Therefore, the class-free correlation rule extraction means 802 outputs only the correlation rule 2. Here, it is assumed that the phenomenon that the visibility is in the range of 0 [km] to 1 [km] corresponds to "visibility: evil".
【0093】
Now consider the overall similarity calculation between the query (16.2, 996, south, 0.7, 0.5) and the case (17.2, 991, south, 2.3, 0.8; rain) recorded in the case holding means 101. The inter-case similarity calculation means 801 calculates the inter-case similarity by weighting and adding the similarity between attribute values and dividing by the sum of the weights as described above.
【0094】
The condition part determination means 803 determines whether or not the query and the case satisfy the condition part of the correlation rule 2 which is the output of the class-free correlation rule extraction means 802. The condition part of the above correlation rule 2 is "temperature: 15-17.5, wind direction: south". The query is (16.2, 996, south, 0.7, 0.5) and the case is (17.2, 991, south, 2.3, 0.8; rain), both of which are the temperatures described in the Conditional Part of Correlation Rule 2. It is included in the range of wind direction. Therefore, since both the query and the case satisfy the condition part of the correlation rule 2, the condition part determination means 803 outputs the condition part determination result meaning affirmation.
【0095】
The conclusion part determination means 804 determines whether or not the query and the case satisfy the conclusion part of the correlation rule 2. The conclusion of Correlation Rule 2 is "Visibility: Evil". On the other hand, the visibility of the query is 0.5 [km] and the visibility of the case is 0.8 [km], both of which correspond to "visibility: evil". Therefore, since it is determined that the query and the case satisfy the conclusion part of the correlation rule 2, the conclusion part determination means 804 outputs the conclusion part determination result which means affirmation.
【0096】
At this time, the similarity control means 805 outputs the condition part determination result meaning affirmative from the condition part determination means 803 and the conclusion part determination result meaning both affirmative from the conclusion part determination means 804. Outputs a similarity control command whose content is to increase. In response to this, the similarity integration means 806 increases the similarity between cases and outputs it as the total similarity.
【0097】
Next, consider the total similarity calculation between the query (16.2, 996, south, 0.7, 10) and the case (17.2, 991, south, 2.3, 20; rain) recorded in the case holding means 101. In this case, the condition unit determination means 803 outputs the condition unit determination result meaning affirmation as in the above example. However, the visibility of the query is 10 [km] and the visibility of the case is 20 [km], both of which do not correspond to the conclusion part of correlation rule 2, "Visibility: Evil". The conclusion part judgment result which means denial is output.
【0098】
At this time, the similarity control means 805 outputs the condition part determination result meaning affirmative from the condition part determination means 803 and the conclusion part determination result meaning both negative from the conclusion part determination means 804. Outputs a similarity control command whose content is to increase. In response to this, the similarity integration means 806 increases the similarity between cases and outputs it as the total similarity.
【0099】
Next, consider the total similarity calculation between the query (16.2, 996, south, 0.7, 0.5) and the case (17.2, 991, south, 2.3, 20; rain) recorded in the case holding means 101. In this case, the condition unit determination means 803 outputs the condition unit determination result meaning affirmation as in the above example. However, the visibility of the query is 0.5 [km] and the visibility of the case is 20 [km]. Regarding the conclusion part of correlation rule 2, "Visibility: Evil", the query corresponds but the case does not, so the conclusion part The determination means 804 outputs a conclusion part determination result meaning a mismatch.
【0100】
At this time, the similarity control means 805 outputs the condition part determination result meaning affirmation from the condition part determination means 803 and the conclusion part determination result meaning disagreement from the conclusion part determination means 804. Outputs a similarity control command whose content is to decrease. In response to this, the similarity integration means 806 reduces the similarity between cases and outputs it as a total similarity.
【0101】
Next, consider the total similarity calculation between the query (16.2, 996, west, 0.7, 0.5) and the case (17.2, 991, south, 2.3, 20; rain) recorded in the case holding means 101. In this case, since the wind direction is west in the query and the wind direction is south in the condition part of correlation rule 2, this query does not satisfy the condition part of correlation rule 2. Therefore, the condition unit determination means 803 outputs the condition unit determination result meaning denial.
【0102】
At this time, the similarity control means 805 outputs the condition part judgment result meaning negative from the condition part judgment result, so that the similarity is not changed regardless of the output of the conclusion part judgment result. Outputs a degree control command. In response to this, the similarity integration means 806 outputs the inter-case similarity output by the inter-case similarity calculation means 801 as the total similarity as it is.
【0103】
As described above, by classifying the states into normal and unusual depending on whether or not the correlation rule is established and reflecting the presence or absence of the difference in the states in the total similarity calculation, it becomes easy to output appropriate similar cases. There is an effect.
【0104】
When both the query and the case are judged to be in the normal state depending on whether or not the correlation rule is established, there is an effect that the total similarity can be appropriately calculated by increasing the similarity between them.
【0105】
When both the query and the case are determined to be in an abnormal state depending on whether or not the correlation rule is established, there is an effect that the total similarity can be appropriately calculated by increasing the similarity between them.
【0106】
When it is determined that one of the query and the case is in the normal state and the other is in the abnormal state depending on whether or not the correlation rule is established, the total similarity can be calculated appropriately by reducing the similarity between them. effective.
【0107】
Now, when generating the similarity control instruction, it is necessary to determine the degree of change in the similarity. At this time, it is determined that the degree of change in similarity is large for the correlation rule having a large correlation strength, and the degree of change in the similarity is small for the correlation rule having a small correlation strength. .. As the strength of the correlation, for example, the degree of certainty may be taken. By calculating the degree of similarity change using the strength of the correlation in this way, there is an effect that the total similarity can be calculated appropriately.
【0108】
In the above, we have described the case where there is only one correlation rule to be considered for the pair of query and case for which the total similarity should be calculated. When there are a plurality of target correlation rules, all the similarity control instructions by each correlation rule may be applied. Alternatively, the correlation rule having the strongest correlation may be selected from the target correlation rules, and only the similarity control instruction according to the selected correlation rule may be applied.
【0109】
[Effect of the invention]
As described above, according to the first and second inventions, the total similarity between the query, which is a new case, and the pre-stored case is calculated by using the correlation rule, and the total similarity is calculated. Since a high case is extracted and output together with the value of the corresponding class, it is possible to provide an appropriate material for the user to predict the value of the class corresponding to the query.
【0110】
Also, according to the third and fourth inventions, if the query and the case satisfy the condition part of the correlation rule that includes the item corresponding to the class in the conclusion part, the similarity is increased, otherwise the similarity is increased. Since the degree is not changed, it has the effect of facilitating the output of appropriate cases.
【0111】
Further, according to the fifth invention, regarding the correlation rule that does not include the item corresponding to the class, the conclusion part determining means means affirmation when both the query and the case satisfy the conclusion part of the correlation rule. If both the query and the case do not satisfy the conclusion part of the correlation rule, the conclusion part judgment result means denial, otherwise the conclusion means disagreement. The unit determination result is output, and the similarity control means either increases the similarity, decreases the similarity, or does not change the similarity according to the condition unit determination result and the conclusion unit determination result. Since it is decided, it has the effect of facilitating the output of appropriate cases.
【0112】
Further, according to the sixth invention, regarding the correlation rule that does not include the item corresponding to the class, for the query and the case, the similarity control means means that the condition part judgment result is affirmative and the conclusion part judgment result. When both means affirmative, a similarity control instruction having the content of increasing the similarity is generated and output, so that an appropriate case can be easily output.
【0113】
Further, according to the seventh invention, regarding the correlation rule that does not include the item corresponding to the class, for the query and the case, the similarity control means means that the condition part judgment result is affirmative and the conclusion part judgment result. When both of them mean denial, a similarity control instruction having the content of increasing the similarity is generated and output, so that an appropriate case can be easily output.
【0114】
Further, according to the eighth invention, regarding the correlation rule that does not include the item corresponding to the class, for the query and the case, the similarity control means means that the condition part judgment result is affirmative and the conclusion part judgment result. When indicates a disagreement, a similarity control instruction containing the content of reducing the similarity is generated and output, so that an appropriate case can be easily output.
【0115】
Further, according to the ninth invention, regarding the correlation rule that does not include the item corresponding to the class, the similarity control means determines the similarity for the query and the case when the condition part determination result means negation. Since the similarity control instruction whose content is not changed is generated and output, it has the effect of facilitating the output of an appropriate case.
【0116】
Further, according to the tenth invention, the similarity control means generates a similarity control instruction so that the greater the correlation strength of the corresponding correlation rule, the greater the degree of similarity change. Has the effect of being able to calculate appropriately.
[Simple explanation of drawings]
[Figure 1]
It is a block diagram of the similar case search apparatus in Embodiment 1 of this invention.
[Figure 2]
It is a figure which shows the example about the meteorological data and the example of the value of the corresponding class.
[Fig. 3]
It is a figure which shows the data after discretizing the continuous value attribute of the meteorological data shown in FIG.
[Fig. 4]
It is a figure which shows the example of the data after all the attribute value and the class value of the meteorological data of FIG. 3 are converted into an item.
[Fig. 5]
It is a figure which shows the similarity between the predetermined attribute values about a wind direction.
[Fig. 6]
It is a figure which shows the similarity between the predetermined attribute values about visibility.
[Fig. 7]
It is a block diagram of the similar case search apparatus in Embodiment 2 of this invention.
[Fig. 8]
It is a block diagram of the similar case search apparatus in Embodiment 3 of this invention.
[Explanation of symbols]
101 Case retention means, 103 Similarity calculation means, 104 Case-to-case similarity calculation means, 105 Class-containing correlation rule extraction means, 106 Condition part judgment means, 107 Similarity control means, 108 Similarity integration means, 109 Correlation rule holding means , 110 Correlation rule generation means, 111 Similarity case extraction means, 704 Class judgment means, 705 Similarity control means, 802 Class-free correlation rule extraction means, 804 Conclusion part judgment means, 805 Similarity control means.
3 sheets
Sheet 1 Sheet 2 Sheet 3
Every citation, both ways
| Document | Relation | Office | Cited during |
|---|---|---|---|
| WO2020026643A1 | Cited by | World Intellectual Property Organization (WIPO) | International search |
| US7734639B2 | Cited by | United States of America | Applicant |
| US9811513B2 | Cited by | United States of America | Applicant |
| US10275348B2 | Cited by | United States of America | Applicant |
| US7444332B2 | Cited by | United States of America | Search report |
| JP2006107299A | Cited by | Japan | Examiner |
| JP2010210245A | Cited by | Japan | Search report |
| JP2009187268A | Cited by | Japan | Search report |
| JP2010092432A | Cited by | Japan | Search report |
| JP2009087141A | Cited by | Japan | Examiner |
| JPWO2020026643A1 | Cited by | Japan | Search report |
| US7440945B2 | Cited by | United States of America | Search report |
| US10229043B2 | Cited by | United States of America | Applicant |
| US12217138B2 | Cited by | United States of America | Applicant |
| JP2019008640A | Cited by | Japan | Search report |
| US7310639B2 | Cited by | United States of America | Search report |
| JP2004206167A | Cited by | Japan | Search report |
1 member in 1 office
Priority claims2
| Document | Office | Kind | Date |
|---|---|---|---|
| 2000339116 | Japan | A | |
| JP20000339116 | – | – | – |
Members1
| Document | Office | Kind | |
|---|---|---|---|
| JP2002149697AThis record | Japan | A |
Numbers
- Publication
- 2002-149697
- Publication, DOCDB
- 2002149697
- Publication, EPODOC
- JP2002149697
- Application
- 339116
- Application, DOCDB
- 2000339116
- Application, EPODOC
- JP20000339116
Titles2
- Japanese
- 【発明の名称】類似事例検索装置
- English
- [Title of Invention] Similar Case Search Device
Classification
- IPC, 3
- G06F17 30
- G06F9 44
- G06N5 04