Evaluating distinctiveness of document
Abstract
This record has no abstract on file.
Term
Term ended
Expired 4 July 2022, 4.2 years ago.
- Priority and filed
- Granted
- Expired
- Today
6 claims: 6 independent, 0 dependent
- 1一つ以上 の文書セグメントから成る比較文書 T に対する一つ以上の文書セグメントから成る着目文書 D に含まれる用語の特有度を評価し て 、特有な用語を選択する方法 であって、 (a)前記比較文書 T と前記着目文書 Dと に含まれる前記文書セグメント毎に、前記文書セグメントに出現する用語の出現頻度に関連した値を成分とする文書セグメントベクトル t k ,d k であって、 前記比較文書Tに含まれるk番目の文書セグメントの文書セグメントベクトルt k は、t k =(t k1 ,...,t kJ ) T と表され(Tはベクトルの転置を表わし、Jは前記着目文書Dと前記比較文書Tに現れる用語の種類数の大きいほうを表す)、 前記着目文書Dに含まれるk番目の文書セグメントの文書セグメントベクトルd k は、d k =(d k1 ,...,d kJ ) T と表される 文書セグメントベクトルを生成するステップと、 (b)前記文書セグメントベクトル d k ,t k より、 前記比較文書 T と前記着目文書 D に対応する平方和行列 S T ,S D それぞれ を生成するステップと、 (c)前記比較文書 T と前記着目文書 Dと に対応する平方和行列 S T ,S D から着目文書 Dのi次のトピック差分因子ベクトルであって、S D α=λS T αなる一般固有値問題のi次の固有ベクトルα i により計算されるi次の トピック差分因子ベクトル を、一定次数Lまで求めるステップと、 (d)前記着目文書 D 及び前記比較文書 T の各文書セグメントに対して、対応する文書セグメントベクトル d k ,t k それぞれと、前記一定次数Lまでのi次のトピック差分因子ベクトルそれぞれと の内積の値 y ki ,z ki それぞれを 求めるステップと、 (e)前記着目文書 Dと前記比較文書Tと に含まれる各用語の 文書セグメントベクトルd k ,t k それぞれと、前記一定次数Lまでのトピック差分因子ベクトルそれぞれとの内積の値y k i ,z ki それぞれの大きさに対応する前記一定次数Lまでのi次の特有度distinc(w j ,i)、及び、前記一定次数までのi次の特有度の合計に対応する総合特有度distinc(w j )を 求めるステップと、 (f) 前記一定次数Lまでのi次の特有度distinc(w j ,i)それぞれ または総合特有度 distinc(w j )に基づいて、 着目文書 D に特有な用語を選択するステップ と をコンピュータに実行させる方法であって、 前記着目文書Dおよび前記比較文書Tに含まれる各用語の前記一定次数Lまでのi次の特有度distinc(w j ,i)それぞれは、各用語の各文書セグメントにおける頻度と、前記文書セグメントベクトルd k ,t k と前記一定次数Lまでのi次の特有度distinc(w j ,i)それぞれとの内積値y ki ,z ki との間の相関係数の絶対値もしくは2乗値により決定され、総合特有度distinc(w j )は各次の特有度を一定次数加えて決定されること を特徴とする方法。
- 2前記着目文書 D において、 前記文書セグメント d k の数がM個であり、 前記kは、k=1,..,Mと表され、 d kj は前記文書セグメントに出現するj番目の用語の出現頻度に関連した値を表わす、 とした場合、 前記着目文書 D の平方和行列 S D が、 となるように求められ、 前記比較文書 T において、前記文書セグメントの数が Nであり、 t kj は前記文書セグメントに出現するj番目の用語の出現頻度に関連した値を表わす、 とした場合、 前記比較文書 T の平方和行列 S T が で計算されること を特徴とする請求項 1 に記載の方法。
- 3一つ以上 の文書セグメントから成る比較文書 T に対する一つ以上の文書セグメントから成る着目文書 D に含まれる文書セグメントの特有度を評価し、特有な文書セグメントを選択する方法 であって、 (a)前記比較文書 T と前記着目文書 Dと に含まれる前記文書セグメント毎に、前記文書セグメントに出現する用語の出現頻度に関連した値を成分とする文書セグメントベクトル t k ,d k であって、 前記比較文書Tに含まれるk番目の文書セグメントの文書セグメントベクトルt k は、t k =(t k1 ,...,t kJ ) T と表され(Tはベクトルの転置を表わし、Jは前記着目文書Dと前記比較文書Tに現れる用語の種類数の大きいほうを表す)、 前記着目文書Dに含まれるk番目の文書セグメントの文書セグメントベクトルd k は、d k =(d k1 ,...,d kJ ) T と表される 文書セグメントベクトルを生成するステップと、 (b)前記着目文書 D の各文書セグメントに対して、対応する文書セグメントベクトル d k と着目文書 D との類似度 sim(D, d k )及び比較文書Tとの類似度sim(T, d k )を、下式により 求めるステップと、 (c)前記着目文書 D の各文書セグメントに対して、前記着目文書 D との類似度 sim(D, d k ) 及び前記比較文書 T との類似度 sim(T, d k ) を用いて、 下式により特有度distinc(d k ) を求めるステップと、 (d) 前記特有度distinc(d k )に基づいて 着目文書 D に特有な文書セグメントを選択するステップ と をコンピュータに実行させる方法。
- 4一つ以上 の文書セグメントから成る比較文書 T に対する一つ以上の文書セグメントから成る着目文書 D に含まれる文書セグメントの特有度を評価し、特有な文書セグメントを選択する方法 であって、 (a)前記比較文書 T と前記着目文書 Dと に含まれる前記文書セグメント毎に、前記文書セグメントに出現する用語の出現頻度に関連した値を成分とする文書セグメントベクトル t k ,d k であって、 前記比較文書Tに含まれるk番目の文書セグメントの文書セグメントベクトルt k は、t k =(t k1 ,...,t kJ ) T と表され(Tはベクトルの転置を表わし、Jは前記着目文書Dと前記比較文書Tに現れる用語の種類数の大きいほうを表す)、 前記着目文書Dに含まれるk番目の文書セグメントの文書セグメントベクトルd k は、d k =(d k1 ,...,d kJ ) T と表される 文書セグメントベクトルを生成するステップと、 (b)前記着目文書 D の各文書セグメントに対して、対応する文書セグメントベクトル d k と着目文書 D との類似度 sim(D, d k )及び比較文書Tとの類似度sim(T, d k )を、下式により求めるステップであって、 前記着目文書Dの平均文ベクトル及び前記比較文書Tの平均文ベクトルそれぞれを、 としたときに、 前記着目文書Dの類似度sim(D, d k )は、下式 により求められ、 前記比較文書Tの類似度sim(T, d k )は、下式 により求められる ステップと、 (c)前記着目文書 D の各文書セグメントに対して、前記着目文書 D との類似度 sim(D, d k ) 及び前記比較文書 T との類似度 sim(T, d k ) を用いて、 下式により特有度distinc(d k ) を求めるステップと、 (d) 前記特有度distinc(d k ) から着目文書 D に特有な文書セグメントを選択するステップ と をコンピュータに実行させる方法。
- 5一つ以上 の文書セグメントから成る比較文書に対して一つ以上の文書セグメントから成る着目文書に含まれる用語の特有度を評価し、特有な用語を選択する方法 であって、 (a) 前記比較文書Tと前記着目文書Dとに含まれる前記文書セグメント毎に、前記文書セグメントに出現する用語の出現頻度に関連した値を成分とする文書セグメントベクトルt k ,d k であって、 前記比較文書Tに含まれるk番目の文書セグメントの文書セグメントベクトルt k は、t k =(t k1 ,...,t kJ ) T と表され(Tはベクトルの転置を表わし、Jは前記着目文書Dと前記比較文書Tに現れる用語の種類数の大きいほうを表す)、 前記着目文書Dに含まれるk番目の文書セグメントの文書セグメントベクトルd k は、d k =(d k1 ,...,d kJ ) T と表される 文書セグメントベクトルを生成するステップと、 (b) 前記着目文書Dの各文書セグメントに対して、対応する文書セグメントベクトルd k と着目文書Dとの類似度sim(D, d k )及び比較文書Tとの類似度sim(T, d k )を、下式により求めるステップと、 (c) 前記着目文書Dの各文書セグメントに対して、前記着目文書Dとの類似度sim(D, d k )及び前記比較文書Tとの類似度sim(T, d k )を用いて、下式により特有度distinc(d k )を求めるステップと、 (d) 前記比較文書Tの各文書セグメントに対して、前記着目文書Dとの類似度sim(D, t k )及び前記比較文書Tとの類似度sim(T, d k )を用いて特有度distinc(t k )を求めるステップと、 (e)Mを着目文書Dの文書セグメントの数とし、Nを前記比較文書Tの文書セグメントの数として、下式に基づいて、用語w j の特有度distinc(w j ) を求めるステップと、 (f) 前記特有度distinc(w j )に基づいて、 着目文書に特有な用語を選択するステップと をコンピュータに実行させる方法。
- 6(a) 前記比較文書Tと前記着目文書Dとに含まれる前記文書セグメント毎に、前記文書セグメントに出現する用語の出現頻度に関連した値を成分とする文書セグメントベクトルt k ,d k であって、 前記比較文書Tに含まれるk番目の文書セグメントの文書セグメントベクトルt k は、t k =(t k1 ,...,t kJ ) T と表され(Tはベクトルの転置を表わし、Jは前記着目文書Dと前記比較文書Tに現れる用語の種類数の大きいほうを表す)、 前記着目文書Dに含まれるk番目の文書セグメントの文書セグメントベクトルd k は、d k =(d k1 ,...,d kJ ) T と表される 文書セグメントベクトルを生成するステップと、 (b) 前記着目文書Dの各文書セグメントに対して、対応する文書セグメントベクトルd k と着目文書Dとの類似度sim(D, d k )及び比較文書Tとの類似度sim(T, d k )を、下式により求めるステップであって、 前記着目文書Dの平均文ベクトル及び前記比較文書Tの平均文ベクトルそれぞれを、 としたときに、 前記着目文書Dの類似度sim(D, d k )は、下式 により求められ、 前記比較文書Tの類似度sim(T, d k )は、下式 により求められる ステップと、 (c) 前記着目文書Dの各文書セグメントに対して、前記着目文書Dとの類似度sim(D, d k )及び前記比較文書Tとの類似度sim(T, d k )を用いて、下式により特有度distinc(d k )を求めるステップと、 (d) 前記比較文書Tの各文書セグメントに対して、前記着目文書Dとの類似度sim(D, t k )及び前記比較文書Tとの類似度sim(T, d k )を用いて特有度distinc(t k )を求めるステップと、 (e) Mを着目文書Dの文書セグメントの数とし、Nを前記比較文書Tの文書セグメントの数として、下式に基づいて、用語w j の特有度distinc(w j )を求めるステップと、 (f) 前記特有度distinc(w j )に基づいて、 着目文書に特有な用語を選択するステップと をコンピュータに実行させる方法。
Independent claims6
1 paragraph, as filed
[0001] [Industrial application field] The present invention relates to natural language processing such as document summarization, and is particularly peculiar to one document or a component (sentence, term, phrase, etc.) of a document set when comparing two documents or a document set. By making it possible to quantitatively evaluate the properties, the performance of the processing is improved. [0002] [Conventional technology] When comparing two documents or a set of documents, the process of extracting the difference between the two documents is an important process in multi-document summarization and the like. Here, a pair of documents will be described, a document in which differences are extracted will be referred to as a document of interest, and a comparison partner thereof will be referred to as a comparison document. In the past, both the document of interest and the comparison document were divided into small elements, and then the elements were matched against each other, and the elements that could not be matched were made into different parts. Elements include sentences, paragraphs, and individual areas when the document is divided by automatically extracted topical change points. In such a case, a vector space model is often used for matching elements. When each element is represented by a vector space model, each component of the vector corresponds to each term appearing in the document, and the value of each component is given the frequency of each term in each element, or the amount associated therewith. [0003] Further, the good or bad of the correspondence between the elements can use the cosine similarity between the vectors, and when the cosine similarity is higher than the predetermined threshold value, it is determined that the elements correspond to each other. Therefore, among the elements in the document of interest, the elements whose similarity is smaller than the threshold value for any of the elements in the comparison document are regarded as different parts. Another method is also known in which both documents are represented using a graph, the correspondence between the graph elements is obtained, and the difference is obtained from the uncorresponding graph elements. [0004] [Problems to be Solved by the Invention] By the way, it seems that there are the following two viewpoints for the extraction of different parts. That is, (A) Extract the parts where the information to be represented is different. (B) Extract the part that reflects the difference in the concept that both documents represent as a whole document. Is. Looking at the conventional multi-document summarization method from this point of view, most of them are based on (A), only the difference between the two documents is extracted, and the importance of the difference in the document of interest is evaluated. Not. Therefore, it was possible that a part that was not very important as information was extracted as a difference part simply because it was different from the comparison document. From the standpoint of (B) above, the present invention makes it possible to extract a difference portion that satisfies the following conditions. [0005] That is, (1) The differences extracted from the document of interest are at the same time important parts of the document of interest. That is, there is a balance between difference and importance. It seems more appropriate to express the difference part that satisfies this condition as a peculiar part of the document of interest rather than just a difference part. Therefore, from now on, the difference part satisfying this condition will be referred to as a peculiar part. (2) For each sentence of the document of interest, an evaluation value is required for the degree of peculiarity. (3) An evaluation value is required for the specificity of terms and term sequences so that it is possible to explain what kind of terms or term sequences are the main factors for the extracted peculiar parts. Is. [0006] [Means for solving problems] The first means for realizing the peculiarity evaluation method of the document of interest that satisfies the above conditions is as follows. As a first embodiment, a method of extracting a document segment having a high degree of specificity from a document of interest will be described. First, both documents are divided into document segments of appropriate units, and a vector of the document segment whose component is the frequency of appearance of terms appearing in the document segment is generated. Since the most natural document segment is a sentence, we will continue to explain the sentence as a document segment. Therefore, both documents are represented as a set of sentence vectors. Next, assuming that the full-text vectors of both documents are projected on a certain projection axis, a projection axis that maximizes (sum of projection values from the document of interest) / (sum of squares of projection values from the comparison document) is set. Ask. For such a projection axis, the sum of squares of the projection values of the sentence vector of the document of interest is large, and that of the comparison document is small, so information that is abundant in the document of interest and difficult to exist in the comparison document is reflected. Will be done. As a result, when a sentence vector is projected on the projection axis, if the content of the sentence is different from that of the comparison document, the absolute value of the projection value becomes large in the document of interest, and it can be used as a base for calculating the peculiarity of each sentence of the document of interest. [0007] As a second embodiment, a method of selecting a term having a high degree of specificity will be described. As for terms, it is assumed that the correlation between the frequency of the word of interest in each sentence and the peculiarity of each sentence is obtained, and the term having a large correlation value is selected. Such terms should appear only in highly specific sentences and can be regarded as unique terms. Therefore, it is possible to obtain the peculiarity of a term based on this correlation value. In addition, the degree of peculiarity of term sequences such as phrases and patterns appearing in the document of interest is evaluated in the same way as for sentences and terms. For each term series, for example, by creating a vector such that the component corresponding to the term included in the term series of interest is 1 and the others are 0, the specificity of each term series is obtained by the method of obtaining the sentence specificity. be able to. Alternatively, if the frequency of each term series in each sentence is obtained, the specificity of the term series can be evaluated by using the frequency of each term series instead of the frequency of each term in the method of obtaining the term specificity. it can. [0008] Further, in the present invention, the following second means will be disclosed as Example 3 in order to evaluate the peculiarity of the document of interest. Here, too, the sentence is explained as a document segment. The second means is the same as the first means until the vector of the document segment is generated, but after that, for each sentence of the document of interest, the degree of similarity with the entire document of interest and the degree of similarity with the entire comparative document are determined. Ask. The important sentences in the document of interest have a large degree of similarity with the entire document of interest, and the sentences having different contents from the comparative document have a small degree of similarity with the entire comparison document. Therefore, by using (similarity with the entire document of interest) / (similarity with the entire comparative document), a balanced peculiarity between difference and importance can be defined. Further, the specificity of the term can be obtained by obtaining the correlation between the specificity of each sentence and the frequency of the term as in the first means. For each term series, the peculiarity can be calculated by obtaining the similarity between the vector obtained from the term sequence and the entire document of interest and the similarity with the entire comparison document in the same manner as in the first means. It is also possible to calculate the specificity of each term series from the correlation between the frequency of each term series in each sentence and the specificity of each sentence. [0009] According to the present invention, when two documents are compared, the degree of peculiarity can be obtained for each sentence, each phrase, and each word constituting one of the documents of interest. If the comparison document and the document of interest are newspaper articles that describe the same case, for example, it is possible to extract sentences that describe a topic different from the comparison document by selecting a sentence with a high degree of specificity from the document of interest. become able to. For example, regarding a certain traffic accident, the comparative document describes "accident summary" and "victim, perpetrator", and the focus document describes "police opinion" in addition to "accident summary". In the document of interest, the specificity of the sentence related to the "police opinion" becomes high, and the part related to the "police opinion" can be extracted. If the user has already read the comparison document, the user can extract and read only the part of the "police opinion" that is unknown to the user. This makes it possible to improve the efficiency of information acquisition. Further, in the questionnaire survey, assuming that the document of interest is a set of responses from a certain population and the comparative document is a set of responses from another population, by applying the present invention, the responses peculiar to the population of the document of interest You can grasp the tendency of. By applying the present invention in this way, it becomes easy to acquire and analyze information from the document of interest. [0010] [Example] A block diagram showing an embodiment of the present invention is shown in FIG. The block 110 is a document input unit, and a comparison document and a document of interest are input. Block 120 is a data processing unit, which performs term detection, morphological analysis, document segment classification, etc. of the input document. Block 130 is a selection engine that selects highly specific document segments or highly specific terms in a document. Block 140 is an output unit, and the selected specific document segment or specific term is output. As a first embodiment, a method of extracting a document segment having a high degree of specificity from a document of interest will be described. FIG. 1 is a flow chart of a first embodiment of the present invention for evaluating the peculiarity of a document segment. The method of the present invention can be carried out by running a program incorporating the present invention on a general-purpose computer. FIG. 1 is a flowchart of a computer running such a program. Block 11 is for comparison / input of the document of interest, block 12 is for term detection, block 13 is for morphological analysis, and block 14 is for document segmentation. Block 15 is for document segment vector creation, block 16 is for topic difference factor analysis, block 17 is for document segment vector projection, block 18 is for document segment specificity calculation for each topic difference factor, and block 19 is for document segment total specificity calculation. , Block 20 is a unique document segment selection. Hereinafter, examples will be described using an English document as an example. [0011] First, the target document and the comparative document are input in the comparison / focused document input 11. Term detection 12 detects words, formulas, symbol sequences, etc. from both documents. Here, words and symbol sequences are collectively referred to as terms. In the case of English sentences, it is easy to detect terms because the orthography method for writing terms separately has been established. Next, the morphological analysis 13 performs a morphological analysis such as part-speech of terms for both documents. Next, in the document segment division 14, both documents are divided into document segments. The most basic unit of a document segment is a sentence. In the case of English sentences, the sentence ends with a period, followed by a space, so the sentence can be easily cut out. As a method of dividing into other document segments, when one sentence consists of multiple sentences, it is divided into a main clause and a subordinate clause, and multiple sentences are grouped into a document segment so that the number of terms is almost the same. There is a method, a method of classifying regardless of the sentence so that the number of terms included from the beginning of the document is the same. [0012] The document segment vector creation 15 first determines the number of dimensions of the vector to be created from the terms appearing in the entire document and the correspondence between each dimension and each term. It is not necessary to correspond the vector components to all the types of terms that appear at this time, and a vector is created using the results of part-speech processing, for example, using only terms determined to be nouns and verbs. You may do so. Next, the types of terms appearing in each document segment and their frequencies are obtained, and the values are multiplied by weights to determine the values of the corresponding components, and a document segment vector is created. As a method of giving weight, a conventional technique can be used. [0013] In the topic difference factor analysis, the projection axis that maximizes the ratio of both documents with respect to the sum of squares of the projection values of all document segment vectors is found. Hereinafter, the case where the sentence is a document segment will be described. The set of terms that appear is {w<sub>1</sub>, .., w<sub>J</sub>Let us consider documents D and T given by} and consisting of M and N sentences, respectively. Let document D be the document of interest and document T be the comparative document. Each document is represented by a set of sentence vectors, and the sentence vector of each k-th sentence is d.<sub>k</sub>= (d<sub>k1</sub>, .., d<sub>kJ</sub>)<sup>T</sup>, T<sub>k</sub>= (t<sub>k1</sub>, .., t<sub>kJ</sub>)<sup>T</sup>Represented by. Figure 5 shows a conceptual diagram when the document segment is a sentence. Document D of interest is composed of M sentences (a), and the sentence vector d from the kth sentence.<sub>k</sub>(b) is generated. d<sub>k</sub>W<sub>j</sub>The component corresponding to is d<sub>kj</sub>It is shown as. d<sub>kj</sub>Is the term w in the kth sentence<sub>j</sub>Since it represents the frequency of, take a value as shown in the example. Figures (c) and (d) describe comparative documents. Let α be the projection axis to be obtained. For the time being, set TheαThe = 1. P is the sum of squares of the projected values when the full-text vector of documents D and T is projected onto α.<sub>D</sub>, P<sub>T</sub>Then, the projection axis to be obtained is the evaluation standard J (α) = P.<sub>D</sub>/ P<sub>T</sub>Is given as α that maximizes. P<sub>D</sub>, P<sub>T</sub>Is [0014] [Number 1]<img file="JP4452012B2_D0001.tif" />[0015] [Number 2]<img file="JP4452012B2_D0002.tif" />[0016] [Number 3]<img file="JP4452012B2_D0003.tif" />[0017] [Number 4]<img file="JP4452012B2_D0004.tif" />Therefore, the evaluation standard J (α) is [0018] [Number 5]<img file="JP4452012B2_D0005.tif" />Can be written as. [0019] Α that maximizes the evaluation criterion J (α) given by Equation 5 can be obtained as α such that the value obtained by differentiating J (α) with α is 0. is this, [0020] [Number 6]<img file="JP4452012B2_D0006.tif" />Given as the eigenvector of the general eigenvalue problem. This is a projection axis that maximizes (sum of projection values from the document of interest) / (sum of squares of projection values from the comparison document), assuming that the full-text vectors of both documents are projected onto a certain projection axis. Is equivalent to asking for. For such a projection axis, the sum of squares of the projection values of the sentence vector of the document of interest is large, and that of the comparison document is small, so information that is abundant in the document of interest and difficult to exist in the comparison document is reflected. Will be done. Generally, a plurality of eigenvalues and eigenvectors of "Equation 6" can be obtained. i-order eigenvalues and eigenvectors are λ<sub>i</sub>, Α<sub>i</sub>Then, it can be considered that the i-th order eigenvector represents the i-th factor that reflects the information that exists in the document of interest D and does not exist in the comparison document T. Therefore, the i-th order eigenvector α<sub>i</sub>Is called the i-th order topic difference factor vector of the document D of interest. In block 16 (topic factor analysis), this topic difference factor vector is obtained. λ<sub>i</sub>= α<sub>i</sub><sup>T</sup>S<sub>D</sub>α<sub>i</sub>/ α<sub>i</sub><sup>T</sup>S<sub>T</sub>α<sub>i</sub>So λ<sub>i</sub>Is α<sub>i</sub>It is the value of the evaluation standard when using. [0021] [0021] Therefore, since the degree of difference between the two documents reflected in each next topic difference factor vector is different from each other, it is better to give weight to each next topic difference factor vector according to the degree of difference. is this<img file="JP4452012B2_D0007.tif" />To be α<sub>i</sub>This is possible by determining the norm of. Then<img file="JP4452012B2_D0008.tif" />Is established, α<sub>i</sub>And the sum of squares of the inner products of each sentence vector of the document D of interest is λ<sub>i</sub>Is equal to. Also, in the case of "Equation 6", the matrix S is required to obtain the eigenvector.<sub>T</sub>Must be a regular matrix. However, in reality, when the number of sentences in the comparison document is smaller than the number of terms, and a specific term pair always co-occurs, S<sub>T</sub>Is not calculated as an invertible matrix. In such a case S<sub>T</sub>The eigenvector can be obtained by regularizing. [0022] [Number 7]<img file="JP4452012B2_D0009.tif" />However, β<sup>2</sup>Is the parameter, I is the identity matrix. When using "Equation 7", the evaluation criterion J (α) is [0023] [Number 8]<img file="JP4452012B2_D0010.tif" />It is equivalent to In the document segment vector projection 17, each sentence vector of the document of interest is projected onto each next topic difference factor vector, and the value is obtained. Sentence vector d of sentence k of the document of interest<sub>k</sub>I Next topic Difference factor vector α<sub>i</sub>Projection value to y<sub>ki</sub>Then this is [0024] [Number 9]<img file="JP4452012B2_D0011.tif" />Demanded by. However, the projection value according to this definition tends to be larger for longer sentences, so to avoid depending on the length of the sentence, Thed<sub>k</sub>Normalization by The may be performed. In this case, the projection value y<sub>ki</sub>Is [0025] [Number 10]<img file="JP4452012B2_D0012.tif" />Given by. In the document segment specificity calculation 18 for each order of the topic difference factor, y<sub>ki</sub>Sentence vector d based on<sub>k</sub>I-order specificity distinc (d<sub>k</sub>, i) is calculated. y<sub>ki</sub>Generally takes a positive or negative value, and its absolute value takes a larger value as the content of the sentence k is closer to that of the document of interest D and different from that of the comparison document T. So distinc (d<sub>k</sub>, i) [0026] [Number 11]<img file="JP4452012B2_D0013.tif" />Or [0027] [Number 12]<img file="JP4452012B2_D0014.tif" />It can be defined as. The i-order peculiarity obtained in this way is, so to speak, the peculiarity of only the i-th factor, and it is necessary to combine the peculiarities of a plurality of factors in order to accurately express the peculiarity of the sentence k. for that reason, [0028] [Number 13]<img file="JP4452012B2_D0015.tif" />Calculates the total peculiarity of the document segment of sentence k by (block 19). Here, L is the number of topic difference factor vectors used for sentence-specific calculations, and an appropriate value needs to be determined experimentally. The maximum value of L is the number of eigenvalues whose value is 1 or more. In the peculiar document segment selection 20, a sentence peculiar to the document of interest is selected based on the obtained next peculiarity and total peculiarity. This can be done as follows. The simplest method is to select sentences with a total specificity of a certain value or more. The following method is also possible. First, for a topic difference factor vector of a specific order, each sentence is divided into a group in which the projection value of each sentence vector and the topic difference factor vector is positive and a group in which the projection value is negative. Then, from each group, select sentences with a certain degree of peculiarity or more. This is done for all topic difference factor vectors up to a certain degree L, and peculiar sentences are selected by eliminating duplication. Either method can be used to select a unique sentence. [0029] Further, in the first embodiment, not only the document segment but also the peculiarity of the combination of terms such as a phrase, a term group having a dependency relationship, and a term sequence pattern can be evaluated by doing the following. For example, taking the expression "soccer match held in Yokohama" as an example, "soccer match" is a noun phrase because "soccer" modifies the noun "match". Furthermore, "done in Yokohama" modifies the noun phrase "soccer match", so the whole expression above is a noun phrase. In more detail, "in Yokohama" modifies the verb phrase "done", so "done in Yokohama" is a group of terms that are dependent. In addition, when the expression "soccer match played in XX" appears many times with various place names in XX, "soccer match played in XX" becomes a pattern of term series. [0030] In block 13, in addition to the morphological analysis, the combination of terms to be evaluated is extracted. In the case of a group of terms that are related to a phrase or dependency, it can be extracted by performing syntactic analysis. In addition, various methods have already been devised for extracting frequently occurring term sequence patterns, and they can be used without problems. In block 15, in addition to the document segment vector used in block 16, the vector p = (p) for the combination of terms to be evaluated.<sub>1</sub>, .., p<sub>J</sub>)<sup>T</sup>To create. p is a vector in which 1 is set for the component corresponding to the term included in the combination of terms, and 0 is set for the other components. Here, a concrete example of the vector p is as follows. In the case of the expression "soccer match played in Yokohama", p is 1 only for the components corresponding to the terms "Yokohama", "played", "soccer", and "match", and 0 for the others. Becomes a vector. Sentence vector d such p in blocks 17, 18, 19<sub>k</sub>By using it instead of, the peculiarity of the combination of terms to be evaluated can be obtained. Therefore, in block 20, a unique combination of terms can be selected as in the case of sentences. [0031] As a second embodiment, a method of selecting a term having a high degree of specificity from the document of interest will be described. For terms, the correlation between the frequency of the word of interest in each sentence and the degree of peculiarity of each sentence is obtained, and the term with the larger correlation value is selected. Based on this correlation value, the peculiarity of the term is obtained. FIG. 2 is a flow chart showing a second embodiment of the present invention for evaluating the specificity of terms. The method of the present invention can be carried out by running a program incorporating the present invention on a general-purpose computer. FIG. 2 is a flowchart of a computer running such a program. Block 11 is for comparison / input of the document of interest, block 12 is for term detection, block 13 is for morphological analysis, and block 14 is for document segmentation. Block 15 is document segment vector creation, block 16 is topic difference factor analysis, block 27 is document segment vector projection, block 28 is term specificity calculation for each topic difference factor, block 29 is term total specificity calculation, block 30 is a specific term selection. Of these, blocks 11 to 16 are exactly the same as those shown in FIG. An example in which a sentence is used as a document segment will be described as in the case of FIG. In the document segment vector projection 27, in addition to the projection of the sentence vector of the document of interest D in FIG. 117, the projection of the full sentence vector of the comparative document T is also performed. Statement vector t of comparative document T<sub>k</sub>I Next topic Difference factor vector α<sub>i</sub>Projection value to z<sub>ki</sub>Then this is [0032] [Number 14]<img file="JP4452012B2_D0016.tif" />Or [0033] [Number 15]<img file="JP4452012B2_D0017.tif" />Demanded by. In calculating the term specificity of each topic difference factor, first, the correlation between the projection value of each sentence and the term frequency in each sentence is obtained. The jth term w of each sentence vector in the document of interest and the comparison document<sub>j</sub>The value of the component corresponding to and the i-order topic of each sentence vector Difference factor vector α<sub>i</sub>Correl (w) the correlation coefficient with the projected value to<sub>j</sub>, I). Statement vector d<sub>k</sub>, T<sub>k</sub>The jth component of is d<sub>kj</sub>, T<sub>kj</sub>, Α<sub>i</sub>The projected value to is y<sub>ki</sub>, Z<sub>ki</sub>So the correlation coefficient is [0034] [Number 16]<img file="JP4452012B2_D0018.tif" />Demanded by. Term w<sub>j</sub>The correlation coefficient is higher than that of d<sub>k</sub>Or t<sub>k</sub>Term in w<sub>j</sub>The value of the component corresponding to and the α of the sentence vector<sub>i</sub>This is when a proportional relationship holds with the projected value of. That is, the term w<sub>j</sub>The correlation coefficient is high when the i-th order singularity of the sentence appears when appears, and decreases when it does not appear. In such cases, the term w<sub>j</sub>Can be regarded as a peculiar term that governs the i-th order peculiarity of each sentence. Therefore, the i-order term specificity is distinc (w).<sub>j</sub>, i), this is [0035] [Number 17]<img file="JP4452012B2_D0019.tif" />Or [0036] [Number 18]<img file="JP4452012B2_D0020.tif" />Can be obtained by (block 28). In the term total specificity calculation, as in the case of Fig. 1, the total specificity for each term is obtained by combining multiple factors. Term w<sub>j</sub>Distinc (w)<sub>j</sub>), This is [0037] [Number 19]<img file="JP4452012B2_D0021.tif" />Can be found at (block 29). In the peculiar term selection 30, a term peculiar to the document of interest is selected based on each of the following peculiarities and total peculiarities obtained. This can be done as follows. The simplest method is to select terms whose overall specificity is above a certain value. The following method is also possible. First, for a topic difference factor vector of a specific order, each term is negative, a group in which the correlation coefficient between each sentence vector and the projection value of the topic difference factor vector and the frequency of each term is positive. Divide into groups. Then, from each group, select terms with a certain degree of peculiarity or more. This is done for all topic difference factor vectors up to a certain degree L, and peculiar terms are selected by eliminating duplication. Either method allows you to select specific terms. [0038] Further, in the second embodiment, it is possible to evaluate not only the terms but also the peculiarity of the combination of terms such as a phrase, a term group having a dependency relationship, and a term series pattern by doing the following. Similar to the first embodiment, in block 13, in addition to the morphological analysis, a combination of terms to be evaluated is extracted. In the case of a group of terms that are related to a phrase or dependency, it can be extracted by performing syntactic analysis. In addition, various methods have already been devised for extracting frequently occurring term sequence patterns, and they can be used without problems. In block 15, in addition to creating the document segment vector used in block 16, the frequency at which the combination of terms to be evaluated appears in each document segment is calculated. The frequency in sentence k of document D of interest is p<sub>Dk</sub>, P frequency in sentence k of comparative document T<sub>Tk</sub>Then, in blocks 28 and 29, d<sub>kj</sub>Instead of p<sub>Dk</sub>, T<sub>ki</sub>Instead of p<sub>Tk</sub>By using, the term w<sub>j</sub>Instead, the peculiarity of the combination of terms to be evaluated can be obtained. As a result, in block 30, it is possible to select a unique combination of terms as in the case of terms. [0039] Next, in order to evaluate the peculiarity of the document of interest, the following second means will be disclosed as Example 3. In the second means, up to the generation 15 of the vector of the document segment is common to the first means (Example 1 and Example 2), but after that, each sentence of the document of interest is similar to the entire document of interest. , And the degree of similarity with the entire comparison document. FIG. 3 is a flow chart showing a third embodiment of the present invention for evaluating the document segment and the specificity of terms. The method of the present invention can be carried out by running a program incorporating the present invention on a general-purpose computer. FIG. 3 is a flowchart of a computer running such a program. [0040] 11 is the comparison / input of the document of interest, block 12 is the term detection, block 13 is the morphological analysis, and block 14 is the document segment classification. Block 15 is for creating a document segment vector, block 36 is for calculating similarity, block 37 is for calculating document segment specificity, and block 38 is for calculating term specificity. Block 39 is a unique document segment / term selection. Of these, blocks 11 to 15 are exactly the same as those shown in FIG. In the similarity calculation, the similarity between each sentence vector of the document of interest and the comparison document and the entire document of interest and the entire comparison document is obtained. Sentence vector d of the document of interest<sub>k</sub>The degree of similarity with the entire document of interest is sim (D, d)<sub>k</sub>), Sim (T, d) the similarity with the entire comparison document<sub>k</sub>), Then these are d<sub>k</sub>Based on the sum of squares of the inner product of the document of interest and the full-text vector of the comparison document [0041] [Number 20]<img file="JP4452012B2_D0022.tif" />[0042] [Number 21]<img file="JP4452012B2_D0023.tif" />Can be obtained as follows. Alternatively, set the average sentence vector of the document of interest and the comparison document, respectively.<img file="JP4452012B2_D0024.tif" />Then, it can be calculated as follows. [0043] [Number 22]<img file="JP4452012B2_D0025.tif" />[0044] [Number 23]<img file="JP4452012B2_D0026.tif" />In the similarity calculation, in order to calculate the term specificity in the latter part, the similarity of the full-text vector of the comparison document with the entire document of interest and the entire comparison document is obtained (block 36). In the document segment specificity calculation, the specificity is obtained for the full-text vector of the document of interest. The important sentences in the document of interest have a large degree of similarity with the entire document of interest, and the sentences having different contents from the comparative document have a small degree of similarity with the entire comparison document. Therefore, by using (similarity with the entire document of interest) / (similarity with the entire comparative document), a balanced peculiarity between difference and importance can be defined. Therefore, the peculiarity of the sentence k of the document of interest distinc (d)<sub>k</sub>) Can be calculated as follows (block 37). [0045] [Number 24]<img file="JP4452012B2_D0027.tif" />The peculiarity of the sentence k obtained in this way increases when the sentence k has a high degree of similarity to the document of interest and a low degree to the comparative document. Since the document segment specificity calculation is used to calculate the following term specificity, the sentence specificity of the comparative document T is also obtained. The peculiarity of sentence k in comparative document T is distinc (t)<sub>k</sub>). In the term specificity calculation, the term specificity is obtained from the correlation coefficient between the specificity of each sentence and the term frequency in each sentence. Term w<sub>j</sub>Distinc (w)<sub>j</sub>) [0046] [Number 25]<img file="JP4452012B2_D0028.tif" />Can be obtained by (block 38). Term w<sub>j</sub>The correlation coefficient is higher than that of d<sub>k</sub>Or t<sub>k</sub>Term in w<sub>j</sub>This is when a proportional relationship holds between the value of the component corresponding to and the peculiarity of the sentence. That is, the term w<sub>j</sub>The correlation coefficient is high when the degree of peculiarity of the sentence is large when appears, and when it does not appear, it decreases. In such cases, the term w<sub>j</sub>Can be regarded as a peculiar term that governs the peculiarity of each sentence. In the peculiar document segment 39 and the peculiar term selection 40, a peculiar sentence and a term can be obtained by selecting a sentence having a sentence peculiarity of a certain value or more and a term having a term peculiarity of a certain value or more. [0047] In this embodiment, not only the document segment but also the peculiarity of the combination of terms such as phrases, dependent term groups, and term sequence patterns can be evaluated by doing the following. In block 13, in addition to the morphological analysis, the combination of terms to be evaluated is extracted. In the case of a group of terms that are related to a phrase or dependency, it can be extracted by performing syntactic analysis. In addition, various methods have already been devised for extracting patterns that frequently appear, and can be used without problems. In block 15, in addition to the document segment vector used in block 16, the vector p = (p) for the combination of terms to be evaluated.<sub>1</sub>, .., p<sub>J</sub>)<sup>T</sup>To create. p is a vector in which 1 is set for the component corresponding to the term included in the combination of terms to be evaluated, and 0 is set for the other components. Next, in blocks 36 and 37, such p is expressed as a statement vector d.<sub>k</sub>By using instead of, the similarity sim (D, p) between p and the document of interest and the similarity sim (T, p) between p and the comparative document are obtained. Like the numbers 20 and 21, these can be defined as follows. [0048] [Number 26]<img file="JP4452012B2_D0029.tif" />[0049] [Number 27]<img file="JP4452012B2_D0030.tif" />Alternatively, it may be defined as follows in the same manner as the numbers 22 and 23. [0050] [Number 28]<img file="JP4452012B2_D0031.tif" />[0051] [Number 29]<img file="JP4452012B2_D0032.tif" />Using these similarities, the specificity distinc (p) of the combination of terms to be evaluated can be obtained as follows. [0052] [Number 30]<img file="JP4452012B2_D0033.tif" />In block 40, a combination of terms having a certain degree of peculiarity or more is selected as a combination of peculiar terms. Further, in this embodiment, the peculiarity of a phrase composed of a plurality of terms, a term group having a dependency relationship, or a term sequence pattern can be obtained as follows. In block 15, in addition to creating the document segment vector used in block 16, the frequency at which the combination of terms to be evaluated appears in each document segment is calculated. The frequency in sentence k of document D of interest is p<sub>Dk</sub>, P frequency in sentence k of comparative document T<sub>Tk</sub>Then, in block 38, d<sub>kj</sub>Instead of p<sub>Dk</sub>, T<sub>ki</sub>Instead of p<sub>Tk</sub>By using, the term w<sub>j</sub>Instead, the peculiarity of the combination of terms to be evaluated can be obtained. In block 39, a combination of terms having a certain degree of peculiarity or more is selected as a combination of peculiar terms. [0053] [Effect of the invention] Here, the experimental results using "Formula 13" are shown in order to explain the effect of the present invention. The data used in the experiment were selected from the first category "acq" of the document classification corpus Reuters-21578 based on the criteria of having an appropriate length and high similarity. These ids are 1836 and 2375. The cosine similarity between them was 0.955. Document 1836 consists of 43 sentences and 2375 consists of 32 sentences. These are news articles on the same day, but we decided to extract specific sentences from the document D of interest, with 2375 as the document D of interest and 1836 as the comparison document T, which seems to have been published later. In terms of content, these are related to the acquisition of US Airways USAir by US Airways TWA, D-1 to D-4 are summarized as articles, D-5 to D-24 are the background of the acquisition drama, D -25 and later are TWA's analysis, and information that is not in document T is included in some sentences in D-1 to D-4, D-5 to D-24, and D-25 and later. ing. The full text of these documents is shown at the end of this specification as "experimental document data". [0054] As a result of conducting an experiment according to Example 1 of the present invention, sentences having high specific values include D-1, D-8, D-11, D-24, D-25, D-27, D-28, and D-. Eight 30 sentences were selected. These were recognized as sentences peculiar to the document of interest because they had little relation to the comparative document even in the human reading and comparison experiment. The results of selecting words with high specificity according to "Formula 19" are shown below. For 10 words with high specificity, the specificity of each word, the frequency of appearance in the document of interest D, and the frequency of appearance in the comparative document T are shown. [0055]<img file="JP4452012B2_D0034.tif" />From these results, it was possible to select words that appear less frequently in the comparative document T and more frequently appear in the document D of interest. The following example can be considered as an application of this. If you read the previous article and understand the content, you can extract keywords that are not described in the previous article from the articles that came in after that. Therefore, you can decide if you need to read the articles that came in later. Different specificities are required for two terms such as "succeed" and "clear" that have exactly the same frequency in the document of interest and the comparison document, and it is possible to determine which is more specific in the present invention. It is a feature of. [0056] [Experimental document data] The documents used in the present invention are described below. Comparative document T (Reuter-id1836) Trans World Airlines Inc complicated the bidding for Piedmont Aviation Inc by offering either to buy Piedmont suitor USAir Group or, optionally, To merge with Piedmont and USAir. Piedmont's board was meeting today, and Wall Street speculated the board was discussing approaching bids from Norfolk Southern Corp and USAir. The TWA offer was announced shortly after the Piedmont board meeting was scheduled to begin. USAir for 52 dlrs cash per share. It also said it was the largest shareholder of USAir and threatened to go directly to USAir shareholders with an offer for 51 pct of the stock at a lower price. TWA also said it believed its offer was a better deal for USAir shareholders than an acquisition of Piedmont, but it said it discrete would discuss a three way combination of the airlines. Market sources and analysts speculated that TWA chairman Carl Icahn made the offer in order We're just wondering if he's not just trying to get TWA into play. [0057] There's speculation on the street he just wants to move onto somthing else, said one arbitrager. We think TWA might just be putting up a trial balloon. Analysts said the offer must be taken seriously by USAir, but that the airline will probably reject it because They also said Icahn must prove his offer credible by revealing financing arrangements. They need to show their commitment and their ability to finance. I think it's a credible offer, said Timothy Pettee, a Bear Stearns analyst. I think it's certainly on the low end of relative values of airline deals, said Pettee. Pettee estimated 58 dlrs would be in a more reasonable range based on other airline mergers . USAir stock soared after TWA made public its offer. [0058] [0058] A spokesman for USAir declined comment, and said USAir had not changed its offer for Piedmont. USAir offered of buy 50 pct of that airline's stock for 71 dlrs cash per share and the balance for 73 dlrs per share in USAir stock. USAir closed up 5-3 / 8 at 49-1 / 8 on volume of 1.9 mln shares. Piedmont , which slipped 1/2 to close at 69-5 / 8, also remained silent on the TWA action. Piedmont has an outstanding 65 dlr cash per share offer from Norfolk Southern Corp. Norfolk Southern declined comment, but said it stuck with its offer for Piedmont. Norfolk owns about 20 pct of Piedmont and opened the bidding when it said it would propose a takeover of Piedmont. Some analysts said Icahn may be trying to acquire USAir to make his own airline a more attractive takeover target. Icahn I think had wanted to sell I think the strategy might have called for making his investment more attractive. [0059] One way to accomplish that specific objective is to go out and acquire other airlines, said Andrew Kim of Eberstadt Fleming. I don't know whose going to buy them, but at least this way it becomes a much more viable package, said Kim. But Icahn's financing ability for such a transaction remains in doubt, in part because of TWA's heavy debt load. Wall street sources said TWA has some cash with which to do the offer. The sources said Icahn has not lined up outside financial advisers and plans to make his own arrangements. Icahn earlier this year abandoned plans to buy USX Corp <X> and still retains 11 pct of that company's stock. Some Wall street sources said the financier's USX plan was impacted by the cloud hanging over his adviser, Drexel Burnham Lambert Inc, because of Wall Street's insider trading scandal. Industry sources also predicted USAir might reject the TWA offer on price and financing concerns. It's littered with contingencies and it doesn't even have a financing arrangement, said one executive at another major airline. But the executive conceded a merged TWA USAir would be a strong contender with USAir's east coast route system and planned west coast presence from PSA. USAir could feed the intenrational flights of TWA, which has a midwest presence in its St. Louis hub. Adding Piedmont, dominant in the southeast, to the mix would develop an even stronger force. The combined entity would also have TWA's pars reservation system. Such a merger would be complex and analysts said it would result in an airline iwth an 18 pct market share .. [0060] Document of interest D (Reuter-id2375) D-1 Carl Icahn's bold takeover bid for USAir Group <u style="single"> has clouded the fate of Piedmont Aviation Inc, which was being courted by USAir.</u><u style="single">D-2 Yesterday, Icahn's Transworld Airlines Inc <TWA> made a 1.4 billion dlr offer for USAir Group.</u><u style="single">D-3 The move complicated a USAir takeover offer for Piedmont, which was believed to be close to accepting the bid.</u><u style="single">D-4 Today, USAir rejected Icahn's 52 dlr per share offer and said the bid was a last minute effort to interfere in its takeover of Piedmont.</u><u style="single">D-5 Icahn was unavailable for comment.</u><u style="single">D-6 Piedmont fell one to 68-5 / 8 on volume of 963,000.</u><u style="single">D-7 TWA was off 3/8 to 31-1 / 2.</u><u style="single">D-8 USAir fell 1-3 / 8 to 47-3 / 4 as doubt spread it would be taken over.</u><u style="single">D-9 Analysts and market sources view the TWA bid as an attempt to either trigger a counter offer from USAir or to attract a suitor who might want both airlines once they merged.</u><u style="single">D-10 The next move is either Icahn starts a tender offer or Piedmont and USAir announce a deal, speculated one arbitrager.</u><u style="single">【0061】</u><u style="single">D-11 Some arbitragers said there is now some risk in the current price of Piedmont since it is not clear that USAir's bid will succeed.</u><u style="single">D-12 Piedmont's largest shareholder and other suitor, Norfolk Southern Corp <NSC> has offered 65 dlrs per share for the company.</u><u style="single">D-13 USAir offered 71 dlrs cash per share for half of Piedmont stock, and 73 dlrs per share in stock for the balance.</u><u style="single">D-14 Some arbitragers, however, believe the depressed price of Piedmont offers a buying opportunity since the airline is destined to be acquired by someone.</u><u style="single">D-15 USAir, they said, is the least likely to be bought.</u><u style="single">D-16 Icahn, who has long talked about further consolidation in the airline industry, also offered USAir the alternative of a three way airline combination, including TWA and Piedmont.</u><u style="single">D-17 But Wall Street has given little credibility to Icahn's offer, which lacked financing and was riddled with contingencies.</u><u style="single">D-18 Still, he has succeeded in holding up a merger of two airlines both of which analysts said would fit well with TWA.</u><u style="single">D-19 You can't discount him, said one arbitrager.</u><u style="single">D-20 Analysts, however, said Icahn would have to prove he is serious by following through with his threats or making a new offer.</u><u style="single">【0062】</u><u style="single">D-21 In making the offer for USAir, Icahn threatened to go directly to shareholders for 51 pct of the stock at a lower price if USAir rejected his offer.</u><u style="single">D-22 It's clear Icahn wants to sell and he's bluffing, said one arbitrager.</u><u style="single">D-23 Analysts said the 52 dlr per share offer was underpriced by about six dlrs per share.</u><u style="single">D-24 Some analysts believe Icahn's proposed three way airline combination might face insurmountable regulatory hurdles, but others believe it could be cleared if the companies are acquired separately.</u><u style="single">D-25 TWA would have to be the surviving company for the deal to work, said one analyst.</u><u style="single">D-26 Analysts said such a merger would be costly and complicated.</u><u style="single">D-27 TWA has the best cost structure, since Icahn succeeded in winning concessions from its unions.</u><u style="single">D-28 In order for the other carriers to come down to TWA's wage scale in a merger, TWA would have to be the surviving entity, analysts said.</u><u style="single">D-29 Such a move does not necessarily free Icahn of TWA, they said.</u><u style="single">D-30 They said he showed skill in reducing Ozark Airlines' costs when he merged it into TWA last year, and he might be a necessary ingredient for a merger to work.</u><u style="single">D-31 However, other analysts speculated the managements of Piedmont and USAir would not tolerate Icahn as head of a new company.</u><u style="single">D-32 They said a USAir acquisition of TWA might be a way for him to exit the company if USAir's airline is then merged into TWA.</u>[Simple explanation of drawings] FIG. 1 is a diagram showing a first embodiment of the present invention, showing a procedure from the stage when a document is input to the determination of the specificity of a document segment. FIG. 2 is a diagram showing a second embodiment of the present invention, showing a procedure from the stage when a document is input to the determination of the specificity of terms. FIG. 3 is a diagram showing a third embodiment of the present invention, showing a procedure from the stage when a document is input to the determination of the specificity of a document segment and terms. FIG. 4 is a block diagram of the present invention. FIG. 5 is a diagram illustrating a sentence vector of a document of interest and a comparative document of the present invention. [Explanation of symbols] 110: Document input section 120: Data processing unit 130: Selected engine 140: Specific document segment / specific term output section
Every citation, both waysCites: the store holds 1 of 2
| Document | Relation | Office |
|---|---|---|
| JP200241543A | Cites | Japan |
| 川谷隆彦、文書集合間の差異検出法と文書分類への応用、情報処理学会研究報告.自然言語処理研究会報告、日本、2002.03.04発行、Vol.2002,No.20、p1-8 | Non-patent | – |
11 members in 5 offices
Priority claims2
| Document | Office | Kind | Date |
|---|---|---|---|
| 2002195375 | Japan | A | |
| JP20020195375 | – | – | – |
Members11
| Document | Office | Kind | |
|---|---|---|---|
| EP1378838A2 | European Patent Office (EPO) | A2 | |
| US2004006736A1 | United States of America | A1 | |
| JP2004038606A | Japan | A | |
| CN1495644A | China | A | |
| EP1378838A3 | European Patent Office (EPO) | A3 | |
| US7200802B2 | United States of America | B2 | |
| EP1378838B1 | European Patent Office (EPO) | B1 | |
| DE60316227D1 | Germany | D1 | |
| DE60316227T2 | Germany | T2 | |
| JP4452012B2This record | Japan | B2 | |
| CN1495644B | China | B |
32 legal events, as the office reported them to INPADOC
Over the term
Point at a mark for the eventEvents
| Event | Code | |
|---|---|---|
| Cancellation because of no payment of annual feesLAPS | LAPS | |
| Renewal fee payment (event date is renewal date of database)FPAY | FPAY | |
| Renewal fee payment (event date is renewal date of database)FPAY | FPAY | |
| Written notification of registration of transferJAPANESE INTERMEDIATE CODE: R350R350 | R350 | |
| Request for change of ownership or part of ownershipJAPANESE INTERMEDIATE CODE: R313113S111 | S111 | |
| Renewal fee payment (event date is renewal date of database)FPAY | FPAY | |
| Written notification for declining of transfer of rightsJAPANESE INTERMEDIATE CODE: R360R360 | R360 | |
| Transfer withdrawnWithdrawnJAPANESE INTERMEDIATE CODE: R371R371 | R371 | |
| Renewal fee payment (event date is renewal date of database)FPAY | FPAY | |
| Written notification for declining of transfer of rightsJAPANESE INTERMEDIATE CODE: R360R360 | R360 | |
| Renewal fee payment (event date is renewal date of database)FPAY | FPAY | |
| Request for change of ownership or part of ownershipJAPANESE INTERMEDIATE CODE: R313113S111 | S111 | |
| Renewal fee payment (event date is renewal date of database)FPAY | FPAY | |
| Certificate of patent or registration of utility modelJAPANESE INTERMEDIATE CODE: R150R150 | R150 | |
| First payment of annual fees (during grant procedure)JAPANESE INTERMEDIATE CODE: A61A61 | A61 | |
| Written decision to grant a patent or to grant a registration (utility model)JAPANESE INTERMEDIATE CODE: A01A01 | A01 | |
| Written amendmentJAPANESE INTERMEDIATE CODE: A523A521 | A521 | |
| Written amendmentJAPANESE INTERMEDIATE CODE: A821A521 | A521 | |
| Decision of refusalJAPANESE INTERMEDIATE CODE: A02A02 | A02 | |
| Written amendmentJAPANESE INTERMEDIATE CODE: A821A521 | A521 | |
| Written amendmentJAPANESE INTERMEDIATE CODE: A523A521 | A521 | |
| Written amendmentJAPANESE INTERMEDIATE CODE: A523A521 | A521 | |
| Written permission of extension of timeJAPANESE INTERMEDIATE CODE: A602A602 | A602 | |
| Written request for extension of timeJAPANESE INTERMEDIATE CODE: A601A601 | A601 | |
| Notification of resignation of power of attorneyJAPANESE INTERMEDIATE CODE: A7424RD04 | RD04 | |
| Notification of acceptance of power of attorneyJAPANESE INTERMEDIATE CODE: A7422RD02 | RD02 | |
| Notification of reasons for refusalJAPANESE INTERMEDIATE CODE: A131A131 | A131 | |
| Written amendmentJAPANESE INTERMEDIATE CODE: A523A521 | A521 | |
| Written amendmentJAPANESE INTERMEDIATE CODE: A523A521 | A521 | |
| Written request for application examinationJAPANESE INTERMEDIATE CODE: A621A621 | A621 | |
| Notification of resignation of power of attorneyJAPANESE INTERMEDIATE CODE: A7424RD04 | RD04 | |
| Notification of acceptance of power of attorneyJAPANESE INTERMEDIATE CODE: A7422RD02 | RD02 |
Numbers
- Publication
- 4452012
- Publication, DOCDB
- 4452012
- Publication, EPODOC
- JP4452012B
- Application
- 195375
- Application, DOCDB
- 2002195375
- Application, EPODOC
- JP20020195375
Titles2
- Japanese
- 文書の特有性評価方法
- English
- Document uniqueness evaluation method
Classification
- CPC, 3
- G06F16/3347
- G06F16/93
- G06F18/2132
- IPC, 5
- G06F17 27
- G06F17 21
- G06F17 30
- G06F17 28
- G06K9 62