US8301642B2

Information retrieval from a collection of information objects tagged with hierarchical keywords

Summary by NHIP

Keyword hierarchy expansion

The method arranges keywords into hierarchical trees and computes association scores based on tree distances between keyword positions. It automatically expands original queries by adding friend keywords whose scores meet or exceed a predetermined value and are not in the original query.

Claim Score by NHIP

Read claim 1, the broadest

Abstract

The present invention can include a data processing system-implemented method or a data processing system readable media having software code for carrying out the method. The method can comprise formulating queries, searching for a plurality of information objects, or a combination thereof. In a specific embodiment, an original query with at least one keyword can be automatically expanded to an expanded query that includes at least one keyword that is not in the original query. The expanded query may be used to search for information objects that are relevant to the expanded query.

US8301642B2, drawing sheet 1
Sheet 1 of 8

Term

Term ended

Expired 20 July 2021, 5.2 years ago.

  1. Priority
  2. Filed
  3. Granted
  4. Expired
  5. Today

17 claims: 3 independent, 14 dependent

  1. 1
    Broadest claimClaim Score 31, narrow(NHIP)A data preparation method useful for information retrieval, comprising:at a server computer, arranging a master list of keywords into one or more trees, at least one of which representing a keyword hierarchy;relating a set of keywords to each information object of a set of information objects stored in a repository, the set of keywords being members of the master list of keywords;determining friend keywords for each keyword in the master list of keywords, the determining comprising computing association scores between a given keyword and all other keywords in the keyword hierarchy based at least in part upon positions of each keyword-friend pair within the keyword hierarchy and a tree distance between the positions, each of the association scores representing a degree of association of the given keyword and a friend keyword in the keyword hierarchy;automatically expanding an original query to produce an expanded query, the original query being generated by end-user activity at a client computer communicatively connected to the server computer over a network connection, the original query comprising a first keyword, the expanded query comprising the first keyword and a second keyword, the second keyword being associated with the first keyword from the original query in a keyword-friend pair according to the keyword hierarchy, the keyword-friend pair having an association score that meets or exceeds a predetermined value, wherein the second keyword is not in the original query;and searching the repository to identify information objects that correspond to the expanded query.
  2. 9
    A computer program product comprising at least one non-transitory computer readable medium storing instructions translatable by a processor of a server computer to perform:arranging a master list of keywords into one or more trees, at least one of which representing a keyword hierarchy;relating a set of keywords to each information object of a set of information objects stored in a repository, the set of keywords being members of the master list of keywords;determining friend keywords for each keyword in the master list of keywords, the determining comprising computing association scores between a given keyword and all other keywords in the keyword hierarchy based at least in part upon positions of each keyword-friend pair within the keyword hierarchy and a tree distance between the positions, each of the association scores representing a degree of association of the given keyword and a friend keyword in the keyword hierarchy;automatically expanding an original query to produce an expanded query, the original query being generated by end-user activity at a client computer communicatively connected to the server computer over a network connection, the original query comprising a first keyword, the expanded query comprising the first keyword and a second keyword, the second keyword being associated with the first keyword from the original query in a keyword-friend pair according to the keyword hierarchy, the keyword-friend pair having an association score that meets or exceeds a predetermined value, wherein the second keyword is not in the original query;and searching the repository to identify information objects that correspond to the expanded query.
  3. 15
    A system, comprising:a processor;and at least one non-transitory computer readable medium storing instructions translatable by the processor to perform: arranging a master list of keywords into one or more trees, at least one of which includes nodes representing a keyword hierarchy;relating a set of keywords to each information object of a set of information objects stored in a repository, the set of keywords being members of the master list of keywords;determining friend keywords for each keyword in the master list of keywords, the determining comprising computing association scores between a given keyword and all other keywords in the keyword hierarchy based at least in part upon positions of each keyword-friend pair within the keyword hierarchy and a tree distance between the positions, each of the association scores representing a degree of association of the given keyword and a friend keyword in the keyword hierarchy;automatically expanding an original query to produce an expanded query, the original query being generated by end-user activity at a client computer communicatively connected to the system over a network connection, the original query comprising a first keyword, the expanded query comprising the first keyword and a second keyword, the second keyword being associated with the first keyword from the original query in a keyword-friend pair according to the keyword hierarchy, the keyword-friend pair having an association score that meets or exceeds a predetermined value, wherein the second keyword is not in the original query;and searching the repository to identify information objects that correspond to the expanded query.