EP3614275A1

Indexing using machine learning techniques

Abstract

Systems and techniques for indexing and/or querying a database are described herein. Discrete sections and/or segments from documents may be determined by a concept indexing system. The segments may be indexed by concept and/or higher-level category of interest to a user. A user may query the segments by one or more concepts. The segments may be analyzed to rank the segments by statistical accuracy and/or relatedness to one or more particular concepts. The rankings may be used for presentation of search results in a user interface. Furthermore, segments and/or documents may be ranked based on recency decay functions that distinguish between segments that maintain their relevance over time in contrast with temporal segments whose relevance decays quicker over time, for example.

EP3614275A1, drawing sheet 1
Sheet 1 of 42

Term

9.2 yearsto projected expiry

Projected expiry 22 December 2035, counted from filing; an application has no term until it is granted.

  1. Priority
  2. Filed
  3. Published
  4. Today
  5. Projected expiry

14 claims: 3 independent, 11 dependent

  1. 1
    A computer-implemented method comprising:identifying a plurality of segments within a plurality of documents;accessing a plurality of concepts of interest to a user, wherein each concept is associated with one or more concept keywords;for each concept: determining statistical likelihoods that respective identified segments are associated with the concept;determining a ranked listing of one or more segments having highest statistical likelihoods of being associated with the concepts;and presenting the ranked listing in a user interface.
  2. 2
    The computer-implemented method according to Claim 1, further comprising indexing the plurality of concepts and the determined statistical likelihoods in a concept indexing database and wherein the ranked listing is presented in response to a user query for a specific concept.
  3. 3
    The computer-implemented method according to Claim 1 or Claim 2, wherein the ranked listing is further determined based on a recency score associated with the one or more segments.
  4. 5
    The computer-implemented method according to Claim 4, wherein traversing the concept hierarchy includes traversing the concept hierarchy in a descending direction.
  5. 6
    The computer-implemented method according to Claim 4, wherein traversing the concept hierarchy includes traversing the concept hierarchy in an ascending direction.
  6. 7
    The computer-implemented method according to any of Claims 4 to 6, wherein accessing a plurality of concepts of interest to the user further comprises presenting, in a user interface, at least the first concept of interest and one or more of the one or more related concepts of interest.
  7. 8
    The computer-implemented method according to Claim 7, wherein accessing a plurality of concepts of interest to the user further comprises:receiving, from the user, a selection of one or more of the one or more related concepts of interest;determining one or more additional concepts of interest by at least: finding the one or more concepts associated with the selection in the concept hierarchy;and traversing the concept hierarchy to determine additional concepts of interest.
  8. 9
    The computer-implemented method according to Claim 2, wherein the user query is received by means of a concept selection area including selectable concepts from a concept hierarchy, the concept hierarchy including a plurality of concepts of interest to the user, and including concept keywords associated with respective concepts.
  9. 11
    The computer-implemented method of Claim 10, wherein generating the statistical likelihood is performed by processing the binary vector using a machine learning algorithm.
  10. 12
    The computer-implemented method of Claim 10 or Claim 11, further comprising merging a first segment and a second segment, wherein a determination to merge the first segment and the second segment is based on a cosine distance calculation above a threshold, the cosine distance calculation based at least on a first word vector associated with the first segment and a second world vector associated with the second segment.
  11. 13
    A computing system for identifying segments of interest in documents, the computing system including:one or more hardware computer processors configured to execute software instructions;and one or more storage devices storing software instructions configured for execution by the one or more hardware computer processors to cause the computing system to perform the computer-implemented method of any preceding claim.
  12. 14
    A non-transitory computer storage medium storing computer executable instructions that when executed by a computer hardware processor perform operations comprising the computer-implemented method of any of claims 1 to 12.