EP0730765B1

Associative text search and retrieval system

Abstract

An associative text search and retrieval system uses one or more front end processors to interacting with a network having one or more user terminals connected thereto to allow a user to provide information to the system and receive information from the system. The system also includes storage for a plurality of text documents, and at least one processor, coupled to the front end processors and the document storage. The processor(s) search the text documents according to a search request provided by the user and provide to the front end processor a predetermined number of retrieved documents containing at least one term of the search request. The retrieved documents have higher ranks than documents not provided to the front end processor. The ranks are calculated using a formula that varies according to the square of the frequency in each of the text documents of each of the search terms.

EP0730765B1, drawing sheet 1
Sheet 1 of 17

Term

Term ended

Expired 22 November 2014, 11.8 years ago.

  1. Priority
  2. Filed
  3. Granted
  4. Expired
  5. Today

41 claims: 41 independent, 0 dependent

  1. 1
    An associative text search and retrieval system, comprising:front end processing means (56, 57, 58) for interacting with a network having one or more user terminals connected thereto to allow a user to provide information to the system and receive information from the system;storage means (46, 47, 48, 49) for storing a plurality of text documents;andprocessor means (32, 33, 34, 42, 43, 44), coupled to the front end processing means and the storage means, for performing a search of the text documents using a plurality of search terms provided by the user, for calculating a score for each of the text documents containing at least one of the search terms, for ranking the text documents based on their scores, and for providing to the front end processing means a predetermined number of retrieved documents that are a subset of the text documents based on the documents' ranks, the retrieved documents having higher ranks than text documents not provided to the front end processing means, wherein the scores are calculated using a formula that varies according to the square of the frequency in each of the text documents of each of the search terms, where the document frequency is defined as the number of documents within a searched collection which contain the search term. Assoziatives Textsuch- und -retrievalsystem mit: Datenübertragungsvorrechnern (56, 57, 58) zum Dialog mit einem Netz mit einer oder mehreren angeschlossenen Benutzerstationen, um Informationen in das System einzugeben und Informationen aus dem System abzurufen;Speichereinrichtungen (46, 47, 48, 49) zur Speicherung einer Anzahl von Textdokumenten;und mit den Datenübertragungsvorrechnern und den Speichereinrichtungen gekoppelter Prozessoreinrichtung (32, 33, 34, 42, 43, 44) zum Durchsuchen der Textdokumente unter Benutzung einer Anzahl von vom Benutzer festgelegten Suchbegriffen, zum Errechnen einer Auswertung für jedes der Textdokumente, in denen mindestens einer der Suchbegriffe enthalten ist, zum Festlegen einer Rangfolge der Textdokumente auf der Grundlage ihrer Auswertungen und zum Bereitstellen einer vorbestimmten Anzahl von abgerufenen Dokumenten als Teilmenge der Textdokumente auf der Basis der Rangfolge der Dokumente an die Datenübertragungsvorrechner, wobei die abgerufenen Dokumente mit höherer Rangfolge als die Textdokumente den Datenübertragungsvorrechnern nicht zur Verfügung gestellt werden, wobei die Auswertungen anhand einer Formel errechnet werden, die in Abhängigkeit vom Quadrat der Häufigkeit eines jeden Suchbegriffs in jedem der Textdokumente variiert, und wobei die Dokumentenhäufigkeit als die Anzahl der Dokumente innerhalb einer durchsuchten Sammlung definiert wird, in denen der Suchbegriff enthalten ist Système associatif de recherche et d'extraction de textes, comprenant : - des moyens de traitement frontal (56, 57, 58), pour interagir avec un réseau ayant un ou plusieurs terminaux d'utilisateurs connectés à celui-ci, afin de permettre à un utilisateur de fournir des informations au système et de recevoir des informations du système ;- des moyens de mémorisation (46, 47, 48, 49) pour mémoriser une pluralité de documents de textes ;et- des moyens processeurs (32, 33, 34, 42, 43, 44), couplés aux moyens de traitement frontal et aux moyens de mémorisation, afin d'effectuer une recherche des documents de textes en utilisant une pluralité de termes de recherche fournis par l'utilisateur, de calculer un compte de points pour chacun des documents de textes contenant au moins un des termes de recherche, de classer des documents de textes en se fondant sur leurs comptes de points et afin de délivrer aux moyens de traitement frontal un nombre prédéterminé de documents extraits qui constituent un sous-ensemble des documents de textes, basé sur les rangs de classement des documents, les documents extraits ayant des rangs plus élevés que les documents de textes n'ayant pas été fournis aux moyens de traitement frontal, dans lequel les comptes de points sont calculés en utilisant une formule qui varie suivant le carré de la fréquence, dans chacun des documents de textes, de chacun des termes de recherche, la fréquence de documents étant définie comme étant le nombre de documents, à l'intérieur d'une collection recherchée, qui contiennent le terme de recherche.
  2. 2
    Assoziatives Textsuch- und -retrievalsystem nach Anspruch 1, wobei die Formel in Abhängigkeit von einer umgekehrten Dokumentenhäufigkeit bei jedem der Suchbegriffe ebenfalls veränderlich ist. Système associatif de recherche et d'extraction de textes, selon la revendication 1, dans lequel la formule varie également suivant l'inverse de la fréquence, dans le document, de chacun des termes de recherche. The associative text search and retrieval system, according to claim 1, wherein the formula also varies according to an inverse document frequency of each of the search terms.
  3. 3
    Assoziatives Textsuch- und -retrievalsystem nach Anspruch 2, wobei die Formel wie folgt lautet:Formel    wobei nt = die Gesamtanzahl der Suchbegriffe, ut = eine Anzahl eindeutiger Suchbegriffe, die in einem bestimmten Textdokument vorkommen, tfi = eine Häufigkeit, mit der Suchbegriff i im Textdokement enthalten ist, oc = ein prozentuales Vorkommen von Suchbegriffen in einem Fliestextfenster mit einer maximalen Anzahl von Suchbegriffen, wobei oc durch Teilen der Häufigkeit von Suchbegriffen im Fenster durch eine Gesamthäufigkeit von Suchbegriffen im Dokument und durch anschliessendes Multiplizieren des Ergebnisses mit 100 errechnet wird, dfi = eine Anzahl der Textdokumente, in denen der Begriff i enthalten ist, maxdfi = eine maximale Anzahl von Textdokumenten, in denen irgendwelche Suchbegriffe vorkommen, und wobei für alle Logarithmen die Zahlenbasis 2 gilt. Système associatif de recherche et d'extraction de textes, selon la revendication 2, dans lequel la formule est : dans laquelle nt représente le nombre total de termes de recherche, ut représente le nombre de termes de recherche uniques apparaissant dans un document particulier des documents de textes, tfi représente le nombre de fois où le terme de recherche i apparaît dans le document de textes, oc représente le pourcentage d'apparition de termes de recherche à l'intérieur d'une fenêtre flottante de textes, contenant un nombre maximum de termes de recherche, et est calculé en divisant le compte d'apparitions de termes de recherche dans la fenêtre par le nombre total d'apparitions de termes de recherche dans le document, et en multipliant le résultat par cent, dfi est le compte de documents de textes contenant le terme i, maxdfi est le nombre maximum de documents de textes dans lesquels apparaît un quelconque des termes de recherche, et tous les logarithmes sont en base deux. The associative text search and retrieval system, according to claim 2, wherein the formula is: wherein nt represents a total number of search terms, ut represents a number of unique search terms that occur in a particular one of the text documents, tfi represents a number of times search term i occurs in the text document, oc represents a percentage of occurrences of search terms in a floating text window containing a maximum number of search terms and is calculated by dividing a count of occurrences of search terms in the window by a total number of occurrences of search terms in the document and then multiplying the result by one hundred, dfi is a count of the text documents that contain term i, maxdfi is a maximum number of the text documents in which any of the search terms occur, and all logs are in base two.
  4. 4
    Assoziatives Textsuch- und -retrievalsystem nach irgendeinem der vorstehenden Ansprüche, wobei die Prozessoreinrichtung umfasst:mindestens einen mit den Datenübertragungsvorrechnem gekoppelten Session Administrator (SA)-Computer (42, 43, 44);und mindestens einen mit dem SA-Computer und den Dokumentenspeichereinrichtungen verbundenen Search and Retrieval [Such- und Retrieval-](SR)-Computer (32, 33, 34), wobei der SR-Computer die Suche in den Dokumentenspeichereinrichtungen durchführt und die abgerufenen Dokumente an den SA-Computer zurückgibt und wobei der SA-Computer den Benutzer auffordert, Suchbegriffe und Suchoptionen einzugeben, die Suchanforderung an den entsprechenden SR-Computer gibt und dem Benutzer die Möglichkeit bietet, die vom SR-Computer an den SA-Computer zurückgegebenen abgerufenen Dokumente einzusehen. Système associatif de recherche et d'extraction de textes, selon l'une quelconque des revendications précédentes, dans lequel lesdits moyens de traitement comprennent : - au moins un ordinateur administrateur de session (SA) (42, 43, 44) couplé aux moyens de traitement frontal ;et- au moins un ordinateur de recherche et d'extraction (SR) (32, 33, 34) couplé à l'ordinateur SA et aux moyens de mémorisation de documents ;- dans lequel l'ordinateur SR effectue la recherche dans les moyens de mémorisation de documents et renvoie les documents extraits à l'ordinateur SA et dans lequel l'ordinateur SA invite l'utilisateur à entrer des termes de recherche et des options de recherche, délivre à l'ordinateur SR approprié la demande de recherche et permet à l'utilisateur de visualiser les documents extraits renvoyés à l'ordinateur SA par l'ordinateur SR. The associative text search and retrieval system, according to any preceding claim, wherein said processing means comprises: at least one Session Administrator (SA) computer (42, 43, 44) coupled to the front end processing means;andat least one Search and Retrieval (SR) computer (32, 33, 34) coupled to the SA computer and to the document storage means,    wherein the SR computer performs the search on the document storage means and returns the retrieved documents to the SA computer and wherein the SA computer    prompts the user to enter search terms and search options, provides the appropriate SR computer with the search request, and allows the user to view the retrieved documents returned to the SA computer by the SR computer.
  5. 5
    Assoziatives Textsuch- und -retrievalsystem nach Anspruch 4, wobei die Suchanforderung vom SA-Computer an mehr als nur einen SR-Computer gegeben wird, wobei die SR-Computer die Dokumentenauswertung für die während der Suche gefundenen Textdokumente errechnen und wobei der SA-Computer die Auswertungen abgleicht und die Reihenfolge der Dokumente in Abhängigkeit von ihrer Auswertung bestimmt und die SR-Computer veranlasst, eine Teilmenge der Textdokumente mit den höchsten Gesamtauswertungen zurückzugeben. Système associatif de recherche et d'extraction de textes, selon la revendication 4, dans lequel la demande de recherche est délivrée par l'ordinateur SA à plus d'un ordinateur SR, les ordinateurs SR calculent des comptes de points de documents pour des documents de textes trouvés au cours de la recherche et renvoient les comptes de points de documents à l'ordinateur SA et l'ordinateur SA fusionne les comptes de points et effectue un classement des documents suivant leurs comptes de points et demande aux ordinateurs SR de renvoyer un sous-ensemble des documents de textes ayant les rangs de classement général les plus élevés. The associative text search and retrieval system, according to claim 4, wherein the search request is provided by the SA computer to more than one SR computer, the SR computers calculate document scores for text documents found in the course of the search and return the document scores to the SA computer, and the SA computer merges the scores and ranks the documents according to their scores and requests the SR computers to return a subset of the text documents having the highest overall ranks.
  6. 6
    Assoziatives Textsuch- und -retrievalsystem nach irgendeinem der vorstehenden Ansprüche, desweiteren mit:einem Thesaurus (52, 53, 54) mit Eintragungen für eine Anzahl von Wörtern zur Korrelation eines jeden Worts mit sowohl Synonymen als auch morphologischen Variationen. Système associatif de recherche et d'extraction de textes, selon l'une quelconque des revendications précédentes, comprenant, en outre : - un thésaurus (52, 53, 54) ayant des entrées pour une pluralité de mots, qui effectue une corrélation entre chaque mot et à la fois des synonymes et des variations morphologiques. The associative text and retrieval system, according to any preceding claim, further comprising: a thesaurus (52, 53, 54) having entries for a plurality of words which correlate each word with both synonyms and morphological variations.
  7. 7
    Assoziatives Textsuch- und -retrievalsystem nach irgendeinem der vorstehenden Ansprüche, desweiteren mit:Einrichtungen, um dem Benutzer unabhängig von den vom ihm vorgegebenen Suchbegriffen die Eingabe obligatorischer Begriffe zu ermöglichen, die in jedem der abgerufenen Dokumente vorkommen müssen, wobei die Prozessoreinrichtung Auswertungen nur für solche Dokumente vornimmt, in denen die obligatorischen Begriffe ggf. enthalten sind. Système associatif de recherche et d'extraction de textes, selon l'une quelconque des revendications précédentes, comprenant, en outre : - des moyens pour permettre à l'utilisateur d'entrer des termes obligatoires, devant être présents dans chacun des documents extraits, séparément des termes de recherche fournis par l'utilisateur, dans lequel les moyens processeurs calculent des comptes de points seulement pour les documents contenant les termes obligatoires, s'il en existe. The associative text search and retrieval system, according to any preceding claim, further comprising: means for allowing the user to enter mandatory terms which must be present in each of the retrieved documents, separately from the search terms provided by the user, wherein the processor means calculates scores only for documents containing the mandatory terms, if any.
  8. 8
    Assoziatives Textsuch- und -retrievalsystem nach irgendeinem der vorstehenden Ansprüche, desweiteren mit:einer Tabelle zur Erfassung von Sätzen, wobei die Tabelle Eintragungen umfasst, die für jedes Wort, das Teil eines Satzes sein kann, eine Position angibt, die das Wort in einem Satz einnehmen kann. Système associatif de recherche et d'extraction de textes, selon une quelconque revendication précédente, comprenant, en outre : - un tableau utilisé pour détecter des phrases, le tableau contenant des entrées qui, pour chaque mot pouvant faire partie d'une phrase, indiquent une position que le mot peut occuper dans une phrase quelconque. The associative text search and retrieval system, according to any preceding claim, further comprising: a table used to detect phrases, the table containing entries which, for each word that can be part of a phrase, indicate a position that the word can occupy in any phrase.
  9. 9
    Assoziatives Textsuch- und -retrievalsystem nach Anspruch 8, wobei in der Tabelle einer jeden Eintragung eine Bitmap zugeordnet ist, welche mögliche Stellungen in einem Satz der zugehörigen Eintragung angibt. Système associatif de recherche et d'extraction de textes, selon la revendication 8, dans lequel le tableau a une matrice de bits associée à chaque entrée, la matrice de bits indiquant les positions possibles, dans une phrase, de l'entrée associée. The associative text search and retrieval system, according to claim 8, wherein the table has a bitmap associated with each entry, the bitmap indicating possible locations in a phrase of the associated entry.
  10. 10
    Assoziatives Textsuch- und -retrievalsystem nach Anspruch 9, wobei die Tabelle desweiteren eine Kennung besitzt, um die Darstellungen eines jeden Wortes durch Zuordnung einer unverwechselbaren beliebigen Nummer zum Darstellen eines jeden Wortes zu komprimieren. Système associatif de recherche et d'extraction de textes, selon la revendication 9, dans lequel le tableau a, en outre, une identification (ID) pour comprimer les représentations de chacun des mots en attribuant un numéro arbitraire unique pour représenter chaque mot. The associative text search and retrieval system, according to claim 9, wherein the table further has an ID for compressing the representations of each of the words by assigning a unique arbitrary number to represent each word.
  11. 11
    Assoziatives Textsuch- und -retrievalsystem nach Anspruch 9 oder Anspruch 10, wobei jede Bitmap-Eintragung eine Länge von einem Byte hat, wobei ein Wert 1 einer bestimmten Bitposition in der Bitmap-Eintragung anzeigt, dass das der Bitmap zugeordnete Wort an der entsprechenden Stelle in einem Satz vorkommen könnte, und wobei ein Wert 0 einer bestimmten Bitposition in der Bitmap-Eintragung bedeutet, dass das der Bitmap zugeordnete Wort an der entsprechenden Stelle in einem Satz nicht vorkommen könnte. Système associatif de recherche et d'extraction de textes, selon la revendication 9 ou la revendication 10, dans lequel chaque entrée de la matrice de bits est une longueur de multiplet, une valeur de un à une position de bit particulière dans l'entrée de la matrice de bits indique que le mot associé à la matrice de bits peut apparaître à la position correspondante dans une phrase, et une valeur de zéro à une position de bit particulière de l'entrée de la matrice de bits indique que le mot associé à la matrice de bits ne peut apparaître à la position correspondante dans une phrase. The associative text search and retrieval system, according to claim 9 or claim 10, wherein each bitmap entry is one byte long, a value of one at a particular bit position in the bitmap entry indicates that the word associated with the bitmap could appear in the corresponding position in a phrase, and a value of zero at a particular bit position in the bitmap entry indicates that the word associated with the bitmap could not appear in the corresponding position in a phrase.
  12. 12
    Assoziatives Textsuch- und -retrievalsystem nach Anspruch 10, wobei die Tabelle einen Baum mit Knoten benutzt, die den Wörtern in einem Satz zugeordneten Kennungen entsprechen, wobei die Knoten in der Reihenfolge miteinander verbunden sind, in welcher die Wörter in einem Satz vorkommen können. Système associatif de recherche et d'extraction de textes, selon la revendication 10, dans lequel le tableau utilise une arborescence ayant des noeuds correspondant aux identifications (ID's) associées à des mots dans une phrase, les noeuds étant reliés selon l'ordre dans lequel les mots peuvent apparaître dans une phrase. The associative text search and retrieval system, according to claim 10, wherein the table uses a tree having nodes corresponding to the ID's associated with words in a phrase, the nodes being connected according to the order that the words can appear in a phrase.
  13. 13
    Assoziatives Textsuch- und -retrievalsystem nach irgendeinem der vorstehenden Ansprüche, desweiteren mit:Einrichtungen, die dem Benutzer die Möglichkeit zum Eingeben obligatorischer Begriffe bieten, die unabhängig von den vom Benutzer vorgegebenen Suchbegriffen in jedern der abgerufenen Dokumente vorkommen müssen, wobei die Prözessoreinrichtung eine Auswertung für jedes der Textdokumente errechnet, in denen die etwaigen obligatorischen Suchbegriffe und mindestens einer der Suchbegriffe enthalten sind. Système associatif de recherche et d'extraction de textes, selon l'une quelconque des revendications précédentes, comprenant, en outre : - des moyens permettant à l'utilisateur d'entrer des termes obligatoires, devant être présents dans chacun des documents extraits, séparément des termes de recherche fournis par l'utilisateur, dans lequel les moyens processeurs calculent un compte de points pour chacun des documents de textes contenant les termes obligatoires de recherche, s'il en existe, et au moins un des termes de recherche. The associative text search and retrieval system, according to any preceding claim, further comprising: means for allowing the user to enter mandatory terms which must be present in each of the retrieved documents, separately from the search terms provided by the user, wherein the processor means calculates a score for each of the text documents containing the mandatory search terms, if any, and at least one of the search terms.
  14. 14
    Assoziatives Textsuch- und -retrievalsystem nach irgendeinem der vorstehenden Ansprüche, desweiteren mit:einem den Textdokumenten zugeordneten Index zur Anzeige der Positionen potentieller Suchbegriffe innerhalb der Textdokumente;Einrichtungen, um bei der Suche Störbegriffe auszuschliessen, indem keine Störbegriffe in den Index aufgenommen werden;undEinrichtungen, um bei der Suche häufig gebrauchte Begriffe auszuschliessen, wobei die häufig benutzten Begriffe im Index enthalten und in einer Liste häufig benutzter Begriffe aufgelistet sind und wobei die häufig benutzten Begriffe dadurch von der Suche ausgeschlossen bleiben, dass bei der Suche keine in der Liste enthaltene Begriffe benutzt werden. Système associatif de recherche et d'extraction de textes, selon l'une quelconque des revendications précédentes, comprenant, en outre : - un index, associé aux documents de textes, indiquant les positions des termes potentiels de recherche dans les documents de textes ;- des moyens pour exclure que les termes erratiques soient considérés pour la recherche, en n'incluant pas les termes erratiques dans l'index ;et- des moyens pour exclure que les termes fréquemment utilisés soient considérés pour la recherche, les termes fréquemment utilisés étant contenus dans l'index et maintenus dans une liste de termes fréquemment utilisés, les termes fréquemment utilisés étant exclus de la recherche en n'utilisant pas les termes dans la liste pour la recherche. The associative text search and retrieval system, according to any preceding claim, further comprising: an index, associated with the text documents, for indicating the locations of potential search terms within the text documents;means for excluding noise terms from being considered for the search by not including noise terms in the index;andmeans for excluding frequently used terms from being considered for the search, the frequently used terms being contained in the index and maintained in a list of frequently used terms, the frequently used terms being excluded from the search by not using terms in the list for the search.
  15. 15
    Assoziatives Textsuch- und -retrievalsystem nach irgendeinem der vorstehenden Ansprüche, desweiteren mit:Einrichtungen, um dem Benutzer eine Bildschirmmaske zur Verfügung zu stellen, die für jedes abgerufene Dokument anzeigt, welche Suchbegriffe in welchen abgerufenen Dokumenten vorkommen. Système associatif de recherche et d'extraction de textes, selon l'une quelconque des revendications précédentes, comprenant, en outre : - des moyens pour fournir à l'utilisateur un écran indiquant, pour chaque document extrait, quels termes de recherche sont présents dans quels documents extraits. The associative text search and retrieval system, according to any preceding claim, further comprising: means for providing the user with a screen indicating for each retrieved document which search terms are present in which retrieved documents.
  16. 16
    Assoziatives Textsuch- und -retrievalsystem nach Anspruch 2, desweiteren mit:Einrichtungen, um dem Benutzer eine Bildschirmmaske zur Verfügung zu stellen, die für jeden der Suchbegriffe eine Bedeutung angibt, wobei die Bedeutung der Begriffe in Abhängigkeit von der umgekehrten Dokumentenhäufigkeit des Suchbegriffs variiert. Système associatif de recherche et d'extraction de textes, selon la revendication 2, comprenant, en outre : - des moyens pour fournir à l'utilisateur un écran indiquant l'importance du terme pour chacun des termes de recherche, l'importance du terme variant suivant l'inverse de la fréquence d'apparition, dans le document, du terme de recherche The associative text search and retrieval system, according to claim 2, further comprising: means for providing the user with a screen indicating a term importance for each of the search terms wherein the term importance varies according to the inverse document frequency of the search term.
  17. 17
    Assoziatives Textsuch- und -retrievalsystem nach Anspruch 16, wobei die Bedeutung der Begriffe in Abhängigkeit von log(maxdfi/dfi) variiert, wobei für den Logarithmus die Basiszahl 2 gilt, dfi = eine Anzahl der abgerufenen Dokumente, in denen der Suchbegriff i enthalten ist, und maxdfi = eine maximale Anzahl der abgerufenen Dokumente, in denen irgendwelche Suchbegriffe vorkommen. Système associatif de recherche et d'extraction de textes, selon la revendication 16, dans lequel l'importance du terme varie suivant la fonction log (maxdfi / dfi), dans laquelle le logarithme est en base deux, dfi est un compte des documents extraits contenant le terme de recherche i et maxdfi est le nombre maximum de documents extraits dans lesquels apparaît un quelconque des termes de recherche. The associative text search and retrieval system, according to claim 16, wherein the term importance varies according to log(maxdfi/dfi), wherein the log is to the base two, dfi is a count of the retrieved documents that contain search term i, and maxdfi is a maximum number of the retrieved documents in which any of the search terms appear.
  18. 18
    Assoziatives Textsuch- und -retrievalsystem nach irgendeinem der vorstehenden Ansprüche, wobei die Speichereinrichtungen zur Speicherung von mindestens einer Dokumentensammlung mit einer Anzahl von Textdokumenten und vorbestimmten Informationen dienen, aus denen hervorgeht, wie die Dokumente in der Dokumentensammlung präsentiert werden können, wobei das System Einrichtungen umfasst, die dem Benutzer die Möglichkeit bieten, einen von vielen möglichen Befehlen zur Darstellung der abgerufenen Dokumente auf der Grundlage der vorbestimmten und in der Dokumentensammlung enthaltenen Informationen anzuwählen. Système associatif de recherche et d'extraction de textes, selon l'une quelconque des revendications précédentes, dans lequel lesdits moyens de mémorisation servent à mémoriser au moins une collection de documents, contenant une pluralité de documents de textes et des informations prédéterminées indiquant la manière dont les documents peuvent être présentés dans la collection de documents, ledit système comportant des moyens permettant à l'utilisateur de sélectionner un ordre, parmi de nombreux ordres possibles, pour la présentation des documents extraits, en se fondant sur les informations prédéterminées contenues dans la collection de documents. The associative text search and retrieval system, according to any preceding claim, wherein said storage means is for storing at least one document collection containing a plurality of text documents and predetermined information indicating how the documents in the document collection can be presented, said system including means for allowing the user to select one of many possible orders for presenting the retrieved documents based on the predetermined information contained in the document collection.
  19. 19
    Assoziatives Textsuch- und -retrievalsystem nach irgendeinem der vorstehenden Ansprüche, desweiteren mit:Einrichtungen zur Anzeige des Textes eines der abgerufenen Dokumente in einem Fenster, wobei das Fenster die höchste Fensterauswertung aller möglichen Fenster des abgerufenen Dokuments hat, wobei die Fensterauswertung auf der Häufigkeit und Unterschiedlichkeit der Suchbegriffe im Fenster basiert und wobei die Vielseitigkeit der Suchbegriffe im Fenster ausgehend von der Anzahl der Suchbegriffe im Fenster errechnet wird, denen ein anderer Suchbegriff im Fenster vorausgeht. Système associatif de recherche et d'extraction de textes, selon l'une quelconque des revendications précédentes, comprenant, en outre : - des moyens pour afficher une fenêtre de texte d'un des documents extraits, la fenêtre ayant le compte de points de fenêtre le plus élevé parmi toutes les fenêtres possibles du document extrait, le compte de points de fenêtre étant fondé sur le nombre d'apparitions et la diversité des termes de recherche dans la fenêtre, la diversité des termes de recherche dans la fenêtre étant calculée en se fondant sur le nombre de termes de recherche dans la fenêtre précédés d'un terme de recherche différent dans la fenêtre. The associative text search and retrieval system, according to any preceding claim, further comprising: means for displaying a window of text of one of the retrieved documents, the window having a highest window score of all possible windows of the retrieved document, the window score being based upon the number of occurrences and diversity of search terms in the window, the diversity of search terms in the window being calculated based on the number of search terms in the window preceded by a different search term in the window.
  20. 20
    Assoziatives Textsuch- und -retrievalsystem nach Anspruch 19, wobei die Fensterauswertung durch Hinzurechnen von 1 zur Fensterauswertung für die Anzahl der Suchbegriffe im Fenster, durch Hinzurechnen von 2 zur Fensterauswertung für jeden Suchbegriff im Fenster, dem ein anderer Suchbegriff vorausgeht, und durch Hinzurechnen von 2 zur Fensterauswertung für jeden Suchbegriff im Fenster errechnet wird, dem ein anderer Suchbegriff vorausgeht, vor dem seinerseits wiederum ein anderer Suchbegriff steht. Système associatif de recherche et d'extraction de textes, selon la revendication 19, dans lequel le compte de points de fenêtre est calculé en ajoutant un au compte de points de fenêtre pour le nombre de termes de recherche dans la fenêtre, en ajoutant deux au compte de points de fenêtre pour chaque terme de recherche, dans la fenêtre, qui est précédé d'un terme de recherche différent et en ajoutant deux au compte de points de fenêtre pour chaque terme de recherche, dans la fenêtre, qui est précédé d'un terme de recherche différent, qui est également précédé d'un terme de recherche différent. The associative text search and retrieval system, according to claim 19, wherein the window score is calculated by adding one to the window score for the number of search terms in the window, adding two to the window score for each search term in the window that is preceded by a different search term, and by adding two to the window score for each search term in the window that is preceded by a different search term that is also preceded by a different search term.
  21. 21
    Assoziatives Textsuch- und -retrievalsystem nach Anspruch 19, desweiteren mit Einrichtungen, um dem Benutzer die Eingabe von obligatorischen Begriffen zu ermöglichen, die in jedem der abgerufenen Dokumente vorkommen müssen, wobei die Fensterauswertung durch Hinzurechnen von 1 zur Fensterauswertung für die Anzahl der Suchbegriffe und obligatorischen Begriffe im Fenster, durch Hinzurechnen von 2 zur Fensterauswertung für jeden Suchbegriff und obligatorischen Begriff im Fenster, dem ein anderer Suchbegriff oder obligatorischer Begriff vorausgeht, und durch Hinzurechnen von 2 zur Fensterauswertung für jeden Suchbegriff und obligatorischen Begriff im Fenster errechnet wird, dem ein anderer Suchbegriff oder obligatorischer Begriff vorausgeht, vor dem seinerseits wiederum ein anderer Suchbegriff oder obligatorischer Begriff steht. Système associatif de recherche et d'extraction de textes, selon la revendication 19, comprenant en outre des moyens permettant à l'utilisateur d'entrer des termes obligatoires devant être présents dans chacun des documents extraits, dans lequel le compte de points de fenêtre est calculé en ajoutant un au compte de points de fenêtre pour le nombre de termes de recherche et de termes obligatoires dans la fenêtre, en ajoutant deux au compte de points de fenêtre pour chaque terme de recherche et chaque terme obligatoire dans la fenêtre qui est précédé d'un terme de recherche différent ou d'un terme obligatoire différent, et en ajoutant deux au compte de points de fenêtre pour chaque terme de recherche et chaque terme obligatoire dans la fenêtre, qui est précédé d'un terme de recherche différent ou d'un terme obligatoire différent, qui est également précédé d'un terme de recherche différent ou d'un terme obligatoire différent. The associative text search and retrieval system, according to claim 19, further comprising means for allowing the user to enter mandatory terms which must be present in each of the retrieved documents, wherein the window score is calculated by adding one to the window score for the number of search terms and mandatory terms in the window, adding two to the window score for each search term and mandatory term in the window that is preceded by a different search term or mandatory term, and by adding two to the window score for each search term and mandatory term in the window that is preceded by a different search term or mandatory term that is also preceded by a different search term or mandatory term.
  22. 22
    A method of operating an associative text search and retrieval system, comprising the steps of:performing a search of text documents (263) using a plurality of search terms provided by a user;calculating a score for each of the text documents containing at least one of the search terms using a formula that varies according to the square of the frequency in each of the text documents of each of the search terms (268);where the document frequency is defined as the number of documents within a searched collection which contain the search term;ranking the text documents based on their scores (269);andproviding the user with a predetermined number of retrieved documents (272) that are a subset of the text documents based on the ranks of the documents, the retrieved documents having higher ranks than text documents not provided. Procédé d'exploitation d'un système associatif de recherche et d'extraction de textes, comprenant les étapes consistant à : - effectuer une recherche de documents de textes (263) en utilisant une pluralité de termes de recherche fournis par l'utilisateur ;- calculer un compte de points pour chacun des documents de textes contenant au moins un des termes de recherche en utilisant une formule qui varie suivant le carré de la fréquence d'apparition, dans chacun des documents de textes de chacun des termes de recherche (268), la fréquence de documents étant définie comme le nombre de documents, à l'intérieur d'une collection faisant l'objet d'une recherche, qui contiennent le terme de recherche ;- effectuer un classement des documents de textes en se fondant sur leurs comptes de points (269) ;et- délivrer à l'utilisateur un nombre prédéterminé de documents extraits (272), qui constituent un sous-ensemble des documents de textes, fondé sur les rangs de classement des documents, les documents extraits ayant des rangs plus élevés que les documents de textes n'ayant pas été fournis. Verfahren zum Betrieb eines assoziativen Textsuch- und-retrievalsystems, zu dem die folgen-den Verfahrensschritte gehören: Durchsuchen von Textdokumenten (263) unter Benutzung einer Anzahl von Suchbegriffen, die von einem Benutzer festgelegt werden;Errechnen einer Auswertung für jedes der Textdokumente, in denen mindestens einer der Suchbegriffe enthalten ist, nach einer Formel, die in Abhängigkeit vom Quadrat der Häufigkeit eines jeden Suchbegriffs (268) in jedem der Textdokumente variiert, wobei die Dokumentenhäufigkeit als die Anzahl der Dokumente innerhalb einer durchsuchten Sammlung definiert wird, in denen der Suchbegriff enthalten ist;Festlegen einer Rangfolge der Textdokumente auf der Grundlage ihrer Auswertungen (269);und Bereitstellen einer vorbestimmten Anzahl von abgerufenen Dokumenten (272) als Teilmenge der Textdokumente auf der Basis der Rangfolge der Dokumente an den Benutzer, wobei die abgerufenen Dokumente mit höherer Rangfolge als die Textdokumente nicht zur Verfügung gestellt werden.
  23. 23
    Procédé d'exploitation d'un système associatif de recherche et d'extraction de textes, selon la revendication 22, dans lequel la formule varie également suivant l'inverse de la fréquence d'apparition, dans le document, de chacun des termes de recherche. The method of operating an associative text search and retrieval system, according to claim 22, wherein the formula also varies according to an inverse document frequency of each of the search terms. Verfahren zum Betrieb eines assoziativen Textsuch- und -retrievalsystems nach Anspruch 22, wobei die Formel in Abhängigkeit von einer umgekehrten Dokumentenhäufigkeit bei jedem der Suchbegriffe ebenfalls veränderlich ist.
  24. 24
    Procédé d'exploitation d'un système associatif de recherche et d'extraction de textes, selon la revendication 23, dans lequel la formule est :dans laquelle nt représente le nombre total de termes de recherche, ut représente le nombre de termes de recherche uniques apparaissant dans un document particulier des documents de textes, tfi représente le nombre de fois où le terme de recherche i apparaît dans le document de textes, oc représente le pourcentage d'apparition de termes de recherche à l'intérieur d'une fenêtre flottante de textes, contenant un nombre maximum de termes de recherche, et est calculé en divisant le compte d'apparitions de termes de recherche dans la fenêtre par le nombre total d'apparitions de termes de recherche dans le document, puis en multipliant le résultat par cent, dfi est le compte de documents de textes contenant le terme i, maxdfi est le nombre maximum de documents de textes dans lesquels apparaît un quelconque des termes de recherche, et tous les logarithmes sont en base deux. The method of operating an associative text search and retrieval system, according to claim 23, wherein the formula is: wherein nt represents a total number of search terms, ut represents a number of unique search terms that occur in a particular one of the text documents, tfi represents a number of times search term i occurs in the text document, oc represents a percentage of occurrences of search terms in a floating text window containing a maximum number of search terms and is calculated by dividing a count of occurrences of search terms in the window by a total number of occurrences of search terms in the document and then multiplying the result by one hundred, dfi is a count of the text documents that contain term i, maxdfi is maximum number of the text documents in which any of the search terms occur, and all logs are in base two. Verfahren zum Betrieb eines assoziativen Textsuch- und -retrievalsystems nach Anspruch 23, wobei die Formel wie folgt lautet:       Formel wobei nt = die Gesamtanzahl der Suchbegriffe, ut = eine Anzahl eindeutiger Suchbegriffe, die in einem bestimmten Textdokument vorkommen, tfi = eine Häufigkeit, mit der Suchbegriff i im Textdokument enthalten ist, oc = ein prozentuales Vorkommen von Suchbegriffen in einem Fliestextfenster mit einer maximalen Anzahl von Suchbegriffen, wobei oc durch Teilen der Häufigkeit von Suchbegriffen im Fenster durch eine Gesamthäufigkeit von Suchbegriffen im Dokument und durch anschliessendes Multiplizieren des Ergebnisses mit 100 errechnet wird, dfi = eine Anzahl der Textdokumente, in denen der Begriff i enthalten ist, maxdfi = eine maximale Anzahl von Textdokumenten, in denen irgendwelche Suchöegriffe vorkommen, und wobei für alle Logarithmen die Zahlenbasis 2 gilt
  25. 25
    Procédé d'exploitation d'un système associatif de recherche et d'extraction de textes, selon l'une quelconque des revendications 22 à 24, comprenant, en outre, l'étape consistant à :- utiliser un thésaurus ayant des entrées pour une pluralité de mots, qui effectue une corrélation entre chaque mot et à la fois des synonymes et des variations morphologiques. The method of operating an associative text search and retrieval system, according to any of claims 22 to 24, further comprising the step of: using a thesaurus having entries for a plurality of words which correlate each word with both synonyms and morphological variations. Verfahren zum Betrieb eines assoziativen Textsuch- und -retrievalsystems nach irgendeinem der Ansprüche 22 bis 24, zu dem desweiteren der folgende Verfahrensschritt gehört: Benutzung eines Thesaurus mit Eintragungen für eine Anzahl von Wörtern zur Korrelation eines jeden Worts mit sowohl Synonymen als auch morphologischen Variationen.
  26. 26
    Procédé d'exploitation d'un système associatif de recherche et d'extraction de textes, selon l'une quelconque des revendications 22 à 25, comprenant, en outre, les étapes consistant à :- permettre à un utilisateur de fournir une pluralité de termes de recherche ;- utiliser un tableau pour détecter des phrases dans les limites des termes de recherche fournis par l'utilisateur, le tableau contenant des entrées qui, pour chaque mot pouvant faire partie d'une phrase, indiquent une position que le mot peut occuper dans une phrase quelconque ;- l'étape de recherche de documents de textes étant effectuée en utilisant la pluralité de termes de recherche fournis par l'utilisateur et les phrases, s'il en existe, détectées par le tableau. The method of operating an associative text search and retrieval system, according to any of claims 22 to 25, further comprising the steps of: allowing a user to provide a plurality of search terms;using a table to detect phrases within the search terms provided by the user, the table containing entries which, for each word that can be part of a phrase, indicate a position that the word can occupy in any phrase;wherein the step of performing the search of text documents is performed using the plurality of search terms provided by the user and the phrases, if any, detected by the table. Verfahren zum Betrieb eines assoziativen Textsuch- und -retrievalsystems nach irgendeinem der Ansprüche 22 bis 25, mit desweiteren den folgenden Verfahrensschritten: Schaffung der Möglichkeit für einen Benutzer, eine Anzahl von Suchbegriffen einzugeben;Benutzung einer Tabelle zur Erfassung von Sätzen innerhalb der vom Benutzer vorgegebenen Suchbegriffe, wobei die Tabelle Eintragungen umfasst, die für jedes Wort, das Teil eines Satzes sein kann, eine Position angibt, die das Wort in einem Satz einnehmen kann;wobei das Durchsuchen der Textdokumente unter Benutzung der vom Benutzer vorgegebenen Anzahl von Suchbegriffen und der ggf. mittels der Tabelle erfassten Sätze erfolgt.
  27. 27
    Procédé d'exploitation d'un système associatif de recherche et d'extraction de textes, selon la revendication 26, dans lequel le tableau a une matrice de bits, associée à chaque entrée, la matrice de bits indiquant les positions possibles dans une phrase de l'entrée associée. The method of operating an associative text search and retrieval system, according to claim 26, wherein the table has a bitmap associated with each entry, the bitmap indicating possible locations in a phrase of the associated entry. Verfahren zum Betrieb eines assoziativen Textsuch- und -retrievalsystems nach Anspruch 26, wobei in der Tabelle einer jeden Eintragung eine Bitmap zugeordnet ist, welche mögliche Stellungen in einem Satz der zugehörigen Eintragung angibt.
  28. 28
    Procédé d'exploitation d'un système associatif de recherche et d'extraction de textes, selon la revendication 27, dans lequel le tableau a en outre une identification (ID) pour comprimer les représentations de chacun des mots en attribuant un numéro arbitraire unique pour représenter chaque mot. The method of operating an associative text search and retrieval system, according to claim 27, wherein the table further has an ID for compressing the representations of each of the words by assigning a unique arbitrary number to represent each word. Verfahren zum Betrieb eines assoziativen Textsuch- und -retrievalsystems nach Anspruch 27, wobei die Tabelle desweiteren eine Kennung besitzt, um die Darstellungen eines jeden Wortes durch Zuordnung einer unverwechselbaren beliebigen Nummer zum Darstellen eines jeden Wortes zu komprimieren.
  29. 29
    Procédé d'exploitation d'un système associatif de recherche et d'extraction de textes, selon la revendication 27 ou la revendication 28, dans lequel chaque entrée de la matrice de bits est une longueur de multiplet, une valeur de un à une position de bit particulière dans l'entrée de la matrice de bits indique que le mot associé à la matrice de bits peut apparaître à la position correspondante dans une phrase, et une valeur de zéro à une position de bit particulière dans l'entrée de la matrice de bits indique que le mot associé à la matrice de bits ne peut apparaître à la position correspondante dans une phrase. The method of operating an associative text search and retrieval system, according to claim 27 or claim 28, wherein each bitmap entry is one byte long, a value of one at a particular bit position in the bitmap entry indicates that the word associated with the bitmap could appear in the corresponding position in a phrase, and a value of zero at a particular bit position in the bitmap entry indicates that the word associated with the bitmap could not appear in the corresponding position in a phrase. Verfahren zum Betrieb eines assoziativen Textsuch- und -retrievalsystems nach Anspruch 27 oder Anspruch 28, wobei jede Bitmap-Eintragung eine Länge von einem Byte hat, wobei ein Wert 1 einer bestimmten Bitposition in der Bitmap-Eintragung anzeigt, dass das der Bitmap zugeordnete Wort an der entsprechenden Stelle in einem Satz zu finden sein könnte, und wobei ein Wert 0 einer bestimmten Bitposition in der Bitmap-Eintragung bedeutet, dass das der Bitmap zugeordnete Wort an der entsprechenden Stelle in einem Satz nicht vorkommen könnte.
  30. 30
    Procédé d'exploitation d'un système associatif de recherche et d'extraction de textes, selon la revendication 28, dans lequel le tableau utilise une arborescence ayant des noeuds correspondant aux identifications (ID's) associées à des mots dans une phrase, les noeuds étant reliés selon l'ordre dans lequel les mots peuvent apparaître dans une phrase. The method of operating an associative text search and retrieval system, according to claim 28, wherein the table uses a tree having nodes corresponding to the ID's associated with words in a phrase, the nodes being connected according to the order that the words can appear in a phrase. Verfahren zum Betrieb eines assoziativen Textsuch- und -retrievalsystems nach Anspruch 28, wobei die Tabelle einen Baum mit Knoten benutzt, die den Wörtern in einem Satz zugeordneten Kennungen entsprechen, wobei die Knoten in der Reihenfolge miteinander verbunden sind, in welcher die Wörter in einem Satz vorkommen können.
  31. 31
    Procédé d'exploitation d'un système associatif de recherche et d'extraction de textes, selon l'une quelconque des revendications 22 à 30, comprenant, en outre, les étapes consistant à :- permettre à l'utilisateur d'entrer des termes obligatoires, devant être présents dans chacun des documents extraits, séparément des termes de recherche fournis par l'utilisateur ;- dans lequel un compte de points est calculé pour chacun des documents de textes contenant les termes obligatoires de recherche, s'il en existe, et au moins un des termes de recherche. The method of operating an associated text search and retrieval system, according to any of claims 22 to 30, further comprising the steps of: allowing the user to enter mandatory terms which must be present in each of the retrieved documents, separately from the search terms provided by the user;wherein a score is calculated for each of the text documents containing the mandatory terms, if any, and at least one of the search terns. Verfahren zum Betrieb eines assoziativen Textsuch- und -retrievalsystems nach irgendeinem der Ansprüche 22 bis 30, zu dem desweiteren die folgenden Verfahrensschritte gehören: Schaffung der Möglichkeit für den Benutzer zum Eingeben obligatorischer Begriffe, die unabhängig von den vom Benutzer vorgegebenen Suchbegriffen in jedem der abgerufenen Dokumente vorkommen milssen;wobei eine Auswertung für jedes der Textdokumente errechnet wird, in denen die etwaigen obligatorischen Suchbegriffe und mindestens einer der Suchbegriffe enthalten sind.
  32. 32
    Procédé d'exploitation d'un système associatif de recherche et d'extraction de textes, selon l'une quelconque des revendications 22 à 31, comprenant, en outre, les étapes consistant à :- indiquer les positions des termes potentiels de recherche dans les documents de textes, contenus dans une collection de documents, en utilisant un index qui est associé aux documents de textes ;- exclure que les termes erratiques soient considérés pour la recherche, en n'incluant pas les termes erratiques dans l'index ;- exclure que soient considérés pour la recherche les termes fréquemment utilisés contenus dans l'index et dans la collection de documents dans une liste de termes fréquemment utilisés, la liste de termes fréquemment utilisés étant dynamique, fondée sur une variété de facteurs fonctionnels incluant la fréquence d'apparition d'un terme dans la collection de documents et la nature de la collection de documents, les termes fréquemment utilisés étant exclus de la recherche en n'utilisant pas les termes dans la liste pour la recherche ;- dans lequel un compte de points est calculé pour chacun des documents de textes contenant au moins un des termes de recherche, à l'exception des termes erratiques et des termes fréquemment utilisés exclus à ladite étape d'exclusion. The method of operating an associative text search and retrieval system, according to any of claims 22 to 31, further comprising the steps of: indicating locations of potential search terms within the text documents contained in a document collection using an index which is associated with the text documents;excluding noise terms from being considered for the search by not including noise terms in the index;excluding from being considered for the search frequently used terms contained in the index and in the document collection in a list of frequently used terms, the list of frequently used terms being dynamic, based upon a variety of functional factors including the frequency of occurrence of a term in the document collection and the nature of the document collection, the frequently used terms being excluded from the search by not using terms in the list for the search;wherein a score is calculated for each of the text documents containing at least one of the search terms except for noise terms and frequently used terms excluded in said excluding step. Verfahren zum Betrieb eines assoziativen Textsuch- und -retrievalsystems nach irgendeinem der Ansprüche 22 bis 31, zu dem desweiteren die folgenden Verfahrensschritte gehören: Anzeige der Positionen potentieller Suchbegriffe innerhalb der in einer Dokumentensammlung enthaltenen Textdokumente mittels eines den Textdokumenten zugeordneten Indexes;Ausschluss von Störbegriffen bei der Suche, indem keine Störbegriffe in den Index aufgenommen werden;Nichtberücksichtigung häufig gebrauchter Begriffe bei der Suche, wobei die häufig benutzten Begriffe im Index enthalten und in der Dokumentensammlung in einer Liste häufig benutzter Begriffe aufgelistet sind, wobei die Liste häufig benutzter Begriffe auf der Grundlage einer Vielzahl von Funktionsfaktoren einschliesslich der Häufigkeit eines Begriffs in der Dokumentensammlung und der Art derDokumentensammlung dynamisch ist und wobei die häufig benutzten Begriffe dadurch von der Suche ausgeschlossen bleiben, däss bei der Suche keine in der Liste enthaltene Begriffe benutzt werden;wobei für jedes der Textdokumente, die mit Ausnahme der nach vorgenannten Verfahrensschritten ausgeschlossenen Störbegriffe und häufig benutzten Begriffe mindestens einen Suchbegriff enthalten, eine Auswertung errechnet wird.
  33. 33
    Procédé d'exploitation d'un système associatif de recherche et d'extraction de textes, selon l'une quelconque des revendications 22 à 32, comprenant, en outre, l'étape consistant à :- indiquer, pour chaque document extrait, quels termes de recherche sont présents dans quels documents extraits. The method of operating an associative text search and retrieval system, according to any of claims 22 to 32, further comprising the step of: indicating for each retrieved document which search terms are present in which retrieved documents. Verfahren zum Betrieb eines assoziativen Textsuch- und -retrievalsystems nach irgendeinem der Ansprüche 22 bis 32, zu dem desweiteren der folgende Verfahrensschritt gehört: Anzeige für jedes abgerufene Dokument, welche Suchbegriffe in welchen abgerufenen Dokumenten vorkommen.
  34. 34
    Procédé d'exploitation d'un système associatif de recherche et d'extraction de textes, selon la revendication 33, dans lequel la présence de chacun des termes de recherche dans les documents extraits est affichée sous une forme lisible à l'oeil. The method of operating an associative text search and retrieval system, according to claim 33, wherein the presence of each of the search terms within the retrieved documents is displayed in eye-readable form. Verfahren zum Betrieb eines assoziativen Textsuch- und -retrievalsystems nach Anspruch 33, wobei das Vorkommen eines jeden Suchbegriffs innerhalb der abgerufenen Dokumente in ablesbarer Form angezeigt wird.
  35. 35
    Procédé d'exploitation d'un système associatif de recherche et d'extraction de textes, selon la revendication 34, comprenant, en outre, l'étape consistant à :- indiquer l'importance d'un terme pour chacun des termes de recherche, l'importance du terme variant suivant l'inverse de la fréquence dans le document du terme de recherche. The method of operating an associative text search and retrieval system, according to claim 34, further comprising the step of: indicating a term importance for each of the search terms wherein the term importance varies according to the inverse document frequency of the search term. Verfahren zum Betrieb eines assoziativen Textsuch- und -retrievalsystems nach Anspruch 34, zu dem desweiteren der folgende Verfahrensschritt gehört: Angabe der Bedeutung für jeden der Suchbegriffe, wobei die Bedeutung der Begriffe in Abhängigkeit von der umgekehrten Dokumentenbäufigkeit des Suchbegriffs variiert.
  36. 36
    Procédé d'exploitation d'un système associatif de recherche et d'extraction de textes, selon la revendication 35, dans lequel l'importance du terme est affichée sous une forme lisible à l'oeil. The method of operating an associative text search and retrieval system, according to claim 35, wherein the term importance is displayed in eye-readable form. Verfahren zum Betrieb eines assoziativen Textsuch- und -retrievalsystems nach Anspruch 35, wobei die Bedeutung des Begriffs in ablesbarer Form angezeigt wird.
  37. 37
    Procédé d'exploitation d'un système associatif de recherche et d'extraction de textes, selon la revendication 35 ou la revendication 36, dans lequel l'importance du terme varie suivant la fonction log (maxdfi / dfi), dans laquelle le logarithme est en base deux, dfi est un compte des documents extraits contenant le terme de recherche i et maxdfi est le nombre maximum de documents extraits dans lesquels apparaît un quelconque des termes de recherche. The method of operating an associative text search and retrieval system, according to claim 35 or claim 36, wherein the term importance varies according to log(maxdfi/dfi), wherein the log is to the base two, dfi is a count of the retrieved documents that contain search term i, and maxdfi is a maximum number of the retrieved documents in which any of the search terms appear. Verfahren zum Betrieb eines assoziativen Textsuch- und -retrievalsystems nach Anspruch 35 oder Anspruch 36, wobei die Bedeutung der Begriffe in Abhängigkeit von log(maxdfi/dfi) variiert, wobei für den Logarithmus die Basiszahl 2 gilt, dfi = eine Anzahl der abgerufenen Dokumente, in denen der Suchbegriff i enthalten ist, und maxdfi = eine maximale Anzahl der abgerufenen Dokumente, in denen irgendwelche Suchbegriffe vorkommen.
  38. 38
    Procédé d'exploitation d'un système associatif de recherche et d'extraction de textes, selon l'une quelconque des revendications 22 à 37, les documents de textes étant contenus dans une collection de documents, le procédé comprenant, en outre, l'étape consistant à :- permettre à l'utilisateur de sélectionner un ordre, parmi de nombreux ordres possibles, pour la présentation des documents extraits, fondé sur des informations prédéterminées contenues dans la collection de documents indiquant la manière dont les documents de la collection de documents peuvent être présentés. The method of operating an associative text search and retrieval system, according to any of claims 22 to 37, the text documents being contained in a document collection, further comprising the step of: allowing the user to select one of many possible orders for presenting the retrieved documents based on predetermined information contained in the document collection indicating how the documents in the document collection can be presented. Verfahren zum Betrieb eines assoziativen Textsuch- und -retrievalsystems nach irgendeinem der Ansprüche 22 bis 37, wobei die Textdokumente in einer Dokumentensammlung enthalten sind und zu dem desweiteren der folgende Verfahrensschritt gehört: Schaffung der Möglichkeit für den Benutzer, einen von vielen möglichen Befehlen zur Darstellung der abgerufenen Dokumente auf der Grundlage der vorbestimmten und in der Dokumentensammlung enthaltenen Informationen anzuwählen, die zeigen, wie die Dokumente in der Dokumentensammlung präsentiert werden können.
  39. 39
    Procédé d'exploitation d'un système associatif de recherche et d'extraction de textes, selon l'une quelconque des revendications 22 à 38, comprenant, en outre, les étapes consistant à :- pour un document sélectionné parmi les documents extraits, calculer un compte de points de fenêtre pour chaque fenêtre parmi toutes les fenêtres possibles du document extrait en se fondant sur le nombre d'apparitions et la diversité des termes de recherche dans la fenêtre, la diversité des termes de recherche dans la fenêtre étant calculée en se fondant sur le nombre de termes de recherche, dans la fenêtre, qui sont précédés d'un terme de recherche différent dans la fenêtre ;et- afficher la fenêtre de texte, pour le document extrait sélectionné, ayant le compte de points de fenêtre le plus élevé parmi toutes les fenêtres possibles du document extrait sélectionné. The method of operating an associative text search and retrieval system, according to any of claims 22 to 38, further comprising the steps of: for a selected one of the retrieved documents, calculating a window score for each of all possible windows of the retrieved document based upon the number of occurrences and diversity of search terms in the window, the diversity of search terms in the window being calculated based on the number of search terms in the window preceded by a different search term in the window;anddisplaying the window of text for the selected retrieved document having the highest window score of all possible windows of the selected retrieved document. Verfahren zum Betrieb eines assoziativen Textsuch- und -retrievalsystems nach irgendeinem der Ansprüche 22 bis 38, zu dem desweiteren die folgenden Verfahrensschritte gehören: für ein bestimmtes abgerufenes Dokument Errechnen einer Fensterauswertung für jedes aller möglichen Fenster des abgerufenen Dokuments auf der Grundlage der Häufigkeit und Unterschiedlichkeitder Suchbegriffe im Fenster, wobei die Vielseitigkeit der Suchbegriffe im Fenster ausgehend von der Anzahl der Suchbegriffe im Fenster errechnet wird, denen ein anderer Suchbegriff im Fenster vorausgeht;undAnzeige des Textes des jeweils gewählten abgerufenen Dokuments in einem Fenster, das die höchste Fensterausweitung aller möglichen Fenster des gewählten abgerufenen Dokuments hat.
  40. 40
    Procédé d'exploitation d'un système associatif de recherche et d'extraction de textes, selon la revendication 39, dans lequel le compte de points de fenêtre est calculé en ajoutant un au compte de points de fenêtre pour le nombre de termes de recherche dans la fenêtre, en ajoutant deux au compte de points de fenêtre pour chaque terme de recherche, dans la fenêtre, qui est précédé d'un terme de recherche différent, et en ajoutant deux au compte de points de fenêtre pour chaque terme de recherche, dans la fenêtre, qui est précédé d'un terme de recherche différent qui est également précédé d'un terme de recherche différent. The method of operating an associative text search and retrieval system, according to claim 39, wherein the window score is calculated by adding one to the window score for the number of search terms in the window, adding two to the window score for each search term in the window that is preceded by a different search term, and by adding two to the window score for each search term in the window that is preceded by a different search term that is also preceded by a different search term. Verfahren zum Betrieb eines assoziativen Textsuch- und -retrievalsystems nach Anspruch 39, wobei die Fensterauswertung durch Hinzurechnen von 1 zur Fensterauswertung für die Anzahl der Suchbegriffe im Fenster, durch Hinzurechnen von 2 zur Fensterauswertung für jeden Suchbegriff im Fenster, dem ein anderer Suchbegriff vorausgeht, und durch Hinzurechnen von 2 zur Fensterauswertung für jeden Suchbegriff im Fenster errechnet wird, dem ein anderer Suchbegriff vorausgeht, vor dem seinerseits wiederum ein anderer Suchbegriff steht.
  41. 41
    Procédé d'exploitation d'un système associatif de recherche et d'extraction de textes, selon la revendication 39, comprenant, en outre, l'étape consistant à permettre à l'utilisateur d'entrer des termes obligatoires, devant être présents dans chacun des documents extraits ;- dans lequel le compte de points de fenêtre est calculé en ajoutant un au compte de points de fenêtre pour le nombre de termes de recherche et de termes obligatoires dans la fenêtre, en ajoutant deux au compte de points de fenêtre pour chaque terme de recherche et chaque terme obligatoire dans la fenêtre qui est précédé d'un terme de recherche différent ou d'un terme obligatoire différent, et en ajoutant deux au compte de points de fenêtre pour chaque terme de recherche et chaque terme obligatoire dans la fenêtre, qui est précédé d'un terme de recherche différent ou d'un terme obligatoire différent qui est également précédé d'un terme de recherche différent ou d'un terme obligatoire différent. The method of operating an associative text search and retrieval system, according to claim 39, further comprising the step of allowing the user to enter mandatory terms which must be present in each of the retrieved documents;and    wherein the window score is calculated by adding one to the window score for the number of search terms and mandatory terms in the window, adding two to the window score for each search term and mandatory term in the window that is preceded by a different search term or mandatory term, and by adding two to the window score for each search term and mandatory term in the window that is preceded by a different search term or mandatory term that is also preceded by a different search term or mandatory term. Verfahren zum Betrieb eines assoziativen Textsuch- und -retrievalsystems nach Anspruch 39, zu dem desweiteren der Verfahrensschritt gehört, dass dem Benutzer die Eingabe von obligatorischen Begriffen ermöglicht wird, die in jedem der abgerufenen Dokumente vorkommen müssen;und wobei die Fensterauswertung durch Hinzurechnen von 1 zur Fensterauswertung für die Anzahl der Suchbegriffe und obligatorischen Begriffe im Fenster, durch Hinzurechnen von 2 zur Fensterauswertung für jeden Suchbegriff und obligatorischen Begriff im Fenster, dem ein anderer Suchbegriff oder obligatorischer Begriff vorausgeht, und durch Hinzurechnen von 2 zur Fensterauswertung für jeden Suchbegriff und obligatorischen Begriff im Fenster errechnet wird, dem ein anderer Suchbegriff oder obligatorischer Begriff vorausgeht, vor dem seinerseits wiederum ein anderer Suchbegriff oder obligatorischer Begriff steht.
Independent claims41