Method to create annotated search index and server used therein
Abstract
FIELD: information technology. SUBSTANCE: invention relates to method and server to create an annotated search index. In method extraction of search session part from history for first search request is performed, which includes first and second resource, relevant to first search request, first resource includes some of first search request search terms, and first resource is indexed by search terms included into it in first search index, second resource does not include any of first search request search terms and has not been indexed by search terms in first search index, creation of communication parameter for second resource, wherein communication parameter is based on first history parameter, which is a number of transitions between first resource and second resource in search session from history and second history parameter, which is time spent by previous user interacting with second resource in search session from history, and in response to the fact that communication parameter for second resource exceeds predetermined threshold, binding of second resource to first resource and search terms included into it, which creates annotated search index for these search terms. EFFECT: technical result consists in increase in search results completeness. 14 cl, 4 dwg

Term
No projected expiry on record.
- Priority and filed
- Granted
- Today
30 claims: 25 independent, 5 dependent
- 1A method for creating a search index annotated, the method is performed on a server and includes:1. Способ создания аннотированного поискового индекса, способ выполняется на сервере и включает в себя: 1. Способ создания аннотированного поискового индекса, способ выполняется на сервере и включает в себя: - извлечение части поисковой сессии из истории для первого поискового запроса, причем эта часть включает в себя первый ресурс и второй ресурс, которые релевантны первому поисковому запросу, первый ресурс включает в себя по меньшей мере некоторые из поисковых терминов из первого поискового запроса, и первый ресурс был проиндексирован по включенным в него поисковым терминам в первом поисковом индексе, второй ресурс не включает в себя ни один из поисковых терминов из первого поискового запроса и не был проиндексирован по поисковым терминам в первом поисковом индексе;- создание параметра связи для второго ресурса, причем параметр связи основан на первом параметре истории и втором параметре истории, первый параметр истории является числом переходов между первым ресурсом и вторым ресурсом в поисковой сессии из истории, и второй параметр истории является временем, проведенным предыдущим пользователем во взаимодействии со вторым ресурсом в поисковой сессии из истории;и - в ответ на то, что параметр связи для второго ресурса превышает заранее определенный порог, связывание второго ресурса с одним или несколькими первыми ресурсами и включенными в него или них поисковыми терминами, что создает аннотированный поисковый индекс для этих поисковых терминов.
- 2- Extraction of the search session in the history of the first search query, which part includes a first resource and a second resource that are relevant to the first search, - извлечение части поисковой сессии из истории для первого поискового запроса, причем эта часть включает в себя первый ресурс и второй ресурс, которые релевантны первому поисковому запросу, 2. Способ по п. 1, в котором параметр связи находится выше предварительно определенного порога, когда первый параметр истории является одним из следующего:1, или 2, или 3 перехода, а второй параметр истории составляет по меньшей мере 30 секунд.
- 3the first resource includes at least some of the search terms from the first search query and the first resource has been indexed by the inclusion of the search terms in the first search index, первый ресурс включает в себя по меньшей мере некоторые из поисковых терминов из первого поискового запроса, и первый ресурс был проиндексирован по включенным в него поисковым терминам в первом поисковом индексе, 3. Способ по п. 1, в котором второй ресурс является одним пунктом из следующего списка:документом, изображением, аудиофайлом, веб-страницей, твитом (записью в Твиттере), ссылкой, заголовком документа или фрагментом документа.
- 4second resource does not include any of the search terms from the first search query and was not indexed for search terms in the first search index;второй ресурс не включает в себя ни один из поисковых терминов из первого поискового запроса и не был проиндексирован по поисковым терминам в первом поисковом индексе;4. Способ по п. 1, в котором на этапе связывания второго ресурса с одним или несколькими первыми ресурсами второй ресурс связан и с первым ресурсом, и с включенными в него или них поисковыми терминами.
- 5- The creation of a communication parameter for the second resource, the communication parameter based on the first parameter and the second parameter is the history of history, - создание параметра связи для второго ресурса, причем параметр связи основан на первом параметре истории и втором параметре истории, 5. Способ по п. 1, в котором на этапе связывания второго ресурса с одним или несколькими первыми ресурсами второй ресурс связан с одним или несколькими первыми ресурсами и включенными в него или них поисковыми терминами во втором поисковом индексе, причем созданный аннотированный поисковый индекс включает в себя второй поисковый индекс и отличается от первого поискового индекса.
- 6the history of the first parameter is the number of transitions between the first resource and the second resource in the search session history, and первый параметр истории является числом переходов между первым ресурсом и вторым ресурсом в поисковой сессии из истории, и 6. Способ по п. 2, в котором параметр связи находится выше предварительно определенного порога, когда первый параметр истории является одним из следующего:1 или 2 перехода, а второй параметр истории составляет по меньшей мере 30 секунд.
- 7the second parameter is the history of the time spent by a previous user in conjunction with the second resource in the search session history;and второй параметр истории является временем, проведенным предыдущим пользователем во взаимодействии со вторым ресурсом в поисковой сессии из истории;и 7. Способ по п. 2, в котором число переходов между первым поисковым запросом и первым ресурсом равно одному.
- 8- In response to that communication parameter for the second resource is greater than a predetermined threshold, the binding of the second resource with one or more first resources, and incorporating the search terms or are, creating a search index for an annotated those search terms. - в ответ на то, что параметр связи для второго ресурса превышает заранее определенный порог, связывание второго ресурса с одним или несколькими первыми ресурсами и включенными в него или них поисковыми терминами, что создает аннотированный поисковый индекс для этих поисковых терминов. 8. Способ по п. 4, в котором первый поисковый индекс является инвертированным индексом;первый ресурс и включенные в него поисковые термины связаны друг с другом в списке (списках) словопозиций в инвертированном индексе;и на этапе связывания второго ресурса с одним или несколькими первыми ресурсами ссылку на второй ресурс вставляют в подходящий список (списки) словопозиций в инвертированном индексе, что создает аннотированный поисковой индекс.
- 114. The method of claim. 1, wherein the step of bonding the second resource with one or more first resource and the second resource is associated with a first resource, and incorporating the search terms or them. 4. Способ по п. 1, в котором на этапе связывания второго ресурса с одним или несколькими первыми ресурсами второй ресурс связан и с первым ресурсом, и с включенными в него или них поисковыми терминами. 11. Сервер для создания аннотированного поискового индекса включает в себя:интерфейс передачи данных для передачи данных по сети передачи данных поисковому кластеру, который имеет доступ к базе данных;память;процессор, функционально соединенный с интерфейсом передачи данных и памятью, причем процессор выполнен с возможностью сохранять объекты в памяти;процессор дополнительно выполнен с возможностью: - извлекать части поисковой сессии из истории для первого поискового запроса, причем эта часть включает в себя первый ресурс и второй ресурс, которые релевантны первому поисковому запросу, - создавать параметр связи для второго ресурса, причем параметр связи основан на первом параметре истории и втором параметре истории, первый параметр истории является числом переходов между первым ресурсом и вторым ресурсом в поисковой сессии из истории, и второй параметр истории является временем, проведенным предыдущим пользователем во взаимодействии со вторым ресурсом в поисковой сессии из истории;и - в ответ на то, что параметр связи для второго ресурса превышает заранее определенный порог, связывать второй ресурс с одним или несколькими первыми ресурсами и включенными в него или них поисковыми терминами, что создает аннотированный поисковый индекс для этих поисковых терминов.
- 125. The method of claim. 1, wherein the step of bonding the second resource with one or more first resources, the second resource is associated with one or more first resources, and incorporating the search terms or are in the second index search, and the search index generated annotated includes a second search index and differs from the first search index. 5. Способ по п. 1, в котором на этапе связывания второго ресурса с одним или несколькими первыми ресурсами второй ресурс связан с одним или несколькими первыми ресурсами и включенными в него или них поисковыми терминами во втором поисковом индексе, причем созданный аннотированный поисковый индекс включает в себя второй поисковый индекс и отличается от первого поискового индекса. 12. Сервер по п. 11, в котором процессор выполнен с возможностью связывать второй ресурс с первым ресурсом и с включенными в него поисковыми терминами для создания аннотированного поискового индекса.
- 147. The method of claim. 2, wherein the number of transitions between the first search query and the first resource is one. 7. Способ по п. 2, в котором число переходов между первым поисковым запросом и первым ресурсом равно одному. 14. Сервер по п. 12, в котором процессор выполнен с возможностью вносить ссылку на второй ресурс в подходящий список (списки) словопозиций в инвертированном индексе на этапе связывания второго ресурса с одним или несколькими первыми ресурсами, что создает аннотированный поисковой индекс.
- 169. The method of claim. 5, wherein the second search index is a three- or four-dimensional data array. 9. Способ по п. 5, в котором второй поисковый индекс является трех- или четырехмерным массивом данных.
- 1811. Server to create an annotated search index includes:11. Сервер для создания аннотированного поискового индекса включает в себя:
- 19data interface for transmission over a data network search cluster data, who has access to the database;интерфейс передачи данных для передачи данных по сети передачи данных поисковому кластеру, который имеет доступ к базе данных;
- 20memory;память;
- 21a processor operatively connected to the data interface and the memory, the processor being configured to store objects in memory; the processor is further operable to:процессор, функционально соединенный с интерфейсом передачи данных и памятью, причем процессор выполнен с возможностью сохранять объекты в памяти;процессор дополнительно выполнен с возможностью:
- 22- Extract the parts of a search session history for a first search query, which part includes a first resource and a second resource that are relevant to the first search, - извлекать части поисковой сессии из истории для первого поискового запроса, причем эта часть включает в себя первый ресурс и второй ресурс, которые релевантны первому поисковому запросу,
- 23- Create a connection option for a second resource, the communication parameter based on the first parameter and the second parameter is the history of history, - создавать параметр связи для второго ресурса, причем параметр связи основан на первом параметре истории и втором параметре истории,
- 24the history of the first parameter is the number of transitions between the first resource and the second resource in the search session history, and первый параметр истории является числом переходов между первым ресурсом и вторым ресурсом в поисковой сессии из истории, и
- 25the second parameter is the history of the time spent by a previous user in conjunction with the second resource in the search session history;второй параметр истории является временем, проведенным предыдущим пользователем во взаимодействии со вторым ресурсом в поисковой сессии из истории;
- 26and и
- 27- In response to that communication parameter for the second resource is greater than a predetermined threshold, the second resource to bind to one or more first resources, and incorporating the search terms or are, creating a search index for an annotated those search terms. - в ответ на то, что параметр связи для второго ресурса превышает заранее определенный порог, связывать второй ресурс с одним или несколькими первыми ресурсами и включенными в него или них поисковыми терминами, что создает аннотированный поисковый индекс для этих поисковых терминов.
- 2812. The server of claim. 11 wherein the processor is configured to bind a second resource with the first resource and incorporating the search terms to create annotated search index. 12. Сервер по п. 11, в котором процессор выполнен с возможностью связывать второй ресурс с первым ресурсом и с включенными в него поисковыми терминами для создания аннотированного поискового индекса.
- 2913. The server on P. 11, wherein the processor is configured to bind a second resource with one or more first resources and included in his or their search terms in the second search index, and created an annotated search index includes a second search index and differs from first search index. 13. Сервер по п. 11, в котором процессор выполнен с возможностью связывать второй ресурс с одним или несколькими первыми ресурсами и включенными в него или них поисковыми терминами во втором поисковом индексе, причем созданный аннотированный поисковый индекс включает в себя второй поисковый индекс и отличается от первого поискового индекса.
- 3014. The server of claim. 12 wherein the processor is configured to make reference to the appropriate resource in the second list (list) slovopozitsy inverted index at block binding of the second resource with one or more first resource, which creates an annotated search index. 14. Сервер по п. 12, в котором процессор выполнен с возможностью вносить ссылку на второй ресурс в подходящий список (списки) словопозиций в инвертированном индексе на этапе связывания второго ресурса с одним или несколькими первыми ресурсами, что создает аннотированный поисковой индекс.
Independent claims25
130 paragraphs in 4 sections, as filed
TECHNICAL FIELD OF THE INVENTION
[01] The present solution relates to the field of search engines in general and in particular to a method and apparatus for creating a keyword index annotated.
BACKGROUND
[02] The Internet provides access to a wide range of resources, such as video files, image files, audio files or web pages that contain content on specific topics, to extracts from books and news articles. The search engine in response to a search query can select one or more resources. The search query is the data that the user sends (or triggers, consciously or unconsciously, their transfer or receipt) search engine for searching to satisfy their information needs. Searches almost always contain the data as a text - such as one or more search query terms - as well as other information. For providing search results associated with the selected resources, the search system selects and evaluates resources based on their correspondence to the search query and the relative importance compared to other resources. The search results are ordered generally in the order in accordance with the rated output and in accordance with this order.
[03] Today's big data centers handle data sets containing billions of data items. In such a large set of search for specific items that meet the conditions of the search request is a task that requires considerable computing resources. This task also requires a significant amount of time, even for the most powerful processing multiprocessor computer systems. In many applications, the response time on a search query is crucial, either due to specific technical requirements or because of user expectations. To reduce the lead time a search query, use a variety of conventional methods.
[04] Generally, in the formation of a set of management system with the ability to search data elements are indexed in accordance with some or all possible search terms that may be contained in the search query. The system is created (and saved and updated) "inverted index" data set to be used when performing searches. The inverted index consists of a series of "slovopozitsy lists." Each list slovopozitsy matches the search term, and provides links to the data elements that contain the search term (or otherwise satisfy certain other conditions, which are expressed in a search term). For example, if data elements are text documents that is common in the Internet search systems, the search terms are individual words (and / or some of the most frequently used combinations thereof), and inverted indexes contain a list slovopozitsy for each word, which met at least one document. In another example, the data set is a database containing one or more very long table. The data elements are individual records (ie, rows in the table), with a number of characteristics, provided certain values in the appropriate columns of the table. Search terms are specific characteristics values or other conditions or characteristics. List slovopozitsy for the search term is a list of references (indices, serial numbers) on the record that match the search term.
[05] In order to reduce the time required to perform a search query, an inverted index is usually stored in the apparatus-speed memory (e.g., RAM), one or more computer systems, while the own data elements stored at high, but the carrier functioning slower (e.g. , on magnetic disks or laser devices or other similar high-capacity). Thus, the search request processing will include a list of one or more lists of inverted indexes slovopozitsy speed memory device, not the data itself search elements (slow-memory device). In general this allows for searches with a much higher speed.
[06] Given the amount of information on the Internet and unsystematic distribution of various resources, the user is not always easy to formulate a search query, which can be easily and quickly will result in the necessary information. In addition, in many cases, a user interested in the resource is not directly related to the search terms in the search query or search query suggestions. A page with a high relevance can not be included in the lists slovopozitsy for keyword and thus can not be found using conventional inverted index. For example, highly relevant document may be a web resource containing only graphics images that do not include any text characters or references to a search query (e.g. URL, title, etc.).
[07] There is a need to improve the existing search engines to provide more comprehensive search results and more satisfactory search user interaction.
[08] In US Patent Application Publication No. US 20070038608 discloses a computer search engine to improve the ranking of web pages and their representation on the basis of additional information related to the contents of the recovered document. Additional information is directly related to the content retrieved Web pages, but not present on the retrieved web pages and / or reference structure. New search engine looks for a common set of web pages, as well as a database containing publications and semantic web data to provide the above additional information. Information relating to the concept, is then used in determining the final page ranking, which results in more relevant and objective ranking pages.
[09] In US Patent Application Publication No. US 20130132381 discloses a system using a markup-phrases descriptions. Maybe defined set of phrases, descriptions associated with the first domain on the basis of the analysis of the first set of documents to determine the co-occurrence of phrases, descriptions of one or more name tags associated with the first domain. It can be obtained record associated with the first domain. Analysis of the second set of documents can be initiated to identify the occurrence of a joint element obtained and references one or more sets of phrases, and descriptions of contexts associated with each of the co-occurrence-references and descriptions of phrases in each of the second plurality of documents. On the basis of the identified context can be defined through the relationship between the description of the mark-obtained element and one of the phrases, descriptions.
[10] U.S. Patent No. US 8095538 discloses a system and method annotated index. This patent describes a method of programming a computer system to retrieve information in the structure of the summary, which has views of the inverted list, including the collection of a group of documents and save them in digital format, the definition of a group of annotations, referring to the group of documents and the formation index of pieces of information by grouping the group annotations annotations using the unique identifier. The method also includes generating dictionary data fragments that a unique identifier for each annotation indicates the corresponding position in the index of pieces of annotation group which has this unique identifier annotations.
SUMMARY OF THE iNVENTION
[11] The object of the present invention is to address at least some of the disadvantages of the prior art.
[12] The first object of this technical solution is the way to create an annotated search index (search index with notes). The method may be performed on the server. The method includes: the extraction of the search session in the history of the first search query, which part includes a first resource and a second resource that are relevant to the first search, the first resource includes at least some of the search terms the first search query and It has been indexed by the inclusion of the search terms in the first search index and the second resource does not include any search term first search query, and it was not indexed for search terms in the first search index; the creation of a communication parameter for the second resource, the communication parameter based on the first parameter and the second parameter history stories; and in response to that communication parameter for the second resource is greater than a predetermined threshold, the binding of the second resource with one or more first resources, and incorporating the search terms or are, creating a search index for an annotated those search terms.
[13] The first parameter is the history of the number of transitions between the first resource and the second resource in the search session history. The second parameter is the history of the time spent by a previous user in conjunction with the second session of the resource in search of stories.
[14] In some embodiments, the technical solutions of the present communication parameter is above a predetermined threshold, when the first parameter is a history of the following: 1, 2 or 3, the transition and the second parameter history is at least 30 seconds. In other embodiments, the technical solutions of the present communication parameter is above a predetermined threshold, when the first parameter is a history of the following: one or two transition history and the second parameter is at least 30 seconds. The number of transitions between the first search query and the first resource is generally 1, but in some embodiments, the technical solutions it can be more than one.
[15] In some embodiments of the present technical solution annotated search index is created by linking the second resource with the first resource and the inclusion of search terms. In alternative embodiments of the present technical solution annotated search index is created by binding the second resource with one or more first resources and included in his or their search terms.
[16] In some embodiments of the present technical solution first and second resource are independently one or more items from the list: document, image, audio file, web page, tweet (post to Twitter), link, title of the document and document fragment.
[17] In an embodiment of this technical solution the first search index is an inverted index; first resource included therein and search terms associated with each other in the list (s) slovopozitsy inverted index; and a link to the second resource is inserted into the appropriate list (lists) slovopozitsy inverted index search engine that creates an annotated index. In an alternative embodiment of the technical solution of the second resource is associated with one or more first resources and included in his or their search terms in the second search index, and created an annotated search index includes a second search index and differs from the first search index. The second search index may be, for example, three- or four-dimensional data set (ie 3 or 4 levels of data). 3 or 4 measurements include one or more items from the list: docID (document ID), breakID (ID rupture), regionID (ID field) and sourceID (source ID).
[18] Another object of the present technical solution is the system for creating an annotated search index, wherein the system includes a server. The server includes a data interface for transmission over a data network search cluster data, which has access to a database, a memory and a processor operably coupled to the data interface and the memory, the processor being configured to store objects in memory. The processor is further configured to carry out: the extraction of the search session in the history of the first search query, which part includes a first resource and a second resource that are relevant to the first search, the first resource includes at least some of the search terms the first search request and has been indexed by the inclusion of the search terms in the first search index and the second resource is not a single search term in the first search query, and it was not indexed for search terms in the first search index; the creation of a communication parameter for the second resource, the communication parameter based on the first parameter and the second parameter history stories; and in response to that communication parameter for the second resource is greater than a predetermined threshold, the binding of the second resource with one or more first resources, and incorporating the search terms or are, creating a search index for an annotated those search terms.
[19] In some embodiments of the present technical solution processor is configured to bind a second resource with the first resource and the inclusion of search terms to create an annotated search index. In other embodiments of the present technical solution processor is configured to bind a second resource with one or more first resource and the inclusion of these search terms, or to create an annotated search index.
[20] In one embodiment of the present technical solution first search index is an inverted index; first resource included therein and search terms associated with each other in the list (s) slovopozitsy inverted index; and a processor configured to insert a link to the second resource in the appropriate list (lists) slovopozitsy inverted index search engine that creates an annotated index. In another embodiment of the present technical solution processor is configured to bind a second resource with one or more first resources and included in his or their search terms in the second search index, and created an annotated search index includes a second search index and differs from the first search index . The second search index may be, for example, three- or four-dimensional data set, with 3 or 4 measurements contain one or more items from the list: docID (document ID), breakID (ID rupture), regionID (ID field) and sourceID (Source ID ).
[21] As used herein, "Server" means by a computer program running on the relevant equipment, which is able to receive requests (for example, from client devices) on the network and to carry out these inquiries or initiate these queries. The equipment can be a single physical computer or a physical computer system, but neither one nor the other is not required for this technical solution. In the context of this technical solution the use of the phrase "server" does not mean that each task (for example, the commands or queries, extract search sessions of stories) or any specific task will be obtained, made or initiated to implement the same server ( that is, the same software and / or hardware); This means that any number of software components or hardware devices may be involved in transmission / reception, execution or initiation of the execution of any request or the consequences of any request associated with a client device, all hardware and software could be a single server or multiple servers Both options are included in the term "server".
[22] As used herein, "client device" means by a hardware device that can work with the software suitable to solving the corresponding problem. Thus, examples of client devices (among others) can serve as personal computers (desktops, laptops, netbooks, etc.), smart phones, tablets, and network equipment such as routers, switches, and gateways. It should be understood that the device itself as the master client device in this context, can behave as a server with respect to other client devices. Using the expression "client device" does not exclude the possibility of using a plurality of client devices to receive / dispatch, execute or perform any task initiation or request, or the consequences of any problem or query, or the steps of any method described above.
[23] As used herein, "database" means itself any structured set of data that is independent of the concrete structure, the software for database management, computer hardware on which data is stored, used or otherwise are available for use. The database may reside on the same machine that is running a process that stores or uses the information stored in the database, or it may be on a separate equipment, for example a dedicated server or multiple servers.
[24] As used herein, "information" includes information of any kind or type, which may be stored in the database. Thus, the information includes, among other things, audiovisual works (image, video, sound, presentation, etc.), data (location data, digital data, etc.), text (reviews, comments, and , messages, etc.), documents, spreadsheets, etc.
[25] As used herein, "component" means by a software (corresponding to specific hardware context), which is necessary and sufficient to perform a specific (s) appear (s) function (s).
[26] As used herein, "computer storage media used by the computer" refers to a carrier for absolutely any type and nature, including RAM, ROM, disks (CD-ROMs, DVD-drives, floppy disks, hard disks, etc.), USB flash drives, solid state drives, tape drives, etc.
[27] As used herein, as previously described, "slovopozitsy list" for the search term - is usually a list of references to data elements in the data set, which includes the search term. In this case, it is clear that the more common search term is, the greater the number of links in slovopozitsy list. For common search terms, for example, the article "the" in English slovopozitsy list will include a reference to each data item in the dataset. Almost all other search terms such does not occur, however, there will be gaps between the data elements in the data set containing the search term, and these data elements are formed by gaps which do not contain the term. For example, assuming that the reference in the list pointed to slovopozitsy document numbers corresponding gaps in the document number will be present in slovopozitsy list.
[28] The list of common slovopozitsy for search term (i.e., a search term, which can be detected in a relatively large number of documents, but not all) will contain reference numbers in the form of documents, those documents in which the search term appears . References in slovopozitsy lists themselves are numbered in the order, but in the absence of a search term in the documents formed gaps between the numbers of documents due to missing documents. Slovopozitsy length list may be different depending on the number of data items in the data set, which includes the search term. Length slovopozitsy list may be zero, in the absence of the data set of documents that meet the search term of the query.
[29] As used herein, as described above, the "inverted index" slovopozitsy contains a number of lists.
[30] In some embodiments, the technical solutions of each set corresponds to a plurality slovopozitsy reference lists of search terms into a plurality of indexed items, the indexed elements are numbered sequentially. As described above, this scheme is used Internet search system, and indexed elements numbered consecutively numbered documents.
[31] In some embodiments, the technical solutions of each set corresponds to a plurality of lists of links slovopozitsy search terms into a plurality of indexed items, the indexed elements are grouped in order of decreasing relevancy, not dependent on the query. This scheme is typically used internet search system in which the data elements are not entered in a data set haphazardly. Usually, the items in the data set are arranged in order of their decreasing relevance, independent of the request. Thus, data elements, the likelihood that be part of search results for any given search query statistically above are arranged so that they can be found in the beginning of the search. Hence more likely that they will be found more quickly than the data in the data set, which were introduced irregularly.
[32] As used herein, the words "first", "second", "third", etc. adjectives used as solely to distinguish nouns to which they belong, from each other and not for purposes of describing a specific relation between the nouns. For example, it should be understood that the terms "first server" and "third server" does not imply any order of reference to a particular type of timeline, a hierarchy or rank (e.g.) server / server to server, as well as their use (by itself) does not imply that a "second server" must exist in a given situation. Further, as described herein in other contexts, reference to "first" element and "second" element does not exclude the possibility that one and the same actual real element. For example, in some cases, the "first" server, "second" server may be one and the same software and / or hardware, while in other cases they may be different software and / or hardware.
[33] Each embodiment of these technical solutions intended for at least one of the above purposes. It should be understood that some aspects of the present technical solution resulting from attempts to achieve the above object, and can satisfy other objectives not specifically mentioned here.
[34] Additional and / or alternative characteristics, aspects and advantages of embodiments of the present technical solution will become apparent from the following description, the accompanying drawings and the appended claims.
BRIEF DESCRIPTION OF DRAWINGS
[35] For a better understanding of the present invention and other embodiments and its characteristics, reference is made to the following description, which must be used in conjunction with the accompanying drawings, wherein:
[36] FIG. 1 is a schematic diagram of a system implemented in accordance with embodiments of the technical solutions which do not restrict its scope.
[37] FIG. 2 illustrates the history of the search session in accordance with embodiments of the technical solutions which do not restrict its scope.
[38] FIG. 3 is a schematic diagram showing annotated search index comprising a set of four-dimensional data in accordance with embodiments of the technical solutions which do not restrict its scope.
[39] FIG. 4 is a flowchart of a method performed in the system shown in FIG. 1 and constructed in accordance with embodiments of the technical solutions which do not restrict its scope.
EMBODIMENTS
[40] FIG. 1 is a schematic diagram of a system 100 configured in accordance with embodiments of the technical solutions which do not restrict its scope. It is important to bear in mind that the following description of system 100 represents a description of illustrative embodiments of the present technical solution. Thus, all of the following description as the description is only an illustrative example of this technical solution. This description is not intended to define the scope or boundaries of the establishment of technical solutions. Some useful examples of modifications of the system 100 may also be covered by the following description. The purpose of this is also the only help in understanding, rather than defining the scope and boundaries of these technical solutions. These modifications do not constitute an exhaustive list and it will be understood that other modifications are possible in the art. Furthermore, it should not be interpreted so that, where it has not yet been made, that is where examples of modifications, any possible modifications and / or that they were not described, as described, is the only member of this embodiment of the present technical solution. As will be appreciated by those skilled in the art, this is probably not the case. Furthermore, it should be appreciated that the system 100 is a specific manifestation in some relatively simple embodiment of these technical solutions, and in such cases it is presented here to facilitate understanding. As will be appreciated by those skilled in the art, many variations of the present technical solution will have a much greater complexity.
[41] In general, system 100 is configured to receive queries and perform searches (e.g., usual and vertical search) in response to these requests, and to create annotated search index in accordance with embodiments of the technical solutions which do not restrict its scope . Therefore, any version of the system adapted to process the user's search query and create an annotated search indexes, can be adapted to a specialist to perform embodiments of the present technical solution after expert herein was read.
[42] The system 100 includes an electronic device 102. The electronic device 102 is typically associated with a user (not shown) and thus may sometimes be referred to as "client device". It should be noted that the fact that the electronic device 102 is associated with the user, does not imply any particular mode of operation, as well as having to login, registration, or the like.
[43] Embodiments of the electronic device 102 is not particularly limited, but in a PC can be used as an example of the electronic device 102 (desktops, laptops, netbooks, etc.), wireless communication devices (smart phones, mobile phones, tablets, etc )., as well as network equipment (routers, switches, or gateways). The electronic device 102 includes hardware and / or software application and / or system software (or combination thereof), as known in the art, to the search application 104. In general, the goal of the search application 104 is to allow the user to (not shown) to perform a search, for example, network search using the search engine.
[44] The implementation of the search application 104 does not specifically limited. One example of performing the search application 104 is to access the search application 104. For example, a search application can be caused by enter URL, associated with search engine Yandex (Yandex ™) user access to the website corresponding to the search engine:<u>www.yandex.ru</u>. It is important to bear in mind that the search application 104 may be caused by any other commercially available or proprietary search engine.
[45] In other embodiments, the technical solutions, not limiting its scope, the search application 104 may be a browser application on a portable device (e.g., wireless communication device). For those cases (but not only) when the electronic device 102 is a portable device such as, for example, Samsung ™ Galaxy ™ Sill, the electronic device may use a browser application Yandex. It is important to bear in mind that any other commercially available or proprietary browser application may be used to implement embodiments of the technical solutions without limiting its scope.
[46] In general, the search application 104 includes a search query interface 106 and interface 108 search results. The main task 106 searches the interface is to enable the user (not shown) to enter your request or "search terms." The main task of the interface 108 is to provide the search results of search results corresponding to the user search query 210, which was introduced in the search query interface 106.
[47] There are also data network connected server 116. The server 116 may be a conventional computer server. In the example embodiment of the present technical solution server 116 may be a Dell ™ PowerEdge ™ server, which uses the operating Microsoft ™ Windows Server ™ system. Needless to say, the server 116 may be any other suitable hardware and / or software application and / or system software. In the illustrated embodiment, the technical solutions, not limiting its scope, the server 116 is a single server. In other embodiments, the technical solutions, not limiting its scope, the server 116 functionality may be divided and may be performed by multiple servers.
[48] The electronic device 102 is adapted to communicate with the server 116 via a data line 112. In general, the data line 112 provides the electronic device 102 to perform access to the server 116 via a communication network (not shown). In some embodiments, the technical solutions, not limiting its scope, the data network (not shown) may be an Internet. In other embodiments, the technical solutions of the present data communication network (not shown) may be implemented differently - as a global data communication network, a local data network, private data networks, etc.
[49] The implementation of the data transmission line 112 is not limited and will depend on how the electronic device 102 is used. By way of example, and not limitation, in these embodiments of the present technical solution when the electronic device 102 is a wireless communication device (e.g., smart phone), a data line 112 is a wireless data network (e.g., inter alia, the line 3G data, 4G data connection, wireless internet wireless Fidelity or short WiFi®, Bluetooth®, etc.). In those instances where the electronic device 102 is a laptop computer, the data line can be a wireless (Wi-Fi or Wireless Fidelity Internet short WiFi®, Bluetooth®, etc.) or wired (Ethernet connection to a network-based).
[50] The server 116 operatively connected (or otherwise has access to) a cluster search 118. According to these embodiments of the technical solutions cluster 118 performs search in general and / or vertical search in response to a user's search query entered via interface 106 search queries and displays the search results for presentation to the user via the search results interface 108. within these embodiments of the present technical solution, without limiting it, the search cluster 118 includes or has access to database 122. As is known in the art, database 122 stores information associated with a plurality of resources potentially available via a data network (e.g., the resources available on the internet). The database 122 also stores information and data, such as the characteristics of the search histories for a particular search query 210 (e.g., history of search sessions), an inverted index, etc. The process of filling and maintenance of the database 122 is commonly known as "data collection" ( "crawling" from the English. "Crawling").
[51] The implementation of the database 122 does not specifically limited. It should be understood that any suitable data storage equipment could be used. In some embodiments, the technical solutions database 122 may be physically combined with the search cluster 118, i.e. not necessary that they are separate pieces of hardware, as shown, although they may be separate pieces of hardware. In the illustrated embodiment, the technical solutions, not limiting its scope, the database 122 is a single database. In alternate non-limiting embodiments, the technical solutions of the present data base 122 can be divided into one or more separate database (not shown). These separate databases may be parts of the same physical database or can be implemented as separate physical units. For example, one database in, for example, database 122 may store an inverted index, and another database in the database 122 may store the available resources, while another database in the database 122 can store characteristics search histories relating to specific search requests (ie search sessions in the history). Needless to say, the above example is illustrative only and other additional features are possible for implementing embodiments of the present technical solution.
[52] It is important to note that to simplify the following description of the search cluster configuration database 118 and 122 has been greatly simplified. It is believed that those skilled in the art can understand the implementation details of the search cluster 118 and its components, and the database 122.
[53] In general, the search query 210 may be viewed as a series of one or more search terms and search terms of the query can be represented as T<sub>1</sub>, T<sub>2</sub>, ... T<sub>n</sub>. Thus the search query 210 may be understood as a request for a search application 104 in the document define each data set (not shown) stored in the database 122, each containing a search term T<sub>1</sub>, T<sub>2</sub>, ... Τ<sub>n</sub> (Logical equivalent of "and" between search terms, ie in each document in the search results must meet at least once word T<sub>i</sub>For each i from 1 to n. This is the simplest form of implementation of the search query 210.
[54] In these embodiments, the technical solutions server 116 is configured to perform access to the search cluster 118 (to implement a regular web search and / or vertical search, for example, in response to a search query 210). Within the embodiment of the present technical solution, shown in FIG. 1, the server 116 is configured to: (i) to conduct searches (via access to the search cluster 118); (Ii) to analyze the search results and ranking of search results; (Iii) to group the results and compile the results page (the SERP) for output to the electronic device 102 in response to a search request 210 (not shown).
[55] In accordance with non-limiting embodiments of the technical solutions server 116 is further configured to create a searchable index annotated using: extracting a portion of a search session 200 (Fig. 2) of the stories for the search query 210 (search session history is stored 200, e.g. in the database 122); create a connection parameter for the associated document 280 identified in a search of the history of the session 200, but not having any of the search query terms 210; and in response to that communication parameter exceeds a pre-defined threshold, the coupling 280 associated with the document search query 210 or 220 the first resource to create annotated search index 300 (Fig. 3) (which may be stored, for example, in database 122) . Annotated search index 300 then becomes available for use in future searches when users search results for a search query 210 in the context of the system 100, as described above.
[56] As is known to those skilled in the art, the index is used to increase search efficiency for large data sets. Thus, one of the technical fields in which the present technical solution is applicable, is the area of search applications that use, for example, Internet search system as described above, although the present solution may be used in other fields (for example, against large databases). Embodiments of the present technical solution, as described here, mention Internet search system, because they are a good example for illustration and understanding, but it is important to remember that this solution is not limited to Internet search-engines.
[57] Search online system will usually have access (through the search application 104 and the server 116) to a set of data in the database 122, including, among other things, a very large number of web pages from the Internet, which, together with related hyperlinks to them, may be referred to as "documents." Typically, the data set contains the other resources available on the Internet, not just documents; For simplicity in understanding the examples described herein involve only documents, but it should be understood that the present solution applies all types of resources in the data set. Non-limiting examples of other resource types include images, audio files, web pages, links, titles of documents and document fragments.
[58] The documents are usually made to the data set by performing the background of web pages indexing process, which is generally referred to in the art as a "crawler" (eng. "Crawler"). The total number of documents in the collection of data and for indexing documents reputed searchable may typically be in the range of 10 to 100 billion, depending on various factors such as the volume data set language (i.e., a set of content data on a document language or more). In a non-limiting embodiment of the present technical solution, shown in FIG. 1, as part of the search cluster 118 can be implemented web search robot that sends its results to the database 122. Typically, web search robot carries out a systematic automated web browsing to find new or newly updated web pages.
[59] The process of indexing a document usually consists of determining what the words (in any language), which web addresses (hyperlinks, also referred to herein as "links"), and / or any other specific terms, which are considered potential search terms, appear in the document. In some cases, certain phrases (for example, a sequence of words) may be considered as the search terms, in which case these phrases themselves will become a part of the indexing process. In some processes the search term indexing documents will include various lexical representations, such as various grammatical forms of the words of the base. What will be used as a search term, and that - no, usually determined by the specific search policy of the search engine. Visitor services general Internet search systems usually consider each word in any language as a valid search term.
[60] For any given search term (for example, words, links, a specific term or phrase) document indexing process creates and maintains a list of links to documents that contain the search term - list slovopozitsy this search term. Thus slovopozitsy list search term data set contains a reference to each document in the set, where the search term occurs at least once. Link to document (generally called "slovopozitsiey" - this term has occurred, the term "slovopozitsy list") may be, for example, the number of the document. Each list slovopozitsy formed by ascending numbers of documents that are referenced. As an example for the list slovopozitsy term in this data set can begin with the number 5 and a document include, in order, document numbers 7, 8, 40, 41, 64 and so on. The list will not include any said number is not less than 64 (in this example, since the search term does not appear in such documents such numbers). Thus, such slovopozitsy list can be represented as {5, 7, 8, 40, 41, 64, ...}.
[61] Regarding the implementation of the search query example query Q = {T<sub>1</sub>, T<sub>2</sub>, T<sub>3</sub>} Refers to as "find all documents that meet each search term (usually word) T<sub>1</sub>, T<sub>2</sub> and T<sub>3</sub>". It should also be understood that slovopozitsy lists that correspond to the search terms will be marked as P<sub>1</sub>, R<sub>2</sub> and R<sub>3</sub> respectively. This is a specific case of a more general search query Q = {T<sub>1</sub>, T<sub>2</sub>, ... T<sub>n</sub>} With n search terms. This particular case will be considered only for purposes of illustration and simplification.
[62] The procedure for performing the search query is an iterative process that will create a new list slovopozitsy the R, having been found by the search results, ie, document numbers of these documents (in ascending order) that meet all the criteria of the search query Q (in which, for example, there is one of the search terms the previous example T<sub>1</sub>, T<sub>2</sub>, T<sub>3</sub>).
[63] Many systems are well-known indexing documents, it is assumed that those skilled in the art can understand the details of the embodiments of the present technical solutions for the creation and maintenance documents indexing system.
[64] For simple illustration and as an aid to understanding FIG. 2 is a schematic diagram of a search session, 200 the history of the first keyword ( "Query1") (210). In search of the history of the session 200 is the first document ( "document 1") 220 was extracted and displayed in search results interface 108 after the introduction of zaprosa1 210 in interface 106 searches for slovopozitsy document1 220 zaprose1 210 in the inverted index 230. After the transition to 220 user document1 reformulated as a zapros2 Query1 210 240. in response to zapros2 240 was extracted and displayed in search results interface 108 of the second document ( "document2") 260. from dok2 260 user switched to the related document ( "linked document") 280.
[65] Thus, in an embodiment of the present technical solution, shown in FIG. 2, the number of transitions between bound dokumentom1 220 and document 280 is three. It should be understood that a search session 200 of the history is shown only for illustrative purposes and may be many other variations and permutations.
[66] The search session 200 of history, though the linked document 280 does not include a single search term from zaprosa1 210 and although the transition to the associated document 280 has been executed after the zaprosa2 240 associated document 280, however, is relevant for relative to the initial search query (Query1 210). Therefore, it is desirable in this example, to create an annotated searchable index of 300, in which the linked document 280 is associated with one or more zaprosom1 dokumentom1 210 and 220 to increase the completeness of future searches.
[67] In some non-limiting embodiments of the present technical solution annotated search index 300 is created entering a reference to the linked document 280 in the appropriate list (lists) slovopozitsy for inclusion in it (them) search term (s) (and the lists slovopozitsy already include document1 220 in the inverted index 230. in such embodiments of the present technical solution annotated searchable index 300 can be considered as an extension of the original inverted index 230.
[68] In some embodiments, the technical solutions are additional codes relating to the session 200 of the search history. For example, in an embodiment of the present technical solution, as shown in FIG. 2, are 250 requests the index ( "Source 1"); Index 270 URL ( "Source 2"); and an index of 290 titles ( "Source 3"). In general, the annotations of the same type (for example, requests, documents and parts thereof) can be collected at source and are shown in the annotated search index. As an example, can be created as follows (without limitation): Source Wikipedia, Wikipedia articles containing headings shown in FIG. 2 as an index of 290 titles; Index references URL containing certain web resources depicted in FIG. 2 as an index of 270 URL; source and related inquiries, shown in FIG. 2 as an index of 250 requests. In some embodiments of the present technical solution, such sources are annotations in the annotated search index, such as an annotated searchable index of 300 (described later).
[69] In alternative embodiments of the present technical solution annotated search index 300 is created as a second search index, which is related ny document 280 is associated with one or more zaprsom1 dokumentom1 210 and 220 in the data array. The second search index may be, for example, a separate database or the index data comprising a plurality of levels or references to the linked document 280, such as a three- or four-dimensional dataset. For ease of illustration and as an aid to understanding the annotated example of a search index 300 containing an array of 305 four-dimensional data (also referred to as "data array 4D"), shown in FIG. 3, which is a schematic diagram annotated search indexes 300, 305 containing an array of data 4D.
[70] An array of 305 data 4D database contains 4 levels: the first level of 310, consisting of documents that contain line by line comparison of the query or query part (docID); a second layer 320 containing a string (for example, the description, the phrases) documents (breakID); a third layer 330 containing the area in which the user is located, when injected their requests (regionID); and a fourth layer 340 containing a certain type of annotation sources (sourceID). During query processing 210 and associated Query1 document1 220 are first determined in a standard inverted index 230, and the retrieved identifiers and doclDI breaklDI document.1 220 to that shown schematically in the cells 310 and 320 respectively in FIG. 3. These identifiers will lead to the recovery of regionID for document1 220 schematically shown in cell 330. In accordance with the identifiers docID, breakID and regionID extracted sources annotations (sourceIDs) and stored as annotations (shown schematically in section 340).
[71] Total of the data array 305 4D database contains 4 levels. In a non-limiting embodiment of the present technical solution, as shown in FIG. 3, these four levels: 1) Level 1 - docID - for example, a document identifier, which can be a string that matches the query, or a part of it (310); 2) Level 2 - break ID - the line (for example, descriptions, phrases) of documents (320); 3) Level 3 - regionID - areas in which the user is located, when injected their requests (330); and 4) the level of 4 - sourceID - annotated data from the source (eg, title, URL, link, etc.) (340).
[72] It is understood that a particular resource may be given a reference or may be identified by many different ways in the search index 300. annotated annotation method does not specifically limited. As an example, a resource such as a linked document 280 may be annotated with one or more items from the following list (without limitation): 1) a request user or part of the request (for example, after receiving a response to Query1 210, the user switched to a related document 280, the request or the keywords can be used as an annotation to any other requests); 2) The text of the link to the associated document 280, which may be associated with a related document 280. A link can contain, for example, keywords, synonyms, URL, similar words zaprosa1 210 mark, and so on; 3) the text located before and / or after the link, the link being associated with a bound document 280; 4) the title of the site description in a web resource directory index or query, and so on; 6) tweet, for example, text, tweet, pertaining to the associated document 280 may be used in a variety of other references and identifiers. It should be understood that each link identifier and making a signal that allows you to find the linked document 280 in response to the Query1 210.
[73] Although the embodiment of the present technical solution, as shown in FIG. 3 depicts an array 305 4D data, it should be understood that the present technical solution is not limited annotated search index containing 4D dataset. For example, in one embodiment of the present technical solution, as mentioned above, annotated search index may comprise an extension of an existing inverted index, for example, where an inverted index 230 is annotated by introducing a reference to a linked document list 280 to the corresponding (lists) slovopozitsy. In some non-limiting embodiments of the technical solutions inverted index 230 is the index of references. reference index is not based on the contents of documents and text-document links. In the case of the presence of the search term in the text references to the document, the document will be placed in slovopozitsy listing for this search term. Typically, each entry in the index reference contains information about the Web page to which reference is made, for example, language, geographic area, the owner of the links to the source, date and other links, and so on. In alternate non-limiting embodiments of the present technical solution annotated search index can contain two-dimensional (2D array of data) or three-dimensional (3D array of data) of the data array. There are other embodiments of the technical solutions that are part of the scope thereof.
[74] FIG. 4 is a schematic diagram of a method 400 implemented in accordance with embodiments of the technical solutions which do not restrict its scope. Method 400 can be executed on the server 116.
[75] Step 402 - extraction of the search session in the history of the first search query, which part includes a first resource and a second resource that are relevant to the first search query
[76] The method 400 begins at step 402 where the server 116 extracts part of a search session, 200 the history of the first keyword ( "Query1") 210. Part 200 of the search session history includes the first resource ( "document 1") 220 and the second resource ( "linked document") 280 that are relevant zaprosu1 210.
[77] The first resource ( "document 1") 220 includes at least some of the search terms from zaprosa1 210, and the first resource has been indexed by the inclusion of the search terms in the first search index. In some embodiments of the present technical solution first search index is an inverted index 230. Inverted index 230 does not specifically limited. For example, it can be inverted index level record (containing a list of links to documents for each search term) word level (further comprising the position of each word in the document), the index references (containing a list of references to link each containing a search term), etc. .
[78] The second resource ( "linked document") 280 does not contain the search terms of zaprosa1 210 and indexed by search terms from zaprosa1 in the first search index, for example, the inverted index 230. Thus, the linked document 280 may not have an obvious connection zaprosom1 with 210, having, for example, any of the search term associated zaprosa1 210. Therefore, the document 280 was not retrieved and displayed in the interface 108 of search results in response to search Query1 210 to 200 of the session history. However search session history 200 indicates that the associated document is relevant zaprosu1 280 210 (relevant is based on a communication parameter for the associated document 280, as described below).
[79] In some embodiments of the present technical solution first resource and a second resource is a document (for example, paper 1 220 and 280 respectively, the linked document). However, the type of resource does not specifically limited. In one non-limiting example, the first and second resource may independently be a document, images, audio files, Web pages, tweet (Twitter account), a link, the document title or document fragment. It should be understood that the first resource and the second resource may be a resource of the same type or may be different types of resources. For example, the first resource and the second resource may both be documents. Alternatively, the first resource may be a document, and the second resource may be an image. Many other options are included in the scope of the technical solutions.
[80] Step 404 - the creation of a communication parameter for the second resource, the communication parameter based on the first parameter and the second parameter is the history of history.
[81] In the continuation of the method in step 404 creates a connection option for a second resource ( "linked document") 280. The communication parameters based on the first parameter and the second parameter is the history of history.
[82] The first parameter is the history of the number of transitions between the first resource ( "document 1") 220 and a second resource ( "linked document") 280 to 200 of the search session history. In an embodiment of the technical solution shown in FIG. 2, the first parameter is 3 stories, because there were three dokumentom1 transition between 220 and 280 related to the document (the first transition from document.1 zaprosu2 220 to 240, the second transition from zaprosa2 dokumentu2 240 to 260; and the third shift from document.2 260 to 280. It Related documents note that the transition from one zaprosa1 document1 220 to 210, and 210 and Query1 document1 220 slovopozitsy indexed together in inverted index list 230. it should be understood that the first parameter is not limited to the history value of 3 and may take on different values depending on many factors, for example a particular search query, a particular search session history, relevant resource to the search query, and so on. In some embodiments of the present technical solution first parameter history is 1. In some embodiments of the present technical solution first parameter history is 2. In some embodiments, technical solutions first parameter history is 3. In certain embodiments of the present technical solution first parameter history is either 1 or 2, or 3.
[83] The second parameter is the history of the time spent by a previous user in conjunction with the second resource ( "linked document") 280 in the search session history. The amount of time the user spent interacting with the associated document 280, gives a measure of the relevance of the associated document 280. In general, the longer a user interface associated with the document 280, the greater the relevance of the associated document 280.
[84] Step 406 - In response to the fact that the second communication parameter to a resource exceeds a predetermined threshold, the binding of the second resource with the first one or more resources and are included in the search terms that creates an annotated search index for these search terms.
[85] The method 400 continues at step 406 where an annotated search index 300 is created in response to that communication parameter for the associated document 280 exceeds a predetermined threshold. In some embodiments, the technical solutions of the present communication parameter is above a predetermined threshold, when a search session history 200 of the first parameter is a history of the following: 1, 2 or 3, the transition and the second parameter history is at least 30 seconds. In other embodiments, the technical solutions of the present communication parameter is above a predetermined threshold, when a search session history 200 of the first parameter is a history of the following: one or two transition history and the second parameter is at least 30 seconds. In some embodiments, the technical solutions in the search history of the session 200 of the number of transitions between zaprosom1 dokumentom1 210 and 220 is equal to unity. However, it should be understood that in alternative embodiments, the number of technical solutions transitions between zaprosom1 dokumentom1 210 and 220 may be greater than one, i.e. two or three, in particular if there are other indicators of high relevance of the associated document 280 (e.g., the second parameter history greatly exceeds 30 seconds and stories or the first parameter is equal to one, indicating a close relationship between the linked document 280 and the document 1220).
[86] In response to the fact that the communication parameter for the associated document 280 exceeds a predetermined threshold, the associated document 280 is associated with one or more first resources and incorporated therein or which search terms that creates an annotated search index 300 for these incorporated therein search terms. In some embodiments of the present technical solution linked document 280 is associated with search terms included in it (example - annotated in them). In some embodiments of the present technical solution linked document 280 is associated with dokumentom1 220 (example - annotated in it). In some embodiments of the present technical solution linked document 280 is associated with incorporating the search terms and dokumentom1 220 (example - annotated in them) in the annotated search index 300.
[87] As mentioned above, the method of binding a document 280 associated with one or more dokumentami1 220 and incorporating the search terms or are not particularly limited in any way. For example, in some embodiments of the present technical solution first search index is an inverted index 230, and 220 document1 and included in it the search terms are linked together in a list (s) slovopozitsy inverted index 230. Annotated search index 300 can be created by inserting links in the linked document 280 in the appropriate list (lists) slovopozitsy inverted index 230. in alternative embodiments of the present technical solution annotated search index 300 is created by coupling the associated document 280 with one or more dokumentami1 220 and included in their search terms in the second search index For example, 305 data array 4D.
[88] A reference or identification information that are used to annotate a particular resource, such as 280 associated document also does not specifically limited. As mentioned above, in some embodiments, the technical solutions or the reference information for identifying one or more items from the list: docID, breakID, regionID and source ID, for example, URL, links, titles of documents, etc.
[89] It will be appreciated that the procedure described above is merely an illustrative embodiment of the technical solution. It is not intended to define or establish the scope of the technical solutions.
[90] Some of the technical effects of non-limiting embodiments of the present technical solution may include the provision of more complete search results to the user in response to a user search query. Resources that are of interest to the user, but it is not obvious related to the search query can be retrieved and displayed on the search results page SERP. Such resources may include, for example, documents that contain the search terms of the search query, such as image-circuits, which include text characters that could be signs of relevance to the search query. The provision of such resources may allow a user to more efficiently find the information that (s) he was looking for (a) and immerse yourself in a subject of interest in greater depth. The ability of a user to find information more effectively reduces the use of traffic. Also, despite the fact that the electronic device 102 is configured as a wireless data transmission device, the user's ability to find information more effectively lead to conserve the battery of the electronic device 102. As can also give the user a more attractive or interesting search interface or a search results page. It is important to note that embodiments of the technical solutions can be implemented with display and other technical results.
[91] Modifications and improvements of the embodiments described above technical solutions will be apparent to those skilled in the art. The foregoing description is only exemplary and shall not have any restrictions. Thus, the scope of the present technical solution is limited only by the scope of the appended claims.
[92] Accordingly, embodiments of the technical solutions described above can be summarized in the following claims.
[93] A method of item 1 (400) creating an annotated search index (300), the method is performed on a server (116) and includes:
[94] a) extraction of the search session (200) from the first history search request (210), which part includes a first resource (220) and the second resource (280) that are relevant to the first search query (210)
[95] The first resource (220) comprises at least some of the search terms of the first search query 210 and indexed by the inclusion of the first search term in a search index (230)
[96] The second resource (280) does not include any search term in the first search query 210 and was not indexed for search terms in the first search index (230);
[97] b) creating a communication parameter to the second resource (280), the communication parameter based on the first parameter and the second parameter history stories,
[98] The first parameter is the number of stories of transitions between the first resource (220) and the second resource (280) in a search session (200) of the stories, and
[99] The second parameter is the history of the time spent by a previous user in conjunction with the second resource (280) in a search session (200) of the stories; and
[100] c) in response to the fact that the parameter of communication for the second resource (280) exceeds a predetermined threshold, the binding of the second resource (280) with one or more first resource (220) and included in it or them search terms that creates annotated searchable index (300) for these search terms.
[101] ITEM 2. The method of claim 1, wherein the communication parameter is above a predetermined threshold, when the first parameter is a history of the following:. 1, 2 or 3, the transition and the second parameter history is at least 30 seconds.
[102] Step 3. Method according to claim 2, wherein the communication parameter is above a predetermined threshold, when the first parameter is a history of the following:. 1 or 2 of the transition, and the second parameter history is at least 30 seconds.
[103] 4. The method of item n. 2 or 3, wherein the number of transitions between the first search term (210) and the first resource (220) is equal to one.
[104] ITEM 5. Method according to any one of claims. 1-4, wherein the second resource (280) is one item from the following list: documents, images, audio files, Web pages, tweet (Twitter account), a link, the document title or document fragment.
[105] ITEM 6. Method according to any one of claims. 1-5, wherein in step c) the second resource (280) associated with the first resource (220) and incorporating the search terms.
[106] ITEM 7. The method of claim 6, wherein the first search index (230) is an inverted index.; first resource (220) included therein and search terms associated with each other in the list (s) slovopozitsy inverted index (230); and in step c) link to the second resource (280) is inserted into an appropriate list (list) slovopozitsy inverted index (230) that creates an annotated search index (300).
[107] Item 8. A method according to claim. 1-7, wherein in step c) the second resource (280) associated with one or more first resource (220) and incorporating the search terms or are in the second index search, and the search index generated annotated (300) includes a second search index and differs from the first search index (230).
[108] 9. The method of item n. 8, wherein the second search index is a three- or four-dimensional array (305) of data.
[109] ITEM 10. The method of claim 9, wherein the three or four measurements contain one or more items from the list:. DocID (310) (ID document), breakID (320) (ID rupture), regionID (330) (ID field) and sourceid (340) (source ID).
[110] ITEM 11. The server (116), comprising:
[111] The data interface for transmission on the data network search data cluster (118) that has access to a database (122);
[112] memory;
[113] processor operably coupled to the data interface and the memory, the processor being configured to store objects in memory; the processor is further configured to:
[114] a) removing part of the search session (200) from the first history search request (210), which part includes a first resource (220) and the second resource (280) that are relevant to the first search query (210)
[115] The first resource (220) includes at least some of the search terms the first search query 210, and the first resource has been indexed by the inclusion of the search terms in the first search index (230),
[116] The second resource (280) does not include any of the search terms from the first search query 210, and has not been indexed by the search terms in the first search index (230);
[117] b) to establish a communication parameter for the second resource (280), and a communication parameter based on the first parameter and the second parameter is the history of history,
[118] The first parameter is the number of stories of transitions between the first resource (220) and the second resource (280) in a search session (200) of the stories, and
[119] The second parameter is a history of the time spent in the previous user interaction with the second resource (280) in a search session (200) of the stories; and
[120] c) in response to the fact that the communication parameter for the second resource (280) exceeds a predetermined threshold, bind a second resource (280) with one or more first resource (220) and included in these search terms, which creates an annotated search index (300) for those search terms.
[121] 12. The item server (116) according to claim 11, wherein the communication parameter is above a predetermined threshold, when the first parameter is a history of the following:. 1, 2 or 3, the transition and the second parameter history is at least 30 seconds.
[122] 13. The item server (116) according to claim 12, wherein the communication parameter is above a predetermined threshold, when the first parameter is a history of the following:. 1 or 2 of the transition, and the second parameter history is at least 30 seconds.
[123] 14. The server item (116) of claim. 12 or 13, wherein the number of transitions between the first search term (210) and the first resource (220) is equal to one.
[124] 15. The item server (116) according to any one of claims. 11-14, wherein the second resource (280) is one item from the following list: documents, images, audio files, Web pages, tweet (Twitter account), a link, the document title or document fragment.
[125] 16. The item server (116) according to any one of claims. 11-15, wherein the processor is configured to bind a second resource (280) from the first resource (220) with the search terms and in step c) included therein.
[126] 17. The item server (116) according to claim 16, wherein the first search index (230) is an inverted index.; first resource (220) included therein and search terms associated with each other in the list (s) slovopozitsy inverted index (230); and wherein the processor is configured to make reference to the second resource (280) in the appropriate list (list) slovopozitsy inverted index (230) in step c), which creates an annotated search index (300).
[127] 18. The item server (116) according to any one of claims. 11-17, wherein the processor is configured to bind a second resource (280) with one or more first resource (220) and incorporating the search terms or are in the second index search in step c), the search index generated annotated (300) It includes a second search index and differs from the first search index (230).
[128] 19. The server according to item n. 18, wherein the second search index is a three- or four-dimensional array (305) of data.
Item 20. The server of claim 19, wherein 3 or 4 measurements contain one or more items from the list:. DocID (310) (ID document), breakID (320) (ID rupture), regionID (330) (ID field) and sourceid (340) (source ID).
Contents4
5 sheets
Sheet 1 Sheet 2 Sheet 3 Sheet 4 Sheet 5
Every citation, both ways
| Document | Relation | Office | Cited during |
|---|---|---|---|
| US2007022135A1 | Cites | United States of America | Search report |
| RU2473119C1 | Cites | Russian Federation | Search report |
| US7308643B1 | Cites | United States of America | Search report |
| US8095538B2 | Cites | United States of America | Search report |
| US20070022135A1 | Cites | United States of America | – |
5 members in 3 offices
Priority claims2
| Document | Office | Kind | Date |
|---|---|---|---|
| 2015121844 | Russian Federation | A | |
| RU20150121844 | – | – | – |
Members5
| Document | Office | Kind | |
|---|---|---|---|
| WO2016198927A1 | World Intellectual Property Organization (WIPO) | A1 | |
| RU2015121844A | Russian Federation | A | |
| RU2606309C2This record | Russian Federation | C2 | |
| US2017262481A1 | United States of America | A1 | |
| US9773035B1 | United States of America | B1 |
Numbers
- Publication
- 0002606309
- Publication, DOCDB
- 2606309
- Publication, EPODOC
- RU2606309
- Application
- 121844
- Application, DOCDB
- 2015121844
- Application, EPODOC
- RU20150121844
Titles2
- English
- METHOD TO CREATE ANNOTATED SEARCH INDEX AND SERVER USED THEREIN
- Russian
- ?????? ???????? ??????????????? ?????????? ??????? ? ??????, ???????????? ? ???
Classification
- CPC, 6
- G06F16/2228
- G06F16/2272
- G06F16/24539
- G06F16/24573
- G06F16/93
- G06F16/319