Contents indexing retrieval system and retrieval result providing method
Abstract
[Task] Provide a content indexing search system that filters and blocks content.
Solution.Users make search queries to search engines through a proxy server and a gateway that acts as a cache and blocking engine. The blocking engine filters and blocks the content of the search results to provide consistency with the user's content search results. Also, to implement a content blocking policy similar to the cache and filtering engine, or to build an index database by searching the cache and engine content, or to have the search engine index itself. Modify the search engine to go through the cache and filtering engines when building the attached database.
Term
Term ended
Projected expiry passed 21 April 2020, 6.4 years ago.
- Priority
- Filed
- Published
- Projected expiry
- Today
12 claims: 3 independent, 9 dependent
- 1【特許請求の範囲】 【請求項1】コンテンツのフィルタリング及びブロッキング制限と一致する検索結果を提供するコンテンツ索引付け検索システムであって、 データベースを含むコンテンツ索引付け検索エンジンと、 キャッシュを含むキャッシュ及びブロッキング・プロキシ・サーバと、 前記コンテンツ索引付け検索エンジンに結合した情報ネットワークと、 前記コンテンツ索引付け検索エンジンに検索問い合わせを行い、前記キャッシュから検索結果を受け取るための手段と、 前記コンテンツ索引付け検索エンジンと結合し、コンテンツのフィルタリング及びブロッキング・ポリシーを実行するブロッキング・エンジンと、 前記コンテンツ索引付け検索エンジンを修正して前記ブロッキング・エンジンと同じコンテンツ・ブロッキング・ポリシーを実行するための手段と、 を備えるコンテンツ索引付け検索システム。
- 2【請求項2】コンテンツ索引付けフェーズの間、前記コンテンツ索引付け検索エンジンで前記ブロッキング・ポリシーを実行するための手段をさらに備える請求項1に記載のシステム。
- 3【請求項3】エンド・ユーザ検索結果表示フェーズの間、前記ブロッキング・ポリシーを実行するための手段をさらに備える請求項1に記載のシステム。
- 4【請求項4】キャッシュ・コンテンツを索引することによって索引付けデータベースを構築するために前記コンテンツ索引付け検索エンジンを修正する手段をさらに備える請求項1に記載のシステム。
- 5【請求項5】前記コンテンツ索引付け検索エンジンが索引付けデータベースを構築するときに前記キャッシュ及びブロッキング・エンジン結果を取り込むために前記コンテンツ索引付け検索エンジンを修正するための手段をさらに備える請求項1に記載のシステム。
- 6【請求項6】データベース及びキャッシュに結合したコンテンツ索引付け検索エンジンと、前記コンテンツ索引付け検索エンジンに結合した情報ネットワークと、前記キャッシュを介してエンド・ユーザに提供される検索結果に対してコンテンツのフィルタリング及びブロッキング制限を実行するブロッキング・エンジンとを有するコンテンツ索引付け検索システムにおいて、コンテンツのフィルタリング及びブロッキング・の実施に一致する検索結果を提供する方法であって、 (a)前記コンテンツ索引付け検索エンジンのプロセスを変えて、除外パターンに一致する任意の情報サイトURLをスキップさせるステップと、 (b)前記コンテンツ索引付け検索エンジンのプロセスを変えて、明確に許容される情報サイトURLリストに一致するサイトまたはルート・コンテンツ・ソースのみを検索するステップと、 (c)前記キャッシュ及びブロッキング・エンジンで定められるフィルタリングポリシーを前記コンテンツ索引付け検索エンジンで実行するステップとを有し、前記(C)のステップで前記フィルタリングポリシーは、 (i)一定の間隔で、または変化が検知されるたびに、前記キャッシュ及びフィルタリングエンジンからコンテンツのフィルタリング規則を読み込むステップと、 (ii)多数索引付けデータベース・ツリーを生成し、各ツリーを前記コンテンツのフィルタリング規則で定められたユーザ・グループと対応付けるステップと、 (iii)除外パターンに一致する任意の情報サイト、URL、またはドキュメントのユーザに対しての表示を避けるステップと、 (iv)明確に許容される情報サイトURLリストに一致するソース由来のドキュメント/コンテンツ・ポインタを表示するステップと、 (v)前記キャッシュ及び前記ブロッキング・エンジンで定められたフィルタリングプロセスに応じる情報ネットワーク/コンテンツ/ドキュメントのみをユーザに表示するステップとによって定められ、さらに前記(v)のステップで前記フィルタリングプロセスは、 (aa)一定の間隔で、または変化が検知されるたびに、前記キャッシュ及びフィルタリングエンジンから前記コンテンツのフィルタリング規則を読み込むステップと、 (bb)ユーザに対して個々のものまたはグループに対して前記フィルタリング規則によって許容される検索結果のみを表示するステップによって定められる方法。
- 7【請求項7】(d)コンテンツ・エンジン走査標的を情報サイト/URLリストよりはむしろコンテンツ・キャッシュ記憶装置に修正するステップと、 (e)API、データベース走査、及び共有ファイル走査によって前記キャッシュ及びブロッキング・エンジンのURL/コンテンツ/ドキュメント・ツリーをトラバースするステップとをさらに有する請求項6に記載の方法。
- 8【請求項8】(f)エンド・ユーザのブラウザと同様にして構成されるように前記検索エンジン・コンテンツ走査及び索引付けプロセスを修正するステップをさらに有する請求項6に記載の方法。
- 9【請求項9】(g)索引付けデータベースを構築する時に前記キャッシュを通過するように前記コンテンツ索引付け検索エンジンを修正するステップをさらに有する請求項6に記載の方法。
- 10【請求項10】(h)前記キャッシュを検索することで索引付けデータベースを構築するために前記コンテンツ索引付け検索エンジンを修正するステップをさらに有する請求項6に記載の方法。
- 11【請求項11】(i)前記コンテンツ索引付け検索エンジンを内部ネットワークに接続するステップと、(j)前記コンテンツ索引付け検索エンジンを内部ネットワーク操作に接続するステップとをさらに有する請求項6に記載の方法。
- 12【請求項12】前記コンテンツ索引付け検索エンジンを外部ネットワークに接続して組織コンテンツ・ブロッキング・ポリシーに対して整合性を与えるステップをさらに有する請求項6に記載の方法。
Independent claims12
126 paragraphs in 1 section, as filed
Description: TECHNICAL FIELD [Detailed description of the invention]
【0001】
[Technical field to which the invention belongs]
The present invention relates to an information retrieval system. In particular, the present invention relates to content indexing search systems and methods that provide search results consistent with policies that filter and block content implemented in blocking engines.
【0002】
[Conventional technology]
With the explosive growth of text and multimedia content available on the Internet and other data networks and systems, end users are increasingly relying on text and keyword search tools to find information of interest. .. The end user enters the key word that represents the information and document they are looking for into a search tool or search engine. A search tool or search engine then searches an existing indexing database for a list of pointers to documents that you may be interested in, along with the title of the document, and often a few lines of text extracted from the body of the document. It will be returned with the description of. The end user then navigates some or all of the pointers returned for search to view the actual document or online content. Search engine indexing databases typically launch automated programs for content sources (eg, Internet websites) and root as well as links to content trees (often navigating to other sites). Automatically or semi-automatically constructed by automatically searching content sources and indexing the information contained in the database for future searches. For large content sources such as websites on the Internet, automated search and indexing is the only practical way to create an index search database.
【0003】
As the variety of information available to online systems and networks, companies, individuals, groups, and network service providers (NSPs) increases, it is considered inappropriate or unnecessary for end users. Increasingly, policies and controls are being enforced to screen content, or to limit the availability of such content. Such content management policies generally block unwanted content from reaching all or some of the end users of a given online service and network. Content blocking is typically performed on a content proxy gateway, data network firewall, or other device inserted between the end user and the intended content source. Often, content filtering is implemented as part of the content cache engine, preventing only the desired content from being cached for the user population and the unwanted content being cached. All users can access network content only through the cache. Content is generally blocked if it is harmful or inappropriate for a user group or business use, or if it is viewed at a particular time of the day. Often, NSPs and businesses will rely on rating systems or services, such as the Platform for Internet Content Selection (PICS), to determine the suitability of content sites or documents for a particular population. End users may choose a set of blocking policies that they impose on some systems.
【0004】
NSPs and data transfers that lie between the need for an automated search engine that automatically indexes large amounts of content and the need for a blocking engine that blocks some of the content from reaching the end user in the end. An important issue emerges for providers. In particular, the problem is the lack of integration and coordination between engines that implement filtering and blocking policies and search engines. The lack of unity arises for several reasons. That is, (a) many organizations have adopted and implemented policies that filter and block content for search engines, such as those organizations' sites or services that rely on search engines available on the Internet. There is. (b) Search engines must deliberately search for and index as much content as possible, and tend to be willing to seek all content. Filtering and blocking engines, on the other hand, deliberately attempt to select documents that are cached and ultimately displayed to the end user.
【0005】
The inherently different roles between search and blocking engines, as well as the demand for execution efficiency and high performance, hinder the integration and coordination of these two information retrieval functions.
【0006】
The problem is that despite content documents that are ultimately inaccessible under filtering / blocking policies, the description and title of such content documents ends at the same time as using search engine services. It is made clear by the fact that it is displayed to the user. In addition to end-user inconvenience and frustration due to inconsistencies, the titles and short descriptions returned by search engines may themselves be very offensive or otherwise undesirable.
【0007】
Therefore, there is a need for an information retrieval system that can obtain search results that match or match the performance blocking method with a slight protocol and performance effect that can be executed.
【0008】
Content indexing search and blocking systems include:
【0009】
U.S. Pat. No. 5,701,469 (Brandli et al.), Issued December 23, 1997, provides a search results collection routine that removes erroneously included saved search results from the results and adds erroneously excluded saved search results. It discloses a content indexing search system to be executed. In this way, search results generated in response to user queries are accurately created even if the content index used to generate the initial search results is not up-to-date.
【0010】
U.S. Pat. No. 5,835,722 (Brandshaw et al.), Filed June 27, 1996 and issued November 10, 1998, comprehensively monitors computer operations for the production or transfer of search-inappropriate materials. By doing so, it discloses a terminal for blocking the use and transfer of improper material, thereby either terminal blocking or only by surveillance intervention.
【0011】
US Pat. No. 5,706,507 (Schloss), filed July 5, 1995 and issued January 6, 1998, is the content of data downloaded from a content server to block or detect unwanted material. Disclosures an advisory server operated by a third party that evaluates.
【0012】
U.S. Pat. No. 5,619,648 (Canale et al.), Issued April 8, 1997, provides e-mail filtering to determine whether e-mail messages should be provided to users based on a user interaction model. It is disclosed.
【0013】
Both prior arts have been implemented in a blocking engine where only the content allowed by the blocking policy is returned to the end user as a result of the content search and the search results match the blocking policy. It does not disclose a content indexing search system that provides search results that meet blocking policies.
【0014】
[Problems to be Solved by the Invention]
A first object of the present invention is to provide an improved information retrieval system and method for operations that provide consistency between search engine results and content blocking policies.
【0015】
A second object of the present invention is to provide an improved content indexing search system and method for operations that give search results consistent with blocking policies.
【0016】
A third object of the present invention is to provide an improved content indexing search system and method for implementing blocking policies in cache and filtering engines.
【0017】
A fourth object of the present invention is to provide an improved content indexing search system and method for operations that implement blocking policies during the content indexing phase.
【0018】
A fifth object of the present invention is to provide an improved content indexing search system and method for operations that implement blocking policies during the phase of displaying end-user search results.
【0019】
A sixth object of the present invention is an improved content indexing search system for operations that search the local repositories of cache and blocking engines instead of searching and index the final content site and content server. And to provide a method.
【0020】
A seventh object of the present invention is to provide an improved content indexing search system and method for operations configured to pass through a cache and filtering engine towards targeted content.
【0021】
[Means for solving problems]
These and other objectives, features and benefits implement a management policy that generally blocks unwanted content so that search results match the end-user's organizational filtering and blocking policies performed in another embodiment. To be achieved in an information retrieval network that includes a content indexing search engine with a cache engine and database that connects between the search engine and the end user.
【0022】
In the first embodiment, only the content allowed by the blocking policy is added to the search engine indexing database. In the second embodiment, the search engine search and display process is modified to enforce the blocking policy. In a third embodiment, the target of the search engine operation and indexing automaton process is modified to build an indexing database by searching the contents of the cache engine. In a fourth embodiment, the search engine scanning and indexing automaton is configured similar to an end-user browser, i.e., passing through a cache and filtering engine to reach the target content.
【0023】
BEST MODE FOR CARRYING OUT THE INVENTION
In FIG. 1, the information retrieval system 100 has a plurality of client devices 102, 104 connected to the Internet or other distributed data network 106 via an internal or controlled network 107. A typical client is a personal computer (PC) with a display 110, keyboard 111, CPU 112, memory 113 and network connectivity input / output device 115. Examples of such clients and networks include business users of PCs connected to the company's internal network and home users of PCs connected to the service provider's network, both of which are ultimately larger. Connect to the internet. Netscape Communicator, IBM Web Browser 116, sold under a trademark such as Explorer, is installed in memory 113 along with standard operating system 117 and application program 118. Browser 116 runs or runs on client devices 102, 104 to load or download content from content server 120 connected to internet 106. Each content server has a database 122 for storing data that can respond to content requests from clients 102, 104, and the like. In one form, the data is stored as a collection of HTML documents containing text and other multimedia content.
【0024】
The gateway 124 is commonly used as an interface between multiple clients or internal network segments 107 and the Internet 106, as shown in the figure. Typically, a proxy server with a cache and content filtering engine 126 is inserted into the connection path from the internal network 107 to the Internet 106 to implement content blocking policies to increase performance and management. The caching and blocking proxy server can connect to the gateway or can connect directly in parallel with the Internet 107 and the external network 106.
【0025】
Client systems 102 and 104 running the web browser 116 request content from the content server 120 using an HTTP (Hypertext Transfer Protocol) request and receive the content in an HTTP response. HTTP requests and responses occur on TCP / IP sockets that travel over the communication link between the client and the content server. The user generates a content request either by explicitly requesting the content stored on the content server or by taking a hyperlink anchor pointing to the content stored on the content server. Upon receipt, the browser loads its content using an HTTP session. For a detailed explanation of HTTP, see "Hypertext Transfer Protocol-HTTP / 1.0" by Berners-Lee et al. Draft IEFT-HTTP-V10-Spec-0.0 Text 1995 (March) As shown in 8), the entire contents of this document are incorporated herein by reference. A detailed description of HTML is given in Berners-Lee's "Hypertext Markup Language (HTML)" Draft IEFT.IIIR-HTML-01, June 1993 (out-of-print draft). Incorporated as a part of this specification. A detailed description of TCP / IP sockets can be found in W. Richard Stevens, "TCP / IP Illustrated, Vol.1--The Protocols", Addison-Westlake, 1994 pages 1-20, 229-262, in the book. In the specification, the entire contents of this document are incorporated as a part of the present specification.
【0026】
Users of client systems using Web browser 116 often access traditional search engine servers 130, 135 and databases 131, 136, respectively, limited to locating Internet content by keyword search means. .. These fixed engine servers will be external search engine servers 130 or internal search engine servers 135 for the managed network 107. While they perform the same basic functions, the internally coupled search engine server 135 is managed independently by the internal network operator and is rather preferred. The invention of this method may be realized by a search engine 135, which is generally managed internally, or by an external search engine 130, which consistently gives the organization a content blocking policy as a service to the organization. Let's go. As a result of a keyword search directed to search engine server 130 or 135, the end user sees a text excerpt displayed as a hyperlink anchor for the final content and a list of matching URLs in web browser 116. You will see it. The user can select and follow a link for one or more content matches using the Web browser 116.
【0027】
In Figure 2, a table 200 configured to filter / block content samples is generated by the client or by the network / service administrator to filter or otherwise limit the effectiveness of content that is considered inappropriate or unnecessary. Installed on proxy server 126 to do this. These content access control schemes generally block unwanted content from reaching all or part of the end user over a given online service or network. The table is installed in the cache and filtering engine 126 and is typically stored in database 127. In one form, the table is for each user or each user group a user or group identifier (ID) 203, a list of keywords for blocking 205, PICS (Platform For Interconnect Content). Selection) It has line 201 containing one or more of Rule 207, a blacklist 209 of URLs that should not be contacted, and a white list 211 of only URLs that can be contacted. The URL description is "Uniform Resource Locators (URL)" by Berners-Lee et al., RFC 1738, As set forth in December 1994, the entire contents of this document are incorporated herein by reference. PICS evaluations are derived from PICS rules that allow or block access to URLs based on the PIC label included in the document that describes the URL. The PICS rules are described in http://www.w3.org/TR/REC-PICSRules-971229 on the Internet issued by W3C etc. In particular, PICS rules are a language for representing filtering rules (profiles) that allow or block access to URLs based on the PICS labels that describe those URLs. The label is PICS available on the internet at http://www.w#.org/PICS/ Generated by using software tools based on Technical specification-1.1. Software tools are used to generate labels in documents that describe specific URLs. Alternatively, instead of labeling in the document, an independent reader distributes the label through another server called the Label Bureau. Filtering software will know to look in the label bureau to find the label, just as consumers know to read a particular magazine commenting on equipment or private cars. Once the label is generated, it is inserted as an additional header in the HTTP header stream that precedes the content of the document sent to the web browser. Alternatively, META tags can be used to embed labels in HTML documents. In this way, labels are sent only in HTML documents, not images, videos, or anything else. PICS-Compatible content server is International Business Machines Corporation Available from (Armonk, NY).
【0028】
With the blocking table installed in the cache and blocking engine, some process choices are available to join the content search with the content blocking engine and are ultimately allowed by the blocking policy. Only the content is returned to the client as a result of the content search. It is possible to have different rules for each individual user, while having a single set of rules that apply to all users, or to a small group of users, each with its own rules. Separating users is easier to handle. Individual users or groups can be identified by using several means, if defined. Such means include, for example, the client system IP address for user / group ID mapping, the use of HTTP basic authentication at the beginning of the browsing section, and the use of HTTP web cookies to track user identification. ..
【0029】
Returning to Figure 3, Process 300 implements the blocking policy during the content indexing phase. In step 302, the content manipulation and indexing automaton process by search engine 135 is modified. In step 304, content filtering rules from the content and filtering engine 126 are filtered through the application program interface (API) or rule definition file transfer at regular intervals or each time a change is detected. Loaded into server 135. In step 306, a large number of indexed database trees are generated as needed, and each tree is associated with a user group as defined in the content filtering rules. For example, one indexing database tree with strict PIC filtering rules for children is defined, and one indexing database tree with more free filtering rules for adults is defined.
【0030】
In step 308, the search engine automaton process initiates scanning and indexing of content from the list of target servers while the content blocking rules are being examined.
【0031】
In step 310, if a white list exists, the search engine will only search for websites or root content sources that match the clearly acceptable site / URL list or white list.
【0032】
In step 312, any website URL that matches the blacklist pattern is excluded if a blacklist of URLs to be excluded is set in the rule.
【0033】
In step 314, the PICS rules that apply to the user set serviced by the indexed database tree apply to the site / content / document being processed, and as a result the document is excluded or included. Is done.
【0034】
In step 316, if a list of keywords to exclude is identified, the document text is scanned to remove the document if one or more keywords are included in the list.
【0035】
In step 318, the document is added to the appropriate indexing database only if the filtering rules allow for that group.
【0036】
The advantage of the process in Figure 3 is that all additional (exclusion) processing is performed during the database indexing phase. There is very little additional processing required during the user search and presentation phases. Perhaps search operations are much more frequent than indexing operations in the search engine life cycle, even if the content is rescanned for possible changes.
【0037】
In Figure 4, another process 400 executes the blocking policy during the end-user search results display phase. In step 402, the search engine scanning and indexing automaton process remains unchanged and a single indexing database tree is maintained. In step 404, the search engine search and display process is modified to apply the blocking policy.
【0038】
In step 406, content filtering rules from the cache engine (via API or via rule definition transfer) are loaded into the search engine at regular intervals or each time a change is detected.
【0039】
In step 407, processing of the user-initiated search request is initiated against the index database.
【0040】
In step 408, a list of all matching documents that satisfy the user request is created and prepared for application of the blocking rule.
【0041】
In step 410, all matching documents not included in the white list are excluded if a white file with a clearly acceptable URL is specified in the rule.
【0042】
In step 412, any website, URL, or document that matches the exclusion pattern list (blacklist) is excluded.
【0043】
In step 414, if the PICS rule is specified, any URL that is not confirmed by the PICS rule is excluded.
【0044】
In step 416, if the keyword list is specified in the rule, any URL that contains one or more of the keywords in the list in the text is excluded.
【0045】
In step 418, the remaining subset of URL pointers that match the user's request and satisfy the blocking rules are returned for display to the client.
【0046】
The first advantage of the process in Figure 4 is that the latest policies can be applied to each search without the impact of rebuilding the indexed database. A single indexed database can be used for all users. This process allows the definition of changing filtering groups down to individual controls with little impact.
【0047】
In Figure 5, process 500 modifies the search engine to build its indexing database by searching for the content of the engine that caches the content. In step 501, the search engine scanning and indexing automaton process is modified. Instead of scanning and indexing the final content source site, the process is set up to search the local repository of cache and blocking engine content. In step 503, the search engine scan target is modified to a storage device that caches the appropriate content rather than the site / URL list. In step 505, the cache and blocking engine URL / content / document tree is traversed via API, database operations or shared field system operations. Step 507 follows a blocking filtering scheme for one or more user groups in the local installation so that any document found in the cache is added to the indexed database.
【0048】
The first advantage of the process in Figure 5 is that the application of filtering and blocking rules is done only once by the engine so designed, namely the cache and blocking engines. Scanning and indexing scanning is performed on a local (high performance) copy of the target content rather than on a more variable Internet content site.
【0049】
In Figure 6, the search engine builds its own indexing database, so process 600 modifies the search engine to pass through the cache and filtering engine. In step 601, the search engine scanning and indexing automaton is modified to build like an end-user browser. That is, it passes through a cache and filtering engine using an HTTP proxy to reach the target content. In step 603, the search engine automaton is configured to use an HTTP proxy configured for the appropriate cache and filtering engine. In step 605, while scanning and indexing the content, the search engine automaton is one of the user groups so that the user receives only a subset containing the sites / contents / documents allowed by the user group's policies. It is configured to simulate the end users who belong to one.
【0050】
The first advantage of the process in Figure 6 is that it does not substantially modify the search engine. Content blocking and filtering is performed by a cache and blocking engine so designed and optimized. Only content allowed by the blocking policy reaches the search engine for indexing. Search engine efficiency and performance will increase as some of the sites / content that should be scanned and indexed will be found in the local cache storage.
【0051】
As described above, the content search and content blocking engines combine so that only pointers to the content that is ultimately allowed by the blocking policy are returned to the end user as a result of the content search. Numerous processes are described to combine the content search and content blocking engines. As such, the present invention provides consistency between end-user content search results and individual organization content filtering and blocking policies. The present invention can be immediately utilized with respect to existing Internet and other networks, namely data protocols and standards, without the need for modification.
【0052】
Although the present invention has described the Internet (HTTP / Web) environment, similar concepts apply to most data and network environments in which data is retrieved. A list of possible matches is displayed to the end user who consumes / lists the data in sequence, if allowed by the access or content management scheme. Various modifications can be made without departing from the spirit and scope of the invention as defined in the claims.
【0053】
In summary, the following matters will be disclosed with respect to the constitution of the present invention. (1) A content indexing search system that provides search results that match the content filtering and blocking limits, including a content indexing search engine that includes a database, a cache and blocking proxy server that includes a cache, and the content. Information networks coupled to indexed search engines, means for making search queries to the content indexed search engine and receiving search results from the cache, and content filtering and blocking combined with the content indexed search engine. A content indexing search system comprising a blocking engine that executes the policy and a means for modifying the content indexing search engine to execute the same content blocking policy as the blocking engine. (2) The system according to (1) above, further comprising means for executing the blocking policy in the content indexing search engine during the content indexing phase. (3) The system according to (1) above, further comprising means for executing the blocking policy during the end-user search result display phase. (4) The system according to (1) above, further comprising means of modifying the content indexing search engine to build an indexing database by indexing cached content. (5) Described in (1) above, further comprising means for modifying the content indexing search engine to capture the cache and blocking engine results when the content indexing search engine builds the indexing database. System. (6) Content indexing search engine combined with database and cache, information network combined with the content indexing search engine, and content filtering and content filtering for search results provided to end users via the cache. In a content indexing search system having a blocking engine that implements blocking restrictions, a method of providing search results that match the execution of content filtering and blocking, (a) the process of the content indexing search engine. And (b) change the process of the content indexing search engine to skip any information site URLs that match the exclusion pattern, and sites or routes that match the explicitly allowed information site URL list. It has a step of searching only the content source and a step of (c) executing the filtering policy defined by the cache and the blocking engine in the content indexing search engine, and the filtering in the step (C). The policy is to (i) read the content filtering rules from the cache and filtering engine at regular intervals or each time a change is detected, and (ii) generate a multi-indexed database tree, each tree. To associate with the user group defined in the content filtering rules, and (iii) avoid displaying any information site, URL, or document that matches the exclusion pattern to users, and (iv) Steps to display document / content pointers from sources that match a clearly tolerated list of information site URLs, and (v) information networks / content / documents according to the filtering process defined by the cache and the blocking engine. It is defined by a step of displaying only to the user, and in the step (v) above, the filtering process is (aa) at regular intervals.In, or each time a change is detected, the step of reading the filtering rules for the content from the cache and filtering engine, and (bb) the search allowed by the filtering rules for individual or groups for the user. The method defined by the step of displaying only the results. (7) (d) Content engine scanning The cache and blocking engine by modifying the target to the content cache storage rather than the information site / URL list and (e) API, database scanning, and shared file scanning. The method described in (6) above, further comprising a step of traversing the URL / content / document tree of. (8) (f) The method according to (6) above, further comprising modifying the search engine content scanning and indexing process to be configured in the same manner as the end user's browser. (9) (g) The method according to (6) above, further comprising modifying the content indexing search engine to pass through the cache when constructing the indexing database. (10) (h) The method according to (6) above, further comprising modifying the content indexing search engine to build an indexed database by searching the cache. The method according to (6) above, further comprising (11) (i) connecting the content indexing search engine to an internal network and (j) connecting the content indexing search engine to an internal network operation. .. (12) The method according to (6) above, further comprising the step of connecting the content indexing search engine to an external network to impart consistency to the organizational content blocking policy.
[Simple explanation of drawings]
[Figure 1]
It is a block diagram for demonstrating the structure of the information retrieval system based on this invention.
[Figure 2]
It is a table of content blocking rules executed in the information retrieval system shown in Fig. 1.
[Fig. 3]
It is a flowchart for demonstrating the operation of the system of FIG. 1 in the 1st Embodiment which executes a blocking policy during a content search and indexing phase.
[Fig. 4]
It is a flowchart for demonstrating the operation of the system of FIG. 1 in the 2nd Embodiment which executes a blocking policy during the phase of displaying the search result of an end user.
[Fig. 5]
It is a flowchart for demonstrating the search engine of FIG. 1 in the 3rd Embodiment example.
[Fig. 6]
It is a flowchart for demonstrating the search engine of FIG. 1 in 4th Embodiment Example.
[Explanation of symbols]
100 Information retrieval system 102 Client device 104 Client device 106 internet 107 Network 110 display 111 keyboard 112 CPU 113 memory 115 Network connectivity I / O devices 116 browser 117 Operating system 118 Application program 120 content server 122 database 124 gateway 126 Cache and content filtering engine 130 search engine server 131 database 135 search engine server 136 database 200 tables 201 lines 203 Identifier 204 User or Group ID 205 List of keywords for blocking 207 PICS rules 209 black list 211 White list 300 processes 400 processes
Every citation, both ways
| Document | Relation | Office | Cited during |
|---|---|---|---|
| US10229453B2 | Cited by | United States of America | Applicant |
| US10346181B2 | Cited by | United States of America | Applicant |
| US7047229B2 | Cited by | United States of America | Applicant |
| US7711956B2 | Cited by | United States of America | Applicant |
| US10367860B2 | Cited by | United States of America | Applicant |
| US11275594B2 | Cited by | United States of America | Applicant |
| JP2002334034A | Cited by | Japan | Search report |
| US9116966B2 | Cited by | United States of America | Applicant |
| US7558805B2 | Cited by | United States of America | Applicant |
| US10580518B2 | Cited by | United States of America | Applicant |
| US7359951B2 | Cited by | United States of America | Applicant |
| US10719334B2 | Cited by | United States of America | Applicant |
| JP2009528786A | Cited by | Japan | Examiner |
| US9898312B2 | Cited by | United States of America | Applicant |
| US9621501B2 | Cited by | United States of America | Applicant |
| US11416778B2 | Cited by | United States of America | Applicant |
| US9122731B2 | Cited by | United States of America | Applicant |
| US10957423B2 | Cited by | United States of America | Applicant |
| JP2001282797A | Cited by | Japan | Search report |
| JP2010049650A | Cited by | Japan | Examiner |
| JP2006048193A | Cited by | Japan | Search report |
| JP2007011777A | Cited by | Japan | Examiner |
| US7984061B1 | Cited by | United States of America | Applicant |
| US10909623B2 | Cited by | United States of America | Applicant |
| US7162526B2 | Cited by | United States of America | Applicant |
| US10572824B2 | Cited by | United States of America | Applicant |
| JP2008204427A | Cited by | Japan | Examiner |
| JP2003283549A | Cited by | Japan | Search report |
| US9128992B2 | Cited by | United States of America | Applicant |
| US10846624B2 | Cited by | United States of America | Applicant |
| JP2012198856A | Cited by | Japan | Search report |
| JP2004151905A | Cited by | Japan | Examiner |
| JP2008197748A | Cited by | Japan | Search report |
| US9621501B2 | Cited by | United States of America | Applicant |
| US7007008B2 | Cited by | United States of America | Applicant |
| US7225180B2 | Cited by | United States of America | Applicant |
| JP2010508592A | Cited by | Japan | Examiner |
| JP2005018751A | Cited by | Japan | Examiner |
| US8346953B1 | Cited by | United States of America | Applicant |
| US10929152B2 | Cited by | United States of America | Applicant |
| US10367860B2 | Cited by | United States of America | Applicant |
7 members in 5 offices
Priority claims5
| Document | Office | Kind | Date |
|---|---|---|---|
| 09302851 | United States of America | – | |
| 30285199 | United States of America | A | |
| 30285199 | United States of America | A | |
| 302851 | – | – | – |
| US19990302851 | – | – | – |
Members7
| Document | Office | Kind | |
|---|---|---|---|
| CA2300239A1 | Canada | A1 | |
| CN1272656A | China | A | |
| JP2000357176AThis record | Japan | A | |
| US6336117B1 | United States of America | B1 | |
| SG96549A1 | Singapore | A1 | |
| CN1253813C | China | C | |
| CA2300239C | Canada | C |
Numbers
- Publication
- 2000-357176
- Publication, DOCDB
- 2000357176
- Publication, EPODOC
- JP2000357176
- Application
- 121247
- Application, DOCDB
- 2000121247
- Application, EPODOC
- JP20000121247
Titles2
- Japanese
- 【発明の名称】コンテンツ索引付け検索システム及び検索結果提供方法
- English
- INDUSTRIAL APPLICABILITY: Content indexing search system and search result providing method
Classification
- CPC, 3
- G06F16/9535
- Y10S707/959
- Y10S707/99945
- IPC, 3
- G06F12 00
- G06F13 00
- G06F17 30