Method and device for composing document
Abstract
[Task] Provided are a document synthesizing method and a document synthesizing device capable of easily and versatilely synthesizing information of a plurality of websites on one web document.
Solution.At a minimum, the location of the first document on the Internet in the markup language on the WWW on the Internet, the range of subdocuments extracted from the first document, and the subdocument on the second document for compositing. Describes the insertion position of, the range in which the document structure on the second document including the partial document inserted in the insertion position should be converted, and the conversion rule for converting the document structure into a desired document structure. The sub-document is extracted from the first document according to the second document in which the identification information of the file is described in the markup language, and the sub-document is placed in the designated composite position on the second document. The conversion rule is used to convert the document structure in the specified range on the second document.

Term
Term ended
Projected expiry passed 18 December 2020, 5.8 years ago.
- Priority and filed
- Published
- Projected expiry
- Today
8 claims: 4 independent, 4 dependent
- 1【特許請求の範囲】 【請求項1】 インターネットにおけるWWW(World Wide web)上のマークアップ言語で記述された複数の第1の文書の内容の一部をWWW上のマークアップ言語で記述された第2の文書に合成するための文書合成方法であって、 少なくとも、前記第1の文書の該インターネット上の所在と、該第1の文書から抽出する部分文書の範囲と、前記第2の文書上の前記部分文書の挿入位置と、前記挿入位置に挿入される前記部分文書を含む前記第2の文書上の文書構造を変換すべき範囲と、前記文書構造を所望の文書構造に変換するための変換ルールを記述したファイルの識別情報とをマークアップ言語により記述した第2の文書に従って、 前記第1の文書から前記部分文書を抽出して、その部分文書を前記第2の文書上の前記指定された合成位置に挿入するとともに、前記変換ルールを用いて前記第2の文書上の前記指定された範囲の文書構造を変換することで、前記第2の文書上に1または複数の前記部分文書を合成することを特徴とする文書合成方法。
- 2【請求項2】 前記第2の文書は、少なくとも、前記第2の文書上の前記部分文書の挿入位置とを指定するとともに、前記第1の文書の所在と、該第1の文書から抽出する部分文書の範囲とを記述するため第1のタグと、 前記変換ルールを用いて文書構造を変換すべき範囲を指定するとともに、前記変換ルールを記述したファイルの識別情報を記述するための第2のタグと、 を用いて記述されていることを特徴とする請求項1記載の文書合成方法。
- 3【請求項3】 前記第2の文書は、XML(Extensible Markup Language)で記述され、前記第1の文書がXMLで記述されていないときは、まず、XMLによる記述型式に変換した後、前記第1の文書から前記部分文書を抽出して、その部分文書を前記第2の文書上の前記指定された挿入位置に挿入することを特徴とする請求項1記載の文書合成方法。
- 4【請求項4】 インターネットにおけるWWW(World Wide web)上のマークアップ言語で記述された複数の第1の文書の内容の一部をWWW上のマークアップ言語で記述された第2の文書に合成する文書合成装置であって、 少なくとも、前記第1の文書の該インターネット上の所在と、該第1の文書から抽出する部分文書の範囲と、前記第2の文書上の前記部分文書の挿入位置と、前記挿入位置に挿入される前記部分文書を含む前記第2の文書上の文書構造を変換すべき範囲と、前記文書構造を所望の文書構造に変換するための変換ルールを記述したファイルの識別情報とをマークアップ言語により記述した第2の文書に従って、前記第1の文書から前記部分文書を抽出して、その部分文書を前記第2の文書上の前記指定された挿入位置に挿入する挿入手段と、 前記第2の文書に従って、該第2の文書上の前記指定された範囲の文書構造を、前記変換ルールを用いて所望の文書構造に変換する変換手段と、 を具備し、 前記第2の文書上に1または複数の前記部分文書を合成することを特徴とする文書合成装置。
- 5【請求項5】 前記第2の文書は、少なくとも、前記第2の文書上の前記部分文書の挿入位置とを指定するとともに、前記第1の文書の所在と、該第1の文書から抽出する部分文書の範囲とを記述するため第1のタグと、 前記変換ルールを用いて文書構造を変換すべき範囲を指定するとともに、前記変換ルールを記述したファイルの識別情報を記述するための第2のタグと、 を用いて記述されていることを特徴とする請求項4記載の文書合成装置。
- 6【請求項6】 前記第2の文書は、XML(Extensible Markup Language)で記述されていることを特徴とする請求項4記載の文書合成装置。
- 7【請求項7】 前記第1の文書がXMLで記述されていないとき、該第1の文書をXMLによる記述型式に変換する第2の変換手段をさらに具備し、 前記挿入手段は、XML文書の前記第1の文書から前記部分文書を抽出して、その部分文書を前記第2の文書上の前記指定された挿入位置に挿入することを特徴とする請求項4記載の文書合成装置。
- 8【請求項8】 インターネットにおけるWWW(World Wide web)上のマークアップ言語で記述された複数の第1の文書の内容の一部をマークアップ言語で記述された第2の文書に合成するための処理をコンピュータに実行させるためのプログラムであって、 少なくとも、前記第1の文書の該インターネット上の所在と、該第1の文書から抽出する部分文書の範囲と、前記第2の文書上の前記部分文書の挿入位置と、前記挿入位置に挿入される前記部分文書を含む前記第2の文書上の文書構造を変換すべき範囲と、前記文書構造を所望の文書構造に変換するための変換ルールを記述したファイルの識別情報とをマークアップ言語により記述した第2の文書に従って、前記第1の文書から前記部分文書を抽出して、その部分文書を前記第2の文書上の前記指定された挿入位置に挿入するための処理と、 前記第2の文書に基づき、該第2の文書上の前記指定された範囲の文書構造を、前記変換ルールを用いて所望の文書構造に変換するための処理と、 をコンピュータに実行させるためのプログラム。
Independent claims8
369 paragraphs in 1 section, as filed
Description: TECHNICAL FIELD [Detailed description of the invention]
【0001】
[Technical field to which the invention belongs]
The present invention relates to a web document synthesizing method for synthesizing a plurality of web documents on one web document, and a web document synthesizing device using the same.
【0002】
[Conventional technology]
The WWW (World Wide Web) has become widespread as an information infrastructure that enables the construction and publication of effective presentations at low cost, and a huge amount of information resources are published on sites around the world. The WWW also has an infrastructure aspect for server-client systems. In particular, it is expected to be applied to electronic commerce and ASP (Application Service Providing) recently, and the number of full-scale commerce sites is increasing rapidly. In electronic commerce, the web page acts as an operating panel that connects the user with the back-end system of the corporate LAN that processes the commerce. The WWW is the only infrastructure that connects computer systems around the world across sites, but it is expected that the trend toward web tops will continue in the future.
【0003】
The information resources exchanged on the WWW will continue to increase, and the processing required for web systems will become more complex and diverse.
【0004】
In particular, companies are actively using the WWW and publish a large amount of their own data such as corporate data, news and product catalog information through web pages, but it is too much to create each web page from scratch. Since it takes too much manpower, we have introduced a technology to statically or dynamically generate a web page containing standard content from a database to improve the efficiency of site construction and operation. The tools for building and operating such websites are provided by many software vendors and are extremely substantial. However, all of these technologies are related to the efficiency and performance of the construction and operation of a single closed website.
【0005】
Now that the construction and operation environment for a single website has been established, the next requirement for the WWW is inter-website collaboration. That is, it is a development from a server-client system to a distributed system. Especially in the era of full-scale e-commerce, cooperation of e-commerce systems of each commerce site is indispensable.
【0006】
Cooperation of electronic commerce systems requires many agreements such as standardization of data formats such as product profiles and vocabulary, common business models, and common message formats and protocols according to them. On the other hand, industry groups such as OASIS and BizTalk are promoting standardization, but there are many barriers such as disagreement of interests between companies and differences in business customs, so it will take time for the results to bear fruit. There is no doubt that.
【0007】
On the other hand, in order to meet the urgent needs, each software vendor provides a package that adds a website cooperation mechanism to the above-mentioned website construction / operation tool.
【0008】
However, the conventional system construction method centered on the database-centered application logic group worked effectively by positioning the web page as a mere user interface for a single website, but multiple websites. It cannot be applied as it is to a system that spans. This is because, although in this building method it is necessary to connect the application logic to implement the system coordination, inter-site are blocked by a firewall, most of the field is because if non-HTTP message can not be replaced.
【0009】
Therefore, a system integration model based on HTTP, which is the only message exchange channel, is required, but many of the packages only add HTTP access function to the conventional site construction technology, and can make full use of HTTP and WWW functions. It is in a situation where it is not.
【0010】
In this way, system cooperation between sites is an essentially difficult task because many agreements are required to connect the logic of each system.
【0011】
Therefore, focusing on the issue of inter-website collaboration using content exchange instead of logic connection, inter-website content linkage requires only adjustment of the degree of structural conversion of web resources, so compared to inter-website system linkage. There are few issues to be solved.
【0012】
However, on the other hand, the effect of content linkage is large enough. As mentioned earlier, the WWW has already published a huge amount of web resources. Web resources are also multimedia and can include any content media. An environment in which such web resources can be easily reused with each other under agreement between sites would make the WWW much more rational and economical, and would make great strides in the application of the WWW.
【0013】
For example, a decentralized management type website construction style such as outsourcing some of the information resources that make up the website, such as book sales information and TV program audience rating information, will be possible, and a large web parts market may be created. There is also. In addition, recently, portal sites that provide intermediary services such as shopping malls that compare and display the product catalogs of each shopping site on one web page and marketplaces that integrate the projects of multiple procurement systems and auction systems have been introduced one after another. It has appeared and is receiving a lot of attention. This is because there is an inevitable need for services that organize web information and act as guides in a situation where web information is extremely flooded, and this is one form of responding to that demand. Creating an environment for reusing web resources with each other will greatly contribute to the construction of such a portal site. From that point of view, it can be said that it is positioned as a steady technology transition that will serve as a foothold for linking systems between websites such as e-commerce systems.
【0014】
By the way, portal sites that provide intermediary services that collect information on multiple websites, such as web page search services and various product comparison services, have appeared one after another and are attracting a great deal of attention. Furthermore, it is showing development in the specialization and diversification of functions such as image collection and MP3 collection. The essence of the task is content coordination between websites that collects and processes distributed web resources and provides them as web pages.
【0015】
In HTML technology, it is possible to jump to an arbitrary web page by using the hyperlink mechanism, and it is possible to display the entire multiple web pages as independent windows by using the frame mechanism, but the product comparison function and total It is completely insufficient to link organic contents such as providing a price estimation function. In order to realize these, it is necessary to have a function to collect arbitrary web pages and process them flexibly. Due to this lack of functionality in HTML, it is possible to have an external program executed by a program startup mechanism such as CGI (Common Gateway Interface) or Servlet, or a daemon program independent of the web server perform these processing processes. Has been done. This processing generally requires the following execution procedure. If a database is used, data registration and retrieval processing to the database will be added.
【0016】
1. The process of getting the HTML page of an external website 2. The process of extracting the required text from the HTML page 3. The process of converting the extracted text into the desired format 4. The process of stitching text together to create a single HTML Such a solution has drawbacks. In other words, although many of these processes are similar in content between intermediary services, it is inferior in production efficiency and maintainability that each site builder creates a program from scratch. In addition, the created program depends on the environment of the site, and inevitably becomes a program asset dedicated to the site, so that it cannot be reused in another site environment.
【0017】
Such a drawback is due to the fact that there is no tool or system for targeting content linkage in WWW technology and easily realizing it.
【0018】
[Problems to be Solved by the Invention]
In this way, conventionally, a general-purpose method for collecting necessary information from multiple web pages, converting it into a specific format, and then synthesizing it on one web page. There was a problem that there was no.
【0019】
In the future, in a situation where intermediary services such as portal sites that collect information from multiple websites will become more active, providing a common platform specializing in content collaboration will be effective in terms of production efficiency and portability. It is one of the means.
【0020】
Therefore, in view of the above problems, the present invention provides a document synthesizing method capable of synthesizing information of a plurality of websites on one web document easily and universally, and a document synthesizing apparatus using the same. The purpose is.
【0021】
[Means for solving problems]
The present invention synthesizes a part of the contents of a plurality of first documents written in a markup language on WWW (World Wide web) on the Internet into a second document written in a markup language on WWW. The location of the first document on the Internet, the range of the partial document extracted from the first document, the insertion position of the partial document on the second document, and the above. The range to be converted in the document structure on the second document including the partial document inserted at the insertion position, and the identification information of the file describing the conversion rule for converting the document structure into the desired document structure. Is extracted from the first document according to the second document described in the markup language, the subdocument is inserted at the designated insertion position on the second document, and the above It is characterized in that the document structure in the specified range on the second document is converted by using a conversion rule.
【0022】
According to the present invention, information of a plurality of websites can be easily and universally combined on one web document.
【0023】
Preferably, the second document specifies the insertion position of the partial document on the second document, the location of the first document, and the range of the partial documents extracted from the first document. In order to describe, the first tag (insert instruction tag pz: targets) and the range in which the document structure should be converted using the conversion rule are specified, and the identification information of the file in which the conversion rule is described is described. It is described using the second tag (conversion instruction tag pz: convert) for.
【0024】
Further, preferably, the second document is described in XML (Extensible Markup Language).
【0025】
Further, preferably, when the first document is not described in XML, first, after converting to the description type in XML, the partial document is extracted from the first document, and the partial document is described as described above. Insert at the specified insertion position on the second document.
【0026】
When the above method was incorporated into a web server on the Internet and a request for the second document was received from a client device (web browser), one or more partial documents were synthesized according to the description in this second document. A server device that provides a second document to the requesting web browser can be configured.
【0027】
BEST MODE FOR CARRYING OUT THE INVENTION
Hereinafter, embodiments of the present invention will be described with reference to the drawings.
【0028】
The following description is given in the order of the following items.
【0029】
(A) Functions required to combine information from multiple websites into a single web document (B) XML-P'z document (B-1) XML-P'z language specifications (B-2) Configuration and operation of XML-P'z language processing system (C) A series of operations for synthesizing multiple web documents on one web document (D) Coordination between XML-P'z servers for compositing web documents (E) Addendum (A) Functions required to combine information from multiple websites into a single web document First, before explaining the embodiment, the functions required for synthesizing the information (web document) of a plurality of websites into one web document will be described.
【0030】
The functions required to combine multiple web documents on one web document are narrowed down to three types: extraction, insertion, and conversion. However, since not all website information, that is, web documents as content (for example, HTML documents) is required, and only a part of them is generally required, the extraction function is required. It is required to capture a partial document of any web document. Further, when synthesizing a plurality of extracted partial documents in combination, a flexible insertion function such as inserting a table in a table is required. Furthermore, that alone is not enough, and when synthesizing the extracted partial documents into a list format, if the formats are non-uniform, a document conversion function is required, such as matching them to the same format. There is also.
【0031】
Based on this analysis, the present invention adopts the following descriptive model. First, in the same way as SSI (Server Side Inclusion) and its developments such as ASP (Active Server Pages) and JSP (Java Server Pages), in a web document for compositing multiple web documents (partial documents). A patchwork-like document processing method is adopted in which a command is placed at an arbitrary position and the command execution result is embedded at that position.
【0032】
Then, as a command to be prepared, a command for inserting a partial document indicating which part of which web page is to be extracted and where to insert is prepared. This method has the advantage that the specified partial document to be extracted and its insertion position can be freely and intuitively described using a synthetic web document as a skeleton. In addition, a conversion command that can perform conversion processing on an arbitrary range of the synthetic web document that serves as the skeleton is prepared. This conversion command takes range information and conversion rules as input and outputs the conversion result document. In summary, we adopted a description format that allows the synthesis logic to be embedded at any position in the synthesis web document, and prepared insertion and conversion as commands for the synthesis logic.
【0033】
Also, one of the execution models adopted is similar to SSI, and this synthetic web document is placed on the web server, and when the browser requests the URL, it is placed on the web server. The language processing system interprets and executes the commands contained in the synthetic web document, and returns the result to the browser. This method has the advantage that the site builder does not have to be aware of the activation of the interpretation execution only by placing the synthetic web document on the web server. However, in principle, it is possible for the user to manually perform the interpretation execution as well as such an execution method. In this case, arbitrary composition can be performed on the client side.
【0034】
By the way, XML (Extensible Markup Language) is the most suitable language for describing such synthetic web documents. In XML, tag names and attribute names can be freely defined, and the application side can give semantics to them. In addition to that, XML is also guaranteed to have a tree-type document structure, so it is only necessary to point to a specific element represented as one node in the document structure represented by the tree structure (document). Range) can be specified.
【0035】
In addition, due to the demand for XML itself as a standard data format at the low level, transformation technology such as XSLT (Extensible Stylesheet Language Transformations) (reference: http://www.w3.org/TR/xslt) is also available. In the future development of XML technology, the above-mentioned web document for synthesis can be described in a language to which this XML language is applied (XML application language according to the present invention) for convenience such as expandability and tool usage. Sex will be promised.
【0036】
In addition, there is an advantage that it is easy to handle as an extraction target when not only HTML documents but also XML documents are often used in the future.
【0037】
Therefore, in the present invention, the description language of the web document for synthesis is specifically designed as an XML application language.
【0038】
In the present invention, a synthetic web document (sometimes called a synthetic web page) that is a base for combining is described in XML, and a specified range part (partial document) is extracted from other specified web documents. Then, insert it at the specified position of the composite web document, and perform conversion processing (conversion process to the desired document structure) to the specified range of the composite web document, two synthesis logics of insertion and conversion. The policy is to have the instruction as an element in the synthetic web document.
【0039】
Such a synthetic web document, that is, an XML document (XML page) is referred to here as an XML-P'z (XML-Pieces) document (XML-P'z page).
【0040】
By incorporating the XML-P'z language processing system into the web server, the operation shown in Fig. 1 becomes possible. A web server incorporating an XML-P'z language processing system is sometimes called an XML-P'z server. Specifically, the case of incorporating into IIS (Internet Information Server), which is Microsoft's web server, will be described as an example.
【0041】
In the basic operating principle shown in Fig. 1, (Step S101) The request (GET / HTTP) of XML-P'z document 2 is transmitted from the web browser of the client terminal B1 to the XML-P'z server A1 (hereinafter, simply referred to as server A1).
【0042】
(Step S102) Server A1 determines if the requested resource is an XML-P'z document.
【0043】
(Step S103) If it is determined that the document is an XML-P'z document, the server A1 starts the XML-P'z language processing system (synthesis processing unit 1 in FIG. 1) and describes it in the XML-P'z document 2. The part (partial document) of the specified range (partial document) is extracted from the web document (page) W2, W3 of the specified web server (for example, web servers A2 and A3 in this case), and it is XML-P'z. It is inserted at the specified position of the document and converted to the specified range described in the XML-P'z document. Finally, the XML document (synthetic web document) W1 as the processing result of the XML-P'z language processing system is obtained.
【0044】
(Step S104) Send the obtained XML document to the browser as a reply to the requester.
【0045】
The above operation is realized by setting the web server. Most web servers have the ability to map URL string patterns (often object extensions) to the add-ins needed to preprocess them (step S102). ) ~ (Step S103) can be realized.
【0046】
Further, if the web browser can display the XML document, the XML document may be displayed, and if the web browser cannot display the XML document, the style sheet may be processed on the server A1 side and the HTML document may be returned.
【0047】
(B) XML-P'z document In the XML-P'z document, the insert instruction element "pz: targets" and the conversion instruction element "pz: convert" are defined.
【0048】
Inserting (synthesizing) a partial document of another XML document or HTML document as a child document under one element on the document structure represented by the tree structure of the XML-P'z document by using the insert command tag. Can be done. The URL with XPointer (reference: http://www.w3.org/TR/WD-xptr#uri-escaping) is used to specify the partial document to be inserted. This makes it possible to concisely specify a partial document of a specific web page in one line. However, since the XPointer standard is for XML, it cannot directly target HTML. For this reason, we will introduce a mechanism for performing structurally equivalent HTML-XML transformation by using HTML-DOM (Document Object Model) and XML-DOM when extracting. As a result, the HTML document can be treated as an XML document, so that all processing can be performed as XML.
【0049】
Moreover, in the XML-P'z document, by using the conversion instruction element, it is possible to execute the conversion operation using XSLT (Extensible Style Language transformations) for each child document under an arbitrary element (node). That is, the specified XSLT is applied to each child document specified by the conversion instruction element and placed as a child node of the conversion instruction element. By utilizing this, the web document inserted by the insertion command tag can be converted by using the conversion command tag.
【0050】
The following is a simple example of an XML-P'z document with insert and transform functions using insert and transform command elements.
【0051】
(First example of XML-P'z document) 1. <? Xml version = 1.0?> 2. <root xmlns: pz = http://www.shiba.co.jp/xmlpz> 3. <category> xxx </ category> 4. <item_holder> 5. <pz: convert href = xxx.xsl> 6. <pz: targets href = http://www.yyy.com/index.xml#xpointer (// item) /> 7. </ pz: convert> 8. </ item_holder> 9. </ root> FIG. 11 (a) schematically shows the document structure of the first example, and FIG. 11 (b) schematically shows the document structure of the XML document after interpreting the first example. It is shown.
【0052】
In the first example above, each XML partial document (http://www.yyy.com/index.xml#xpointer (// item) to be inserted specified by the insert instruction element "pz: targets" on the 6th line. ), Hereinafter referred to simply as the partial document PD1) is converted by applying the XSLT conversion rule specified by the conversion instruction element "pz: convert" on the 5th line, and the 4th to 8th lines. It is inserted as a child element of a certain "item_holder" element, as shown in Figure 11 (b). However, the web document specified by "pz: targets" on the 6th line is all the partial documents that match XPointer (in the case of the first example above, all the partial documents whose root is the "item" tag. ), Generally multiple web documents.
【0053】
The above-mentioned web document synthesis method for distributed web resources has the following advantages.
【0054】
One of the advantages is ease of construction. Unlike the conventional method centered on a database, this method can concisely describe the synthesis logic of information resources without a programming language, so it is easy to construct and change the configuration of web document integration. In addition, since an interpreter-type execution model that is interpreted and processed when requested by the browser is adopted, changes in the synthesis logic are reflected immediately.
【0055】
Another advantage is high reusability. In the XML-P'z framework, all components such as content, conversion rules, and synthesis logic are provided as web resources. Unlike the conventional method of having synthetic logic as a program outside the web document, this method allows access to all these components via a URL, so in principle it can be re-installed from web systems around the world. It can be used. This means that each resource required for a distributed system beyond the website can be freely allocated, and it is possible to flexibly construct and change the system according to the operation.
【0056】
Furthermore, if the XML-P'z document targets the XML-P'z document of another site as the synthesis target, the synthesis logic can be divided (coordinated) between the websites.
【0057】
In addition, it does not use any special protocol other than HTTP, and the website that provides the web resource does not need to introduce a special processing system. Therefore, the information resources of any website can be reused. In other words, existing websites can utilize system resources as they are, and can be synthesized simply by creating XML-P'z resources separately.
【0058】
However, such high accessibility involves practical problems related to usage such as copyright issues. For example, XML-P'z technology makes it easy to provide a meta-search page that synthesizes search results from multiple websites that offer web search services, but it violates copyright issues. Such a problem has become a problem even in the current WWW regarding the permission of hyperlinks, and the current situation is that it has been overcome in operation. On the other hand, while WWW technology related to access control such as Extranet construction technology is being provided, legislation regarding the handling of copyrighted works published on WWW is being developed at a rapid pace. Also, in the XML-P'z framework, we would like to introduce a model that comprehensively handles copyright issues as a future issue.
【0059】
Next, the web document synthesis method of the distributed web resource described above will be described in the following two parts.
【0060】
(B-1) XML-P'z language specifications (B-2) Configuration and operation of XML-P'z language processing system The XML-P'z language is a web page description language that includes synthetic logic and forms the core of this system. First, the language specifications will be described in (B-1). Next, the configuration and operation of a language processing system as a language engine that interprets and processes an XML-P'z document written in the XML-P'z language and returns the result will be described in (B-2).
【0061】
(B-1) XML-P'z language specifications The XML-P'z language is one of the XML application languages in which semantics are given to specific tag names, and is a web document description language for the purpose of synthesizing distributed web resources. In addition to being able to describe content as in a normal XML document, synthetic logic can be included internally by describing the tag name for the instruction that operates the web resource for any element. .. The description of this synthetic logic is as simple as an HTML hyperlink.
【0062】
An XML-P'z document written in the XML-P'z language that includes synthetic logic in this way is interpreted as a web document that virtually integrates and synthesizes distributed resources according to the synthetic logic.
【0063】
Two instruction elements related to web resource operations, "targets" and "convert", are prepared, and "pz" is reserved as the XML namespace. By using these instruction elements in combination, it is possible to extract arbitrary partial documents including other web documents, insert own documents, and perform structural conversion using XSLT. Each instruction element (pz: convert element, pz: targets element) will be described below.
【0064】
Also, these instruction elements must be interpreted in a depth-first search order. For example, in the document structure of the XML-P'z document shown in Fig. 12, if there are multiple pz: targets elements as child elements of the pz: convert element, after each pz: targets element is interpreted in order from brother to brother. , The pz: convert element is interpreted.
【0065】
Also, as explained in the section of each instruction tag, the web document inserted by the insert instruction element and the web document converted by the conversion instruction element must be interpreted as an XML-P'z document before being synthesized and converted. Must be. That is, if an instruction element (insertion, conversion instruction element) is included in the web document to be inserted or converted by the instruction element, they are preferentially interpreted in the above-mentioned order, and then the insertion destination, this XML- The flow of recursive interpretation processing is such that the interpretation execution of the P'z document is continued.
【0066】
It also introduces a URL with XPointer as a web resource specifier. This conforms to the XPointer standard (reference: http://www.w3.org/TR/WD-xptr), but since this standard does not define the relative specification of URLs with XPointer, XML- The P'z language has its own standards.
【0067】
The standard is shown below.
【0068】
(XML namespace) In order to use each instruction tag of XML-P'z, the following namespace must be declared.
【0069】
Namespace name pz -Namespace URI http://shiba.co.jp/xmlpz (pz: targets element) Extract / insert arbitrary web resources grammar <pz: targetshref = web-resources-url> </ pz: targets> ·attribute href URLs to multiple web resources to insert. If the URL has an XPointer, all partial documents that match the XPointer pattern are specified in the web document in the body of the URL.
【0070】
Structural constraints Parent element: optional Child element: None Annotation The pz: targets element interprets the single or multiple web resources specified by the href attribute as an XML-P'z document, inserts it into the context of the element, and the pz: targets element itself disappears. If the URL indicated by the href attribute has an XPointer, all partial documents that match the XPointer pattern are specified in the web document in the body part of the URL.
【0071】
·sample The following example is an XML-P'z document that captures all the book data contained in the "http://www.xxx.com/booklist.xml" page in addition to the book data contained in the document. Is.
【0072】
1. <? Xml version = 1.0?> 2. <bookstore specialty = novel 3. xmlns: pz = http://www.shiba.co.jp/xmlpz> 4. <book style = textbook> 5. <author> 6. <first-name> Shinichiro </ first-name> 7. <last-name> Hamada </ last-name> 8. <publication> Selected Short Stories of 9. <first-name> Shinichiro </ first-name> 10. <last-name> Hamada </ last-name> 11. </ publication> 12. </ author> 13. <price> 55 </ price> 14, </ book> 15. 15. <pz: targets href = http://www.xxx.com/booklist.xml#xpointer (// book) /> 16. </ bookstore> (pz: convert element) Convert any subdocuments using XSLT documents grammar <pz: converthref = xslt-url> </ pz: targets> attribute href The URL to the XSLT document that defines the conversion rules. When the URL has XPointer, the first partial document in the document order is specified among the partial documents that match the XPointer pattern in the web document in the body part of the URL.
【0073】
Structural constraints Parent element: optional Child element: optional Annotation The pz: convert element applies the XSLT document specified by the href attribute to each child document under the element and converts it. Each converted child document is interpreted as an XML-P'z document and then inserted into the context of the pz: convert element, and the pz: convert element itself disappears. When the URL indicated by the href attribute has XPointer, the first sub-document in the document order is specified among the sub-documents that match the XPointer pattern in the web document of the body part of the URL.
【0074】
sample The following example shows all the textbooks contained in the "http://www.xxx.com/booklist.xml" page, in addition to the textbook data contained in the self-document represented by the "textbook" element. The data is converted to the common book format according to the conversion rules described in the XSLT document "textbook-book.xsl", and it is also published on the "http://www.yyy.com/index.html" page. This is an XML-P'z document that captures all the data converted into a common book format.
【0075】
1. <? Xml version = 1.0?> 2. <bookstore specialty = novel xmlns: pz = http://www.shiba.co.jp/xmlpz> 3. <pz: convert href = textbook-book.xsl> 4. <textbook> 5. <author> 6. <first-name> Shinichiro </ first-name> 7. <last-name> Hamada </ last-name> 8. <publication> Selected Short Stories of 9. <first-name> Shinichiro </ first-name> 10. <last-name> Hamada </ last-name> 11. </ publication> 12. </ author> 13. <price> 55 </ price> 14. </ textbook> 15. 15. <pz: targets href = http://www.xxx.com/booklist.xml#xpointer (// textbook) /> 16. </ pz: convert> 17. <pz: convert href = html-book.xsl> 18. <pz: targets href = http://www.yyy.com/index.html#xpointer (// TABLE [2] // TR) /> 19. </ pz: convert> 20. </ bookstore> (Relative specification of URL with XPointer) When a web resource references another web resource, a relative URL can be used based on the URL of the own web resource. This is called a relative URL. In order to uniquely distinguish resources, the processing system must expand relative URLs to absolute URLs. The solution is shown below. However, in the following explanation, the terms are based on the IETF (http://www.ietf.org/rfc/rfc1738.txt).
【0076】
1.) When the object of the base URL and the object of the relative URL are different Between the body part of the base URL with the XPointer fragment removed (if any) and the body part of the relative URL with the XPointer fragment removed (if any), the IETF (http://www.ietf.org/rfc/rfc) Give the XPointer fragment of the relative URL (if any) to the result of the relative URL resolution based on (/rfc1808.txt). The XPointer fragment is, for example, "# xpointer (/ node1 / node2)" or "# xpointer (./node3// node4)" in the part below "#xpointer" in the description of the sample below. ..
【0077】
·sample (Base URL) http://aaa.com/dir1/xxx.xml#xpointer (/ node1/node2) (Relative URL) ./dir2/yyy.xml#xpointer (./node3//node4) (Solution result) http://aaa.com/dir1/dir2/yyy.xml#xpointer(./node3//node4) 2.) When the object of the base URL and the object of the relative URL are the same Determine the relative URL XPointer node (if any) starting from the document node indicated by XPointer if the base URL contains an XPointer fragment, or the root document node if it does not contain an XPointer fragment, and its node path. Give an XPointer fragment to the URL of the object.
【0078】
·sample (Base URL) http://aaa.com/dir1/xxx.xml#xpointer (/ node1/node2) (Relative URL) http://aaa.com/dir1/xxx.xml#xpointer (./node3//node4) (Solution result) http://aaa.com/dir1/xxx.xml#xpointer (/ node1/node2/node3 // node4) 3.) When the object is not specified in the relative URL Determine the relative URL XPointer node (if any) starting from the document node indicated by XPointer if the base URL contains an XPointer fragment, or the root document node if it does not contain an XPointer fragment, and its node path. Give an XPointer fragment to the URL of the object in the base URL.
【0079】
sample (Base URL) http://aaa.com/dir1/xxx.xml#xpointer (/ node1/node2) (Relative URL) #xpointer (./node3//node4) (Solution result) http://aaa.com/dir1/xxx.xml#xpointer (/ node1/node2/node3 // node4) (B-2) Configuration and operation of XML-P'z language processing system Next, the interpretation processing system of the XML-P'z language will be described.
【0080】
The XML-P'z language processing system is a software component that inputs a URL or source indicating the location of an XML-P'z document and outputs the XML document source of the interpretation result. This processing system uses a method that interprets the XML-P'z language in two passes. In the first pass, parsing is performed as XML to create an XML-DOM tree, and then in the second pass, XML is used. -While tracing the DOM tree with priority given to depth, it interprets the instruction elements (the part surrounded by the insertion and conversion instruction tags) peculiar to the XML-P'z language. In this language processing, even if a grammatical deviation is found or a runtime error such as a network trouble occurs, the processing policy is adopted to output the best possible result by continuing the interpretation processing as it is.
【0081】
Also, in XML-P'z language, it is possible to specify a web resource using a URL with XPointer, but in this processing system, after downloading the entire document indicated by the URL, the partial document specified by XPointer is cut out. Take a two-step process. This makes it possible to request web resources from most web servers that do not support URLs with XPointer.
【0082】
The above is the basic processing policy. An example of a system configuration of this processing system based on this processing policy will be described.
【0083】
FIG. 2 is an overall configuration example of the XML-P'z language processing system 100 (corresponding to the synthesis processing unit 1 in FIG. 1). In FIG. 2, the language processing system 100 is roughly divided into an interpretation buffer factory 101, which is a processing module related to XML-P'z document reading, and a processing module that returns XML as a result of interpreting the read document. , Interpreter 102. These basically work independently. The two interpretation buffer factories 101 in FIG. 2 are the same, but they are written separately for easy viewing.
【0084】
The interpretation buffer factory 101 starts operation triggered by the input of the URL or source indicating the location of the XML-P'z document. First, in the XML normalizer 111, if the input document is XML, the structure is the same as if it is HTML. After performing the equivalent conversion process to XML with, an XML-DOM tree is created using the XML-DOM parser 114, and the XPointer processor 115 extracts partial documents according to the XPointer fragment contained in the URL. Based on the result, the interpretation buffer initializer 116 generates interpretation buffers 103 and 104.
【0085】
Further, when the input of the URL or the source is from outside the processing system 100, the interpretation buffer to be generated is registered as the default interpretation buffer 103. Here, the interpretation buffer is a state memory of the XML-P'z language interpretation process, and is actively updated during the interpretation process of the interpreter 102.
【0086】
On the other hand, the interpreter 102 starts operation when there is a request for the interpretation result from the outside of the processing system 100, and while tracing the interpretation XML-DOM tree 131 of the default interpretation buffer 103 with priority on depth, the pz: targets element and Interpretation execution of two instruction elements of pz: convert element is performed, and the XML document of the final interpretation result is output.
【0087】
However, in order to perform XML-P'z interpretation processing on the partial document temporarily generated during the interpretation of the instruction element, the interpretation buffer factory 101 is used to generate the temporary interpretation buffer 104.
【0088】
Next, the processing operation of each component (module) constituting the interpretation buffer factory 101 will be described.
【0089】
The XML normalizer 111 that constitutes the interpretation buffer factory 101 is composed of an HTML determiner 112 and an HTML-XML converter 113.
【0090】
The HTML determiner 112 determines whether the web resource (web document) pointed to by the given URL is an HTML document or an XML document. The judgment is made by a two-step test, one is to use the "Content-type" of the HTTP header and the other is to use the extension included in the URL. This processing operation is shown in Fig. 3.
【0091】
In FIG. 3, first, the "Content-Type" is acquired (step S1). The most direct way to get this is to make a HEAD request to the URL. However, there are many web servers out there that do not understand HEAD requests. A GET request can also be used as an alternative. Next, it is determined whether or not an HTTP connection can be made to the URL (step S2). If the connection is successful, the process proceeds to step S3, and if the connection is unsuccessful, the process proceeds to step S5.
【0092】
In step S3, the "Content-Type" header is taken out, and it is determined whether or not the character string "text / html" is included in the header. If it is included, it is judged as HTML and finished (step S6), and if not, it is tentatively judged as XML and finished (step S4).
【0093】
In step S5, it is determined whether the extension of the object field in the URL is "html" or "htm". If so, it is determined to be HTML and terminated (step S6), otherwise it is tentatively determined to be XML and terminated (step S7).
【0094】
The HTML-XML converter 113 converts the web resource determined to be an HTML document by the HTML determiner 112 into a structurally equivalent XML document. This can be achieved by sequentially moving from the HTML-DOM tree to the XML-DOM tree using the methods of each DOM. Figure 4 shows the processing operation of the HTML-XML converter 113.
【0095】
First, in step S11, the given HTML document is read into the HTML parser, and an HTML-DOM tree is constructed. It is desirable that the HTML parser is used internally by the web browser. This is because the HTML parser used by web browsers has an error recovery function for HTML grammar deviations.
【0096】
Next, in step S12, an empty XML-DOM tree is constructed using the XML-DOM parser. Then, in step S13, while searching the HTML-DOM tree completely, the values of the nodes that stopped by are taken out and inserted as nodes in the XML-DOM tree.
【0097】
Through the above processing, the XML normalizer 111 outputs all the web resources input as URLs in the interpretation buffer factory 101 as XML documents. On the other hand, all web resources input as sources are treated assuming that they are XML documents.
【0098】
The XML document that has passed through the XML normalizer 111 or the XML document that is input as the source is input to the XML-DOM parser 114 and is converted into an XML-DOM tree. In addition, the XPointer processor 115 is used to obtain the XML-DOM tree of the partial document within the XML document indicated by the XPointer fragment of the URL. Figure 5 shows the processing operation of the XPointer processor 115 for the XPointer fragment.
【0099】
First, in step S21, it is determined whether the given web resource was due to a URL or a source. If it is from the source, the URL does not exist, so it ends at this point.
【0100】
Next, in step S22, the XPointer fragment is extracted from the URL fragment. However, if XPointer is not specified, it will be an empty string. Then, in step S23, the node pointed to by XPointer is identified with the root element of the XML-DOM tree as the base point. A general XPointer processing system may be used for this.
【0101】
Next, it is determined whether or not the node pointed to in step S24 is an element. If it is not an element, it terminates abnormally. Then, in step S25, the XML-DOM tree of the partial document with the obtained element as the root element is cut out. Further, in step S26, the cut out XML-DOM tree is used as the XML-DOM tree of the new XML document.
【0102】
Now, based on the obtained XML-DOM tree, the interpretation buffer initializer 116 creates an interpretation buffer. If the web resource given at this time is input from the outside of the language processing system 100, the interpretation buffer is registered as the default interpretation buffer 103. Figure 6 shows the initialization processing operation of this interpretation buffer (consisting of memory). In the case of the XML-DOM tree of the partial document, the temporary interpretation buffer 104 is initialized in the same manner as in FIG.
【0103】
First, in step S31, the given XML-DOM tree is copied to the source XML-DOM tree 134. The source XML-DOM tree 134 is a buffer that stores the initial state of the XML-DOM tree before it is changed by the subsequent interpretation process of the XML-P'z language, and is provided as the source of the XML-P'z language. However, it is not used in this embodiment.
【0104】
Next, in step S32, the given XML-DOM tree is copied to the XML-DOM tree 131 for interpretation. The XML-DOM tree 131 for interpretation is used by the interpreter 102 to read the structure and write the interpretation result in the interpretation process.
【0105】
In step S33, the program counter 132 is set in the root element of the XML-DOM tree 131 for interpretation. The program counter 132 is a pointer that stores the progress of the interpretation process of the interpreter 102.
【0106】
Finally, in step S34, the load flag 133 is set to "false". The load flag 133 is a flag indicating whether or not the interpretation buffer 103 has already been interpreted. The interpreter 102 does not reinterpret the interpretation buffer that has been interpreted in the past by using the flag 133.
【0107】
The above is the description of the processing operation of the interpretation buffer factory 101.
【0108】
Next, the processing operation of the interpreter 102 will be described.
【0109】
The context manager 121, which constitutes the interpreter 102, plays a central role in the interpretation process. According to the program counters 132 and 142 of the interpretation buffers 103 and 104, when each node of the XML-DOM tree for interpretation 131 and 141 is visited with priority on depth, if an instruction element is found, it is interpreted by the corresponding processing module (targets command processor 122, convert command processor 123). Request processing. When the interpretation processing of the instruction element is completed, the stop-by processing is continued. When all the processing is completed, the XML document is output as the interpretation result. This processing operation is shown in FIG. Hereinafter, the case of interpretation processing using the default interpretation buffer 103 will be described, but the same applies to the case of the temporary interpretation buffer 104.
【0110】
First, in step S41, the load flag 133 of the interpretation buffer 103 is examined. If the load flag is "true", it has already been interpreted, and if it is "false", it means that the interpretation process has not been performed yet. If "true", the process proceeds to step S49, and if "false", the process proceeds to step S42.
【0111】
In step S42, the program counter 132 is read to determine the element to be interpreted (referred to as the current element).
【0112】
In step S43, it is checked whether the element name of the current element is "pz: targets", and if it is "pz: targets", the process proceeds to step S4 and the targets command processor 122 is requested to interpret the pz: targets element. To do.
【0113】
Next, in step S45, it is checked whether the element name of the current element is "pz: convert", and if it is "pz: convert", the process proceeds to step S46, and the interpretation process of the pz: convert element is performed by the convert command processor 123. Ask to.
【0114】
Then, in step S47, the destination element is determined with priority given to depth and set in the program counter. If there is an element that has not been interpreted yet among the child elements of the current element, the eldest brother element of them is set in the program counter. If all child elements have been interpreted, set the parent element to the program counter. However, if there is no parent element, the program counter is set to "NULL".
【0115】
In step S8, the program counter 132 is checked for "NULL", and if it is not "NULL", the process returns to step S42. If it is "NULL", the interpretation of the XML-DOM tree 131 for interpretation is completed, and the process proceeds to step S49.
【0116】
In step S49, the XML-DOM parser 151 is used to generate and output an XML document based on the XML-DOM tree 131 of the interpretation buffer 103, and the process ends.
【0117】
The targets command processor 122 that constitutes the interpreter 102 interprets the pz: targets element and writes the result to the current element. This processing operation is shown in FIG.
【0118】
First, in step S51, the href attribute value of the pz: targets element, which is the current element, is extracted, and in step S52, the attribute value is used as the input URL of the interpretation buffer factory 101, and is processed by the interpretation buffer initializer 116 from the XML normalizer 111 described above. The temporary interpretation buffer 104 is generated via. However, if the target URL is a relative URL, it will be converted to an absolute URL based on the URL of the interpretation buffer of the insertion destination based on the explanation of "Relative specification of URL with XPointer" described above.
【0119】
Next, the process proceeds to step S53, the generated temporary interpretation buffer 104 is interpreted using the interpreter 102, and the resulting XML document is obtained.
【0120】
Finally, in step S54, the DOM parser 152 is used to convert the resulting XML document into an XML-DOM tree and replace it with the current element "pz: targets" element. Also, the generated temporary interpretation buffer 104 is discarded.
【0121】
The convert command processor 123 that constitutes the interpreter 102 interprets the convert element and writes the result to the current element. This processing operation is shown in FIG.
【0122】
First, in step S61, the href attribute value of the pz: convert element, which is the current element, is extracted, and in step S62, the attribute value is used as the input URL of the interpretation buffer factory 101, and is processed by the interpretation buffer initializer 116 from the XML normalizer 111 described above. The temporary interpretation buffer 104 is generated via. However, if the target URL is a relative URL, it will be converted to an absolute URL based on the URL of the interpretation buffer of the insertion destination based on the above explanation (relative specification of URL with XPointer).
【0123】
Next, the process proceeds to step S63, and the generated temporary interpretation buffer 104 is interpreted using the interpreter 102, and as a result, an XSLT document is obtained. It should be noted that such processing is performed because the XSLT document itself may be written in the XML-P'z language (that is, the XSLT document may be composed as a synthesis result). is there).
【0124】
Then, the process proceeds to step S64, and the XSLT processor 124 applies the child element of the current element "pz: convert" element to the eldest brother element (and the partial document including its descendant element) to which XLST has not been applied yet. Using the obtained XSLT document, the document structure of the partial document is converted using the conversion rules described in the XSLT document, and the XML-DOM tree obtained by the conversion is converted into the compositing web in step S65. Replace with the unconverted child element (and the subdocument containing its descendant elements) on the document.
【0125】
In step S66, if there are unprocessed child elements, the process returns to step S64. If all child elements have been processed, the process proceeds to step S67, where the pz: convert element is replaced with the converted document structure that is each child subdocument of the pz: convert element.
【0126】
The above is the processing operation of the interpreter 102, and the explanation of each component of the XML-P'z language processing system is completed.
【0127】
(C) A series of operations for synthesizing multiple web documents on one web document Next, the XML-P'z language processing system 100 having the configuration shown in Fig. 2 is incorporated into the web server, the basic operation shown in Fig. 1 is performed, and the actual operation is performed from the web document W2 of the web server A2. Flow charts shown in FIGS. 13 to 15 show a series of operations for extracting a part, synthesizing each extracted partial document on one web document, and outputting the synthesized web document (XML document) W1. Will be described with reference to.
【0128】
Here, it is assumed that the XML-P'z document 2 as a web document for synthesis is shown in FIG. The XML-P'z document shown in FIG. 16 is an excerpt of a part of the XML-P'z document 2 shown in FIG.
【0129】
The XML-P'z document shown in Fig. 16 is the textbook data contained in the own document represented by the "textbook" element E1 and the "http: // www" inserted by the pz: targets element E2. All the textbook data contained in the web document ".xxx.com/booklist.xml" is converted to the common book format according to the conversion rule described in the XSLT document "textbook-book.xsl" and synthesized. It is for outputting the created web document (XML document) W1.
【0130】
In FIG. 1, it is assumed that the request for XML-P'z document 2 is made from the web browser of the client terminal B1 to the XML-P'z server A1 (hereinafter, simply referred to as server A1) (step S201).
【0131】
Since the language processing system 100 of the server A1 is a synthetic web document (XML-P'z document) 2 that the requested document has, the XML-DOM parser 114 is used to display the XML-P'z document. Create an XML-DOM tree (step S202). The part of the created XML-DOM tree corresponding to FIG. 16 is shown in FIG. 17, for example. It should be noted that FIG. 17 is shown schematically for the sake of simplicity of explanation.
【0132】
Copy this created XML-DOM tree to the source and interpretation DOM trees 134,131 of the default interpretation buffer 103, and initialize the default interpretation buffer 103 as shown in FIG. 6 (step S203).
【0133】
Next, the interpreter 102 performs the interpretation process of the default interpretation buffer 103. Here, for example, it is assumed that the XML-DOM tree as shown in FIG. 17 is interpreted.
【0134】
As described above, the interpreter 102 determines the destination element by giving priority to the depth of the instruction element. Therefore, in the DOM tree shown in FIG. 17, the pz: targets element E2 is first interpreted (step S204). ~ Step S205). After that, the pz: convert element E3, which is the parent element of the elements E1 and E2, is interpreted (step S206 to step S207). After that, although not shown in FIG. 17, the program counter 132 is moved to the younger brother element or the parent element of the pz: convert element E3, and the default interpretation buffer 103 of this default interpretation buffer 103 is moved until the program counter becomes "NULL". The interpretation process proceeds (step S208).
【0135】
By the way, in step S205, the interpretation processing of the pz: targets element E2 is performed, and the processing operation here is shown in FIG.
【0136】
The targets command processor 122 extracts the href attribute value of the pz: targets element E3, that is, "http://www.xxx.com/booklist.xml#xpointer (// textbook)" and interprets the attribute value. The input URL is 101. If the document specified in this input URL is not an XML document, the XML normalizer 111 converts it to an XML document (step S212), and then creates an XML-DOM tree for this XML document in the XML-DOM parser 114. (Step S213). Since the specified document is an XML document here, the XML-DOM parser 114 creates an XML-DOM tree of this XML document as it is.
【0137】
In this case, since the above input URL is a URL with XPointer indicating the web document W2 of server A2, the XPointer processor 115 takes out the XPointer fragment, that is, "#xpointer (// textbook)" and creates it in step S213. Cut out the XML-DOM tree of the "textbook" element (partial document including its descendant elements) pointed to by the XPointer from the XML-DOM tree. If there are multiple "textbook" elements, do so for each. This cut out XML-DOM tree is the XML-DOM tree of the partial document to be inserted (step S214).
【0138】
Next, the interpretation buffer initializer 116 initializes the temporary interpretation buffer 104, and if a pz: targets element or pz: convert element is described in this subdocument, interprets them and performs the interpretation processing of the subdocument. Get the XML document.
【0139】
If it is not described, the interpretation process of the temporary interpretation buffer 104 is terminated as it is, and the context manager 121 generates an XML document from the XML-DOM tree of the partial document using the DOM parser 151 (step S221). The targets command processor 122 uses the DOM parser 152 to create an XML-DOM tree of the XML document of the subdocument, and uses this as the subdocument group E2 ́ to interpret the XML-DOM tree 131 of the default interpretation buffer 103. Replace with the pz: targets element E2, which is the current element of. As a result, as shown in Figure 18, this subdocument group E2 ́ becomes a child element of the pz: convert element E3 and the XML-DOM tree is updated. The generated temporary interpretation buffer 104 is discarded (step S222). After that, the process returns to step S208 in FIG.
【0140】
As shown in Fig. 18, since there are multiple textbook data in the web document of "http://www.xxx.com/booklist.xml", all of them are XML-DOM of the partial document of the web document. It is inserted as a tree.
【0141】
On the other hand, in step S207, the interpretation processing of the pz: convert element E3 is performed, and the processing operation here is shown in FIG.
【0142】
The convert command processor 123 extracts the href attribute value of the pz: convert element E3, that is, the URL to the XSLT document, "textbook-book.xsl", and uses that attribute value as the input URL of the interpretation buffer factory 101. The following steps S232 to S240 are processes for obtaining an XSLT document as an XML document, and are shown in FIG. 19 in step S241 of FIG. 15 in the same manner as in steps S212 to S220 of FIG. Get an XSLT document as an XML document like.
【0143】
The XSLT document shown in Figure 19 is a conversion to convert the "publication" element, "price" element, and "author" element of the current partial document to the "title" element, "price" element, and "author" element, respectively. It describes the rules.
【0144】
Using an XSLT document as shown in FIG. 19, the XSLT processor 124 uses a subdocument (also a child subdocument) contained in the pz: convert element, which is the current element of the XML-DOM tree 131 for interpretation of the default interpretation buffer 103. Transform each child element on the XML-DOM tree (to call) (step S242).
【0145】
Here, the textbook data contained in the own document and the textbook data extracted from the web document of "http://www.xxx.com/booklist.xml" have the same structure, so element E1 Taking the case of the textbook data included in the own document as an example, the case of converting the structure will be described using the XSLT document of FIG.
【0146】
As shown in FIG. 16, the value of the "publication" element, which is a child element of the element E1, is "Selected Short Stories of Shinichiro Hamada", which becomes the value of the "title" element after conversion. Further, in FIG. 16, the value of the "author" element, which is a child element of the element E1, is "Shinichiro Hamada", which becomes the "author" element after conversion. Further, as shown in FIG. 16, the value of the "price" element, which is a child element of the element E1, is "55", which is the same after the conversion.
【0147】
The convert command processor 123 replaces the converted XML-DOM tree of the partial document as a new element E3 ́ with the pz: convert element E3, which is the current element of the XML-DOM tree 131 for interpretation of the default interpretation buffer 103. , An XML-DOM tree with a document structure as shown in Figure 20 is generated.
【0148】
The generated temporary interpretation buffer 104 is discarded (step S243). After that, the process returns to step S208 in FIG.
【0149】
As described above, when the program counter 132 of the default interpretation buffer 103 becomes "NULL" and the interpretation of the XML-DOM tree 131 is completed, the context manager 121 uses the XML-DOM parser 151 and is shown in FIG. Generates and outputs an XML document as the target web document W1 based on the XML-DOM tree 131 of the interpretation buffer 103 including the XML-DOM tree.
【0150】
If the web browser of the client terminal B1 can display the XML document, the web document W1 of the XML document is returned to the web browser of the client terminal B1 as it is, but if it cannot be displayed, the style sheet is processed on the server A1 side. , Convert the web document W1 to an HTML document and then return it to the web browser of the client terminal B1 (step S209 in FIG. 13).
【0151】
(D) Coordination between XML-P'z servers for compositing web documents Next, a case where the composition processing of the Web document is performed cooperatively between the XML-P'z servers will be described.
【0152】
For example, if you want to insert an XML-P'z document from another XML-P'z server while interpreting an XML-P'z document on one XML-P'z server, the XML-P to be inserted There is a question of which server interprets the'z document. In other words, when there is a request by the GET command, it is necessary to determine whether to return the XML-P'z document itself or the XML document as a result of interpretation processing.
【0153】
If the HTTP client cannot interpret and process the XML-P'z document between the HTTP server (the side that requests the XML-P'z document) and the HTTP client (the side that requests the XML-P'z document) , There is a restriction that the XML-P'z document must be interpreted and processed on the HTTP server side.
【0154】
In order to introduce this constraint into the judgment material, when the interpretation buffer factory 101 of the XML-P'z language processing system 100 requests an XML-P'z document, "XML-P" is added to the header of the request by the GET command. 'z: enable' shall be added.
【0155】
Also, as an HTTP server, there is an advantage that the load on the server can be reduced by delegating the interpretation processing of the XML-P'z document to the HTTP client, but there is something that you do not want to publish the XML-P'z document. There may be a reason (for example, you don't want to expose the included synthetic logic), so it's up to you to interpret the XML-P'z language on the server side.
【0156】
Based on the above, the operation of determining whether or not the HTTP server interprets and executes is described with reference to the flowchart shown in FIG.
【0157】
First, in step S71, it is checked whether "XML-P'z: enable" is included in the header of the GET request, and if it is not included, the process proceeds to step S72 and XML-P'z on the HTTP server. Interpret the document and finish. If so, go to step S73 and check if the HTTP server is set to process XML-P'z documents, and if so, go to step S74 and XML-P on the HTTP server. Interpret and exit the'z document, otherwise proceed to step S75, send the XML-P'z document as is to the HTTP client without interpreting and exit.
【0158】
(E) Addendum As described above, according to the above embodiment, the synthesis web document that is the base for synthesis is described in XML, and the part (partial document) in the specified range is extracted from the other specified web documents. , Inserts it at the specified position of the compositing web document and performs conversion processing on the specified range of the compositing web document. It has two compositing logic instructions of insertion and conversion as elements in the compositing web document. Define the set XML-P'z (XML-Pieces) document. The language processing system 100 is a part of the range specified from the web documents (pages) W2 and W3 of the specified web server (for example, in this case, web servers A2 and A3) described in the XML-P'z document. (Partial document) is extracted, inserted at the specified position of the XML-P'z document, and converted to the specified range described in the XML-P'z document. Finally, by obtaining the XML document (synthesized web document) W1 as the processing result of the XML-P'z language processing system 100, it is possible to synthesize the information of multiple websites on one web document. It can be done easily and universally.
【0159】
The method described in the above embodiment can also be stored and distributed in a recording medium such as a DVD, a CD-ROM, a floppy disk, an individual memory, or an optical disk as a program that can be executed by a computer.
【0160】
[Effect of the invention]
As described above, according to the present invention, it is possible to easily and universally combine information from a plurality of websites on one web document.
[Simple explanation of drawings]
[Figure 1]
The figure for demonstrating the basic operation of the web server (XML-P'z server) which incorporated the XML-P'z language processing system of this invention.
[Figure 2]
The figure which showed the whole structure example of the XML-P'z language processing system.
[Fig. 3]
A flowchart showing a processing operation for determining whether a web document specified by a given URL is an HTML document or an XML document in an HTML judgment device.
[Fig. 4]
Flowchart for explaining the conversion process operation from HTML document to XML document of HTML-XML converter.
[Fig. 5]
A flowchart for explaining the processing operation of the XPointer processor for XPointer fragments.
[Fig. 6]
A flowchart for explaining the initialization processing operation of the interpretation buffer of the interpretation buffer initializer.
[Fig. 7]
A flowchart for explaining the processing operation of the context manager.
[Fig. 8]
A flowchart for explaining the interpretation processing operation of the targets element of the targets command processor.
[Fig. 9]
A flowchart for explaining the interpretation processing operation of the convert element of the convert command processor.
[Fig. 10]
A flowchart for explaining the judgment processing operation for determining whether the interpretation processing of the XML-P'z document is performed on the server side or the client side.
[Fig. 11]
The figure (a) is a diagram schematically showing the document structure of the first example of the XML-P'z document, and the figure (b) is the document structure of the XML document after the interpretation of the XML-P'z document. The figure which showed.
[Fig. 12]
Diagram for explaining the interpretation order of XML-P'z documents.
[Fig. 13]
A flowchart for explaining the operation of a series for synthesizing a plurality of web documents on one web document by the language processing system having the configuration shown in FIG.
[Fig. 14]
A flowchart for explaining the operation of a series for synthesizing a plurality of web documents on one web document by the language processing system having the configuration shown in FIG.
[Fig. 15]
A flowchart for explaining the operation of a series for synthesizing a plurality of web documents on one web document by the language processing system having the configuration shown in FIG.
[Fig. 16]
A diagram showing a part of an XML-P'z document, which is an example of an XML-P'z document as a web document for compositing.
[Fig. 17]
A schematic diagram of the XML-DOM tree that corresponds to the XML-P'z document in Figure 16.
[Fig. 18]
A schematic diagram of the XML-DOM tree resulting from the interpretation of the pz: targets element in Figure 16.
[Fig. 19]
The figure which showed an example of the XSLT document described in the XML-P'z document of FIG.
[Fig. 20]
A schematic diagram of the XML-DOM tree resulting from the interpretation of the pz: targets and pz: convert elements in Figure 16.
[Explanation of symbols]
A1, A2, A3 ... server B1 ... Client terminal W1 ... Synthesized web document (XML document) W2 ~ W3 ... Web documents 1 ... XML-P'z language processing system (synthesis processing part) 2 ... XML-P'z document 100 ... XML-P'z language processing system 101 ... Interpretation Buffer Factory 102 ... interpreter 103 ... default interpretation buffer 104 ... Temporary interpretation buffer 111 ... XML normalizer 112 ... HTML Judge 113 ... HTML-XML converter 114 ... XML-DOM parser 115 ... XPointer processor 116 ... Interpretation Buffer Initializer 121 ... Context Manager 122 ... targets Command Manager 123 ... convert command manager 124 ... XSLT processor 131 ... XML-DOM tree for interpretation 132 ... Program counter 133 ... load flag 134 ... Source XML-DOM Tree 141 ... XML-DOM tree for interpretation 142 ... Program counter 143 ... load flag 144 ... Source XML-DOM Tree 151 ~ 153 ... DOM parser
6 sheets
Sheet 1 Sheet 2 Sheet 3 Sheet 4 Sheet 5 Sheet 6
Every citation, both ways
| Document | Relation | Office | Cited during |
|---|---|---|---|
| JP2005032230A | Cited by | Japan | Examiner |
| US7661063B2 | Cited by | United States of America | Applicant |
| JP2015108874A | Cited by | Japan | Search report |
| KR101251686B1 | Cited by | Republic of Korea | Search report |
| JP2007310564A | Cited by | Japan | Search report |
| JP2008305180A | Cited by | Japan | Examiner |
| US7900136B2 | Cited by | United States of America | Applicant |
| JP2004046357A | Cited by | Japan | Search report |
| WO2004013765A1 | Cited by | World Intellectual Property Organization (WIPO) | International search |
| JP2007249619A | Cited by | Japan | Examiner |
| KR101065937B1 | Cited by | Republic of Korea | Search report |
| JP2008538841A | Cited by | Japan | Examiner |
| WO2006001392A1 | Cited by | World Intellectual Property Organization (WIPO) | International search |
| JP2006154952A | Cited by | Japan | Examiner |
| JP2007122609A | Cited by | Japan | Examiner |
| JP2015108874A | Cited by | Japan | Search report |
| US8086954B2 | Cited by | United States of America | Applicant |
3 members in 2 offices
Priority claims2
| Document | Office | Kind | Date |
|---|---|---|---|
| 2000383625 | Japan | A | |
| JP20000383625 | – | – | – |
Members3
| Document | Office | Kind | |
|---|---|---|---|
| US2002078105A1 | United States of America | A1 | |
| JP2002183116AThis record | Japan | A | |
| JP3943830B2 | Japan | B2 |
7 legal events, as the office reported them to INPADOC
Over the term
Point at a mark for the eventEvents
| Event | Code | |
|---|---|---|
| Cancellation because of no payment of annual feesLAPS | LAPS | |
| First payment of annual fees (during grant procedure)JAPANESE INTERMEDIATE CODE: A61A61 | A61 | |
| Written decision to grant a patent or to grant a registration (utility model)JAPANESE INTERMEDIATE CODE: A01A01 | A01 | |
| Decision of grant or rejection writtenTRDD | TRDD | |
| Written amendmentJAPANESE INTERMEDIATE CODE: A523A521 | A521 | |
| Notification of reasons for refusalJAPANESE INTERMEDIATE CODE: A131A131 | A131 | |
| Report on retrievalJAPANESE INTERMEDIATE CODE: A971007A977 | A977 |
Numbers
- Publication
- 2002-183116
- Publication, DOCDB
- 2002183116
- Publication, EPODOC
- JP2002183116
- Application
- 383625
- Application, DOCDB
- 2000383625
- Application, EPODOC
- JP20000383625
Titles2
- Japanese
- 【発明の名称】文書合成方法および文書合成装置
- English
- [Title of Invention] Document Synthesis Method and Document Synthesis Device
Classification
- CPC, 1
- G06F16/958
- IPC, 3
- G06F17 21
- G06F12 00
- G06F17 30