Compiler, compiler program, recording medium, control method and central processor
Abstract
Problem to be solved.To provide a compiler, a compiler program, a recording medium, a control method and a central processor which realize optimized conversion of a character code system.
Solution.The compiler optimizes the conversion of the character code system of characters to be stored in a character variable in an object program to be optimized and is provided with: a conversion instruction generation part which reads characters of the character variable written by a first character code system, converts the characters from the first character code system into a second character code system and generates a conversion instructions to be stored in the character variable prior to each of a plurality of processings using the above read characters in the second character code system; and a conversion instruction removal part which removes the conversion instructions about each conversion instruction generated by the conversion instruction generation part when the characters in the second character code system are stored in the character variable in all execution paths to be executed prior to the conversion instructions.
Copyright (C)2006,JPO&NCIPI

Term
No projected expiry on record.
- Priority and filed
- Published
- Today
21 claims: 9 independent, 12 dependent
- 1A compiler that optimizes the conversion of the character code system of the characters stored in the character variable in the target program to be optimized. It reads the characters of the character variable written by the first character code system and reads the second character. Prior to each of a plurality of processes using the character in the code system, a conversion instruction for converting the character from the first character code system to the second character code system and storing it in the character variable is generated. For the conversion instruction generation unit and each conversion instruction generated by the conversion instruction generation unit, the characters of the second character code system are stored in the character variables in all the execution paths executed prior to the conversion instruction. A compiler including a conversion instruction removing unit that removes the conversion instruction when the conversion instruction is performed. 最適化対象の対象プログラムにおいて文字変数に格納される文字の文字コード体系の変換を最適化するコンパイラであって、 第1の文字コード体系により書き込まれた文字変数の文字を読み出して第2の文字コード体系において当該文字を使用する複数の処理の各々に先立って、当該文字を前記第1の文字コード体系から前記第2の文字コード体系に変換して当該文字変数に格納する変換命令を生成する変換命令生成部と、 前記変換命令生成部により生成された各変換命令について、当該変換命令に先立って実行される全ての実行パスにおいて、前記文字変数に前記第2の文字コード体系の文字が格納される場合に、当該変換命令を除去する変換命令除去部とを備えるコンパイラ。
- 8A compiler that optimizes the conversion of the character code system of the characters stored in the character variable in the target program to be optimized. It reads the characters of the character variable written by the first character code system and reads the second character. As each of a plurality of processes using the character in the code system, the process is executed when the character stored in the character variable is the second character code system, and the character is the first character code. A character operation processing generator that generates an instruction sequence with an exception that raises an exception when it is a system, and a character that is executed when the exception occurs and is stored in the character variable from the first character code system. A compiler including an exception handler generator that converts to the second character code system, stores it in the character variable, and generates an exception handler that returns to the process that caused the exception. 最適化対象の対象プログラムにおいて文字変数に格納される文字の文字コード体系の変換を最適化するコンパイラであって、 第1の文字コード体系により書き込まれた文字変数の文字を読み出して第2の文字コード体系において当該文字を使用する複数の処理の各々として、前記文字変数に格納される文字が前記第2の文字コード体系である場合に当該処理を実行し、当該文字が前記第1の文字コード体系である場合に例外を発生させる例外付命令列を生成する文字操作処理生成部と、 前記例外が発生した場合に実行され、前記文字変数に格納される文字を前記第1の文字コード体系から前記第2の文字コード体系に変換して当該文字変数に格納し、前記例外を発生させた当該処理に復帰させる例外ハンドラを生成する例外ハンドラ生成部とを備えるコンパイラ。
- 11A compiler that optimizes the conversion process that converts the characters stored in the character variables in the target program to be optimized from the first character code system to the second character code system, and converts the character code of the characters to be converted. The acquisition process generation unit that generates the instruction sequence of the process to be acquired, each of the plurality of conversion processes that should be selected and executed according to the range of the character code values, and the range of the character code values. Based on the detection result of the detection process, the conversion detection process generation unit that generates an instruction sequence that executes the detection process for detecting the above, and the process result of any one of the plurality of conversion processes. A compiler equipped with a selection processing generator that generates an instruction sequence to be selected and output. 最適化対象の対象プログラムにおいて文字変数に格納される文字を第1の文字コード体系から第2の文字コード体系に変換する変換処理を最適化するコンパイラであって、 変換対象の文字の文字コードを取得する処理の命令列を生成する取得処理生成部と、 前記文字コードの値の範囲に応じて何れかを選択して実行すべき複数の前記変換処理の各々と、前記文字コードの値の範囲を検出する検出処理とを並行して実行する命令列を生成する変換検出処理生成部と、 前記複数の変換処理のうち何れかの変換処理の処理結果を、前記検出処理の検出結果に基づいて選択して出力する命令列を生成する選択処理生成部とを備えるコンパイラ。
- 15It is a central processing device that executes a conversion process that converts characters in the first character code system, which have different data sizes according to the type of characters, into characters in the second character code system, and is the data size of the character code of the characters. Each time, in parallel with the conversion process to be executed when a plurality of character codes of the data size are continuously arranged in the register, a plurality of character codes of the data size are continuously arranged in the register. A central processing device having an instruction to execute a detection process for detecting whether or not a data is generated. 文字の種類に応じてデータサイズの異なる第1の文字コード体系の文字を第2の文字コード体系の文字に変換する変換処理を実行する中央処理装置であって、 前記文字の文字コードのデータサイズ毎に、当該データサイズの複数の文字コードがレジスタに連続して配列される場合に実行すべき前記変換処理と並行して、当該データサイズの複数の文字コードが前記レジスタに連続して配列されるか否かを検出する検出処理を実行する命令を有する中央処理装置。
- 16A compiler program that makes a computer function as a compiler that optimizes the conversion of the character code system of characters stored in character variables in the target program to be optimized. The computer is written by the first character code system. Prior to each of a plurality of processes in which the character of the character variable is read and the character is used in the second character code system, the character is converted from the first character code system to the second character code system. For the conversion instruction generation unit that generates the conversion instruction to be stored in the character variable and each conversion instruction generated by the conversion instruction generation unit, the character variable is set to the character variable in all the execution paths executed prior to the conversion instruction. A compiler program that functions as a conversion instruction removal unit that removes the conversion instruction when the characters of the second character code system are stored. 最適化対象の対象プログラムにおいて文字変数に格納される文字の文字コード体系の変換を最適化するコンパイラとしてコンピュータを機能させるコンパイラプログラムであって、 前記コンピュータを、 第1の文字コード体系により書き込まれた文字変数の文字を読み出して第2の文字コード体系において当該文字を使用する複数の処理の各々に先立って、当該文字を前記第1の文字コード体系から前記第2の文字コード体系に変換して当該文字変数に格納する変換命令を生成する変換命令生成部と、 前記変換命令生成部により生成された各変換命令について、当該変換命令に先立って実行される全ての実行パスにおいて、前記文字変数に前記第2の文字コード体系の文字が格納される場合に、当該変換命令を除去する変換命令除去部と して機能させるコンパイラプログラム。
- 17A compiler program that functions a computer as a compiler that optimizes the conversion of the character code system of characters stored in character variables in the target program to be optimized. The computer is written by the first character code system. As each of a plurality of processes for reading the character of the character variable and using the character in the second character code system, the process is performed when the character stored in the character variable is the second character code system. A character operation processing generator that executes and generates an instruction string with an exception that raises an exception when the character is the first character code system, and is executed when the exception occurs and stored in the character variable. Exception handler generator that converts the character to be executed from the first character code system to the second character code system, stores it in the character variable, and generates an exception handler that returns to the process that caused the exception. A compiler program that functions as. 最適化対象の対象プログラムにおいて文字変数に格納される文字の文字コード体系の変換を最適化するコンパイラとして、コンピュータを機能させるコンパイラプログラムであって、 前記コンピュータを、 第1の文字コード体系により書き込まれた文字変数の文字を読み出して第2の文字コード体系において当該文字を使用する複数の処理の各々として、前記文字変数に格納される文字が前記第2の文字コード体系である場合に当該処理を実行し、当該文字が前記第1の文字コード体系である場合に例外を発生させる例外付命令列を生成する文字操作処理生成部と、 前記例外が発生した場合に実行され、前記文字変数に格納される文字を前記第1の文字コード体系から前記第2の文字コード体系に変換して当該文字変数に格納し、前記例外を発生させた当該処理に復帰させる例外ハンドラを生成する例外ハンドラ生成部として機能させるコンパイラプログラム。
- 18A compiler program that makes a computer function as a compiler that optimizes the conversion process that converts the characters stored in character variables in the target program to be optimized from the first character code system to the second character code system. A plurality of the conversions to be executed by selecting and executing one of the acquisition processing generation unit for generating the instruction sequence of the processing for acquiring the character code of the character to be converted and the range of the value of the character code. A conversion detection processing generation unit that generates an instruction sequence that executes each of the processing and a detection processing that detects a range of the character code values in parallel, and a conversion processing process of any one of the plurality of conversion processes. A compiler program that functions as a selection process generator that generates an instruction sequence that selects and outputs the result based on the detection result of the detection process. 最適化対象の対象プログラムにおいて文字変数に格納される文字を第1の文字コード体系から第2の文字コード体系に変換する変換処理を最適化するコンパイラとして、コンピュータを機能させるコンパイラプログラムであって、 前記コンピュータを、 変換対象の文字の文字コードを取得する処理の命令列を生成する取得処理生成部と、 前記文字コードの値の範囲に応じて何れかを選択して実行すべき複数の前記変換処理の各々と、前記文字コードの値の範囲を検出する検出処理とを並行して実行する命令列を生成する変換検出処理生成部と、 前記複数の変換処理のうち何れかの変換処理の処理結果を、前記検出処理の検出結果に基づいて選択して出力する命令列を生成する選択処理生成部として機能させるコンパイラプログラム。
- 20This is a control method in which a computer controls a compiler that optimizes the conversion of the character code system of characters stored in character variables in the target program to be optimized, and is written by the computer by the first character code system. The character of the character variable is read out, and the character is converted from the first character code system to the second character code system prior to each of a plurality of processes using the character in the second character code system. In the conversion instruction generation stage for generating the conversion instruction to be stored in the character variable, and for each conversion instruction generated in the conversion instruction generation stage, the character variable is executed in all the execution paths prior to the conversion instruction. A control method including a conversion instruction removal step of removing the conversion instruction when the characters of the second character code system are stored in the second character code system. 最適化対象の対象プログラムにおいて文字変数に格納される文字の文字コード体系の変換を最適化するコンパイラを、コンピュータにより制御する制御方法であって、 前記コンピュータにより、 第1の文字コード体系により書き込まれた文字変数の文字を読み出して第2の文字コード体系において当該文字を使用する複数の処理の各々に先立って、当該文字を前記第1の文字コード体系から前記第2の文字コード体系に変換して当該文字変数に格納する変換命令を生成する変換命令生成段階と、 前記変換命令生成段階において生成された各変換命令について、当該変換命令に先立って実行される全ての実行パスにおいて、前記文字変数に前記第2の文字コード体系の文字が格納される場合に、当該変換命令を除去する変換命令除去段階と を備える制御方法。
- 21This is a control method in which a computer controls a compiler that optimizes the conversion process that converts the characters stored in the character variables in the target program to be optimized from the first character code system to the second character code system. The computer generates an instruction sequence for a process of acquiring the character code of the character to be converted, and a plurality of the conversions to be executed by selecting one according to the value range of the character code. A conversion detection process generation stage that generates an instruction sequence that executes each of the processes and a detection process that detects a range of the character code values in parallel, and a process of any of the plurality of conversion processes. A control method including a selection process generation step of generating an instruction sequence that selects and outputs a result based on the detection result of the detection process. 最適化対象の対象プログラムにおいて文字変数に格納される文字を第1の文字コード体系から第2の文字コード体系に変換する変換処理を最適化するコンパイラを、コンピュータにより制御する制御方法であって、 前記コンピュータにより、 変換対象の文字の文字コードを取得する処理の命令列を生成する取得処理生成段階と、 前記文字コードの値の範囲に応じて何れかを選択して実行すべき複数の前記変換処理の各々と、前記文字コードの値の範囲を検出する検出処理とを並行して実行する命令列を生成する変換検出処理生成段階と、 前記複数の変換処理のうち何れかの変換処理の処理結果を、前記検出処理の検出結果に基づいて選択して出力する命令列を生成する選択処理生成段階とを備える制御方法。
Independent claims9
143 paragraphs, as filed
The present invention relates to a compiler, a compiler program, a recording medium, a control method, and a central processing unit. In particular, the present invention relates to a compiler, a compiler program, a recording medium, a control method, and a central processing unit that convert a character code system of a character variable.
In recent years, XML (eXtensible Markup Language) has been attracting attention as a technology for structuring and handling various types of data in a unified manner. In XML, it is recommended to use UTF8 (8-bit UCS Transformation Format), which is a character code system that handles characters used in various countries in a unified manner. UTF8 can express the alphabet, which is expected to be used frequently, in 1 byte, while Japanese characters can be expressed in about 3 bytes. In this way, UTF8 has different data sizes depending on the type of characters and the like.
In recent years, many APIs for efficiently analyzing and editing XML documents have been prepared in the Java (registered trademark) language. However, the Java language usually treats characters as UTF16 (16-bit UCS Transformation Format) data. Therefore, in order for a program written in Java (registered trademark) language to operate an XML document, it is necessary to convert UTF8 to UTF16. Furthermore, in order to output the characters processed by the program written in Java (registered trademark) language as an XML document, it is necessary to convert UTF16 to UTF8.
Conventionally, a technique for converting a continuously arranged UTF8 character string into a UTF16 character string has been used (see Non-Patent Document 1). This technology reads UTF8 characters one by one from a character string, determines the data length and character type of the read characters, and performs different conversion processing depending on the determination result. Further, conventionally, in the Java (registered trademark) language, a library program for manipulating UTF8 characters without converting them to UTF16 has been proposed (see Non-Patent Document 2).
Non-patent documents 3 and 4 will be described later.<nplcit num="1"><text>Internet URL "http://cvs.apache.org/viewcvs.cgi/xml-xerces/java/src/org/apache/xerces/impl/io/UTF8Reader.java?rev=1.7&content-type=text/vnd. viewcvs-markup "</text></nplcit><nplcit num="2"><text>S. Makino, K. Tamura, T. Imamura, and Y. Nakamura. Implementation and Performance of WS-Security, IBM Research Report RT0546, 2003.</text></nplcit><nplcit num="3"><text>J. Knoop, O. Ruthing, and B. Steffen. Optimal code motion: theory & practice. ACM TOPLAS, 18 (3): 300-324, 1996.</text></nplcit><nplcit num="4"><text>R. Bodik, R. Gupta, and ML Soffa. Complete removal of redundant expressions. In Proceedings of the ACM SIGPLAN 1998 Conference on Programming Language Design and Implementation, pages 1-14, 1998.</text></nplcit>
<p> According to the technique of Non-Patent Document 2, the process of converting UTF8 characters to UTF16 can be omitted. However, many APIs for manipulating UTF16 characters have already been widely developed, and it is not possible to efficiently develop programs using these existing APIs. In addition, the processing targeting UTF8 is often inefficient as compared with the processing targeting UTF16. Therefore, the first issue is to improve the efficiency of the conversion process from UTF8 to UTF16 while effectively utilizing the API that has already been developed.</p><p> Moreover, if the technique of Non-Patent Document 1 is used, the character string of UTF8 can be appropriately converted into the character string of UTF16. However, it is not efficient to convert all the characters to be manipulated to UTF16. For example, when the character input in UTF8 is output as it is, it may be a redundant process of converting from UTF8 to UTF16 and then returning to UTF8 again. Therefore, the second task is to properly select the characters to be converted.</p><p> Further, according to the program example of Non-Patent Document 1, in order to convert each character of UTF8 into each character of UTF16, different processing is required depending on the conditions such as the data size of UTF8. Therefore, a conditional branch instruction for determining conditions such as UTF8 data size is generated. In recent central processing units, the instruction pipeline may be flushed by a conditional branch instruction and the processing efficiency may deteriorate, which is not preferable. Therefore, the third issue is to reduce conditional branching as much as possible in the conversion process.</p><p> Therefore, an object of the present invention is to provide a compiler, a compiler program, a recording medium, a control method, and a central processing unit capable of solving the above problems. This purpose is achieved by a combination of the features described in the independent section in the claims. Dependent terms also define further advantageous embodiments of the present invention.</p>
<p> In order to solve the above problems, in the first embodiment of the present invention, the first embodiment is a compiler that optimizes the conversion of the character code system of the characters stored in the character variables in the target program to be optimized. Prior to each of a plurality of processes in which the character of the character variable written by the character code system is read and the character is used in the second character code system, the character is transferred from the first character code system to the second character code. For each conversion instruction generated by the conversion instruction generation unit that converts to the system and generates the conversion instruction to be stored in the character variable, and in all the execution paths executed prior to the conversion instruction, A compiler having a conversion instruction removing unit that removes the conversion instruction when a character of the second character code system is stored in a character variable, a control method of the compiler, a program that makes a computer function as the compiler, and the program. Provide a recording medium on which the above is recorded.</p><p> Further, in the second embodiment of the present invention, it is a compiler that optimizes the conversion of the character code system of the characters stored in the character variables in the target program to be optimized, and is written by the first character code system. As each of the plurality of processes that read the character of the character variable and use the character in the second character code system, the process is executed when the character stored in the character variable is the second character code system. , The character operation processing generator that generates an exception instruction string that raises an exception when the character is the first character code system, and the character that is executed when an exception occurs and is stored in the character variable. A compiler having an exception handler generator that converts one character code system to a second character code system, stores it in the character variable, and generates an exception handler that returns to the process that caused the exception. Provided are a control method, a program for operating a computer as the compiler, and a recording medium on which the program is recorded.</p><p> Further, in the third embodiment of the present invention, a compiler that optimizes the conversion process for converting the characters stored in the character variables from the first character code system to the second character code system in the target program to be optimized. The acquisition process generator that generates the instruction sequence of the process that acquires the character code of the character to be converted, and the plurality of conversion processes that should be selected and executed according to the range of the character code values. A conversion detection processing generator that generates an instruction sequence that executes each and a detection process that detects a range of character code values in parallel, and a processing result of any one of a plurality of conversion processes is detected. Provided are a compiler having a selection process generator that generates an instruction sequence to be selected and output based on the detection result of the process, a control method of the compiler, a program that makes a computer function as the compiler, and a recording medium that records the program. To do.</p><p> The outline of the above invention does not enumerate all the necessary features of the present invention, and subcombinations of these feature groups can also be inventions.</p>
<p> According to the present invention, the conversion of the character code system can be optimized.</p>
Hereinafter, the present invention will be described through embodiments of the invention, but the following embodiments do not limit the invention within the scope of the claims, and all combinations of features described in the embodiments are included. It is not always essential for the means of solving the invention.
FIG. 1 is a block diagram of compiler 10 (Example 1). Compiler 10 aims to optimize the conversion of the character code system of characters stored in character variables. In particular, Compiler 10 converts from UTF8, which is the character code system of characters in XML documents, to UTF16, which is the character code system used when processing described as a Java (registered trademark) program manipulates character strings. The purpose is to optimize the instructions.
The compiler 10 includes a character code system determination unit 100, a conversion instruction generation unit 105, a conversion instruction removal unit 110, a method information storage unit 120, a method recursive determination unit 130, a character operation processing generation unit 140, and an exception handler. It includes a generation unit 145 and an output processing instruction generation unit 150. The character code system determination unit 100 receives the target program to be optimized written in the Java (registered trademark) language for each method to be optimized, for example. Then, the character code system determination unit 100 determines whether or not the probability that the character input to the method is UTF16 is higher than the predetermined reference probability. When the probability of being UTF16 is higher than the reference probability, the character code system determination unit 100 sends the method to the character operation processing generation unit 140. On the other hand, when the probability is lower than the reference probability, the character code system determination unit 100 sends the method to the conversion instruction generation unit 105.
The conversion instruction generator 105 reads the character of the character variable written by UTF8, converts the character from UTF8 to UTF16, and stores it in the character variable prior to each of a plurality of processes using the character in UTF16. Generate a conversion instruction to do. When the conversion instruction removing unit 110 stores UTF16 characters in the character variable in all the execution paths executed prior to the conversion instruction for each conversion instruction generated by the conversion instruction generation unit 105, the conversion instruction removing unit 110 performs. Remove the conversion instruction.
Specifically, the conversion instruction removal unit 110 determines that UTF16 characters are stored in the character variable in the execution path when any of the following conditions is satisfied for each execution path. 1. When the conversion instruction is executed in the execution path 2. When the constructor that secures the storage area for UTF16 characters for the character variable and sets the character code system to UTF16 is executed in the execution path. 3. When a method that stores UTF16 characters as a return value is executed in the execution path For example, the method information storage unit 120 stores the identification information of a method that returns UTF16 characters as a return value. The conversion instruction removal unit 110 may determine whether or not the return value of the method corresponding to the identification information stored in the method information storage unit 120 is stored in the character variable in each execution path.
Further, the conversion instruction removal unit 110 does not store UTF16 characters in the character variable in any execution path executed prior to the conversion instruction for each conversion instruction generated by the conversion instruction generation unit 105. To remove the partial redundancy of the conversion instruction by generating a new conversion instruction in the execution path and removing the conversion instruction.
When the method to be optimized uses UTF16 as a return value, the method recursive judgment unit 130 stores the identification information of the method in the method information storage unit 120. Specifically, the method recursive judgment unit 130 stores the characters whose optimization target method is stored in the character variable set in UTF16 by the constructor, the characters already converted to UTF16 by the conversion instruction, and the method information. When any of the return value of the method corresponding to the identification information stored in the unit 120 is further used as the return value, the identification information of the method is stored in the method information storage unit 120.
The character operation processing generation unit 140 reads the character of the character variable written by UTF8 and uses the character in UTF16 as each of the plurality of processing, when the character stored in the character variable is UTF16. To generate a sequence of instructions with exceptions that raises an exception if the character is UTF8. Then, the exception handler generator 145 is executed when an exception occurs, converts the character stored in the character variable from UTF8 to UTF16, stores it in the character variable, and returns to the process that caused the exception. Generate an exception handler to make it.
The output processing command generator 150 outputs the character stored in the character variable as the process of outputting the character when the character stored in the character variable is UTF8, and the character stored in the character variable is output. If it is UTF16, return the character to UTF8 and generate an instruction sequence to output. Then, the output processing instruction generation unit 150 outputs the generated instruction sequence together with each of the other methods optimized by the above processing as a result program which is an optimized program.
FIG. 2 is an example of the data structure of the character variable 20 (Example 1). The character variable 20 is a code system information 210 indicating whether the character stored in the character variable 20 is UTF8 or UTF16, a first character storage area 225 for storing UTF8 characters, and a th-order for storing UTF16 characters. Manages the two-character storage area 235. Specifically, the character variable 20 includes the code system information 210, the first pointer 220 indicating the address of the first character storage area 225, the second pointer 230 indicating the address of the second character storage area 235, and the first character variable 20. It has data length information 240 indicating the length of character data stored in the character storage area 225 or the second character storage area 235.
In the target program, the constructor that generates the character variable 20 secures the first character storage area 225 and / or the second character storage area 235, which are storage areas for storing characters, and sets the character code system for the character variable. To do. For example, the constructor generates code system information 210 indicating that the character of the character variable is UTF8 or UTF16. As a result, the conversion instruction removal unit 110 of the compiler 10 can determine whether the constructor of the character variable sets the character variable to UTF8 or UTF16.
As a concrete implementation method of the character variable 20 according to this figure, the definition of the Java (registered trademark) language standard String object is described in the JVM (Java Virtual Machine), which is the execution environment of the Java (registered trademark) program. Replace with the definition of the String object related to the figure. Then, define UTF8String, which is a class that manipulates UTF8 characters. The UTF8String class contains, for example, the following methods. Boolean isUTF8 (String str); API to read code system information 210. Byte [] getRawBytes (String str); API to read the first character storage area 225 or the second character storage area 235. Int getRawLength (String str); API to read data length information 240.
Also, in the JVM, define create, which is a method to create a UTF8 String object, as follows. String create (byte [] date, int start, int length) This method extracts the part specified by the 2nd and 3rd arguments from the character data acquired as the 1st argument and has UTF8 characters. Create as an object.
FIG. 3 shows a processing flow in which the compiler 10 optimizes the conversion processing (Example 1). The character code system determination unit 100 determines whether or not the probability that the character input for each method is UTF8 is higher than the reference probability (S300). For example, the character code system determination unit 100 determines that the probability that the character input to the method is UTF16 is lower than the reference probability for the method in which the character of UTF8 is predetermined to be input. More specifically, since it is predetermined that each of the following method groups inputs an XML document, the character code system determination unit 100 determines that UTF8 is input to these method groups.
The following methods in the SAX application program, or methods that are further called from these methods. (Including subclasses) org.xml.sax.Parser.parse (..) org.xml.sax.XMLReader.parse (..) javax.xml.parsers.SAXParser.parse (..) The following in the DOM application program Or a method that is further called from these methods. (Including subclasses) Javax.xml.parsers.DocumentBuilder.parse (..) Org.apache.xml.serialize.XMLserializer.serialize (..) Javax.xml.parsers.DocumentBuilderFactory.newInstance (..) Pull Parser Application Program The following methods in, or methods that are further called from these methods. (Including subclasses) Org.xmlpull.v1.XmlPullParser.next () Org.xmlpull.v1.XmlPullParser.nextToken () Org.xmlpull.v1.XmlPullParser.nextTag () Org.xmlpull.v1.XmlPullParser.nextText ()
When the probability of being UTF16 is lower than the reference probability, the conversion instruction generator 105 reads the character of the character variable written by UTF8 and prints the character prior to each of the plurality of processes that use the character in UTF16. Generate a conversion instruction that converts from UTF8 to UTF16 and stores it in that character variable (S310). Here, the above-mentioned plurality of processes are, for example, a group of methods in a predetermined class having a UTF16 character variable as an instance. As an example, in the Java (registered trademark) language, the above-mentioned multiple processes are all methods in the java / lang / String class except for methods that do not need to manipulate the character contents in UTF16 format, such as character output methods. To say. That is, the conversion instruction generation unit 105 generates a conversion instruction at a position executed prior to each of the call processing of these methods, for example, a position executed immediately before the call processing.
When the conversion instruction removal unit 110 stores UTF16 characters in the character variable in all the execution paths executed prior to the conversion instruction for each conversion instruction generated by the conversion instruction generation unit 105 ( Condition A), and remove the conversion instruction (S320). Further, when the conversion instruction removal unit 110 does not store characters in UTF16 in the character variable in any execution path executed prior to the conversion instruction for each conversion instruction generated by the conversion instruction generation unit 105. (Condition A is not satisfied), a new conversion instruction is generated in the execution path, and the conversion instruction is removed to remove the partial redundancy of the conversion instruction. In this case, it is a condition that UTF16 characters are stored in the character variable at any execution position of the execution path in all execution paths starting from the execution position (condition B).
As a method for realizing this process, for example, the conversion instruction removing unit 110 may remove partial redundancy by the method exemplified in Non-Patent Documents 3 and 4. Here, an example of a process for removing partial redundancy of a conversion instruction by applying the method described in Non-Patent Document 3 will be described. First, the conversion instruction removal unit 110 calculates the following predicates for each basic block in the method to be optimized. However, the processing part of the conversion instruction that checks whether UTF8 characters are stored in the character variable to be converted is called the format check process.
TRANSP (bb): Whether or not the contents of the character variable to be processed by the format check process are destroyed in the basic block bb N-COMP (bb): The format check process exists in the basic block bb and its format check Whether it is safe to place the process at the entrance of the basic block bb (that is, whether the meaning of the target program does not change) X-COMP (bb): There is a format check process in the basic block bb, and its Whether it is safe to place the format check process at the exit of the basic block (that is, whether the meaning of the target program does not change)
Next, the conversion instruction removal unit 110 uses Busy Code Motion to obtain a portion in which the format check process executed first among the format check processes executed twice or more is arranged in the same character variable. Specifically, first, the conversion command removal unit 110 obtains the upper placement safety and the lower placement safety of the format check process.
Upper placement at the entrance of each basic block Safety:<maths num="1"><img file="JP2005293386A_D0001.tif" /></maths> Upper placement safety at the exit of each basic block:<maths num="2"><img file="JP2005293386A_D0002.tif" /></maths> Downward placement at the entrance of each basic block Safety:<maths num="3"><img file="JP2005293386A_D0003.tif" /></maths> Lower placement at the exit for each basic block Safety:<maths num="4"><img file="JP2005293386A_D0004.tif" /></maths>
Next, the conversion command removal unit 110 finds the fixed point solutions of ND-SAFE (), XD-SAFE (), NU-SAFE (), and XU-SAFE () by solving the above equations. Next, the conversion instruction removal unit 110 uses the floating-point solution to determine the position where the format check process should be placed as follows. At this time, since all the format check processes originally arranged are redundant in principle, the conversion instruction removal unit 110 deletes these format check processes.
Basic block where format check processing should be placed at the entrance = Basic block where the following formula (5) is true.<maths num="5"><img file="JP2005293386A_D0005.tif" /></maths> Basic block where format check processing should be placed at the exit = Basic block where the following formula (6) is true.<maths num="6"><img file="JP2005293386A_D0006.tif" /></maths>
Further, the conversion instruction removal unit 110 may move the format check process unnecessarily arranged upward by Lazy Code Motion downward as long as the effect of redundancy removal is not lost. Since this method is an application of the method described in Non-Patent Document 3 in the same manner as in the case of Busy Code Motion described above, the description thereof will be omitted.
As yet another example, the conversion instruction removing unit 110 may modify the target program so that the control flow is changed in order to remove more conversion instructions. For example, the conversion instruction removal unit 110 does not store the UTF16 character in the character variable in a part of the execution paths leading to the conversion instruction for each conversion instruction (condition A is not satisfied), and executes any of the execution paths. Even at the position, if the character of UTF16 is not stored in the character variable in all the execution paths starting from the execution position (condition B is not satisfied), the following processing is performed.
First, the conversion instruction removal unit 110 finds a confluence point between an execution path in which UTF16 characters are stored and an execution path in which UTF16 characters are not stored. Next, the conversion instruction removing unit 110 finds a branch point between the execution path in which the UTF16 character is stored and the execution path in which the UTF16 character is not stored in the path from the confluence point to the conversion instruction. Next, the conversion command removal unit 110 copies the path between the confluence point and the branch point. Next, the conversion command removing unit 110 connects each of the paths merging at the merging point to each of the copy source path and the copy destination path. As a result, since the condition B is satisfied, the conversion instruction removing unit 110 can remove the partial redundancy of the conversion instruction by, for example, the above method applying busy code motion.
Subsequently, the method recursive determination unit 130 stores the identification information of the method in the method information storage unit 120 when the method to be optimized uses UTF16 as the return value (S330). Specifically, the method recursive judgment unit 130 stores the characters whose optimization target method is stored in the character variable set in UTF16 by the constructor, the characters already converted to UTF16 by the conversion instruction, and the method information. When any of the return value of the method corresponding to the identification information stored in the unit 120 is further used as the return value, the identification information of the method is stored in the method information storage unit 120.
On the other hand, when the probability of being UTF16 is higher than the reference probability, the character operation processing generation unit 140 reads the character of the character variable written by UTF8 and uses that character in UTF16 as each of the plurality of processes that use the character. Executes the process when the character stored in the variable is UTF16, and generates an instruction sequence with exception that raises an exception when the character is UTF8 (S340).
Specifically, the character operation processing generation unit 140 generates an instruction sequence for reading the second character storage area 235 in which UTF16 characters are stored as an instruction sequence with exceptions. In this case, the character operation processing generation unit 140 invalidates the second pointer 230 indicating the address of the second character storage area 235, which is an instruction sequence executed when the UTF8 character is stored in the character variable. By setting to the address, an instruction sequence that raises an exception is generated in the instruction sequence with exception. Further, the character operation processing generation unit 140 generates an instruction sequence that is executed when a UTF16 character is stored in a character variable and sets the second pointer 230 to a valid address. In the Java (registered trademark) language, for example, the methods in the java / lang / String class for manipulating character variables are rewritten in advance, and the compiler 10 compiles these rewritten methods to perform the above. An instruction string may be generated.
Then, the exception handler generator 145 is executed when an exception occurs, converts the character stored in the character variable from UTF8 to UTF16, stores it in the character variable, and returns to the process that caused the exception. Generate an exception handler to make it (S350). If the exception handlers already generated for other methods can be used, the exception handler generation unit 145 does not have to generate the exception handler again in S350.
Next, when the optimization target is a method that performs output processing that outputs the characters stored in the character variable, the output processing instruction generation unit 150 determines that the characters stored in the character variable are UTF8. Outputs that character, and if the character stored in that character variable is UTF16, returns that character to UTF8 and generates an instruction sequence to output (S360).
FIG. 4 is an example of a program showing the operation of the instruction sequence generated by the output processing instruction generation unit 150 as output processing (Example 1). This program is written using the API described in Figure 2. Specifically, the output process outputs the UTF8 character as it is when the character code system to be output is UTF8 and the output character is UTF8. On the other hand, if the character code system to be output is UTF8 and the character to be output is UTF16, UTF16 is returned to UTF8 and output.
As a method for the output processing instruction generation unit 150 to generate the instruction sequence according to this figure, specifically, among the library programs compiled by the compiler 10 together with the target program, a method for performing output processing (for example, OutputStreamWriter.write). Is changed in advance. As a result, the output processing instruction generation unit 150 can generate an appropriate instruction sequence that realizes the output processing according to the present figure by compiling the modified library program.
FIG. 5 shows an example of a program in which the compiler 10 according to this embodiment is applied to the SAX input library (Example 1). (a) shows the interface of the SAX (The Simple API for XML) handler, and (b) shows an example of the SAX parser in this embodiment.
The SAX parser reads the startElement of the object that implements the method defined in the ContentHandler interface for each element of the input XML document. Then, the SAX parser in this embodiment creates a String object having UTF8 character data when the character code system of the input XML document is UTF8, and calls the method startElement with this generated object as an argument. ..
(c) shows a fragment of the application program using the SAX parser in this example. The argument qName of StartElement is passed to OutputStreamWriter.writer () via PrintWriter.print without being converted as UTF8. As a result, UTF8 characters are output as UTF8 without being converted to UTF16. In this way, the conversion process between UTF16 and UTF8 can be optimized without changing the application program that uses the SAX parser.
FIG. 6 shows an example of a program in which the compiler 10 according to this embodiment is applied to the DOM input library (Example 1). In the DOM (Document Object Model), the application program acquires a character string such as an element name by calling a method defined in the Node interface. The DOM input library according to this embodiment creates a UTF8 String object as a component of the object that implements the Node interface when creating a Document object. This figure is an implementation example of a library that creates an object of a class that implements an Element interface that inherits a Node interface.
As described above, according to the first embodiment shown in FIGS. 1 to 6, the conversion process of the character code system from UTF8 to UTF16 can be reduced to the minimum necessary by the cooperation of the compiler program and the library program. As a result, for example, when the input UTF8 is output as it is, unnecessary conversion processing can be omitted. In addition, the method for manipulating characters in UTF16 can omit the format check process for checking whether the input characters have already been converted to UTF16.
Furthermore, in the method where the probability that UTF16 characters are input is higher, the conversion process can be generated only in the exception handler without generating it in the normal execution path, which is more efficient. Further, the above processing is realized by the library program and / or the compiler program without changing the application program. Therefore, the existing application program can be effectively used.
FIG. 7 is a block diagram of the compiler 70 and the central processing unit 80 (Example 2). The purpose of the compiler 70 is to optimize the conversion process of converting the characters stored in the character variables from UTF8 to UTF16 in the target program to be optimized, or the conversion process of converting from UTF16 to UTF8. The compiler 70 includes an acquisition process generation unit 700, a conversion detection process generation unit 710, and a selection process generation unit 720.
When the target program is input, the acquisition process generation unit 700 first generates an instruction sequence for processing for acquiring the character code of the character to be converted. Hereinafter, the processing by this instruction sequence is referred to as an acquisition processing. Next, the conversion detection process generation unit 710 includes each of a plurality of conversion processes to be executed by selecting one according to the character code value range, and a detection process that detects a value range in the character code. Generates a sequence of instructions that executes in parallel. Hereinafter, the processing by this instruction sequence is referred to as conversion detection processing.
The selection process generation unit 720 generates a selection process instruction sequence that selects and outputs the process result of any of the above-mentioned conversion processes based on the detection result of the detection process. Then, the central processing unit 80 executes the output instruction sequence, generates a character code converted to UTF8, and outputs the character code.
FIG. 8 shows a processing flow in which the compiler 10 optimizes the conversion processing (Example 2). The compiler 10 repeats the following processing for each conversion processing for converting UTF8 characters to UTF16 characters. First, the acquisition process generation unit 700 generates an instruction sequence for acquisition processing (S800). For example, the acquisition process generation unit 700 may generate an instruction sequence for acquiring unit data of a predetermined size including a character code. More specifically, when the character code has a variable length of 8 to 32 bits, 32-bit or 64-bit unit data including the character code may be acquired. Next, the conversion detection processing generation unit 710 generates an instruction sequence for conversion detection processing (S810). Then, the selection process generation unit 720 generates an instruction sequence for the selection process (S820).
FIG. 9 shows a comparison of UTF16 and UTF8 character code systems (Example 2). UTF8, which is an example of the first character code system according to the present invention, represents a character code from 0 to 7F in hexadecimal as the lower 7 bits of a character code having a data size of 1 byte. On the other hand, UTF16, which is an example of the second character code system according to the present invention, represents a character code from 0 to 7F in hexadecimal as the lower 7 bits of a character code having a data size of 2 bytes. In this way, UTF8 has better memory usage efficiency than UTF16 for character codes from 0 to 7F in hexadecimal.
UTF8 divides the character code from 80 to 7FF in hexadecimal into the lower 5 bits of the 1st byte and the lower 6 bits of the 2nd byte in the character code of the data size of 2 bytes. The other bits are predetermined control data indicating the data size of the character code, the type of character, and the like. UTF16, on the other hand, represents this character code as the lower 11 bits of a character code with a 2-byte data size.
UTF8 divides the character code from 800 to 0FFF in hexadecimal into the lower 5 bits of the 2nd byte and the lower 6 bits of the 3rd byte in the character code of the data size of 3 bytes. The data of the first byte and the other bits of the second and third bytes are predetermined control data indicating the data size of the character code, the type of the character, and the like. UTF16, on the other hand, represents this character code as the lower 12 bits of a character code with a 2-byte data size.
UTF8 is a hexadecimal character code from 1000 to DF77 or E000 to FFFF in a character code with a data size of 3 bytes, the lower 4 bits of the 1st byte, the lower 6 bits of the 2nd byte, and 3 bytes. It is divided into the lower 6 bits of the eye. On the other hand, UTF16 expresses this character code as 16 bits in a character code having a data size of 2 bytes.
UTF8 divides the character code from 10000 to 3FFFF in hexadecimal into the lower 6 bits of the 2nd byte, the lower 6 bits of the 3rd byte, and the lower 6 bits of the 4th byte in the character code of the data size of 4 bytes. And express. The 1st byte and the other bits of the 2nd to 4th bytes are predetermined control data indicating the data size of the character code, the type of the character, and the like. UTF16, on the other hand, represents this character code as the lower 18 bits of the 20-bit character code.
UTF8 uses a hexadecimal character code from 40000 to FFFFF as the lower 2 bits of the 1st byte, the lower 4 bits of the 2nd byte, the lower 6 bits of the 3rd byte, and 4 in the character code of the data size of 4 bytes. It is divided into the lower 6 bits of the byte and represented. The other bits are predetermined control data indicating the data size of the character code, the type of character, and the like. On the other hand, UTF16 expresses this character code as 20 bits in the 20-bit character code.
UTF8 divides the character code from 100000 to 10FFFF in hexadecimal into the lower 4 bits of the 2nd byte, the lower 6 bits of the 3rd byte, and the lower 6 bits of the 4th byte in the character code of the data size of 4 bytes. And express. The data of the first byte and other bits are predetermined control data indicating the data size of the character code, the type of character, and the like. On the other hand, UTF16 expresses this character code as 16 bits in the 24-bit character code.
As described above, in UTF8, the data size of the character code differs depending on the value of the character code to be stored or the type of the character. On the other hand, in UTF16, the data size is constant at 2 bytes except for characters exceeding 10000 in hexadecimal. As a result, UTF8 has higher memory usage efficiency when characters with small data size are consecutive. On the other hand, UTF16 is less likely to change its data size depending on the type of character than UTF8. For this reason, processing that targets UTF16 is often faster than processing that targets UTF8.
Then, in order to convert UTF8 to UTF16, it is necessary to read the control data in the UTF8 character code and select an appropriate conversion process according to the control data. Subsequently, FIG. 10 shows an example of this process of converting UTF16 to UTF8. In the following description of this figure, the constants used in the calculation are shown in hexadecimal unless otherwise specified.
The computer's central processing unit, in collaboration with memory and other hardware, first reads the first 8 bits of the character code into the variable b0 (S1000). In the following description, "the central processing unit of the computer cooperates with the memory and other hardware" is simply abbreviated as "the computer is". Then, when the result of masking the read character code with 80 is 0 (S1005: YES), the computer calculates the logical sum of the variable b0 and the binary number 00000000, that is, b0 itself (S1010), and S1140. Move the process to.
On the other hand, when the result of masking the variable b0 with 80 is a value other than 0 (S1005: NO), the computer reads it as b0 when the result of masking the variable b0 with e0 is c0 (S1015: YES). The next 8 bits consecutive to the character code are read into the variable b1 (S1020). Then, if the result of masking the variable b1 with c0 does not become 80 (S1025: NO), the computer notifies the user or the like that an error has occurred in the conversion process (S1030). For example, this is the case when the input characters are not UTF8 compliant.
On the other hand, if the result of masking the variable b1 with c0 is 80 (S1025: YES), the computer shifts the variable b0 to the left by 6 bits and masks it with 7c0, and the variable b1 is masked with 3f. The logical sum is stored in the variable c and output in S1140 (S1035). If the result of masking the variable b0 with e0 is not c0 (S1015: NO), the computer determines whether the result of masking the variable b0 with f0 is e0 (S1040). If e0 (S1040: YES), the computer reads the next 8 bits following the character code read as b0 into the variable b1 (S1045).
Then, if the result of masking the variable b1 with c0 does not become 80 (S1050: NO), the computer notifies the user or the like that an error has occurred in the conversion process (S1055). On the other hand, when the result of masking the variable b1 with c0 is 80 (S1050: YES), the computer reads the next 8 bits consecutive to the character code read as b1 into the variable b2 (S1060).
If the result of masking the variable b2 with c0 is not 80 (S1065: NO), the computer notifies the user etc. that an error has occurred in the conversion process (S1070). If the result of masking the variable b2 with c0 is 80 (S1065: YES), the computer left-shifts the variable b0 12 bits and masks it with f000, and shifts the variable b1 6 bits left with fc0. The logical sum of the masked value and the value obtained by masking the variable b2 with 3f is stored in the variable c (S1075).
The computer determines if the result of masking the variable b0 with f0 is not e0 (S1040: NO) and the result of masking the variable b0 with f8 is f0 (S1080). If it is not f0 (S1080: NO), the computer notifies the user or the like that an error has occurred in the conversion process (S1085). If the result of masking the variable b0 with f8 is f0 (S1080: YES), the computer reads the next 8 bits consecutive to the character code read as b0 into the variable b1 (S1090).
If the result of masking the variable b1 with c0 is not 80 (S1095: NO), the computer notifies the user or the like that an error has occurred in the conversion process (S1100). On the other hand, when the result of masking the variable b1 with c0 is 80 (S1095: YES), the computer reads the next 8 bits consecutive to the character code read as b1 into the variable b2 (S1105).
Subsequently, when the result of masking the variable b2 with c0 is not 80 (S1110: NO), the computer notifies the user or the like that an error has occurred in the conversion process (S1115). If the result of masking the variable b2 with c0 is 80 (S1110: YES), the computer reads the next 8 bits consecutive to the character code read as b2 into the variable b3 (S1120).
Then, if the result of masking the variable b3 with c0 is not 80 (S1125: NO), the computer notifies the user that an error has occurred in the conversion process (S1130). If the result of masking the variable b3 with c0 is 80 (S1125: YES), the computer does the following calculation (S1135). First, the logical sum of the value obtained by shifting the variable b0 to the left by 2 bits and masking with 1c and the value obtained by shifting the variable b1 to the right by 4 bits and masking with 3 are obtained, and the value subtracted by 1 is stored in the variable w. ..
Next, the computer shifts the constant d800, the variable w left-shifted 6 bits and masked with 3c0, the variable b1 left-shifted 2 bits and masked with 3c, and the variable b2 right-shifted 4 bits. The logical sum with the value masked in 3 is stored in the variable c_high. Further, the computer stores the logical sum of the constant dc00, the value of the variable b2 left-shifted by 6 bits and masked by 3c0, and the value of the variable b3 masked by 3f in the variable c_low.
Finally, the computer outputs the value stored in the variable c or the value in which c_high is the high-order bit and c_low is the low-order bit as the UTF16 character code (S1140).
In this way, in order to convert UTF8 to UTF16, it is necessary to read the control data in the UTF8 character code and select an appropriate conversion process according to the control data. That is, for example, the computer needs to determine what kind of value the control data is and conditionally branch. When conditional branching is frequently performed as shown in this figure, it is difficult for the computer to anticipate and execute the instruction to be executed after the instruction being executed because the instruction group to be executed next cannot be accurately predicted. .. Therefore, the degree of parallelism of the instructions that can be processed by the central processing unit of the computer cannot be effectively used, and the execution efficiency of the target program deteriorates. Further, while a recent central processing unit can store 64-bit data in a register, in the processing of this example, only 8-bit data is often stored in a register, which is not efficient. On the other hand, the compiler 10 in this embodiment can reduce the number of conditional branches and improve the execution efficiency of the conversion instruction. Hereinafter, an example of the conversion instruction generated by the compiler 10 in this embodiment will be described.
FIG. 11 shows a processing flow diagram of the conversion process generated by the compiler 70 in this embodiment (Example 2). Also in this figure, unless otherwise specified, the constants to be calculated are shown in hexadecimal. The computer first reads 32-bit data and stores it in the variable w as an acquisition process (S1145). Then, the computer performs the eight processes shown in S1150 to S1190 in parallel as the conversion detection process.
Specifically, as the first conversion process, the computer shifts the variable w to the right by 24 bits and stores the value masked by 7f in the variable c0 (S1150). In addition, as the second conversion process, the computer performs the logical sum of the value obtained by shifting the variable w to the right by 18 bits and masking it with 7c0 and the value obtained by shifting the variable w to the right by 16 bits and masking it with 3f. Store in (S1155). In addition, as the third conversion process, the computer shifts the variable w to the right by 12 bits and masks it with f000, shifts the variable w to the right by 10 bits and masks it with fc0, and shifts the variable w to the right by 8 bits. The logical sum of the shifted value and the value masked by 3f is stored in the variable c2 (S1160).
Further, as the fourth conversion process, the computer first obtains the logical sum of the value obtained by right-shifting the variable w by 22 bits and masking with 1c, and the value obtained by shifting the variable w to the right by 20 bits and masking with 3. , The value subtracted by 1 is stored in the variable x (S1165). Next, as the fourth conversion process, the computer uses the constant d800, the value of the variable x left-shifted by 6 bits and masked with 3c0, the variable w with the variable w shifted 14 bits to the right and masked with 3c, and the variable. Shift w to the right by 12 bits and store the logical sum with the value masked by 3 in the variable c_high. Next, the computer stores the logical sum of the constant dc00, the value of the variable w shifted to the right by 2 bits and masked with 3c0, and the value of the variable w masked with 3f in the variable c_low.
In addition, as a detection process, the computer stores the true in the variable f0 when the value obtained by masking the variable w with 80000000 is 0 (S1170). Also, the computer stores the true in the variable f1 when the value of the variable w masked with e0c00000 is c0800000 (S1180). The computer also stores the true in the variable f2 when the value of the variable w masked with f0c0c000 is e0808000 (S1185). The computer also stores the true in the variable f3 when the value of the variable w masked by f8c0c0c0 is f0808080 (S1190).
The computer may execute the first to fourth conversion processes and the detection process in parallel with each other, or at least two of these processes may be executed in parallel. Subsequently, the computer executes the processes S1195 to S1220 as the selection process.
Specifically, the computer first stores the variable c0 in the variable c if the variable f0 is true (S1195). Also, if the variable f1 is true, the variable c1 is stored in the variable c (S1200). Also, if the variable f2 is true, the variable c2 is stored in the variable c (S1205). In these cases, the computer outputs the contents of the variable c as the conversion result (S1220). If the variable f3 is true, the value with c_high as the high-order bit and c_low as the low-order bit is output as the UTF16 character code (S1220). If neither f0 to f3 is true, an error is notified in the conversion process (S1215).
As described above, according to the instruction sequence generated by the compiler 70 in the present embodiment, a plurality of conversion processes to be executed by selecting one of them can be speculatively executed prior to the conditional determination. Then, the conversion process and the detection process can be executed in parallel with each other without processing the conditional branch. As a result, the character code system can be efficiently converted by effectively utilizing the performance of the central processing unit capable of executing a plurality of instructions in parallel.
Hereinafter, as a method for executing this conversion process more efficiently, a process of collectively converting a plurality of character codes having the same data size arranged continuously will be described. First, in order to execute this process more efficiently, it is preferable that the central processing unit 80 has an instruction group suitable for the conversion process. Hereinafter, the instruction group included in the central processing unit 80 will be described with reference to FIGS. 12 and 13. Then, in FIG. 14, an example in which the compiler 70 generates an instruction sequence including these instruction groups will be described.
FIG. 12 shows an example of a group of instructions for converting UTF8 to UTF16, which is possessed by the central processing unit 80 (Example 2). (a) shows the details of the instruction UTF81toUTF16. The conversion detection processing generation unit 710, as the instruction UTF81toUTF16, performs data in parallel with the conversion processing to be executed when a plurality of character codes having a data size of 8 bits are continuously arranged in the 64-bit register rs. Generates an instruction to execute a detection process that detects whether or not a plurality of character codes having a size of 8 bits are arranged consecutively.
More specifically, by the instruction UTF81toUTF16, the central processing unit 80 is set to the 1st to 7th bits, the 9th to 15th bits, the 17th to 23rd bits, and the 25th to the 25th bits of the register rs. The 31 bits are copied to the 9th to 15th bits, the 25th to 31st bits, the 41st to 47th bits, and the 57th to 63rd bits of the register rt, respectively.
Furthermore, according to the instruction UTF81toUTF16, the central processing unit 80 determines that the 0th bit of the register rs is 0, the 8th bit is 0, the 16th bit is 0, and the 24th bit is 0. 1 is stored in the register cr to indicate the detection result that a plurality of character codes having a data size of 8 bits are continuously arranged.
(b) shows the details of the instructions UTF82toUTF16. The conversion detection processing generation unit 710, as the instruction UTF82toUTF16, performs data in parallel with the conversion processing to be executed when a plurality of character codes having a data size of 16 bits are continuously arranged in the 64-bit register rs. Generates an instruction to execute a detection process that detects whether or not a plurality of character codes having a size of 16 bits are arranged consecutively.
More specifically, by the instruction UTF82toUTF16, the central processing unit 80 combines the 3rd to 7th bits of the register rs and the 10th to 15th bits to change the 5th to 15th bits of the register rt. make a copy. Further, the central processing unit 80 combines the 19th to 23rd bits and the 26th to 31st bits of the register rs and copies them from the 21st bit to the 31st bit of the register rt.
Further, the central processing unit 80 combines the 35th to 39th bits and the 42nd to 47th bits of the register rs and copies them from the 37th bit to the 47th bit of the register rt. Further, the central processing unit 80 combines the 51st to 55th bits and the 58th to 63rd bits of the register rs and copies them from the 53rd bit to the 63rd bit of the register rt.
Further, by the instruction UTF82toUTF16, the central processing unit 80 has 2 bits each of the 0th to 2nd bits, the 16th to 18th bits, the 32nd to 34th bits, and the 48th to 50th bits of the register rs. It is determined whether or not the condition of the base number 110 is satisfied. Further, the central processing unit 80 determines whether the 8th and 9th bits, the 24th and 25th bits, the 40th and 41st bits, and the 56th and 57th bits satisfy the condition of the binary number 10. ..
Further, the central processing unit 80 determines whether the third bit of the register rs is not 0, or whether the condition that the fourth to seventh bits are neither 000 nor 001 is satisfied. Further, the central processing unit 80 determines whether the 19th bit of the register rs is not 0, or whether the condition that the 20th to 23rd bits are neither 000 nor 001 is satisfied. Further, the central processing unit 80 determines whether the 35th bit of the register rs is not 0, or whether the condition that the 36th to 39th bits are neither 000 nor 001 is satisfied. Further, the central processing unit 80 determines whether the 51st bit of the register rs is not 0, or whether the condition that the 52nd to 55th bits are neither 000 nor 001 is satisfied. When all the above conditions are satisfied, the central processing unit 80 stores 1 in the register cr.
(c) shows the details of the instruction UTF83toUTF16. The conversion detection processing generation unit 710, as the instruction UTF83toUTF16, performs data in parallel with the conversion processing to be executed when a plurality of character codes having a data size of 24 bits are continuously arranged in the 64-bit register rs. Generates an instruction to execute a detection process that detects whether or not a plurality of character codes having a size of 24 bits are arranged consecutively.
More specifically, by the instruction UTF83toUTF16, the central processing unit 80 combines the 4th to 7th bits, the 10th to 15th bits, and the 18th to 23rd bits of the register rs to the 32nd of the register rt. Copy from bit to 47th bit. Further, the central processing unit 80 combines the 28th to 31st bits, the 34th to 39th bits, and the 42nd to 47th bits of the register rs and copies them from the 48th bit to the 63rd bit of the register rt. To do.
Further, by the instruction UTF83toUTF16, the central processing unit 80 determines whether or not each of the 0th to 3rd bits and the 24th to 27th bits of the register rs satisfies the condition of the binary number 1110. Furthermore, according to the instruction UTF83toUTF16, the central processing unit 80 has the 8th and 9th bits, the 16th and 17th bits, the 32nd and 33rd bits of the register rs, and the 40th and 41st bits, respectively, in binary. Judge whether or not the condition of 10 is satisfied.
Further, the central processing unit 80 determines whether the fourth to seventh bits of the register rs are not the binary 1101 or the tenth bit is not 1. Further, the central processing unit 80 determines whether the 28th to 31st bits of the register rs satisfy the condition that the binary 1101 or the 34th bit is not 1. When all the above conditions are satisfied, the central processing unit 80 stores 1 in the register cr.
As described above, as shown in this figure, the conversion detection processing generation unit 710 has a plurality of character codes of the data size consecutively in the unit data for every 8-bit, 16-bit, and 24-bit character code data sizes. In parallel with the conversion process to be executed when the data is arranged, an instruction to execute the detection process to detect whether or not multiple character codes of the data size are continuously arranged in the unit data is generated. be able to.
Furthermore, when the processing of these instructions is realized by a logic circuit, a NAND gate of 2 to 3 stages is required. Therefore, even if the signal delay when the NAND gate is implemented is taken into consideration, these instructions are executed with the latency of one cycle and the throughput of one cycle, respectively, as in the case of a simple addition instruction or the like. That is, according to the instructions in this figure, the computer can convert a UTF8 character string of up to 4 characters into a UTF16 character string in one cycle.
FIG. 13 shows other instructions possessed by the central processing unit 80, and (a) shows details of instructions UTF16 to UTF81, which is one of a group of instructions for converting UTF16 to UTF8 (Example 2). The conversion detection process generation unit 710 uses the instruction UTF16toUTF81 in parallel with the conversion process to be executed when the character code to be converted to UTF8 with a data size of 8 bits is continuously arranged in the 64-bit register rs. , Generates an instruction to execute detection processing to detect whether or not the character code to be converted to UTF8 whose data size is 8 bits is continuously arranged.
More specifically, by the instruction UTF16toUTF81, the central processing unit 80 is set to the 9th to 15th bits, the 25th to 31st bits, the 41st bit to the 47th bit, and the 57th bit to the 57th bit of the register rt. The 63 bits are copied to the 33rd to 39th bits, the 41st to 47th bits, the 49th to 55th bits, and the 57th to 63rd bits of the register rt, respectively.
Further, by the instruction UTF16toUTF81, the central processing unit 80 has all 0s of the 0th to 8th bits, the 16th to 24th bits, the 32nd to 40th bits, and the 48th to 56th bits of the register rs. When the condition is satisfied, 1 is stored in the register cr.
(b) shows the details of the instruction UTF16 to UTF82, which is one of the instruction group for converting UTF16 to UTF8, which is possessed by the central processing unit 80. The conversion detection processing generation unit 710, as the instruction UTF16toUTF82, performs the conversion processing to be executed in parallel with the conversion processing to be executed when the UTF16 character code to be converted to 16-bit UTF8 is continuously arranged in the 64-bit register rs. Generates an instruction to execute detection processing to detect whether the UTF16 character code to be converted to 16-bit UTF8 is continuously arranged.
More specifically, according to the instruction UTF16toUTF82, the central processing unit 80 divides the 5th to 15th bits of the register rs into the 3rd to 7th bits and the 10th to 15th bits of the register rt. Copy to. Further, the central processing unit 80 divides the 21st bit to the 31st bit of the register rs and copies them from the 19th bit to the 23rd bit and the 26th bit to the 31st bit of the register rt.
Further, by the instruction UTF16toUTF82, the central processing unit 80 divides the 37th bit to the 47th bit of the register rs and copies them from the 35th bit to the 39th bit and the 42nd bit to the 47th bit of the register rt. .. Further, the central processing unit 80 divides the 53rd bit to the 63rd bit of the register rs and copies them from the 51st bit to the 55th bit and the 58th bit to the 63rd bit of the register rt.
Further, by the instruction UTF16toUTF82, the central processing apparatus 80 is in the register rs, in which the 0th to 4th bits, the 16th to 20th bits, the 32nd to 36th bits, and the 48th to 52nd bits are all all. Judge whether or not the condition of 0 is satisfied. Further, the central processing unit 80 satisfies the condition that none of the 5th to 7th bits, the 21st to 23rd bits, the 37th to 39th bits, and the 53rd to 55th bits are binary numbers of 000. Judge whether or not. When all the above conditions are satisfied, the central processing unit 80 stores 1 in the register cr.
(c) shows the details of the instruction UTF16 to UTF82, which is one of the instruction group for converting UTF16 to UTF8, which is possessed by the central processing unit 80. The conversion detection processing generation unit 710 uses the instruction UTF16toUTF83 to execute the conversion processing in parallel with the conversion processing to be executed when the UTF16 character code to be converted to 24-bit UTF8 is continuously arranged in the 64-bit register rs. Generates an instruction to execute detection processing to detect whether the UTF16 character code to be converted to 24-bit UTF8 is continuously arranged.
More specifically, with UTF16toUTF83, the central processor 80 divides the 32nd to 47th bits of the register rs into the 4th to 7th bits, the 10th to 15th bits, and the 17th bit of the register rt. Copy from 18th to 23rd bit. Further, the central processing unit 80 divides the 48th bit to the 63rd bit of the register rs into the 28th bit to the 31st bit, the 34th bit to the 39th bit, and the 42nd to 47th bit of the register rt. make a copy.
Further, according to the instruction UTF16toUTF83, the central processing unit 80 satisfies the condition that the 32nd to 36th bits of the register rs are binary 00001, or the 32nd to 35th bits are not binary 0000. to decide. Further, the central processing unit 80 determines whether the 48th to 52nd bits of the register rs satisfy the condition that the binary number 00001, or the 48th to 51st bits are not the binary number 0000. When all the above conditions are satisfied, the central processing unit 80 stores 1 in the register cr.
(d) shows an example of the conditional addition instruction included in the central processing unit 80. By this instruction, the computer stores the result of adding the constant imm to the register rs in the register rt when the register cr is 1.
FIG. 14 shows the processing flow of the instruction sequence generated by the compiler 10 by optimizing the conversion processing (Example 2). The computer reads the unit data of a predetermined size, for example, 64-bit data into the register by the instruction sequence of the acquisition process generated by the acquisition process generation unit 700 (S1500). The read-destination register is the register r6. Next, the computer performs the following processes in parallel by the conversion detection process generated by the acquisition process generation unit 700.
First, the computer determines whether or not 1-byte UTF8 is continuously arranged in register r6 in parallel with the conversion process to be executed when 1-byte UTF8 is continuously arranged in register r6. Execute the detection process to be detected (S1510). In addition, the computer determines whether or not 2-byte UTF8 is continuously arranged in register r6 in parallel with the conversion process to be executed when 2-byte UTF8 is continuously arranged in register r6. Execute the detection process to be detected (S1550).
Furthermore, the computer determines whether or not the 3-byte UTF8 is continuously arranged in the register r6 in parallel with the conversion process to be executed when the 3-byte UTF8 is continuously arranged in the register r6. Detection Executes detection processing (S1570). The processing of S1500, S1510, S1550, and S1570 will be described in detail in FIG.
When it is detected that 1-byte characters are arranged consecutively in register r6, the computer outputs 4 converted characters one by one each time a call is received from another method, for example (S1520). ). In this case, the computer converts the unconverted 32-bit data in register r6 to UTF16 (S1530), and outputs the converted characters one by one for another four characters (S1540).
If it detects that 2-byte characters are arranged consecutively in register r6, the computer outputs 4 converted characters one by one each time it receives a call from another method, for example (S1560). ). On the other hand, when it is detected that 3-byte characters are continuously arranged in the register r6, the computer outputs two converted characters one by one each time a call is received from another method, for example. (S1580).
On the other hand, when the character codes of any data size are not continuously arranged in the register r6, the computer processes a pre-prepared instruction sequence for converting the character code system without using instructions such as UTF81toUTF16. (S1590), the conversion result is output character by character (S1595).
As described above, according to the processing of this figure, when UTF8 character codes having the same data size are continuously arranged in 64-bit unit data, these character codes are collectively converted to UTF16. As a result, the efficiency of the conversion process can be improved when the character codes having the same data size are continuous. For example, when 1-byte UTF8 characters are arranged consecutively, the processing proceeds according to the route shown by the thick line in this figure, so the 8-character character code can be converted by reading the data once, which is efficient. .. As a result, a document such as an XML document in which 1-byte UTF8 characters are continuously arranged in the tag part and a specific code indicating Japanese characters is continuously arranged in the text part. It can be converted especially efficiently.
Further, according to the process of this figure, even if a plurality of characters are converted at once, the converted characters are output one by one each time a call is received from another method. As a result, it is not necessary to call the process of this figure and change other methods, so that the affinity with the conventional program is high.
FIG. 15 shows an example of the instruction group 50 generated by the compiler 10 by optimizing the conversion process (Example 2). The acquisition process generation unit 700 generates a process instruction for acquiring the character code of the character to be converted on the second line. This instruction reads 64-bit data into register r6 from the address pointed to by register r3. Then, the conversion detection process generation unit 710 performs each of the plurality of conversion processes to be executed by selecting one according to the character code value range, and the detection process for detecting the character code value range. Generate the instruction strings to be executed in parallel on the 4th, 5th, and 7th rows.
For each data size that UTF8 can take, these instructions execute multiple conversion processes that should be executed when the character code of that data size is continuously arranged in register r6, and the conversion result is registered in register r7, It is stored in register r8 and register r9, respectively. Furthermore, these instructions store the detection result of detecting whether or not the character code of the data size is continuously arranged in the register r6 in the register cr0.lt, the register cr1.lt, and the register cr2.lt, respectively. To do.
Subsequently, the selection processing generation unit 720 generates instruction sequences for detecting whether character codes of any data size are continuously arranged in the register r6 on the 11th, 13th, and 22nd lines. To do. That is, for example, this instruction sequence determines whether a plurality of character codes having different data sizes are arranged in the register r6, and if such a determination is made, the character code system is converted without using an instruction such as UTF81toUTF16. Move the process to the pre-prepared instruction sequence.
Further, the selection processing generation unit 720 selects and outputs the processing result of any of the conversion processing among the plurality of conversion processing based on the detection result of the detection processing, and outputs the ninth, fifteenth, and thirteenth instruction sequences. Generate on the 20th line. These instruction sequences are stored in register r4 by selecting one of registers r7, r8, and r8 according to the values of registers cr0, cr1, and cr2. That is, the conversion result is output to the register r4.
Further, preferably, the selection process generation unit 720 gives an instruction to increment the register r3 indicating the reading destination of the character code on the 14th, 18th, and 25th lines in preparation for the subsequent instruction to acquire the character code. Generate. Further, preferably, the selection processing generation unit 720 issues commands for calculating and outputting the number of characters of the converted characters in the 10th, 17th, and 21st lines in preparation for the processing using the converted character code. To generate. These instructions output the number of converted characters to register r5.
The number of cycles required to execute the above instructions when executed by the POWER4 (registered trademark) architecture of IBM Corporation (registered trademark) will be described. In this architecture, two fixed-point instruction instructions, two branch instructions, and one conditional register instruction are executed in parallel in one cycle. Also, in this architecture, the latency of the load instruction is 3 cycles.
Also, in this architecture, the latency of a normal addition instruction is 2 cycles, and the throughput is 1 cycle. Therefore, the UTF81to16, UTF82to16, and UTF83to16 instructions are treated as having similar throughput and latency. Under this premise, the instruction sequence in this figure is executed in 11 cycles. That is, according to the instruction string in this figure, a UTF8 character string of 2 to 4 characters can be converted into a UTF16 character string in 11 cycles.
On the other hand, for example, according to the example of FIG. 10, it takes a minimum of 10 cycles to convert one character. Specifically, when converting 1-byte UTF8 to UTF16, a load instruction (3 cycles), a comparison instruction (2 cycles), a branch instruction (1 cycle), a shift instruction (2 cycles), and an AND instruction (2 cycles). Cycle) is executed. In this way, the instruction sequence in this figure can improve the conversion efficiency per number of characters as compared with the example in FIG.
Instead of the above processing, the conversion detection processing generation unit 710 further detects which character code of the unit data has that data size as the detection processing to be executed for each character code data size. May be generated. In this case, the selection processing generation unit 720 selects and outputs only the processing result of the character code detected to be the data size by the detection processing from the processing results of the conversion processing corresponding to the detection processing. Generate a sequence of instructions to do. In this case, the processing can be speeded up not only when the character codes of the same size are arranged in the unit data but also when the character codes of different sizes are arranged.
Further, in this figure, the conversion result of each conversion process is selected according to the values of the register cr0.lt, the register cr1.lt, and the register cr2.lt, and is output to the register r4. Instead of this, the processing result of each conversion processing may be stored in the continuous address of the storage area, and the processing result acquired from the address determined according to the processing result of the detection processing may be output to the register r4. That is, in this case, the conversion detection processing generation unit 710 generates an instruction sequence for storing the processing result of the conversion processing in a predetermined storage area corresponding to the conversion processing as each of the plurality of conversion processes. Then, the selection process generation unit 720 generates an address of a storage area for storing the process result to be read based on the range of character code values detected by the detection process, and an instruction to read the process result from the generated address. Generate a column. In this case, the process of reading the process result into the register r4 can be aggregated into one instruction.
FIG. 16 shows an example of the hardware configuration of the computer 500 that functions as the compiler 10 or the compiler 70. The computer 500 includes a CPU peripheral part having a CPU 80, RAM 1720, a graphic controller 1775, and a display device 1780 connected to each other by a host controller 1782, a communication interface 1730 connected to the host controller 1782 by an input / output controller 1784, and a hard disk drive. It includes an input / output unit having 1740 and a CD-ROM drive 1760, and a legacy input / output unit having a ROM 1710 connected to an input / output controller 1784, a flexible disk drive 1750, and an input / output chip 1770.
The host controller 1782 connects the RAM 1720 to the CPU 80 and the graphic controller 1775 that access the RAM 1720 at a high transfer rate. The CPU 80 operates based on the programs stored in the ROM 1710 and the RAM 1720, and controls each part. The graphic controller 1775 acquires the image data generated by the CPU 80 or the like on the frame buffer provided in the RAM 1720, and displays the image data on the display device 1780. Instead of this, the graphic controller 1775 may internally include a frame buffer for storing image data generated by the CPU 80 or the like.
The input / output controller 1784 connects the host controller 1782 to the communication interface 1730, the hard disk drive 1740, and the CD-ROM drive 1760, which are relatively high-speed input / output devices. The communication interface 1730 communicates with an external device via a network. Further, the communication interface 1730 communicates with the semiconductor test apparatus 10. Hard disk drive 1740 stores programs and data used by computer 500. The CD-ROM drive 1760 reads a program or data from the CD-ROM 1795 and provides it to the input / output chip 1770 via the RAM 1720.
Further, the ROM 1710 is connected to the input / output controller 1784 to a relatively low-speed input / output device such as a flexible disk drive 1750 or an input / output chip 1770. The ROM 1710 stores a boot program executed by the CPU 80 when the computer 500 starts up, a program that depends on the hardware of the computer 500, and the like. The flexible disk drive 1750 reads a program or data from the flexible disk 1790 and provides it to the I / O chip 1770 via RAM 1720. The input / output chip 1770 connects various input / output devices via a flexible disk 1790, for example, a parallel port, a serial port, a keyboard port, a mouse port, and the like.
The program provided to the computer 500 is stored in a recording medium such as a flexible disk 1790, a CD-ROM1795, or an IC card and provided by the user. The program is read from the recording medium via the input / output chip 1770 and / or the input / output controller 1784, installed in the computer 500, and executed. Since the operation performed by the compiler program installed and executed on the computer 500 on the computer 500 is the same as the operation on the compiler 10 or the compiler 70 described with reference to FIGS. 1 to 15, the description thereof will be omitted.
The program shown above may be stored in an external storage medium. As the storage medium, in addition to the flexible disk 1790 and CD-ROM1795, an optical recording medium such as a DVD or PD, an optical magnetic recording medium such as an MD, a tape medium, a semiconductor memory such as an IC card, or the like can be used. Further, a storage device such as a hard disk or RAM provided in a dedicated communication network or a server system connected to the Internet may be used as a recording medium, and a program may be provided to the computer 500 via the network.
Although the present invention has been described above using the embodiments, the technical scope of the present invention is not limited to the scope described in the above embodiments. It will be apparent to those skilled in the art that various changes or improvements can be made to the above embodiments. It is clear from the description of the claims that the form with such modifications or improvements may also be included in the technical scope of the present invention.
According to the above embodiment, the compiler, the compiler program, the recording medium, the control method, and the central processing unit shown in the following items are realized. (Item 1) A compiler that optimizes the conversion of the character code system of the characters stored in the character variable in the target program to be optimized, and reads the characters of the character variable written by the first character code system. Prior to each of a plurality of processes using the character in the second character code system, the character is converted from the first character code system to the second character code system and stored in the character variable. For the conversion instruction generation unit that generates the instruction and each conversion instruction generated by the conversion instruction generation unit, the second character code system is applied to the character variable in all the execution paths executed prior to the conversion instruction. A compiler including a conversion instruction removing unit that removes the conversion instruction when the character of is stored. (Item 2) The conversion command generator is an instruction to convert from UTF8, which is a character code system of characters in an XML document, to UTF16, which is a character code system used when the plurality of processes described as a Java program operate a character string. The compiler according to item 1, which generates the above conversion instruction.
(Item 3) The conversion instruction removing unit sets the second character in the character variable in any execution path executed prior to the conversion instruction for each conversion instruction generated by the conversion instruction generation unit. The compiler described in item 1 that generates a conversion instruction in the execution path when the characters of the code system are not stored and removes the partial redundancy of the conversion instruction by removing the conversion instruction. (Item 4) The character variable is generated by a constructor that secures a storage area for storing characters and sets the character code system, and the conversion instruction removing unit is each conversion generated by the conversion instruction generation unit. Regarding the instruction, when the constructor of the character variable sets the character variable in the second character code system in each execution path executed prior to the conversion instruction, the character variable is set to the first character variable in the execution path. The compiler described in item 1 that determines that the characters of the character code system of 2 are stored.
(Item 5) The conversion instruction removing unit returns the characters of the second character code system for each conversion instruction generated by the conversion instruction generation unit in each execution path executed prior to the conversion instruction. The compiler according to item 1, which determines that the characters of the second character code system are stored in the character variable in the execution path when the method stored in the character variable is executed as a value. (Item 6) A method information storage unit that stores identification information of a method whose return value is a character of the second character code system, a character stored in a character variable set in the second character code system by the constructor, and a character. A method in which any of the characters already converted into the second character code system by the conversion instruction and the return value of the method corresponding to the identification information stored in the method information storage unit is further used as the return value. The method recursive determination unit that stores the identification information in the method information storage unit is further provided, and the conversion instruction removal unit executes each conversion instruction generated by the conversion instruction generation unit prior to the conversion instruction. When the return value of the method corresponding to the identification information stored in the method information storage unit is stored in the character variable in each execution path, the character of the second character code system is described in the execution path. The compiler described in item 5 that is judged to be stored in a character variable.
(Item 7) As a process for outputting the character stored in the character variable, when the character stored in the character variable is the first character code system, the character is output and stored in the character variable. The compiler according to item 1, further comprising an output processing instruction sequence generator that generates an instruction sequence for returning the character to the first character code system and outputting the character when the character is the second character code system. (Item 8) A compiler that optimizes the conversion of the character code system of the characters stored in the character variable in the target program to be optimized. It reads the characters of the character variable written by the first character code system and reads the second character. As each of a plurality of processes using the character in the code system, the process is executed when the character stored in the character variable is the second character code system, and the character is the first character code. A character operation processing generator that generates an instruction sequence with an exception that raises an exception when it is a system, and a character that is executed when the exception occurs and is stored in the character variable from the first character code system. A compiler including an exception handler generator that converts to the second character code system, stores it in the character variable, and generates an exception handler that returns to the process that caused the exception.
(Item 9) For each of the plurality of processes in which the character operation process generation unit reads and operates the characters of the character variable, the probability that the character input to the process is the second character code system is determined. The compiler according to item 8, which generates the exception instruction sequence when the probability is higher than a predetermined reference probability. (Item 10) The character variable stores code system information indicating whether the character variable is the first character code system or the second character code system, and characters of the first character code system. The first character storage area and the second character storage area for storing the characters of the second character code system are managed, and the character operation processing generation unit uses the characters of the first character code system as the character variables. When stored, the instruction sequence that causes an exception to the exception instruction sequence by setting the pointer indicating the address of the second character storage area to an invalid address, and the characters of the second character code system are Item 8 that generates an instruction sequence that sets the address of the second character storage area as a valid address when stored in the character variable, and an instruction sequence that reads the second character storage area as the exception instruction sequence. Described compiler.
(Item 11) A compiler that optimizes the conversion process that converts the characters stored in the character variables in the target program to be optimized from the first character code system to the second character code system, and is the character to be converted. The acquisition process generation unit that generates the instruction sequence of the process for acquiring the character code of, each of the plurality of conversion processes to be executed by selecting one according to the range of the value of the character code, and the character code. A conversion detection process generator that generates an instruction sequence that executes a detection process that detects a range of values in parallel, and a process result of any of the plurality of conversion processes are detected by the detection process. A compiler equipped with a selection processing generator that generates an instruction sequence that selects and outputs based on the result. (Item 12) The data size of the character code differs depending on the range of values of the character code, and the acquisition processing generation unit generates an instruction sequence for acquiring unit data of a predetermined size including the character code, and the above-mentioned The conversion detection process generation unit performs the unit data in parallel with the conversion process to be executed when a plurality of character codes of the data size are continuously arranged in the unit data for each data size of the character code. Item 11. The compiler according to item 11, which generates an instruction sequence for executing a detection process for detecting whether or not a plurality of character codes of the data size are continuously arranged in the data.
(Item 13) The conversion detection process generation unit further detects which character code of the unit data has the data size as the detection process to be executed for each data size of the character code. An item for generating an instruction string to be generated, and the selection processing generation unit selects and outputs a processing result for a character code detected by the detection processing to be the data size among the processing results of the conversion processing. 12 Described compilers. (Item 14) The conversion detection process generation unit generates, as each of the plurality of conversion processes, an instruction sequence for storing the processing result of the conversion process in a predetermined storage area corresponding to the conversion process. The selection process generation unit generates an address of a storage area for storing the processing result of converting the character code based on the range of the value of the character code detected by the detection process, and processes from the generated address. The compiler according to item 11 that generates an instruction sequence for reading the result.
(Item 15) A central processing device that executes a conversion process that converts characters in the first character code system, which have different data sizes according to the type of characters, into characters in the second character code system, and is the character of the character. For each data size of the code, in parallel with the conversion process to be executed when a plurality of character codes of the data size are continuously arranged in the register, a plurality of character codes of the data size are continuously arranged in the register. A central processing device having an instruction to execute a detection process for detecting whether or not the data is arranged. (Item 16) A compiler program that makes a computer function as a compiler that optimizes the conversion of the character code system of characters stored in character variables in the target program to be optimized, and the computer is written by the first character code system. Prior to each of a plurality of processes in which the character of the character variable is read and the character is used in the second character code system, the character is converted from the first character code system to the second character code system. For the conversion instruction generation unit that generates the conversion instruction to be stored in the character variable and each conversion instruction generated by the conversion instruction generation unit, the character variable is set to the character variable in all the execution paths executed prior to the conversion instruction. A compiler program that functions as a conversion instruction removal unit that removes the conversion instruction when the characters of the second character code system are stored.
(Item 17) A compiler program that functions a computer as a compiler that optimizes the conversion of the character code system of characters stored in character variables in the target program to be optimized. When the character stored in the character variable is the second character code system as each of a plurality of processes in which the character of the character variable written by the system is read and the character is used in the second character code system. A character operation process generator that executes the process and generates an instruction sequence with an exception that generates an exception when the character is the first character code system, and is executed when the exception occurs. A character stored in a character variable is converted from the first character code system to the second character code system, stored in the character variable, and an exception handler for returning to the process in which the exception is generated is generated. A compiler program that functions as an exception handler generator. (Item 18) A compiler program that makes a computer function as a compiler that optimizes the conversion process that converts the characters stored in character variables in the target program to be optimized from the first character code system to the second character code system. A plurality of the conversions to be executed by selecting and executing the acquisition processing generation unit for generating the instruction sequence of the processing for acquiring the character code of the character to be converted and the value range of the character code. A conversion detection processing generation unit that generates an instruction sequence that executes each of the processing and a detection processing that detects a range of the character code values in parallel, and a conversion processing process of any one of the plurality of conversion processes. A compiler program that functions as a selection process generator that generates an instruction sequence that selects and outputs the result based on the detection result of the detection process.
(Item 19) A recording medium on which the compiler program described in any of items 16 to 18 is recorded. (Item 20) A control method in which a computer controls a compiler that optimizes the conversion of the character code system of characters stored in character variables in the target program to be optimized, and the first character code is controlled by the computer. Prior to each of a plurality of processes in which the character of the character variable written by the system is read and the character is used in the second character code system, the character is transferred from the first character code system to the second character code. In the conversion instruction generation step of converting to a system and generating the conversion instruction to be stored in the character variable, and for each conversion instruction generated in the conversion instruction generation stage, in all the execution paths executed prior to the conversion instruction. , A control method including a conversion instruction removing step of removing the conversion instruction when the character of the second character code system is stored in the character variable. (Item 21) This is a control method in which a computer controls a compiler that optimizes the conversion process that converts the characters stored in the character variables in the target program to be optimized from the first character code system to the second character code system. The computer generates an instruction sequence for a process of acquiring the character code of the character to be converted, and a plurality of the conversions to be executed by selecting one according to the value range of the character code. A conversion detection process generation stage that generates an instruction sequence that executes each of the processes and a detection process that detects a range of the character code values in parallel, and a process of any of the plurality of conversion processes. A control method including a selection process generation step of generating an instruction sequence that selects and outputs a result based on the detection result of the detection process.
<figref num="1">FIG. 1 is a block diagram of compiler 10 (Example 1).</figref><figref num="2">FIG. 2 is an example of the data structure of the character variable 20 (Example 1).</figref><figref num="3">FIG. 3 shows a processing flow in which the compiler 10 optimizes the conversion processing (Example 1).</figref><figref num="4">FIG. 4 is an example of a program showing the operation of the instruction sequence generated by the output processing instruction generation unit 150 as output processing (Example 1).</figref><figref num="5">FIG. 5 shows an example of a program in which the compiler 10 according to this embodiment is applied to the SAX input library (Example 1).</figref><figref num="6">FIG. 6 shows an example of a program in which the compiler 10 according to this embodiment is applied to the DOM input library (Example 1).</figref><figref num="7">FIG. 7 is a block diagram of the compiler 70 and the central processing unit 80 (Example 2).</figref><figref num="8">FIG. 8 shows a processing flow in which the compiler 10 optimizes the conversion processing (Example 2).</figref><figref num="9">FIG. 9 shows a comparison of UTF16 and UTF8 character code systems (Example 2).</figref><figref num="10">FIG. 10 shows another example of the process of converting UTF16 to UTF8 (Example 2).</figref><figref num="11">FIG. 11 shows a processing flow diagram of the conversion process generated by the compiler 70 in this embodiment (Example 2).</figref><figref num="12">FIG. 12 shows an example of a group of instructions for converting UTF8 to UTF16, which is possessed by the central processing unit 80 (Example 2).</figref><figref num="13">FIG. 13 shows other instructions that the central processing unit 80 has (Example 2).</figref><figref num="14">FIG. 14 shows the processing flow of the instruction sequence generated by the compiler 10 by optimizing the conversion processing (Example 2).</figref><figref num="15">FIG. 15 shows an example of the instruction group 50 generated by the compiler 10 by optimizing the conversion process (Example 2).</figref><figref num="16">FIG. 16 shows an example of the hardware configuration of the computer 500 that functions as the compiler 10 or the compiler 70.</figref>
Code description
10 Compiler 20 Character variable 50 Instruction group 70 Compiler 80 Central processing device 100 Character code system Judgment unit 105 Conversion instruction generation unit 110 Conversion instruction removal unit 120 Method information storage unit 130 Method recursive judgment unit 140 Character operation processing generation unit 145 Exception handler generation Part 150 Output processing instruction generation unit 210 Code system information 220 1st pointer 225 1st character storage area 230 2nd pointer 235 2nd character storage area 240 Data length information 500 Computer 700 Acquisition processing generation unit 710 Conversion detection processing generation unit 720 Selection Processing generator
23 sheets
Sheet 1 Sheet 2 Sheet 3 Sheet 4 Sheet 5 Sheet 6 Sheet 7 Sheet 8 Sheet 9 Sheet 10 Sheet 11 Sheet 12 Sheet 13 Sheet 14 Sheet 15 Sheet 16 Sheet 17 Sheet 18 Sheet 19 Sheet 20 Sheet 21 Sheet 22 Sheet 23
Every citation, both ways
| Document | Relation | Office | Cited during |
|---|---|---|---|
| WO2016121509A1 | Cited by | World Intellectual Property Organization (WIPO) | International search |
| JP2016139294A | Cited by | Japan | Search report |
| US8633936B2 | Cited by | United States of America | Applicant |
| JP2016139294A | Cited by | Japan | Search report |
| JP2011518398A | Cited by | Japan | Search report |
| JP2016139294A | Cited by | Japan | Search report |
| JP2010108076A | Cited by | Japan | Examiner |
| CN107209672A | Cited by | China | Search report |
| US9880612B2 | Cited by | United States of America | Applicant |
2 priority claims, no other members on record
Priority claims2
| Document | Office | Kind | Date |
|---|---|---|---|
| 2004109650 | Japan | A | |
| JP20040109650 | – | – | – |
12 legal events, as the office reported them to INPADOC
Over the term
Point at a mark for the eventEvents
| Event | Code | |
|---|---|---|
| Cancellation because of no payment of annual feesLAPS | LAPS | |
| Renewal fee payment (event date is renewal date of database)FPAY | FPAY | |
| Certificate of patent or registration of utility modelR150 | R150 | |
| First payment of annual fees (during grant procedure)A61 | A61 | |
| Notification of resignation of power of sub attorneyRD14 | RD14 | |
| Written decision to grant a patent or to grant a registration (utility model)A01 | A01 | |
| Written decision to grant a patent or to grant a registration (utility model)A01 | A01 | |
| Decision of grant or rejection writtenTRDD | TRDD | |
| Written amendmentA521 | A521 | |
| Notification of reasons for refusalA131 | A131 | |
| Report on accelerated examinationA975 | A975 | |
| Explanation of circumstances concerning accelerated examinationA871 | A871 |
Numbers
- Publication
- 2005293386
- Publication, DOCDB
- 2005293386
- Publication, EPODOC
- JP2005293386
- Application
- 109650
- Application, DOCDB
- 2004109650
- Application, EPODOC
- JP20040109650
Titles3
- Japanese
- コンパイラ、コンパイラプログラム、記録媒体、制御方法、及び中央処理装置
- English
- Compilers, compiler programs, recording media, control methods, and central processing units
- English
- COMPILER, COMPILER PROGRAM, RECORDING MEDIUM, CONTROL METHOD AND CENTRAL PROCESSOR
Classification
- CPC, 4
- G06F17/2264
- G06F40/126
- G06F40/151
- G06F17/2217
- IPC, 3
- G06F9 45
- G06F17 22
- H03M7 00