System and method of organizing community intelligent information by using organic matter data model
Abstract
A system and method for organizing community intelligence information using an organic object data model. The method includes: receiving one or more web pages containing community intelligence data; Word hyphenation in the content; identify the Named Entity in the hyphenated content of the one or more webpages; identify the themes in the hyphenated content of one or more webpages; identify one or more Opinions in the content of the segmented words of the webpage; integrate the identified named entities, topics and opinions to construct the organic object data model; and store the organic object data associated with the constructed organic object data model in the organic object In the database.

Term
4.1 yearsto projected expiry
Projected expiry 25 October 2030, counted from filing; an application has no term until it is granted.
- Priority
- Filed
- Published
- Today
- Projected expiry
18 claims: 3 independent, 15 dependent
- 1一种用于使用有机物件数据模型来撷取及组织线上收集的社群智能数据的方法,所 述方法包括: 通过用以撷取及管理社群智能信息的一计算机来接收含有社群智能数据的一个或多 个网页; 通过所述计算机来对含有社群智能数据的所述一个或多个网页的内容进行断词; 通过所述计算机来识别所述一个或多个网页的所断词的所述内容中的附名实体; 通过所述计算机来识别所述一个或多个网页的所断词的所述内容中的主题; 通过所述计算机来识别所述一个或多个网页的所断词的所述内容中的意见; 通过所述计算机来整合所识别的所述附名实体、所述主题及所述意见,以建构一有机 物件数据模型;以及 通过所述计算机来将与所建构的所述有机物件数据模型相关联的有机物件数据存储 于一有机物件数据库中。
- 2如权利要求1所述的方法,其中所述识别所述附名实体的步骤还包括: 通过所述计算机,使用一以条件随机域为基础的演算法来训练一物件辨识模块。
- 3如权利要求2所述的方法,其中所述识别所述附名实体的步骤还包括: 通过所述计算机,根据一预定标准来对所识别的所述附名实体进行分类,并将所分类 的所述附名实体存储于一专用名词词典中。
- 4如权利要求3所述的方法,其中所述识别所述主题的步骤还包括: 通过所述计算机,根据主题之间的语意相似性与以机器为基础的分类来训练一主题分 类及辨识模块。
- 5如权利要求4所述的方法,其中所述识别所述主题的步骤还包括: 通过所述计算机,根据存储于所述专用名词词典中的主题样式与语意相似性来对所识 别的所述主题进行分类。
- 6如权利要求5所述的方法,其中所述识别所述意见的步骤还包括: 通过所述计算机,根据一以机器学习为基础的演算法来训练一意见探勘模块,其中该 以机器学习为基础的演算法包括一支持向量机。
- 7如权利要求6所述的方法,其中所述识别所述意见的步骤还包括: 通过所述计算机,使用含有语言样式或语法规则的一外挂模块来对所识别的所述意见 进行分类。 如权利要求1所述的方法,其中所述识别所述附名实体的步骤包括: 通过所述计算机,使用一以条件随机域为基础的演算法来训练一物件辨识模块;以及 通过所述计算机,根据一预定标准来对所识别的所述附名实体进行分类,并将所分类 的所述附名实体存储于一专用名词词典中。
- 89. 如权利要求8所述的方法,其中所述识别所述附名实体的步骤还包括: 通过所述计算机来选择在一特定时间周期内出现频率高于一临限值的附名实体。
- 910. 如权利要求1所述的方法,其中所述识别所述主题的步骤包括: 通过所述计算机,根据主题之间的语意相似性来训练一主题分类及辨识模块。
- 1011. 如权利要求10所述的方法,其中所述识别所述主题的步骤还包括: 通过所述计算机,根据存储于所述专用名词词典中的主题样式及语意相似性来对所识 别的所述主题进行分类。
- 1112. 如权利要求1所述的方法,其中所述识别所述意见的步骤包括: 通过所述计算机,根据一以机器学习为基础的演算法来训练一意见探勘模块,其中所 述以机器学习为基础的演算法包括一支持向量机。
- 1213. 如权利要求12所述的方法,其中所述识别所述意见还包括: 通过所述计算机,使用含有语言样式或语法规则之外挂模块来对所识别的所述意见进 行分类。
- 1314. 一种用于使用有机物件数据模型来撷取及组织线上收集的社群智能数据的系统, 所述系统由一个或多个计算机处理器实施,所述一个或多个计算机处理器执行存储于计算 机可读存储介质上的计算机程序,所述系统包括: 一断词及整合模块,耦接至一训练数据库,所述断词及整合模块用以接收含有社群智 能数据的网页; 一物件辨识模块,耦接至所述断词及整合模块,所述物件辨识模块用以识别包含于所 接收到的所述网页中的经分类的附名实体; 一主题分类及辨识模块,耦接至所述断词及整合模块,所述主题分类及辨识模块用以 识别所接收到的所述网页的每一句子及段落的主题; 一意见探勘及情感分析模块,耦接至所述断词及整合模块,所述意见探勘及情感分析 模块用以判定所接收到的所述网页的句子中的意见及与所识别的所述附名实体或所识别 的所述主题相关联的意见;以及 一物件关系建构模块,耦接至所述断词及整合模块,所述物件关系建构模块用以界定 附名实体之间的关系。
- 1415. 如权利要求14所述的系统,其中所识别的所述附名实体为有机物件,且所识别的 所述主题及意见为与其对应的物件相关联的社会属性。
- 1516. 如权利要求14所述的系统,所述物件辨识模块包括: 一附名实体辨识模块,用以根据一以条件随机域为基础的机器学习程序来识别附名实 体; 一后处理分类器模块,用以根据一预定标准对所识别的所述附名实体进行分类;以及 一智能附名实体过滤模块,用以更新一专用名词词典及所述训练数据库。
- 1617. 如权利要求14所述的系统,所述主题分类及辨识模块包括: 一训练模块,用以应用以语意向量为基础的机器学习方法来训练一主题分类器,以识 别主题样式及新的主题。 1 如权利要求14所述的系统,所述意见探勘及情感分析模块包括: 一意见探勘分类器,用以实施一机器学习演算法,且从含有语法规则或语言样式的一 外挂模块中撷取数据,以判定所述意见。 19.如权利要求14所述的系统,所述断词及整合模块包括: 一断词模块,用以根据一以条件随机域为基础的演算法及从一专用名词词典中撷取的 数据来对所接收到的所述网页的内容进行断词;以及 一整合模块,用以整合从所述物件辨识模块中接收到的所识别的所述附名实体、从所 述主题分类及辨识模块中接收到的所识别的所述主题以及从所述意见探勘及情感分析模 块接收到的所识别的所述意见,以建立一有机物件数据模型。
- 1720. 如权利要求19所述的系统,其中所述有机物件模型包含一有机物件、与所述有机 物件相关联的自产生属性、与所述有机物件相关联的领域专用属性以及与所述有机物件相 关联的社会属性。
- 1821. 一种用于撷取及组织线上收集的社群智能数据的系统,所述系统由一个或多个计 算机处理器实施,所述一个或多个计算机处理器执行存储于计算机可读存储介质上的计算 机程序,所述系统包括: 一断词及整合模块,耦接至一训练数据库,所述断词及整合模块用以接收含有社群智 能数据的网页,并支持一有机物件模型,其中该有机物件模型包含一有机物件、与所述有机 物件相关联的自产生属性、与所述有机物件相关联的领域专用属性以及与所述有机物件相 关联的社会属性; 一物件辨识模块,耦接至所述断词及整合模块,所述物件辨识模块用以识别包含于所 接收到的所述网页中的附名实体,其中所判定的所述附名实体为有机物件; 一主题分类及辨识模块,其耦接至所述断词及整合模块,所述主题分类及辨识模块用 以识别所接收到的所述网页的每一句子及段落的主题,其中所识别的所述主题为与其对应 的有机物件相关联的社会属性; 一意见探勘及情感分析模块,耦接至所述断词及整合模块,所述意见探勘及情感分析 模块用以判定所接收到的所述网页的句子中的意见及与所识别的附名实体相关联的意见, 其中所识别的所述意见为与其对应的有机物件相关联的社会属性;以及 一物件关系建构模块,耦接至所述断词及整合模块,所述物件关系建构模块用以界定 有机物件之间的关系。
Independent claims18
165 paragraphs, as filed
System and method for organizing community intelligent information using organic object data modelTechnical field
[0001] The present disclosure relates to the field of capturing and analyzing online collective intelligence information, and more specifically, it relates to collecting and managing data from an online social community , And use organic object architecture to provide high-quality search results systems and methods.
Background technique
[0002] Web 2.0 websites allow their users to interact with each other to become content providers of the website, and on some websites, users are restricted to passively viewing the information provided to them. Because of the ability to create and update content, many web authors can collaborate on creation together. For example, in wikis, users can expand, cancel, and redo each other's creations. In the blog, personal posts and comments will gradually accumulate over time.
[0003] Social intelligence (SI) refers to the concept of analyzing data collected from a group of Internet users, which enables people to understand opinions and past and future behaviors in social groups. In order for an online search engine to provide responsive online search results, the search system must effectively capture and manage SI information from various sources.
[0004] Keyword search in Web 2.0 websites is one of the commonly used online search methods. However, keyword search has several disadvantages. Keyword search is easy to over search, that is, to find irrelevant documents; and easy to search under, that is, to find some related documents. Moreover, the results of a keyword search usually do not distinguish the same keywords in different contexts. Therefore, Internet users may need to spend minutes or even hours scanning search results to identify useful information. These shortcomings of keyword search are even more significant when dealing with SI information.
[0005] The embodiments of the present disclosure are directed to managing the collected social intelligence information by using an organic object data model, so as to promote effective online search and overcome one or more of the above-mentioned problems.
Summary of the invention
[0006] In one aspect of the present invention, the present disclosure is directed to a method of using an organic object data model to capture and organize data collected online. The disclosed method includes: receiving one or more webpages containing community intelligence data; hyphenating the content of the one or more webpages containing community intelligence data; identifying the interruption of the one or more webpages The named entity in the content of the word; identify the topic in the content of the segmented word of one or more web pages; identify the opinions in the content of the segmented word of one or more web pages; integrate the identified named entities and topics And opinions to construct an organic object data model; and store the organic object data associated with the constructed organic object data model in the organic object database.
[0007] In another aspect of the present invention, the present disclosure is directed to a system for capturing and organizing social intelligence data collected online, the system is actually operated by one or more computer processors, The computer processor executes a computer program stored on a computer-readable storage medium. The system includes a word segmentation and integration module, an object identification module, a subject classification and identification module, an opinion exploration and sentiment analysis module, and an object relationship construction module. The word segmentation and integration module is coupled to the training database and used to receive web pages containing community intelligence data. The object recognition module is coupled to hyphenation and
The integration module is used to identify the named entity included in the received webpage. The topic classification and identification module is coupled to the word segmentation and integration module, and is used to identify the topic of each sentence and paragraph of the received webpage. The opinion exploration and sentiment analysis module is coupled to the word segmentation and integration module, and is used to determine the opinions in the sentences of the received webpage and the opinions associated with the identified named entity. The object relationship construction module is coupled to the hyphenation and integration module, and is used to define the relationship between the named entities.
[0008] In another aspect of the present invention, the present disclosure is directed to a system for capturing and organizing social intelligence data collected online. The system can be actually operated by one or more computer processors, and the computer processors execute computer programs stored on a computer-readable storage medium. The system includes a word segmentation and integration module, an object identification module, a subject classification and identification module, an opinion exploration and sentiment analysis module, and an object relationship construction module. The word segmentation and integration module is coupled to the training database and used to receive web pages containing community intelligence data. The word segmentation and integration module supports an organic object model including organic objects, self-generated attributes associated with the organic objects, Domain-specific attributes associated with the organic object and social attributes associated with the organic object. The object identification module is coupled to the word segmentation and integration module, and is used to identify the named entity included in the received webpage, wherein the determined named entity is an organic object. The topic classification and identification module is coupled to the word segmentation and integration module, and is used to identify the topic of each sentence and paragraph of the received webpage, wherein the identified topic is the social attribute associated with its corresponding organic object. The opinion exploration and sentiment analysis module is coupled to the word segmentation and integration module, and is used to determine the opinions in the sentences of the received webpage and the opinions associated with the identified named entities, where the identified opinions are corresponding to them Organic objects related Social attributes. The object relationship construction module is coupled to the hyphenation and integration module, and is used to define the relationship between organic objects.
Description of the drawings
[0009] FIG. 1a is a block diagram showing an example of the hardware architecture of an online search engine.
[0010] FIG. 1b is a block diagram showing an example of an organic object data model.
[0011] FIG. 2 is a block diagram showing an example of an organic data object.
[0012] FIG. 3 is a block diagram showing an example of an information capture and management system based on an organic object data model.
[0013] FIG. 4 is an exemplary flow chart of the procedure of the object recognition module of the information retrieval and management system shown in FIG. 3.
[0014] FIG. 5 is a flowchart illustrating an example of a procedure for applying an N letter combination algorithm through the object recognition module shown in FIG. 3.
[0015] FIG. 6 is a schematic diagram illustrating an example of a program that applies the N letter combination and algorithm.
[0016] FIG. 7 is a schematic diagram illustrating an example of the calculation of the trust value used in the object recognition module.
[0017] FIG. 8 is a block diagram illustrating an example of the topic classification and recognition module shown in FIG. 3.
[0018] FIG. 9 shows an example of semantic similarity calculation applied by the topic classification and recognition module.
[0019] FIG. 10 is a flowchart illustrating an example of a procedure for collecting and improving the quality of training data implemented by the topic classification and recognition module.
[0020] FIG. 11 is a more detailed example block diagram showing a procedure for collecting and improving the quality of training data implemented by the topic classification and recognition module.
[0021] FIG. 12a is a block diagram illustrating an example of the opinion mining and sentiment analysis module shown in FIG. 3.
[0022] FIG. 12b is a block diagram illustrating an example of the testing procedure implemented by the opinion exploration and sentiment analysis module.
[0023] FIG. 12c is a block diagram showing an example of the architecture that can be used to implement the topic classification and recognition module, and the opinion exploration and sentiment analysis module.
[0024] FIG. 13 is a block diagram illustrating an example of the word segmentation and integration module shown in FIG. 3.
[0025] [Description of main component symbols]
[0026] 10: Internet
[0027] 20: Load balancing server
[0028] 30: Web server
[0029] 40: Advertising server
[0030] 50: data collection server
[0031] 60: File database
[0032] 70: Online search engine
<td>[0033]</td><td>100</td>
<td>[0034]</td><td>110</td>
<td>[0035]</td><td>120</td>
<td>[0036]</td><td>130</td>
<td>[0037]</td><td>140</td>
<td>[0038]</td><td>150</td>
<td>[0039]</td><td>160</td>
<td>[0040]</td><td>170</td>
<td>[0041]</td><td>200</td>
<td>[0042]</td><td>210</td>
<td>[0043]</td><td>221</td>
<td>[0044]</td><td>222</td>
<td>[0045]</td><td>223</td>
<td>[0046]</td><td>224</td>
<td>[0047]</td><td>225</td>
<td>[0048]</td><td>231</td>
<td>[0049]</td><td>232</td>
<td>[0050]</td><td>241</td>
<td>[0051]</td><td>242</td>
<td>[0052]</td><td>243</td>
<td>[0053]</td><td>244</td>
<td>[0054]</td><td>245</td>
<td>[0055]</td><td>300</td>
<td>[0056]</td><td>310</td>
<td>[0057]</td><td>320</td>
<td>[0058]</td><td>330</td>
<td>[0059]</td><td>340</td>
<td>[0060]</td><td>350</td>
Organic Objects Data Model Organic Objects (Parent Objects) Self-generated Attribute Domain Special Attributes Social Attributes Child Objects Time Stamp Affirmative or Negative Opinions Organic Objects Named Restaurant Price Address Promotions Free Gifts Discount Dishes Types Parking Space User Comments Atmosphere Service Price Food Taste Information retrieval and management system word segmentation and integration module Object identification module Object relationship construction module Subject classification and identification module Opinion exploration and sentiment analysis module
<td>[0061]</td><td>360 : Training database</td>
<td>[0062]</td><td>370 :Web page</td>
<td>[0063]</td><td>380a: Organic Object Database</td>
<td>[0064]</td><td>380b: A dictionary of specific nouns</td>
<td>[0065]</td><td>440 : Smart NE filter module</td>
<td>[0066]</td><td>450 : Automatic word breaker training data generation module</td>
<td>[0067]</td><td>452 : Automatic NER training data generation module</td>
<td>[0068]</td><td>460 : CRF-based word breaker training module</td>
<td>[0069]</td><td>470 : Hyphenation module</td>
<td>[0070]</td><td>480 :NE identification module</td>
<td>[0071]</td><td>485 : CRF-based NER training module</td>
<td>[0072]</td><td>490 : Post-processing classifier</td>
<td>[0073]</td><td>495 : Hyphenation program</td>
<td>[0074]</td><td>496 : Object recognition procedure</td>
<td>[0075]</td><td>861 : Theme style table</td>
<td>[0076]</td><td>862 : Topic semantic vector table</td>
<td>[0077]</td><td>863 : Topic similarity table</td>
<td>[0078]</td><td>870 : Topic classifier module</td>
<td>[0079]</td><td>1010, 1020, 1030, 1040, 1050, 1060: programs used to collect and improve the quality of training data sets</td>
<td>[0080]</td><td>1110: Manually labeled data collection</td>
<td>[0081]</td><td>1111 : Sentence group/labeled data set</td>
<td>[0082]</td><td>1112 : Sentence group/labeled data set</td>
<td>[0083]</td><td>1113 : Sentence group/labeled data set</td>
<td>[0084]</td><td>1114 : Sentence group/labeled data set</td>
<td>[0085]</td><td>1115 : Sentence group/labeled data set</td>
<td>[0086]</td><td>1116 : Training data set</td>
<td>[0087]</td><td>1117: Test data collection</td>
<td>[0088]</td><td>1120 : SVM Trainer</td>
<td>[0089]</td><td>1130 : SVM model</td>
<td>[0090]</td><td>1140 : SVM classifier</td>
<td>[0091]</td><td>1150 : Sentence group/data collection</td>
<td>[0092]</td><td>1160 : Validator</td>
<td>[0093]</td><td>1210 : Opinion exploration procedure</td>
<td>[0094]</td><td>1220: CRF-based opinion word and style detector module</td>
<td>[0095]</td><td>1222 :form</td>
<td>[0096]</td><td>1224 :form</td>
<td>[0097]</td><td>1226 :form</td>
<td>[0098]</td><td>1240 : Machine learning classifier/opinion mining classifier</td>
<td>[0099]</td><td>1250 : Classifier based on grammar and rules / opinion exploration classifier</td>
<td>[0100]</td><td>1260 : Opinion decision score</td>
<td>[0101]</td><td>1270 : Opinion decision score</td>
<td>[0102]</td><td>1280 : Opinion Mining Classifier</td>
<td>[0103]</td><td>1310 : The result of the word segmentation, objects found, themes and opinions</td>
<td>[0104]</td><td>1330 : Module interface</td>
<td>[0105]</td><td>1340 :Integration module</td>
Detailed ways
[0106] The system and method of the present disclosure capture and manage the collected community intelligence information, so as to provide faster and more accurate online search results in response to user inquiries. The embodiments of the present disclosure use an organic object data model to provide a framework to capture and analyze information collected from online social networks and other online communities and other web pages. The organic object data model reflects the heterogeneous nature of intelligent information established by online social networks and communities. By applying the organic object data model, the information acquisition and management system of the present disclosure can efficiently classify a large amount of information, and present the searched information according to the request.
[0107] The embodiments of the present disclosure include software modules and databases, which can be actually operated by various configurations of computer software and hardware components. The configuration of each software and hardware can be various computer storage media, various computers used to perform certain disclosed functions, various third-party software applications, and software applications that implement the disclosed system functionality.
[0108] FIG. 1a is a block diagram showing an example hardware architecture of an online search engine 70. The online search engine 70 refers to any software and hardware used to provide search results of online content after receiving a user's search request. A well-known example of an online search engine is the Google search engine. As shown in FIG. 1a, the online search engine 70 receives user inquiries, such as search requests, from the Internet 10. The online search engine 70 can also collect SI information from online communities. The online search engine 70 can actually operate by using one or more servers (such as one or more 2X 300MHz Dual Pentium II servers produced by Intel). A server refers to a computer running a server operating system, but it can also be any software or dedicated hardware that can provide services.
[0109] The online search engine 70 includes one or more load balancing servers 20, which can receive search requests from the Internet 10 and forward the requests to one of the plurality of web servers 30. The network server 30 can coordinate the execution of the query received from the Internet 10, format the corresponding search results received from the data gathering server 50, retrieve the advertisement list from the ad server 40, and The search result is generated in response to the user's search request received from the Internet 10. The advertisement server 40 is used to manage advertisements associated with the online search engine 70. The data collection server 50 is used to collect SI information from the Internet 10 and organize the collected data by indexing the data or using various data structures. The data collection server 50 stores the organized data in the file database 60 and retrieves the organized data from the file database 60. In an exemplary embodiment, the data collection server 50 can host the information retrieval and management system according to the organic object data model. Hereinafter, the organic object data model will be described in conjunction with FIG. 1b and FIG. 2, and the information capture and management system will be described in conjunction with FIG. 3.
[0110] FIG. 1b is a block diagram of the organic object data model 100. As shown in FIG. 1b, the organic object 110 may be a named entity (for example, a named restaurant) with sub-objects 150. The child object 150 may be a named entity that inherits the characteristics of its parent object 110. The organic object 110 may have at least three types of attributes: self-producing attributes (self-producing attributes)
attribute) 120>domain-specific attribute (domain-specific attribute) 130 and social attribute (social attribute) 140. The self-generated attributes 120 include attributes generated by the object 110 itself. The domain-specific attributes 130 include attributes that describe the subject domain of the object 110. The social attribute 140 includes classified intelligence information contributed by an online community related to the object 110. In an exemplary instance, the intelligent information contributed by the online community may be user opinions, such as positive or negative opinions 170 about the object 110 or its attributes. Each category of classified intelligent information can be a topic associated with one or more opinions. The theme can also be a social attribute.
[0111] The organic object 110 includes a time stamp 160 (TS 160), which can associate the object 110 with a time period or time. TS 160 may indicate the life cycle of the object, which may be the time period between the creation and deletion of the object 110, or the effective time period of the object 110. In another example, the TS 160 may be the creation time of the information entry (entry) related to the object 110. As shown in FIG. 1b, all attributes (120, 130, and 140) and sub-objects (150) associated with object 110 may also have time stamps associated with them.
[0112] FIG. 2 provides an example of the organic object 200. As shown in FIG. 2, a restaurant 210 with a name (for example, McDonalds) may be an organic object. The sub-items of the restaurant 210 (not shown in FIG. 2 ), for example, include different types of food served in the restaurant 210, such as burgers, French fries, and the like. The self-generated attribute 120 of the organic restaurant 210 includes a lot of information, such as the address 222 of the restaurant 210, the price 221 set by the restaurant 210, and the promotional activities 223 of the restaurant 210 (for example, free gifts 224 and discounts 225). The domain-specific attributes 130 of the restaurant 210 include the type of dishes 231 served by the restaurant 210, the parking space 232 of the restaurant 210, and the like. The social attributes 140 of the restaurant 210 include user reviews 241 of the restaurant 210 and user opinions on topics such as atmosphere 242, service 243, price 244, and food taste 245. User opinions can be negative (for example, the price is too expensive) or positive (for example, the service is excellent). As shown in Figure 2, an attribute can be associated with a time stamp (TS) to indicate its effective time.
[0113] FIG. 3 illustrates an information retrieval and management system 300 for retrieving information from the Internet and using organic object models to organize the information. The information acquisition and management system 300 collects community intelligence information provided by online social networks and other communities, and classifies and stores the collected community intelligence information by applying an organic object data model. The information retrieval and management system 300 will receive user inquiries requesting to search for certain information (for example, restaurant reviews on specific restaurants). The information capture and management system 300 responds to user inquiries by capturing information captured and organized based on the organic object model.
[0114] The information retrieval and management system 300 includes a word segmentation and integration module 310, an object recognition module 320, an object relation construction module 330>the subject classification and recognition module 340, and an opinion exploration and sentiment analysis module 350. The information retrieval and management system 300 may further include a training database 360, an organic object database 380a, and a lexicon dictionary 380b. The training database 360 stores data records, for example, NE (named entity), topic or topic style, opinion word, and opinion style. The training database 360 can provide training data sets for the object recognition module 320, the topic classification and recognition module 340, and the opinion exploration and sentiment analysis module 350 to facilitate the machine learning process. The training database 360 can receive training data from the object recognition module 320, the topic classification and recognition module 340, and the opinion exploration and sentiment analysis module 350 to facilitate the machine learning process. The organic object database 380a may store organic objects (for example, 200 in FIG. 2). The noun dictionary 380b stores the recognized NE (organic objects), themes (social attributes), thematic styles (social attributes), opinions (social attributes), and opinion styles (social attributes). Sex) and other information classified by one or more modules of the information retrieval and management system 300.
[0115] The word segmentation and integration module 310 will receive a web page 370 from the Internet. The web page 370 can be any web page containing community intelligence data collected from an online community. The word segmentation and integration module 310 will also perform analysis on the content in the web page 370
Hyphenate words and identify the boundaries of proper nouns in each sentence. For example, one difference between Chinese and English is that proper nouns in Chinese sentences do not have clear boundaries. Therefore, before processing any Chinese language content from the web page 370, the word segmentation and integration module 310 needs to segment the specific nouns in the sentence. Traditionally, software applications use plug-in modules containing various language styles/grammatical rules to segment text. The Linear Chain Conditional Random Field (CRF) algorithm is one of the improved algorithms used to segment the text, and it is widely used in the segmentation of Chinese words.
[0116] One of the disadvantages of the CRF method is its poor performance in processing rapidly changing input data. However, the social intelligence information provided by online social networks and communities is rapidly changing data. Therefore, in this exemplary embodiment, the word segmentation and integration module 310 uses an improved machine learning method, which benefits from the machine learning of other modules (object recognition module 320, topic classification and recognition module 340, and opinion exploration module 350). Function to implement improved machine learning and word segmentation procedures. Examples of the improved machine learning program are further disclosed in FIGS. 4 to 13 below.
[0117] In an illustrative example, the training database 360 is updated by the training procedures in the object recognition module 320, the topic classification and recognition module 340, and the opinion exploration module 350 to improve the quality of the training data. The high-quality training data from the training database 360 can improve the accuracy of the word segmentation performed by the word segmentation and integration module 310.
[0118] FIG. 4 shows the object recognition module 320. The object identification module 320 is used to identify NE, classify the identified NE, and store the classified NE in the proper noun dictionary 380b. The proper noun dictionary 380b contains multiple named entity proper nouns, for example, food NE, restaurant NE, and geographic location NE. The word segmentation program 495 and the Object Recognition (NER) program 496 respectively include two programs: a learning program and a testing program. During the learning process, the module (for example, the training module) of the information retrieval and management system 300 reads the labeled data from the training database (for example, the database 360), and calculates the parameters for the mathematical model related to machine learning . During the learning process, the training module can also configure the classifier according to the calculated parameters and the mathematical model related to machine learning. A classifier refers to a software module that maps multiple sets of input data to multiple categories based on one or more attributes of the input data. For example, Category refers to the subject, opinion, or any other classification based on one or more attributes of the input data. After that, the module (ie, the test module) of the information retrieval and management system 300 uses the classifier to test the new data, and this operation can be called a test procedure. During the test procedure, the test module marks the newly read data as different NEs, such as restaurants, food types, or geographic locations. The training database 360 contains domain-specific training files, which can be labeled for different NEs.
[0119] As shown in FIG. 4, the object recognition module 320 retrieves data from the proper noun dictionary 380b and the training database 360. The segmentation program 495 includes an autosegmenter training data producing module 450> CRF-based segmenter training module (CRF-based segmenter training module) 460 and a segmenter testing module (segmenter testing) module) 470<sub>o</sub>The word segmentation program 495 can be actually operated as a part of the word segmentation and integration module 310, or the actual operation can be a part of the object recognition module 320. When the information retrieval and management system 300 retrieves the web page 370, the system 300 first executes the word segmentation program 495 to segment the content of the web page 370. The system 300 then executes the named object recognition program 496 in the object recognition module 320 to identify NEo in the content
[0120] Next, the object identification module 320 uses a post-processing classifier 490 to classify the identified NE. The post-processing classifier 490 will use the context of the sentences surrounding the NE to determine the NE category. For example, the web page 370 may include community comments discussing several restaurants in different geographic locations. The post-processing classifier 490 classifies the identified NE into at least three entity classes: food, restaurant, and geographic location.
[0121] As shown in FIG. 4, the word segmentation program 495 and the object recognition program 496 both include automatic training data generation modules (450 and 452). The automatic training data generation modules 450 and 452 will receive the identified NE from the intelligent NE filtering module 440, and store the received NE in the training database 360. The automatic training data generating modules 450 and 452 can also access the NE stored in the training database 360 and send the extracted NE to the training modules 460 and 485. The word segmentation program 495 and the object recognition program 496 both include CRF-based training modules 460 and 485. In addition, CRF-based training modules 460 and 485 will use N-gram-based NE recognition training. CRF refers to a discriminative probability model commonly used to label or analyze continuous data (for example, natural language text or biological sequences). The N letter group refers to a subsequence of n items (such as letters, syllables, etc.) from a given order.
[0122] Moreover, both the word segmentation program 495 and the object recognition program 496 can use the training data from the training database 360 to train the word breaker training module 460 and the NE recognition training module 485 to better recognize the NE. The quality of the training data in the database 360 (for example, as well as the completeness and balance of the training data set (the smooth distribution of the data between categories) will affect the performance of the modules 310 and 320 (Figure 3). The quality of the training data can be determined by each The precision and recall values achieved by a module are measured.
[0123] After repeated training procedures, CRF-based word segmentation or NE recognition can achieve a high degree of precision and completeness (recall). The word segmentation module 470 will then segment the content in the web page 370 and send the segmented word content to the NE recognition (NER) module 480. The NE identification module 480 includes parallel identification sub-modules. For example, each identification sub-module can identify one type of NE. If the NE includes three types of NE (such as food, restaurant, and geographic location), the NE identification module 480 can actually operate three sub-modules to identify each type of NE (food name, restaurant name, and geographic location). The NE identification module 480 will then identify the NE, and then send the NE to the post-processing classifier 490.
[0124] If the output from the NE identification module 480 is ambiguous, the post-processing classifier 490 will arbitrate the result. For example, if two NE recognition sub-modules (for example, one for food and one for restaurants) respectively map one NE (for example, American ravioli) to the organic object data model, the post-processing classifier 490 The sentence context surrounding NE is used to determine its correct category (for example, "American Wonton" refers to the food itself, or a dish served by the restaurant in the sentence). The post-processing classifier 490 classifies the NE into multiple categories (for example, food name, restaurant name, and geographic location), and sends the identified NE to the intelligent NE filtering module 440.
[0125] As shown in FIG. 4, the intelligent NE filtering module 440 determines the best quality object identified by the NE identification module 480, and sends the newly identified NE (object) to be stored in the training database 360. The intelligent NE filtering module 440 may also add the newly identified NE to the proper noun dictionary 380b. The intelligent NE filtering module 440 may further send the identified NE to the NE identification module 480. FIG. 5 shows a block diagram of a program executed by an exemplary implementation of the intelligent NE filter module 440 (including its interface with other components of the system 300).
[0126] As shown in FIG. 5, the intelligent NE filtering module 440 uses the N letter combination and algorithm 510 to identify the NE pattern. NE style refers to the placement of NE in various sentences, including its word length (for example, the number of characters in a word) and its relative position to other words adjacent to it. The intelligent NE filtering module 440 may determine the frequency (term frequeue, TF) of various NE patterns by checking the timestamp and location in the sentence associated with the NE (520). TF refers to the frequency of appearance of NE or NE pattern in a specific time period. As shown in FIG. 5, the intelligent NE filtering module 440 determines the TF of each NE pattern in the current time period (530) and all time history (540) to filter out the outdated NE. Next, based on the calculated TF, the intelligent NE filtering module 440 can determine which NE patterns are correct (for example, higher than the threshold
TF), and send the selected NE pattern for further inspection by the subsequent program (step 550). The intelligent NE filtering module 440 may also group (560 and 575) the ambiguous NE patterns to be monitored (for example, TFs below the threshold). The intelligent NE filtering module 440 will then use the monitoring results (575 and 550) when it recognizes the correct NE pattern.
[0127] In order to further analyze the correct NE pattern (570), the intelligent NE filter module 440 calculates the confidence value (580), the trustworthiness value (582), and detects the boundary of the NE pattern (584). The following will be further described in conjunction with FIG. 6 and FIG. 7. The smart NE filter module 440 will then check the confidence value of the NE pattern, and for example, if the confidence value is higher than the threshold, it will send the NE pattern to be stored in the proper noun dictionary 380b or added to the training database 360. The intelligent NE filtering module 440 will similarly check the trustworthiness value of the NE pattern (582), and send the NE pattern to the automatic NER training data generation module 452 to be stored as part of the training data stored in the training database 360. The intelligent NE filtering module 440 will also determine the NE boundary, calculate the confidence value of the NE boundary (584), and use this boundary to identify the correct NE (496) in the sentence. The intelligent NE filtering module 440 then sends the identified NE to the post-processing classifier 490, and the post-processing classifier 490 can classify the NE and send the NE to be stored in the specific noun dictionary 380b. Alternatively, the intelligent NE filtering module 440 may also directly send and store the correct NE to the proper noun dictionary 380b (586).
[0128] FIG. 6 shows an example of a procedure 600 for calculating the trustworthiness value and the confidence value. As shown in FIG. 6, the intelligent NE filtering module 440 will recognize the N letter group pattern having a pattern length between 2 characters and 6 characters (610). The intelligent NE filtering module 440 will sort all NE patterns according to the length of the NE pattern, and then sort the result list according to the frequency of appearance in the file (620). The intelligent NE filtering module 440 can also sort the result list according to the appearance of the NE pattern. Frequency to calculate the NE style confidence value (see Figure 6, 660). According to the confidence value of the NE pattern, the intelligent NE filter module 440 checks the time stamp of the first occurrence of the NE pattern and its frequency of occurrence in a certain time period. For example, if the NE pattern is out of date, the intelligent NE filter module deletes the expired NE from the training database 360 to improve the quality of the training data.
[0129] The smart NE filter module 440 will then check whether certain NE patterns can be merged (640). For the merged NE pattern, the intelligent NE filter module 440 will determine the trustworthiness value according to the frequency of occurrence of the pre-merged NE (640). FIG. 7 shows a calculation example of the NE pattern trust value, which reflects the reliability of the NE identification in a certain time period. As shown in FIG. 7, in order to determine the trustworthy value, the intelligent NE filtering module 440 first extracts the initial code, the intermediate code, and the final code N letter group features from the NE (710). For example, Chinese NΕ "Italian noodles" has the initial code "Italy", the middle code "Big profit" and the ending code "Noodles" as its two-letter group features. Next, the intelligent NE filtering module 440 may determine whether the extracted feature belongs to a feature group of a specific field (for example, dining) (720). After that, the intelligent NE filtering module 440 calculates the weight of each extracted feature according to the length of the N letter group feature and its appearance frequency (730). Next, the intelligent NE filtering module 440 determines the trustworthy value according to the weight of the N letter group feature (740). In addition, the intelligent NE filtering module 440 can also determine the boundary of the new NE by calculating the trustworthy values of the prefix code, the intermediate code, and the suffix code. As shown in Figure 7, if the trustworthy value of a particular NE pattern is low, a manual data processing staff (for example, a data entry clerk) is used to view the data. According to and correct the N letter group feature or the appearance frequency of the feature (750).
[0130] FIG. 8 shows an example block diagram of the topic classification and recognition module 340. The topic classification and recognition module 340 will analyze the word-breaking web content received from the word-breaking and integration module 310 to identify topics discussed in online communities, mark each sentence and paragraph with the identified topics, and mark The identified and marked topics are sent to the segmentation and integration module 310 for further analysis. As shown in FIG. 8, the topic classification and recognition module 340 extracts topic patterns from sentences in the training database 360 based on the organic object data stored in the organic object database 380a and the topics and opinions in the term dictionary 380b (810). Next, the topic classification and recognition module 340 can remove the usual
The topic-independent stop words and other common words are used to reduce the length of the extracted topic style (820). Next, the topic classification and recognition module 340 can create a hierarchical topic style grouping by manual labeling (step 830). For example, referring to FIG. 2, the user view 241 can be a broad theme, which includes more specific themes: atmosphere 242, service 243, price 244, and taste 245. The theme classification and identification module 340 can group atmosphere 242, service 243, price 244, and taste 245 into four theme style groups.
[0131] Next, the topic classification and recognition module 340 calculates semantic similarity between the two topics (840). Figure 9 shows an example of semantic similarity calculation. As shown in Figure 9, topics i and j can be represented by topic semantic vectors V: and Vj, where the semantic similarity between topics i and j can be defined as:
[0132] Similarity, V'=cos(X, Vj)=cos θ
[0133] Assumption d<sub>ave</sub>Is the average similarity between topics in a group of topics, then when the topic classification and recognition module 340 determines the semantic similarity d between topic 1 and topic η<sub>n</sub>Greater than d<sub>ave</sub>At that time, it can determine that topic η is a new topic. In the disclosed example, the topic classification and recognition module 340 groups topic styles (830) before calculating semantic similarity (840), so as to improve the accuracy of new topic detection.
[0134] Please refer to FIG. 8 again, after calculating the semantic similarity (840), the topic classification and recognition module 340 stores the topic style, the topic semantic vector, and the semantic similarity in one or more tables (860). As shown in FIG. 8, the theme classification and recognition module 340 adds the recognized theme pattern to the training database 360 for use as training data.
[0135] As shown in FIG. 8, the topic classifier module 870 matches the topic styles stored in the topic style table 861, and checks the semantic similarity according to the data stored in the topic semantic vector table 862 and the semantic similarity table 863 , To process the segmented web page 370 (word segmentation and integration module 310). After that, the topic classifier module 870 classifies the topics in the content of the web page 370 and detects new topics in the content. Finally, the topic classification and identification module 340 will mark and compose topics related to each sentence on the web page 370, and determine the topic of each paragraph according to the topic of the sentence in the paragraph (880). The topic classification and recognition module 340 sends the sentence topic and paragraph topic to the word segmentation and integration module 310 for further processing.
[0136] FIG. 10 shows an example of the program 1000 for collecting and improving the quality of the training data set actually operated by the topic classification and recognition module 340. Other modules, such as the object recognition module 320 and the opinion exploration module 350, can use similar procedures to improve the quality of the training data. As shown in FIG. 10, the information retrieval and management system 300 starts with a collection of original training data (1010), such as a larger number of sentences and paragraphs collected from web pages of an online social network. For example, the original data set may include 50,000 sentences. Next, the data acquisition and management system 300 will sample the sentences from the original data set (for example, sample one of every 10 sentences) (1020). For example, manual data processing personnel (for example, data entry Member) will pass mark 5, Thousands of topics in the sample sentences are used to label the sampled data set, and the labeled data are stored in the training database 360 (1030). After that, the data acquisition and management system 300 will verify and correct the manually marked data set (1040) [0137] FIG. 11 shows an example of the verification and correction process 1040 actually operated by the subject classification and identification module 340. The data acquisition and management system 300 receives the manually marked data set 1110, in which one or more topics are marked in each sentence. The marked data set 1110 includes one or more marked sentences. The topic classification and recognition module 340 then recognizes five groups of sentences, for example, sentence groups 1111 to 1115. Each sentence data set (1111 to 1115) includes one or more sentences. The topic classification and recognition module 340 will then use the four labeled data sets 1111 to 1114 as the training data set 1116, and use the fifth data set 1115 as the test data set 1117. The data acquisition and management system 300 will pass through the support vector machine (Support Vector Machine, SVM) trainer 1120 to
The Π/14 page processes the four sentence data sets in 1116 to process the training data set 1116. The SVM trainer 1120 can use the SVM model 1130. The SVM model 1130 may be a presentation of data samples as points in space, which are mapped so that samples of individual categories can be distinguished by clear gaps. Next, the topic classification and identification module 340 will use the SVM parameters calculated from the training data set 1116 to configure the SVM classifier 1140. The topic classification and identification module 340 will use the configured SVM classifier 1140 to predict the fifth data set 1115 Whether the sentence is about one or more predetermined topics. The SVM classifier 1140 generates a predicted sentence group 1150, which includes sentences in the data set 1115 and topics predicted for the sentences in the data set 1115. The SVM classifier 1140 will mark the topic predicted for the sentence in the predicted group 1150. The predicted group 1150 includes the reliability scores of one or more topics predicted for the sentences in the data set 1115.
[0138] As shown in FIG. 11, the topic classification and recognition module 340 will use the validator 1160 to compare the test data set 1117 (which is the same as the data set 1115) with the predicted data set 1150 to determine the manually labeled Whether the fifth data set 1115 is the same subject as the subject in the predicted data set. The validator 1160 sorts the data in 1117 that is different from the predicted answer of 1150 according to the confidence value predicted by the SVM, and generates a sorted set 1170. Next, the human data processing personnel will review and correct the inconsistent sets in the sequence of the sorted confidence scores (1180). That is, the human data processing personnel will first review and correct the incorrectly predicted data points (for example, the predicted subject) with the highest confidence score. The manual data processing staff then transmits the corrected data back to the marked data sample file.
[0139] The example of the procedure described in FIG. 11 may be repeated in various groups of the labeled data set 1110. For example, the topic classification and recognition module 340 can divide the labeled data set 1111 into five groups (for example, 11111, 11112, 11113, 11114, and 11115). The topic classification and recognition module 340 can use the aforementioned procedures (1120, 1130, 1149, 1150, 1160, 1170, and 1180), by using the data sets 11111, 11112, 11113, and 11114 as the training data set 1116, and using the data set 11115 as the training data set 1116. The data set 1117 is tested to cross-validate the labeled data set 1111 to verify whether the data set 1111 is correctly labeled.
[0140] Returning to FIG. 10, after verifying and correcting the marked data set, the topic classification and identification module 340 will check the cross-validation results (for example, the correction percentage of topic prediction) to evaluate the SVM prediction and the manually marked sample The accuracy of the comparison of the data set is used to evaluate the quality of the data set (1050). For example, the subject classification and identification module 340 can set a threshold for the cross-validation correction percentage. When the cross-validation between the labeled data set and the predicted set is below the threshold, the topic classification and identification module 340 will sample more input data (1020) and reprocess the sampled data (1030 and 1040). ). If the cross-validation correction percentage reaches a given threshold, the topic classification and identification module 340 will output the marked data set 1060 to the training database 360. Therefore, through the above procedures to test and improve the quality of the training data.
[0141] FIG. 12a shows an example of the opinion mining program 1210 actually operated by the opinion mining and sentiment analysis module 350. The opinion exploration and sentiment analysis module 350 can receive the segmented documents and sentence topics from the word segmentation and integration module 310 (FIG. 3) for further processing. Opinion exploration and sentiment analysis module 350 includes CRF-based opinion words and patterns explorer module (CRF-based opinionwords and patterns explorer module) 1220<sub>o</sub> The opinion word and style detector module 1220 will use the topic style and NE stored in the special noun dictionary 380b (Figure 4) in the CRF-based algorithm to identify opinion words and opinion styles in the word-breaking file And negative words/styles. The opinion word and style detector module 1220 stores the opinion words, opinion styles, and negative words/styles in tables 1222, 1224, and 1226 (which may be part of the training database 360). In each table, the opinion word and style detector module 1220
Words/styles are also classified into: V: (independent verbs), Vd (verbs that need to be followed by opinion words), Ad j (adjectives that need to be followed by opinion words), and Adv (emphasis or reduction of an opinion )adverb. The tables 1222.1224 and 1226 can also store opinions, opinion styles/phrase tendencies marked by human data processing personnel.
[0142] As shown in FIG. 12a, the opinion exploration and sentiment analysis module 350 will be based on the topic patterns stored in the noun dictionary 380b, the opinion words 1222, the opinion patterns/phrases 1224 and the negative words 1226 stored in the database 360. Identify sentences that are topic-based and opinion-based. According to the identified opinion words, opinion styles, and negative words, the opinion mining and sentiment analysis module 350 can use an opinion mining classifier (opinion mining classifier) 1280 to determine whether the opinion in the sentence is positive or negative, and based on the name, ν. , Adj and Adv strength to calculate the opinion decision score (1260), the opinion mining classifier 1280 includes machine learning classifier 1240 (for example, the classifier that actually operates SVM or NaTve Bayes algorithm) and grammar and rule-based classificationDevice1250. The SVM classifier 1140 described in conjunction with the discussion of FIG. 11 is one example of the machine classifier 1240.
[0143] The rule-based classifier 1250 will use one or more plug-in modules containing language patterns and grammatical rules (for example, language patterns stored in the organic object database 380a and the term dictionary 380b (Figure 3)), To help determine the tendency of opinions. The opinion mining classifier 1280 may also calculate the confidence value of the opinion word or opinion style. For opinions or opinion styles with lower reliability scores, manual data processing personnel can be used to view and possibly correct the opinion tendency, and the corrected opinion words or styles are added to the training stored in tables 1222.1224 and 1226 Data collection.
[0144] Next, the opinion mining and sentiment analysis module 350 will calculate the opinion decision score of the paragraph according to the decision score of each sentence in the paragraph (for example, the average score of the sentences in a paragraph). FIG. 12b shows an example of the opinion mining test program actually operated by the opinion mining and sentiment analysis module 350. The test web page 370 is sent to the opinion mining classifier (1240 and 1250) through the word segmentation and integration module 310. According to the identified topic-based and opinion-based sentences 1230, the opinion mining classifiers 1240 and 1250 can determine whether the opinions in the sentences are positive or negative, and calculate opinion decisions based on the strength of Vi, Vd, Adj, and Adv Scoring (1310) ο Next, the opinion mining and sentiment analysis module 350 calculates the opinion decision score of the paragraph according to the decision score of the identified opinion in each sentence of the paragraph (1320). The opinion exploration and sentiment analysis module 350 outputs the opinions associated with sentences and paragraphs and the opinions associated with organic objects to the segmentation and integration module 310 for further processing.
[0145] Please refer to FIG. 3 again, the object relationship construction module 330 constructs two types of relationships: a relationship between a parent object and a child object, and a relationship between two child objects. In an example, the object relationship construction module 330 uses the layout and content of the web page to determine the relationship between the parent object and the child objects. The object relationship construction module 330 may also use a natural language parser (Parser) to analyze the relationship between the two sub-objects.
[0146] The topic classification and identification module 340 (FIG. 8) and the opinion exploration and sentiment analysis module 350 (FIG. 12a) can be actually operated by using a similar software architecture. FIG. 12c provides an example of a software architecture that can be used to actually operate the topic classification and identification module 340 and the opinion exploration and sentiment analysis module 350. As shown in FIG. 12c, the topic classification and identification module 340 or the opinion exploration and sentiment analysis module 350 extracts topics or opinion words based on the topic patterns and opinion words stored in the organic object database 380a and the noun dictionary 380b.
[0147] According to the extracted opinion words and opinion patterns, for example, the opinion mining classifier 1280 can match the opinion words and opinion patterns stored in the opinion word table 1222 or the opinion style table 1224, and according to the opinions stored in the table 1226 Data check negative words or special grammatical rules to process the segmented web pages (by the segmentation and integration module 310)
word). The tables 1222.1224 and 1226 may be part of the training database 360. According to the recognized opinion words, opinion styles, and negative words, the opinion mining and sentiment analysis module 350 can use machine learning classifiers 1240 (for example, classifiers that implement SVM or Naive Bayes algorithm) and grammar and rule-based The opinion of the classifier 1250 explores the classifier 1280 to determine whether the opinion in the sentence is positive or negative, and calculates the opinion decision score according to the strength of Vi, Vd, Adj, and Adv (1260). The rule-based classifier 1250 can use one or more plug-in modules containing language styles and grammatical rules (for example, data stored in the organic object database 380a and the term dictionary 380b (Figure 3)) to help determine opinions. tendency. The opinion mining classifier 1280 may also calculate the confidence value of the opinion word or opinion style. For opinions or opinion styles with lower reliability scores, manual data processing personnel can view and possibly correct the opinion tendency, and the corrected opinion words or styles can be added to the training stored in tables 1222.1224 and 1226 Data collection.
[0148] According to the extracted theme, the theme classifier 870 can check semantic similarity by matching the theme patterns stored in the theme style table 861 and checking the data stored in the theme semantic vector table 862 and the semantic similarity table 863 Sex, to process the segmented web page (word segmentation and integration module 310). The tables 86K862 and 863 may be part of the training database 360. Then, the topic classifier module 870 classifies the topics in the content of the webpage, and detects new topics in the content. Finally, the topic classification and identification module 340 will mark and compose topics related to each sentence on the webpage, and determine the topic of each paragraph according to the topic of the sentence in the paragraph (880). The topic classification and identification module 340 sends the sentence topic and paragraph topic to the word segmentation and integration module 310 for further processing.
[0149] In FIG. 3, the word segmentation and integration module 310 receives and processes input data from all other modules, and stores the extracted organic object data in the organic object database 380a. FIG. 13 shows an example of the word segmentation and integration module 310.
[0150] As shown in FIG. 13, the word segmentation and integration module 310 will use the special noun dictionary 380b (store NE, topic, opinion style, etc.) as the CRF-based word breaker training module 460 and word breaker 470 (see Figure 4) plug-in program to improve the accuracy of word segmentation. The plug-in program of the special noun dictionary 380b will provide the word breaker 470 with NE, topic, and opinion styles to help the word breaker 470 recognize the styles. As described above, the content in the noun dictionary 380b can be updated by the object recognition module 320, the topic classification and recognition module 340, and the opinion exploration module 350 (via the module interface 1330). As shown in FIG. 13, these modules can also send the result of the word segmentation, the found objects, themes, and opinions 1310 to the word segmentation and integration module 310 via the module interface 1330. The integration module 1340 monitors the working status of other modules (1342) and provides updates to other modules (1344). The integration module 1340 also integrates the data (NE, topic, opinion style, etc.) received from other modules via the module interface 1330 into the organic object data model 100, and stores the object data in the specific noun dictionary 380b.
[0151] Those skilled in the art will understand that various modifications and changes can be made in the system and method for extracting community intelligence from online communities and communities. For example, after considering the disclosed embodiments, those skilled in the art will understand that different configurations of the database can be used to store the training data for the organic object data model and the term dictionary. In addition, after considering the disclosed embodiments, those skilled in the art will understand that various machine learning algorithms can be used to identify NE, topics, and opinions defined in the organic object data model. In addition, after considering the disclosed embodiments, those skilled in the art will also understand that the disclosed organic object data model can be applied to information other than online community intelligence (for example, in standby databases or paper publications). Large amounts of data). Moreover, after considering the disclosed embodiments, those skilled in the art will further understand that various software/hardware configurations can be used to implement the disclosed embodiments by using various computer servers, computer storage media, and software applications. Therefore, although the present invention has been disclosed in embodiments
As above, however, it is not used to limit the present invention. Those skilled in the art can make some changes and modifications without departing from the spirit and scope of the present invention. Therefore, the scope of protection of the present invention should be regarded as the appended claims. Those defined shall prevail.
16 sheets
Sheet 1 Sheet 2 Sheet 3 Sheet 4 Sheet 5 Sheet 6 Sheet 7 Sheet 8 Sheet 9 Sheet 10 Sheet 11 Sheet 12 Sheet 13 Sheet 14 Sheet 15 Sheet 16
Every citation, both ways
| Document | Relation | Office | Cited during |
|---|---|---|---|
| US11275753B2 | Cited by | United States of America | Applicant |
| CN105608091A | Cited by | China | Search report |
| US11126489B2 | Cited by | United States of America | Applicant |
| US10942947B2 | Cited by | United States of America | Applicant |
| US11138279B1 | Cited by | United States of America | Applicant |
| US11269906B2 | Cited by | United States of America | Applicant |
| US10484407B2 | Cited by | United States of America | Applicant |
| US11113298B2 | Cited by | United States of America | Applicant |
| US11928211B2 | Cited by | United States of America | Applicant |
| US10452651B1 | Cited by | United States of America | Applicant |
| US10572487B1 | Cited by | United States of America | Applicant |
| US10929436B2 | Cited by | United States of America | Applicant |
| US10706434B1 | Cited by | United States of America | Applicant |
| US10460486B2 | Cited by | United States of America | Applicant |
| US11769096B2 | Cited by | United States of America | Applicant |
| US10628834B1 | Cited by | United States of America | Applicant |
| US10726507B1 | Cited by | United States of America | Applicant |
| US10970261B2 | Cited by | United States of America | Applicant |
| GB2513472A | Cited by | United Kingdom | Search report |
| US11281726B2 | Cited by | United States of America | Applicant |
| US10909159B2 | Cited by | United States of America | Applicant |
| US10795918B2 | Cited by | United States of America | Applicant |
| US10223429B2 | Cited by | United States of America | Applicant |
| US12079357B2 | Cited by | United States of America | Applicant |
| US9619557B2 | Cited by | United States of America | Applicant |
| US12105719B2 | Cited by | United States of America | Applicant |
| US10885021B1 | Cited by | United States of America | Applicant |
| US11599369B1 | Cited by | United States of America | Applicant |
| US11126638B1 | Cited by | United States of America | Applicant |
| US10579950B1 | Cited by | United States of America | Applicant |
| US10360702B2 | Cited by | United States of America | Applicant |
| US11182204B2 | Cited by | United States of America | Applicant |
| US10877654B1 | Cited by | United States of America | Applicant |
| US11302426B1 | Cited by | United States of America | Applicant |
| US10452678B2 | Cited by | United States of America | Applicant |
| US10732803B2 | Cited by | United States of America | Applicant |
| US10853338B2 | Cited by | United States of America | Applicant |
| US11074277B1 | Cited by | United States of America | Applicant |
| US9727622B2 | Cited by | United States of America | Applicant |
| US10728262B1 | Cited by | United States of America | Applicant |
| US10678860B1 | Cited by | United States of America | Applicant |
| US10839144B2 | Cited by | United States of America | Applicant |
| US10474326B2 | Cited by | United States of America | Applicant |
| US10545982B1 | Cited by | United States of America | Applicant |
| US10581954B2 | Cited by | United States of America | Applicant |
| US10728277B2 | Cited by | United States of America | Applicant |
| US11308117B2 | Cited by | United States of America | Applicant |
| US10909130B1 | Cited by | United States of America | Applicant |
| US11907175B2 | Cited by | United States of America | Applicant |
| US11106692B1 | Cited by | United States of America | Applicant |
| US12147657B2 | Cited by | United States of America | Applicant |
| US10924362B2 | Cited by | United States of America | Applicant |
| US10796318B2 | Cited by | United States of America | Applicant |
| US11373752B2 | Cited by | United States of America | Applicant |
| US11526471B2 | Cited by | United States of America | Applicant |
| US10762102B2 | Cited by | United States of America | Applicant |
| US10698938B2 | Cited by | United States of America | Applicant |
| US9501761B2 | Cited by | United States of America | Applicant |
| US10579647B1 | Cited by | United States of America | Applicant |
| US9671776B1 | Cited by | United States of America | Applicant |
| US10956508B2 | Cited by | United States of America | Applicant |
| US10444941B2 | Cited by | United States of America | Applicant |
| US11004039B2 | Cited by | United States of America | Applicant |
| US10437450B2 | Cited by | United States of America | Applicant |
| US11392550B2 | Cited by | United States of America | Applicant |
| US11106701B2 | Cited by | United States of America | Applicant |
| US10523787B2 | Cited by | United States of America | Applicant |
| US10719527B2 | Cited by | United States of America | Applicant |
| US10444940B2 | Cited by | United States of America | Applicant |
| US10346410B2 | Cited by | United States of America | Applicant |
| US11954300B2 | Cited by | United States of America | Applicant |
| US10956406B2 | Cited by | United States of America | Applicant |
| US11507657B2 | Cited by | United States of America | Applicant |
| US10915536B2 | Cited by | United States of America | Applicant |
| US10552994B2 | Cited by | United States of America | Applicant |
| US11250425B1 | Cited by | United States of America | Applicant |
| US11829928B2 | Cited by | United States of America | Applicant |
| US11294928B1 | Cited by | United States of America | Applicant |
| US10795909B1 | Cited by | United States of America | Applicant |
| US10942627B2 | Cited by | United States of America | Applicant |
| US11269931B2 | Cited by | United States of America | Applicant |
| US9760556B1 | Cited by | United States of America | Applicant |
| US10866685B2 | Cited by | United States of America | Applicant |
| US11150629B2 | Cited by | United States of America | Applicant |
| USRE48589E | Cited by | United States of America | Applicant |
| US10360238B1 | Cited by | United States of America | Applicant |
| US11789931B2 | Cited by | United States of America | Applicant |
| US11250027B2 | Cited by | United States of America | Applicant |
| US10754822B1 | Cited by | United States of America | Applicant |
| US10719621B2 | Cited by | United States of America | Applicant |
| US11216762B1 | Cited by | United States of America | Applicant |
| US10509844B1 | Cited by | United States of America | Applicant |
| US10635276B2 | Cited by | United States of America | Applicant |
| US10504067B2 | Cited by | United States of America | Applicant |
| US11035690B2 | Cited by | United States of America | Applicant |
| US9734217B2 | Cited by | United States of America | Applicant |
| US11595492B2 | Cited by | United States of America | Applicant |
| US11934847B2 | Cited by | United States of America | Applicant |
| US9639580B1 | Cited by | United States of America | Applicant |
| US10120857B2 | Cited by | United States of America | Applicant |
10 members in 3 offices
Priority claims10
| Document | Office | Kind | Date |
|---|---|---|---|
| 25549409 | United States of America | P | |
| 25549409 | United States of America | P | |
| 61255494 | United States of America | – | |
| 12801777 | United States of America | – | |
| 80177710 | United States of America | A | |
| 80177710 | United States of America | A | |
| 12801777 | – | – | – |
| 61255494 | – | – | – |
| US20090255494P | – | – | – |
| US20100801777 | – | – | – |
Members10
| Document | Office | Kind | |
|---|---|---|---|
| US2011099133A1 | United States of America | A1 | |
| TW201115370A | Taiwan Province of China | A | |
| TW201115371A | Taiwan Province of China | A | |
| CN102054015AThis record | China | A | |
| CN102054016A | China | A | |
| US2011112995A1 | United States of America | A1 | |
| TWI424325B | Taiwan Province of China | B | |
| CN102054015B | China | B | |
| TWI438637B | Taiwan Province of China | B | |
| CN102054016B | China | B |
3 legal events, as the office reported them to INPADOC
Over the term
Point at a mark for the eventEvents
| Event | Code | |
|---|---|---|
| Grant of patent or utility modelGrantedC14 | C14 | |
| Entry into substantive examinationC10 | C10 | |
| PublicationC06 | C06 |
Numbers
- Publication
- 102054015
- Publication, DOCDB
- 102054015
- Publication, EPODOC
- CN102054015
- Application
- 105269618
- Application, DOCDB
- 201010526961
- Application, EPODOC
- CN20101526961
Titles2
- Chinese
- 使用有机物件数据模型来组织社群智能信息的系统及方法
- English
- System and method for organizing community intelligent information using organic object data model
Classification
- IPC, 1
- G06F17 30