Group extraction device, group extraction method, and program
Abstract
Problem to be solved.To extract a dynamic friend group or a group keyword in consideration of time series.
Solution.A stationary feature of a friend set is calculated from an action history, and based on a stationary feature amount and an action history within a specific period, it does not appear constantly, but appears only within a specific period. The characteristics to be used are calculated, the relationships between friends in the friend set are weighted, and a group of friends who have similar interests or interests within a specific period is extracted from the weighted friend relationships. [Selection diagram] Fig. 2

Term
Projected expiry 30 November 2032.
- Priority and filed
- Published
- Today
- Projected expiry
14 claims: 3 independent, 11 dependent
- 1ソーシャルメディア上の友人集合と友人の行動履歴から、特定の期間の興味あるいは関心が類似するグループを抽出するグループ抽出装置であって、 前記行動履歴から友人集合の定常的な特徴を算出する大域特徴算出部と、 前記定常的な特徴量と特定の期間の行動履歴とに基づいて、定常的に出現せず、特定の期間にのみ出現する特徴を算出し、前記友人集合における友人間の関係性を重み付けする重み算出部と、 前記重み付けされた友人関係から、特定の期間内の興味あるいは関心が類似する友人のグループを抽出するグループ抽出部と、 を備えることを特徴とするグループ抽出装置。
- 2前記グループ毎に、特定の期間内の友人の投稿内容からグループを特徴付けるキーワードを抽出するグループキーワード抽出部を備えたことを特徴とする請求項1に記載のグループ抽出装置。
- 3前記グループキーワードをユーザに提示し、ユーザの興味あるいは関心に合致するグループの指定をユーザから受け付け、指定されたグループに所属する友人の特定の期間内の投稿内容のみを取得し、出力するグループフィルタリング部を備えたことを特徴とする請求項2に記載のグループ抽出装置。
- 4前記グループキーワードのうち、ユーザがプロフィール(属性)情報として登録する興味あるいは関心を表す単語、あるいはユーザの投稿内容から抽出された単語に合致するグループキーワードを友人毎に求めるキーワードマッチング部を備えたことを特徴とする請求項2に記載のグループ抽出装置。
- 5前記行動履歴は、ソーシャルメディア上の友人の投稿内容あるいは、コミュニケーション履歴であり、 前記大域特徴算出部は、投稿内容に定常的に出現する単語群とその出現頻度、および定常的にコミュニケーションを行う友人の組み合わせとその頻度を抽出し、 前記重み算出部は、前記定常的に出現する単語群に比して特定の期間内で出現頻度が向上した単語を、共に投稿内容に含む友人間の重みを高く設定し、更に定常的なコミュニケーション頻度に比して特定の期間内でコミュニケーション頻度が向上した友人間の重みを高く設定することを特徴とする、請求項1から4のいずれかに記載のグループ抽出装置。
- 6前記重み算出部は、さらに、プロフィール属性の類似度が高い友人間の重みを高く設定することを特徴とする、請求項1から4のいずれかに記載のグループ抽出装置。
- 7ソーシャルメディア上の友人集合と友人の行動履歴から、特定の期間内の興味あるいは関心が類似するグループを抽出するグループ抽出装置におけるグループ抽出方法であって、 前記行動履歴から友人集合の定常的な特徴を算出する第1のステップと、 前記定常的な特徴量と特定の期間内の行動履歴を基に、定常的には出現せず、特定の期間内に出現する特徴を算出し、前記友人集合における友人間の関係性を重み付けする第2のステップと、 前記重み付けされた友人関係から、特定の期間内の興味あるいは関心が類似する友人のグループを抽出する第3のステップと、 を備えることを特徴とするグループ抽出方法。
- 8前記グループ毎に、特定の期間内の友人の投稿内容からグループを特徴付けるキーワードを抽出する第4のステップを備えたことを特徴とする請求項7に記載のグループ抽出方法。
- 9前記第4のステップにおいて、前記グループキーワードをユーザに提示し、ユーザの興味あるいは関心に合致するグループの指定をユーザから受け付け、指定されたグループに所属する友人の特定の期間内の投稿内容のみを取得し、出力することを特徴とする請求項8に記載のグループ抽出方法。
- 10前記第4のステップにおいて、前記グループキーワードのうち、ユーザがプロフィール(属性)情報として登録する興味あるいは関心を表す単語、あるいはユーザの投稿内容から抽出された単語に合致するグループキーワードを友人毎に求めることを特徴とする請求項8に記載のグループ抽出方法。
- 11ソーシャルメディア上の友人集合と友人の行動履歴から、特定の期間内の興味あるいは関心が類似するグループを抽出するグループ抽出装置におけるグループ抽出方法をコンピュータに実行させるためのプログラムであって、 前記行動履歴から友人集合の定常的な特徴を算出する第1のステップと、 前記定常的な特徴量と特定の期間内の行動履歴を基に、定常的には出現せず、特定の期間内に出現する特徴を算出し、前記友人集合における友人間の関係性を重み付けする第2のステップと、 前記重み付けされた友人関係から、特定の期間内の興味あるいは関心が類似する友人のグループを抽出する第3のステップと、 をコンピュータに実行させるためのプログラム。
- 12前記グループ毎に、特定の期間内の友人の投稿内容からグループを特徴付けるキーワードを抽出する第4のステップを備えたことを特徴とする請求項11に記載のプログラム。
- 13前記第4のステップにおいて、前記グループキーワードをユーザに提示し、ユーザの興味あるいは関心に合致するグループの指定をユーザから受け付け、指定されたグループに所属する友人の特定の期間内の投稿内容のみを取得し、出力することを特徴とする請求項12に記載のプログラム。
- 14前記第4のステップにおいて、前記グループキーワードのうち、ユーザがプロフィール(属性)情報として登録する興味あるいは関心を表す単語、あるいはユーザの投稿内容から抽出された単語に合致するグループキーワードを友人毎に求めることを特徴とする請求項12に記載のプログラム。
Independent claims14
82 paragraphs, as filed
The present invention extracts interests or account groups with similar interests within a specific period from a set of accounts of others to which users are connected on social media, and further, keywords expressing interests / interests for each group. It relates to a group extraction device, a group extraction method and a program to be extracted.
There is known a technology for providing a community formation system capable of forming a community that includes users who have no contact with each other as members. The outline of the processing procedure of this prior art is as follows: (1) A set of users who frequently communicate with a certain topic is extracted as a cluster, and a cluster pair candidate which is the basis of new cluster formation is selected from a plurality of clusters. Next, (2) detect common topic attributes to characterize the community from user information belonging to the cluster pair (mainly profile attributes such as affiliation and hobbies). (3) A community (corresponding to a group in the present invention) is formed and presented by selectively merging target cluster pairs by topic attribute (see, for example, Patent Document 1).
In addition, there is known a technique of dividing accounts connected on social media into a plurality of groups and extracting keywords representing the characteristics of each group from the posted contents in the group. The outline of the processing procedure of this prior art is as follows: (1) Based on the presence or absence of friendship, a set of accounts with close friendship is extracted as a creek (corresponding to a group in the present invention). Next, (2) increase the importance of the words uttered by the influential accounts in each creek (accounts with many friendships in the group or accounts that seem to be news sources), and comprehensively from the posted contents in the creek. Words with high importance are extracted and presented as keywords (see, for example, Non-Patent Document 1).
<p><patcit num="1"><text>Japanese Unexamined Patent Publication No. 2010-286868</text></patcit><nplcit num="1"><text>Maike Erdmann, Tomoya Takeyoshi, Gen Hattori and Chihiro Ono, Extraction and Annotation of Personal Cliques from Social Networks, The 12th IEEE / IPSJ International Symposium on Application and the Internet (SAINT2012), July, 2012.</text></nplcit></p>
<p> However, in the techniques described in Patent Document 1 and Non-Patent Document 1 described above, topics from the past to the present and topics within a specific period (for example, the latest one week will be described as an example below). Since interests and changes in interests (time series) such as differences are not taken into consideration, groups based on interests and interests within a specific period are not preferentially formed. In addition, since the keywords that characterize the group are based on profile attributes, there is a problem that interests / interests within a specific period are not correctly expressed as keywords.</p><p> Therefore, the present invention has been made in view of the above-mentioned problems, and provides a group extraction device, a group extraction method, and a program capable of dynamically extracting a friend group in consideration of time series or extracting a group keyword. The purpose is to provide.</p>
<p> The present invention proposes the following items in order to solve the above problems. In addition, in order to facilitate understanding, the description will be given with reference numerals corresponding to the embodiments of the present invention, but the present invention is not limited thereto.</p><p> (1) The present invention is a group extraction device that extracts a group having similar interests or interests within a specific period from a friend set on social media and a friend's action history, and the constant state of the friend set from the action history. Appear constantly based on the global feature calculation unit (for example, corresponding to the global feature calculation unit 202 in FIG. 1) that calculates a specific feature, the constant feature amount, and the behavior history within a specific period. A weight calculation unit (for example, corresponding to the weight calculation unit 203 in FIG. 1) that calculates features that appear only within a specific period and weights the relationships between friends in the friend set, and the weighted friend. We propose a group extraction device characterized by including a group extraction unit (for example, corresponding to the group extraction unit 203 in FIG. 1) that extracts a group of friends who have similar interests or interests within a specific period from the relationship. doing.</p><p> According to the present invention, the global feature calculation unit calculates the stationary features of the friend set from the behavior history, and the weight calculation unit is steady based on the steady feature amount and the behavior history within a specific period. The features that do not appear in a specific period but appear only within a specific period are calculated, the relationships between friends in the friend set are weighted, and the group extraction unit uses the weighted friend relationships to determine the interest within a specific period or Extract a group of friends with similar interests. Therefore, by providing the group extraction result, it is possible to overview the interests / interests within a specific period and the friends who have the interests / interests regarding the friends connected on the social media.</p><p> (2) The present invention is a group keyword extraction unit (for example, the group keyword of FIG. 1) for extracting keywords that characterize a group from the contents posted by friends within a specific period for each group of the group extraction device of (1). We are proposing a group extraction device characterized by having an extraction unit 205). </p><p> According to the present invention, the group keyword extraction unit extracts keywords that characterize a group from the posted contents of friends within a specific period for each group. That is, since the keywords that characterize the group are extracted from the posted contents of friends within a specific period for each group, it is possible to extract accurate keywords based on the interests and interests within the specific period.</p><p> (3) The present invention presents the group keyword to the user for the group extraction device of (2), accepts the user's interest or designation of a group that matches the interest, and of a friend who belongs to the designated group. We are proposing a group extraction device characterized by having a group filtering unit (for example, corresponding to the group filtering unit 206 in FIG. 9) that acquires and outputs only the posted content within a specific period. </p><p> According to the present invention, the group filtering unit presents a group keyword to the user, accepts the user's interest or designation of a group that matches the interest, and posts within a specific period of time of friends belonging to the designated group. Get only the content and output it. Therefore, by providing a timeline filter function based on the group extraction result, the user can efficiently follow a certain interest and the posted content based on the interest.</p><p> (4) The present invention relates to the group extraction device of (2) as a word representing interest or interest registered by the user as profile (attribute) information among the group keywords, or a word extracted from the content posted by the user. We are proposing a group extraction device characterized by having a keyword matching unit (for example, corresponding to the keyword matching unit 207 in FIG. 12) for obtaining matching group keywords for each friend.</p><p> According to the present invention, among the group keywords, the keyword matching unit selects a word representing interest or interest registered by the user as profile (attribute) information, or a group keyword matching a word extracted from the user's posted content as a friend. Ask every time. Therefore, if keyword matching is performed, matching friends and keywords are listed, and these lists are presented, it is possible to support the user's destination setting.</p><p> (5) In the present invention, regarding the group extraction devices (1) to (4), the action history is the posted content or communication history of a friend on social media, and the global feature calculation unit is used as the posted content. A group of words that appear constantly and their frequency of appearance, and a combination of friends that communicate constantly and their frequency are extracted, and the weight calculation unit performs a specific period as compared with the group of words that appear constantly. Words that appear more frequently in the post are included in the posted content with a higher weight between friends, and the weight between friends who have improved communication frequency within a specific period is higher than the regular communication frequency. We are proposing a group extraction device characterized by setting. </p><p> According to the present invention, the behavior history is the posted content or communication history of a friend on social media, and the global feature calculation unit constantly appears in the posted content, its frequency of appearance, and constantly. The combination of friends who communicate and their frequency are extracted, and the weight calculation unit includes words that appear more frequently within a specific period than the words that appear constantly, among the friends who both include in the posted content. The weight is set high, and the weight between friends whose communication frequency is improved within a specific period is set higher than that of the regular communication frequency. By performing such processing, it becomes easy to extract a group in which similar topics are transmitted and communication is easily activated. </p><p> (6) The present invention relates to the group extraction devices (1) to (4), wherein the weight calculation unit further sets a high weight between friends having a high degree of similarity in profile attributes. We are proposing a device. </p><p> According to the present invention, the weight calculation unit further sets a high weight between friends having a high degree of similarity in profile attributes. This facilitates the extraction of groups to which friends from similar or the same place of employment or similar or the same school of origin belong.</p><p> (7) The present invention is a group extraction method in a group extraction device that extracts groups having similar interests or interests within a specific period from a group of friends on social media and the behavior history of friends, and is based on the behavior history. Appears constantly based on the first step (for example, corresponding to step S100 in FIG. 2) for calculating the stationary characteristics of the friend set, the stationary characteristics, and the behavior history within a specific period. Instead, the second step (e.g., corresponding to step S200 in FIG. 2), which calculates the features that appear within a specific period and weights the relationships between friends in the friend set, and the weighted friend relationships. Proposes a group extraction method characterized by comprising a third step (for example, corresponding to step S300 in FIG. 2) to extract a group of friends who have similar interests or interests within a specific period. There is. </p><p> According to the present invention, the stationary characteristics of a friend set are calculated from the behavior history, and based on the stationary features and the behavior history within a specific period, they do not appear constantly and within a specific period. Only the features that appear are calculated, the relationships between friends in the friend set are weighted, and a group of friends with similar interests or interests within a specific period is extracted from the weighted friend relationships. Therefore, by providing the group extraction result, it is possible to overview the interests / interests within a specific period and the friends who have the interests / interests regarding the friends connected on the social media.</p><p> (8) The present invention relates to the group extraction method of (7), in which a fourth step (for example, step S400 in FIG. 2) is used to extract keywords that characterize a group from the contents posted by friends within a specific period for each group. We are proposing a group extraction method characterized by having (corresponding to). </p><p> According to the present invention, further, for each group, keywords that characterize the group are extracted from the posted contents of friends within a specific period. That is, since the keywords that characterize the group are extracted from the posted contents of friends within a specific period for each group, it is possible to extract accurate keywords based on the interests and interests within the specific period.</p><p> (9) Regarding the group extraction method of (8), the present invention presents the group keyword to the user in the fourth step, accepts the user's interest or designation of a group that matches the interest, and designates the group. We are proposing a group extraction method that is characterized by acquiring and outputting only the posted content of friends who belong to the same group within a specific period. </p><p> According to the present invention, a group keyword is presented to a user, a group designation that matches the user's interest or interest is accepted from the user, and only the posted content of a friend belonging to the designated group within a specific period is acquired. ,Output. Therefore, by providing a timeline filter function based on the group extraction result, the user can efficiently follow a certain interest and the posted content based on the interest.</p><p> (10) The present invention relates to the group extraction method of (8), in the fourth step, among the group keywords, a word representing interest or interest registered by the user as profile (attribute) information, or a user's post. We are proposing a group extraction method characterized by asking each friend for a group keyword that matches the word extracted from the content.</p><p> According to the present invention, among the group keywords, a word representing interest or interest registered by the user as profile (attribute) information, or a group keyword matching a word extracted from the content posted by the user is sought for each friend. Therefore, if keyword matching is performed, matching friends and keywords are listed, and these lists are presented, it is possible to support the user's destination setting.</p><p> (11) The present invention is a program for causing a computer to execute a group extraction method in a group extraction device that extracts groups having similar interests or interests within a specific period from a group of friends on social media and the behavior history of friends. The first step (for example, corresponding to step S100 in FIG. 2) for calculating the stationary characteristics of the friend set from the behavior history, and the stationary characteristics and the behavior history within a specific period are Based on this, a second step (for example, corresponding to step S200 in FIG. 2) of calculating features that do not appear constantly but appear within a specific period and weighting the relationships between friends in the friend set. To have the computer perform a third step (e.g., corresponding to step S300 in FIG. 2) of extracting a group of friends with similar interests or interests within a specific period from the weighted friendship. We are proposing a program for. </p><p> According to the present invention, the stationary characteristics of a friend set are calculated from the behavior history, and based on the stationary features and the behavior history within a specific period, they do not appear constantly and within a specific period. Only the features that appear are calculated, the relationships between friends in the friend set are weighted, and a group of friends with similar interests or interests within a specific period is extracted from the weighted friend relationships. Therefore, by providing the group extraction result, it is possible to overview the interests / interests within a specific period and the friends who have the interests / interests regarding the friends connected on the social media.</p><p> (12) The present invention corresponds to the fourth step (for example, step S400 in FIG. 2) of extracting the keywords that characterize a group from the posted contents of friends within a specific period for each of the groups in the program of (11). ) Is provided. </p><p> According to the present invention, further, for each group, keywords that characterize the group are extracted from the posted contents of friends within a specific period. That is, since the keywords that characterize the group are extracted from the posted contents of friends within a specific period for each group, it is possible to extract accurate keywords based on the interests and interests within the specific period.</p><p> (13) The present invention presents the group keyword to the user in the fourth step of the program (12), accepts the user's interest or designation of a group that matches the interest, and the designated group. We are proposing a program that is characterized by acquiring and outputting only the posted content of a friend who belongs to the company within a specific period. </p><p> According to the present invention, a group keyword is presented to a user, a group designation that matches the user's interest or interest is accepted from the user, and only the posted content of a friend belonging to the designated group within a specific period is acquired. ,Output. Therefore, by providing a timeline filter function based on the group extraction result, the user can efficiently follow a certain interest and the posted content based on the interest.</p><p> (14) The present invention relates to the program of (12) from the words expressing interest or interest that the user registers as profile (attribute) information among the group keywords in the fourth step, or the content posted by the user. We are proposing a program that features group keywords that match the extracted words for each friend. </p><p> According to the present invention, among the group keywords, a word representing interest or interest registered by the user as profile (attribute) information, or a group keyword matching a word extracted from the content posted by the user is sought for each friend. Therefore, if keyword matching is performed, matching friends and keywords are listed, and these lists are presented, it is possible to support the user's destination setting.</p>
<p> According to the present invention, by providing the group extraction result, it is possible to give an overview of the interests / interests of friends connected on social media within a specific period and the friends who have the interests / interests. There is. In addition, by providing a timeline filter function based on the group extraction result, there is an effect that the user can efficiently follow a certain interest and the posted content based on the interest.</p><p> Furthermore, by presenting the group extraction results (presentation of group composition accounts and keywords), it is possible to provide new awareness to known friends, which has the effect of activating communication. In addition, when a user shares or sends a certain content (example: Web news), it is possible to support the user's destination setting by enabling sharing and sending to only the above group members. There is an effect.</p>
<figref num="1">It is a figure which shows the structure of the group extraction apparatus which concerns on 1st Embodiment of this invention.</figref><figref num="2">It is a figure which shows the process of the group extraction apparatus which concerns on 1st Embodiment of this invention.</figref><figref num="3">It is a figure which shows the process of weight calculation which concerns on 1st Embodiment of this invention.</figref><figref num="4">It is a figure which shows the calculation process of the similarity degree of the posted content which concerns on 1st Embodiment of this invention.</figref><figref num="5">It is a figure which shows the calculation process of the communication frequency which concerns on 1st Embodiment of this invention.</figref><figref num="6">It is a figure which shows the calculation process of the attribute similarity which concerns on 1st Embodiment of this invention.</figref><figref num="7">It is a figure which shows the outline of this invention.</figref><figref num="8">It is a figure which shows the output content of this invention.</figref><figref num="9">It is a figure which shows the structure of the group extraction apparatus which concerns on 2nd Embodiment of this invention.</figref><figref num="10">It is a figure which shows the process of the group extraction apparatus which concerns on 2nd Embodiment of this invention.</figref><figref num="11">It is a figure which shows the relationship of input / output which concerns on 2nd Embodiment of this invention.</figref><figref num="12">It is a figure which shows the structure of the group extraction apparatus which concerns on 3rd Embodiment of this invention.</figref><figref num="13">It is a figure which shows the process of the group extraction apparatus which concerns on 3rd Embodiment of this invention.</figref><figref num="14">It is a figure which shows the structure of the group extraction apparatus which concerns on 4th Embodiment of this invention.</figref><figref num="15">It is a figure which shows the process of the group extraction apparatus which concerns on 4th Embodiment of this invention.</figref>
<Outline of the present invention> Before the detailed description of the embodiment, the outline of the present invention will be described with reference to FIG. 7. First, the activities (posts) on social media included in the latter period are divided into the specific period (for example, the last week) targeted for group extraction and the past to before the group extraction period (see Fig. 7). And communication) From the history, the topical words that appear constantly and the features that represent the friends who communicate constantly are extracted.
Next, in order to make it possible to extract friends with similar interests / interests within a specific period as members of the same group, the following two types (of the following two types among friends from the comparison of the above global characteristics and the characteristics within a specific period ( Or 3 types) of index values are calculated, and the weight of friendship is determined using these index values.
As parameters that determine the weight of friendship, the similarity of posted contents: S (A, B), the communication frequency: F (A, B), or the similarity of profile attributes: P (A, B) can be considered.
Post content similarity: S (A, B) is the similarity between post content between Mr. A and Mr. B. If the posted content is similar within a specific period, increase the weight of friendship. At that time, if a word that does not appear constantly and appears frequently within a specific period appears in both the posted contents of Mr. A and Mr. B, the weight of the friendship is particularly increased. This makes it easier for friends who mention the same topic to be included in the same group.
Communication frequency: F (A, B) is the communication frequency between Mr. A and Mr. B. If the frequency of communication within a certain period is high, the weight of friendship is increased. This makes it easier for friends who communicate (= show interest) under a certain topic to be included in the same group. For example, if Mr. B responds to Mr. A's post "I was impressed by the match of the Japanese national futsal team", I think that Mr. B is also interested in "futsal" or "Japan national team".
Profile attribute similarity: P (A, B) is the similarity between Mr. A and Mr. B for profile attributes. If the profile attributes have a high degree of similarity, the weight of friendship is increased. If you want to add attributes such as work place and school of origin to a group of friends, adopt the similarity of profile attributes and indicate the degree of importance for weighting. Similarity of posted content: S (A, B), Communication frequency: F ( Similar to A, B), or similar to posted content: S (A, B), communication frequency: higher than F (A, B) (that is, γ> for α, β, γ described later = α, β).
For example, "I want to divide the same interests and interests such as soccer into a group of only colleagues and a group of university classmates as much as possible." However, in an environment where profile attributes can be acquired, it is possible to perform filtering using only this, so it is positioned as an option in the present invention.
From the above values, the "friendship weight" W (A, B) between Mr. A and Mr. B is determined by Eq. (1). W (A, B) = α × S (A, B) + β × F (A, B) + γ × P (A, B) ... (1) Here, α, β, γ (all 0) The above) indicates the degree of influence on the weight of each index value, and can be freely set when the present invention is provided as a service or an application.
Specifically, the following three types of patterns and effects on the extracted groups can be considered. High degree of influence of α: Emphasis is placed on the similarity of posted content, and it becomes easier to extract groups that send similar topics. High degree of influence of β: Focusing on the occurrence of communication, it becomes easier to extract groups in which communication is likely to be activated. High degree of influence of γ: Emphasis is placed on the similarity of profiles, and it becomes easier to extract groups to which friends who have similar / same workplaces or schools of origin belong.
Next, from the weighted friendships (social graph), the number of friendships in the group or the group having a high total weight is obtained. Then, the keywords that characterize the group are extracted from the posted contents of the friends who belong to each group. At this time, topical words that do not appear constantly but appear within a specific period are preferentially selected.
Hereinafter, embodiments of the present invention will be described in detail with reference to the drawings. The components in the present embodiment can be appropriately replaced with existing components and the like, and various variations including combinations with other existing components are possible. Therefore, the description of this embodiment does not limit the content of the invention described in the claims.
<First Embodiment> An embodiment of the present invention will be described with reference to FIGS. 1 to 8.
<Structure of Group Extractor> As shown in FIG. 1, the group extractor according to the present embodiment includes social media, API 100, data collection unit 201, global feature calculation unit 202, weight calculation unit 203, and so on. It is composed of a group extraction unit 204, a group keyword extraction unit 205, a profile DB211, a posting content DB212, a communication history DB213, a friendship DB214, and a display unit 300.
Social media is a service that is based on relationships between accounts (eg friendships). API100 is an interface provided by various social media that enables data on social media to be acquired from the outside. The display unit 300 presents the output information to the user.
The data collection unit 201 collects necessary data from various social media through API100. Global feature calculation unit 202. Global features are calculated from the collected and accumulated data. Here, the global feature refers to a feature that is not subject to group extraction in time series and is calculated based on data accumulated from the past (period and data corresponding to the past in FIG. 7). .. Group extraction targets data within a specific time period (eg, last week).
The weight calculation unit 203 calculates the weight for each friendship on the social graph from the global feature and the data within a specific period. The group extraction unit 204 extracts a dynamic friend group from the social graph and the weight of the friendship. For each group, the group keyword extraction unit 205 extracts keywords that characterize the group from the posted contents of friends in the group.
Profile DB211 is a database that stores profile (attribute) information registered by a user's friend on social media. The posted content DB212 is a database that stores the content posted on social media by a user's friend. The communication history DB 213 is a database that stores a history of user's friends reacting on social media. Here, the reaction means, for example, a comment input action or a "like" button pressing action in the case of Facebook. The friendship DB 214 is a database that stores a social graph (a network showing who and who are friends).
<Processing of Group Extractor> The processing of the group extraction device according to the present embodiment will be described with reference to FIGS. 2 to 6.
Figure 1 shows the major flow of processing in the group extraction device. According to this figure, first, the global feature calculation unit 202 calculates the global feature amount from various data accumulated in the DB group (step S100 in FIG. 2). The data to be calculated is the data obtained by excluding the data within a specific period to be extracted by the group from all the data accumulated from the past (the period and data corresponding to the "past" in FIG. 7). By comparing the feature amount calculated here with the data within a specific period, the weight of the friendship is determined in step S200 of FIG.
Here, in step S100, the global features to be calculated are A) global features for each word, B) global features for each friend, and A) global features for each word, the number of posts df (i) in which the word i appears. ), That is, the total number of posts in which a certain word appears in the past data is totaled for each word.
On the other hand, in B) the global feature of each friend, the number of posts by a friend A who is a user P (A), that is, the total number of posts by a friend is calculated for each friend based on past data. Also, the total number of posts by a common friend between the user and a friend A G (A), the total number of communications by Mr. A with respect to the posts by a common friend between the user and a friend A, C (A), and others (for example, Find the number of times C (A, B) of A's communication to the post of common friend B). Here, the common friend refers to, for example, a common friend between the user and Mr. A when focusing on the user's friend A. At this time, the total number of communications made by Mr. A for the posts of common friends is totaled and calculated. (Between Mr. A and Mr. B, between Mr. A and Mr. C, ...)
In step S200, the weight calculation unit 203 determines the weight of the friendship by comparing the global feature calculated in step S100 with the data within a specific period.
Here, the process in step S200 will be described in detail with reference to FIG. First, in step S210, friendships are extracted. Specifically, for each friend of the user, an explicit friendship on social media is acquired from the friendship DB, and a pair of friends who are friends is generated. For example, if the user's friends A, B, C, and D are present, the following pairs of friends are listed. (Hereafter, the list below is referred to as List 1) 1) Mr. A and Mr. B 2) Mr. A and Mr. C 3) Mr. B and Mr. C
In step S220, the similarity of the posted contents is calculated. The specific contents will be described with reference to FIG. First, in step S221, a target friend list is created. That is, a list of friends to be processed in step S220 is generated from the friend relationships (pairs of friends) acquired in step S210. For example, Mr. A, B, and Mr. C are extracted from List 1 to generate a list (hereinafter, this list is referred to as List 2).
In step S222, if there is a friend who has not performed the process from step S223 to step S226, the process proceeds to step S223. If it does not exist, the process proceeds to step S227. In step S223, the posted content of the unprocessed friend is extracted. That is, the posted content (text) within a specific period of the target friend is extracted from the posted content DB.
In step S224, morphological analysis is performed. That is, morphological analysis is applied to the posted content to acquire the words that appear in the posted content. In step S225, the importance of the word is calculated. That is, for each word, the importance w (i, A) of a certain word i in the post of friend A is calculated by comparing with the global feature. The calculation formula is as follows. w (i, A) = df (i, A) × idf (i) ... (2) Where df (i, A) is the word i in the content posted by friend A within a specific period. Represents the number of posts including, idf (i) is the Inversed Document Frequency calculated by the following formula using the above-mentioned df (i). idf (i) = log {| D | / df (i)} ... (3) where | D | is the total number of posts by the user's friends in the past data. If there is a possibility that df (i) = 0, read df (i) = 1 or replace df (i) with df (i) + 1.
In step S226, a feature vector is generated. Specifically, a feature vector is generated for each friend from the word i appearing in the posted content within a specific period of friend A and the importance w (i, A) of the word. In step S227, if there is a friendship (pair) that has not been processed from step S228 to step S229, the process proceeds to step S228. If it does not exist, this processing flow ends.
In step S228, we get the feature vector of the unprocessed 2 friends. That is, for the unprocessed friend relationships (pairs), the feature vectors of each target friend (generated from step S223 to step S226) are acquired.
In step S229, the similarity between the feature vectors is calculated. Specifically, the similarity between the feature vectors is calculated and used as the similarity S (A, B) of the posted content. For example, the cosine similarity corresponding to the distance between the vectors is calculated. This similarity corresponds to the similarity of the posted content between two friends.
In step S230, the communication frequency is calculated. This detail will be described with reference to FIG. First, in step S231, if there is a friendship (pair) that has not been processed from step S232 to step S233, the process proceeds to step S232. If it does not exist, this processing flow ends.
In step S232, the communication history / global features of the unprocessed 2 friends are acquired. Specifically, let's assume that the two unprocessed friends are Mr. A and Mr. B. From past data, the above-mentioned global features P (A), P (B), C (A), C (B), C (A, B), C (B, A), G (A), G ( B) Get. Next, the numerical value R (A, B) representing the rarity of communication from Mr. A to Mr. B in the past data is calculated by the following formula. R (A, B) = (C (A) / G (A)) × (C (A, B) / P (B)) ... (4)
In step S233, the frequency of communication between two friends is calculated. Here, the communication frequency F (A, B) between two friends is calculated by the following formula. F (A, B) = (iR (A, B) x R'(A, B) + iR (B, A) x R'(B, A)) / 2 ... (5) Where iR (A, B) represents the reciprocal of equation (4), and R'(A, B) represents the value obtained by calculating R (A, B) using only the posts of friends within a specific period. When applying equations (4) and (5), if there is a possibility that the denominator will be 0, read it as 1 or add 1 by default to apply.
In step S241, the similarity of attributes is calculated. The specific contents will be described with reference to FIG.
In step S241, if there is a friendship (pair) that has not been processed from step S242 to step S244, the process proceeds to step S242. On the other hand, if it does not exist, this processing flow is terminated.
In step S242, the attribute information of the unprocessed 2 friend is acquired. In other words, if the unprocessed 2 friends are Mr. A and Mr. B, the profile and attribute information of both will be acquired from the profile DB.
In step S243, an attribute vector for each friend is generated. Specifically, the profile attributes of Mr. A and Mr. B are vectorized according to the available profiles and attribute items. For example, if the gender, place of origin, and place of employment can be obtained, the attribute vector is set to "male, XX prefecture, XX company".
In step S244, the similarity between the attribute vectors is calculated. For example, the similarity between the attribute vectors of Mr. A and Mr. B is calculated and used as the profile attribute similarity P (A, B). At this time, the Simpson coefficient may be calculated to obtain the similarity P (A, B) of the profile attributes.
In step S250, the weight of the friendship is output. That is, substituting S (A, B), F (A, B), and P (A, B) calculated in steps S220 to S240 into equation (1), and the friendship between Mr. A and Mr. B. Weight W (A, B).
In step S300, the group is extracted. The group extraction unit 204 weights the friendship on the social graph with the value output in step S200, and dynamically extracts the group from the weighted graph. Here, extracting a group having a high total number of edges and weights in the group corresponds to extracting a group having similar interests and interests within a specific period. For example, by applying the method described in the literature "A. Clauset, MEJ Newman and C. Moore," Finding community structure in very large networks, "Phys. Rev. E70,066111 (2004)." Such groups can be extracted.
In step S400, the group keywords are extracted. The group keyword extraction unit 205 extracts the keywords that characterize the group from the posted contents of the friends in the group within a specific period for each group extracted in step S300. For example, if Mr. A belongs to the group to be calculated, the importance w (i, A) of the words is totaled for each word included in the posted content within the specific period of Mr. A, and the value is high. Extract words as group keywords.
That is, as shown in Fig. 8, accounts with similar affiliations and interests are grouped from the structure of the social graph, and keywords that characterize the group are extracted from the posted contents of friends who belong to the group. Therefore, by automatically generating a friend list, it is possible to check the posted content for each account with similar interests.
As described above, according to the present embodiment, by providing the group extraction result, an overview of the interests / interests within a specific period and the friends who have the interests / interests regarding the friends connected on social media. can do. In addition, since the keywords that characterize the group are extracted from the posted contents of friends within a specific period for each group, it is possible to extract accurate keywords based on the interests and interests within the specific period.
<Second Embodiment> The embodiment of the present invention will be described with reference to FIGS. 9 to 11.
<Structure of Group Extractor> As shown in FIG. 9, the group extractor according to the present embodiment includes social media, API 100, data collection unit 201, global feature calculation unit 202, weight calculation unit 203, and so on. It is composed of a group extraction unit 204, a group keyword extraction unit 205, a group filter unit 206, a profile DB211, a post content DB212, a communication history DB213, a friendship DB214, and a display unit 300. Since the components having the same reference numerals as those in the first embodiment have the same functions, detailed description thereof will be omitted.
The group filter unit 206 acquires and presents only the posted contents of friends who belong to the group specified by the user. In addition, this embodiment provides a function of filtering a social stream (flow of update information of a friend) that can be viewed by a user based on the group acquired in the first embodiment. A specific example from the group designation to the filtering application result is shown in FIG. 11 as a screen image.
<Processing of Group Extractor> The processing of the group extraction device according to the present embodiment will be described with reference to FIG.
First, the processing contents of steps S100 to S400 are the same as those in the first embodiment.
In step S500, the group designation is accepted. That is, in response to the user's filtering request, a group keyword list is presented for each group extracted in step S400, and the user can select a keyword to specify the group. Here, when a plurality of group keywords exist, a predetermined number (example: 1) is presented, and the importance adopted at the time of group keyword extraction is used as the selection criterion.
In step S600, the group filter is applied. Specifically, only the posted contents of friends belonging to the group specified by the user are acquired and presented.
As described above, according to the present embodiment, the function of filtering the social stream (flow of updated information of a friend) that can be viewed by the user is provided based on the group acquired in the first embodiment. By selecting the extracted and presented group keywords, the user can select a group that matches his / her interests and interests, and can view only the posted contents of friends who belong to the specified group. ..
<Third Embodiment> An embodiment of the present invention will be described with reference to FIGS. 12 and 13.
<Structure of Group Extractor> As shown in FIG. 12, the group extractor according to the present embodiment includes social media, API 100, data collection unit 201, global feature calculation unit 202, weight calculation unit 203, and so on. It is composed of a group extraction unit 204, a group keyword extraction unit 205, a keyword matching unit 207, a profile DB211, a posting content DB212, a communication history DB213, a friendship DB214, and a display unit 300. Since the components having the same reference numerals as those in the first embodiment have the same functions, detailed description thereof will be omitted.
The keyword matching unit 207 acquires keywords expressing the user's interests / interests from the user's profile information or the user's posted contents. In this embodiment, the group acquired in the first embodiment and the group keywords are used to label the interests / interests of friends within a specific period, and the friends who are likely to match the interests / interests of the user and their friends. Provide a function to present the reason (group keyword).
<Processing of Group Extractor> The processing of the group extraction device according to the present embodiment will be described with reference to FIG.
First, the processing contents of steps S100 to S400 are the same as those in the first embodiment.
In step S700, the user keyword is acquired. Specifically, keywords expressing the user's interests are acquired from the user's profile information or the user's posted content. If the user has registered a specific keyword (for example, a proper noun such as an artist name) for a category such as "hobby", "interest", or "interest", the corresponding keyword is listed. In addition, if the user's posted content can be obtained, the posted content is converted to the posted content by the same process as when calculating the importance of the word for the posted content of the friend (steps S224 to S225 in FIG. 4). Assign importance to each included word, and use the word with high importance as a keyword.
In step S800, keyword matching is performed. Specifically, keywords that express the interests of friends are assigned to each of the friends who belong to the group. For example, when a word extracted as a group keyword is included in the posted content within a specific period, the word is assigned to the friend as a keyword. Next, keywords are matched between the user and each friend, and the matched friends and keywords are listed.
In step S900, a pair of friends and keywords is presented. That is, in step S800, a list of matching friends and keywords is presented.
As described above, according to the present embodiment, the interest / interest of the user is labeled with respect to the friend within a specific period from the group acquired in the first embodiment and the group keyword. By providing a function to present friends who are likely to match and the reason (group keyword), it is possible to easily know friends who can be expected to activate communication with themselves while their interests and interests change. It will be possible.
<Fourth Embodiment> An embodiment of the present invention will be described with reference to FIGS. 14 and 15.
In this embodiment, the group extraction means and the group keyword extraction means shown in the first embodiment are accumulated while shifting the target period for group extraction (for example, one week ago, two weeks ago, ...). Apply to completed data and acquire the group extraction result for each period. Furthermore, by evaluating the burst property (degree of excitement) of each group, groups related to special events (for example, soccer World Cup and friend's anniversary) that occur and are shared among friends are extracted and listed. That is, in the first to third embodiments, groups are extracted from a specific period focusing on common topics and improvement of communication frequency, but in this embodiment, individuals or the general public are further extracted. When there is a group related to a topic that is particularly exciting in, it is characterized by extracting it as a special event-related group.
<Structure of Group Extractor> As shown in FIG. 14, the group extractor according to the present embodiment includes social media, API 100, data collection unit 201, global feature calculation unit 202, weight calculation unit 203, and so on. It is composed of a group extraction unit 204, a group keyword extraction unit 205, a group evaluation unit 208, a profile DB211, a posting content DB212, a communication history DB213, a friendship DB214, and a display unit 300. Since the components having the same reference numerals as those in the first embodiment have the same functions, detailed description thereof will be omitted.
The group evaluation unit 208 determines the burst property (excitement) from the extracted data of each group, and determines that the particularly exciting group is an event-linked group.
<Processing of Group Extractor> The processing of the group extraction device according to the present embodiment will be described with reference to FIG.
First, the processing contents of steps S100 to S400 are the same as those in the first embodiment.
In step S1100, if there is a period in which the group extraction process has not been executed, the process proceeds to step S1200, and if it does not exist, the present process flow is terminated.
In step S1200, the data is divided. Specifically, the data related to the period to be extracted by the group and the data related to the period to be calculated by the global feature are divided. The division unit may be arbitrary (for example, one week) and may be specified by the service or application provider or the user.
In step S1300, the keyword group is evaluated. Specifically, the burst property (excitement) is determined from the extracted data of each group, and the particularly exciting group is determined to be an event-linked group. Here, threshold processing is applied to the determination of excitement, for example, "when the value of the excitement index of a certain group is X times or more the average value of the excitement index calculated from all groups, it is determined to be an event-related group. , Etc. are performed. As the excitement index, for example, any one of "total number of posts in the group" and "total number of communications in the group", or a combination of a plurality of them is determined.
As described above, according to the present embodiment, the group extraction shown in the first embodiment is performed while shifting the target period of the group extraction (for example, one week ago, two weeks ago, ...). By applying the means and group keyword extraction means to the accumulated data, acquiring the group extraction results in each period, and evaluating the burst property (degree of excitement) of each group, it was generated and shared among friends. By extracting and enumerating groups related to special events (for example, World Cup Soccer and Friend's Anniversary), we will give an overview of groups related to individual (friends) or special events that occur in the general public. It will be possible.
The group extraction device of the present invention is realized by recording the processing of the group extraction device on a recording medium readable by a computer system, causing the group extraction device to read the program recorded on the recording medium, and executing the program. Can be done. The computer system referred to here includes hardware such as an OS and peripheral devices.
In addition, the "computer system" shall include the homepage provision environment (or display environment) if a WWW (World Wide Web) system is used. Further, the program may be transmitted from a computer system in which this program is stored in a storage device or the like to another computer system via a transmission medium or by a transmission wave in the transmission medium. Here, the "transmission medium" for transmitting a program refers to a medium having a function of transmitting information, such as a network (communication network) such as the Internet or a communication line (communication line) such as a telephone line.
Further, the above program may be for realizing a part of the above-mentioned functions. Further, it may be a so-called difference file (difference program) that can realize the above-mentioned function in combination with a program already recorded in the computer system.
Although the embodiments of the present invention have been described in detail with reference to the drawings, the specific configuration is not limited to this embodiment, and includes designs and the like within a range that does not deviate from the gist of the present invention.
100; API 201; Data collection unit 202; Global feature calculation unit 203; Weight calculation unit 204; Group extraction unit 205; Group keyword extraction unit 206; Group filter unit 207; Keyword matching unit 208: Group evaluation unit 211; Profile DB 212 Posted content DB 213; Communication history DB 214; Friendship DB 300; Display
16 sheets
Sheet 1 Sheet 2 Sheet 3 Sheet 4 Sheet 5 Sheet 6 Sheet 7 Sheet 8 Sheet 9 Sheet 10 Sheet 11 Sheet 12 Sheet 13 Sheet 14 Sheet 15 Sheet 16
Every citation, both ways
| Document | Relation | Office | Cited during |
|---|---|---|---|
| CN104965878A | Cited by | China | Search report |
| JP2021043878A | Cited by | Japan | Search report |
| US10509835B2 | Cited by | United States of America | Applicant |
| JPWO2023013122A1 | Cited by | Japan | Search report |
| WO2023013122A1 | Cited by | World Intellectual Property Organization (WIPO) | International search |
| JP2022134801A | Cited by | Japan | Search report |
| JP2021026551A | Cited by | Japan | Search report |
| JP2021500659A | Cited by | Japan | Search report |
| JP2009140075A | Cites | Japan | Examiner |
| US2009164926A1 | Cites | United States of America | Examiner |
2 members in 1 office
Members2
| Document | Office | Kind | |
|---|---|---|---|
| JP2014106943AThis record | Japan | A | |
| JP6042188B2 | Japan | B2 |
9 legal events, as the office reported them to INPADOC
Over the term
Point at a mark for the eventEvents
| Event | Code | |
|---|---|---|
| Cancellation because of no payment of annual feesLAPS | LAPS | |
| Certificate of patent or registration of utility modelJAPANESE INTERMEDIATE CODE: R150R150 | R150 | |
| First payment of annual fees (during grant procedure)JAPANESE INTERMEDIATE CODE: A61A61 | A61 | |
| Written decision to grant a patent or to grant a registration (utility model)JAPANESE INTERMEDIATE CODE: A01A01 | A01 | |
| Decision of grant or rejection writtenTRDD | TRDD | |
| Request for written amendment filedJAPANESE INTERMEDIATE CODE: A523A521 | A521 | |
| Notification of reasons for refusalJAPANESE INTERMEDIATE CODE: A131A131 | A131 | |
| Report on retrievalJAPANESE INTERMEDIATE CODE: A971007A977 | A977 | |
| Written request for application examinationJAPANESE INTERMEDIATE CODE: A621A621 | A621 |
Numbers
- Publication
- 2014106943
- Application
- 262005
Titles2
- Japanese
- グループ抽出装置、グループ抽出方法およびプログラム
- English
- Group extraction device, group extraction method and program
Classification
- IPC, 4
- G06F17 30
- G06F13 00
- G06F15 00
- G06Q50 10