Searching for Relevant Tweets Based on Topic-Related User Activities

Searching for Relevant Tweets Based on Topic-Related User Activities
复制标题

DOI:
--
复制
发表时间:
2016-07
期刊:
J. Web Eng.
影响因子:
--
通讯作者:
T. Noro;T. Tokuda
T. Noro;T. Tokuda
中科院分区:
其他
文献类型:
--
作者:
T. Noro;T. Tokuda

文献摘要

相似文献

Twitter是最大的社交媒体之一。虽然它可以用来获取感兴趣的主题的信息,但由于大量的tweet和每条tweet的小尺寸,我们不容易找到与主题相关的tweet。一些相关推文可能不包括与主题明确相关的任何术语,并且一般的基于内容的关键字搜索技术和查询扩展技术对于查找这样的相关推文是无效的。为了解决这个问题,我们提出了一种方法,用于找到一个感兴趣的主题上的Twitter用户活动的基础上,如推文,转发和回复的主题,推文。该方法包括两个阶段:准备阶段和主要阶段。在准备阶段,我们根据过去与主题相关的用户活动,创建一个表示用户和推文关系的用户-推文参考图,计算每个用户和推文在主题中的影响力,然后定义每个用户的两种力量,称为“Voice”和“Impact”,指示“用户对该主题有多少声音”和“用户对其他用户关于该主题的推文有多少影响”。在主阶段,我们根据发布,转发或回复每条推文的用户的声音和影响力得分来计算新到达的推文与主题的相关性,然后根据相关性得分对推文进行排名。这两个阶段独立处理。一旦准备阶段完成,主阶段可以随时返回最终结果。实验结果表明,“谁转发或回复了该推文”比“谁发布了该推文”更能有效地判断每条推文与主题的相关性,并且我们的方法可以找到不包含任何与主题明确相关的术语的相关推文。我们比较了我们的方法与基于度的方法和基于PageRank的方法,并表明我们的方法优于比较的方法。
Twitter is one of the largest social media. Although it can be used to get information on a topic of interest, it is not easy for us to find tweets relevant to the topic due to a massive amount of tweets and the small size of each tweet. Some relevant tweets may not include any terms explicitly related to the topic, and general content-based keyword search techniques and query expansion techniques are not effective for finding such relevant tweets. To solve this problem, we present a method for finding tweets on a topic of interest based on the Twitter user activities related to the topic such as tweet, retweet, and reply. The method consists of two phases: the preparation phase and the main phase. In the preparation phase, we create a user-tweet reference graph representing the relation between users and tweets based on the past user activities related to the topic, calculate the influence of each user and tweet in the topic, then define two types of each user's power, called "Voice" and "Impact", indicating "how much voice the user has on the topic" and "how much impact the user has on the other users' tweets on the topic". In the main phase, we calculate the relevance of newly-arrived tweets to the topic according to the Voice and the Impact score of the users who posted, retweeted, or replied to each of the tweets, then rank the tweets by the relevance score. The two phases are processed independently. Once the preparation phase is completed, the main phase can return the final result any time. Experimental results show that "who retweeted or replied to the tweet" is more effective for judging the relevance of each tweet to the topic than "who posted the tweet", and our method can find relevant tweets which do not include any terms explicitly related to the topic. We compare our method with an indegree-based method and a PageRank-based method, and show that our method outperforms the methods compared.