Microblog Retrieval Using Ensemble of Feature Sets through Supervised Feature Selection

Microblog Retrieval Using Ensemble of Feature Sets through Supervised Feature Selection
复制标题

DOI:
10.1587/transinf.2016dap0032
复制
发表时间:
2017-04
期刊:
IEICE Trans. Inf. Syst.
影响因子:
--
通讯作者:
Abu Nowshed Chy;Md. Zia Ullah;Masaki Aono
Abu Nowshed Chy;Md. Zia Ullah;Masaki Aono
中科院分区:
其他
文献类型:
--
作者:
Abu Nowshed Chy;Md. Zia Ullah;Masaki Aono

文献摘要

相似文献

微博,特别是twitter,已经成为我们日常生活中不可或缺的一部分,人们可以通过微博搜索最新的新闻和事件信息。由于推文长度较短,且频繁使用非常规缩略语,基于内容相关性的搜索不能满足用户的信息需求。最近的研究表明,在这方面考虑时间和上下文方面的检索性能显着提高。在本文中,我们专注于微博检索,强调减轻词汇不匹配,并利用时间(例如,新近度和突发性质)和推文的上下文特征。为了解决时间和上下文方面的推文,我们提出了新的功能的基础上查询推文时间,词嵌入,查询推文的情感相关性。我们还引入了一些流行度特征来估计推文的重要性。一个三阶段的查询扩展技术,以提高相关性的推文。此外,为了确定查询的时间和情感敏感度,我们引入了查询类型确定技术。在有监督的特征选择之后,我们应用随机森林作为特征排序方法来估计所选特征的重要性。然后,我们利用集成的学习排名(L2 R)框架来估计查询推文对的相关性。我们在TREC微博2011和2012测试集上进行了实验,这些测试集是在TREC Tweets 2011语料库上进行的。实验结果表明,我们的方法的有效性在基线和已知的相关工作的精度在30(P@30),平均平均精度(MAP),归一化折扣累积增益在30(NDCG@30),和R精度(R-Prec)指标。关键词:微博搜索,时态信息检索,查询扩展,特征选择,学习排序,时间感知排序
Microblog, especially twitter, has become an integral part of our daily life for searching latest news and events information. Due to the short length characteristics of tweets and frequent use of unconventional abbreviations, content-relevance based search cannot satisfy user’s information need. Recent research has shown that considering temporal and contextual aspects in this regard has improved the retrieval performance significantly. In this paper, we focus on microblog retrieval, emphasizing the alleviation of the vocabulary mismatch, and the leverage of the temporal (e.g., recency and burst nature) and contextual characteristics of tweets. To address the temporal and contextual aspect of tweets, we propose new features based on query-tweet time, word embedding, and query-tweet sentiment correlation. We also introduce some popularity features to estimate the importance of a tweet. A three-stage query expansion technique is applied to improve the relevancy of tweets. Moreover, to determine the temporal and sentiment sensitivity of a query, we introduce query type determination techniques. After supervised feature selection, we apply random forest as a feature ranking method to estimate the importance of selected features. Then, we make use of ensemble of learning to rank (L2R) framework to estimate the relevance of query-tweet pair. We conducted experiments on TREC Microblog 2011 and 2012 test collections over the TREC Tweets2011 corpus. Experimental results demonstrate the effectiveness of our method over the baseline and known related works in terms of precision at 30 (P@30), mean average precision (MAP), normalized discounted cumulative gain at 30 (NDCG@30), and R-precision (R-Prec) metrics. key words: microblog search, temporal information retrieval, query expansion, feature selection, learning to rank, time-aware ranking