Extracting Fine-Grained Location with Temporal Awareness in Tweets: A Two-Stage Approach

Extracting Fine-Grained Location with Temporal Awareness in Tweets: A Two-Stage Approach
复制标题

在推文中提取具有时间感知的细粒度位置:两阶段方法

DOI:
10.1002/asi.23816
复制
发表时间:
2017
期刊:
JOURNAL OF THE ASSOCIATION FOR INFORMATION SCIENCE AND TECHNOLOGY (CCF-B类)
影响因子:
--
通讯作者:
Sun Aixin
Sun Aixin
中科院分区:
其他
文献类型:
--
作者:
Li Chenliang;Sun Aixin

文献摘要

被引文献

相似文献

Twitter已经吸引了数十亿用户进行生活日志记录和分享活动和观点。在他们的推文中,用户经常会透露他们的位置信息和短期访问历史或计划。捕获用户的短期活动可以帮助许多应用程序在正确的时间和地点提供正确的上下文。在本文中,我们感兴趣的是在具有时间感知的细粒度上提取推文中提到的位置。具体地说,我们识别推文中提到的兴趣点(POI),并预测用户是否已经访问、正在访问或即将访问所提到的POI。POI可以是餐厅、购物中心、书店或任何其他细粒度的位置。该框架被称为TS-PETAR(Two-Stage POI Extractor with Time Aware),它由两个主要部分组成:APOI库存和两阶段时间感知POI标签器。POI清单是通过利用Foursquare社区的群体智慧建立的。它既包含POI的正式名称,也包含它们的非正式缩写,这在Foursquare签到中很常见。基于条件随机场(CRF)模型的时间感知POI标记器被设计来消除POI提及的歧义,并相应地解析它们关联的时间感知。三组语境特征(语言、时间和清单特征)和两个标记图式特征(OP和Bilou图式)被用于时间感知POI提取任务。我们的实证研究表明,POI歧义消解的子任务和时间意识消解的子任务需要不同的特征设置才能获得最佳性能。我们还针对几种强基线方法对所提出的TS-Petar进行了评估。实验结果表明,两阶段方法的准确率最高,并且在效率和有效性上都优于所有的基线方法。
Twitter has attracted billions of users for life logging and sharing activities and opinions. In their tweets, users often reveal their location information and short‐term visiting histories or plans. Capturing user's short‐term activities could benefit many applications for providing the right context at the right time and location. In this paper we are interested in extracting locations mentioned in tweets at fine‐grained granularity, with temporal awareness. Specifically, we recognize the points‐of‐interest (POIs) mentioned in a tweet and predict whether the user has visited, is currently at, or will soon visit the mentioned POIs. A POI can be a restaurant, a shopping mall, a bookstore, or any other fine‐grained location. Our proposed framework, named TS‐Petar(Two‐Stage POI Extractor with Temporal Awareness), consists of two main components: aPOI inventoryand atwo‐stage time‐aware POI tagger. The POI inventory is built by exploiting the crowd wisdom of the Foursquare community. It contains both POIs' formal names and their informal abbreviations, commonly observed in Foursquare check‐ins. The time‐aware POI tagger, based on the Conditional Random Field (CRF) model, is devised to disambiguate the POI mentions and to resolve their associated temporal awareness accordingly. Three sets of contextual features (linguistic, temporal, and inventory features) and two labeling schema features (OP and BILOU schemas) are explored for the time‐aware POI extraction task. Our empirical study shows that the subtask of POI disambiguation and the subtask of temporal awareness resolution call for different feature settings for best performance. We have also evaluated the proposed TS‐Petaragainst several strong baseline methods. The experimental results demonstrate that the two‐stage approach achieves the best accuracy and outperforms all baseline methods in terms of both effectiveness and efficiency.