Nowcasting Events from the Social Web with Statistical Learning

Nowcasting Events from the Social Web with Statistical Learning
复制标题

DOI:
10.1145/2337542.2337557
复制
发表时间:
2012-01-01
影响因子:
5
通讯作者:
Cristianini, Nello
Cristianini, Nello
中科院分区:
计算机科学3区
文献类型:
--
作者:
Lampos, Vasileios;Cristianini, Nello

文献摘要

被引文献

相似文献

我们提出了一个一般的方法来推断事件或现象的发生和规模,通过探索丰富的非结构化的文本信息的社会部分的Web。有地理标记的用户帖子的微博服务Twitter作为我们的输入数据,我们调查两个案例研究。第一个是基准问题,即从推文的内容推断给定地点和时间的实际降雨量。第二个是现实生活中的任务,我们推断区域流感样疾病的发病率,以及时发现新出现的流行病。我们的分析建立在一个统计学习框架上,该框架通过LASSO的自举版本执行稀疏学习,从大量候选人中选择一致的文本特征子集。在这两个案例研究中,选定的功能表明密切的语义相关性与目标主题和推理,进行回归,具有显着的性能,特别是考虑到短的长度-约一年-Twitter的数据时间序列。
We present a general methodology for inferring the occurrence and magnitude of an event or phenomenon by exploring the rich amount of unstructured textual information on the social part of the Web. Having geotagged user posts on the microblogging service of Twitter as our input data, we investigate two case studies. The first consists of a benchmark problem, where actual levels of rainfall in a given location and time are inferred from the content of tweets. The second one is a real-life task, where we infer regional Influenza-like Illness rates in the effort of detecting timely an emerging epidemic disease. Our analysis builds on a statistical learning framework, which performs sparse learning via the bootstrapped version of LASSO to select a consistent subset of textual features from a large amount of candidates. In both case studies, selected features indicate close semantic correlation with the target topics and inference, conducted by regression, has a significant performance, especially given the short length -approximately one year- of Twitter's data time series.