CAREER: Scaling up Modeling and Statistical Inference for Massive Collections of Time Series
CAREER: Scaling up Modeling and Statistical Inference for Massive Collections of Time Series
批准号:
1350133
负责人:
Emily Fox
金额:
$54.92万
依托单位:
依托单位国家:
美国
项目类别:
Continuing Grant
财政年份:
2014
资助国家:
美国
项目状态:
已结题
起止时间:
2014-06-15 至 2021-05-31
中文摘要
考虑一下在一组非常大的空间位置预测流感发病率的任务。单独对每个区域进行建模不会利用相关区域的信息,可能会导致预测不佳,特别是在存在遗漏观测的情况下。同样,想象一下估算美国每一栋房子的价值。捕捉社区内的趋势是关键;然而,每个社区最近只有几次房屋销售。这些日益普遍的海量时间序列所带来的挑战是广泛应用的普遍问题,从警察资源分配的犯罪建模到预测消费趋势和社交网络:个别数据流往往只包括不频繁的观察,因此每个数据流本身都不能提供足够的数据来进行准确的推断。然而,它们之间的结构化关系提供了共享信息的机会。一个关键问题是如何发现这些关系。这个项目采用了一种计算驱动的贝叶斯非参数方法,权衡了灵活性和可伸缩性,以解决大量不经常观察到的时间序列的挑战。我们的方法利用数据流之间的相关性,例如相关区域之间的相关性,同时支持数据驱动的稀疏依赖发现。多分辨率和模块化形式还允许合并不同的辅助信息。提出的方法成功的关键是可伸缩的贝叶斯后验推断。我们专注于(I)利用稀疏图依赖的并行计算,(Ii)多分辨率推理,以及(Iii)相关数据的在线算法。这个项目代表了一项雄心勃勃的跨学科努力,整合了机器学习、系统、工程和统计学的思想。这项工作解决了大数据讨论中一个基本上被忽视的问题:当数据具有随时间变化的关键结构时,如何处理建模和计算问题,特别是来自个别稀疏和不同的测量源。所开发的工具将极大地拓宽可以解决的科学问题的范围。这项工作的结果将被公开传播,包括通过开放源码软件,我们的行业合作伙伴的目标是将该技术转化为现实世界的系统。该项目还涉及开发(I)利用现有基础设施UW DawgBytes的令人兴奋和密集的计划,以增加K-12学生,特别是女孩,接触机器学习的机会;以及(Ii)培训学生统计和计算思维的课程。有关更多信息,请参阅项目网站http://www.stat.washington.edu/~ebfox/CAREER.html.
英文摘要
Consider the task of predicting influenza rates at a very large set of spatial locations. Modeling each region independently does not leverage the information from related regions and can lead to poor predictions, especially in the presence of missing observations. Likewise, imagine estimating the value of every house in the United States. Capturing trends within a neighborhood is key; however, each neighborhood only has a few recent house sales. The challenges presented by these increasingly prevalent massive time series are endemic to a wide range of applications, from crime modeling for police resource allocation to forecasting consumer trends and social networks: the individual data streams often include only infrequent observations such that each alone does not provide sufficient data for accurate inferences. However, the structured relationships between them offer an opportunity to share information. A key question is how to discover these relationships. This project takes a computationally-driven Bayesian nonparametric approach, trading off flexibility and scalability, to address the challenges of massive collections of infrequently observed time series. Our approaches exploit correlation among the data streams, e.g., among related regions, while enabling data-driven discovery of sparse dependencies. The multi-resolution and modular forms also allow incorporation of heterogeneous side information. Key to the success of the proposed methods is scalable Bayesian posterior inference. We focus on (i) parallel computations exploiting sparse graph dependencies, (ii) multi-resolution inference, and (iii) online algorithms for dependent data.This project represents an ambitious cross-disciplinary effort, integrating ideas from machine learning, systems, engineering, and statistics. The work addresses a largely ignored question in the discussion on big data: How to cope with modeling and computational issues when the data has crucial structure across time, especially arising from individually sparse and disparate measurement sources. The tools developed will significantly broaden the scope of scientific questions that can be addressed. Results from this work will be publicly disseminated, including through open source software, and our industry partners aim to transition the technology into real-world systems. This project also involves developing (i) exciting and intensive programs harnessing existing infrastructure, UW DawgBytes, to increase the exposure of K-12 students, and especially girls, to machine learning; and (ii) curriculum training students in both statistical and computational thinking.For further information, see the project website at http://www.stat.washington.edu/~ebfox/CAREER.html.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
PostDoctoral Research Fellowship
-
批准号:0903022
-
项目类别:Fellowship Award
-
资助金额:$13.5万
-
财政年份:2009
-
负责人:Emily Fox
-
依托单位:
海外基金