CAREER: Scaling up Modeling and Statistical Inference for Massive Collections of Time Series
CAREER: Scaling up Modeling and Statistical Inference for Massive Collections of Time Series
批准号:
1350133
负责人:
Emily Fox
金额:
$54.92万
依托单位:
依托单位国家:
美国
项目类别:
Continuing Grant
财政年份:
2014
资助国家:
美国
项目状态:
已结题
起止时间:
2014-06-15 至 2021-05-31
中文摘要
考虑在一个非常大的空间位置集合上预测流感发生率的任务。对每个区域独立建模无法利用相关区域的信息,可能导致预测结果不佳,特别是在存在缺失观测值的情况下。同样,想象一下估算美国每栋房子的价值。捕捉社区内的趋势是关键;然而,每个社区只有几个最近的房屋销售。这些日益普遍的大规模时间序列所带来的挑战是广泛的应用所特有的,从用于警察资源分配的犯罪建模到预测消费者趋势和社交网络:单个数据流通常仅包括不频繁的观察结果,使得每个单独的数据都不能提供足够的数据来进行准确的推断。然而,它们之间的结构化关系提供了分享信息的机会。 一个关键问题是如何发现这些关系。 该项目采用计算驱动的贝叶斯非参数方法,权衡灵活性和可扩展性,以解决大量收集不经常观察的时间序列的挑战。我们的方法利用数据流之间的相关性,例如,在相关区域之间,同时实现稀疏依赖性的数据驱动发现。多分辨率和模块化形式还允许合并异构辅助信息。所提出的方法的成功的关键是可扩展的贝叶斯后验推理。我们专注于(i)利用稀疏图依赖的并行计算,(ii)多分辨率推理,以及(iii)依赖数据的在线算法。该项目代表了一个雄心勃勃的跨学科努力,整合了机器学习,系统,工程和统计学的想法。这项工作解决了大数据讨论中一个很大程度上被忽视的问题:当数据在时间上具有关键结构时,如何科普建模和计算问题,特别是来自单独稀疏和不同的测量源。所开发的工具将大大扩大可解决的科学问题的范围。这项工作的结果将公开传播,包括通过开源软件,我们的行业合作伙伴旨在将该技术转化为现实世界的系统。该项目还涉及开发(i)利用现有基础设施的令人兴奋和密集的计划,UW Dawgestion,以增加K-12学生,特别是女孩,对机器学习的接触;以及(ii)在统计和计算思维方面对学生进行课程培训。http://www.stat.washington.edu/~ebfox/CAREER.html
英文摘要
Consider the task of predicting influenza rates at a very large set of spatial locations. Modeling each region independently does not leverage the information from related regions and can lead to poor predictions, especially in the presence of missing observations. Likewise, imagine estimating the value of every house in the United States. Capturing trends within a neighborhood is key; however, each neighborhood only has a few recent house sales. The challenges presented by these increasingly prevalent massive time series are endemic to a wide range of applications, from crime modeling for police resource allocation to forecasting consumer trends and social networks: the individual data streams often include only infrequent observations such that each alone does not provide sufficient data for accurate inferences. However, the structured relationships between them offer an opportunity to share information. A key question is how to discover these relationships. This project takes a computationally-driven Bayesian nonparametric approach, trading off flexibility and scalability, to address the challenges of massive collections of infrequently observed time series. Our approaches exploit correlation among the data streams, e.g., among related regions, while enabling data-driven discovery of sparse dependencies. The multi-resolution and modular forms also allow incorporation of heterogeneous side information. Key to the success of the proposed methods is scalable Bayesian posterior inference. We focus on (i) parallel computations exploiting sparse graph dependencies, (ii) multi-resolution inference, and (iii) online algorithms for dependent data.This project represents an ambitious cross-disciplinary effort, integrating ideas from machine learning, systems, engineering, and statistics. The work addresses a largely ignored question in the discussion on big data: How to cope with modeling and computational issues when the data has crucial structure across time, especially arising from individually sparse and disparate measurement sources. The tools developed will significantly broaden the scope of scientific questions that can be addressed. Results from this work will be publicly disseminated, including through open source software, and our industry partners aim to transition the technology into real-world systems. This project also involves developing (i) exciting and intensive programs harnessing existing infrastructure, UW DawgBytes, to increase the exposure of K-12 students, and especially girls, to machine learning; and (ii) curriculum training students in both statistical and computational thinking.For further information, see the project website at http://www.stat.washington.edu/~ebfox/CAREER.html.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
PostDoctoral Research Fellowship
-
批准号:0903022
-
项目类别:Fellowship Award
-
资助金额:$13.5万
-
财政年份:2009
-
负责人:Emily Fox
-
依托单位:
海外基金