CI-SUSTAIN: Stan for the Long Run
CI-SUSTAIN: Stan for the Long Run
批准号:
1730414
负责人:
Andrew Gelman
金额:
$98.39万
依托单位:
依托单位国家:
美国
项目类别:
Standard Grant
财政年份:
2017
资助国家:
美国
项目状态:
已结题
起止时间:
2017-08-01 至 2020-07-31
中文摘要
Stan是一个软件包,通过允许科学家快速轻松地探索,评估和完善针对其特定研究问题和数据收集机制的丰富科学假设,从而改变科学发现。由于计算的原因,对数据(无论是大数据还是其他数据)的分析往往是简单的,并且更多地关注操作数据的困难,而不是现实的科学模型。下一代贝叶斯推理可以通过复杂的模型(可以调整样本和总体之间的差异,以及实验组和对照组之间的差异)将科学家带出这个僵局,以加入具有统计调整的严峻性和力量的大型数据集的好处。Stan通过自然、易于学习和可移植的建模语言,以及强大、实用的推理工具,进一步帮助培养下一代数据科学家。这个项目的具体目标是巩固Stan代码库,以支持Stan软件的应用、维护和开发。Stan被应用于物理、生物和社会科学的许多领域,从中微子到超新星,从细胞生物学到种群生态学,从人类反应时间到社会网络进化,数以百计的尺度。在这个项目中,PI的目标是记录和加固Stan的核心基础设施,使其能够被更广泛的科学家使用,由更广泛的软件开发人员维护,并可扩展以允许未来开发新的科学应用程序和统计算法。从技术上讲,该项目致力于为用户和开发人员提供完整的文档,在单元、集成和功能级别进行测试,并不可避免地重构Stan组件的应用程序编程接口(API)。这些组件包括:(1)一个自动可微的数学、统计和矩阵代数库;(2)一个用于表达统计数据/参数结构和科学/测量模型的命令式概率编程语言;(3)提供精确和近似的全贝叶斯参数估计和预测推理的核心推理算法;(4)一个高级命令、I/O回调和中断的服务层;(5)接口,将Stan的概率建模、分析和可视化功能集成到现有的数据科学工作流中,语言包括R、Python、Julia、Stata、MATLAB、Mathematica以及用于云和集群计算的命令行。重构的主要目标是在系统组件和文档中实现足够的模块化,使开发人员能够专注于单个组件,如新功能、算法或可视化,同时允许Stan被用作算法开发的研究工具。这些目标补充了记录语言和接口的目标,以便为应用科学家促进严格的统计方法和可重复的计算工作流程。
英文摘要
Stan is a software package that transforms scientific discovery by allowing scientists to quickly and easily explore, evaluate, and refine rich scientific hypotheses tailored to their particular research question and data collection mechanism. For computational reasons, analyses of data (big or otherwise) have tended to be simple and focused more on the difficulties of manipulating the data than on realistic scientific models. The next generation of Bayesian inference can take scientists beyond this impasse, via sophisticated models that can adjust for differences between sample and population, and between treatment and control groups, to join the benefits of large datasets with the rigor and power of statistical adjustment. Stan further helps educate the next generation of data scientists, with a natural, easy-to-learn and portable modeling language, coupled with robust, practical inference tools. The specific goal of this project is to solidify the Stan code base to enable application, maintenance, and development of the Stan software. Stan is being applied in many corners of the physical, biological, and social sciences, hundreds of at scales ranging from the neutrinos to supernovas, from cellular biology to population ecology, and from human reaction times to social network evolution. In this project the PI aims to document and ruggedize the core infrastructure of Stan to enable it to be used by a wider audience of scientists, to be maintained by a wider group of software developers, and to be extensible to allow for the future development of new scientific applications and statistical algorithms.Technically, this project is devoted to thoroughly documenting for users and developers, testing at the unit, integration, and functional levels, and inevitably refactoring the application programming interfaces (API) of Stan's components. These components include (1) an automatically differentiable mathematics, statistics, and matrix algebra library, (2) an imperative probabilistic programming language for expressing statistical data/parameter structures and scientific/measurement models, (3) core inference algorithms for providing exact and approximate full Bayesian parameter estimation and predictive inference, (4) a service layer of high-level commands, I/O callbacks, and interrupts, and (5) interfaces integrating Stan's probability modeling, analysis and visualization capabilities into existing data science workflows in languages including R, Python, Julia, Stata, MATLAB, Mathematica and the command-line for cloud and cluster computing. The main goal of the refactoring is to achieve enough modularity in system components and documentation that developers will be able to concentrate on a single component, such as a new function, algorithm, or visualization, as well as to allow Stan to be used as a research tool for algorithm development. These goals complement the goals of documenting the language and interfaces in order to promote rigorous statistical methodology and reproducible computational workflows for applied scientists.
期刊论文(4)
专著(0)
科研奖励(0)
会议论文
登录
查看更多内容
DOI:
10.1214/17-ba1091
发表时间:
2018-09-01
期刊:
BAYESIAN ANALYSIS
影响因子:
4.4
作者:
[Yao, Yuling, Vehtari, Aki, Tonellato, Stefano]
通讯作者:
Tonellato, Stefano
DOI:
10.1080/00031305.2018.1549100
发表时间:
2019-07-03
期刊:
AMERICAN STATISTICIAN
影响因子:
1.8
作者:
[Gelman, Andrew, Goodrich, Ben, Vehtari, Aki]
通讯作者:
Vehtari, Aki
DOI:
10.1111/rssa.12378
发表时间:
2019-02-01
期刊:
JOURNAL OF THE ROYAL STATISTICAL SOCIETY SERIES A-STATISTICS IN SOCIETY
影响因子:
2
作者:
[Gabry, Jonah, Simpson, Daniel, Gelman, Andrew]
通讯作者:
Gelman, Andrew
DOI:
10.1214/17-aoas1122
发表时间:
2018-09-01
期刊:
ANNALS OF APPLIED STATISTICS
影响因子:
1.8
作者:
[Weber, Sebastian, Gelman, Andrew, Racine-Poon, Amy]
通讯作者:
Racine-Poon, Amy
Scalable Bayesian regression: Analytical and numerical tools for efficient Bayesian analysis in the large data regime
-
批准号:2311354
-
项目类别:Standard Grant
-
资助金额:$29.99万
-
财政年份:2023
-
负责人:Andrew Gelman
-
依托单位:
RAPID: Flexible, Efficient, and Available Bayesian Computation for Epidemic Models
-
批准号:2055251
-
项目类别:Standard Grant
-
资助金额:$18.7万
-
财政年份:2020
-
负责人:Andrew Gelman
-
依托单位:
Collaborative Research: PPoSS: Planning: Scalable Systems for Probabilistic Programming
-
批准号:2029022
-
项目类别:Standard Grant
-
资助金额:$11.72万
-
财政年份:2020
-
负责人:Andrew Gelman
-
依托单位:
RIDIR: Collaborative Research: Bayesian analytical tools to improve survey estimates for subpopulations and small areas
-
批准号:1926578
-
项目类别:Standard Grant
-
资助金额:$63.22万
-
财政年份:2019
-
负责人:Andrew Gelman
-
依托单位:
Collaborative Research: Multilevel Regression and Poststratification: A Unified Framework for Survey Weighted Inference
-
批准号:1534414
-
项目类别:Standard Grant
-
资助金额:$9.13万
-
财政年份:2015
-
负责人:Andrew Gelman
-
依托单位:
CI-ADDO-NEW: Stan, Scalable Software for Bayesian Modeling
-
批准号:1205516
-
项目类别:Standard Grant
-
资助金额:$49.96万
-
财政年份:2012
-
负责人:Andrew Gelman
-
依托单位:
CMG: Reconstructing Climate from Tree Ring Data
-
批准号:0934516
-
项目类别:Standard Grant
-
资助金额:$59.81万
-
财政年份:2009
-
负责人:Andrew Gelman
-
依托单位:
Design and Analysis of "How many X's do you know" surveys for the study of polarization in social networks
-
批准号:0532231
-
项目类别:Standard Grant
-
资助金额:$60.0万
-
财政年份:2005
-
负责人:Andrew Gelman
-
依托单位:
Multilevel Modeling for the Study of Public Opinion and Voting
-
批准号:0318115
-
项目类别:Continuing Grant
-
资助金额:$21.49万
-
财政年份:2003
-
负责人:Andrew Gelman
-
依托单位:
Doctoral Dissertation Research: Estimating Congressional District-Level Opinions from National Surveys using a Bayesian Hierarchical Logistic Regression Model
-
批准号:0241709
-
项目类别:Standard Grant
-
资助金额:$1.2万
-
财政年份:2003
-
负责人:Andrew Gelman
-
依托单位:
Collaborative Research: Combining Expert Judgments for Environmental Risk Analysis.
-
批准号:0084368
-
项目类别:Standard Grant
-
资助金额:$5.75万
-
财政年份:2000
-
负责人:Andrew Gelman
-
依托单位:
Bayesian Analysis of Sample Surveys
-
批准号:9987748
-
项目类别:Continuing Grant
-
资助金额:$25.47万
-
财政年份:2000
-
负责人:Andrew Gelman
-
依托单位:
Models and Model Checking for Spatially-Varying Environmental Hazards and Decision Problems
-
批准号:9708424
-
项目类别:Standard Grant
-
资助金额:$22.71万
-
财政年份:1997
-
负责人:Andrew Gelman
-
依托单位:
NSF Young Investigator
-
批准号:9796129
-
项目类别:Continuing Grant
-
资助金额:$14.1万
-
财政年份:1996
-
负责人:Andrew Gelman
-
依托单位:
NSF Young Investigator
-
批准号:9457824
-
项目类别:Continuing Grant
-
资助金额:$9.73万
-
财政年份:1994
-
负责人:Andrew Gelman
-
依托单位:
Mathematical Sciences: Using Inference from Simulation to Improve Efficiency of Simulations
-
批准号:9404305
-
项目类别:Standard Grant
-
资助金额:$4.5万
-
财政年份:1994
-
负责人:Andrew Gelman
-
依托单位:
Mathematical Sciences: Postdoctoral Research Fellowship
-
批准号:9007223
-
项目类别:Fellowship Award
-
资助金额:$7.5万
-
财政年份:1990
-
负责人:Andrew Gelman
-
依托单位:
海外基金