课题基金 / 基金详情

AF: Medium: Collaborative Proposal: Foundations of Adaptive Data Analysis

AF: Medium: Collaborative Proposal: Foundations of Adaptive Data Analysis
AF:媒介:协作提案:自适应数据分析的基础
批准号:
1763665
负责人:
Cynthia Dwork
金额:
$28.6万
依托单位:
依托单位国家:
美国
项目类别:
Continuing Grant
财政年份:
2018
资助国家:
美国
项目状态:
已结题
起止时间:
2018-03-01 至 2021-05-31

项目摘要

项目成果

相似基金

相关文献

中文摘要
翻译
点击翻译按钮获取中文摘要
英文摘要
Classical tools for rigorously analyzing data make the assumption that the analysis is static: the models and the hypotheses to be tested are fixed independently of the data, and preliminary analysis of the data does not feed back into the data gathering procedure. On the other hand, modern data analysis is highly adaptive. Large parts of modern machine learning perform model selection as a function of the data by iteratively tuning hyper-parameters, and exploratory data analysis is conducted to suggest hypotheses, which are then validated on the same data sets used to discover them. This kind of adaptivity is often referred to as p-hacking, and blamed in part for the surprising prevalence of non-reproducible science in some empirical fields. This project aims to develop rigorous tools and methodologies to perform statistically valid data analysis in the adaptive setting, drawing on techniques from statistics, information theory, differential privacy, and stable algorithm design. The technical goals of this project include coming up with: 1) information-theoretic measures that characterize the degree to which a worst-case data analysis can over-fit, given an interaction with a dataset; 2) models for data analysts that move beyond the worst-case setting, and; 3) empirical investigations that bridge the gap between theory and practice. The problem of adaptive data analysis (also called post-selection inference, or selective inference) has attracted attention in both computer science and statistics over the past several years, but from relatively disjoint communities. Part of the aim of this project is to integrate these two lines of work. The team of researchers on this project span departments of computer science, statistics, and biomedical data science. In addition to attempting to unify these two areas, the broader impacts of this research will be to make science more reliable, and reduce the prevalence of "over-fitting" and "false discovery." The project also has a significant outreach and education component, and will educate graduate students, organize workshops, and produce expository materials.This award reflects NSF's statutory mission and has been deemed worthy of support through evaluation using the Foundation's intellectual merit and broader impacts review criteria.
期刊论文(14)
专著(0)
科研奖励(0)
会议论文
DOI: --
发表时间: 2020-10
期刊: ArXiv
影响因子: --
作者: [Linjun Zhang;Zhun Deng;Kenji Kawaguchi;Amirata Ghorbani;James Y. Zou]
通讯作者: Linjun Zhang;Zhun Deng;Kenji Kawaguchi;Amirata Ghorbani;James Y. Zou
DOI: --
发表时间: 2020-07
期刊: ArXiv
影响因子: --
作者: [Zhun Deng;C. Dwork;Jialiang Wang;Linjun Zhang]
通讯作者: Zhun Deng;C. Dwork;Jialiang Wang;Linjun Zhang
Abstracting Fairness: Oracles, Metrics, and Interpretability
抽象公平性:预言、指标和可解释性
DOI: --
发表时间: 2020
期刊: Proceedings of the 1st Symposium on Foundations of Responsible Computing (FORC 2020
影响因子: --
作者: [Dwork, C, Ilvento, C, Rothblum, G, Sur, P]
通讯作者: Sur, P
DOI: 10.1145/3188745.3188946
发表时间: 2018-06
期刊: Proceedings of the 50th Annual ACM SIGACT Symposium on Theory of Computing
影响因子: --
作者: [Mark Bun;C. Dwork;G. Rothblum;T. Steinke]
通讯作者: Mark Bun;C. Dwork;G. Rothblum;T. Steinke
13
    海外基金