BIGDATA: F: Reliable Inference with Big Data: Reproducibility, Data Sharing, Heterogeneity
BIGDATA: F: Reliable Inference with Big Data: Reproducibility, Data Sharing, Heterogeneity
批准号:
1741162
负责人:
Andrea Montanari
金额:
$65.0万
依托单位:
依托单位国家:
美国
项目类别:
Standard Grant
财政年份:
2017
资助国家:
美国
项目状态:
已结题
起止时间:
2017-09-01 至 2021-08-31
中文摘要
在过去的十年中,“大数据”技术已经允许获取大量数据(例如通过智能手机)并将其积累到大型数据库中。已经开发出强大的硬件和软件系统来处理这些数据并提取统计模型。例如,可以根据患者的特征对某一医疗程序的结果进行建模,从而原则上为该程序提供个性化的风险评分。不幸的是,这些数据和所用算法日益复杂,使得统计模型的透明度大大降低。我们对这些统计预测有多确定呢?它们的有效期限是什么?最终模型的偏差有多大?该项目侧重于大数据中普遍存在的四个主要挑战,这些挑战对于提取可靠的见解至关重要:可重复性;数据共享;缺失的数据;数据的异质性。(1)再现性要求能够比较从不同数据集提取的两个模型(例如,在积累了额外数据之后)。这反过来是不可能的,除非我们有可靠的程序来量化复杂的高维模型的不确定性和信心。最近在这个方向上提出的想法仍然不足以应付现实的大规模应用。(2)数据共享是现代数据分析的一个关键特征,即单个海量数据集由数百名独立研究人员进行研究。这样一群研究人员毫无防备的统计推断不可避免地导致大量错误的发现。该项目建立在错误发现率控制方法的基础上,为分散的数据分析提出了安全的方法。(3)缺失数据在大数据中普遍存在。虽然过去已经开发了几种方法来处理丢失的数据,但尚不清楚它们在多大程度上适用于现代情景。该项目旨在根据各种方法的严格比较制定原则性指导方针,并开发基于最大似然的新算法。(4)数据异质性。大数据通常是由多个数据源的聚合产生的。我们怎样才能防止标准的统计程序受到这种异质性的严重影响?该项目使用新的正则化方案来融合多源信息。
英文摘要
Over the last decade, 'big data' technologies have allowed the acquisition of vast amount of data (e.g. through smartphones) and their accumulation into large scale databases. Powerful hardware and software systems have been developed to crunch these data and extract statistical models. For instance, the outcome of a certain medical procedure can be modeled in terms of the features of the patient, thus in principle providing a personalized risk score for that procedure. Unfortunately, the increasing complexity of these data and of the algorithms used has made statistical models significantly less transparent. How certain are we of these statistical predictions? What is their limit of validity? How biased is the resulting model?This project focuses on four main challenges that are ubiquitous in big-data, and are crucial to extract reliable insights: reproducibility; data sharing; missing data; data heterogeneity. (1) Reproducibility requires being able to compare two models extracted from different data sets (e.g. after additional data have been accumulated). This is in turn impossible unless we have reliable procedures to quantify uncertainty and confidence in complex high-dimensional models. Recently proposed ideas in this direction are still insufficient to cope with realistic large-scale applications.(2) Data sharing is a key feature of modern data analysis, whereby a single massive data set is being studied by hundreds of independent researchers. Unguarded statistical inference by such a population of researchers unavoidably leads to large numbers of false discoveries. The project builds on false discovery rate-controlling methods to propose safe approaches for decentralized data analysis.(3) Missing data are ubiquitous in big data. While several methods have been developed in the past to deal with missing data, it is unclear to what extent they are applicable to modern scenarios. The project aims at developing principled guidelines based on a rigorous comparison of various approaches, and developing new algorithms based on maximum likelihood.(4) Data heterogeneity. Big data are often produced by the aggregation of multiple data sources. How can we prevent standard statistical procedures to be critically affected by such heterogeneities? The project uses new regularization schemes to fusion information across multiple sources.
期刊论文(28)
专著(0)
科研奖励(0)
会议论文
登录
查看更多内容
DOI:
--
发表时间:
2021-02
期刊:
ArXiv
影响因子:
--
作者:
[Song Mei;Theodor Misiakiewicz;A. Montanari]
通讯作者:
Song Mei;Theodor Misiakiewicz;A. Montanari
DOI:
10.1214/19-aos1910
发表时间:
2020-08
期刊:
The Annals of Statistics
影响因子:
--
作者:
[B. Ghorbani;Song Mei;Theodor Misiakiewicz;A. Montanari]
通讯作者:
B. Ghorbani;Song Mei;Theodor Misiakiewicz;A. Montanari
DOI:
10.1088/1742-5468/ac3a81
发表时间:
2020-06
期刊:
Journal of Statistical Mechanics: Theory and Experiment
影响因子:
--
作者:
[B. Ghorbani;Song Mei;Theodor Misiakiewicz;A. Montanari]
通讯作者:
B. Ghorbani;Song Mei;Theodor Misiakiewicz;A. Montanari
DOI:
--
发表时间:
2021-06
期刊:
ArXiv
影响因子:
--
作者:
[Yuchen Wu;M. Bateni;André Linhares;Filipe Almeida;A. Montanari;A. Norouzi-Fard;Jakab Tardos]
通讯作者:
Yuchen Wu;M. Bateni;André Linhares;Filipe Almeida;A. Montanari;A. Norouzi-Fard;Jakab Tardos
Optimization of the Sherrington--Kirkpatrick Hamiltonian
Sherrington--Kirkpatrick 哈密顿量的优化
DOI:
10.1137/20m132016x
发表时间:
2021
期刊:
SIAM Journal on Computing
影响因子:
1.6
作者:
[Montanari, Andrea]
通讯作者:
Montanari, Andrea
共 24 条
CIF: Small: Learning and estimation with rough non-convex objectives: Fundamental limits and efficient algorithms
-
批准号:2006489
-
项目类别:Standard Grant
-
资助金额:$33.0万
-
财政年份:2020
-
负责人:Andrea Montanari
-
依托单位:
Workshop: Advances in Asymptotic Probability
-
批准号:1839440
-
项目类别:Standard Grant
-
资助金额:$3.5万
-
财政年份:2018
-
负责人:Andrea Montanari
-
依托单位:
CIF:Small:Information-theoretic and Computational Thresholds in Statistical Learning
-
批准号:1714305
-
项目类别:Standard Grant
-
资助金额:$45.0万
-
财政年份:2017
-
负责人:Andrea Montanari
-
依托单位:
CIF: Small: Optimal Iterative Estimation in Signal Processing, Information Theory and Machine Learning
-
批准号:1319979
-
项目类别:Standard Grant
-
资助金额:$41.62万
-
财政年份:2013
-
负责人:Andrea Montanari
-
依托单位:
The game dynamics of social interaction: Algorithms and applications
-
批准号:0915145
-
项目类别:Standard Grant
-
资助金额:$49.98万
-
财政年份:2009
-
负责人:Andrea Montanari
-
依托单位:
CAREER: New Information Processing Techniques from Statistical Physics and Probability Theory
-
批准号:0743978
-
项目类别:Continuing Grant
-
资助金额:$32.0万
-
财政年份:2008
-
负责人:Andrea Montanari
-
依托单位:
海外基金