High-dimensional M-estimation: Understanding risk, improving performance and assessing resampling
High-dimensional M-estimation: Understanding risk, improving performance and assessing resampling
批准号:
1510172
负责人:
Noureddine El Karoui
金额:
$39.42万
依托单位国家:
美国
项目类别:
Continuing Grant
财政年份:
2015
资助国家:
美国
项目状态:
已结题
起止时间:
2015-08-01 至 2019-10-31
中文摘要
学术界或工业界的科学家目前正在处理的数据集的性质正在以非常快的速度发生变化。数据比以往任何时候都更加复杂、更大、维度更高。实践中使用的许多方法都基于这样一种想法,即它们在测量精度或对未来结果的预测方面是最优的,至少对于某些数据生成机制模型是这样。最广泛应用和由来已久的数据分析原则是使用所谓的“最大似然法”。PI和合作者最近发现,在现代和大型数据集经常遇到的情况下,这些最大似然方法是次优的,可以改进。即使对于非常基本和广泛使用的技术(例如,线性回归)也是如此。该项目的目的之一是了解在机器/统计学习实践中广泛使用的其他方法是否也存在同样的现象,进而为数据科学家和数据分析师开发更好的工具。目前,对这些估计量的精度评估通常是通过数据驱动程序(如Bootstrap)进行的。该项目的另一个目标是了解相应的准确性评估对于具有许多预测值的数据集是否具有误导性。如果是这样的话,PI正计划制定方法来纠正现有程序,以便它们产生值得信赖的准确性评估。高维统计对经典统计学提出了深刻的挑战,无论是在应用方面还是在理论方面。在实践中使用的一大类方法是基于求解非平凡的优化问题来估计感兴趣的参数。这就产生了所谓的M-估计量。当这种估计器的维度与实践者拥有的观察数量相比很小时,可以应用标准的经验过程技术来理解这些估计器的统计特性。在PI考虑的情况下,这些技术失败了,需要开发新的技术。PI计划使用来自随机矩阵理论、凸分析和测量结果集中的混合工具来研究这些估计器。基于使用凸性分析的工具,新的优化方法的发展是可预期的。另一个令人兴奋的研究方向是,PI开发的技术应该允许我们研究高维的重采样方法(如Bootstrap)。它们被广泛用于从观察到的数据集中评估统计意义,而不必诉诸理论论证。虽然低维理论已经建立得很好,而且相对容易,并表明这些数值方法应该很好地工作,但高维的情况还没有被理解。PI计划彻底研究这些问题,并提出实际相关的解决方案,如果这些在实践中广泛使用的方法被证明提供了在统计上具有误导性的准确性评估。
英文摘要
The nature of datasets that scientists in academia or industry are currently working with is changing at a very rapid pace. The data is more complex, larger and higher-dimensional than it has ever been before. A lot of the methods used in practice are based on the idea that they are somehow optimal, in terms of measurement accuracy or prediction of future outcomes, at least for certain models of data generating mechanism. A most widely applied and time-honored principle of data analysis is the use of so-called "maximum likelihood methods". It has recently been discovered by the PI and collaborators that in a setting often encountered with modern and large datasets, these maximum likelihood methods are suboptimal and can be improved upon. This is true even for an extremely basic and widely used technique (e.g.,linear regression). One of the aims of the project is to understand if the same phenomena occur for other methods that are widely used in machine/statistical learning practice and in turn develop better tools for data scientists and data analysts. Currently, accuracy assessment for these estimators are often performed through data driven procedures (such as the bootstrap). Another aim of the project is to understand if the corresponding accuracy assessment are misleading for datasets with many predictors. If that is the case, the PI is planning to work on methods to correct the existing procedures so they yield trustworthy accuracy assessments. High-dimensional statistics offers a profound challenge to classical statistics, both on the applied and the theoretical end. A broad class of methods used in practice is based on solving nontrivial optimization problems to estimate parameters of interest. This yields a so-called M-estimator. When the dimension of this estimator is small compared to the number of observations the practitioner has, standard empirical process techniques can be applied to understand the statistical properties of those estimators. In the setting the PI considers, these techniques fail and new techniques need to be developed. The PI plans on using a mix of tools inspired from random matrix theory, convex analysis and concentration of measure results to study those estimators. The development of new optimal methods is expected - based on using tools from convex analysis. Another exciting research line is that the techniques developed by the PI should allow us to study resampling methods in high-dimension (such as the bootstrap). Those are widely used to assess statistical significance from the observed dataset, without having to appeal to theoretical arguments. While the low-dimensional theory is well-established and relatively easy, and suggests that these numerical methods should work well, the high-dimensional case has yet to be understood. The PI plans on studying these problems thoroughly and propose practically relevant solutions if these widely used-in-practice methods are shown to provide statistically misleading accuracy assessments.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
CAREER: Random matrices and High-dimensional statistics
-
批准号:0847647
-
项目类别:Continuing Grant
-
资助金额:$40.0万
-
财政年份:2009
-
负责人:Noureddine El Karoui
-
依托单位:
Random Matrices in Multivariate Statistics: Theoretical Developments and Applications
-
批准号:0605169
-
项目类别:Standard Grant
-
资助金额:$24.0万
-
财政年份:2006
-
负责人:Noureddine El Karoui
-
依托单位:
国内基金
海外基金
肌肉挫伤后组织中时间相关基因表达与损伤经历时间研究
-
批准号:81001347
-
项目类别:青年科学基金项目
-
资助金额:20.0万元
-
批准年份:2010
-
负责人:孙俊红
-
依托单位:
基于计算和存储感知的运动估计算法与结构研究
-
批准号:60803013
-
项目类别:青年科学基金项目
-
资助金额:18.0万元
-
批准年份:2008
-
负责人:邓磊
-
依托单位:
多用户MIMO-OFDM系统中的同步和信道估计的研究
-
批准号:60302025
-
项目类别:联合基金项目
-
资助金额:30.0万元
-
批准年份:2003
-
负责人:张建华
-
依托单位: