Resampling Methods for High-Dimensional and Large-Scale Data
Resampling Methods for High-Dimensional and Large-Scale Data
批准号:
1613218
负责人:
Miles Lopes
金额:
$15.0万
依托单位国家:
美国
项目类别:
Standard Grant
财政年份:
2016
资助国家:
美国
项目状态:
已结题
起止时间:
2016-07-01 至 2020-06-30
中文摘要
呼吸方法是一类广泛的工具,用于测量统计结果的变异性,例如,允许研究人员确定实验结果是否显著。 在过去的几十年里,这些方法得到了广泛的研究,它们已经成为统计实践的基础-在很大程度上是因为它们可以解决复杂的问题,同时依赖于相对较少的假设。尽管如此,在现代数据分析的背景下,仍然有很多关于reservation方法的性能有待理解,其中观测往往具有大量的特征(高维数据),或者数据量如此之大,以至于超过了计算资源(大规模数据)。在这两个具有挑战性的设置中,拟议的研究将扩展reservation方法的适用性,这些努力将由下面讨论的两个研究主题指导:首先,在高维数据的设置中,推理问题的理解,包括测试和置信区间,与估计和预测问题相比,仍然不发达。考虑到reservation方法是一种通用的推理方法,重要的是要知道它们如何受到低维结构和正则化的影响。特别是,拟议的研究将研究涉及结构化协方差矩阵的高维模型中重采样方法的性能。其次,在大规模数据的设置中,随机算法由于其产生快速近似解的能力而受到越来越多的关注。虽然这种算法的输出是随机的,但它们的波动通常可以以更大的计算为代价来减少。随机算法的这个一般特性导致了在精度和计算成本之间进行优化的问题。为了解决这个问题,本研究将探讨如何使用重新排序方法来衡量一系列流行的随机算法的权衡。
英文摘要
Resampling methods are a broad class of tools that serve to measure the variability of statistical results, for example, allowing a researcher to determine whether or not the outcome of an experiment is significant. Over the course of the last few decades, these methods have been extensively studied, and they have become fundamental to the practice of statistics - in large part because they can solve complex problems while relying on relatively few assumptions. Nevertheless, much remains to be understood about the performance of resampling methods in the context of modern data analysis, where observations tend to have large numbers of features (high-dimensional data), or where the quantity of data is so large that it outstrips computational resources (large-scale data). In both of these challenging settings, the proposed research will extend the applicability of resampling methods, and these efforts will be guided by two research themes discussed below.First, in the setting of high-dimensional data, the understanding of inference problems, including tests and confidence intervals, remains underdeveloped in comparison with estimation and prediction problems. Given that resampling methods are a general-purpose approach to inference, it is important to know how they are influenced by the effects of low-dimensional structure and regularization. In particular, the proposed research will study the performance of resampling methods in high-dimensional models involving structured covariance matrices. Second, in the setting of large-scale data, randomized algorithms have received growing attention for their ability to produce fast approximate solutions. Although the outputs of such algorithms are random, their fluctuations can often be reduced at the expense of greater computation. This general trait of randomized algorithms leads to the problem of optimizing a tradeoff between precision and computational cost. Towards a solution, the proposed research will investigate how resampling methods can be used to measure this tradeoff for a collection of popular randomized algorithms.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
Bootstrap Methods in Modern Settings: Inference and Computation
-
批准号:1915786
-
项目类别:Continuing Grant
-
资助金额:$22.0万
-
财政年份:2019
-
负责人:Miles Lopes
-
依托单位:
国内基金
海外基金
Computational Methods for Analyzing Toponome Data
-
批准号:60601030
-
项目类别:青年科学基金项目
-
资助金额:17.0万元
-
批准年份:2006
-
负责人:Axel Mosig
-
依托单位: