课题基金 / 基金详情

BIGDATA: Collaborative Research: F: Statistical Theory and Methods Beyond the Dimensionality Barrier

BIGDATA: Collaborative Research: F: Statistical Theory and Methods Beyond the Dimensionality Barrier
BIGDATA:协作研究:F:超越维度障碍的统计理论和方法
批准号:
1633212
负责人:
Kai Zhang
金额:
$20.0万
依托单位国家:
美国
项目类别:
Standard Grant
财政年份:
2016
资助国家:
美国
项目状态:
已结题
起止时间:
2016-09-01 至 2019-08-31

项目摘要

项目成果

Kai Zhang的其他基金

相似基金

相关文献

中文摘要
翻译
随着技术的最新进步,现在可以测量和记录单个人的大量特征。 大数据的体积、速度和多样性,即“3V”,对这些海量数据集的建模和分析提出了重大挑战。 例如,为了在遗传水平上了解癌症,研究人员需要从有限数量的受试者中获得的数千甚至数百万个候选遗传标记中检测到罕见和微弱的信号。 现有的方法通常假设受试者的数量非常大,这是一个在实践中经常违反的假设。 该项目的主要目标是为非常大的维度,小样本量的数据开发有效的方法。 方法上的进步对于解决医学研究、生物信息学、金融分析和天文图像分析等不同领域的大数据挑战将是非常有价值的。 高效的软件包和算法来实现所提出的方法将被开发和公开,激励这项研究的关键创新思想是从一个新的包装的角度来看,它允许变量的数量,p,是任意大的和观察的数量,n,是有限的高维问题。 本研究将系统地研究在这种“有限n,任意大p”范式下的三个基本问题:(1)伪相关的渐近理论,(2)低秩相关结构的快速检测,以及(3)检测边界和检测稀有和弱信号的最佳测试程序。 这项研究将改变目前的渐近框架,从“大n,小p”和“大n,大p”的制度过渡到“有限n,任意大p”的制度。
英文摘要
With recent advances in technology, it is now possible to measure and record significant numbers of features on a single individual. The volume, velocity, and variety, the "3Vs", of Big Data pose significant challenges for modeling and analysis of these massive datasets. For example, to understand cancer at the genetic level, researchers need to detect rare and weak signals from thousands, or even millions, of candidate genetic markers obtained from a limited number of subjects. Existing methods typically assume that the number of subjects is very large, an assumption often violated in practice. The main goal of this project is to develop efficient methods for extremely large-dimensional, small sample size data. The methodological advances will be extremely valuable in addressing Big Data challenges in different areas such as medical research, bioinformatics, financial analysis, and astronomic image analysis. Efficient software packages and algorithms to implement the proposed methods will be developed and made publicly available.The key innovative idea motivating this research is viewing a high-dimensional problem from a novel packing perspective, which allows the number of variables, p, to be arbitrarily large and the number of observations, n, to be finite. The proposed research will systematically investigate three fundamental problems under this "finite n, arbitrarily large p" paradigm: (1) asymptotic theory of spurious correlations, (2) fast detection of low-rank correlation structures, and (3) detection boundary and optimal testing procedures for detecting rare and weak signals. This research will transform the current asymptotic framework, transitioning from the regimes of "large n, small p" and "large n, larger p" to the regime of "finite n, arbitrarily large p".
期刊论文(5)
专著(0)
科研奖励(0)
会议论文
DOI: 10.1109/tit.2017.2700202
发表时间: 2017
期刊: IEEE Transactions on Information Theory
影响因子: 2.5
作者: [Zhang, Kai]
通讯作者: Zhang, Kai
An Exploratory Statistical Cusp Catastrophe Model
探索性统计尖点灾难模型
DOI: 10.1109/dsaa.2016.17
发表时间: 2016
期刊: 2016 IEEE International Conference on Data Science and Advanced Analytics
影响因子: --
作者: [Chen, Ding-Geng Din, Chen, Xinguang Jim, Zhang, Kai]
通讯作者: Zhang, Kai
JIVE integration of imaging and behavioral data
JIVE 整合影像和行为数据
DOI: 10.1016/j.neuroimage.2017.02.072
发表时间: 2017
期刊: NeuroImage
影响因子: 5.7
作者: [Yu, Qunqun, Risk, Benjamin B., Zhang, Kai, Marron, J.S.]
通讯作者: Marron, J.S.
Calibrated percentile double bootstrap for robust linear regression inference
用于稳健线性回归推理的校准百分位双引导
DOI: 10.5705/ss.202016.0546
发表时间: 2018
期刊: Statistica sinica
影响因子: 1.4
作者: [Daniel McCarthy, Kai Zhang]
通讯作者: Daniel McCarthy, Kai Zhang
FRG: Collaborative Research: Mathematical and Statistical Analysis of Compressible Data on Compressive Networks
Binary Expansion Statistics: A Nonparametric Inference Framework for Big Data
Geometric Perspectives on the Correlation
Collaborative Research: Inference for Linear Model Parameters in Model-free Populations
海外基金