BIGDATA: Collaborative Research: F: Statistical Theory and Methods Beyond the Dimensionality Barrier
BIGDATA: Collaborative Research: F: Statistical Theory and Methods Beyond the Dimensionality Barrier
批准号:
1633212
负责人:
Kai Zhang
金额:
$20.0万
依托单位国家:
美国
项目类别:
Standard Grant
财政年份:
2016
资助国家:
美国
项目状态:
已结题
起止时间:
2016-09-01 至 2019-08-31
中文摘要
随着最近技术的进步,现在可以测量和记录单个个体的大量特征。大数据的数量、速度和种类,即“3V”,给这些海量数据集的建模和分析带来了巨大的挑战。例如,为了在基因水平上理解癌症,研究人员需要从从有限数量的受试者那里获得的数千、甚至数百万个候选遗传标记中检测出罕见而微弱的信号。现有的方法通常假设受试者的数量非常大,这一假设在实践中经常被违反。这个项目的主要目标是为极大的维度、小样本的数据开发有效的方法。这些方法上的进步将在应对不同领域的大数据挑战方面具有极其宝贵的价值,例如医学研究、生物信息学、金融分析和天文图像分析。这项研究的关键创新想法是从一个新的布局角度来看待高维问题,它允许变量p的数量任意大,而观测的数量n有限。本研究将系统地研究这一“有限n,任意大p”范式下的三个基本问题:(1)伪相关的渐近理论,(2)低阶相关结构的快速检测,以及(3)检测稀有和微弱信号的检测边界和最优测试方法。本研究将改变目前的渐近框架,从“大n,小p”和“大n,大p”的体制过渡到“有限n,任意大p”的体制。
英文摘要
With recent advances in technology, it is now possible to measure and record significant numbers of features on a single individual. The volume, velocity, and variety, the "3Vs", of Big Data pose significant challenges for modeling and analysis of these massive datasets. For example, to understand cancer at the genetic level, researchers need to detect rare and weak signals from thousands, or even millions, of candidate genetic markers obtained from a limited number of subjects. Existing methods typically assume that the number of subjects is very large, an assumption often violated in practice. The main goal of this project is to develop efficient methods for extremely large-dimensional, small sample size data. The methodological advances will be extremely valuable in addressing Big Data challenges in different areas such as medical research, bioinformatics, financial analysis, and astronomic image analysis. Efficient software packages and algorithms to implement the proposed methods will be developed and made publicly available.The key innovative idea motivating this research is viewing a high-dimensional problem from a novel packing perspective, which allows the number of variables, p, to be arbitrarily large and the number of observations, n, to be finite. The proposed research will systematically investigate three fundamental problems under this "finite n, arbitrarily large p" paradigm: (1) asymptotic theory of spurious correlations, (2) fast detection of low-rank correlation structures, and (3) detection boundary and optimal testing procedures for detecting rare and weak signals. This research will transform the current asymptotic framework, transitioning from the regimes of "large n, small p" and "large n, larger p" to the regime of "finite n, arbitrarily large p".
期刊论文(5)
专著(0)
科研奖励(0)
会议论文
登录
查看更多内容
DOI:
10.1109/tit.2017.2700202
发表时间:
2017
期刊:
IEEE Transactions on Information Theory
影响因子:
2.5
作者:
[Zhang, Kai]
通讯作者:
Zhang, Kai
DOI:
10.1109/dsaa.2016.17
发表时间:
2016
期刊:
2016 IEEE International Conference on Data Science and Advanced Analytics
影响因子:
--
作者:
[Chen, Ding-Geng Din, Chen, Xinguang Jim, Zhang, Kai]
通讯作者:
Zhang, Kai
JIVE integration of imaging and behavioral data
JIVE 整合影像和行为数据
DOI:
10.1016/j.neuroimage.2017.02.072
发表时间:
2017
期刊:
NeuroImage
影响因子:
5.7
作者:
[Yu, Qunqun, Risk, Benjamin B., Zhang, Kai, Marron, J.S.]
通讯作者:
Marron, J.S.
Calibrated percentile double bootstrap for robust linear regression inference
用于稳健线性回归推理的校准百分位双引导
DOI:
10.5705/ss.202016.0546
发表时间:
2018
期刊:
Statistica sinica
影响因子:
1.4
作者:
[Daniel McCarthy, Kai Zhang]
通讯作者:
Daniel McCarthy, Kai Zhang
FRG: Collaborative Research: Mathematical and Statistical Analysis of Compressible Data on Compressive Networks
-
批准号:2152289
-
项目类别:Continuing Grant
-
资助金额:$80.0万
-
财政年份:2022
-
负责人:Kai Zhang
-
依托单位:
Binary Expansion Statistics: A Nonparametric Inference Framework for Big Data
-
批准号:1916237
-
项目类别:Standard Grant
-
资助金额:$15.0万
-
财政年份:2019
-
负责人:Kai Zhang
-
依托单位:
Geometric Perspectives on the Correlation
-
批准号:1613112
-
项目类别:Standard Grant
-
资助金额:$12.0万
-
财政年份:2016
-
负责人:Kai Zhang
-
依托单位:
Collaborative Research: Inference for Linear Model Parameters in Model-free Populations
-
批准号:1309619
-
项目类别:Standard Grant
-
资助金额:$5.0万
-
财政年份:2013
-
负责人:Kai Zhang
-
依托单位:
海外基金