课题基金 / 基金详情

BIGDATA: Collaborative Research: F: Statistical Theory and Methods Beyond the Dimensionality Barrier

BIGDATA: Collaborative Research: F: Statistical Theory and Methods Beyond the Dimensionality Barrier
BIGDATA:协作研究:F:超越维度障碍的统计理论和方法
批准号:
1633212
负责人:
Kai Zhang
金额:
$20.0万
依托单位国家:
美国
项目类别:
Standard Grant
财政年份:
2016
资助国家:
美国
项目状态:
已结题
起止时间:
2016-09-01 至 2019-08-31

项目摘要

项目成果

Kai Zhang的其他基金

相似基金

相关文献

中文摘要
翻译
随着最近技术的进步,现在可以测量和记录单个人的大量特征。大数据的数量、速度和多样性,即“3v”,对这些海量数据集的建模和分析提出了重大挑战。例如,为了在基因水平上了解癌症,研究人员需要从从有限数量的受试者中获得的数千甚至数百万个候选遗传标记中检测出罕见和微弱的信号。现有方法通常假设受试者数量非常大,但在实践中经常违反这一假设。这个项目的主要目标是为极大的大维度,小样本数据开发有效的方法。方法上的进步对于解决不同领域的大数据挑战,如医学研究、生物信息学、金融分析和天文图像分析,将是非常有价值的。将开发有效的软件包和算法来实现所提出的方法并使其公开可用。推动这项研究的关键创新思想是从一种新颖的包装角度来看待一个高维问题,它允许变量的数量p是任意大的,而观察的数量n是有限的。本研究将系统地研究“有限n,任意大p”范式下的三个基本问题:(1)伪相关的渐近理论;(2)低秩相关结构的快速检测;(3)检测稀有和微弱信号的检测边界和最优测试程序。本研究将改变现有的渐近框架,从“大n,小p”和“大n,大p”的模式过渡到“有限n,任意大p”的模式。
英文摘要
With recent advances in technology, it is now possible to measure and record significant numbers of features on a single individual. The volume, velocity, and variety, the "3Vs", of Big Data pose significant challenges for modeling and analysis of these massive datasets. For example, to understand cancer at the genetic level, researchers need to detect rare and weak signals from thousands, or even millions, of candidate genetic markers obtained from a limited number of subjects. Existing methods typically assume that the number of subjects is very large, an assumption often violated in practice. The main goal of this project is to develop efficient methods for extremely large-dimensional, small sample size data. The methodological advances will be extremely valuable in addressing Big Data challenges in different areas such as medical research, bioinformatics, financial analysis, and astronomic image analysis. Efficient software packages and algorithms to implement the proposed methods will be developed and made publicly available.The key innovative idea motivating this research is viewing a high-dimensional problem from a novel packing perspective, which allows the number of variables, p, to be arbitrarily large and the number of observations, n, to be finite. The proposed research will systematically investigate three fundamental problems under this "finite n, arbitrarily large p" paradigm: (1) asymptotic theory of spurious correlations, (2) fast detection of low-rank correlation structures, and (3) detection boundary and optimal testing procedures for detecting rare and weak signals. This research will transform the current asymptotic framework, transitioning from the regimes of "large n, small p" and "large n, larger p" to the regime of "finite n, arbitrarily large p".
期刊论文(5)
专著(0)
科研奖励(0)
会议论文
DOI: 10.1109/tit.2017.2700202
发表时间: 2017
期刊: IEEE Transactions on Information Theory
影响因子: 2.5
作者: [Zhang, Kai]
通讯作者: Zhang, Kai
An Exploratory Statistical Cusp Catastrophe Model
探索性统计尖点灾难模型
DOI: 10.1109/dsaa.2016.17
发表时间: 2016
期刊: 2016 IEEE International Conference on Data Science and Advanced Analytics
影响因子: --
作者: [Chen, Ding-Geng Din, Chen, Xinguang Jim, Zhang, Kai]
通讯作者: Zhang, Kai
JIVE integration of imaging and behavioral data
JIVE 整合影像和行为数据
DOI: 10.1016/j.neuroimage.2017.02.072
发表时间: 2017
期刊: NeuroImage
影响因子: 5.7
作者: [Yu, Qunqun, Risk, Benjamin B., Zhang, Kai, Marron, J.S.]
通讯作者: Marron, J.S.
Calibrated percentile double bootstrap for robust linear regression inference
用于稳健线性回归推理的校准百分位双引导
DOI: 10.5705/ss.202016.0546
发表时间: 2018
期刊: Statistica sinica
影响因子: 1.4
作者: [Daniel McCarthy, Kai Zhang]
通讯作者: Daniel McCarthy, Kai Zhang
FRG: Collaborative Research: Mathematical and Statistical Analysis of Compressible Data on Compressive Networks
Binary Expansion Statistics: A Nonparametric Inference Framework for Big Data
Geometric Perspectives on the Correlation
Collaborative Research: Inference for Linear Model Parameters in Model-free Populations
海外基金