课题基金 / 基金详情

项目摘要

项目成果

相关文献

中文摘要
翻译
点击翻译按钮获取中文摘要
英文摘要
This proposal develops novel statistics and machine learning methods for distributed analysis of big data in biomedical studies and precision medicine and for selecting a small group of molecules that are associated with biological and clinical outcomes from high-throughput data such as microarray, proteomic, and next generation sequence from biomedical research, especially for autism studies and Alzheimer’s disease research. It focuses on developing efficient distributed statistical methods for Big Data computing, storage, and communication, and for solving distributed health data collected at different locations that are hard to aggregate in meta-analysis due to privacy and ownership concerns. It develops both computationally and statistically efficient methods and valid statistical tools for exploring heterogeneity of big data in precision medicine, for studying associations of genomics and genetic information with clinical and biological outcomes, and for feature selection and model building in presence of errors-in- variables, endogeneity, and heavy-tail error distributions, and for predicting clinical outcomes and understanding molecular mechanisms. It introduces more robust and powerful statistical tests for selection of significant genes, SNPs, and proteins in presence of dependence of data, valid control of false discovery rate for dependent test statistics, and evaluation of treatment effects on a group of molecules. The strength and weakness of each proposed method will be critically analyzed via theoretical investigations and simulation studies. Related software will be developed for free dissemination. Data sets from ongoing autism research, Alzheimer’s disease, and other biomedical studies will be analyzed by using the newly developed methods and the results will be further biologically confirmed and investigated. The research findings will have strong impact on statistical analysis of high throughput big data for biomedical research and on understanding heterogeneity for precision medicine and molecular mechanisms of autism, Alzheimer’s disease, and other diseases.
期刊论文(89)
专著(0)
科研奖励(0)
会议论文
DOI: 10.1080/01621459.2012.656041
发表时间: 2012
期刊: Journal of the American Statistical Association
影响因子: 3.7
作者: [Fan J, Li Y, Yu K]
通讯作者: Yu K
DOI: --
发表时间: 2021-08
期刊: Journal of machine learning research : JMLR
影响因子: --
作者: [Fan J, Jiang B, Sun Q]
通讯作者: Sun Q
PROJECTED PRINCIPAL COMPONENT ANALYSIS IN FACTOR MODELS.
在因子模型中预计主成分分析。
DOI: 10.1214/15-aos1364
发表时间: 2016-02
期刊: Annals of statistics
影响因子: 4.5
作者: [Fan J, Liao Y, Wang W]
通讯作者: Wang W
DOI: 10.1198/jasa.2011.tm09779
发表时间: 2011-06
期刊: Journal of the American Statistical Association
影响因子: 3.7
作者: [Fan J, Feng Y, Song R]
通讯作者: Song R
53