课题基金 / 基金详情

Modeling, Inference, and Optimization for Genomic and Biomedical Big Data

Modeling, Inference, and Optimization for Genomic and Biomedical Big Data
基因组和生物医学大数据的建模、推理和优化
批准号:
10205870
负责人:
Kenneth L Lange
金额:
$53.92万
依托单位国家:
美国
项目类别:
财政年份:
2021
资助国家:
美国
项目状态:
未结题
起止时间:
2021-07-01 至 2026-05-31

项目摘要

项目成果

Kenneth L Lange的其他基金

相似基金

相关文献

中文摘要
翻译
点击翻译按钮获取中文摘要
英文摘要
Abstract The biomedical sciences are drowning in big data. Progress in fields such as genomics and medical imaging is being stymied by the lack of ap- propriate computational tools. This grant promotes the development of algorithms, statistical methods, and software for the analysis of the big datasets encountered in the biomedical sciences. The NIH All of Us Pro- gram, the Million Veteran Project (MVP) sponsored by US Department of Veterans Affairs (VA), and the UK Biobank are three prime examples of recent massive datasets. These datasets require terabytes of storage on sample sizes ranging from 105 to 106 and above subjects. The datasets are also dynamic, growing over time in size and complexity. In addition, the datasets are heterogeneous; for example, the UK Biobank offers ge- nomic data, electronic health record (EHR) data, and imaging data on the same study individuals. Finally, as with most real-world data, the data are fraught with missingness and inaccuracy. We propose attacking the issues of parameter estimation and model selection raised by such massive datasets. We will be guided by princi- ples of parsimony and high-dimensional optimization. Most of the specific applications we have in mind involve imaging and genomics, particularly genomewide association discovery. Fortunately, most of the tools and soft- ware we construct will be more generically useful. Our successful algo- rithms will be coded in the modern scientific programming language Julia and posted on publicly available websites. We will focus on constrained and sparse regression, EM and MM algorithms for optimization, variance components models, bootstrapping of linear mixed models, a copula-like model for correlated data, and sensitivity analysis in epidemic models. These are all subjects of paramount importance in modern genomics, bio- statistics and data mining.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
Modeling, Inference, and Optimization for Genomic and Biomedical Big Data
Modeling, Inference, and Optimization for Genomic and Biomedical Big Data
Statistical Methods for Gene Mapping
Training Grant in Genomic Analysis and Interpretation
海外基金