DMS/NIGMS 2: Statistical Methods and Computational Algorithms for Biobank Data
DMS/NIGMS 2: Statistical Methods and Computational Algorithms for Biobank Data
批准号:
2054253
负责人:
Hua Zhou
金额:
$95.58万
依托单位国家:
美国
项目类别:
Continuing Grant
财政年份:
2021
资助国家:
美国
项目状态:
未结题
起止时间:
2021-07-01 至 2025-06-30
中文摘要
生物样本库数据的特点是其体积、速度、种类和准确性(4V)。两个典型的例子是美国退伍军人事务部(VA)的百万老兵项目(MVP)和英国生物银行(Biobank)。数据非常大,有多达100万个主题,占用了tb的存储空间(容量)。它们的样本量和数据内容不断增加(速度)。它们包含异构信息源:基因组、电子健康记录(EHR)、可穿戴设备、图像以及最近的COVID-19数据(品种)。此外,它们充满了缺失和不准确(真实性)。该项目旨在开发新的统计方法和计算算法,以解决4V的特定方面。这些方法是由主要研究人员最近分析MVP和UK Biobank数据的经验所激发的,并且可以推广到任何生物银行或其他通用大数据。这些方法为生物银行数据分析中一些最紧迫的问题提供了解决方案。这项工作将推动统计学、优化和遗传学的几个前沿。这项研究将与大量的教育和推广活动结合起来,包括开发新的课程和软件以及指导学生。这些活动旨在让包括女性和少数族裔在内的各类学生接触到用于大数据分析的最先进的统计和计算技术。要研究三组问题。(1)电子健康记录和可穿戴设备在生物银行中产生大量的纵向数据。在许多研究中,纵向结果的受试者内部变异性是主要的科学兴趣。受血压变异性和血糖变异性对糖尿病并发症影响研究的启发,pi提出了一种稳健且可扩展的方法,用于估计和推断时变和时不变预测因子对受试者内方差的影响。与现有方法相比,该方法对分布不规范具有较强的鲁棒性,且速度更快。计算可扩展性使其成为研究生物库中基于大量纵向数据的性状变异的有力工具。(2) pi将开发一类新的在线学习算法,该算法将统计学中的最大最小化原理与随机近端迭代算法相结合。新算法适用于更广泛的模型类别,并且被证明更稳定和健壮。它们有助于解决数量问题,并将应用于大量生物银行数据的全基因组关联研究。(3) pi提出了一种小自助(BLB)方法来估计在遗传学和生物统计学中起核心作用的大方差成分模型。由于巨大协方差矩阵的反转,拟合这样的模型对生物样本库数据是禁止的。BLB方法将大量的方差分量模型分解成许多较小的模型,这些模型并行启动,然后平均。新方法将使生物库数据中复杂性状的遗传力和遗传相关性量化成为可能。该奖项反映了美国国家科学基金会的法定使命,并通过使用基金会的知识价值和更广泛的影响审查标准进行评估,被认为值得支持。
英文摘要
Biobank data is characterized by its volume, velocity, variety, and veracity (4V). Two prime examples are the Million Veteran Project (MVP) at US Veterans Affairs (VA) and UK Biobank. The data are big, with up to a million subjects and occupying terabytes of storage (volume). Their sample sizes and data content keep increasing (velocity). They contain heterogeneous sources of information: genome, electronic health record (EHR), wearable devices, images, and most recently, COVID-19 data (variety). Furthermore, they are fraught with missingness and inaccuracy (veracity). This project seeks to develop novel statistical methods and computational algorithms that address specific aspects of 4V. The methods are motivated by the principal investigators' recent experience in analyzing MVP and UK Biobank data, and are generalizable to any biobank or other generic big data. The methods provide solutions to some of the most pressing issues in biobank data analysis. The work will push forward several frontiers in statistics, optimization, and genetics. The research will be integrated with substantial education and outreach activities, including developing new courses and software and mentoring students. These activities aim to expose a diverse set of students, including women and minorities, to state-of-the-art statistical and computational techniques for big data analysis. Three sets of problems are to be investigated. (1) Electronic health records and wearable devices generate a vast amount of longitudinal data in biobanks. In many studies, the within-subject variability of a longitudinal outcome is the primary scientific interest. Motivated by studies of the impacts of blood pressure variability and glycemic variability on diabetes complications, the PIs propose a robust and scalable method for the estimation and inference of the effects of both time-varying and time-invariant predictors on within-subject variance. Compared to existing approaches, the method is robust to the distribution misspecification and orders of magnitude faster. Computational scalability makes it a powerful tool for studying trait variability based on massive longitudinal data in biobanks. (2) The PIs will develop a new class of online learning algorithms, which combine the majorization-minimization principle in statistics and the stochastic proximal iteration algorithm. The new algorithms apply to a broader class of models and are demonstrably more stable and robust. They help solve the volume issue and will be applied to genome-wide association studies of massive biobank data. (3) The PIs propose a bag of little bootstraps (BLB) approach for estimating massive variance component models, which play a central role in genetics and biostatistics. Fitting such models is prohibitive for biobank data because of the inversion of the giant covariance matrix. The BLB approach breaks the massive variance component model into many smaller ones, which are bootstrapped in parallel and then averaged. The new method will enable quantifying heritability and genetic correlation of complex traits in biobank data.This award reflects NSF's statutory mission and has been deemed worthy of support through evaluation using the Foundation's intellectual merit and broader impacts review criteria.
期刊论文(35)
专著(0)
科研奖励(0)
会议论文
登录
查看更多内容
Efficient Algorithms and Implementation of a Semiparametric Joint Model for Longitudinal and Competing Risk Data: With Applications to Massive Biobank Data.
有效的算法和实施纵向和竞争风险数据的半参数联合模型:与大规模生物库数据的应用。
DOI:
10.1155/2022/1362913
发表时间:
2022
期刊:
Computational and mathematical methods in medicine
影响因子:
--
作者:
[Li S, Li N, Wang H, Zhou J, Zhou H, Li G]
通讯作者:
Li G
DOI:
10.1002/sim.9253
发表时间:
2022-02-20
期刊:
Statistics in medicine
影响因子:
2
作者:
[Doubleday K, Zhou J, Zhou H, Fu H]
通讯作者:
Fu H
DOI:
10.1214/21-aoas1491
发表时间:
2021-12
期刊:
The annals of applied statistics
影响因子:
--
作者:
[Kim J, Shen J, Wang A, Mehrotra DV, Ko S, Zhou JJ, Zhou H]
通讯作者:
Zhou H
ORTHOGONAL TRACE-SUM MAXIMIZATION: TIGHTNESS OF THE SEMIDEFINITE RELAXATION AND GUARANTEE OF LOCALLY OPTIMAL SOLUTIONS.
正交迹和最大化:半定松弛的严格性和局部最优解的保证。
DOI:
10.1137/21m1422707
发表时间:
2022
期刊:
SIAM journal on optimization : a publication of the Society for Industrial and Applied Mathematics
影响因子:
--
作者:
[Won,Joong-Ho, Zhang,Teng, Zhou,Hua]
通讯作者:
Zhou,Hua
DOI:
10.1093/jamiaopen/ooad006
发表时间:
2023-04
期刊:
JAMIA open
影响因子:
2.1
作者:
[]
通讯作者:
共 20 条
SCH: Statistical Foundation and Predictive Modeling for Personalized Diabetes Management: Continuous Glucose Monitoring (CGM), Electronic Health Records (EHR), and Biobanks
-
批准号:2205441
-
项目类别:Standard Grant
-
资助金额:$120.0万
-
财政年份:2022
-
负责人:Hua Zhou
-
依托单位:
Tensor Regressions and Applications in Neuroimaging Data Analysis
-
批准号:1645093
-
项目类别:Continuing Grant
-
资助金额:$8.56万
-
财政年份:2015
-
负责人:Hua Zhou
-
依托单位:
Tensor Regressions and Applications in Neuroimaging Data Analysis
-
批准号:1310319
-
项目类别:Continuing Grant
-
资助金额:$12.0万
-
财政年份:2013
-
负责人:Hua Zhou
-
依托单位:
海外基金