Data science tools to identify robust exposure-phenotype associations for precision medicine
Data science tools to identify robust exposure-phenotype associations for precision medicine
批准号:
10874056
负责人:
ARJUN KUMAR MANRAI
金额:
$14.8万
依托单位:
依托单位国家:
美国
项目类别:
财政年份:
2021
资助国家:
美国
项目状态:
未结题
起止时间:
2021-09-10 至 2025-06-30
关键词:
AddressAll of Us Research ProgramBig DataBiologicalBiological FactorsBiological MarkersCardiologyCatalogsCohort StudiesCommunitiesComplexCountryDataData ScienceData SetDemographic FactorsDepositionDiabetes MellitusDiet and NutritionDisadvantagedDiseaseDisparityEnvironmentEnvironmental ExposureEnvironmental Risk FactorEpidemiologyEtiologyExhibitsGoalsHealthHeart DiseasesHumanIncidenceLeadLibrariesLinkLiteratureMachine LearningMalignant NeoplasmsMeasurementMeasuresMeta-AnalysisMetadataMethodsModelingNational Health and Nutrition Examination SurveyNational Institute of Environmental Health SciencesObservational StudyPhenotypePollutionPopulationPopulation HeterogeneityProcessReproducibilityResearch DesignResearch PersonnelResourcesRisk FactorsRoleSample SizeSamplingTestingTimeTranslationsUnited States National Institutes of HealthVariantanalytical methodbiobankcohortdata resourcedeep learningdisease disparitydisease phenotypedisorder riskenvironmental health disparityfeature selectiongenetic risk factorhealth differencehealth disparityhypercholesterolemiamachine learning methodnovelphenomeprecision medicinescale uptoolvibration
中文摘要
项目概要/摘要
在人口统计学上不同的人群中,表型变异是由环境因素驱动的。的
该提案的总体目标是部署数据科学方法,以推动发现
暴露(E)和表型(P)在人口统计学上不同的人群。我们缺乏数据科学方法,
将表型(P)和疾病中的干扰素组(E)的暴露变量关联、复制和优先化
发病率(D),需要提供精确的医疗。观察性研究充满了4个未解决的问题
数据科学挑战首先,基于电子的研究:(1)仅限于将一些假设的暴露-
表型对(E-P)的时间,导致一个支离破碎的文献的环境协会。机
然而,用于特征选择和预测的学习(ML)方法有希望,(2)大多数现存的基于E的
队列包含缺失的数据,挑战使用ML来检测复杂的E-P关联,第三,(3)偏差,
如混淆和研究设计影响联想和阻碍翻译。第四,(4)数量少
强大的数据资源,系统地记录纵向E-P和E-D关联,
大规模精准医疗将多个风险中的多个风险系统地联系起来是一个挑战,
表型,并在整个队列中复制这些关联。(Aim 1)。"效应的振动",或程度
其中关联作为研究设计的函数而改变(例如,分析方法、样本量)和模型
选择是观察性研究中的一种隐藏偏倚(目标2)。第三,一个突出的问题是,
环境差异导致健康差距。为了应对这些挑战和差距,我们建议Aim
1:开发和测试机器学习方法,将多种环境暴露指标与
多种表型:EP-WAS。我们假设,暴露将解释大量的变化,
表型,并将所有数据和模型存款在一个新的EP-WAS目录。目的2:定量
研究设计如何影响暴露生物标志物和表型之间的关联。我们会扩大规模,
扩展并测试一种称为"效应振动"(VoE)的方法,以衡量研究标准如何影响
关联的稳定性(关联的可复制性如何作为分析选择的函数)。目标3。杠杆
EP-WAS和VoE来解开表型的生物学、人口统计学和环境影响,
高胆固醇血症的差异。我们将在最大的队列中部署EP-WAS和VoE打包库
研究划分高胆固醇血症因素中人口统计学群体的表型变异。我们将
为生物医学界提供数据科学方法,以实现强大的数据驱动发现,
在观察数据集中解释确定表型因素,需要识别
环境卫生差距。调查人员将首次确定
心脏病的大规模环境正好赶上我们所有人的计划。
英文摘要
Project Summary/Abstract
Phenotypic variability across demographically diverse populations are driven by environmental factors. The
overall goal of this proposal is to deploy data science approaches to drive discovery of associations between
exposures (E) and phenotypes (P) in demographically diverse populations. We lack data science methods to
associate, replicate, and prioritize exposure variables of the exposome (E) in phenotypes (P) and disease
incidence (D), required for the delivery of precision medicine. Observational studies are fraught with 4 unsolved
data science challenges. First, E-based studies are: (1) limited to associating a few hypothesized exposure-
phenotype pairs (E-P) at a time, leading to a fragmented literature of environmental associations. Machine
learning (ML) approaches for feature selection and prediction hold promise, however, (2) most extant E-based
cohorts contain missing data, challenging the use of ML to detect complex E-P associations, Third, (3) biases,
such as confounding and study design influence associations and hinder translation. Fourth, (4) there are few
well-powered data resources that systematically document longitudinal E-P and E-D associations across
massive precision medicine. It is a challenge to systematically associate a number of exposures in multiple
phenotypes and replicate these associations across cohorts. (Aim 1). The “vibration of effects”, or the degree
to which associations change as a function of study design (e.g., analytic method, sample size) and model
choice is a hidden bias in observational studies (Aim 2). Third, an outstanding question is the degree to which
environmental differences lead to health disparities. To address these challenges and gaps, we propose to Aim
1: develop and test machine learning methods to associate multiple environmental exposure indicators with
multiple phenotypes: EP-WAS. We hypothesize that exposures will explain a significant amount of variation in
phenotype in populations and will deposit all data and models in a novel EP-WAS Catalog. Aim 2: Quantitate
how study design influences associations between exposure biomarkers and phenotype. We will scale up,
extend, and test a method called “vibration of effects” (VoE) to measure how study criteria influences the
stability of associations (how reproducible associations are as a function of analytic choice). Aim 3. Leverage
EP-WAS and VoE to disentangle biological, demographic, and environmental influences of phenotypic
disparities in hypercholesterolemia. We will deploy EP-WAS and VoE packaged libraries in the largest cohort
study to partition phenotypic variation across demographic groups in factors for hypercholesterolemia. We will
equip the biomedical community with data science approaches for robust data-driven discovery and
interpretation of exposure-phenotype factors in observational datasets, required for the identification of
environmental health disparities. For the first time, investigators will ascertain the collective role of the
environment in heart disease at scale just in time for the All of Us program.
期刊论文(6)
专著(0)
科研奖励(0)
会议论文
登录
查看更多内容
DOI:
10.1093/ije/dyaa164
发表时间:
2021-03-03
期刊:
International journal of epidemiology
影响因子:
7.7
作者:
[Klau S, Hoffmann S, Patel CJ, Ioannidis JP, Boulesteix AL]
通讯作者:
Boulesteix AL
Reply.
回复。
DOI:
10.1002/art.40923
发表时间:
2019
期刊:
Arthritis & rheumatology (Hoboken, N.J.)
影响因子:
--
作者:
[Kim,AlfredHJ, Strand,Vibeke, Atkinson,JohnP]
通讯作者:
Atkinson,JohnP
DOI:
10.1002/cncr.33341
发表时间:
2021-04-01
期刊:
CANCER
影响因子:
6.2
作者:
[Sohlberg, Ericka M., Thomas, I-Chun, Yang, Jaden, Kapphahn, Kristopher, Velaer, Kyla N., Goldstein, Mary K., Wagner, Todd H., Chertow, Glenn M., Brooks, James D., Patel, Chirag J., Desai, Manisha, Leppert, John T.]
通讯作者:
Leppert, John T.
DOI:
10.1038/s41598-022-08050-1
发表时间:
2022-03-08
期刊:
Scientific reports
影响因子:
4.6
作者:
[Poveda A, Pomares-Millan H, Chen Y, Kurbasic A, Patel CJ, Renström F, Hallmans G, Johansson I, Franks PW]
通讯作者:
Franks PW
DOI:
10.1016/j.urolonc.2021.08.011
发表时间:
2022-01
期刊:
UROLOGIC ONCOLOGY-SEMINARS AND ORIGINAL INVESTIGATIONS
影响因子:
2.7
作者:
[Velaer, Kyla, Thomas, I-Chun, Yang, Jaden, Kapphahn, Kristopher, Metzner, Thomas J., Golla, Abhinav, Hoerner, Christian R., Fan, Alice C., Master, Viraj, Chertow, Glenn M., Brooks, James D., Patel, Chirag J., Desai, Manisha, Leppert, John T.]
通讯作者:
Leppert, John T.
Data science tools to identify robust exposure-phenotype associations for precision medicine
-
批准号:10705899
-
项目类别:
-
资助金额:$9.25万
-
财政年份:2022
-
负责人:ARJUN KUMAR MANRAI
-
依托单位:
Precision Cardiovascular Medicine for Multi-Ethnic Populations
-
批准号:10582991
-
项目类别:
-
资助金额:$15.54万
-
财政年份:2022
-
负责人:ARJUN KUMAR MANRAI
-
依托单位:
Data science tools to identify robust exposure-phenotype associations for precision medicine
-
批准号:10653214
-
项目类别:
-
资助金额:$62.02万
-
财政年份:2021
-
负责人:ARJUN KUMAR MANRAI
-
依托单位:
Data science tools to identify robust exposure-phenotype associations for precision medicine
-
批准号:10487388
-
项目类别:
-
资助金额:$65.2万
-
财政年份:2021
-
负责人:ARJUN KUMAR MANRAI
-
依托单位:
Data science tools to identify robust exposure-phenotype associations for precision medicine
-
批准号:10095924
-
项目类别:
-
资助金额:$69.78万
-
财政年份:2021
-
负责人:ARJUN KUMAR MANRAI
-
依托单位:
Precision Cardiovascular Medicine for Multi-Ethnic Populations
-
批准号:9917879
-
项目类别:
-
资助金额:$16.1万
-
财政年份:2018
-
负责人:ARJUN KUMAR MANRAI
-
依托单位:
海外基金