Causal graphical methods for high-dimensional heterogeneous biomedical data
Causal graphical methods for high-dimensional heterogeneous biomedical data
批准号:
10625257
负责人:
Tyler Lovelace
金额:
$4.77万
依托单位国家:
美国
项目类别:
财政年份:
2022
资助国家:
美国
项目状态:
未结题
起止时间:
2022-03-21 至 2025-03-20
关键词:
AddressAlgorithmsAutomobile DrivingBiologicalCellsClinicalComplexDataData AnalyticsData SetDevelopmentDimensionsEventExplosionGenerationsGenesGraphHeterogeneityImmunologicsIndividualIntensive Care UnitsInterventionIntubationInvestigationLearningLifeMalignant NeoplasmsMeasuresMedicalMedicineMethodologyMethodsMiningModelingMortality DeterminantsMotivationOutcomePatientsPerformancePeriodicityPhenotypeProcessPropertyRNAResearchResearch PersonnelResolutionSkeletonStructureSystemTestingThe Cancer Genome AtlasTimeValidationVentilatorWorkanalytical methodcancer carecausal modelcell typeclinically relevantcohortcomplex datacopingexperimental studyflexibilitygene regulatory networkgraph learninghigh dimensionalityimprovedlearning algorithmlearning strategymachine learning methodmalignant breast neoplasmmethod developmentmodel designmortalitymultidimensional datamultiple omicsnovelpredictive modelingprognostic modelsingle-cell RNA sequencingtoolvector
中文摘要
在过去的十年里,从生物和生物医学系统收集的数据呈爆炸式增长
在类型和数量方面。挖掘这些高维、异质且通常是动态的数据集以
做出生物或医学上重要的推论或开发预测模型需要新的复杂
数据分析方法。新的机器学习方法开始填补这一空白,但这些方法中的大多数
生成缺乏清晰可解释性的“黑箱”模型。此外,这些方法是关联的,并且
因此不能梳理出数据集中特征之间的复杂因果关系。定向
因果图形模型(DCGM)是填补这一空白的有力工具。DCGMS,从观测中学习
数据集,可以表示变量之间的因果关系。这使得DCGMS可以生成以下假设
机制,并构建简约、因果信息的预测模型。然而,生物医学数据集
通常具有使在整个数据集上构建因果图形模型变得困难的功能。示例
包括:数据类型异构性、高维性、多重共线性、周期性和非平稳性。致信地址
对于这些问题,我建议开发学习数据集中因果关系图的方法,这些数据集中包含(1)一个
连续、分类和删失变量的异质混合,(2)高维和
多重共线性,以及(3)周期性和非平稳性。在目标1中,我将开发一个新的因果发现算法
这包括连续的、绝对的和被审查的变量(例如,生存)。在目标2中,我将测试和
比较各种矩阵分解和降维方法的学习能力
用于图学习方法的有意义的低维潜在特征空间。在目标3中,我将发展
一种在单细胞分辨率下在动态的、可能是循环的基因调控网络中发现因果关系的新方法。
在所有情况下,测试和验证都将在合成的和现实生活中公开可用的数据集上进行。这些
方法论的改进是在因果发现领域向前迈出的重要一步,它们可以
一起使用或单独使用,以提供灵活而强大的平台来分析各种
生物医学数据集。一旦可用,它们将使研究人员能够对因果关系做出推断
机制,生成假设,并建立健壮、简约的预测模型。
英文摘要
In the past decade, there has been an explosion of data collected from biological and biomedical systems, both
in terms of type and volume. Mining these high-dimensional, heterogeneous, and often dynamic datasets to
make biologically or medically important inferences or develop predictive models requires new sophisticated
data analytics methods. New machine learning methods have begun filling this gap, but most of these methods
generate “black box” models that lack clear interpretability. Additionally, these methods are associative, and are
thus incapable of teasing out the complex cause-effect relationships among features in the dataset. Directed
causal graphical models (DCGMs) are a powerful tool for filling this gap. DCGMs, learned from observational
datasets, can represent causal relationships between variables. This allows DCGMs to generate hypotheses of
mechanisms and construct parsimonious, causally informed predictive models. However, biomedical datasets
often have features that make it difficult to construct causal graphical models over the full dataset. Examples
include: data type heterogeneity, high dimensionality, multicollinearity, cyclicity, and nonstationarity. To address
these problems, I propose to develop methods for learning causal graphs in datasets containing (1) a
heterogeneous mixture of continuous, categorical, and censored variables, (2) high dimensionality and
multicollinearity, and (3) cyclicity and nonstationarity. In Aim 1, I will develop a new causal discovery algorithm
that accommodates continuous, categorical and censored variables (e.g., survival). In Aim 2, I will test and
compare various methods for matrix decomposition and dimensionality reduction in their ability to learn a
meaningful low-dimensional latent feature space to be used in graph learning methods. In Aim 3, I will develop
a new method for causal discovery in dynamic, possibly cyclic, gene regulatory networks at single cell resolution.
In all cases, testing and validation will be performed on synthetic and real-life publicly available datasets. These
methodological improvements constitute important steps forward in the field of causal discovery and they can
be utilized together or independently to provide a flexible and powerful platform for analysis of a wide range of
biomedical datasets. Once made available, they will enable researchers to make inferences about causal
mechanisms, generate hypotheses, and build robust, parsimonious predictive models.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
Causal graphical methods for high-dimensional heterogeneous biomedical data
-
批准号:10388447
-
项目类别:
-
资助金额:$4.68万
-
财政年份:2022
-
负责人:Tyler Lovelace
-
依托单位:
海外基金