课题基金 / 基金详情

Probabilistic methods towards understanding complex human phenotypes using genomic and healthcare data

Probabilistic methods towards understanding complex human phenotypes using genomic and healthcare data
使用基因组和医疗数据理解复杂人类表型的概率方法
批准号:
RGPIN-2019-06216
负责人:
Li, Yue
金额:
$2.84万
依托单位:
依托单位国家:
加拿大
项目类别:
Discovery Grants Program - Individual
财政年份:
2022
资助国家:
加拿大
项目状态:
已结题
起止时间:
2022-01-01 至 2023-12-31

项目摘要

项目成果

Li, Yue的其他基金

相似基金

相关文献

中文摘要
翻译
大规模生物数据集的出现对现有的分析框架提出了挑战。大量的基因组图谱数据提供了将突变与特定组织中基因表达变化联系起来的分子基础。EHR系统的广泛采用创造了丰富的表型数据,包括诊断代码、实验室测试和问卷调查。这些数据为开发新的机器学习方法提供了有希望的场所,以阐明导致表型多样性和相互依赖的生物学机制。然而,由于缺乏可扩展的推理方法,现有的研究往往仅限于分析整个数据集的一小部分快照,而无法解释数据的稀疏、多模式、纵向、不规则采样和非随机缺失的性质。我们的长期愿景是开发新的机器学习方法,以人类可以理解的方式,根据遗传变异、细胞类型特异性、基因组调节元件、基因和途径功能以及它们与环境的相互作用来破译不同表型的病因学。为了实现这一愿景,我们提出了四个短期目标。我们将开发:1.贝叶斯模型,通过将基因、组织、实验室结果、诊断代码通过潜在的表型主题关联起来,解释异质数据分布的多形态,并预测复合生物标志物;2.生成模型,将相关的非随机缺失的实验室结果和对患者EHR中自我报告的问卷的答案以及新患者不可接触的组织样本中的基因表达归因于生成模型;3.无监督模型,根据不同患者的纵向和不规则抽样的门诊和住院病历来推断不同患者健康状态的潜在轨迹;4.分层贝叶斯网络,它利用从基因组数据推断出的序列突变的功能影响,并联合推断从遗传变异中定向的路径,致病基因和致病途径,以及表型。与现有的特别方法相比,我们建议的方法的关键创新是,我们同时学习我们建议的模型的所有组件(尽管它们很复杂),因此协调不同的数据集和互补的信息。我们通过利用概率论和深度学习技术的可扩展变分推理算法来实现这一点。这项拟议的研究将推动贝叶斯学习在挖掘海量异质数据方面的应用,并在医学上产生影响,包括复合生物标记物的发现、基于归因的临床建议、预测健康轨迹、个性化风险预测、推断因果突变和疾病风险的深度可解释模型。我们共同提出了通过对海量数据进行高效的贝叶斯集成来弥合基因组和表型组之间的差距的一步,从而提高了我们对从基因突变到广泛表型谱的级联事件的理解。
英文摘要
The advent of massive biological datasets challenge existing analytic frameworks. Large genomic profiling data confer the molecular basis to link mutations to gene expression changes in specific tissues. The broad adoption of EHR systems creates rich phenotypic data including diagnostic code, lab tests, and questionnaires. These data provide promising venues for developing novel machine learning methods to elucidate the biological mechanisms that give rise to the phenotypic diversities and interdependence. However, due to the lack of scalable inference methods, existing research is often limited to analyzing only a small snapshot of the entire datasets and unable to account for the sparse, multimodal, longitudinal, irregularly sampled, and non-missing-at-random nature of the data. Our long-term vision is to develop novel machine learning methods to decipher, in a human-understandable manner, the etiology of diverse phenotypes based on genetic variants, cell-type specificities, genomic regulatory elements, gene and pathway functions, and their interactions with environments. In pursuing this vision, we propose four short-term objectives. We will develop: 1. Bayesian model to account for the multi-modality of the heterogeneous data distributions and predict composite biomarkers by associating genes, tissues, lab results, diagnosis codes via latent phenotypic topics, 2. generative model to impute correlated non-randomly missing lab results and answers to self-reported questionnaires in patients' EHR and gene expression in inaccessible tissue samples of new patients, 3. unsupervised model to infer latent trajectory of diverse patients' health states based on their longitudinal and irregularly sampled outpatient and inpatient medical records, 4. hierarchical Bayesian network that leverages the functional impacts of sequence mutations inferred from genomic data and jointly infer the directed paths from driver genetic variants, causal genes and pathways, and to phenotypes. The key innovation of our proposed methods is that, in contrast to the existing ad hoc methods, we learn all components of our proposed models simultaneously (despite their complexity) and therefore harmonize diverse datasets with complementary information. We achieve this by scalable variational inference algorithms that leverage probability theory and deep learning techniques. The proposed research will advance Bayesian learning for mining massive heterogenous data with impactful applications in medicine including composite biomarker discovery, imputation-based clinical recommendations, forecasting health trajectories, personalized risk predictions, deep interpretable models for inferring causal mutations and disease risks. Together, we present a step towards bridging the gap between the genome and the phenome by efficient Bayesian integrations of massive data, thereby improving our understanding of the cascading events from genetic mutations to a broad phenotypic spectrum.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
Probabilistic methods towards understanding complex human phenotypes using genomic and healthcare data
  • 批准号:
    RGPIN-2019-06216
  • 项目类别:
    Discovery Grants Program - Individual
  • 资助金额:
    $2.84万
  • 财政年份:
    2021
  • 负责人:
    Li, Yue
  • 依托单位:
Probabilistic methods towards understanding complex human phenotypes using genomic and healthcare data
  • 批准号:
    RGPIN-2019-06216
  • 项目类别:
    Discovery Grants Program - Individual
  • 资助金额:
    $2.84万
  • 财政年份:
    2020
  • 负责人:
    Li, Yue
  • 依托单位:
Probabilistic methods towards understanding complex human phenotypes using genomic and healthcare data
  • 批准号:
    DGECR-2019-00253
  • 项目类别:
    Discovery Launch Supplement
  • 资助金额:
    $0.91万
  • 财政年份:
    2019
  • 负责人:
    Li, Yue
  • 依托单位:
Probabilistic methods towards understanding complex human phenotypes using genomic and healthcare data
  • 批准号:
    RGPIN-2019-06216
  • 项目类别:
    Discovery Grants Program - Individual
  • 资助金额:
    $2.84万
  • 财政年份:
    2019
  • 负责人:
    Li, Yue
  • 依托单位:
国内基金
海外基金
复杂图像处理中的自由非连续问题及其水平集方法研究
  • 批准号:
    60872130
  • 项目类别:
    面上项目
  • 资助金额:
    28.0万元
  • 批准年份:
    2008
  • 负责人:
    刘国才
  • 依托单位:
Computational Methods for Analyzing Toponome Data