课题基金 / 基金详情

Collaborative Research: ABI Innovation: Interpretable Machine Learning to Identify Molecular Markers for Complex Phenotypes

Collaborative Research: ABI Innovation: Interpretable Machine Learning to Identify Molecular Markers for Complex Phenotypes
合作研究:ABI 创新:可解释的机器学习来识别复杂表型的分子标记
批准号:
1759487
负责人:
Su-In Lee
金额:
$149.93万
依托单位:
依托单位国家:
美国
项目类别:
Continuing Grant
财政年份:
2018
资助国家:
美国
项目状态:
已结题
起止时间:
2018-06-01 至 2024-05-31

项目摘要

项目成果

Su-In Lee的其他基金

相似基金

相关文献

中文摘要
翻译
生物学家现在能够从特定组织中收集完整的基因表达数据和特定目标的蛋白质浓度。这些分子的存在和浓度在确定特定发育状态或疾病的诊断模式时可以作为特征。本研究中采用的生物标志物鉴定方法试图找到一组最能预测结果的特征(这里是基因表达水平)(疾病中发生的蛋白质水平)。识别的特征,生物标记物,可以帮助确定这种情况的分子基础。不幸的是,假阳性的生物标记物是非常常见的,正如在独立数据集中复制的低成功率所证明的那样,因此这些标记物在临床实践中的诊断等应用中变得很重要。我们试图通过使用新颖的、理论上有充分基础的机器学习(ML)方法从数据中学习可解释的模型,解决当前方法的基本问题,从而从根本上改变生物标志物发现的当前范式,并在模式生物中进行系统的实验验证系统。我们正在使用的疾病模型是针对阿尔茨海默病(AD)的,这是一个紧迫的国家和国际研究重点。淀粉样斑块和神经原纤维缠结是阿尔茨海默病的标志,它们的组成部分分别是淀粉样蛋白和tau蛋白。这些蛋白质可以从人类脑组织中精确测量,就像全球基因表达值一样。目前,我们缺乏对影响斑块和缠结形成的一组基因的理解,或者对这些有毒肽的任何保护性或病理反应。利用高通量分子数据(如基因表达数据)发现生物标志物,大大提高了我们对分子生物学和遗传学的认识。目前的方法试图找到一组最能预测表型的特征(例如,基因表达水平),并使用选定的特征,分子标记,来确定表型的分子基础。然而,在独立数据中复制的低成功率表明这种方法存在三个基本问题。首先,高维、隐变量和特征相关性在可预测性(即统计关联)和真正的生物相互作用之间造成了差异;我们需要新的特征选择标准来使模型更好地解释而不是简单地预测表型。其次,复杂模型(如深度学习或集成模型)可以比简单的线性模型更准确地描述基因和表型之间的复杂关系,但它们缺乏可解释性。第三,在不进行干预性实验的情况下分析观测数据并不能证明因果关系。为了解决这些问题,我们提出了一种集成的机器学习方法,通过1)选择可解释的特征,2)做出可解释的预测,以及3)通过干预实验验证和改进预测,从数据中学习可解释的模型。这种方法有以下目的:开发NEBULA(基于网络的无监督特征学习)框架,以学习可解释的特征,这些特征可能会从公开可用的多组数据集中提供有意义的表型解释。2. 开发一个统一的框架,称为SHAP (Shapley加性解释),通过估计每个特征对特定预测的重要性来解释复杂模型的预测。通过介入实验验证和完善预测,使用高通量基因敲除分析强大的线虫蛋白质毒性模型。欲了解更多信息,请参阅该项目的网站:http://suinlee.cs.washington.edu/projects/im3.This该奖项反映了美国国家科学基金会的法定使命,并通过基金会的智力价值和更广泛的影响审查标准进行评估,认为值得支持。
英文摘要
Biologists are now able to gather complete sets of gene expression data and protein concentrations for particular targets from specific tissues. The presence and concentrations of these molecules serve as features when determining a diagnostic pattern for specific states of development or disease. The approach to biomarker identification taken in this research attempts to find a set of features (here, gene expression levels) that best predict an outcome (protein levels occurring in the condition). The identified features, biomarkers, can help determine the molecular basis for the condition. Unfortunately, false positive biomarkers are very common, as evidenced by low success rates of replication in independent data sets and therefore low success in such markers becoming important in applications such as diagnostics in clinical practice. We seek to radically shift the current paradigm in biomarker discovery by resolving fundamental problems with the current approach by using novel, theoretically well-founded machine learning (ML) methods to learn interpretable models from data, and follow this up with a systematic experimental validation system in model organisms. The disease model we are using is for Alzheimer's disease (AD), an urgent national and international research priority. Amyloid plaques and neurofibrillary tangles are the hallmark of AD, and their building blocks are Amyloid-alpha and tau proteins, respectively. These proteins can be measured accurately from human brain tissues, as can global gene expression values. At present, we lack an understanding of the set of genes that affect formation of plaques and tangles, or any protective or pathological responses to these toxic peptides. Biomarker discovery using high-throughput molecular data (e.g., gene expression data) has significantly advanced our knowledge of molecular biology and genetics. The current approach attempts to find a set of features (e.g., gene expression levels) that best predict a phenotype and use the selected features, molecular markers, to determine the molecular basis for the phenotype. However, the low success rates of replication in independent data indicate three fundamental problems with this approach. First, high-dimensionality, hidden variables, and feature correlations create a discrepancy between predictability (i.e., statistical associations) and true biological interactions; we need new feature selection criteria to make the model better explain rather than simply predict phenotypes. Second, complex models (e.g., deep learning or ensemble models) can more accurately describe intricate relationships between genes and phenotypes than simpler, linear models, but they lack interpretability. Third, analyzing observational data without conducting interventional experiments does not prove causal relations. To address these problems, we propose an integrated machine learning methodology for learning interpretable models from data by 1) selecting interpretable features, 2) making interpretable predictions, and 3) validating and refining predictions through interventional experiments. This approach has the following aims:1. Develop NEBULA (network-based unsupervised feature learning) framework to learn interpretable features that will likely provide meaningful phenotype explanations from publicly available multi-omic data sets. 2. Develop a unified framework, called SHAP (Shapley additive explanation), to interpret the predictions of complex models by estimating the importance of each feature to a particular prediction.3. Validate and refine predictions through interventional experiments using high-throughput assays of gene knockdown on powerful nematode models of proteotoxicity. For further information see the project website at: http://suinlee.cs.washington.edu/projects/im3.This award reflects NSF's statutory mission and has been deemed worthy of support through evaluation using the Foundation's intellectual merit and broader impacts review criteria.
期刊论文(3)
专著(0)
科研奖励(0)
会议论文
DOI: 10.1038/s42256-019-0138-9
发表时间: 2020-01-01
期刊: NATURE MACHINE INTELLIGENCE
影响因子: 23.8
作者: [Lundberg, Scott M., Erion, Gabriel, Lee, Su-In]
通讯作者: Lee, Su-In
DOI: 10.1038/s41467-021-25680-7
发表时间: 2021-09-10
期刊: Nature communications
影响因子: 16.6
作者: [Beebe-Wang N, Celik S, Weinberger E, Sturmfels P, De Jager PL, Mostafavi S, Lee SI]
通讯作者: Lee SI
DOI: 10.1038/s42256-021-00343-w
发表时间: 2021-05-31
期刊: NATURE MACHINE INTELLIGENCE
影响因子: 23.8
作者: [Erion, Gabriel, Janizek, Joseph D., Lee, Su-In]
通讯作者: Lee, Su-In
CAREER: Learning the Chromatin Network from ChIP-Seq Data
  • 批准号:
    1552309
  • 项目类别:
    Continuing Grant
  • 资助金额:
    $76.83万
  • 财政年份:
    2016
  • 负责人:
    Su-In Lee
  • 依托单位:
ABI Innovation: A Probabilistic Approach to Meta-Analysis of Biological Network Interface
  • 批准号:
    1355899
  • 项目类别:
    Standard Grant
  • 资助金额:
    $68.97万
  • 财政年份:
    2014
  • 负责人:
    Su-In Lee
  • 依托单位:
国内基金
海外基金
Research on Quantum Field Theory without a Lagrangian Description
  • 批准号:
    24ZR1403900
  • 项目类别:
    省市级项目
  • 资助金额:
    --
  • 批准年份:
    2024
  • 负责人:
    SATOSHI NAWATA
  • 依托单位:
Cell Research
Cell Research
Cell Research (细胞研究)