课题基金 / 基金详情

BIGDATA: F: Scalable and Interpretable Machine Learning: Bridging Mechanistic and Data-Driven Modeling in the Biological Sciences

BIGDATA: F: Scalable and Interpretable Machine Learning: Bridging Mechanistic and Data-Driven Modeling in the Biological Sciences
BIGDATA:F:可扩展和可解释的机器学习:桥接生物科学中的机械和数据驱动建模
批准号:
1741340
负责人:
Bin Yu
金额:
$90.0万
依托单位国家:
美国
项目类别:
Standard Grant
财政年份:
2017
资助国家:
美国
项目状态:
已结题
起止时间:
2017-10-01 至 2022-09-30

项目摘要

项目成果

Bin Yu的其他基金

相似基金

相关文献

中文摘要
翻译
随着信息技术的飞速发展,几乎每个科学领域都进入了一个数据丰富的时代。这些数据有可能指导决策,加速对人类发育和疾病进展等复杂过程的理解。例如,关于基因表达和其他分子过程的大量数据库可以用来建立预测疾病驱动因素的模型。预测模型是理解这些复杂系统的重要一步,但同样重要的是此类模型的人类可解释性,例如,获得对哪些因素导致疾病发病的机械性见解,以便确定适当的疗程。下一代测序(NGS)技术已经导致了生物数据收集方式的深刻变化,分析了作为有组织的、立体特异的群体的一部分的单个基因组元素,以驱动新兴的生物现象。这些现代数据需要新的统计/数据科学原理和可扩展的算法来推动科学的前沿。该项目致力于开发新的可扩展的统计机器学习算法,这些算法是可预测的、稳定的和可解释的,可以用于指导生物系统的决策和发现。该项目旨在通过开发具有最先进预测准确性的可解释和稳定的监督学习算法以及可扩展的开放源代码软件,来深入了解单个基因组元素是如何协同工作的。许多具有最先进预测精度的机器学习算法能够学习复杂的规则,这些规则可能管理复杂的系统,但人类很难解释。这项研究建立在迭代随机森林(IRF)的基础上,IRF是PI最近开发的一种算法,它恢复了高阶的、人类可解释的布尔型相互作用,这些相互作用是随机森林预测精度的重要组成部分。拟议的工作将开发和验证改进IRF恢复的相互作用的方法,以便为后续研究产生可检验的假设,以及评估与这些假设相关的不确定性的推理方法。这些途径和方法将在ApacheSpark中实现,以确保基因组学和其他领域的海量数据集的可扩展性。针对大规模应用的方法的实施将利用商业云服务提供商与NSF之间的协议提供的云计算资源,以进行BigData招标。
英文摘要
With the rapid advances in information technology, an age of rich data has dawned in nearly every scientific field. Such data hold the potential to guide decision-making and accelerate understanding of complex processes such as human development and disease progression. For instance, massive databases on gene expression and other molecular processes can be used to build models to predict the drivers of a disease. Predictive models are an important step in understanding these complex systems, but equally important is the human interpretability of such models, e.g. to derive mechanistic insights into what factors drive disease onset in order to identify an appropriate course of treatment. Next Generation Sequencing (NGS) technologies have led to a profound shift in how biological data are collected, assaying individual genomic elements that act as part of organized, stereospecific groups to drive emergent biological phenomena. These modern data call for new statistics/data science principles and scalable algorithms to advance the frontier of science.This project focuses on developing novel scalable statistical machine learning algorithms that are predictable, stable and interpretable, and can be used to guide decision-making and discovery in biological systems. This project aims to build insights into how individual genomic elements act in concert by developing interpretable and stable supervised learning algorithms with state of the art predictive accuracy along with scalable, open source software. Many machine learning algorithms with state of the art predictive accuracies are capable of learning complicated rules that might govern complex systems but are difficult for humans to interpret. The research builds on iterative Random Forests (iRF), an algorithm recently developed by the PIs that recovers the high-order, human interpretable, Boolean type interactions that are important parts of the state-of-the-art predictive accuracy in Random Forests. The proposed work will develop and validate approaches for refining interactions recovered by iRF to produce testable hypotheses for follow-up studies, along with inference methods to assess the uncertainty associated with these hypotheses. These approaches and methods will be implemented in Apache Spark to ensure scalability to massive datasets in genomics and beyond. Implementation of the methods for the large-scale applications will leverage cloud computing resources provided through an agreement between commercial cloud service providers and NSF for the BIGDATA solicitation.
期刊论文(20)
专著(0)
科研奖励(0)
会议论文
Fast Interpretable Greedy-Tree Sums (FIGS)
快速可解释的贪婪树和(FIGS)
DOI: --
发表时间: 2023
期刊: ArXivorg
影响因子: --
作者: [Tan, Yan Shuo, Singh, Chandan, Nasseri, Keyan, Agarwal, Abhineet, Duncan, James, Ronen, Omer, Epland, Matthew, Kornblith, Aaron, Yu, Bin]
通讯作者: Yu, Bin
DOI: 10.1016/j.jbi.2021.103872
发表时间: 2021-09-14
期刊: JOURNAL OF BIOMEDICAL INFORMATICS
影响因子: 4.5
作者: [Altieri, Nicholas, Park, Briton, Yu, Bin]
通讯作者: Yu, Bin
The Data Science Process: One Culture
数据科学过程:一种文化
DOI: 10.1080/01621459.2020.1762615
发表时间: 2020
期刊: Journal of the American Statistical Association
影响因子: 3.7
作者: [Yu, Bin, Barter, Rebecca]
通讯作者: Barter, Rebecca
Unique Sharp Local Minimum in L1-Minimization Complete Dictionary Learning
L1-最小化完整字典学习中独特的尖锐局部最小值
DOI: --
发表时间: 2020
期刊: Journal of machine learning research
影响因子: 6
作者: [Wang, Yu, Wu, Siqi, Yu, Bin]
通讯作者: Yu, Bin
13
    Advancing Theory and Methodology for Tree-Based Algorithms in High Dimensions
    • 批准号:
      2209975
    • 项目类别:
      Standard Grant
    • 资助金额:
      $33.0万
    • 财政年份:
      2022
    • 负责人:
      Bin Yu
    • 依托单位:
    Understanding Complexity and the Bias-Variance Tradeoff in High Dimensions: Theory and Data Evidence
    • 批准号:
      2015341
    • 项目类别:
      Standard Grant
    • 资助金额:
      $30.0万
    • 财政年份:
      2020
    • 负责人:
      Bin Yu
    • 依托单位:
    Parallel Ensemble Learning and Feature Interaction Discovery: High Volume Dynamic Data
    • 批准号:
      1953191
    • 项目类别:
      Standard Grant
    • 资助金额:
      $45.2万
    • 财政年份:
      2020
    • 负责人:
      Bin Yu
    • 依托单位:
    Understand the functional mechanism of the DSP1 complex in the 3' end maturation of plant small nuclear RNAs
    • 批准号:
      1818082
    • 项目类别:
      Standard Grant
    • 资助金额:
      $68.26万
    • 财政年份:
      2018
    • 负责人:
      Bin Yu
    • 依托单位:
    国内基金
    海外基金
    Scalable Learning and Optimization: High-dimensional Models and Online Decision-Making Strategies for Big Data Analysis