课题基金 / 基金详情

BIGDATA: F: Scalable and Interpretable Machine Learning: Bridging Mechanistic and Data-Driven Modeling in the Biological Sciences

BIGDATA: F: Scalable and Interpretable Machine Learning: Bridging Mechanistic and Data-Driven Modeling in the Biological Sciences
BIGDATA:F:可扩展和可解释的机器学习:桥接生物科学中的机械和数据驱动建模
批准号:
1741340
负责人:
Bin Yu
金额:
$90.0万
依托单位国家:
美国
项目类别:
Standard Grant
财政年份:
2017
资助国家:
美国
项目状态:
已结题
起止时间:
2017-10-01 至 2022-09-30

项目摘要

项目成果

Bin Yu的其他基金

相似基金

相关文献

中文摘要
翻译
随着信息技术的快速发展,几乎每个科学领域都出现了丰富的数据时代。这些数据有可能指导决策,并加速对人类发育和疾病进展等复杂过程的理解。例如,关于基因表达和其他分子过程的大量数据库可用于构建模型,以预测疾病的驱动因素。预测模型是理解这些复杂系统的重要一步,但同样重要的是这些模型的人类可解释性,例如,为了确定适当的治疗过程,对什么因素驱动疾病发作的机制进行深入了解。 下一代测序(NGS)技术已经导致了生物数据收集方式的深刻转变,分析了作为有组织的立体特异性组的一部分的单个基因组元件,以驱动新兴的生物现象。这些现代数据需要新的统计/数据科学原理和可扩展的算法来推进科学前沿。该项目专注于开发新型可扩展的统计机器学习算法,这些算法是可预测的,稳定的和可解释的,可用于指导生物系统的决策和发现。该项目旨在通过开发具有最先进预测准确性的可解释和稳定的监督学习算法沿着可扩展的开源软件,深入了解单个基因组元素如何协同作用。许多具有最先进预测精度的机器学习算法能够学习复杂的规则,这些规则可能会管理复杂的系统,但人类很难解释。该研究建立在迭代随机森林(iRF)的基础上,这是PI最近开发的一种算法,可以恢复高阶,人类可解释的布尔型交互,这些交互是随机森林中最先进的预测准确性的重要组成部分。拟议的工作将开发和验证方法,用于完善由iRF恢复的相互作用,以产生可检验的假设,用于后续研究,沿着推理方法,以评估与这些假设相关的不确定性。这些方法将在Apache Spark中实现,以确保可扩展到基因组学及其他领域的大规模数据集。 大规模应用程序方法的实施将利用商业云服务提供商和NSF之间的协议提供的云计算资源,用于BIGDATA征集。
英文摘要
With the rapid advances in information technology, an age of rich data has dawned in nearly every scientific field. Such data hold the potential to guide decision-making and accelerate understanding of complex processes such as human development and disease progression. For instance, massive databases on gene expression and other molecular processes can be used to build models to predict the drivers of a disease. Predictive models are an important step in understanding these complex systems, but equally important is the human interpretability of such models, e.g. to derive mechanistic insights into what factors drive disease onset in order to identify an appropriate course of treatment. Next Generation Sequencing (NGS) technologies have led to a profound shift in how biological data are collected, assaying individual genomic elements that act as part of organized, stereospecific groups to drive emergent biological phenomena. These modern data call for new statistics/data science principles and scalable algorithms to advance the frontier of science.This project focuses on developing novel scalable statistical machine learning algorithms that are predictable, stable and interpretable, and can be used to guide decision-making and discovery in biological systems. This project aims to build insights into how individual genomic elements act in concert by developing interpretable and stable supervised learning algorithms with state of the art predictive accuracy along with scalable, open source software. Many machine learning algorithms with state of the art predictive accuracies are capable of learning complicated rules that might govern complex systems but are difficult for humans to interpret. The research builds on iterative Random Forests (iRF), an algorithm recently developed by the PIs that recovers the high-order, human interpretable, Boolean type interactions that are important parts of the state-of-the-art predictive accuracy in Random Forests. The proposed work will develop and validate approaches for refining interactions recovered by iRF to produce testable hypotheses for follow-up studies, along with inference methods to assess the uncertainty associated with these hypotheses. These approaches and methods will be implemented in Apache Spark to ensure scalability to massive datasets in genomics and beyond. Implementation of the methods for the large-scale applications will leverage cloud computing resources provided through an agreement between commercial cloud service providers and NSF for the BIGDATA solicitation.
期刊论文(20)
专著(0)
科研奖励(0)
会议论文
Fast Interpretable Greedy-Tree Sums (FIGS)
快速可解释的贪婪树和(FIGS)
DOI: --
发表时间: 2023
期刊: ArXivorg
影响因子: --
作者: [Tan, Yan Shuo, Singh, Chandan, Nasseri, Keyan, Agarwal, Abhineet, Duncan, James, Ronen, Omer, Epland, Matthew, Kornblith, Aaron, Yu, Bin]
通讯作者: Yu, Bin
DOI: 10.1016/j.jbi.2021.103872
发表时间: 2021-09-14
期刊: JOURNAL OF BIOMEDICAL INFORMATICS
影响因子: 4.5
作者: [Altieri, Nicholas, Park, Briton, Yu, Bin]
通讯作者: Yu, Bin
The Data Science Process: One Culture
数据科学过程:一种文化
DOI: 10.1080/01621459.2020.1762615
发表时间: 2020
期刊: Journal of the American Statistical Association
影响因子: 3.7
作者: [Yu, Bin, Barter, Rebecca]
通讯作者: Barter, Rebecca
Unique Sharp Local Minimum in L1-Minimization Complete Dictionary Learning
L1-最小化完整字典学习中独特的尖锐局部最小值
DOI: --
发表时间: 2020
期刊: Journal of machine learning research
影响因子: 6
作者: [Wang, Yu, Wu, Siqi, Yu, Bin]
通讯作者: Yu, Bin
13
    Advancing Theory and Methodology for Tree-Based Algorithms in High Dimensions
    • 批准号:
      2209975
    • 项目类别:
      Standard Grant
    • 资助金额:
      $33.0万
    • 财政年份:
      2022
    • 负责人:
      Bin Yu
    • 依托单位:
    Understanding Complexity and the Bias-Variance Tradeoff in High Dimensions: Theory and Data Evidence
    • 批准号:
      2015341
    • 项目类别:
      Standard Grant
    • 资助金额:
      $30.0万
    • 财政年份:
      2020
    • 负责人:
      Bin Yu
    • 依托单位:
    Parallel Ensemble Learning and Feature Interaction Discovery: High Volume Dynamic Data
    • 批准号:
      1953191
    • 项目类别:
      Standard Grant
    • 资助金额:
      $45.2万
    • 财政年份:
      2020
    • 负责人:
      Bin Yu
    • 依托单位:
    Understand the functional mechanism of the DSP1 complex in the 3' end maturation of plant small nuclear RNAs
    • 批准号:
      1818082
    • 项目类别:
      Standard Grant
    • 资助金额:
      $68.26万
    • 财政年份:
      2018
    • 负责人:
      Bin Yu
    • 依托单位:
    国内基金
    海外基金
    Scalable Learning and Optimization: High-dimensional Models and Online Decision-Making Strategies for Big Data Analysis