课题基金 / 基金详情

DMS/NIGMS 2: Deep learning for repository-scale analysis of tandem mass spectrometry proteomics data

DMS/NIGMS 2: Deep learning for repository-scale analysis of tandem mass spectrometry proteomics data
DMS/NIGMS 2:用于串联质谱蛋白质组数据存储库规模分析的深度学习
批准号:
2245300
负责人:
William Noble
金额:
$119.98万
依托单位:
依托单位国家:
美国
项目类别:
Continuing Grant
财政年份:
2023
资助国家:
美国
项目状态:
未结题
起止时间:
2023-06-15 至 2027-05-31

项目摘要

项目成果

William Noble的其他基金

相似基金

相关文献

中文摘要
翻译
蛋白质组学研究细胞中的主要功能分子,鉴定和定量复杂生物样品中的蛋白质,目的是了解它们在健康和疾病中的作用。蛋白质组学也是研究从土壤样品到海水样品等不同环境下微生物的基础。推动该领域快速发展的主要技术是串联质谱法。除了质谱硬件的技术进步外,对串联质谱仪产生的复杂数据进行准确有效的分析需要越来越复杂的算法工具。该项目将开发这些工具。特别是,该项目团队将开发机器学习软件,旨在提高科学家推断复杂样品中数千种蛋白质的身份和数量的能力。蛋白质组学研究界成功采用该项目开发的工具将影响广泛的研究,包括用于了解基本分子功能的模式生物蛋白质组学、人类疾病队列研究和环境蛋白质组学分析。这个项目产生的工具将使科学家能够检测更多的蛋白质,并更准确地量化它们的丰度在健康和疾病以及不同环境条件下的变化。推动该项目的中心假设是,通过使用深度神经网络来利用公共存储库中的数据,可以提高解释自下而上串联质谱数据的统计能力。该项目解决了一系列项目任务,每个任务都使用深度神经网络来解决质谱分析中的不同核心问题,并且每个任务都可以通过利用大量快速增长的公共质谱数据存储库(如PRIDE和massive)来改进。这四项任务涉及光谱的大规模聚类,以从头开始的方式将肽分配到观察到的光谱中,在定量质谱数据队列中输入缺失值,以及去噪质谱测量。这些任务很重要,因为(1)每个任务都代表了一个基本的分析挑战,一个解决方案有可能影响质谱蛋白质组学的各种下游应用,(2)每个任务都允许从存储库规模数据中进行机器学习的创新应用,(3)项目团队现有的质谱合作将直接受益于这些问题的解决方案。该奖项反映了美国国家科学基金会的法定使命,并通过使用基金会的知识价值和更广泛的影响审查标准进行评估,被认为值得支持。
英文摘要
The field of proteomics studies the primary functional molecules in the cell, identifying and quantifying proteins in complex biological samples with the goal of understanding their roles in health and disease. Proteomics is also fundamental to studies of microorganisms in diverse environment, ranging from soil samples to oceanwater samples. The primary technology driving the rapid growth of this field is tandem mass spectrometry. In addition to technological advances in mass spectrometry hardware, accurate and efficient analysis of the complex data produced by a tandem mass spectrometer requires increasingly sophisticated algorithmic tools. The project will develop these tools. In particular, the project team will develop machine learning software that aims to improve scientists' ability to infer the identities and quantities of thousands of proteins in a complex sample. Successful adoption by the proteomics research community of the tools developed by this project will impact a huge range of studies, including model organism proteomics to understand basic molecular function, human disease cohort studies, and environmental proteomics analyses. The tools produced by this project will allow scientists to to detect more proteins and to more accurately quantify how their abundances change in health and disease and across different environmental conditions.The central hypothesis driving this project is that statistical power in interpreting bottom-up tandem mass spectrometry data can be increased by using deep neural networks to leverage data in public repositories. The project addresses a series of project tasks, each of which uses deep neural networks to solve a different core problem in mass spectrometry analysis, and each of which can be improved by making use of massive and rapidly growing repositories of public mass spectrometry data, such as PRIDE and MassIVE. The four tasks address large-scale clustering of spectra, assigning peptides to observed spectra in a de novo fashion, imputing missing values in cohorts of quantitative mass spectrometry data, and de-noising mass spectrometry measurements. These tasks are important because (1) each one represents a fundamental analysis challenge, a solution for which has the potential to impact a wide variety of downstream applications in mass spectrometry proteomics, (2) each task allows for innovative applications of machine learning from repository-scale data, and (3) the project team has existing mass spectrometry collaborations that will directly benefit from solutions to these problems.This award reflects NSF's statutory mission and has been deemed worthy of support through evaluation using the Foundation's intellectual merit and broader impacts review criteria.
期刊论文(1)
专著(0)
科研奖励(0)
会议论文
DOI: 10.1021/acs.jproteome.3c00205
发表时间: 2023-10-20
期刊: JOURNAL OF PROTEOME RESEARCH
影响因子: 4.4
作者: [Harris,Lincoln, Fondrie,William E., Noble,William S.]
通讯作者: Noble,William S.
EAGER: Cloud-based analysis of mass spectrometry proteomics data
  • 批准号:
    1549932
  • 项目类别:
    Standard Grant
  • 资助金额:
    $30.0万
  • 财政年份:
    2015
  • 负责人:
    William Noble
  • 依托单位:
CAREER: Support Vector Methods for Functional Genomic Analysis
  • 批准号:
    0431725
  • 项目类别:
    Continuing Grant
  • 资助金额:
    $17.91万
  • 财政年份:
    2004
  • 负责人:
    William Noble
  • 依托单位:
Generative and Discriminative Methods for Gene Finding and Functional Annotation
  • 批准号:
    0243257
  • 项目类别:
    Standard Grant
  • 资助金额:
    $29.96万
  • 财政年份:
    2002
  • 负责人:
    William Noble
  • 依托单位:
CAREER: Support Vector Methods for Functional Genomic Analysis
  • 批准号:
    0093302
  • 项目类别:
    Continuing Grant
  • 资助金额:
    $44.51万
  • 财政年份:
    2001
  • 负责人:
    William Noble
  • 依托单位:
海外基金