Machine learning analysis of tandem mass spectra
Machine learning analysis of tandem mass spectra
批准号:
8288063
负责人:
William Stafford Noble
金额:
$62.17万
依托单位:
依托单位国家:
美国
项目类别:
财政年份:
2007
资助国家:
美国
项目状态:
已结题
起止时间:
2007-03-01 至 2015-05-31
关键词:
Amino Acid SequenceAreaBiologicalBiological MarkersC-PeptideCellsClinicalCollectionComplexComplex MixturesComputer softwareComputing MethodologiesDataData SetDatabasesDependencyDiagnosticDimensionsDisciplineDiseaseFertilizationGoalsGraphHealthHumanIonsKnowledgeLifeLiquid ChromatographyMachine LearningMass Spectrum AnalysisMethodsModelingMolecularMonitorNatural Language ProcessingPathway interactionsPeptide Sequence DeterminationPeptidesPhasePost-Translational Protein ProcessingProbabilityProcessProtein IsoformsProteinsProteomicsProtocols documentationReactionRelative (related person)ReproducibilityResearch PersonnelSamplingScanningSeriesSet proteinShotgunsSpecific qualifier valueStatistical ModelsTimeVariantWorkbasecomputer based statistical methodscomputerized toolsdesigndisease phenotypeenvironmental stressorimprovedinterestliquid chromatography mass spectrometrymass spectrometernovelprognosticresearch studyresponsespeech recognitionstatisticstandem mass spectrometrytool
中文摘要
描述(申请人提供):蛋白质是活细胞中的主要功能分子,串联质谱仪为高通量研究蛋白质提供了最有效的手段。该提案旨在使用机器学习、统计学和自然语言处理领域的最先进方法来提高我们理解大型串联质谱学数据集的能力。该提案的核心是一种被称为动态贝叶斯网络的概率模型,它允许我们高效而准确地对复杂的序列数据集进行推理。该建模框架利用了自然语言处理和语音识别领域的大量相关工作。许多以前的工作还没有被计算生物学家利用,所以这项提议代表了一种有价值的跨学科交叉。更具体地说,这个项目使用了一组协作的动态贝叶斯网络来联合建模整个质谱学实验。与现有的大多数分析质谱学数据的方法相比,该方法倾向于将实验分析分为一系列小的独立子任务,所提出的统一模型联合考虑了所有可用的数据。因此,这种方法可以利用光谱之间以及沿着数据的不同维度的有价值的相关性。动态贝叶斯网络还提供了一个严格的框架,用于从观察数据和定性专家知识的组合中执行推理。该项目分为五个目标,每个目标都涉及一种特定类型的质谱学实验。这些实验包括:(1)使用标准的质谱学方法识别给定复杂生物样本中的所有蛋白质;(2)使用改进的方法识别蛋白质,其中质谱仪以系统而不是依赖数据的方式对数据进行采样,目的是识别丰度较低的蛋白质;(3)量化生物样本内或生物样本之间的蛋白质的相对丰度;(4)识别翻译后修饰的蛋白质或包含序列变异的蛋白质;以及(5)对特定的一组蛋白质进行有针对性的定量,例如感兴趣的途径中的蛋白质或蛋白质生物标记物。这项提案中描述的方法有可能极大地提高我们从高通量的猎枪蛋白质组学实验中得出结论并提出假设的能力。例如,如上所述的实验可以识别基础疾病过程中涉及的蛋白质,识别以前未知的蛋白质异构体,或者量化蛋白质对环境应激源或疾病状态的反应。
英文摘要
DESCRIPTION (provided by applicant): Proteins are the primary functional molecules in living cells, and tandem mass spectrometry provides the most efficient means of studying proteins in a high-throughput fashion. The proposal aims to use state-of-the-art methods from the fields of machine learning, statistics and natural language processing to improve our ability to make sense of large tandem mass spectrometry data sets. The core of the proposal is a type of probabilistic model, known as a dynamic Bayesian network that allows us to reason efficiently and accurately about complex sequential data sets. This modeling framework leverages a large body of related work from the fields of natural language processing and speech recognition. Much of this prior work has not yet been exploited by computational biologists, so the proposal represents a valuable cross-fertilization across disciplines. More specifically, this project employs a collection of cooperating dynamic Bayesian networks to model jointly an entire mass spectrometry experiment. Relative to most existing methods for analyzing mass spectrometry data, which tend to divide the analysis of an experiment into a series of small independent subtasks, the proposed unified model jointly, considers all of the available data. This approach can thus exploit valuable dependencies among spectra and along various dimensions of the data. Dynamic Bayesian networks also provide a rigorous framework for performing inference from a combination of observed data and qualitative expert knowledge. The project is divided into five aims, each of which concerns a particular type of mass spectrometry experiment. These experiments involve (1) identifying all of the proteins in a given complex biological sample using a standard mass spectrometry protocol; (2) identifying proteins using a modified protocol in which the mass spectrometer samples the data in a systematic, rather than data-dependent, fashion, with the goal of identifying lower abundance proteins; (3) quantifying the relative abundance of proteins within or between biological samples; (4) identifying post-translational modified proteins or proteins that contain sequence variation; and (5) performing targeted quantification of a specified set of proteins, such as proteins in a pathway of interest or protein biomarkers. The methods described in this proposal have the potential to dramatically improve our ability to draw conclusions from and formulate hypotheses on the basis of high-throughput shotgun proteomics experiments. Experiments like the ones described above can, for example, identify proteins involved in fundamental disease processes, identify previously unknown protein isoforms, or quantify the re- sponses of proteins to environmental stressors or disease states.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
Deep tensor genomic imputation
-
批准号:10557916
-
项目类别:
-
资助金额:$38.38万
-
财政年份:2021
-
负责人:William Stafford Noble
-
依托单位:
Deep tensor genomic imputation
-
批准号:10096947
-
项目类别:
-
资助金额:$39.86万
-
财政年份:2021
-
负责人:William Stafford Noble
-
依托单位:
Optimization and joint modeling for peptide detection by tandem mass spectrometry
-
批准号:9214942
-
项目类别:
-
资助金额:$33.23万
-
财政年份:2017
-
负责人:William Stafford Noble
-
依托单位:
Project 2: UW-CNOF Data Analysis and Modeling
-
批准号:9021413
-
项目类别:
-
资助金额:$63.28万
-
财政年份:2015
-
负责人:William Stafford Noble
-
依托单位:
University of Washington Center for Nuclear Organization and Function
-
批准号:9983850
-
项目类别:
-
资助金额:$27.7万
-
财政年份:2015
-
负责人:William Stafford Noble
-
依托单位:
University of Washington Center for Nuclear Organization and Function
-
批准号:9353379
-
项目类别:
-
资助金额:$229.07万
-
财政年份:2015
-
负责人:William Stafford Noble
-
依托单位:
University of Washington Center for Nuclear Organization and Function
-
批准号:9916567
-
项目类别:
-
资助金额:$8.44万
-
财政年份:2015
-
负责人:William Stafford Noble
-
依托单位:
Machine learning methods to impute and annotate epigenomic maps
-
批准号:8814095
-
项目类别:
-
资助金额:$28.51万
-
财政年份:2014
-
负责人:William Stafford Noble
-
依托单位:
Machine learning methods to impute and annotate epigenomic maps
-
批准号:8925082
-
项目类别:
-
资助金额:$28.29万
-
财政年份:2014
-
负责人:William Stafford Noble
-
依托单位:
BIGDATA: DA: Interpreting massive genomic data sets via summarization
-
批准号:8642168
-
项目类别:
-
资助金额:$20.78万
-
财政年份:2013
-
负责人:William Stafford Noble
-
依托单位:
BIGDATA: DA: Interpreting massive genomic data sets via summarization
-
批准号:8840551
-
项目类别:
-
资助金额:$21.9万
-
财政年份:2013
-
负责人:William Stafford Noble
-
依托单位:
BIGDATA: DA: Interpreting massive genomic data sets via summarization
-
批准号:8599826
-
项目类别:
-
资助金额:$21.48万
-
财政年份:2013
-
负责人:William Stafford Noble
-
依托单位:
The MEME suite of motif-based sequence analysis tools
-
批准号:8324604
-
项目类别:
-
资助金额:$32.59万
-
财政年份:2009
-
负责人:William Stafford Noble
-
依托单位:
The MEME suite of motif-based sequence analysis tools
-
批准号:8129528
-
项目类别:
-
资助金额:$32.59万
-
财政年份:2009
-
负责人:William Stafford Noble
-
依托单位:
Machine learning analysis of tandem mass spectra
-
批准号:8038072
-
项目类别:
-
资助金额:$62.76万
-
财政年份:2007
-
负责人:William Stafford Noble
-
依托单位:
Machine learning analysis of tandem mass spectra
-
批准号:7194479
-
项目类别:
-
资助金额:$62.39万
-
财政年份:2007
-
负责人:William Stafford Noble
-
依托单位:
Machine learning analysis of tandem mass spectra
-
批准号:7797540
-
项目类别:
-
资助金额:$59.36万
-
财政年份:2007
-
负责人:William Stafford Noble
-
依托单位:
Machine learning analysis of tandem mass spectra
-
批准号:7581004
-
项目类别:
-
资助金额:$60.77万
-
财政年份:2007
-
负责人:William Stafford Noble
-
依托单位:
Machine learning analysis of tandem mass spectra
-
批准号:8470188
-
项目类别:
-
资助金额:$60.22万
-
财政年份:2007
-
负责人:William Stafford Noble
-
依托单位:
Machine learning analysis of tandem mass spectra
-
批准号:7365198
-
项目类别:
-
资助金额:$60.25万
-
财政年份:2007
-
负责人:William Stafford Noble
-
依托单位:
国内基金
海外基金
层出镰刀菌氮代谢调控因子AreA 介导伏马菌素 FB1 生物合成的作用机理
-
批准号:2021JJ40433
-
项目类别:省市级项目
-
资助金额:--
-
批准年份:2021
-
负责人:孙磊
-
依托单位:
寄主诱导梢腐病菌AreA和CYP51基因沉默增强甘蔗抗病性机制解析
-
批准号:32001603
-
项目类别:青年科学基金项目
-
资助金额:24.0万元
-
批准年份:2020
-
负责人:段真珍
-
依托单位:
AREA国际经济模型的移植.改进和应用
-
批准号:18870435
-
项目类别:面上项目
-
资助金额:2.0万元
-
批准年份:1988
-
负责人:史树中
-
依托单位: