A Computational Framework for Protein Identification and Quantification in Metaproteomics Using Data-Independent Acquisition
A Computational Framework for Protein Identification and Quantification in Metaproteomics Using Data-Independent Acquisition
批准号:
10047086
负责人:
Xuan Guo
金额:
$36.13万
依托单位国家:
美国
项目类别:
财政年份:
2020
资助国家:
美国
项目状态:
已结题
起止时间:
2020-07-01 至 2024-06-30
关键词:
AddressAffectAmino Acid SubstitutionBindingBioinformaticsBiological SciencesBiologyClinicComplexComputational BiologyDataDetectionDevelopmentDiabetes MellitusDiseaseEducational CurriculumFrequenciesGenomeGrainHealthHeterogeneityHomologous ProteinHumanHuman MicrobiomeHybridsImmune systemIn SituInflammatory Bowel DiseasesKnowledgeLeadLibrariesLinear ProgrammingMass Spectrum AnalysisMeasurementMedicalMetabolicMethodologyMethodsMicrobeModelingModernizationNatureObesityOrganismOutcomePeptidesPhysiologicalPhysiologyPlayPositioning AttributeProcessProtein DatabasesProteinsProteomicsReportingReproducibilityResearchResearch ActivityRoleSamplingScienceShotgunsTechniquesTrainingVariantbacterial communitybasebioinformatics toolcareercomputer frameworkcomputer sciencecomputerized toolsdata acquisitiondesigndysbiosisexperimental studygraduate studenthost microbiomehost-microbe interactionshuman microbiotaimprovedinsightmass spectrometermetaproteomicsmicrobialmicrobial communitymicrobiomemicrobiome analysismicrobiome researchmicrobiotamicroorganismnovelnovel diagnosticsnovel therapeuticspathogenprogramsstemsuccesstheoriestranscriptome sequencingundergraduate student
中文摘要
摘要
对人类微生物组特征的广泛努力极大地增加了对
微生物组的多样性及其在健康和疾病中的组成。人类微生物区系中的微生态失调
是许多疾病发展的基础,如肥胖症、糖尿病和炎症性肠病。
基于质谱学(MS)的蛋白质组学已被广泛应用于微生物组研究中
对微生物群落功能状态的洞察。数据相关采集的质谱学
(DDA)是在代谢蛋白质组学中鉴定和定量微生物蛋白质的最常见的选择方法,
但这项技术在重复性和全面性方面基本上是有限的。蛋白质组学使用
从理论上讲,数据独立采集(DIA)可以解决与DDA相关的基本问题
方法。然而,缺乏生物信息学工具仍然是DIA方面尚未解决的挑战,以及
DIA在微生物群或宿主-微生物相互作用方面的应用还很少有报道。
基于MS的代谢蛋白质组学是一项具有挑战性的测量方法,因为它涉及数千个物种,具有高度的复杂性
丰度大不相同。以获得对微生物功能状态的全面表征
群落不仅需要考虑来自优势微生物的蛋白质,还需要考虑来自低丰度的蛋白质
微生物。这一建议解决了通过
一套使用DIA数据识别和量化多肽及其变体的计算工具的可用性
在微生物菌株水平上。通过新提出的假冒伪肽鉴定方法对假冒伪肽进行控制
多个粒度的发现率评估。蛋白质的推断和定量通过以下方式优化
线性规划模型,包含来自基因组/转录组测序数据和
代谢蛋白质组样本复制品。这一改进将增加已识别的蛋白质变体的数量,
尤其是来自低丰度微生物的那些,可以帮助准确地表征功能
微生物群落的组成,并揭示了功能冗余。
英文摘要
Summary
Extensive efforts to characterize the human microbiome have tremendously increased the knowledge about the
diversity of the microbiome and about its composition in health and in disease. Dysbiosis in human microbiota
underlies the development of many diseases, such as obesity, diabetes, and inflammatory bowel disease.
Metaproteomics based on mass spectrometry (MS) has become widely used in microbiome research for gaining
insights into the functional states of microbial communities. Mass spectrometry with data-dependent acquisition
(DDA) is the most common method of choice for identifying and quantifying microbial proteins in metaproteomics,
but this technique is fundamentally limited in terms of reproducibility and comprehensiveness. Proteomics using
data-independent acquisition (DIA) can, in theory, resolve the fundamental problems associated with the DDA
method. However, the lack of bioinformatics tools still presents unresolved challenges in the context of DIA, and
only few DIA applications on microbiome or host-microbe interactions have been reported.
MS-based metaproteomics is a challenging measurement due to the high complexity with thousands of species
at vastly different abundances. To obtain a comprehensive characterization of the functional state of microbial
communities requires considering proteins not just from dominant microorganisms but also low-abundance
microorganisms. This proposal addresses the need for identifying and quantifying proteins through the
availability of a set of computational tools that use DIA data to identify and quantify peptides and their variants
at the microbial strain level. The false peptide identifications are controlled by newly proposed methods for false
discovery rate assessment at multiple granularities. The protein inference and quantification are optimized by
linear programming models that contain information from genome/transcriptome sequencing data and
metaproteome sample replicas. The improvement will increase the number of identified protein variants,
especially those from the low-abundance microorganisms, which can help accurately characterize the functional
composition in microbial communities and reveal the functional redundancy.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
海外基金