Bayesian Modeling of Mass-Spec Proteomics Data to Advance Studies of the Genetic Regulation of Proteins
Bayesian Modeling of Mass-Spec Proteomics Data to Advance Studies of the Genetic Regulation of Proteins
批准号:
10391171
负责人:
Gregory R Keele
金额:
$0.25万
依托单位:
依托单位国家:
美国
项目类别:
财政年份:
2021
资助国家:
美国
项目状态:
已结题
起止时间:
2021-09-01 至 2022-08-31
关键词:
AddressAgingAnimal ModelAutomobile DrivingBayesian MethodBayesian ModelingBiochemistryBiologicalBiological ProcessBiologyCellsCodeCommunitiesComplexComputer softwareDataData SetDiseaseDisease ProgressionEnvironmental ExposureError SourcesExperimental DesignsFoundationsGeneticGenetic ResearchGenetic studyGoalsHealthHumanIndividualInterest GroupIntuitionIslet CellIslets of LangerhansKnowledgeLabelMapsMass Spectrum AnalysisMeasurementMeasuresMetabolic PathwayMethodsModelingModernizationModificationMolecular AnalysisMotivationMusOutcomePatternPeptide FragmentsPeptidesPhenotypePlayPopulationPredispositionPreventionProceduresProcessProtein DynamicsProtein IsoformsProteinsProteomeProteomicsProxyRegulationRoleSamplingScientistShotgunsSignal TransductionSourceStatistical AlgorithmStatistical Data InterpretationStatistical MethodsStatistical ModelsStructureSystemSystematic BiasTechnologyTranscriptTransport ProcessUncertaintyVariantWorkanalytical toolbasebiological systemsdata resourcedesigndisease phenotypeexperimental studyextracellularflexibilitygenetic analysisgenetic makeupglucose metabolismheart metabolismhuman diseaseimprovedinsightkidney metabolismmouse modelnovelprogramsprotein complexprotein protein interactionsuccesstooltranscriptome sequencingtranscriptomics
中文摘要
项目摘要/摘要
蛋白质在基本上所有的生物系统中都扮演着重要的功能角色,参与了复杂的
在人类群体中观察到的表型和疾病。所有蛋白质的定量研究,即
蛋白质组学,有可能直接评估蛋白质动力学如何在个体、治疗和
曝光,最好是不偏不倚的方式,不需要预先形成和有针对性的候选人。从历史上看,
由于原始质谱学的局限性,蛋白质组学方法一直受到限制。
可用的技术。转录组学经常被用来代替蛋白质组学,尽管值得注意的是,
蛋白质的调控可以从它们的转录中分离出来,使它们成为不完美的代理人。可行性
准确和可靠的蛋白质组学的发展得益于MS技术的快速发展。目前,
蛋白质组学的统计工具落后,阻碍了这些丰富数据的充分利用
资源。
MS蛋白质组学数据具有许多独特且具有挑战性的特征,需要在其
统计分析。蛋白质不是直接测量的,而是预先分段成更小的多肽。一个
蛋白质的丰度必须从其组成的多肽中重建出来。这一过程的复杂性
包括具有编码变体的多肽(我们的数据集中约10%的多肽)、映射
与至少一个样品中没有观察到的多种蛋白质(~50%)和高水平的多肽有关
(~50%)。MS实验的设计特征,如等压标记的使用,会影响观察到的
丢失数据的模式以及变化的技术来源的程度,促使需要灵活
分析工具。为了做到这一点,我将使用贝叶斯方法对MS蛋白质组学数据进行建模,以灵活地
整合多个误差源,并解决MS实验性测试的这些具有挑战性的功能
程序。由此产生的统计软件将用于来自
遗传多样性的小鼠种群具有与人类种群相似的遗传变异性水平。
有了来自我的软件的改进的蛋白质丰度估计,我将执行遗传分析来
确定蛋白质及其复合体及其相互作用网络丰度的新的遗传调节器。
具体到每个数据集的实验背景,我将把这些监管签名与重要的
生物过程,如肾脏和心脏的衰老以及胰岛细胞的葡萄糖代谢。
该项目将产生新的统计工具,将增加MS蛋白质组学数据的实用性和能力
下游的基因分析,这将在真实的数据中得到证明。新的基因调控机制
将确定潜在的蛋白质动力学和功能网络之间的关系。这些工具和
方法将涉及不同的利益集团,跨越人类,模型生物体系统,以及
各种以疾病为主的社区。
英文摘要
PROJECT SUMMARY / ABSTRACT
Proteins play vital functional roles in essentially all biological systems, factoring into the complex expression of
phenotypes and diseases observed in human populations. The quantitative study of all proteins, i.e.
proteomics, has the potential to directly assess how protein dynamics vary across individuals, treatments, and
exposures, ideally in an unbiased fashion not requiring pre-formed and targeted candidates. Historically a
proteomics approach has been constrained due to limitations of the original mass spectrometry (MS)
technology available. Transcriptomics has often been used in place of proteomics, though notably, the
regulation of proteins can be decoupled from their transcripts, rendering them imperfect proxies. The feasibility
of accurate and reliable proteomics has been aided by rapid advancement in MS technology. Currently the
statistical tools for proteomics lag behind and present an impediment to the full use of these rich data
resources.
MS proteomics data possess a number of unique and challenging features that need to be addressed in their
statistical analysis. Proteins are not directly measured, but instead pre-fragmented into smaller peptides. A
protein's abundance must then be reconstructed from its component peptides. Complications to this process
includes peptides that possess coding variants (~10% of peptides in one of our data sets), peptides that map
to multiple proteins (~50%) and high levels of peptides that are unobserved in at least one of the samples
(~50%). Desing features of the MS experiment, such as the use of isobaric labels, can influence the observed
pattern of missing data as well as the extent of technical sources of variation, motivating the need for flexible
analytical tools. To accomplish this, I will use Bayesian approaches to model MS proteomics data to flexibly
incorporate multiple sources of error, as well as address these challenging features of the MS experimental
procedure. The resulting statistical software will be employed on multiple large proteomics data sets from
genetically diverse mouse populations that possess similar levels of genetic variability as human populations.
With the improved protein abundance estimates from my software, I will then perform genetic analyses to
identify novel genetic regulators of the abundance of proteins, their complexes, and their interaction networks.
Specific the experimental context of each data set, I will connect these regulatory signatures to important
biological processes, such as aging in the kidney and heart and glucose metabolism in pancreatic islet cells.
This project will produce new statistical tools that will increase the utility of MS proteomics data and the power
of downstream genetic analyses, which will be demonstrated in real data. Novel genetic regulatory
relationships underlying protein dynamics and functional networks will be identified. These tools and
approaches will be relevant across diverse interest groups, spanning humans, model organism systems, and
various disease-focused communities.
期刊论文(1)
专著(0)
科研奖励(0)
会议论文
DOI:
10.15252/msb.202110240
发表时间:
2021-08
期刊:
Molecular systems biology
影响因子:
9.9
作者:
[Čuklina J, Lee CH, Williams EG, Sajic T, Collins BC, Rodríguez Martínez M, Sharma VS, Wendt F, Goetze S, Keele GR, Wollscheid B, Aebersold R, Pedrioli PGA]
通讯作者:
Pedrioli PGA
Bayesian Modeling of Mass-Spec Proteomics Data to Advance Studies of the Genetic Regulation of Proteins
-
批准号:10337036
-
项目类别:
-
资助金额:$7.11万
-
财政年份:2020
-
负责人:Gregory R Keele
-
依托单位:
海外基金