Bayesian Modeling of Mass-Spec Proteomics Data to Advance Studies of the Genetic Regulation of Proteins
Bayesian Modeling of Mass-Spec Proteomics Data to Advance Studies of the Genetic Regulation of Proteins
批准号:
10337036
负责人:
Gregory R Keele
金额:
$7.11万
依托单位:
依托单位国家:
美国
项目类别:
财政年份:
2020
资助国家:
美国
项目状态:
已结题
起止时间:
2020-09-01 至 2022-08-31
关键词:
AddressAgingAnimal ModelAutomobile DrivingBayesian MethodBayesian ModelingBiochemistryBiologicalBiological ProcessBiologyCellsCodeCommunitiesComplexComputer softwareDataData SetDiseaseDisease ProgressionEnvironmental ExposureError SourcesExperimental DesignsFoundationsGeneticGenetic ResearchGenetic studyGoalsHealthHumanIndividualInterest GroupIntuitionIslet CellIslets of LangerhansKnowledgeLabelMapsMass Spectrum AnalysisMeasurementMeasuresMetabolic PathwayMethodsModelingModernizationModificationMolecular AnalysisMotivationMusOutcomePatternPeptide FragmentsPeptidesPhenotypePlayPopulationPredispositionPreventionProceduresProcessProtein DynamicsProtein IsoformsProteinsProteomeProteomicsProxyRegulationRoleSamplingScientistShotgunsSignal TransductionSourceStatistical AlgorithmStatistical Data InterpretationStatistical MethodsStatistical ModelsStructureSystemSystematic BiasTechnologyTranscriptTransport ProcessUncertaintyVariantWorkanalytical toolbasebiological systemsdata resourcedesigndisease phenotypeexperimental studyextracellularflexibilitygenetic analysisgenetic makeupglucose metabolismheart metabolismhuman diseaseimprovedinsightkidney metabolismmouse modelnovelprogramsprotein complexprotein protein interactionsuccesstooltranscriptome sequencingtranscriptomics
中文摘要
项目概要/摘要
蛋白质在基本上所有的生物系统中发挥着重要的功能作用,包括复杂的
在人群中观察到的表型和疾病。所有蛋白质的定量研究,即
蛋白质组学,有可能直接评估蛋白质动力学如何在个体,治疗和
曝光,理想情况下以不需要预先形成和有针对性的候选人的公正的方式。历史上是一个
蛋白质组学方法由于原始质谱(MS)的局限性而受到限制
技术可用。转录组学经常被用来代替蛋白质组学,但值得注意的是,
蛋白质的调节可以与它们的转录物分离,使它们成为不完美的代理。可行性
准确可靠的蛋白质组学的发展得益于质谱技术的快速发展。目前
蛋白质组学的统计工具落后,阻碍了这些丰富数据的充分利用
资源
MS蛋白质组学数据具有许多独特和具有挑战性的特征,需要在其研究中加以解决。
统计分析蛋白质不直接测量,而是预先片段化成较小的肽。一
蛋白质的丰度必须由其组成肽来重建。这个过程的复杂性
包括具有编码变体的肽(在我们的一个数据集中约10%的肽),
多种蛋白质(~50%)和至少一种样本中未观察到的高水平肽
(~50%)。MS实验的设计特征,例如同量异位素标记的使用,可以影响观察到的结果。
缺失数据的模式以及技术来源的变化程度,促使需要灵活
分析工具。为了实现这一点,我将使用贝叶斯方法来建模MS蛋白质组学数据,
结合多种误差来源,并解决MS实验的这些挑战性特征
procedure.由此产生的统计软件将被用于多个大型蛋白质组数据集,
遗传多样性的小鼠群体具有与人类群体相似的遗传变异性水平。
通过我的软件改进的蛋白质丰度估计,然后我将进行遗传分析,
确定蛋白质丰度的新遗传调节因子,它们的复合物,以及它们的相互作用网络。
具体到每个数据集的实验背景,我将把这些监管签名与重要的
生物过程,如肾脏和心脏的衰老以及胰岛细胞的葡萄糖代谢。
该项目将产生新的统计工具,这将增加MS蛋白质组学数据的实用性,
下游基因分析的结果,这将在真实的数据中得到证明。新型基因调控
蛋白质动力学和功能网络的基础关系将被确定。这些工具和
方法将与不同的利益集团相关,包括人类、模式生物系统和
各种以疾病为重点的社区。
英文摘要
PROJECT SUMMARY / ABSTRACT
Proteins play vital functional roles in essentially all biological systems, factoring into the complex expression of
phenotypes and diseases observed in human populations. The quantitative study of all proteins, i.e.
proteomics, has the potential to directly assess how protein dynamics vary across individuals, treatments, and
exposures, ideally in an unbiased fashion not requiring pre-formed and targeted candidates. Historically a
proteomics approach has been constrained due to limitations of the original mass spectrometry (MS)
technology available. Transcriptomics has often been used in place of proteomics, though notably, the
regulation of proteins can be decoupled from their transcripts, rendering them imperfect proxies. The feasibility
of accurate and reliable proteomics has been aided by rapid advancement in MS technology. Currently the
statistical tools for proteomics lag behind and present an impediment to the full use of these rich data
resources.
MS proteomics data possess a number of unique and challenging features that need to be addressed in their
statistical analysis. Proteins are not directly measured, but instead pre-fragmented into smaller peptides. A
protein's abundance must then be reconstructed from its component peptides. Complications to this process
includes peptides that possess coding variants (~10% of peptides in one of our data sets), peptides that map
to multiple proteins (~50%) and high levels of peptides that are unobserved in at least one of the samples
(~50%). Desing features of the MS experiment, such as the use of isobaric labels, can influence the observed
pattern of missing data as well as the extent of technical sources of variation, motivating the need for flexible
analytical tools. To accomplish this, I will use Bayesian approaches to model MS proteomics data to flexibly
incorporate multiple sources of error, as well as address these challenging features of the MS experimental
procedure. The resulting statistical software will be employed on multiple large proteomics data sets from
genetically diverse mouse populations that possess similar levels of genetic variability as human populations.
With the improved protein abundance estimates from my software, I will then perform genetic analyses to
identify novel genetic regulators of the abundance of proteins, their complexes, and their interaction networks.
Specific the experimental context of each data set, I will connect these regulatory signatures to important
biological processes, such as aging in the kidney and heart and glucose metabolism in pancreatic islet cells.
This project will produce new statistical tools that will increase the utility of MS proteomics data and the power
of downstream genetic analyses, which will be demonstrated in real data. Novel genetic regulatory
relationships underlying protein dynamics and functional networks will be identified. These tools and
approaches will be relevant across diverse interest groups, spanning humans, model organism systems, and
various disease-focused communities.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
Bayesian Modeling of Mass-Spec Proteomics Data to Advance Studies of the Genetic Regulation of Proteins
-
批准号:10391171
-
项目类别:
-
资助金额:$0.25万
-
财政年份:2021
-
负责人:Gregory R Keele
-
依托单位:
海外基金