Protein and gene model inference based on statistical modeling in k-partite graphs

Protein and gene model inference based on statistical modeling in k-partite graphs
复制标题

DOI:
10.1073/pnas.0907654107
复制
发表时间:
2010-07-06
影响因子:
11.1
通讯作者:
Buehlmann, Peter
Buehlmann, Peter
中科院分区:
综合性期刊1区
文献类型:
--
作者:
Gerster, Sarah;Qeli, Ermir;Buehlmann, Peter

文献摘要

被引文献

相似文献

蛋白质组学的主要目标之一是全面准确地描述蛋白质组。鸟枪蛋白质组学是分析复杂蛋白质混合物的首选方法,它要求将实验观察到的肽映射回它们所来源的蛋白质。这个过程也被称为蛋白质推理。我们提出了马尔可夫推理的蛋白质和基因模型(MIPGEM),统计模型的基础上明确规定的假设,以解决问题的蛋白质和基因模型的推理鸟枪蛋白质组学数据。特别是,我们正在处理之间的依赖性肽和蛋白质使用马尔可夫假设的k-部图。我们还通过对编码基因模型进行评分来解决共享肽和模糊蛋白质的问题。两个控制数据集与合成的混合物的蛋白质和复杂的蛋白质样品的酿酒酵母,黑腹果蝇,拟南芥的经验结果表明,与MIPGEM的结果与现有的工具蛋白质推理的竞争力。
One of the major goals of proteomics is the comprehensive and accurate description of a proteome. Shotgun proteomics, the method of choice for the analysis of complex protein mixtures, requires that experimentally observed peptides are mapped back to the proteins they were derived from. This process is also known as protein inference. We present Markovian Inference of Proteins and Gene Models (MIPGEM), a statistical model based on clearly stated assumptions to address the problem of protein and gene model inference for shotgun proteomics data. In particular, we are dealing with dependencies among peptides and proteins using a Markovian assumption on k-partite graphs. We are also addressing the problems of shared peptides and ambiguous proteins by scoring the encoding gene models. Empirical results on two control data-sets with synthetic mixtures of proteins and on complex protein samples of Saccharomyces cerevisiae, Drosophila melanogaster, and Arabidopsis thaliana suggest that the results with MIPGEM are competitive with existing tools for protein inference.