Addressing statistical biases in nucleotide-derived protein databases for proteogenomic search strategies.

Addressing statistical biases in nucleotide-derived protein databases for proteogenomic search strategies.
复制标题

DOI:
10.1021/pr300411q
复制
发表时间:
2012-11-02
影响因子:
4.4
通讯作者:
Hubbard SJ
Hubbard SJ
中科院分区:
生物学2区
文献类型:
--
作者:
Blakeley P;Overton IM;Hubbard SJ

文献摘要

参考文献

被引文献

相似文献

蛋白质组学有可能通过来自质谱学实验的高质量多肽鉴定来推进基因组注释,这些实验证明特定的基因或异构体在蛋白质水平上表达和翻译。这可以促进我们对基因组功能的理解,发现尚未被识别或验证的新基因和基因结构。由于大多数蛋白质组学实验的高通量猎枪性质,仔细控制假阳性并防止任何潜在的错误注释是至关重要的。许多处理这一问题的统计程序在蛋白质组学中被广泛使用,计算群体和单个肽谱匹配的错误发现率(FDR)和后验误差概率(PEP)值。这些方法控制多个测试,并利用诱饵数据库来估计统计意义。在这里,我们表明数据库的选择对这些置信度估计有很大的影响,从而导致报告的PSM数量的显著差异。我们注意到,使用核苷酸序列的六帧翻译的标准靶标:诱饵方法,例如组装的转录组数据,显然低估了分配给PSM的置信度。这一错误的来源源于六帧数据库的夸大和不同寻常的性质,其中每个靶标序列都存在五个不太可能编码蛋白质的“不正确”靶标。随之而来的FDR和PEP估计导致固定阈值下被接受的PSM较少,我们表明这种影响是数据库和统计建模的产物,而不是搜索引擎的结果。根据产生的改变的统计估计和报告的PSM,检查和讨论了限制数据库大小和移除非编码靶序列的各种方法。这些结果对进行蛋白质组学研究的团队具有重要意义,他们的目标是最大限度地验证和发现测序基因组中的基因结构,同时仍然控制假阳性。
Proteogenomics has the potential to advance genome annotation through high quality peptide identifications derived from mass spectrometry experiments, which demonstrate a given gene or isoform is expressed and translated at the protein level. This can advance our understanding of genome function, discovering novel genes and gene structure that have not yet been identified or validated. Because of the high-throughput shotgun nature of most proteomics experiments, it is essential to carefully control for false positives and prevent any potential misannotation. A number of statistical procedures to deal with this are in wide use in proteomics, calculating false discovery rate (FDR) and posterior error probability (PEP) values for groups and individual peptide spectrum matches (PSMs). These methods control for multiple testing and exploit decoy databases to estimate statistical significance. Here, we show that database choice has a major effect on these confidence estimates leading to significant differences in the number of PSMs reported. We note that standard target:decoy approaches using six-frame translations of nucleotide sequences, such as assembled transcriptome data, apparently underestimate the confidence assigned to the PSMs. The source of this error stems from the inflated and unusual nature of the six-frame database, where for every target sequence there exists five “incorrect” targets that are unlikely to code for protein. The attendant FDR and PEP estimates lead to fewer accepted PSMs at fixed thresholds, and we show that this effect is a product of the database and statistical modeling and not the search engine. A variety of approaches to limit database size and remove noncoding target sequences are examined and discussed in terms of the altered statistical estimates generated and PSMs reported. These results are of importance to groups carrying out proteogenomics, aiming to maximize the validation and discovery of gene structure in sequenced genomes, while still controlling for false positives.
DOI: 10.1002/pmic.200900445
发表时间: 2010-03
期刊: PROTEOMICS
影响因子: 3.4
作者:
Blakeley, Paul;Siepen, Jennifer A.;Lawless, Craig;Hubbard, Simon J.
通讯作者: Hubbard, Simon J.
DOI: 10.1021/pr900256v
发表时间: 2010-02-01
影响因子: 4.4
作者:
Everett, Logan J.;Bierl, Charlene;Master, Stephen R.
通讯作者: Master, Stephen R.
DOI: 10.1093/bioinformatics/btp024
发表时间: 2009-03-01
期刊: Bioinformatics (Oxford, England)
影响因子: --
作者:
Gouzy J;Carrere S;Schiex T
通讯作者: Schiex T
DOI: 10.1021/pr200876c
发表时间: 2012-02-01
影响因子: 4.4
作者:
Ching, Ana T. C.;Paes Leme, Adriana F.;Junqueira-de-Azevedo, Inacio L. M.
通讯作者: Junqueira-de-Azevedo, Inacio L. M.
DOI: 10.1093/bioinformatics/bth092
发表时间: 2004-06-12
期刊: BIOINFORMATICS
影响因子: 5.8
作者:
Craig, R;Beavis, RC
通讯作者: Beavis, RC