TagRecon: high-throughput mutation identification through sequence tagging.

TagRecon: high-throughput mutation identification through sequence tagging.
复制标题

DOI:
10.1021/pr900850m
复制
发表时间:
2010-04-05
影响因子:
4.4
通讯作者:
Tabb, David L.
Tabb, David L.
中科院分区:
生物学2区
文献类型:
--
作者:
Dasari, Surendra;Chambers, Matthew C.;Slebos, Robbert J.;Zimmerman, Lisa J.;Ham, Amy-Joan L.;Tabb, David L.

文献摘要

参考文献

被引文献

相似文献

鸟枪蛋白质组学产生串联质谱的集合,其包含从临床样品中鉴定突变肽所需的所有数据。然而,识别这些序列变异用传统的数据库搜索策略是不可行的,这需要观察到的序列和预期的序列之间的精确匹配。通过数据库搜索在特定残基上搜索作为质量偏移的突变可能会导致显著的性能损失并产生大量的假阳性率。在这里,我们描述了TagRecon,一种利用推断的序列标签来识别临床蛋白质组数据集中意外突变的算法。TagRecon识别未修饰肽的灵敏度与相关MyriMatch数据库搜索引擎相同。在LTQ和Orbitrap数据集中,TagRecon在识别具有已知变体的数据集的序列错配方面优于最先进的软件。我们制定了从临床样本中筛选推定突变的指南,并将其应用于癌细胞系分析和结肠组织检查。在高达6%的鉴定肽中发现突变,并且只有一小部分对应于dbSNP条目。DNA错配修复缺陷的RKO细胞系比错配修复熟练的SW480细胞系产生更多的突变肽。结肠癌肿瘤和邻近组织的分析揭示了与细胞外基质降解相关的羟脯氨酸修饰。这些结果证明了使用序列标记算法来充分询问临床蛋白质组数据集的价值。
Shotgun proteomics produces collections of tandem mass spectra that contain all the data needed to identify mutated peptides from clinical samples. Identifying these sequence variations, however, has not been feasible with conventional database search strategies, which require exact matches between observed and expected sequences. Searching for mutations as mass shifts on specified residues through database search can incur significant performance penalties and generate substantial false positive rates. Here we describe TagRecon, an algorithm that leverages inferred sequence tags to identify unanticipated mutations in clinical proteomic data sets. TagRecon identifies unmodified peptides as sensitively as the related MyriMatch database search engine. In both LTQ and Orbitrap data sets, TagRecon outperformed state of the art software in recognizing sequence mismatches from data sets with known variants. We developed guidelines for filtering putative mutations from clinical samples, and we applied them in an analysis of cancer cell lines and an examination of colon tissue. Mutations were found in up to 6% of identified peptides, and only a small fraction corresponded to dbSNP entries. The RKO cell line, which is DNA mismatch repair deficient, yielded more mutant peptides than the mismatch repair proficient SW480 line. Analysis of colon cancer tumor and adjacent tissue revealed hydroxyproline modifications associated with extracellular matrix degradation. These results demonstrate the value of using sequence tagging algorithms to fully interrogate clinical proteomic data sets.
DOI: 10.1021/pr0700908
发表时间: 2007-01-01
影响因子: 4.4
作者:
Bunger, Maureen K.;Cargile, Benjamin J.;Stephenson, James L., Jr.
通讯作者: Stephenson, James L., Jr.
DOI: 10.1093/bioinformatics/btl226
发表时间: 2006-07-01
期刊: BIOINFORMATICS
影响因子: 5.8
作者:
Liu, Chunmei;Yan, Bo;Cai, Liming
通讯作者: Cai, Liming
DOI: 10.1158/0008-5472.can-08-3543
发表时间: 2009-02-01
期刊: Cancer research
影响因子: 11.2
作者:
Bacolod MD;Schemmann GS;Giardina SF;Paty P;Notterman DA;Barany F
通讯作者: Barany F
DOI: 10.1073/pnas.89.22.10915
发表时间: 1992-11-15
影响因子: 11.1
作者:
HENIKOFF, S;HENIKOFF, JG
通讯作者: HENIKOFF, JG
DOI: 10.1021/pr049781j
发表时间: 2005-03-01
影响因子: 4.4
作者:
Searle, BC;Dasari, S;Nagalla, SR
通讯作者: Nagalla, SR