A fast coarse filtering method for protein identification by mass spectrometry

A fast coarse filtering method for protein identification by mass spectrometry
复制标题

DOI:
--
复制
发表时间:
--
期刊:
--
影响因子:
--
通讯作者:
Smriti R. Ramakrishnan;Rui Mao;Aleksey A. Nakorchevskiy;J. Prince;Willard S. Willard-Willard-S.-Willard-2668707;Weijia Xu;E. Marcotte;Daniel P. Miranker
Smriti R. Ramakrishnan;Rui Mao;Aleksey A. Nakorchevskiy;J. Prince;Willard S. Willard-Willard-S.-Willard-2668707;Weijia Xu;E. Marcotte;Daniel P. Miranker
中科院分区:
其他
文献类型:
--
作者:
Smriti R. Ramakrishnan;Rui Mao;Aleksey A. Nakorchevskiy;J. Prince;Willard S. Willard-Willard-S.-Willard-2668707;Weijia Xu;E. Marcotte;Daniel P. Miranker

文献摘要

被引文献

相似文献

动机:我们通过将谱映射到文档检索中常用的向量空间模型来重新表述质谱比较问题。因此,文件相似度和文件索引的措施可能适用于蛋白质鉴定。在我们的方法中,使用一种利用度量空间索引算法的快速粗过滤方法来产生初始候选集。然后,我们使用ProFound的贝叶斯评分方案对该约简集中的光谱进行排序。理想情况下,与该领域大多数领先工具提供的线性性能相比,粗过滤器搜索的复杂性接近O(log n)。结果:我们考虑了基于余弦距离和汉明距离的三种距离度量,修改它们以适应质谱固有的峰移,并研究了它们与多优势点指数结构的集成。其中,使用肽质量约束的半度量模糊余弦距离表现最好。我们实现了一个近似的半度量搜索,并证明了这比标准度量空间搜索提高了索引修剪能力。我们测量结果的准确性和指数性能测试集的肽片段谱从大肠杆菌蛋白。我们还报告了在1000个血管紧张素- ii串联质谱的更全面基准上的灵敏度(召回率)和特异性(精度)得分,表明在实践中,在这个高维稀疏空间中,当伴随大量增加时,近似搜索是可以接受的
Motivation: We reformulate the problem of comparing massspectra by mapping spectra to the vector space model commonly used in document retrieval. It follows that measures of document similarity and document indexing may be adapted for protein identification. In our approach a fast coarse filtering method leveraging a metric space indexing algorithm is used to produce an initial candidate set. We then rank the spectra in this reduced set using ProFound’s Bayesian scoring scheme. Ideally,the complexity of the coarse filter search approaches O(log n), as compared to the linear performance provided by most leading tools in the field. Results: We consider three distance measures based on cosine and hamming distances, modifying them to accommodate the peak shifts intrinsic to mass spectra and investigate their integration with the multivantage-point index structure. Of these, a semi-metric, fuzzy-cosine distance using peptide mass constraints performs the best. We implement an approximate semi-metric search, and show that this improves index pruning power over a standard metric space search. We measure accuracy of results and index performance on a test set of peptide fragmentation spectra from E.coli proteins. We also report sensitivity(recall) and specificity(precision) scores on a more comprehensive benchmark of 1000 Angiotensin-II tandem mass spectra, showing that, in practice, approximate searches in this high dimensional sparse space are acceptable when accompanied by substantial increase in