A fast coarse filtering method for protein identification by mass spectrometry
A fast coarse filtering method for protein identification by mass spectrometry
复制标题
DOI:
--
复制
发表时间:
--
期刊:
影响因子:
--
通讯作者:
Smriti R. Ramakrishnan;Rui Mao;Aleksey A. Nakorchevskiy;J. Prince;Willard S. Willard-Willard-S.-Willard-2668707;Weijia Xu;E. Marcotte;Daniel P. Miranker
中科院分区:
文献类型:
--
作者:
Smriti R. Ramakrishnan;Rui Mao;Aleksey A. Nakorchevskiy;J. Prince;Willard S. Willard-Willard-S.-Willard-2668707;Weijia Xu;E. Marcotte;Daniel P. Miranker
Motivation: We reformulate the problem of comparing massspectra by mapping spectra to the vector space model commonly used in document retrieval. It follows that measures of document similarity and document indexing may be adapted for protein identification. In our approach a fast coarse filtering method leveraging a metric space indexing algorithm is used to produce an initial candidate set. We then rank the spectra in this reduced set using ProFound’s Bayesian scoring scheme. Ideally,the complexity of the coarse filter search approaches O(log n), as compared to the linear performance provided by most leading tools in the field. Results: We consider three distance measures based on cosine and hamming distances, modifying them to accommodate the peak shifts intrinsic to mass spectra and investigate their integration with the multivantage-point index structure. Of these, a semi-metric, fuzzy-cosine distance using peptide mass constraints performs the best. We implement an approximate semi-metric search, and show that this improves index pruning power over a standard metric space search. We measure accuracy of results and index performance on a test set of peptide fragmentation spectra from E.coli proteins. We also report sensitivity(recall) and specificity(precision) scores on a more comprehensive benchmark of 1000 Angiotensin-II tandem mass spectra, showing that, in practice, approximate searches in this high dimensional sparse space are acceptable when accompanied by substantial increase in