Refining similarity scoring to enable decoy-free validation in spectral library searching

Refining similarity scoring to enable decoy-free validation in spectral library searching
复制标题

DOI:
10.1002/pmic.201300232
复制
发表时间:
2013-11-01
期刊:
影响因子:
3.4
通讯作者:
Lam, Henry
Lam, Henry
中科院分区:
生物学3区
文献类型:
--
作者:
Shao, Wenguang;Zhu, Kan;Lam, Henry

文献摘要

被引文献

相似文献

光谱库搜索是一种成熟的MS/MS肽鉴定方法,为传统的序列数据库搜索提供了一种替代方法。光谱库搜索依赖于查询数据和光谱库之间的直接光谱到光谱匹配,这提供了对真匹配和假匹配的更好区分,从而提高灵敏度。然而,由于真实的光谱的峰位置和强度轮廓的固有多样性,所得到的相似性分数分布通常呈现不可预测的形状。这使得很难准确地对错误匹配的分数进行建模,需要使用诱饵搜索来对错误匹配的分数分布进行采样。在这里,我们改进了光谱库搜索中的相似性评分,以在不使用诱饵的情况下验证光谱搜索结果。我们对峰强度进行了等级变换以标准化所有光谱,从而可以将参数分布拟合到非最高得分光谱匹配的分数。最高得分匹配的统计显著性然后可以根据极值理论以严格的方式估计。总体结果是光谱匹配质量的更鲁棒和可解释的测量,其可以在没有诱饵的情况下获得。我们在真实的数据集上测试了这种改进的相似性评分函数,并证明了其有效性。这种方法减少了搜索时间,提高了灵敏度,并扩展了光谱库搜索的情况下,诱饵光谱不能很容易地产生,如在搜索身份不明和非肽光谱库。
Spectral library searching is a maturing approach for peptide identification from MS/MS, offering an alternative to traditional sequence database searching. Spectral library searching relies on direct spectrum-to-spectrum matching between the query data and the spectral library, which affords better discrimination of true and false matches, leading to improved sensitivity. However, due to the inherent diversity of the peak location and intensity profiles of real spectra, the resulting similarity score distributions often take on unpredictable shapes. This makes it difficult to model the scores of the false matches accurately, necessitating the use of decoy searching to sample the score distribution of the false matches. Here, we refined the similarity scoring in spectral library searching to enable the validation of spectral search results without the use of decoys. We rank-transformed the peak intensities to standardize all spectra, making it possible to fit a parametric distribution to the scores of the nontop-scoring spectral matches. The statistical significance of the top-scoring match can then be estimated in a rigorous manner according to Extreme Value Theory. The overall result is a more robust and interpretable measure of the quality of the spectral match, which can be obtained without decoys. We tested this refined similarity scoring function on real datasets and demonstrated its effectiveness. This approach reduces search time, increases sensitivity, and extends spectral library searching to situations where decoy spectra cannot be readily generated, such as in searching unidentified and nonpeptide spectral libraries.