OPTIMIZATION AND TESTING OF MASS-SPECTRAL LIBRARY SEARCH ALGORITHMS FOR COMPOUND IDENTIFICATION

OPTIMIZATION AND TESTING OF MASS-SPECTRAL LIBRARY SEARCH ALGORITHMS FOR COMPOUND IDENTIFICATION
复制标题

DOI:
10.1016/1044-0305(94)87009-8
复制
发表时间:
1994-09-01
影响因子:
3.2
通讯作者:
SCOTT, DR
SCOTT, DR
中科院分区:
化学3区
文献类型:
--
作者:
STEIN, SE;SCOTT, DR

文献摘要

被引文献

相似文献

通过将测试光谱与NIST-EPA-NIH质谱库中的参比光谱进行匹配,对文献中提出的五种从低分辨率质谱库中识别未知化合物的算法进行了优化和测试。这些算法是基于概率的匹配(PBM)、点积、Hertz等。相似性指数、欧几里得距离和绝对值距离。测试集由数据库中代表的约8000种化合物的12,592个交替光谱组成。大多数算法通过改变它们的质量权重和强度比例因子进行了优化。以候选化合物列表中的排名作为准确度的标准。性能最好的算法(对于等级1的准确率为75%)是点积函数,它测量以向量表示的光谱之间的夹角的余弦。按执行顺序排列的其他方法有欧几里德距离(72%)、绝对值距离(68%)、PBM(65%)和Hertz等。(%)。在优化算法中,强度标度和质量加权是重要的,强度标度的平方根接近最优,平方或立方体的质量加权能力最好。还测试了几个更复杂的方案,但对结果几乎没有影响。对点积算法的性能进行了适度的改进,增加了一个项,赋予具有许多共同峰的光谱的相对峰强度额外的权重。
Five algorithms proposed in the literature for library search identification of unknown compounds from their low resolution mass spectra were optimized and tested by matching test spectra against reference spectra in the NIST-EPA-NIH Mass Spectral Database. The algorithms were probability-based matching (PBM), dot-product, Hertz et al. similarity index, Euclidean distance, and absolute value distance. The test set consisted of 12,592 alternate spectra of about 8000 compounds represented in the database. Most algorithms were optimized by varying their mass weighting and intensity scaling factors. Rank in the list of candidate compounds was used as the criterion for accuracy. The best performing algorithm (75% accuracy for rank 1) was the dot-product function that measures the cosine of the angle between spectra represented as vectors. Other methods in order of performance were the Euclidean distance (72%), absolute value distance (68%), PBM (65%), and Hertz et al. (64%). Intensity scaling and mass weighting were important in the optimized algorithms with the square root of the intensity scale nearly optimal and the square or cube the best mass weighting power. Several more complex schemes also were tested, but had little effect on the results. A modest improvement in the performance of the dot-product algorithm was made by adding a term that gave additional weight to relative peak intensities for spectra with many peaks in common.