Using Inverted Indices for Accelerating LINGO Calculations

Using Inverted Indices for Accelerating LINGO Calculations
复制标题

使用倒排索引加速 LINGO 计算

DOI:
--
复制
发表时间:
2011
影响因子:
5.6
通讯作者:
Christian N. S. Pedersen
Christian N. S. Pedersen
中科院分区:
化学2区
文献类型:
--
作者:
T. Kristensen;J. Nielsen;Christian N. S. Pedersen

文献摘要

被引文献

相似文献

不断增长的化学数据库的规模要求开发新的方法来表示和比较分子。其中一种称为LINGO的方法是基于对分子的SMILES字符串表示进行分段。然后可以通过计算Tanimoto系数来进行分子的比较,当在LINGO多集上使用时称为LINGOsim。本文介绍了一种用于存储LINGO多集的详细表示,这使得有可能将它们转换为稀疏指纹,以便指纹数据结构和算法可以用于加速查询。以前快速计算LINGOsim相似性矩阵的最佳方法需要专门的硬件来产生比现有方法显著的加速。通过在详细表示中表示LINGO多集并使用倒排索引,可以计算LINGOsim相似性矩阵,其速度比现有方法快约2.6倍,而无需依赖于专门的硬件。
The ever growing size of chemical databases calls for the development of novel methods for representing and comparing molecules. One such method called LINGO is based on fragmenting the SMILES string representation of molecules. Comparison of molecules can then be performed by calculating the Tanimoto coefficient, which is called LINGOsim when used on LINGO multisets. This paper introduces a verbose representation for storing LINGO multisets, which makes it possible to transform them into sparse fingerprints such that fingerprint data structures and algorithms can be used to accelerate queries. The previous best method for rapidly calculating the LINGOsim similarity matrix required specialized hardware to yield a significant speedup over existing methods. By representing LINGO multisets in the verbose representation and using inverted indices, it is possible to calculate LINGOsim similarity matrices roughly 2.6 times faster than existing methods without relying on specialized hardware.