Utilizing Low-Dimensional Molecular Embeddings for Rapid Chemical Similarity Search.

Utilizing Low-Dimensional Molecular Embeddings for Rapid Chemical Similarity Search.
复制标题

利用低维分子嵌入进行快速化学相似性搜索。

DOI:
10.1007/978-3-031-56060-6_3
复制
发表时间:
2024
期刊:
Advances in information retrieval : ... European Conference on IR Research, ECIR ... proceedings. European Conference on IR Research
影响因子:
--
通讯作者:
Tropsha,Alexander
Tropsha,Alexander
中科院分区:
--
文献类型:
--
作者:
Kirchoff,KathrynE;Wellnitz,James;Hochuli,JoshuaE;Maxfield,Travis;Popov,KonstantinI;Gomez,Shawn;Tropsha,Alexander

文献摘要

相似文献

基于最近邻的相似性搜索是化学中的一项常见任务,在药物发现中具有显著的用例。然而,这项任务中最常用的一些方法仍然利用蛮力方法。在实践中,这可能是计算成本高,过于耗时,部分原因是现代化学数据库的庞大规模。以前的这项任务的计算进步通常依赖于对硬件的改进或缺乏普遍性的特定于网络的技巧。利用低复杂度搜索算法的方法仍然相对未被探索。然而,这些算法中的许多是近似解和/或与典型的高维化学嵌入斗争。在这里,我们评估低维化学嵌入和ak-d树数据结构的组合是否可以实现快速最近邻查询,同时保持标准化学相似性搜索基准的性能。我们研究不同的降维标准化学嵌入以及学习,结构感知嵌入-SmallSA-这项任务。有了这个框架,在单个CPU核心上搜索超过10亿种化学物质的时间不到一秒,比蛮力方法快了五个数量级。我们还证明了SmallSA在化学相似性基准上具有竞争力的性能。
Nearest neighbor-based similarity searching is a common task in chemistry, with notable use cases in drug discovery. Yet, some of the most commonly used approaches for this task still leverage a brute-force approach. In practice this can be computationally costly and overly time-consuming, due in part to the sheer size of modern chemical databases. Previous computational advancements for this task have generally relied on improvements to hardware or dataset-specific tricks that lack generalizability. Approaches that leverage lower-complexity searching algorithms remain relatively underexplored. However, many of these algorithms are approximate solutions and/or struggle with typical high-dimensional chemical embeddings. Here we evaluate whether a combination of low-dimensional chemical embeddings and ak-d tree data structure can achieve fast nearest neighbor queries while maintaining performance on standard chemical similarity search benchmarks. We examine different dimensionality reductions of standard chemical embeddings as well as a learned, structurally-aware embedding—SmallSA—for this task. With this framework, searches on over one billion chemicals execute in less than a second on a single CPU core, five orders of magnitude faster than the brute-force approach. We also demonstrate that SmallSA achieves competitive performance on chemical similarity benchmarks.