AMBIT-SMARTS: Efficient Searching of Chemical Structures and Fragments

AMBIT-SMARTS: Efficient Searching of Chemical Structures and Fragments
复制标题

DOI:
10.1002/minf.201100028
复制
发表时间:
2011-08-01
影响因子:
3.6
通讯作者:
Kochev, Nikolay
Kochev, Nikolay
中科院分区:
医学4区
文献类型:
--
作者:
Jeliazkova, Nina;Kochev, Nikolay

文献摘要

被引文献

相似文献

我们介绍了用于有效搜索化学结构和结构片段的AMBIT开源软件包的新进展。Ambit-Smarts是一个基于Java的软件,建立在化学开发工具包的基础上。Ambit-Smarts解析器使用多个语法扩展实现了整个Smarts语言规范,支持由OpenEye、Moe和OpenBabel等第三方软件包引入的自定义修改。另一个开源Smarts解析器实现的目标是实现更好的性能和与多种现有风格的Smarts语言的兼容性,以及提供在大型结构化数据库中运行高效Smarts查询的实用程序。我们描述了降低计算成本和改善子结构查询的响应时间的方法的组合。对该算法和几种子图同构实现进行了详尽的比较。为了从终端用户的角度展示整个系统的性能,还报告了针对具有4.5M个结构的数据库的Web服务子结构搜索查询的响应时间统计。该程序包在各种化学信息学任务的实施中具有广泛的适用性。它已经在几个涉及描述符计算和预测算法、数据库查询、Web应用程序和Web服务的项目中使用。
We present new developments in the AMBIT open source software package for efficient searching of chemical structures and structural fragments. AMBIT-SMARTS is a Java based software built on top of The Chemistry Development Kit. The AMBIT-SMARTS parser implements the entire SMARTS language specification with several syntax extensions that enable support for custom modifications introduced by third party software packages such as OpenEye, MOE and OpenBabel. The goal of yet another open-source SMARTS parser implementation is to achieve better performance and compatibility with multiple existing flavours of the SMARTS language, as well as to provide utilities for running efficient SMARTS queries in large structural databases. We describe a combination of approaches towards lowering the computational cost and improving the response time of substructure queries. An exhaustive comparison of the AMBIT algorithm with several subgraph isomorphism implementations is performed. To demonstrate the performance of the entire system from an end-user point of view, response time statistics for Web service substructure search queries against a database of 4.5 M structures are also reported. The package has wide applicability in the implementation of various chemoinformatics tasks. It has already been used in several projects dealing with descriptor calculation and predictive algorithms, database queries, web applications and web services.