Ultrafast shape recognition for similarity search in molecular databases

Ultrafast shape recognition for similarity search in molecular databases
复制标题

DOI:
10.1098/rspa.2007.1823
复制
发表时间:
2007-05-08
影响因子:
3.5
通讯作者:
Richards, W. Graham
Richards, W. Graham
中科院分区:
综合性期刊3区
文献类型:
--
作者:
Ballester, Pedro J.;Richards, W. Graham

文献摘要

被引文献

相似文献

分子数据库定期筛选与已知生物活性分子最相似的化合物,以提供新的药物线索。人们普遍认为,三维分子形状是生物活性的最具区分性的模式,因为它与类药物分子与其大分子靶标之间相互作用势的陡峭排斥部分直接相关。然而,分子形状的有效比较目前是一个挑战。在这里,我们展示了一种基于距离分布矩的新方法能够识别分子形状的速度比目前的方法至少快三个数量级。这种超快方法允许在最大的分子数据库中识别相似形状的化合物。此外,由于建议的分布与分子取向无关,因此避免了比较分子排列的有问题的要求。我们的方法也可以用来解决其他领域的类似困难问题,例如为三维几何对象设计基于内容的互联网搜索引擎,或者执行蛋白质之间的快速相似性比较。从更广泛的角度来看,我们预计超快模式识别将很快变得不仅有用,而且对于解决目前大多数科学学科中经历的数据爆炸问题也是必不可少的。
Molecular databases are routinely screened for compounds that most closely resemble a molecule of known biological activity to provide,novel drug leads. It is widely believed that three-dimensional molecular shape is the most discriminating pattern for biological activity as it is directly related to the steep repulsive part of the interaction potential between the drug-like molecule and its macromolecular target. However, efficient comparison of molecular shape is currently a challenge. Here, we show that a new approach based on moments of distance distributions is able to recognize molecular shape at least three orders of magnitude faster than current methodologies. Such an ultrafast method permits the identification of similarly shaped compounds within the largest molecular databases. In addition, the problematic requirement of aligning molecules for comparison is circumvented, as the proposed distributions are independent of molecular orientation. Our methodology could be also adapted to tackle similar hard problems in other fields, such as designing content-based Internet search engines for three-dimensional geometrical objects or performing fast similarity comparisons between proteins. From a broader perspective, we anticipate that ultrafast pattern recognition will soon become not only useful, but also essential to address the data explosion currently experienced in most scientific disciplines.