Efficient calculation of compound similarity based on maximum common subgraphs and its application to prediction of gene transcript levels.

Efficient calculation of compound similarity based on maximum common subgraphs and its application to prediction of gene transcript levels.
复制标题

DOI:
10.1504/ijbra.2013.054688
复制
发表时间:
2013-01-01
影响因子:
--
通讯作者:
Ridder, Dick De
Ridder, Dick De
中科院分区:
其他
文献类型:
--
作者:
Berlo, Rogier J P Van;Winterbach, Wynand;Ridder, Dick De

文献摘要

被引文献

相似文献

化学实体的性质,包括物理和生物性质,都与其结构有关。由于化合物相似度可以用来推断新化合物的性质,因此在化学信息学中,计算结构相似度的方法受到了极大的关注。衡量化合物之间结构相似性的一个有用的度量是最大公共子图(MCS)的相对大小。当用图形表示时,MCS是一对化合物中存在的最大的子结构。然而,在实践中很难使用这样的度量,因为当MCS很大时,计算MCS变得难以计算。我们提出了一种新的算法,与一些最先进的方法相比,它显著减少了寻找大型MCS的计算时间。该算法的使用在预测乳腺癌细胞系对不同类药物化合物的转录反应的应用程序中得到了演示,其规模对于迄今最有效的MCS算法来说是具有挑战性的。在这一应用中,对714种化合物进行了比较。
Properties of a chemical entity, both physical and biological, are related to its structure. Since compound similarity can be used to infer properties of novel compounds, in chemoinformatics much attention has been paid to ways of calculating structural similarity. A useful metric to capture the structural similarity between compounds is the relative size of the Maximum Common Subgraph (MCS). The MCS is the largest substructure present in a pair of compounds, when represented as graphs. However, in practice it is difficult to employ such a metric, since calculation of the MCS becomes computationally intractable when it is large. We propose a novel algorithm that significantly reduces computation time for finding large MCSs, compared to a number of state-of-the-art approaches. The use of this algorithm is demonstrated in an application predicting the transcriptional response of breast cancer cell lines to different drug-like compounds, at a scale which is challenging for the most efficient MCS-algorithms to date. In this application 714 compounds were compared.