A Combination Method of the Tanimoto Coefficient and Proximity Measure of Random Forest for Compound Activity Prediction

A Combination Method of the Tanimoto Coefficient and Proximity Measure of Random Forest for Compound Activity Prediction
复制标题

DOI:
10.2197/ipsjdc.4.238
复制
发表时间:
2008-03
期刊:
Ipsj Digital Courier
影响因子:
--
通讯作者:
G. Kawamura;S. Seno;Y. Takenaka;H. Matsuda
G. Kawamura;S. Seno;Y. Takenaka;H. Matsuda
中科院分区:
其他
文献类型:
--
作者:
G. Kawamura;S. Seno;Y. Takenaka;H. Matsuda

文献摘要

相似文献

化合物的化学和生物活性为发现新药提供了有价值的信息。由活动的结构信息表示的复合指纹被用于考察相似度的候选对象。然而,从化合物结构相似性的要求来看,预测精度存在一些问题。尽管化合物数据的数量正在迅速增长,但经过良好注释的化合物的数量,例如MDL药物数据报告(MDDR)数据库中的化合物,并没有迅速增加。由于已知具有靶标生物类某些活性的化合物在药物发现过程中很少,因此随着活性的降低,预测的准确性应该增加,或者应该在具有大量未注释化合物和少量注释生物活性化合物的数据库中保持假阳性率。本文提出了一种新的相似性评分方法,该方法将TAnimoto系数和随机森林的邻近度相结合。分数包含两个属性,它们来自化合物的无监督和有监督的部分依赖方法。因此,所提出的方法有望显示具有准确活性的化合物。通过与TAnimoto系数和邻近度两种评分方法的预测性能比较,证明了该评分方法的预测结果优于线性判别分析(LDA)方法的预测结果。使用该方法对从MDDR中提取的复合数据集的预测精度进行了评估。实验还表明,该方法可以从包含多个未标注化合物的数据集中识别出活性化合物。
Chemical and biological activities of compounds provide valuable information for discovering new drugs. The compound fingerprint that is represented by structural information of the activities is used for candidates for investigating similarity. However, there are several problems with predicting accuracy from the requirement in the compound structural similarity. Although the amount of compound data is growing rapidly, the number of well-annotated compounds, e.g., those in the MDL Drug Data Report (MDDR)database, has not increased quickly. Since the compounds that are known to have some activities of a biological class of the target are rare in the drug discovery process, the accuracy of the prediction should be increased as the activity decreases or the false positive rate should be maintained in databases that have a large number of un-annotated compounds and a small number of annotated compounds of the biological activity. In this paper, we propose a new similarity scoring method composed of a combination of the Tanimoto coefficient and the proximity measure of random forest. The score contains two properties that are derived from unsupervised and supervised methods of partial dependence for compounds. Thus, the proposed method is expected to indicate compounds that have accurate activities. By evaluating the performance of the prediction compared with the two scores of the Tanimoto coefficient and the proximity measure, we demonstrate that the prediction result of the proposed scoring method is better than those of the two methods by using the Linear Discriminant Analysis (LDA) method. We estimate the prediction accuracy of compound datasets extracted from MDDR using the proposed method. It is also shown that the proposed method can identify active compounds in datasets including several un-annotated compounds.