Ligand Prediction for Orphan Targets Using Support Vector Machines and Various Target-Ligand Kernels Is Dominated by Nearest Neighbor Effects

Ligand Prediction for Orphan Targets Using Support Vector Machines and Various Target-Ligand Kernels Is Dominated by Nearest Neighbor Effects
复制标题

DOI:
10.1021/ci9002624
复制
发表时间:
2009-10-01
影响因子:
5.6
通讯作者:
Bajorath, Juergen
Bajorath, Juergen
中科院分区:
化学2区
文献类型:
--
作者:
Wassermann, Anne Mai;Geppert, Hanna;Bajorath, Juergen

文献摘要

被引文献

相似文献

结合蛋白质和小分子信息的支持向量机(SVM)计算已被应用于识别模拟孤儿靶标(即没有配体可用的靶标)的配体。通过目标配体核函数的设计促进了蛋白质和配体信息的结合,该核函数考虑了配体和靶标的成对相似性。这些核函数的设计和生物信息含量有望在靶向配体预测中发挥重要作用。因此,实现了多种靶配体核来捕获不同类型的靶标信息,包括序列、二级结构、三级结构、生物物理性质、本体或结构分类。这些核在两个靶蛋白系统中模拟孤儿靶标的配体预测中进行了测试,其特征是存在不同的靶标间关系。令人惊讶的是,尽管不同目标配体核的预测率存在目标特异性和集合特异性差异,但这些核的性能总体上是相似的,也与SVM线性组合相似。为了更好地理解这些观察结果的可能原因而设计的测试计算表明,孤儿目标的最近邻居提供的配体信息显著影响SVM的性能,比包含蛋白质信息的影响要大得多。无论核函数的类型和复杂度如何,只要孤儿目标的近邻配体可以用于SVM学习,就可以很好地预测孤儿目标配体。这些发现为基于svm的孤儿靶点配体预测提供了简化策略。
Support vector machine (SVM) calculations combining protein and small molecule information have been applied to identify ligands for simulated orphan targets (i.e., targets for which no ligands were available). The combination of protein and ligand information was facilitated through the design of target-ligand kernel functions that account for pairwise ligand and target similarity. The design and biological information content of such kernel functions was expected to play a major role for target-directed ligand prediction. Therefore, a variety of target-ligand kernels were implemented to capture different types of target information including sequence, secondary structure, tertiary structure, biophysical properties, ontologies, or structural taxonomy. These kernels were tested in ligand predictions for simulated orphan targets in two target protein systems characterized by the presence of different inter-target relationships. Surprisingly, although there were target- and set-specific differences in prediction rates for alternative target-ligand kernels, the performance of these kernels was overall similar and also similar to SVM linear combinations. Test calculations designed to better understand possible reasons for these observations revealed that ligand information provided by nearest neighbors of orphan targets significantly influenced SVM performance, much more so than the inclusion of protein information. As long as ligands of closely related neighbors of orphan targets were available for SVM learning, orphan target ligands could be well predicted, regardless of the type and sophistication of the kernel function that was used. These findings suggest simplified strategies for SVM-based ligand prediction for orphan targets.