Support vector machines in HTS data mining: Type I MetAPs inhibition study

Support vector machines in HTS data mining: Type I MetAPs inhibition study
复制标题

DOI:
10.1177/1087057105284334
复制
发表时间:
2006-03-01
影响因子:
--
通讯作者:
Georg, GI
Georg, GI
中科院分区:
化学3区
文献类型:
--
作者:
Fang, JW;Dong, YH;Georg, GI

文献摘要

被引文献

相似文献

本文报道了支持向量机在高通量筛选(HTS)数据挖掘中的成功应用。以含有43,736个有机小分子的文库为研究对象,其中1355个具有40%或更高抑制率的化合物被认为具有活性。数据集被随机分为训练集和测试集(比例为3:1)。作者能够使用建立在训练集上的支持向量机模型预测的决策值对测试集中的化合物进行排名。他们定义了一个新的分数PT50,即需要筛选的测试集的百分比,以回收50%的活性物质,以衡量模型的性能。通过精心选择参数,支持向量机模型显著提高了命中率,只需对7%的测试集进行筛选,就可以回收50%的活性化合物。作者发现,训练集的大小对模型的性能起着重要的作用。一个包含10,000个成员化合物的训练集很可能是建立具有合理预测能力的模型所需的最小规模。
This article reports a successful application of support vector machines (SVMs) in mining high-throughput screening (HTS) data of a type I methionine aminopeptidases (MetAPs) inhibition study. A library with 43,736 small organic molecules was used in the study, and 1355 compounds in the library with 40% or higher inhibition activity were considered as active. The data set was randomly split into a training set and a test set (3:1 ratio). The authors were able to rank compounds in the test set using their decision values predicted by SVM models that were built on the training set. They defined a novel score PT50, the percentage of the test set needed to be screened to recover 50% of the actives, to measure the performance of the models. With carefully selected parameters, SVM models increased the hit rates significantly, and 50% of the active compounds could be recovered by screening just 7% of the test set. The authors found that the size of the training set played a significant role in the performance of the models. A training set with 10,000 member compounds is likely the minimum size required to build a model with reasonable predictive power.