Enrichment of high-throughput screening data with increasing levels of noise using support vector machines, recursive partitioning, and Laplacian-modified naive Bayesian classifiers

Enrichment of high-throughput screening data with increasing levels of noise using support vector machines, recursive partitioning, and Laplacian-modified naive Bayesian classifiers
复制标题

DOI:
10.1021/ci050374h
复制
发表时间:
2006-01-01
影响因子:
5.6
通讯作者:
Davies, JW
Davies, JW
中科院分区:
化学2区
文献类型:
--
作者:
Glick, M;Jenkins, JL;Davies, JW

文献摘要

被引文献

相似文献

高通量筛选(HTS)在制药行业的铅发现中起着关键作用。在串联,化学信息学方法被用来增加的概率,通过挖掘HTS数据的新的生物活性化合物的鉴定。HTS数据是出了名的嘈杂,因此,选择最佳的数据挖掘方法是很重要的成功,这样的分析。在这里,我们描述了一个回顾性分析的四个HTS数据集,使用三种挖掘方法:拉普拉斯修正的朴素贝叶斯,递归分区,支持向量机(SVM)分类器,增加随机噪声的形式,假阳性和假阴性。所有这三种数据挖掘方法都容忍了越来越多的假阳性,即使在训练集中错误分类的化合物与真正的活性化合物的比例为5:1。1:1比例的假阴性也是可以容忍的。SVM在捕获前1%的活性化合物和支架方面优于其他两种方法。Murcko支架分析可以解释四个数据集之间富集的差异。这项研究表明,数据挖掘方法可以添加一个真正的价值,即使在屏幕上的数据是污染了高水平的随机噪声。
High-throughput screening (HTS) plays a pivotal role in lead discovery for the pharmaceutical industry. In tandem, cheminformatics approaches are employed to increase the probability of the identification of novel biologically active compounds by mining the HTS data. HTS data is notoriously noisy, and therefore, the selection of the optimal data mining method is important for the success of such an analysis. Here, we describe a retrospective analysis of four HTS data sets using three mining approaches: Laplacian-modified naive Bayes, recursive partitioning, and support vector machine (SVM) classifiers with increasing stochastic noise in the form of false positives and false negatives. All three of the data mining methods at hand tolerated increasing levels of false positives even when the ratio of misclassified compounds to true active compounds was 5: 1 in the training set. False negatives in the ratio of 1: 1 were tolerated as well. SVM outperformed the other two methods in capturing active compounds and scaffolds in the top 1%. A Murcko scaffold analysis could explain the differences in enrichments among the four data sets. This study demonstrates that data mining methods can add a true value to the screen even when the data is contaminated with a high level of stochastic noise.