Influence relevance voting: an accurate and interpretable virtual high throughput screening method.

Influence relevance voting: an accurate and interpretable virtual high throughput screening method.
复制标题

DOI:
10.1021/ci8004379
复制
发表时间:
2009-04
影响因子:
5.6
通讯作者:
Baldi P
Baldi P
中科院分区:
化学2区
文献类型:
--
作者:
Swamidass SJ;Azencott CA;Lin TW;Gramajo H;Tsai SC;Baldi P

文献摘要

参考文献

被引文献

相似文献

给定来自高通量筛选(HTS)实验的活性训练数据,虚拟高通量筛选(vHTS)方法旨在通过计算机模拟预测未经测试的化学品的活性。我们提出了一种新的方法,影响相关选民(IRV),专门为vHTS任务。IRV是一种低参数神经网络,它通过非线性组合训练集中化学品邻居的影响来改进k-最近邻分类器。影响被分解,也是非线性的,成一个相关性组件和一个投票组件。IRV使用两个大型公开比赛的数据和规则进行基准测试,并将其性能与其他参与方法以及内部支持向量机(SVM)方法的性能进行比较。在这些基准数据集上,IRV实现了最先进的结果,在一种情况下与SVM相当,在另一种情况下明显优于SVM,在其预测排序列表的前1%中检索到三倍的活性。与SVM和其他方法相比,IRV具有其他几个重要的优点:(1)输出预测具有概率语义;(2)底层推理是可解释的;(3)训练时间非常短,即使对于非常大的数据集也只有几分钟的数量级;(4)由于自由参数的数量很少,过拟合的风险很小;以及(5)附加信息可以容易地并入IRV架构中。结合其性能,这些品质使IRV特别适合vHTS。
Given activity training data from Hight-Throughput Screening (HTS) experiments, virtual High-Throughput Screening (vHTS) methods aim to predict in silico the activity of untested chemicals. We present a novel method, the Influence Relevance Voter (IRV), specifically tailored for the vHTS task. The IRV is a low-parameter neural network which refines a k-nearest neighbor classifier by non-linearly combining the influences of a chemical's neighbors in the training set. Influences are decomposed, also non-linearly, into a relevance component and a vote component. The IRV is benchmarked using the data and rules of two large, open, competitions, and its performance compared to the performance of other participating methods, as well as of an in-house Support Vector Machine (SVM) method. On these benchmark datasets, IRV achieves state-of-the-art results, comparable to the SVM in one case, and significantly better than the SVM in the other, retrieving three times as many actives in the top 1% of its prediction-sorted list. The IRV presents several other important advantages over SVMs and other methods: (1) the output predictions have a probabilistic semantic; (2) the underlying inferences are interpretable; (3) the training time is very short, on the order of minutes even for very large data sets; (4) the risk of overfitting is minimal, due to the small number of free parameters; and (5) additional information can easily be incorporated into the IRV architecture. Combined with its performance, these qualities make the IRV particularly well suited for vHTS.
DOI: 10.1021/ci600397p
发表时间: 2007-05-01
影响因子: 5.6
作者:
Azencott, Chloe-Agathe;Ksikes, Alexandre;Baldi, Pierre
通讯作者: Baldi, Pierre
DOI: 10.1021/ci970437z
发表时间: 1998-05-01
期刊: JOURNAL OF CHEMICAL INFORMATION AND COMPUTER SCIENCES
影响因子: --
作者:
Flower, DR
通讯作者: Flower, DR
DOI: 10.1109/72.788642
发表时间: 1999-09-01
影响因子: --
作者:
Kwok, JTY
通讯作者: Kwok, JTY
DOI: 10.1021/ci0601160
发表时间: 2006-11-27
影响因子: 5.6
作者:
Cannon, Edward O.;Bender, Andreas;Mitchell, John B. O.
通讯作者: Mitchell, John B. O.
DOI: 10.1007/s10822-008-9181-z
发表时间: 2008-03-01
影响因子: 3.5
作者:
Clark, Robert D.;Webster-Clark, Daniel J.
通讯作者: Webster-Clark, Daniel J.