Benchmark Data Set for in Silico Prediction of Ames Mutagenicity

Benchmark Data Set for in Silico Prediction of Ames Mutagenicity
复制标题

DOI:
10.1021/ci900161g
复制
发表时间:
2009-09-01
影响因子:
5.6
通讯作者:
Mueller, Klaus-Robert
Mueller, Klaus-Robert
中科院分区:
化学2区
文献类型:
--
作者:
Hansen, Katja;Mika, Sebastian;Mueller, Klaus-Robert

文献摘要

被引文献

相似文献

到目前为止,用于构建和评价艾姆斯致突变性预测工具的公开可用数据集在规模和涵盖的化学空间方面非常有限。在本报告中,我们描述了一个新的独特的公共艾姆斯致突变性数据集,包括约6500种非机密化合物(可作为SMILES字符串和SDF获得)及其生物活性。三个商业工具(DEREK,MultiCASE和现成的贝叶斯机器学习器在管道试点)进行了比较,与四个非商业机器学习实现(支持向量机,随机森林,k-最近邻,高斯过程)在新的基准数据集。
Up to now, publicly available data sets to build and evaluate Ames mutagenicity prediction tools have been very limited in terms of size and chemical space covered. In this report we describe a new unique public Ames mutagenicity data set comprising about 6500 nonconfidential compounds (available as SMILES strings and SDF) together with their biological activity. Three commercial tools (DEREK, MultiCASE, and an off-the-shelf Bayesian machine learner in Pipeline Pilot) are compared with four noncommercial machine learning implementations (Support Vector Machines, Random Forests, k-Nearest Neighbors, and Gaussian Processes) on the new benchmark data set.