Are bigger data sets better for machine learning? Fusing single-point and dual-event dose response data for Mycobacterium tuberculosis.

Are bigger data sets better for machine learning? Fusing single-point and dual-event dose response data for Mycobacterium tuberculosis.
复制标题

DOI:
10.1021/ci500264r
复制
发表时间:
2014-07-28
影响因子:
5.6
通讯作者:
Reynolds RC
Reynolds RC
中科院分区:
化学2区
文献类型:
--
作者:
Ekins S;Freundlich JS;Reynolds RC

文献摘要

被引文献

相似文献

结核病是一种被忽视的主要疾病,人们仍在继续寻找新的治疗方法。公共领域有大量针对结核分枝杆菌 (Mtb) 的大型表型筛选数据。由于机器学习方法可以从过去的数据中学习,因此我们有兴趣解决更多数据是否可以构建更好的模型的问题。我们现在描述使用贝叶斯机器学习来评估我们是否可以通过将大量单点数据与更小(更高质量)的双事件数据集相结合来改进我们的模型,该数据集使用全细胞抗结核活性和 Vero 细胞细胞毒性的剂量反应数据。我们评估了 12 个模型,包括不同的单点、双事件剂量反应、单点和双事件剂量反应以及来自同一实验室的三个不同数据集的组合数据集。我们使用来自同一组的活性和非活性化合物的第四个数据集以及来自 GlaxoSmithKline 的 177 种活性化合物作为测试集。我们的数据表明,与较小数量级的双事件模型(内部 ROC 范围 0.6-0.83 和外部 ROC 0.54-0.83)相比,将单点与双事件剂量反应数据相结合不会削弱基于这些模型的受试者工作曲线(ROC)(内部 ROC 范围 0.83-0.91,外部 ROC 范围 0.62-0.83)的模型的内部或外部预测能力。总之,使用 1200-5000 种化合物开发的模型似乎与使用 25,000 到 350,000 种分子生成的模型具有相同的预测能力。我们的结果对于证明进一步的 HTS 与基于模型预测的集中测试的合理性具有重要意义。
Tuberculosis is a major neglected disease for which the quest to find new treatments continues. There is an abundance of data from large phenotypic screens in the public domain against Mycobacterium tuberculosis (Mtb). Since machine learning methods can learn from past data, we were interested in addressing whether more data builds better models. We now describe using Bayesian machine learning to assess whether we can improve our models by combining the large quantities of single-point data with the much smaller (higher quality) dual-event datasets, which use both dose-response data for both whole-cell antitubercular activity and Vero cell cytotoxicity. We have evaluated 12 models ranging from different single-point, dual-event dose response, single-point and dual-event dose response as well as combined datasets for three distinct datasets from the same laboratory. We used a fourth dataset of active and inactive compounds from the same group as well as a smaller set of 177 active compounds from GlaxoSmithKline as test sets. Our data suggest combining single-point with dual-event dose response data does not diminish the internal or external predictive ability of the models based on the receiver operator curve (ROC) for these models (internal ROC range 0.83-0.91, external ROC range 0.62-0.83) compared to the orders of magnitude smaller dual event models (internal ROC range 0.6-0.83 and external ROC 0.54-0.83). In conclusion, models developed with 1200-5000 compounds appear to be as predictive as those generated with 25,000 to 350,000 molecules. Our results have implications for justifying further HTS versus focused testing based on model predictions.