In silico prediction of toxic action mechanisms of phenols for imbalanced data with Random Forest learner

In silico prediction of toxic action mechanisms of phenols for imbalanced data with Random Forest learner
复制标题

使用随机森林学习器对不平衡数据的酚类毒性作用机制进行计算机预测

DOI:
10.1016/j.jmgm.2012.01.002
复制
发表时间:
2012-05-01
影响因子:
2.9
通讯作者:
Guo, Chang
Guo, Chang
中科院分区:
生物学4区
文献类型:
--
作者:
Chen, Jing;Tang, Yuan Yan;Guo, Chang

文献摘要

被引文献

相似文献

随着对工业和民用产品中化合物安全性快速有效评价的需求日益增加,二氧化硅毒性探测技术为环境危害评价提供了一种经济的途径。在过去的十年中,已有大量的计算机模拟研究建立了定量构效关系模型来预测毒性机制。这些方法大多受益于数据分析和机器学习技术,这些技术严重依赖于数据集的特征。对于梨形四膜虫毒性数据集,存在着很大的技术挑战数据不平衡。数据类分布的偏态会严重影响稀有类的预测性能。以往对苯酚毒性作用机理预测的研究大多没有考虑这一实际问题。在这项工作中,我们通过考虑两种类型的错误分类之间的差异来处理这个问题。在代价敏感学习框架中采用随机森林学习器,根据选定的分子描述符构建预测模型。在计算实验中,全局和局部模型都获得了可观的整体预测精度。尤其是稀有职业的表现,确实得到了提升。此外,对于这些模型的实际使用,可以通过根据应用目标使用不同的成本矩阵来调整两个错误分类的平衡。(C)2012 Elsevier Inc. All rights reserved.
With an increasing need for the rapid and effective safety assessment of compounds in industrial and civil-use products, in silica toxicity exploration techniques provide an economic way for environmental hazard assessment. The previous in silico researches have developed many quantitative structure-activity relationships models to predict toxicity mechanisms for last decade. Most of these methods benefit from data analysis and machine learning techniques, which rely heavily on the characteristics of data sets. For Tetrahymena pyriformis toxicity data sets, there is a great technical challenge data imbalance. The skewness of data class distribution would greatly deteriorate the prediction performance on rare classes. Most of the previous researches for phenol mechanisms of toxic action prediction did not consider this practical problem. In this work, we dealt with the problem by considering the difference between the two types of misclassifications. Random Forest learner was employed in cost-sensitive learning framework to construct prediction models based on selected molecular descriptors. In computational experiments, both the global and local models obtained appreciable overall prediction accuracies. Particularly, the performance on rare classes was indeed promoted. Moreover, for practical usage of these models, the balance of the two misclassifications can be adjusted by using different cost matrices according to the application goals. (C) 2012 Elsevier Inc. All rights reserved.