Random forest: A classification and regression tool for compound classification and QSAR modeling

Random forest: A classification and regression tool for compound classification and QSAR modeling
复制标题

DOI:
10.1021/ci034160g
复制
发表时间:
2003-11-01
期刊:
JOURNAL OF CHEMICAL INFORMATION AND COMPUTER SCIENCES
影响因子:
--
通讯作者:
Feuston, BP
Feuston, BP
中科院分区:
其他
文献类型:
--
作者:
Svetnik, V;Liaw, A;Feuston, BP

文献摘要

被引文献

相似文献

一种新的分类和回归工具,随机森林,介绍和研究用于预测化合物的定量或分类的生物活性的基础上的化合物的分子结构的定量描述。随机森林是一个未修剪的分类或回归树的集合,通过使用训练数据的自助样本和树归纳中的随机特征选择来创建。预测是通过聚集(多数投票或平均)集合的预测来进行的。我们为六个化学信息学数据集建立了预测模型。我们的分析表明,随机森林是一个强大的工具,能够提供性能,是迄今为止最准确的方法之一。我们还提出了随机森林的三个额外的功能:内置的性能评估,描述符的相对重要性的措施,和复合相似性的措施,是由描述符的相对重要性加权。相对较高的预测精度和所需特征的集合相结合,使得随机森林特别适合于化学信息学建模。
A new classification and regression tool, Random Forest, is introduced and investigated for predicting a compound's quantitative or categorical biological activity based on a quantitative description of the compound's molecular structure. Random Forest is an ensemble of unpruned classification or regression trees created by using bootstrap samples of the training data and random feature selection in tree induction. Prediction is made by aggregating (majority vote or averaging) the predictions of the ensemble. We built predictive models for six cheminformatics data sets. Our analysis demonstrates that Random Forest is a powerful tool capable of delivering performance that is among the most accurate methods to date. We also present three additional features of Random Forest: built-in performance assessment, a measure of relative importance of descriptors, and a measure of compound similarity that is weighted by the relative importance of descriptors. It is the combination of relatively high prediction accuracy and its collection of desired features that makes Random Forest uniquely suited for modeling in cheminformatics.