Recursive feature elimination in random forest classification supports nanomaterial grouping

Recursive feature elimination in random forest classification supports nanomaterial grouping
复制标题

DOI:
10.1016/j.impact.2019.100179
复制
发表时间:
2019-03-01
期刊:
影响因子:
4.9
通讯作者:
Haase, Andrea
Haase, Andrea
中科院分区:
环境科学与生态学3区
文献类型:
--
作者:
Bahl, Aileen;Hellack, Bryan;Haase, Andrea

文献摘要

被引文献

相似文献

纳米材料(NM)可以在相同化学物质的许多不同变体中生产。通过生成测试数据对每个变体进行深入的安全性评估根本不可行。因此,NM分组方法,将显着减少时间和数量的测试新NM的迫切需要。然而,识别结构相似的NM变体仍然具有挑战性,因为许多物理化学性质可能是相关的。在这里,我们的目的是强调机器学习模型在NM分组过程中的价值,通过考虑对11个选定的,表征良好的NM的案例研究。为此,我们将这些NM的理化性质与吸入毒性的特征联系起来。我们应用了无监督和有监督的机器学习技术来确定哪种属性组合最具预测性。首先,我们评估NM相似性在无监督的方式使用主成分分析(PCA),随后叠加的活动标签与k-最近邻的方法相结合。然后,我们使用随机森林(RFs)作为一种监督机器学习技术,它直接使用活动类的知识来定义NM相似性。因此,相似性仅定义在与活性相关性最高的那些性质上,因此具有最高的区分能力。为了提高性能,我们然后使用递归特征消除(RFE)删除无信息的功能偏置的结果。基于RFE的简化RF模型实现了最佳性能,其中获得了0.82的平衡精度。在11种不同的性质中,我们确定zeta电位、氧化还原电位和溶解速率对本数据集中的生物NM活性具有最强的预测影响。虽然数据集相对于所研究的NM的数量太小,并且由于仅涵盖少数材料类别,因此适用性域预计非常有限,但我们的研究展示了如何实施机器学习和特征选择方法来识别最相关的物理化学NM属性毒性。我们建议,一旦最相关的属性已被检测到的模型建立在足够数量的不同NM和跨多个NM类,他们应该得到特别强调,在未来的分组方法。
Nanomaterials (NMs) can be produced in numerous different variants of the same chemical substance. An in-depth safety assessment for each variant by generating test data will simply not be feasible. Thus, NM grouping approaches that would significantly reduce the time and amount of testing for novel NMs are urgently needed. However, identifying structurally similar NM variants remains challenging as many physico-chemical properties could be relevant.Here, we aimed at emphasizing on the value of machine learning models in the process of NM grouping by considering a case study on eleven selected, well-characterized NMs. To that end, we linked physico-chemical properties of these NMs to characterized hallmarks for inhalation toxicity. We applied unsupervised and supervised machine learning techniques to determine which combination of properties is most predictive. First, we assessed NM similarity in an unsupervised manner using principal component analysis (PCA) followed by subsequent superposition of activity labels combined with a k-nearest neighbors approach. Then, we used random forests (RFs) as a supervised machine learning technique which directly uses the knowledge on the activity class in the process of defining NM similarity. Thus, similarity was defined only on those properties showing the highest correlation with the activity and therefore had the highest discriminative power. In order to improve the performance, we then used recursive feature elimination (RFE) to delete uninformative features biasing the results. The best performance was achieved by the reduced RF model based on RFE where a balanced accuracy of 0.82 was obtained. Out of eleven different properties we determined zeta potential, redox potential and dissolution rate to have the strongest predicting impact on biological NM activity in the present dataset. Though the dataset is too small with respect to the number of NMs studied and the applicability domain is expected to be very limited due to the fact that only few material classes were covered, our study demonstrates how machine learning and feature selection methods can be implemented for identifying the most relevant physico-chemical NM properties with respect to toxicity. We suggest that once the most relevant properties have been detected in a model built on a sufficient number of different NMs and across multiple NM classes, they should obtain special emphasis in future grouping approaches.