Prediction of bioconcentration factors in fish and invertebrates using machine learning.

Prediction of bioconcentration factors in fish and invertebrates using machine learning.
复制标题

DOI:
10.1016/j.scitotenv.2018.08.122
复制
发表时间:
2019-01-15
期刊:
The Science of the total environment
影响因子:
--
通讯作者:
Barron LP
Barron LP
中科院分区:
其他
文献类型:
--
作者:
Miller TH;Gallidabino MD;MacRae JI;Owen SF;Bury NR;Barron LP

文献摘要

参考文献

被引文献

相似文献

机器学习的应用最近引起了生态毒理学领域的兴趣,因为它能够模拟和预测化学和/或生物过程,例如生物浓度的预测。然而,不同模型的比较和无脊椎动物生物浓度的预测以前没有进行过评估。本文介绍了用于预测鱼类生物浓度的24种线性和机器学习模型的比较,并确定了影响积累的重要因素。检验数据(n = 110例)的R2和均方根误差(RMSE)范围分别为0.23-0.73和0.34-1.20。通过神经网络和基于树的学习器对模型性能进行严格评估,显示出最佳性能。选择一个优化的4层多层感知器(14个描述符)进行进一步测试。该模型应用于淡水无脊椎动物Gammarus pulex的跨物种生物浓度预测。验证和试验数据的R2分别为0.99和0.93,模型表现出良好的性能。确定影响生物浓度的重要分子描述符是分子质量(MW)、辛醇-水分布系数(logD)、拓扑极性表面积(TPSA)和氮原子数(nN)等。对PBT等危害标准的建模显示出有可能取代对动物试验的需求。然而,迄今为止,机器学习模型在监管环境中的使用一直很少,本文对此进行了批判性讨论。从积累的实验估计转向计算机模拟,将能够迅速确定可能对环境健康和食物链构成风险的污染物的优先次序。对24种预测鱼类生物浓度因子的模型进行了评价。机器学习显示出良好的预测性能。第一个机器学习应用于预测无脊椎动物的生物浓度跨物种建模受到案例相似性和生物变异性的限制。TPSA、LogD和Mw是模拟积累过程的重要描述符。
The application of machine learning has recently gained interest from ecotoxicological fields for its ability to model and predict chemical and/or biological processes, such as the prediction of bioconcentration. However, comparison of different models and the prediction of bioconcentration in invertebrates has not been previously evaluated. A comparison of 24 linear and machine learning models is presented herein for the prediction of bioconcentration in fish and important factors that influenced accumulation identified. R2 and root mean square error (RMSE) for the test data (n = 110 cases) ranged from 0.23–0.73 and 0.34–1.20, respectively. Model performance was critically assessed with neural networks and tree-based learners showing the best performance. An optimised 4-layer multi-layer perceptron (14 descriptors) was selected for further testing. The model was applied for cross-species prediction of bioconcentration in a freshwater invertebrate, Gammarus pulex. The model for G. pulex showed good performance with R2 of 0.99 and 0.93 for the verification and test data, respectively. Important molecular descriptors determined to influence bioconcentration were molecular mass (MW), octanol-water distribution coefficient (logD), topological polar surface area (TPSA) and number of nitrogen atoms (nN) among others. Modelling of hazard criteria such as PBT, showed potential to replace the need for animal testing. However, the use of machine learning models in the regulatory context has been minimal to date and is critically discussed herein. The movement away from experimental estimations of accumulation to in silico modelling would enable rapid prioritisation of contaminants that may pose a risk to environmental health and the food chain. Evaluation of 24 models to predict bioconcentration factors in fish is presented. Machine learning showed good predictive performance. First machine learning application to predict bioconcentration in invertebrates Cross-species modelling is limited by case similarity and biological variability. TPSA, LogD, and Mw were important descriptors for modelling accumulation processes.
DOI: 10.1016/s0003-2670(03)00468-9
发表时间: 2003-06-11
影响因子: 6.2
作者:
Fatemi, MH;Jalali-Heravi, M;Konuze, E
通讯作者: Konuze, E
DOI: 10.1021/acs.est.7b01265
发表时间: 2017-06-20
影响因子: 11.4
作者:
Karlsson, Maja V.;Carter, Laura J.;Boxall, Alistair B. A.
通讯作者: Boxall, Alistair B. A.
DOI: 10.1016/j.chemosphere.2015.12.022
发表时间: 2016-03-01
期刊: CHEMOSPHERE
影响因子: 8.8
作者:
de Solla, S. R.;Gilroy, E. A. M.;Gillis, P. L.
通讯作者: Gillis, P. L.
DOI: 10.1016/j.scitotenv.2013.03.104
发表时间: 2013-07-01
影响因子: 9.8
作者:
Gissi, Andrea;Nicolotti, Orazio;Benfenati, Emilio
通讯作者: Benfenati, Emilio
DOI: 10.1023/a:1015040217741
发表时间: 1999-10-01
影响因子: 3.7
作者:
Kelder, J;Grootenhuis, PDJ;Ploemen, JP
通讯作者: Ploemen, JP