Prediction of bioconcentration factor using genetic algorithm and artificial neural network

Prediction of bioconcentration factor using genetic algorithm and artificial neural network
复制标题

DOI:
10.1016/s0003-2670(03)00468-9
复制
发表时间:
2003-06-11
影响因子:
6.2
通讯作者:
Konuze, E
Konuze, E
中科院分区:
化学1区
文献类型:
--
作者:
Fatemi, MH;Jalali-Heravi, M;Konuze, E

文献摘要

被引文献

相似文献

在本文中,遗传算法(GA)和逐步多元回归变量选择方法被用作特征选择工具和神经网络用于特征映射。为了提供对这些混合方法的扩展测试,选择了由53种分子的生物浓缩因子(BCF)组成的数据集。通过遗传算法和逐步多元回归方法计算出合适的分子描述符集,并筛选出重要的分子描述符。这些变量用作生成的神经网络的输入。在对网络进行优化和训练之后,将其用于计算预测集的生物富集系数。结果表明,遗传算法在特征选择方面优于逐步多元回归方法。对于网络,使用遗传算法的特征选择方法的平均相对误差之间的预测值和实验值的生物富集系数的训练和预测集分别为0.07和0.13,分别。该模型的校正标准误差和预测标准误差分别为0.26%和0.40%。遗传算法选择的描述符的分析表明,他们是必不可少的,以暗示所研究的分子的空间和静电属性,以获得一个令人满意的定量构效关系(QSAR)模型。(C)2003 Elsevier Science B.V.保留所有权利。
In this paper, genetic algorithm (GA) and stepwise multiple regression variable selection methods were used as a feature-selection tools and neural network was employed for feature mapping. To provide an extended test of these hybrid methods, a data set consists of the bioconcentration factors (BCF) for 53 molecules were selected. Suitable set of molecular descriptors were calculated and the important descriptors were selected by genetic algorithm and stepwise multiple regression methods. These variables serve as inputs to generated neural networks. After optimization and training of the networks, they were used for the calculation of bioconcentration factors for the prediction set. Comparison between results obtained showed the superiority of genetic algorithm over stepwise multiple regression method in feature-selection. For network that used the genetic algorithm for feature-selection methods the average relative error between predicted and experimental values of bioconcentration factor for training and prediction set are 0.07 and 0.13, respectively. Also the standard error of calibration and standard error of prediction are 0.26 and 0.40%, respectively, for this model. An analysis of the descriptor selected by genetic algorithm showed that they are essential to implies the steric and electrostatic attributes of studied molecules to obtain a satisfactory quantitative structure-activity relationship (QSAR) model. (C) 2003 Elsevier Science B.V. All rights reserved.