Maximal Information Coefficient and Support Vector Regression Based Nonlinear Feature Selection and QSAR Modeling on Toxicity of Alcohol Compounds to Tadpoles of Rana temporaria

Maximal Information Coefficient and Support Vector Regression Based Nonlinear Feature Selection and QSAR Modeling on Toxicity of Alcohol Compounds to Tadpoles of Rana temporaria
复制标题

基于最大信息系数和支持向量回归的非线性特征选择和QSAR建模醇类化合物对林蛙蝌蚪的毒性

DOI:
10.21577/0103-5053.20180176
复制
发表时间:
2018
影响因子:
1.4
通讯作者:
Bai Lianyang
Bai Lianyang
中科院分区:
化学4区
文献类型:
--
作者:
Wang Lifeng;Xing Pengwei;Wang Cong;Zhou Xiaomao;Dai Zhijun;Bai Lianyang

文献摘要

相似文献

有机物生物毒性的有效评价对资源利用和环境保护具有重要意义。以110种醇类化合物对中国林蛙蝌蚪的毒性为因变量,用PCLIENT计算的1388个理化参数(特征)代表每种化合物。提出了一种特征选择流程,通过最大信息系数(MIC)法初步筛选出与化合物生物毒性显著相关的282个特征,经支持向量回归(SVR)后向剔除后保留对模型性能有正贡献的138个描述子,并通过聚类分析得到了与化合物生物毒性显著相关的138个特征。最后通过综合最小冗余最大相关性(mRMR)、MIC和SVR的前向选择过程选出18个描述符。针对不同数量变量的特征子集,分别采用多元线性回归(MLR)、偏最小二乘回归(PLS)和支持向量回归(SVR)建立定量构效关系(QSAR)模型。三个回归模型的独立预测评价指数Q分别从-74.787,0.824和0.868增加到0.892,0.878和0.940。结果表明,MIC和SVR中的非线性特征选择方法可以有效地消除不相关的描述符。对于特征间存在非线性关系的高维数据,支持向量回归比传统的统计模型具有更好的QSAR建模性能。该方法在生物毒性化合物等QSAR研究领域具有潜在的应用价值。
Efficient evaluation of biotoxicity of organics is of vital significance to resource utilization and environmental protection. In this study, toxicity of 110 alcohol compounds to tadpoles of Rana temporaria is adopted as the dependent variable and 1388 physiochemical parameters (features) calculated by PCLIENT is used for representing each compound. A feature selection pipeline with three steps is developed to refine the feature subset: 282 features that significantly correlated with biotoxicity of chemical compounds are preliminarily selected via the maximum information coefficient (MIC); 138 descriptors that have positive contribution to the model’s performance are reserved after a support vector regression (SVR) based backward elimination; 18 descriptors are finally selected via a forward selection process that integrated minimal redundancy maximal relevance (mRMR), MIC and SVR. In terms of feature subsets with different numbers of variables, quantitative structure activity relationship (QSAR) models are built using multiple linear regression (MLR), partial least square regression (PLS) and SVR, respectively. The independent prediction evaluation index, Q, increases from −74.787, 0.824 and 0.868 to 0.892, 0.878 and 0.940, for the three regression models, respectively. Results suggest that nonlinear feature selection methods involved in MIC and SVR can effectively eliminate irrelevant descriptors. SVR outperforms classical statistical models to QSAR modeling on high-dimensional data containing nonlinear relationship between features. The methods proposed in this study have a potential application in the QSAR research field such as biotoxicity compounds.