Optimizing artificial neural network models for metabolomics and systems biology: an example using HPLC retention index data

Optimizing artificial neural network models for metabolomics and systems biology: an example using HPLC retention index data
复制标题

DOI:
10.4155/bio.15.1
复制
发表时间:
2015-01-01
期刊:
影响因子:
1.8
通讯作者:
Grant, David F.
Grant, David F.
中科院分区:
医学4区
文献类型:
--
作者:
Hall, L. Mark;Hill, Dennis W.;Grant, David F.

文献摘要

被引文献

相似文献

背景:人工神经网络(ANN)被广泛用于组学数据建模。不同的建模方法和可调参数组合会影响模型性能,并使模型优化复杂化。方法:我们使用保留指数(RI)数据对390个化合物的四个ANN建模参数(学习率退火、停止标准、数据分割方法、网络架构)进行优化评估。使用新测量的1492个化合物的RI值对模型进行独立验证(I-Val)评估。结论:最佳模型显示I-Val标准误差为55 RI单位,并使用Ward聚类数据分割和最小非线性网络架构构建。在停止和最终模型选择中使用验证统计数据比使用测试集统计数据产生更好的独立验证性能。
Background: Artificial Neural Networks (ANN) are extensively used to model omics' data. Different modeling methodologies and combinations of adjustable parameters influence model performance and complicate model optimization. Methodology: We evaluated optimization of four ANN modeling parameters (learning rate annealing, stopping criteria, data split method, network architecture) using retention index (RI) data for 390 compounds. Models were assessed by independent validation (I-Val) using newly measured RI values for 1492 compounds. Conclusion: The best model demonstrated an I-Val standard error of 55 RI units and was built using a Ward's clustering data split and a minimally nonlinear network architecture. Use of validation statistics for stopping and final model selection resulted in better independent validation performance than the use of test set statistics.