Genetic algorithm optimization for pre-processing and variable selection of spectroscopic data

Genetic algorithm optimization for pre-processing and variable selection of spectroscopic data
复制标题

DOI:
10.1093/bioinformatics/bti102
复制
发表时间:
2005-04-01
期刊:
影响因子:
5.8
通讯作者:
Goodacre, R
Goodacre, R
中科院分区:
生物学3区
文献类型:
--
作者:
Jarvis, RM;Goodacre, R

文献摘要

被引文献

相似文献

动机:与光谱数据数学建模相关的主要困难是光谱再现性的不一致和建模技术的黑箱性质。对于生物样本的分析,第一个问题是由于生物、实验和机器的变异性,这可能导致样本量差异和不可避免的基线偏移。因此,如果要形成最佳模型,通常需要对原始数据进行数学修正。第二个问题阻碍了对结果的解释,因为对分析贡献最大的变量不容易揭示;结果,失去了从此类数据中获取新知识的机会。方法:我们使用遗传算法(GA)为傅里叶变换红外(FT-IR)光谱数据选择光谱预处理步骤。我们展示了一种通过 GA 从 FT-IR 光谱中选择重要判别变量的新方法,以便通过判别函数分析 (DFA) 进行多类识别。结果:GA 从总共类似 10(10) 种可能的数学变换中选择合理的预处理步骤。与原始数据模型相比,这些算法的应用使模型误差减少了 16%。 GA-DFA 从全套 882 个谱变量中恢复出 6 个变量,根据这些变量可以形成令人满意的 DFA 模型;因此可以推断这些光谱带反映的生化差异。
Motivation: The major difficulties relating to mathematical modelling of spectroscopic data are inconsistencies in spectral reproducibility and the black box nature of the modelling techniques. For the analysis of biological samples the first problem is due to biological, experimental and machine variability which can lead to sample size differences and unavoidable baseline shifts. Consequently, there is often a requirement for mathematical correction(s) to be made to the raw data if the best possible model is to be formed. The second problem prevents interpretation of the results since the variables that most contribute to the analysis are not easily revealed; as a result, the opportunity to obtain new knowledge from such data is lost.Methods: We used genetic algorithms (GAs) to select spectral pre-processing steps for Fourier transform infrared (FT-IR) spectroscopic data. We demonstrate a novel approach for the selection of important discriminatory variables by GA from FT-IR spectra for multi-class identification by discriminant function analysis (DFA).Results: The GA selects sensible pre-processing steps from a total of similar to 10(10) possible mathematical transformations. Application of these algorithms results in a 16% reduction in the model error when compared against the raw data model. GA-DFA recovers six variables from the full set of 882 spectral variables against which a satisfactory DFA model can be formed; thus inferences can be made as to the biochemical differences that are reflected by these spectral bands.