Enhancement of Plant Metabolite Fingerprinting by Machine Learning

Enhancement of Plant Metabolite Fingerprinting by Machine Learning
复制标题

DOI:
10.1104/pp.109.150524
复制
发表时间:
2010-08-01
期刊:
影响因子:
7.4
通讯作者:
Beale, Michael H.
Beale, Michael H.
中科院分区:
生物学1区
文献类型:
--
作者:
Scott, Ian M.;Vermeer, Cornelia P.;Beale, Michael H.

文献摘要

被引文献

相似文献

利用H-1-核磁共振、傅立叶变换红外光谱和流动注射电喷雾-质谱法对已知或预测代谢损伤的拟南芥突变体进行了代谢产物指纹图谱分析。指纹图谱能够处理比传统色谱分析多五倍的植物,并且在区分突变方面具有竞争力,而不是那些只影响低丰度代谢物的突变。尽管指纹快速而复杂,但它们产生了对新陈代谢的洞察力(例如,单个病变的影响通常不限于个别途径)。在指纹技术中,H-1核磁共振区分野生型和突变表型最多,傅立叶变换红外分辨最少。为了最大限度地利用指纹提供的信息,数据分析至关重要。如果数据模型被限制在主成分分析得分曲线图上,三分之一的独特表型可能会被忽视。在测试的几种方法中,机器学习(ML)算法,即支持向量机或随机森林(RF)分类器,在表型识别方面是无与伦比的。支持向量机通常是性能最好的分类器,但RFS产生了一些特别有用的度量。首先,RFS估计了突变表型之间的边际,然后可以通过Sammon映射或等级聚类来可视化它们的关系。其次,RFS为区分突变的指纹中的特征提供了重要性分数。这些分数与方差分析F值相关(Kruskal-Wallis检验、真阳性和假阳性测量、互信息和救济特征选择算法也是如此)。ML分类器作为在一个数据集上训练以预测另一个数据集的模型,对于集中的代谢学查询是理想的,例如突变表型的独特性和一致性。重点介绍了在植物生理学中使用ML的易用软件。
Metabolite fingerprinting of Arabidopsis (Arabidopsis thaliana) mutants with known or predicted metabolic lesions was performed by H-1-nuclear magnetic resonance, Fourier transform infrared, and flow injection electrospray-mass spectrometry. Fingerprinting enabled processing of five times more plants than conventional chromatographic profiling and was competitive for discriminating mutants, other than those affected in only low-abundance metabolites. Despite their rapidity and complexity, fingerprints yielded metabolomic insights (e. g. that effects of single lesions were usually not confined to individual pathways). Among fingerprint techniques, H-1-nuclear magnetic resonance discriminated the most mutant phenotypes from the wild type and Fourier transform infrared discriminated the fewest. To maximize information from fingerprints, data analysis was crucial. One-third of distinctive phenotypes might have been overlooked had data models been confined to principal component analysis score plots. Among several methods tested, machine learning (ML) algorithms, namely support vector machine or random forest (RF) classifiers, were unsurpassed for phenotype discrimination. Support vector machines were often the best performing classifiers, but RFs yielded some particularly informative measures. First, RFs estimated margins between mutant phenotypes, whose relations could then be visualized by Sammon mapping or hierarchical clustering. Second, RFs provided importance scores for the features within fingerprints that discriminated mutants. These scores correlated with analysis of variance F values (as did Kruskal-Wallis tests, true-and false-positive measures, mutual information, and the Relief feature selection algorithm). ML classifiers, as models trained on one data set to predict another, were ideal for focused metabolomic queries, such as the distinctiveness and consistency of mutant phenotypes. Accessible software for use of ML in plant physiology is highlighted.