Feature Selection Methods for Early Predictive Biomarker Discovery Using Untargeted Metabolomic Data.

Feature Selection Methods for Early Predictive Biomarker Discovery Using Untargeted Metabolomic Data.
复制标题

DOI:
10.3389/fmolb.2016.00030
复制
发表时间:
2016
影响因子:
5
通讯作者:
Pujos-Guillot E
Pujos-Guillot E
中科院分区:
生物学3区
文献类型:
--
作者:
Grissa D;Pétéra M;Brandolini M;Napoli A;Comte B;Pujos-Guillot E

文献摘要

被引文献

相似文献

非靶向代谢组学是一种强大的表型分析工具,可用于更好地理解人类病理学发展中涉及的生物学机制,并识别早期预测生物标志物。这种方法基于多个分析平台,如质谱(MS),化学计量学和生物信息学,产生大量复杂的数据,需要适当的分析来提取有生物意义的信息。尽管有各种可用的工具,但在不冒过拟合风险的情况下处理具有有限数量个体的如此大且有噪声的数据集仍然是一个挑战。此外,当目标集中在发生前几年识别临床结果的早期预测标志物时,使用适当的算法和工作流程以便能够在大量数据中发现微妙的影响变得至关重要。在这种情况下,这项工作包括研究工作流程描述的一般特征选择过程中,使用知识发现和数据挖掘方法,提出先进的解决方案,预测生物标志物的发现。该策略的重点是评估特征选择的数字-符号方法的组合,目的是获得产生有效和准确预测模型的代谢物的最佳组合。首先依靠数值方法,特别是机器学习方法(SVM-RFE,RF,RF-RFE)和单变量统计分析(ANOVA),对原始代谢组学数据集和减少的子集进行比较研究。为了降低过拟合的风险,采用LOOCV作为拟合方法。从这些不同的方法的组合的重要性不同的分数获得的最佳k-功能进行了比较,并允许使用形式概念分析确定变量的稳定性。结果表明,RF-Gini结合ANOVA进行特征选择的兴趣,因为这两种互补的方法允许选择48个最佳候选预测。在这个简化的数据集上使用线性逻辑回归使我们能够在预测准确性和假阳性数量方面获得最佳性能,其中模型包括5个顶级变量。因此,这些结果突出了特征选择方法的兴趣以及在减少的数据集上工作的重要性,以用于识别从非靶向代谢组学数据发出的预测性生物标志物。
Untargeted metabolomics is a powerful phenotyping tool for better understanding biological mechanisms involved in human pathology development and identifying early predictive biomarkers. This approach, based on multiple analytical platforms, such as mass spectrometry (MS), chemometrics and bioinformatics, generates massive and complex data that need appropriate analyses to extract the biologically meaningful information. Despite various tools available, it is still a challenge to handle such large and noisy datasets with limited number of individuals without risking overfitting. Moreover, when the objective is focused on the identification of early predictive markers of clinical outcome, few years before occurrence, it becomes essential to use the appropriate algorithms and workflow to be able to discover subtle effects among this large amount of data. In this context, this work consists in studying a workflow describing the general feature selection process, using knowledge discovery and data mining methodologies to propose advanced solutions for predictive biomarker discovery. The strategy was focused on evaluating a combination of numeric-symbolic approaches for feature selection with the objective of obtaining the best combination of metabolites producing an effective and accurate predictive model. Relying first on numerical approaches, and especially on machine learning methods (SVM-RFE, RF, RF-RFE) and on univariate statistical analyses (ANOVA), a comparative study was performed on an original metabolomic dataset and reduced subsets. As resampling method, LOOCV was applied to minimize the risk of overfitting. The best k-features obtained with different scores of importance from the combination of these different approaches were compared and allowed determining the variable stabilities using Formal Concept Analysis. The results revealed the interest of RF-Gini combined with ANOVA for feature selection as these two complementary methods allowed selecting the 48 best candidates for prediction. Using linear logistic regression on this reduced dataset enabled us to obtain the best performances in terms of prediction accuracy and number of false positive with a model including 5 top variables. Therefore, these results highlighted the interest of feature selection methods and the importance of working on reduced datasets for the identification of predictive biomarkers issued from untargeted metabolomics data.