Feature selection and classifier performance on diverse bio- logical datasets.

Feature selection and classifier performance on diverse bio- logical datasets.
复制标题

DOI:
10.1186/1471-2105-15-s13-s4
复制
发表时间:
2014
期刊:
影响因子:
3
通讯作者:
Nelson CE
Nelson CE
中科院分区:
生物学4区
文献类型:
--
作者:
Hemphill E;Lindsay J;Lee C;Măndoiu II;Nelson CE

文献摘要

被引文献

相似文献

技术范围不断扩大,可以产生大量用于研究和临床应用的生物标志物。从高维数据集中选择信息最丰富的生物标志物,并结合确定与该生物标志物集一起使用的最可靠和最准确的分类算法,可能是一项艰巨的任务。现有的特征选择和分类算法的调查通常集中于单一数据类型,例如基因表达微阵列,而很少探索模型在多种生物数据类型上的性能。本文介绍了大规模实证研究的结果,其中使用大量流行的特征选择和分类算法来识别 NCI-60 癌细胞系的起源组织。实施计算流程以最大限度地提高 NCI-60 细胞系可用的五种不同数据类型的所有参数下所有模型的预测准确性。使用外部数据进行验证实验以证明稳健性。正如预期的那样,生物标志物的数据类型和数量对预测模型的性能有显着影响。尽管没有任何模型或数据类型在整个测试标记数量范围内均优于其他模型或数据类型,但可以看到一些明显的趋势。在生物标志物数量较少的情况下,基因和蛋白质表达数据类型能够比其他三种数据类型(即 SNP、阵列比较基因组杂交 (aCGH) 和 microRNA 数据)更好地区分癌细胞系。有趣的是,随着所选生物标志物数量的增加,基于 SNP 数据匹配的表现最佳的分类器或略优于基于基因和蛋白质表达的分类器,而基于 aCGH 和 microRNA 数据的分类器仍然表现最差。据观察,一类特征选择和分类器在数据类型和标记数量方面始终表现最佳,这表明无论分析中使用的数据类型如何,性能良好的特征选择/分类器配对在生物分类问题中可能是稳健的。
There is an ever-expanding range of technologies that generate very large numbers of biomarkers for research and clinical applications. Choosing the most informative biomarkers from a high-dimensional data set, combined with identifying the most reliable and accurate classification algorithms to use with that biomarker set, can be a daunting task. Existing surveys of feature selection and classification algorithms typically focus on a single data type, such as gene expression microarrays, and rarely explore the model's performance across multiple biological data types. This paper presents the results of a large scale empirical study whereby a large number of popular feature selection and classification algorithms are used to identify the tissue of origin for the NCI-60 cancer cell lines. A computational pipeline was implemented to maximize predictive accuracy of all models at all parameters on five different data types available for the NCI-60 cell lines. A validation experiment was conducted using external data in order to demonstrate robustness. As expected, the data type and number of biomarkers have a significant effect on the performance of the predictive models. Although no model or data type uniformly outperforms the others across the entire range of tested numbers of markers, several clear trends are visible. At low numbers of biomarkers gene and protein expression data types are able to differentiate between cancer cell lines significantly better than the other three data types, namely SNP, array comparative genome hybridization (aCGH), and microRNA data. Interestingly, as the number of selected biomarkers increases best performing classifiers based on SNP data match or slightly outperform those based on gene and protein expression, while those based on aCGH and microRNA data continue to perform the worst. It is observed that one class of feature selection and classifier are consistently top performers across data types and number of markers, suggesting that well performing feature-selection/classifier pairings are likely to be robust in biological classification problems regardless of the data type used in the analysis.