A Systematic Evaluation of Feature Selection and Classification Algorithms Using Simulated and Real miRNA Sequencing Data.

A Systematic Evaluation of Feature Selection and Classification Algorithms Using Simulated and Real miRNA Sequencing Data.
复制标题

使用模拟和真实 miRNA 测序数据对特征选择和分类算法进行系统评估

DOI:
10.1155/2015/178572
复制
发表时间:
2015
影响因子:
--
通讯作者:
Chen F
Chen F
中科院分区:
工程技术4区
文献类型:
--
作者:
Yang S;Guo L;Shao F;Zhao Y;Chen F

文献摘要

被引文献

相似文献

测序被广泛用于发现microRNA(miRNAs)与疾病之间的关联。然而,使用测序获得的数据的负二项分布(NB)和高维数可导致低功效结果和低再现性。已经提出了几种统计学习算法来处理测序数据,尽管对这些方法的评估是必不可少的,但这样的研究相对较少。通过仿真,比较了baySeq、DESeq、edgeR、秩和检验、lasso、粒子群乐观决策树和随机森林等7种特征选择算法在不同条件下的性能,并分析了均值、NB的离散度和信噪比的差异.使用真实的数据来评估RF、逻辑回归和支持向量机的性能。基于模拟和真实的数据,我们讨论的FS和分类算法的行为。Apriori算法从来自癌症基因组学图谱的六个数据集的去调控的miRNA中识别出频繁项集(mir-133 a、mir-133 b、mir-183、mir-937和mir-96)。综合考虑这些发现并考虑计算内存需求,我们提出了一种结合edgeR和DESeq的策略,以实现大样本量。
Sequencing is widely used to discover associations between microRNAs (miRNAs) and diseases. However, the negative binomial distribution (NB) and high dimensionality of data obtained using sequencing can lead to low-power results and low reproducibility. Several statistical learning algorithms have been proposed to address sequencing data, and although evaluation of these methods is essential, such studies are relatively rare. The performance of seven feature selection (FS) algorithms, including baySeq, DESeq, edgeR, the rank sum test, lasso, particle swarm optimistic decision tree, and random forest (RF), was compared by simulation under different conditions based on the difference of the mean, the dispersion parameter of the NB, and the signal to noise ratio. Real data were used to evaluate the performance of RF, logistic regression, and support vector machine. Based on the simulation and real data, we discuss the behaviour of the FS and classification algorithms. The Apriori algorithm identified frequent item sets (mir-133a, mir-133b, mir-183, mir-937, and mir-96) from among the deregulated miRNAs of six datasets from The Cancer Genomics Atlas. Taking these findings altogether and considering computational memory requirements, we propose a strategy that combines edgeR and DESeq for large sample sizes.