STatistical Inference Relief (STIR) feature selection.

STatistical Inference Relief (STIR) feature selection.
复制标题

统计推断浮雕(搅拌)特征选择。

DOI:
10.1093/bioinformatics/bty788
复制
发表时间:
2019-04-15
期刊:
Bioinformatics (Oxford, England)
影响因子:
--
通讯作者:
McKinney BA
McKinney BA
中科院分区:
其他
文献类型:
--
作者:
Le TT;Urbanowicz RJ;Moore JH;McKinney BA

文献摘要

参考文献

被引文献

相似文献

Relief是机器学习算法的一个家族,它使用最近邻来选择与结果相关的特征,这些特征可能是由于上位性或与高维数据中其他特征的统计相互作用。从统计学意义上讲,基于浮雕的估计量是非参数的,因为它们没有参数化模型,而参数化模型具有估计量的潜在概率分布,因此难以确定基于浮雕的属性估计量的统计显著性。因此,需要一种统计推断的形式主义,以避免施加任意阈值来选择最重要的特征。我们重新概念化的救济为基础的功能选择算法,以创建一个新的家庭的统计推断救济(STIR)估计,保留识别相互作用的能力,同时将最近邻距离的样本方差到属性的重要性估计。该方差允许计算特征的统计学显著性并调整基于缓解的评分的多重检验。具体来说,我们开发了一个伪t-检验版本的救济为基础的算法的病例对照数据。我们展示了统计功率和控制的STIR家族的功能选择方法的I型错误的面板上的模拟数据,表现出的属性反映在真实的基因表达数据,包括主效应和网络相互作用的影响。我们比较了STIR的性能时,自适应半径方法作为最近邻构造器与STIR时,固定k最近邻构造器使用。我们将STIR应用于来自重性抑郁症研究的真实的RNA-Seq数据,并讨论STIR直接扩展到全基因组关联研究。代码和数据可在http://insilico.utulsa.edu/software/STIR获得。 补充数据可在Bioinformatics在线获得。
Relief is a family of machine learning algorithms that uses nearest-neighbors to select features whose association with an outcome may be due to epistasis or statistical interactions with other features in high-dimensional data. Relief-based estimators are non-parametric in the statistical sense that they do not have a parameterized model with an underlying probability distribution for the estimator, making it difficult to determine the statistical significance of Relief-based attribute estimates. Thus, a statistical inferential formalism is needed to avoid imposing arbitrary thresholds to select the most important features. We reconceptualize the Relief-based feature selection algorithm to create a new family of STatistical Inference Relief (STIR) estimators that retains the ability to identify interactions while incorporating sample variance of the nearest neighbor distances into the attribute importance estimation. This variance permits the calculation of statistical significance of features and adjustment for multiple testing of Relief-based scores. Specifically, we develop a pseudo t-test version of Relief-based algorithms for case-control data. We demonstrate the statistical power and control of type I error of the STIR family of feature selection methods on a panel of simulated data that exhibits properties reflected in real gene expression data, including main effects and network interaction effects. We compare the performance of STIR when the adaptive radius method is used as the nearest neighbor constructor with STIR when the fixed-k nearest neighbor constructor is used. We apply STIR to real RNA-Seq data from a study of major depressive disorder and discuss STIR’s straightforward extension to genome-wide association studies. Code and data available at http://insilico.utulsa.edu/software/STIR. Supplementary data are available at Bioinformatics online.
DOI: 10.3389/fgene.2011.00109
发表时间: 2011
影响因子: 3.7
作者:
McKinney BA;Pajewski NM
通讯作者: Pajewski NM
DOI: 10.1093/bioinformatics/btx298
发表时间: 2017-09-15
期刊: BIOINFORMATICS
影响因子: 5.8
作者:
Le, Trang T.;Simmons, W. Kyle;McKinney, Brett A.
通讯作者: McKinney, Brett A.
DOI: 10.1038/msb.2013.2
发表时间: 2013
影响因子: 9.9
作者:
通讯作者: --
DOI: 10.1371/journal.pgen.1000432
发表时间: 2009-03
期刊: PLoS genetics
影响因子: 4.5
作者:
McKinney BA;Crowe JE;Guo J;Tian D
通讯作者: Tian D
DOI: 10.1023/a:1025667309714
发表时间: 2003-10-01
期刊: MACHINE LEARNING
影响因子: 7.5
作者:
Robnik-Sikonja, M;Kononenko, I
通讯作者: Kononenko, I