Robustness and accuracy of methods for high dimensional data analysis based on Student's t-statistic

Robustness and accuracy of methods for high dimensional data analysis based on Student's t-statistic
复制标题

DOI:
10.1111/j.1467-9868.2010.00761.x
复制
发表时间:
2011-01-01
影响因子:
5.8
通讯作者:
Jin, Jiashun
Jin, Jiashun
中科院分区:
数学1区
文献类型:
--
作者:
Delaigle, Aurore;Hall, Peter;Jin, Jiashun

文献摘要

被引文献

相似文献

学生的T统计目前正在发现当今的应用程序是一个多世纪前引入时从未设想的。这些应用中的许多依赖于属性,例如对重尾抽样分布的鲁棒性,直到最近才明确考虑。我们将T统计量的这些特征在其应用于非常高维问题的上下文中,包括特征选择和排名,许多不同假设的同时测试以及稀疏的高维信号检测。突出显示了T比率的鲁棒性特性,并确定这些特性在引导程序的应用下保存。特别是,引导方法对偏度正确,因此即使在极端的尾巴中也会导致二阶精度。确实,这表明,引导程序以及更流行但更准确的T分布和正常近似值在尾部比分布的中间更有效。这些属性激发了新方法,例如基于引导的信号检测技术,将注意力局限于统计量的重要尾巴。
Student's t-statistic is finding applications today that were never envisaged when it was introduced more than a century ago. Many of these applications rely on properties, e.g. robustness against heavy-tailed sampling distributions, that were not explicitly considered until relatively recently. We explore these features of the t-statistic in the context of its application to very high dimensional problems, including feature selection and ranking, the simultaneous testing of many different hypotheses and sparse, high dimensional signal detection. Robustness properties of the t-ratio are highlighted, and it is established that those properties are preserved under applications of the bootstrap. In particular, bootstrap methods correct for skewness and therefore lead to second-order accuracy, even in the extreme tails. Indeed, it is shown that the bootstrap and also the more popular but less accurate t-distribution and normal approximations are more effective in the tails than towards the middle of the distribution. These properties motivate new methods, e.g. bootstrap-based techniques for signal detection, that confine attention to the significant tail of a statistic.