The projack: a resampling approach to correct for ranking bias in high-throughput studies.

The projack: a resampling approach to correct for ranking bias in high-throughput studies.
复制标题

Projack:一种重采样方法,用于纠正高通量研究中的排名偏差。

DOI:
10.1093/biostatistics/kxv022
复制
发表时间:
2016
期刊:
Biostatistics (Oxford, England)
影响因子:
--
通讯作者:
Wright,FredA
Wright,FredA
中科院分区:
--
文献类型:
--
作者:
Zhou,Yi-Hui;Wright,FredA

文献摘要

相似文献

排序推理的问题出现在许多设置中,研究者希望在对一组统计数据进行排序后执行参数推理。与单一假设的推断相比,排序过程引入了相当大的偏差,这是遗传关联中被称为“赢家诅咒”的问题。本文介绍了一种基于重序折刀和交叉验证的预测模型。projack是一种基于重采样的程序,为一组可能相关的统计数据提供低偏差的预期排序效应大小参数估计。该方法灵活,广泛适用于高维数据集,包括基因组学平台产生的数据集。最初,对于原始数据可用于重新采样的设置,projack可以扩展到只有值向量可用的情况。我们举例说明了在遗传关联中纠正赢家诅咒的projack,尽管它可以更普遍地使用。
The problem of ranked inference arises in a number of settings, for which the investigator wishes to perform parameter inference after ordering a set ofstatistics. In contrast to inference for a single hypothesis, the ranking procedure introduces considerable bias, a problem known as the “winner's curse” in genetic association. We introduce theprojack(for Prediction by Re- Ordered Jackknife and Cross-Validation,-fold). The projack is a resampling-based procedure that provides low-bias estimates of the expected ranked effect size parameter for a set of possibly correlatedstatistics. The approach is flexible, and has wide applicability to high-dimensional datasets, including those arising from genomics platforms. Initially, motivated for the setting where original data are available for resampling, the projack can be extended to the situation where only the vector ofvalues is available. We illustrate the projack for correction of the winner's curse in genetic association, although it can be used much more generally.