Large numbers of explanatory variables, a semi-descriptive analysis

Large numbers of explanatory variables, a semi-descriptive analysis
复制标题

DOI:
10.1073/pnas.1703764114
复制
发表时间:
2017-08-08
影响因子:
11.1
通讯作者:
Battey, H. S.
Battey, H. S.
中科院分区:
综合性期刊1区
文献类型:
--
作者:
Cox, D. R.;Battey, H. S.

文献摘要

被引文献

相似文献

研究个体数量相对较少,但具有大量潜在解释特征的数据尤其出现在基因组学中,但绝不仅仅出现在基因组学中。一种强有力的分析方法,套索[Tibshirani R(1996)J Roy Stat Soc B 58:267-288],考虑了假设的效应稀疏性,即大多数特征是无效的。模型拟合的标准准则,如最小二乘法,通过对所使用的每个解释变量施加惩罚来修改。结果是一个单一的模型,留下了其他稀疏的解释性特征选择几乎同样适合的可能性。本文提出的方法旨在指定基本上同样有效的简单模型,将详细解释留给特定研究的细节。该方法取决于最初进行大量单独分析的能力,允许将每个解释性特征与许多其他此类特征结合起来进行评估。进一步的阶段允许评估更复杂的模式,如非线性和交互依赖性。该方法与80年前引入的所谓部分平衡不完全区组设计[Yates F(1936)J Agric Sci 26:424-455]具有形式上的相似之处,用于研究大规模植物育种试验。本文的重点是探索性分析,在理想化假设下获得的更正式的统计特性将单独报告。
Data with a relatively small number of study individuals and a very large number of potential explanatory features arise particularly, but by no means only, in genomics. A powerful method of analysis, the lasso [Tibshirani R (1996) J Roy Stat Soc B 58: 267-288], takes account of an assumed sparsity of effects, that is, that most of the features are nugatory. Standard criteria for model fitting, such as the method of least squares, are modified by imposing a penalty for each explanatory variable used. There results a single model, leaving open the possibility that other sparse choices of explanatory features fit virtually equally well. The method suggested in this paper aims to specify simple models that are essentially equally effective, leaving detailed interpretation to the specifics of the particular study. The method hinges on the ability to make initially a very large number of separate analyses, allowing each explanatory feature to be assessed in combination with many other such features. Further stages allow the assessment of more complex patterns such as nonlinear and interactive dependences. The method has formal similarities to so-called partially balanced incomplete block designs introduced 80 years ago [Yates F (1936) J Agric Sci 26: 424-455] for the study of large-scale plant breeding trials. The emphasis in this paper is strongly on exploratory analysis; the more formal statistical properties obtained under idealized assumptions will be reported separately.