Selection-adjusted inference: an application to confidence intervals for cis-eQTL effect sizes

Selection-adjusted inference: an application to confidence intervals for cis-eQTL effect sizes
复制标题

DOI:
10.1093/biostatistics/kxz024
复制
发表时间:
2021-01-01
期刊:
影响因子:
2.1
通讯作者:
Sabatti, Chiara
Sabatti, Chiara
中科院分区:
数学2区
文献类型:
--
作者:
Panigrahi, Snigdha;Zhu, Junjie;Sabatti, Chiara

文献摘要

被引文献

相似文献

表达数量性状位点 (eQTL) 研究的目标是识别影响生物体基因表达水平的遗传变异。高通量技术使此类研究成为可能:在给定的组织样本中,它使我们能够量化大约 20 000 个基因的表达水平,并记录数百万个遗传多态性中存在的等位基因。虽然一旦有样本,获取这些数据的成本相对较低,但获取人体组织仍然是一项成本高昂的工作:eQTL 研究仍然基于相对较小的样本量,这种限制对于脑、肝脏等组织(通常是具有最直接医学相关性的器官)尤其严重。鉴于这些数据集的高维性质和测试的大量假设,科学界很早就采用了多重性调整程序。这些测试程序主要控制识别影响表达水平的遗传变异的错误发现率。相比之下,迄今为止尚未受到太多关注的一个问题是,以考虑大量选择的方式提供与这些变体相关的效应大小的估计。然而,考虑到获取更多样品的困难,这一挑战具有实际意义。我们在这项工作中说明了如何部署最近开发的条件推理方法来获得具有可靠覆盖范围的 eQTL 效应大小的置信区间。我们提出的程序基于具有 2 倍贡献的随机分层策略:(1)它反映了最先进的研究中通常采用的选择步骤;(2)它引入了随机性的使用而不是数据分割,以最大限度地利用可用数据。 GTEx 肝脏数据集 (v6) 的分析表明,单纯获得的置信区间可能无法涵盖效应大小的真实值,并且影响基因表达水平的局部遗传多态性的数量可能被低估。
The goal of expression quantitative trait loci (eQTL) studies is to identify the genetic variants that influence the expression levels of the genes in an organism. High throughput technology has made such studies possible: in a given tissue sample, it enables us to quantify the expression levels of approximately 20 000 genes and to record the alleles present at millions of genetic polymorphisms. While obtaining this data is relatively cheap once a specimen is at hand, obtaining human tissue remains a costly endeavor: eQTL studies continue to be based on relatively small sample sizes, with this limitation particularly serious for tissues as brain, liver, etc.-often the organs of most immediate medical relevance. Given the high-dimensional nature of these datasets and the large number of hypotheses tested, the scientific community has adopted early on multiplicity adjustment procedures. These testing procedures primarily control the false discoveries rate for the identification of genetic variants with influence on the expression levels. In contrast, a problem that has not received much attention to date is that of providing estimates of the effect sizes associated with these variants, in a way that accounts for the considerable amount of selection. Yet, given the difficulty of procuring additional samples, this challenge is of practical importance. We illustrate in this work how the recently developed conditional inference approach can be deployed to obtain confidence intervals for the eQTL effect sizes with reliable coverage. The procedure we propose is based on a randomized hierarchical strategy with a 2-fold contribution: (1) it reflects the selection steps typically adopted in state of the art investigations and (2) it introduces the use of randomness instead of data-splitting to maximize the use of available data. Analysis of the GTEx Liver dataset (v6) suggests that naively obtained confidence intervals would likely not cover the true values of effect sizes and that the number of local genetic polymorphisms influencing the expression level of genes might be underestimated.