Design of experiments for fine-mapping quantitative trait loci in livestock populations

Design of experiments for fine-mapping quantitative trait loci in livestock populations
复制标题

DOI:
10.1186/s12863-020-00871-1
复制
发表时间:
2020-06-29
期刊:
影响因子:
2.9
通讯作者:
Reyer, Henry
Reyer, Henry
中科院分区:
生物学3区
文献类型:
--
作者:
Wittenburg, Doerte;Bonk, Sarah;Reyer, Henry

文献摘要

被引文献

相似文献

背景单核苷酸多态(SNPs)可以通过全基因组关联研究来识别,它捕捉到了对一个性状的重大影响。SNPs之间的高度连锁不平衡(LD)使正确识别致病变异变得困难。因此,报告的往往是目标区域,而不是单个SNPs。样本量不仅对参数估计的精度有至关重要的影响,而且还能确保达到期望的统计能力水平。我们研究了在这样一个靶区精细定位数量性状基因座信号的实验设计。方法多基因座模型可以同时识别致病变异,更准确地描述它们的位置,并解释现有的相关性。基于常用的SNP-BLUP方法,我们确定了局部检验非零SNP效应的z-Score统计量,并考察了其在另一种假设下的分布。这一数量使用的是SNPs之间的理论相关性,而不是观察到的相关性;对于任何给定的人口结构,它可以设置为父亲和母亲的LD的函数。结果我们模拟了多个父系半同胞家系,并考虑了1个MBP的靶区。观察到估计样本量的双峰分布,特别是当假设两个以上的原因变量时。估计的中位数构成了最优样本量的最终建议;它始终少于作为基线方法的单一SNP调查所估计的样本量。第二种模式指向样本大小膨胀,可以用不同连锁阶段的区块来解释,从而导致SNP之间的负相关。最优样本量几乎与待识别信号的数量呈线性增加,但更依赖于遗传力的假设。例如,如果遗传度为0.1,则需要的样本数量是0.3的三倍。提供了一个包含所有必需工具的R包。结论我们的方法将有关种群结构的信息纳入到实验设计中。与传统方法相比,这导致减少了样本大小的估计,从而能够节省资源设计未来的实验,以精细绘制候选变异。
Background Single nucleotide polymorphisms (SNPs) which capture a significant impact on a trait can be identified with genome-wide association studies. High linkage disequilibrium (LD) among SNPs makes it difficult to identify causative variants correctly. Thus, often target regions instead of single SNPs are reported. Sample size has not only a crucial impact on the precision of parameter estimates, it also ensures that a desired level of statistical power can be reached. We study the design of experiments for fine-mapping of signals of a quantitative trait locus in such a target region. Methods A multi-locus model allows to identify causative variants simultaneously, to state their positions more precisely and to account for existing dependencies. Based on the commonly applied SNP-BLUP approach, we determine the z-score statistic for locally testing non-zero SNP effects and investigate its distribution under the alternative hypothesis. This quantity employs the theoretical instead of observed dependence between SNPs; it can be set up as a function of paternal and maternal LD for any given population structure. Results We simulated multiple paternal half-sib families and considered a target region of 1 Mbp. A bimodal distribution of estimated sample size was observed, particularly if more than two causative variants were assumed. The median of estimates constituted the final proposal of optimal sample size; it was consistently less than sample size estimated from single-SNP investigation which was used as a baseline approach. The second mode pointed to inflated sample sizes and could be explained by blocks of varying linkage phases leading to negative correlations between SNPs. Optimal sample size increased almost linearly with number of signals to be identified but depended much stronger on the assumption on heritability. For instance, three times as many samples were required if heritability was 0.1 compared to 0.3. An R package is provided that comprises all required tools. Conclusions Our approach incorporates information about the population structure into the design of experiments. Compared to a conventional method, this leads to a reduced estimate of sample size enabling the resource-saving design of future experiments for fine-mapping of candidate variants.