An efficient study design to test parent-of-origin effects in family trios.

An efficient study design to test parent-of-origin effects in family trios.
复制标题

一种有效的研究设计,用于测试家庭三人组中的父母效应。

DOI:
10.1002/gepi.22060
复制
发表时间:
2017
影响因子:
2.1
通讯作者:
Feng,Rui
Feng,Rui
中科院分区:
医学4区
文献类型:
--
作者:
Yu,Xiaobo;Chen,Gao;Feng,Rui

文献摘要

相似文献

越来越多的证据表明,基因可能导致产前、新生儿和儿科疾病,这取决于它们的父母起源。纳入亲本起源效应(poe)的统计模型可以提高检测疾病相关基因的能力,并有助于解释疾病缺失的遗传性。在许多研究中,对儿童进行全基因组关联检测测序。但是,对他们的父母进行测序和评估poe可能变得负担不起。基于现实情况,我们提出了一个预算友好的研究设计,对儿童进行测序,仅通过单核苷酸多态性阵列对其父母进行基因分型。我们开发了一种强大的基于似然的方法,该方法考虑了序列读取和连锁不平衡来推断儿童等位基因的亲本起源,并估计其结果的poe。通过广泛的模拟,我们评估了我们提出的方法的性能,并将其与仅使用基因型的现有方法进行了比较。我们的方法显示出比基于基因型的方法更高的功效。当平均读取深度或对端长度相当大时,我们的方法达到了理想的功率。当无法获得单亲父母的基因型或测试位点上的亲本基因型时,两种方法都比获得完整数据时无效;但我们方法的功率损失小于基于基因型的方法。我们还扩展了我们的方法,以适应来自儿童及其父母的混合基因型、低覆盖率和高覆盖率的序列数据。在存在序列错误的情况下,低覆盖率亲本序列数据可能导致比亲本基因型数据更低的功效。
Increasing evidence has shown that genes may cause prenatal, neonatal, and pediatric diseases depending on their parental origins. Statistical models that incorporate parent‐of‐origin effects (POEs) can improve the power of detecting disease‐associated genes and help explain the missing heritability of diseases. In many studies, children have been sequenced for genome‐wide association testing. But it may become unaffordable to sequence their parents and evaluate POEs. Motivated by the reality, we proposed a budget‐friendly study design of sequencing children and only genotyping their parents through single nucleotide polymorphism array. We developed a powerful likelihood‐based method, which takes into account both sequence reads and linkage disequilibrium to infer the parental origins of children's alleles and estimate their POEs on the outcome. We evaluated the performance of our proposed method and compared it with an existing method using only genotypes, through extensive simulations. Our method showed higher power than the genotype‐based method. When either the mean read depth or the pair‐end length was reasonably large, our method achieved ideal power. When single parents’ genotypes were unavailable or parental genotypes at the testing locus were not typed, both methods lost power compared with when complete data were available; but the power loss from our method was smaller than the genotype‐based method. We also extended our method to accommodate mixed genotype, low‐, and high‐coverage sequence data from children and their parents. At presence of sequence errors, low‐coverage parental sequence data may lead to lower power than parental genotype data.