A method for detecting IBD regions simultaneously in multiple individuals-with applications to disease genetics

A method for detecting IBD regions simultaneously in multiple individuals-with applications to disease genetics
复制标题

DOI:
10.1101/gr.115360.110
复制
发表时间:
2011-07-01
期刊:
影响因子:
7
通讯作者:
Nielsen, Rasmus
Nielsen, Rasmus
中科院分区:
生物学1区
文献类型:
--
作者:
Moltke, Ida;Albrechtsen, Anders;Nielsen, Rasmus

文献摘要

被引文献

相似文献

如果追溯到足够长的时间,有限种群中的所有个体都是相关的,因此将共享其基因组中相同的血统区域(IBD)。检测这些区域有几个重要的应用——从回答有关人类进化的问题到定位人类基因组中包含致病变异的区域。然而,IBD 区域可能难以检测,特别是在没有谱系信息的常见情况下。特别是,所有现有的非基于谱系的方法只能推断两个个体之间共享 IBD。在这里,我们提出了一种新的马尔可夫链蒙特卡罗方法来检测 IBD 区域,该方法不依赖于任何谱系信息。它基于适用于非定相 SNP 数据的概率模型。它可以考虑近亲繁殖、等位基因频率、基因分型错误和基因组距离。最重要的是,它可以同时推断多个个体之间的 IBD 共享情况。通过模拟,我们表明,多个个体的同时建模使得该方法比其他几种非基于谱系的方法更加强大和准确。我们通过将其应用于乳腺癌和/或卵巢癌个体的数据来说明该方法的潜力,并表明使用来自仅五个看似无关的受影响个体的 SNP 数据可以将已知的致病突变映射到 2.2 Mb 的区域。使用经典的链接映射或关联映射是不可能实现这一点的。
All individuals in a finite population are related if traced back long enough and will, therefore, share regions of their genomes identical by descent (IBD). Detection of such regions has several important applications-from answering questions about human evolution to locating regions in the human genome containing disease-causing variants. However, IBD regions can be difficult to detect, especially in the common case where no pedigree information is available. In particular, all existing non-pedigree based methods can only infer IBD sharing between two individuals. Here, we present a new Markov Chain Monte Carlo method for detection of IBD regions, which does not rely on any pedigree information. It is based on a probabilistic model applicable to unphased SNP data. It can take inbreeding, allele frequencies, genotyping errors, and genomic distances into account. And most importantly, it can simultaneously infer IBD sharing among multiple individuals. Through simulations, we show that the simultaneous modeling of multiple individuals makes the method more powerful and accurate than several other non-pedigree based methods. We illustrate the potential of the method by applying it to data from individuals with breast and/or ovarian cancer, and show that a known disease-causing mutation can be mapped to a 2.2-Mb region using SNP data from only five seemingly unrelated affected individuals. This would not be possible using classical linkage mapping or association mapping.