Conflation of Short Identity-by-Descent Segments Bias Their Inferred Length Distribution

Conflation of Short Identity-by-Descent Segments Bias Their Inferred Length Distribution
复制标题

DOI:
10.1534/g3.116.027581
复制
发表时间:
2016-05-01
影响因子:
2.6
通讯作者:
Novembre, John
Novembre, John
中科院分区:
生物学3区
文献类型:
--
作者:
Chiang, Charleston W. K.;Ralph, Peter;Novembre, John

文献摘要

被引文献

相似文献

血统认同(IBD)是遗传学中的一个基本概念,有着广泛的应用。在通常的定义中,如果一个IBD片段是从最近共享的共同祖先继承而来的,而没有干预重组,则两个单倍型被称为共享IBD片段。几厘米长的片段可以通过使用来自群体样本的高密度SNP阵列数据的许多算法有效地检测到,并且目前正在努力从测序中检测较短的片段。这里,我们研究可识别性问题:由于现有方法基于逐个状态身份的连续片段来检测IBD,推测IBD的长片段可能来自较小的、邻近的IBD片段的合并。我们使用合并模拟来量化这种影响,发现很大一部分1-2厘米长的推断片段是两个或更多较短片段的合并结果,每个片段至少0.2厘米或更长,在所有测试的程序中,现代人典型的人口场景下。这种合并的影响对于更长(>2厘米)的细分市场要小得多。这偏向了推断的IBD片段长度分布,因此可能会影响下游推断,这些推断依赖于IBD的每个片段来自单一共同祖先的假设。作为一个例子,我们提出并分析了一种使用IBD片段的从头突变率的估计器,并证明了未建模的合并导致对这些片段上共同祖先的年龄的低估,从而显著高估了突变率。详细了解合并效应将使其在未来方法中的修正更容易处理。
Identity-by-descent (IBD) is a fundamental concept in genetics with many applications. In a common definition, two haplotypes are said to share an IBD segment if that segment is inherited from a recent shared common ancestor without intervening recombination. Segments several cM long can be efficiently detected by a number of algorithms using high-density SNP array data from a population sample, and there are currently efforts to detect shorter segments from sequencing. Here, we study a problem of identifiability: because existing approaches detect IBD based on contiguous segments of identity-by-state, inferred long segments of IBD may arise from the conflation of smaller, nearby IBD segments. We quantified this effect using coalescent simulations, finding that significant proportions of inferred segments 1-2 cM long are results of conflations of two or more shorter segments, each at least 0.2 cM or longer, under demographic scenarios typical for modern humans for all programs tested. The impact of such conflation is much smaller for longer (> 2 cM) segments. This biases the inferred IBD segment length distribution, and so can affect downstream inferences that depend on the assumption that each segment of IBD derives from a single common ancestor. As an example, we present and analyze an estimator of the de novo mutation rate using IBD segments, and demonstrate that unmodeled conflation leads to underestimates of the ages of the common ancestors on these segments, and hence a significant overestimate of the mutation rate. Understanding the conflation effect in detail will make its correction in future methods more tractable.