Fast and accurate shared segment detection and relatedness estimation in un-phased genetic data using TRUFFLE

Fast and accurate shared segment detection and relatedness estimation in un-phased genetic data using TRUFFLE
复制标题

使用 TRUFFLE 在非定相遗传数据中快速准确地检测共享片段并进行相关性估计

DOI:
--
复制
发表时间:
2018
期刊:
bioRxiv
影响因子:
--
通讯作者:
Lei Sun
Lei Sun
中科院分区:
--
文献类型:
--
作者:
A. Dimitromanolakis;A. Paterson;Lei Sun

文献摘要

参考文献

被引文献

相似文献

个体间的亲缘关系估计和片段检测是疾病基因定位的一个重要方面。现有的方法要么是为计算效率而定制的,要么需要定相以提高精度。我们开发了TRUFFLE,这是一种集成了计算技术和统计原理的方法,用于使用非阶段性数据识别和可视化身份下降(IBD)片段。通过跳过单倍型定相步骤,而是依赖于更简单的基于区域的方法,我们的方法在保持推理准确性的同时具有计算效率。此外,误差模型校正由于基因分型误差而发生的片段断裂。TRUFFLE可以在一台典型的笔记本电脑上在几分钟内从1000个基因组项目的数据中估计出310万对的相关性。与预期一致,我们在不同人群中发现了三对远房表亲或更近的配对,而常用的方法发现了超过15,000对这样的配对。同样,在人群中,我们发现了更少的相关配对。TRUFFLE以依赖于阶段数据的方法为基准,具有良好的准确性,但速度要快得多。我们还确定了特定的局部基因组区域,这些区域通常在种群中共享,这表明选择。当应用于系谱数据时,我们观察到在检测1至5度关系时的准确率为99.7%。随着基因组数据集变得越来越大,TRUFFLE可以通过准确的IBD片段检测,通过隐含的共享单倍型进行疾病基因定位。
Relationship estimation and segment detection between individuals is an important aspect of disease gene mapping. Existing methods are either tailored for computational efficiency, or require phasing to improve accuracy. We developed TRUFFLE, a method that integrates computational techniques and statistical principles for the identification and visualization of identity-by-descent (IBD) segments using un-phased data. By skipping the haplotype phasing step and, instead, relying on a simpler region-based approach, our method is computationally efficient while maintaining inferential accuracy. In addition, an error model corrects for segment break-ups that occur as a consequence of genotyping errors. TRUFFLE can estimate relatedness for 3.1 million pairs from the 1000 Genomes Project data in a few minutes on a typical laptop computer. Consistent with expectation, we identified three second cousin or closer pairs across different populations, while commonly used methods identified over 15,000 such pairs. Similarly, within populations, we identified much fewer related pairs. Benchmarking to methods relying on phased data, TRUFFLE has a favorable accuracy profile but is drastically faster. We also identified specific local genomic regions that are commonly shared within populations, suggesting selection. When applied to pedigree data, we observed 99.7% accuracy in detecting 1st to 5th degree relationships. As genomic datasets become much larger, TRUFFLE can enable disease gene mapping through implicit shared haplotypes by accurate IBD segment detection.
DOI: 10.1101/gr.115972.110
发表时间: 2011-05-01
期刊: GENOME RESEARCH
影响因子: 7
作者:
Huff, Chad D.;Witherspoon, David J.;Jorde, Lynn B.
通讯作者: Jorde, Lynn B.
DOI: 10.1093/molbev/msr133
发表时间: 2012-02-01
影响因子: 10.7
作者:
Gusev, Alexander;Palamara, Pier Francesco;Pe'er, Itsik
通讯作者: Pe'er, Itsik
DOI: 10.1016/j.ajhg.2010.02.021
发表时间: 2010-04-09
影响因子: 9.8
作者:
Browning, Sharon R.;Browning, Brian L.
通讯作者: Browning, Brian L.
DOI: 10.1086/302800
发表时间: 2000-03-01
影响因子: 9.8
作者:
McPeek, MS;Sun, L
通讯作者: Sun, L
DOI: 10.1016/j.ajhg.2011.01.010
发表时间: 2011-02-11
影响因子: 9.8
作者:
Browning, Brian L.;Browning, Sharon R.
通讯作者: Browning, Sharon R.