Identifying structural variants using linked-read sequencing data

Identifying structural variants using linked-read sequencing data
复制标题

DOI:
10.1093/bioinformatics/btx712
复制
发表时间:
2018-01-15
期刊:
影响因子:
5.8
通讯作者:
Raphael, Benjamin J.
Raphael, Benjamin J.
中科院分区:
生物学3区
文献类型:
--
作者:
Elyanow, Rebecca;Wu, Hsin-Ta;Raphael, Benjamin J.

文献摘要

被引文献

相似文献

动机:结构变异,包括大缺失、重复、倒位、易位和其他重排,在人类和癌症基因组中很常见。已经开发了许多方法来从Illumina短读段测序数据中鉴定结构变体。然而,结构变体的可靠鉴定仍然具有挑战性,因为许多变体在基因组的重复区域中具有断点,因此难以用短读段鉴定。最近开发的来自10X Genomics的连接阅读测序技术将新型条形码策略与Illumina测序相结合。该技术用相同的分子条形码标记源自少量(类似于5至10个)长度类似于50 Kbp的DNA分子的所有读段。这些条形码读段包含有利于识别结构变体的远程序列信息。结果:我们提出了新型条形码读段邻接识别(NAIBR),这是一种识别连接读段测序数据中结构变体的算法。NAIBR使用概率模型预测由结构变异引起的个体基因组中的新邻接,该概率模型在条形码读取中组合了多个信号。我们表明,NAIBR优于几种现有的结构变异识别方法,包括最近的两种方法,也分析了连接读取模拟测序数据和10倍全基因组测序数据从NA12878人类基因组和HCC 1954乳腺癌细胞系。在HCC 1954中鉴定的几种新的体细胞结构变体与已知的癌症基因重叠。
Motivation: Structural variation, including large deletions, duplications, inversions, translocations and other rearrangements, is common in human and cancer genomes. A number of methods have been developed to identify structural variants from Illumina short-read sequencing data. However, reliable identification of structural variants remains challenging because many variants have breakpoints in repetitive regions of the genome and thus are difficult to identify with short reads. The recently developed linked-read sequencing technology from 10X Genomics combines a novel bar-coding strategy with Illumina sequencing. This technology labels all reads that originate from a small number (similar to 5 to 10) DNA molecules similar to 50 Kbp in length with the same molecular barcode. These barcoded reads contain long-range sequence information that is advantageous for identification of structural variants.Results: We present Novel Adjacency Identification with Barcoded Reads (NAIBR), an algorithm to identify structural variants in linked-read sequencing data. NAIBR predicts novel adjacencies in an individual genome resulting from structural variants using a probabilistic model that combines multiple signals in barcoded reads. We show that NAIBR outperforms several existing methods for structural variant identification-including two recent methods that also analyze linked-reads-on simulated sequencing data and 10X whole-genome sequencing data from the NA12878 human genome and the HCC1954 breast cancer cell line. Several of the novel somatic structural variants identified in HCC1954 overlap known cancer genes.