CLEVER: clique-enumerating variant finder

CLEVER: clique-enumerating variant finder
复制标题

DOI:
10.1093/bioinformatics/bts566
复制
发表时间:
2012-11-15
期刊:
影响因子:
5.8
通讯作者:
Schonhuth, Alexander
Schonhuth, Alexander
中科院分区:
生物学3区
文献类型:
--
作者:
Marschall, Tobias;Costa, Ivan G.;Schonhuth, Alexander

文献摘要

被引文献

相似文献

动机:下一代测序技术促进了人类遗传变异的大规模分析。尽管测序速度取得了进步,但结构变异的计算发现尚未成为标准。在大多数测序个体中,许多变异很可能仍未被发现。 结果:在这里,我们提出了一种新颖的基于内部片段大小的方法,该方法将所有读段(包括一致读段)组织到读段比对图中,其中最大派代表最大无矛盾比对组。然后,一种新颖的算法会枚举所有最大派系,并统计评估它们反映插入或删除的潜力。我们在文献中首次使用来自完全注释的基因组的模拟 Illumina 读数来比较大量最先进的方法,并提供相关的性能统计数据。我们实现了卓越的性能,特别是对于长度为 20-100 nt 的删除或插入 (indel)。这之前已被认为是结构变异发现中剩余的主要挑战,特别是对于基于插入尺寸的方法。在这个尺寸范围内,我们甚至优于 split-read 对齐器。我们在生物数据上也取得了有竞争力的结果,我们的方法是唯一能够做出大量正确预测的方法,此外,这与分割读取对齐器的预测不相交。
Motivation: Next-generation sequencing techniques have facilitated a large-scale analysis of human genetic variation. Despite the advances in sequencing speed, the computational discovery of structural variants is not yet standard. It is likely that many variants have remained undiscovered in most sequenced individuals.Results: Here, we present a novel internal segment size based approach, which organizes all, including concordant, reads into a read alignment graph, where max-cliques represent maximal contradiction-free groups of alignments. A novel algorithm then enumerates all max-cliques and statistically evaluates them for their potential to reflect insertions or deletions. For the first time in the literature, we compare a large range of state-of-the-art approaches using simulated Illumina reads from a fully annotated genome and present relevant performance statistics. We achieve superior performance, in particular, for deletions or insertions (indels) of length 20-100 nt. This has been previously identified as a remaining major challenge in structural variation discovery, in particular, for insert size based approaches. In this size range, we even outperform split-read aligners. We achieve competitive results also on biological data, where our method is the only one to make a substantial amount of correct predictions, which, additionally, are disjoint from those by split-read aligners.