Deep repeat resolutionthe assembly of the Drosophila Histone Complex

Deep repeat resolutionthe assembly of the Drosophila Histone Complex
复制标题

DOI:
10.1093/nar/gky1194
复制
发表时间:
2019-02-20
影响因子:
14.9
通讯作者:
Schloissnig, Siegfried
Schloissnig, Siegfried
中科院分区:
生物学2区
文献类型:
--
作者:
Bongartz, Philipp;Schloissnig, Siegfried

文献摘要

被引文献

相似文献

尽管长阅读测序技术的出现导致了从头基因组组装的邻接性的飞跃,但目前高等生物的参考基因组仍然不能提供完整的染色体序列。尽管读数超过3万个碱基对,但仍然有一些重复结构不能被当前最先进的组装者解决。这些结构中最具挑战性的是排列有序的重复序列,这种重复序列存在于所有真核生物的基因组中。解开串联重复簇是非常困难的,因为重复拷贝之间的罕见差异被长时间读取的高错误率所掩盖。解决这个问题将是计算完全组装的基因组的重要一步。在这里,我们通过果蝇组蛋白复合体的例子证明,通过机器学习算法,有可能从非常嘈杂的数据中利用重复序列的单核苷酸变体的潜在区分模式来解析大型且高度保守的重复序列簇。本文探索的思想是迈向复杂重复结构自动组装的第一步,并有望适用于广泛的真核基因组。
Though the advent of long-read sequencing technologies has led to a leap in contiguity of de novo genome assemblies, current reference genomes of higher organisms still do not provide unbroken sequences of complete chromosomes. Despite reads in excess of 30 000 base pairs, there are still repetitive structures that cannot be resolved by current state-of-the-art assemblers. The most challenging of these structures are tandemly arrayed repeats, which occur in the genomes of all eukaryotes. Untangling tandem repeat clusters is exceptionally difficult, since the rare differences between repeat copies are obscured by the high error rate of long reads. Solving this problem would constitute a major step towards computing fully assembled genomes. Here, we demonstrate by example of the Drosophila Histone Complex that via machine learning algorithms, it is possible to exploit the underlying distinguishing patterns of single nucleotide variants of repeats from very noisy data to resolve a large and highly conserved repeat cluster. The ideas explored in this paper are a first step towards the automated assembly of complex repeat structures and promise to be applicable to a wide range of eukaryotic genomes.