The UCSC repeat browser allows discovery and visualization of evolutionary conflict across repeat families

The UCSC repeat browser allows discovery and visualization of evolutionary conflict across repeat families
复制标题

DOI:
10.1186/s13100-020-00208-w
复制
发表时间:
2020-03-31
期刊:
影响因子:
4.9
通讯作者:
Haeussler, Maximilian
Haeussler, Maximilian
中科院分区:
生物学3区
文献类型:
--
作者:
Fernandes, Jason D.;Zamudio-Hurtado, Armando;Haeussler, Maximilian

文献摘要

被引文献

相似文献

背景资料:近一半的人类基因组由重复元件组成,其中大多数是反转录转座子,其中许多起着重要的生物学作用。然而,重复元件对当前的生物信息学分析和可视化工具提出了几个独特的挑战,因为短重复序列可以映射到多个基因组位点,导致它们的错误分类和误解。事实上,映射到重复元素的序列数据经常被从分析管道中丢弃。因此,有一个持续的需要标准化的工具和技术来解释基因组数据的repeat.Results:我们提出了UCSC重复浏览器,它由一套完整的人类重复参考序列来自注释的常用程序RepeatMasker。UCSC Repeat Browser还提供了从人类基因组到这些参考的比对,使用它来映射标准人类基因组注释轨迹,并将所有这些作为一个综合界面来提供,以方便重复元素的工作。它还提供了重复序列社区特别感兴趣的多个公开数据集的处理轨迹,包括KRAB锌指蛋白(KZNF)的ChIP-seq数据集-已知结合和抑制某些类型的重复序列的蛋白质家族。我们使用UCSC RepeatBrowser结合这些数据集,以及几种非人类灵长类动物的RepeatMasker注释,来追踪LINE 1 retroelements及其阻遏物之间物种特异性进化斗争的独立轨迹。此外,我们的文件在研究人员如何可以映射自己的人类基因组注释这些reference repeat sequences.Conclusions:UCSC重复浏览器允许简单直观的可视化基因组数据的共识重复元素,规避多映射的问题,其中重复元素的测序读取映射到人类基因组上的多个位置。通过开发参考共识,多个数据集和注释轨迹可以很容易地覆盖,以在单个交互式窗口中揭示重复序列的复杂进化历史。具体来说,我们使用这种方法追溯了几种灵长类动物特定LINE-1家族在猿类中的历史,并发现了几种与KZNF的出现和结合相关的物种特异性进化途径。
Background: Nearly half the human genome consists of repeat elements, most of which are retrotransposons, and many of which play important biological roles. However repeat elements pose several unique challenges to current bioinformatic analyses and visualization tools, as short repeat sequences can map to multiple genomic loci resulting in their misclassification and misinterpretation. In fact, sequence data mapping to repeat elements are often discarded from analysis pipelines. Therefore, there is a continued need for standardized tools and techniques to interpret genomic data of repeats.Results: We present the UCSC Repeat Browser, which consists of a complete set of human repeat reference sequences derived from annotations made by the commonly used program RepeatMasker. The UCSC Repeat Browser also provides an alignment from the human genome to these references, uses it to map the standard human genome annotation tracks, and presents all of them as a comprehensive interface to facilitate work with repetitive elements. It also provides processed tracks of multiple publicly available datasets of particular interest to the repeat community, including ChIP-seq datasets for KRAB Zinc Finger Proteins (KZNFs) - a family of proteins known to bind and repress certain classes of repeats. We used the UCSC Repeat Browser in combination with these datasets, as well as RepeatMasker annotations in several non-human primates, to trace the independent trajectories of species-specific evolutionary battles between LINE 1 retroelements and their repressors. Furthermore, we document at how researchers can map their own human genome annotations to these reference repeat sequences.Conclusions: The UCSC Repeat Browser allows easy and intuitive visualization of genomic data on consensus repeat elements, circumventing the problem of multi-mapping, in which sequencing reads of repeat elements map to multiple locations on the human genome. By developing a reference consensus, multiple datasets and annotation tracks can easily be overlaid to reveal complex evolutionary histories of repeats in a single interactive window. Specifically, we use this approach to retrace the history of several primate specific LINE-1 families across apes, and discover several species-specific routes of evolution that correlate with the emergence and binding of KZNFs.