SkewIT: The Skew Index Test for large-scale GC Skew analysis of bacterial genomes.
SkewIT: The Skew Index Test for large-scale GC Skew analysis of bacterial genomes.
复制标题
DOI:
10.1371/journal.pcbi.1008439
复制
发表时间:
2020-12
影响因子:
4.3
通讯作者:
Salzberg SL
中科院分区:
文献类型:
--
作者:
Lu J;Salzberg SL
GC skew is a phenomenon observed in many bacterial genomes, wherein the two replication strands of the same chromosome contain different proportions of guanine and cytosine nucleotides. Here we demonstrate that this phenomenon, which was first discovered in the mid-1990s, can be used today as an analysis tool for the 15,000+ complete bacterial genomes in NCBI’s Refseq library. In order to analyze all 15,000+ genomes, we introduce a new method, SkewIT (Skew Index Test), that calculates a single metric representing the degree of GC skew for a genome. Using this metric, we demonstrate how GC skew patterns are conserved within certain bacterial phyla, e.g. Firmicutes, but show different patterns in other phylogenetic groups such as Actinobacteria. We also discovered that outlier values of SkewIT highlight potential bacterial mis-assemblies. Using our newly defined metric, we identify multiple mis-assembled chromosomal sequences in previously published complete bacterial genomes. We provide a SkewIT web app https://jenniferlu717.shinyapps.io/SkewIT/ that calculates SkewI for any user-provided bacterial sequence. The web app also provides an interactive interface for the data generated in this paper, allowing users to further investigate the SkewI values and thresholds of the Refseq-97 complete bacterial genomes. Individual scripts for analysis of bacterial genomes are provided in the following repository: https://github.com/jenniferlu717/SkewIT. Even though every guanine (G) is paired with a cytosine (C) in double-stranded DNA molecules, bacterial genomes have more G’s than C’s when we focus only on a single strand in the direction of replication, called the leading strand. This phenomenon, called GC skew, is so ubiquitous that it has been used reliably to identify the replication origin (the location from which DNA begins the process of replicating itself) in thousands of bacteria. Here we describe a new method that automatically captures the “skewness” of a genome by finding the origin and terminus of replication, and then reporting skewness as a single number. We calculated this value for over 15,000 genomes, and found that most phylogenetic groups have a characteristic amount of skewness. We also observed that an unusually low value for skewness sometimes indicated that the genome was incorrectly assembled. To assist others in this type of analysis, we developed a graphical tool to compute and display GC-skew for any genome of interest.
登录
查看更多内容
影响因子:
48
作者:
Langmead, Ben;Salzberg, Steven L.
通讯作者:
Salzberg, Steven L.
影响因子:
3.6
作者:
Picardeau, M;Lobry, JR;Hinnebusch, BJ
通讯作者:
Hinnebusch, BJ
DOI:
10.1073/pnas.59.2.598
发表时间:
1968-01-01
影响因子:
11.1
作者:
OKAZAKI, R;OKAZAKI, T;SUGINO, A
通讯作者:
SUGINO, A
影响因子:
3.5
作者:
Frank, AC;Lobry, JR
通讯作者:
Lobry, JR
影响因子:
7
作者:
Breitwieser, Florian P.;Pertea, Mihaela;Salzberg, Steven L.
通讯作者:
Salzberg, Steven L.