Parallel computation of genome-scale RNA secondary structure to detect structural constraints on human genome.

Parallel computation of genome-scale RNA secondary structure to detect structural constraints on human genome.
复制标题

DOI:
10.1186/s12859-016-1067-9
复制
发表时间:
2016-05-06
期刊:
影响因子:
3
通讯作者:
Kiryu H
Kiryu H
中科院分区:
生物学4区
文献类型:
--
作者:
Kawaguchi R;Kiryu H

文献摘要

被引文献

相似文献

已知剪接位点周围的RNA二级结构通过促进剪接体识别来帮助正常剪接。然而,到目前为止,由于实验和计算的严重限制,如阅读覆盖率低和数值问题,分析整个内含子区域或前mRNA序列的结构特性一直是困难的。我们的新型软件ParasoR可以在计算机集群上运行,能够在最大碱基配对距离的约束下精确计算长RNA序列的各种结构特征。ParasoR将动态规划(DP)矩阵分成更小的片段,使得每个片段可以由单独的计算机节点计算,而不会丢失片段之间的连接信息。ParasoR直接计算DP变量的比值,以避免由于大量的Boltzmann因子被抵消而导致的数值精度的降低。ParasoR计算的mRNAs结构偏好与高通量测序分析确定的结构偏好高度一致。使用ParasoR,我们研究了人类基因组中转录区域的全球结构偏好。全基因组折叠模拟表明,在去除重复序列和k-mer频率偏差后,转录区域明显比基因间区更具结构性。特别是,我们观察到与其反义序列以及基因间隔区相比,碱基配对对整个内含子区域具有非常显著的偏好。前mRNAs和mRNAs之间的比较表明,剪接后编码区更容易接近,这表明翻译效率受到限制。这种变化与基因表达水平以及GC含量相关,并在与细胞骨架和激酶功能相关的基因中得到丰富。我们已经证明,ParasoR对于分析长RNA序列的结构特性非常有用,例如mRNAs、前mRNAs和长的非编码RNA,这些RNA的长度可以超过一百万个碱基。在我们的分析中,包括内含子在内的转录区域被表明受到各种类型的结构限制,这些限制不能用简单的序列组成偏差来解释。ParasoR可在https://github.com/carushi/ParasoR.上免费获得本文的在线版本(doi:10.1186/s12859-0161067-9)包含补充材料,授权用户可以使用。
RNA secondary structure around splice sites is known to assist normal splicing by promoting spliceosome recognition. However, analyzing the structural properties of entire intronic regions or pre-mRNA sequences has been difficult hitherto, owing to serious experimental and computational limitations, such as low read coverage and numerical problems. Our novel software, “ParasoR”, is designed to run on a computer cluster and enables the exact computation of various structural features of long RNA sequences under the constraint of maximal base-pairing distance. ParasoR divides dynamic programming (DP) matrices into smaller pieces, such that each piece can be computed by a separate computer node without losing the connectivity information between the pieces. ParasoR directly computes the ratios of DP variables to avoid the reduction of numerical precision caused by the cancellation of a large number of Boltzmann factors. The structural preferences of mRNAs computed by ParasoR shows a high concordance with those determined by high-throughput sequencing analyses. Using ParasoR, we investigated the global structural preferences of transcribed regions in the human genome. A genome-wide folding simulation indicated that transcribed regions are significantly more structural than intergenic regions after removing repeat sequences and k-mer frequency bias. In particular, we observed a highly significant preference for base pairing over entire intronic regions as compared to their antisense sequences, as well as to intergenic regions. A comparison between pre-mRNAs and mRNAs showed that coding regions become more accessible after splicing, indicating constraints for translational efficiency. Such changes are correlated with gene expression levels, as well as GC content, and are enriched among genes associated with cytoskeleton and kinase functions. We have shown that ParasoR is very useful for analyzing the structural properties of long RNA sequences such as mRNAs, pre-mRNAs, and long non-coding RNAs whose lengths can be more than a million bases in the human genome. In our analyses, transcribed regions including introns are indicated to be subject to various types of structural constraints that cannot be explained from simple sequence composition biases. ParasoR is freely available at https://github.com/carushi/ParasoR. The online version of this article (doi:10.1186/s12859-016-1067-9) contains supplementary material, which is available to authorized users.