Finding noncoding RNA transcripts from low abundance expressed sequence tags

Finding noncoding RNA transcripts from low abundance expressed sequence tags
复制标题

从低丰度表达序列标签中寻找非编码 RNA 转录本

DOI:
10.1038/cr.2008.59
复制
发表时间:
2008-06-01
期刊:
影响因子:
44.1
通讯作者:
Li, Fei
Li, Fei
中科院分区:
生物学1区
文献类型:
--
作者:
Xue, Chenghai;Li, Fei;Li, Fei

文献摘要

被引文献

相似文献

非编码RNA(noncoding RNA,ncRNA)基因的数量比预期的要多得多。然而,它仍然是一个困难的任务,以确定ncRNA的计算算法或生物实验。最近的报告表明,ncRNA也可能出现在表达序列标签(EST)数据库中。然而,基因间EST很少受到关注,并且由于其丰度低而注释不足。在这里,我们已经开发了一个计算策略,从人类EST中发现ncRNA基因。我们首先收集了位于基因间区域并且没有详细注释的EST。将基因间区域分成非重叠的50-nt窗口,并将从UCSC数据库获得的PhastCons评分分配给这些窗口。我们保留了PhastCons得分超过0.8的保守窗口,并且至少有三个支持EST作为种子。将对应于种子的每个EST簇组装成长重叠群。我们使用两个标准从这些重叠群中筛选ncRNA转录本:第一个是预测的最长开放阅读框小于300 nt;第二个是可能的Pol-II启动子存在于重叠群上游或下游2 000 nt内。结果,从人类低丰度EST中鉴定出118个新的ncRNA基因。在7个随机选择的候选者中,如RT-PCR所示,6个在人2BS细胞中转录。我们的工作证明了EST是发现新的ncRNA基因的“宝藏”。
It has been proved that noncoding RNA (ncRNA) genes are much more numerous than expected. However, it remains a difficult task to identify ncRNAs with either computational algorithms or biological experiments. Recent reports have suggested that ncRNAs may also appear in the expressed sequence tags (EST's) database. Nevertheless, intergenic ESTs have received little attention and are poorly annotated owing to their low abundance. Here, we have developed a computational strategy for discovering ncRNA genes from human ESTs. We first collected ESTs that are located in the intergenic regions and do not have detailed annotations. The intergenic regions were divided into non-overlapping 50-nt windows and PhastCons scores obtained from the UCSC database were assigned to these windows. We kept conserved windows that had PhastCons scores of over 0.8 and that had at least three supporting ESTs to act as seeds. Each cluster of ESTs corresponding to the seeds was assembled into a long contig. We used two criteria to screen for ncRNA transcripts from these contigs: the first was that the longest predicted open reading frame was less than 300 nt and the second was that the likely Pol-II promoters exist within 2 000 nt upstream or downstream of the contigs. As a result, 118 novel ncRNA genes were identified from human low abundance ESTs. Of seven randomly selected candidates, six were transcribed in human 2BS cells as shown by RT-PCR. Our work proves that the EST is a'hidden treasure'for detecting novel ncRNA genes.