HOPPSIGEN: a database of human and mouse processed pseudogenes.

HOPPSIGEN: a database of human and mouse processed pseudogenes.
复制标题

DOI:
10.1093/nar/gki084
复制
发表时间:
2005-01-01
影响因子:
14.9
通讯作者:
Mouchiroud D
Mouchiroud D
中科院分区:
生物学2区
文献类型:
--
作者:
Khelifi A;Duret L;Mouchiroud D

文献摘要

参考文献

被引文献

相似文献

加工后的假基因由逆转录 mRNA 产生。一般来说,由于加工后的假基因缺乏启动子,它们从插入基因组的那一刻起就不再具有功能。随后,它们自由地累积替换、插入和删除。此外,使用其功能同源基因的序列可以很容易地推断出加工后的假基因的祖先结构。由于这些特征,处理后的假基因代表了研究基因组进化的良好中性标记。最近,人们对这些标记越来越感兴趣,特别是在基因组注释、功能基因组学和基因组进化分析(替代模式)领域中帮助基因预测。由于这些原因,我们开发了一种方法来注释完整基因组中经过处理的假基因。为了使它们对不同的研究领域有用,我们在对它们进行注释后将它们存储在核酸数据库中。在这项工作中,我们从 ENSEMBL 中筛选了小鼠和人类的完整基因组,以找到由带有内含子的功能基因生成的加工假基因。我们使用保守的方法来检测处理后的假基因,以尽量减少假阳性序列的发生率。在经过处理的假基因中,有些仍然具有保守的开放阅读框,有些具有重叠的基因位置。我们将所有逆转录序列指定为逆转录元件,更严格地说,我们将所有不属于前两个类别的逆转录元件指定为经过处理的假基因(具有保守的开放阅读或重叠的基因位置)。我们注释了人类基因组中的 5823 个逆转录元件(5206 个经过处理的假基因)和小鼠基因组中的 3934 个逆转录元件(3428 个经过处理的假基因)。与之前的估计相比,处理的假基因的总数被低估,但此过程的目的是生成高质量的数据集。为了便于在研究基因组结构和进化中使用加工后的假基因,加工后的假基因的 DNA 序列及其功能性逆转录同源物现在存储在核酸数据库 HOPPSIGEN 中。 HOPPSIGEN 可以在 PBIL (Pôle Bioinformatique Lyonnais) 万维网服务器 (http://pbil.univ-lyon1.fr/) 上浏览,也可以完全下载进行本地安装。
Processed pseudogenes result from reverse transcribed mRNAs. In general, because processed pseudogenes lack promoters, they are no longer functional from the moment they are inserted into the genome. Subsequently, they freely accumulate substitutions, insertions and deletions. Moreover, the ancestral structure of processed pseudogenes could be easily inferred using the sequence of their functional homologous genes. Owing to these characteristics, processed pseudogenes represent good neutral markers for studying genome evolution. Recently, there is an increasing interest for these markers, particularly to help gene prediction in the field of genome annotation, functional genomics and genome evolution analysis (patterns of substitution). For these reasons, we have developed a method to annotate processed pseudogenes in complete genomes. To make them useful to different fields of research, we stored them in a nucleic acid database after having annotated them. In this work, we screened both mouse and human complete genomes from ENSEMBL to find processed pseudogenes generated from functional genes with introns. We used a conservative method to detect processed pseudogenes in order to minimize the rate of false positive sequences. Within processed pseudogenes, some are still having a conserved open reading frame and some have overlapping gene locations. We designated as retroelements all reverse transcribed sequences and more strictly, we designated as processed pseudogenes, all retroelements not falling in the two former categories (having a conserved open reading or overlapping gene locations). We annotated 5823 retroelements (5206 processed pseudogenes) in the human genome and 3934 (3428 processed pseudogenes) in the mouse genome. Compared to previous estimations, the total number of processed pseudogenes was underestimated but the aim of this procedure was to generate a high-quality dataset. To facilitate the use of processed pseudogenes in studying genome structure and evolution, DNA sequences from processed pseudogenes, and their functional reverse transcribed homologs, are now stored in a nucleic acid database, HOPPSIGEN. HOPPSIGEN can be browsed on the PBIL (Pôle Bioinformatique Lyonnais) World Wide Web server (http://pbil.univ-lyon1.fr/) or fully downloaded for local installation.
DOI: 10.1093/nar/gkg530
发表时间: 2003-07-01
影响因子: 14.9
作者:
Perrière, G;Combet, C;Deléage, G
通讯作者: Deléage, G
DOI: 10.1093/nar/22.12.2360
发表时间: 1994-06-25
影响因子: 14.9
作者:
DURET, L;MOUCHIROUD, D;GOUY, M
通讯作者: GOUY, M
DOI: 10.1101/gr.331902
发表时间: 2002-10-01
期刊: GENOME RESEARCH
影响因子: 7
作者:
Zhang, ZL;Harrison, P;Gerstein, M
通讯作者: Gerstein, M
DOI: 10.1093/oxfordjournals.molbev.a040454
发表时间: 1987-07-01
影响因子: 10.7
作者:
SAITOU, N;NEI, M
通讯作者: NEI, M
DOI: 10.1101/gr.1455503
发表时间: 2003-12-01
期刊: GENOME RESEARCH
影响因子: 7
作者:
Torrents, D;Suyama, M;Bork, P
通讯作者: Bork, P