Reference sequence (RefSeq) database at NCBI: current status, taxonomic expansion, and functional annotation.

Reference sequence (RefSeq) database at NCBI: current status, taxonomic expansion, and functional annotation.
复制标题

DOI:
10.1093/nar/gkv1189
复制
发表时间:
2016-01-04
影响因子:
14.9
通讯作者:
Pruitt KD
Pruitt KD
中科院分区:
生物学2区
文献类型:
--
作者:
O'Leary NA;Wright MW;Brister JR;Ciufo S;Haddad D;McVeigh R;Rajput B;Robbertse B;Smith-White B;Ako-Adjei D;Astashyn A;Badretdin A;Bao Y;Blinkova O;Brover V;Chetvernin V;Choi J;Cox E;Ermolaeva O;Farrell CM;Goldfarb T;Gupta T;Haft D;Hatcher E;Hlavina W;Joardar VS;Kodali VK;Li W;Maglott D;Masterson P;McGarvey KM;Murphy MR;O'Neill K;Pujar S;Rangwala SH;Rausch D;Riddick LD;Schoch C;Shkeda A;Storz SS;Sun H;Thibaud-Nissen F;Tolstoy I;Tully RE;Vatsan AR;Wallin C;Webb D;Wu W;Landrum MJ;Kimchi A;Tatusova T;DiCuccio M;Kitts P;Murphy TD;Pruitt KD

文献摘要

被引文献

相似文献

国家生物技术信息中心(NCBI)的RefSeq项目维护并管理着一个公开的基因组、转录本和蛋白质序列记录数据库(http://www.ncbi.nlm.nih.gov/refseq/)。RefSeq项目利用提交给国际核苷酸序列数据库协作(INSDC)的数据,结合计算、人工管理和协作,生成一套稳定、无冗余的标准参考序列。RefSeq项目用现有的知识(包括出版物、功能特性和信息命名法)来扩充这些参考序列。该数据库目前包含超过55000种生物体的序列(>4800种病毒,> 40000种原核生物,> 10000种真核生物;RefSeq版本71),范围从单个记录到完整基因组。本文总结了RefSeq项目的病毒、原核和真核分支的现状,报告了对数据访问的改进,并详细介绍了进一步扩大该集合的分类代表性的工作。我们还强调了多种功能管理计划,支持RefSeq数据的多种用途,包括分类验证、基因组注释、比较基因组学和临床测试。我们总结了我们在脊椎动物、植物和其他物种的人工管理过程中利用现有RNA-Seq和其他数据类型的方法,并描述了原核基因组和蛋白质名称管理的新方向。
The RefSeq project at the National Center for Biotechnology Information (NCBI) maintains and curates a publicly available database of annotated genomic, transcript, and protein sequence records (http://www.ncbi.nlm.nih.gov/refseq/). The RefSeq project leverages the data submitted to the International Nucleotide Sequence Database Collaboration (INSDC) against a combination of computation, manual curation, and collaboration to produce a standard set of stable, non-redundant reference sequences. The RefSeq project augments these reference sequences with current knowledge including publications, functional features and informative nomenclature. The database currently represents sequences from more than 55 000 organisms (>4800 viruses, >40 000 prokaryotes and >10 000 eukaryotes; RefSeq release 71), ranging from a single record to complete genomes. This paper summarizes the current status of the viral, prokaryotic, and eukaryotic branches of the RefSeq project, reports on improvements to data access and details efforts to further expand the taxonomic representation of the collection. We also highlight diverse functional curation initiatives that support multiple uses of RefSeq data including taxonomic validation, genome annotation, comparative genomics, and clinical testing. We summarize our approach to utilizing available RNA-Seq and other data types in our manual curation process for vertebrate, plant, and other species, and describe a new direction for prokaryotic genomes and protein name management.