The consensus coding sequence (CCDS) project: Identifying a common protein-coding gene set for the human and mouse genomes

The consensus coding sequence (CCDS) project: Identifying a common protein-coding gene set for the human and mouse genomes
复制标题

DOI:
10.1101/gr.080531.108
复制
发表时间:
2009-07-01
期刊:
影响因子:
7
通讯作者:
Lipman, David
Lipman, David
中科院分区:
生物学1区
文献类型:
--
作者:
Pruitt, Kim D.;Harrow, Jennifer;Lipman, David

文献摘要

被引文献

相似文献

有效利用人类和小鼠基因组需要可靠地鉴定基因及其产物。虽然多个公共资源提供注释,但使用不同的方法可以导致基因,转录本和蛋白质的相似但不相同的表示。协同一致编码序列(CCDS)项目使用稳定标识符(CCDS ID)跟踪参考小鼠和人类基因组上相同的蛋白质注释,并确保它们在NCBI、Ensembl和UCSC基因组浏览器上一致地呈现。重要的是,该项目协调手动审查站点之间不一致的蛋白质注释,以及新证据表明需要修订的注释,以逐步收敛于人类和小鼠参考基因组的完整蛋白质编码集,同时保持高标准的可靠性和生物准确性。迄今为止,该项目已经从17,052个人类和16,893个小鼠基因中确定了20,159个人类和17,707个小鼠共有编码区。三种评价方法表明,CCDS集中的条目极有可能代表真实的蛋白质,比未纳入CCDS的贡献组的注释更有可能。因此,CCDS数据库集中了识别支持良好的、相同注释的蛋白质编码区的功能。
Effective use of the human and mouse genomes requires reliable identification of genes and their products. Although multiple public resources provide annotation, different methods are used that can result in similar but not identical representation of genes, transcripts, and proteins. The collaborative consensus coding sequence (CCDS) project tracks identical protein annotations on the reference mouse and human genomes with a stable identifier (CCDS ID), and ensures that they are consistently represented on the NCBI, Ensembl, and UCSC Genome Browsers. Importantly, the project coordinates on manually reviewing inconsistent protein annotations between sites, as well as annotations for which new evidence suggests a revision is needed, to progressively converge on a complete protein-coding set for the human and mouse reference genomes, while maintaining a high standard of reliability and biological accuracy. To date, the project has identified 20,159 human and 17,707 mouse consensus coding regions from 17,052 human and 16,893 mouse genes. Three evaluation methods indicate that the entries in the CCDS set are highly likely to represent real proteins, more so than annotations from contributing groups not included in CCDS. The CCDS database thus centralizes the function of identifying well-supported, identically-annotated, protein-coding regions.