Integrative annotation of 21,037 human genes validated by full-length cDNA clones.

Integrative annotation of 21,037 human genes validated by full-length cDNA clones.
复制标题

DOI:
10.1371/journal.pbio.0020162
复制
发表时间:
2004-06
期刊:
影响因子:
9.8
通讯作者:
Sugano S
Sugano S
中科院分区:
生物学1区
文献类型:
--
作者:
Imanishi T;Itoh T;Suzuki Y;O'Donovan C;Fukuchi S;Koyanagi KO;Barrero RA;Tamura T;Yamaguchi-Kabata Y;Tanino M;Yura K;Miyazaki S;Ikeo K;Homma K;Kasprzyk A;Nishikawa T;Hirakawa M;Thierry-Mieg J;Thierry-Mieg D;Ashurst J;Jia L;Nakao M;Thomas MA;Mulder N;Karavidopoulou Y;Jin L;Kim S;Yasuda T;Lenhard B;Eveno E;Suzuki Y;Yamasaki C;Takeda J;Gough C;Hilton P;Fujii Y;Sakai H;Tanaka S;Amid C;Bellgard M;Bonaldo Mde F;Bono H;Bromberg SK;Brookes AJ;Bruford E;Carninci P;Chelala C;Couillault C;de Souza SJ;Debily MA;Devignes MD;Dubchak I;Endo T;Estreicher A;Eyras E;Fukami-Kobayashi K;Gopinath GR;Graudens E;Hahn Y;Han M;Han ZG;Hanada K;Hanaoka H;Harada E;Hashimoto K;Hinz U;Hirai M;Hishiki T;Hopkinson I;Imbeaud S;Inoko H;Kanapin A;Kaneko Y;Kasukawa T;Kelso J;Kersey P;Kikuno R;Kimura K;Korn B;Kuryshev V;Makalowska I;Makino T;Mano S;Mariage-Samson R;Mashima J;Matsuda H;Mewes HW;Minoshima S;Nagai K;Nagasaki H;Nagata N;Nigam R;Ogasawara O;Ohara O;Ohtsubo M;Okada N;Okido T;Oota S;Ota M;Ota T;Otsuki T;Piatier-Tonneau D;Poustka A;Ren SX;Saitou N;Sakai K;Sakamoto S;Sakate R;Schupp I;Servant F;Sherry S;Shiba R;Shimizu N;Shimoyama M;Simpson AJ;Soares B;Steward C;Suwa M;Suzuki M;Takahashi A;Tamiya G;Tanaka H;Taylor T;Terwilliger JD;Unneberg P;Veeramachaneni V;Watanabe S;Wilming L;Yasuda N;Yoo HS;Stodolsky M;Makalowski W;Go M;Nakai K;Takagi T;Kanehisa M;Sakaki Y;Quackenbush J;Okazaki Y;Hayashizaki Y;Hide W;Chakraborty R;Nishikawa K;Sugawara H;Tateno Y;Chen Z;Oishi M;Tonellato P;Apweiler R;Okubo K;Wagner L;Wiemann S;Strausberg RL;Isogai T;Auffray C;Nomura N;Gojobori T;Sugano S

文献摘要

参考文献

被引文献

相似文献

人类基因组序列定义了我们固有的生物潜力;实现其中编码的生物学需要了解每个基因的功能。目前,我们在这方面的知识仍然有限。几条研究路线已被用于阐明人类基因组中基因的结构和功能。即便如此,基因预测仍然是一项艰巨的任务,因为基因的转录本的种类可能在很大程度上变化。因此,我们对41,118个全长cDNA进行了详尽的综合表征,这些cDNA将基因转录物捕获为完整的功能盒,提供了基因水平上结构和功能多样性的明确报告。我们的国际合作已经通过使用统一标准的策展分析高质量的全长cDNA克隆,验证了21,037个人类候选基因。这导致了5,155个新的候选基因的鉴定。这也是控制cDNA克隆质量最可靠的方法。我们开发了一个人类基因数据库,称为H-Invitational Database(H-InvDB; http://www.h-invitational.jp/)。它规定如下:人类基因的综合注释、基因结构的描述、新的可变剪接异构体的细节、非蛋白质编码RNA、功能结构域、亚细胞定位、代谢途径、蛋白质三维结构的预测、已知单核苷酸多态性(SNP)的作图、人类基因内多态性微卫星重复序列的鉴定以及与小鼠全长cDNA的比较结果。H-InvDB分析表明,高达4%的人类基因组序列(国家生物技术信息中心构建34组装)可能包含错误组装或缺失的区域。我们发现,6.5%的人类候选基因(1,377个位点)不具有良好的蛋白质编码开放阅读框架,其中296个位点是非蛋白质编码RNA基因的强候选基因。此外,在72,027个位于人类基因内的独特定位的SNP和插入/缺失中,13,215个非同义SNP,315个无义SNP和452个indel发生在编码区。与编码区中存在的25个多态性微卫星重复序列一起,它们可以改变蛋白质结构,引起表型效应或导致疾病。H-InvDB平台代表了对人类生物学和病理学探索所需资源的重大贡献。一个国际团队使用全长cDNA系统地验证和注释了21,000多个人类基因,从而为人类遗传学界提供了宝贵的新资源
The human genome sequence defines our inherent biological potential; the realization of the biology encoded therein requires knowledge of the function of each gene. Currently, our knowledge in this area is still limited. Several lines of investigation have been used to elucidate the structure and function of the genes in the human genome. Even so, gene prediction remains a difficult task, as the varieties of transcripts of a gene may vary to a great extent. We thus performed an exhaustive integrative characterization of 41,118 full-length cDNAs that capture the gene transcripts as complete functional cassettes, providing an unequivocal report of structural and functional diversity at the gene level. Our international collaboration has validated 21,037 human gene candidates by analysis of high-quality full-length cDNA clones through curation using unified criteria. This led to the identification of 5,155 new gene candidates. It also manifested the most reliable way to control the quality of the cDNA clones. We have developed a human gene database, called the H-Invitational Database (H-InvDB; http://www.h-invitational.jp/). It provides the following: integrative annotation of human genes, description of gene structures, details of novel alternative splicing isoforms, non-protein-coding RNAs, functional domains, subcellular localizations, metabolic pathways, predictions of protein three-dimensional structure, mapping of known single nucleotide polymorphisms (SNPs), identification of polymorphic microsatellite repeats within human genes, and comparative results with mouse full-length cDNAs. The H-InvDB analysis has shown that up to 4% of the human genome sequence (National Center for Biotechnology Information build 34 assembly) may contain misassembled or missing regions. We found that 6.5% of the human gene candidates (1,377 loci) did not have a good protein-coding open reading frame, of which 296 loci are strong candidates for non-protein-coding RNA genes. In addition, among 72,027 uniquely mapped SNPs and insertions/deletions localized within human genes, 13,215 nonsynonymous SNPs, 315 nonsense SNPs, and 452 indels occurred in coding regions. Together with 25 polymorphic microsatellite repeats present in coding regions, they may alter protein structure, causing phenotypic effects or resulting in disease. The H-InvDB platform represents a substantial contribution to resources needed for the exploration of human biology and pathology. An international team has systematically validated and annotated just over 21,000 human genes using full-length cDNA, thereby providing a valuable new resource for the human genetics community
DOI: 10.1073/pnas.201182798
发表时间: 2001-10-09
影响因子: 11.1
作者:
Camargo, AA;Samaia, HPB;de Souza, SJ
通讯作者: de Souza, SJ
DOI: 10.1101/gr.186901
发表时间: 2001-09-01
期刊: GENOME RESEARCH
影响因子: 7
作者:
Clark, MD;Hennig, S;Johnson, SL
通讯作者: Johnson, SL
DOI: 10.1038/ng0893-398
发表时间: 1993-08-01
期刊: NATURE GENETICS
影响因子: 30.8
作者:
ANDREW, SE;GOLDBERG, YP;HAYDEN, MR
通讯作者: HAYDEN, MR
DOI: 10.1073/pnas.152324199
发表时间: 2002-08-20
影响因子: 11.1
作者:
Boon, K;Osório, EC;Riggins, GJ
通讯作者: Riggins, GJ
DOI: 10.1038/ng0893-387
发表时间: 1993-08-01
期刊: NATURE GENETICS
影响因子: 30.8
作者:
DUYAO, M;AMBROSE, C;MACDONALD, M
通讯作者: MACDONALD, M