Generation and analysis of 280,000 human expressed sequence tags

Generation and analysis of 280,000 human expressed sequence tags
复制标题

DOI:
10.1101/gr.6.9.807
复制
发表时间:
1996-09-01
期刊:
影响因子:
7
通讯作者:
Marra, M
Marra, M
中科院分区:
生物学1区
文献类型:
--
作者:
Hillier, L;Lennon, G;Marra, M

文献摘要

被引文献

相似文献

我们报告了从194,031个人cDNA克隆的5'和3'末端获得的319,311个单程测序反应(称为表达序列抽头或EST)。我们的目标是从许多不同的基因中获得标签序列,并将其存款可公开访问的表达序列标签数据库。数据的高效自动筛选允许无延迟地沉积注释序列。从26个oligo(dT)引物定向克隆文库中产生序列,其中18个被标准化。使用从代表三种发育状态的17种不同组织分离的mRNA构建文库。我们的数据与非冗余的人类mRNA和蛋白质数据库的一个子集的比较表明,EST代表许多已知的序列,并包含许多是新的。使用隐马尔可夫模型的蛋白质家族的分析证实了这一观察结果,并支持这样的论点,即虽然归一化显着减少冗余的cDNA克隆的相对丰度,它不会导致基因家族的成员完全删除。
We report the generation of 319,311 single-pass sequencing reactions (known as expressed sequence taps, or ESTs) obtained from the 5' and 3' ends of 194,031 human cDNA clones. Our goal has been to obtain tag sequences From many different genes and to deposit these in the publicly accessible Data Base for Expressed Sequence Taps. Highly efficient automatic screening of the data allows deposition of the annotated sequences without delay. Sequences have been generated From 26 oligo(dT) primed directionally cloned libraries, of which 18 were normalized. The libraries were constructed using mRNA isolated From 17 different tissues representing three developmental states. Comparisons of a subset of our data with nonredundant human mRNA and protein data bases show that the ESTs represent many known sequences and contain many that are novel. Analysis of protein families using Hidden Markov Models confirms this observation and supports the contention that although normalization reduces significantly the relative abundance of redundant cDNA clones, it does not result in the complete removal of members of gene families.