Sequencing Medicago truncatula expressed sequenced tags using 454 Life Sciences technology.

Sequencing Medicago truncatula expressed sequenced tags using 454 Life Sciences technology.
复制标题

测序的Medicago truncatula使用454个生命科学技术表达了测序标签。

DOI:
10.1186/1471-2164-7-272
复制
发表时间:
2006-10-24
期刊:
影响因子:
4.4
通讯作者:
Town, Christopher D
Town, Christopher D
中科院分区:
生物学2区
文献类型:
--
作者:
Cheung, Foo;Haas, Brian J;Goldberg, Susanne M D;May, Gregory D;Xiao, Yongli;Town, Christopher D

文献摘要

被引文献

相似文献

在这项研究中,我们解决了一个单一的454生命科学GS 20测序运行是否提供了新的基因发现从一个标准化的cDNA文库,以及通过这种技术产生的短读段是否在基因结构注释的价值。对来自标准化cDNA文库的衔接子连接的cDNA进行单次454个GS 20测序,产生292,465个读段,其在清洁后减少至252,384个读段,平均读段长度为92个核苷酸。经过聚类和组装,总共产生了184,599个独特的序列,包含超过400个SSR。这454个序列比来自MtGI的相当数量的序列产生更多基因的命中。虽然很短,但454个读段具有足够的长度以与通过常规测序产生的较长EST一样有效地映射到独特的基因组位置。通过基因本体论分配从匹配到拟南芥进行序列的功能解释,并且显示覆盖广泛的GO类别。53,796个组装件和单件(29%)在现有MtGI中没有匹配。在以前未观察到的苜蓿转录本中,数千个在综合蛋白质数据库和一个或多个TIGR植物基因索引中匹配。这些新序列中大约20%可以在苜蓿基因组序列中找到。使用PASA将454技术产生的总共70,026个读段映射到785个苜蓿成品BAC,并且超过1,000个基因模型需要修饰。与454测序平行,通过使用相同文库的常规测序产生4,445个5 ′-引物读段,并且从组装的序列显示含有约52%的全长cDNA,其编码长度为50至超过500个氨基酸的蛋白质。由于454 DNA测序技术提供了大量的读取,它有效地揭示了广泛的GO类别的转录本的表达,并在标准化cDNA文库中包含许多罕见的转录本,尽管只有有限的一部分序列未被发现。与较长的EST一样,454个读段可以唯一地映射到基因组序列上,以提供对基因预测的支持和修改。
In this study, we addressed whether a single 454 Life Science GS20 sequencing run provides new gene discovery from a normalized cDNA library, and whether the short reads produced via this technology are of value in gene structure annotation. A single 454 GS20 sequencing run on adapter-ligated cDNA, from a normalized cDNA library, generated 292,465 reads that were reduced to 252,384 reads with an average read length of 92 nucleotides after cleaning. After clustering and assembly, a total of 184,599 unique sequences were generated containing over 400 SSRs. The 454 sequences generated hits to more genes than a comparable amount of sequence from MtGI. Although short, the 454 reads are of sufficient length to map to a unique genome location as effectively as longer ESTs produced by conventional sequencing. Functional interpretation of the sequences was carried out by Gene Ontology assignments from matches to Arabidopsis and was shown to cover a broad range of GO categories. 53,796 assemblies and singletons (29%) had no match in the existing MtGI. Within the previously unobserved Medicago transcripts, thousands had matches in a comprehensive protein database and one or more of the TIGR Plant Gene Indices. Approximately 20% of these novel sequences could be found in the Medicago genome sequence. A total of 70,026 reads generated by the 454 technology were mapped to 785 Medicago finished BACs using PASA and over 1,000 gene models required modification. In parallel to 454 sequencing, 4,445 5'-prime reads were generated by conventional sequencing using the same library and from the assembled sequences it was shown to contain about 52% full length cDNAs encoding proteins from 50 to over 500 amino acids in length. Due to the large number of reads afforded by the 454 DNA sequencing technology, it is effective in revealing the expression of transcripts from a broad range of GO categories and contains many rare transcripts in normalized cDNA libraries, although only a limited portion of their sequence is uncovered. As with longer ESTs, 454 reads can be mapped uniquely onto genomic sequence to provide support for, and modifications of, gene predictions.