De novo assembly and characterization of Camelina sativa transcriptome by paired-end sequencing.

De novo assembly and characterization of Camelina sativa transcriptome by paired-end sequencing.
复制标题

DOI:
10.1186/1471-2164-14-146
复制
发表时间:
2013-03-05
期刊:
影响因子:
4.4
通讯作者:
Lim BL
Lim BL
中科院分区:
生物学2区
文献类型:
--
作者:
Liang C;Liu X;Yiu SM;Lim BL

文献摘要

参考文献

被引文献

相似文献

从亚麻荠种子中提取的生物燃料最近已成功地用作环保喷气燃料,以减少温室气体排放。亚麻荠在遗传上与拟南芥非常接近,两者都是拟南芥科的成员。虽然目前已有一些菊科植物的公共数据库,如A. thaliana,A. lyrata、甘蓝型油菜、B. juncea和B. rapa,没有公开的表达序列标签(EST)或亚麻荠的基因组数据。在这项研究中,一个高通量,大规模的RNA测序(RNA-seq)的亚麻荠转录组进行了生成一个数据库,这将是有用的进一步的功能分析。通过去除衔接子、模糊读段和低质量读段(2.42千兆碱基对)从原始读段过滤的大约2700万个干净的“读段”通过Illumina配对末端RNA-seq技术生成。使用SOAPdenovo和Trinity分别将所有这些干净读段从头组装成83,493个单基因和103,196个转录物。Trinity产生的转录本的平均长度为697 bp(N50 = 976),大于单基因的平均长度(319 bp,N50 = 346 bp)。尽管如此,在拟南芥CDS序列数据库(TAIR)的BLAstrom搜索中,由SOAPdenovo产生的组装产生了与Trinity(22,433)相似数量的非冗余命中(22,435)。四个公共数据库,基因和基因组的京都百科全书(KEGG),Swiss-prot,NCBI非冗余蛋白(NR),和集群的Orthogonal Groups(COG),被用于unigene注释; 67,791的83,493 unigene(81.2%),最终与基因描述或保守的蛋白质结构域被映射到25,329个非冗余蛋白质序列。我们将83,493个unigenes中的27,042个(32.4%)映射到119个KEGG代谢途径。这是亚麻荠转录组数据库的第一个报告,亚麻荠,一个环境上重要的成员。我们证明了C. savita与拟南芥属(Arabidopsis spp.)与芸苔属(Brassica spp.)虽然大多数注释基因与A.在拟南芥中,相当大比例的抗病基因(编码NBS的LRR基因)与其他菊科的基因更接近;这些基因包括B中的BrCN、BrCNL、BrNL、BrTN、BrTNL。拉帕。由于植物基因组长期处于环境胁迫的选择压力之下,这些抗病基因在C. sativa和B.拉帕基因组表明,它们在自然生境中受到密切相关的病原体的威胁。
Biofuels extracted from the seeds of Camelina sativa have recently been used successfully as environmentally friendly jet-fuel to reduce greenhouse gas emissions. Camelina sativa is genetically very close to Arabidopsis thaliana, and both are members of the Brassicaceae. Although public databases are currently available for some members of the Brassicaceae, such as A. thaliana, A. lyrata, Brassica napus, B. juncea and B. rapa, there are no public Expressed Sequence Tags (EST) or genomic data for Camelina sativa. In this study, a high-throughput, large-scale RNA sequencing (RNA-seq) of the Camelina sativa transcriptome was carried out to generate a database that will be useful for further functional analyses. Approximately 27 million clean “reads” filtered from raw reads by removal of adaptors, ambiguous reads and low-quality reads (2.42 gigabase pairs) were generated by Illumina paired-end RNA-seq technology. All of these clean reads were assembled de novo into 83,493 unigenes and 103,196 transcripts using SOAPdenovo and Trinity, respectively. The average length of the transcripts generated by Trinity was 697 bp (N50 = 976), which was longer than the average length of unigenes (319 bp, N50 = 346 bp). Nonetheless, the assembly generated by SOAPdenovo produced similar number of non-redundant hits (22,435) with that of Trinity (22,433) in BLASTN searches of the Arabidopsis thaliana CDS sequence database (TAIR). Four public databases, the Kyoto Encyclopedia of Genes and Genomes (KEGG), Swiss-prot, NCBI non-redundant protein (NR), and the Cluster of Orthologous Groups (COG), were used for unigene annotation; 67,791 of 83,493 unigenes (81.2%) were finally annotated with gene descriptions or conserved protein domains that were mapped to 25,329 non-redundant protein sequences. We mapped 27,042 of 83,493 unigenes (32.4%) to 119 KEGG metabolic pathways. This is the first report of a transcriptome database for Camelina sativa, an environmentally important member of the Brassicaceae. We showed that C. savita is closely related to Arabidopsis spp. and more distantly related to Brassica spp. Although the majority of annotated genes had high sequence identity to those of A. thaliana, a substantial proportion of disease-resistance genes (NBS-encoding LRR genes) were instead more closely similar to the genes of other Brassicaceae; these genes included BrCN, BrCNL, BrNL, BrTN, BrTNL in B. rapa. As plant genomes are under long-term selection pressure from environmental stressors, conservation of these disease-resistance genes in C. sativa and B. rapa genomes implies that they are exposed to the threats from closely-related pathogens in their natural habitats.
DOI: 10.1007/s11103-006-9086-y
发表时间: 2007-01-01
影响因子: 5.1
作者:
Kachroo, Aardra;Shanklin, John;Kachroo, Pradeep
通讯作者: Kachroo, Pradeep
DOI: 10.3168/jds.2007-0031
发表时间: 2007-11-01
影响因子: 3.5
作者:
Hurtaud, C.;Peyraud, J. L.
通讯作者: Peyraud, J. L.
DOI: 10.1186/1471-2164-11-180
发表时间: 2010-03-16
期刊: BMC genomics
影响因子: 4.4
作者:
Parchman TL;Geist KS;Grahnen JA;Benkman CW;Buerkle CA
通讯作者: Buerkle CA
DOI: 10.1016/s0926-6690(02)00098-5
发表时间: 2003-05-01
影响因子: 5.9
作者:
Bernardo, A;Howard-Hildige, R;Leahy, JJ
通讯作者: Leahy, JJ
DOI: 10.1186/1471-2229-10-233
发表时间: 2010-10-27
期刊: BMC plant biology
影响因子: 5.3
作者:
Hutcheon C;Ditt RF;Beilstein M;Comai L;Schroeder J;Goldstein E;Shewmaker CK;Nguyen T;De Rocher J;Kiser J
通讯作者: Kiser J