De novo assembly of the pepper transcriptome (Capsicum annuum): a benchmark for in silico discovery of SNPs, SSRs and candidate genes.

De novo assembly of the pepper transcriptome (Capsicum annuum): a benchmark for in silico discovery of SNPs, SSRs and candidate genes.
复制标题

DOI:
10.1186/1471-2164-13-571
复制
发表时间:
2012-10-30
期刊:
影响因子:
4.4
通讯作者:
Van Deynze A
Van Deynze A
中科院分区:
生物学2区
文献类型:
--
作者:
Ashrafi H;Hill T;Stoffel K;Kozik A;Yao J;Chin-Wo SR;Van Deynze A

文献摘要

被引文献

相似文献

辣椒分子育种研究进展可以通过在育种种质中开发与转录组相关的DNA标记来加速。在下一代测序(NGS)技术出现之前,大多数测序数据由桑格测序方法产生。通过利用桑格EST数据,我们已经产生了丰富的辣椒遗传信息,包括数千个SNP和单位点多态性(SPP)标记。为了补充和增强这些资源,我们将NGS应用于三种辣椒基因型:Maor,Early Jalapeño和Criollo de莫雷洛斯-334(CM 334),以鉴定这三种基因型组装中的SNP和SSR。两个辣椒转录组组装体开发了不同的目的。第一参考序列由CAP 3软件组装,包含来自> 125,000个Sanger-EST序列的31,196个重叠群,所述Sanger-EST序列主要来自韩国F1杂交系Bukang。为30,815个单基因设计重叠探针,以构建用于全基因组分析的辣椒AffyesterGeneChip ®微阵列。此外,自定义Python脚本用于识别组装体的重叠群中的4,236个SNP。从组装中总共鉴定了2,489个简单序列重复(SSR),并为SSR设计了引物。使用Blast 2GO软件对重叠群进行注释,得到组装中60%的单基因的信息。使用Velvet、CLC workbench和CAP 3软件包的组合,从超过2亿个Illumina Genome Analyzer II读段(80-120 nt)构建第二转录组组装体。BWA,SAMtools和内部Perl脚本被用来确定三个辣椒基因型之间的SNPs。将SNP过滤为距离任何内含子-外显子连接点以及侧翼SNP至少50 bp。超过22,000个高质量的推定SNP被确定。使用MISA软件,还在Illumina转录组组装中鉴定了10,398个SSR标记,并为鉴定的标记设计了引物。组装通过Blast 2GO注释,并且14,740(12%)个注释的重叠群与功能蛋白相关。在辣椒基因组序列可用之前,需要组装这种经济上重要的作物的转录组以产生数千个可用于育种计划的高质量分子标记。为了更好地理解组装的序列并鉴定潜在QTL的候选基因,我们注释了Sanger-EST和Illumina转录组组装的重叠群。这些信息和其他信息已经在一个数据库中整理,我们已经专门为胡椒项目。
Molecular breeding of pepper (Capsicum spp.) can be accelerated by developing DNA markers associated with transcriptomes in breeding germplasm. Before the advent of next generation sequencing (NGS) technologies, the majority of sequencing data were generated by the Sanger sequencing method. By leveraging Sanger EST data, we have generated a wealth of genetic information for pepper including thousands of SNPs and Single Position Polymorphic (SPP) markers. To complement and enhance these resources, we applied NGS to three pepper genotypes: Maor, Early Jalapeño and Criollo de Morelos-334 (CM334) to identify SNPs and SSRs in the assembly of these three genotypes. Two pepper transcriptome assemblies were developed with different purposes. The first reference sequence, assembled by CAP3 software, comprises 31,196 contigs from >125,000 Sanger-EST sequences that were mainly derived from a Korean F1-hybrid line, Bukang. Overlapping probes were designed for 30,815 unigenes to construct a pepper Affymetrix GeneChip® microarray for whole genome analyses. In addition, custom Python scripts were used to identify 4,236 SNPs in contigs of the assembly. A total of 2,489 simple sequence repeats (SSRs) were identified from the assembly, and primers were designed for the SSRs. Annotation of contigs using Blast2GO software resulted in information for 60% of the unigenes in the assembly. The second transcriptome assembly was constructed from more than 200 million Illumina Genome Analyzer II reads (80–120 nt) using a combination of Velvet, CLC workbench and CAP3 software packages. BWA, SAMtools and in-house Perl scripts were used to identify SNPs among three pepper genotypes. The SNPs were filtered to be at least 50 bp from any intron-exon junctions as well as flanking SNPs. More than 22,000 high-quality putative SNPs were identified. Using the MISA software, 10,398 SSR markers were also identified within the Illumina transcriptome assembly and primers were designed for the identified markers. The assembly was annotated by Blast2GO and 14,740 (12%) of annotated contigs were associated with functional proteins. Before availability of pepper genome sequence, assembling transcriptomes of this economically important crop was required to generate thousands of high-quality molecular markers that could be used in breeding programs. In order to have a better understanding of the assembled sequences and to identify candidate genes underlying QTLs, we annotated the contigs of Sanger-EST and Illumina transcriptome assemblies. These and other information have been curated in a database that we have dedicated for pepper project.