Annotation-based genome-wide SNP discovery in the large and complex Aegilops tauschii genome using next-generation sequencing without a reference genome sequence.

Annotation-based genome-wide SNP discovery in the large and complex Aegilops tauschii genome using next-generation sequencing without a reference genome sequence.
复制标题

DOI:
10.1186/1471-2164-12-59
复制
发表时间:
2011-01-25
期刊:
影响因子:
4.4
通讯作者:
Anderson OD
Anderson OD
中科院分区:
生物学2区
文献类型:
--
作者:
You FM;Huo N;Deal KR;Gu YQ;Luo MC;McGuire PE;Dvorak J;Anderson OD

文献摘要

参考文献

被引文献

相似文献

许多植物都有庞大而复杂的基因组,其中包含大量的重复序列。许多植物也是多倍体的。这两个属性都代表了小麦族的基因组结构,其成员包括经济上重要的小麦、黑麦和大麦。庞大的基因组大小、丰富的重复序列和多倍体给使用下一代基因组DNA测序(NGS)的全基因组SNP发现带来了挑战,因为这使得NGS平台产生的短阅读序列的比对和聚类变得困难,特别是在缺乏参考基因组序列的情况下。报道了一种基于注释的、全基因组范围的SNP发现管道,该管道使用针对大型和复杂基因组的NGS数据,而不需要参考基因组序列。对一种基因具有低基因组覆盖率的Roche 454鸟枪读数进行注释,以便将单拷贝序列和重复连接与重复序列和由近缘基因共享的序列区分开来。然后,用Solid或Solexa产生的另一个基因的猎枪读数的多个基因组等价物被映射到注释的罗氏454读数,以识别推定的SNP。开发了一个流水线程序包AGSNP,用于小麦D基因组的二倍体来源节节麦全基因组SNP的发现,其基因组大小为4.02 GB,其中90%是重复序列。Ae.基因组DNA。用罗氏454NGS平台对节节杆菌的AL8/78进行了测序。Ae.基因组DNA和c DNA序列。虽然也产生了一些Solexa和Roche 454基因组序列,但主要使用Solid对tauschii Access AS75进行了测序。在基因序列中共发现195,631个可能的SNPs,155,580个可能的SNPs在未鉴定的单复制区中被发现,另外145,907个可能的SNPs在重复连接中被发现。这些SNPs分布在整个Ae地区。结节菌基因组。为了评估假阳性SNP的发现率,用聚合酶链式反应从AL8/78和AS75中扩增了含有可能的SNP的DNA,并与ABI 3730 XL进行了重新测序。在随机选取的302个SNPs样本中,84.0%位于基因区,88.0%位于重复连接,81.3%位于未特定区。为NGS平台开发了一个基于注释的全基因组SNP发现流水线。该管道适合于在复杂基因组的基因组文库中发现SNP,并且不需要参考基因组序列。该管道适用于目前所有的NGS平台,前提是至少有一个这样的平台产生相对较长的读取。管道包AGSNP和已发现的497,118 Ae。Tauschii SNPs可以在(http://avena.pw.usda.gov/wheatD/agsnp.shtml).上获得
Many plants have large and complex genomes with an abundance of repeated sequences. Many plants are also polyploid. Both of these attributes typify the genome architecture in the tribe Triticeae, whose members include economically important wheat, rye and barley. Large genome sizes, an abundance of repeated sequences, and polyploidy present challenges to genome-wide SNP discovery using next-generation sequencing (NGS) of total genomic DNA by making alignment and clustering of short reads generated by the NGS platforms difficult, particularly in the absence of a reference genome sequence. An annotation-based, genome-wide SNP discovery pipeline is reported using NGS data for large and complex genomes without a reference genome sequence. Roche 454 shotgun reads with low genome coverage of one genotype are annotated in order to distinguish single-copy sequences and repeat junctions from repetitive sequences and sequences shared by paralogous genes. Multiple genome equivalents of shotgun reads of another genotype generated with SOLiD or Solexa are then mapped to the annotated Roche 454 reads to identify putative SNPs. A pipeline program package, AGSNP, was developed and used for genome-wide SNP discovery in Aegilops tauschii-the diploid source of the wheat D genome, and with a genome size of 4.02 Gb, of which 90% is repetitive sequences. Genomic DNA of Ae. tauschii accession AL8/78 was sequenced with the Roche 454 NGS platform. Genomic DNA and cDNA of Ae. tauschii accession AS75 was sequenced primarily with SOLiD, although some Solexa and Roche 454 genomic sequences were also generated. A total of 195,631 putative SNPs were discovered in gene sequences, 155,580 putative SNPs were discovered in uncharacterized single-copy regions, and another 145,907 putative SNPs were discovered in repeat junctions. These SNPs were dispersed across the entire Ae. tauschii genome. To assess the false positive SNP discovery rate, DNA containing putative SNPs was amplified by PCR from AL8/78 and AS75 and resequenced with the ABI 3730 xl. In a sample of 302 randomly selected putative SNPs, 84.0% in gene regions, 88.0% in repeat junctions, and 81.3% in uncharacterized regions were validated. An annotation-based genome-wide SNP discovery pipeline for NGS platforms was developed. The pipeline is suitable for SNP discovery in genomic libraries of complex genomes and does not require a reference genome sequence. The pipeline is applicable to all current NGS platforms, provided that at least one such platform generates relatively long reads. The pipeline package, AGSNP, and the discovered 497,118 Ae. tauschii SNPs can be accessed at (http://avena.pw.usda.gov/wheatD/agsnp.shtml).
DOI: 10.1186/gb-2007-8-7-r143
发表时间: 2007
期刊: Genome biology
影响因子: 12.3
作者:
Huse SM;Huber JA;Morrison HG;Sogin ML;Welch DM
通讯作者: Welch DM
DOI: 10.1186/1471-2164-11-702
发表时间: 2010-12-14
期刊: BMC genomics
影响因子: 4.4
作者:
Akhunov ED;Akhunova AR;Anderson OD;Anderson JA;Blake N;Clegg MT;Coleman-Derr D;Conley EJ;Crossman CC;Deal KR;Dubcovsky J;Gill BS;Gu YQ;Hadam J;Heo H;Huo N;Lazo GR;Luo MC;Ma YQ;Matthews DE;McGuire PE;Morrell PL;Qualset CO;Renfro J;Tabanao D;Talbert LE;Tian C;Toleno DM;Warburton ML;You FM;Zhang W;Dvorak J
通讯作者: Dvorak J
DOI: 10.1093/oxfordjournals.jhered.a105590
发表时间: 1946-01-01
影响因子: 3.1
作者:
MCFADDEN, ES;SEARS, ER
通讯作者: SEARS, ER
DOI: 10.1139/g88-115
发表时间: 1988-10-01
期刊: GENOME
影响因子: 3.1
作者:
DVORAK, J;MCGUIRE, PE;CASSIDY, B
通讯作者: CASSIDY, B
DOI: 10.1007/bf00229502
发表时间: 1992-07-01
影响因子: 5.4
作者:
DVORAK, J;ZHANG, HB
通讯作者: ZHANG, HB