Sampling Variation of RAD-Seq Data from Diploid and Tetraploid Potato (Solanum tuberosum L.).

Sampling Variation of RAD-Seq Data from Diploid and Tetraploid Potato (Solanum tuberosum L.).
复制标题

二倍体和四倍体马铃薯 (Solanum tuberosumL.) 的 RAD-Seq 数据的采样变异

DOI:
10.3390/plants10020319
复制
发表时间:
2021-02-07
期刊:
Plants (Basel, Switzerland)
影响因子:
--
通讯作者:
Luo Z
Luo Z
中科院分区:
其他
文献类型:
--
作者:
Dang Z;Yang J;Wang L;Tao Q;Zhang F;Zhang Y;Luo Z

文献摘要

参考文献

被引文献

相似文献

新的测序技术能够在种群水平上以具有竞争力的低成本鉴定基于全基因组序列的变异。基于序列变异的分子标记引起了人们对群体和定量遗传分析的极大兴趣。序列数据的生成涉及一个复杂的实验过程,其中包含丰富的非生物变异。从统计学上讲,测序过程确实包括从单个序列中取样DNA片段。充分了解序列数据生成的抽样变化是任何下游数据分析和实施统计适当方法的关键统计特性之一。本文对二倍体和同源四倍体马铃薯(Solanum tuberosum L.)亲本及其后代进行了优化后的RAD-seq(限制性内切相关DNA测序)实验,并对其序列数据的采样变异进行了建模研究。分析表明,序列数据的抽样变化与文献中广泛假设的多项分布下的预期差异显著,并提供了对变化建模和计算模型参数的统计方法,这很容易在实际序列数据集中实现。RAD-seq实验的优化设计能够有效控制产生的序列数据中不需要的叶绿体DNA和RNA基因的呈现。
The new sequencing technology enables identification of genome-wide sequence-based variants at a population level and a competitively low cost. The sequence variant-based molecular markers have motivated enormous interest in population and quantitative genetic analyses. Generation of the sequence data involves a sophisticated experimental process embedded with rich non-biological variation. Statistically, the sequencing process indeed involves sampling DNA fragments from an individual sequence. Adequate knowledge of sampling variation of the sequence data generation is one of the key statistical properties for any downstream analysis of the data and for implementing statistically appropriate methods. This paper reports a thorough investigation on modeling the sampling variation of the sequence data from the optimized RAD-seq (Restriction sit associated DNA sequencing) experiments with two parents and their offspring of diploid and autotetraploid potato (Solanum tuberosum L.). The analysis shows significant dispersion in sampling variation of the sequence data over that expected under multinomial distribution as widely assumed in the literature and provides statistical methods for modeling the variation and calculating the model parameters, which may be easily implemented in real sequence datasets. The optimized design of RAD-seq experiments enabled effective control of presentation of undesirable chloroplast DNA and RNA genes in the sequence data generated.
DOI: 10.1038/nmeth.3582
发表时间: 2015-11
期刊: Nature methods
影响因子: 48
作者:
van de Geijn B;McVicker G;Gilad Y;Pritchard JK
通讯作者: Pritchard JK
DOI: 10.1007/s00122-014-2347-2
发表时间: 2014-09
影响因子: 5.4
作者:
Hackett, Christine A.;Bradshaw, John E.;Bryan, Glenn J.
通讯作者: Bryan, Glenn J.
DOI: 10.1038/nmeth.1923
发表时间: 2012-03-04
期刊: NATURE METHODS
影响因子: 48
作者:
Langmead, Ben;Salzberg, Steven L.
通讯作者: Salzberg, Steven L.
DOI: 10.1093/bioinformatics/btx133
发表时间: 2017-08-01
期刊: Bioinformatics (Oxford, England)
影响因子: --
作者:
Wu SH;Schwartz RS;Winter DJ;Conrad DF;Cartwright RA
通讯作者: Cartwright RA
DOI: 10.1371/journal.pone.0003376
发表时间: 2008
期刊: PloS one
影响因子: 3.7
作者:
Baird NA;Etter PD;Atwood TS;Currey MC;Shiver AL;Lewis ZA;Selker EU;Cresko WA;Johnson EA
通讯作者: Johnson EA