Homoplasy corrected estimation of genetic similarity from AFLP bands, and the effect of the number of bands on the precision of estimation

Homoplasy corrected estimation of genetic similarity from AFLP bands, and the effect of the number of bands on the precision of estimation
复制标题

DOI:
10.1007/s00122-009-1047-9
复制
发表时间:
2009-08-01
影响因子:
5.4
通讯作者:
van Eeuwijk, Fred
van Eeuwijk, Fred
中科院分区:
农林科学1区
文献类型:
--
作者:
Gort, Gerrit;van Hintum, Theo;van Eeuwijk, Fred

文献摘要

被引文献

相似文献

AFLP是一种DNA指纹技术,产生具有已知或未知条带位置的二进制条带存在-不存在模式,称为谱。我们将AFLP建模为片段的采样过程,长度从分布中采样。条带代表特定长度的片段。我们专注于估计成对的遗传相似性,定义为共同片段的平均分数,AFLP。最小估计量是Dice(D)或Jaccard系数。D高估了遗传相似性,因为图谱对中相同的条带可能对应于不同的片段(同质性)。另一个复杂的因素是在一个剖面中出现不同的等长片段,表现为一个单一的条带,我们称之为碰撞。D的偏差随着带数的增加和遗传相似性的降低而增加。我们提出了两个同质性和碰撞校正估计的遗传相似性。第一个是D的修改,用估计的片段计数代替条带计数。第二个是最大似然估计,仅适用于如果波段位置可用。通过仿真研究了估计量的性质。第一个标准误差和置信区间是通过自举法获得的,第二个是通过似然理论获得的。该估计量几乎是无偏的,并且在大多数实际情况下具有比D更小的标准误差。基于似然的估计通常给出最高的精度。利用仿真研究了碎片数与精度之间的关系。通常的条带计数范围(50-100)似乎接近最佳。该方法说明使用的数据从莴苣的系统发育研究。
AFLP is a DNA fingerprinting technique, resulting in binary band presence-absence patterns, called profiles, with known or unknown band positions. We model AFLP as a sampling procedure of fragments, with lengths sampled from a distribution. Bands represent fragments of specific lengths. We focus on estimation of pairwise genetic similarity, defined as average fraction of common fragments, by AFLP. Usual estimators are Dice (D) or Jaccard coefficients. D overestimates genetic similarity, since identical bands in profile pairs may correspond to different fragments (homoplasy). Another complicating factor is the occurrence of different fragments of equal length within a profile, appearing as a single band, which we call collision. The bias of D increases with larger numbers of bands, and lower genetic similarity. We propose two homoplasy- and collision-corrected estimators of genetic similarity. The first is a modification of D, replacing band counts by estimated fragment counts. The second is a maximum likelihood estimator, only applicable if band positions are available. Properties of the estimators are studied by simulation. Standard errors and confidence intervals for the first are obtained by bootstrapping, and for the second by likelihood theory. The estimators are nearly unbiased, and have for most practical cases smaller standard error than D. The likelihood-based estimator generally gives the highest precision. The relationship between fragment counts and precision is studied using simulation. The usual range of band counts (50-100) appears nearly optimal. The methodology is illustrated using data from a phylogenetic study on lettuce.