PCR Strategies for Complete Allele Calling in Multigene Families Using High-Throughput Sequencing Approaches.

PCR Strategies for Complete Allele Calling in Multigene Families Using High-Throughput Sequencing Approaches.
复制标题

DOI:
10.1371/journal.pone.0157402
复制
发表时间:
2016
期刊:
影响因子:
3.7
通讯作者:
Godoy JA
Godoy JA
中科院分区:
综合性期刊3区
文献类型:
--
作者:
Marmesat E;Soriano L;Mazzoni CJ;Sommer S;Godoy JA

文献摘要

被引文献

相似文献

具有高拷贝数变异的多基因家族的表征通常通过使用高度简并引物的PCR扩增来进行,以解释感兴趣区域侧翼的所有预期变体。这种方法通常引入PCR偏差,其导致高通量测序文库中靶标的不平衡代表,最终导致靶向等位基因的不完全检测。在这里,我们证实了这一结果,并提出了两种不同的扩增策略,以减轻这个问题。第一种策略(称为池化PCR)使用不同的中度简并引物对在多个独立PCR中靶向等位基因的不同子集,而第二种方法(称为池化引物)在单个PCR中使用定制的非简并引物池。我们比较他们的表现,一个单一的PCR与高度简并引物的伊比利亚猞猁的MHC I类作为一个模型的共同使用。我们发现这两种新方法的工作效果相似,而且比传统方法更好。他们对每个个体的等位基因评分显著更高(11.33 ± 1.38和11.72 ± 0.89 vs 7.94 ± 1.95),产生了更完整的等位基因谱(96.28 ± 8.46和99.50 ± 2.12 vs 63.76 ± 15.43),并在群体水平上揭示了更多的等位基因(13 vs 12)。最后,我们可以将每个等位基因的扩增效率与其侧翼序列中的引物错配联系起来,并表明高通量技术提供的超深覆盖并不能完全补偿这种偏差,特别是当真实的等位基因可能比人工制品达到更低的覆盖时。采用所提出的扩增方法中的任一种提供了在较低覆盖率下获得更完整的等位基因谱的机会,从而提高了下游分析和后续应用的置信度。
The characterization of multigene families with high copy number variation is often approached through PCR amplification with highly degenerate primers to account for all expected variants flanking the region of interest. Such an approach often introduces PCR biases that result in an unbalanced representation of targets in high-throughput sequencing libraries that eventually results in incomplete detection of the targeted alleles. Here we confirm this result and propose two different amplification strategies to alleviate this problem. The first strategy (called pooled-PCRs) targets different subsets of alleles in multiple independent PCRs using different moderately degenerate primer pairs, whereas the second approach (called pooled-primers) uses a custom-made pool of non-degenerate primers in a single PCR. We compare their performance to the common use of a single PCR with highly degenerate primers using the MHC class I of the Iberian lynx as a model. We found both novel approaches to work similarly well and better than the conventional approach. They significantly scored more alleles per individual (11.33 ± 1.38 and 11.72 ± 0.89 vs 7.94 ± 1.95), yielded more complete allelic profiles (96.28 ± 8.46 and 99.50 ± 2.12 vs 63.76 ± 15.43), and revealed more alleles at a population level (13 vs 12). Finally, we could link each allele’s amplification efficiency with the primer-mismatches in its flanking sequences and show that ultra-deep coverage offered by high-throughput technologies does not fully compensate for such biases, especially as real alleles may reach lower coverage than artefacts. Adopting either of the proposed amplification methods provides the opportunity to attain more complete allelic profiles at lower coverages, improving confidence over the downstream analyses and subsequent applications.