Quantifying sequence proportions in a DNA-based diet study using Ion Torrent amplicon sequencing: which counts count?

Quantifying sequence proportions in a DNA-based diet study using Ion Torrent amplicon sequencing: which counts count?
复制标题

DOI:
10.1111/1755-0998.12103
复制
发表时间:
2013-07-01
影响因子:
7.7
通讯作者:
Jarman, Simon N.
Jarman, Simon N.
中科院分区:
生物学1区
文献类型:
--
作者:
Deagle, Bruce E.;Thomas, Austen C.;Jarman, Simon N.

文献摘要

被引文献

相似文献

许多环境DNA条形码研究的目标是基于由高通量测序产生的序列读段比例来推断关于不同分类群的相对丰度的定量信息。然而,与这种方法相关的潜在偏见才刚刚开始受到审查。我们对捕获的斑海豹(Phoca vitulina)的粪便(粪便)扩增的DNA进行测序,以研究序列计数是否可以用来量化海豹的饮食。海豹喂食鱼在固定的比例,脊索动物特异性线粒体16S标记扩增的粪便DNA和扩增子测序使用离子激流PGM。对于一组给定的生物信息学参数,粪便样本之间的比例猎物物种序列回收的变异性一般较低。然而,根据测序方向、质量过滤水平(由于物种之间的序列质量差异)和考虑的最小读取长度,比例变化很大。用于识别单个样品的短引物标签也影响物种比例。此外,各因素之间存在复杂的交互作用,如引物标签和测序方向对质量过滤效果的影响。样本子集的重新测序显示,运行之间存在一些(但不是全部)偏倚。不太严格的数据过滤(基于质量评分或读取长度)通常产生更一致的比例数据,但序列的总体比例与膳食质量比例非常不同,表明存在额外的技术或生物学偏倚。我们的研究结果强调,通过高通量测序产生的序列比例的定量解释将需要仔细的实验设计和周到的数据分析。
A goal of many environmental DNA barcoding studies is to infer quantitative information about relative abundances of different taxa based on sequence read proportions generated by high-throughput sequencing. However, potential biases associated with this approach are only beginning to be examined. We sequenced DNA amplified from faeces (scats) of captive harbour seals (Phoca vitulina) to investigate whether sequence counts could be used to quantify the seals' diet. Seals were fed fish in fixed proportions, a chordate-specific mitochondrial 16S marker was amplified from scat DNA and amplicons sequenced using an Ion Torrent PGM. For a given set of bioinformatic parameters, there was generally low variability between scat samples in proportions of prey species sequences recovered. However, proportions varied substantially depending on sequencing direction, level of quality filtering (due to differences in sequence quality between species) and minimum read length considered. Short primer tags used to identify individual samples also influenced species proportions. In addition, there were complex interactions between factors; for example, the effect of quality filtering was influenced by the primer tag and sequencing direction. Resequencing of a subset of samples revealed some, but not all, biases were consistent between runs. Less stringent data filtering (based on quality scores or read length) generally produced more consistent proportional data, but overall proportions of sequences were very different than dietary mass proportions, indicating additional technical or biological biases are present. Our findings highlight that quantitative interpretations of sequence proportions generated via high-throughput sequencing will require careful experimental design and thoughtful data analysis.