FDSTools: A software package for analysis of massively parallel sequencing data with the ability to recognise and correct STR stutter and other PCR or sequencing noise

FDSTools: A software package for analysis of massively parallel sequencing data with the ability to recognise and correct STR stutter and other PCR or sequencing noise
复制标题

DOI:
10.1016/j.fsigen.2016.11.007
复制
发表时间:
2017-03-01
影响因子:
3.1
通讯作者:
Laros, Jeroen F. J.
Laros, Jeroen F. J.
中科院分区:
医学2区
文献类型:
--
作者:
Hoogenboom, Jerry;van der Gaag, Kristiaan J.;Laros, Jeroen F. J.

文献摘要

被引文献

相似文献

大规模并行测序(MPS)在法医学研究和案件工作中的广泛应用即将到来。经常提到的是,这种技术的主要优点之一是提高了分析代表不平衡混合物的证据痕迹的能力。然而,大多数分析法医短串联重复序列(STR)测序数据的可用软件包不太适合这种混合痕迹的高通量分析。最大的挑战是STR扩增中存在的断续伪影,这些伪影不容易从较小的贡献中辨别出来。FDSTools是为此目的开发的开源软件解决方案。口吃形成的水平受到序列的各个方面的影响,例如STR中发生的最长不间断延伸的长度。当使用MPS时,STR被评估为序列变体,每个变体具有可以精确确定的特定口吃特征。FDSTools使用参考样本的数据库来确定每个等位基因的断续和其他系统性PCR或测序伪影。此外,为每个重复元件创建断续模型,以便预测未包括在参考组中的等位基因的断续伪影。该信息随后用于识别和补偿序列谱中的噪声。结果是更好地代表样品的真实组成。使用Promega PowerseqTM自动系统数据从450个参考样本和31个两人混合,我们表明,FDSTools校正模块减少口吃率超过20%,低于3%。因此,检测到混合痕迹中的贡献水平低得多。FDSTools包含以交互式格式可视化数据的模块,允许用户使用自己的首选阈值过滤数据。(C)2016作者出版社:Elsevier爱尔兰Ltd.
Massively parallel sequencing (MPS) is on the advent of a broad scale application in forensic research and casework. The improved capabilities to analyse evidentiary traces representing unbalanced mixtures is often mentioned as one of the major advantages of this technique. However, most of the available software packages that analyse forensic short tandem repeat (STR) sequencing data are not well suited for high throughput analysis of such mixed traces. The largest challenge is the presence of stutter artefacts in STR amplifications, which are not readily discerned from minor contributions. FDSTools is an open-source software solution developed for this purpose. The level of stutter formation is influenced by various aspects of the sequence, such as the length of the longest uninterrupted stretch occurring in an STR. When MPS is used, STRs are evaluated as sequence variants that each have particular stutter characteristics which can be precisely determined. FDSTools uses a database of reference samples to determine stutter and other systemic PCR or sequencing artefacts for each individual allele. In addition, stutter models are created for each repeating element in order to predict stutter artefacts for alleles that are not included in the reference set. This information is subsequently used to recognise and compensate for the noise in a sequence profile. The result is a better representation of the true composition of a sample. Using Promega PowerseqTM Auto System data from 450 reference samples and 31 two-person mixtures, we show that the FDSTools correction module decreases stutter ratios above 20% to below 3%. Consequently, much lower levels of contributions in the mixed traces are detected. FDSTools contains modules to visualise the data in an interactive format allowing users to filter data with their own preferred thresholds. (C) 2016 The Authors. Published by Elsevier Ireland Ltd.