Determining the quality and complexity of next-generation sequencing data without a reference genome

Determining the quality and complexity of next-generation sequencing data without a reference genome
复制标题

DOI:
10.1186/s13059-014-0555-3
复制
发表时间:
2014-01-01
期刊:
影响因子:
12.3
通讯作者:
Laros, Jeroen F. J.
Laros, Jeroen F. J.
中科院分区:
生物学1区
文献类型:
--
作者:
Anvar, Seyed Yahya;Khachatryan, Lusine;Laros, Jeroen F. J.

文献摘要

被引文献

相似文献

我们描述了一个开源 kPAL 包,该包通过分析 k-mer 频率来促进对测序数据集的质量和可比性进行无比对评估。我们证明 kPAL 可以检测技术伪影,例如高重复率、文库嵌合体、污染和文库制备方案的差异。 kPAL 还成功捕获了微生物组的复杂性和多样性,并为研究微生物群落的变化提供了强大的手段。这些功能共同使 kPAL 成为一种有吸引力且广泛适用的工具,即使在没有参考序列的情况下也可以确定序列库的质量和可比性。 kPAL 可在 https://github.com/LUMC/kPAL 免费获取。
We describe an open-source kPAL package that facilitates an alignment-free assessment of the quality and comparability of sequencing datasets by analysing k-mer frequencies. We show that kPAL can detect technical artifacts such as high duplication rates, library chimeras, contamination, and differences in library preparation protocols. kPAL also successfully captures the complexity and diversity of microbiomes and provides a powerful means to study changes in microbial communities. Together, these features make kPAL an attractive and broadly applicable tool to determine the quality and comparability of sequence libraries even in the absence of a reference sequence. kPAL is freely available at https://github.com/LUMC/kPAL.