Alignment-Free Sequence Analysis and Applications

Alignment-Free Sequence Analysis and Applications
复制标题

DOI:
10.1146/annurev-biodatasci-080917-013431
复制
发表时间:
2018-01-01
期刊:
ANNUAL REVIEW OF BIOMEDICAL DATA SCIENCE, VOL 1
影响因子:
--
通讯作者:
Sun, Fengzhu
Sun, Fengzhu
中科院分区:
其他
文献类型:
--
作者:
Ren, Jie;Bai, Xin;Sun, Fengzhu

文献摘要

被引文献

相似文献

由于庞大的数据量和相对较短的读取长度,基于大量下一代测序(NGS)数据的基因组和元基因组比较给基于比对的方法带来了巨大的挑战。基于NGS数据中单词模式计数的无比对方法不依赖于完整的基因组,通常在计算上是高效的。因此,它们对基因组和元基因组的比较有很大贡献。最近,已经发展了新的统计方法来比较长序列和猎枪序列。这些方法已经被应用于许多问题,包括比较基因调节区、基因组序列、元基因组、元基因组数据中的重叠群、识别病毒与宿主的相互作用以及水平基因转移的检测。我们提供了这些应用和基于单词计数的无比对序列分析方法的其他相关发展的最新综述。
Genome and metagenome comparisons based on large amounts of nextgeneration sequencing (NGS) data pose significant challenges for alignmentbased approaches due to the huge data size and the relatively short length of the reads. Alignment-free approaches based on the counts of word patterns in NGS data do not depend on the complete genome and are generally computationally efficient. Thus, they contribute significantly to genome and metagenome comparison. Recently, novel statistical approaches have been developed for the comparison of both long and shotgun sequences. These approaches have been applied to many problems, including the comparison of gene regulatory regions, genome sequences, metagenomes, binning contigs in metagenomic data, identification of virus-host interactions, and detection of horizontal gene transfers. We provide an updated review of these applications and other related developments of word count-based approaches for alignment-free sequence analysis.