tuple_plot: Fast pairwise nucleotide sequence comparison with noise suppression

tuple_plot: Fast pairwise nucleotide sequence comparison with noise suppression
复制标题

DOI:
10.1093/bioinformatics/btl277
复制
发表时间:
2006-08-01
期刊:
影响因子:
5.8
通讯作者:
Platzer, Matthias
Platzer, Matthias
中科院分区:
生物学3区
文献类型:
--
作者:
Szafranski, Karol;Jahn, Niels;Platzer, Matthias

文献摘要

被引文献

相似文献

程序tuple_lot通过应用众所周知的点图原理识别并可视化两个基因组序列之间的局部相似性,通常是100kb或更长。从输入序列构建的序列单词词典用于构建特定于任务的期望模型,该模型用于将重要性值赋予成对单词命中。基于词典的方法允许快速计算,根据输入序列的大小,计算时间可扩展到O(N Log N)。所提出的评分方案显著地提高了信噪比,并可能有助于改进其他基于单词的序列比较方法。
The program tuple_plot identifies and visualizes local similarities between two genomic sequences, typically 100 kb or longer, by applying the well-known dotplot principle. A dictionary of sequence words built from the input sequences serves to construct a task-specific expectancy model that is used to attribute significance values to pairwise word hits. The dictionary-based approach allows fast computation, the computation time scaling to O(N log N), depending on the size of the input sequences. The proposed scoring scheme appreciably increases the signal-to-noise ratio and may help to improve other word-based sequence comparison approaches.