WindowMasker:: window-based masker for sequenced genomes

WindowMasker:: window-based masker for sequenced genomes
复制标题

DOI:
10.1093/bioinformatics/bti774
复制
发表时间:
2006-01-15
期刊:
影响因子:
5.8
通讯作者:
Agarwala, R
Agarwala, R
中科院分区:
生物学3区
文献类型:
--
作者:
Morgulis, A;Gertz, EM;Agarwala, R

文献摘要

被引文献

相似文献

动机:在DNA数据库搜索的输出中,与重复序列的匹配通常是不希望的。如果重复序列可以在数据库中被屏蔽,则不需要与查询匹配。RepeatMasker/Maskeraid(RM),目前最广泛使用的软件用于DNA序列掩蔽,是缓慢的,需要一个重复的模板序列库,如手动策划的RepBase库,可能不存在新测序的genomes.Results:我们已经开发出一种软件工具,称为WindowMasker(WM),识别和掩蔽基因组中高度重复的DNA序列,只使用基因组本身的序列。WM比RM快几个数量级,因为WM使用基因组序列的一些线性时间扫描,而不是将每个文库序列与基因组的每个片段进行比较的局部比对方法。我们验证WM比较BLAST输出从两个版本的相同的基因组,一个由WM掩蔽,另一个由RM掩蔽的查询。即使对于基因组,如人类基因组,其中一个很好的RepBase库是可用的,搜索数据库与WM屏蔽产生更多的匹配,显然是非重复的和更少的匹配重复序列。我们发现,这些结果也适用于转录区域。WM在基因组中也表现良好,因为在分析时大部分序列都是草稿形式。
Motivation: Matches to repetitive sequences are usually undesirable in the output of DNA database searches. Repetitive sequences need not be matched to a query, if they can be masked in the database. RepeatMasker/Maskeraid (RM), currently the most widely used software for DNA sequence masking, is slow and requires a library of repetitive template sequences, such as a manually curated RepBase library, that may not exist for newly sequenced genomes.Results: We have developed a software tool called WindowMasker (WM) that identifies and masks highly repetitive DNA sequences in a genome, using only the sequence of the genome itself. WM is orders of magnitude faster than RM because WM uses a few linear-time scans of the genome sequence, rather than local alignment methods that compare each library sequence with each piece of the genome. We validate WM by comparing BLAST outputs from large sets of queries applied to two versions of the same genome, one masked by WM, and the other masked by RM. Even for genomes such as the human genome, where a good RepBase library is available, searching the database as masked with WM yields more matches that are apparently non-repetitive and fewer matches to repetitive sequences. We show that these results hold for transcribed regions as well. WM also performs well on genomes for which much of the sequence was in draft form at the time of the analysis.