Gentle Masking of Low-Complexity Sequences Improves Homology Search

Gentle Masking of Low-Complexity Sequences Improves Homology Search
复制标题

DOI:
10.1371/journal.pone.0028819
复制
发表时间:
2011-12-09
期刊:
影响因子:
3.7
通讯作者:
Frith, Martin C.
Frith, Martin C.
中科院分区:
综合性期刊3区
文献类型:
--
作者:
Frith, Martin C.

文献摘要

被引文献

相似文献

同源序列的检测是计算生物学中的一项基本任务。这个任务被低复杂度的片段(如atatatatatat)混淆了,这些片段频繁且独立地出现,导致了不是同源的强烈相似性。对于低复杂度片段的识别已经有了很多研究,但是对于在同源性搜索过程中如何处理低复杂度片段的研究很少。我们建议通过对齐序列与低复杂度域的“温和”掩蔽来找到同源性。温和掩蔽意味着涉及掩蔽字母的匹配分数是min(0,S),其中S是未掩蔽的分数。温和的掩蔽稍微但明显地提高了同源性搜索的灵敏度(与“苛刻”掩蔽相比),而不损害特异性。我们展示了三个有用的同源性搜索问题的例子:NUMTs(线粒体DNA的核拷贝)的检测,宏基因组DNA读取参考基因组的招聘,和假基因检测。温和的掩蔽是目前在同源性搜索过程中处理低复杂性片段的最佳方法。
Detection of sequences that are homologous, i.e. descended from a common ancestor, is a fundamental task in computational biology. This task is confounded by low-complexity tracts (such as atatatatatat), which arise frequently and independently, causing strong similarities that are not homologies. There has been much research on identifying low-complexity tracts, but little research on how to treat them during homology search. We propose to find homologies by aligning sequences with "gentle" masking of low-complexity tracts. Gentle masking means that the match score involving a masked letter is min(0, S), where S is the unmasked score. Gentle masking slightly but noticeably improves the sensitivity of homology search (compared to "harsh" masking), without harming specificity. We show examples in three useful homology search problems: detection of NUMTs (nuclear copies of mitochondrial DNA), recruitment of metagenomic DNA reads to reference genomes, and pseudogene detection. Gentle masking is currently the best way to treat low-complexity tracts during homology search.