keeSeek: searching distant non-existing words in genomes for PCR-based applications

keeSeek: searching distant non-existing words in genomes for PCR-based applications
复制标题

DOI:
10.1093/bioinformatics/btu312
复制
发表时间:
2014-09-15
期刊:
影响因子:
5.8
通讯作者:
Lavezzo, Enrico
Lavezzo, Enrico
中科院分区:
生物学3区
文献类型:
--
作者:
Falda, Marco;Fontana, Paolo;Lavezzo, Enrico

文献摘要

被引文献

相似文献

一个总结:寻找一个或多个生物体基因组中不存在的短词(neverwords,也称为nullomers)正吸引越来越多的兴趣,因为它们可能在最近的分子生物学应用中产生影响。keeSeek能够找到具有引物样特征的缺失序列,这些序列可用作外源插入DNA片段的独特标记,以使用PCR技术恢复其在基因组中的确切位置。相对于先前开发的用于neverwords生成的工具的主要差异是(i)计算与参考基因组的距离,根据错配的数量,并选择具有低概率非特异性退火的最远序列;(ii)应用一系列过滤器以丢弃不适合用作PCR引物的候选物。KeeSeek已在C++和CUDA(计算统一设备架构)中实现,可在图形处理单元(GPGPU)环境中工作。
A Summary: The search for short words that are absent in the genome of one or more organisms (neverwords, also known as nullomers) is attracting growing interest because of the impact they may have in recent molecular biology applications. keeSeek is able to find absent sequences with primer-like features, which can be used as unique labels for exogenously inserted DNA fragments to recover their exact position into the genome using PCR techniques. The main differences with respect to previously developed tools for neverwords generation are (i) calculation of the distance from the reference genome, in terms of number of mismatches, and selection of the most distant sequences that will have a low probability to anneal unspecifically; (ii) application of a series of filters to discard candidates not suitable to be used as PCR primers. KeeSeek has been implemented in C++ and CUDA (Compute Unified Device Architecture) to work in a General-Purpose Computing on Graphics Processing Units (GPGPU) environment.