Efficient Algorithms for Finding the Closest l-mers in Biological Data
Efficient Algorithms for Finding the Closest l-mers in Biological Data
复制标题
寻找生物数据中最接近的 l-mers 的有效算法
DOI:
10.1109/tcbb.2018.2843364
复制
发表时间:
2018
期刊:
影响因子:
--
通讯作者:
Rajasekaran, Sanguthevar
中科院分区:
文献类型:
--
作者:
Cai, Xingyu;Mamun, Abdullah-Al;Rajasekaran, Sanguthevar
With the advances in the next generation sequencing technology, huge amounts of data have been and get generated in biology. A bottleneck in dealing with such datasets lies in developing effective algorithms for extracting useful information from them. Algorithms for finding patterns in biological data pave the way for extracting crucial information from the voluminous datasets. In this paper, we focus on a fundamental pattern, namely, the closest-mers. Given a set ofbiological stringsand an integer, the problem of interest is that of finding an-mer from each string such that the distance among them is the least. For example we want to find-merssuch thatis an-mer in(for) and the Hamming distance among these-mers is the least (from among all such possible-mers). This problem has many applications. An application of great importance is motif search. Algorithms for finding the closest-mers have been used in solving the-motif search problem (see e.g., , ). In this paper novel exact and approximate algorithms are proposed for this problem for the case of. In particular, a comprehensive experimental evaluation is performed for, along with a further empirical study ofand 5. We also extend our solution to euclidean distance measurement metric if the sequences contain real numbers.