Novel algorithms for finding the closest l-mers in biological data
Novel algorithms for finding the closest l-mers in biological data
复制标题
用于查找生物数据中最接近的 l-mers 的新算法
DOI:
10.1109/bibm.2017.8217702
复制
发表时间:
2017
期刊:
影响因子:
--
通讯作者:
S. Rajasekaran
中科院分区:
文献类型:
--
作者:
Xingyu Cai;A. Mamun;S. Rajasekaran
With the advances in the next generation sequencing technology, huge amounts of data have been and get generated in biology. A bottleneck in dealing with such datasets lies in developing effective algorithms for extracting useful information from them. Algorithms for finding patterns in biological data pave the way for extracting crucial information from voluminous datasets. In this paper we focus on a fundamental pattern, namely, the closest l-mers. Given a set of m biological strings S1, S2, …, Sm and an integer l, the problem of interest is that of finding an l-mer from each string such that the distance among them is the least. I.e., we want to find m l-mers X1, X2, …, Xm such that Xi is an l-mer in Si (for 1 ≤ i ≤ m) and the Hamming distance among these m l-mers is the least (from among all such possible l-mers). This problem has many applications. An application of great importance is motif search. Algorithms for finding the closest l-mers have been used in solving the (l, d)-motif search problem (see e.g., [1], [2]). In this paper novel exact and approximate algorithms are proposed for this problem for the special case of m = 3. We consider the Euclidean distance metric if the sequences contain real numbers.