Discovering Motifs in Biological Sequences Using the Micron Automata Processor

Discovering Motifs in Biological Sequences Using the Micron Automata Processor
复制标题

使用微米自动机处理器发现生物序列中的基序

DOI:
--
复制
发表时间:
2016
期刊:
IEEE/ACM Transactions on Computational Biology & Bioinformatics
影响因子:
--
通讯作者:
S. Aluru
S. Aluru
中科院分区:
--
文献类型:
--
作者:
Indranil Roy;S. Aluru

文献摘要

参考文献

被引文献

相似文献

在多个DNA或蛋白质序列之间找到大约保守的序列,称为基序是计算生物学中的重要问题。在本文中,我们考虑了(L; d)基序搜索问题,即在至少n给定序列的至少q中识别一个或多个长度为l的基序,每种序列的发生与大多数替代的基序都不同。该问题已知是NP完整的,迄今为止报告的最大解决实例为(26; 11)。我们使用大量非确定性有限自动机(NFA)上的流执行(NFA)提出了一种针对(L; d)基序搜索问题的新型算法。该解决方案旨在利用微米自动机处理器,这是一种靠近部署的新技术,可以同时并行执行多个NFA。我们通过估计问题实例(39; 18)和(40; 17)的运行时间来证明(L; d)图案搜索问题的更大实例,以解决(L; d)图案搜索问题的更大实例。该论文是使用这种新的加速器技术解决问题的有用指南。
Finding approximately conserved sequences, called motifs, across multiple DNA or protein sequences is an important problem in computational biology. In this paper, we consider the (l; d) motif search problem of identifying one or more motifs of length l present in at least q of the n given sequences, with each occurrence differing from the motif in at most d substitutions. The problem is known to be NP-complete, and the largest solved instance reported to date is (26;11). We propose a novel algorithm for the (l; d) motif search problem using streaming execution over a large set of non-deterministic finite automata (NFA). This solution is designed to take advantage of the micron automata processor, a new technology close to deployment that can simultaneously execute multiple NFA in parallel. We demonstrate the capability for solving much larger instances of the (l; d) motif search problem using the resources available within a single automata processor board, by estimating run-times for problem instances (39; 18) and (40; 17). The paper serves as a useful guide to solving problems using this new accelerator technology.
DOI: 10.1126/science.8211139
发表时间: 1993-10-08
期刊: SCIENCE
影响因子: 56.9
作者:
LAWRENCE, CE;ALTSCHUL, SF;WOOTTON, JC
通讯作者: WOOTTON, JC