Parallel implementation of motif-based clustering for HT-SELEX dataset

Parallel implementation of motif-based clustering for HT-SELEX dataset
复制标题

HT-SELEX 数据集基于主题的聚类的并行实现

DOI:
10.1109/bibe.2019.00018
复制
发表时间:
2019
期刊:
Proceedings of IEEE International Conference on Bioinformatics and Bioengineering
影响因子:
--
通讯作者:
Takayoshi Ono
Takayoshi Ono
中科院分区:
--
文献类型:
--
作者:
川原田美雪;唐澤弘明;坂本美沙子;天野宗佑;山肩洋子;相澤清晴;Takayoshi Ono

文献摘要

相似文献

利用SELEX池进行高通量测序的聚类方法(HT-SELEX)对于选择不同类型的适体候选体至关重要。对于HT-SELSEX产生的海量序列数据,快速准确的聚类方法是必不可少的。我们已经开发了一种基于R语言实现的基于motif的HT-SELEX数据快速聚类方法(FMBC)。与传统方法相比,FMBC的序列聚类精度较高,但处理时间较AptaCluster长。为了提高FMBC的性能,本文提出了使用Python多线程并行实现FMBC。利用来自BioProject PRJNA315881的SRR3279661的NCBI SRA数据进行的实验评估表明,并行FMBC比传统方法具有更高的聚类精度和更短的处理时间。
A clustering method for high-throughput sequencing with SELEX pools (HT-SELEX) is crucial for selecting different types of aptamer candidates. The fast and accurate clustering method is indispensable for an enormous sequence data produced by HT-SELSEX. We have already developed a fast motif-based clustering (FMBC) method for HT-SELEX data implemented by R language. FMBC exhibited high accuracy of sequence clustering compared with conventional methods, while the processing time of FMBC is longer than AptaCluster. This paper proposes the parallel implementation of FMBC using Python with multi-threading to improve the performance of FMBC. Experimental evaluation using the NCBI SRA data of SRR3279661 from BioProject PRJNA315881 demonstrated that parallel FMBC exhibited higher accuracy of clustering and shorter processing time than conventional methods.