Parallel implementation of motif-based clustering for HT-SELEX dataset
Parallel implementation of motif-based clustering for HT-SELEX dataset
复制标题
HT-SELEX 数据集基于主题的聚类的并行实现
DOI:
10.1109/bibe.2019.00018
复制
发表时间:
2019
期刊:
影响因子:
--
通讯作者:
Takayoshi Ono
中科院分区:
文献类型:
--
作者:
川原田美雪;唐澤弘明;坂本美沙子;天野宗佑;山肩洋子;相澤清晴;Takayoshi Ono
A clustering method for high-throughput sequencing with SELEX pools (HT-SELEX) is crucial for selecting different types of aptamer candidates. The fast and accurate clustering method is indispensable for an enormous sequence data produced by HT-SELSEX. We have already developed a fast motif-based clustering (FMBC) method for HT-SELEX data implemented by R language. FMBC exhibited high accuracy of sequence clustering compared with conventional methods, while the processing time of FMBC is longer than AptaCluster. This paper proposes the parallel implementation of FMBC using Python with multi-threading to improve the performance of FMBC. Experimental evaluation using the NCBI SRA data of SRR3279661 from BioProject PRJNA315881 demonstrated that parallel FMBC exhibited higher accuracy of clustering and shorter processing time than conventional methods.