Simrank: Rapid and sensitive general-purpose k-mer search tool.

Simrank: Rapid and sensitive general-purpose k-mer search tool.
复制标题

DOI:
10.1186/1472-6785-11-11
复制
发表时间:
2011-04-27
期刊:
影响因子:
--
通讯作者:
Larsen N
Larsen N
中科院分区:
环境科学与生态学3区
文献类型:
--
作者:
DeSantis TZ;Keller K;Karaoz U;Alekseyenko AV;Singh NN;Brodie EL;Pei Z;Andersen GL;Larsen N

文献摘要

被引文献

相似文献

字符串编码数据的TB级集合预计来自诸如人类微生物组项目http://nihroadmap.nih.gov/hmp等联盟的努力。项目内和项目间数据相似性搜索通过快速k-mer匹配策略实现。用于序列数据库划分、指导树估计、分子分类和比对加速的软件应用已经受益于作为子例程的嵌入式k-mer搜索。然而,一个快速的,通用的,开源的,灵活的,独立的k-mer工具还没有。在这里,我们提出了一个独立的实用程序,Simrank,它允许用户快速识别数据库字符串最相似的查询字符串。Simrank和相关工具针对DNA、RNA、蛋白质和人类语言的性能测试发现,Simrank的速度快了10倍到928倍,具体取决于数据集。Simrank为分子生态学家提供了一个高通量、开源的选择,用于比较大型序列集以找到相似性。
Terabyte-scale collections of string-encoded data are expected from consortia efforts such as the Human Microbiome Project http://nihroadmap.nih.gov/hmp. Intra- and inter-project data similarity searches are enabled by rapid k-mer matching strategies. Software applications for sequence database partitioning, guide tree estimation, molecular classification and alignment acceleration have benefited from embedded k-mer searches as sub-routines. However, a rapid, general-purpose, open-source, flexible, stand-alone k-mer tool has not been available. Here we present a stand-alone utility, Simrank, which allows users to rapidly identify database strings the most similar to query strings. Performance testing of Simrank and related tools against DNA, RNA, protein and human-languages found Simrank 10X to 928X faster depending on the dataset. Simrank provides molecular ecologists with a high-throughput, open source choice for comparing large sequence sets to find similarity.