课题基金 / 基金详情

Developing fast algorithm for analyzing Giga-sequence data

Developing fast algorithm for analyzing Giga-sequence data
开发用于分析千兆序列数据的快速算法
批准号:
22700319
负责人:
SHIMIZU Kana
金额:
$2.08万
依托单位国家:
日本
项目类别:
Grant-in-Aid for Young Scientists (B)
财政年份:
2010
资助国家:
日本
项目状态:
已结题
起止时间:
2010 至 2011

项目摘要

项目成果

相似基金

相关文献

中文摘要
翻译
下一代测序(NGS)技术需要快速、准确的算法来评估海量数据的序列相似性。在这项研究中,我们设计并实现了精确算法SlideSort,它可以从NGS数据中找到编辑距离不超过给定阈值的所有相似对,这有助于许多重要的分析,如从头组装基因组、识别频繁出现的序列模式和准确的聚类。与最先进的方法相比,我们的方法在查找远程匹配方面要快得多,很容易扩展到数千万个序列。我们的软件还增加了单链路聚类功能,这对于汇总NGS数据进行进一步处理很有用。
英文摘要
Next Generation Sequencing(NGS) technology calls for fast and accurate algorithms that can evaluate sequence similarity for a huge amount data. In this study, we designed and implemented exact algorithm SlideSort that finds all similar pairs whose edit-distance does not exceed a given threshold from NGS data, which helps many important analyses, such as de novo genome assembly, identification of frequently appearing sequence patterns and accurate clustering. In comparison to state-of-the-art methods, our method is much faster in finding remote matches, scaling easily to tens of millions of sequences. Our software has an additional function of single link clustering, which is useful in summarizing NGS data for further processing.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
SlideSort: all pairs similarity search for short reads.
SlideSort:所有对相似性搜索短读。
DOI: 10.1093/bioinformatics/btq677
发表时间: 2011-02-15
期刊: Bioinformatics (Oxford, England)
影响因子: --
作者: [Shimizu K, Tsuda K]
通讯作者: Tsuda K
海外基金