课题基金 / 基金详情

Developing fast algorithm for analyzing Giga-sequence data

Developing fast algorithm for analyzing Giga-sequence data
开发用于分析千兆序列数据的快速算法
批准号:
22700319
负责人:
SHIMIZU Kana
金额:
$2.08万
依托单位国家:
日本
项目类别:
Grant-in-Aid for Young Scientists (B)
财政年份:
2010
资助国家:
日本
项目状态:
已结题
起止时间:
2010 至 2011

项目摘要

项目成果

相似基金

相关文献

中文摘要
翻译
下一代测序(NGS)技术需要快速准确的算法来评估大量数据的序列相似性。在这项研究中,我们设计并实现了精确的算法SlideSort,从NGS数据中找到编辑距离不超过给定阈值的所有相似对,这有助于许多重要的分析,如从头基因组组装,识别频繁出现的序列模式和准确聚类。与最先进的方法相比,我们的方法在寻找远程匹配方面要快得多,可以轻松扩展到数千万个序列。我们的软件增加了一个单链路聚类功能,这对总结NGS数据进行进一步处理很有用。
英文摘要
Next Generation Sequencing(NGS) technology calls for fast and accurate algorithms that can evaluate sequence similarity for a huge amount data. In this study, we designed and implemented exact algorithm SlideSort that finds all similar pairs whose edit-distance does not exceed a given threshold from NGS data, which helps many important analyses, such as de novo genome assembly, identification of frequently appearing sequence patterns and accurate clustering. In comparison to state-of-the-art methods, our method is much faster in finding remote matches, scaling easily to tens of millions of sequences. Our software has an additional function of single link clustering, which is useful in summarizing NGS data for further processing.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
SlideSort: all pairs similarity search for short reads.
SlideSort:所有对相似性搜索短读。
DOI: 10.1093/bioinformatics/btq677
发表时间: 2011-02-15
期刊: Bioinformatics (Oxford, England)
影响因子: --
作者: [Shimizu K, Tsuda K]
通讯作者: Tsuda K
海外基金