课题基金 / 基金详情

USE THE EDIT DISTANCE IN THE ND-TREE FOR EFFICIENT BIOINFORMATICS QUERIES

USE THE EDIT DISTANCE IN THE ND-TREE FOR EFFICIENT BIOINFORMATICS QUERIES
使用 ND 树中的编辑距离进行高效的生物信息学查询
批准号:
7725103
负责人:
GANG QIAN
金额:
$2.64万
依托单位国家:
美国
项目类别:
财政年份:
2008
资助国家:
美国
项目状态:
已结题
起止时间:
2008-05-01 至 2009-04-30

项目摘要

项目成果

GANG QIAN的其他基金

相似基金

相关文献

中文摘要
翻译
这个子项目是许多研究子项目中利用 资源由NIH/NCRR资助的中心拨款提供。子项目和 调查员(PI)可能从NIH的另一个来源获得了主要资金, 并因此可以在其他清晰的条目中表示。列出的机构是 该中心不一定是调查人员的机构。 随着生物数据量的快速增长,基于索引的方法搜索数据变得比基于顺序扫描的方法更有利。这个子项目调查ND-tree在生物信息学查询中的应用,ND-tree是一种专门设计的多维结构,用于索引具有生物信息学数据典型的离散和无序成分的子串/Q图。这个子项目的目的是扩展ND-树以支持编辑距离,编辑距离是同源区域查询的一种广泛使用的相似性度量。扩展的目标是提高生物信息学数据库查询过滤阶段的敏感度。为了结合使用比汉明距离更多的插入和删除操作的编辑距离,ND-树必须支持具有相对较大搜索范围的高效相似性查询。在这个项目的第一阶段,我们将设计和评估新的算法,这些算法可以有效地处理ND-树中搜索范围相对较大的查询。我们计划研究基于近似的技术,通过修剪大量前景较差的索引分支来提高查询性能。在第二阶段,将开发一种基于编辑距离的查询算法。为了进一步提高性能,还将检查和调整ND-树的构建和批量加载算法,使索引中的数据组织变得更适合编辑距离查询。为了评估新算法的有效性,我们将通过实验将它们与现有算法进行比较。该项目将导致一个基于ND-树的新型生物信息学搜索引擎的设计。
英文摘要
This subproject is one of many research subprojects utilizing the resources provided by a Center grant funded by NIH/NCRR. The subproject and investigator (PI) may have received primary funding from another NIH source, and thus could be represented in other CRISP entries. The institution listed is for the Center, which is not necessarily the institution for the investigator. As the volume of biological data increases rapidly, index-based approaches to searching the data becomes more favorable than sequential-scan-based approaches. This subproject investigates the application of the ND-tree, a multidimensional structure specifically designed to index substrings/q-grams with discrete and non-ordered components typical of bioinformatics data, to bioinformatics queries. The aim of this subproject is to extend the ND-tree to support the edit distance, a widely-used similarity measure for homologous region queries. The goal of the extension is to enhance the sensitivity in the filtering stage of a bioinformatics database query. To incorporate the edit distance which employs the extra insertion and deletion operations than the Hamming distance, the ND-tree must support efficient similarity queries with a relatively large search range. In the first phase of this project, we will design and evaluate novel algorithms that efficiently process queries with relatively large search ranges in the ND-tree. We plan to investigate approximation-based techniques that can improve query performance by pruning a large amount of less-promising index branches. In the second phase, a query algorithm based on the edit distance will be developed. To further enhance the performance, the construction and bulk-loading algorithms of the ND-tree will also be examined and adapted so that the data organization within the index becomes more suitable for edit distance queries. To evaluate the effectiveness of the new algorithms, we will experimentally compare them with existing algorithms. The project will lead to the design of a novel bioinformatics search engine based on the ND-tree.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
SUBSTITUTION MATRICES INTO THE NSP-TREE IN BIOLOGICAL SEQUENCE DATABASES
USE THE EDIT DISTANCE IN THE ND-TREE FOR EFFICIENT BIOINFORMATICS QUERIES
BULK-LOADING & PERFORMANCE STUDIES OF THE ND-TREE FOR LARGE GENOME DATABASES
海外基金