课题基金 / 基金详情

BULK-LOADING & PERFORMANCE STUDIES OF THE ND-TREE FOR LARGE GENOME DATABASES

BULK-LOADING & PERFORMANCE STUDIES OF THE ND-TREE FOR LARGE GENOME DATABASES
散装
批准号:
7610287
负责人:
GANG QIAN
金额:
$4.0万
依托单位国家:
美国
项目类别:
财政年份:
2007
资助国家:
美国
项目状态:
已结题
起止时间:
2007-05-01 至 2008-04-30

项目摘要

项目成果

GANG QIAN的其他基金

相似基金

相关文献

中文摘要
翻译
点击翻译按钮获取中文摘要
英文摘要
This subproject is one of many research subprojects utilizing the resources provided by a Center grant funded by NIH/NCRR. The subproject and investigator (PI) may have received primary funding from another NIH source, and thus could be represented in other CRISP entries. The institution listed is for the Center, which is not necessarily the institution for the investigator. The subproject seeks to provide an efficient indexing method to speed up the search of large biological information databases. In particular, the research is based on a multi-dimensional disk-based index structure, called the ND-tree, which is designed to support similarity queries on vectors/q-grams of large non-ordered discrete data sets. The current method used to construct the ND-tree is incremental, which may significantly affect the effective use of the index due to the huge amount of data to be indexed. The subproject focuses on finding an efficient algorithm to bulkload the ND-tree. Unlike the incremental method, the new bulkloading algorithm assumes that there is some memory space available for bulkloading. Therefore, it is possible for the algorithm to load thousands of vectors into the index structure without incurring a single disk I/O, resulting in a significant reduction in the loading time. The algorithm is also designed such that a bulkloaded ND-tree has a comparable query performance to those incrementally constructed. To evaluate the effectiveness of the new algorithm, it will be experimentally compared with the incremental method and other existing bulkloading methods in terms of both loading and querying efficiency. A theoretical analysis of the bulkloading algorithm is planned for future research. Furthermore, the bulkloading algorithm will become an integrated part of a planned index-based bioinformatics search engine in future research.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
SUBSTITUTION MATRICES INTO THE NSP-TREE IN BIOLOGICAL SEQUENCE DATABASES
USE THE EDIT DISTANCE IN THE ND-TREE FOR EFFICIENT BIOINFORMATICS QUERIES
USE THE EDIT DISTANCE IN THE ND-TREE FOR EFFICIENT BIOINFORMATICS QUERIES
海外基金