BULK-LOADING & PERFORMANCE STUDIES OF THE ND-TREE FOR LARGE GENOME DATABASES
BULK-LOADING & PERFORMANCE STUDIES OF THE ND-TREE FOR LARGE GENOME DATABASES
批准号:
7610287
负责人:
GANG QIAN
金额:
$4.0万
依托单位国家:
美国
项目类别:
财政年份:
2007
资助国家:
美国
项目状态:
已结题
起止时间:
2007-05-01 至 2008-04-30
关键词:
AffectAlgorithmsBioinformaticsBiologicalComputer Retrieval of Information on Scientific Projects DatabaseDataData SetDatabasesEffectivenessFundingFutureGrantInstitutionMemoryMethodsPerformanceResearchResearch PersonnelResourcesSourceSpeedStructureTechniquesTimeTreesUnited States National Institutes of Healthbasedesigndiscrete datagenome databaseindexingvector
中文摘要
点击翻译按钮获取中文摘要
英文摘要
This subproject is one of many research subprojects utilizing the
resources provided by a Center grant funded by NIH/NCRR. The subproject and
investigator (PI) may have received primary funding from another NIH source,
and thus could be represented in other CRISP entries. The institution listed is
for the Center, which is not necessarily the institution for the investigator.
The subproject seeks to provide an efficient indexing method to speed up the search of large biological information databases. In particular, the research is based on a multi-dimensional disk-based index structure, called the ND-tree, which is designed to support similarity queries on vectors/q-grams of large non-ordered discrete data sets. The current method used to construct the ND-tree is incremental, which may significantly affect the effective use of the index due to the huge amount of data to be indexed. The subproject focuses on finding an efficient algorithm to bulkload the ND-tree. Unlike the incremental method, the new bulkloading algorithm assumes that there is some memory space available for bulkloading. Therefore, it is possible for the algorithm to load thousands of vectors into the index structure without incurring a single disk I/O, resulting in a significant reduction in the loading time. The algorithm is also designed such that a bulkloaded ND-tree has a comparable query performance to those incrementally constructed. To evaluate the effectiveness of the new algorithm, it will be experimentally compared with the incremental method and other existing bulkloading methods in terms of both loading and querying efficiency. A theoretical analysis of the bulkloading algorithm is planned for future research. Furthermore, the bulkloading algorithm will become an integrated part of a planned index-based bioinformatics search engine in future research.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
SUBSTITUTION MATRICES INTO THE NSP-TREE IN BIOLOGICAL SEQUENCE DATABASES
-
批准号:8167540
-
项目类别:
-
资助金额:$2.97万
-
财政年份:2010
-
负责人:GANG QIAN
-
依托单位:
USE THE EDIT DISTANCE IN THE ND-TREE FOR EFFICIENT BIOINFORMATICS QUERIES
-
批准号:7960025
-
项目类别:
-
资助金额:$2.91万
-
财政年份:2009
-
负责人:GANG QIAN
-
依托单位:
USE THE EDIT DISTANCE IN THE ND-TREE FOR EFFICIENT BIOINFORMATICS QUERIES
-
批准号:7725103
-
项目类别:
-
资助金额:$2.64万
-
财政年份:2008
-
负责人:GANG QIAN
-
依托单位:
海外基金