BULK-LOADING & PERFORMANCE STUDIES OF THE ND-TREE FOR LARGE GENOME DATABASES
BULK-LOADING & PERFORMANCE STUDIES OF THE ND-TREE FOR LARGE GENOME DATABASES
批准号:
7610287
负责人:
GANG QIAN
金额:
$4.0万
依托单位国家:
美国
项目类别:
财政年份:
2007
资助国家:
美国
项目状态:
已结题
起止时间:
2007-05-01 至 2008-04-30
关键词:
AffectAlgorithmsBioinformaticsBiologicalComputer Retrieval of Information on Scientific Projects DatabaseDataData SetDatabasesEffectivenessFundingFutureGrantInstitutionMemoryMethodsPerformanceResearchResearch PersonnelResourcesSourceSpeedStructureTechniquesTimeTreesUnited States National Institutes of Healthbasedesigndiscrete datagenome databaseindexingvector
中文摘要
这个子项目是许多研究子项目中利用
资源由NIH/NCRR资助的中心拨款提供。子项目和
调查员(PI)可能从NIH的另一个来源获得了主要资金,
并因此可以在其他清晰的条目中表示。列出的机构是
该中心不一定是调查人员的机构。
该分项目寻求提供一种有效的索引方法,以加快对大型生物信息数据库的搜索。特别是,该研究基于一种称为ND-树的多维磁盘索引结构,该结构被设计用于支持对大型无序离散数据集的向量/Q-gram的相似性查询。目前用于构建ND-树的方法是增量的,由于要索引的数据量巨大,这可能会显著影响索引的有效使用。该子项目的重点是寻找一种有效的算法来批量加载ND-树。与增量方法不同,新的批量加载算法假设有一些内存空间可用于批量加载。因此,该算法可以在不引起单个磁盘I/O的情况下将数千个向量加载到索引结构中,从而显著减少加载时间。该算法还被设计成使得块加载的ND-树具有与增量构造的ND-树相当的查询性能。为了评估新算法的有效性,将其与增量法和其他现有的批量加载方法在加载和查询效率方面进行了实验比较。计划对批量加载算法进行理论分析,以供进一步研究。此外,在未来的研究中,批量加载算法将成为计划中的基于索引的生物信息学搜索引擎的组成部分。
英文摘要
This subproject is one of many research subprojects utilizing the
resources provided by a Center grant funded by NIH/NCRR. The subproject and
investigator (PI) may have received primary funding from another NIH source,
and thus could be represented in other CRISP entries. The institution listed is
for the Center, which is not necessarily the institution for the investigator.
The subproject seeks to provide an efficient indexing method to speed up the search of large biological information databases. In particular, the research is based on a multi-dimensional disk-based index structure, called the ND-tree, which is designed to support similarity queries on vectors/q-grams of large non-ordered discrete data sets. The current method used to construct the ND-tree is incremental, which may significantly affect the effective use of the index due to the huge amount of data to be indexed. The subproject focuses on finding an efficient algorithm to bulkload the ND-tree. Unlike the incremental method, the new bulkloading algorithm assumes that there is some memory space available for bulkloading. Therefore, it is possible for the algorithm to load thousands of vectors into the index structure without incurring a single disk I/O, resulting in a significant reduction in the loading time. The algorithm is also designed such that a bulkloaded ND-tree has a comparable query performance to those incrementally constructed. To evaluate the effectiveness of the new algorithm, it will be experimentally compared with the incremental method and other existing bulkloading methods in terms of both loading and querying efficiency. A theoretical analysis of the bulkloading algorithm is planned for future research. Furthermore, the bulkloading algorithm will become an integrated part of a planned index-based bioinformatics search engine in future research.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
SUBSTITUTION MATRICES INTO THE NSP-TREE IN BIOLOGICAL SEQUENCE DATABASES
-
批准号:8167540
-
项目类别:
-
资助金额:$2.97万
-
财政年份:2010
-
负责人:GANG QIAN
-
依托单位:
USE THE EDIT DISTANCE IN THE ND-TREE FOR EFFICIENT BIOINFORMATICS QUERIES
-
批准号:7960025
-
项目类别:
-
资助金额:$2.91万
-
财政年份:2009
-
负责人:GANG QIAN
-
依托单位:
USE THE EDIT DISTANCE IN THE ND-TREE FOR EFFICIENT BIOINFORMATICS QUERIES
-
批准号:7725103
-
项目类别:
-
资助金额:$2.64万
-
财政年份:2008
-
负责人:GANG QIAN
-
依托单位:
海外基金