EAGER: Autonomous Data Partitioning Using Data Mining for High End Computing
EAGER: Autonomous Data Partitioning Using Data Mining for High End Computing
批准号:
0954310
负责人:
Sudarshan Dhall
金额:
$12.5万
依托单位国家:
美国
项目类别:
Standard Grant
财政年份:
2009
资助国家:
美国
项目状态:
已结题
起止时间:
2009-09-01 至 2016-08-31
中文摘要
当涉及到数据库和文件访问性能时,查询响应时间和系统吞吐量是最重要的指标。由于数据激增,高效的访问方法和数据存储技术对于保持可接受的查询响应时间和系统吞吐量变得越来越关键。减少磁盘I/O从而提高查询响应时间的常用方法之一是数据库集群,它是对数据库/文件进行垂直(属性集群)和/或水平(记录集群)分区的过程。为了利用并行性来提高系统吞吐量,可以将集群放置在集群机器中的不同节点上。该项目开发了一种用于数据库/文件聚类的新算法AutoCluust,该算法基于从针对数据库/文件运行的查询中发现的属性和记录集挖掘的封闭项目集来动态和自动地生成属性和记录聚类。该算法能够对数据库/文件进行重新聚集,以便在数据和/或查询集合发生变化的情况下继续实现良好的系统性能。然后,该项目开发了实施AutoCluust的创新方法,使用集群计算范例,通过并行性和数据冗余进一步减少查询响应时间和系统吞吐量。这些算法是在俄克拉荷马大学拥有486个计算节点的戴尔Linux集群计算机上构建的原型。对于更广泛的影响,性能研究不仅使用决策支持系统数据库基准(TPC-H),而且与包括俄克拉荷马大学风暴分析和预测中心(CAPS)的科学家在内的领域专家合作,使用从科学和医疗保健应用程序收集的数据库和文件格式记录的真实数据。该项目还对教育产生了重要影响,因为它为从事该项目的研究生和本科生提供了国家关键需求领域的培训:数据库和文件管理系统,以及高端计算和应用程序。开发的算法和原型、实际数据集和性能评估结果已在以下网站向公众公布:http://www.cs.ou.edu/~database/AutoClust.html.
英文摘要
Query response time and system throughput are the most important metrics when it comes to database and file access performance. Because of data proliferation, efficient access methods and data storage techniques have become increasingly critical to maintain an acceptable query response time and system throughput. One of the common ways to reduce disk I/Os and therefore improve query response time is database clustering, which is a process that partitions the database/file vertically (attribute clustering) and/or horizontally (record clustering). To take advantage of parallelism to improve system throughput, clusters can be placed on different nodes in a cluster machine. This project develops a novel algorithm, AutoClust, for database/file clustering that dynamically and automatically generates attribute and record clusters based on closed item sets mined from the attributes and records sets found in the queries running against the database/files. The algorithm is capable of re-clustering the database/file in order to continue achieving good system performance despite changes in the data and/or query sets. The project then develops innovative ways to implement AutoClust using the cluster computing paradigm to reduce query response time and system throughput even further through parallelism and data redundancy. The algorithms are prototyped on a Dell Linux Cluster computer with 486 compute nodes available at the University of Oklahoma. For broader impacts, performance studies are conducted using not only the decision support system database benchmark (TPC-H) but also real data recorded in database and file formats collected from science and healthcare applications in collaboration with domain experts, including scientists at the Center for Analysis and Prediction of Storms (CAPS) at the University of Oklahoma. The project also makes important impacts on education as it provides training for graduate and undergraduate students working on this project in the areas of national critical needs: database and file management systems, and high-end computing and applications. The developed algorithm and prototype, real datasets and performance evaluation results are made available to the public at the Website: http://www.cs.ou.edu/~database/AutoClust.html.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
A Power-Aware Technique to Manage Real-Time Database Transactions in Mobile Ad-Hoc Networks
-
批准号:0312746
-
项目类别:Standard Grant
-
资助金额:$40.0万
-
财政年份:2003
-
负责人:Sudarshan Dhall
-
依托单位:
A Workshop on Parallel Processing Using the Heterogeneous Element Processor (HEP), March 20-21, 1985, at the University of Oklahoma, Norman, Oklahoma
-
批准号:8500481
-
项目类别:Standard Grant
-
资助金额:$0.5万
-
财政年份:1985
-
负责人:Sudarshan Dhall
-
依托单位:
海外基金