课题基金 / 基金详情

SOFTWARE: Framework for Mining Large and Complex Scientific Datasets

SOFTWARE: Framework for Mining Large and Complex Scientific Datasets
软件:挖掘大型复杂科学数据集的框架
批准号:
0234273
负责人:
Raghu Machiraju
金额:
$37.3万
依托单位国家:
美国
项目类别:
Continuing Grant
财政年份:
2003
资助国家:
美国
项目状态:
已结题
起止时间:
2003-09-15 至 2006-08-31

项目摘要

项目成果

Raghu Machiraju的其他基金

相似基金

相关文献

中文摘要
翻译
在深入了解复杂的物理现象方面,数值模拟正在取代传统的实验。鉴于计算机硬件和数值方法的最新进展,现在有可能以非常精细的时间和空间分辨率模拟物理现象。因此,产生的数据量是巨大的。科学家们对分析和可视化此类模拟产生的数据感兴趣,以更好地理解正在模拟的过程。分析如此大规模的数据是很困难的。不仅使用的方法计算昂贵,而且目前的编程工具使分析难以指定和修改。因此,迫切需要一种系统的方法,以及支持灵活并行实施的算法和方法,以实现对大型科学数据集的可伸缩和交互分析。在这个项目中,我们建议构建这样一个可扩展的工具包,即计算分析工具包(CAT)。该工具包建议利用特征分析、可伸缩数据挖掘和并行编程环境中正在进行的工作。该方法的关键是特征挖掘;这是一个通过检测、验证、去噪和跟踪兴趣点的不同阶段来按区域划分的过程。此外,我们建议使用一些关键的数据挖掘算法来实现增强的和健壮的特征挖掘算法的实现。我们的目标是,CAT工具包不仅应该允许检测特征,而且还应该提供一种在交互环境中控制分析的手段。例如,由用户/科学家确定的对某些关键特征的人口统计和生命周期分析可能是理解所模拟的潜在过程的重要方式。一旦通过合适的界面标记了这些关键特征,就可以对其进行描述,然后根据需要向用户呈现该描述的简明表示。我们认为,对于特征和数据挖掘工具的长期使用,a)算法在各种平台上是并行的,b)并行实现易于维护和修改,c)用户可以使用API来快速创建新挖掘算法的可伸缩实现,这一点很重要。我们建议通过使用和扩展本地开发的并行化框架来实现这些目标。这个框架被称为数据挖掘引擎快速实现框架(Freeride),它提供了高级API和运行时技术,以实现数据挖掘和相关任务的算法并行化。它允许分布式内存和共享内存配置上的并行化,并进一步支持磁盘驻留数据集的高效处理。该提议除了提供有用的工具包外,还可能产生用于大数据探索的方法。我们的努力可能会对可伸缩数据和特征挖掘算法以及特征轮廓摘要方面的文献做出贡献。
英文摘要
Numerical simulations are replacing traditional experiments in gaining insights into complex physical phenomena. Given recent advances in computer hardware and numerical methods, it is now possible to simulate physical phenomena at very fine temporal and spatial resolutions. As a result, the the amount of data generated is overwhelming. Scientists are interested in analyzing and visualizing the data produced by such simulations to better understand the process that is being simulated. Analyzing such large scale data is hard. Not only the methods used are computationally expense, current programming tools make the analysis difficult to specify and modify. Thus, there is a dire need for a systematic approach, along with supporting algorithms and methodologies for flexible parallel implementations, to achieve scalable and interactive analysis on large scientific datasets. In this project, we propose the construction of such a scalable toolkit, namely the Computational Analysis Toolkit (CAT). This toolkit proposes to exploit ongoing work in feature analysis, scalable data mining and parallel programing environments. The crux of the approach is feature-mining; a process where by regions are delineated through various stages of detection, verification, de-noising, and tracking of points of interest. Additionally, we propose the use of some key data mining mining algorithms for achieving enhanced and robust implementations of feature-mining algorithms. It is our objective that the CAT toolkit should not only allow for the detection of features but also provide for a means to control the analysis in an interactive setting. For example, demographic and lifetime analysis of certain critical features as determined by the user/scientist may be an important way of understanding the underlying process being simulated. These critical features, once tagged via a suitable interface, can be profiled and a concise representation this profile can then be presented to the user as needed. We believe that for long-term use of a tool for feature and data mining, it is important that a) the algorithms are parallelized on a variety of platforms, b) the parallel implementations are easy to maintain and modify, and c) APIs are available for users to rapidly create scalable implementations of new mining algorithms. We are proposing to achieve these goals by using and extending a parallelization framework developed locally. This framework, referred to as FRamework for Rapid Implementations of Datamining Engines (FREERIDE), offers high-level APIs and runtime techniques to enable parallelization of algorithms for data mining and related tasks. It allows parallelization on both distributed memory and shared memory configurations, and further supports efficient processing of disk-resident datasets.The proposal, besides providing a useful toolkit, is likely engender the use of methodologies for large data exploration. Our efforts are likely to contribute to literature in scalable data and feature mining algorithms, and feature profile summarization.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
Collaborative Research: Autonomous Computing Materials
  • 批准号:
    1940168
  • 项目类别:
    Continuing Grant
  • 资助金额:
    $35.66万
  • 财政年份:
    2019
  • 负责人:
    Raghu Machiraju
  • 依托单位:
Spokes: MEDIUM: MIDWEST: Collaborative: Community-Driven Data Engineering for Substance Abuse Prevention in the Rural Midwest
  • 批准号:
    1761969
  • 项目类别:
    Standard Grant
  • 资助金额:
    $65.1万
  • 财政年份:
    2018
  • 负责人:
    Raghu Machiraju
  • 依托单位:
SCC-Planning: Using Innovations in Big Data and Technology to Address the High Rate of Infant Mortality in Greater Columbus Ohio
  • 批准号:
    1737560
  • 项目类别:
    Standard Grant
  • 资助金额:
    $10.0万
  • 财政年份:
    2017
  • 负责人:
    Raghu Machiraju
  • 依托单位:
BCSP: ABI Innovation: Collaborative Research: Predicting changes in protein activity from changes in sequence by identifying the underlying Biophysical Conditional Random Field
  • 批准号:
    1262469
  • 项目类别:
    Standard Grant
  • 资助金额:
    $41.14万
  • 财政年份:
    2014
  • 负责人:
    Raghu Machiraju
  • 依托单位:
海外基金