课题基金 / 基金详情

SOFTWARE: Framework for Mining Large and Complex Scientific Datasets

SOFTWARE: Framework for Mining Large and Complex Scientific Datasets
软件:挖掘大型复杂科学数据集的框架
批准号:
0234273
负责人:
Raghu Machiraju
金额:
$37.3万
依托单位国家:
美国
项目类别:
Continuing Grant
财政年份:
2003
资助国家:
美国
项目状态:
已结题
起止时间:
2003-09-15 至 2006-08-31

项目摘要

项目成果

Raghu Machiraju的其他基金

相似基金

相关文献

中文摘要
翻译
数值模拟正在取代传统的实验,以深入了解复杂的物理现象。鉴于计算机硬件和数值方法的最新进展,现在可以以非常精细的时间和空间分辨率模拟物理现象。因此,产生的数据量是压倒性的。科学家们有兴趣分析和可视化这些模拟产生的数据,以更好地理解正在模拟的过程。分析如此大规模的数据是非常困难的。不仅使用的方法是计算费用,目前的编程工具使分析难以指定和修改。因此,迫切需要一种系统的方法,沿着支持灵活的并行实现的算法和方法,以实现对大型科学数据集的可扩展和交互式分析。在这个项目中,我们提出了这样一个可扩展的工具包,即计算分析工具包(CAT)的建设。该工具包建议利用正在进行的工作,功能分析,可扩展的数据挖掘和并行编程环境。该方法的关键是特征挖掘;通过检测,验证,去噪和跟踪感兴趣点的各个阶段来划分区域的过程。 此外,我们建议使用一些关键的数据挖掘挖掘算法,以实现增强和强大的实现功能挖掘算法。这是我们的目标,CAT工具包应该不仅允许检测的功能,但也提供了一种手段来控制在一个互动的设置分析。 例如,由用户/科学家确定的某些关键特征的人口统计和寿命分析可能是理解正在模拟的基础过程的重要方式。一旦通过合适的界面标记这些关键特征,就可以对其进行剖析,然后可以根据需要将该剖析的简明表示呈现给用户。我们认为,对于长期使用的功能和数据挖掘工具,重要的是a)算法在各种平台上并行化,B)并行实现易于维护和修改,以及c)API可供用户快速创建新的挖掘算法的可扩展实现。我们建议通过使用和扩展本地开发的并行化框架来实现这些目标。这个框架被称为数据挖掘引擎快速实现框架(FRamework for Rapid Implementations of Datamining Engines,FREERIDE),提供了高级API和运行时技术,以实现数据挖掘和相关任务的算法并行化。 它允许在分布式内存和共享内存配置上并行化,并进一步支持磁盘驻留数据集的有效处理。 使用大数据探索的方法。 我们的努力很可能有助于文献中的可扩展数据和特征挖掘算法,和功能配置文件总结。
英文摘要
Numerical simulations are replacing traditional experiments in gaining insights into complex physical phenomena. Given recent advances in computer hardware and numerical methods, it is now possible to simulate physical phenomena at very fine temporal and spatial resolutions. As a result, the the amount of data generated is overwhelming. Scientists are interested in analyzing and visualizing the data produced by such simulations to better understand the process that is being simulated. Analyzing such large scale data is hard. Not only the methods used are computationally expense, current programming tools make the analysis difficult to specify and modify. Thus, there is a dire need for a systematic approach, along with supporting algorithms and methodologies for flexible parallel implementations, to achieve scalable and interactive analysis on large scientific datasets. In this project, we propose the construction of such a scalable toolkit, namely the Computational Analysis Toolkit (CAT). This toolkit proposes to exploit ongoing work in feature analysis, scalable data mining and parallel programing environments. The crux of the approach is feature-mining; a process where by regions are delineated through various stages of detection, verification, de-noising, and tracking of points of interest. Additionally, we propose the use of some key data mining mining algorithms for achieving enhanced and robust implementations of feature-mining algorithms. It is our objective that the CAT toolkit should not only allow for the detection of features but also provide for a means to control the analysis in an interactive setting. For example, demographic and lifetime analysis of certain critical features as determined by the user/scientist may be an important way of understanding the underlying process being simulated. These critical features, once tagged via a suitable interface, can be profiled and a concise representation this profile can then be presented to the user as needed. We believe that for long-term use of a tool for feature and data mining, it is important that a) the algorithms are parallelized on a variety of platforms, b) the parallel implementations are easy to maintain and modify, and c) APIs are available for users to rapidly create scalable implementations of new mining algorithms. We are proposing to achieve these goals by using and extending a parallelization framework developed locally. This framework, referred to as FRamework for Rapid Implementations of Datamining Engines (FREERIDE), offers high-level APIs and runtime techniques to enable parallelization of algorithms for data mining and related tasks. It allows parallelization on both distributed memory and shared memory configurations, and further supports efficient processing of disk-resident datasets.The proposal, besides providing a useful toolkit, is likely engender the use of methodologies for large data exploration. Our efforts are likely to contribute to literature in scalable data and feature mining algorithms, and feature profile summarization.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
Collaborative Research: Autonomous Computing Materials
  • 批准号:
    1940168
  • 项目类别:
    Continuing Grant
  • 资助金额:
    $35.66万
  • 财政年份:
    2019
  • 负责人:
    Raghu Machiraju
  • 依托单位:
Spokes: MEDIUM: MIDWEST: Collaborative: Community-Driven Data Engineering for Substance Abuse Prevention in the Rural Midwest
  • 批准号:
    1761969
  • 项目类别:
    Standard Grant
  • 资助金额:
    $65.1万
  • 财政年份:
    2018
  • 负责人:
    Raghu Machiraju
  • 依托单位:
SCC-Planning: Using Innovations in Big Data and Technology to Address the High Rate of Infant Mortality in Greater Columbus Ohio
  • 批准号:
    1737560
  • 项目类别:
    Standard Grant
  • 资助金额:
    $10.0万
  • 财政年份:
    2017
  • 负责人:
    Raghu Machiraju
  • 依托单位:
BCSP: ABI Innovation: Collaborative Research: Predicting changes in protein activity from changes in sequence by identifying the underlying Biophysical Conditional Random Field
  • 批准号:
    1262469
  • 项目类别:
    Standard Grant
  • 资助金额:
    $41.14万
  • 财政年份:
    2014
  • 负责人:
    Raghu Machiraju
  • 依托单位:
海外基金