Scalable Visualization and Model Building
Scalable Visualization and Model Building
批准号:
0937123
负责人:
William Cleveland
金额:
$50.0万
依托单位:
依托单位国家:
美国
项目类别:
Standard Grant
财政年份:
2009
资助国家:
美国
项目状态:
已结题
起止时间:
2009-09-01 至 2013-08-31
中文摘要
标题:可扩展的可视化和模型构建PI:克利夫兰,威廉S。开发能够预测和解释数据模式的新算法、可视化工具和数学模型是机器学习和统计学的基础。 它们实现了对科学和工程至关重要的预测建模。可视化在数据分析的所有阶段都是至关重要的,从需要数据检查和清理时收集数据的那一刻,到结果的最终呈现。可视化通过允许分析师批判性地评估模型的预测能力,并诊断在拟合数据中的模式时存在的问题,从而促进了模型的构建。研究者们正在进行数据模式描述的方法、方法和模型的研究,重点是可视化和大规模数据集的综合分析。研究涉及两大主题。一个是可视化分析和统计建模的集成框架。 我们设想一个系统,设施的迭代建模过程。 建模周期包括多个阶段,从描述性可视化开始,然后是模型选择、模型拟合、诊断和评估,最后是迭代模型细化。 第二个主题是从小型数据集扩展到大型数据集的可视化和建模的一般方法,以及专门用于扩展数据可视化的新方法的开发。 我们通过将数据划分为子集,对子集进行采样,并对每个子集应用建模和可视化来实现缩放。调查人员正在国土安全领域两个具有挑战性的数据分析项目的背景下进行研究:(1)印第安纳州公共卫生应急监测系统76个紧急部门的每日投诉数量;(2)我们在普渡大学和斯坦福大学校园收集的网络安全互联网数据包跟踪。
英文摘要
Title: Scalable Visualization and Model BuildingPI: Cleveland, William S. [Purdue University]Developing new algorithms, visualization tools, and mathematical models that can predict and explain patterns in data is fundamental to machine learning and statistics. They enable a predictive modeling that is fundamental to science and engineering. Visualization is critical in all phases of data analysis, from the moment the data are collected when data checking and cleaning are needed, to the final presentation of results. Visualization facilitates model building by allowing the analyst to critically assess the predictive power of a model, and to diagnose problems in fitting the patterns in the data. The investigators are carrying out research in approaches, methods, and models for describing patterns in data with a strong emphasis on visualization and on comprehensive analysis of massive datasets.The research is addressing two broad topics. One is a framework for the integration of visual analysis and statistical modeling. We envision a system that facilities an iterative modeling process. The modeling cycle includes multiple stages, starting with descriptive visualization, then model selection, model fitting, diagnosis and evaluation, and finally iterative model refinement. The second topic is a general approach to visualization and modeling that scales from small to massive datasets, and the development of new methods specifically for the scaling of data visualization. We approach scaling by partitioning the data into subsets, sampling the subsets, and applying modeling and visualization to each subset. The investigators are carrying out the research in the context of two challenging data analysis projects in homeland security: (1) Daily counts of chief complaints from 76 emergency departments of the Indiana Public Health Emergency Surveillance System; and (2) Internet packet traces for network security that we collect on the campuses of Purdue University and Stanford University.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
Statistical Theory and Methods for D&R Analysis of Large Complex Data
-
批准号:1228348
-
项目类别:Standard Grant
-
资助金额:$31.5万
-
财政年份:2012
-
负责人:William Cleveland
-
依托单位:
Data Mining, Statistical Learning, and Data Visualization for Complex Data
-
批准号:0532217
-
项目类别:Standard Grant
-
资助金额:$0.0万
-
财政年份:2005
-
负责人:William Cleveland
-
依托单位:
海外基金