课题基金 / 基金详情

Guided Analytics for the Visual Exploration of Higher Dimensional Data

Guided Analytics for the Visual Exploration of Higher Dimensional Data
高维数据可视化探索的引导分析
批准号:
RGPIN-2022-03894
负责人:
Oldford, Richmond
金额:
$1.31万
依托单位:
依托单位国家:
加拿大
项目类别:
Discovery Grants Program - Individual
财政年份:
2022
资助国家:
加拿大
项目状态:
已结题
起止时间:
2022-01-01 至 2023-12-31

项目摘要

项目成果

相似基金

相关文献

中文摘要
翻译
对于现代数据,大小很重要。我们可以测量的几乎每一个变量都有大量的,也许是数百万的观测数据,而且,越来越多地,在每个观测数据上测量100个,甚至1000个变量已经成为惯例。例如,想象一下标准普尔500指数股票在五年内的每日收盘价。数据包括约1,250(= 5 x 250)天(或观察)和约500个股票(或变量)。例如,如果记录每小时的价格,并记录每只股票每小时的开盘价、收盘价和该小时的平均价格,则这些计数可能会大得多。在检查这些数据时,我们希望找到模式,发现一些以前没有预料到的东西,遇到一个"啊哈! "科学发现的时刻为此,人类的视觉系统已经进化到可以从字面上发现异常,注意模式,并看到关系。计算机交互式数据可视化字面上允许一个"看到"数据中发生了什么。但数据的规模是压倒性的。要简单地看每对股票之间的关系,我们必须看124,750个图!这是不可能的。拟议中的研究是将数学结构、计算资源和统计模型应用于问题,让计算机引导分析师只查看那些可能"有趣"的少数几个图(即使是100个也是一种节省)。它不仅提供了查找这些信息的方法,还提供了有效查看这些信息并将其纳入报告的技术。这需要确定哪些地块是“有趣的”,以及我们可以根据数据计算什么来判断地块是有趣的。在某些分析中有趣的东西在另一些分析中并不有趣,因此必须考虑许多不同的"有趣性"衡量标准。这项研究提出了开发几个这样的措施,具有广泛的适用性,以及一些特定的应用程序。一旦我们在124,750个可能的图上计算出不同的兴趣度,我们必须从中选择。这项研究正在为这种选择提供工具。一旦我们选择了10或100个有趣的图的子集,我们需要了解它们是如何相关的,一对变量与另一对变量,以及我们可能从这些关系中推断出什么。在这里,研究带来了数学图论,以提供一个结构,将情节相互关联。对这一结构的分析可能有助于了解变量之间的关系。这些结构的概率和统计模型将被开发,以帮助确保我们的推断是可靠的。在整个过程中,将设计软件并向公众传播(通过开放源码许可证)。通过将软件交给数据分析师,我们希望使探索性可视化,数据分析和科学发现变得更容易。
英文摘要
With modern data, size matters. Large numbers, perhaps millions, of observations are available for almost every variable we might measure, and, increasingly, it has become routine to measure 100s, even 1000s of variables on each observation. For example, imagine looking at daily closing prices of S&P 500 stocks over a five-year period. The data consists of about 1,250 (= 5 x 250) days (or observations) and about 500 stocks (or variables). These counts could be dramatically larger, if, for example, hourly prices were recorded and for each hour the opening, closing, and average prices for that hour were recorded for every stock. In examining such data, we hope to find patterns, to uncover something not previously anticipated, to encounter an "aha!" moment of scientific discovery. To this end, the human visual system has evolved to literally spot the unusual, to notice patterns, and to see relations. Computer interactive data visualization literally allows one "to see" what is going on in the data. But the size of the data is overwhelming. To simply look at the relation between every pair of stocks, we would have to look at 124,750 plots! It is not possible. The proposed research is to bring mathematical structure, computational resources, and statistical modelling to bear on the problem, to have the computer guide the analyst to view only those few plots (even 100 would be a saving) that might be "interesting" to look at. It will provide not just the means to find these, but the technology to efficiently view them, and incorporate them into a report. This requires determining which plots are "interesting", and what we might calculate on the data to tell that a plot was interesting. What is interesting in some analyses is not interesting in others, so many different measures of "interestingness" must be considered. This research proposes to develop several such measures that are of wide applicability as well as some that are peculiar to selected applications. Once we have calculated different measures of interestingness on our 124,750 possible plots, we must select amongst them. The research is providing tools for such selection. Once we have selected our subsets of 10s or 100s of interesting plots, we need to understand how they are related, one pair of variables to another, and what we might infer from those relations if anything. Here, the research brings mathematical graph theory to provide a structure relating the plots to one another. Analysis of that structure might then shed light on the relations between variables. Probability and statistical models of these structures will be developed to help ensure our inferences are reliable. Throughout, software will be designed and disseminated (through open-source licensing) to the general public. By putting the software in the hands of the data analysts we hope to make exploratory visualization, data analysis, and scientific discovery a little easier.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
海外基金