课题基金 / 基金详情

Guided Analytics for the Visual Exploration of Higher Dimensional Data

Guided Analytics for the Visual Exploration of Higher Dimensional Data
高维数据可视化探索的引导分析
批准号:
RGPIN-2022-03894
负责人:
Oldford, Richmond
金额:
$1.31万
依托单位:
依托单位国家:
加拿大
项目类别:
Discovery Grants Program - Individual
财政年份:
2022
资助国家:
加拿大
项目状态:
已结题
起止时间:
2022-01-01 至 2023-12-31

项目摘要

项目成果

相似基金

相关文献

中文摘要
翻译
在现代数据中,规模很重要。对于我们可能测量的几乎每个变量,都有大量甚至数百万的观测数据,而且,越来越多地,在每个观测数据上测量数百个甚至数千个变量已经成为例行公事。例如,想象一下S标准普尔500指数成份股在五年内的每日收盘价。数据由大约1,250(=5×250)天(或观察)和约500只股票(或变量)组成。例如,如果记录了每小时的价格,并记录了每一小时的开盘、收盘和每只股票的平均价格,那么这些计数可能会显著增加。在研究这些数据时,我们希望找到规律,发现一些以前没有预料到的东西,遇到一个惊喜!科学发现的时刻。为此,人类的视觉系统已经进化到真正地发现不寻常的东西,注意到模式,并看到关系。计算机交互数据可视化确实可以让人“看到”数据中正在发生的事情。但数据的规模是压倒性的。简单地看每一对股票之间的关系,我们必须看124,750块!这不可能。建议的研究是将数学结构、计算资源和统计建模应用到问题上,让计算机引导分析师只查看那些可能看起来“有趣”的少数几个曲线图(即使是100个也是节省下来的)。它不仅提供了找到这些信息的手段,还提供了有效查看它们并将它们合并到报告中的技术。这需要确定哪些情节是“有趣的”,以及我们可以根据数据计算出哪些情节是有趣的。在一些分析中有趣的东西在另一些分析中并不有趣,所以必须考虑许多不同的“有趣”衡量标准。这项研究建议开发几种具有广泛适用性的此类措施,以及一些特定应用所特有的措施。一旦我们对124,750个可能的情节计算出不同的趣味性衡量标准,我们就必须从中进行选择。这项研究正在为这样的选择提供工具。一旦我们选择了10个或100个有趣的情节的子集,我们需要了解它们是如何相关的,一对变量与另一对变量,以及我们可以从这些关系中推断出什么。在这里,这项研究引入了数学图论,提供了一种将这些情节彼此联系起来的结构。对这种结构的分析可能会揭示变量之间的关系。将开发这些结构的概率和统计模型,以帮助确保我们的推断是可靠的。在整个过程中,软件将被设计并(通过开放源码许可)向普通公众传播。通过将软件交给数据分析师,我们希望使探索性可视化、数据分析和科学发现变得更容易。
英文摘要
With modern data, size matters. Large numbers, perhaps millions, of observations are available for almost every variable we might measure, and, increasingly, it has become routine to measure 100s, even 1000s of variables on each observation. For example, imagine looking at daily closing prices of S&P 500 stocks over a five-year period. The data consists of about 1,250 (= 5 x 250) days (or observations) and about 500 stocks (or variables). These counts could be dramatically larger, if, for example, hourly prices were recorded and for each hour the opening, closing, and average prices for that hour were recorded for every stock. In examining such data, we hope to find patterns, to uncover something not previously anticipated, to encounter an "aha!" moment of scientific discovery. To this end, the human visual system has evolved to literally spot the unusual, to notice patterns, and to see relations. Computer interactive data visualization literally allows one "to see" what is going on in the data. But the size of the data is overwhelming. To simply look at the relation between every pair of stocks, we would have to look at 124,750 plots! It is not possible. The proposed research is to bring mathematical structure, computational resources, and statistical modelling to bear on the problem, to have the computer guide the analyst to view only those few plots (even 100 would be a saving) that might be "interesting" to look at. It will provide not just the means to find these, but the technology to efficiently view them, and incorporate them into a report. This requires determining which plots are "interesting", and what we might calculate on the data to tell that a plot was interesting. What is interesting in some analyses is not interesting in others, so many different measures of "interestingness" must be considered. This research proposes to develop several such measures that are of wide applicability as well as some that are peculiar to selected applications. Once we have calculated different measures of interestingness on our 124,750 possible plots, we must select amongst them. The research is providing tools for such selection. Once we have selected our subsets of 10s or 100s of interesting plots, we need to understand how they are related, one pair of variables to another, and what we might infer from those relations if anything. Here, the research brings mathematical graph theory to provide a structure relating the plots to one another. Analysis of that structure might then shed light on the relations between variables. Probability and statistical models of these structures will be developed to help ensure our inferences are reliable. Throughout, software will be designed and disseminated (through open-source licensing) to the general public. By putting the software in the hands of the data analysts we hope to make exploratory visualization, data analysis, and scientific discovery a little easier.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
海外基金