A general purpose tool-set for representing data relationships: Converting data into knowledge

A general purpose tool-set for representing data relationships: Converting data into knowledge
复制标题

用于表示数据关系的通用工具集:将数据转换为知识

DOI:
10.1109/nysds.2016.7747809
复制
发表时间:
2016
期刊:
IEEE 2016 New York Scientific Data Summit (NYSDS
影响因子:
--
通讯作者:
Wright, John
Wright, John
中科院分区:
--
文献类型:
--
作者:
Stillerman, Joshua;Fredian, Thomas;Greenwald, Martin;Wright, John

文献摘要

参考文献

相似文献

需要丰富的元数据来查找和理解现代实验中巨大而复杂的数据存储所记录的测量结果。随着时间的推移,存储和管理这些元数据的系统得到了改进,但在大多数情况下是数据关系的临时集合,通常在特定于域或站点的应用程序代码中表示。我们正在开发一套通用的工具来存储、管理和检索数据关系元数据。这些工具对底层数据存储机制和存储在其中的数据是不可知的,这使得系统可以应用于广泛的科学领域。数据管理工具通常通过隐式或显式元数据表示至少一种关系范式。这些元数据的添加允许更大的用户组在更长的时间内搜索和理解数据。使用这些系统,研究人员不太依赖于与参与实验的科学家进行一对一的交流,也不依赖于他们记住数据细节的能力。在磁聚变研究界,MDSplus系统被广泛用于记录实验的原始和处理数据。用户为他们的每个实验实例创建一个层次关系树,允许他们记录所记录内容的含义。该系统的大多数用户都添加了一组特别的工具来帮助用户定位特定的实验运行,然后他们可以通过这个分层组织访问这些工具。然而,MDSplus树只是记录的一种可能的组织,而这些将实验“拍摄”与运行天数、实验建议、日志条目、运行摘要、分析工作流程、出版物等联系起来的附加应用程序到目前为止,都是在一个接一个实验的基础上实现的。元数据来源本体项目(MPO)是一个用于记录计算结果的数据来源信息的系统。它允许用户记录其计算工作流程的每个步骤的输入和输出,特别是使用哪些原始数据和处理过的数据作为输入,运行了哪些代码以及产生了哪些结果。生成的来源图集合可以进行注释、分组、搜索、过滤和浏览。这为记录、理解和定位计算结果提供了一个强大的工具。然而,这可以被理解为一个更具体的数据关系,它可以被解释为更一般的东西的实例。在这些项目中开发的概念的基础上,我们正在开发一个通用系统,可用于将所有这些类型的数据关系表示为数学图形。正如MDSplus和MPO是用户集合的数据管理需求的一般化,这个新系统将一般化数据之间关系的存储、位置和检索。系统将数据关系存储为数据,而不是编码在一组应用程序特定的程序或特别的数据结构中。存储的数据将通过uri引用,从而允许系统对底层数据表示不可知。然后用户可以遍历这些图。该系统将允许用户构建一个图形集合,描述数据项之间的任何或所有关系,定位感兴趣的数据,查看这些数据是哪些其他图形的成员,并导航到和通过它们。
Rich metadata is required to find and understand the recorded measurements from modern experiments with their immense and complex data stores. Systems to store and manage these metadata have improved over time, but in most cases are ad-hoc collections of data relationships, often represented in domain or site specific application code. We are developing a general set of tools to store, manage, and retrieve data-relationship metadata. These tools will be agnostic to the underlying data storage mechanisms, and to the data stored in them, making the system applicable across a wide range of science domains. Data management tools typically represent at least one relationship paradigm through implicit or explicit metadata. The addition of these metadata allows the data to be searched and understood by larger groups of users over longer periods of time. Using these systems, researchers are less dependent on one on one communication with the scientists involved in running the experiments, nor to rely on their ability to remember the details of their data. In the magnetic fusion research community, the MDSplus system is widely used to record raw and processed data from experiments. Users create a hierarchical relationship tree for each instance of their experiment, allowing them to record the meanings of what is recorded. Most users of this system, add to this a set of ad-hoc tools to help users locate specific experiment runs, which they can then access via this hierarchical organization. However, the MDSplus tree is only one possible organization of the records, and these additional applications that relate the experiment `shots' into run days, experimental proposals, logbook entries, run summaries, analysis work flow, publications, etc. have up until now, been implemented on an experiment by experiment basis. The Metadata Provenance Ontology project, MPO, is a system built to record data provenance information about computed results. It allows users to record the inputs and outputs from each step of their computational workflows, in particular, what raw and processed data were used as inputs, what codes were run and what results were produced. The resulting collections of provenance graphs can be annotated, grouped, searched, filtered and browsed. This provides a powerful tool to record, understand, and locate computed results. However, this can be understood as one more specific data relationship, which can be construed as an instance of something more general. Building on concepts developed in these projects, we are developing a general system that could be used to represent all of these kinds of data relationships as mathematical graphs. Just as MDSplus and MPO were generalizations of data management needs for a collection of users, this new system will generalize the storage, location, and retrieval of the relationships between data. The system will store data relationships as data, not encoded in a set of application specific programs or ad hoc data structures. Stored data, would be referred to by URIs allowing the system to be agnostic to the underlying data representations. Users can then traverse these graphs. The system will allow users to construct a collection of graphs describing ANY OR ALL OF the relationships between data items, locate interesting data, see what other graphs these data are members of and navigate into and through them.
用于聚变模拟数据的组织和系统化的元数据目录
DOI: --
发表时间: 2012
期刊:
影响因子: --
作者:
M. Greenwald;T. Fredian;D. Schissel;J. Stillerman
通讯作者: J. Stillerman