FAIR Science for Social Machines: Let's Share Metadata Knowlets in the Internet of FAIR Data and Services

FAIR Science for Social Machines: Let's Share Metadata Knowlets in the Internet of FAIR Data and Services
复制标题

DOI:
10.1162/dint_a_00002
复制
发表时间:
2019-03-01
期刊:
影响因子:
3.9
通讯作者:
Mons, Barend
Mons, Barend
中科院分区:
计算机科学4区
文献类型:
--
作者:
Mons, Barend

文献摘要

被引文献

相似文献

在一个充斥着碎片化数据和工具的世界里,开放科学的概念已经获得了很大的动力,但同时,它也引起了很大的焦虑。一些焦虑可能与王国的崩溃有关,但也有非常合理的担忧,特别是关于机器和算法与人类相比的相对作用以及两者的结合(即,社会机器)。人们还严重关切“开放”一词的含义,以及不必要的副作用以及新方法发展的早期采用者所提倡的方法的可扩展性。其中许多问题与人机交互以及计算机在我们日常科学实践中所扮演的关键角色有关。在这里,我们解决了其中的一些问题,并提供了一些可能的解决方案。FAIR(机器可操作)数据和服务显然是开放科学(或者说FAIR科学)的核心。数据、工具和计算(运行工具)的可扩展和透明路由是公平数据和服务互联网(IFDS)的关键核心特征。欧盟委员会在其关于欧洲开放科学云的宣言中,G7和美国数据共享都确定了确保开放科学坚实和可持续基础设施的必要性。在这里,我们首先定义术语公平科学相对于开放科学。在FAIR科学中,数据和相关工具都是可查找的,在明确定义的条件下可扩展的,可互操作和可重用的,但不一定是“开放的”;没有限制,当然也不总是“免费的”。“开放”这个模棱两可的术语已经引起了相当大的混乱,也引起了研究人员和其他数据密集型专业人员的选择退出反应,他们出于非常好的理由不能开放他们的数据,比如病人隐私或国家安全。虽然开放科学是一种工作方式的定义,而不是明确要求所有数据都以完全开放获取的方式提供,但开放科学所涉及的数据的开放性内涵非常强烈。在FAIR科学中,数据和相关服务在数据管理周期中运行所有流程,从实验设计到捕获到策展,处理,链接和分析,都具有最低限度的FAIR元数据,这些元数据指定了实际基础研究对象可重用的条件,首先是机器,然后也是人类。这实际上意味着,正确地进行开放科学是公平科学的一部分。然而,公平科学也可以用部分封闭的、敏感的和专有的数据来完成。正如前面所强调的,公平不等于“开放”。在FAIR/开放科学中,数据应该尽可能开放,并在必要时封闭。如果数据是使用公共资金生成的,则默认情况通常是,对于研究产生的FAIR数据,可访问性将尽可能高,并且必须明确说明和描述对这些数据的更严格的访问和许可政策。然而,在所有情况下,即使重用受到限制,数据和相关服务也应该可以为它们的主要用途(机器)找到,这将使人类用户更容易找到它们。随着良好的数据管理成为常态的趋势,分布式数据分析和学习的一个非常重要的新市场正在开放,大量的工具和可重用的数据对象正在开发和发布。这些都需要FAIR元数据相互路由并有效。
In a world awash with fragmented data and tools, the notion of Open Science has been gaining a lot of momentum, but simultaneously, it caused a great deal of anxiety. Some of the anxiety may be related to crumbling kingdoms, but there are also very legitimate concerns, especially about the relative role of machines and algorithms as compared to humans and the combination of both (i.e., social machines). There are also grave concerns about the connotations of the term "open", but also regarding the unwanted side effects as well as the scalability of the approaches advocated by early adopters of new methodological developments. Many of these concerns are associated with mind-machine interaction and the critical role that computers are now playing in our day to day scientific practice. Here we address a number of these concerns and provide some possible solutions. FAIR (machine-actionable) data and services are obviously at the core of Open Science (or rather FAIR science). The scalable and transparent routing of data, tools and compute (to run the tools on) is a key central feature of the envisioned Internet of FAIR Data and Services (IFDS). Both the European Commission in its Declaration on the European Open Science Cloud, the G7, and the USA data commons have identified the need to ensure a solid and sustainable infrastructure for Open Science. Here we first define the term FAIR science as opposed to Open Science. In FAIR science, data and the associated tools are all Findable, Accessible under well defined conditions, Interoperable and Reusable, but not necessarily "open"; without restrictions and certainly not always "gratis". The ambiguous term "open" has already caused considerable confusion and also opt-out reactions from researchers and other data-intensive professionals who cannot make their data open for very good reasons, such as patient privacy or national security. Although Open Science is a definition for a way of working rather than explicitly requesting for all data to be available in full Open Access, the connotation of openness of the data involved in Open Science is very strong. In FAIR science, data and the associated services to run all processes in the data stewardship cycle from design of experiment to capture to curation, processing, linking and analytics all have minimally FAIR metadata, which specify the conditions under which the actual underlying research objects are reusable, first for machines and then also for humans. This effectively means that-properly conducted-Open Science is part of FAIR science. However, FAIR science can also be done with partly closed, sensitive and proprietary data. As has been emphasized before, FAIR is not identical to "open". In FAIR/Open Science, data should be as open as possible and as closed as necessary. Where data are generated using public funding, the default will usually be that for the FAIR data resulting from the study the accessibility will be as high as possible, and that more restrictive access and licensing policies on these data will have to be explicitly justified and described. In all cases, however, even if the reuse is restricted, data and related services should be findable for their major uses, machines, which will make them also much better findable for human users. With a tendency to make good data stewardship the norm, a very significant new market for distributed data analytics and learning is opening and a plethora of tools and reusable data objects are being developed and released.These all need FAIR metadata to be routed to each other and to be effective.