Evaluation-as-a-Service for the Computational Sciences

Evaluation-as-a-Service for the Computational Sciences
复制标题

计算科学评估即服务

DOI:
10.1145/3239570
复制
发表时间:
2018
期刊:
Journal of Data and Information Quality
影响因子:
--
通讯作者:
Kalpathy-Cramer Jayashree
Kalpathy-Cramer Jayashree
中科院分区:
--
文献类型:
--
作者:
Hopfgartner Frank;Kando Noriko;Kato Makoto P.;Krithara Anastasia;Gollub Tim;Potthast Martin;Viegas Evelyne;Mercer Simon;Hanbury Allan;M?ller Henning;Eggel Ivan;Balog Krisztian;Brodt Torben;Cormack Gordon V.;Lin Jimmy;Kalpathy-Cramer Jayashree

文献摘要

相似文献

经验计算机科学的评估对于显示进展和评估所开发的技术至关重要。一些研究领域,如信息检索,长期以来一直依赖于系统评估来衡量进展:在这里,创建共享测试集,定义搜索任务,并收集这些任务的地面真相的克兰菲尔德范式一直持续到现在。然而,近年来,出现了几个新的挑战,不适合这种模式很好:非常大的数据集,在医疗领域中发现的机密数据集,以及在工业中经常遇到的快速变化的数据集。众包也改变了行业解决问题的方式,公司现在组织挑战并发放奖金,以激励人们应对挑战,特别是在机器学习领域。本文基于评估即服务(EaaS)研讨会上的讨论。EaaS是一种范例,它不向参与者提供数据集,让他们在本地处理数据,而是将数据保持在中心,并允许通过应用程序编程接口(API)、虚拟机(VM)或其他可能性进行访问,以发布可执行文件。本文的目的是总结和比较当前的方法,并整合这些方法的经验,以概述EaaS的下一步,特别是对可持续的研究基础设施。本文总结了几种现有的EaaS方法,并分析了它们的使用场景和优缺点。总结了影响EaaS的许多因素,以及各种利益相关者的动机环境,从资助机构到挑战组织者,研究人员和参与者,到有兴趣提供他们需要解决方案的现实世界问题的行业。EaaS解决了当前研究环境中的许多问题,其中许多研究人员通常无法访问数据集。已发布工具的可执行文件同样经常不可用,使得结果的再现性不可能。然而,EaaS创建了可重用/可引用的数据集以及可用的可执行文件。尽管仍存在许多挑战,但这样的研究框架也可以促进研究人员之间的更多合作,从而有可能加快获得研究成果的速度。
Evaluation in empirical computer science is essential to show progress and assess technologies developed. Several research domains such as information retrieval have long relied on systematic evaluation to measure progress: here, the Cranfield paradigm of creating shared test collections, defining search tasks, and collecting ground truth for these tasks has persisted up until now. In recent years, however, several new challenges have emerged that do not fit this paradigm very well: extremely large data sets, confidential data sets as found in the medical domain, and rapidly changing data sets as often encountered in industry. Crowdsourcing has also changed the way in which industry approaches problem-solving with companies now organizing challenges and handing out monetary awards to incentivize people to work on their challenges, particularly in the field of machine learning.This article is based on discussions at a workshop on Evaluation-as-a-Service (EaaS). EaaS is the paradigm of not providing data sets to participants and have them work on the data locally, but keeping the data central and allowing access via Application Programming Interfaces (API), Virtual Machines (VM), or other possibilities to ship executables. The objectives of this article are to summarize and compare the current approaches and consolidate the experiences of these approaches to outline the next steps of EaaS, particularly toward sustainable research infrastructures.The article summarizes several existing approaches to EaaS and analyzes their usage scenarios and also the advantages and disadvantages. The many factors influencing EaaS are summarized, and the environment in terms of motivations for the various stakeholders, from funding agencies to challenge organizers, researchers and participants, to industry interested in supplying real-world problems for which they require solutions.EaaS solves many problems of the current research environment, where data sets are often not accessible to many researchers. Executables of published tools are equally often not available making the reproducibility of results impossible. EaaS, however, creates reusable/citable data sets as well as available executables. Many challenges remain, but such a framework for research can also foster more collaboration between researchers, potentially increasing the speed of obtaining research results.