Shuffler: A Large Scale Data Management Tool for Machine Learning in Computer Vision

Shuffler: A Large Scale Data Management Tool for Machine Learning in Computer Vision
复制标题

Shuffler:计算机视觉机器学习的大规模数据管理工具

DOI:
--
复制
发表时间:
2019
期刊:
Practice and Experience in Advanced Research Computing
影响因子:
--
通讯作者:
José M. F. Moura
José M. F. Moura
中科院分区:
--
文献类型:
--
作者:
E. Toropov;Paola A. Buitrago;José M. F. Moura

文献摘要

被引文献

相似文献

计算机视觉学术研究社区中的数据集主要是静态的。一旦一个数据集被接受为计算机视觉任务的基准,从事这项任务的研究人员就不会为了使他们的结果具有可重复性而改变它。与此同时,在探索新任务和新应用程序时,数据集往往是一个不断变化的实体。从业者可以联合收割机组合现有的公共数据集,过滤其中的图像或对象,更改注释或添加新的注释以适应手头的任务,可视化样本图像,或者可能以文本或绘图的形式输出统计数据。事实上,随着从业者对数据和算法进行大量实验,试图充分利用机器学习模型,数据集也会发生变化。鉴于ML和深度学习需要大量数据才能产生令人满意的结果,因此与处理实时数据集相关的数据和软件管理可能非常复杂也就不足为奇了。据我们所知,没有灵活的、公开可用的工具来促进在整个ML管道中操作图像数据及其注释。在这项工作中,我们介绍了Shuffler,这是一个开源工具,可以轻松管理大型计算机视觉数据集。它将注释存储在一个人类可读的关系数据库中。Shuffler定义了超过40种带有注释的数据处理操作,这些操作在应用于计算机视觉的监督学习中非常有用,并支持一些最知名的计算机视觉数据集。最后,它易于扩展,使得添加新操作和数据集成为一项快速且易于完成的任务。
Datasets in the computer vision academic research community are primarily static. Once a dataset is accepted as a benchmark for a computer vision task, researchers working on this task will not alter it in order to make their results reproducible. At the same time, when exploring new tasks and new applications, datasets tend to be an ever changing entity. A practitioner may combine existing public datasets, filter images or objects in them, change annotations or add new ones to fit a task at hand, visualize sample images, or perhaps output statistics in the form of text or plots. In fact, datasets change as practitioners experiment with data as much as with algorithms, trying to make the most out of machine learning models. Given that ML and deep learning call for large volumes of data to produce satisfactory results, it is no surprise that the resulting data and software management associated to dealing with live datasets can be quite complex. As far as we know, there is no flexible, publicly available instrument to facilitate manipulating image data and their annotations throughout a ML pipeline. In this work, we present Shuffler, an open source tool that makes it easy to manage large computer vision datasets. It stores annotations in a relational, human-readable database. Shuffler defines over 40 data handling operations with annotations that are commonly useful in supervised learning applied to computer vision and supports some of the most well-known computer vision datasets. Finally, it is easily extensible, making the addition of new operations and datasets a task that is fast and easy to accomplish.