Reusable Templates and Guides For Documenting Datasets and Models for Natural Language Processing and Generation: A Case Study of the HuggingFace and GEM Data and Model Cards

Reusable Templates and Guides For Documenting Datasets and Models for Natural Language Processing and Generation: A Case Study of the HuggingFace and GEM Data and Model Cards
复制标题

用于记录自然语言处理和生成的数据集和模型的可重用模板和指南:HuggingFace 和 GEM 数据和模型卡的案例研究

DOI:
--
复制
发表时间:
2021
期刊:
IEEE Games Entertainment Media Conference
影响因子:
--
通讯作者:
Yacine Jernite
Yacine Jernite
中科院分区:
--
文献类型:
--
作者:
Angelina McMillan;Salomey Osei;Juan Diego Rodriguez;Pawan Sasanka Ammanamanchi;Sebastian Gehrmann;Yacine Jernite

文献摘要

被引文献

相似文献

为数据集和模型开发文档指南和易于使用的模板是一项具有挑战性的任务,特别是考虑到参与构建自然语言处理(NLP)工具的人员的背景、技能和动机的多样性。然而,在整个NLP领域采用标准文档实践促进了对NLP数据集和模型的更易于访问和详细描述,同时支持研究人员和开发人员反思他们的工作。为了帮助文档的标准化,我们介绍了两个旨在开发可重用文档模板的案例研究——HuggingFace数据卡,NLP中用于数据集的通用卡,以及GEM基准数据和模型卡,重点是自然语言生成。我们描述了开发这些模板的过程,包括相关涉众组的识别、一组指导原则的定义、现有模板作为基础的使用,以及基于反馈的迭代修订。
Developing documentation guidelines and easy-to-use templates for datasets and models is a challenging task, especially given the variety of backgrounds, skills, and incentives of the people involved in the building of natural language processing (NLP) tools. Nevertheless, the adoption of standard documentation practices across the field of NLP promotes more accessible and detailed descriptions of NLP datasets and models, while supporting researchers and developers in reflecting on their work. To help with the standardization of documentation, we present two case studies of efforts that aim to develop reusable documentation templates – the HuggingFace data card, a general purpose card for datasets in NLP, and the GEM benchmark data and model cards with a focus on natural language generation. We describe our process for developing these templates, including the identification of relevant stakeholder groups, the definition of a set of guiding principles, the use of existing templates as our foundation, and iterative revisions based on feedback.