Predicting Performance for Natural Language Processing Tasks

Predicting Performance for Natural Language Processing Tasks
复制标题

DOI:
10.18653/v1/2020.acl-main.764
复制
发表时间:
2020-05
期刊:
ArXiv
影响因子:
--
通讯作者:
Mengzhou Xia;Antonios Anastasopoulos;Ruochen Xu;Yiming Yang;Graham Neubig
Mengzhou Xia;Antonios Anastasopoulos;Ruochen Xu;Yiming Yang;Graham Neubig
中科院分区:
其他
文献类型:
--
作者:
Mengzhou Xia;Antonios Anastasopoulos;Ruochen Xu;Yiming Yang;Graham Neubig

文献摘要

被引文献

相似文献

考虑到自然语言处理(NLP)研究中任务、语言和领域组合的复杂性,在每个可能的实验设置上详尽地测试新提出的模型在计算上是令人望而却步的。在这项工作中,我们试图探索在没有实际训练或测试模型的情况下,获得关于NLP模型在实验环境下表现如何的合理判断的可能性。为此,我们建立了回归模型,以实验设置为输入来预测NLP实验的评价分数。在大约9个不同的NLP任务上进行实验,我们发现我们的预测器可以在未见过的语言和不同的建模架构上产生有意义的预测,优于合理的基线和人类专家。我们用一组特征来表示实验设置。进一步,我们概述了如何使用我们的预测器来找到应该运行的代表性实验的一小部分,以便为所有其他实验设置获得合理的预测。
Given the complexity of combinations of tasks, languages, and domains in natural language processing (NLP) research, it is computationally prohibitive to exhaustively test newly proposed models on each possible experimental setting. In this work, we attempt to explore the possibility of gaining plausible judgments of how well an NLP model can perform under an experimental setting, without actually training or testing the model. To do so, we build regression models to predict the evaluation score of an NLP experiment given the experimental settings as input. Experimenting on~9 different NLP tasks, we find that our predictors can produce meaningful predictions over unseen languages and different modeling architectures, outperforming reasonable baselines as well as human experts. %we represent experimental settings using an array of features. Going further, we outline how our predictor can be used to find a small subset of representative experiments that should be run in order to obtain plausible predictions for all other experimental settings.