Representing Numbers in NLP: a Survey and a Vision

Representing Numbers in NLP: a Survey and a Vision
复制标题

DOI:
10.18653/v1/2021.naacl-main.53
复制
发表时间:
2021-03
期刊:
ArXiv
影响因子:
--
通讯作者:
Avijit Thawani;J. Pujara;Pedro A. Szekely;Filip Ilievski
Avijit Thawani;J. Pujara;Pedro A. Szekely;Filip Ilievski
中科院分区:
其他
文献类型:
--
作者:
Avijit Thawani;J. Pujara;Pedro A. Szekely;Filip Ilievski

文献摘要

被引文献

相似文献

自然语言处理系统很少会特别考虑文本中的数字。这与神经科学中的共识形成了鲜明对比,在大脑中,数字与文字的表示方式不同。我们将最近关于计算的NLP工作整理成一个全面的任务和方法分类。我们将算术的主观概念分解为7个子任务,按两个维度排列:粒度(精确与近似)和单位(抽象与扎根)。我们分析了十多个以前发表的数字编码者和解码者所做出的无数代表性选择。我们综合了在文本中表示数字的最佳实践,并阐明了NLP中整体计算的愿景,包括设计权衡和统一评估。
NLP systems rarely give special consideration to numbers found in text. This starkly contrasts with the consensus in neuroscience that, in the brain, numbers are represented differently from words. We arrange recent NLP work on numeracy into a comprehensive taxonomy of tasks and methods. We break down the subjective notion of numeracy into 7 subtasks, arranged along two dimensions: granularity (exact vs approximate) and units (abstract vs grounded). We analyze the myriad representational choices made by over a dozen previously published number encoders and decoders. We synthesize best practices for representing numbers in text and articulate a vision for holistic numeracy in NLP, comprised of design trade-offs and a unified evaluation.