Who’s on First?: Probing the Learning and Representation Capabilities of Language Models on Deterministic Closed Domains

Who’s on First?: Probing the Learning and Representation Capabilities of Language Models on Deterministic Closed Domains
复制标题

DOI:
10.18653/v1/2021.conll-1.16
复制
发表时间:
2021
期刊:
--
影响因子:
--
通讯作者:
David Demeter;Doug Downey
David Demeter;Doug Downey
中科院分区:
其他
文献类型:
--
作者:
David Demeter;Doug Downey

文献摘要

相似文献

当今的自然语言处理系统的能力通常是使用策划的问题和答案的大型数据集,而这些问题是进步的重要基准,由于人工分布和不完整的知识,他们也遭受了弱点。在简化的语言轨迹上(SPLAT)是某些封闭域中的语言编码(我们研究了这项工作中的国际象棋和棒球游戏的痕迹)。我们的方法只有动词般的编码。
The capabilities of today’s natural language processing systems are typically evaluated using large datasets of curated questions and answers. While these are critical benchmarks of progress, they also suffer from weakness due to artificial distributions and incomplete knowledge. Artifacts arising from artificial distributions can overstate language model performance, while incomplete knowledge limits fine-grained analysis. In this work, we introduce a complementary benchmarking approach based on SimPlified Language Activity Traces (SPLAT). SPLATs are corpora of language encodings of activity in some closed domain (we study traces from chess and baseball games in this work). SPLAT datasets use naturally-arising distributions, allow the generation of question-answer pairs at scale, and afford complete knowledge in their closed domains. We show that language models of three different architectures can answer questions about world states using only verb-like encodings of activity. Our approach is extensible to new language models and additional question-answering tasks.