Interpreting and improving natural-language processing (in machines) with natural language-processing (in the brain)

Interpreting and improving natural-language processing (in machines) with natural language-processing (in the brain)
复制标题

DOI:
--
复制
发表时间:
2019-05
期刊:
--
影响因子:
--
通讯作者:
Mariya Toneva;Leila Wehbe
Mariya Toneva;Leila Wehbe
中科院分区:
其他
文献类型:
--
作者:
Mariya Toneva;Leila Wehbe

文献摘要

被引文献

相似文献

NLP 的神经网络模型通常在没有语言规则的显式编码的情况下实现,但它们能够打破一个又一个的性能记录。这引起了人们对解释这些网络学习到的表征的大量研究兴趣。我们在这里提出了一种新颖的解释方法,该方法依赖于我们唯一能够理解语言的处理系统:人脑。我们使用阅读复杂自然文本的受试者的大脑成像记录来解释来自 4 个最新 NLP 模型(ELMo、USE、BERT 和 Transformer-XL)的单词和序列嵌入。我们研究它们的表示在层深度、上下文长度和注意力类型方面有何不同。我们的结果揭示了这些模型中上下文相关表示的差异。此外,在 Transformer 模型中,我们发现层深度和上下文长度之间以及层深度和注意力类型之间的相互作用。我们最终假设,改变 BERT 以更好地与大脑记录保持一致,将使其能够更好地理解语言。使用句法 NLP 任务探索改变后的 BERT 表明,大脑对齐增强的模型优于原始模型。认知神经科学家已经开始使用 NLP 网络来研究大脑,这项工作闭合了循环,让 NLP 和认知神经科学之间的相互作用成为真正的异花授粉。
Neural networks models for NLP are typically implemented without the explicit encoding of language rules and yet they are able to break one performance record after another. This has generated a lot of research interest in interpreting the representations learned by these networks. We propose here a novel interpretation approach that relies on the only processing system we have that does understand language: the human brain. We use brain imaging recordings of subjects reading complex natural text to interpret word and sequence embeddings from 4 recent NLP models - ELMo, USE, BERT and Transformer-XL. We study how their representations differ across layer depth, context length, and attention type. Our results reveal differences in the context-related representations across these models. Further, in the transformer models, we find an interaction between layer depth and context length, and between layer depth and attention type. We finally hypothesize that altering BERT to better align with brain recordings would enable it to also better understand language. Probing the altered BERT using syntactic NLP tasks reveals that the model with increased brain-alignment outperforms the original model. Cognitive neuroscientists have already begun using NLP networks to study the brain, and this work closes the loop to allow the interaction between NLP and cognitive neuroscience to be a true cross-pollination.