Verb Argument Structure Alternations in Word and Sentence Embeddings

Verb Argument Structure Alternations in Word and Sentence Embeddings
复制标题

单词和句子嵌入中的动词参数结构变化

DOI:
10.7275/q5js-4y86
复制
发表时间:
2018
期刊:
ArXiv
影响因子:
--
通讯作者:
Samuel R. Bowman
Samuel R. Bowman
中科院分区:
--
文献类型:
--
作者:
Katharina Kann;Alex Warstadt;Adina Williams;Samuel R. Bowman

文献摘要

参考文献

被引文献

相似文献

动词出现在不同的句法环境或框架中。我们研究人工神经网络是否编码语法的区别,推断动词的特质框架选择属性。我们介绍了五个数据集,统称为FAVA,包含在聚合近10K的句子标记为语法的可接受性,说明不同的口头论点结构的交替。然后,我们测试模型是否可以区分可接受的英语动词框架组合,从不可接受的单独使用一个句子嵌入。对于收敛的证据,我们进一步构建了LaVA,一个相应的词级数据集,并调查是否可以从词嵌入中提取相同的句法特征。我们的模型对某些语言变化进行了可靠的分类,但对其他语言变化则没有,这表明虽然这些表示确实编码了细粒度的词汇信息,但它是不完整的,或者很难提取。此外,单词级和句子级模型之间的差异表明,单词嵌入中存在的一些信息不会传递给下游的句子嵌入。
Verbs occur in different syntactic environments, or frames. We investigate whether artificial neural networks encode grammatical distinctions necessary for inferring the idiosyncratic frame-selectional properties of verbs. We introduce five datasets, collectively called FAVA, containing in aggregate nearly 10k sentences labeled for grammatical acceptability, illustrating different verbal argument structure alternations. We then test whether models can distinguish acceptable English verb-frame combinations from unacceptable ones using a sentence embedding alone. For converging evidence, we further construct LaVA, a corresponding word-level dataset, and investigate whether the same syntactic features can be extracted from word embeddings. Our models perform reliable classifications for some verbal alternations but not others, suggesting that while these representations do encode fine-grained lexical information, it is incomplete or can be hard to extract. Further, differences between the word- and sentence-level models show that some information present in word embeddings is not passed on to the down-stream sentence embeddings.
DOI: 10.1162/tacl_a_00290
发表时间: 2019-01-01
影响因子: 10.9
作者:
Warstadt, Alex;Singh, Amanpreet;Bowman, Samuel R.
通讯作者: Bowman, Samuel R.