A Corpus of Adpositional Supersenses for Mandarin Chinese

A Corpus of Adpositional Supersenses for Mandarin Chinese
复制标题

DOI:
--
复制
发表时间:
2020-03
期刊:
ArXiv
影响因子:
--
通讯作者:
Siyao Peng;Yang Janet Liu;Yilun Zhu;Austin Blodgett;Yushi Zhao;Nathan Schneider
Siyao Peng;Yang Janet Liu;Yilun Zhu;Austin Blodgett;Yushi Zhao;Nathan Schneider
中科院分区:
其他
文献类型:
--
作者:
Siyao Peng;Yang Janet Liu;Yilun Zhu;Austin Blodgett;Yushi Zhao;Nathan Schneider

文献摘要

相似文献

形容词是语义关系的常用标记,但它具有高度的歧义性,并且在不同的语言中有很大的差异。此外,还缺乏用于研究介词语义的跨语言变异或用于构建多语言消歧系统的注释语料库。本文介绍了一个语料库,其中所有的形容词都有语义注释的汉语普通话,据我们所知,这是第一个汉语语料库,以广泛的形容词语义注释。我们的方法采用了一个框架,该框架根据表面上独立于语言的语义标准定义了一组一般的超义,尽管其发展主要集中在英语介词上(Schneider et al.,2018年)。我们发现,尽管汉语和英语在句法上存在差异,但超义范畴非常适合汉语形容词。在《小王子》的中文译本中,我们获得了较高的注释者间一致性,并分析了双文本中介词标记的语义对应关系。
Adpositions are frequent markers of semantic relations, but they are highly ambiguous and vary significantly from language to language. Moreover, there is a dearth of annotated corpora for investigating the cross-linguistic variation of adposition semantics, or for building multilingual disambiguation systems. This paper presents a corpus in which all adpositions have been semantically annotated in Mandarin Chinese; to the best of our knowledge, this is the first Chinese corpus to be broadly annotated with adposition semantics. Our approach adapts a framework that defined a general set of supersenses according to ostensibly language-independent semantic criteria, though its development focused primarily on English prepositions (Schneider et al., 2018). We find that the supersense categories are well-suited to Chinese adpositions despite syntactic differences from English. On a Mandarin translation of The Little Prince, we achieve high inter-annotator agreement and analyze semantic correspondences of adposition tokens in bitext.