If all you have is a bit of the Bible: Learning POS taggers for truly low-resource languages

If all you have is a bit of the Bible: Learning POS taggers for truly low-resource languages
复制标题

如果你只有一点圣经:学习真正的低资源语言的词性标注器

DOI:
10.3115/v1/p15-2044
复制
发表时间:
2015
期刊:
The American Historical Review
影响因子:
--
通讯作者:
Anders Søgaard
Anders Søgaard
中科院分区:
--
文献类型:
--
作者:
Zeljko Agic;Dirk Hovy;Anders Søgaard

文献摘要

被引文献

相似文献

我们提出了一个简单的方法来学习像Akawaio,Aukan或Cakchiquel这样的语言的词性标记-这些语言只存在圣经部分的翻译。通过聚合来自几种注释语言的标签,并通过诗句上的单词对齐来传播它们,我们学习了100种语言的POS标签,使用这些语言相互引导。我们评估我们的跨语言模型的测试集存在的25种语言,以及另外10个,我们有标签字典。我们的方法比从圣经翻译中诱导的最先进的无监督POS标记器表现得更好(20-30%),并且通常与弱监督方法竞争,这些方法假设高质量的平行语料库,具有完美标记化的代表性单语语料库和/或标签词典。我们为所有100种语言提供模型。
We present a simple method for learning part-of-speech taggers for languages like Akawaio, Aukan, or Cakchiquel – languages for which nothing but a translation of parts of the Bible exists. By aggregating over the tags from a few annotated languages and spreading them via wordalignment on the verses, we learn POS taggers for 100 languages, using the languages to bootstrap each other. We evaluate our cross-lingual models on the 25 languages where test sets exist, as well as on another 10 for which we have tag dictionaries. Our approach performs much better (20-30%) than state-of-the-art unsupervised POS taggers induced from Bible translations, and is often competitive with weakly supervised approaches that assume high-quality parallel corpora, representative monolingual corpora with perfect tokenization, and/or tag dictionaries. We make models for all 100 languages available.