Weakly Supervised Word Segmentation for Computational Language Documentation

Weakly Supervised Word Segmentation for Computational Language Documentation
复制标题

计算语言文档的弱监督分词

DOI:
10.18653/v1/2022.acl-long.510
复制
发表时间:
2022
期刊:
ArXiv
影响因子:
--
通讯作者:
François Yvon
François Yvon
中科院分区:
--
文献类型:
--
作者:
Shu Okabe;L. Besacier;François Yvon

文献摘要

参考文献

被引文献

相似文献

单词和语素分割是语言文档的基本步骤,因为它们允许发现词典未知的语言中的词汇单元。然而,在大多数语言文档场景中,语言学家并不是从空白页开始:他们可能已经有一本预先存在的词典,或者已经开始对其一小部分数据进行手动分段。本文研究了如何在贝叶斯非参数分割模型中利用这种弱监督。我们对两种资源非常低的语言(Mboshi 和 Japhug)进行的实验(其文档仍在进行中)表明弱监督可能有利于分割质量。此外,我们研究了增量学习场景,其中以顺序方式提供手动分段。这项工作为纪录片语言学家的交互式注释工具开辟了道路。
Word and morpheme segmentation are fundamental steps of language documentation as they allow to discover lexical units in a language for which the lexicon is unknown. However, in most language documentation scenarios, linguists do not start from a blank page: they may already have a pre-existing dictionary or have initiated manual segmentation of a small part of their data. This paper studies how such a weak supervision can be taken advantage of in Bayesian non-parametric models of segmentation. Our experiments on two very low resource languages (Mboshi and Japhug), whose documentation is still in progress, show that weak supervision can be beneficial to the segmentation quality. In addition, we investigate an incremental learning scenario where manual segmentations are provided in a sequential manner. This work opens the way for interactive annotation tools for documentary linguists.
DOI: 10.18653/v1/2021.americasnlp-1.10
发表时间: 2021-06
期刊: Proceedings of the First Workshop on Natural Language Processing for Indigenous Languages of the Americas
影响因子: --
作者:
Zoey Liu;Robert Jimerson;Emily Prudhommeaux
通讯作者: Zoey Liu;Robert Jimerson;Emily Prudhommeaux