Morphological Segmentation for Seneca

Morphological Segmentation for Seneca
复制标题

DOI:
10.18653/v1/2021.americasnlp-1.10
复制
发表时间:
2021-06
期刊:
Proceedings of the First Workshop on Natural Language Processing for Indigenous Languages of the Americas
影响因子:
--
通讯作者:
Zoey Liu;Robert Jimerson;Emily Prudhommeaux
Zoey Liu;Robert Jimerson;Emily Prudhommeaux
中科院分区:
其他
文献类型:
--
作者:
Zoey Liu;Robert Jimerson;Emily Prudhommeaux

文献摘要

被引文献

相似文献

这项研究承担了塞内卡语的低资源形态分割任务,塞内卡语是一种极度濒危且形态复杂的美洲原住民语言,主要在现在的纽约州和安大略省使用。我们实验中的标记数据来自两个来源:一个是从公开的语法书中数字化的,另一个是从非正式来源收集的。我们将这两个来源视为不同的领域,并研究模型选择的不同评估设计。第一种设计遵循标准实践,使用域内开发集评估模型,而第二种设计使用开发域或域外开发集进行评估。在一系列单语和跨语言训练设置中,我们的结果证明了神经编码器-解码器架构与多任务学习相结合的实用性。
This study takes up the task of low-resource morphological segmentation for Seneca, a critically endangered and morphologically complex Native American language primarily spoken in what is now New York State and Ontario. The labeled data in our experiments comes from two sources: one digitized from a publicly available grammar book and the other collected from informal sources. We treat these two sources as distinct domains and investigate different evaluation designs for model selection. The first design abides by standard practices and evaluate models with the in-domain development set, while the second one carries out evaluation using a development domain, or the out-of-domain development set. Across a series of monolingual and crosslinguistic training settings, our results demonstrate the utility of neural encoder-decoder architecture when coupled with multi-task learning.