Morphological Segmentation for Seneca
Morphological Segmentation for Seneca
复制标题
DOI:
10.18653/v1/2021.americasnlp-1.10
复制
发表时间:
2021-06
期刊:
影响因子:
--
通讯作者:
Zoey Liu;Robert Jimerson;Emily Prudhommeaux
中科院分区:
文献类型:
--
作者:
Zoey Liu;Robert Jimerson;Emily Prudhommeaux
This study takes up the task of low-resource morphological segmentation for Seneca, a critically endangered and morphologically complex Native American language primarily spoken in what is now New York State and Ontario. The labeled data in our experiments comes from two sources: one digitized from a publicly available grammar book and the other collected from informal sources. We treat these two sources as distinct domains and investigate different evaluation designs for model selection. The first design abides by standard practices and evaluate models with the in-domain development set, while the second one carries out evaluation using a development domain, or the out-of-domain development set. Across a series of monolingual and crosslinguistic training settings, our results demonstrate the utility of neural encoder-decoder architecture when coupled with multi-task learning.