SigMoreFun Submission to the SIGMORPHON Shared Task on Interlinear Glossing

SigMoreFun Submission to the SIGMORPHON Shared Task on Interlinear Glossing
复制标题

DOI:
10.18653/v1/2023.sigmorphon-1.22
复制
发表时间:
2023
期刊:
--
影响因子:
--
通讯作者:
Taiqi He;Lindia Tjuatja;Nathaniel R. Robinson;Shinji Watanabe;David R. Mortensen;Graham Neubig;Lori Levin
Taiqi He;Lindia Tjuatja;Nathaniel R. Robinson;Shinji Watanabe;David R. Mortensen;Graham Neubig;Lori Levin
中科院分区:
其他
文献类型:
--
作者:
Taiqi He;Lindia Tjuatja;Nathaniel R. Robinson;Shinji Watanabe;David R. Mortensen;Graham Neubig;Lori Levin

文献摘要

被引文献

相似文献

在我们提交给SIGMORPHON 2023共享任务的线间注释(IGT)中,我们探索了跨七种低资源语言的数据增强和建模方法。对于数据增强,我们探索了两种方法:从提供的训练数据创建人工数据和利用其他语言的现有IGT资源。在建模方面,我们测试了所提供的令牌分类基线的增强版本以及预训练的多语言seq2seq模型。此外,我们使用Gitksan(数据量最小的语言)的字典进行后校正。我们发现,我们的标记分类模型是性能最好的,在所有提交的数据中,阿拉帕霍的词级准确度最高,Gitksan的词素级准确度最高。我们还表明,数据增强是一种有效的策略,尽管应用人工数据预训练在两个测试模型中具有非常不同的效果。
In our submission to the SIGMORPHON 2023 Shared Task on interlinear glossing (IGT), we explore approaches to data augmentation and modeling across seven low-resource languages. For data augmentation, we explore two approaches: creating artificial data from the provided training data and utilizing existing IGT resources in other languages. On the modeling side, we test an enhanced version of the provided token classification baseline as well as a pretrained multilingual seq2seq model. Additionally, we apply post-correction using a dictionary for Gitksan, the language with the smallest amount of data. We find that our token classification models are the best performing, with the highest word-level accuracy for Arapaho and highest morpheme-level accuracy for Gitksan out of all submissions. We also show that data augmentation is an effective strategy, though applying artificial data pretraining has very different effects across both models tested.