CMU-01 at the SIGMORPHON 2019 Shared Task on Crosslinguality and Context in Morphology

CMU-01 at the SIGMORPHON 2019 Shared Task on Crosslinguality and Context in Morphology
复制标题

DOI:
10.18653/v1/w19-4208
复制
发表时间:
2019-07
期刊:
ArXiv
影响因子:
--
通讯作者:
Aditi Chaudhary;Elizabeth Salesky;G. Bhat;David R. Mortensen;J. Carbonell;Yulia Tsvetkov
Aditi Chaudhary;Elizabeth Salesky;G. Bhat;David R. Mortensen;J. Carbonell;Yulia Tsvetkov
中科院分区:
其他
文献类型:
--
作者:
Aditi Chaudhary;Elizabeth Salesky;G. Bhat;David R. Mortensen;J. Carbonell;Yulia Tsvetkov

文献摘要

相似文献

本文介绍了CMU-01团队提交给SIGMORPHON 2019任务2的上下文中的形态分析和词形化。这个任务要求我们为107个树库生成一个序列中每个标记的引理和形态句法描述。我们使用分层神经条件随机场(CRF)模型来处理这个任务,该模型预测每个粗粒度特征(例如,POS、Case等)独立地。然而,大多数树库资源不足,因此为它们训练深度神经模型具有挑战性。因此,我们提出了一个多语言迁移培训制度,我们从共享类似类型的多个相关语言迁移。
This paper presents the submission by the CMU-01 team to the SIGMORPHON 2019 task 2 of Morphological Analysis and Lemmatization in Context. This task requires us to produce the lemma and morpho-syntactic description of each token in a sequence, for 107 treebanks. We approach this task with a hierarchical neural conditional random field (CRF) model which predicts each coarse-grained feature (eg. POS, Case, etc.) independently. However, most treebanks are under-resourced, thus making it challenging to train deep neural models for them. Hence, we propose a multi-lingual transfer training regime where we transfer from multiple related languages that share similar typology.