CoNLL-SIGMORPHON 2017 Shared Task: Universal Morphological Reinflection in 52 Languages

CoNLL-SIGMORPHON 2017 Shared Task: Universal Morphological Reinflection in 52 Languages
复制标题

DOI:
10.18653/v1/k17-2001
复制
发表时间:
2017-06
期刊:
--
影响因子:
--
通讯作者:
Ryan Cotterell;Christo Kirov;John Sylak-Glassman;Géraldine Walther;Ekaterina Vylomova;Patrick Xia;Manaal Faruqui;Sandra Kübler;David Yarowsky;Jason Eisner;Mans Hulden
Ryan Cotterell;Christo Kirov;John Sylak-Glassman;Géraldine Walther;Ekaterina Vylomova;Patrick Xia;Manaal Faruqui;Sandra Kübler;David Yarowsky;Jason Eisner;Mans Hulden
中科院分区:
其他
文献类型:
--
作者:
Ryan Cotterell;Christo Kirov;John Sylak-Glassman;Géraldine Walther;Ekaterina Vylomova;Patrick Xia;Manaal Faruqui;Sandra Kübler;David Yarowsky;Jason Eisner;Mans Hulden

文献摘要

被引文献

相似文献

CoNLL-SIGMORPHON 2017共享的监督形态生成任务要求系统在52种类型学上不同的语言中进行训练和测试。在子任务1中,提交的系统被要求预测给定引理的特定变形形式。在子任务2中,系统被给予一个引理及其一些特定的屈折形式,并被要求通过预测所有剩余的屈折形式来完成屈折范式。这两个子任务包括高、中和低资源条件。子任务1收到24个系统提交,而子任务2收到3个系统提交。继神经序列到序列模型在SIGMORPHON 2016共享任务中取得成功之后,除了一个提交的文件外,所有提交的文件都包含了神经组件。结果表明,只要模型具有适当的归纳偏差或利用额外的未标记数据或合成数据,就可以在较小的训练数据集上实现高性能。然而,不同的偏置和数据增强导致正确预测了不相交的变形形式集,这表明未来还有改进的空间。
The CoNLL-SIGMORPHON 2017 shared task on supervised morphological generation required systems to be trained and tested in each of 52 typologically diverse languages. In sub-task 1, submitted systems were asked to predict a specific inflected form of a given lemma. In sub-task 2, systems were given a lemma and some of its specific inflected forms, and asked to complete the inflectional paradigm by predicting all of the remaining inflected forms. Both sub-tasks included high, medium, and low-resource conditions. Sub-task 1 received 24 system submissions, while sub-task 2 received 3 system submissions. Following the success of neural sequence-to-sequence models in the SIGMORPHON 2016 shared task, all but one of the submissions included a neural component. The results show that high performance can be achieved with small training datasets, so long as models have appropriate inductive bias or make use of additional unlabeled data or synthetic data. However, different biasing and data augmentation resulted in disjoint sets of inflected forms being predicted correctly, suggesting that there is room for future improvement.