Zero-Shot Cross-Lingual Machine Reading Comprehension via Inter-Sentence Dependency Graph

Zero-Shot Cross-Lingual Machine Reading Comprehension via Inter-Sentence Dependency Graph
复制标题

DOI:
10.1609/aaai.v36i10.21407
复制
发表时间:
2021-12
期刊:
ArXiv
影响因子:
--
通讯作者:
Liyan Xu;Xuchao Zhang;Bo Zong;Yanchi Liu;Wei Cheng;Jingchao Ni;Haifeng Chen;Liang Zhao;Jinho D. Choi
Liyan Xu;Xuchao Zhang;Bo Zong;Yanchi Liu;Wei Cheng;Jingchao Ni;Haifeng Chen;Liang Zhao;Jinho D. Choi
中科院分区:
其他
文献类型:
--
作者:
Liyan Xu;Xuchao Zhang;Bo Zong;Yanchi Liu;Wei Cheng;Jingchao Ni;Haifeng Chen;Liang Zhao;Jinho D. Choi

文献摘要

相似文献

我们的目标是直接零样本设置中的跨语言机器阅读理解(MRC)任务,通过结合通用依赖关系(UD)的句法特征,我们使用的关键特征是每个句子内的句法关系。虽然之前的工作已经证明了有效的语法引导 MRC 模型,但我们建议除了基本的句子内关系之外,还采用句子间句法关系,以进一步利用 MRC 任务的多句子输入中的句法依赖关系。在我们的方法中,我们构建了句子间依存图(ISDG),连接依存树以形成跨句子的全局句法关系。然后,我们提出了对全局依赖图进行编码的 ISDG 编码器,通过一跳和多跳依赖路径显式地解决句子间关系。在三个多语言 MRC 数据集(XQuAD、MLQA、TyDiQA-GoldP)上的实验表明,我们仅接受英语训练的编码器能够提高涵盖 8 种语言的所有 14 个测试集的零样本性能,平均提高 3.8 F1 / 5.2 EM,在某些语言上提高 5.2 F1 / 11.2 EM。进一步的分析表明,这种改进可归因于对跨语言一致句法路径的关注。我们的代码位于 https://github.com/lxucs/multilingual-mrc-isdg。
We target the task of cross-lingual Machine Reading Comprehension (MRC) in the direct zero-shot setting, by incorporating syntactic features from Universal Dependencies (UD), and the key features we use are the syntactic relations within each sentence. While previous work has demonstrated effective syntax-guided MRC models, we propose to adopt the inter-sentence syntactic relations, in addition to the rudimentary intra-sentence relations, to further utilize the syntactic dependencies in the multi-sentence input of the MRC task. In our approach, we build the Inter-Sentence Dependency Graph (ISDG) connecting dependency trees to form global syntactic relations across sentences. We then propose the ISDG encoder that encodes the global dependency graph, addressing the inter-sentence relations via both one-hop and multi-hop dependency paths explicitly. Experiments on three multilingual MRC datasets (XQuAD, MLQA, TyDiQA-GoldP) show that our encoder that is only trained on English is able to improve the zero-shot performance on all 14 test sets covering 8 languages, with up to 3.8 F1 / 5.2 EM improvement on-average, and 5.2 F1 / 11.2 EM on certain languages. Further analysis shows the improvement can be attributed to the attention on the cross-linguistically consistent syntactic path. Our code is available at https://github.com/lxucs/multilingual-mrc-isdg.