Time for More Languages

Time for More Languages
复制标题

是时候使用更多语言了

DOI:
--
复制
发表时间:
2014
期刊:
ACM Transactions on Asian Language Information Processing
影响因子:
--
通讯作者:
Michael Gertz
Michael Gertz
中科院分区:
--
文献类型:
--
作者:
Jannik Strotgen;Ayser Armiti;T. V. Canh;Julian Zell;Michael Gertz

文献摘要

被引文献

相似文献

目前对时态标注的研究大多是针对英语文本文档进行的。几乎没有支持两种以上语言的多语言时态标记器。最近,时间标记器HeidelTime已经公开可用,它通过开发依赖于语言的资源来支持新语言的集成,而无需修改源代码。在本文中,我们描述了我们为两种亚洲语言和两种罗曼语(阿拉伯语、越南语、西班牙语和意大利语)开发此类资源的工作。虽然对两种罗曼语的时态标注已有研究,但对阿拉伯语和越南语时态标注的研究迄今几乎没有。此外,我们分析了时态标注的语言相关挑战,并解释了我们遵循的解决这些挑战的策略。我们对公开的和新标注的语料库的评估结果表明,我们为这四种语言提供了高质量的新资源,我们向研究界公开了这些资源。
Most of the research on temporal tagging so far is done for processing English text documents. There are hardly any multilingual temporal taggers supporting more than two languages. Recently, the temporal tagger HeidelTime has been made publicly available, supporting the integration of new languages by developing language-dependent resources without modifying the source code. In this article, we describe our work on developing such resources for two Asian and two Romance languages: Arabic, Vietnamese, Spanish, and Italian. While temporal tagging of the two Romance languages has been addressed before, there has been almost no research on Arabic and Vietnamese temporal tagging so far. Furthermore, we analyze language-dependent challenges for temporal tagging and explain the strategies we followed to address them. Our evaluation results on publicly available and newly annotated corpora demonstrate the high quality of our new resources for the four languages, which we make publicly available to the research community.