Time for More Languages
Time for More Languages
复制标题
是时候使用更多语言了
DOI:
--
复制
发表时间:
2014
期刊:
影响因子:
--
通讯作者:
Michael Gertz
中科院分区:
文献类型:
--
作者:
Jannik Strotgen;Ayser Armiti;T. V. Canh;Julian Zell;Michael Gertz
Most of the research on temporal tagging so far is done for processing English text documents. There are hardly any multilingual temporal taggers supporting more than two languages. Recently, the temporal tagger HeidelTime has been made publicly available, supporting the integration of new languages by developing language-dependent resources without modifying the source code. In this article, we describe our work on developing such resources for two Asian and two Romance languages: Arabic, Vietnamese, Spanish, and Italian. While temporal tagging of the two Romance languages has been addressed before, there has been almost no research on Arabic and Vietnamese temporal tagging so far. Furthermore, we analyze language-dependent challenges for temporal tagging and explain the strategies we followed to address them. Our evaluation results on publicly available and newly annotated corpora demonstrate the high quality of our new resources for the four languages, which we make publicly available to the research community.