What's missing in geographical parsing?

What's missing in geographical parsing?
复制标题

DOI:
10.1007/s10579-017-9385-8
复制
发表时间:
2018-01-01
影响因子:
2.7
通讯作者:
Collier, Nigel
Collier, Nigel
中科院分区:
计算机科学4区
文献类型:
--
作者:
Gritta, Milan;Pilehvar, Mohammad Taher;Collier, Nigel

文献摘要

被引文献

相似文献

可通过将地名从自由格式文本转换为地理坐标来获取地理数据。在文本报告中地理定位事件的能力代表了许多现实世界应用中的宝贵信息来源,例如紧急响应,实时社交媒体地理事件分析,理解自动响应系统中的位置指令等。然而,地理解析仍然被广泛认为是一个挑战,因为域语言的多样性,地名歧义,转喻语言和有限的利用上下文,我们在我们的分析中显示。迄今为止的结果虽然很有希望,但都是基于实验室数据,与更广泛的NLP不同,它们通常不会进行交叉比较。在这项研究中,我们评估和分析了一些领先的地理分析器的性能,一些语料库,并强调详细的挑战。我们还发布了一个自动地理标记的维基百科语料库,以减轻缺乏(开源)语料库在这个领域。
Geographical data can be obtained by converting place names from free-format text into geographical coordinates. The ability to geo-locate events in textual reports represents a valuable source of information in many real-world applications such as emergency responses, real-time social media geographical event analysis, understanding location instructions in auto-response systems and more. However, geoparsing is still widely regarded as a challenge because of domain language diversity, place name ambiguity, metonymic language and limited leveraging of context as we show in our analysis. Results to date, whilst promising, are on laboratory data and unlike in wider NLP are often not cross-compared. In this study, we evaluate and analyse the performance of a number of leading geoparsers on a number of corpora and highlight the challenges in detail. We also publish an automatically geotagged Wikipedia corpus to alleviate the dearth of (open source) corpora in this domain.