Which Melbourne? Augmenting Geocoding with Maps

Which Melbourne? Augmenting Geocoding with Maps
复制标题

DOI:
10.18653/v1/p18-1119
复制
发表时间:
2018-07
期刊:
--
影响因子:
--
通讯作者:
Milan Gritta;Mohammad Taher Pilehvar;Nigel Collier
Milan Gritta;Mohammad Taher Pilehvar;Nigel Collier
中科院分区:
其他
文献类型:
--
作者:
Milan Gritta;Mohammad Taher Pilehvar;Nigel Collier

文献摘要

被引文献

相似文献

文本地理定位的目的是将文档中包含的地理信息与一组(或多组)坐标关联起来,可以隐式地使用语言特征,也可以显式地使用与启发式相结合的地理元数据。我们介绍了一个地理编码器(位置提及消歧器),它通过利用隐式词汇线索在三个不同的数据集上实现最先进的(SOTA)结果。此外,我们提出了一种对地理元数据进行系统编码的新方法,以生成同一文本的两种不同视图。为此,我们引入了地图向量(MapVec),这是一种稀疏表示,通过绘制从世界地图上的人口数据导出的先验地理概率获得。然后,我们将隐式(语言)和显式(映射)功能集成在一起,以显著改进一系列度量标准。我们还引入了一个开源数据集,用于地质分析涵盖全球疾病爆发和流行病的新闻事件,以帮助将来对地质分析进行评估。
The purpose of text geolocation is to associate geographic information contained in a document with a set (or sets) of coordinates, either implicitly by using linguistic features and/or explicitly by using geographic metadata combined with heuristics. We introduce a geocoder (location mention disambiguator) that achieves state-of-the-art (SOTA) results on three diverse datasets by exploiting the implicit lexical clues. Moreover, we propose a new method for systematic encoding of geographic metadata to generate two distinct views of the same text. To that end, we introduce the Map Vector (MapVec), a sparse representation obtained by plotting prior geographic probabilities, derived from population figures, on a World Map. We then integrate the implicit (language) and explicit (map) features to significantly improve a range of metrics. We also introduce an open-source dataset for geoparsing of news events covering global disease outbreaks and epidemics to help future evaluation in geoparsing.