TopoBERT: a plug and play toponym recognition module harnessing fine-tuned BERT

TopoBERT: a plug and play toponym recognition module harnessing fine-tuned BERT
复制标题

DOI:
10.1080/17538947.2023.2239794
复制
发表时间:
2023-08
影响因子:
5.1
通讯作者:
Bing Zhou;Lei Zou;Yingjie Hu;Yi Qiang;Daniel Goldberg
Bing Zhou;Lei Zou;Yingjie Hu;Yi Qiang;Daniel Goldberg
中科院分区:
地球科学1区
文献类型:
--
作者:
Bing Zhou;Lei Zou;Yingjie Hu;Yi Qiang;Daniel Goldberg

文献摘要

被引文献

相似文献

摘要从文本内容中提取精确的地理信息,即地名识别,是地理信息检索的基础,也是大量空间分析的关键,例如从社交媒体、新闻报道和各种应用调查中挖掘基于位置的信息。然而,现有地名识别方法和工具的性能不足以支持依赖于从文本中提取细粒度地理信息的任务,例如,在灾害期间通过社交媒体发送带有地址的求助请求的人的定位。新兴的预训练语言模型彻底改变了机器对自然语言的处理和理解,为优化地名识别提供了一条有前途的途径,以支持实际应用。本文提出了一种基于一维卷积神经网络(CNN1D)和变压器双向编码器表示(BERT)的地名识别模块TopoBERT,并对其进行了优化。利用三个数据集来调整超参数并发现训练模型的最佳策略。另外七个数据集用于评估性能。与七个基线模型相比,TopoBERT实现了最先进的性能(平均f1分数= 0.854)。它被封装到易于使用的python脚本中,可以无缝地应用于各种地名识别任务,而无需额外的培训。
ABSTRACT Extracting precise geographical information from the textual content, referred to as toponym recognition, is fundamental in geographical information retrieval and crucial in a plethora of spatial analyses, e.g. mining location-based information from social media, news reports, and surveys for various applications. However, the performance of existing toponym recognition methods and tools is deficient in supporting tasks that rely on extracting fine-grained geographic information from texts, e.g. locating people sending help requests with addresses through social media during disasters. The emerging pretrained language models have revolutionized natural language processing and understanding by machines, offering a promising pathway to optimize toponym recognition to underpin practical applications. In this paper, TopoBERT, a uniquely designed toponym recognition module based on a one-dimensional Convolutional Neural Network (CNN1D) and Bidirectional Encoder Representation from Transformers (BERT), is proposed and fine-tuned. Three datasets are leveraged to tune the hyperparameters and discover the best strategy to train the model. Another seven datasets are used to evaluate the performance. TopoBERT achieves state-of-the-art performance (average f1-score = 0.854) compared to the seven baseline models. It is encapsulated into easy-to-use python scripts and can be seamlessly applied to diverse toponym recognition tasks without additional training.