Towards a foundation model for geospatial artificial intelligence (vision paper)

Towards a foundation model for geospatial artificial intelligence (vision paper)
复制标题

DOI:
10.1145/3557915.3561043
复制
发表时间:
2022-11
期刊:
Proceedings of the 30th International Conference on Advances in Geographic Information Systems
影响因子:
--
通讯作者:
Gengchen Mai;Chris Cundy;Kristy Choi;Yingjie Hu;Ni Lao;Stefano Ermon
Gengchen Mai;Chris Cundy;Kristy Choi;Yingjie Hu;Ni Lao;Stefano Ermon
中科院分区:
其他
文献类型:
--
作者:
Gengchen Mai;Chris Cundy;Kristy Choi;Yingjie Hu;Ni Lao;Stefano Ermon

文献摘要

相似文献

大型预训练模型,也称为基础模型(FM),在大规模数据上以任务不可知的方式进行训练,并且可以通过微调,少量甚至零次学习来适应各种下游任务。尽管他们在语言和视觉任务方面取得了成功,但我们还没有看到为地理空间人工智能(GeoAI)开发基础模型的尝试。在这项工作中,我们探讨了为GeoAI开发多模式基础模型的前景和挑战。我们首先通过测试现有的大型预训练语言模型(LLM)(例如GPT-2和GPT-3)在两个地理空间语义任务上的性能来展示这一想法的优势。结果表明,这些任务不可知的LLM可以在两个任务上都优于特定于任务的全监督模型,在几次学习设置中有2- 9%的改进。然而,鉴于GeoAI的多模态性质,我们也展示了这些现有基础模型的局限性,特别是在与其他模态结合处理几何形状时。因此,我们讨论了一个多模式的基础模型,它可以通过地理空间对齐的各种类型的地理空间数据的原因的可能性。我们通过讨论为GeoAI开发这种模型的独特风险和挑战来结束本文。
Large pre-trained models, also known as foundation models (FMs), are trained in a task-agnostic manner on large-scale data and can be adapted to a wide range of downstream tasks by fine tuning, few-shot, or even zero-shot learning. Despite their successes in language and vision tasks, we have yet to see an attempt to develop foundation models for geospatial artificial intelligence (GeoAI). In this work, we explore the promises and challenges for developing multimodal foundation models for GeoAI. We first show the advantages of this idea by testing the performance of existing Large pre-trained Language Models (LLMs) (e.g. GPT-2 and GPT-3) on two geospatial semantics tasks. Results indicate that these task-agnostic LLMs can outperform task-specific fully-supervised models on both tasks with 2--9% improvement in a few-shot learning setting. However, we also show the limitations of these existing foundation models given the multimodality nature of GeoAI, especially when dealing with geometries in conjunction with other modalities. So we discuss the possibility of a multimodal foundation model which can reason over various types of geospatial data through geospatial alignments. We conclude this paper by discussing the unique risks and challenges to develop such model for GeoAI.