Parsing Natural Language Queries for Extracting Data from Large-Scale Geospatial Transportation Asset Repositories

Parsing Natural Language Queries for Extracting Data from Large-Scale Geospatial Transportation Asset Repositories
复制标题

DOI:
10.1061/9780784481295.008
复制
发表时间:
2018-03
期刊:
--
影响因子:
--
通讯作者:
Tuyen Le;H. D. Jeong;Stephen B Gilbert;E. Chukharev-Hudilainen
Tuyen Le;H. D. Jeong;Stephen B Gilbert;E. Chukharev-Hudilainen
中科院分区:
其他
文献类型:
--
作者:
Tuyen Le;H. D. Jeong;Stephen B Gilbert;E. Chukharev-Hudilainen

文献摘要

相似文献

摘要数据和信息技术的最新进展使广泛的数字数据集在交通运输项目的整个生命周期中可供决策者使用。然而,由于为特定目的提取所需数据的过程既具有挑战性又耗时,因此这些数据中的大多数尚未得到充分重用。数字数据集仅以计算机可读格式呈现,并且大多数都很复杂。从复杂的大型数据源中提取数据非常耗时,并且需要大量的专业知识。因此,需要一种用户友好的数据探索框架,允许用户以人类语言呈现他们的数据兴趣。为了满足这一需求,本研究采用自然语言处理(NLP)技术来开发一个自然语言接口(NLI),它可以理解用户的意图,并自动将他们的人类语言输入转换为正式的查询。本文介绍了这样一个NLI的发展,这是建立一个专门的查询标记根据其语义贡献相应的正式查询的分类方法的一个重要任务的结果。该方法在一个由专家手动注释的30个简单英语问题的小测试集上进行了验证。结果显示,令人印象深刻的准确率超过95%。本文中提出的令牌分类,预计将提供一个基本的手段,开发一个有效的NLI运输资产数据库。
Abstract Recent advances in data and information technologies have enabled extensive digital datasets to be available to decision makers throughout the life cycle of a transportation project. However, most of these data are not yet fully reused due to the challenging and time-consuming process of extracting the desired data for a specific purpose. Digital datasets are presented only in computer-readable formats and they are mostly complicated. Extracting data from complex and large data sources is significantly time-consuming and requires considerable expertise. Thus, there is a need for a user-friendly data exploration framework that allows users to present their data interests in human language. To fulfill that demand, this study employs natural language processing (NLP) techniques to develop a natural language interface (NLI) which can understand users’ intent and automatically convert their inputs in the human language into formal queries. This paper presents the results of an important task of the development of such a NLI that is to establish a method for classifying the tokens of an ad-hoc query in accordance with their semantic contribution to the corresponding formal query. The method was validated on a small test set of 30 plain English questions manually annotated by an expert. The result shows an impressive accuracy of over 95%. The token classification presented in this paper is expected to provide a fundamental means for developing an effective NLI to transportation asset databases.