Technical Term Similarity Model for Natural Language Based Data Retrieval in Civil Infrastructure Projects

Technical Term Similarity Model for Natural Language Based Data Retrieval in Civil Infrastructure Projects
复制标题

DOI:
10.22260/isarc2016/0126
复制
发表时间:
2016-07
期刊:
--
影响因子:
--
通讯作者:
Tuyen Le;H. D. Jeong
Tuyen Le;H. D. Jeong
中科院分区:
其他
文献类型:
--
作者:
Tuyen Le;H. D. Jeong

文献摘要

被引文献

相似文献

数据和信息技术的最新进展使决策者能够在民用基础设施项目的生命周期内获得广泛的数字数据集。然而,由于为特定目的提取所需数据的过程具有挑战性且耗时,因此许多数据尚未完全重用。数字数据集仅以计算机可读格式呈现,并且大多数都很复杂。为了准确地提取所需的数据子集,最终用户需要深入了解数据模式的结构、每个数据实体的含义和查询语言。因此,为了真正促进数字项目数据的重用,需要一个计算平台,允许用户以自然语言表达他们的数据需求。计算机执行此任务的关键要求之一是理解和解释用户的自然语言输入的能力,其中关键字是基本的语言成分。本研究旨在收集民用基础设施领域中常用的技术术语,并开发一个语义相似度模型,可以衡量术语之间的意义相关性/相似性。自然语言处理(NLP)技术和C值方法被用来自动提取文本文档中的术语。然后,使用一个机器学习模型称为跳跃语法模型学习的技术术语之间的语义相关性使用未标记的公路语料库作为输入数据。输入语料库包括1000万字,主要从美国各地的道路设计指南收集的模型进行评估,通过比较计算机和人类执行的映射结果。
Recent advances in data and information technologies have enabled extensive digital datasets to be available to decision makers during the life cycle of a civil infrastructure project. However, much of the data is not yet fully reused due to the challenging and time consuming process of extracting the desired data for a specific purpose. Digital datasets are presented only in computer-readable formats and they are mostly complicated. In order to accurately extract a required subset of data, end users need to have deep understanding of the structure of the data schema, the meaning of each data entity and a query language. Thus, to truly facilitate the reuse of digital project data, a computational platform is needed to allow users to present their data needs in natural language. One of the critical requirements for a computer to perform this task is the ability to understand and interpret users' natural language inputs where keywords are a basic linguistic component. This research aims to collect technical terms commonly used in the civil infrastructure domain and develop a semantic similarity model that can measure the meaning relatedness/similarity between terms. Natural Language Processing (NLP) techniques and C-value method are used to automatically extract terms from text documents. A machine learning model called Skip-gram model is then employed to learn the semantic relatedness between technical terms using an unlabeled highway corpora as the input data. The input corpus includes 10 million words mainly collected from roadway design guidelines across the U.S. The model is evaluated by comparing the mapping results performed by a computer and a human.