Recognizing Quantity Names for Tabular Data

Recognizing Quantity Names for Tabular Data
复制标题

识别表格数据的数量名称

DOI:
--
复制
发表时间:
2018
期刊:
ProfS/KG4IR/Data:Search@SIGIR
影响因子:
--
通讯作者:
Brian D. Davison
Brian D. Davison
中科院分区:
--
文献类型:
--
作者:
Yang Yi;Zhiyu Chen;J. Heflin;Brian D. Davison

文献摘要

被引文献

相似文献

在这个大数据时代,有许多公共网络存储库供人们访问、检索和存储数据。数据集搜索查询很自然地包含数量及其单位。当要使用数据时,用单位描述数据是一个重要特征,即使数据模式中通常不存在这样的单位。然而,数量名称通常在列名称中或以缩写格式提供或不带相应单位。数量名称(例如长度、重量和时间)可以与一组相关单位相匹配。因此,非常需要自动确定列值的数量名称。我们研究了识别单位所属数量名称的潜力。我们为每一列分配一个与数量名称相对应的类标签,从而将问题配置为多类分类任务,然后根据列名称和列内容建立各种特征。使用随机森林,我们证明这些特征对于预测表中列的数量名称非常有用。
In this era of Big Data, there are many public web repositories for people to access, retrieve, and store data. It is natural for dataset search queries to include quantities along with their units. Describing data in terms of units is an important characteristic when that data is to be used, even though such units are often not present in the data schema. However, quantity names are often provided with or without corresponding units in the column name or in an abbreviated format. Quantity names (e.g., length, weight and time ) can be matched to a set of relevant units. Therefore, there is a significant need to automatically determine the quantity names for column values. We investigate the potential to recognize quantity names to which units belong. We assign each column a class label corresponding to the quantity name and thus configure the problem as a multi-class classification task, and then establish a variety of features based on the column name and column content. Using a random forest, we show that these features are useful for predicting quantity names for columns in tables.