Detecting Temporal Dependencies in Data

Detecting Temporal Dependencies in Data
复制标题

DOI:
--
复制
发表时间:
2021
期刊:
--
影响因子:
--
通讯作者:
--
中科院分区:
其他
文献类型:
--
作者:

文献摘要

相似文献

组织从各种来源收集数据,这些数据集可能具有未知的特征。为数据分析目的选择适当的统计和机器学习算法得益于理解这些特征,例如它是否包含时间属性。本文提出了一个理论基础,自动确定存在的时态数据的数据集给出了没有先验知识的属性。我们使用一种方法将属性分类为时态、非时态或隐藏时态。一个隐藏的(分组的)时态属性只有在它的值被分类在组中时才能被视为时态属性。我们的方法使用的Ljung-Box测试的自相关性,以及一组指标,我们提出的分类统计的基础上。我们的方法检测所有的时间和隐藏的时间属性在15个不同领域的数据集。
Organizations collect data from various sources, and these datasets may have characteristics that are unknown. Selecting the appropriate statistical and machine learning algorithm for data analytical purposes benefits from understanding these characteristics, such as if it contains temporal attributes or not. This paper presents a theoretical basis for automatically determining the presence of temporal data in a dataset given no prior knowledge about its attributes. We use a method to classify an attribute as temporal, non-temporal, or hidden temporal. A hidden (grouping) temporal attribute can only be treated as temporal if its values are categorized in groups. Our method uses a Ljung-Box test for autocorrelation as well as a set of metrics we proposed based on the classification statistics. Our approach detects all temporal and hidden temporal attributes in 15 datasets from various domains.