TurboLift: fast accuracy lifting for historical data recovery

TurboLift: fast accuracy lifting for historical data recovery
复制标题

TurboLift:历史数据恢复的快速精度提升

DOI:
10.1007/s00778-020-00609-6
复制
发表时间:
2020
期刊:
The VLDB Journal
影响因子:
--
通讯作者:
V. Zadorozhny
V. Zadorozhny
中科院分区:
--
文献类型:
--
作者:
Fan Yang;Faisal M. Almutairi;H. Song;C. Faloutsos;N. Sidiropoulos;V. Zadorozhny

文献摘要

参考文献

被引文献

相似文献

历史数据经常涉及关于时间序列的现有报告在不同级别进行临时汇总的情况,例如每月感染麻疹的人数。在实际数据库中,不同报告所涵盖的时间段可能有重叠(即,多个报告覆盖的时间刻度)或差距(即,任何报告都没有覆盖的时间刻度)。然而,数据分析和机器学习模型需要以更精细的粒度重建历史事件,例如每周的患者计数,以进行详细的分析和预测。因此,数据分解算法在各个领域中变得越来越重要。时间序列分解方法通常利用关于数据的领域知识,例如,平滑性、周期性或稀疏性,以提高重建精度。在本文中,我们提出了一种新的方法,称为Turbolift,其目的是改善现有解集方法提供的解的质量。Turbolift从特定方法产生的解决方案开始,找到一个新的解决方案,减少解聚误差,并接近最初的解决方案。我们推导出了所提出的涡轮升力公式的封闭形式的解,使我们能够获得精确的解析重建,而无需执行耗费资源和时间的迭代。在不同领域的真实数据上的实验表明,Turbolift在分解误差、孤立点和异常检测方面是有效的。
Historical data are frequently involved in situations where the available reports on time series are temporally aggregated at different levels, e.g., the monthly counts of people infected with measles. In real databases, the time periods covered by different reports can have overlaps (i.e., time-ticks covered by more than one reports) or gaps (i.e., time-ticks not covered by any report). However, data analysis and machine learning models require reconstructing the historical events in a finer granularity, e.g., the weekly patient counts, for elaborate analysis and prediction. Thus, data disaggregation algorithms are becoming increasingly important in various domains. Time series disaggregation methods commonly utilize domain knowledge about the data, e.g., smoothness, periodicity, or sparsity, to improve the reconstruction accuracy. In this paper, we propose a novel approach, called TurboLift, which aims to improve the quality of the solutions provided by existing disaggregation methods. Starting from a solution produced by a specific method, TurboLift finds a new solution that reduces the disaggregation error and is close to the initial one. We derive a closed-form solution to the proposed formulation of TurboLift that enables us to obtain an accurate reconstruction analytically, without performing resource and time-consuming iterations. Experiments on real data from different domains showcase the effectiveness of TurboLift in terms of disaggregation error, and outlier and anomaly detection.
DOI: 10.1145/3035918.3035951
发表时间: 2015-12
期刊: Proceedings of the 2017 ACM International Conference on Management of Data
影响因子: --
作者:
Theodoros Rekatsinas;Manas R. Joglekar;H. Garcia-Molina;Aditya G. Parameswaran;Christopher Ré
通讯作者: Theodoros Rekatsinas;Manas R. Joglekar;H. Garcia-Molina;Aditya G. Parameswaran;Christopher Ré
DOI: 10.1056/nejmms1215400
发表时间: 2013-11-28
期刊: The New England journal of medicine
影响因子: --
作者:
van Panhuis WG;Grefenstette J;Jung SY;Chok NS;Cross A;Eng H;Lee BY;Zadorozhny V;Brown S;Cummings D;Burke DS
通讯作者: Burke DS
Prema:从多个聚合视图恢复有原则的张量数据
DOI: 10.1109/jstsp.2021.3056918
发表时间: 2021
影响因子: 7.5
作者:
Almutairi, Faisal M.;Kanatsoulis, Charilaos I.;Sidiropoulos, Nicholas D.
通讯作者: Sidiropoulos, Nicholas D.