Time Series FeatuRe Extraction on basis of Scalable Hypothesis tests (tsfresh - A Python package)

Time Series FeatuRe Extraction on basis of Scalable Hypothesis tests (tsfresh - A Python package)
复制标题

DOI:
10.1016/j.neucom.2018.03.067
复制
发表时间:
2018-09-13
期刊:
影响因子:
6
通讯作者:
Kempa-Liehr, Andreas W.
Kempa-Liehr, Andreas W.
中科院分区:
计算机科学2区
文献类型:
--
作者:
Christ, Maximilian;Braun, Nils;Kempa-Liehr, Andreas W.

文献摘要

被引文献

相似文献

时间序列特征工程是一个耗时的过程,因为科学家和工程师必须考虑各种各样的信号处理和时间序列分析算法,以识别和提取有意义的特征。Python包tsfresh(基于可扩展假设测试的时间序列特征提取)通过结合63种时间序列特征化方法(默认情况下计算总共794个时间序列特征)和基于自动配置的假设测试的特征选择来加速这一过程。通过在数据科学过程的早期阶段识别统计上显著的时间序列特征,tsfresh关闭了与领域专家的反馈循环,并促进了早期领域特定功能的开发。该软件包实现了时间序列和机器学习库的标准API(例如pandas和scikit-learn)它被设计用于探索性分析以及直接集成到操作数据科学应用程序中。(C)2018作者由爱思唯尔公司出版
Time series feature engineering is a time-consuming process because scientists and engineers have to consider the multifarious algorithms of signal processing and time series analysis for identifying and extracting meaningful features from time series. The Python package tsfresh (Time Series FeatuRe Extraction on basis of Scalable Hypothesis tests) accelerates this process by combining 63 time series characterization methods, which by default compute a total of 794 time series features, with feature selection on basis automatically configured hypothesis tests. By identifying statistically significant time series characteristics in an early stage of the data science process, tsfresh closes feedback loops with domain experts and fosters the development of domain specific features early on. The package implements standard APIs of time series and machine learning libraries (e.g. pandas and scikit-learn) and is designed for both exploratory analyses as well as straightforward integration into operational data science applications. (C) 2018 The Authors. Published by Elsevier B.V.