The UCR Time Series Archive

The UCR Time Series Archive
复制标题

DOI:
10.1109/jas.2019.1911747
复制
发表时间:
2019-11-01
影响因子:
11.8
通讯作者:
Keogh, Eamonn
Keogh, Eamonn
中科院分区:
计算机科学1区
文献类型:
--
作者:
Hoang Anh Dau;Bagnall, Anthony;Keogh, Eamonn

文献摘要

被引文献

相似文献

UCR 时间序列档案 - 于 2002 年推出,已成为时间序列数据挖掘社区的重要资源,至少有一千篇发表的论文使用了档案中的至少一个数据集。档案馆的最初版本有十六个数据集,但从那时起,它经历了定期扩展。最后一次扩展发生在 2015 年夏天,当时存档数据集从 45 个增加到 85 个。本文介绍并将重点关注从 85 个数据集到 128 个数据集的新数据扩展。除了扩展这一宝贵资源之外,本文还为任何希望评估存档中的新算法的人提供了务实的建议。最后,本文提出了一个新颖但可操作的主张:在数百篇显示出比标准基线(1-最近邻分类)有所改进的论文中,一小部分可能错误地归因了其改进的原因。此外,这些论文所声称的改进可能可以通过更简单的修改来实现,只需要几行代码。
The UCR time series archive - introduced in 2002, has become an important resource in the time series data mining community, with at least one thousand published papers making use of at least one data set from the archive. The original incarnation of the archive had sixteen data sets but since that time, it has gone through periodic expansions. The last expansion took place in the summer of 2015 when the archive grew from 45 to 85 data sets. This paper introduces and will focus on the new data expansion from 85 to 128 data sets. Beyond expanding this valuable resource, this paper offers pragmatic advice to anyone who may wish to evaluate a new algorithm on the archive. Finally, this paper makes a novel and yet actionable claim: of the hundreds of papers that show an improvement over the standard baseline (1-nearest neighbor classification), a fraction might be mis-attributing the reasons for their improvement. Moreover, the improvements claimed by these papers might have been achievable with a much simpler modification, requiring just a few lines of code.