Anomaly Detection in COVID-19 Time-Series Data.

Anomaly Detection in COVID-19 Time-Series Data.
复制标题

DOI:
10.1007/s42979-021-00658-w
复制
发表时间:
2021
期刊:
SN computer science
影响因子:
--
通讯作者:
Kahn MG
Kahn MG
中科院分区:
其他
文献类型:
--
作者:
Homayouni H;Ray I;Ghosh S;Gondalia S;Kahn MG

文献摘要

被引文献

相似文献

在大量真实世界的医疗数据中进行异常检测和解释,例如与COVID-19有关的数据,带来了一些挑战。首先,我们处理的是时间序列数据。典型的时间序列数据描述单个对象随时间的行为。在医疗数据中,我们处理的是属于多个实体的时间序列数据。因此,可能存在多个记录子集,使得属于单个实体的每个子集中的记录是时间相关的,但是不同子集中的记录是不相关的。此外,子集中的记录包含不同类型的属性,其中一些必须以特定的方式进行分组,以使分析有意义。异常检测技术需要为属于多个实体的时间序列数据定制。其次,异常检测技术无法向专家解释异常值的原因。这对于目前知识不足的新疾病和流行病至关重要。我们建议通过扩展我们现有的工作称为IDEAL,这是一个基于LSTM的自动编码器的顺序记录的数据质量测试的方法来解决这些问题,并提供解释的约束违反的方式,是可以理解的最终用户。该扩展(1)使用了一种新的两级整形技术,将COVID-19数据集拆分为多个时间依赖性的异常,(2)添加了一个数据可视化图,以进一步解释异常并评估IDEAL检测到的异常水平。我们进行了两个系统的评估研究,我们的异常子序列检测。一项研究使用了汇总数据,包括病例数、死亡数、康复数和住院率百分比,这些数据是从一个COVID跟踪项目、纽约时报和约翰霍普金斯医院收集的同一时期的数据。另一项研究使用了从安舒茨医疗中心健康数据仓库获得的COVID-19患者医疗记录。结果是有希望的,并表明我们的技术可以用来检测异常大量的真实世界的未标记的数据,其准确性或有效性是未知的。
Anomaly detection and explanation in big volumes of real-world medical data, such as those pertaining to COVID-19, pose some challenges. First, we are dealing with time-series data. Typical time-series data describe behavior of a single object over time. In medical data, we are dealing with time-series data belonging to multiple entities. Thus, there may be multiple subsets of records such that records in each subset, which belong to a single entity are temporally dependent, but the records in different subsets are unrelated. Moreover, the records in a subset contain different types of attributes, some of which must be grouped in a particular manner to make the analysis meaningful. Anomaly detection techniques need to be customized for time-series data belonging to multiple entities. Second, anomaly detection techniques fail to explain the cause of outliers to the experts. This is critical for new diseases and pandemics where current knowledge is insufficient. We propose to address these issues by extending our existing work called IDEAL, which is an LSTM-autoencoder based approach for data quality testing of sequential records, and provides explanations of constraint violations in a manner that is understandable to end-users. The extension (1) uses a novel two-level reshaping technique that splits COVID-19 data sets into multiple temporally-dependent subsequences and (2) adds a data visualization plot to further explain the anomalies and evaluate the level of abnormality of subsequences detected by IDEAL. We performed two systematic evaluation studies for our anomalous subsequence detection. One study uses aggregate data, including the number of cases, deaths, recovered, and percentage of hospitalization rate, collected from a COVID tracking project, New York Times, and Johns Hopkins for the same time period. The other study uses COVID-19 patient medical records obtained from Anschutz Medical Center health data warehouse. The results are promising and indicate that our techniques can be used to detect anomalies in large volumes of real-world unlabeled data whose accuracy or validity is unknown.