Drift in a popular metal oxide sensor dataset reveals limitations for gas classification benchmarks

Drift in a popular metal oxide sensor dataset reveals limitations for gas classification benchmarks
复制标题

DOI:
10.1016/j.snb.2022.131668
复制
发表时间:
2022-03-15
影响因子:
8.4
通讯作者:
Schmuker, Michael
Schmuker, Michael
中科院分区:
化学1区
文献类型:
--
作者:
Dennler, Nik;Rastogi, Shavika;Schmuker, Michael

文献摘要

被引文献

相似文献

金属氧化物(MOx)气体传感器由于其可调灵敏度、空间效率和低成本而成为许多应用的热门选择。公开的传感器数据集对研究界特别有价值,因为它们加速了气体传感器数据分析新算法的开发和评估。Vergara及其同事在2013年发表的一个数据集包含了风洞中MOx气体传感器阵列的记录。此后,它已成为该领域的标准基准。在这里,我们报告了这个数据集的一个潜在属性,限制了它对气体分类研究的适用性。测量时间戳显示,气体被记录在单独的,时间上聚集的批次。气体暴露前的传感器基线响应与记录批次密切相关,以至于基线响应在很大程度上足以推断给定试验中使用的气体。零偏移基线补偿没有解决这个问题,因为残余短期漂移仍然包含足够的信息,可以使用机器学习分类器进行气体/试验识别。在短时间内记录的数据的子集受漂移的影响最小,并且在偏移补偿后适合于气体分类基准,但是与完整数据集相比,分类性能降低得多。我们发现有18篇出版物在使用该数据集时没有对我们描述的情况采取预防措施,因此可能高估了气体分类算法的准确性。这些观察结果突出了使用先前记录的气体传感器数据的潜在陷阱,这些数据可能会扭曲广泛报道的结果。
Metal oxide (MOx) gas sensors are a popular choice for many applications, due to their tunable sensitivity, space efficiency and low cost. Publicly available sensor datasets are particularly valuable for the research community as they accelerate the development and evaluation of novel algorithms for gas sensor data analysis. A dataset published in 2013 by Vergara and colleagues contains recordings from MOx gas sensor arrays in a wind tunnel. It has since become a standard benchmark in the field. Here we report a latent property of this dataset that limits its suitability for gas classification studies. Measurement timestamps show that gases were recorded in separate, temporally clustered batches. Sensor baseline response before gas exposure were strongly correlated with the recording batch, to the extent that baseline response was largely sufficient to infer the gas used in a given trial. Zero-offset baseline compensation did not resolve the issue, since residual short-term drift still contained enough information for gas/trial identification using a machine learning classifier. A subset of the data recorded within a short period of time was minimally affected by drift and suitable for gas classification benchmarking after offset compensation, but with much reduced classification performance compared to the full dataset. We found 18 publications where this dataset was used without precautions against the circumstances we describe, thus potentially overestimating the accuracy of gas classification algorithms. These observations highlight potential pitfalls in using previously recorded gas sensor data, which may have distorted widely reported results.