Selective Sampling for Sensor Type Classification in Buildings

Selective Sampling for Sensor Type Classification in Buildings
复制标题

DOI:
10.1109/ipsn48710.2020.00028
复制
发表时间:
2020-04
期刊:
2020 19th ACM/IEEE International Conference on Information Processing in Sensor Networks (IPSN)
影响因子:
--
通讯作者:
Jing Ma;Dezhi Hong;Hongning Wang
Jing Ma;Dezhi Hong;Hongning Wang
中科院分区:
其他
文献类型:
--
作者:
Jing Ma;Dezhi Hong;Hongning Wang

文献摘要

被引文献

相似文献

将任何智能技术应用于建筑物的一个关键障碍是需要在数千个传感和控制点中定位并连接到必要的资源,即元数据映射问题。现有的解决方案依赖于对传感器元数据进行详尽的手动注释——这是一个费力、昂贵且难以扩展的过程。为了减少所需的手动工作量,本文提出了一种多预言机选择性采样框架,以利用来自可靠性未知的信息源(例如现有建筑物,我们称之为弱预言机)的噪声标签来进行元数据映射。该框架涉及一个交互过程,其中逐步选择并标记一小组传感器实例,以学习如何聚合噪声标签以及预测传感器类型。设计框架时出现两个关键挑战,即弱预言机可靠性估计和查询的实例选择。为了解决第一个挑战,我们开发了一种基于集群的弱预言机可靠性估计方法,以利用弱预言机在不同实例组中表现不同的观察结果。对于第二个挑战,我们提出了一种基于分歧的查询选择策略,以结合标记实例对减少分类器不确定性和提高标签聚合质量的潜在影响。我们根据来自 5 座建筑物的大量真实建筑传感器数据来评估我们的解决方案,这些数据包含 18 种不同类型的 11, 000 多个传感器。实验结果验证了我们解决方案的有效性,该解决方案优于一组最先进的基线。
A key barrier to applying any smart technology to a building is the requirement of locating and connecting to the necessary resources among the thousands of sensing and control points, i.e., the metadata mapping problem. Existing solutions depend on exhaustive manual annotation of sensor metadata — a laborious, costly, and hardly scalable process. To reduce the amount of manual effort required, this paper presents a multi-oracle selective sampling framework to leverage noisy labels from information sources with unknown reliability such as existing buildings, which we refer to as weak oracles, for metadata mapping. This framework involves an interactive process, where a small set of sensor instances are progressively selected and labeled for it to learn how to aggregate the noisy labels as well as to predict sensor types.Two key challenges arise in designing the framework, namely, weak oracle reliability estimation and instance selection for querying. To address the first challenge, we develop a clustering-based approach for weak oracle reliability estimation to capitalize on the observation that weak oracles perform differently in different groups of instances. For the second challenge, we propose a disagreement-based query selection strategy to combine the potential effect of a labeled instance on both reducing classifier uncertainty and improving the quality of label aggregation. We evaluate our solution on a large collection of real-world building sensor data from 5 buildings with more than 11, 000 sensors of 18 different types. The experiment results validate the effectiveness of our solution, which outperforms a set of state-of-the-art baselines.