Task-Driven Privacy-Preserving Data-Sharing Framework for the Industrial Internet

Task-Driven Privacy-Preserving Data-Sharing Framework for the Industrial Internet
复制标题

DOI:
10.1109/bigdata55660.2022.10020861
复制
发表时间:
2022-12
期刊:
2022 IEEE International Conference on Big Data (Big Data)
影响因子:
--
通讯作者:
Parshin Shojaee;Yingyan Zeng;Muntasir Wahed;Avi Seth;Ran Jin;Ismini Lourentzou
Parshin Shojaee;Yingyan Zeng;Muntasir Wahed;Avi Seth;Ran Jin;Ismini Lourentzou
中科院分区:
其他
文献类型:
--
作者:
Parshin Shojaee;Yingyan Zeng;Muntasir Wahed;Avi Seth;Ran Jin;Ismini Lourentzou

文献摘要

相似文献

工业互联网为参与的企业提供了一个协同计算平台,允许收集大数据用于机器学习任务。尽管这些技术有望加速培训和部署,并有可能通过数据共享来优化决策过程,但对信息隐私的日益关注对这些技术的采用产生了影响。由于企业倾向于保持数据的私密性,这限制了互操作性。虽然之前的工作在很大程度上探索了隐私保护机制,但所提出的方法天真地对所有参与者共享的数据进行平均或随机采样,而不是为特定的下游学习任务选择最适合的子集。由于工业互联网中异构机器学习任务缺乏有效的数据共享机制,我们提出了PriED,一个任务驱动的数据共享框架,它有选择性地融合来自参与者的共享数据和本地数据,以提高监督学习的性能。PriED利用保护隐私的数据蒸馏来促进数据交换,并利用动态数据选择来优化下游机器学习任务。我们在一个真实的半导体制造案例研究中展示了性能改进。
Industrial Internet provides a collaborative computational platform for participating enterprises, allowing the collection of big data for machine learning tasks. Despite the promise of training and deployment acceleration, and the potential to optimize decision-making processes through data-sharing, the adoption of such technologies is impacted by the increasing concerns about information privacy. As enterprises prefer to keep data private, this limits interoperability. While prior work has largely explored privacy-preserving mechanisms, the proposed methods naively average or randomly sample data shared from all participants instead of selecting the most well-suited subsets for a particular downstream learning task. Motivated by the lack of effective data-sharing mechanisms for heterogeneous machine learning tasks in Industrial Internet, we propose PriED, a task-driven data-sharing framework that selectively fuses shared data and local data from participants to improve supervised learning performance. PriED utilizes privacy-preserving data distillation to facilitate data exchange, and dynamic data selection to optimize downstream machine learning tasks. We demonstrate performance improvements on a real semiconductor manufacturing case study.