Reinforcement Learning Enabled Autonomous Manufacturing Using Transfer Learning and Probabilistic Reward Modeling

Reinforcement Learning Enabled Autonomous Manufacturing Using Transfer Learning and Probabilistic Reward Modeling
复制标题

DOI:
10.1109/lcsys.2022.3188014
复制
发表时间:
2023-01-01
影响因子:
3
通讯作者:
Hoelzle, David
Hoelzle, David
中科院分区:
其他
文献类型:
--
作者:
Alam, Md Ferdous;Shtein, Max;Hoelzle, David

文献摘要

被引文献

相似文献

在这里,我们提出了一个强化学习使物理自主制造系统(AMS),能够学习的制造工艺参数自主制造具有所需的性能特性的复杂几何工件。由于原材料、机器利用率和劳动力成本的可变成本很高,传统RL算法的样本效率很差,这对现实世界的制造决策提出了挑战。为了使决策样本有效,我们建议利用基于第一性原理的源任务进行训练,从训练的知识中转移有效的表示,然后使用这些表示与物理系统进行交互,以学习目标奖励函数的概率模型。我们将这个想法部署到一个新的数据集,从一个自定义的物理AMS机器,可以自主制造声子晶体,一个复杂的几何工件与光谱响应的性能特征。我们证明了我们的方法使用低至25工件建模的目标奖励函数的有趣的部分,并找到一个工件与高奖励。这项任务通常需要人工设计声子晶体和大量的经验迭代,数量级为数百。
Here we propose a reinforcement learning enabled physical autonomous manufacturing system (AMS) that is capable of learning the manufacturing process parameters to autonomously fabricate a complex-geometry artifact with desired performance characteristics. The poor sample efficiency of traditional RL algorithms challenges real-world manufacturing decision making due to a high variable cost from raw material, machine utilization, and labor costs. To make decision making sample efficient, we propose to leverage a first-principles based source task for training, transfer effective representations from trained knowledge, and then use these representations to interact with the physical system to learn a probabilistic model of the target reward function. We deploy this idea to a novel dataset obtained from a custom physical AMS machine that can autonomously manufacture phononic crystals, a complex geometry artifact with spectral response as performance characteristic. We demonstrate that our method uses as low as 25 artifacts to model the interesting part of the target reward function and find an artifact with high reward. This task typically requires manual design of phononic crystals and extensive empirical iterations on the order of hundreds.