An Efficient Meta-Reinforcement Learning Approach for Circuit Linearity Calibration via Style Injection

An Efficient Meta-Reinforcement Learning Approach for Circuit Linearity Calibration via Style Injection
复制标题

DOI:
10.1109/mwscas57524.2023.10406135
复制
发表时间:
2023-08
期刊:
2023 IEEE 66th International Midwest Symposium on Circuits and Systems (MWSCAS)
影响因子:
--
通讯作者:
Chao Rong;J. Paramesh;L. R. Carley
Chao Rong;J. Paramesh;L. R. Carley
中科院分区:
其他
文献类型:
--
作者:
Chao Rong;J. Paramesh;L. R. Carley

文献摘要

相似文献

在可观测性有限的情况下,电路线性校正可以表示一组高维搜索问题。例如,数字时间转换器(DTC)的线性度校准是现代数字锁相环(DPLL)的重要组成部分,它是高维搜索问题的一个例子,因为测量ps延迟的困难阻碍了现有逐级校准的方法。而且,由于温度(T)和电源电压(V)的变化,校准的DTC可能再次变为非线性。以前的工作报告了一种深度强化学习框架,该框架能够使用非线性校准库执行DTC线性校准;然而,这项先前的工作并没有解决在面对温度和电源电压变化时保持校准的问题。本文提出了一种元强化学习(RL)方法,使元强化学习主体能够在温度和/或电压变化时快速适应新的环境。受Style生成性对抗网络(StyleGANs)的启发,我们提出将温度和电压的变化作为电路的样式来处理。与使用电路传感器来检测T和V变化的传统方法不同,我们利用机器学习(ML)传感器来隐含地推断广泛的环境变化。来自ML传感器的样式信息随后被注入策略网络的一小部分,调整其权重。作为概念验证,我们首先设计了一个在常压(1V)和常温(27°C)拐角(NVNT)下工作的5位DTC。RL代理在NVNT环境中开始其培训。在这一初始阶段之后,该试剂的任务是适应不同温度和电源电压的环境。实验结果表明,在环境变化的情况下,该方法可以在1万步搜索步长内将积分非线性(INL)降低到0.5LSB以下。与从随机初始化策略和训练策略开始学习相比,Meta-RL方法完成线性校正的步骤分别减少了63%和47%。我们的方法也适用于许多其他类型的模拟电路和射频电路的校准。
Circuit linearity calibration can represent a set of high-dimensional search problems if the observability is limited. For example, linearity calibration of digital-to-time converters (DTC), an essential building block of modern digital phase-locked loops (DPLLs), is an example of a high-dimensional search problem as difficulty of measuring ps delays hinders prior methods that calibrate stage by stage. And, a calibrated DTC can become nonlinear again due to changes in temperature (T) and power supply voltage (V). Prior work reports a deep reinforcement learning framework that is capable of performing DTC linearity calibration with nonlinear calibration banks; however, this prior work does not address maintaining calibration in the face of temperature and supply voltage variations. In this paper, we present a meta-reinforcement learning (RL) method that can enable the RL agent to quickly adapt to a new environment when the temperature and/or voltage change. Inspired by the Style Generative Adversarial Networks (StyleGANs), we propose to treat temperature and voltage changes as the styles of the circuits. In contrast to traditional methods employing circuit sensors to detect changes in T and V, we utilize a machine learning (ML) sensor, to implicitly infer a wide range of environmental changes. The style information from the ML sensor is subsequently injected into a small portion of the policy network, modulating its weights. As a proof of concept, we first designed a 5-bit DTC at the normal voltage (1V) and normal temperature (27°C) corner (NVNT) as the environment. The RL agent begins its training in the NVNT environment. Following this initial phase, the agent is then tasked with adapting to environments with different temperature and supply voltages. Our results show that the proposed technique can reduce the Integral Non-Linearity (INL) to less than 0.5 LSB within 10, 000 search steps in a changed environment. Compared to starting learning from a random initialized policy and a trained policy, the proposed meta-RL approach takes 63% and 47% fewer steps to complete the linearity calibration, respectively. Our method is also applicable to the calibration of many other kinds of analog and RF circuits.