ARLPE: A meta reinforcement learning framework for glucose regulation in type 1 diabetics

ARLPE: A meta reinforcement learning framework for glucose regulation in type 1 diabetics
复制标题

DOI:
10.1016/j.eswa.2023.120156
复制
发表时间:
2023-05
期刊:
Expert Syst. Appl.
影响因子:
--
通讯作者:
Xuehui Yu;Yi Guan;Lian Yan;Shulang Li;Xuelian Fu;Jingchi Jiang
Xuehui Yu;Yi Guan;Lian Yan;Shulang Li;Xuelian Fu;Jingchi Jiang
中科院分区:
其他
文献类型:
--
作者:
Xuehui Yu;Yi Guan;Lian Yan;Shulang Li;Xuelian Fu;Jingchi Jiang

文献摘要

相似文献

具有自主控制算法的体外人工胰腺已证明其在1型糖尿病血糖调节中的有效性。尽管如此,大多数现有的算法不能适应未知的患者有限的临床数据。为了实现未知患者的自动血糖调节,即使有不确定性和噪声,我们提出了个性化嵌入主动强化学习(ARLPE)用于正常血糖维持。我们的框架包含一个元训练期和一个微调期。元训练期旨在学习:(1)葡萄糖调节的通用策略和(2)将个性化信息和上下文汇总到嵌入中的概率编码器。微调期旨在借助主动学习模块为未知患者生成个性化策略,以探索有价值的经验。对多个患者的实验表明,我们的算法不仅可以收敛血糖的正常范围,避免低血糖,但也实现了一个新的陌生患者的血糖调节使用有限的血糖数据(只有25个样本)。ARLPE在成人和青少年队列中分别达到98.63%和97.93%的范围内时间(TIR)评分,显著优于最先进的竞争性血糖调节方法。它显示了为糖尿病患者生成个性化临床策略的巨大潜力。
External artificial pancreas with autonomous control algorithms has proved its effectiveness in glucose regulation for type 1 diabetes. Nonetheless, most existing algorithms cannot adapt to unknown patients with limited clinical data. To achieve the automatic glucose regulation of unknown patients even with uncertainties and noises, we propose Active Reinforcement Learning with Personalized Embeddings (ARLPE) for normoglycemia maintenance. Our framework contains a meta-training period and a fine-tuning period. The meta-training period aims to learn: (1) a generalized policy for glucose regulation and (2) a probabilistic encoder that summarizes the personalized information and context into an embedding. The fine-tuning period is designed to generate a personalized policy for the unknown patient with the help of an active learning module to explore valuable experiences. Experiments on multiple patients demonstrate that our algorithm can not only converge blood glucose to the normoglycemic bounds and avoid hypoglycemia but also achieve the glucose regulation of a new unacquainted patient using limited BG data (only 25 samples). ARLPE achieves the time in range (TIR) score of 98.63% and 97.93% in adult and adolescent cohorts, respectively, significantly outperforming the state-of-the-art competing methods for glucose regulation. It shows the great potential of generating personalized clinical strategies for diabetics.