课题基金 / 基金详情

CAREER: Hierarchical Reinforcement Learning Framework for Safe Dynamic Bipedal Locomotion

CAREER: Hierarchical Reinforcement Learning Framework for Safe Dynamic Bipedal Locomotion
职业:安全动态双足运动的分层强化学习框架
批准号:
2144156
负责人:
Ayonga Hereid
金额:
$60.0万
依托单位:
依托单位国家:
美国
项目类别:
Continuing Grant
财政年份:
2022
资助国家:
美国
项目状态:
未结题
起止时间:
2022-04-15 至 2027-03-31

项目摘要

项目成果

相似基金

相关文献

中文摘要
翻译
双足机器人的拟人化设计使这些机器具有独特的优势,可以在具有挑战性的地形(例如,天然不平的地面,山丘和楼梯)中导航,在受限的环境中操作(例如,为人类操作员设计的房屋或仓库中的狭窄垂直空间),并以自然的方式与其他人互动。然而,由于对运动控制器的基本机制缺乏了解,两足机器人的安全和动态行为的技术实现仍然具有挑战性。这种理解也将有助于改进辅助设备,如下肢外骨骼,它可以帮助恢复由于中风或其他运动障碍而失去的活动能力。这个学院早期职业发展(Career)项目旨在通过一个新的基于学习的反馈运动控制框架,显著推进双足机器人和下肢外骨骼的技术,特别关注在现实世界中实验实现安全和动态的双足运动。受人类如何分层学习复杂任务的启发,该项目将通过物理启发的分层结构对运动控制进行重大创新,以实现安全的双足运动。此外,我们将在我们的协调教育和推广计划中利用双足机器人的先天吸引力,吸引STEM教育和研究项目中各级代表性不足的学生。该项目的成功完成有可能加速双足机器人在工业和公共卫生领域的实际应用,并提高STEM人才库的多样性。双足机器人具有固有的不稳定性,并且由多个自由度组成。尽管目前在机器人操作和移动机器人方面取得了成就,但典型的强化学习(RL)算法在用于双足机器人时容易失败且难以扩展。如果没有适当的控制,机器人将会下降,导致非常稀疏和不连续的奖励,导致RL算法不收敛。双足机器人的高维也增加了经典“平面”强化学习算法的搜索空间,导致采样效率低下和冗长(可能不成功)的训练过程。现有的许多强化学习在两足机器人上的应用没有考虑到机器人的物理限制,因此无法在机器人硬件上实现。因此,本研究将通过追求以下四个研究目标来解决这些科学挑战:(G1)构建分层学习结构,通过对不同层次技能的时间抽象,降低任务复杂度,使机器人能够高效地探索高维行为空间;(G2)设计概率安全的“安全过滤器”,通过控制障碍函数确保和引导安全的策略学习;(G3)通过分层方式模仿物理启发的模板模型,引导策略探索,提高学习效率。最后(G4)弥合现实世界部署中“模拟到真实”的差距。最终的框架将在现实环境中通过3D双足机器人(Digit)和下肢外骨骼(ATALANTE)在无脊髓损伤健康人类受试者的临床前试验中进行实验验证。该项目由跨部门机器人基础研究项目支持,由工程(ENG)和计算机与信息科学与工程(CISE)联合管理和资助。该奖项反映了美国国家科学基金会的法定使命,并通过使用基金会的知识价值和更广泛的影响审查标准进行评估,被认为值得支持。
英文摘要
The anthropomorphic design of bipedal robots equips these machines with unique advantages navigating challenging terrains (e.g., natural uneven grounds, hills, and stairs), operating in restricted environments (e.g., narrow vertical spaces in house or warehouse designed for human operators), and interacting with other humans in a natural manner. However, the technological realization of safe and dynamic behaviors in bipedal robots remains challenging due to the fundamental lack of understanding of underlying mechanisms of locomotion controllers. Such understanding will also help improve assistive devices such as lower-limb exoskeletons, which could help restore mobility lost due to stroke or other movement disorders. This Faculty Early Career Development (CAREER) project aims to significantly advance the technology of bipedal robots and lower-limb exoskeletons through a novel learning-based feedback motion control framework, with a specific focus on experimentally realizing safe and dynamic bipedal locomotion in real-world settings. Inspired by how humans learn complex tasks hierarchically, this project will make major innovations to motion control for safe bipedal locomotion through a physics-inspired hierarchical structure. Moreover, we will leverage the innate appeal of bipedal robots in our coordinated education and outreach plans to engage underrepresented students of various levels in STEM education and research programs. The successful completion of this project has the potential to accelerate real-world applications of bipedal robots in industry and public health and improve the diversity in the STEM talent pool.Bipedal robots are inherently unstable and consist of many degrees of freedom. Despite current achievements in robot manipulation and mobile robots, typical reinforcement learning (RL) algorithms are prone to fail and are difficult to scale when used for bipedal robots. The robot will fall without proper control, resulting in very sparse and discontinuous rewards, causing the RL algorithms not to converge. The high dimensionality of bipedal robots also increases the search space of the classic “flat” RL algorithms, leading to sampling inefficiency and a lengthy (potentially unsuccessful) training process. Many existing applications of RL on bipedal robots do not respect the physical limitations of the robot, and consequently, cannot be implemented on robot hardware. This research will therefore address these scientific challenges by pursuing the following four research goals: (G1) develop a hierarchical learning structure that enables the robot to efficiently explore the high-dimensional behavior space by reducing the task complexity through the temporal abstraction of skills at different levels, (G2) design a probabilistically safe “safety filter” to ensure and guide safe policy learning via control barrier functions, (G3) improve learning efficiency through guided policy exploration that imitates physics-inspired template models in a layered fashion, and finally (G4) bridge the “sim-to-real” gap for real-world deployments. The resulting framework will be experimentally demonstrated with a 3D bipedal robot (Digit) in real-world settings and a lower-limb exoskeleton (ATALANTE) in pre-clinical trials with non-Spinal-Cord-Injury healthy human subjects.This project is supported by the cross-directorate Foundational Research in Robotics program, jointly managed and funded by the Directorates for Engineering (ENG) and Computer and Information Science and Engineering (CISE).This award reflects NSF's statutory mission and has been deemed worthy of support through evaluation using the Foundation's intellectual merit and broader impacts review criteria.
期刊论文(9)
专著(0)
科研奖励(0)
会议论文
MELP: Model Embedded Linear Policies for Robust Bipedal Hopping
MELP:用于鲁棒双足跳跃的嵌入式线性策略模型
DOI: 10.1109/iros55552.2023.10342023
发表时间: 2023
期刊: IEEE
影响因子: --
作者: [Soni, Raghav, Castillo, Guillermo A., Krishna, Lokesh, Hereid, Ayonga, Kolathaya, Shishir]
通讯作者: Kolathaya, Shishir
Time-Varying ALIP Model and Robust Foot-Placement Control for Underactuated Bipedal Robotic Walking on a Swaying Rigid Surface
摇摆刚性表面欠驱动双足机器人行走的时变 ALIP 模型和鲁棒足部放置控制
DOI: 10.23919/acc55779.2023.10156254
发表时间: 2023
期刊: Proceedings of the American Control Conference
影响因子: --
作者: [Gao, Yuan, Gong, Yukai, Paredes, Victor, Hereid, Ayonga, Gu, Yan]
通讯作者: Gu, Yan
DOI: 10.1109/icra48891.2023.10160671
发表时间: 2023-05
期刊: 2023 IEEE International Conference on Robotics and Automation (ICRA)
影响因子: --
作者: [Chengyang Peng;Octavian A. Donca;Guillermo A. Castillo;Ayonga Hereid]
通讯作者: Chengyang Peng;Octavian A. Donca;Guillermo A. Castillo;Ayonga Hereid
DOI: 10.48550/arxiv.2309.15740
发表时间: 2023-09
期刊: ArXiv
影响因子: --
作者: [Guillermo A. Castillo;Bowen Weng;Wei Zhang;Ayonga Hereid]
通讯作者: Guillermo A. Castillo;Bowen Weng;Wei Zhang;Ayonga Hereid
共 8 条
    国内基金
    海外基金
    丙烷脱氢Pt@hierarchical zeolite催化剂的设计制备与反应调控
    • 批准号:
      22178062
    • 项目类别:
      面上项目
    • 资助金额:
      60万元
    • 批准年份:
      2021
    • 负责人:
      朱海波
    • 依托单位: