Attention-Based Interrelation Modeling for Explainable Automated Driving

Attention-Based Interrelation Modeling for Explainable Automated Driving
复制标题

DOI:
10.1109/tiv.2022.3229682
复制
发表时间:
2023-02
影响因子:
8.2
通讯作者:
Zhengming Zhang;Renran Tian;Rini Sherony;Joshua E. Domeyer;Zhengming Ding
Zhengming Zhang;Renran Tian;Rini Sherony;Joshua E. Domeyer;Zhengming Ding
中科院分区:
工程技术2区
文献类型:
--
作者:
Zhengming Zhang;Renran Tian;Rini Sherony;Joshua E. Domeyer;Zhengming Ding

文献摘要

被引文献

相似文献

自动驾驶希望在运动规划和混合交通环境中与行人互动等任务上有更好的性能。深度学习算法可以在这些任务中实现高性能,具有显着的视觉场景理解和泛化能力。然而,当使用常见的场景解析方法来训练端到端模型时,这些算法中的可解释性限制抑制了它们在全自动驾驶中的实现。主要挑战包括算法性能缺陷和不一致性,人工智能透明度不足,用户信任度下降,以及破坏人类与人工智能的交互。这项研究通过提供多模态解释,特别是在与行人互动时,有助于自动驾驶系统的决策性能和透明度。该算法结合全局视觉特征和相互关系特征,通过解析场景图像作为自构造的图,并使用基于注意力的模块来捕获自我车辆和其他交通相关对象之间的相互关系。输出模块在生成语义文本解释的同时做出决策。结果表明,融合的功能,从全球框架和相互关系图改善决策和解释预测相比,两个国家的最先进的基准算法。相互关系模块还通过公开用于决策的视觉注意力来增强算法的透明度。沿着揭示了两个预测任务的相关性特征的重要性以及在具有层次标签的数据集上进行多任务学习的潜在机制。所提出的模型改善了行人互动过程中的驾驶决策,具有可理解的推理线索,用于为人类用户构建适当的自动驾驶性能心理模型。
Automated driving desires better performance on tasks like motion planning and interacting with pedestrians in mixed-traffic environments. Deep learning algorithms can achieve high performance in these tasks with remarkable visual scene understanding and generalization abilities. However, when common scene-parsing methods are used to train end-to-end models, limitations of explainability in such algorithms inhibit their implementations in fully automated driving. The main challenges include algorithm performance deficiencies and inconsistencies, insufficient AI transparency, degraded user trust, and undermining human-AI interactions. This research aids the decision-making performance and transparency of automated driving systems by providing multi-modal explanations, especially when interacting with pedestrians. The proposed algorithm combines global visual features and interrelation features by parsing scene images as self-constructed graphs and using an attention-based module to capture the interrelationship among the ego-vehicle and other traffic-related objects. The output modules make decisions while simultaneously generating semantic text explanations. The results show that the fusion of the features from global frames and interrelational graphs improves decision-making and explanation predictions compared to two state-of-the-art benchmark algorithms. The interrelation module also enhances algorithm transparency by disclosing the visual attention used for decision-making. The importance of interrelation features on the two prediction tasks is further revealed along with the underlying mechanism of multitask learning on the datasets with hierarchical labels. The proposed model improves driving decision-making during pedestrian interactions with intelligible reasoning cues for building an appropriate mental model of automated driving performance for human users.