That’s the Wrong Lung! Evaluating and Improving the Interpretability of Unsupervised Multimodal Encoders for Medical Data

That’s the Wrong Lung! Evaluating and Improving the Interpretability of Unsupervised Multimodal Encoders for Medical Data
复制标题

DOI:
10.48550/arxiv.2210.06565
复制
发表时间:
2022-10
期刊:
Proceedings of the Conference on Empirical Methods in Natural Language Processing. Conference on Empirical Methods in Natural Language Processing
影响因子:
--
通讯作者:
Denis Jered McInerney;Geoffrey S. Young;Jan-Willem van de Meent;Byron Wallace
Denis Jered McInerney;Geoffrey S. Young;Jan-Willem van de Meent;Byron Wallace
中科院分区:
其他
文献类型:
--
作者:
Denis Jered McInerney;Geoffrey S. Young;Jan-Willem van de Meent;Byron Wallace

文献摘要

相似文献

电子健康记录(EHR)上的预训练多模态模型提供了一种学习表示的方法,可以在最少的监督下转移到下游任务。最近的多模态模型诱导图像区域和句子之间的软局部对齐。这在医学领域中特别令人感兴趣,其中对齐可能会突出显示图像中与自由文本中描述的特定现象相关的区域。虽然过去的工作表明,注意力“热图”可以用这种方式来解释,但很少有人对这种对齐进行评估。我们比较了来自最先进的多模态(图像和文本)EHR模型的对齐与将图像区域链接到句子的人类注释。我们的主要发现是,文本对注意力的影响往往很弱或不直观;对齐并不能始终如一地反映基本的解剖信息。此外,合成修改-例如用“左”替换“右”-基本上不会影响高光。简单的技术,如允许模型选择不关注图像和少数镜头微调,显示出在很少或没有监督的情况下改善对齐的能力。我们将代码和检查点开源。
Pretraining multimodal models on Electronic Health Records (EHRs) provides a means of learning representations that can transfer to downstream tasks with minimal supervision. Recent multimodal models induce soft local alignments between image regions and sentences. This is of particular interest in the medical domain, where alignments might highlight regions in an image relevant to specific phenomena described in free-text. While past work has suggested that attention “heatmaps” can be interpreted in this manner, there has been little evaluation of such alignments. We compare alignments from a state-of-the-art multimodal (image and text) model for EHR with human annotations that link image regions to sentences. Our main finding is that the text has an often weak or unintuitive influence on attention; alignments do not consistently reflect basic anatomical information. Moreover, synthetic modifications — such as substituting “left” for “right” — do not substantially influence highlights. Simple techniques such as allowing the model to opt out of attending to the image and few-shot finetuning show promise in terms of their ability to improve alignments with very little or no supervision. We make our code and checkpoints open-source.