Vision Transformer and Language Model Based Radiology Report Generation

Vision Transformer and Language Model Based Radiology Report Generation
复制标题

DOI:
10.1109/access.2022.3232719
复制
发表时间:
2023
期刊:
影响因子:
3.9
通讯作者:
M. Mohsan;M. Akram;G. Rasool;N. Alghamdi;Muhammad Abdullah Aamer Baqai;Muhammad Abbas
M. Mohsan;M. Akram;G. Rasool;N. Alghamdi;Muhammad Abdullah Aamer Baqai;Muhammad Abbas
中科院分区:
计算机科学3区
文献类型:
--
作者:
M. Mohsan;M. Akram;G. Rasool;N. Alghamdi;Muhammad Abdullah Aamer Baqai;Muhammad Abbas

文献摘要

相似文献

变压器的最新进展利用了计算机视觉问题,从而产生了最先进的模型。基于transformer的模型在各种序列预测任务中,如语言翻译、情感分类和字幕生成,表现出了卓越的性能。通过字幕生成模型自动生成医学影像报告是语言模型的应用场景之一,具有很强的社会影响力。在这些模型中,卷积神经网络被用作编码器来获得空间信息,递归神经网络被用作解码器来生成字幕或医疗报告。然而,使用Transformer架构作为编码器和解码器在字幕或报告编写任务中仍然是未开发的。在这项研究中,我们探讨了在编码器中丢失空间偏差信息的影响,通过使用预训练的香草图像Transformer架构和联合收割机与不同的预训练语言transformer作为解码器。为了评价所提出的方法,使用了印第安纳州大学胸部X射线数据集,其中还针对不同的评价进行了消融研究。对比分析表明,所提出的方法具有显着的性能相比,现有的技术在不同的性能参数。
Recent advancements in transformers exploited computer vision problems which results in state-of-the-art models. Transformer-based models in various sequence prediction tasks such as language translation, sentiment classification, and caption generation have shown remarkable performance. Auto report generation scenarios in medical imaging through caption generation models is one of the applied scenarios for language models and have strong social impact. In these models, convolution neural networks have been used as encoder to gain spatial information and recurrent neural networks are used as decoder to generate caption or medical report. However, using transformer architecture as encoder and decoder in caption or report writing task is still unexplored. In this research, we explored the effect of losing spatial biasness information in encoder by using pre-trained vanilla image transformer architecture and combine it with different pre-trained language transformers as decoder. In order to evaluate the proposed methodology, the Indiana University Chest X-Rays dataset is used where ablation study is also conducted with respect to different evaluations. The comparative analysis shows that the proposed methodology has represented remarkable performance when compared with existing techniques in terms of different performance parameters.