XrayGPT: Chest Radiographs Summarization using Medical Vision-Language Models

XrayGPT: Chest Radiographs Summarization using Medical Vision-Language Models
复制标题

XrayGPT:使用医学视觉语言模型总结胸部 X 光片

DOI:
--
复制
发表时间:
2023
期刊:
arXiv.org
影响因子:
--
通讯作者:
F. Khan
F. Khan
中科院分区:
--
文献类型:
--
作者:
Omkar Thawakar;Abdelrahman M. Shaker;Sahal Shaji Mullappilly;Hisham Cholakkal;R. Anwer;Salman Siddique Khan;J. Laaksonen;F. Khan

文献摘要

被引文献

相似文献

大型视觉语言模型的最新突破,如巴德和GPT-4,展示了在执行广泛任务方面的非凡能力。这些模型是在包含数十亿公共图像-文本对的海量数据集上进行训练的,这些数据集具有不同的任务。然而,它们在特定任务领域(如放射学)的表现仍未得到充分研究,并且由于对生物医学图像的理解缺乏复杂性而可能受到限制。另一方面,会话医学模型取得了显著的成功,但主要集中在基于文本的分析上。在本文中,我们介绍了XrayGPT,一种新的会话医学视觉语言模型,可以分析和回答关于胸片的开放式问题。具体来说,我们使用简单的线性变换将医学视觉编码器(MedClip)与微调的大型语言模型(Vicuna)对齐。这种一致性使我们的模型具有卓越的视觉对话能力,基于对x光片和医学领域知识的深刻理解。为了提高法学硕士在医学背景下的表现,我们从自由文本放射学报告中生成了约217k交互式高质量摘要。这些总结有助于通过微调过程提高法学硕士的绩效。我们的方法为推进胸片自动分析开辟了新的研究途径。我们的开源演示、模型和指令集可以在:https://github.com/mbzuai-oryx/XrayGPT上获得。
The latest breakthroughs in large vision-language models, such as Bard and GPT-4, have showcased extraordinary abilities in performing a wide range of tasks. Such models are trained on massive datasets comprising billions of public image-text pairs with diverse tasks. However, their performance on task-specific domains, such as radiology, is still under-investigated and potentially limited due to a lack of sophistication in understanding biomedical images. On the other hand, conversational medical models have exhibited remarkable success but have mainly focused on text-based analysis. In this paper, we introduce XrayGPT, a novel conversational medical vision-language model that can analyze and answer open-ended questions about chest radiographs. Specifically, we align both medical visual encoder (MedClip) with a fine-tuned large language model (Vicuna), using a simple linear transformation. This alignment enables our model to possess exceptional visual conversation abilities, grounded in a deep understanding of radiographs and medical domain knowledge. To enhance the performance of LLMs in the medical context, we generate ~217k interactive and high-quality summaries from free-text radiology reports. These summaries serve to enhance the performance of LLMs through the fine-tuning process. Our approach opens up new avenues the research for advancing the automated analysis of chest radiographs. Our open-source demos, models, and instruction sets are available at: https://github.com/mbzuai-oryx/XrayGPT.