Qwen-VL: A Frontier Large Vision-Language Model with Versatile Abilities

Qwen-VL: A Frontier Large Vision-Language Model with Versatile Abilities
复制标题

DOI:
10.48550/arxiv.2308.12966
复制
发表时间:
2023
期刊:
ArXiv
影响因子:
--
通讯作者:
Jinze Bai;Shuai Bai;Shusheng Yang;Shijie Wang;Sinan Tan;Peng Wang;Junyang Lin;Chang Zhou;Jingren Zhou
Jinze Bai;Shuai Bai;Shusheng Yang;Shijie Wang;Sinan Tan;Peng Wang;Junyang Lin;Chang Zhou;Jingren Zhou
中科院分区:
其他
文献类型:
--
作者:
Jinze Bai;Shuai Bai;Shusheng Yang;Shijie Wang;Sinan Tan;Peng Wang;Junyang Lin;Chang Zhou;Jingren Zhou

文献摘要

被引文献

相似文献

我们介绍Qwen-VL系列,这是一套大规模的视觉语言模型,旨在感知和理解文本和图像。这些模型包括Qwen-VL和Qwen-VL-Chat,在图像字幕、问题回答、视觉定位和灵活交互等任务中表现出显著的性能。评估涵盖了广泛的任务,包括零镜头字幕、视觉或文档视觉问题回答和接地。我们证明了QWEN-VL的性能优于现有的大型视觉语言模型(LVLM)。我们介绍了他们的架构、培训、能力和性能,强调了他们对推进多模式人工智能的贡献。代码、演示和模型可在https://github.com/QwenLM/Qwen-VL上找到。
We introduce the Qwen-VL series, a set of large-scale vision-language models designed to perceive and understand both text and images. Comprising Qwen-VL and Qwen-VL-Chat, these models exhibit remarkable performance in tasks like image captioning, question answering, visual localization, and flexible interaction. The evaluation covers a wide range of tasks including zero-shot captioning, visual or document visual question answering, and grounding. We demonstrate the Qwen-VL outperforms existing Large Vision Language Models (LVLMs). We present their architecture, training, capabilities, and performance, highlighting their contributions to advancing multimodal artificial intelligence. Code, demo and models are available at https://github.com/QwenLM/Qwen-VL .