BraIN: A Bidirectional Generative Adversarial Networks for image captions

BraIN: A Bidirectional Generative Adversarial Networks for image captions
复制标题

DOI:
10.1145/3446132.3446406
复制
发表时间:
2020-12
期刊:
Proceedings of the 2020 3rd International Conference on Algorithms, Computing and Artificial Intelligence
影响因子:
--
通讯作者:
Yuhui Wang;D. Cook
Yuhui Wang;D. Cook
中科院分区:
其他
文献类型:
--
作者:
Yuhui Wang;D. Cook

文献摘要

被引文献

相似文献

尽管在图像字幕方面取得了进展,但机器生成的字幕和人类生成的字幕仍然相当不同。基于自动指标,机器生成的字幕表现良好。然而,它们缺乏自然性,这是人类语言的一个基本特征,因为它们最大化了训练样本的可能性。我们提出了一种新的模型来生成比以前的方法更接近人类的字幕。我们的模型包括注意机制、双向语言生成模型和条件生成对抗性网络。具体地说,注意力机制通过将重要信息分割成更小的片段来捕捉图像细节。双向语言生成模型考虑了多个角度,生成了类似人类的句子。同时,条件生成对抗性网络通过比较一组字幕来提高句子质量。为了评估我们模型的性能,我们比较了人类对大脑生成字幕的偏好和基线方法。我们还使用自动度量将结果与实际的人工生成的字幕进行比较。结果表明,与基线方法相比,我们的模型能够生成更多的类似于人类的字幕。
Although progress has been made in image captioning, machine-generated captions and human-generated captions are still quite distinct. Machine-generated captions perform well based on automated metrics. However, they lack naturalness, an essential characteristic of human language, because they maximize the likelihood of training samples. We propose a novel model to generate more human-like captions than has been accomplished with prior methods. Our model includes an attention mechanism, a bidirectional language generation model, and a conditional generative adversarial network. Specifically, the attention mechanism captures image details by segmenting important information into smaller pieces. The bidirectional language generation model produces human-like sentences by considering multiple perspectives. Simultaneously, the conditional generative adversarial network increases sentence quality by comparing a set of captions. To evaluate the performance of our model, we compare human preferences for BraIN-generated captions with baseline methods. We also compare results with actual human-generated captions using automated metrics. Results show our model is capable of producing more human-like captions than baseline methods.