Zero-Shot Text-to-Image Generation

Zero-Shot Text-to-Image Generation
复制标题

DOI:
--
复制
发表时间:
2021-02
期刊:
ArXiv
影响因子:
--
通讯作者:
A. Ramesh;Mikhail Pavlov;Gabriel Goh;S. Gray;Chelsea Voss;Alec Radford;Mark Chen;I. Sutskever-I.-Sutske
A. Ramesh;Mikhail Pavlov;Gabriel Goh;S. Gray;Chelsea Voss;Alec Radford;Mark Chen;I. Sutskever-I.-Sutske
中科院分区:
其他
文献类型:
--
作者:
A. Ramesh;Mikhail Pavlov;Gabriel Goh;S. Gray;Chelsea Voss;Alec Radford;Mark Chen;I. Sutskever-I.-Sutske

文献摘要

被引文献

相似文献

传统上,文本到图像生成的重点是为固定数据集的训练找到更好的建模假设。这些假设可能涉及复杂的体系结构、辅助损失或在训练期间提供的诸如对象部分标签或分割掩码之类的侧信息。我们描述了一种基于转换器的简单方法,该转换器将文本和图像标记自回归地建模为单个数据流。有了足够的数据和规模,我们的方法在以零射击方式评估时,与以前的领域特定模型具有竞争力。
Text-to-image generation has traditionally focused on finding better modeling assumptions for training on a fixed dataset. These assumptions might involve complex architectures, auxiliary losses, or side information such as object part labels or segmentation masks supplied during training. We describe a simple approach for this task based on a transformer that autoregressively models the text and image tokens as a single stream of data. With sufficient data and scale, our approach is competitive with previous domain-specific models when evaluated in a zero-shot fashion.