Expressive Text-to-Image Generation with Rich Text

Expressive Text-to-Image Generation with Rich Text
复制标题

DOI:
10.1109/iccv51070.2023.00694
复制
发表时间:
2023-04
期刊:
2023 IEEE/CVF International Conference on Computer Vision (ICCV)
影响因子:
--
通讯作者:
Songwei Ge;Taesung Park;Jun-Yan Zhu;Jia-Bin Huang
Songwei Ge;Taesung Park;Jun-Yan Zhu;Jia-Bin Huang
中科院分区:
其他
文献类型:
--
作者:
Songwei Ge;Taesung Park;Jun-Yan Zhu;Jia-Bin Huang

文献摘要

相似文献

纯文本已经成为文本到图像合成的流行接口。然而,其有限的定制选项阻碍了用户准确地描述所需的输出。例如,纯文本很难指定连续的数量,例如精确的RGB颜色值或每个单词的重要性。此外,为复杂场景创建详细的文本提示对于人类来说是乏味的,并且对于文本编码器来说是具有挑战性的。为了应对这些挑战,我们建议使用一个富文本编辑器,它支持字体样式、大小、颜色和脚注等格式。我们从富文本中提取每个单词的属性,以实现本地样式控制,显式标记重新加权,精确的颜色渲染和详细的区域合成。我们通过基于区域的传播过程实现这些能力。我们首先获得每个单词的区域的基础上的注意力地图的扩散过程中使用纯文本。对于每个区域,我们通过创建特定于区域的详细提示和应用特定于区域的指导来执行其文本属性,并通过基于区域的注入来保持其对纯文本生成的保真度。我们提出了从富文本生成图像的各种例子,并证明了我们的方法优于强基线与定量评估。
Plain text has become a prevalent interface for text-to-image synthesis. However, its limited customization options hinder users from accurately describing desired outputs. For example, plain text makes it hard to specify continuous quantities, such as the precise RGB color value or importance of each word. Furthermore, creating detailed text prompts for complex scenes is tedious for humans to write and challenging for text encoders to interpret. To address these challenges, we propose using a rich-text editor supporting formats such as font style, size, color, and footnote. We extract each word’s attributes from rich text to enable local style control, explicit token reweighting, precise color rendering, and detailed region synthesis. We achieve these capabilities through a region-based diffusion process. We first obtain each word’s region based on attention maps of a diffusion process using plain text. For each region, we enforce its text attributes by creating region-specific detailed prompts and applying region-specific guidance, and maintain its fidelity against plain-text generation through region-based injections. We present various examples of image generation from rich text and demonstrate that our method outperforms strong baselines with quantitative evaluations.