计算机科学
编码器
任务(项目管理)
领域(数学分析)
图像编辑
多样性(控制论)
编码(集合论)
图像(数学)
人工智能
自然语言处理
质量(理念)
人机交互
计算机视觉
程序设计语言
工程类
操作系统
认识论
数学分析
哲学
集合(抽象数据类型)
系统工程
数学
作者
Katherine Crowson,Stella Biderman,Daniel Kornis,Dashiell Stander,Eric Hallahan,Louis Castricato,Edward Raff
标识
DOI:10.48550/arxiv.2204.08583
摘要
Generating and editing images from open domain text prompts is a challenging task that heretofore has required expensive and specially trained models. We demonstrate a novel methodology for both tasks which is capable of producing images of high visual quality from text prompts of significant semantic complexity without any training by using a multimodal encoder to guide image generations. We demonstrate on a variety of tasks how using CLIP [37] to guide VQGAN [11] produces higher visual quality outputs than prior, less flexible approaches like DALL-E [38], GLIDE [33] and Open-Edit [24], despite not being trained for the tasks presented. Our code is available in a public repository.
科研通智能强力驱动
Strongly Powered by AbleSci AI