计算机科学
图像编辑
发电机(电路理论)
人工智能
鉴别器
图像(数学)
可扩展性
深度学习
可视化
计算机视觉
安全性令牌
视频编辑
领域(数学)
编码器
图像处理
语言模型
自然语言
上下文图像分类
机器学习
模式识别(心理学)
对抗制
像素
图像检索
图像翻译
自然语言处理
语义学(计算机科学)
作者
Bing Yang,Xueqin Xiang,Yong Peng,Wanzeng Kong,Jianhai Zhang,Jinliang Yao
标识
DOI:10.1109/tvcg.2026.3678876
摘要
Automatic real image editing offers unprecedented freedom to modify the appearance of the image or to edit a few objects through natural language. Recent scalable model families such as diffusion models have showcased remarkable proficiency in editing highly realistic images due to the introduction of vast amounts of training data and large pretrained language models. However, these large diffusion models require iterative evaluation that would significantly hinder the pace of image editing. Moreover, the pioneering work in this field necessitates the learning of a unique textual token that corresponds to each input image, or a group of images containing the same object, leading to the generation of redundant and fragmented models. Given the aforementioned problems, we suggest a novel Large pretrained models Assistant text-guided image Editing adversarial Network (LAE-Net) in this paper. More concretely, we introduce a deep semantic editing network to globally transfer text information among different isolated editing blocks, which would extract features from the source image to differentiate text-required areas from text-irrelevant ones. Furthermore, based on idea that the multi-modal CLIP model, leveraging vision-language alignment, captures comprehensive global semantic cues, whereas the vision-centric DINO model specializes in delivering intricate, fine-grained pixel-level details, the powerful discriminator of LAE-Net is designed by harnessing the visual embeddings derived from both the CLIP and DINO models separately to boost the visual discriminant capability and facilitate training a strong generator for conditioning image generation. Comprehensive experimental evaluations show that our LAE-Net not only delivers outstanding performance but also surpasses several cutting-edge models.
科研通智能强力驱动
Strongly Powered by AbleSci AI