亲爱的研友该休息了!由于当前在线用户较少,发布求助请尽量完整地填写文献信息,科研通机器人24小时在线,伴您度过漫漫科研夜!身体可是革命的本钱,早点休息,好梦!

LAE-Net: Large Pretrained Models Assistant Text-Guided Image Editing Adversarial Network

计算机科学 图像编辑 发电机(电路理论) 人工智能 鉴别器 图像(数学) 可扩展性 深度学习 可视化 计算机视觉 安全性令牌 视频编辑 领域(数学) 编码器 图像处理 语言模型 自然语言 上下文图像分类 机器学习 模式识别(心理学) 对抗制 像素 图像检索 图像翻译 自然语言处理 语义学(计算机科学)
作者
Bing Yang,Xueqin Xiang,Yong Peng,Wanzeng Kong,Jianhai Zhang,Jinliang Yao
出处
期刊:IEEE Transactions on Visualization and Computer Graphics [Institute of Electrical and Electronics Engineers]
卷期号:32 (7): 6226-6239
标识
DOI:10.1109/tvcg.2026.3678876
摘要

Automatic real image editing offers unprecedented freedom to modify the appearance of the image or to edit a few objects through natural language. Recent scalable model families such as diffusion models have showcased remarkable proficiency in editing highly realistic images due to the introduction of vast amounts of training data and large pretrained language models. However, these large diffusion models require iterative evaluation that would significantly hinder the pace of image editing. Moreover, the pioneering work in this field necessitates the learning of a unique textual token that corresponds to each input image, or a group of images containing the same object, leading to the generation of redundant and fragmented models. Given the aforementioned problems, we suggest a novel Large pretrained models Assistant text-guided image Editing adversarial Network (LAE-Net) in this paper. More concretely, we introduce a deep semantic editing network to globally transfer text information among different isolated editing blocks, which would extract features from the source image to differentiate text-required areas from text-irrelevant ones. Furthermore, based on idea that the multi-modal CLIP model, leveraging vision-language alignment, captures comprehensive global semantic cues, whereas the vision-centric DINO model specializes in delivering intricate, fine-grained pixel-level details, the powerful discriminator of LAE-Net is designed by harnessing the visual embeddings derived from both the CLIP and DINO models separately to boost the visual discriminant capability and facilitate training a strong generator for conditioning image generation. Comprehensive experimental evaluations show that our LAE-Net not only delivers outstanding performance but also surpasses several cutting-edge models.
最长约 10秒,即可获得该文献文件

科研通智能强力驱动
Strongly Powered by AbleSci AI
科研通是完全免费的文献互助平台,具备全网最快的应助速度,最高的求助完成率。 对每一个文献求助,科研通都将尽心尽力,给求助人一个满意的交代。
实时播报
3秒前
勤恳代云完成签到,获得积分10
5秒前
tao完成签到 ,获得积分10
8秒前
16秒前
22秒前
万能图书馆应助秋收冬藏采纳,获得10
22秒前
yeye发布了新的文献求助10
23秒前
Bo完成签到 ,获得积分10
26秒前
科研通AI6.2应助张二十八采纳,获得10
26秒前
矢志不渝发布了新的文献求助10
27秒前
31秒前
秋收冬藏发布了新的文献求助10
36秒前
BUBU完成签到,获得积分10
37秒前
呆萌的鞯完成签到,获得积分10
39秒前
LL完成签到 ,获得积分10
41秒前
人间天星完成签到,获得积分10
44秒前
45秒前
shanfeng完成签到 ,获得积分10
45秒前
46秒前
Shanglinqin完成签到,获得积分10
47秒前
49秒前
张二十八发布了新的文献求助10
50秒前
52秒前
Shanglinqin发布了新的文献求助10
52秒前
53秒前
李健应助zayne采纳,获得10
1分钟前
sevakdumpling完成签到 ,获得积分10
1分钟前
1分钟前
1分钟前
wwwwww发布了新的文献求助10
1分钟前
zayne发布了新的文献求助10
1分钟前
stresm完成签到,获得积分10
1分钟前
PhysicsXX完成签到,获得积分10
1分钟前
张二十八发布了新的文献求助10
1分钟前
Lucas应助PS采纳,获得30
1分钟前
大模型应助929采纳,获得10
1分钟前
wwwwww完成签到,获得积分20
1分钟前
mmyhn发布了新的文献求助10
1分钟前
科研打工狗完成签到 ,获得积分10
1分钟前
v0id应助科研通管家采纳,获得10
1分钟前
高分求助中
(应助此贴封号)【重要!!请各用户(尤其是新用户)详细阅读】【科研通的精品贴汇总】 10000
APA handbook of comparative psychology: Basic concepts, methods, neural substrate, and behavior 1000
Health Psychology 1000
全员动态考核,锚定高质量发展:读懂同济大学教师人事改革新政的深层价值 900
Matrix Methods in Data Mining and Pattern Recognition Second Edition 510
The fast track to determining transfer functions of linear circuits: The student guide 500
Römisch-Germanische Forschungen 500
热门求助领域 (近24小时)
化学 材料科学 医学 生物 纳米技术 工程类 有机化学 化学工程 生物化学 计算机科学 内科学 物理 复合材料 催化作用 细胞生物学 无机化学 光电子学 物理化学 电极 基因
热门帖子
关注 科研通微信公众号,转发送积分 7597269
求助须知:如何正确求助?哪些是违规求助? 9173991
关于积分的说明 19640007
捐赠科研通 7174293
什么是DOI,文献DOI怎么找? 3268225
关于科研通互助平台的介绍 2432790
邀请新用户注册赠送积分活动 2261408