Domain-RAG: Retrieval-Guided Compositional Image Generation for Cross-Domain Few-Shot Object Detection

人工智能 计算机科学 目标检测 对象(语法) 计算机视觉 图像(数学) 领域(数学分析) 生成语法 生成模型 特征(语言学) 背景图像 模式识别(心理学) 图像分割 特征提取 图像处理 语义学(计算机科学) 视觉对象识别的认知神经科学 图像处理 混合模型 组分(热力学)
作者
Li Yu,Xingyu Qiu,Yuqian Fu,Jie Chen,Tianwen Qian,Zheng Xu,Danda Pani Paudel,Yanwei Fu,Xuanjing Huang,Luc Van Gool,Yu–Gang Jiang
标识
DOI:10.48550/arxiv.2506.05872
摘要

Cross-Domain Few-Shot Object Detection (CD-FSOD) aims to detect novel objects with only a handful of labeled samples from previously unseen domains. While data augmentation and generative methods have shown promise in few-shot learning, their effectiveness for CD-FSOD remains unclear due to the need for both visual realism and domain alignment. Existing strategies, such as copy-paste augmentation and text-to-image generation, often fail to preserve the correct object category or produce backgrounds coherent with the target domain, making them non-trivial to apply directly to CD-FSOD. To address these challenges, we propose Domain-RAG, a training-free, retrieval-guided compositional image generation framework tailored for CD-FSOD. Domain-RAG consists of three stages: domain-aware background retrieval, domain-guided background generation, and foreground-background composition. Specifically, the input image is first decomposed into foreground and background regions. We then retrieve semantically and stylistically similar images to guide a generative model in synthesizing a new background, conditioned on both the original and retrieved contexts. Finally, the preserved foreground is composed with the newly generated domain-aligned background to form the generated image. Without requiring any additional supervision or training, Domain-RAG produces high-quality, domain-consistent samples across diverse tasks, including CD-FSOD, remote sensing FSOD, and camouflaged FSOD. Extensive experiments show consistent improvements over strong baselines and establish new state-of-the-art results. Codes will be released upon acceptance.
最长约 10秒,即可获得该文献文件

科研通智能强力驱动
Strongly Powered by AbleSci AI
科研通是完全免费的文献互助平台,具备全网最快的应助速度,最高的求助完成率。 对每一个文献求助,科研通都将尽心尽力,给求助人一个满意的交代。
实时播报
Nole应助宇宙第一甜妹采纳,获得10
2秒前
youwenjing11发布了新的文献求助10
3秒前
aajhajkahna应助Debra采纳,获得10
3秒前
芝诺的乌龟完成签到 ,获得积分0
4秒前
函_完成签到 ,获得积分10
5秒前
明理耷完成签到 ,获得积分10
8秒前
亿眼万年完成签到,获得积分10
8秒前
充电宝应助dingdingdang9770采纳,获得10
9秒前
9秒前
9秒前
传奇3应助科研通管家采纳,获得10
9秒前
田様应助科研通管家采纳,获得10
9秒前
所所应助科研通管家采纳,获得10
10秒前
CodeCraft应助科研通管家采纳,获得10
10秒前
所所应助科研通管家采纳,获得10
10秒前
aajhajkahna应助科研通管家采纳,获得10
10秒前
小蘑菇应助科研通管家采纳,获得10
10秒前
maxlovol应助科研通管家采纳,获得10
11秒前
在水一方应助科研通管家采纳,获得10
11秒前
GYK应助科研通管家采纳,获得10
11秒前
爆米花应助科研通管家采纳,获得10
11秒前
Orange应助科研通管家采纳,获得10
11秒前
11秒前
Lucas应助科研通管家采纳,获得10
11秒前
东方元语应助科研通管家采纳,获得20
12秒前
12秒前
12秒前
研友_VZG7GZ应助科研通管家采纳,获得30
12秒前
12秒前
欣慰羊完成签到,获得积分10
12秒前
今后应助科研通管家采纳,获得10
12秒前
loading完成签到,获得积分10
12秒前
12秒前
桐桐应助科研通管家采纳,获得10
13秒前
伶俐的老头完成签到 ,获得积分10
13秒前
13秒前
爆米花应助科研通管家采纳,获得10
13秒前
爱你沛沛完成签到 ,获得积分10
13秒前
lo完成签到,获得积分10
13秒前
orixero应助科研通管家采纳,获得10
13秒前
高分求助中
(应助此贴封号)【重要!!请各用户(尤其是新用户)详细阅读】【科研通的精品贴汇总】 10000
China Pluperfect I: Epistemology of Past and Outside in Chinese Art 520
Management and the Arts 510
Matrix Methods in Data Mining and Pattern Recognition Second Edition 510
基于锂离子电池正极材料回收的绿色溶剂开发及工程化应用研究 500
Auslegungsgeschichte 500
Transdermal drug delivery systems market size report 500
热门求助领域 (近24小时)
化学 材料科学 医学 生物 纳米技术 工程类 有机化学 化学工程 生物化学 计算机科学 内科学 物理 复合材料 催化作用 细胞生物学 无机化学 光电子学 物理化学 电极 基因
热门帖子
关注 科研通微信公众号,转发送积分 7642408
求助须知:如何正确求助?哪些是违规求助? 9215419
关于积分的说明 19768587
捐赠科研通 7207644
什么是DOI,文献DOI怎么找? 3276367
关于科研通互助平台的介绍 2438115
邀请新用户注册赠送积分活动 2274102