计算机科学
块(置换群论)
人工智能
光学(聚焦)
目标检测
突出
特征(语言学)
比例(比率)
对象(语法)
计算机视觉
解码方法
集合(抽象数据类型)
模式识别(心理学)
任务(项目管理)
语义学(计算机科学)
特征提取
骨干网
桥(图论)
任务分析
融合
图像融合
上下文模型
隐马尔可夫模型
传感器融合
图像分割
分割
视觉对象识别的认知神经科学
作者
Xiandong Wang,Tianqi Guo,Fengqin Yao,Qi Guo,Shengke Wang,Qing Cai,Xinzhe Li,Guoqiang Zhong
标识
DOI:10.1109/tcsvt.2026.3670168
摘要
Referring camouflaged object detection (Ref-COD) is an emerging and challenging task that aims to localize camouflaged objects in complex scenes based on a small set of referring images with salient objects. However, existing methods primarily focus on semantic alignment between the referring and camouflaged objects while overlooking scale discrepancies, leading to under-response when small references guide large objects and over-response when large references guide small ones. To overcome this limitation, we propose a novel Multi-scale Interaction Network (MINet), explicitly designed to handle feature interactions across different scales in Ref-COD. MINet begins with a Dual-Source Fusion Block (DSFB) for semantic fusion between the referring and camouflaged features. Then, the Intra-scale Interaction Block (IIB) enhances local saliency within each scale by modeling contextual importance. Next, the Cross-scale Interaction Block (CIB) performs offset-guided alignment to bridge spatial gaps in multiscale feature fusion. Finally, the Cross-scale Aggregation Decoder (CAD) integrates multiscale features, effectively decoding the aggregated information to produce accurate predictions. Extensive experiments on Ref-COD datasets demonstrate that our method achieves state-of-the-art performance, highlighting the importance of scale interaction in Ref-COD.
科研通智能强力驱动
Strongly Powered by AbleSci AI