计算机科学
人工智能
分割
计算机视觉
伪装
RGB颜色模型
对象(语法)
融合
频域
模态(人机交互)
特征(语言学)
解析
模式识别(心理学)
图像分割
钥匙(锁)
特征提取
领域(数学分析)
图像融合
目标检测
水准点(测量)
一般化
特征向量
像素
语音识别
投影(关系代数)
透视图(图形)
可视化
空间频率
传感器融合
尺度空间分割
作者
Peng Ren,Cheng Jiang,Fuming Sun,Tian Bai
标识
DOI:10.1109/tmm.2026.3668692
摘要
Recently, some studies have introduced depth cues to solve camouflaged object segmentation (COS) tasks and significantly improve segmentation performance. However, current methods still have two limitations: 1) they are confined to first-order or second-order interaction modeling; 2) they neglect the frequency-domain complementary characteristics between RGB and depth modalities. In this work, we propose a Frequency-domain Fusion and High-order Interaction framework, named F$^{2}$HI, to alleviate the above limitations. Specifically, F$^{2}$HI consists of two key components: Frequency Domain Interaction Fusion (FDIF) and the High-order Interaction Unit (HOIU). The FDIF decouples the frequency information of RGB and depth modalities into high-frequency and low-frequency components, utilizing the dominant frequency components of one modality to enhance the weaker frequency components of the other modality, thereby achieving complementary enhancement in the frequency domain. Moreover, it employs cross-attention to facilitate global multimodal interaction. The HOIU employs cascaded self-attention to parse cross-region features generated by multi-object camouflage scenarios, thereby achieving high-order feature interaction. Compared with 27 state-of-the-art COS methods, F$^{2}$HI achieves competitive performance on four mainstream COS benchmarks.
科研通智能强力驱动
Strongly Powered by AbleSci AI