计算机科学
人工智能
计算机视觉
目标检测
模式识别(心理学)
对象类检测
突出
对象(语法)
图像分割
图像处理
特征提取
可视化
Viola–Jones对象检测框架
边缘检测
特征(语言学)
变更检测
数据压缩
上下文图像分类
图像(数学)
直方图
噪音(视频)
视觉对象识别的认知神经科学
算法设计
作者
Yunping Zheng,Haobo Li,Shiqiang Shu,Ke Xu,Mudar Sarem
标识
DOI:10.1109/tcsvt.2026.3673067
摘要
The advent of scribble-supervised learning has opened new frontiers in weakly supervised RGB-D salient object detection (SOD), representing a novel paradigm that significantly reduces annotation costs while maintaining detection efficacy. The SOD methods predominantly based on convolutional neural networks (CNNs) exhibit notable limitations in capturing global contextual dependencies and effectively utilizing hierarchical multiscale representations. Furthermore, conventional edge extraction approaches based on Canny edge detection demonstrate inherent limitations in preserving structural coherence across diverse object boundaries. To address these critical challenges, we propose an NESS-Net model, an innovative weakly supervised architecture that synergistically integrates our previously developed NAMLab framework with Swin-Transformer-based feature learning. The core contributions of our work are threefold: Firstly, building upon our prior work in perceptual-aware image segmentation, we employ the NAMLab algorithm—a hierarchical segmentation framework inspired by Nonsymmetry and Anti-packing pattern representation Model in the Lab color space (NAMLab)—to generate edge maps that better align with human visual perception. Secondly, we devise a U-shaped Edge Construction Module (ECM) that systematically refines edge features through progressive refinement of the NAMLab-generated edge priors. Thirdly, leveraging Swin-Transformer’s hierarchical attention mechanism, we put forward a lightweight Cross-Attention Fusion Module (CAFM) that establishes long-range dependencies across modalities while maintaining computational efficiency. This architectural innovation naturally extends to our Cross-Modal multi-scale Attention-weighting Module (CMAM), which explicitly models inter-modal relationships through transformer-based attention weighting. The extensive experimental results on nine benchmark RGB-D SOD datasets demonstrate that our proposed NESS-Net model outperforms all the current state-of-the-art scribble-supervised models and even achieves performance competitive to leading fully supervised models. The source code and pre-trained models will be made publicly available at https://github.com/EricLie-c/NESS-Net.
科研通智能强力驱动
Strongly Powered by AbleSci AI