自编码
透视图(图形)
突出
人工智能
计算机视觉
计算机科学
遥感
目标检测
对象(语法)
光学成像
模式识别(心理学)
地质学
光学
深度学习
物理
作者
Yuxiang Fu,Wei Fang,Victor S. Sheng
标识
DOI:10.1109/tgrs.2025.3566228
摘要
Recently masked autoencoder (MAE) has achieved great success in visual representation learning and delivered promising potential in many downstream vision tasks. However, due to the lack of saliency supervision signal in original MAE, almost no saliency information can be learned from masked image reconstruction process. Therefore, salient object detection (SOD) can hardly benefit from MAE pretraining. To address this issue, we integrate SOD model and saliency supervision in MAE and propose a simple and effective framework DHMMAE utilizing MAE with dynamic hybrid masking ratios to pretrain SOD model. Specifically, we treat MAE as an online data augmenter to generate endless pseudo images for SOD model to predict saliency maps, where saliency supervision is employed to ensure the model can learn robust saliency prior knowledge from the reconstructed images. Besides, we propose a simple and novel network EANet driven by DHMMAE pretraining for SOD in optical remote sensing images (ORSIs). Two key modules are designed for EANet to further enhance the performance: enhanced diverse feature aggregation module (EDFAM) and adjacent-context shuffle spatial attention module (ASSAM). EDFAM aggregates diverse features via three different types of convolution layer and enhances them by convolution block attention module. ASSAM captures spatial location information of salient objects by employing channel shuffle operation and weighted spatial attention mechanism on the fused adjacent context. Experiments on three ORSI-SOD datasets demonstrate that our proposed method outperforms the cutting-edge methods. Code is available at: https://github.com/Voruarn/EANet.
科研通智能强力驱动
Strongly Powered by AbleSci AI