计算机科学
分割
人工智能
棱锥(几何)
编码器
图像分割
特征(语言学)
地理空间分析
计算机视觉
背景(考古学)
遥感
特征提取
遥感应用
空间分析
模式识别(心理学)
尺度空间分割
目标检测
骨干网
空间语境意识
基于分割的对象分类
网络体系结构
上下文模型
高光谱成像
变压器
图像处理
作者
Cheng Zhou,Bofeng Cai,Chengyi Xiong,Renfeng Liu
标识
DOI:10.1109/tgrs.2026.3660697
摘要
Accurate semantic segmentation of remote sensing imagery is fundamental to geospatial analysis, yet the task remains daunting due to intrinsic complexities such as blurred object boundaries, severe occlusion, and high inter-class similarity. To overcome these persistent bottlenecks, this study presents an engineering-integrated framework that strategically synthesizes advanced vision components to balance segmentation accuracy with model efficiency. Rather than relying on a single core innovation, we rationally combine a Pyramid Vision Transformer (PVT) backbone with specialized edge and alignment modules to reconcile the trade-off between global context and local detail. Specifically, the architecture features an edge-guided encoder that augments the PVT backbone with a learnable Edge-guided Feature Module (EFM), explicitly targeting the recovery of boundary information often lost during downsampling. Complementing this, the decoder employs a Feature-aligned Pyramid Network (FaPN) reinforced by Adaptive Spatial Feature Fusion (ASFF). This combination ensures precise cross-scale alignment and dynamic feature selection, effectively suppressing background noise while preserving fine-grained spatial structures. Experimental results on two large-scale high-resolution remote sensing datasets demonstrate that the proposed network offers superior segmentation performance while maintaining an efficient balance between computational complexity and model parameters.
科研通智能强力驱动
Strongly Powered by AbleSci AI