计算机科学
编码器
分割
人工智能
语义压缩
自然语言处理
计算机视觉
语义计算
语义技术
语义网
操作系统
作者
Qingsong Tang,Shitong Min,Xiaomeng Shi,Qi Zhang,Liu Yang
标识
DOI:10.1088/1361-6501/ad9106
摘要
Abstract Real-time semantic segmentation is widely used in various domains such as autonomous driving and medical imaging. Most real-time semantic segmentation networks adopt an encoder–decoder structure. During encoding, feature maps undergo multiple downsampling stages to enhance computational efficiency and enlarge the receptive field. However, this process may lead to the loss of detailed information, which may not be fully recovered during upsampling. Moreover, upsampling semantic features may cause blurring and deviations due to the lack of spatial detail information. To mitigate these issues, we use spatially guided cross-resolution self-attention to improve the upsampling of the semantic features by supplementing detailed information from the spatial branch. Furthermore, we incorporate an inductive bias into the cross-resolution attention mechanism to enhance its ability to learn generalized features. Additionally, we design a semantic feature extraction block, and a spatial feature extraction branch to construct a lightweight backbone. The results on Cityscapes and CamVid show that the proposed model achieves a good balance between accuracy and parameter size. Specifically, it obtains 74.2% and 71.5% mIoU on the two test datasets with 1.31 M parameters, respectively. Code is available on https://github.com/clearwater753/DECENet.
科研通智能强力驱动
Strongly Powered by AbleSci AI