计算机科学
人工智能
分割
图像分割
棱锥(几何)
模式识别(心理学)
计算机视觉
背景(考古学)
卷积(计算机科学)
特征(语言学)
像素
尺度空间分割
特征提取
语义学(计算机科学)
目标检测
图像处理
匹配(统计)
编码(集合论)
可视化
基于分割的对象分类
可分离空间
直方图
适应(眼睛)
噪音(视频)
特征选择
作者
Xiaoqing Zhao,Chaojun Zhang,Yuan Gao,Jing Yang,Laurence T. Yang,Jieming Yang
标识
DOI:10.1109/tcsvt.2026.3651347
摘要
Recently, most image segmentation methods exhibit an extreme trade-off between performance and efficiency, resulting in approaches with high performance typically having low computational efficiency, while efficient methods compromise on segmentation accuracy. To address this dual challenge, this study introduces a simple yet efficient segmentation framework based on Multi-scale Prototype matching and visual Sparse Attention mechanisms (MPSA), which is a transformer-based architecture designed to optimize the balance between performance and efficiency. The proposed MPSA integrates a novel lightweight cross-attention mechanism and prototype selection and filtering strategy to accurately correlate category queries with corresponding visual objects with a multi-scale Feature Pyramid Network (FPN). Within the pixel decoder, our Axial Convolution Enhanced (ACE) module mitigates lost global context by combining depth-wise separable convolutions with deformable convolutions, thereby recovering global semantics while preserving fine-grained spatial details. Through this innovative design, MPSA demonstrates outstanding performance in both semantic and panoptic segmentation tasks across multiple datasets. Remarkably, MPSA achieves surprising 83.9% mIoU with only 114M parameters on the Cityscapes dataset while compared to some state-of-the-art architectures, highlighting its ability to deliver exceptional results with significantly reduced resource consumption. Our code is released at https://github.com/zxqing01/MPSA.
科研通智能强力驱动
Strongly Powered by AbleSci AI