计算机视觉
人工智能
计算机科学
图像分割
分割
变压器
工程类
电气工程
电压
作者
Zhaoxuan Gong,Meas Chanmean,Wenwen Gu
标识
DOI:10.1109/aipmv62663.2024.10691911
摘要
In this paper, we present an advanced approach for image segmentation that enhances Vision Transformers (ViTs) by integrating multi-scale hybrid attention mechanisms. While ViTs use self-attention to capture global dependencies in images, they often miss local details crucial for segmentation tasks. Our method addresses this by incorporating convolutional layers with different kernel sizes to extract local features at multiple scales. These features are then merged with the global self-attention outputs, creating a hybrid attention map that captures both fine details and overall structures. Testing on standard benchmarks shows our approach significantly improves segmentation accuracy and robustness compared to traditional ViTs, demonstrating its potential for handling complex visual scenes.
科研通智能强力驱动
Strongly Powered by AbleSci AI