姿势
人工智能
变压器
融合
计算机科学
计算机视觉
模式识别(心理学)
工程类
电气工程
电压
语言学
哲学
作者
Han Xu Sun,Zhenning Zhou,Yizhao Wang,Zhuangzhuang Zhang,Qixin Cao
标识
DOI:10.1109/lra.2024.3381016
摘要
The 6D pose estimation for metal parts is essential in industrial robotic applications. The color homogeneity, texture-less and light-reflecting properties of metal parts raise great challenges. Current 6D pose estimation methods have gained extensive concern using CNNs. However, these CNN-based methods lack Transformer's ability to focus on extracting low-frequency features and long-range context information. In the study, we explore taking full advantage of CNN and Transformer from a frequency-domain perspective to enhance the performance of metal parts' 6D pose estimation. Specifically, we propose a frequency-guided CNN-Transformer fusion 6D pose estimation network (FGCT6D). First, we construct a novel pixel attention residual module to improve the high-frequency attention of CNN. Then, we design a dual-branch CNN-Transformer encoder: the Swin-Transformer extracts global information and low-frequency features, and the CNN captures local information and high-frequency features. Second, the frequency-guided feature fusion module is proposed to fuse the extracted multi-spectral features. Third, to maximize the utilization of the rich frequency-domain feature representation, we propose a feature fusion decoder with Conv-MSA modules. Additionally, we leverage optimal transport theory, treating dense correspondences as spatial probability distributions, and design the optimal transport loss function. Experiments show that our method can extract rich frequency-domain features, and achieve competitive performance on the MP6D and LINEMOD datasets.
科研通智能强力驱动
Strongly Powered by AbleSci AI