计算机科学
推论
嵌入
软件部署
人工智能
机器学习
感知
网格
质量(理念)
传感器融合
无线
面子(社会学概念)
语义学(计算机科学)
分布式计算
网络数据包
无线传感器网络
建筑
注释
接头(建筑物)
利用
遮罩(插图)
作者
Nicanor Mayumu,Xiaoheng Deng,Bigomokero Antoine Bagula,Saif ur Rehman Khan,Patrick Mukala
标识
DOI:10.1109/jiot.2026.3660030
摘要
Autonomous vehicles face perception challenges due to occlusions, limited sensor ranges, and adverse weather. Vehicle-to-Everything (V2X) cooperative perception mitigates these limitations by enabling vehicles to share sensor data. However, existing methods rely on supervised learning, requiring costly manual 3D annotations, exhibiting limited generalization, and employing static fusion strategies that fail under communication disruptions. We propose V2X-JEPA, a self-supervised framework that extends the joint-embedding predictive architecture (JEPA) to V2X cooperative perception for both Vehicle-to-Vehicle (V2V) and Vehicle-to-Infrastructure scenarios. V2X-JEPA learns semantic representations via latent embedding prediction, eliminating the need for manual annotations during pretraining. We introduce three cooperative masking strategies and grid spatial attention fusion, which adapt to communication quality and agent reliability. Extensive evaluation on OPV2V and DAIR-V2X benchmarks shows that V2X-JEPA reduces annotation requirements by 85% while achieving competitive performance. On OPV2V (LiDAR, V2V), V2X-JEPA achieves 94.5% mAP@0.5 and 89.2% mAP@0.7. On DAIR-V2X (camera, V2I), it achieves 24.8% AP@0.5 and 11.2% AP@0.7. It outperforms V2X-ViT by 4.7%, CoCa3D by 2.2%, and CooPre by 2.8%. V2X-JEPA is efficient, with 98M parameters and 78-ms inference time, and demonstrates high robustness, degrading only 5.1% under 20% packet loss, supporting practical deployment in bandwidth-constrained environments.
科研通智能强力驱动
Strongly Powered by AbleSci AI