计算机科学
判别式
异常检测
人工智能
嵌入
模式识别(心理学)
利用
子空间拓扑
语义鸿沟
语义映射
领域(数学分析)
模态(人机交互)
支持向量机
点(几何)
Boosting(机器学习)
图形
边界(拓扑)
自然语言
交错
上下文图像分类
机器学习
歧管(流体力学)
目标检测
交叉口(航空)
数据挖掘
边界判定
计算机视觉
机器人
异常(物理)
超平面
分割
作者
Yunfeng Ma,Ming Liu,Shuai Jiang,Jingyu Zhou,Yuan Bian,Xueping Wang,Yaonan Wang
标识
DOI:10.1109/tpami.2026.3658856
摘要
Multimodal anomaly detection (MAD) aims to exploit both texture and spatial attributes to identify deviations from normal patterns in complex scenarios. However, zero-shot (ZS) settings arising from privacy concerns or confidentiality constraints present significant challenges to existing MAD methods. To address this issue, we introduce ZUMA, a training-free, Zero-shot Unified Multimodal Anomaly detection framework that unleashes CLIP's cross-modal potential to perform ZS MAD. To mitigate the domain gap between CLIP's pretraining space and point clouds, we propose cross-domain calibration (CDC), which efficiently bridges the manifold misalignment through source-domain semantic transfer and establishes a hybrid semantic space, enabling a joint embedding of 2D and 3D representations. Subsequently, ZUMA performs dynamic semantic interaction (DSI) to enable structural decoupling of anomaly regions in the high-dimensional embedding space constructed by CDC, where natural languages serve as semantic anchors to help DSI establish discriminative hyperplanes within hybrid modality representations. Within this framework, ZUMA enables plug-and-play detection of 2D, 3D or multimodal anomalies, without training or fine-tuning even for cross-dataset or incomplete-modality scenarios. Additionally, to further investigate the potential of the training-free ZUMA within the training-based paradigm, we develop ZUMA-FT, a fine-tuned variant that achieves notable improvements with minimal parameter trade-off. Extensive experiments are conducted on two MAD benchmarks, MVTec 3D-AD and Eyecandies. Notably, the training-free ZUMA achieves state-of-the-art (SOTA) performance on both datasets, outperforming existing ZS MAD methods, including training-based approaches. Moreover, ZUMA-FT further extends the performance boundary of ZUMA with only 6.75 M learnable parameters.
科研通智能强力驱动
Strongly Powered by AbleSci AI