对抗制
异常检测
计算机科学
人工智能
异常(物理)
计算机视觉
模式识别(心理学)
物理
凝聚态物理
作者
Yuchen Qiang,Jiuxin Cao,Siyang Zhou,Junyang Yang,Lijia Yu,Bo Liu
标识
DOI:10.1109/tii.2025.3598429
摘要
Industrial anomaly detection aims to identify and localize defective regions in images. Among various architectures, reconstruction-based methods have demonstrated exceptional performance. These methods reconstruct anomalous samples into normal ones and identify anomalies through a comparison between them. However, reconstruction process within these methods often focuses on PL similarity, overlooking the high-frequency consistency between the input and output, which constrains the model’s accuracy. This article proposes tGARD, a novel text-guided adversarial reconstruction method for anomaly detection. Specifically, we introduce feature aggregation module, using nonlocal block and dilated convolution to handle complex anomaly patterns. Subsequently, text-guided reconstruction module is meticulously designed to harness CLIP’s multimodal alignment capabilities, allowing for a semantically controllable reconstruction process. This control is achieved by incorporating dynamic text embeddings derived from the CLIP encoder within discriminator. Meanwhile, during reconstruction process, high-frequency details are preserved through a convolutional adversarial discriminator. Finally, category-aware loss weighting strategy is conceived to balance similarity and adversarial loss. Experiments demonstrate that our model achieves significant improvements in anomaly localization, surpassing all reconstruction-based models on MVTec-AD. It also establishes a new state-of-the-art on VisA dataset, outperforming all existing architectures.
科研通智能强力驱动
Strongly Powered by AbleSci AI