Automated sewer defect recognition technology based on machine vision is crucial for modern urban sewage systems. However, existing recognition models face two significant challenges: (1) inadequate performance in identifying rare but high-risk defects, and (2) complex interrelations among co-occurring defects that hinder the extraction of discriminative features. To tackle these issues, we propose a label-guided multimodal sewer defect recognition method incorporating a tail-aware knowledge distillation (KD) strategy. This strategy involves fine-tuning the teacher model on tail data to guide the student model’s learning process, enhancing its ability to identify rare defect features. Furthermore, our proposed Discrete Image-Text Interaction Module (DITIM) explores the semantic relationships between image patches and text through an interactive mechanism, which helps uncover co-occurrence relationships within multi-label information. This improves the model’s capability to capture complex correlations between different defects. Experimental validation on the Sewer-ML and QV-Pipe datasets demonstrates that our approach not only boosts overall recognition accuracy but also excels in detecting rare, high-risk defects, offering an effective technical solution for sewer defect identification and management.