情态动词
计算机科学
结核(地质)
人工智能
分割
计算机视觉
图像分割
甲状腺
超声波
放射科
模式识别(心理学)
医学
地质学
内科学
古生物学
化学
高分子化学
作者
Xia Sun,Boxiong Wei,Yalong Jiang,Linjie Mao,Qi Zhao
标识
DOI:10.1109/lsp.2025.3556789
摘要
Thyroid nodule segmentation in ultrasound images is crucial for accurate diagnosis and treatment planning. However, existing methods struggle with segmentation accuracy, interpretability, and generalization. This letter proposes CLIP-TNseg, a novel framework that integrates a multimodal large model with a neural network architecture to address these challenges. We innovatively divide visual features into coarse-grained and fine-grained components, leveraging textual integration with coarse-grained features for enhanced semantic understanding. Specifically, the Coarse-grained Branch extracts high-level semantic features from a frozen CLIP model, while the Fine-grained Branch refines spatial details using U-Net-style residual blocks. Extensive experiments on the newly collected PKTN dataset and other public datasets demonstrate the competitive performance of CLIP-TNseg. Additional ablation experiments confirm the critical contribution of textual inputs, particularly highlighting the effectiveness of our carefully designed textual prompts compared to fixed or absent textual information.
科研通智能强力驱动
Strongly Powered by AbleSci AI