人工智能
计算机科学
图像分割
模式识别(心理学)
计算机视觉
分割
医学影像学
尺度空间分割
特征提取
图像处理
图像(数学)
特征(语言学)
迭代重建
基于分割的对象分类
图像纹理
作者
Wenhui Huang,Zhen Pan,Xiaoyan Wang,Yedi Zhang,Jingzhen He,J. Bernard L. Gee,Yuanjie Zheng
标识
DOI:10.1109/tmi.2026.3674592
摘要
Fully supervised polyp segmentation relies on costly pixel-level annotations. Although semi- and weakly supervised methods reduce annotation requirements, they still depend on partial mask supervision. Text-supervised segmentation is a promising alternative; however, for polyps, the key challenge is to ground instance-specific phrases to the correct lesion region under cluttered backgrounds and large appearance variations. Existing approaches often rely on coarse text-image alignment, limiting precise region-level semantic correspondence. In this paper, we propose Text-Image Co-Alignment (TICoA), a text-supervised framework for polyp segmentation. TICoA leverages large language models (LLMs)-generated structured clinical descriptions as weak supervision and formulates segmentation as a fine-grained phrase-region co-alignment problem. Through contrastive learning, TICoA explicitly associates query phrases with corresponding image regions to achieve robust semantic grounding under weak supervision. Architecturally, we adopt a State-Space Model (Mamba) to efficiently model long-range dependencies with linear computational complexity. To support effective cross-modal interaction, we further design a dedicated Mamba Fusion module with a Bi-Dimension Fusion (BiDF) strategy, which progressively propagates information along spatial and channel dimensions. Experiments on polyp datasets, with additional validation on skin lesion segmentation, demonstrate that TICoA is competitive with state-of-the-art weakly supervised methods. Our code and data are available at https://github.com/silentyuchen/TICoA.
科研通智能强力驱动
Strongly Powered by AbleSci AI