计算机科学
计算机视觉
人工智能
图像增强
图像(数学)
图像处理
医学影像学
图像分割
对比度增强
图像复原
迭代重建
图像融合
目标检测
作者
Hong Wang,Zhijian Wu,Haodu Fang,Dong Wei,Jinghan Sun,Yefeng Zheng,Jianhua Ma
标识
DOI:10.1109/tmi.2026.3694909
摘要
Low light conditions in endoscopic imaging would lead to poor visibility, reduced contrast, and increased noise, which may hinder accurate diagnosis and surgical guidance. Against this low-light endoscopic image enhancement (LLEIE) task, inspired by the remarkable performance of pretrained CLIP in downstream vision tasks, in this paper, we carefully investigate the pretrained priors of CLIP and embed them into a text-modulated semantic-aware discriminator (TMSD). Through the adversarial learning mechanism, the discriminator can be easily integrated into different low-light enhancement baselines for helping them accomplish better visual restoration effects without incurring any extra inference cost. Specifically, to make the foundation model CLIP suitable for the LLEIE task, we initially propose a prompt learning procedure to obtain the text embedding and image semantics corresponding to the normal-light endoscopic imaging scenario. Building upon the acquired text prior and image semantic priors, we devise a text modulator to synergize these two priors, yielding a richer semantic representation. Leveraging the convolutional modulation and cross-attention mechanisms, we blend this semantic guidance information into the discriminator, thereby fostering the fine-grained distribution learning of normal-light endoscopic images in visual semantics and guiding different enhancement baselines achieving higher visual quality. Based on five public benchmark datasets, including three synthetic datasets, one real clinical dataset, and one clinical downstream segmentation dataset, we comprehensively evaluate the effectiveness of our proposed TMSD. Extensive experiments substantiate that the integration of the proposed TMSD enables seven representative baselines to obtain better perceptual quality, especially in the cross-domain clinical generalization scenario. Besides, the downstream segmentation accuracy can be evidently improved, showing the favorable application potential of the proposed TMSD. Moreover, to comprehensively evaluate the generality of our TMSD framework, we successfully apply it to a new and classic metal artifact reduction task. It is worth mentioning that our TMSD does not incur any extra computational cost during inference.
科研通智能强力驱动
Strongly Powered by AbleSci AI