计算机科学
概率逻辑
多样性(控制论)
透视图(图形)
人工智能
生成语法
扩散
控制(管理)
数据科学
机器学习
降噪
理论计算机科学
生成模型
芯(光纤)
作者
Pu Cao,Feng Zhou,Qing Song,Lu Yang
标识
DOI:10.1109/tpami.2025.3646548
摘要
In the rapidly advancing realm of visual generation, diffusion models have revolutionized the landscape, marking a significant shift in capabilities with their impressive text-guided generative functions. However, relying solely on text for conditioning these models does not fully cater to the varied and complex requirements of different applications and scenarios. Acknowledging this shortfall, a variety of studies aim to control pre-trained text-to-image (T2I) models to support novel conditions. In this survey, we undertake a thorough review of the literature on controllable generation with T2I diffusion models, covering both the theoretical foundations and practical advancements in this domain. Our review begins with a brief introduction to the basics of denoising diffusion probabilistic models (DDPMs) and widely used T2I diffusion models. Additionally, we provide a detailed overview of research in this area, categorizing it from the condition perspective into three directions: generation with specific conditions, generation with multiple conditions, and universal controllable generation. For each category, we analyze the underlying control mechanisms and review representative methods based on their core techniques.
科研通智能强力驱动
Strongly Powered by AbleSci AI