亲爱的研友该休息了!由于当前在线用户较少,发布求助请尽量完整地填写文献信息,科研通机器人24小时在线,伴您度过漫漫科研夜!身体可是革命的本钱,早点休息,好梦!

Learned representation-guided diffusion models for large-image generation

计算机科学 人工智能 稳健性(进化) 忠诚 概化理论 编码 模式识别(心理学) 分类器(UML) 图像(数学) 数学 电信 生物化学 化学 统计 基因
作者
Alexandros Graikos,Srikar Yellapragada,Minh-Quan Le,Saarthak Kapse,Prateek Prasanna,Joel Saltz,Dimitris Samaras
出处
期刊:Cornell University - arXiv [Cornell University]
标识
DOI:10.48550/arxiv.2312.07330
摘要

To synthesize high-fidelity samples, diffusion models typically require auxiliary data to guide the generation process. However, it is impractical to procure the painstaking patch-level annotation effort required in specialized domains like histopathology and satellite imagery; it is often performed by domain experts and involves hundreds of millions of patches. Modern-day self-supervised learning (SSL) representations encode rich semantic and visual information. In this paper, we posit that such representations are expressive enough to act as proxies to fine-grained human labels. We introduce a novel approach that trains diffusion models conditioned on embeddings from SSL. Our diffusion models successfully project these features back to high-quality histopathology and remote sensing images. In addition, we construct larger images by assembling spatially consistent patches inferred from SSL embeddings, preserving long-range dependencies. Augmenting real data by generating variations of real images improves downstream classifier accuracy for patch-level and larger, image-scale classification tasks. Our models are effective even on datasets not encountered during training, demonstrating their robustness and generalizability. Generating images from learned embeddings is agnostic to the source of the embeddings. The SSL embeddings used to generate a large image can either be extracted from a reference image, or sampled from an auxiliary model conditioned on any related modality (e.g. class labels, text, genomic data). As proof of concept, we introduce the text-to-large image synthesis paradigm where we successfully synthesize large pathology and satellite images out of text descriptions.
最长约 10秒,即可获得该文献文件

科研通智能强力驱动
Strongly Powered by AbleSci AI
科研通是完全免费的文献互助平台,具备全网最快的应助速度,最高的求助完成率。 对每一个文献求助,科研通都将尽心尽力,给求助人一个满意的交代。
实时播报
大模型应助汪小楠吖采纳,获得10
6秒前
会撒娇的思萱完成签到,获得积分10
15秒前
16秒前
汪小楠吖发布了新的文献求助10
19秒前
科研通AI6.2应助XIAOEN采纳,获得10
19秒前
未闻星名完成签到,获得积分10
32秒前
33秒前
Ricky_Ao发布了新的文献求助10
37秒前
失眠紫完成签到,获得积分10
1分钟前
Lucas应助刘智舰采纳,获得10
1分钟前
1分钟前
刘智舰发布了新的文献求助10
1分钟前
酷波er应助汪小楠吖采纳,获得10
1分钟前
1分钟前
汪小楠吖完成签到,获得积分10
1分钟前
汪小楠吖发布了新的文献求助10
1分钟前
光亮如容完成签到,获得积分10
1分钟前
1分钟前
psyche完成签到,获得积分10
1分钟前
华仔应助微光采纳,获得10
2分钟前
科研通AI2S应助刘智舰采纳,获得10
2分钟前
2分钟前
李紫月发布了新的文献求助10
2分钟前
2分钟前
忧虑的如雪完成签到,获得积分10
2分钟前
追寻的不言完成签到,获得积分20
2分钟前
du完成签到,获得积分10
2分钟前
3分钟前
3分钟前
3分钟前
3分钟前
坚定的白云完成签到,获得积分10
3分钟前
3分钟前
4分钟前
完美世界应助Ricky_Ao采纳,获得10
4分钟前
连安阳完成签到,获得积分10
4分钟前
4分钟前
生活完成签到 ,获得积分10
4分钟前
4分钟前
Ricky_Ao发布了新的文献求助10
4分钟前
高分求助中
(应助此贴封号)【重要!!请各用户(尤其是新用户)详细阅读】【科研通的精品贴汇总】 10000
China Pluperfect I: Epistemology of Past and Outside in Chinese Art 520
Matrix Methods in Data Mining and Pattern Recognition Second Edition 510
基于锂离子电池正极材料回收的绿色溶剂开发及工程化应用研究 500
Auslegungsgeschichte 500
Cosmos as Art Object: Studies in Plato's Timaeus and Other Dialogues 500
Middle East Patterns 444
热门求助领域 (近24小时)
化学 材料科学 医学 生物 纳米技术 工程类 有机化学 化学工程 生物化学 计算机科学 内科学 物理 复合材料 催化作用 细胞生物学 无机化学 光电子学 物理化学 电极 基因
热门帖子
关注 科研通微信公众号,转发送积分 7640070
求助须知:如何正确求助?哪些是违规求助? 9213146
关于积分的说明 19763406
捐赠科研通 7206274
什么是DOI,文献DOI怎么找? 3276074
关于科研通互助平台的介绍 2437673
邀请新用户注册赠送积分活动 2273458