亲爱的研友该休息了!由于当前在线用户较少,发布求助请尽量完整地填写文献信息,科研通机器人24小时在线,伴您度过漫漫科研夜!身体可是革命的本钱,早点休息,好梦!

VILA: Learning Image Aesthetics from User Comments with Vision-Language Pretraining

计算机科学 隐藏字幕 人工智能 图像(数学) 自然语言处理 自然语言 语义学(计算机科学) 情报检索 程序设计语言
作者
Junjie Ke,Keren Ye,Jiahui Yu,Yonghui Wu,Peyman Milanfar,Feng Yang
出处
期刊: 卷期号:: 10041-10051 被引量:46
标识
DOI:10.1109/cvpr52729.2023.00968
摘要

Assessing the aesthetics of an image is challenging, as it is influenced by multiple factors including composition, color, style, and high-level semantics. Existing image aesthetic assessment (IAA) methods primarily rely on human-labeled rating scores, which oversimplify the visual aesthetic information that humans perceive. Conversely, user comments offer more comprehensive information and are a more natural way to express human opinions and preferences regarding image aesthetics. In light of this, we propose learning image aesthetics from user comments, and exploring vision-language pretraining methods to learn multimodal aesthetic representations. Specifically, we pretrain an image-text encoder-decoder model with image-comment pairs, using contrastive and generative objectives to learn rich and generic aesthetic semantics without human labels. To efficiently adapt the pretrained model for downstream IAA tasks, we further propose a lightweight rank-based adapter that employs text as an anchor to learn the aesthetic ranking concept. Our results show that our pretrained aesthetic vision-language model outperforms prior works on image aesthetic captioning over the AVA-Captions dataset, and it has powerful zero-shot capability for aesthetic tasks such as zero-shot style classification and zero-shot IAA, surpassing many supervised baselines. With only minimal finetuning parameters using the proposed adapter module, our model achieves state-of-the-art IAA performance over the AVA dataset. 1 1 Our model is available at https://github.com/google-research/google-research/tree/master/VILA
最长约 10秒,即可获得该文献文件

科研通智能强力驱动
Strongly Powered by AbleSci AI
科研通是完全免费的文献互助平台,具备全网最快的应助速度,最高的求助完成率。 对每一个文献求助,科研通都将尽心尽力,给求助人一个满意的交代。
实时播报
10秒前
seiya发布了新的文献求助10
14秒前
nano_grid完成签到,获得积分10
22秒前
49秒前
54秒前
jimmy发布了新的文献求助10
55秒前
研友_LMo56Z完成签到,获得积分10
1分钟前
1分钟前
nian发布了新的文献求助10
1分钟前
hmhu完成签到,获得积分10
1分钟前
hmhu发布了新的文献求助60
1分钟前
枫叶53完成签到 ,获得积分10
1分钟前
Criminology34举报xuxu213求助涉嫌违规
2分钟前
2分钟前
2分钟前
2分钟前
qiuqiu发布了新的文献求助10
2分钟前
cc发布了新的文献求助10
2分钟前
lixinglei应助Anna采纳,获得20
2分钟前
LL完成签到,获得积分10
2分钟前
YU完成签到,获得积分10
2分钟前
Criminology34举报七七求助涉嫌违规
3分钟前
3分钟前
Criminology34举报一条狗求助涉嫌违规
3分钟前
善良安荷完成签到,获得积分10
3分钟前
3分钟前
Kao应助科研通管家采纳,获得10
3分钟前
Kao应助科研通管家采纳,获得10
3分钟前
ataybabdallah完成签到,获得积分10
4分钟前
领导范儿应助张张张采纳,获得10
4分钟前
4分钟前
5分钟前
张张张发布了新的文献求助10
5分钟前
5分钟前
葱饼完成签到 ,获得积分10
5分钟前
彭于晏应助细腻友安采纳,获得30
5分钟前
hailang完成签到 ,获得积分10
5分钟前
Kao应助科研通管家采纳,获得10
5分钟前
科研通AI2S应助qiuqiu采纳,获得10
5分钟前
狂野的含烟完成签到 ,获得积分10
6分钟前
高分求助中
Markov Chain Monte Carlo 10000
(应助此贴封号)【重要!!请各用户(尤其是新用户)详细阅读】【科研通的精品贴汇总】 10000
Common Foundations of American and East Asian Modernisation: From Alexander Hamilton to Junichero Koizumi 5000
Pediatric Dermoscopy Trichoscopy & Onychoscopy 1000
悉尼大学博士学位论文,题目:Modelling and testing of one-sided stitched laminated composites. 作者:Kristopher P. Plain 700
Matrix Methods in Data Mining and Pattern Recognition Second Edition 610
International Security Studies and Technology :Approaches, Assessments, and Frontiers 500
热门求助领域 (近24小时)
化学 材料科学 医学 生物 纳米技术 工程类 有机化学 化学工程 生物化学 计算机科学 内科学 物理 复合材料 催化作用 细胞生物学 无机化学 光电子学 物理化学 电极 基因
热门帖子
关注 科研通微信公众号,转发送积分 7571827
求助须知:如何正确求助?哪些是违规求助? 9151287
关于积分的说明 19572901
捐赠科研通 7156705
什么是DOI,文献DOI怎么找? 3264050
关于科研通互助平台的介绍 2429422
邀请新用户注册赠送积分活动 2254238