感知
美学
计算机视觉
计算机科学
图像(数学)
心理学
人工智能
认知心理学
艺术
神经科学
作者
Lanjun Wang,Zhi Qiao,Ruidong Chen,Jingqiu Li,Wenjie Wang,Xiaoqiong Wang,Wei Rao,Shuai Chen,An‐An Liu
出处
期刊:
日期:2025-03-12
卷期号:: 1-5
标识
DOI:10.1109/icassp49660.2025.10889477
摘要
Image Aesthetic Assessment (IAA) aims to rate the aesthetic quality of images and has many practical applications. However, existing methods typically rely on limited annotated data for training, leading to two key issues: 1) score-only predictions lack interpretability, making it hard for users to understand the reasoning behind ratings; 2) the assessment ability learned through supervised training struggles to generalize to scenarios beyond the training data. To address these challenges, we leverage Multi-Modal Large Language Models (MLLMs) for interpretable image aesthetic assessment. Drawing inspiration from the human aesthetic perception process, we propose two key components: aesthetic attribute assessment (AAA) and scene-aware in-context learning (ICL). AAA is to provide detailed attribute-based analysis by prompting with evaluation criteria. Meanwhile, scene-aware ICL is to improve the model’s understanding of the aesthetic scoring principles across different scenes with given corresponding references. The output of these two components is used to guide the model to provide ratings and interpretations for more reliable and understandable results. Experiments across multiple datasets show the effectiveness of our framework in enhancing MLLM’s aesthetic perception ability and underscore its potential for interpretable IAA.
科研通智能强力驱动
Strongly Powered by AbleSci AI