Food-500 Cap: A Fine-Grained Food Caption Benchmark for Evaluating Vision-Language Models

可解释性 计算机科学 水准点(测量) 领域(数学分析) 人工智能 机器学习 一般化 自然语言处理 地理 数学分析 数学 大地测量学
作者
Zheng Ma,Mianzhi Pan,Wenhan Wu,Kanzhi Cheng,Jianbing Zhang,Shujian Huang,Jiajun Chen
标识
DOI:10.1145/3581783.3611994
摘要

Vision-language models (VLMs) have shown impressive performance in substantial downstream multi-modal tasks. However, only comparing the fine-tuned performance on downstream tasks leads to the poor interpretability of VLMs, which is adverse to their future improvement. Several prior works have identified this issue and used various probing methods under a zero-shot setting to detect VLMs' limitations, but they all examine VLMs using general datasets instead of specialized ones. In practical applications, VLMs are usually applied to specific scenarios, such as e-commerce and news fields, so the generalization of VLMs in specific domains should be given more attention. In this paper, we comprehensively investigate the capabilities of popular VLMs in a specific field, the food domain. To this end, we build a food caption dataset, Food-500 Cap, which contains 24,700 food images with 494 categories. Each image is accompanied by a detailed caption, including fine-grained attributes of food, such as the ingredient, shape, and color. We also provide a culinary culture taxonomy that classifies each food category based on its geographic origin in order to better analyze the performance differences of VLM in different regions. Experiments on our proposed datasets demonstrate that popular VLMs underperform in the food domain compared with their performance in the general domain. Furthermore, our research reveals severe bias in VLMs' ability to handle food items from different geographic regions. We adopt diverse probing methods and evaluate nine VLMs belonging to different architectures to verify the aforementioned observations. We hope that our study will bring researchers' attention to VLM's limitations when applying them to the domain of food or culinary cultures, and spur further investigations to address this issue.

科研通智能强力驱动
Strongly Powered by AbleSci AI
科研通是完全免费的文献互助平台,具备全网最快的应助速度,最高的求助完成率。 对每一个文献求助,科研通都将尽心尽力,给求助人一个满意的交代。
实时播报
1秒前
123完成签到,获得积分10
2秒前
popopanda完成签到,获得积分10
2秒前
在水一方应助Fangyu采纳,获得10
3秒前
3秒前
3秒前
4秒前
Kobe完成签到,获得积分10
4秒前
无花果应助11采纳,获得10
5秒前
科目三应助忆枫采纳,获得10
5秒前
小宋发布了新的文献求助10
5秒前
123发布了新的文献求助10
6秒前
999999完成签到,获得积分20
6秒前
7秒前
凯zdwx发布了新的文献求助10
8秒前
8秒前
8秒前
高山发布了新的文献求助10
8秒前
wang完成签到,获得积分10
9秒前
隐形完成签到,获得积分20
9秒前
11414发布了新的文献求助10
9秒前
999999发布了新的文献求助10
10秒前
10秒前
yutingting完成签到,获得积分10
11秒前
yyuchen发布了新的文献求助10
11秒前
小Q发布了新的文献求助10
12秒前
rs完成签到 ,获得积分10
12秒前
12秒前
12秒前
CaiLing完成签到,获得积分20
12秒前
大头完成签到 ,获得积分10
13秒前
领导范儿应助去阿中午采纳,获得10
13秒前
wwwww发布了新的文献求助10
14秒前
海大鱼完成签到,获得积分10
15秒前
15秒前
16秒前
17秒前
17秒前
imp完成签到,获得积分10
17秒前
阔达故事发布了新的文献求助10
17秒前
高分求助中
(应助此贴封号)【重要!!请各用户(尤其是新用户)详细阅读】【科研通的精品贴汇总】 10000
Rosenblum, Global Change Biology 800
Essentials of Carbohydrate Chemistry and Biochemistry, 4th Edition 800
Organizational Behavior 510
Management and the Arts 510
Matrix Methods in Data Mining and Pattern Recognition Second Edition 510
Physiologic specialization in Peronospora manshurica 500
热门求助领域 (近24小时)
化学 材料科学 医学 生物 纳米技术 计算机科学 化学工程 工程类 有机化学 物理 复合材料 生物化学 内科学 细胞生物学 基因 遗传学 免疫学 冶金 光电子学 癌症研究
热门帖子
关注 科研通微信公众号,转发送积分 7777229
求助须知:如何正确求助?哪些是违规求助? 9318314
关于积分的说明 20363610
捐赠科研通 7364312
什么是DOI,文献DOI怎么找? 3318873
关于科研通互助平台的介绍 2466526
邀请新用户注册赠送积分活动 2334092