Is AI Ground Truth Really True? The Dangers of Training and Evaluating AI Tools Based on Experts’ Know-What

基本事实 人工智能 计算机科学 质量(理念) 共同点 领域(数学) 工作(物理) 培训(气象学) 知识管理 心理学 工程类 数学 认识论 社会心理学 哲学 物理 气象学 纯数学 机械工程
作者
Sarah Lebovitz,Natalia Levina,Hila Lifshitz‐Assaf
出处
期刊:Management Information Systems Quarterly [MIS Quarterly]
卷期号:45 (3): 1501-1526 被引量:49
标识
DOI:10.25300/misq/2021/16564
摘要

Organizational decision-makers need to evaluate AI tools in light of increasing claims that such tools outperform human experts. Yet, measuring the quality of knowledge work is challenging, raising the question of how to evaluate AI performance in such contexts. We investigate this question through a field study of a major U.S. hospital, observing how managers evaluated five different machine-learning (ML) based AI tools. Each tool reported high performance according to standard AI accuracy measures, which were based on ground truth labels provided by qualified experts. Trying these tools out in practice, however, revealed that none of them met expectations. Searching for explanations, managers began confronting the high uncertainty of experts’ know-what knowledge captured in ground truth labels used to train and validate ML models. In practice, experts address this uncertainty by drawing on rich know-how practices, which were not incorporated into these ML-based tools. Discovering the disconnect between AI’s know-what and experts’ know-how enabled managers to better understand the risks and benefits of each tool. This study shows dangers of treating ground truth labels used in ML models objectively when the underlying knowledge is uncertain. We outline implications of our study for developing, training, and evaluating AI for knowledge work.
最长约 10秒,即可获得该文献文件

科研通智能强力驱动
Strongly Powered by AbleSci AI
科研通是完全免费的文献互助平台,具备全网最快的应助速度,最高的求助完成率。 对每一个文献求助,科研通都将尽心尽力,给求助人一个满意的交代。
实时播报
1秒前
5秒前
6秒前
6秒前
6秒前
张欢馨应助pengyufen采纳,获得10
7秒前
稳重的灵安完成签到,获得积分10
7秒前
我是老大应助zym采纳,获得10
8秒前
8秒前
hyl完成签到,获得积分10
8秒前
7badada馒困的完成签到 ,获得积分10
8秒前
Bond完成签到 ,获得积分10
9秒前
大和湘子发布了新的文献求助10
10秒前
小马甲应助清茶旧友采纳,获得10
10秒前
星辰大海应助饱满的恋风采纳,获得10
11秒前
gwy发布了新的文献求助10
11秒前
12秒前
12秒前
16秒前
在水一方应助yzkyg采纳,获得10
18秒前
bkagyin应助lijingyi采纳,获得10
18秒前
shhs发布了新的文献求助10
18秒前
19秒前
靓丽翩跹完成签到,获得积分10
19秒前
科研通AI6.4应助yutao采纳,获得10
20秒前
斯文败类应助李涵霖采纳,获得10
21秒前
21秒前
nast1c完成签到,获得积分10
22秒前
shen发布了新的文献求助10
23秒前
LX发布了新的文献求助10
24秒前
24秒前
25秒前
酷波er应助科研通管家采纳,获得10
25秒前
科目三应助科研通管家采纳,获得10
26秒前
26秒前
肖肖发布了新的文献求助10
26秒前
今后应助坚强的冷荷采纳,获得10
26秒前
26秒前
隐形曼青应助科研通管家采纳,获得10
26秒前
咕哒发布了新的文献求助10
26秒前
高分求助中
(应助此贴封号)【重要!!请各用户(尤其是新用户)详细阅读】【科研通的精品贴汇总】 10000
Matrix Methods in Data Mining and Pattern Recognition Second Edition 510
Discerning Saints: Moralization of Intrinsic Motivation and Selective Prosociality at Work 500
Handbuch Trainingswissenschaft – Trainingslehre 500
Additive Manufacturing Design and Applications (ASM Handbook, Volume 24A) 500
Variations: A More Diverse Picture of Contemporary Art 400
Induction Heating and Heat Treatment (ASM Handbook, Volume 4C) 300
热门求助领域 (近24小时)
化学 材料科学 医学 生物 纳米技术 工程类 有机化学 化学工程 生物化学 计算机科学 内科学 物理 复合材料 催化作用 细胞生物学 无机化学 光电子学 物理化学 电极 基因
热门帖子
关注 科研通微信公众号,转发送积分 7589998
求助须知:如何正确求助?哪些是违规求助? 9167502
关于积分的说明 19622389
捐赠科研通 7169320
什么是DOI,文献DOI怎么找? 3267205
关于科研通互助平台的介绍 2432112
邀请新用户注册赠送积分活动 2259412