已入深夜,您辛苦了!由于当前在线用户较少,发布求助请尽量完整地填写文献信息,科研通机器人24小时在线,伴您度过漫漫科研夜!祝你早点完成任务,早点休息,好梦!

Direct Preference Optimization: Your Language Model is Secretly a Reward Model

自动汇总 超参数 计算机科学 强化学习 人工智能 机器学习 偏爱 偏好学习 质量(理念) 时差学习 简单(哲学) 无监督学习 控制(管理) 数学 统计 哲学 认识论
作者
Rafael Rafailov,Archit Sharma,Eric Mitchell,Stefano Ermon,Christopher D. Manning,Chelsea Finn
出处
期刊:Cornell University - arXiv [Cornell University]
被引量:167
标识
DOI:10.48550/arxiv.2305.18290
摘要

While large-scale unsupervised language models (LMs) learn broad world knowledge and some reasoning skills, achieving precise control of their behavior is difficult due to the completely unsupervised nature of their training. Existing methods for gaining such steerability collect human labels of the relative quality of model generations and fine-tune the unsupervised LM to align with these preferences, often with reinforcement learning from human feedback (RLHF). However, RLHF is a complex and often unstable procedure, first fitting a reward model that reflects the human preferences, and then fine-tuning the large unsupervised LM using reinforcement learning to maximize this estimated reward without drifting too far from the original model. In this paper we introduce a new parameterization of the reward model in RLHF that enables extraction of the corresponding optimal policy in closed form, allowing us to solve the standard RLHF problem with only a simple classification loss. The resulting algorithm, which we call Direct Preference Optimization (DPO), is stable, performant, and computationally lightweight, eliminating the need for sampling from the LM during fine-tuning or performing significant hyperparameter tuning. Our experiments show that DPO can fine-tune LMs to align with human preferences as well as or better than existing methods. Notably, fine-tuning with DPO exceeds PPO-based RLHF in ability to control sentiment of generations, and matches or improves response quality in summarization and single-turn dialogue while being substantially simpler to implement and train.
最长约 10秒,即可获得该文献文件

科研通智能强力驱动
Strongly Powered by AbleSci AI
科研通是完全免费的文献互助平台,具备全网最快的应助速度,最高的求助完成率。 对每一个文献求助,科研通都将尽心尽力,给求助人一个满意的交代。
实时播报
隐形曼青应助张艺凡采纳,获得10
刚刚
1秒前
1秒前
初景发布了新的文献求助10
3秒前
su完成签到 ,获得积分10
3秒前
海上星完成签到,获得积分10
3秒前
eily完成签到,获得积分10
4秒前
明理汲发布了新的文献求助10
4秒前
家俊发布了新的文献求助10
5秒前
ljm完成签到,获得积分20
6秒前
7秒前
8秒前
晕晕鲨发布了新的文献求助10
8秒前
斯文败类应助我是KJ采纳,获得10
9秒前
江枫渔火VC完成签到 ,获得积分10
9秒前
李侑勳完成签到,获得积分10
11秒前
kawayifenm完成签到,获得积分10
11秒前
hiraabb完成签到 ,获得积分10
12秒前
雷欣儿发布了新的文献求助10
12秒前
ljm发布了新的文献求助10
12秒前
tfli发布了新的文献求助10
13秒前
13秒前
科研通AI6.2应助Jocelyn采纳,获得10
15秒前
FashionBoy应助好日子谁在过采纳,获得10
17秒前
17秒前
18秒前
18秒前
18秒前
18秒前
18秒前
Lucas应助科研通管家采纳,获得10
18秒前
雷欣儿完成签到,获得积分10
18秒前
GingerF应助科研通管家采纳,获得10
18秒前
19秒前
19秒前
在水一方应助zzz采纳,获得10
21秒前
舒适的如萱应助忧虑的勒采纳,获得10
21秒前
天天快乐应助xiao采纳,获得10
21秒前
21秒前
远志发布了新的文献求助10
23秒前
高分求助中
(应助此贴封号)【重要!!请各用户(尤其是新用户)详细阅读】【科研通的精品贴汇总】 10000
Principles of town planning: translating concepts to applications 1000
Navigating Normative Orders. Interdisciplinary Perspectives 800
1 Peter and Christ's Descent to the Dead in Its Early Christian Reception 700
Organizational Behavior 510
Management and the Arts 510
Matrix Methods in Data Mining and Pattern Recognition Second Edition 510
热门求助领域 (近24小时)
化学 材料科学 医学 生物 纳米技术 工程类 有机化学 化学工程 生物化学 计算机科学 内科学 物理 复合材料 催化作用 细胞生物学 无机化学 光电子学 物理化学 电极 基因
热门帖子
关注 科研通微信公众号,转发送积分 7738383
求助须知:如何正确求助?哪些是违规求助? 9287511
关于积分的说明 20183613
捐赠科研通 7316252
什么是DOI,文献DOI怎么找? 3305861
关于科研通互助平台的介绍 2458182
邀请新用户注册赠送积分活动 2315722