Training effective deep reinforcement learning agents for real-time life-cycle production optimization

强化学习 马尔可夫决策过程 数学优化 计算机科学 增强学习 时间范围 最优控制 生产(经济) 贝尔曼方程 动态规划 任务(项目管理) 人工智能 马尔可夫过程 工程类 数学 统计 宏观经济学 经济 系统工程
作者
Kai Zhang,Zhongzheng Wang,Guodong Chen,Liming Zhang,Yongfei Yang,Chuanjin Yao,Jian Wang,Jun Yao
出处
期刊:Journal of Petroleum Science and Engineering [Elsevier BV]
卷期号:208: 109766-109766 被引量:143
标识
DOI:10.1016/j.petrol.2021.109766
摘要

Life-cycle production optimization aims to obtain the optimal well control scheme at each time control step to maximize financial profit and hydrocarbon production. However, searching for the optimal policy under the limited number of simulation evaluations is a challenging task. In this paper, a novel production optimization method is presented, which maximizes the net present value (NPV) over the entire life-cycle and achieves real-time well control scheme adjustment. The proposed method models the life-cycle production optimization problem as a finite-horizon Markov decision process (MDP), where the well control scheme can be viewed as sequence decisions. Soft actor-critic, known as the state-of-the-art model-free deep reinforcement learning (DRL) algorithm, is subsequently utilized to train DRL agents that can solve the above MDP. The DRL agent strives to maximize long-term NPV rewards as well as the control scheme randomness by training a stochastic policy that maps reservoir states to well control variables and an action-value function that estimates the objective value of the current policy. Since the trained policy is an explicit function structure, the DRL agent can adjust the well control scheme in real-time under different reservoir states. Different from most existing methods that introduce task-specific sensitive parameters or construct complex supplementary structures, the DRL agent learns adaptively by executing goal-directed interactions with an uncertain reservoir environment and making use of accumulated well control experience, which is similar to the actual field well control mode. The key insight here is that the DRL method's ability to utilize gradients information (well-control experience) for higher sample efficiency. The simulation results based on two reservoir models indicate that compared to other optimization methods, the proposed method can attain higher NPV and access excellent performance in terms of oil displacement. • A novel production optimization framework that incorporating advanced deep reinforcement leaning technologies is presented. • The proposed method models the life-cycle production optimization problem as a finite-horizon Markov decision process. • The trained policy is an explicit function structure that utilizing powerful gradient information for higher sample efficiency. • The proposed method achieves excellent performance on one classic control task and two reservoir models.
最长约 10秒,即可获得该文献文件

科研通智能强力驱动
Strongly Powered by AbleSci AI
科研通是完全免费的文献互助平台,具备全网最快的应助速度,最高的求助完成率。 对每一个文献求助,科研通都将尽心尽力,给求助人一个满意的交代。
实时播报
linzx009完成签到,获得积分10
3秒前
冷艳的太君完成签到 ,获得积分10
7秒前
不开花花完成签到 ,获得积分10
7秒前
13秒前
儒雅的夏翠完成签到,获得积分10
13秒前
风中芷容完成签到 ,获得积分10
17秒前
嘎嘎嘎嘎完成签到,获得积分10
19秒前
萂昕完成签到 ,获得积分10
22秒前
万1完成签到,获得积分10
25秒前
25秒前
nkuwangkai完成签到,获得积分10
26秒前
追风hyzhang完成签到,获得积分10
27秒前
趙途嘵生完成签到,获得积分10
29秒前
活力尔柳完成签到 ,获得积分10
29秒前
YoungLee完成签到 ,获得积分10
29秒前
zzz完成签到 ,获得积分20
31秒前
论文我一直读完成签到 ,获得积分10
31秒前
Dain完成签到,获得积分10
31秒前
花花2024完成签到 ,获得积分10
34秒前
坚强的卷儿完成签到 ,获得积分10
38秒前
ding应助112采纳,获得10
38秒前
李爱国应助放牧星空采纳,获得20
39秒前
dan完成签到 ,获得积分10
43秒前
Nicho驳回了Owen应助
43秒前
旅行者N0501完成签到,获得积分10
43秒前
47秒前
舒适刺猬完成签到 ,获得积分10
48秒前
49秒前
112发布了新的文献求助10
52秒前
Dr_Li1741完成签到,获得积分10
53秒前
放牧星空发布了新的文献求助20
53秒前
gf完成签到 ,获得积分10
54秒前
shizhiheng完成签到 ,获得积分10
58秒前
QXS完成签到 ,获得积分10
1分钟前
1分钟前
verymiao完成签到 ,获得积分10
1分钟前
冷静1等待完成签到 ,获得积分10
1分钟前
徐团伟完成签到 ,获得积分10
1分钟前
1分钟前
chen发布了新的文献求助10
1分钟前
高分求助中
(应助此贴封号)【重要!!请各用户(尤其是新用户)详细阅读】【科研通的精品贴汇总】 10000
China Pluperfect I: Epistemology of Past and Outside in Chinese Art 520
Matrix Methods in Data Mining and Pattern Recognition Second Edition 510
Cosmos as Art Object: Studies in Plato's Timaeus and Other Dialogues 500
What is the Future of Psychotherapy in Digital Age? Technology, AI Bots, and Psychotherapy after Covid 444
Management and the Arts 310
Teaching Social and Emotional Learning in Physical Education 300
热门求助领域 (近24小时)
化学 材料科学 医学 生物 纳米技术 工程类 有机化学 化学工程 生物化学 计算机科学 内科学 物理 复合材料 催化作用 细胞生物学 无机化学 光电子学 物理化学 电极 基因
热门帖子
关注 科研通微信公众号,转发送积分 7634336
求助须知:如何正确求助?哪些是违规求助? 9208374
关于积分的说明 19748423
捐赠科研通 7202566
什么是DOI,文献DOI怎么找? 3275029
关于科研通互助平台的介绍 2436932
邀请新用户注册赠送积分活动 2271934