强化学习
马尔可夫决策过程
计算机科学
智能电网
随机性
贝尔曼方程
马尔可夫过程
数学优化
调度(生产过程)
需求响应
动态定价
电价
电
人工智能
工程类
电力市场
数学
统计
业务
营销
电气工程
作者
Zhiqiang Wan,Hepeng Li,Haibo He,Danil Prokhorov
标识
DOI:10.1109/tsg.2018.2879572
摘要
Driven by the recent advances in electric vehicle (EV) technologies, EVs become important for smart grid economy. When EVs participate in demand response program which has real-time pricing signals, the charging cost can be greatly reduced by taking full advantage of these pricing signals. However, it is challenging to determine an optimal charging strategy due to the existence of randomness in traffic conditions, user’s commuting behavior, and pricing process of the utility. Conventional model-based approaches require a model of forecast on the uncertainty and optimization for the scheduling process. In this paper, we formulate this scheduling problem as a Markov Decision Process (MDP) with unknown transition probability. A model-free approach based on deep reinforcement learning is proposed to determine the optimal strategy for this problem. The proposed approach can adaptively learn the transition probability and does not require any system model information. The architecture of the proposed approach contains two networks: a representation network to extract discriminative features from the electricity prices and a Q network to approximate the optimal action-value function. Numerous experimental results demonstrate the effectiveness of the proposed approach.
科研通智能强力驱动
Strongly Powered by AbleSci AI