已入深夜,您辛苦了!由于当前在线用户较少,发布求助请尽量完整地填写文献信息,科研通机器人24小时在线,伴您度过漫漫科研夜!祝你早点完成任务,早点休息,好梦!

Self-supervised reinforcement learning for multi-step object manipulation skills

强化学习 对象(语法) 计算机科学 钢筋 人工智能 心理学 社会心理学
作者
Jiaqi Wang,Chuxin Chen,Jingwei Liu,Guanglong Du,Xiaojun Zhu,Quanlong Guan,Xiaojian Qiu
出处
期刊:Industrial Robot-an International Journal [Emerald Publishing Limited]
卷期号:52 (6): 853-865
标识
DOI:10.1108/ir-12-2024-0534
摘要

Purpose The purpose of this study is to address the challenge of object manipulation in scenarios where the target is not explicitly defined, requiring robots to engage in efficient planning to determine the sequence of actions for picking, placing and positioning objects. The aim is to develop a multistep skill learning method that integrates perception with a set of primitive actions, including a novel action of orienting, to enable robots to perform complex tasks that require multistep planning and interaction with various objects in cluttered and unstructured environments. Design/methodology/approach To achieve the purpose, the authors propose a pipeline that decomposes the object manipulation task into three independent stages, each trained end-to-end with raw visual inputs using off-policy reinforcement learning algorithms. The Q-learning algorithm is used to simultaneously train three fully convolutional neural networks for each primitive action – grasping, pushing, placing and orienting – from scratch. The framework is designed to be modular, allowing for easy extension to multistep manipulation tasks. Findings The findings demonstrate that robots can learn complex behaviors through both simulated and real-world experiments. In simulation, the robot achieved an efficient block-stacking success rate of up to 98% during testing. When transferring the model to a real universal robots UR3 (UR3) robot using effective domain randomization, the robot achieved a 100% completion rate with convex objects and a 92% completion rate with various objects not seen during training. Originality/value We develop a novel multistep skill learning method that integrates perception with multiple primitive actions, including a new action of orienting, and the use of off-policy reinforcement learning algorithms for end-to-end training. The modular design of the framework allows for easy extension to more complex manipulation tasks, and the encouraging results in both simulated and real-world experiments demonstrate significant improvements over current long-term planning methods.
最长约 10秒,即可获得该文献文件

科研通智能强力驱动
Strongly Powered by AbleSci AI
科研通是完全免费的文献互助平台,具备全网最快的应助速度,最高的求助完成率。 对每一个文献求助,科研通都将尽心尽力,给求助人一个满意的交代。
实时播报
Zora完成签到 ,获得积分10
3秒前
团长完成签到 ,获得积分10
3秒前
爱听歌的半凡完成签到,获得积分10
4秒前
5秒前
6秒前
Su发布了新的文献求助30
6秒前
vv完成签到,获得积分20
6秒前
7秒前
7秒前
7秒前
orixero应助Zzzzz采纳,获得10
7秒前
沫茶发布了新的文献求助10
8秒前
CipherSage应助飞天小猫采纳,获得10
9秒前
诸军则应助LOYC采纳,获得10
10秒前
CH完成签到,获得积分10
10秒前
shuaizhou发布了新的文献求助10
10秒前
子川发布了新的文献求助20
11秒前
顾矜应助kkk采纳,获得10
11秒前
果子发布了新的文献求助10
11秒前
11秒前
CH发布了新的文献求助10
13秒前
田様应助accept采纳,获得10
15秒前
15秒前
vv发布了新的文献求助10
16秒前
16秒前
17秒前
18秒前
科研通AI6.4应助dracovu采纳,获得10
18秒前
少卿发布了新的文献求助10
18秒前
吴大王发布了新的文献求助10
19秒前
互助棍哥完成签到,获得积分10
20秒前
七七完成签到,获得积分10
21秒前
21秒前
Zzzzz发布了新的文献求助10
21秒前
21秒前
21秒前
23DD完成签到,获得积分10
22秒前
22秒前
kkk发布了新的文献求助10
23秒前
23秒前
高分求助中
(应助此贴封号)【重要!!请各用户(尤其是新用户)详细阅读】【科研通的精品贴汇总】 10000
HYDROLYSE ACIDE DE QUELQUES DIOXASPIROCYCLANES 1314
Essentials of Carbohydrate Chemistry and Biochemistry, 4th Edition 800
Navigating Normative Orders. Interdisciplinary Perspectives 800
1 Peter and Christ's Descent to the Dead in Its Early Christian Reception 700
Organizational Behavior 510
Management and the Arts 510
热门求助领域 (近24小时)
化学 材料科学 医学 生物 纳米技术 工程类 有机化学 化学工程 生物化学 计算机科学 内科学 物理 复合材料 催化作用 细胞生物学 无机化学 光电子学 物理化学 电极 基因
热门帖子
关注 科研通微信公众号,转发送积分 7749586
求助须知:如何正确求助?哪些是违规求助? 9297320
关于积分的说明 20239682
捐赠科研通 7330885
什么是DOI,文献DOI怎么找? 3309225
关于科研通互助平台的介绍 2460806
邀请新用户注册赠送积分活动 2321503