This paper addresses the problem of legged locomotion in non-flat terrain. As\nlegged robots such as quadrupeds are to be deployed in terrains with geometries\nwhich are difficult to model and predict, the need arises to equip them with\nthe capability to generalize well to unforeseen situations. In this work, we\npropose a novel technique for training neural-network policies for\nterrain-aware locomotion, which combines state-of-the-art methods for\nmodel-based motion planning and reinforcement learning. Our approach is\ncentered on formulating Markov decision processes using the evaluation of\ndynamic feasibility criteria in place of physical simulation. We thus employ\npolicy-gradient methods to independently train policies which respectively plan\nand execute foothold and base motions in 3D environments using both\nproprioceptive and exteroceptive measurements. We apply our method within a\nchallenging suite of simulated terrain scenarios which contain features such as\nnarrow bridges, gaps and stepping-stones, and train policies which succeed in\nlocomoting effectively in all cases.\n