In this paper we propose a reinforcement learning scheme for finding optimal and sub-optimal policies for the finite state Markov decision problem (MDP) with the infinite horizon discounted cost criterion. Online learning is utilized along with temporal difference schemes for approximating value functions to obtain a direct adaptive control scheme for the MDP. The approach features the approximation of stationary deterministic policies with randomized policies. We provide convergence results of the algorithm under very reasonable assumptions, in particular without aperiodicity assumptions.