We consider evaluating and training a new policy for the evaluation data by\nusing the historical data obtained from a different policy. The goal of\noff-policy evaluation (OPE) is to estimate the expected reward of a new policy\nover the evaluation data, and that of off-policy learning (OPL) is to find a\nnew policy that maximizes the expected reward over the evaluation data.\nAlthough the standard OPE and OPL assume the same distribution of covariate\nbetween the historical and evaluation data, a covariate shift often exists,\ni.e., the distribution of the covariate of the historical data is different\nfrom that of the evaluation data. In this paper, we derive the efficiency bound\nof OPE under a covariate shift. Then, we propose doubly robust and efficient\nestimators for OPE and OPL under a covariate shift by using a nonparametric\nestimator of the density ratio between the historical and evaluation data\ndistributions. We also discuss other possible estimators and compare their\ntheoretical properties. Finally, we confirm the effectiveness of the proposed\nestimators through experiments.\n