Tomographic SAR (TomoSAR) at low frequency, i.e., P/L band, has become a promising tool for forest structure study. Forest canopy height and underlying topography are two of the most important parameters one can estimate using TomoSAR technique. One simple way to estimate these two parameters is to detecting the peaks of the tomographic profile, which, however, can lead to large biases due to complicated forest structure, sidelobes or insufficient TomoSAR resolution. Polarimetric TomoSAR (Pol-TomoSAR) provides a solution to this by exploring the polarimetric diversity to separate the ground and canopy components and then conduct independent TomoSAR analysis. However, Pol-TomoSAR technique suffers from low ground-to-volume ratio (GVR), which often leads to unsuccessful ground and canopy separation. To mitigate this propblem, in this paper, we provide a deep-learning solution to ground and canopy height estimation from 3D tomographic profile through the identification of the patterns of ground and canopy peaks. A 3D U-net model is introduced in our solution to grasp as much three-dimensional characteristics of the tomographic profile as possible. Moreover, our model can be well trained using only synthetic TomoSAR dataset, making it easy to implement when we don’t have enough real data with LiDAR references. The proposed method is validated on P-band real TomoSAR dataset from multiple test sites in AfriSAR campaign, showing that it can achieve more accurate ground and canopy height estimation than the state-of-the-art Pol-TomoSAR techniques. The maximum RMSE improvement reaches as high as 66.4% and 63.2% for ground and canopy top height, respectively.