Deploying the Unmanned Aerial Vehicle (UAV) formation as aerial base stations to construct aerial communication networks in hotspot areas holds considerable promise. However, planning the trajectories of UAV formation in complex and unknown environments, while ensuring obstacle avoidance and formation maintenance, presents unpreceding challenges. In this paper, we propose a hierarchical reinforcement learning-based trajectory planning algorithm for UAV formation. This algorithm implements two-timescale trajectory planning within a leader-follower control framework, where the leader UAV (LUAV) plans the shortest safe trajectory to the hotspot area on a large timescale and the follower UAVs (FUAVs) are responsible for obstacle avoidance and formation maintenance on a small timescale. The LUAV and FUAVs collaborate across different timescales to achieve joint trajectory optimization. To tackle the sparse reward problem in existing learning-based trajectory planning algorithms, we introduce an intrinsic curiosity-driven module that integrates historical information to enhance the exploration of the UAV formation in unknown environments. Our algorithm enhances the UAV formation’s capability to handle complex obstacles and maintain the formation, thereby improving overall performance. Simulation results demonstrate that our algorithm achieves complete obstacle avoidance with a 100% UAV survival rate. Compared to existing algorithms, our proposed algorithm reduces the trajectory length by 14% and improves the formation maintaining performance by over 90%.