Optical remote sensing images (ORSI) feature unique scenes and complex imaging conditions. Specifically, they exhibit substantial variations in object scale, quantity, structure, and distribution. Consequently, salient object detection in ORSI (ORSI-SOD) is pivotal in ORSI content perception and understanding. Additionally, the limitations of the single modality impede the advancement of ORSI-SOD. To tackle these issues, we propose a Depth-guided and Iterative Refinement Network (DINet) for ORSI-SOD. By incorporating depth information as auxiliary cues, we introduce a multi-modal strategy for ORSI-SOD, resulting in improved accuracy in the localization and segmentation of salient objects. To address the variability of salient objects, we design an Aggregation Perception Enhancement (APE) Module. This module integrates complementary cues from cross-modal features using multi-dimensional attention mechanisms. By fostering cross-modal interactions, the APE module effectively preserves both detail and spatial location information. Furthermore, we propose an Iterative Guidance Refinement Decoder to handle boundary uncertainty. The decoder uses initial predictions to guide the decoding phase and iteratively refine results. Simultaneously, it minimizes noise from depth cues, yielding predictions with more accurate boundaries. Experimental comparisons with 22 state-of-the-art methods show that DINet exhibits superior performance while maintaining lightweight (11.98M) and real-time (55FPS) capabilities.