Spatio-Temporal Saliency Prediction for Panoramic Videos via Motion-Aware Cues and Viewing Priors
作者
Y Wang,Chunyi Chen
标识
DOI:10.1109/eiecs67708.2025.11283447
摘要
Accurate saliency prediction in panoramic video is essential for virtual reality (VR) and other immersive media. Existing approaches, however, still struggle to capture both temporal dependencies and gaze biases simultaneously. Our approach introduces a unified spherical spatio-temporal framework, the core of which is a spherical ConvLSTM encoder-decoder architecture. At the network bottleneck, we have embedded a feature-level temporal extraction layer composed of spherical motion excitation (SME). This layer computes the differences between frames to generate temporally adaptive channel weights, thereby highlighting motion-related saliency cues. Additionally, we introduce a learnable Gaussian prior, fusing it with the decoder prediction in the pixel domain to model viewing bias and low-frequency global structure in panoramic content. The framework is trained from start to finish using a spherical weighted Kullback-Leibler (KL) divergence that is geometrically consistent with the sphere. Experiments on public panoramic video datasets demonstrate consistent improvements over strong baselines across multiple metrics, as well as closer alignment with human fixations. This establishes an effective and useful solution for spherical saliency prediction.