地标
计算机视觉
计算机科学
人工智能
眼动
噪音(视频)
眼球运动
工程类
特征(语言学)
窗口(计算)
作者
Anima Rahman,Roger Woodman,V. Donzella
标识
DOI:10.1109/itsc60802.2025.11423463
摘要
Accurate facial landmark detection is fundamental to many video-based driver monitoring applications. Identifying specific facial points such as the eyes, nose, and mouth is crucial for tracking facial movements and assessing the driver's state. While widely used models like MediaPipe, Dlib, and FAN, perform well on static images, they often struggle with video data, where consistency across frames and robustness to motion blur and head pose variation are essential. In this work, we address these challenges with the newly proposed EyeTrackNet, a two-stage spatiotemporal neural network designed to improve eye landmark prediction stability and blink detection in video sequences. Our method employs a convolutional LSTM to refine predictions over time by learning residual temporal corrections on top of features from a pre-trained spatial backbone. We extensively evaluate EyeTrackNet on multiple video datasets and demonstrate it outperforms common baselines in landmark localisation and blink detection. Notably, EyeTrackNet maintains over 90 % of frames below a Normalised Mean Error of 0.4, a key threshold for reliable blink detection, and achieves the highest overall F1 score on blink detection benchmarks.
科研通智能强力驱动
Strongly Powered by AbleSci AI