挖掘机
计算机科学
变压器
人工智能
提取器
推论
编码器
动作识别
特征提取
计算机视觉
模式识别(心理学)
工程类
机械工程
电压
工艺工程
电气工程
班级(哲学)
操作系统
作者
A.H. Martin,Andrew J. Hill,Konstantin M. Seiler,Mehala Balamurali
标识
DOI:10.1080/17480930.2023.2290364
摘要
In mining and construction, excavators are integral to earth-moving operations. Accurate knowledge of excavator activities may be used in productivity analysis to streamline delivery. This paper presents a computer vision-based method for excavator action detection which can automatically inference the occurrence and time duration of excavator actions from untrimmed video captured from the excavator cab. The model uses a three-stage architecture consisting of a VGG16 feature extractor, a four-stage Transformer Encoder-Long Short-Term Memory (LSTM) module, and a post-processing component. The model's predictive performance has been validated on the largest dataset among similar studies, comprising 567,000 frames filmed on-site at day and night. When tested on night and daytime videos, the model achieves accuracies of 90% and 70%, respectively, highlighting strong potential for practical implementation of the Transformer-LSTM network in excavator action detection. This study presents the first application of the combined Transformer-LSTM network for action detection in computer vision.
科研通智能强力驱动
Strongly Powered by AbleSci AI