计算机科学
稳健性(进化)
编码器
变压器
建筑
水准点(测量)
人工智能
语音识别
手势
自然语言处理
操作系统
地理
基因
视觉艺术
电压
量子力学
大地测量学
物理
化学
生物化学
艺术
作者
Alexandros Koumparoulis,Gerasimos Potamianos
出处
期刊:
日期:2022-04-27
卷期号:: 8467-8471
被引量:24
标识
DOI:10.1109/icassp43922.2022.9747729
摘要
We present a novel resource-efficient end-to-end architecture for lipreading that achieves state-of-the-art results on a popular and challenging benchmark. In particular, we make the following contributions: First, inspired by the recent success of the EfficientNet architecture in image classification and our earlier work on resource-efficient lipreading models (MobiLipNet), we introduce Efficient-Nets to the lipreading task. Second, we show that the currently most popular in the literature 3D front-end contains a max-pool layer that prohibits networks from reaching superior performance and propose its removal. Finally, we improve our system's back-end robustness by including a Transformer encoder. We evaluate our proposed system on the "Lipreading In-The-Wild" (LRW) corpus, a database containing short video segments from BBC TV broadcasts. The proposed network (T-variant) attains 88.53% word accuracy, a 0.17% absolute improvement over the current state-of-the-art, while being five times less computationally intensive. Further, an up-scaled version of our model (L-variant) achieves 89.52%, a new state-of-the-art result on the LRW corpus.
科研通智能强力驱动
Strongly Powered by AbleSci AI