人工智能
计算机科学
计算机视觉
变压器
面子(社会学概念)
面部识别系统
估计
模式识别(心理学)
工程类
社会学
电气工程
系统工程
电压
社会科学
作者
Ahmed Chaouki Chami,Riadh Ajgou
标识
DOI:10.1109/isnib64820.2025.10983229
摘要
Transformers have emerged as powerful models featuring self-attention mechanisms that dynamically weigh the significance of different input elements. While traditionally utilized in Natural Language Processing, their application in image classification remains relatively novel, with Convolutional Neural Networks historically dominating this domain. This study conducts a comparative analysis between Vision Transformers and Convolutional Neural Networks, specifically in the context of face age regression on the MORPH II dataset. By examining existing literature, we identify key architectural differences and performance metrics. Our experiments demonstrate that Vision Transformers outperform Convolutional Neural Networks in terms of accuracy and generalization ability, achieving a Mean Absolute Error of 3.44 compared to 4.85 for Convolutional Neural Networks. Furthermore, we analyze how dataset characteristics, image dimensions, and computational resources impact model performance. This work contributes to understanding the potential of Vision Transformers in image classification tasks, particularly in age estimation, and provides insights for future research in this evolving field.
科研通智能强力驱动
Strongly Powered by AbleSci AI