残余物
计算机科学
人工智能
代表(政治)
深度学习
模式识别(心理学)
残差神经网络
编码(内存)
理论(学习稳定性)
特征学习
特征(语言学)
人工神经网络
限制
变化(天文学)
外部数据表示
信号(编程语言)
机器学习
算法
卷积神经网络
方案(数学)
图像(数学)
网络体系结构
作者
Sheela Ghoshal,Himanshu Buckchash
出处
期刊:Applied sciences
[Multidisciplinary Digital Publishing Institute]
日期:2026-05-24
卷期号:16 (11): 5252-5252
摘要
Deep ConvNets suffer from gradient signal degradation as network depth increases, limiting effective feature learning in complex architectures. ResNet addressed this through residual connections, but these fixed short circuits cannot adapt to varying input complexity or selectively emphasize task-relevant features across network hierarchies. This study introduces GradAttn, a variation of the residual approach in CNNs that replaces the fixed residual connections with attention-controlled gradient flow. By extracting multi-scale CNN features at different depths and regulating them through self-attention, GradAttn dynamically weights shallow texture features and deep semantic representations. For representational analysis, we evaluated three GradAttn variants across eight diverse datasets: from natural images and medical imaging to fashion recognition. The results demonstrate that GradAttn outperforms ResNet-18 on five of eight datasets, achieving up to +11.07% accuracy improvement on FashionMNIST while maintaining a comparable network size. Gradient flow analysis reveals that controlled instabilities, introduced by attention, often coincide with improved generalization, challenging the assumption that perfect stability is optimal. Furthermore, positional encoding’s effectiveness turned out to be dataset-dependent, with CNN hierarchies frequently encoding sufficient spatial structure. These findings render attention mechanisms as enablers of learnable gradient control, offering a new way for adaptive representation learning in deep neural architectures.
科研通智能强力驱动
Strongly Powered by AbleSci AI