计算机科学
语音识别
理论(学习稳定性)
班级(哲学)
可塑性
人工智能
弹丸
模式识别(心理学)
机器学习
热力学
物理
有机化学
化学
作者
Yongjie Si,Yanxiong Li,Jiaxin Tan,Guoqing Chen,Qianqian Li,Mladen Russo
标识
DOI:10.1109/taslpro.2025.3527147
摘要
In the problem of Few-shot Class-incremental Audio Classification (FCAC), training samples per class in the base session are required to be abundant. However, in many scenarios, it is difficult to collect abundant training samples in the base session because of data scarcity and high collection cost. In this paper, we explore a new FCAC problem, namely Fully FCAC (FFCAC), in which training samples for all classes in both the base and incremental sessions are few. Moreover, we propose a FFCAC method by adaptively improving the model's stability for seen classes and plasticity for unseen classes. The model consists of an Embedding Extractor (EE) and an Evolvable Classifier (EC). The EE consists of an encoder of pretrained Audio Spectrogram Transformer (AST), an encoder of finetuned AST and three fusion modules. In each incremental session, the two encoders are frozen for memorizing the knowledge learned by the model and thus can improve the model's stability. The three fusion modules are used to fuse the embeddings output by the two encoders. The EC is composed of a fully-connected layer and a Softmax layer. The fusion modules and EC are updated in incremental sessions to improve the model's plasticity for unseen classes. Besides, two losses are defined to train the model in the base and incremental sessions. Results on three public audio datasets (LS-100, NSynth-100, and FSC-89) show that our FFCAC method exceeds previous methods in accuracy under many conditions. The code is at https://github.com/YongjieSi/AISP.
科研通智能强力驱动
Strongly Powered by AbleSci AI