计算机科学
恶意软件
JPEG格式
计算机视觉
人工智能
变压器
计算机图形学(图像)
模式识别(心理学)
计算机安全
数据压缩
工程类
电气工程
电压
作者
Binghui Zou,Chunjie Cao,Fangjian Tao,Yang Sun,Longjuan Wang,Yuqing Zhang,Jingzhang Sun
标识
DOI:10.1109/tdsc.2025.3596831
摘要
Malware is proliferating at an exponential rate in cyberspace, posing serious threats to on-device systems characterized by limited computational capabilities. In this work, we address the critical challenge posed by data imbalance—where rare malware families receive inadequate representation—by proposing MalViT, a lightweight Vision Transformer (ViT) architecture that directly operates in the JPEG frequency domain. Rather than converting Huffman-coded signals into RGB spatial images, MalViT leverages Discrete Cosine Transform (DCT) coefficients to reduce data redundancy and computational overhead. We further improve the model’s generalization through both pre-training and fine-tuning workflows. Comprehensive evaluations on two large-scale, real-world malware datasets, MalNet-Image (1.26 M samples) and BODMAS (51 K samples), demonstrate that MalViT accelerates data loading by nearly threefold compared to existing methods. On GPU and CPU, MalViT achieves approximately 2.0× and 4.7× faster inference throughput than MobileViT, respectively, while incurring minimal or even improved accuracy loss. When processing 224 × 224-pixel JPEG images, MalViT completes inference within an average of 3.12ms per sample, which is 8.79× faster than VisMal and 5.91× faster than ViT4Mal. Furthermore, its compact design comprises only 1.1M parameters and requires 10M MACs, making it particularly suitable for resource-constrained on-device deployment.
科研通智能强力驱动
Strongly Powered by AbleSci AI