计算机科学
克隆(Java方法)
抽象语法树
精确性和召回率
人工智能
图形
语义学(计算机科学)
范围(计算机科学)
源代码
背景(考古学)
语法
自然语言处理
机器学习
代表(政治)
程序设计语言
理论计算机科学
政治
生物
古生物学
法学
DNA
遗传学
政治学
作者
Yifan Zhang,Junwen Yang,Haoyu Dong,Qingchen Wang,Huajie Shao,Kevin Leach,Yu Huang
标识
DOI:10.48550/arxiv.2208.08067
摘要
Large Language Models (LLMs) are transforming software engineering tasks, including code vulnerability detection-a critical area of software security. However, existing methods often rely on resource-intensive models or graph-based techniques, limiting their accessibility and practicality. This paper introduces K-ASTRO, a lightweight Transformer model that combines semantic embeddings from LLMs with structural features of Abstract Syntax Trees (ASTs) to improve both efficiency and accuracy in code vulnerability detection. Our approach introduces an AST-based augmentation technique inspired by mutation testing, a structure-aware attention mechanism that incorporates augmented AST features, and a joint adaptation pipeline to unify code semantics and syntax. Experimental results on three large-scale datasets, including BigVul, DiverseVul, and PrimeVul-demonstrate state-of-the-art performance while enabling rapid inference on CPUs with minimal training time. By offering a scalable, interpretable, and efficient solution, K-ASTRO bridges the gap between LLM advancements and practical software vulnerability detection, providing open-sourced tools to foster further research.
科研通智能强力驱动
Strongly Powered by AbleSci AI