计算机科学
还原(数学)
内存管理
计算机体系结构
GSM演进的增强数据速率
边缘设备
功率(物理)
人工智能
语音增强
性能增强
建筑
计算机工程
并行计算
深度学习
软件部署
控制(管理)
内存体系结构
张量(固有定义)
资源(消歧)
平面存储模型
记忆模型
并行处理
人工神经网络
循环(图论)
分布式计算
绩效改进
作者
Yonghun Lee,Minjung Kim,Daejin Park
标识
DOI:10.1109/icaiic68212.2026.11454348
摘要
This paper presents a real-time speech enhancement AI model based on TRU-Net, implemented in Pure C for deployment on resource constrained edge devices. High-level machine learning frameworks such as PyTorch and TensorFlow can impose significant limitations in edge environments due to their memory footprint, power consumption, and real-time control constraints. Accordingly, this study emphasizes the importance of architecture design that prioritizes efficient memory utilization in addition to computational performance optimization. By implementing the model in Pure C, memory access patterns, buffer sizes, computational methods, and parallel processing structures are precisely controlled, enabling an experimental analysis of the trade-offs between performance and memory usage. Based on these analyses, an optimized NPU architecture is proposed that minimizes memory consumption, improves the efficiency of intermediate tensor and buffer management, and enhances parallel processing performance through loop unrolling. Experimental results demonstrate that the proposed Pure-C-based model significantly reduces per-frame memory usage and achieves more than a 70% reduction in execution time, validating its effectiveness for real-time, low-power edge AI systems.
科研通智能强力驱动
Strongly Powered by AbleSci AI