计算机科学
可扩展性
程序员
氙气
至强融核
架空(工程)
吞吐量
嵌入式系统
软件
计算机体系结构
操作系统
并行计算
无线
作者
Reese Kuper,Ipoom Jeong,Yifan Yuan,Ren Wang,Narayan Ranganathan,Nikhil Rao,Jiayu Hu,Sanjay Kumar,Philip Lantz,Nam Sung Kim
标识
DOI:10.1145/3620665.3640401
摘要
As semiconductor power density is no longer constant with the technology process scaling down, we need different solutions if we are to continue scaling application performance. To this end, modern CPUs are integrating capable data accelerators on the chip, aiming to improve performance and efficiency for a wide range of applications and usages. One such accelerator is the Intel® Data Streaming Accelerator (DSA) introduced since Intel® 4th Generation Xeon® Scalable CPUs (Sapphire Rapids). DSA targets data movement operations in memory that are common sources of overhead in datacenter workloads and infrastructure. In addition, it supports a wider range of operations on streaming data, such as CRC32 calculations, computation of deltas between data buffers, and data integrity field (DIF) operations. This paper aims to introduce the latest features supported by DSA, dive deep into its versatility, and analyze its throughput benefits through a comprehensive evaluation with both microbenchmarks and real use cases. Along with the analysis of its characteristics and the rich software ecosystem of DSA, we summarize several insights and guidelines for the programmer to make the most out of DSA, and use an in-depth case study of DPDK Vhost to demonstrate how these guidelines benefit a real application.
科研通智能强力驱动
Strongly Powered by AbleSci AI