计算机科学
卷积(计算机科学)
熵(时间箭头)
可分离空间
人工智能
图像压缩
压缩(物理)
人工神经网络
数学
图像(数学)
图像处理
数学分析
材料科学
物理
复合材料
量子力学
作者
Dongjian Yang,Xiaopeng Fan,Xiandong Meng,Debin Zhao
标识
DOI:10.1109/dcc62719.2025.00098
摘要
Recently, neural image compression (NIC) has made remarkable progress. Two key parts of NIC are the encoder-decoder and the entropy model. For the encoder-decoder, a larger effective receptive field (ERF) means a stronger transformation ability. Existing methods usually enlarge the ERF at the expense of complexity, which is intolerable. To address this issue, we propose a multi-scale depthwise separable dilated convolution (MSDSDC) to build the encoder-decoder. Specifically, we first construct a depthwise separable dilated convolution (DSDC) by using the depthwise separable strategy in dilated convolution to reduce its complexity. Subsequently, multi-scale features extracted by three DSDCs with varying dilation rates are fused to expand the ERF of the encoder-decoder, consequently enhancing its transformation capability. Besides, we design a multi-distribution mixture entropy model (MDMEM) to further enhance the flexibility of latent representation probability modeling. The experimental results demonstrate that our proposed method achieves the best balance between rate-distortion performance and complexity.
科研通智能强力驱动
Strongly Powered by AbleSci AI