计算机科学
传输(电信)
人工智能
计算机视觉
电信
作者
Guangyi Zhang,Hanlei Li,Yunlong Cai,Qiyu Hu,Guanding Yu,Zhijin Qin
标识
DOI:10.1109/tccn.2025.3546935
摘要
We present a novel framework called progressive learned image transmission (PLIT) that utilizes a hierarchical variational autoencoder (VAE) to catalyze semantic communication. PLIT employs autoregressive generation through bottom-up and top-down paths to create several feature representations of the transmitted image, effectively capturing contextual information. In this paper, we investigate the progressive transmission of these representations, particularly in the context of successive refinement. In this scenario, the representations are sent to the receiver in phases, with representations in later phases serving to enhance the image quality. Diverging from previous works, our proposed PLIT offers improved flexibility as it is able to dynamically determine the transmission rate. Specifically, PLIT can transform each representation into different numbers of channel symbols, guided by the hierarchical VAE’s learned priors that indicate the entropy of each representation. In addition, we devise a rate attention mechanism to help adjust the encoding strategy to realize different transmission rates. Furthermore, we introduce a spatial grouping strategy to reduce communication overhead for rate matching without compromising image fidelity. Extensive experiments show that our proposed approach outperforms existing baseline methods in terms of rate-distortion performance and maintains robust performance against channel noise.
科研通智能强力驱动
Strongly Powered by AbleSci AI