Dong Li, Zou Jian, Jiang Feibo, et al. Vision-Language Model- and World Model-Empowered Video Semantic Communication[J/OL]. Journal on Communications, 2026.
DOI:
Dong Li, Zou Jian, Jiang Feibo, et al. Vision-Language Model- and World Model-Empowered Video Semantic Communication[J/OL]. Journal on Communications, 2026.DOI: 10.11959/j.issn.1000-436x.TXXB260376.
Vision-Language Model- and World Model-Empowered Video Semantic Communication
The rapid development of intelligent video services has imposed increasingly stringent requirements on the efficiency
robustness
and semantic fidelity of wireless video transmission. However
existing video transmission systems still face several challenges
including substantial semantic redundancy
difficulties in generative reconstruction
and vulnerability to data distribution shifts. To address these challenges
a multi-task video semantic communication system empowered by a vision-language model (VLM) and a world model (WM) was proposed. At the transmitter
the original video was compressed by the VLM into keyframes and textual semantic descriptions
such that transmission redundancy was reduced while essential semantic content was preserved. At the receiver
high-quality videos were reconstructed by the WM based on the received keyframes and textual semantic descriptions. To mitigate task-specific data distribution shifts in dynamic environments
a continual semantic adaptation mechanism was further developed
through which the semantic and channel encoders and decoders were stably updated using cross-task optimization to preserve semantic understanding across different tasks. High semantic consistency
stability
and generalization capability under different video tasks and channel conditions were demonstrated by the experimental results.
Zhang P , Xu X D , Niu K , et al . Theoretical and technical system of modern semantic communications and 6G intelligent-simple networks [J ] . Journal of Beijing University of Posts and Telecommunications , 2025 , 48 ( 5 ): 1 - 16 .
Nan G , Li Z , Zhai J , et al . Physical-layer adversarial robustness for deep learning-based semantic communications [J ] . IEEE Journal on Selected Areas in Communications , 2023 , 41 ( 8 ): 2592 – 2608 .
Zhang P , Xu W , Liu Y , et al . Intellicise wireless networks from semantic communications: A survey, research issues, and challenges [J ] . IEEE Communications Surveys & Tutorials , 2024 , 27 ( 3 ): 2051 – 2084 .
Zhang P , Niu K , Yao S S , et al . Semantic communications for future: basic principle and implementation methodology [J ] . Journal on Communications , 2023 , 44 ( 5 ): 1 - 14 .
Cui Q , You X , Wei N , et al . Overview of AI and communication for 6G network: Fundamentals, challenges, and future research opportunities [J ] . Science China Information Sciences , 2025 , 68 ( 7 ): 171301 .
Jiang F , Tu S , Dong L , et al . FlashSAM: Lightweight Vision Model for Multi-UAV Token Communication in Low-Altitude Wireless Networks [J ] . IEEE Journal of Selected Topics in Signal Processing , 2026 .
Kirkpatrick J , Pascanu R , Rabinowitz N , et al . Overcoming catastrophic forgetting in neural networks [J ] . Proceedings of the national academy of sciences , 2017 , 114 ( 13 ): 3521 – 3526 .
Jiang F , Tang C , Dong L , et al . Visual language model based cross-modal semantic communication systems [J ] . IEEE Transactions on Wireless Communications , 2025 , 24 ( 5 ): 3937 – 3948 .
Zhao Y , Yue Y , Hou S , et al . LaMoSC: Large language model-driven semantic communication system for visual transmission [J ] . IEEE Transactions on Cognitive Communications and Networking , 2024 , 10 ( 6 ): 2005 – 2018 .
Liang S D . Vision language models for massive MIMO semantic communication [C ] // Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops . 2025 : 1675 - 1685 .
Jiang P , Guo J , Wen C-K , et al . Semantic Communications with World Models [J ] . `arXiv preprint arXiv: 2510.24785 , 2025 .
Lokumarambage M , Sivalingam T , Dong F , et al . Latent Prediction based Generative Semantic Communication for Video Transmission in Wireless Networks [J ] . IEEE Open Journal of the Communications Society , 2026 , 7 : 3974 – 3986 .
Jiang F , Tu S , Dong L , et al . Tokencom: Vision-language model for multimodal and multitask token communications [J ] . arXiv preprint arXiv: 260300482 , 2026 .
Bertasius G , Wang H , Torresani L . Is space-time attention all you need for video understanding? [C ] // Proceedings of the 38th International Conference on Machine Learning . 2021 : 813 - 824 .
Peebles W , Xie S . Scalable diffusion models with transformers [C ] // Proceedings of the IEEE/CVF International Conference on Computer Vision . 2023 : 4195 - 4205 .
Zhang H , Lei Y , Gui L , et al . CPPO: continual learning for reinforcement learning with human feedback [C ] // Proceedings of the 12th International Conference on Learning Representations (ICLR) . Vienna : OpenReview , 2024 .
Soomro K , Zamir A R , Shah M . Ucf101: A dataset of 101 human actions classes from videos in the wild [J ] . arXiv preprint arXiv: 1212.0402 , 2012 .
Carreira J , Noland E , Hillier C , et al . A short note on the kinetics-700 human action dataset [J ] . arXiv preprint arXiv: 1907.06987 , 2019 .
Sigurdsson G A , Varol G , Wang X , et al . Hollywood in homes: Crowdsourcing data collection for activity understanding [C ] // European Conference on Computer Vision . 2016 : 510 - 526 .
Wang S , Dai J , Liang Z , et al . Wireless deep video semantic transmission [J ] . IEEE Journal on Selected Areas in Communications , 2023 , 41 ( 1 ): 214 - 229 .
Yin H , Qiao L , Ma Y , et al . Generative video semantic communication via multimodal semantic fusion with large model [J ] . IEEE Transactions on Vehicular Technology , 2025 .
Wang Y , Yu S , Yang X . Error robustness scheme for H.264 based on LDPC code [C ] // 2006 12th International Multi-Media Modelling Conference . 2006 : 79 .
Maaz M , Rasheed H , Khan S , et al . Videogpt+: Integrating image and video encoders for enhanced video understanding [J ] . arXiv preprint arXiv: 2406.09418 , 2024 .
Wan T , Wang A , Ai B , et al . Wan: Open and advanced large-scale video generative models [J ] . arXiv preprint arXiv: 2503.20314 , 2025 .
Li K , Wang Y , He Y , et al . MVBench: A comprehensive multi-modal video understanding benchmark [C ] // Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition . 2024 : 22195 - 22206 .
Cheng Z , Leng S , Zhang H , et al . Videollama 2: Advancing spatial-temporal modeling and audio understanding in video-llms [J ] . arXiv preprint arXiv: 2406.07476 , 2024 .
Yang Z , Teng J , Zheng W , et al . CogVideoX: Text-to-video diffusion models with an expert transformer [C ] // International Conference on Learning Representations . 2025 .
Hessel J , Holtzman A , Forbes M , et al . CLIPScore: A reference-free evaluation metric for image captioning [C ] // Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing . 2021 : 7514 - 7528 .
Venugopalan S , Rohrbach M , Donahue J , et al . Sequence to sequence-video to text [C ] // Proceedings of the 2015 IEEE International Conference on Computer Vision (ICCV) . Piscataway : IEEE Press , 2015 : 4534 - 4542 .
Gao Z , Tan C , Wu L , et al . SimVP: simpler yet better video prediction [C ] // Proceedings of the 2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) . Piscataway : IEEE Press , 2022 : 3170 - 3180 .
Papineni K , Roukos S , Ward T , et al . Bleu: a method for automatic evaluation of machine translation [C ] // Proceedings of the 40th Annual Meeting of the Association for Computational Linguistics . 2002 : 311 - 318 .
Jiang F , Tu S , Dong L , et al . Large generative model-assisted talking-face semantic communication system [J ] . IEEE Journal on Selected Areas in Communications , 2025 .