1.湖南工商大学计算机学院,湖南 长沙 410205
2.湖南师范大学信息科学与工程学院,湖南 长沙 410081
3.英国伦敦布鲁内尔大学计算机科学系,英国 伦敦 UB8 3PH
4.东南大学移动通信国家重点实验室,中国 南京 210096
收稿:2026-06-27,
修回:2026-08-25,
录用:2026-08-27,
移动端阅览
董莉, 邹健, 江沸菠, 等. 视觉语言模型和世界模型赋能的视频语义通信[J/OL]. 通信学报, 2026.
Dong Li, Zou Jian, Jiang Feibo, et al. Vision-Language Model- and World Model-Empowered Video Semantic Communication[J/OL]. Journal on Communications, 2026.
董莉, 邹健, 江沸菠, 等. 视觉语言模型和世界模型赋能的视频语义通信[J/OL]. 通信学报, 2026. DOI: 10.11959/j.issn.1000-436x.TXXB260376.
Dong Li, Zou Jian, Jiang Feibo, et al. Vision-Language Model- and World Model-Empowered Video Semantic Communication[J/OL]. Journal on Communications, 2026. DOI: 10.11959/j.issn.1000-436x.TXXB260376.
智能视频业务的快速发展对无线视频传输的效率、鲁棒性和语义保真度提出了更高要求。然而,现有视频传输仍面临语义冗余高、生成式重构失真以及数据分布易偏移等问题。为此,本文提出一种视觉语言模型和世界模型赋能的多任务视频语义通信系统。在发送端,视觉语言模型将原始视频压缩为关键帧与文本语义,以减少冗余传输并保留核心内容。在接收端,世界模型基于接收的关键帧和文本语义重建高质量视频。针对动态环境中的任务数据分布偏移问题,本文进一步提出持续语义适应机制,通过跨任务优化对语义/信道编解码器进行稳定更新,从而保持跨任务语义理解能力。多种视频任务和信道条件下的实验结果表明,所提出的视频语义通信系统具有较高的语义一致性、稳定性和泛化能力。
The rapid development of intelligent video services has imposed increasingly stringent requirements on the efficiency
robustness
and semantic fidelity of wireless video transmission. However
existing video transmission systems still face several challenges
including substantial semantic redundancy
difficulties in generative reconstruction
and vulnerability to data distribution shifts. To address these challenges
a multi-task video semantic communication system empowered by a vision-language model (VLM) and a world model (WM) was proposed. At the transmitter
the original video was compressed by the VLM into keyframes and textual semantic descriptions
such that transmission redundancy was reduced while essential semantic content was preserved. At the receiver
high-quality videos were reconstructed by the WM based on the received keyframes and textual semantic descriptions. To mitigate task-specific data distribution shifts in dynamic environments
a continual semantic adaptation mechanism was further developed
through which the semantic and channel encoders and decoders were stably updated using cross-task optimization to preserve semantic understanding across different tasks. High semantic consistency
stability
and generalization capability under different video tasks and channel conditions were demonstrated by the experimental results.
张平 , 许晓东 , 牛凯 等 . 现代语义通信与6G智简网络理论技术体系 [J ] . 北京邮电大学学报 , 2025 , 48 ( 5 ): 1 – 16 .
Zhang P , Xu X D , Niu K , et al . Theoretical and technical system of modern semantic communications and 6G intelligent-simple networks [J ] . Journal of Beijing University of Posts and Telecommunications , 2025 , 48 ( 5 ): 1 - 16 .
Nan G , Li Z , Zhai J , et al . Physical-layer adversarial robustness for deep learning-based semantic communications [J ] . IEEE Journal on Selected Areas in Communications , 2023 , 41 ( 8 ): 2592 – 2608 .
Zhang P , Xu W , Liu Y , et al . Intellicise wireless networks from semantic communications: A survey, research issues, and challenges [J ] . IEEE Communications Surveys & Tutorials , 2024 , 27 ( 3 ): 2051 – 2084 .
张平 , 牛凯 , 姚圣时 等 . 面向未来的语义通信:基本原理与实现方法 [J ] . 通信学报 , 2023 , 44 ( 5 ): 1 – 14 .
Zhang P , Niu K , Yao S S , et al . Semantic communications for future: basic principle and implementation methodology [J ] . Journal on Communications , 2023 , 44 ( 5 ): 1 - 14 .
Cui Q , You X , Wei N , et al . Overview of AI and communication for 6G network: Fundamentals, challenges, and future research opportunities [J ] . Science China Information Sciences , 2025 , 68 ( 7 ): 171301 .
Jiang F , Tu S , Dong L , et al . FlashSAM: Lightweight Vision Model for Multi-UAV Token Communication in Low-Altitude Wireless Networks [J ] . IEEE Journal of Selected Topics in Signal Processing , 2026 .
Kirkpatrick J , Pascanu R , Rabinowitz N , et al . Overcoming catastrophic forgetting in neural networks [J ] . Proceedings of the national academy of sciences , 2017 , 114 ( 13 ): 3521 – 3526 .
Jiang F , Tang C , Dong L , et al . Visual language model based cross-modal semantic communication systems [J ] . IEEE Transactions on Wireless Communications , 2025 , 24 ( 5 ): 3937 – 3948 .
Zhao Y , Yue Y , Hou S , et al . LaMoSC: Large language model-driven semantic communication system for visual transmission [J ] . IEEE Transactions on Cognitive Communications and Networking , 2024 , 10 ( 6 ): 2005 – 2018 .
Liang S D . Vision language models for massive MIMO semantic communication [C ] // Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops . 2025 : 1675 - 1685 .
Jiang P , Guo J , Wen C-K , et al . Semantic Communications with World Models [J ] . `arXiv preprint arXiv: 2510.24785 , 2025 .
Lokumarambage M , Sivalingam T , Dong F , et al . Latent Prediction based Generative Semantic Communication for Video Transmission in Wireless Networks [J ] . IEEE Open Journal of the Communications Society , 2026 , 7 : 3974 – 3986 .
Jiang F , Tu S , Dong L , et al . Tokencom: Vision-language model for multimodal and multitask token communications [J ] . arXiv preprint arXiv: 260300482 , 2026 .
Bertasius G , Wang H , Torresani L . Is space-time attention all you need for video understanding? [C ] // Proceedings of the 38th International Conference on Machine Learning . 2021 : 813 - 824 .
Peebles W , Xie S . Scalable diffusion models with transformers [C ] // Proceedings of the IEEE/CVF International Conference on Computer Vision . 2023 : 4195 - 4205 .
Zhang H , Lei Y , Gui L , et al . CPPO: continual learning for reinforcement learning with human feedback [C ] // Proceedings of the 12th International Conference on Learning Representations (ICLR) . Vienna : OpenReview , 2024 .
Soomro K , Zamir A R , Shah M . Ucf101: A dataset of 101 human actions classes from videos in the wild [J ] . arXiv preprint arXiv: 1212.0402 , 2012 .
Carreira J , Noland E , Hillier C , et al . A short note on the kinetics-700 human action dataset [J ] . arXiv preprint arXiv: 1907.06987 , 2019 .
Sigurdsson G A , Varol G , Wang X , et al . Hollywood in homes: Crowdsourcing data collection for activity understanding [C ] // European Conference on Computer Vision . 2016 : 510 - 526 .
Wang S , Dai J , Liang Z , et al . Wireless deep video semantic transmission [J ] . IEEE Journal on Selected Areas in Communications , 2023 , 41 ( 1 ): 214 - 229 .
Yin H , Qiao L , Ma Y , et al . Generative video semantic communication via multimodal semantic fusion with large model [J ] . IEEE Transactions on Vehicular Technology , 2025 .
Wang Y , Yu S , Yang X . Error robustness scheme for H.264 based on LDPC code [C ] // 2006 12th International Multi-Media Modelling Conference . 2006 : 79 .
Maaz M , Rasheed H , Khan S , et al . Videogpt+: Integrating image and video encoders for enhanced video understanding [J ] . arXiv preprint arXiv: 2406.09418 , 2024 .
Wan T , Wang A , Ai B , et al . Wan: Open and advanced large-scale video generative models [J ] . arXiv preprint arXiv: 2503.20314 , 2025 .
Li K , Wang Y , He Y , et al . MVBench: A comprehensive multi-modal video understanding benchmark [C ] // Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition . 2024 : 22195 - 22206 .
Cheng Z , Leng S , Zhang H , et al . Videollama 2: Advancing spatial-temporal modeling and audio understanding in video-llms [J ] . arXiv preprint arXiv: 2406.07476 , 2024 .
Yang Z , Teng J , Zheng W , et al . CogVideoX: Text-to-video diffusion models with an expert transformer [C ] // International Conference on Learning Representations . 2025 .
Hessel J , Holtzman A , Forbes M , et al . CLIPScore: A reference-free evaluation metric for image captioning [C ] // Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing . 2021 : 7514 - 7528 .
Venugopalan S , Rohrbach M , Donahue J , et al . Sequence to sequence-video to text [C ] // Proceedings of the 2015 IEEE International Conference on Computer Vision (ICCV) . Piscataway : IEEE Press , 2015 : 4534 - 4542 .
Gao Z , Tan C , Wu L , et al . SimVP: simpler yet better video prediction [C ] // Proceedings of the 2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) . Piscataway : IEEE Press , 2022 : 3170 - 3180 .
Papineni K , Roukos S , Ward T , et al . Bleu: a method for automatic evaluation of machine translation [C ] // Proceedings of the 40th Annual Meeting of the Association for Computational Linguistics . 2002 : 311 - 318 .
Jiang F , Tu S , Dong L , et al . Large generative model-assisted talking-face semantic communication system [J ] . IEEE Journal on Selected Areas in Communications , 2025 .
0
浏览量
0
下载量
0
CSCD
关联资源
相关文章
相关作者
相关机构
京公网安备11010602201714号