
- Home
高级 检索
Chinese
English



1.电子科技大学 航空航天学院,成都 611731
2.自适应光学全国重点实验室,成都 610209
3.飞行器集群智能感知与协同控制四川省重点实验室,成都 611731
Received:31 January 2026,
Revised:2026-07-08,
Accepted:13 July 2026,
移动端阅览
刘羽航,李文博,杨轲,等. 双光图像融合与目标检测并行网络及轻量化部署[J].光子学报,2026,55(8):0810002
LIU Yuhang, LI Wenbo, YANG Ke, et al. Parallel Network for Dual-band Image Fusion and Object Detection with Lightweight Deployment[J]. Acta Photonica Sinica, 2026, 55(8):0810002
刘羽航,李文博,杨轲,等. 双光图像融合与目标检测并行网络及轻量化部署[J].光子学报,2026,55(8):0810002 DOI: 10.3788/gzxb20265508.0810002. CSTR: 32255.14.gzxb20265508.0810002.
LIU Yuhang, LI Wenbo, YANG Ke, et al. Parallel Network for Dual-band Image Fusion and Object Detection with Lightweight Deployment[J]. Acta Photonica Sinica, 2026, 55(8):0810002 DOI: 10.3788/gzxb20265508.0810002. CSTR: 32255.14.gzxb20265508.0810002.
双光图像融合与目标检测因其在遥感、安防等领域的应用价值受到广泛关注,然而现有方法多采用串行级联结构,存在顺序依赖、难以平衡任务间关系、计算复杂度高等问题。提出一种并行网络实现双光图像融合与目标检测的多任务学习与轻量化部署。首先,设计伪孪生特征提取器与通道注意力特征融合模块,实现任务适配的高效特征提取;其次,引入跨尺度跳跃连接以增强融合细节表征,并采用自适应任务损失权重缓解任务间的优化冲突;最后,通过结构化剪枝与混合量化对网络进行压缩,得到适用于边缘设备的轻量化模型。在公开数据集M³FD上的对比实验表明,所提方法在保持较高检测与融合性能的同时,显著降低了参数量与计算开销。嵌入式平台上的部署结果进一步验证了该模型在效率与精度间的良好平衡,体现了其在实际应用中的有效性与实用性。
The purpose of this study is to address the limitations of cascade-based dual light object detection frameworks by developing a parallel multi-task learning network capable of jointly performing dual-band image fusion and object detection. Instead of enforcing a strict sequential dependency between fusion and detection tasks, the proposed approach aims to achieve collaborative optimization through shared feature learning while preserving task-specific representations. By mitigating the common error propagation and task interference in traditional cascade architectures, this study aims to leverage the advantages of dual-optical feature complementarity to improve detection accuracy and fusion quality in complex lighting scenarios, while maintaining deployability on resource-constrained edge computing platforms.To accomplish the stated objective, a novel parallel multi-task network architecture is designed, in which dual light image fusion and object detection are simultaneously learned within a unified framework. Unlike conventional cascade-based approaches, the proposed architecture enables both tasks to share intermediate representations while retaining task-specific feature pathways, thereby reducing error accumulation caused by sequential processing. A pseudo-twin feature extraction module is first introduced to process RGB and thermal inputs separately, enabling the network to capture modality-specific characteristics while maintaining structural consistency between feature streams. This design ensures that each modality preserves its inherent representational advantages during early-stage feature encoding. To further enhance cross-modal representation learning, a channel-attention-based feature fusion module is incorporated, allowing the network to adaptively emphasize informative channels and suppress redundant or noisy features across modalities. Through learned channel-wise weighting, the fusion module facilitates effective inter-modal interaction and improves the discriminative capability of the shared feature space. Considering the heterogeneous feature requirements of the two tasks, the proposed network explicitly differentiates between shared and task-oriented feature representations. The shared backbone focuses on extracting robust cross-modal representations, while dedicated task heads are designed to accommodate the specific objectives of image fusion and object detection. For the image fusion branch, cross-scale skip connections are employed to preserve fine-grained spatial details and high-frequency structural information that are critical for generating perceptually informative fused images. These connections enable low-level texture features to be effectively propagated and integrated with higher-level semantic representations, thus enhancing visual quality and structural consistency in the fused outputs. In parallel, the object detection branch prioritizes higher-level semantic features and multi-scale contextual information to ensure robust object localization and classification under varying illumination and environmental conditions. To mitigate potential optimization conflicts arising from joint learning, an adaptive task loss weighting strategy is further designed. Instead of relying on fixed loss coefficients, this mechanism dynamically adjusts the relative contributions of fusion and detection losses during training based on task learning dynamics. By continuously balancing the optimization objectives of both tasks, the proposed strategy improves convergence stability, prevents dominance of a single task, and promotes coordinated multi-task learning throughout the training process.Comprehensive qualitative and quantitative experiments are conducted on the M³FD dual light dataset to evaluate the effectiveness of the proposed parallel multi-task learning framework. Experimental results demonstrate that the proposed method consistently outperforms baseline cascade-based models in both object detection accuracy and image fusion quality. In terms of detection performance, the proposed network achieves a relative improvement of 2.8% on visible-light images and 4.2% on thermal images compared with the baseline method. Qualitative evaluations further indicate that the fused images generated by the proposed approach exhibit clearer structural contours and enhanced target saliency, which are beneficial for downstream detection tasks. Furthermore, to reduce model complexity and improve deployment efficiency, structured pruning is applied to the trained network, followed by sufficient retraining to obtain a lightweight model. Experimental results show that the pruned model retains only 52.1% of the original parameters, while reducing the GPU inference latency from 22.9 ms to 16.3 ms, with a detection accuracy drop of only 1.7%. Compared with the baseline YOLOv5s model, the proposed lightweight network achieves a detection accuracy improvement of 2.4% while reducing the parameter count by approximately 4.5 million, demonstrating its superior compactness and efficiency. In addition to algorithmic performance, deployment experiments are carried out on an edge computing platform equipped with the RK3588 chip to assess real-world applicability. The results show that the proposed model can be successfully deployed and executed under limited computational and memory resources, while maintaining satisfactory detection accuracy.These research findings validate that the proposed parallel multi-task architecture achieves a good balance between performance and computational efficiency, making it suitable for practical complex lighting scenarios. The proposed parallel multi-task learning network effectively realizes the joint optimization of dual light image fusion and object detection, demonstrating superior performance and deployment feasibility on resource-constrained edge platforms.
SONG Kechen , ZHAO Ying , HUANG Liming , et al . RGB-T image analysis technology and application: A survey [J]. Engineering Applications of Artificial Intelligence , 2023 , 120 : 105919 .
GUO Cuixia , XU Yongtao , ZOU Zhanghuang , et al . Lightweight pedestrian vehicle detection algorithm based on visible and infrared bimodal fusion [J]. Acta Photonica Sinica , 2025 , 54 ( 6 ): 0610001 .
郭翠霞 , 徐永涛 , 邹章煌 , 等 . 基于可见和红外双模态融合的轻量化行人车辆检测算法 [J]. 光子学报 , 2025 , 54 ( 6 ): 0610001 .
YANG Chen , HOU Zhiqiang , LI Xinyue , et al . Object detection algorithm based on CNN-transformer dual modal feature fusion [J]. Acta Photonica Sinica , 2024 , 53 ( 3 ): 0310001 .
杨晨 , 侯志强 , 李新月 , 等 . 基于CNN-Transformer双模态特征融合的目标检测算法 [J]. 光子学报 , 2024 , 53 ( 3 ): 0310001 .
YU Zhirui , YIN Zhanpeng , WANG Junyu , et al . Multimodal object detection method using adaptive dusion of infrared and visible features [J]. National Remote Sensing Bulletin , 2025 , 29 ( 10 ): 3006 - 3019 .
喻智睿 , 尹展鹏 , 王俊宇 , 等 . 可见光和红外特征自适应融合的多模态目标检测方法 [J]. 遥感学报 , 2025 , 29 ( 10 ): 3006 - 3019 .
SUN Bin , YOU Hang , LI Wenbo , et al . Dual-band payload image fusion and its applications in low-altitude remote sensing [J]. Acta Aeronautica et Astronautica Sinica , 2025 , 46 ( 11 ): 531343 .
孙彬 , 游航 , 李文博 , 等 . 双光载荷图像融合及其在低空遥感中的应用 [J]. 航空学报 , 2025 , 46 ( 11 ): 531343 .
HAO Yongping , CAO Zhaorui , BAI Fan , et al . Research on infrared visible image fusion and target recognition algorithm based on region of interest mask convolution neural network [J]. Acta Photonica Sinica , 2021 , 50 ( 2 ): 0210002 .
郝永平 , 曹昭睿 , 白帆 , 等 . 基于兴趣区域掩码卷积神经网络的红外-可见光图像融合与目标识别算法研究 [J]. 光子学报 , 2021 , 50 ( 2 ): 0210002 .
ZHANG Xiaodong , WANG Shuo , GAO Shaoshu , et al . Infrared and visible image fusion method based on information enhancement and mask loss [J]. Acta Photonica Sinica , 2024 , 53 ( 9 ): 0910003 .
张晓东 , 王硕 , 高绍姝 , 等 . 基于信息增强和掩码损失的红外与可见光图像融合方法 [J]. 光子学报 , 2024 , 53 ( 9 ): 0910003 .
TANG Linfeng , YUAN Jiteng , MA Jiayi . Image fusion in the loop of high-level vision tasks: a semantic-aware real-time infrared and visible image fusion network [J]. Information Fusion , 2022 , 82 : 28 - 42 .
LIU Jinyuan , FAN Xin , HUANG Zhanbo , et al . Target-aware dual adversarial learning and a multi-scenario multi-modality benchmark to fuse infrared and visible for object detection [C]. Proceedings of the IEEE/CVF conference on computer vision and pattern recognition , 2022 : 5792 - 5801 .
SUN Yiming , CAO Bing , ZHU Pengfei , et al . Detfusion: a detection-driven infrared and visible image fusion network [C]. Proceedings of the 30th ACM international conference on multimedia , 2022 : 4003 - 4011 .
ZHAO Wenda , XIE Shigeng , ZHAO Fan , et al . Metafusion: infrared and visible image fusion via meta-feature embedding from object detection [C]. Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2023 : 13955 - 13965 .
WU Yuhui , LIU Zhu , LIU Jinyuan , et al . Breaking free from fusion rule: A fully semantic-driven infrared and visible image fusion [J]. IEEE Signal Processing Letters , 2023 , 30 : 418 - 422 .
LIU Xiaowen , HUO Hongtao , LI Jing , et al . A semantic-driven coupled network for infrared and visible image fusion [J]. Information Fusion , 2024 , 108 : 102352 .
TANG Linfeng , ZHANG Hao , XU Han , et al . Rethinking the necessity of image fusion in high-level vision tasks: a practical infrared and visible image fusion network based on progressive semantic injection and scene fidelity [J]. Information Fusion , 2023 , 99 : 101870 .
CHEN Xiaoxuan , XU Shuwen , HU Shaohai , et al . UIFD: a unified interactive network for image fusion and traffic object detection under low-light conditions [J]. IEEE Transactions on Intelligent Transportation Systems , 2025 , 26 ( 11 ): 19182 - 19196 .
YANG Zengyi , ZHANG Yafei , LI Huafeng , et al . Instruction-driven fusion of Infrared–visible images: Tailoring for diverse downstream tasks [J]. Information Fusion , 2025 , 121 : 103148 .
BAKIRCI M . Performance evaluation of low-power and lightweight object detectors for real-time monitoring in resource-constrained drone systems [J]. Engineering Applications of Artificial Intelligence , 2025 , 159 : 111775 .
DENG Jian , LI Mingyue , CHEN Yongxin , et al . Cross-guided feature fusion with intra-modality reweighting for multi-spectral pedestrian detection [C]. 2022 26th International Conference on Pattern Recognition (ICPR) , IEEE , 2022 : 4864 - 4870 .
MA J , TANG L , XU M , et al . STDFusionNet: an infrared and visible image fusion network based on salient target detection [J]. IEEE Transactions on Instrumentation and Measurement , 2021 , 70 : 1 - 13 .
杨轲 . 可见光与红外图像融合和目标检测多任务学习研究 [D]. 成都 : 电子科技大学 , 2024 .
JOCHER G . YOLOv5 by Ultralytics [CP/OL]. San Francisco : Ultralytics , 2020 ( 2020-05 )[ 2026-01-31 ]. https://github.com/ultralytics/yolov5 https://github.com/ultralytics/yolov5 .
HU Jie , SHEN Li , SUN Gang . Squeeze-and-excitation networks [C]. Proceedings of the IEEE conference on computer vision and pattern recognition . 2018 : 7132 - 7141 .
HE Yang , LIU Ping , WANG Ziwei , et al . Filter pruning via geometric median for deep convolutional neural networks acceleration [C]. Proceedings of the IEEE/CVF conference on computer vision and pattern recognition , 2019 : 4340 - 4349 .
YANG C , LUO X , ZHANG Z , et al . KDFuse: a high-level vision task-driven infrared and visible image fusion method based on cross-domain knowledge distillation [J]. Information Fusion , 2025 , 118 : 102944 .
MA Jiayi , TANG Linfeng , FAN Fan , et al . SwinFusion: cross-domain long-range learning for general image fusion via swin transformer [J]. IEEE/CAA Journal of Automatica Sinica , 2022 , 9 ( 7 ): 1200 - 1217 .
WANG Zhou , BOVIK A C , SHEIKH H R , et al . Image quality assessment: from error visibility to structural similarity [J]. IEEE Transactions on Image Processing , 2004 , 13 ( 4 ): 600 - 612 .
GONG Ting , LEE T , STEPHENSON C , et al . A comparison of loss weighting strategies for multi task learning in deep neural networks [J]. IEEE Access , 2019 , 7 : 141627 - 141632 .
LIU Shikun , JOHNS E , DAVISON A J . End-to-end multi-task learning with attention [C]. Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2019 : 1871 - 1880 .
ZHANG Xingchen , DEMIRIS Y . Visible and infrared image fusion using deep learning [J]. IEEE Transactions on Pattern Analysis and Machine Intelligence , 2023 , 45 ( 8 ): 10535 - 10554 .
SUN Bin , GAO Yunxiang , ZHUGE Wuwei , et al . Analysis of quality objective assessment metrics for visible and infrared image fusion [J]. Journal of Image and Graphics , 2023 , 28 ( 1 ): 144 - 155 .
孙彬 , 高云翔 , 诸葛吴为 , 等 . 可见光与红外图像融合质量评价指标分析 [J]. 中国图象图形学报 , 2023 , 28 ( 1 ): 144 - 155 . DOI: 10.11834/jig.210719 http://dx.doi.org/10.11834/jig.210719
LI Hao , XU Zheng , TAYLOR G , et al . Visualizing the loss landscape of neural nets [J]. Advances in neural information processing systems , 2018 , 31 : 6389 – 6399 .
0
Views
0
下载量
0
CSCD
Publicity Resources
Related Articles
Related Author
Related Institution
京公网安备11010602201714号