
- Home
高级 检索
Chinese
English



中国人民解放军国防科技大学 电子科学学院, 长沙 410028
Revised:2026-06-25,
Accepted:30 June 2026,
移动端阅览
LI Jiale, HUI Bingwei, WANG Kaiyue, et al. Zero-Shot Infrared Object Detection via Frequency-Domain Feature Decoupling of Visible Images[J/OL]. Acta Photonica Sinica, 2026, gz26-0175
LI Jiale, HUI Bingwei, WANG Kaiyue, et al. Zero-Shot Infrared Object Detection via Frequency-Domain Feature Decoupling of Visible Images[J/OL]. Acta Photonica Sinica, 2026, gz26-0175 DOI: 10.3788/gzxb20265508.0810001. CSTR: 32255.14.gzxb20265508.0810001.
针对典型应用场景下红外图像训练样本匮乏导致的零样本检测难题,提出一种基于可见光图像频域特征解耦的零样本红外目标检测方法。首先,设计频域处理模块,通过削弱高频纹理以降低源域特有风格干扰,并对低频分量实施扰动以提升源域样本多样性,从而为跨模态域不变特征学习提供先验约束。其次,构建权重共享的双分支特征提取框架,并在浅层嵌入特征解耦模块,以显式分离域不变特征与域私有特征,增强模型对跨波段共性表征的建模能力。最后,在颈部网络中引入注意力引导的特征拼接策略,优化多尺度信息融合过程,提升小目标感知与定位性能。在DroneVehicle和FLIR公开数据集上的实验结果表明,本方法与最新的单域泛化目标检测算法相比,mAP@50分别提升了2.9%和5%,验证了所提方法在跨波段零目标域样本条件下的有效性与鲁棒性,为单域泛化目标检测提供了新的解决思路。
Infrared imaging systems are widely used in battlefield reconnaissance, disaster warning, autonomous driving, and industrial security because of their all-weather sensing capability and robustness to illumination variations. However, in non-cooperative environments, infrared samples and annotations are often difficult or costly to obtain, making zero-shot infrared object detection highly challenging. Traditional domain adaptation and fine-tuning methods usually require target-domain samples, which limits their applicability in infrared scenarios. Single-domain generalization (SDG), which learns from only one source domain and does not rely on target-domain data, provides a feasible solution. Existing SDG methods mainly focus on generic visual feature generalization or cross-modal semantic alignment, but they rarely exploit the frequency characteristics of infrared imaging and are often insufficient for small-target detection under cross-modal domain shifts. Moreover, few SDG methods are specifically designed for the lightweight YOLO architecture required by real-time applications.To address the above challenges, this paper proposes a zero-shot infrared object detection method based on visible-image frequency-domain feature decoupling. First, a High-Frequency Suppression and Low-Frequency Enhancement (HSLE) module is designed to process visible images in the frequency domain. High-frequency components are adaptively suppressed to reduce source-domain-specific texture interference, while low-frequency statistics are perturbed to increase source-domain diversity and to narrow the visible-to-infrared gap without severely damaging semantic structure. During training, a weight-shared dual-branch backbone is adopted, where the original visible image is used as the primary branch and the HSLE-enhanced image is used as the auxiliary branch. A Feature Decoupling (FD) module is embedded in shallow backbone stages to separate domain-invariant features from domain-private features, so that deeper layers focus more on cross-spectral common semantics rather than source-style cues. To reduce the loss of weak target information, conventional downsampling is replaced by Space-to-Depth Convolution (SPD_Conv), which preserves more fine-grained details during feature transformation. In addition, a Guided Attention Concatenation (GA_Concat) module is introduced in the neck to enhance multi-scale feature fusion by using shallow detail-rich features to guide deep semantic aggregation. During inference, the auxiliary branch is removed, and the detector remains a single-branch structure, thereby preserving deployment efficiency.Experiments are conducted on the DroneVehicle dataset under a strict zero-shot visible-to-infrared protocol, where only labeled visible images are used for training and infrared images are used only for validation. To ensure a rigorous evaluation of scene-level generalization, the visible training set and infrared validation set are constructed with scene-level non-overlap. On the infrared validation set, the proposed method achieves an mAP@50 of 87.2%, which is 17.4 percentage points higher than YOLOv8n. Compared with SW, IBN-Net, ISW, CDSD, G-NAS, OA-DG, and PDOC, the proposed method improves mAP@50 by 6.8, 10.7, 13.7, 4.6, 8.6, 2.9, and 7.2 percentage points, respectively. Meanwhile, the detector remains lightweight, with only 3.7 M parameters and 14.1 GFLOPs.To further evaluate multi-view infrared generalization, experiments are also conducted on the FLIR dataset. The proposed method achieves an mAP@50 of 72.1%, outperforming YOLOv8n by 17.5 percentage points and surpassing SW, IBN-Net, ISW, CDSD, G-NAS, OA-DG, and PDOC by 16.1, 13.6, 12.5, 11.0, 6.3, 5.0, and 5.8 percentage points, respectively. Ablation studies verify the effectiveness of HSLE, FD, and the small-object detection components. Qualitative results further show that the proposed method reduces missed detections and false alarms in cluttered infrared scenes. These results demonstrate that the proposed method improves zero-shot visible-to-infrared object detection and small infrared target detection while maintaining the efficiency of a lightweight one-stage detector.
ZHENG Guangtao , HUAI Mengdi , ZHANG Aidong , et al . AdvST: revisiting data augmentations for single domain generalization [C]. Proceedings of the AAAI Conference on Artificial Intelligence. Palo Alto : AAAI Press , 2024 , 38 ( 19 ): 21832 - 21840 .
SU Zixian , YAO Kai , YANG Xi , et al . Rethinking data augmentation for single-source domain generalization in medical image segmentation [C]. Proceedings of the AAAI Conference on Artificial Intelligence. Palo Alto : AAAI Press , 2023 , 37 ( 2 ): 2366 - 2374 .
SHI Caijuan , ZHENG Yuanfan , REN Bijuan , et al . Single-domain generalized breast tumor detection in X-ray images [J]. Journal of Image and Graphics , 2024 , 29 ( 3 ): 725 - 740 .
史彩娟 , 郑远帆 , 任弼娟 , 等 . 单域泛化X-ray乳腺肿瘤检测 [J]. 中国图象图形学报 , 2024 , 29 ( 3 ): 725 - 740 .
ZHANG Chen , LI Yunping , TANG Xin , et al . Review of domain generalization person re-identification based on deep learning [J]. Journal of Kunming University of Science and Technology (Natural Science) , 2024 , 49 ( 6 ): 86 - 99 .
张臣 , 李云平 , 唐鑫 , 等 . 基于深度学习的域泛化行人重识别综述 [J]. 昆明理工大学学报(自然科学版) , 2024 , 49 ( 6 ): 86 - 99 .
JIN Xin , LAN Cuiling , ZENG Wenjun , et al . Style normalization and restitution for generalizable person re-identification [C]. Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition . Piscataway : IEEE , 2020 : 3143 - 3152 .
ZHENG Kecheng , LIU Jiawei , WU Wei , et al . Calibrated feature decomposition for generalizable person re-identification [J/OL]. arXiv : 2111 . 13945 , ( 2021-11-27 )[ 2026-06-24 ]. https://doi.org/10.48550/arXiv.2111.13945 https://doi.org/10.48550/arXiv.2111.13945 .
LIU Jiawei , HUANG Zhipeng , LI Liang , et al . Debiased batch normalization via gaussian process for generalizable person re-identification [C]. Proceedings of the AAAI Conference on Artificial Intelligence. Palo Alto : AAAI Press , 2022 , 36 ( 2 ): 1729 - 1737 .
WU Aming , DENG Cheng . Single-domain generalized object detection in urban scene via cyclic-disentangled self-distillation [C]. Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition . Piscataway : IEEE , 2022 : 847 - 856 .
LIN Chuang , YUAN Zehuan , ZHAO Sicheng , et al . Domain-invariant disentangled network for generalizable object detection [C]. Proceedings of the IEEE International Conference on Computer Vision . Piscataway : IEEE , 2021 : 8771 - 8780 .
YU Shijie , ZHU Feng , CHEN Dapeng , et al . Multiple domain experts collaborative learning: multi-source domain generalization for person re-identification [J/OL]. arXiv : 2105 . 12355 , ( 2021-05-26 )[ 2026-06-24 ]. https://doi.org/10.48550/arXiv.2105.12355 https://doi.org/10.48550/arXiv.2105.12355 .
LI Da , YANG Yongxin , SONG Yizhe , et al . Learning to generalize: meta-learning for domain generalization [C]. Proceedings of the AAAI Conference on Artificial Intelligence. Palo Alto : AAAI Press , 2018 , 32 ( 1 ): 3490 - 3497 .
ZHAO Yuyang , ZHONG Zhun , YANG Fengxiang , et al . Learning to generalize unseen domains via memory-based multi-source meta-learning for person re-identification [C]. Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition . Piscataway : IEEE , 2021 : 6273 - 6282 .
ZHAO Yuyang , ZHONG Zhun , ZHAO Na , et al . Style-hallucinated dual consistency learning: a unified framework for visual domain generalization [J]. International Journal of Computer Vision , 2024 , 132 ( 3 ): 837 - 853 .
QI Lei , DONG Peng , XIONG Tan , et al . DoubleAUG: single-domain generalized object detector in urban via color perturbation and dual-style memory [J]. ACM Transactions on Multimedia Computing, Communications, and Applications , 2024 , 20 ( 5 ): 1 - 20 .
FAN Qi , SEGU M , TAI Y W , et al . Towards robust object detection invariant to real-world domain shifts [C]. The Eleventh International Conference on Learning Representations . Kigali : OpenReview , 2023 .
LIAO Shengcai , SHAO Ling . Interpretable and generalizable person re-identification with query-adaptive convolution and temporal lifting [J/OL]. arXiv : 1904 . 10424 , ( 2019-04-23 )[ 2026-06-24 ]. https://doi.org/10.48550/arXiv.1904.10424 https://doi.org/10.48550/arXiv.1904.10424 .
VIDIT V , ENGILBERGE M , SALZMANN M . CLIP the gap: a single domain generalization approach for object detection [C]. Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition . Piscataway : IEEE , 2023 : 3219 - 3229 .
DENG Li , WU Aming , WANG Yaowei , et al . Prompt-driven dynamic object-centric learning for single domain generalization [C]. Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition . Piscataway : IEEE , 2024 : 17606 - 17615 .
LI Hao , WANG Wei , WANG Cong , et al . Phrase grounding-based style transfer for single-domain generalized object detection [J]. IEEE Transactions on Circuits and Systems for Video Technology , 2026 , 36 ( 1 ): 106 - 118 .
SUNKARA R , LUO Tie . No more strided convolutions or pooling: a new CNN building block for low-resolution images and small objects [C]. Machine Learning and Knowledge Discovery in Databases . Cham : Springer Nature Switzerland , 2023 : 443 - 459 .
TONG Yunfei , LIU Jing , FU Zhiling , et al . Guided attention and joint loss for infrared dim small target detection [J]. IEEE Transactions on Geoscience and Remote Sensing , 2024 , 62 : 1 - 14 .
HWANG S , HAN D , JEON M . DG-DETR: toward domain generalized detection transformer [J]. Pattern Recognition Letters , 2026 , 199 : 128 - 134 .
JEON S , HONG K , LEE P , et al . Feature stylization and domain-aware contrastive learning for domain generalization [C]. Proceedings of the 29th ACM International Conference on Multimedia . New York : Association for Computing Machinery , 2021 : 22 - 31 .
GUO Jintao , WANG Na , QI Lei , et al . ALOFT: a lightweight MLP-like architecture with dynamic low-frequency transform for domain generalization [C]. Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition . Piscataway : IEEE , 2023 : 24132 - 24141 .
YOO J , UH Y , CHUN S , et al . Photorealistic style transfer via wavelet transforms [C]. Proceedings of the IEEE/CVF International Conference on Computer Vision . Piscataway : IEEE , 2019 : 9035 - 9044 .
YUAN Yuxuan , TANG Luyao , XU Ying , et al . Filling and disentanglement: toward low- and high-order parallel single-domain generalization for SAR ship detection [J]. IEEE Transactions on Aerospace and Electronic Systems , 2025 , 61 ( 2 ): 3668 - 3682 .
HOU Qibin , ZHOU Daquan , FENG Jiashi . Coordinate attention for efficient mobile network design [C]. Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition . Piscataway : IEEE , 2021 : 13713 - 13722 .
HU Jie , SHEN Li , SUN Gang . Squeeze-and-excitation networks [C]. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition . Piscataway : IEEE , 2018 : 7132 - 7141 .
PAN Xingang , ZHAN Xiaohang , SHI Jianping , et al . Switchable whitening for deep representation learning [C]. Proceedings of the IEEE/CVF International Conference on Computer Vision . Piscataway : IEEE , 2019 : 1863 - 1871 .
PAN Xingang , LUO Ping , SHI Jianping , et al . Two at once: enhancing learning and generalization capacities via IBN-Net [C]. Proceedings of the European Conference on Computer Vision . Cham : Springer , 2018 : 464 - 479 .
CHOI S , JUNG S , YUN H , et al . RobustNet: improving domain generalization in urban-scene segmentation via instance selective whitening [C]. Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition . Piscataway : IEEE , 2021 : 11575 - 11585 .
WU Fan , GAO Jinling , HONG Lanqing , et al . G-NAS: generalizable neural architecture search for single domain generalization object detection [C]. Proceedings of the AAAI Conference on Artificial Intelligence. Palo Alto : AAAI Press , 2024 , 38 ( 6 ): 5958 - 5966 .
LEE W , HONG D , LIM H , et al . Object-aware domain generalization for object detection [C]. Proceedings of the AAAI Conference on Artificial Intelligence. Palo Alto : AAAI Press , 2024 , 38 ( 4 ): 2947 - 2955 .
0
Views
0
下载量
0
CSCD
Publicity Resources
Related Articles
Related Author
Related Institution
京公网安备11010602201714号