
- Home
高级 检索
Chinese
English



1.宁夏大学信息工程学院,宁夏 银川 750021
2.宁夏“东数西算”人工智能与信息安全重点实验室,宁夏 银川 750021
3.宁夏大学 前沿交叉学院,宁夏 中卫 755000
Received:30 January 2026,
Revised:2026-05-11,
Accepted:30 June 2026,
移动端阅览
ZHAO Yongqing, ZHANG Ruonan, LIU Libo. Open Vocabulary Object Detection via Global-Local Context Synergy[J/OL]. Chinese Journal on Internet of Things, 2026.
ZHAO Yongqing, ZHANG Ruonan, LIU Libo. Open Vocabulary Object Detection via Global-Local Context Synergy[J/OL]. Chinese Journal on Internet of Things, 2026. DOI: 10.11959/j.issn.2096-3750.WLW26015.
在开放词汇目标检测中,传统的区域提议范式对候选区域进行独立的特征裁剪与提取,忽略了区域间潜在的依赖关系,导致特征表征中的上下文信息丢失。针对这一问题,本文提出一种全局-局部上下文协同网络(GLCS-Net)。首先,构建全局上下文聚合模块(GCA),通过引入自注意力机制聚合全局语义,以重构候选区域间的语义关联;其次,设计局部空间整合模块(LSI),利用动态邻域采样整合目标周边空间线索,以补全物体缺失的局部几何细节。最后,为了实现全局与局部上下文的协同,进一步提出上下文协同净化机制(CSOP)。该机制融合GCA和LSI增强后的特征,通过重新校准伪标签置信度来剔除噪声,从而构建可靠的监督信号以指导模型训练。实验结果表明,该方法在COCO数据集上,核心指标新类检测精度达到40.5%,较现有先进方法提升1.8%。在更具挑战性的细粒度LVIS数据集上,各项指标稀有类、常见类、频繁类及整体AP依次达到25.4%、32.9%、37.7%和33.5%,相较于此前的先进方法分别提升了0.8%、0.4%、2.1%和1.1%。结果证明GLCS-Net有效缓解了候选区域独立裁剪带来的信息瓶颈,为基于区域提议的检测框架提供了一种通用的特征增强范式。
In open vocabulary object detection
the traditional region proposal paradigm performs independent feature cropping and extraction on candidate regions
ignoring potential dependencies between regions
which leads to the loss of contextual information in feature representation. To address this issue
this paper proposes a Global-Local Context Synergy Network (GLCS-Net). First
a Global Context Aggregation (GCA) module is constructed
which aggregates global semantics by introducing a self-attention mechanism to reconstruct semantic dependencies between candidate regions. Second
a Local Spatial Integration (LSI) module is designed
which integrates spatial cues around the target by dynamic neighborhood sampling to complete the missing local geometric details of the object. Finally
to achieve the synergy between global and local contexts
a Context Synergy Purification (CSOP) mechanism is further proposed. This mechanism fuses the enhanced features from GCA and LSI
and eliminates noise by recalibrating the confidence of pseudo-labels
thereby constructing reliable supervision signals to guide model training. Experimental results show that on the COCO dataset
the core metric of novel class detection accuracy reaches 40.5%
which is 1.8% higher than the existing state-of-the-art methods. On the more challenging fine-grained LVIS dataset
the AP of rare
common
frequent categories and overall AP reach 25.4%
32.9%
37.7% and 33.5% respectively
which are 0.8%
0.4%
2.1% and 1.1% higher than the previous state-of-the-art methods. The results demonstrate that GLCS-Net effectively alleviates the information bottleneck caused by independent cropping of candidate regions
and provides a general feature enhancement paradigm for region-proposal-based detection frameworks.
ZAREIAN A , ROSA K D , HU D H , et al . Open-vocabulary object detection using captions [C ] // Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) . Nashville, TN, USA : IEEE , 2021 : 14393 - 14402 .
REN S Q , HE K M , GIRSHICK R , et al . Faster R-CNN: towards real-time object detection with region proposal networks [J ] . IEEE Transactions on Pattern Analysis and Machine Intelligence , 2017 , 39 ( 6 ): 1137 - 1149 .
HE K M , GKIOXARI G , DOLLÁR P , et al . Mask R-CNN [C ] // Proceedings of the IEEE International Conference on Computer Vision (ICCV) . Venice, Italy : IEEE , 2017 : 2980 - 2988 .
刘立波 , 郗思宇 , 邓箴 . 结合改进 ConvNeXt 网络与知识蒸馏的天气识别 [J ] . 光学精密工程 , 2023 , 31 ( 14 ): 2123 - 2134 .
LIU L B , XI S Y , DENG Z . Weather recognition combining improved ConvNeXt models with knowledge distillation [J ] . Opt. Precision Eng. , 2023 , 31 ( 14 ): 2123 - 2134 .
聂秀山 , 赵润虎 , 宁阳 , 等 . 开放词汇目标检测方法综述 [J ] . 山东大学学报(工学版) , 2025 , 55 ( 1 ): 1 - 14 .
NIE X S , ZHAO R H , NING Y , et al . Survey of open vocabulary object detection methods [J ] . Journal of Shandong University(Engineering Science) , 2025 , 55 ( 1 ): 1 - 14 .
RADFORD A , KIM J W , HALLACY C , et al . Learning transferable visual models from natural language supervision [C ] // Proceedings of the 38th International Conference on Machine Learning (ICML) . PMLR , 2021 : 8748 - 8763 .
GU X , LIN T Y , KUO W , et al . Open-vocabulary object detection via vision and language knowledge distillation [C ] // The Tenth International Conference on Learning Representations (ICLR) . 2022 : 1 - 21 .
CHEN Y , WANG C , LI Z H , et al . Enhancing open-vocabulary object detection through region-word and region-vision matching [J ] . Multimedia Systems , 2025 , 31 ( 3 ): 232 .
金友 , 邓箴 , 刘立波 . 知识嵌入引导的双分支融合增强开放词汇目标检测 [J ] . 光学精密工程 , 2025 , 33 ( 18 ): 2929 - 2943 .
JIN Y , DENG Z , LIU L B . Open-vocabulary object detection enhanced by knowledge embedding-guided dual-branch fusion [J ] . Opt. Precision Eng. , 2025 , 33 ( 18 ): 2929 - 2943 .
单飞龙 , 吕鹏远 , 李梦晨 . 基于CNN-Transformer半监督交叉学习的遥感图像场景分类方法 [J ] . 宁夏大学学报(自然科学版) , 2024 , 45 ( 3 ): 325 - 332 .
SHAN F L , LV P Y , LI M C . Remote sensing image scene classification based on semi-supervised cross learning of CNN and transformer [J ] . Journal of Ningxia University (Natural Science Edition) , 2024 , 45 ( 3 ): 325 - 332 .
ZHOU X , GIRDAR R , JOULIN A , et al . Detecting twenty-thousand classes using image-level supervision [C ] // European Conference on Computer Vision (ECCV) . Cham : Springer Nature Switzerland , 2022 : 350 - 368 .
ZHAO S , ZHANG Z , SCHULTER S , et al . Exploiting unlabeled data with vision and language models for object detection [C ] // European Conference on Computer Vision (ECCV) . Cham : Springer Nature Switzerland , 2022 : 159 - 175 .
WANG K , CHENG L C , CHEN W K , et al . MarvelOVD: Marrying object recognition and vision-language models for robust open-vocabulary object detection [C ] // European Conference on Computer Vision (ECCV) . Cham : Springer Nature Switzerland , 2024 : 106 - 122 .
TISHBY N , PEREIRA F C , BIALEK W . The information bottleneck method [J ] . arXiv preprint physics/0004057 , 2000 .
TISHBY N , ZASLAVSKY N . Deep learning and the information bottleneck principle [C ] // IEEE Information Theory Workshop (ITW) . Jerusalem, Israel : IEEE , 2015 : 1 - 5 .
SHWARTZ-ZIV R , TISHBY N . Opening the black box of deep neural networks via information [J ] . arXiv preprint arXiv: 1703.00810 , 2017 .
OLIVA A , TORRALBA A . The role of context in object recognition [J ] . Trends in Cognitive Sciences , 2007 , 11 ( 12 ): 520 - 527 .
WU H , SUN H , LIU K , et al . TSMamba: Multi-drone feature interaction via temporal-spatial mamba networks for aerial object tracking [J ] . Information Fusion , 2026 , 133 : 104279 .
WU H , SUN H , LIU K , et al . Temporal-Spatial Feature Interaction Network for Multi-Drone Multi-Object Tracking [J ] . IEEE Transactions on Circuits and Systems for Video Technology , 2025 , 35 ( 2 ): 1 - 14 .
伍瀚 , 孙浩 , 刘奎 , 等 . 无人机视频多目标特征关联技术研究进展 [J ] . 航空学报 , 2026 , 47 ( 4 ): 331967 .
WU H , SUN H , LIU K , et al . Multi-object feature association in UAV videos: Recent progress and perspectives [J ] . Acta Aeronautica et Astronautica Sinica , 2026 , 47 ( 4 ): 331967 .
伍瀚 , 孙浩 , 计科峰 , 等 . 时序信息引导跨视角特征融合的多无人机多目标跟踪方法 [J ] . 电子学报 , 2025 , 53 ( 3 ): 729 .
WU H , SUN H , JI K F , et al . Temporal-Guided Cross-View Feature Fusion Network for Multi-Drone Multi-Object Tracking [J ] . ACTA ELECTRONICA SINICA , 2025 , 53 ( 3 ): 729 .
LIU Y , WANG R , SHAN S , et al . Structure inference net: object detection using scene-level context and instance-level relationships [C ] // Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR) . Piscataway : IEEE , 2018 : 6985 - 6994 .
GIDARIS S , KOMODAKIS N . Inside-outside net: detecting objects in context with skip pooling and recurrent neural networks [C ] // Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR) . Piscataway : IEEE , 2016 : 2874 - 2883 .
HU H , GU J , ZHANG Z , et al . Relation networks for object detection [C ] // Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR) . Piscataway : IEEE , 2018 : 3588 - 3597 .
LIN T Y , MAIRE M , BELONGIE S , et al . Microsoft COCO: common objects in context [C ] // Proceedings of the European Conference on Computer Vision (ECCV) . Cham : Springer International Publishing , 2014 : 740 - 755 .
GUPTA A , DOLLÁR P , GIRSHICK R . LVIS: a dataset for large vocabulary instance segmentation [C ] // Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) . Long Beach, CA, USA : IEEE , 2019 : 5351 - 5359 .
WU Y , KIRILLOV A , MASSA F , et al . Detectron2 [EB/OL ] . https://github.com/facebookresearch/detectron2 https://github.com/facebookresearch/detectron2 , 2019 .
HE K , ZHANG X , REN S , et al . Deep residual learning for image recognition [C ] // Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR) . Piscataway : IEEE , 2016 : 770 - 778 .
LIN T Y , DOLLÁR P , GIRSHICK R , et al . Feature pyramid networks for object detection [C ] // Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR) . Piscataway : IEEE , 2017 : 2117 - 2125 .
ZHOU X , KOLTUN V , KRÄHENBÜHL P . Probabilistic two-stage detection [J ] . arXiv preprint arXiv: 2103.07461 , 2021 .
WU S , ZHANG W , JIN S , et al . Aligning bag of regions for open-vocabulary object detection [C ] // Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) . Piscataway : IEEE , 2023 : 15254 - 15264 .
YANG L X , CHEN D P , CHEN Y F , et al . A neuroinspired contrast mechanism enables few-shot object detection [J ] . Pattern Recognition , 2024 , 156 : 110766 .
ZHANG H L , GUAN D Y , KE X R , et al . Open-vocabulary object detection via debiased curriculum self-training [J ] . Expert Systems with Applications , 2024 , 255 : 124762 .
ZHONG Y W , YANG J W , ZHANG P C , et al . RegionCLIP: region-based language-image pretraining [C ] // Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) . New Orleans, LA, USA : IEEE , 2022 : 16772 - 16782 .
LI Y . Federated fine-grained prompts for vision-language models based on open-vocabulary object detection [J ] . Applied Intelligence , 2025 , 55 ( 7 ): 1 - 15 .
XU S L , LI X T , WU S Z , et al . DST-Det: open-vocabulary object detection via dynamic self-training [J ] . IEEE Transactions on Circuits and Systems for Video Technology , 2025 , 35 ( 5 ): 5037 - 5050 .
0
Views
0
下载量
0
CSCD
Publicity Resources
Related Articles
Related Author
Related Institution
京公网安备11010602201714号