
- Home
高级 检索
Chinese
English



1.中国科学院半导体研究所全固态光源实验室,北京 100083
2.中国科学院大学电子电气与通信工程学院,北京 101407
3.中国科学院大学现代产业学院,北京 101407
4.中国科学院大学材料科学与光电技术学院,北京 101407
5.中国科学院大学材料科学与光电工程中心,北京 100049
6.北京市全固态激光先进制造工程技术研究中心,北京 100083
7.厦门大学海洋与地球学院,海洋生物地球化学全国重点实验室,福建 厦门 361102
Received:29 March 2026,
Revised:2026-06-14,
Accepted:15 June 2026,
移动端阅览
ZHAO Hongwei, XIA Wentao, LIU Wenlong, et al. Research Progress on LiDAR-Based 3D Object Detection Algorithms[J/OL]. Acta Photonica Sinica, 2026, gz26-0135
ZHAO Hongwei, XIA Wentao, LIU Wenlong, et al. Research Progress on LiDAR-Based 3D Object Detection Algorithms[J/OL]. Acta Photonica Sinica, 2026, gz26-0135 DOI: 10.3788/gzxb20265508.0814001. CSTR: 32255.14.gzxb20265508.0814001.
激光雷达三维目标检测可为自动驾驶和铁路异物入侵监测等任务提供目标三维位置和尺寸信息。受点云稀疏、遮挡造成局部缺失以及语义信息不足影响,现有方法在远距离目标、小目标和复杂环境下的检测结果仍不够稳定。本文从传统几何处理、深度学习检测和激光雷达—相机融合三类方法展开分析。传统几何方法流程清晰,适用于规则场景中的候选筛选和误差控制,但对阈值和环境条件较敏感;深度学习方法提升了检测精度,但在远距离稀疏点云和小目标检测中仍易漏检;多模态融合方法能够利用图像语义信息补充点云特征,但标定误差和时间同步误差会降低检测可靠性。后续研究仍需解决稀疏点云下检测不稳定、小目标易漏检和三维标注成本高等问题。
LiDAR-based three-dimensional object detection is a key perception task for autonomous driving, railway intrusion monitoring, and other safety-critical applications. This review aims to clarify the main technical routes, representative methods, applicable scenarios, and remaining limitations of LiDAR-based 3D object detection. Instead of treating existing algorithms as a simple chronological sequence, this paper reorganizes them according to their processing mechanisms and practical constraints. Particular attention is paid to how different methods handle point cloud sparsity, incomplete object observations, semantic insufficiency, computational cost, and multimodal alignment errors. The purpose is to provide a structured reference for method selection, performance interpretation, and future research on reliable 3D object detection.The reviewed studies are organized into three parts: traditional geometric processing, deep learning-based detection, and LiDAR-camera fusion. For traditional methods, the review discusses ground segmentation, range-image segmentation, Euclidean clustering, density-based clustering, geometric fitting, and handcrafted descriptors, with emphasis on their roles in candidate generation, noise suppression, and error control. For deep learning methods, the review analyzes two closely related aspects: input representation and detection structure. The input representations include point-based methods, voxel-based sparse convolution, pillar and bird’s-eye-view representations, range-view projection, and hybrid point-voxel schemes. The detection structures include two-stage, one-stage, and center-based paradigms. For multimodal fusion, this review covers calibration-based alignment, targetless and learning-based self-calibration, projection-guided fusion, misalignment compensation, cross-modal feature interaction, and tightly coupled mapping. These methods are compared in terms of geometric detail preservation, feature representation ability, detection accuracy, inference efficiency, robustness, and deployment complexity.The review shows that traditional geometric methods remain useful when interpretability, simple implementation, and low computational cost are required. Ground fitting, clustering, L-shape fitting, and local geometric descriptors can reduce the search space and provide explicit shape constraints, but their performance is strongly affected by threshold settings, ground assumptions, target priors, and point density variation. When the scene contains uneven terrain, sparse distant objects, occlusion, or non-convex structures, errors in early segmentation and clustering are likely to propagate to bounding-box fitting and classification. Deep learning methods reduce this dependence on handcrafted rules by integrating point organization, feature extraction, candidate generation, and box regression into trainable frameworks. Point-based methods preserve fine-grained geometry and avoid discretization loss, but neighborhood search and local aggregation increase computational burden. Voxel-based sparse convolution improves spatial regularity and reduces invalid computation in empty regions, but voxel resolution still affects small-object representation and memory cost. Pillar and BEV methods convert 3D detection into efficient 2D dense prediction, making them suitable for real-time deployment, although height compression weakens vertical geometric expression. Range-view methods are compact and fast because they follow the native scanning form of LiDAR, but they are sensitive to occlusion, scale variation, and projection distortion. Hybrid methods combine voxel-level context with point-level detail and often achieve high accuracy, but their pipelines are more complex and harder to deploy. From the detection-structure perspective, two-stage methods usually provide stronger candidate refinement and localization accuracy, one-stage methods offer shorter inference paths and higher efficiency, and center-based methods reduce anchor matching complexity while maintaining a balance between accuracy and speed. LiDAR-camera fusion can improve recognition in sparse, distant, or semantically ambiguous scenes by introducing image texture and semantic cues. However, its effectiveness depends on reliable extrinsic calibration, temporal synchronization, spatial alignment, and appropriate cross-modal feature weighting. Misalignment, sensor degradation, or inaccurate semantic projection may introduce additional noise rather than improve detection.Reliable LiDAR-based 3D object detection cannot be achieved by improving network accuracy alone. Method selection should be guided by target scale, point cloud density, scene regularity, computing resources, and system maintenance requirements. For distant sparse targets, small objects, and occluded objects, future work should strengthen multi-frame information use, scale-adaptive feature learning, point cloud completion, and uncertainty estimation. For data annotation, low-cost supervision such as weak supervision, semi-supervised learning, click-level annotation, and cross-domain self-training deserves further attention. For multimodal systems, future research should focus less on simply adding fusion modules and more on calibration drift, time synchronization error, modality degradation, and failure detection. A practical 3D detection system should finally balance accuracy, real-time performance, robustness, annotation cost, and deployability under real operating conditions.
ZHAO Zengxu , HU Lianqing , REN Bin , et al . PointPillars-S 3D object detection algorithm based on LiDAR [J]. Acta Photonica Sinica , 2025 , 54 ( 6 ): 0614002 .
赵增旭 , 胡连庆 , 任彬 , 等 . 基于激光雷达的PointPillars-S三维目标检测算法 [J]. 光子学报 , 2025 , 54 ( 6 ): 0614002 .
TANG Jie , CHEN Wenwu , ZHOU Xinran , et al . Deep Learning-based Point Cloud Registration via Neighborhood Geometric Centroids (Invited) [J]. Acta Photonica Sinica , 2025 , 54 ( 9 ): 0954211 .
汤洁 , 陈文武 , 周昕然 , 等 . 基于邻域几何质心的深度学习点云配准(特邀) [J]. 光子学报 , 2025 , 54 ( 9 ): 0954211 .
VALVERDE M , MOUTINHO A , ZACCHI J V . A survey of deep learning-based 3D object detection methods for autonomous driving across different sensor modalities [J]. Sensors , 2025 , 25 ( 17 ): 5264 .
GEIGER A , LENZ P , URTASUN R . Are we ready for autonomous driving? The KITTI vision benchmark suite [C]// Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition . Piscataway : IEEE , 2012 : 3354 - 3361 .
CHEN Xiaozhi , MA Huimin , WAN Ji , et al . Multi-view 3D object detection network for autonomous driving [C]// Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition . Piscataway : IEEE , 2017 : 1907 - 1915 .
ZERMAS D , IZZAT I , PAPANIKOLOPOULOS N . Fast segmentation of 3D point clouds: A paradigm on LiDAR data for autonomous vehicle applications [C]// 2017 IEEE International Conference on Robotics and Automation (ICRA) . Piscataway : IEEE , 2017 : 5067 - 5073 .
HIMMELSBACH M , VON HUNDELSHAUSEN F , WUENSCHE H J . Fast segmentation of 3D point clouds for ground vehicles [C]// 2010 IEEE Intelligent Vehicles Symposium . Piscataway : IEEE , 2010 : 560 - 565 .
BOGOSLAVSKYI I , STACHNISS C . Fast range image-based segmentation of sparse 3D laser scans for online operation [C]// 2016 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS) . Piscataway : IEEE , 2016 : 163 - 169 .
CHEN Tongtong , DAI Bin , LIU Daxue , et al . Sparse Gaussian process regression based ground segmentation for autonomous land vehicles [C]// The 27th Chinese Control and Decision Conference (2015 CCDC) . Piscataway : IEEE , 2015 : 3993 - 3998 .
FU Huijin , MA Zhen , YANG Qi , et al . Lidar-based foreign object detection for railway intrusion under adverse weather conditions [C]// International Conference on Optical Communication, Signal Processing, and Optical Engineering (OCSPOE 2025) . SPIE , 2025 , 13797 : 153 - 159 .
ZHANG Wuming , QI Jianbo , WAN Peng , et al . An easy-to-use airborne LiDAR data filtering method based on cloth simulation [J]. Remote Sensing , 2016 , 8 ( 6 ): 501 .
ZEYBEK M . Spatio-Temporal Detection and Filtering of Dynamic Objects in Mobile LiDAR Point Clouds [J]. The International Archives of the Photogrammetry, Remote Sensing and Spatial Information Sciences , 2026 , 48 : 371 - 378 .
RUSU R B , COUSINS S . 3D is here: Point Cloud Library (PCL) [C]// 2011 IEEE International Conference on Robotics and Automation . Piscataway : IEEE , 2011 : 1 - 4 .
YANG Yuxing , ZHANG Bowen , YANG Boyu , et al . End-to-end railway obstacle detection enhanced by point cloud segmentation [J]. Engineering Applications of Artificial Intelligence , 2026 , 168 : 114008 .
ESTER M , KRIEGEL H P , SANDER J , et al . A density-based algorithm for discovering clusters in large spatial databases with noise [C]// Proceedings of the 2nd International Conference on Knowledge Discovery and Data Mining . Menlo Park : AAAI Press , 1996 : 226 - 231 .
WANG Heng , WANG Bin , LIU Bingbing , et al . Pedestrian recognition and tracking using 3D LiDAR for autonomous vehicle [J]. Robotics and Autonomous Systems , 2017 , 88 : 71 - 78 .
HELD D , GUILLORY D , REBSAMEN B , et al . A probabilistic framework for real-time 3D segmentation using spatial, temporal, and semantic cues [C]// Robotics: Science and Systems . 2016 .
IVARSEN M F , ST-MAURICE J P , HUSSEY G C , et al . Point-cloud clustering and tracking algorithm for radar interferometry [J]. Physical Review E , 2024 , 110 ( 4 ): 045207 .
GU Bo , LIU Jianxu , XIONG Huiyuan , et al . ECPC-ICP: A 6D vehicle pose estimation method by fusing the roadside lidar point cloud and road feature [J]. Sensors , 2021 , 21 ( 10 ): 3489 .
WANG C C , WANG Mudan , SUN Jun , et al . A safety warning algorithm based on axis aligned bounding box method to prevent onsite accidents of mobile construction machineries [J]. Sensors , 2021 , 21 ( 21 ): 7075 .
ZHANG Xiao , XU Wenda , DONG Chiyu , et al . Efficient L-shape fitting for vehicle detection using laser scanners [C]// 2017 IEEE Intelligent Vehicles Symposium (IV) . Piscataway : IEEE , 2017 : 54 - 59 .
RUSU R B , BLODOW N , BEETZ M . Fast point feature histograms (FPFH) for 3D registration [C]// 2009 IEEE International Conference on Robotics and Automation . Piscataway : IEEE , 2009 : 3212 - 3217 .
ZongliangNAN , ZHU Guoan , ZHANG Xu , et al . A novel high-precision railway obstacle detection algorithm based on 3D LiDAR [J]. Sensors , 2024 , 24 ( 10 ): 3148 .
TEICHMAN A , LEVINSON J , THRUN S . Towards 3D object recognition via classification of arbitrary object tracks [C]// 2011 IEEE International Conference on Robotics and Automation . Piscataway : IEEE , 2011 : 4034 - 4041 .
GOSWAMI P , VAISHNAV R , ANAND T , et al . A comprehensive review on LiDAR based 3D deep learning object detection algorithms [C]// 2025 International Conference on Computer , Electrical Communication Engineering (ICCECE) . Piscataway : IEEE , 2025 : 1 - 6 .
ZHANG Xiang , WANG Hai , DONG Haoran . A survey of deep learning-driven 3D object detection: Sensor modalities, technical architectures, and applications [J]. Sensors , 2025 , 25 ( 12 ): 3668 .
WU Shuwen , LI Yanxi , ZHANG Shaochen , et al . Review of deep learning-based 3D point cloud object detection [J]. Journal of Telemetry, Tracking and Command , 2024 , 45 ( 5 ): 1 - 18 .
武淑文 , 李燕烯 , 张少琛 , 等 . 基于深度学习的3D点云目标检测研究综述 [J]. 遥测遥控 , 2024 , 45 ( 5 ): 1 - 18 .
QI C R , SU Hao , MO Kaichun , et al . PointNet: Deep learning on point sets for 3D classification and segmentation [C]// Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition . Piscataway : IEEE , 2017 : 652 - 660 .
QI C R , YI Li , SU Hao , et al . PointNet++: Deep hierarchical feature learning on point sets in a metric space [J]. Advances in Neural Information Processing Systems , 2017 , 30 .
SHI Shaoshuai , WANG Xiaogang , LI Hongsheng . PointRCNN: 3D object proposal generation and detection from point cloud [C]// Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition . Piscataway : IEEE , 2019 : 770 - 779 .
YANG Zetong , SUN Yanan , LIU Shu , et al . 3DSSD: Point-based 3D single stage object detector [C]// Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition . Piscataway : IEEE , 2020 : 11040 - 11048 .
YAN Yan , MAO Yuxing , LI Bo . SECOND: Sparsely embedded convolutional detection [J]. Sensors , 2018 , 18 ( 10 ): 3337 .
CUI Yubo , LI Zhiheng , FANG Zheng , et al . Dynamic clustering transformer for LiDAR-based 3D object detection [J]. Pattern Recognition , 2026 , 172 : 112444 .
LIAN Lirong , QIN Yong , CAO Zhiwei , et al . RVSA-3D: Voxel-based fully sparse attention 3D object detection for rail transit obstacle perception [J]. Pattern Recognition , 2026 , 171 : 112324 .
CHEN Yukang , LI Yanwei , ZHANG Xiangyu , et al . Focal sparse convolutional networks for 3D object detection [C]// Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition . Piscataway : IEEE , 2022 : 5428 - 5437 .
LANG A H , VORA S , CAESAR H , et al . PointPillars: Fast encoders for object detection from point clouds [C]// Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition . Piscataway : IEEE , 2019 : 12697 - 12705 .
AO Lei , WAN Wenkang , OuyangNAN , et al . SVP: Stratified vertical priors for LiDAR-based 3D object detection [J]. Neurocomputing , 2026 , 659 : 131737 .
LI Junru , WANG Zhiling , GONG Diancheng , et al . SCNet3D: Rethinking the feature extraction process of pillar-based 3D object detection [J]. IEEE Transactions on Intelligent Transportation Systems , 2025 , 26 ( 1 ): 770 - 784 .
NOH J , LEE J , PARK H , et al . 3DPillars: Pillar-based two-stage 3D object detection [J]. Expert Systems With Applications , 2025 , 289 : 128349 .
MEYER G P , LADDHA A , KEE E , et al . LaserNet: An efficient probabilistic 3D object detector for autonomous driving [C]// Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition . Piscataway : IEEE , 2019 : 12677 - 12686 .
SHI Shaoshuai , GUO Chaoxu , JIANG Li , et al . PV-RCNN: Point-voxel feature set abstraction for 3D object detection [C]// Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition . Piscataway : IEEE , 2020 : 10529 - 10538 .
REN Junfen , WEN Changji , ZHANG Long , et al . High performance point-voxel feature set abstraction with mamba for 3D object detection [J]. Expert Systems With Applications , 2025 , 286 : 128127 .
SHI Shaoshuai , WANG Zhe , SHI Jianping , et al . From points to parts: 3D object detection from point cloud with part-aware and part-aggregation network [J]. IEEE Transactions on Pattern Analysis and Machine Intelligence , 2020 , 43 ( 8 ): 2647 - 2664 .
YANG Bin , LUO Wenjie , URTASUN R . PIXOR: Real-time 3D object detection from point clouds [C]// Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition . Piscataway : IEEE , 2018 : 7652 - 7660 .
YIN Tianwei , ZHOU Xingyi , KRAHENBUHL P . Center-based 3D object detection and tracking [C]// Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition . Piscataway : IEEE , 2021 : 11784 - 11793 .
CHEN Shitao , ZHANG Haolin , ZHENG Nanning . Leveraging anchor-based LiDAR 3D object detection via point assisted sample selection [J]. IEEE Transactions on Intelligent Transportation Systems , 2025 , 26 ( 6 ): 7939 - 7952 .
XIA Qiming , LIN Hongwei , YE Wei , et al . Label-efficient outdoor 3D object detection via single click annotation from LiDAR point cloud [J]. ISPRS Journal of Photogrammetry and Remote Sensing , 2026 , 234 : 151 - 166 .
MI Zeng , LIAN Zhe . Review of YOLO methods for general object detection [J]. Computer Engineering and Applications , 2024 , 60 ( 21 ): 38 - 54 .
米增 , 连哲 . 面向通用目标检测的YOLO方法研究综述 [J]. 计算机工程与应用 , 2024 , 60 ( 21 ): 38 - 54 .
HUANG Zhe , WANG Yongcai , LI Deying . Review of 3D object detection methods [J]. Chinese Journal of Intelligent Science and Technology , 2023 , 5 ( 1 ): 7 - 31 .
黄哲 , 王永才 , 李德英 . 3D目标检测方法研究综述 [J]. 智能科学与技术学报 , 2023 , 5 ( 1 ): 7 - 31 .
YANG Boquan , LI Jixiong , ZENG Ting . A review of environmental perception technology based on multi-sensor information fusion in autonomous driving [J]. World Electric Vehicle Journal , 2025 , 16 ( 1 ): 20 .
ZHAO Xuanyi , DOU Xiaohan , ZHENG Jihong , et al . A Lightweight Multi-Stage Visual Detection Approach for Complex Traffic Scenes [J]. Sensors , 2025 , 25 ( 16 ): 5014 .
YU Shuainan , WANG Xichao , YanCI , et al . MSPE-Fusion: A multimodal 3D object detection method with multi-sensor perception enhanced fusion [J]. Neurocomputing , 2025 , 645 : 130486 .
ZHANG Yongsheng , TU Chen , GAO Kun , et al . Multisensor information fusion: Future of environmental perception in intelligent vehicles [J]. Journal of Intelligent and Connected Vehicles , 2024 , 7 ( 3 ): 163 - 176 .
FAN Zheng , ZHANG Lele , WANG Xueyi , et al . LiDAR, IMU, and camera fusion for simultaneous localization and mapping: A systematic review [J]. Artificial Intelligence Review , 2025 , 58 ( 6 ): 174 .
SHI Pengtao , WEI Kangle . Extrinsic calibration method of LiDAR and camera based on semantic segmentation for autonomous driving environment [J]. Laser Optoelectronics Progress , 2024 , 61 ( 24 ): 2428011 .
史鹏涛 , 危康乐 . 自动驾驶环境下基于语义分割的激光雷达与相机外参标定方法 [J]. 激光与光电子学进展 , 2024 , 61 ( 24 ): 2428011 .
LIU Yang , JI Jie , PAN Deng , et al . Agricultural robot localization method based on LiDAR and IMU fusion [J]. Smart Agriculture , 2024 , 6 ( 3 ): 94 - 106 .
刘洋 , 冀杰 , 潘登 , 赵立军 , 李明生 . 基于激光雷达与IMU融合的农业机器人定位方法 [J]. 智慧农业(中英文) , 2024 , 6 ( 3 ): 94 - 106 .
FU Ting , XIE Shuke , HUI Weichao , et al . LiDAR-camera fusion: Dual-scale correction for vehicle multi-object detection and trajectory extraction [J]. Journal of Intelligent Transportation Systems , 2024 : 1 - 14 .
ZongliangNAN , LIU Wenlong , ZHU Guoan , et al . LiDAR-camera joint obstacle detection algorithm for railway track area [J]. Expert Systems with Applications , 2025 , 275 : 127089 .
SINGANDHUPE A , LA H M . Single frame lidar and stereo camera calibration using registration of 3D planes [C]// 2021 Fifth IEEE International Conference on Robotic Computing (IRC) . Piscataway : IEEE , 2021 : 115 - 118 .
JEONG S , KIM S . O3 LiDAR-camera calibration: One-shot, one-target and overcoming LiDAR limitations [J]. IEEE Sensors Journal , 2024 , 24 ( 11 ): 18659 - 18671 .
HE Zhanyuan , CHENG Guangzhi , ZHANG Yongbo , et al . Design of a LiDAR-camera self-calibration algorithm based on railway geometric features [C]// 2025 IEEE 8th International Conference on Signal Processing and Machine Learning (SPML) . Piscataway : IEEE , 2025 : 454 - 459 .
XU Zejing , LIU Yiqing , GAO Ruipeng , et al . KFCalibNet: A KansFormer-based self-calibration network for camera and LiDAR [C]// 2025 IEEE International Conference on Robotics and Automation (ICRA) . Piscataway : IEEE , 2025 : 627 - 633 .
WANG Zhangyu , YU Guizhen , WU Xinkai , et al . A camera and LiDAR data fusion method for railway object detection [J]. IEEE Sensors Journal , 2021 , 21 ( 12 ): 13442 - 13454 .
ZHAO Lin , ZHOU Hui , ZHU Xinge , et al . Lif-Seg: LiDAR and camera image fusion for 3D LiDAR semantic segmentation [J]. IEEE Transactions on Multimedia , 2023 , 26 : 1158 - 1168 .
LIU Wentao , WANG Yunpeng , YU Guizhen , et al . RailFusion: A LiDAR-camera data interaction network for 3-D railway object detection [J]. IEEE Transactions on Intelligent Transportation Systems , 2025 , 26 ( 8 ): 12761 - 12773 .
XIA Yu , WU Hongwei , ZHU Liucun , et al . A multi-sensor fusion framework with tight coupling for precise positioning and optimization [J]. Signal Processing , 2024 , 217 : 109343 .
FAWOLE O A , RAWAT D B . Recent advances in 3D object detection for self-driving vehicles: A survey [J]. AI , 2024 , 5 ( 3 ): 1255 - 1285 .
ZHU Minling , GONG Yadong , TIAN Chunwei , et al . A systematic survey of transformer-based 3D object detection for autonomous driving: Methods, challenges and trends [J]. Drones , 2024 , 8 ( 8 ): 41 .
0
Views
0
下载量
0
CSCD
Publicity Resources
Related Articles
Related Author
Related Institution
京公网安备11010602201714号