Innovations in 3D Object Detection: A Comprehensive Review of Methods, Sensor Fusion, and Future Directions
Review Article  ·  Published: 12 October 2024
Issue cover
ICCK Transactions on Sensing, Communication, and Control
Volume 1, Issue 1, 2024: 3-29
Review Article Free to Read

Innovations in 3D Object Detection: A Comprehensive Review of Methods, Sensor Fusion, and Future Directions

1 Interdisciplinary Research Centre for Aviation and Space Exploration (IRCASE), King Fahd University of Petroleum and Minerals (KFUPM), Dhahran 31261, Kingdom of Saudi Arabia
2 Electronic Engineering Department, Maynooth International Engineering College (MIEC), Maynooth University, Maynooth, Co. Kildare, Ireland
3 Department of Telecommunication Engineering, Mehran University of Engineering and Technology (MUET), Pakistan
* Corresponding Author: Ghulam E Mustafa Abro, [email protected]
Volume 1, Issue 1
You have access to this article · Limited-Time Free Access

Abstract

This review paper offers a thorough assessment of three-dimensional object recognition methods, an essential element in the perception frameworks of autonomous systems. This analysis emphasises the integration of LiDAR and camera sensors, providing a distinctive contrast with more economical alternatives like camera-only or camera-Radar combinations. This study objectively evaluates performance and practical implementation issues, such as cost and operational efficiency, thereby elucidating the limitations of existing systems and proposing avenues for further research. The insights provided render it a significant asset for enhancing 3D object recognition and autonomy in intelligent systems.

Graphical Abstract

Innovations in 3D Object Detection: A Comprehensive Review of Methods, Sensor Fusion, and Future Directions

Keywords

autonomous systems camera fusion methods LiDAR object detection radar and three-dimensional

Data Availability Statement

Not applicable.

Funding

This work was supported by the Interdisciplinary Research Centre for Aviation and Space Exploration (IRCASE), King Fahd University of Petroleum and Minerals (KFUPM, Kingdom of Saudi Arabia).

Conflicts of Interest

The authors declare no conflicts of interest.

Ethical Approval and Consent to Participate

Not applicable.

References

  1. Wang, L., Zhang, X., Song, Z., Bi, J., Zhang, G., Wei, H., ... & Zhao, L. (2023). Multi-modal 3d object detection in autonomous driving: A survey and taxonomy. IEEE Transactions on Intelligent Vehicles, 8(7), 3781-3798.
    [CrossRef] [Google Scholar]
  2. Yurtsever, E., Lambert, J., Carballo, A., & Takeda, K. (2020). A survey of autonomous driving: Common practices and emerging technologies. IEEE access, 8, 58443-58469.
    [CrossRef] [Google Scholar]
  3. Grigorescu, S., Trasnea, B., Cocias, T., & Macesanu, G. (2020). A survey of deep learning techniques for autonomous driving. Journal of field robotics, 37(3), 362-386.
    [CrossRef] [Google Scholar]
  4. Li, J., Yang, B., Chen, C., Huang, R., Dong, Z., & Xiao, W. (2018). Automatic registration of panoramic image sequence and mobile laser scanning data using semantic features. ISPRS Journal of Photogrammetry and Remote Sensing, 136, 41-57.
    [CrossRef] [Google Scholar]
  5. Liao, Y., Li, J., Kang, S., Li, Q., Zhu, G., Yuan, S., ... & Yang, B. (2023). SE-Calib: Semantic Edge-Based LiDAR–Camera Boresight Online Calibration in Urban Scenes. IEEE Transactions on Geoscience and Remote Sensing, 61, 1-13.
    [CrossRef] [Google Scholar]
  6. Wang, J. G., & Zhou, L. B. (2018). Traffic light recognition with high dynamic range imaging and deep learning. IEEE Transactions on Intelligent Transportation Systems, 20(4), 1341-1352.
    [CrossRef] [Google Scholar]
  7. Bijelic, M., Gruber, T., Mannan, F., Kraus, F., Ritter, W., Dietmayer, K., & Heide, F. (2020, June). Seeing Through Fog Without Seeing Fog: Deep Multimodal Sensor Fusion in Unseen Adverse Weather. In 2020 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) (pp. 11679-11689). IEEE.
    [CrossRef] [Google Scholar]
  8. Zang, S., Ding, M., Smith, D., Tyler, P., Rakotoarivelo, T., & Kaafar, M. A. (2019). The impact of adverse weather conditions on autonomous vehicles: How rain, snow, fog, and hail affect the performance of a self-driving car. IEEE vehicular technology magazine, 14(2), 103-111.
    [CrossRef] [Google Scholar]
  9. Kurihata, H., Takahashi, T., Ide, I., Mekada, Y., Murase, H., Tamatsu, Y., & Miyahara, T. (2005, June). Rainy weather recognition from in-vehicle camera images for driver assistance. In IEEE Proceedings. Intelligent Vehicles Symposium, 2005. (pp. 205-210). IEEE.
    [CrossRef] [Google Scholar]
  10. Hahner, M., Sakaridis, C., Dai, D., & Van Gool, L. (2021). Fog simulation on real LiDAR point clouds for 3D object detection in adverse weather. In Proceedings of the IEEE/CVF international conference on computer vision (pp. 15283-15292).
    [CrossRef] [Google Scholar]
  11. Fan, L., Pang, Z., Zhang, T., Wang, Y. X., Zhao, H., Wang, F., ... & Zhang, Z. (2022, June). Embracing Single Stride 3D Object Detector with Sparse Transformer. In 2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) (pp. 8448-8458). IEEE.
    [CrossRef] [Google Scholar]
  12. Filgueira, A., González-Jorge, H., Lagüela, S., Díaz-Vilariño, L., & Arias, P. (2017). Quantifying the influence of rain in LiDAR performance. Measurement, 95, 143-148.
    [CrossRef] [Google Scholar]
  13. Rasshofer, R. H., Spies, M., & Spies, H. (2011). Influences of weather phenomena on automotive laser radar systems. Advances in radio science, 9, 49-60.
    [CrossRef] [Google Scholar]
  14. Tang, X., Zhang, Z., & Qin, Y. (2021). On-road object detection and tracking based on radar and vision fusion: A review. IEEE Intelligent Transportation Systems Magazine, 14(5), 103-128.
    [CrossRef] [Google Scholar]
  15. Feng, D., Haase-Schütz, C., Rosenbaum, L., Hertlein, H., Glaeser, C., Timm, F., ... & Dietmayer, K. (2020). Deep multi-modal object detection and semantic segmentation for autonomous driving: Datasets, methods, and challenges. IEEE Transactions on Intelligent Transportation Systems, 22(3), 1341-1360.
    [CrossRef] [Google Scholar]
  16. Wei, Z., Zhang, F., Chang, S., Liu, Y., Wu, H., & Feng, Z. (2022). Mmwave radar and vision fusion for object detection in autonomous driving: A review. Sensors, 22(7), 2542.
    [CrossRef] [Google Scholar]
  17. Svenningsson, P., Fioranelli, F., & Yarovoy, A. (2021, May). Radar-pointgnn: Graph based object recognition for unstructured radar point-cloud data. In 2021 IEEE Radar Conference (RadarConf21) (pp. 1-6). IEEE.
    [CrossRef] [Google Scholar]
  18. Ulrich, M., Braun, S., Köhler, D., Niederlöhner, D., Faion, F., Gläser, C., & Blume, H. (2022, October). Improved orientation estimation and detection with hybrid object detection networks for automotive radar. In 2022 IEEE 25th International Conference on Intelligent Transportation Systems (ITSC) (pp. 111-117). IEEE.
    [CrossRef] [Google Scholar]
  19. Kim, Y., Choi, J. W., & Kum, D. (2020, October). Grif net: Gated region of interest fusion network for robust 3d object detection from radar point cloud and monocular image. In 2020 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS) (pp. 10857-10864). IEEE.
    [CrossRef] [Google Scholar]
  20. Chadwick, S., Maddern, W., & Newman, P. (2019, May). Distant vehicle detection using radar and vision. In 2019 International Conference on Robotics and Automation (ICRA) (pp. 8311-8317). IEEE.
    [CrossRef] [Google Scholar]
  21. Nobis, F., Geisslinger, M., Weber, M., Betz, J., & Lienkamp, M. (2019, October). A deep learning-based radar and camera sensor fusion architecture for object detection. In 2019 Sensor Data Fusion: Trends, Solutions, Applications (SDF) (pp. 1-7). IEEE.
    [CrossRef] [Google Scholar]
  22. John, V., & Mita, S. (2019). RVNet: Deep sensor fusion of monocular camera and radar for image-based obstacle detection in challenging environments. In Image and Video Technology: 9th Pacific-Rim Symposium, PSIVT 2019, Sydney, NSW, Australia, November 18–22, 2019, Proceedings 9 (pp. 351-364). Springer International Publishing. https://link.springer.com/chapter/10.1007/978-3-030-34879-3_27
    [Google Scholar]
  23. Li, L. Q., & Xie, Y. L. (2020, December). A feature pyramid fusion detection algorithm based on radar and camera sensor. In 2020 15th IEEE International Conference on Signal Processing (ICSP) (Vol. 1, pp. 366-370). IEEE.
    [CrossRef] [Google Scholar]
  24. Chang, S., Zhang, Y., Zhang, F., Zhao, X., Huang, S., Feng, Z., & Wei, Z. (2020). Spatial attention fusion for obstacle detection using mmwave radar and vision sensor. Sensors, 20(4), 956.
    [CrossRef] [Google Scholar]
  25. Nabati, R., & Qi, H. (2021, January). CenterFusion: Center-based Radar and Camera Fusion for 3D Object Detection. In 2021 IEEE Winter Conference on Applications of Computer Vision (WACV) (pp. 1526-1535). IEEE.
    [CrossRef] [Google Scholar]
  26. Li, Y., Zeng, K., & Shen, T. (2023). CenterTransFuser: radar point cloud and visual information fusion for 3D object detection. EURASIP Journal on Advances in Signal Processing, 2023(1), 7.
    [CrossRef] [Google Scholar]
  27. Long, Y., Kumar, A., Morris, D., Liu, X., Castro, M., & Chakravarty, P. (2023, June). RADIANT: Radar-image association network for 3D object detection. In Proceedings of the AAAI Conference on Artificial Intelligence (Vol. 37, No. 2, pp. 1808-1816).
    [CrossRef] [Google Scholar]
  28. Li, P., Zhao, H., Liu, P., & Cao, F. (2020, August). Rtm3d: Real-time monocular 3d detection from object keypoints for autonomous driving. In European Conference on Computer Vision (pp. 644-660). Cham: Springer International Publishing.
    [CrossRef] [Google Scholar]
  29. Zhang, Y., Lu, J., & Zhou, J. (2021, June). Objects are Different: Flexible Monocular 3D Object Detection. In 2021 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) (pp. 3288-3297). IEEE.
    [CrossRef] [Google Scholar]
  30. Simonelli, A., Bulò, S. R., Porzi, L., Lopez-Antequera, M., & Kontschieder, P. (2019, October). Disentangling Monocular 3D Object Detection. In 2019 IEEE/CVF International Conference on Computer Vision (ICCV) (pp. 1991-1999). IEEE.
    [CrossRef] [Google Scholar]
  31. Brazil, G., & Liu, X. (2019, October). M3D-RPN: Monocular 3D Region Proposal Network for Object Detection. In 2019 IEEE/CVF International Conference on Computer Vision (ICCV) (pp. 9286-9295). IEEE.
    [CrossRef] [Google Scholar]
  32. Cai, Y., Li, B., Jiao, Z., Li, H., Zeng, X., & Wang, X. (2020, April). Monocular 3d object detection with decoupled structured polygon estimation and height-guided depth estimation. In Proceedings of the AAAI Conference on Artificial Intelligence (Vol. 34, No. 07, pp. 10478-10485).
    [CrossRef] [Google Scholar]
  33. Chen, H., Huang, Y., Tian, W., Gao, Z., & Xiong, L. (2021). Monorun: Monocular 3d object detection by reconstruction and uncertainty propagation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (pp. 10379-10388).
    [CrossRef] [Google Scholar]
  34. Chen, Y., Tai, L., Sun, K., & Li, M. (2020, June). MonoPair: Monocular 3D Object Detection Using Pairwise Spatial Relationships. In 2020 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) (pp. 12090-12099). IEEE.
    [CrossRef] [Google Scholar]
  35. Heylen, J., De Wolf, M., Dawagne, B., Proesmans, M., Van Gool, L., Abbeloos, W., ... & Reino, D. O. (2021, October). MonoCInIS: Camera Independent Monocular 3D Object Detection using Instance Segmentation. In 2021 IEEE/CVF International Conference on Computer Vision Workshops (ICCVW) (pp. 923-934). IEEE.
    [CrossRef] [Google Scholar]
  36. Liu, Z., Wu, Z., & Tóth, R. (2020, June). SMOKE: Single-Stage Monocular 3D Object Detection via Keypoint Estimation. In 2020 IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops (CVPRW) (pp. 4289-4298). IEEE.
    [CrossRef] [Google Scholar]
  37. Liu, L., Lu, J., Xu, C., Tian, Q., & Zhou, J. (2019, June). Deep Fitting Degree Scoring Network for Monocular 3D Object Detection. In 2019 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) (pp. 1057-1066). IEEE.
    [CrossRef] [Google Scholar]
  38. Luo, S., Dai, H., Shao, L., & Ding, Y. (2021, June). M3DSSD: Monocular 3D Single Stage Object Detector. In 2021 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) (pp. 6141-6150). IEEE.
    [CrossRef] [Google Scholar]
  39. Barabanau, I., Artemov, A., Burnaev, E., & Murashkin, V. (2019). Monocular 3d object detection via geometric reasoning on keypoints. arXiv preprint arXiv:1905.05618.
    [Google Scholar]
  40. Lu, Y., Ma, X., Yang, L., Zhang, T., Liu, Y., Chu, Q., ... & Ouyang, W. (2021, October). Geometry Uncertainty Projection Network for Monocular 3D Object Detection. In 2021 IEEE/CVF International Conference on Computer Vision (ICCV) (pp. 3091-3101). IEEE.
    [CrossRef] [Google Scholar]
  41. You, Y., Wang, Y., Chao, W. L., Garg, D., Pleiss, G., Hariharan, B., ... & Weinberger, K. Q. (2019). Pseudo-lidar++: Accurate depth for 3d object detection in autonomous driving. arXiv preprint arXiv:1906.06310.
    [Google Scholar]
  42. Brazil, G., Pons-Moll, G., Liu, X., & Schiele, B. (2020, August). Kinematic 3d object detection in monocular video. In European Conference on Computer Vision (pp. 135-152). Cham: Springer International Publishing.
    [CrossRef] [Google Scholar]
  43. Simonelli, A., Bulo, S. R., Porzi, L., Ricci, E., & Kontschieder, P. (2020). Towards generalization across depth for monocular 3d object detection. In Computer Vision–ECCV 2020: 16th European Conference, Glasgow, UK, August 23–28, 2020, Proceedings, Part XXII 16 (pp. 767-782). Springer International Publishing.
    [CrossRef] [Google Scholar]
  44. Li, B., Ouyang, W., Sheng, L., Zeng, X., & Wang, X. (2019, June). GS3D: An Efficient 3D Object Detection Framework for Autonomous Driving. In 2019 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) (pp. 1019-1028). IEEE.
    [CrossRef] [Google Scholar]
  45. Ding, M., Huo, Y., Yi, H., Wang, Z., Shi, J., Lu, Z., & Luo, P. (2020, June). Learning Depth-Guided Convolutions for Monocular 3D Object Detection. In 2020 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) (pp. 11669-11678). IEEE.
    [CrossRef] [Google Scholar]
  46. Weng, X., & Kitani, K. (2019, October). Monocular 3D Object Detection with Pseudo-LiDAR Point Cloud. In 2019 IEEE/CVF International Conference on Computer Vision Workshop (ICCVW) (pp. 857-866). IEEE.
    [CrossRef] [Google Scholar]
  47. Hu, H. N., Cai, Q. Z., Wang, D., Lin, J., Sun, M., Kraehenbuehl, P., ... & Yu, F. (2019, October). Joint Monocular 3D Vehicle Detection and Tracking. In 2019 IEEE/CVF International Conference on Computer Vision (ICCV) (pp. 5389-5398). IEEE.
    [CrossRef] [Google Scholar]
  48. Ku, J., Pon, A. D., & Waslander, S. L. (2019, June). Monocular 3D Object Detection Leveraging Accurate Proposals and Shape Reconstruction. In 2019 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) (pp. 11859-11868). IEEE.
    [CrossRef] [Google Scholar]
  49. Lian, Q., Ye, B., Xu, R., Yao, W., & Zhang, T. (2022, June). Exploring Geometric Consistency for Monocular 3D Object Detection. In 2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) (pp. 1675-1684). IEEE.
    [CrossRef] [Google Scholar]
  50. Zia, M. Z., Stark, M., & Schindler, K. (2014, June). Are Cars Just 3D Boxes? Jointly Estimating the 3D Shape of Multiple Objects. In 2014 IEEE Conference on Computer Vision and Pattern Recognition (pp. 3678-3685). IEEE.
    [CrossRef] [Google Scholar]
  51. Chabot, F., Chaouch, M., Rabarisoa, J., Teulière, C., & Chateau, T. (2017, July). Deep MANTA: A Coarse-to-Fine Many-Task Network for Joint 2D and 3D Vehicle Analysis from Monocular Image. In 2017 IEEE Conference on Computer Vision and Pattern Recognition (CVPR) (pp. 1827-1836). IEEE.
    [CrossRef] [Google Scholar]
  52. He, T., & Soatto, S. (2019, July). Mono3d++: Monocular 3d vehicle detection with two-scale 3d hypotheses and task priors. In Proceedings of the AAAI Conference on Artificial Intelligence (Vol. 33, No. 01, pp. 8409-8416).
    [CrossRef] [Google Scholar]
  53. Tang, Y., Dorn, S., & Savani, C. (2020, September). Center3d: Center-based monocular 3d object detection with joint depth understanding. In DAGM German Conference on Pattern Recognition (pp. 289-302). Cham: Springer International Publishing.
    [CrossRef] [Google Scholar]
  54. Manhardt, F., Kehl, W., & Gaidon, A. (2019, June). ROI-10D: Monocular Lifting of 2D Detection to 6D Pose and Metric Shape. In 2019 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) (pp. 2064-2073). IEEE.
    [CrossRef] [Google Scholar]
  55. Beker, D., Kato, H., Morariu, M. A., Ando, T., Matsuoka, T., Kehl, W., & Gaidon, A. (2020). Monocular differentiable rendering for self-supervised 3d object detection. In Computer Vision–ECCV 2020: 16th European Conference, Glasgow, UK, August 23–28, 2020, Proceedings, Part XXI 16 (pp. 514-529). Springer International Publishing.
    [CrossRef] [Google Scholar]
  56. Zakharov, S., Kehl, W., Bhargava, A., & Gaidon, A. (2020, June). Autolabeling 3D Objects With Differentiable Rendering of SDF Shape Priors. In 2020 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) (pp. 12221-12230). IEEE.
    [CrossRef] [Google Scholar]
  57. Xu, Z., Zhang, W., Ye, X., Tan, X., Yang, W., Wen, S., ... & Huang, L. (2020, April). Zoomnet: Part-aware adaptive zooming neural network for 3d object detection. In Proceedings of the AAAI Conference on Artificial Intelligence (Vol. 34, No. 07, pp. 12557-12564).
    [CrossRef] [Google Scholar]
  58. Naiden, A., Paunescu, V., Kim, G., Jeon, B., & Leordeanu, M. (2019, September). Shift r-cnn: Deep monocular 3d object detection with closed-form geometric constraints. In 2019 IEEE international conference on image processing (ICIP) (pp. 61-65). IEEE.
    [CrossRef] [Google Scholar]
  59. Shi, X., Ye, Q., Chen, X., Chen, C., Chen, Z., & Kim, T. K. (2021, October). Geometry-based Distance Decomposition for Monocular 3D Object Detection. In 2021 IEEE/CVF International Conference on Computer Vision (ICCV) (pp. 15152-15161). IEEE.
    [CrossRef] [Google Scholar]
  60. Wang, T., Pang, J., & Lin, D. (2022, October). Monocular 3d object detection with depth from motion. In European Conference on Computer Vision (pp. 386-403). Cham: Springer Nature Switzerland.
    [CrossRef] [Google Scholar]
  61. Yan, L., Yan, P., Xiong, S., Xiang, X., & Tan, Y. (2024, June). MonoCD: Monocular 3D Object Detection with Complementary Depths. In 2024 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) (pp. 10248-10257). IEEE.
    [CrossRef] [Google Scholar]
  62. Reading, C., Harakeh, A., Chae, J., & Waslander, S. L. (2021, June). Categorical Depth Distribution Network for Monocular 3D Object Detection. In 2021 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) (pp. 8551-8560). IEEE.
    [CrossRef] [Google Scholar]
  63. Qin, Z., Wang, J., & Lu, Y. (2019, July). Monogrnet: A geometric reasoning network for monocular 3d object localization. In Proceedings of the AAAI conference on artificial intelligence (Vol. 33, No. 01, pp. 8851-8858).
    [CrossRef] [Google Scholar]
  64. Dong, S., Kong, X., Pan, X., Tang, F., Li, W., Chang, Y., & Dong, W. (2023). Semantic-context graph network for point-based 3D object detection. IEEE Transactions on Circuits and Systems for Video Technology, 33(11), 6474-6486.
    [CrossRef] [Google Scholar]
  65. Wang, L., Du, L., Ye, X., Fu, Y., Guo, G., Xue, X., ... & Zhang, L. (2021). Depth-conditioned dynamic message propagation for monocular 3d object detection. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (pp. 454-463).
    [CrossRef] [Google Scholar]
  66. Chang, J., & Wetzstein, G. (2019). Deep optics for monocular depth estimation and 3d object detection. In Proceedings of the IEEE/CVF international conference on computer vision (pp. 10193-10202).
    [Google Scholar]
  67. Park, D., Ambruş, R., Guizilini, V., Li, J., & Gaidon, A. (2021, October). Is Pseudo-Lidar needed for Monocular 3D Object detection?. In 2021 IEEE/CVF International Conference on Computer Vision (ICCV) (pp. 3122-3132). IEEE.
    [CrossRef] [Google Scholar]
  68. Wang, T., Zhu, X., Pang, J., & Lin, D. (2021). Fcos3d: Fully convolutional one-stage monocular 3d object detection. In 2Proceedings of the IEEE/CVF international conference on computer vision (pp. 913-922).
    [Google Scholar]
  69. Poggi, M., Tosi, F., Batsos, K., Mordohai, P., & Mattoccia, S. (2021). On the synergies between machine learning and binocular stereo for depth estimation from images: A survey. IEEE Transactions on Pattern Analysis and Machine Intelligence, 44(9), 5314-5334.
    [CrossRef] [Google Scholar]
  70. Li, P., Chen, X., & Shen, S. (2019, June). Stereo R-CNN Based 3D Object Detection for Autonomous Driving. In 2019 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) (pp. 7636-7644). IEEE.
    [CrossRef] [Google Scholar]
  71. Sun, J., Chen, L., Xie, Y., Zhang, S., Jiang, Q., Zhou, X., & Bao, H. (2020, June). Disp R-CNN: Stereo 3D Object Detection via Shape Prior Guided Instance Disparity Estimation. In 2020 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) (pp. 10545-10554). IEEE.
    [CrossRef] [Google Scholar]
  72. Liu, Y., Wang, L., & Liu, M. (2021, May). Yolostereo3d: A step back to 2d for efficient stereo 3d detection. In 2021 IEEE international conference on Robotics and automation (ICRA) (pp. 13018-13024). IEEE.
    [CrossRef] [Google Scholar]
  73. Qin, Z., Wang, J., & Lu, Y. (2019, June). Triangulation Learning Network: From Monocular to Stereo 3D Object Detection. In 2019 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) (pp. 7607-7615). IEEE.
    [CrossRef] [Google Scholar]
  74. Chen, Y., Liu, S., Shen, X., & Jia, J. (2020, June). DSGN: Deep Stereo Geometry Network for 3D Object Detection. In 2020 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) (pp. 12533-12542). IEEE.
    [CrossRef] [Google Scholar]
  75. Guo, X., Shi, S., Wang, X., & Li, H. (2021, October). LIGA-Stereo: Learning LiDAR Geometry Aware Representations for Stereo-based 3D Detector. In 2021 IEEE/CVF International Conference on Computer Vision (ICCV) (pp. 3133-3143). IEEE.
    [CrossRef] [Google Scholar]
  76. Gao, A., Pang, Y., Nie, J., Shao, Z., Cao, J., Guo, Y., & Li, X. (2022). ESGN: Efficient stereo geometry network for fast 3D object detection. IEEE Transactions on Circuits and Systems for Video Technology, 34(4), 2000-2009.
    [CrossRef] [Google Scholar]
  77. Su, K., Yan, W., Wei, X., & Gu, M. (2022). Stereo VoVNet-CNN for 3D object detection. Multimedia Tools and Applications, 81(25), 35803-35813.
    [CrossRef] [Google Scholar]
  78. Mao, J., Niu, M., Bai, H., Liang, X., Xu, H., & Xu, C. (2021, October). Pyramid R-CNN: Towards Better Performance and Adaptability for 3D Object Detection. In 2021 IEEE/CVF International Conference on Computer Vision (ICCV) (pp. 2703-2712). IEEE.
    [CrossRef] [Google Scholar]
  79. Shi, Y., Guo, Y., Mi, Z., & Li, X. (2022). Stereo CenterNet-based 3D object detection for autonomous driving. Neurocomputing, 471, 219-229.
    [CrossRef] [Google Scholar]
  80. Chen, L., Sun, J., Xie, Y., Zhang, S., Shuai, Q., Jiang, Q., ... & Zhou, X. (2021). Shape prior guided instance disparity estimation for 3d object detection. IEEE Transactions on Pattern Analysis and Machine Intelligence, 44(9), 5529-5540.
    [CrossRef] [Google Scholar]
  81. Li, Y., Bao, H., Ge, Z., Yang, J., Sun, J., & Li, Z. (2023, June). Bevstereo: Enhancing depth estimation in multi-view 3d object detection with temporal stereo. In Proceedings of the AAAI conference on artificial intelligence (Vol. 37, No. 2, pp. 1486-1494).
    [CrossRef] [Google Scholar]
  82. Peng, X., Zhu, X., Wang, T., & Ma, Y. (2022, January). SIDE: Center-based Stereo 3D Detector with Structure-aware Instance Depth Estimation. In 2022 IEEE/CVF Winter Conference on Applications of Computer Vision (WACV) (pp. 225-234). IEEE.
    [CrossRef] [Google Scholar]
  83. Qian, R., Garg, D., Wang, Y., You, Y., Belongie, S., Hariharan, B., ... & Chao, W. L. (2020, June). End-to-End Pseudo-LiDAR for Image-Based 3D Object Detection. In 2020 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) (pp. 5880-5889). IEEE.
    [CrossRef] [Google Scholar]
  84. Chen, Y., Huang, S., Liu, S., Yu, B., & Jia, J. (2022). Dsgn++: Exploiting visual-spatial relation for stereo-based 3d detectors. IEEE Transactions on Pattern Analysis and Machine Intelligence, 45(4), 4416-4429.
    [CrossRef] [Google Scholar]
  85. He, Q., Wang, Z., Zeng, H., Zeng, Y., Liu, Y., Liu, S., & Zeng, B. (2022). Stereo RGB and deeper LiDAR-based network for 3D object detection in autonomous driving. IEEE Transactions on Intelligent Transportation Systems, 24(1), 152-162.
    [CrossRef] [Google Scholar]
  86. Peng, W., Pan, H., Liu, H., & Sun, Y. (2020, June). IDA-3D: Instance-Depth-Aware 3D Object Detection From Stereo Vision for Autonomous Driving. In 2020 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) (pp. 13012-13021). IEEE.
    [CrossRef] [Google Scholar]
  87. Bai, L., Zhao, Y., Elhousni, M., & Huang, X. (2020). DepthNet: Real-time LiDAR point cloud depth completion for autonomous vehicles. IEEE access, 8, 227825-227833.
    [CrossRef] [Google Scholar]
  88. Ma, C., Pei, S., Sun, G., Meng, R., & Luo, K. (2022). Disparity estimation based on fusion of vision and LiDAR. International Journal of Wavelets, Multiresolution and Information Processing, 20(05), 2250014.
    [CrossRef] [Google Scholar]
  89. Wang, L., Li, R., Sun, J., Liu, X., Zhao, L., Seah, H. S., ... & Tandianus, B. (2019). Multi-view fusion-based 3D object detection for robot indoor scene perception. Sensors, 19(19), 4092.
    [CrossRef] [Google Scholar]
  90. Liu, Y., Yan, J., Jia, F., Li, S., Gao, A., Wang, T., & Zhang, X. (2023, October). PETRv2: A Unified Framework for 3D Perception from Multi-Camera Images. In 2023 IEEE/CVF International Conference on Computer Vision (ICCV) (pp. 3239-3249). IEEE.
    [CrossRef] [Google Scholar]
  91. Chen, D., Li, J., Guizilini, V., Ambruş, R., & Gaidon, A. (2023, June). Viewpoint Equivariance for Multi-View 3D Object Detection. In 2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) (pp. 9213-9222). IEEE.
    [CrossRef] [Google Scholar]
  92. Liu, X., Zheng, C., Qian, M., Xue, N., Chen, C., Zhang, Z., ... & Wu, T. (2024, June). Multi-View Attentive Contextualization for Multi-View 3D Object Detection. In 2024 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) (pp. 16688-16698). IEEE.
    [CrossRef] [Google Scholar]
  93. Liu, Y., Wang, T., Zhang, X., & Sun, J. (2022, October). Petr: Position embedding transformation for multi-view 3d object detection. In European conference on computer vision (pp. 531-548). Cham: Springer Nature Switzerland.
    [CrossRef] [Google Scholar]
  94. Deng, J., & Czarnecki, K. (2019, October). MLOD: A multi-view 3D object detection based on robust feature fusion method. In 2019 IEEE intelligent transportation systems conference (ITSC) (pp. 279-284). IEEE.
    [CrossRef] [Google Scholar]
  95. Choy, C. B., Xu, D., Gwak, J., Chen, K., & Savarese, S. (2016, September). 3d-r2n2: A unified approach for single and multi-view 3d object reconstruction. In European conference on computer vision (pp. 628-644). Cham: Springer International Publishing.
    [CrossRef] [Google Scholar]
  96. Rubino, C., Crocco, M., & Del Bue, A. (2017). 3d object localisation from multi-view image detections. IEEE transactions on pattern analysis and machine intelligence, 40(6), 1281-1294.
    [CrossRef] [Google Scholar]
  97. Yang, Z., & Wang, L. (2019, October). Learning Relationships for Multi-View 3D Object Recognition. In 2019 IEEE/CVF International Conference on Computer Vision (ICCV) (pp. 7504-7513). IEEE.
    [CrossRef] [Google Scholar]
  98. Wang, C., Pelillo, M., & Siddiqi, K. (2019). Dominant set clustering and pooling for multi-view 3d object recognition. arXiv preprint arXiv:1906.01592.
    [Google Scholar]
  99. Chen, X., Ma, H., Wan, J., Li, B., & Xia, T. (2017, July). Multi-view 3D Object Detection Network for Autonomous Driving. In 2017 IEEE Conference on Computer Vision and Pattern Recognition (CVPR) (pp. 6526-6534). IEEE.
    [CrossRef] [Google Scholar]
  100. Zhou, Y., Sun, P., Zhang, Y., Anguelov, D., Gao, J., Ouyang, T., ... & Vasudevan, V. (2020, May). End-to-end multi-view fusion for 3d object detection in lidar point clouds. In Conference on Robot Learning (pp. 923-932). PMLR.
    [Google Scholar]
  101. Qi, C. R., Su, H., Mo, K., & Guibas, L. J. (2017). Pointnet: Deep learning on point sets for 3d classification and segmentation. In Proceedings of the IEEE conference on computer vision and pattern recognition (pp. 652-660).
    [Google Scholar]
  102. Lang, A. H., Vora, S., Caesar, H., Zhou, L., Yang, J., & Beijbom, O. (2019). Pointpillars: Fast encoders for object detection from point clouds. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition (pp. 12697-12705).
    [CrossRef] [Google Scholar]
  103. Qi, C. R., Yi, L., Su, H., & Guibas, L. J. (2017). Pointnet++: Deep hierarchical feature learning on point sets in a metric space. Advances in neural information processing systems, 30.
    [Google Scholar]
  104. Wang, Y., Guizilini, V. C., Zhang, T., Wang, Y., Zhao, H., & Solomon, J. (2022, January). Detr3d: 3d object detection from multi-view images via 3d-to-2d queries. In Conference on robot learning (pp. 180-191). PMLR.
    [Google Scholar]
  105. Lin, J., Rickert, M., & Knoll, A. (2021, May). Deep hierarchical rotation invariance learning with exact geometry feature representation for point cloud classification. In 2021 IEEE international conference on robotics and automation (ICRA) (pp. 9529-9535). IEEE.
    [CrossRef] [Google Scholar]
  106. Zhang, K., Hao, M., Wang, J., Chen, X., Leng, Y., de Silva, C. W., & Fu, C. (2021, November). Linked dynamic graph cnn: Learning through point cloud by linking hierarchical features. In 2021 27th international conference on mechatronics and machine vision in practice (M2VIP) (pp. 7-12). IEEE.
    [CrossRef] [Google Scholar]
  107. Zhang, J., Liu, J., Liu, X., Wei, J., Cao, J., & Tang, K. (2021). Feature interpolation convolution for point cloud analysis. Computers & Graphics, 99, 182-191.
    [CrossRef] [Google Scholar]
  108. Shi, S., Wang, X., & Li, H. (2019, June). PointRCNN: 3D Object Proposal Generation and Detection From Point Cloud. In 2019 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) (pp. 770-779). IEEE.
    [CrossRef] [Google Scholar]
  109. Liu, Z., Tang, H., Lin, Y., & Han, S. (2019). Point-voxel cnn for efficient 3d deep learning. Advances in neural information processing systems, 32.
    [Google Scholar]
  110. Chen, C., Chen, Z., Zhang, J., & Tao, D. (2022, June). Sasa: Semantics-augmented set abstraction for point-based 3d object detection. In Proceedings of the AAAI Conference on Artificial Intelligence (Vol. 36, No. 1, pp. 221-229).
    [CrossRef] [Google Scholar]
  111. Ngiam, J., Caine, B., Han, W., Yang, B., Chai, Y., Sun, P., ... & Vasudevan, V. (2019). Starnet: Targeted computation for object detection in point clouds. arXiv preprint arXiv:1908.11069.
    [Google Scholar]
  112. Yang, Z., Sun, Y., Liu, S., & Jia, J. (2020, June). 3DSSD: Point-Based 3D Single Stage Object Detector. In 2020 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) (pp. 11037-11045). IEEE.
    [CrossRef] [Google Scholar]
  113. Yang, H., Liu, Z., Wu, X., Wang, W., Qian, W., He, X., & Cai, D. (2022, October). Graph r-cnn: Towards accurate 3d object detection with semantic-decorated local graph. In European conference on computer vision (pp. 662-679). Cham: Springer Nature Switzerland.
    [CrossRef] [Google Scholar]
  114. Shi, W., & Rajkumar, R. (2020, June). Point-GNN: Graph Neural Network for 3D Object Detection in a Point Cloud. In 2020 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) (pp. 1708-1716). IEEE.
    [CrossRef] [Google Scholar]
  115. Yang, B., Luo, W., & Urtasun, R. (2018, June). PIXOR: Real-time 3D Object Detection from Point Clouds. In 2018 IEEE/CVF Conference on Computer Vision and Pattern Recognition (pp. 7652-7660). IEEE.
    [CrossRef] [Google Scholar]
  116. He, Q., Wang, Z., Zeng, H., Zeng, Y., & Liu, Y. (2022, June). Svga-net: Sparse voxel-graph attention network for 3d object detection from point clouds. In Proceedings of the AAAI Conference on Artificial Intelligence (Vol. 36, No. 1, pp. 870-878).
    [CrossRef] [Google Scholar]
  117. Zarzar, J., Giancola, S., & Ghanem, B. (2019). PointRGCN: Graph convolution networks for 3D vehicles detection refinement. arXiv preprint arXiv:1911.12236.
    [Google Scholar]
  118. Feng, M., Gilani, S. Z., Wang, Y., Zhang, L., & Mian, A. (2020). Relation graph network for 3D object detection in point clouds. IEEE Transactions on Image Processing, 30, 92-107.
    [CrossRef] [Google Scholar]
  119. Xie, Q., Lai, Y. K., Wu, J., Wang, Z., Lu, D., Wei, M., & Wang, J. (2021, October). VENet: Voting Enhancement Network for 3D Object Detection. In 2021 IEEE/CVF International Conference on Computer Vision (ICCV) (pp. 3692-3701). IEEE.
    [CrossRef] [Google Scholar]
  120. Liu, Z., Zhang, Z., Cao, Y., Hu, H., & Tong, X. (2021, October). Group-Free 3D Object Detection via Transformers. In 2021 IEEE/CVF International Conference on Computer Vision (ICCV) (pp. 2929-2938). IEEE.
    [CrossRef] [Google Scholar]
  121. Piewak, F., Rehfeld, T., Weber, M., & Zöllner, J. M. (2017, June). Fully convolutional neural networks for dynamic object detection in grid maps. In 2017 IEEE Intelligent Vehicles Symposium (IV) (pp. 392-398). IEEE.
    [CrossRef] [Google Scholar]
  122. Shi, S., Guo, C., Jiang, L., Wang, Z., Shi, J., Wang, X., & Li, H. (2020, June). PV-RCNN: Point-Voxel Feature Set Abstraction for 3D Object Detection. In 2020 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) (pp. 10526-10535). IEEE.
    [CrossRef] [Google Scholar]
  123. Yang, B., Liang, M., & Urtasun, R. (2018, October). Hdnet: Exploiting hd maps for 3d object detection. In Conference on Robot Learning (pp. 146-155). PMLR.
    [Google Scholar]
  124. Engelcke, M., Rao, D., Wang, D. Z., Tong, C. H., & Posner, I. (2017, May). Vote3deep: Fast object detection in 3d point clouds using efficient convolutional neural networks. In 2017 IEEE International Conference on Robotics and Automation (ICRA) (pp. 1355-1361). IEEE.
    [CrossRef] [Google Scholar]
  125. Beltrán, J., Guindel, C., Moreno, F. M., Cruzado, D., Garcia, F., & De La Escalera, A. (2018, November). Birdnet: a 3d object detection framework from lidar information. In 2018 21st International Conference on Intelligent Transportation Systems (ITSC) (pp. 3517-3523). IEEE.
    [CrossRef] [Google Scholar]
  126. Simon, M., Amende, K., Kraus, A., Honer, J., Sämann, T., Kaulbersch, H., ... & Gross, H. M. (2019, June). Complexer-YOLO: Real-Time 3D Object Detection and Tracking on Semantic Point Clouds. In 2019 IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops (CVPRW) (pp. 1190-1199). IEEE.
    [CrossRef] [Google Scholar]
  127. Yang, B., Wang, J., Clark, R., Hu, Q., Wang, S., Markham, A., & Trigoni, N. (2019). Learning object bounding boxes for 3d instance segmentation on point clouds. Advances in neural information processing systems, 32.
    [Google Scholar]
  128. Shi, G., Li, R., & Ma, C. (2022, October). Pillarnet: Real-time and high-performance pillar-based 3d object detection. In European conference on computer vision (pp. 35-52). Cham: Springer Nature Switzerland.
    [CrossRef] [Google Scholar]
  129. Zhou, Y., & Tuzel, O. (2018, June). VoxelNet: End-to-End Learning for Point Cloud Based 3D Object Detection. In 2018 IEEE/CVF Conference on Computer Vision and Pattern Recognition (pp. 4490-4499). IEEE.
    [CrossRef] [Google Scholar]
  130. Lu, B., Sun, Y., Yang, Z., Song, R., Jiang, H., & Liu, Y. (2024). HRNet: 3D object detection network for point cloud with hierarchical refinement. Pattern Recognition, 149, 110254.
    [CrossRef] [Google Scholar]
  131. Yin, T., Zhou, X., & Krahenbuhl, P. (2021). Center-based 3d object detection and tracking. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition (pp. 11784-11793).
    [CrossRef] [Google Scholar]
  132. Wang, Y., Fathi, A., Kundu, A., Ross, D. A., Pantofaru, C., Funkhouser, T., & Solomon, J. (2020, August). Pillar-based object detection for autonomous driving. In European Conference on Computer Vision (pp. 18-34). Cham: Springer International Publishing.
    [CrossRef] [Google Scholar]
  133. Rukhovich, D., Vorontsova, A., & Konushin, A. (2022, October). Fcaf3d: Fully convolutional anchor-free 3d object detection. In European Conference on Computer Vision (pp. 477-493). Cham: Springer Nature Switzerland.
    [CrossRef] [Google Scholar]
  134. Li, B. (2017, September). 3d fully convolutional network for vehicle detection in point cloud. In 2017 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS) (pp. 1513-1518). IEEE.
    [CrossRef] [Google Scholar]
  135. Rebut, J., Ouaknine, A., Malik, W., & Pérez, P. (2022). Raw high-definition radar for multi-task learning. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (pp. 17021-17030).
    [Google Scholar]
  136. Mao, J., Xue, Y., Niu, M., Bai, H., Feng, J., Liang, X., ... & Xu, C. (2021, October). Voxel Transformer for 3D Object Detection. In 2021 IEEE/CVF International Conference on Computer Vision (ICCV) (pp. 3144-3153). IEEE.
    [CrossRef] [Google Scholar]
  137. Deng, J., Shi, S., Li, P., Zhou, W., Zhang, Y., & Li, H. (2021, May). Voxel r-cnn: Towards high performance voxel-based 3d object detection. In Proceedings of the AAAI conference on artificial intelligence (Vol. 35, No. 2, pp. 1201-1209).
    [CrossRef] [Google Scholar]
  138. Song, Z., Wei, H., Jia, C., Xia, Y., Li, X., & Zhang, C. (2023). VP-Net: Voxels as points for 3-D object detection. IEEE Transactions on Geoscience and Remote Sensing, 61, 1-12.
    [CrossRef] [Google Scholar]
  139. Wang, H., Chen, Z., Cai, Y., Chen, L., Li, Y., Sotelo, M. A., & Li, Z. (2022). Voxel-RCNN-complex: An effective 3-D point cloud object detector for complex traffic conditions. IEEE Transactions on Instrumentation and Measurement, 71, 1-12.
    [CrossRef] [Google Scholar]
  140. Liu, Z., Tang, H., Zhao, S., Shao, K., & Han, S. (2021). Pvnas: 3d neural architecture search with point-voxel convolution. IEEE Transactions on Pattern Analysis and Machine Intelligence, 44(11), 8552-8568.
    [CrossRef] [Google Scholar]
  141. Li, Z., Wang, F., & Wang, N. (2021, June). LiDAR R-CNN: An Efficient and Universal 3D Object Detector. In 2021 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) (pp. 7542-7551). IEEE.
    [CrossRef] [Google Scholar]
  142. Miao, Z., Chen, J., Pan, H., Zhang, R., Liu, K., Hao, P., ... & Zhan, X. (2021, June). PVGNet: A Bottom-Up One-Stage 3D Object Detector with Integrated Multi-Level Features. In 2021 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) (pp. 3278-3287). IEEE.
    [CrossRef] [Google Scholar]
  143. Li, P., Su, S., & Zhao, H. (2021, May). Rts3d: Real-time stereo 3d detection from 4d feature-consistency embedding space for autonomous driving. In Proceedings of the AAAI Conference on Artificial Intelligence (Vol. 35, No. 3, pp. 1930-1939).
    [CrossRef] [Google Scholar]
  144. Guan, T., Wang, J., Lan, S., Chandra, R., Wu, Z., Davis, L., & Manocha, D. (2022, January). M3DETR: Multi-representation, Multi-scale, Mutual-relation 3D Object Detection with Transformers. In 2022 IEEE/CVF Winter Conference on Applications of Computer Vision (WACV) (pp. 2293-2303). IEEE.
    [CrossRef] [Google Scholar]
  145. Thakur, A., & Rajalakshmi, P. (2023, July). Lidar and camera raw data sensor fusion in real-time for obstacle detection. In 2023 IEEE Sensors Applications Symposium (SAS) (pp. 1-6). IEEE.
    [CrossRef] [Google Scholar]
  146. Deng, J., Zhou, W., Zhang, Y., & Li, H. (2021). From multi-view to hollow-3D: Hallucinated hollow-3D R-CNN for 3D object detection. IEEE Transactions on Circuits and Systems for Video Technology, 31(12), 4722-4734.
    [CrossRef] [Google Scholar]
  147. Shi, S., Wang, Z., Shi, J., Wang, X., & Li, H. (2020). From points to parts: 3d object detection from point cloud with part-aware and part-aggregation network. IEEE transactions on pattern analysis and machine intelligence, 43(8), 2647-2664.
    [CrossRef] [Google Scholar]
  148. Wu, D., Liang, Z., & Chen, G. (2022). Deep learning for LiDAR-only and LiDAR-fusion 3D perception: A survey. Intelligence & Robotics, 2(2), 105-129.
    [CrossRef] [Google Scholar]
  149. Vora, S., Lang, A. H., Helou, B., & Beijbom, O. (2020). Pointpainting: Sequential fusion for 3d object detection. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition (pp. 4604-4612).
    [CrossRef] [Google Scholar]
  150. Tu, J., Wang, P., & Liu, F. (2021, July). Pp-rcnn: Point-pillars feature set abstraction for 3d real-time object detection. In 2021 International Joint Conference on Neural Networks (IJCNN) (pp. 1-8). IEEE.
    [CrossRef] [Google Scholar]
  151. Liang, M., Yang, B., Wang, S., & Urtasun, R. (2018). Deep continuous fusion for multi-sensor 3d object detection. In Proceedings of the European conference on computer vision (ECCV) (pp. 641-656).
    [CrossRef] [Google Scholar]
  152. Shi, S., Jiang, L., Deng, J., Wang, Z., Guo, C., Shi, J., ... & Li, H. (2023). PV-RCNN++: Point-voxel feature set abstraction with local vector representation for 3D object detection. International Journal of Computer Vision, 131(2), 531-551.
    [CrossRef] [Google Scholar]
  153. Wu, P., Gu, L., Yan, X., Xie, H., Wang, F. L., Cheng, G., & Wei, M. (2023). PV-RCNN++: semantical point-voxel feature interaction for 3D object detection. The visual computer, 39(6), 2425-2440.
    [CrossRef] [Google Scholar]
  154. Li, S., Geng, K., Yin, G., Wang, Z., & Qian, M. (2023). MVMM: Multiview multimodal 3-D object detection for autonomous driving. IEEE Transactions on Industrial Informatics, 20(1), 845-853.
    [CrossRef] [Google Scholar]
  155. Li, J., Luo, C., & Yang, X. (2023, June). PillarNeXt: Rethinking Network Designs for 3D Object Detection in LiDAR Point Clouds. In 2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) (pp. 17567-17576). IEEE.
    [CrossRef] [Google Scholar]
  156. Farag, W. (2021). Kalman-filter-based sensor fusion applied to road-objects detection and tracking for autonomous vehicles. Proceedings of the Institution of Mechanical Engineers, Part I: Journal of Systems and Control Engineering, 235(7), 1125-1138.
    [CrossRef] [Google Scholar]
  157. Shan, M., Narula, K., Worrall, S., Wong, Y. F., Perez, J. S. B., Gray, P., & Nebot, E. (2022, October). A novel probabilistic V2X data fusion framework for cooperative perception. In 2022 IEEE 25th International Conference on Intelligent Transportation Systems (ITSC) (pp. 2013-2020). IEEE.
    [CrossRef] [Google Scholar]
  158. Hafeez, F., Sheikh, U. U., Alkhaldi, N., Al Garni, H. Z., Arfeen, Z. A., & Khalid, S. A. (2020). Insights and strategies for an autonomous vehicle with a sensor fusion innovation: A fictional outlook. IEEE Access, 8, 135162-135175.
    [CrossRef] [Google Scholar]
  159. Bocu, R., Bocu, D., & Iavich, M. (2021). Objects detection using sensors data fusion in autonomous driving scenarios. Electronics, 10(23), 2903.
    [CrossRef] [Google Scholar]
  160. Zhang, D., Yuan, J., Meng, H., Wang, W., He, R., & Li, S. (2024). Extrinsic calibration method for integrating infrared thermal imaging camera and 3D LiDAR. Sensor Review, 44(4), 490-504.
    [CrossRef] [Google Scholar]
  161. Frossard, D., & Urtasun, R. (2018, May). End-to-end learning of multi-sensor 3d tracking by detection. In 2018 IEEE international conference on robotics and automation (ICRA) (pp. 635-642). IEEE.
    [CrossRef] [Google Scholar]
  162. Kim, K. E., Lee, C. J., Pae, D. S., & Lim, M. T. (2017, October). Sensor fusion for vehicle tracking with camera and radar sensor. In 2017 17th International Conference on Control, Automation and Systems (ICCAS) (pp. 1075-1077). IEEE.
    [CrossRef] [Google Scholar]
  163. Choi, J. D., & Kim, M. Y. (2023). A sensor fusion system with thermal infrared camera and LiDAR for autonomous vehicles and deep learning based object detection. ICT Express, 9(2), 222-227.
    [CrossRef] [Google Scholar]
  164. Xie, S., Yang, D., Jiang, K., & Zhong, Y. (2018). Pixels and 3-D points alignment method for the fusion of camera and LiDAR data. IEEE Transactions on Instrumentation and Measurement, 68(10), 3661-3676.
    [CrossRef] [Google Scholar]
  165. Roriz, R., Cabral, J., & Gomes, T. (2021). Automotive LiDAR technology: A survey. IEEE Transactions on Intelligent Transportation Systems, 23(7), 6282-6297.
    [CrossRef] [Google Scholar]
  166. Mielle, M., Magnusson, M., & Lilienthal, A. J. (2019, September). A comparative analysis of radar and lidar sensing for localization and mapping. In 2019 European conference on mobile robots (ECMR) (pp. 1-6). IEEE.
    [CrossRef] [Google Scholar]
  167. Geiger, A., Lenz, P., & Urtasun, R. (2012, June). Are we ready for autonomous driving? the kitti vision benchmark suite. In 2012 IEEE conference on computer vision and pattern recognition (pp. 3354-3361). IEEE.
    [CrossRef] [Google Scholar]
  168. Khan, M. U., Zaidi, S. A. A., Ishtiaq, A., Bukhari, S. U. R., Samer, S., & Farman, A. (2021, July). A comparative survey of lidar-slam and lidar based sensor technologies. In 2021 Mohammad Ali Jinnah University International Conference on Computing (MAJICC) (pp. 1-8). IEEE.
    [CrossRef] [Google Scholar]
  169. Yin, J., Shen, J., Guan, C., Zhou, D., & Yang, R. (2020). Lidar-based online 3d video object detection with graph-based message passing and spatiotemporal transformer attention. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (pp. 11495-11504).
    [CrossRef] [Google Scholar]
  170. Pandharipande, A., Cheng, C. H., Dauwels, J., Gurbuz, S. Z., Ibanez-Guzman, J., Li, G., ... & Santra, A. (2023). Sensing and machine learning for automotive perception: A review. IEEE Sensors Journal, 23(11), 11097-11115.
    [CrossRef] [Google Scholar]
  171. Liang, M., Yang, B., Chen, Y., Hu, R., & Urtasun, R. (2019). Multi-task multi-sensor fusion for 3d object detection. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition (pp. 7345-7353).
    [Google Scholar]
  172. Ouaknine, A., Newson, A., Rebut, J., Tupin, F., & Pérez, P. (2021, January). Carrada dataset: Camera and automotive radar with range-angle-doppler annotations. In 2020 25th International Conference on Pattern Recognition (ICPR) (pp. 5068-5075). IEEE.
    [CrossRef] [Google Scholar]
  173. Mendez, J., Molina, M., Rodriguez, N., Cuellar, M. P., & Morales, D. P. (2021). Camera-LiDAR multi-level sensor fusion for target detection at the network edge. Sensors, 21(12), 3992.
    [CrossRef] [Google Scholar]
  174. Bijelic, M., Gruber, T., & Ritter, W. (2018, June). A benchmark for lidar sensors in fog: Is detection breaking down?. In 2018 IEEE intelligent vehicles symposium (IV) (pp. 760-767). IEEE.
    [Google Scholar]
  175. Wang, Y., Chen, X., You, Y., Li, L. E., Hariharan, B., Campbell, M., ... & Chao, W. L. (2020). Train in germany, test in the usa: Making 3d object detectors generalize. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition (pp. 11713-11723).
    [CrossRef] [Google Scholar]

Cited By (18)

  1. Peng Shi, Jingjing Guo, Lu Deng, Yingkai Liu, Lizhi Long, Shaopeng Xu. Enhanced recognition of low-discernibility Railway Sleeper Serial Numbers via dual-stage adaptive image enhancement and position prior-guided detection. Engineering Applications of Artificial Intelligence, 2026 , 169 .
    [CrossRef]
  2. Yu Wang, Junhao Qu, Ruifeng Liu, Zibo Wang, Qinying Wang, Kunpeng Shi, Qijie Zhou, Shiwen Xie. Near-field seismic wave attenuation and local magnitude calibration based on controlled blasting experiments. Journal of Vibroengineering, 2026 , 28 (5).
    [CrossRef]
  3. Liangliang Li, Guihua Liu, Feng Xu, Wenjin Liao, Wenjie Zhao. Mamba-enhanced multi-view stereo: Geometry-aware feature fusion and 3D cost volume regularization. Applied Soft Computing, 2026 , 195 .
    [CrossRef]
  4. Wenbo Gu, Ruizhe Ma, Yang Yu, Zongmin Ma. FPSO-Time: A Fuzzy Time Series Forecasting Method Based on Fuzzy Particle Swarm Optimization and Reprogrammed Large Language Models. IEEE Transactions on Fuzzy Systems, 2026 , 34 (2).
    [CrossRef]
  5. Lucija Žužić, Franko Hržić, Jonatan Lerga. Collision course detection for personal watercrafts using models based on recurrent neural networks. Knowledge-Based Systems, 2026 , 333 .
    [CrossRef]
  6. Hasan Abbasi, Marzieh Amini, Abdelhamid Mammeri. Comparative Analysis of Sensor Performance in Autonomous Vehicles: Evaluating RGB, Thermal Cameras, and LiDAR Under Adverse Weather Conditions. IEEE Access, 2026 , 14 .
    [CrossRef]
  7. Chenghong Zhang, Wei Wang, Bo Yu, Hanting Wei. A Pseudo-Point-Based Adaptive Fusion Network for Multi-Modal 3D Detection. Electronics, 2025 , 15 (1).
    [CrossRef]
  8. Xin Wen, Xiao Zheng, Yu He. MSCM-Net: Rail Surface Defect Detection Based on a Multi-Scale Cross-Modal Network. Computers, Materials & Continua, 2025 , 82 (3).
    [CrossRef]
  9. Ruixuan Cong, Hao Sheng, Da Yang, Rongshan Chen, Zhenglong Cui. Pseudo 5D hyperspectral light field for image semantic segmentation. Information Fusion, 2025 , 121 .
    [CrossRef]
  10. Yubo Ni, Xiangjun Wang, Zonghua Zhang, Zhaozong Meng, Nan Gao. Position and orientation determination using quadric fitting optimization for object’s motion analyzing. Measurement Science and Technology, 2025 , 36 (2).
    [CrossRef]
  11. Samuel Appleby, Giacomo Bergami, Gary Ushaw. From Camera Image to Active Target Tracking: Modelling, Encoding and Metrical Analysis for Unmanned Underwater Vehicles. AI, 2025 , 6 (4).
    [CrossRef]
  12. Abdullah Nader Alkhater, Ghulam E Mustafa Abro, Ayman M Abdallah. . 2025 21st IEEE International Colloquium on Signal Processing & Its Applications (CSPA), 2025 .
    [CrossRef]
  13. Xinghong Yang, Lizhen Shao. Green and efficient path planning framework based on multi-granularity view fusion for amphibious unmanned aerial vehicles. Engineering Applications of Artificial Intelligence, 2025 , 151 .
    [CrossRef]
  14. Adrieana Maisarah Shamsul Hamran, Ghulam E Mustafa Abro, Nur Batrisyia Bt Zulkafli, Md Mushfiqur Rahman, Asma Ayuni Binti Mohd Akhir, Hifza Mustafa. . 2025 IEEE Symposium on Wireless Technology & Applications (ISWTA), 2025 .
    [CrossRef]
  15. Xin Li, Hui Zhao, Bingxin Xu, Hongzhe Liu, Vincenzo Conti. Wavelet‐Based Texture Mining and Enhancement for Face Forgery Detection. IET Biometrics, 2025 , 2025 (1).
    [CrossRef]
  16. Jianfeng Han, Xiongwei Gao, Lili Song, Jiandong Fang, Yongzhao Tao, Haixin Deng, Jie Yao. A New Algorithm for Visual Navigation in Unmanned Aerial Vehicle Water Surface Inspection. Sensors, 2025 , 25 (8).
    [CrossRef]
  17. Long Zhao, Jinhui Su, Yusheng Zhong, Weiwei Xie, Jinya Su, Xisong Chen, Congyan Chen, Shihua Li. BeltLineNet: A Shape-Prior-Guided Lightweight Network for Real-Time Deviation Detection in Circular Pipe Conveyors. IEEE Sensors Journal, 2025 , 25 (11).
    [CrossRef]
  18. Yingzi Wang, Ce Yu, Xianglei Zhu, Hongcan Gao, Jie Shang. Stacked neural filtering network for reliable NEV monitoring. Displays, 2025 , 87 .
    [CrossRef]
* Citation data provided by Crossref Cited-by.

Cite This Article

APA Style
Abro, G. E. M., Ali, Z. A., & Rajput, S. (2024). Innovations in 3D Object Detection: A Comprehensive Review of Methods, Sensor Fusion, and Future Directions. ICCK Transactions on Sensing, Communication, and Control, 1(1), 3-29. https://doi.org/10.62762/TSCC.2024.989358
Export Citation
RIS Format
Compatible with EndNote, Zotero, Mendeley, and other reference managers
TY  - JOUR
AU  - Abro, Ghulam E Mustafa
AU  - Ali, Zain Anwar
AU  - Rajput, Summaiya
PY  - 2024
DA  - 2024/10/12
TI  - Innovations in 3D Object Detection: A Comprehensive Review of Methods, Sensor Fusion, and Future Directions
JO  - ICCK Transactions on Sensing, Communication, and Control
T2  - ICCK Transactions on Sensing, Communication, and Control
JF  - ICCK Transactions on Sensing, Communication, and Control
VL  - 1
IS  - 1
SP  - 3
EP  - 29
DO  - 10.62762/TSCC.2024.989358
UR  - https://www.icck.org/article/abs/TSCC.2024.989358
KW  - autonomous systems
KW  - camera
KW  - fusion methods
KW  - LiDAR
KW  - object detection
KW  - radar and three-dimensional
AB  - This review paper offers a thorough assessment of three-dimensional object recognition methods, an essential element in the perception frameworks of autonomous systems. This analysis emphasises the integration of LiDAR and camera sensors, providing a distinctive contrast with more economical alternatives like camera-only or camera-Radar combinations. This study objectively evaluates performance and practical implementation issues, such as cost and operational efficiency, thereby elucidating the limitations of existing systems and proposing avenues for further research. The insights provided render it a significant asset for enhancing 3D object recognition and autonomy in intelligent systems.
SN  - 3068-9287
PB  - Institute of Central Computation and Knowledge
LA  - English
ER  - 
BibTeX Format
Compatible with LaTeX, BibTeX, and other reference managers
@article{Abro2024Innovation,
  author = {Ghulam E Mustafa Abro and Zain Anwar Ali and Summaiya Rajput},
  title = {Innovations in 3D Object Detection: A Comprehensive Review of Methods, Sensor Fusion, and Future Directions},
  journal = {ICCK Transactions on Sensing, Communication, and Control},
  year = {2024},
  volume = {1},
  number = {1},
  pages = {3-29},
  doi = {10.62762/TSCC.2024.989358},
  url = {https://www.icck.org/article/abs/TSCC.2024.989358},
  abstract = {This review paper offers a thorough assessment of three-dimensional object recognition methods, an essential element in the perception frameworks of autonomous systems. This analysis emphasises the integration of LiDAR and camera sensors, providing a distinctive contrast with more economical alternatives like camera-only or camera-Radar combinations. This study objectively evaluates performance and practical implementation issues, such as cost and operational efficiency, thereby elucidating the limitations of existing systems and proposing avenues for further research. The insights provided render it a significant asset for enhancing 3D object recognition and autonomy in intelligent systems.},
  keywords = {autonomous systems, camera, fusion methods, LiDAR, object detection, radar and three-dimensional},
  issn = {3068-9287},
  publisher = {Institute of Central Computation and Knowledge}
}

Article Metrics

Citations
Crossref
18
Scopus
22
Views
12100
PDF Downloads
836

Publisher's Note

ICCK stays neutral with regard to jurisdictional claims in published maps and institutional affiliations.

Rights and Permissions

Institute of Central Computation and Knowledge (ICCK) or its licensor holds exclusive rights to this article under a publishing agreement with the author(s) or other rightsholder(s); author self-archiving of the accepted manuscript version of this article is solely governed by the terms of such publishing agreement and applicable law.
ICCK Transactions on Sensing, Communication, and Control
ICCK Transactions on Sensing, Communication, and Control
ISSN: 3068-9287 (Online) | ISSN: 3068-9279 (Print)
Portico
Preserved at
Portico