Lightweight YOLO Object Detection Algorithms for Vision-Based Robotic Arms: A Review
Article Information
Abstract
To address the requirements of vision-based robotic manipulators for object detection accuracy, real-time performance, and lightweight deployment, this paper reviews lightweight You Only Look Once (YOLO) object detection techniques and their applications in robotic vision. First, the main techniques, including structural lightweighting, model compression, and edge deployment, are summarized. Representative algorithms for vision-based robotic manipulators are then categorized into three groups: backbone lightweighting, local structural optimization, and model compression, with their detection performance and lightweighting characteristics analyzed. On this basis, the limitations of existing methods are discussed in terms of the trade-off between accuracy and efficiency, cross-scenario generalization, coordination between visual perception and grasp execution, and comprehensive evaluation. Finally, future research directions are outlined, including multi-strategy collaborative optimization, cross-domain adaptation, and system-level evaluation for robotic manipulators.
Graphical Abstract
Keywords
Data Availability Statement
Funding
Conflicts of Interest
AI Use Statement
Ethical Approval and Consent to Participate
References
- Du, G., Wang, K., Lian, S., & Zhao, K. (2021). Vision-based robotic grasping from object localization, object pose estimation to grasp estimation for parallel grippers: a review. Artificial Intelligence Review, 54(3), 1677-1734.
[CrossRef] [Google Scholar] - Redmon, J., Divvala, S., Girshick, R., & Farhadi, A. (2016). You only look once: Unified, real-time object detection. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (pp. 779--788).
[CrossRef] [Google Scholar] - Terven, J., C{\'ordova-Esparza, D.-M., & Romero-Gonz{\'alez, J.-A. (2023). A comprehensive review of YOLO architectures in computer vision: From YOLOv1 to YOLOv8 and YOLO-NAS. Machine Learning and Knowledge Extraction, 5(4), 1680--1716.
[CrossRef] [Google Scholar] - Wang, A., Chen, H., Liu, L., Chen, K., Lin, Z., Han, J., & Ding, G. (2024). YOLOv10: Real-time end-to-end object detection. Advances in Neural Information Processing Systems, 37, 107984-108011.
[CrossRef] [Google Scholar] - Tian, Y., Ye, Q., & Doermann, D. (2026). YOLOv12: Attention-centric real-time object detectors. Advances in Neural Information Processing Systems, 38, 78433-78457.
[CrossRef] [Google Scholar] - Wang, L., Wang, H., Letchmunan, S., Xiao, R., Ahmed, O. H., & Liu, Z. (2025). A systematic literature review of lightweight YOLO models for object detection. PeerJ Computer Science, 11, e3357.
[CrossRef] [Google Scholar] - Sandler, M., Howard, A., Zhu, M., Zhmoginov, A., & Chen, L. C. (2018, June). Mobilenetv2: Inverted residuals and linear bottlenecks. In 2018 IEEE/CVF conference on computer vision and pattern recognition (pp. 4510-4520). IEEE.
[CrossRef] [Google Scholar] - Zhang, X., Zhou, X., Lin, M., & Sun, J. (2018, June). Shufflenet: An extremely efficient convolutional neural network for mobile devices. In 2018 IEEE/CVF conference on computer vision and pattern recognition (pp. 6848-6856). IEEE.
[CrossRef] [Google Scholar] - Han, K., Wang, Y., Tian, Q., Guo, J., Xu, C., & Xu, C. (2020, June). Ghostnet: More features from cheap operations. In 2020 IEEE/CVF conference on computer vision and pattern recognition (CVPR) (pp. 1577-1586). IEEE.
[CrossRef] [Google Scholar] - Ganesh, P., Chen, Y., Yang, Y., Chen, D., & Winslett, M. (2022, January). YOLO-ReT: Towards high accuracy real-time object detection on edge GPUs. In 2022 IEEE/CVF Winter Conference on Applications of Computer Vision (WACV) (pp. 1311-1321). IEEE.
[CrossRef] [Google Scholar] - Li, H., Li, J., Wei, H., Liu, Z., Zhan, Z., & Ren, Q. (2024). Slim-neck by GSConv: A lightweight-design for real-time detector architectures. Journal of real-time image processing, 21(3), 62.
[CrossRef] [Google Scholar] - Liu, Y., Huang, Z., Song, Q., & Bai, K. (2025). PV-YOLO: A lightweight pedestrian and vehicle detection model based on improved YOLOv8. Digital Signal Processing, 156, 104857.
[CrossRef] [Google Scholar] - Bacea, D. S., & Oniga, F. (2026). Boosting lightweight object detection with enhanced feature fusion and optimized receptive field. Neurocomputing, 132728.
[CrossRef] [Google Scholar] - Baicheng, Y. E., Youpan, Z. H. U., Yongkang, Z. H. O. U., Chenhao, D. U. A. N., Yudong, Z. H. A. N. G., Zhigang, T. A. O., & Zhiyu, F. U. (2025). Review of Lightweight Target Detection Algorithms. Infrared Technology, 47(3), 289-298. http://hwjs.nvir.cn/en/article/pdf/preview/df48da80-f152-4563-95fb-801ceebcbded.pdf
[Google Scholar] - Liu, Z., Li, J., Shen, Z., Huang, G., Yan, S., & Zhang, C. (2017, October). Learning efficient convolutional networks through network slimming. In 2017 IEEE international conference on computer vision (ICCV) (pp. 2755-2763). IEEE.
[CrossRef] [Google Scholar] - Fang, G., Ma, X., Song, M., Mi, M. B., & Wang, X. (2023, June). Depgraph: Towards any structural pruning. In 2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) (pp. 16091-16101). IEEE.
[CrossRef] [Google Scholar] - Salah, O. B. H., Messaoud, S., Hajjaji, M. A., Atri, M., & Liouane, N. (2026). Low-latency QYOLOv10-based FPGA implementation for real-time object detection. Integration, 106, 102591.
[CrossRef] [Google Scholar] - Tamura, K., & Todo, Y. (2026). Dual-head self-distillation via an auxiliary lightweight neck for YOLO-based object detection. Neurocomputing, 686, 133750.
[CrossRef] [Google Scholar] - Hyung, J. W., Na, W., Won, K., & Kim, D. J. (2025). Robotic arm control by augmented reality-assisted object detection. Scientific reports, 15(1), 35678.
[CrossRef] [Google Scholar] - Han, T., & Yu, D. (2024). Jensen–Shannon divergence you only look once: a real‐time robotic grasp detection network. Advanced Intelligent Systems, 6(5), 2300497.
[CrossRef] [Google Scholar] - Fan, C., Li, J., Zhang, Z., Li, F., & Wang, B. (2026). Two-phase collaborative model compression training for joint pruning and quantization. Neural Networks, 197, 108506.
[CrossRef] [Google Scholar] - Wang, X., Tang, Z., Guo, J., Meng, T., Wang, C., Wang, T., & Jia, W. (2025). Empowering edge intelligence: A comprehensive survey on on-device AI models. ACM Computing Surveys, 57(9), 1-39.
[CrossRef] [Google Scholar] - Song, Q., Li, S., Bai, Q., Yang, J., Zhang, X., Li, Z., & Duan, Z. (2021). Object detection method for grasping robot based on improved YOLOv5. Micromachines, 12(11), 1273.
[CrossRef] [Google Scholar] - Almaliki, H. H., Mazinan, A. H., & Modaresi, S. M. (2026). Design and implementation of a 6-DoF robot arm control with object detection based on machine learning using mini microcontroller. Scientific Reports, 16(1), 6842.
[CrossRef] [Google Scholar] - Kong, C., Li, F., Yan, X., Yang, J., Mo, P., Luo, Q., & Mao, R. (2026). Object detection on low-compute edge SoCs: A reproducible benchmark and deployment guidelines. Scientific Reports, 16(1), 5875.
[CrossRef] [Google Scholar] - Khan, T. M., Ul Haq, Q. E., Iqbal, S., & Soomro, T. A. (2026). Edge-based artificial intelligence: Understanding the evolution of hardware and software and future trends. Engineering Applications of Artificial Intelligence, 174, 114526.
[CrossRef] [Google Scholar] - Wang, A., Jiao, H., Chen, Z., & Yang, J. (2025). An improved YOLO-based waste detection model and its integration to robotic gripping systems. Computers, Materials \textnormal{& Continua, 84(3), 5773--5790.
[CrossRef] [Google Scholar] - Howard, A., Sandler, M., Chu, G., Chen, L. C., Chen, B., Tan, M., Wang, W., Zhu, Y., Pang, R., Vasudevan, V., Le, Q. V., & Adam, H. (2019). Searching for MobileNetV3. In 2019 IEEE/CVF International Conference on Computer Vision (ICCV) (pp. 1314-1324). IEEE.
[CrossRef] [Google Scholar] - Zhang, P., Dai, N., Liu, X., Yuan, J., & Xin, Z. (2024). A novel lightweight model HGCA-YOLO: Application to recognition of invisible spears for white asparagus robotic harvesting. Computers and Electronics in Agriculture, 220, 108852.
[CrossRef] [Google Scholar] - Chen, J., Chen, H., Xu, F., Lin, M., Zhang, D., & Zhang, L. (2024). Real-time detection of mature table grapes using ESP-YOLO network on embedded platforms. Biosystems Engineering, 246, 122--134.
[CrossRef] [Google Scholar] - Chen, Y., Shu, A., Liu, Z., Chen, Y., Lee, W. S., & Zhang, Y. (2025). SP-RTSD: A lightweight real-time strawberry detection on edge devices for onboard robotic harvesting. Journal of Field Robotics, 42(7), 3361--3379.
[CrossRef] [Google Scholar] - Jiao, J., Li, Z., Xia, G., Wang, G., Chen, Y., & Gao, R. (2026). A lightweight object detection approach for precision gripping in multiple peg-in-hole assembly tasks. Robotics and Computer-Integrated Manufacturing, 98, 103185.
[CrossRef] [Google Scholar] - Cao, H., Yang, S., Zhang, Y., Zhang, H., & Ye, L. (2026). SRF-YOLO: A lightweight detection model and robotic system integration for protected tomato harvesting. Smart Agricultural Technology, 14, 102036.
[CrossRef] [Google Scholar] - Dong, J., Deng, L., Wan, D., Liu, C., Yin, J., Guo, M., Zhang, H., Lin, S., Liu, H., & Liu, L. (2026). Oriented bounding box detection algorithm for dense scenarios of robotic arm operation. Expert Systems with Applications, 298, 129678.
[CrossRef] [Google Scholar] - Chiu, Y. J., Yuan, Y. Y., & Jian, S. R. (2024). Design of and research on the robot arm recovery grasping system based on machine vision. Journal of King Saud University Computer and Information Sciences, 36(4), 102014.
[CrossRef] [Google Scholar] - Han, D., Li, H., Li, Y., & Chen, S. (2026). Real-Time Target-Oriented Grasping Framework for Resource-Constrained Robots. Sensors, 26(2), 645.
[CrossRef] [Google Scholar] - He, J., Jiang, J., & Zhang, C. (2026). A survey of lightweight methods for object detection networks. Array, 29, 100589.
[CrossRef] [Google Scholar] - Song, X., Li, Y., Zhang, Y., Liu, Y., & Jiang, L. (2025). An overview of learning-based dexterous grasping: recent advances and future directions. Artificial Intelligence Review, 58(10), 300.
[CrossRef] [Google Scholar] - Wang, Q., Tu, Y., Xu, W., Zhang, J., Knoll, A., Zhou, M., & Ying, Y. (2025). Towards Damage‐Less Robotic Fragile Fruit Grasping: A Systematic Review on System Design, End Effector, and Visual and Tactile Feedback. Journal of Field Robotics, 42(8), 4521-4543.
[CrossRef] [Google Scholar] - Le, Q. N. N., Chembakasseril, M. T., & Hartanto, R. (2026). Benchmark analysis of deep learning algorithms for edge-based robotic applications. Array, 100818.
[CrossRef] [Google Scholar]
Cite This Article
TY - JOUR AU - Guo, Fajun PY - 2026 DA - 2026/09/20 TI - Lightweight YOLO Object Detection Algorithms for Vision-Based Robotic Arms: A Review JO - ICCK Transactions on Intelligent Cyber-Physical Systems T2 - ICCK Transactions on Intelligent Cyber-Physical Systems JF - ICCK Transactions on Intelligent Cyber-Physical Systems VL - 1 IS - 3 SP - 97 EP - 103 DO - 10.62762/TICPS.2026.107157 UR - https://www.icck.org/article/abs/TICPS.2026.107157 KW - YOLO KW - object detection KW - lightweight model KW - vision-based robotic manipulator AB - To address the requirements of vision-based robotic manipulators for object detection accuracy, real-time performance, and lightweight deployment, this paper reviews lightweight You Only Look Once (YOLO) object detection techniques and their applications in robotic vision. First, the main techniques, including structural lightweighting, model compression, and edge deployment, are summarized. Representative algorithms for vision-based robotic manipulators are then categorized into three groups: backbone lightweighting, local structural optimization, and model compression, with their detection performance and lightweighting characteristics analyzed. On this basis, the limitations of existing methods are discussed in terms of the trade-off between accuracy and efficiency, cross-scenario generalization, coordination between visual perception and grasp execution, and comprehensive evaluation. Finally, future research directions are outlined, including multi-strategy collaborative optimization, cross-domain adaptation, and system-level evaluation for robotic manipulators. SN - 3071-2947 PB - Institute of Central Computation and Knowledge LA - English ER -
@article{Guo2026Lightweigh,
author = {Fajun Guo},
title = {Lightweight YOLO Object Detection Algorithms for Vision-Based Robotic Arms: A Review},
journal = {ICCK Transactions on Intelligent Cyber-Physical Systems},
year = {2026},
volume = {1},
number = {3},
pages = {97-103},
doi = {10.62762/TICPS.2026.107157},
url = {https://www.icck.org/article/abs/TICPS.2026.107157},
abstract = {To address the requirements of vision-based robotic manipulators for object detection accuracy, real-time performance, and lightweight deployment, this paper reviews lightweight You Only Look Once (YOLO) object detection techniques and their applications in robotic vision. First, the main techniques, including structural lightweighting, model compression, and edge deployment, are summarized. Representative algorithms for vision-based robotic manipulators are then categorized into three groups: backbone lightweighting, local structural optimization, and model compression, with their detection performance and lightweighting characteristics analyzed. On this basis, the limitations of existing methods are discussed in terms of the trade-off between accuracy and efficiency, cross-scenario generalization, coordination between visual perception and grasp execution, and comprehensive evaluation. Finally, future research directions are outlined, including multi-strategy collaborative optimization, cross-domain adaptation, and system-level evaluation for robotic manipulators.},
keywords = {YOLO, object detection, lightweight model, vision-based robotic manipulator},
issn = {3071-2947},
publisher = {Institute of Central Computation and Knowledge}
}
Article Metrics
Publisher's Note
ICCK stays neutral with regard to jurisdictional claims in published maps and institutional affiliations.
Rights and Permissions
Portico