GeoGaze: A Real-time, Lightweight Gaze Estimation Framework via Geometric Landmark Analysis
Article Information
Abstract
Gaze estimation plays a vital role in human-computer interaction, driver monitoring, and psychological analysis. While state-of-the-art appearance-based methods achieve high accuracy using deep learning, they often demand substantial computational resources, including GPU acceleration and extensive training, limiting their use in resource-constrained or real-time scenarios. This paper introduces GeoGaze, a novel, lightweight, training-free framework that infers categorical gaze direction (“Left”, “Center”, “Right”) solely from geometric analysis of facial landmarks. Leveraging the high-precision 478-point face mesh and iris landmarks provided by MediaPipe, GeoGaze computes a simple normalized iris-to-eye-corner ratio and applies intuitive thresholds, eliminating the need for model training or GPU support. Evaluated on a simulated 1,500-image dataset (SGDD-1500), GeoGaze delivers competitive directional classification accuracy while achieving real-time performance (~66 FPS on CPU), outperforming typical deep learning baselines by more than eight-fold in speed. These results position GeoGaze as an efficient, interpretable alternative for edge devices and applications where precise angular gaze is unnecessary and directional intent suffices.
Graphical Abstract
Keywords
Data Availability Statement
Funding
Conflicts of Interest
AI Use Statement
Ethical Approval and Consent to Participate
References
- Kleinke, C. L. (1986). Gaze and eye contact: A research review. Psychological Bulletin, 100(1), 78.
[CrossRef] [Google Scholar] - Recasens, A., Khosla, A., Vondrick, C., & Torralba, A. (2015). Where are they looking?. Advances in neural information processing systems, 28.
[Google Scholar] - Majaranta, P., & Räihä, K. J. (2002). Twenty years of eye typing: Systems and design issues. In Proceedings of the 2002 symposium on Eye tracking research & applications (pp. 15-22).
[CrossRef] [Google Scholar] - Ji, Q., & Yang, X. (2002). Real-time eye, gaze, and face pose tracking for monitoring driver vigilance. Real-Time Imaging, 8(5), 357-377.
[CrossRef] [Google Scholar] - Deng, T., Yan, H., Qin, L., Ngo, T., & Manjunath, B. S. (2019). How do drivers allocate their potential attention? Driving fixation prediction via convolutional neural networks. IEEE Transactions on Intelligent Transportation Systems, 21(5), 2146-2154.
[CrossRef] [Google Scholar] - Duchowski, A. T. (2017). Eye tracking methodology: Theory and practice (3rd ed.). Springer.
[CrossRef] [Google Scholar] - Jacob, R. J., & Karn, K. S. (2003). Eye tracking in human-computer interaction and usability research: Ready to deliver the promises. In The mind's eye (pp. 573-605). North-Holland.
[CrossRef] [Google Scholar] - Kar, A., & Corcoran, P. (2017). A review and analysis of eye-gaze estimation systems, algorithms and performance evaluation methods in consumer platforms. IEEE Access, 5, 16495-16519.
[CrossRef] [Google Scholar] - Hansen, D. W., & Ji, Q. (2010). In the eye of the beholder: A survey of models for eyes and gaze. IEEE Transactions on Pattern Analysis and Machine Intelligence, 32(3), 478-500.
[CrossRef] [Google Scholar] - Funes-Mora, K. A., & Odobez, J. M. (2016). Gaze estimation in the 3D space using RGB-D sensors: towards head-pose and user invariance. International Journal of Computer Vision, 118(2), 194-216.
[CrossRef] [Google Scholar] - Zhu, Z., & Ji, Q. (2007). Novel eye gaze tracking techniques under natural head movement. IEEE Transactions on biomedical engineering, 54(12), 2246-2260.
[CrossRef] [Google Scholar] - Bhatt, A., Watanabe, K., Dengel, A., & Ishimaru, S. (2024). Appearance-based gaze estimation with deep neural networks: From data collection to evaluation. International Journal of Activity and Behavior Computing, 2024(1), 1-15.
[CrossRef] [Google Scholar] - Lu, F., Sugano, Y., Okabe, T., & Sato, Y. (2014). Adaptive linear regression for appearance-based gaze estimation. IEEE transactions on pattern analysis and machine intelligence, 36(10), 2033-2046.
[CrossRef] [Google Scholar] - Krafka, K., Khosla, A., Kellnhofer, P., Kannan, H., Bhandarkar, S., Matusik, W., & Torralba, A. (2016, June). Eye Tracking for Everyone. In 2016 IEEE Conference on Computer Vision and Pattern Recognition (CVPR) (pp. 2176-2184). IEEE.
[CrossRef] [Google Scholar] - Zhang, X., Sugano, Y., Fritz, M., & Bulling, A. (2015, June). Appearance-based gaze estimation in the wild. In 2015 IEEE Conference on Computer Vision and Pattern Recognition (CVPR) (pp. 4511-4520). IEEE.
[CrossRef] [Google Scholar] - Zhang, X., Sugano, Y., Fritz, M., & Bulling, A. (2017, July). It’s Written All Over Your Face: Full-Face Appearance-Based Gaze Estimation. In 2017 IEEE Conference on Computer Vision and Pattern Recognition Workshops (CVPRW) (pp. 2299-2308). IEEE.
[CrossRef] [Google Scholar] - Fischer, T., Chang, H. J., & Demiris, Y. (2018, September). RT-GENE: Real-Time Eye Gaze Estimation in Natural Environments. In European Conference on Computer Vision (pp. 339-357). Cham: Springer International Publishing.
[CrossRef] [Google Scholar] - Sugano, Y., Matsushita, Y., & Sato, Y. (2014, June). Learning-by-Synthesis for Appearance-Based 3D Gaze Estimation. In 2014 IEEE Conference on Computer Vision and Pattern Recognition (pp. 1821-1828). IEEE.
[CrossRef] [Google Scholar] - Kellnhofer, P., Recasens, A., Stent, S., Matusik, W., & Torralba, A. (2019, October). Gaze360: Physically Unconstrained Gaze Estimation in the Wild. In 2019 IEEE/CVF International Conference on Computer Vision (ICCV) (pp. 6911-6920). IEEE.
[CrossRef] [Google Scholar] - Bodini, M. (2019). A review of facial landmark extraction in 2D images and videos using deep learning. Big Data and Cognitive Computing, 3(1), 14.
[CrossRef] [Google Scholar] - Cheng, Y., Wang, H., Bao, Y., & Lu, F. (2024). Appearance-based gaze estimation with deep learning: A review and benchmark. IEEE Transactions on Pattern Analysis and Machine Intelligence, 46(12), 7509-7528.
[CrossRef] [Google Scholar] - Ye, E. E., Ye, J. E., Ye, J., Ye, J., & Ye, R. (2023). Low-cost Geometry-based Eye Gaze Detection using Facial Landmarks Generated through Deep Learning. arXiv preprint arXiv:2401.00406.
[CrossRef] [Google Scholar] - Wood, E., Baltrušaitis, T., Morency, L. P., Robinson, P., & Bulling, A. (2016, March). Learning an appearance-based gaze estimator from one million synthesised images. In Proceedings of the ninth biennial ACM symposium on eye tracking research & applications (pp. 131-138).
[CrossRef] [Google Scholar] - Sesma, L., Villanueva, A., & Cabeza, R. (2012, March). Evaluation of pupil center-eye corner vector for gaze estimation using a web cam. In Proceedings of the symposium on eye tracking research and applications (pp. 217-220).
[CrossRef] [Google Scholar] - Valenti, R., Sebe, N., & Gevers, T. (2011). Combining head pose and eye location information for gaze estimation. IEEE Transactions on Image Processing, 21(2), 802-815.
[CrossRef] [Google Scholar] - Yan, G., & Grishchenko, I. (2022). Model Card: MediaPipe Face Mesh V2. Google. Retrieved from https://storage.googleapis.com/mediapipe-assets/Model%20Card%20MediaPipe%20Face%20Mesh%20V2.pdf
[Google Scholar] - Zhao, R., Wang, Y., Luo, S., Shou, S., & Tang, P. (2024). Gaze-swin: Enhancing gaze estimation with a hybrid cnn-transformer network and dropkey mechanism. Electronics, 13(2), 328.
[CrossRef] [Google Scholar] - Chen, J., & Ji, Q. (2011, June). Probabilistic gaze estimation without active personal calibration. In CVPR 2011 (pp. 609-616). IEEE.
[CrossRef] [Google Scholar] - Huang, Q., Veeraraghavan, A., & Sabharwal, A. (2017). Tabletgaze: dataset and analysis for unconstrained appearance-based gaze estimation in mobile tablets. Machine Vision and Applications, 28(5), 445-461.
[CrossRef] [Google Scholar] - Wang, K., & Ji, Q. (2017, October). Real Time Eye Gaze Tracking with 3D Deformable Eye-Face Model. In 2017 IEEE International Conference on Computer Vision (ICCV) (pp. 1003-1011). IEEE.
[CrossRef] [Google Scholar] - Bazarevsky, V., Kartynnik, Y., Vakunov, A., Raveendran, K., & Grundmann, M. (2019). Blazeface: Sub-millisecond neural face detection on mobile gpus. arXiv preprint arXiv:1907.05047.
[CrossRef] [Google Scholar] - Zhang, X., Park, S., Beeler, T., Bradley, D., Tang, S., & Hilliges, O. (2020, August). Eth-xgaze: A large scale dataset for gaze estimation under extreme head pose and gaze variation. In European conference on computer vision (pp. 365-381). Cham: Springer International Publishing.
[CrossRef] [Google Scholar]
Cited By (1)
-
Mengyun Wang, Manru Xun, Qinglin Yun. Fine-Grained Human Action Recognition Via Laban Movement Analysis and Attentive Bi-LSTM.
International Journal of Pattern Recognition and Artificial Intelligence, 2026 .
[CrossRef]
Cite This Article
TY - JOUR AU - Khalid, Muhammad Imran AU - Komal, Asma AU - Hussain, Nasir AU - Idrees, Muhammad AU - Wagan, Atif Ali AU - Hussain, Syed Akif PY - 2026 DA - 2026/02/11 TI - GeoGaze: A Real-time, Lightweight Gaze Estimation Framework via Geometric Landmark Analysis JO - ICCK Transactions on Advanced Computing and Systems T2 - ICCK Transactions on Advanced Computing and Systems JF - ICCK Transactions on Advanced Computing and Systems VL - 2 IS - 2 SP - 107 EP - 115 DO - 10.62762/TACS.2025.798133 UR - https://www.icck.org/article/abs/TACS.2025.798133 KW - gaze estimation KW - geometric landmark analysis KW - facial landmarks KW - MediaPipe KW - real-time inference KW - training-free AB - Gaze estimation plays a vital role in human-computer interaction, driver monitoring, and psychological analysis. While state-of-the-art appearance-based methods achieve high accuracy using deep learning, they often demand substantial computational resources, including GPU acceleration and extensive training, limiting their use in resource-constrained or real-time scenarios. This paper introduces GeoGaze, a novel, lightweight, training-free framework that infers categorical gaze direction (“Left”, “Center”, “Right”) solely from geometric analysis of facial landmarks. Leveraging the high-precision 478-point face mesh and iris landmarks provided by MediaPipe, GeoGaze computes a simple normalized iris-to-eye-corner ratio and applies intuitive thresholds, eliminating the need for model training or GPU support. Evaluated on a simulated 1,500-image dataset (SGDD-1500), GeoGaze delivers competitive directional classification accuracy while achieving real-time performance (~66 FPS on CPU), outperforming typical deep learning baselines by more than eight-fold in speed. These results position GeoGaze as an efficient, interpretable alternative for edge devices and applications where precise angular gaze is unnecessary and directional intent suffices. SN - 3068-7969 PB - Institute of Central Computation and Knowledge LA - English ER -
@article{Khalid2026GeoGaze,
author = {Muhammad Imran Khalid and Asma Komal and Nasir Hussain and Muhammad Idrees and Atif Ali Wagan and Syed Akif Hussain},
title = {GeoGaze: A Real-time, Lightweight Gaze Estimation Framework via Geometric Landmark Analysis},
journal = {ICCK Transactions on Advanced Computing and Systems},
year = {2026},
volume = {2},
number = {2},
pages = {107-115},
doi = {10.62762/TACS.2025.798133},
url = {https://www.icck.org/article/abs/TACS.2025.798133},
abstract = {Gaze estimation plays a vital role in human-computer interaction, driver monitoring, and psychological analysis. While state-of-the-art appearance-based methods achieve high accuracy using deep learning, they often demand substantial computational resources, including GPU acceleration and extensive training, limiting their use in resource-constrained or real-time scenarios. This paper introduces GeoGaze, a novel, lightweight, training-free framework that infers categorical gaze direction (“Left”, “Center”, “Right”) solely from geometric analysis of facial landmarks. Leveraging the high-precision 478-point face mesh and iris landmarks provided by MediaPipe, GeoGaze computes a simple normalized iris-to-eye-corner ratio and applies intuitive thresholds, eliminating the need for model training or GPU support. Evaluated on a simulated 1,500-image dataset (SGDD-1500), GeoGaze delivers competitive directional classification accuracy while achieving real-time performance (~66 FPS on CPU), outperforming typical deep learning baselines by more than eight-fold in speed. These results position GeoGaze as an efficient, interpretable alternative for edge devices and applications where precise angular gaze is unnecessary and directional intent suffices.},
keywords = {gaze estimation, geometric landmark analysis, facial landmarks, MediaPipe, real-time inference, training-free},
issn = {3068-7969},
publisher = {Institute of Central Computation and Knowledge}
}
Publisher's Note
ICCK stays neutral with regard to jurisdictional claims in published maps and institutional affiliations.
Rights and Permissions
Copyright © 2026 by the Author(s). Published by Institute of Central Computation and Knowledge. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license (https://creativecommons.org/licenses/by/4.0/), which permits use, sharing, adaptation, distribution and reproduction in any medium or format, as long as you give appropriate credit to the original author(s) and the source, provide a link to the Creative Commons licence, and indicate if changes were made.
Portico