Evaluation Metrics in Digital Image Processing: A Task-Oriented Review and Practical Selection Framework
Article Information
Abstract
The rapid development of digital image processing and machine learning has increased the need for reliable and task-appropriate evaluation methods. However, the wide range of available metrics, each with different assumptions, strengths, and limitations, makes metric selection challenging across diverse image-processing tasks. This article presents a task-oriented review of commonly used evaluation metrics and proposes a practical framework for selecting appropriate metrics across five major categories: image quality assessment, image classification, image segmentation, image restoration, and object detection. Representative metrics are reviewed based on their mathematical basis, interpretation, applications, advantages, and limitations. Particular attention is given to pixel-wise measures, including MAE, MSE, and RMSE, perceptual and structural measures such as PSNR, SSIM, MS-SSIM, FSIM, VIF, and LPIPS, and classification and detection measures including accuracy, precision, recall, F1-score, ROC-AUC, confusion matrices, and mean average precision. Overlap-based metrics such as Dice and IoU are also examined for segmentation. The review emphasizes that no single metric adequately captures all dimensions of image-processing performance. Therefore, metric selection should consider task characteristics, data properties, reference availability, and perceptual requirements, with complementary metrics used where appropriate. The proposed framework aims to support researchers and practitioners in selecting appropriate and interpretable evaluation metrics for digital image processing and machine learning applications.
Graphical Abstract
Keywords
Data Availability Statement
Funding
Conflicts of Interest
AI Use Statement
Ethical Approval and Consent to Participate
References
- Trigka, M., & Dritsas, E. (2025). A comprehensive survey of deep learning approaches in image processing. Sensors, 25(2), 531.
[CrossRef] [Google Scholar] - Valente, J., António, J., Mora, C., & Jardim, S. (2023). Developments in image processing using deep learning and reinforcement learning. Journal of imaging, 9(10), 207.
[CrossRef] [Google Scholar] - Archana, R., & Jeevaraj, P. S. E. (2024). Deep learning models for digital image processing: A review. Artificial Intelligence Review, 57(1), 11.
[CrossRef] [Google Scholar] - Gomathi, S., & Roopa Chandrika, R. (2025). Advancing medical image processing with deep learning: Innovations and impact. ICTACT Journal on Image and Video Processing, 15(3), 3489-3494.
[CrossRef] [Google Scholar] - Alnaggar, O. A. M. F., Jagadale, B. N., Saif, M. A. N., Ghaleb, O. A., Ahmed, A. A., Aqlan, H. A. A., & Al-Ariki, H. D. E. (2024). Efficient artificial intelligence approaches for medical image processing in healthcare: comprehensive review, taxonomy, and analysis. Artificial Intelligence Review, 57(8), 221.
[CrossRef] [Google Scholar] - Gorai, S. K., Sarangi, A., & Pradhan, S. (2025). Deep learning for image classification: Methods, challenges, and future directions. International Journal of Scientific Research in Computer Science, Engineering and Information Technology, 11, 484-496.
[CrossRef] [Google Scholar] - Saravanan, V., Adaikkammai, A., Sasikala, G., Iyswariya, A., Babisha, A., & Latha, M. (2024). Advancements in artificial intelligence for better image analysis, recognition, and interpretation in the field of image processing. In 2024 Global Conference on Communications and Information Technologies (GCCIT) (pp. 1-6). IEEE.
[CrossRef] [Google Scholar] - Jiao, L., & Zhao, J. (2019). A survey on the new generation of deep learning in image processing. IEEE Access, 7, 172231-172263.
[CrossRef] [Google Scholar] - Zhu, X. X., Tuia, D., Mou, L., Xia, G. S., Zhang, L., Xu, F., & Fraundorfer, F. (2017). Deep learning in remote sensing: A comprehensive review and list of resources. IEEE Geoscience and Remote Sensing Magazine, 5(4), 8-36.
[CrossRef] [Google Scholar] - Hsieh, W., Bi, Z., Liu, J., Peng, B., Zhang, S., Pan, X., ... & Liu, M. (2024). Deep learning, machine learning - Digital signal and image processing: From theory to application. arXiv preprint arXiv:2410.20304.
[CrossRef] [Google Scholar] - Shorten, C., & Khoshgoftaar, T. M. (2019). A survey on image data augmentation for deep learning. Journal of big data, 6(1), 60.
[CrossRef] [Google Scholar] - Varga, D. (2025). AI-driven digital image processing: Advancements, challenges, and future directions. Electronics, 14(9), 1754.
[CrossRef] [Google Scholar] - Chen, J. (2025). Research on image recognition model based on deep learning. Applied and Computational Engineering, 177, 159-165.
[CrossRef] [Google Scholar] - Wanjari, A., & Verma, S. (2025). A review of machine learning and deep learning-based approaches for image recognition and its applications. In 2025 4th International Conference on Sentiment Analysis and Deep Learning (ICSADL) (pp. 459-464). IEEE.
[CrossRef] [Google Scholar] - Gui, X., Wang, W., & Tao, D. (2024). A survey of self-supervised learning from multiple perspectives: Algorithms, theory, applications and future trends. IEEE Transactions on Pattern Analysis and Machine Intelligence, 47(1), 248-267.
[CrossRef] [Google Scholar] - Eapen, B. T., & George, S. T. (2024). Exploring self-supervised learning architectures for image processing. International Journal on Recent and Innovation Trends in Computing and Communication, 12(2), 358-372.
[CrossRef] [Google Scholar] - Wang, Y., Lv, F., Li, S., Liu, X., & Peng, H. (2025). A comparative survey of vision transformers for image classification. Neurocomputing, 629, 129752.
[CrossRef] [Google Scholar] - Khan, S., Naseer, M., Hayat, M., Zamir, S. W., Khan, F. S., & Shah, M. (2022). Transformers in vision: A survey. ACM Computing Surveys, 54(10s), 1-41.
[CrossRef] [Google Scholar] - Saeed, W., & Omlin, C. (2023). Explainable AI (XAI): A systematic meta-survey of current challenges and future opportunities. Knowledge-Based Systems, 263, 110273.
[CrossRef] [Google Scholar] - Akhtar, N. (2023). Revolutionizing digital image processing: Advancements and applications of deep visual models explainability. arXiv preprint arXiv:2301.13445.
[CrossRef] [Google Scholar] - Salmanpour, M. R., Shamsaei, M., Hajianfar, G., Askari, A., Maghsudi, M., Bloise, I., ... & Rahmim, A. (2025). Machine learning metrics: Assessing performance differences across modalities and tasks in medical imaging. Physica Medica, 130, 104890.
[CrossRef] [Google Scholar] - Rainio, O., Teuho, J., & Klén, R. (2024). Evaluation metrics and statistical tests for machine learning. Scientific reports, 14(1), 6086.
[CrossRef] [Google Scholar] - Zhai, G., & Min, X. (2020). Perceptual image quality assessment: a survey. Science China Information Sciences, 63(11), 211301.
[CrossRef] [Google Scholar] - Wang, Z., Bovik, A. C., Sheikh, H. R., & Simoncelli, E. P. (2004). Image quality assessment: from error visibility to structural similarity. IEEE Transactions on Image Processing, 13(4), 600-612.
[CrossRef] [Google Scholar] - Urrea, C., & Vélez, J. M. (2025). Advances in semantic segmentation using deep learning techniques: A systematic review. IEEE Access, 13, 4501-4523.
[CrossRef] [Google Scholar] - Cho, J. (2024). Weighted Intersection over Union (wIoU): A metric for evaluating image segmentation. IEEE Access, 12, 116030-116045.
[CrossRef] [Google Scholar] - Taciuc, I. A., Dandu, A. R., Badiger, R., Kalyana Sundaram, M. N., & Marinas, M. C. (2025). Enhancing AI-driven detection of lymph nodes in ultrasound imaging: Comparative analysis of DICE and IoU metrics. Sensors, 25(9), 2727.
[CrossRef] [Google Scholar] - Merkulova, N., & Jayakumar, A. (2025). Evaluation framework for image segmentation. arXiv preprint arXiv:2504.04435.
[CrossRef] [Google Scholar] - Chin, S. P., Yap, W. S., Tee, C. A. T. H., & Kim, Y. J. (2025). Challenges and evaluation metrics in medical image enhancement. arXiv preprint arXiv:2510.13638.
[CrossRef] [Google Scholar] - Russakovsky, O., Deng, J., Su, H., Krause, J., Satheesh, S., Ma, S., ... & Fei-Fei, L. (2015). Imagenet large scale visual recognition challenge. International Journal of Computer Vision, 115, 211-252.
[CrossRef] [Google Scholar] - Lin, T. Y., Maire, M., Belongie, S., Hays, J., Perona, P., Ramanan, D., ... & Zitnick, C. L. (2014). Microsoft coco: Common objects in context. In Computer Vision–ECCV 2014: 13th European Conference, Zurich, Switzerland, September 6-12, 2014, Proceedings, Part V 13 (pp. 740-755). Springer International Publishing.
[CrossRef] [Google Scholar] - Caron, M., Touvron, H., Misra, I., Jégou, H., Mairal, J., Bojanowski, P., & Joulin, A. (2021). Emerging properties in self-supervised vision transformers. In Proceedings of the IEEE/CVF international conference on computer vision (pp. 9650-9660). IEEE.
[CrossRef] [Google Scholar] - Shao, R., Shi, Z., Du, J., Tang, X., & Qiao, Y. (2022). On the open-world classification challenge in image classification. arXiv preprint arXiv:2202.00179.
[CrossRef] [Google Scholar] - Selvaraju, R. R., Cogswell, M., Das, A., Vedantam, R., Parikh, D., & Batra, D. (2017). Grad-CAM: Visual explanations from deep networks via gradient-based localization. In Proceedings of the IEEE International Conference on Computer Vision (pp. 618-626). IEEE.
[CrossRef] [Google Scholar] - Dong, Y., Fu, Q. A., Yang, X., Pang, T., Su, H., Xiao, Z., & Zhu, J. (2020). Benchmarking adversarial robustness on image classification. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (pp. 321-331). IEEE.
[CrossRef] [Google Scholar] - Wang, J., Ai, J., Lu, M., Jiang, J., Yang, T., Bai, Z., ... & Wang, D. (2026). A survey of neural network robustness assessment in image recognition. ACM Computing Surveys, 58(12), 1-40.
[CrossRef] [Google Scholar] - Pokkuluri, K. S., Latha, Y. M., Kumar, P. S., & Narayana, B. L. (2024). Enhancing the accuracy of image segmentation using deep learning techniques. International Journal of Intelligent Systems and Applications in Engineering, 12(4), 474-482. https://ijisae.org/index.php/IJISAE/article/view/6128
[Google Scholar] - Dehdashtian, M., Rafiei, M. H., & Adeli, H. (2024). Fairness and bias mitigation in computer vision: A survey. arXiv preprint arXiv:2408.02464.
[CrossRef] [Google Scholar] - Wang, T., Zhao, J., Yatskar, M., Chang, K. W., & Ordonez, V. (2020). Balanced datasets are not enough: Estimating and mitigating gender bias in deep image representations. In Proceedings of the IEEE/CVF International Conference on Computer Vision (pp. 5310-5319). IEEE.
[CrossRef] [Google Scholar] - Hänel, A., Tschernezki, V., Wiegand, T., & Samek, W. (2022). Enhancing the fairness of visual attribute predictors. In Proceedings of the Asian Conference on Computer Vision (pp. 2707-2724).
[Google Scholar] - Halder, A., Gharami, S., Sadhu, P., Singh, P. K., Woźniak, M., & Ijaz, M. F. (2024). Implementing vision transformer for classifying 2D biomedical images. Scientific Reports, 14(1), 12567.
[CrossRef] [Google Scholar] - Sara, U., Akter, M., & Uddin, M. S. (2019). Image quality assessment through FSIM, SSIM, MSE and PSNR—a comparative study. Journal of Computer and Communications, 7(3), 8-18.
[CrossRef] [Google Scholar] - Al Najjar, Y. (2024). Comparative analysis of image quality assessment metrics: PSNR, SSIM, FSIM, and NAE. International Journal of Advanced Natural Sciences and Engineering Researches, 8(4), 128-133.
[CrossRef] [Google Scholar] - Mason, A., Rioux, J., Clarke, S. E., Costa, A., Schmidt, M., Keough, V., ... & Beyea, S. (2020). Comparison of objective image quality metrics to expert radiologists' scoring of diagnostic quality of MR images. IEEE Transactions on Medical Imaging, 39(4), 1064-1072.
[CrossRef] [Google Scholar] - Arabboev, M., Muksimov, A., Mukhtorov, D., Muminov, B., Abdullaev, S., & Djuraev, O. (2024). A comprehensive survey on image super-resolution: Evaluation metrics from signal processing to deep learning perspectives. Journal of Imaging, 10(10), 252.
[CrossRef] [Google Scholar] - Sharma, A. (2024). Image quality assessment using distortion metrics for hyperspectral image compression. Remote Sensing Letters, 15(9), 899-910.
[CrossRef] [Google Scholar] - Sokolova, M., & Lapalme, G. (2009). A systematic analysis of performance measures for classification tasks. Information Processing & Management, 45(4), 427-437.
[CrossRef] [Google Scholar] - Khan, A., Rauf, Z., Sohail, A., Khan, A. R., Asif, H., Asif, A., & Farooq, U. (2024). Early disease identification of rice and wheat crops using a transfer learning approach on smart handheld devices. Frontiers in Plant Science, 15, 1383748.
[CrossRef] [Google Scholar] - Kim, H., Seo, H., & Kim, E. (2025). A systematic review of hybrid vision transformers for radiological image analysis: Architectures, clinical applications, and future directions. Diagnostics, 15(2), 221.
[CrossRef] [Google Scholar] - Vlăsceanu, G., Botezan, A., Dinu, A., & Acatrinei, C. (2024). Selecting metrics to evaluate image segmentation performance. Applied Sciences, 14(22), 10605.
[CrossRef] [Google Scholar] - Müller, D., Soto-Rey, I., & Kramer, F. (2022). Towards a guideline for evaluation metrics in medical image segmentation. BMC Research Notes, 15(1), 210.
[CrossRef] [Google Scholar] - Taha, A. A., & Hanbury, A. (2015). Metrics for evaluating 3D medical image segmentation: analysis, selection, and tool. BMC Medical Imaging, 15(1), 29.
[CrossRef] [Google Scholar] - Khan, A., Sohail, A., Ali, A., & Khan, A. (2024). Hybrid deep learning model for efficient classification and segmentation of chest X-ray images to detect tuberculosis. Signal, Image and Video Processing, 18(1), 713-723.
[CrossRef] [Google Scholar] - Zhang, R., Isola, P., Efros, A. A., Shechtman, E., & Wang, O. (2018). The unreasonable effectiveness of deep features as a perceptual metric. In Proceedings of the IEEE conference on computer vision and pattern recognition (pp. 586-595). IEEE.
[CrossRef] [Google Scholar] - Padilla, R., Passos, W. L., Dias, T. L., Netto, S. L., & da Silva, E. A. (2021). A comparative analysis of object detection metrics with a companion open-source toolkit. Electronics, 10(3), 279.
[CrossRef] [Google Scholar] - Everingham, M., Van Gool, L., Williams, C. K., Winn, J., & Zisserman, A. (2010). The pascal visual object classes (voc) challenge. International Journal of Computer Vision, 88, 303-338.
[CrossRef] [Google Scholar] - Oksuz, K., Cam, B. C., Akbas, E., & Kalkan, S. (2022). One metric to measure them all: Localisation recall precision (LRP) for evaluating visual detection tasks. IEEE Transactions on Pattern Analysis and Machine Intelligence, 44(9), 4942-4959.
[CrossRef] [Google Scholar] - Rodríguez-Lira, A., Mata, C., Solano-Rojas, B., Rodríguez-Bermúdez, G., Villanueva-Polanco, R., Martínez-Arroyo, M., & Quesada-López, C. (2025). Recent advances in deep learning-based 3D reconstruction: A survey. Electronics, 14(15), 3032.
[CrossRef] [Google Scholar] - Wang, Z., Simoncelli, E. P., & Bovik, A. C. (2003, November). Multiscale structural similarity for image quality assessment. In The thirty-seventh Asilomar Conference on Signals, Systems & Computers, 2003 (Vol. 2, pp. 1398-1402). IEEE.
[CrossRef] [Google Scholar] - Celaya, A., Actor, J. A., Muthusivarajan, R., Gates, E., Chung, C., Schellingerhout, D., ... & Fuentes, D. (2023). A generalized surface loss for reducing the Hausdorff distance in medical imaging segmentation. arXiv preprint arXiv:2302.03868.
[CrossRef] [Google Scholar] - Mittal, A., Soundararajan, R., & Bovik, A. C. (2012). Making a "completely blind" image quality analyzer. IEEE Signal Processing Letters, 20(3), 209-212.
[CrossRef] [Google Scholar] - Sheikh, H. R., & Bovik, A. C. (2006). Image information and visual quality. IEEE Transactions on Image Processing, 15(2), 430-444.
[CrossRef] [Google Scholar] - Al-Khafaji, S. L., & Ramaha, G. (2025). Hybrid deep learning model for image compression. Baghdad Science Journal, 22(3), 1030-1040.
[CrossRef] [Google Scholar] - Chen, T., Liu, H., Ma, Z., Shen, Q., Cao, X., & Wang, Y. (2020). End-to-end learnt image compression via non-local attention optimization and improved context quantization. arXiv preprint arXiv:2007.02711.
[CrossRef] [Google Scholar] - Gu, S., Zhang, L., Zuo, W., Feng, X., & Sang, N. (2016). A no-reference perceptual sharpness assessment for ultra-high-definition 4k displays. In 2016 IEEE International Conference on Image Processing (ICIP) (pp. 2369-2373). IEEE.
[CrossRef] [Google Scholar] - Zhao, H., Gallo, O., Frosio, I., & Kautz, J. (2017). Loss functions for image restoration with neural networks. IEEE Transactions on Computational Imaging, 3(1), 47-57.
[CrossRef] [Google Scholar] - Babu, V. R., Chelladurai, D., & Jeyabharathi, S. (2024). Image compression and reconstruction via enhanced quality using deep-learning. Journal of Intelligent & Fuzzy Systems, 46(4), 9591-9606.
[CrossRef] [Google Scholar] - Ohayon, G., Michaeli, T., & Elad, M. (2025, May). Posterior-mean rectified flow: Towards minimum MSE photo-realistic image restoration. In International Conference on Learning Representations (Vol. 2025, pp. 63101-63140).
[Google Scholar] - Arumugam, P., & Banothu, B. (2025). Integrating multi-scale attention network for image restoration task with noise estimation. Image and Vision Computing, 156, 105419.
[CrossRef] [Google Scholar] - Seif, G., & Androutsos, D. (2018). Edge-based loss function for single image super-resolution. In 2018 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) (pp. 1468-1472). IEEE.
[CrossRef] [Google Scholar] - Seitzer, M., Yang, G., Schlemper, J., Oktay, O., Würfl, T., Christlein, V., ... & Maier, A. (2018). Adversarial and perceptual refinement for compressed sensing MRI reconstruction. In Medical Image Computing and Computer Assisted Intervention–MICCAI 2018: 21st International Conference, Granada, Spain, September 16-20, 2018, Proceedings, Part I (pp. 232-240). Springer International Publishing.
[CrossRef] [Google Scholar] - Khan, A., Sohail, A., Zahoora, U., & Qureshi, A. S. (2023). Fast imaging enhancement network: Iterative refinement for noise removal in coronary angiography images. Biomedical Signal Processing and Control, 84, 104827.
[CrossRef] [Google Scholar] - Dobre-Baron, O., Nicola, C., Nicola, M., & Istratoaie, O. (2025). Evaluating image quality metrics as loss functions for deep learning-based image dehazing. Electronics, 14(5), 1040.
[CrossRef] [Google Scholar] - Keleş, O., Tarhan, C., & Bartan, O. (2021). On the computation of PSNR for a set of images or video. In 2021 Picture Coding Symposium (PCS) (pp. 1-5). IEEE.
[CrossRef] [Google Scholar] - Dong, C., Loy, C. C., He, K., & Tang, X. (2015). Image super-resolution using deep convolutional networks. IEEE Transactions on Pattern Analysis and Machine Intelligence, 38(2), 295-307.
[CrossRef] [Google Scholar] - Hore, A., & Ziou, D. (2010). Image quality metrics: PSNR vs. SSIM. In 2010 20th International Conference on Pattern Recognition (pp. 2366-2369). IEEE.
[CrossRef] [Google Scholar] - Hodson, T. O. (2022). Root-mean-square error (RMSE) or mean absolute error (MAE): When to use them or not. Geoscientific Model Development, 15(14), 5481-5487.
[CrossRef] [Google Scholar] - Radhabai, B. R., Subramani, B., & Karuppasamy, P. (2024). An intelligent deep learning approach for diagnosis of brain MRI images. Computers and Electrical Engineering, 115, 109123.
[CrossRef] [Google Scholar] - Maliński, Ł. (2023). Decomposed dissimilarity measures for image denoising assessment. Signal, Image and Video Processing, 17(4), 945-952.
[CrossRef] [Google Scholar] - Wu, J. (2025). Optimizing MRI signal-to-noise ratio using the Fourier algorithm. Journal of Imaging, 11(3), 88.
[CrossRef] [Google Scholar] - Gnanasambandam, A., & Chan, S. H. (2022). Exposure-referred signal to noise ratio for digital image sensors. IEEE Transactions on Image Processing, 31, 2653-2665.
[CrossRef] [Google Scholar] - Ayde, R., Engelbert, T., Nikulin, P., Schär, M., Pruessmann, K., & Weiss, K. (2025). MRI SNR: A software solution review. Magnetic Resonance in Medicine, 93(3), 933-947.
[CrossRef] [Google Scholar] - Tabuchi, A., Nakamura, Y., Seki, S., Ueda, T., & Ikeda, D. (2022). Estimation of signal-to-noise ratio for a clinical X-ray CT image. Diagnostics, 12(12), 3133.
[CrossRef] [Google Scholar] - Zhang, Q., Jiang, X., & Li, L. (2025). Self-supervised denoising methods for medical images: A survey. Computers in Biology and Medicine, 186, 109624.
[CrossRef] [Google Scholar] - Kojima, S., Shinohara, H., Ueda, T., Ohira, S., Akino, Y., Shimamoto, H., ... & Konishi, K. (2025). Practical MRI signal-to-noise ratio mapping: A comparative study of methods and clinical applications. Radiological Physics and Technology, 18(2), 393-404.
[CrossRef] [Google Scholar] - Fiete, R. D., & Tantalo, T. (2001). Comparison of SNR image quality metrics for remote sensing systems. Optical Engineering, 40(4), 574-585.
[CrossRef] [Google Scholar] - Xu, X., Wang, R., Fu, C. W., & Jia, J. (2022). SNR-aware low-light image enhancement. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (pp. 17714-17724). IEEE.
[CrossRef] [Google Scholar] - Sundarrajan, B., Ramakrishnan, S., Narayana, S., Sathishkumar, V. E., & Prabhu, J. (2024). Enhancing low-light medical images using a deep learning-based noise reduction approach. Scientific Reports, 14(1), 30590.
[CrossRef] [Google Scholar] - Usamentiaga, R., García, D. F., & Molleda, J. (2018). More than fifty shades of grey: quantitative characterization of defects and anomalies in infrared images for nondestructive testing. Sensors, 18(10), 3396.
[CrossRef] [Google Scholar] - Liu, Y., Zhang, L., Lu, Y., Liao, Y., Zhao, Q., Yan, Q., ... & Ma, J. (2025). Joint super-resolution and denoising of MRI images via a multi-branch feature-extraction network. IEEE Transactions on Instrumentation and Measurement, 74, 5014014.
[CrossRef] [Google Scholar] - Zhang, L., Zhang, L., Mou, X., & Zhang, D. (2011). FSIM: A feature similarity index for image quality assessment. IEEE Transactions on Image Processing, 20(8), 2378-2386.
[CrossRef] [Google Scholar] - Wandile, K., Dhore, N., Suryawanshi, A., Wanjari, G., Deshmukh, P., & Wasankar, S. (2025). Filtering methods for image denoising and performance evaluation criteria. International Journal for Research in Applied Science and Engineering Technology, 13(3), 1505-1520.
[CrossRef] [Google Scholar] - Afnan, Ullah, F., Yaseen, Lee, J., Jamil, S., & Kwon, O. J. (2023). Subjective assessment of objective image quality metrics range guaranteeing visually lossless compression. Sensors, 23(3), 1297.
[CrossRef] [Google Scholar] - Singh, G., Sehgal, P., Kaur, A., Bhatt, R., Panchal, B., & Khan, I. R. (2025). An innovative deep learning method based on transformer for denoising of digital mammographic images. Scientific Reports, 15(1), 3416.
[CrossRef] [Google Scholar] - Robeson, S. M., & Willmott, C. J. (2023). Decomposition of the mean absolute error (MAE) into systematic and unsystematic components. PLOS ONE, 18(2), e0279774.
[CrossRef] [Google Scholar] - Karunasingha, D. S. K. (2022). Root mean square error or mean absolute error? Use their ratio as well. Information Sciences, 585, 609-629.
[CrossRef] [Google Scholar] - Eskicioglu, A. M., & Fisher, P. S. (1995). Image quality measures and their performance. IEEE Transactions on Communications, 43(12), 2959-2965.
[CrossRef] [Google Scholar] - Asamoah, D., Effah-Baah, J., Owusu-Banahene, W., & Danso, O. (2018). Measuring image contrast in digital images. Journal of Engineering and Applied Sciences, 13, 5347-5355.
[CrossRef] [Google Scholar] - Rainio, O., & Klén, R. (2026). Modified dice coefficients for evaluation of tumor segmentation from PET images. Journal of Imaging Informatics in Medicine, 39(1), 785-793.
[CrossRef] [Google Scholar] - Zou, K. H., Warfield, S. K., Bharatha, A., Tempany, C. M., Kaus, M. R., Haker, S. J., ... & Kikinis, R. (2004). Statistical validation of image segmentation quality based on a spatial overlap index: scientific reports. Academic radiology, 11(2), 178-189.
[CrossRef] [Google Scholar] - Eelbode, T., Bertels, J., Berman, M., Vandermeulen, D., Maes, F., Bisschops, R., & Blaschko, M. B. (2020). Optimization for medical image segmentation: Theory and practice when evaluating with dice coefficient and dice loss. IEEE Transactions on Medical Imaging, 39(11), 3679-3690.
[CrossRef] [Google Scholar] - Bertels, J., Eelbode, T., Berman, M., Vandermeulen, D., Maes, F., Bisschops, R., & Blaschko, M. B. (2019). Optimizing the dice score and jaccard index for medical image segmentation: Theory and practice. In Medical Image Computing and Computer Assisted Intervention–MICCAI 2019: 22nd International Conference, Shenzhen, China, October 13–17, 2019, Proceedings, Part II 22 (pp. 92-100). Springer International Publishing.
[CrossRef] [Google Scholar] - Kato, H., & Hotta, K. (2022). Adaptive t-vMF dice loss: An effective expansion of dice loss for medical image segmentation. arXiv preprint arXiv:2207.07842.
[CrossRef] [Google Scholar] - Cordts, M., Omran, M., Ramos, S., Rehfeld, T., Enzweiler, M., Benenson, R., ... & Schiele, B. (2016). The cityscapes dataset for semantic urban scene understanding. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (pp. 3213-3223). IEEE.
[CrossRef] [Google Scholar] - Ataş, M. (2023). Performance evaluation of the Jaccard-Dice coefficient in building segmentation from high-resolution satellite images. Sensors, 23(9), 4505.
[CrossRef] [Google Scholar] - Pont-Tuset, J., & Marques, F. (2016). Supervised evaluation of image segmentation and object proposal techniques. IEEE Transactions on Pattern Analysis and Machine Intelligence, 38(7), 1465-1478.
[CrossRef] [Google Scholar] - Maxwell, A. E., Warner, T. A., & Fang, F. (2021). Implementation of machine-learning classification in remote sensing: An applied review. International Journal of Remote Sensing, 39(9), 2784-2817.
[CrossRef] [Google Scholar] - Oksuz, K., Cam, B. C., Akbas, E., & Kalkan, S. (2018). Localization recall precision (LRP): A new performance metric for object detection. In Proceedings of the European conference on computer vision (ECCV) (pp. 504-519). Springer.
[CrossRef] [Google Scholar] - Martin, D. R., Fowlkes, C. C., & Malik, J. (2004). Learning to detect natural image boundaries using local brightness, color, and texture cues. IEEE Transactions on Pattern Analysis and Machine Intelligence, 26(5), 530-549.
[CrossRef] [Google Scholar] - Tharwat, A. (2021). Classification assessment methods. Applied Computing and Informatics, 17(1), 168-192.
[CrossRef] [Google Scholar] - Chicco, D., & Jurman, G. (2020). The advantages of the Matthews correlation coefficient (MCC) over F1 score and accuracy in binary classification evaluation. BMC Genomics, 21(1), 6.
[CrossRef] [Google Scholar] - Ramli, R., Arof, H., Ibrahim, F., Syamsul Khair Mokhtar, H., & Ayob, M. A. (2022). Confusion matrix-based performance measure of detection using a corner detector. Pertanika Journal of Science & Technology, 30(1).
[CrossRef] [Google Scholar] - de Zarzà, I., de Curtò, J., & Calafate, C. T. (2022). Detection and assessment of glaucoma using EfficientNet deep learning model. Informatics in Medicine Unlocked, 31, 100977.
[CrossRef] [Google Scholar] - Chicco, D., & Jurman, G. (2022). An invitation to greater use of Matthews correlation coefficient in robotics and artificial intelligence. Frontiers in Robotics and AI, 9, 876814.
[CrossRef] [Google Scholar] - Farhadpour, M., Farajzadeh, K., & Hashemzadeh, M. (2024). Selecting loss functions and evaluation metrics for multi-class imbalanced problems. Applied Soft Computing, 167, 112406.
[CrossRef] [Google Scholar] - Hemmatian, B., Kargar, M., & Ghaderpour, E. (2025). Addressing class imbalance via cluster-based noise reduction and SMOTE for classification. Scientific Reports, 15(1), 6988.
[CrossRef] [Google Scholar] - Obi, J. C. (2023). A comparative study of classification metrics. International Journal of Computational and Applied Mathematics & Computer Science, 3, 19-25.
[CrossRef] [Google Scholar] - Nahm, F. S. (2022). Receiver operating characteristic curve: overview and practical use for clinicians. Korean Journal of Anesthesiology, 75(1), 25-36.
[CrossRef] [Google Scholar] - Richardson, B., Bhatt, D., Bhatt, P., Bhatt, D., Bhatt, S., & Tiwari, P. (2024). The ROC curve: An accurate assessment of imbalanced datasets in machine learning. Plos one, 19(12), e0315331.
[CrossRef] [Google Scholar] - Chang, C. I. (2020). An analysis of 3D hyperspectral target detection using receiver operating characteristic curves with three-dimensional ROC (3DROC) analysis. Remote Sensing, 12(16), 2604.
[CrossRef] [Google Scholar] - Aguilar-Ruiz, J. S., & Michalak, K. (2024). Classification performance evaluation with imbalanced multi-class data. Neurocomputing, 594, 127856.
[CrossRef] [Google Scholar] - Saito, T., & Rehmsmeier, M. (2015). The precision-recall plot is more informative than the ROC plot when evaluating binary classifiers on imbalanced datasets. PLOS ONE, 10(3), e0118432.
[CrossRef] [Google Scholar]
Cite This Article
TY - JOUR AU - Khan, Muzammil AU - Hamza, Muhammad AU - Ali, Sikandar AU - Ali, Ibrar AU - Khan, Zaheer PY - 2026 DA - 2026/10/11 TI - Evaluation Metrics in Digital Image Processing: A Task-Oriented Review and Practical Selection Framework JO - ICCK Journal of Image Analysis and Processing T2 - ICCK Journal of Image Analysis and Processing JF - ICCK Journal of Image Analysis and Processing VL - 2 IS - 4 SP - 231 EP - 251 DO - 10.62762/JIAP.2026.838036 UR - https://www.icck.org/article/abs/JIAP.2026.838036 KW - digital image processing KW - evaluation metrics KW - image quality assessment KW - image classification KW - image segmentation KW - image restoration KW - object detection AB - The rapid development of digital image processing and machine learning has increased the need for reliable and task-appropriate evaluation methods. However, the wide range of available metrics, each with different assumptions, strengths, and limitations, makes metric selection challenging across diverse image-processing tasks. This article presents a task-oriented review of commonly used evaluation metrics and proposes a practical framework for selecting appropriate metrics across five major categories: image quality assessment, image classification, image segmentation, image restoration, and object detection. Representative metrics are reviewed based on their mathematical basis, interpretation, applications, advantages, and limitations. Particular attention is given to pixel-wise measures, including MAE, MSE, and RMSE, perceptual and structural measures such as PSNR, SSIM, MS-SSIM, FSIM, VIF, and LPIPS, and classification and detection measures including accuracy, precision, recall, F1-score, ROC-AUC, confusion matrices, and mean average precision. Overlap-based metrics such as Dice and IoU are also examined for segmentation. The review emphasizes that no single metric adequately captures all dimensions of image-processing performance. Therefore, metric selection should consider task characteristics, data properties, reference availability, and perceptual requirements, with complementary metrics used where appropriate. The proposed framework aims to support researchers and practitioners in selecting appropriate and interpretable evaluation metrics for digital image processing and machine learning applications. SN - 3068-6679 PB - Institute of Central Computation and Knowledge LA - English ER -
@article{Khan2026Evaluation,
author = {Muzammil Khan and Muhammad Hamza and Sikandar Ali and Ibrar Ali and Zaheer Khan},
title = {Evaluation Metrics in Digital Image Processing: A Task-Oriented Review and Practical Selection Framework},
journal = {ICCK Journal of Image Analysis and Processing},
year = {2026},
volume = {2},
number = {4},
pages = {231-251},
doi = {10.62762/JIAP.2026.838036},
url = {https://www.icck.org/article/abs/JIAP.2026.838036},
abstract = {The rapid development of digital image processing and machine learning has increased the need for reliable and task-appropriate evaluation methods. However, the wide range of available metrics, each with different assumptions, strengths, and limitations, makes metric selection challenging across diverse image-processing tasks. This article presents a task-oriented review of commonly used evaluation metrics and proposes a practical framework for selecting appropriate metrics across five major categories: image quality assessment, image classification, image segmentation, image restoration, and object detection. Representative metrics are reviewed based on their mathematical basis, interpretation, applications, advantages, and limitations. Particular attention is given to pixel-wise measures, including MAE, MSE, and RMSE, perceptual and structural measures such as PSNR, SSIM, MS-SSIM, FSIM, VIF, and LPIPS, and classification and detection measures including accuracy, precision, recall, F1-score, ROC-AUC, confusion matrices, and mean average precision. Overlap-based metrics such as Dice and IoU are also examined for segmentation. The review emphasizes that no single metric adequately captures all dimensions of image-processing performance. Therefore, metric selection should consider task characteristics, data properties, reference availability, and perceptual requirements, with complementary metrics used where appropriate. The proposed framework aims to support researchers and practitioners in selecting appropriate and interpretable evaluation metrics for digital image processing and machine learning applications.},
keywords = {digital image processing, evaluation metrics, image quality assessment, image classification, image segmentation, image restoration, object detection},
issn = {3068-6679},
publisher = {Institute of Central Computation and Knowledge}
}
Article Metrics
Publisher's Note
ICCK stays neutral with regard to jurisdictional claims in published maps and institutional affiliations.
Rights and Permissions
Copyright © 2026 by the Author(s). Published by Institute of Central Computation and Knowledge. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license (https://creativecommons.org/licenses/by/4.0/), which permits use, sharing, adaptation, distribution and reproduction in any medium or format, as long as you give appropriate credit to the original author(s) and the source, provide a link to the Creative Commons licence, and indicate if changes were made.
Portico