Comparative Study of Transfer Learning Strategies for Multi-Class Skin Lesion Classification: Architectures, Fine-Tuning, and Data Augmentation
Research Article  ·  Published: 03 June 2026
Issue cover
ICCK Journal of Image Analysis and Processing
Volume 2, Issue 3, 2026: 153-167
Research Article Open Access

Comparative Study of Transfer Learning Strategies for Multi-Class Skin Lesion Classification: Architectures, Fine-Tuning, and Data Augmentation

1 School of Optoelectronic Engineering, Xi'an Technological University, Xi'an 710021, China
* Corresponding Author: Dingguo Wang, [email protected]
Volume 2, Issue 3

Article Information

Abstract

Skin lesion classification is critical in dermatological diagnosis, where early and accurate identification of malignant lesions can significantly improve patient outcomes. Deep learning approaches, particularly transfer learning with pre-trained CNNs, have demonstrated remarkable performance in automated dermoscopic image analysis. However, the optimal configuration of transfer learning components---including backbone architecture, fine-tuning strategy, and data augmentation intensity---remains an open question. In this paper, we present a systematic comparative study on the HAM10000 dataset, evaluating three CNN architectures (ResNet50, DenseNet121, EfficientNet-B0), three fine-tuning strategies (full, partial, classifier-only), and three data augmentation strategies (basic, moderate, aggressive). Our experiments reveal: (1) all three architectures achieve comparable per-class F1-scores under basic augmentation, with no statistically significant differences (Welch's t-test, p > 0.05), despite EfficientNet-B0 reaching 100% overall validation accuracy; (2) full fine-tuning yields the highest accuracy (86.19%) and AUC (99.42%) at increased computational cost; (3) basic augmentation achieves the best performance (accuracy=91.43%, AUC=99.30%), while aggressive augmentation degrades results due to excessive distortion of medical image features. Ablation studies further demonstrate: label smoothing and inverse class frequency weights produce a small positive interaction effect; gradient clipping at norm 1.0 is essential for training stability (without it, training collapses); and a backbone learning rate of $5 \times 10^{-4}$ yields optimal partial fine-tuning performance. McNemar's test confirms no significant difference in BKL vs.\ MEL misclassification patterns between ResNet50 and EfficientNet-B0 ($p > 0.05$). These findings provide practical guidelines for configuring transfer learning pipelines in medical image classification.

Graphical Abstract

Comparative Study of Transfer Learning Strategies for Multi-Class Skin Lesion Classification: Architectures, Fine-Tuning, and Data Augmentation

Keywords

transfer learning skin lesion classification convolutional neural networks fine-tuning strategies data augmentation HAM10000 medical image analysis ablation study

Data Availability Statement

Data will be made available on request.

Funding

This work was supported without any funding.

Conflicts of Interest

The authors declare no conflicts of interest.

AI Use Statement

The authors declare that no generative AI was used in the preparation of this manuscript.

Ethical Approval and Consent to Participate

This study used only publicly available, de-identified datasets (HAM10000) that were collected and ethically approved in prior studies. No new human subjects were involved; therefore, additional ethical approval was not required.

References

  1. Esteva, A., Kuprel, B., Novoa, R. A., Ko, J., Swetter, S. M., Blau, H. M., & Thrun, S. (2017). Dermatologist-level classification of skin cancer with deep neural networks. nature, 542(7639), 115-118.
    [CrossRef] [Google Scholar]
  2. Tschandl, P., Rosendahl, C., & Kittler, H. (2018). The HAM10000 dataset, a large collection of multi-source dermatoscopic images of common pigmented skin lesions. Scientific data, 5(1), 180161.
    [CrossRef] [Google Scholar]
  3. Litjens, G., Kooi, T., Bejnordi, B. E., Setio, A. A. A., Ciompi, F., Ghafoorian, M., ... & Sánchez, C. I. (2017). A survey on deep learning in medical image analysis. Medical image analysis, 42, 60-88.
    [CrossRef] [Google Scholar]
  4. Deng, J., Dong, W., Socher, R., Li, L. J., Li, K., & Fei-Fei, L. (2009, June). Imagenet: A large-scale hierarchical image database. In 2009 IEEE conference on computer vision and pattern recognition (pp. 248-255). IEEE.
    [CrossRef] [Google Scholar]
  5. Tajbakhsh, N., Shin, J. Y., Gurudu, S. R., Hurst, R. T., Kendall, C. B., Gotway, M. B., & Liang, J. (2016). Convolutional neural networks for medical image analysis: Full training or fine tuning?. IEEE transactions on medical imaging, 35(5), 1299-1312.
    [CrossRef] [Google Scholar]
  6. Brinker, T. J., Hekler, A., Utikal, J. S., Grabe, N., Schadendorf, D., Klode, J., ... & Von Kalle, C. (2018). Skin cancer classification using convolutional neural networks: systematic review. Journal of medical Internet research, 20(10), e11936.
    [CrossRef] [Google Scholar]
  7. Raghu, M., Zhang, C., Kleinberg, J., & Bengio, S. (2019). Transfusion: Understanding transfer learning for medical imaging. Advances in neural information processing systems, 32.
    [Google Scholar]
  8. Shorten, C., & Khoshgoftaar, T. M. (2019). A survey on image data augmentation for deep learning. Journal of big data, 6(1), 1-48.
    [CrossRef] [Google Scholar]
  9. He, K., Zhang, X., Ren, S., & Sun, J. (2016). Deep residual learning for image recognition. In Proceedings of the IEEE conference on computer vision and pattern recognition (pp. 770-778).
    [Google Scholar]
  10. Huang, G., Liu, Z., Van Der Maaten, L., & Weinberger, K. Q. (2017). Densely connected convolutional networks. In Proceedings of the IEEE conference on computer vision and pattern recognition (pp. 4700-4708).
    [Google Scholar]
  11. Tan, M., & Le, Q. (2019, May). Efficientnet: Rethinking model scaling for convolutional neural networks. In International conference on machine learning (pp. 6105-6114). PMLR.
    [Google Scholar]
  12. Gutman, D., Codella, N. C., Celebi, E., Helba, B., Marchetti, M., Mishra, N., & Halpern, A. (2016). Skin lesion analysis toward melanoma detection: A challenge at the international symposium on biomedical imaging (ISBI) 2016, hosted by the international skin imaging collaboration (ISIC). arXiv preprint arXiv:1605.01397.
    [CrossRef] [Google Scholar]
  13. Gessert, N., Nielsen, M., Shaikh, M., Werner, R., & Schlaefer, A. (2020). Skin lesion classification using ensembles of multi-resolution EfficientNets with meta data. MethodsX, 7, 100864.
    [CrossRef] [Google Scholar]
  14. Chaturvedi, S. S., Tembhurne, J. V., & Diwan, T. (2020). A multi-class skin Cancer classification using deep convolutional neural networks. Multimedia Tools and Applications, 79(39), 28477-28498.
    [CrossRef] [Google Scholar]
  15. Shin, H. C., Roth, H. R., Gao, M., Lu, L., Xu, Z., Nogues, I., ... & Summers, R. M. (2016). Deep convolutional neural networks for computer-aided detection: CNN architectures, dataset characteristics and transfer learning. IEEE transactions on medical imaging, 35(5), 1285-1298.
    [CrossRef] [Google Scholar]
  16. Yosinski, J., Clune, J., Bengio, Y., & Lipson, H. (2014). How transferable are features in deep neural networks?. Advances in neural information processing systems, 27.
    [Google Scholar]
  17. Zhong, Z., Zheng, L., Kang, G., Li, S., & Yang, Y. (2020, April). Random erasing data augmentation. In Proceedings of the AAAI conference on artificial intelligence (Vol. 34, No. 07, pp. 13001-13008).
    [CrossRef] [Google Scholar]
  18. Goyal, M., & Rajapakse, J. C. (2018). Deep neural network ensemble by data augmentation and bagging for skin lesion classification. arXiv preprint arXiv:1807.05496.
    [CrossRef] [Google Scholar]
  19. Han, S. S., Kim, M. S., Lim, W., Park, G. H., Park, I., & Chang, S. E. (2018). Classification of the clinical images for benign and malignant cutaneous tumors using a deep learning algorithm. Journal of Investigative Dermatology, 138(7), 1529-1538.
    [CrossRef] [Google Scholar]
  20. Loshchilov, I., & Hutter, F. (2017). Decoupled weight decay regularization. arXiv preprint arXiv:1711.05101.
    [CrossRef] [Google Scholar]
  21. Welch, B. L. (1947). The generalization of ‘STUDENT'S’problem when several different population varlances are involved. Biometrika, 34(1-2), 28-35.
    [CrossRef] [Google Scholar]
  22. Brown, M. B., & Forsythe, A. B. (1974). Robust tests for the equality of variances. Journal of the American statistical association, 69(346), 364-367.
    [CrossRef] [Google Scholar]
  23. McNemar, Q. (1947). Note on the sampling error of the difference between correlated proportions or percentages. Psychometrika, 12(2), 153-157.
    [CrossRef] [Google Scholar]
  24. Zhang, J., Zhong, F., He, K., Ji, M., Li, S., & Li, C. (2023). Recent advancements and perspectives in the diagnosis of skin diseases using machine learning and deep learning: A review. Diagnostics, 13(23), 3506.
    [CrossRef] [Google Scholar]
  25. Shakya, M., Patel, R., & Joshi, S. (2025). A comprehensive analysis of deep learning and transfer learning techniques for skin cancer classification. Scientific reports, 15(1), 4633.
    [CrossRef] [Google Scholar]
  26. Naeem, M. A., Yang, S., Saleem, M. A., Javed, I., & Javeed, A. (2025). Automated skin cancer detection using MedFusionNet with attention-based fusion of ConvNeXt and vision transformer. Scientific Reports.
    [CrossRef] [Google Scholar]
  27. Gulzar, Y., Ya’u, B. I., Alkanan, M., & Onn, C. W. (2025). ScNet: a lightweight CNN with depthwise and SE modules for skin lesion classification. Computer Methods in Biomechanics and Biomedical Engineering: Imaging & Visualization, 13(1), 2576198.
    [CrossRef] [Google Scholar]
  28. Meswal, H., Kumar, D., Gupta, A., & Roy, S. (2024). A weighted ensemble transfer learning approach for melanoma classification from skin lesion images. Multimedia Tools and Applications, 83(11), 33615-33637.
    [CrossRef] [Google Scholar]
  29. Ozdemir, Z., Keles, H. Y., & ozgur Tanriover, O. (2025). Meta-transfer derm-diagnosis: exploring few-shot learning and Transfer learning for skin disease classification in long-tail distribution. IEEE Journal of Biomedical and Health Informatics.
    [CrossRef] [Google Scholar]
  30. Ali, M. S., Miah, M. S., Haque, J., Rahman, M. M., & Islam, M. K. (2021). An enhanced technique of skin cancer classification using deep convolutional neural network with transfer learning models. Machine Learning with Applications, 5, 100036.
    [CrossRef] [Google Scholar]
  31. Phuntsho, K., Lee, K., Lee, I., & Ahn, E. (2025). Adaptation of Foundation Models for Medical Image Analysis: Strategies, Challenges, and Future Directions. arXiv preprint arXiv:2511.01284.
    [CrossRef] [Google Scholar]
  32. Kim, M., Yoo, J., Kwon, S., Kim, B. J., Pak, C. J., Won, C. H., ... & Park, K. H. (2025). Diffusion-based skin disease data augmentation with fine-grained detail preservation and interpolation for data diversity. Plos one, 20(10), e0331404.
    [CrossRef] [Google Scholar]
  33. Musthafa, M. M., TR, M., V, V. K., & Guluwadi, S. (2024). Enhanced skin cancer diagnosis using optimized CNN architecture and checkpoints for automated dermatological lesion classification. BMC Medical Imaging, 24(1), 201.
    [CrossRef] [Google Scholar]
  34. Gabani, V., Navamani, T. M., Shyamala, K., & Vaswani Rajpal, V. K. (2026). Multimodal skin lesion classification for early cancer diagnosis using deep learning. Frontiers in Physiology, 17, 1717517.
    [CrossRef] [Google Scholar]
  35. Wang, Y., Chang, Y., Qin, Y., Zhao, Y., & Wei, S. (2025). Unbiased sample selection and label improvement for mitigating noisy labels in class-imbalanced datasets. IEEE Transactions on Circuits and Systems for Video Technology.
    [CrossRef] [Google Scholar]
  36. Dosovitskiy, A., Beyer, L., Kolesnikov, A., Weissenborn, D., Zhai, X., Unterthiner, T., ... & Houlsby, N. (2020). An image is worth 16x16 words: Transformers for image recognition at scale. arXiv preprint arXiv:2010.11929.
    [CrossRef] [Google Scholar]
  37. Liu, Z., Mao, H., Wu, C. Y., Feichtenhofer, C., Darrell, T., & Xie, S. (2022). A convnet for the 2020s. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition (pp. 11976-11986).
    [Google Scholar]
  38. Guo, L., Andriopoulos, G., Zhao, Z., Ling, S., Dong, Z., & Ross, K. (2024). Cross entropy versus label smoothing: A neural collapse perspective. arXiv preprint arXiv:2402.03979.
    [CrossRef] [Google Scholar]
  39. Xia, G., Laurent, O., Franchi, G., & Bouganis, C. S. (2025, May). Towards understanding why label smoothing degrades selective classification and how to fix it. In International conference on learning representations (Vol. 2025, pp. 62954-62987).
    [Google Scholar]
  40. Chhabra, S., Venkateswara, H., & Li, B. (2025). Label Smoothing++: Enhanced Label Regularization for Training Neural Networks. arXiv preprint arXiv:2509.05307.
    [CrossRef] [Google Scholar]
  41. Wang, G., Li, S., Chen, C., Zeng, J., Yang, J., Yu, D., ... & Shen, L. (2025). Adagc: Improving training stability for large language model pretraining. arXiv preprint arXiv:2502.11034.
    [CrossRef] [Google Scholar]
  42. Micikevicius, P., Narang, S., Alben, J., Diamos, G., Elsen, E., Garcia, D., ... & Wu, H. (2017). Mixed precision training. arXiv preprint arXiv:1710.03740.
    [CrossRef] [Google Scholar]

Cite This Article

APA Style
Wang, D., & Chen, Y. (2026). Comparative Study of Transfer Learning Strategies for Multi-Class Skin Lesion Classification: Architectures, Fine-Tuning, and Data Augmentation. ICCK Journal of Image Analysis and Processing, 2(3), 153-167. https://doi.org/10.62762/JIAP.2026.390206
Export Citation
RIS Format
Compatible with EndNote, Zotero, Mendeley, and other reference managers
TY  - JOUR
AU  - Wang, Dingguo
AU  - Chen, Yudi
PY  - 2026
DA  - 2026/06/03
TI  - Comparative Study of Transfer Learning Strategies for Multi-Class Skin Lesion Classification: Architectures, Fine-Tuning, and Data Augmentation
JO  - ICCK Journal of Image Analysis and Processing
T2  - ICCK Journal of Image Analysis and Processing
JF  - ICCK Journal of Image Analysis and Processing
VL  - 2
IS  - 3
SP  - 153
EP  - 167
DO  - 10.62762/JIAP.2026.390206
UR  - https://www.icck.org/article/abs/JIAP.2026.390206
KW  - transfer learning
KW  - skin lesion classification
KW  - convolutional neural networks
KW  - fine-tuning strategies
KW  - data augmentation
KW  - HAM10000
KW  - medical image analysis
KW  - ablation study
AB  - Skin lesion classification is critical in dermatological diagnosis, where early and accurate identification of malignant lesions can significantly improve patient outcomes. Deep learning approaches, particularly transfer learning with pre-trained CNNs, have demonstrated remarkable performance in automated dermoscopic image analysis. However, the optimal configuration of transfer learning components---including backbone architecture, fine-tuning strategy, and data augmentation intensity---remains an open question. In this paper, we present a systematic comparative study on the HAM10000 dataset, evaluating three CNN architectures (ResNet50, DenseNet121, EfficientNet-B0), three fine-tuning strategies (full, partial, classifier-only), and three data augmentation strategies (basic, moderate, aggressive). Our experiments reveal: (1) all three architectures achieve comparable per-class F1-scores under basic augmentation, with no statistically significant differences (Welch's t-test, p > 0.05), despite EfficientNet-B0 reaching 100% overall validation accuracy; (2) full fine-tuning yields the highest accuracy (86.19%) and AUC (99.42%) at increased computational cost; (3) basic augmentation achieves the best performance (accuracy=91.43%, AUC=99.30%), while aggressive augmentation degrades results due to excessive distortion of medical image features. Ablation studies further demonstrate: label smoothing and inverse class frequency weights produce a small positive interaction effect; gradient clipping at norm 1.0 is essential for training stability (without it, training collapses); and a backbone learning rate of $5 \times 10^{-4}$ yields optimal partial fine-tuning performance. McNemar's test confirms no significant difference in BKL vs.\ MEL misclassification patterns between ResNet50 and EfficientNet-B0 ($p > 0.05$). These findings provide practical guidelines for configuring transfer learning pipelines in medical image classification.
SN  - 3068-6679
PB  - Institute of Central Computation and Knowledge
LA  - English
ER  - 
BibTeX Format
Compatible with LaTeX, BibTeX, and other reference managers
@article{Wang2026Comparativ,
  author = {Dingguo Wang and Yudi Chen},
  title = {Comparative Study of Transfer Learning Strategies for Multi-Class Skin Lesion Classification: Architectures, Fine-Tuning, and Data Augmentation},
  journal = {ICCK Journal of Image Analysis and Processing},
  year = {2026},
  volume = {2},
  number = {3},
  pages = {153-167},
  doi = {10.62762/JIAP.2026.390206},
  url = {https://www.icck.org/article/abs/JIAP.2026.390206},
  abstract = {Skin lesion classification is critical in dermatological diagnosis, where early and accurate identification of malignant lesions can significantly improve patient outcomes. Deep learning approaches, particularly transfer learning with pre-trained CNNs, have demonstrated remarkable performance in automated dermoscopic image analysis. However, the optimal configuration of transfer learning components---including backbone architecture, fine-tuning strategy, and data augmentation intensity---remains an open question. In this paper, we present a systematic comparative study on the HAM10000 dataset, evaluating three CNN architectures (ResNet50, DenseNet121, EfficientNet-B0), three fine-tuning strategies (full, partial, classifier-only), and three data augmentation strategies (basic, moderate, aggressive). Our experiments reveal: (1) all three architectures achieve comparable per-class F1-scores under basic augmentation, with no statistically significant differences (Welch's t-test, p > 0.05), despite EfficientNet-B0 reaching 100\% overall validation accuracy; (2) full fine-tuning yields the highest accuracy (86.19\%) and AUC (99.42\%) at increased computational cost; (3) basic augmentation achieves the best performance (accuracy=91.43\%, AUC=99.30\%), while aggressive augmentation degrades results due to excessive distortion of medical image features. Ablation studies further demonstrate: label smoothing and inverse class frequency weights produce a small positive interaction effect; gradient clipping at norm 1.0 is essential for training stability (without it, training collapses); and a backbone learning rate of \$5 \times 10^{-4}\$ yields optimal partial fine-tuning performance. McNemar's test confirms no significant difference in BKL vs.\ MEL misclassification patterns between ResNet50 and EfficientNet-B0 (\$p > 0.05\$). These findings provide practical guidelines for configuring transfer learning pipelines in medical image classification.},
  keywords = {transfer learning, skin lesion classification, convolutional neural networks, fine-tuning strategies, data augmentation, HAM10000, medical image analysis, ablation study},
  issn = {3068-6679},
  publisher = {Institute of Central Computation and Knowledge}
}

Article Metrics

Citations
Crossref
0
Scopus
0
Views
923
PDF Downloads
466

Publisher's Note

ICCK stays neutral with regard to jurisdictional claims in published maps and institutional affiliations.

Rights and Permissions

CC BY Copyright © 2026 by the Author(s). Published by Institute of Central Computation and Knowledge. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license (https://creativecommons.org/licenses/by/4.0/), which permits use, sharing, adaptation, distribution and reproduction in any medium or format, as long as you give appropriate credit to the original author(s) and the source, provide a link to the Creative Commons licence, and indicate if changes were made.
ICCK Journal of Image Analysis and Processing
ICCK Journal of Image Analysis and Processing
ISSN: 3068-6679 (Online)
Portico
Preserved at
Portico