FocusNet: Feature Oriented Contextual Understanding via Bidirectional Supervision for Surface Defect Detection
Research Article  ·  Published: 17 September 2026
Issue cover
ICCK Transactions on Sensing, Communication, and Control
Volume 3, Issue 3, 2026: 160-175
Research Article Free to Read

FocusNet: Feature Oriented Contextual Understanding via Bidirectional Supervision for Surface Defect Detection

1 Department of Computer Science, Virtual University of Pakistan, Islamabad 44000, Pakistan
2 Department of Computer Science, Graz University of Technology, Graz 8010, Austria
3 Global Degree College, Peshawar 25120, Pakistan
* Corresponding Author: Zaid Muhammad, [email protected]
Volume 3, Issue 3
You have access to this article · Limited-Time Free Access

Article Information

Abstract

Surface defect segmentation in metallic materials presents significant challenges due to irregular defect shapes, extreme scale variations ranging from small blowholes to large uneven regions, and low contrast with complex background textures. While Convolutional Neural Networks excel at local feature extraction, their limited receptive fields hinder effective global context modeling. Conversely, Vision Transformers capture long-range dependencies but struggle with fine-grained boundary details critical for accurate defect localization. To address these limitations, we propose FocusNet, a novel architecture integrating multi-scale feature refinement, hybrid attention mechanisms, and progressive decoding with bidirectional supervision for robust metallic surface defect segmentation. FocusNet introduces a Feature Refinement Module combining Residual Feature Enhancement Blocks and Atrous Spatial Pyramid Pooling to capture multi-scale contextual information while preserving fine-grained details. A Hybrid Attention Mechanism synergistically fuses channel, spatial, and coordinate attention to enhance discriminative feature representations across both global and local scales. The Bidirectional Feature Pyramid Network enables efficient weighted multi-scale feature fusion, facilitating robust detection across diverse defect characteristics. Our progressive decoder employs deep supervision at multiple hierarchical levels, enforcing intermediate representations to learn defect patterns and improving boundary localization accuracy. Extensive experiments on the MT-Defect benchmark demonstrate state-of-the-art performance with 81.6% mIoU, surpassing LGGFormer (80.0%) and SegFormer (66.1%). Notably, FocusNet achieves particularly strong gains on challenging defect categories, including Crack (70.7% vs. 64.9%) and Break (74.3% vs. 72.4%) compared with the best-performing prior method LGGFormer, demonstrating robust segmentation across defects with diverse scales, orientations, and boundary complexities.

Graphical Abstract

FocusNet: Feature Oriented Contextual Understanding via Bidirectional Supervision for Surface Defect Detection

Keywords

surface defect segmentation metallic tile inspection hybrid attention mechanism bidirectional feature pyramid network deep supervision industrial quality control

Data Availability Statement

Data will be made available on request.

Funding

This work was supported without any funding.

Conflicts of Interest

The authors declare no conflicts of interest.

AI Use Statement

The authors declare that Claude (Anthropic) was used for language editing and grammar refinement of the manuscript. The authors have carefully reviewed, revised, and verified the AI-assisted output and take full responsibility for the content of the manuscript.

Ethical Approval and Consent to Participate

Not applicable.

References

  1. Tabernik, D., Šela, S., Skvarč, J., & Skočaj, D. (2020). Segmentation-based deep-learning approach for surface-defect detection. Journal of Intelligent Manufacturing, 31(3), 759-776.
    [CrossRef] [Google Scholar]
  2. Aydin, I., Akin, E., & Karakose, M. (2021). Defect classification based on deep features for railway tracks in sustainable transportation. Applied Soft Computing, 111, 107706.
    [CrossRef] [Google Scholar]
  3. Gao, Y., Lin, J., Xie, J., & Ning, Z. (2020). A real-time defect detection method for digital signal processing of industrial inspection applications. IEEE Transactions on Industrial Informatics, 17(5), 3450-3459.
    [CrossRef] [Google Scholar]
  4. Zeng, N., Wu, P., Wang, Z., Li, H., Liu, W., & Liu, X. (2022). A small-sized object detection oriented multi-scale feature fusion approach with application to defect detection. IEEE Transactions on Instrumentation and Measurement, 71, 1-14.
    [CrossRef] [Google Scholar]
  5. Ye, B., Lai, J., & Xie, X. (2024). Teacher-student collaboration: Effective semi-supervised model for defect instance segmentation. IEEE Transactions on Automation Science and Engineering, 22, 6932-6943.
    [CrossRef] [Google Scholar]
  6. Ma, Y., Liu, M., Jiang, S., Wang, X., Bian, Y., & Wang, Y. (2025). Multi-context aggregation network with foreground correction for automated few-shot defect segmentation. IEEE Transactions on Automation Science and Engineering, 22, 13777-13787.
    [CrossRef] [Google Scholar]
  7. Huang, Y., Jing, J., & Wang, Z. (2021). Fabric defect segmentation method based on deep learning. IEEE Transactions on Instrumentation and Measurement, 70, 1-15.
    [CrossRef] [Google Scholar]
  8. Liu, T., He, Z., Lin, Z., Cao, G. Z., Su, W., & Xie, S. (2022). An adaptive image segmentation network for surface defect detection. IEEE Transactions on Neural Networks and Learning Systems, 35(6), 8510-8523.
    [CrossRef] [Google Scholar]
  9. Bao, Y., Song, K., Liu, J., Wang, Y., Yan, Y., Yu, H., & Li, X. (2021). Triplet-graph reasoning network for few-shot metal generic surface defect segmentation. IEEE Transactions on Instrumentation and Measurement, 70(5), 1-11.
    [CrossRef] [Google Scholar]
  10. Yu, X., Lyu, W., Zhou, D., Wang, C., & Xu, W. (2022). ES-Net: Efficient scale-aware network for tiny defect detection. IEEE Transactions on Instrumentation and Measurement, 71, 1-14.
    [CrossRef] [Google Scholar]
  11. Pan, Y., & Zhang, L. (2022). Dual attention deep learning network for automatic steel surface defect segmentation. Computer‐Aided Civil and Infrastructure Engineering, 37(11), 1468-1487.
    [CrossRef] [Google Scholar]
  12. Song, K., Feng, H., Cao, T., Cui, W., & Yan, Y. (2024). MFANet: Multifeature aggregation network for cross-granularity few-shot seamless steel tubes surface defect segmentation. IEEE Transactions on Industrial Informatics, 20(7), 9725-9735.
    [CrossRef] [Google Scholar]
  13. Liang, W., Sun, Y., Zhang, S., Bai, L., & Yang, J. (2024). SmallNet: A small defects detection network for magnetic chips based on context-weighted aggregation and feature multiscale loop fusion. IEEE Transactions on Automation Science and Engineering, 22, 10095-10106.
    [CrossRef] [Google Scholar]
  14. Ding, X., Zhang, X., Han, J., & Ding, G. (2022, June). Scaling up your kernels to 31× 31: Revisiting large kernel design in CNNs. In 2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) (pp. 11953-11965). IEEE.
    [CrossRef] [Google Scholar]
  15. Cao, Y., Xu, J., Lin, S., Wei, F., & Hu, H. (2019, October). Gcnet: Non-local networks meet squeeze-excitation networks and beyond. In 2019 IEEE/CVF International Conference on Computer Vision Workshop (ICCVW) (pp. 1971-1980). IEEE.
    [CrossRef] [Google Scholar]
  16. Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, Ł., & Polosukhin, I. (2017). Attention is all you need. Advances in Neural Information Processing Systems, 30.
    [Google Scholar]
  17. Dosovitskiy, A., Beyer, L., Kolesnikov, A., Weissenborn, D., Zhai, X., Unterthiner, T., ... & Houlsby, N. (2020). An image is worth 16x16 words: Transformers for image recognition at scale. arXiv preprint arXiv:2010.11929.
    [CrossRef] [Google Scholar]
  18. Yao, H., Luo, W., Yu, W., Zhang, X., Qiang, Z., Luo, D., & Shi, H. (2023). Dual-attention transformer and discriminative flow for industrial visual anomaly detection. IEEE Transactions on Automation Science and Engineering, 21(4), 6126-6140.
    [CrossRef] [Google Scholar]
  19. Strudel, R., Garcia, R., Laptev, I., & Schmid, C. (2021, October). Segmenter: Transformer for semantic segmentation. In 2021 IEEE/CVF International Conference on Computer Vision (ICCV) (pp. 7242-7252). IEEE.
    [CrossRef] [Google Scholar]
  20. Wu, H., Xiao, B., Codella, N., Liu, M., Dai, X., Yuan, L., & Zhang, L. (2021, October). Cvt: Introducing convolutions to vision transformers. In 2021 IEEE/CVF International Conference on Computer Vision (ICCV) (pp. 22-31). IEEE.
    [CrossRef] [Google Scholar]
  21. Xu, Z., Wu, D., Yu, C., Chu, X., Sang, N., & Gao, C. (2024, March). Sctnet: Single-branch cnn with transformer semantic information for real-time segmentation. In Proceedings of the AAAI conference on artificial intelligence (Vol. 38, No. 6, pp. 6378-6386).
    [CrossRef] [Google Scholar]
  22. Xia, C., Wang, X., Lv, F., Hao, X., & Shi, Y. (2024, June). Vit-comer: Vision transformer with convolutional multi-scale feature interaction for dense predictions. In 2024 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) (pp. 5493-5502). IEEE.
    [CrossRef] [Google Scholar]
  23. Shi, D. (2024, June). Transnext: Robust foveal visual perception for vision transformers. In 2024 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) (pp. 17773-17783). IEEE.
    [CrossRef] [Google Scholar]
  24. Guo, J., Zhou, H. Y., Wang, L., & Yu, Y. (2022, June). UNet-2022: Exploring dynamics in non-isomorphic architecture. In International conference on medical imaging and computer-aided diagnosis (pp. 465-476). Singapore: Springer Nature Singapore.
    [CrossRef] [Google Scholar]
  25. Peng, Z., Huang, W., Gu, S., Xie, L., Wang, Y., Jiao, J., & Ye, Q. (2021). Conformer: Local features coupling global representations for visual recognition. In Proceedings of the IEEE/CVF international conference on computer vision (pp. 367-376).
    [Google Scholar]
  26. Lou, M., Zhang, S., Zhou, H. Y., Yang, S., Wu, C., & Yu, Y. (2025). TransXNet: learning both global and local dynamics with a dual dynamic token mixer for visual recognition. IEEE Transactions on Neural Networks and Learning Systems, 36(6), 11534-11547.
    [CrossRef] [Google Scholar]
  27. Xu, G., Jia, W., Wu, T., Chen, L., & Gao, G. (2024). HAFormer: Unleashing the power of hierarchy-aware features for lightweight semantic segmentation. IEEE Transactions on Image Processing, 33, 4202-4214.
    [CrossRef] [Google Scholar]
  28. Shen, X., Liu, J., Zhang, H., Jiang, L., Zhao, H., & Yang, H. (2024). A novel incremental defect detection method via elastic heterogeneous distillation network. IEEE Transactions on Automation Science and Engineering, 22, 10149-10161.
    [CrossRef] [Google Scholar]
  29. Zhang, Y., Wu, J., Li, Q., Zhao, X., & Tan, M. (2021). Beyond crack: Fine-grained pavement defect segmentation using three-stream neural networks. IEEE Transactions on Intelligent Transportation Systems, 23(9), 14820-14832.
    [CrossRef] [Google Scholar]
  30. Liu, T., & He, Z. (2022). TAS 2-Net: Triple-attention semantic segmentation network for small surface defect detection. IEEE Transactions on Instrumentation and Measurement, 71, 1-12.
    [CrossRef] [Google Scholar]
  31. Zhang, J., Ding, R., Ban, M., & Guo, T. (2022, May). FDSNeT: An accurate real-time surface defect segmentation network. In ICASSP 2022-2022 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) (pp. 3803-3807). IEEE.
    [CrossRef] [Google Scholar]
  32. Yeung, C. C., & Lam, K. M. (2023). Attentive boundary-aware fusion for defect semantic segmentation using transformer. IEEE Transactions on Instrumentation and Measurement, 72, 1-13.
    [CrossRef] [Google Scholar]
  33. Zuo, L., Xiao, H., Wen, L., & Gao, L. (2023). A pixel-level segmentation convolutional neural network based on global and local feature fusion for surface defect detection. IEEE Transactions on Instrumentation and Measurement, 72, 1-10.
    [CrossRef] [Google Scholar]
  34. Yin, Z., Qin, L., Han, G., Shi, X., Zhang, F., Xu, G., & Bi, Y. (2024). Ddsnet: Deep dual-branch networks for surface defect segmentation. IEEE Transactions on Instrumentation and Measurement, 73, 1-16.
    [CrossRef] [Google Scholar]
  35. Cao, J., Yang, G., & Yang, X. (2020). A pixel-level segmentation convolutional neural network based on deep feature fusion for surface defect detection. IEEE Transactions on Instrumentation and Measurement, 70, 1-12.
    [CrossRef] [Google Scholar]
  36. Qiu, Y., Liu, H., Liu, J., Shi, B., & Li, Y. (2024). Region and edge-aware network for rail surface defect segmentation. IEEE Transactions on Instrumentation and Measurement, 73, 1-13.
    [CrossRef] [Google Scholar]
  37. Wang, C., Chen, H., & Zhao, S. (2023). RERN: Rich edge features refinement detection network for polycrystalline solar cell defect segmentation. IEEE Transactions on Industrial Informatics, 20(2), 1408-1419.
    [CrossRef] [Google Scholar]
  38. Yang, L., Xu, S., Fan, J., Li, E., & Liu, Y. (2023). A pixel-level deep segmentation network for automatic defect detection. Expert Systems with Applications, 215, 119388.
    [CrossRef] [Google Scholar]
  39. Yang, H., Hu, J., Yin, Z., & Wang, Z. (2022). A semantic information decomposition network for accurate segmentation of texture defects. IEEE Transactions on Industrial Informatics, 19(7), 8319-8327.
    [CrossRef] [Google Scholar]
  40. Yu, R., Guo, B., & Yang, K. (2022). Selective prototype network for few-shot metal surface defect segmentation. IEEE Transactions on Instrumentation and Measurement, 71, 1-10.
    [CrossRef] [Google Scholar]
  41. Wang, W., Xie, E., Li, X., Fan, D. P., Song, K., Liang, D., ... & Shao, L. (2021, October). Pyramid vision transformer: A versatile backbone for dense prediction without convolutions. In 2021 IEEE/CVF international conference on computer vision (ICCV) (pp. 548-558). IEEE.
    [CrossRef] [Google Scholar]
  42. Zheng, S., Lu, J., Zhao, H., Zhu, X., Luo, Z., Wang, Y., ... & Zhang, L. (2021, June). Rethinking semantic segmentation from a sequence-to-sequence perspective with transformers. In 2021 IEEE/CVF conference on computer vision and pattern recognition (CVPR) (pp. 6877-6886). IEEE.
    [CrossRef] [Google Scholar]
  43. Liu, Z., Lin, Y., Cao, Y., Hu, H., Wei, Y., Zhang, Z., ... & Guo, B. (2021, October). Swin transformer: Hierarchical vision transformer using shifted windows. In 2021 IEEE/CVF international conference on computer vision (ICCV) (pp. 9992-10002). Ieee.
    [CrossRef] [Google Scholar]
  44. Xie, E., Wang, W., Yu, Z., Anandkumar, A., Alvarez, J. M., & Luo, P. (2021). SegFormer: Simple and efficient design for semantic segmentation with transformers. Advances in neural information processing systems, 34, 12077-12090.
    [Google Scholar]
  45. Dong, X., Bao, J., Chen, D., Zhang, W., Yu, N., Yuan, L., ... & Guo, B. (2022, June). Cswin transformer: A general vision transformer backbone with cross-shaped windows. In 2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) (pp. 12114-12124). IEEE.
    [CrossRef] [Google Scholar]
  46. Pan, X., Ye, T., Xia, Z., Song, S., & Huang, G. (2023, June). Slide-transformer: Hierarchical vision transformer with local self-attention. In 2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) (pp. 2082-2091). IEEE.
    [CrossRef] [Google Scholar]
  47. Zhu, W., Zhang, H., Zhang, C., Zhu, X., Guan, Z., & Jia, J. (2023). Surface defect detection and classification of steel using an efficient Swin Transformer. Advanced Engineering Informatics, 57, 102061.
    [CrossRef] [Google Scholar]
  48. Wu, Y. H., Liu, Y., Zhan, X., & Cheng, M. M. (2022). P2T: Pyramid pooling transformer for scene understanding. IEEE transactions on pattern analysis and machine intelligence, 45(11), 12760-12771.
    [CrossRef] [Google Scholar]
  49. Wang, W., Xie, E., Li, X., Fan, D. P., Song, K., Liang, D., ... & Shao, L. (2022). Pvt v2: Improved baselines with pyramid vision transformer. Computational visual media, 8(3), 415-424.
    [CrossRef] [Google Scholar]
  50. Shang, H., Sun, C., Liu, J., Chen, X., & Yan, R. (2023). Defect-aware transformer network for intelligent visual surface defect detection. Advanced Engineering Informatics, 55, 101882.
    [CrossRef] [Google Scholar]
  51. Zhou, H., Yang, R., Hu, R., Shu, C., Tang, X., & Li, X. (2023). ETDNet: Efficient transformer-based detection network for surface defect detection. IEEE transactions on instrumentation and measurement, 72, 1-14.
    [CrossRef] [Google Scholar]
  52. Zhang, G., Lu, Y., Jiang, X., Jin, S., Li, S., & Xu, M. (2025). LGGFormer: A dual-branch local-guided global self-attention network for surface defect segmentation. Advanced Engineering Informatics, 64, 103099.
    [CrossRef] [Google Scholar]
  53. Jiang, X., Guo, K., Lu, Y., Yan, F., Liu, H., Cao, J., ... & Tao, D. (2023). CINFormer: Transformer network with multi-stage CNN feature injection for surface defect segmentation. arXiv preprint arXiv:2309.12639.
    [CrossRef] [Google Scholar]
  54. Zhang, D., Song, K., Xu, J., He, Y., Niu, M., & Yan, Y. (2020). MCnet: Multiple context information segmentation network of no-service rail surface defects. IEEE Transactions on Instrumentation and Measurement, 70, 1-9.
    [CrossRef] [Google Scholar]
  55. Tan, M., & Le, Q. (2019, May). Efficientnet: Rethinking model scaling for convolutional neural networks. In International conference on machine learning (pp. 6105-6114). PmLR.
    [Google Scholar]
  56. Deng, J., Dong, W., Socher, R., Li, L. J., Li, K., & Fei-Fei, L. (2009, June). Imagenet: A large-scale hierarchical image database. In 2009 IEEE conference on computer vision and pattern recognition (pp. 248-255). IEEE.
    [CrossRef] [Google Scholar]
  57. Chen, L. C., Papandreou, G., Kokkinos, I., Murphy, K., & Yuille, A. L. (2017). Deeplab: Semantic image segmentation with deep convolutional nets, atrous convolution, and fully connected crfs. IEEE transactions on pattern analysis and machine intelligence, 40(4), 834-848.
    [CrossRef] [Google Scholar]
  58. Hu, J., Shen, L., & Sun, G. (2018). Squeeze-and-excitation networks. In Proceedings of the IEEE conference on computer vision and pattern recognition (pp. 7132-7141).
    [CrossRef] [Google Scholar]
  59. Woo, S., Park, J., Lee, J. Y., & Kweon, I. S. (2018, September). Cbam: Convolutional block attention module. In European conference on computer vision (pp. 3-19). Cham: Springer International Publishing.
    [CrossRef] [Google Scholar]
  60. Hou, Q., Zhou, D., & Feng, J. (2021, June). Coordinate attention for efficient mobile network design. In 2021 IEEE/CVF conference on computer vision and pattern recognition (CVPR) (pp. 13708-13717). IEEE.
    [CrossRef] [Google Scholar]
  61. Huang, Y., Qiu, C., & Yuan, K. (2020). Surface defect saliency of magnetic. The Visual Computer, 36(1), 85-96.
    [CrossRef] [Google Scholar]
  62. Zhang, S., Jiang, L., Ge, L., Wu, C., Wang, Y., Lu, K., & Zhang, C. (2025). SGTP-Net: semantic guidance and texture priors-based dual-branch segmentation network for surface defect detection. IEEE Transactions on Automation Science and Engineering, 23, 417-432.
    [CrossRef] [Google Scholar]
  63. Dong, H., Song, K., He, Y., Xu, J., Yan, Y., & Meng, Q. (2019). PGA-Net: Pyramid feature fusion and global context attention network for automated surface defect detection. IEEE Transactions on Industrial Informatics, 16(12), 7448-7458.
    [CrossRef] [Google Scholar]
  64. Zhou, X., Fang, H., Liu, Z., Zheng, B., Sun, Y., Zhang, J., & Yan, C. (2021). Dense attention-guided cascaded network for salient object detection of strip steel surface defects. IEEE Transactions on Instrumentation and Measurement, 71, 1-14.
    [CrossRef] [Google Scholar]
  65. Guo, M. H., Lu, C. Z., Hou, Q., Liu, Z., Cheng, M. M., & Hu, S. M. (2022). Segnext: Rethinking convolutional attention design for semantic segmentation. Advances in neural information processing systems, 35, 1140-1156.
    [Google Scholar]
  66. Hong, Y., Pan, H., Sun, W., & Jia, Y. (2021). Deep dual-resolution networks for real-time and accurate semantic segmentation of road scenes. arXiv preprint arXiv:2101.06085.
    [CrossRef] [Google Scholar]
  67. Xu, J., Xiong, Z., & Bhattacharyya, S. P. (2023, June). PIDNet: A real-time semantic segmentation network inspired by PID controllers. In 2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) (pp. 19529-19539). IEEE.
    [CrossRef] [Google Scholar]
  68. Xie, L., Xiang, X., Xu, H., Wang, L., Lin, L., & Yin, G. (2020). FFCNN: A deep neural network for surface defect detection of magnetic tile. IEEE Transactions on Industrial Electronics, 68(4), 3506-3516.
    [CrossRef] [Google Scholar]
  69. Dong, B., Wang, P., & Wang, F. (2023, June). Head-free lightweight semantic segmentation with linear transformer. In Proceedings of the AAAI conference on artificial intelligence (Vol. 37, No. 1, pp. 516-524).
    [CrossRef] [Google Scholar]
  70. Ma, S., Song, K., Niu, M., Tian, H., & Yan, Y. (2024). Cross-scale fusion and domain adversarial network for generalizable rail surface defect segmentation on unseen datasets. Journal of Intelligent Manufacturing, 35(1), 367-386.
    [CrossRef] [Google Scholar]
  71. Li, K., Wang, Y., Gao, P., Song, G., Liu, Y., Li, H., & Qiao, Y. (2022). Uniformer: Unified transformer for efficient spatiotemporal representation learning. arXiv preprint arXiv:2201.04676.
    [CrossRef] [Google Scholar]
  72. Si, C., Yu, W., Zhou, P., Zhou, Y., Wang, X., & Yan, S. (2022). Inception transformer. Advances in neural information processing systems, 35, 23495-23509.
    [Google Scholar]
  73. Wan, Q., Huang, Z., Lu, J., Yu, G., & Zhang, L. (2025). SeaFormer++: Squeeze-enhanced axial transformer for mobile visual recognition. International journal of computer vision, 133(6), 3645-3666.
    [CrossRef] [Google Scholar]
  74. Zhang, W., Huang, Z., Luo, G., Chen, T., Wang, X., Liu, W., ... & Shen, C. (2022, June). Topformer: Token pyramid transformer for mobile semantic segmentation. In 2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) (pp. 12073-12083). IEEE.
    [CrossRef] [Google Scholar]

Cite This Article

APA Style
Ali, M. H., Ali, F., & Muhammad, Z. (2026). FocusNet: Feature Oriented Contextual Understanding via Bidirectional Supervision for Surface Defect Detection. ICCK Transactions on Sensing, Communication, and Control, 3(3), 160-175. https://doi.org/10.62762/TSCC.2026.978086
Export Citation
RIS Format
Compatible with EndNote, Zotero, Mendeley, and other reference managers
TY  - JOUR
AU  - Ali, Muhammad Hamza
AU  - Ali, Farhan
AU  - Muhammad, Zaid
PY  - 2026
DA  - 2026/09/17
TI  - FocusNet: Feature Oriented Contextual Understanding via Bidirectional Supervision for Surface Defect Detection
JO  - ICCK Transactions on Sensing, Communication, and Control
T2  - ICCK Transactions on Sensing, Communication, and Control
JF  - ICCK Transactions on Sensing, Communication, and Control
VL  - 3
IS  - 3
SP  - 160
EP  - 175
DO  - 10.62762/TSCC.2026.978086
UR  - https://www.icck.org/article/abs/TSCC.2026.978086
KW  - surface defect segmentation
KW  - metallic tile inspection
KW  - hybrid attention mechanism
KW  - bidirectional feature pyramid network
KW  - deep supervision
KW  - industrial quality control
AB  - Surface defect segmentation in metallic materials presents significant challenges due to irregular defect shapes, extreme scale variations ranging from small blowholes to large uneven regions, and low contrast with complex background textures. While Convolutional Neural Networks excel at local feature extraction, their limited receptive fields hinder effective global context modeling. Conversely, Vision Transformers capture long-range dependencies but struggle with fine-grained boundary details critical for accurate defect localization. To address these limitations, we propose FocusNet, a novel architecture integrating multi-scale feature refinement, hybrid attention mechanisms, and progressive decoding with bidirectional supervision for robust metallic surface defect segmentation. FocusNet introduces a Feature Refinement Module combining Residual Feature Enhancement Blocks and Atrous Spatial Pyramid Pooling to capture multi-scale contextual information while preserving fine-grained details. A Hybrid Attention Mechanism synergistically fuses channel, spatial, and coordinate attention to enhance discriminative feature representations across both global and local scales. The Bidirectional Feature Pyramid Network enables efficient weighted multi-scale feature fusion, facilitating robust detection across diverse defect characteristics. Our progressive decoder employs deep supervision at multiple hierarchical levels, enforcing intermediate representations to learn defect patterns and improving boundary localization accuracy. Extensive experiments on the MT-Defect benchmark demonstrate state-of-the-art performance with 81.6% mIoU, surpassing LGGFormer (80.0%) and SegFormer (66.1%). Notably, FocusNet achieves particularly strong gains on challenging defect categories, including Crack (70.7% vs. 64.9%) and Break (74.3% vs. 72.4%) compared with the best-performing prior method LGGFormer, demonstrating robust segmentation across defects with diverse scales, orientations, and boundary complexities.
SN  - 3068-9287
PB  - Institute of Central Computation and Knowledge
LA  - English
ER  - 
BibTeX Format
Compatible with LaTeX, BibTeX, and other reference managers
@article{Ali2026FocusNet,
  author = {Muhammad Hamza Ali and Farhan Ali and Zaid Muhammad},
  title = {FocusNet: Feature Oriented Contextual Understanding via Bidirectional Supervision for Surface Defect Detection},
  journal = {ICCK Transactions on Sensing, Communication, and Control},
  year = {2026},
  volume = {3},
  number = {3},
  pages = {160-175},
  doi = {10.62762/TSCC.2026.978086},
  url = {https://www.icck.org/article/abs/TSCC.2026.978086},
  abstract = {Surface defect segmentation in metallic materials presents significant challenges due to irregular defect shapes, extreme scale variations ranging from small blowholes to large uneven regions, and low contrast with complex background textures. While Convolutional Neural Networks excel at local feature extraction, their limited receptive fields hinder effective global context modeling. Conversely, Vision Transformers capture long-range dependencies but struggle with fine-grained boundary details critical for accurate defect localization. To address these limitations, we propose FocusNet, a novel architecture integrating multi-scale feature refinement, hybrid attention mechanisms, and progressive decoding with bidirectional supervision for robust metallic surface defect segmentation. FocusNet introduces a Feature Refinement Module combining Residual Feature Enhancement Blocks and Atrous Spatial Pyramid Pooling to capture multi-scale contextual information while preserving fine-grained details. A Hybrid Attention Mechanism synergistically fuses channel, spatial, and coordinate attention to enhance discriminative feature representations across both global and local scales. The Bidirectional Feature Pyramid Network enables efficient weighted multi-scale feature fusion, facilitating robust detection across diverse defect characteristics. Our progressive decoder employs deep supervision at multiple hierarchical levels, enforcing intermediate representations to learn defect patterns and improving boundary localization accuracy. Extensive experiments on the MT-Defect benchmark demonstrate state-of-the-art performance with 81.6\% mIoU, surpassing LGGFormer (80.0\%) and SegFormer (66.1\%). Notably, FocusNet achieves particularly strong gains on challenging defect categories, including Crack (70.7\% vs. 64.9\%) and Break (74.3\% vs. 72.4\%) compared with the best-performing prior method LGGFormer, demonstrating robust segmentation across defects with diverse scales, orientations, and boundary complexities.},
  keywords = {surface defect segmentation, metallic tile inspection, hybrid attention mechanism, bidirectional feature pyramid network, deep supervision, industrial quality control},
  issn = {3068-9287},
  publisher = {Institute of Central Computation and Knowledge}
}

Article Metrics

Citations
Crossref
0
Scopus
0
Views
25
PDF Downloads
2

Publisher's Note

ICCK stays neutral with regard to jurisdictional claims in published maps and institutional affiliations.

Rights and Permissions

Institute of Central Computation and Knowledge (ICCK) or its licensor holds exclusive rights to this article under a publishing agreement with the author(s) or other rightsholder(s); author self-archiving of the accepted manuscript version of this article is solely governed by the terms of such publishing agreement and applicable law.
ICCK Transactions on Sensing, Communication, and Control
ICCK Transactions on Sensing, Communication, and Control
ISSN: 3068-9287 (Online) | ISSN: 3068-9279 (Print)
Portico
Preserved at
Portico