TECHNOLOGICAL BASIS OF “INDUSTRY 4.0”

Using AI agents to create work instructions for NDT and VT product inspection in heavy industry.

  • 1 Technical University in Kosice, Faculty of Mechanical Engineering, Department of Technology and Materials engineering, Košice, Slovakia

Abstract

This paper investigates the efficacy of generative artificial intelligence in automating the creation of technical documentation for non-destructive testing (NDT). The research focuses on a comparative performance analysis between a specialized AI agent, built on the Gemini 3.1 Flash platform, and certified human NDT Level 2 experts. A comparative cross-sectional study was conducted using three distinct industrial forging products. Both the AI agent and two human experts were tasked with generating Visual Testing (VT) work instructions based on a 14-point framework derived from the STN EN 13018 and ISO 9712 standards. The outputs were evaluated by a blind-reviewing Level 3 expert using a modified HEAT (Expertise, Accuracy, Trust) rubric, focusing on regulatory compliance, technical precision, and readability. he results demonstrate that the AI agent achieved a 98.4% reduction in generation time, averaging 54.5 seconds per instruction compared to 57.3 minutes for human experts. While the AI agent consistently outperformed humans in text clarity and structural consistency (scoring 5.0 in usability), it exhibited “conservative technical hallucinations,” such as prescribing unnecessary magnification tools and incorrect defect-coding standards (ISO 6520-1 instead of CSN 421240). Furthermore, cloud-based API instabilities (HTTP 503 errors) were identified as a critical reliability risk for real-time industrial deployment. The study concludes that while AI agents are highly effective as rapid drafting tools, they cannot currently replace certified personnel due to lack of situational engineering judgment and legal accountability. A “Human-in-the-loop” model remains mandatory, where a certified Level 2 or 3 professional must verify and approve all AI-generated NDT documentation to ensure industrial safety and regulatory compliance…

Keywords

References

  1. EN ISO 9712:2022. Non-destructive testing — Qualification and certification of NDT personnel. International Organization for Standardization. (2022).
  2. STN EN 13018. Non-destructive testing. Visual inspection. General principles. Bratislava: Office for Standardization, Metrology and Testing SR, (2016).
  3. ISO 11971:2020. Steel and iron castings – Visual testing of surface quality. 3rd ed. Geneva: INTERNATIONAL ORGANIZATION FOR STANDARDIZATION, (2020)
  4. Wang, R., Guo, J., Gao, C., Fan, G., Chong, C., & Xia, X. ,Can LLMs Replace Human Evaluators? An Empirical Study of LLM-as-a-Judge in Software Engineering. Proceedings of the ACM on Software Engineering, 2, 1955 - 1977. (2025). https://doi.org/10.1145/3728963.
  5. Asgari, E., Brown, N., Dubois, M., Khalil, S., Balloch, J., Yeung, J., & Pimenta, D. A framework to assess clinical safety and hallucination rates of LLMs for medical text summarisation. NPJ Digital Medicine, 8. (2025). https://doi.org/10.1038/s41746-025-01670-7.
  6. Cengiz, A., Sever, A., Umutlu, E., Erdem, N., Aytan, B., Tufan, B., Topraksoy, A., Darici, E., & Toraman, C. Evaluating the Quality of Benchmark Datasets for Low-Resource Languages: A Case Study on Turkish. ArXiv, abs/2504.09714. (2025). https://doi.org/10.48550/arxiv.2504.09714.
  7. Hu, X., Chen, Q., Wang, H., Xia, X., Lo, D., & Zimmermann, T. Correlating Automated and Human Evaluation of Code Documentation Generation Quality. ACM Transactions on Software Engineering and Methodology (TOSEM), 31, 1 - 28. (2022). https://doi.org/10.1145/3502853.
  8. Gregori-Giralt, E., & Menéndez-Varela, J. The content aspect of validity in a rubric-based assessment system for course syllabuses. Studies in Educational Evaluation, 68, 100971. (2021). https://doi.org/10.1016/j.stueduc.2020.100971.
  9. Kroll, M., & Kraus, K. Optimizing the role of human evaluation in LLM-based spoken document summarization systems. ArXiv, abs/2410.18218. (2024).https://doi.org/10.21437/interspeech.2024-2268.
  10. Teterevenkov, D. EXPERT-ORIENTED METHODS FOR EVALUATING THE QUALITY OF TEXTUAL GENERATION OF LARGE LANGUAGE MODELS. SOFT MEASUREMENTS AND COMPUTING. (2025). https://doi.org/10.36871/2618-9976.2025.05.003.
  11. Verhulsdonck, G., Weible, J., Stambler, D., Howard, T., & Tham, J. Incorporating Human Judgment in AI-Assisted Content Development: The HEAT Heuristic. Technical Communication.. (2024). https://doi.org/10.55177/tc286621.
  12. Tam, T.Y.C., Sivarajkumar, S., Kapoor, S. et al. A framework for human evaluation of large language models in healthcare derived from literature review. npj Digit. Med. 7, 258 (2024). https://doi.org/10.1038/s41746-024-01258-7
  13. ISO 6520-1:2007. Welding and allied processes — Classification of geometric imperfections in metallic materials — Part 1: Fusion welding. Geneva: International Organization for Standardization (2007)
  14. ČSN 42 1240, Casting defects. Nomenclature and classification of defects. Prague: Office for Standardization and Measurement. (1965)

Article full text

Download PDF