This paper evaluates the performance of multimodal large language models (MLLMs), ChatGPT-4o and Gemini in structural engineering with the image-input approach for the first time in the literature. The study uses 90 real visual-based fundamental mechanics questions that span five subtopics to examine the accuracy of ChatGPT and Gemini and compare them to the response of students. The results show that the students perform better than two MLLMs, although the performance of ChatGPT-4o is closely aligned with student performance or exceeds their performance on topics of kinematics and kinetics, whereas Gemini has the worst performance. However, the study also reveals the limitations of LLMs in accurately extracting mathematical and mechanical information from structural images, as well as their inadequacy in calculations and ensuring logical relationships. Evaluating the capabilities of MLLMs shows immense potential for structural engineering applications, while fine-tuning for specific tasks and developing artificial intelligence agents with MLLMs as the core represent critical directions for future research.
Article navigation
1 November 2025
Research Article|
October 21 2025
Evaluating performance of large language models on fundamental structural knowledge
Yi Zhang
;
Yi Zhang
Department of Civil Engineering,
The University of Hong Kong
, Hong Kong, PR China
Search for other works by this author on:
Ray Kai Leung Su
Department of Civil Engineering,
The University of Hong Kong
, Hong Kong, PR China
Corresponding author Ray Kai Leung Su (klsu@hku.hk)
Search for other works by this author on:
Corresponding author Ray Kai Leung Su (klsu@hku.hk)
Publisher: Emerald Publishing
Received:
August 24 2024
Accepted:
August 26 2025
Online ISSN: 1751-7702
Print ISSN: 0965-0911
© 2025 Emerald Publishing Limited
2025
Emerald Publishing Limited
Licensed re-use rights only
Proceedings of the Institution of Civil Engineers - Structures and Buildings (2025) 178 (11): 1009–1017.
Article history
Received:
August 24 2024
Accepted:
August 26 2025
Citation
Zhang Y, Su RKL (2025), "Evaluating performance of large language models on fundamental structural knowledge". Proceedings of the Institution of Civil Engineers - Structures and Buildings, Vol. 178 No. 11 pp. 1009–1017, doi: https://doi.org/10.1680/jstbu.24.00139
Download citation file:
New and popular articles
Suggested Reading
Welcome to the Gemini era: Google DeepMind and the information industry
Library Hi Tech News (December,2023)
Evaluating generative AI tools in Arabic academic libraries: a comparative analysis of ChatGPT-4o and Gemini’s performance and practical queries
Global Knowledge, Memory and Communication (April,2025)
Analytical seismic performance assessment of hollow reinforced-concrete bridge columns
Magazine of Concrete Research (May,2018)
Numerical modelling of interfacial transition zone influence on elastic modulus of concrete
Magazine of Concrete Research (April,2019)
Partitioning based reduced order modelling approach for transient analyses of large structures
Engineering Computations (January,2009)
Related Chapters
DIFFUSION COEFFICIENT OF CHLORIDE IONS UNDER SIMULATED CONDITIONS
Cement Combinations for Durable Concrete: Proceedings of the International Conference held at the University of Dundee, Scotland, UK on 5–7 July 2005
AI-assisted Programming and AI Literacy in Computer Science Education
Effective Practices in AI Literacy Education: Case Studies and Reflections
The Dark Side of Artificial Intelligence in Retail Innovation
Retail Futures: The Good, the Bad and the Ugly of the Digital Transformation
Recommended for you
These recommendations are informed by your reading behaviors and indicated interests.
