Article navigation

This paper evaluates the performance of multimodal large language models (MLLMs), ChatGPT-4o and Gemini in structural engineering with the image-input approach for the first time in the literature. The study uses 90 real visual-based fundamental mechanics questions that span five subtopics to examine the accuracy of ChatGPT and Gemini and compare them to the response of students. The results show that the students perform better than two MLLMs, although the performance of ChatGPT-4o is closely aligned with student performance or exceeds their performance on topics of kinematics and kinetics, whereas Gemini has the worst performance. However, the study also reveals the limitations of LLMs in accurately extracting mathematical and mechanical information from structural images, as well as their inadequacy in calculations and ensuring logical relationships. Evaluating the capabilities of MLLMs shows immense potential for structural engineering applications, while fine-tuning for specific tasks and developing artificial intelligence agents with MLLMs as the core represent critical directions for future research.

Licensed re-use rights only
You do not currently have access to this content.
Don't already have an account? Register

Purchased this content as a guest? Enter your email address to restore access.

Pay-Per-View Access
$41.00
Rental

or Create an Account

Close Modal
Close Modal