Error Taxonomy and Failure Analysis of Large Language Models for BIM Scripting

Document Type

Conference Proceeding

Publication Date

1-1-2026

Abstract

Building Information Modeling (BIM) has improved design and construction processes, yet modeling workflows remain heavily manual and error-prone. Large language models (LLMs) offer new opportunities to automate BIM tasks through natural language prompts, but their behavior in practice is not well understood. This study examines errors in LLM-generated BIM scripts using GPT-4 in a Revit-based Python scripting environment. Twenty modeling tasks of varying complexity were tested using an iterative, user-guided debugging process. Errors were systematically recorded and consolidated into a structured taxonomy. The analysis identified ten recurring error categories, with API Misuse the most frequent. Failures were more common and diverse in higher-complexity tasks, and many required extensive correction cycles. While all tasks were ultimately completed, results show that current LLMs remain dependent on human-in-the-loop validation. These findings provide an early structured assessments of LLM error patterns in BIM automation and offer practical insights for designing error-aware BIM workflows.

Publication Title

Construction Research Congress 2026 Advanced Technologies Artificial Intelligence and Data Analytics in Construction Selected Papers from Construction Research Congress 2026

ISBN

[9780784486979]

Share

COinS