Large Language and Vision Models for Construction Workflows: A Literature Review
Document Type
Conference Proceeding
Publication Date
1-1-2026
Abstract
The advancement of large multimodal language models (MLLMs) is reshaping innovation across sectors, with the construction industry emerging as a promising yet underexplored one. Characterized by high-risk environments, visual-rich data, and complex stakeholder dynamics, construction industry presents unique challenges and opportunities for MLLM integration. This paper presents a literature review of 73 recent studies, offering a structured mapping of how MLLMs are being applied across construction use cases. Key domains include safety monitoring, regulatory compliance, project planning, robotics, and generative design. The review highlights prevailing models, practical implementation strategies (e.g., zero-shot prompting, fine-tuning, and retrieval-augmented generation), and emerging trends such as autonomous agents. Beyond applications, the paper addresses core limitations around trust, context adaptation, and domain-specific evaluation. The findings offer both a snapshot of current practice and a strategic roadmap for researchers and practitioners aiming to harness MLLMs in construction workflows.
Publication Title
Construction Research Congress 2026 Advanced Technologies Artificial Intelligence and Data Analytics in Construction Selected Papers from Construction Research Congress 2026
ISBN
[9780784486962]
Recommended Citation
Erfani, A.,
Mansouri, A.,
&
Naghdi, M.
(2026).
Large Language and Vision Models for Construction Workflows: A Literature Review.
Construction Research Congress 2026 Advanced Technologies Artificial Intelligence and Data Analytics in Construction Selected Papers from Construction Research Congress 2026,
1, 421-431.
http://doi.org/10.1061/9780784486962.041
Retrieved from: https://digitalcommons.mtu.edu/michigantech-p2/3005