Benchmarking Vision-Language Models for Traffic Scene Understanding in Inclement Winter Weather: The AWDB Benchmark

Document Type

Conference Proceeding

Publication Date

1-1-2026

Abstract

Autonomous driving systems face significant challenges in winter conditions, and while vision-language models (VLMs) are increasingly being considered for inclusion in advanced driver assistance system (ADAS) pipelines for scene understanding, their reliability in adverse winter scenarios remains poorly understood. We introduce the Autonomous Winter Driving Benchmark (AWDB), a video benchmark comprising 509 dashcam-style clips with human annotations following a strict JSON schema. Each annotation includes a free-text scene description, a weather block with a winter weather flag, and a hazardous event block. The dataset is organized into four categories based on two binary dimensions: presence or absence of snow, and presence or absence of hazardous events. We evaluate closedsource and open-source VLMs at two frame sampling rates (1 FPS and 5 FPS), measuring schema compliance, entitygrounding, and response quality through standard captioning metrics (BLEU, ROUGE, METEOR, CIDEr, SPICE, BERTScore). Our results reveal substantial differences across models. While winter weather detection achieves near-ceiling accuracy (0.97-0.99), accident detection remains challenging (0.51-0.79 accuracy). Higher frame rates (5 FPS) generally improve both schema compliance and response quality, particularly for open-source models. Benchmark resources including video data, humanannotations, and model captions are available on https://github.com/Ali-Awad/Autonomous-Winter-Driving-Benchmark-AWDB-/tree/main.

Publication Title

Proceedings 2026 IEEE Cvf Winter Conference on Applications of Computer Vision Workshops Wacvw 2026

ISBN

[9798331591496]

Share

COinS