PaperScope
LIVE · 2026-10-01 05:40 UTC

Comparative study of adapting pre-trained models for driving behavior video captioning

Sayak Mallick, Philipp Geiger, Augustin Kelava

Latestcs.CLcs.LGcs.AIcs.CV
arXiv ID
2609.39542 v1
Category
Submitted
2026-09-30

Abstract

This report examines and compares some of the many fine tuning and prompting methods existing, applying them within the domain of autonomous driving. The idea is to compare these methods by adapting a Large Language Model (LLM) on a video dataset. LLM's have become extremely good at achieving a good understanding of different forms of data and this study aims to induce a low dimensional understanding of driving situations into our primary test model SpaceTimeGPT. Experiments on BDD-X (Berkeley DeepDrive eXplanation) dataset demonstrate good performance of the full fine tuning framework on some automatic metrics, and in some metrics, it even surpasses the baseline. We also try Low-Rank Adaptation (LoRA) and prompt engineering on VideoLLaVA model and discuss its limitations.

arXiv abs page · PDF · same-day batch