PaperScope
LIVE · 2026-10-02 05:40 UTC

Skeleton-and-Strategy Prompting: Training-Free Negation Understanding for Vision-Language Models

Yuliang Cai, Mohammad Rostami, Jesse Thomason

Latestcs.CLcs.LGcs.AIcs.CV
arXiv ID
2610.01180 v1
Category
Submitted
2026-10-01

Abstract

Despite the strong performance of Vision-Language Models (VLMs) on a wide range of visual question answering (VQA) tasks, these models consistently struggle to understand negation and produce incorrect answers when questions involve negated clauses. To address this limitation, we propose Skeleton-and-Strategy Prompting (\textbf{SSP}), a training-free, in-context learning method that improves VLM negation understanding capabilities without any parameter updates. Given a negation question, our method first abstracts the underlying question structure into a skeleton, retrieves a small set of same-skeleton questions from a lightweight question pool, then prompts the VLM to analyze their shared negation pattern and synthesize a single-sentence answering strategy. The skeleton and strategy are prepended to the test sample to guide the model correctly tackle the negation problems. Experiments on multiple negation VQA benchmarks show that SSP achieves state-of-the-art performance on negation-focused VQA tasks while remaining computationally efficient.

arXiv abs page · PDF · same-day batch