PaperScope
LIVE · 2026-09-17 05:40 UTC

SKIP: a Self-knowledge-guided Step-wise Preference Learning Framework for Concise Reasoning

Qinhong Lin, Yuhao Zhang, Yinglun Feng, Zhongliang Yang, Linna Zhou

Latestcs.CLcs.LGcs.AIcs.CV
arXiv ID
2609.17019 v1
Category
Submitted
2026-09-15

Abstract

While Chain-of-Thought (CoT) reasoning has been proven to be effective, it often leads to overthinking, resulting in computational overhead, inference latency, and even degraded performance in large language models (LLMs). Existing concise reasoning frameworks significantly compromise accuracy while compressing the length of output. In this paper, we propose SKIP, a self-knowledge-guided step-wise preference learning framework. Starting with lightweight fine-tuning to adjust the model's output style, SKIP introduces a carefully designed knowledge probing mechanism to guide model to output an answer at each reasoning step. Based on the correctness of intermediate steps, we construct preference data that guide the model toward more efficient and correct reasoning by leveraging DPO. Experimental results demonstrate that our method effectively improves reasoning compression while mitigating performance degradation after fine-tuning. Besides, SKIP shows strong generalization ability on out-of-distribution datasets. We further conducted ablation studies on the component parameters of our framework.

Comment: 8 pages,3 figures. Accepted at IJCNN 2026

arXiv abs page · PDF · same-day batch