PaperScope
LIVE · 2026-10-01 05:40 UTC

Debias It Yourself: Teaching LLMs Cognitive Bias Mitigation Interventions

Chahat Raj, Sina Mansouri, Aylin Caliskan, Antonios Anastasopoulos, Ziwei Zhu

Latestcs.CLcs.LGcs.AIcs.CV
arXiv ID
2609.40124 v1
Category
Submitted
2026-09-30

Abstract

Bias has long been studied in social psychology and cognitive science, where decades of research have produced a body of validated interventions that reduce stereotypical thinking and prejudiced responses in humans. We propose Debias It Yourself (DIY), a cognitively grounded framework that translates five such interventions into debiasing procedures for large language models and delivers them through three established paradigms: Show (in-context examples), Train (instruction tuning), and Revise (guided self-revision). Across three models, five bias benchmarks, eleven debiasing baselines, and three reasoning benchmarks, Train+Revise and Revise alone attain the top two average ranks, lead the bias-reasoning tradeoff (mean bias as low as 2% at 90% reasoning accuracy), and reduce bias on unseen dimensions by up to 14.8%. Our code and data are publicly available.

Comment: Under Review

arXiv abs page · PDF · same-day batch