PaperScope
LIVE · 2026-10-01 05:40 UTC

CodeMimicry: Exploiting Safety Generalization Lag in Large Language Models via Structured Code Completion

Zhen Liang, Hai Huang, Wentao Chen

Latestcs.CLcs.LGcs.AIcs.CV
arXiv ID
2609.39902 v1
Category
Submitted
2026-09-30

Abstract

Large language models have achieved remarkable capabilities across diverse domains, yet their safety alignment remains vulnerable to jailbreak attacks. In this work, we identify a previously underexplored failure mode - safety generalization lag - where alignment trained predominantly on natural language fails to transfer to the code domain. We show that this lag induces a code-completion blind spot, allowing malicious intent embedded within syntactically valid code to evade safety mechanisms. To exploit this vulnerability, we propose CodeMimicry, a fully automated black-box jailbreak framework that generates structured, object-oriented code prompts to induce harmful outputs via code completion. Experiments on 8 state-of-the-art commercial LLMs demonstrate that CodeMimicry achieves a 96.25% attack success rate with 1.51 queries on average, significantly outperforming both template-based and optimization-based baselines. Beyond empirical performance, we provide a mechanistic analysis of code-based jailbreaks through latent space representations, including projection onto refusal-related directions and activation steering. This analysis offers an explanation of how CodeMimicry bypasses safety mechanisms in code-related domains. Our findings reveal a weakness in current safety alignment and highlight the need for robust alignments in structured domains such as code.

Comment: This paper will be accepted at NeurIPS 2026

arXiv abs page · PDF · same-day batch