PaperScope
LIVE · 2026-10-02 05:40 UTC

IrekoGPT: Turning Structured Pruning into Post-Hoc Slimmable LLMs

Pietro Moriello, Pietro Buzzega, Angelo Porrello, Simone Calderara

Latestcs.CLcs.LGcs.AIcs.CV
arXiv ID
2610.00426 v1
Category
Submitted
2026-09-30

Abstract

We introduce IrekoGPT, a post-hoc method for converting pretrained LLMs into slimmable models whose width can be adjusted at inference time. Building on SliceGPT, we retain its projection matrices without pruning them, allowing a single model to expose nested subnetworks at different widths. We improve robustness by calibrating each layer across multiple compression ratios, and correct downstream linear layers through gradient-free ridge regression. Across Llama and Qwen models, preliminary results show improvements over naive PCA-based slimming, with the largest gains at high compression. Code is available at https://github.com/aimagelab/IrekoGPT

Comment: Accepted at the NeurIPS 2026 Workshop "AXIOM: Foundations of Efficient Deep Learning". 8 pages, 5 figures

arXiv abs page · PDF · same-day batch