PaperScope
LIVE · 2026-09-29 05:40 UTC

Precision As You Need: Stochastic Computing Is a Dense Adaptive Quantizer

Haoran Jin, Kangqi Zhang, Jirong Yang, Barry Lyu, Qiuyi Ding, Ruijie Gao, Nathan Bleier

Latestcs.CLcs.LGcs.AIcs.CV
arXiv ID
2609.32922 v1
Category
Submitted
2026-09-26

Abstract

Matrix multiplications dominate the inference cost of modern transformer-based vision models, yet existing efficiency techniques such as post-training quantization and mixed-precision inference are largely limited to the small set of fixed-width formats (INT4, INT8, BF16, and FP16) supported by conventional accelerators. We revisit stochastic computing (SC) as a way to lift this constraint: viewed as a dense adaptive quantizer, SC controls precision by bit-stream length L rather than a fixed datapath, while each multiplication reduces to a single AND/XNOR gate. We build a GPU library that emulates SC matrix multiplication at scale, exposes stream lengths as first-class kernel arguments, and evaluates SC end-to-end on image classification, object detection and instance segmentation, class-conditional image generation, and visual world-model planning. On top of this substrate, we develop a dynamic per-row mixed-precision policy that assigns stream length per token or group at matched average budget, requires no retraining, and uses the same SC hardware across schedules. Across tasks, SC remains competitive with fixed-format INT quantization at matched bit budgets, while per-row mixed precision helps maintain accuracy at lower average stream lengths. These results provide software-level feasibility evidence that SC can serve as a dense-precision substrate for fine-grained mixed-precision inference on modern vision transformers.

Comment: NeurIPS 2026

arXiv abs page · PDF · same-day batch