PaperScope
LIVE · 2026-09-29 05:40 UTC

JET: Justification Evaluation in Transformer

Shenghao Ding

Latestcs.CLcs.LGcs.AIcs.CV
arXiv ID
2609.33874 v1
Category
Submitted
2026-09-27

Abstract

JET uses pretrained language and vision-language models to select among a finite set of answers without additional training. It evaluates candidate likelihoods directly and shares computation across candidates. Experiments on desktop CPUs and consumer GPUs assess decision accuracy and execution cost. Qwen3.6-35B-A3B achieves 87.48% accuracy on the full MMLU test set and 3.69 requests per second on a separately timed MMLU subset. The accuracy-throughput comparison covers model, hardware, and reasoning choices, with Jev as an external reference. Controlled execution experiments show 2.18-2.23-fold speedups from prefix reuse and cache management, and a 30.8% reduction in process time from input preparation optimizations, with unchanged outputs. Optional reasoning has a task-dependent accuracy-throughput trade-off. These results support local decision inference from existing models.

Comment: 11 pages, 2 figures. Code: https://github.com/yet-another-ai/jet

arXiv abs page · PDF · same-day batch