PaperScope
LIVE · 2026-09-29 05:40 UTC

Jev Matches 7B Language Models for Speech-Neuroprosthesis Rescoring

Gabriele Cinà

Latestcs.CLcs.LGcs.AIcs.CV
arXiv ID
2609.33538 v1
Category
Submitted
2026-09-27

Abstract

A speech neuroprosthesis decodes attempted speech from brain activity and ends by rescoring the decoder's candidate sentences with a language model of several billion parameters, the only component that needs a GPU. Replacing that model with a cheaper one is hard: general language models asked to pick one sentence from a list answer from where a label sits in the list rather than from the sentence itself. We pose rescoring as a single typed decision, one call that returns a probability for every candidate, served by Jev, a hosted model trained for calibrated decisions, and combine it with the decoder's own score. On 978 held-out sentences from a participant with ALS, where the published decoder alone reaches 8.1% word error, Jev reaches 7.5% against 7.8% for both OPT-6.7b and Qwen2.5-7B; with the decoder's weight re-tuned, 6.9% against 7.2% and 7.4%. Jev is ahead in all four comparisons and at most 0.2 points behind at the 95% bound. It costs 0.07 USD per thousand sentences and needs no GPU; a dedicated GPU running a 7B model is cheaper per sentence only above 43% utilisation, far beyond what one user generates. End-to-end latency over the internet is 262 ms, of which 62 ms is spent at the provider, the same order as a 7B model on a local GPU (27 ms) but not faster.

Comment: 8 pages, 1 figure, 4 tables. Code and data: https://github.com/gabrycina/how-much-language-model

arXiv abs page · PDF · same-day batch