KinyaMed: Seeds, Not Rows -- What a Corpus Requirement Written in the Wrong Unit Fails to Constrain
Marius Bayizere
Abstract
Triage decides who is seen first. Building an urgency classifier for patient-voice Kinyarwanda, we found our specification could be met without producing anything it was meant to secure. We report that, and the instruments that detect it, instead of a classifier. Designed for the four languages a Rwandan health centre receives, with every instrument per-language: sentences are authored in all four arms and rows generate in one, because the frame slots those three need do not exist. Our requirement asked for one million examples; generation produced them in 130 seconds. It fails four of the nine quality gates in that specification, six when rows are attributed to their source sentence. The binding gate counts distinct authored seed phrases, not rows: at our 165, no corpus passes at any row count. Two shortfalls follow and differ: 2,835 further sentences to pass the seed-count gate, 19,835 to reach the stated million rows, because a separate gate caps a seed at 50 rows. Rows come from a machine at 7,700 per second; seeds from clinicians. A row count constrains the cheap quantity, leaves the expensive one free, and so does not constrain quality at all. Two further negatives follow. An evaluation set of 17,942 rows built from nine distinct sentences supports no verdict: our gate, which counts distinct sentences, refuses 38 of its cells and reports nothing. A model trained on a corpus of uniform surface form changes its predicted urgency for 31.5% of inputs under capitalisation and 21.0% under a single typo, reported as measurement and not attribution. A model directory shipped without its tokenizer loads without error and answers its class prior on input it cannot read, with well-formed probabilities. No figure here is evidence of model quality; the contribution is the apparatus and the negative results it produced, reproducible from a clean clone except where marked NOT REPRODUCIBLE.