JEV task-classification benchmark: interim status
We are measuring how well two task classifiers answer the same routing requests about real engineering work. This page describes the method and where the measurement stands. It reports no accuracy: no figure has yet passed the controls we set for ourselves.
measurement in progress Status as of 30 September 2026. The page is updated as the measurement moves; nothing already recorded is rewritten.
The status table below is the whole state of the measurement: no component is complete, and none has produced a result that could support an accuracy figure.
- components tracked
- 7
- complete
- 0
What is measured
- Complete routing requests derived from the source of the Arcanada and Talomnia projects, in English and Russian. Each prompt-routing request carries six named questions of three typed kinds — Choice (pick one declared option), Score (a value on an ordered scale) and Noul (the probability that a statement is true).
- The planned selection is 1,000 requests and 5,900 decisions per solver: 490 prompt-routing and 10 pre-tool risk requests per project. These are selection requirements, not a frozen set: no request has yet passed independent admission.
- The pre-tool risk requests are the high-risk stratum, and that stratum is not measured (see the census history below). Any figure will therefore belong to the named estimand “BENCH classification-v1 ROUTE-ONLY (risk stratum excluded)” — by our arithmetic from the contract, 980 requests and 5,880 decisions per solver — and will carry the banner “risk stratum: NOT_MEASURED”. It will never be presented as the 1,000-request result.
- Two solvers receive byte-identical requests: JEV, a hosted classification service called through its public primitives, and Julia, a separately deployed classifier that uses the same request format. Neither sees the expected answer or its rationale. Expected answers (references) are admitted and frozen before any solver call and are never edited after solver output has been seen.
- An earlier exploratory set (V1) was different: 1,000 English cases with one question each, built from a factor design of work objects crossed with recurring patterns, frozen on 28 September. Its figures are withheld (see the controls below).
Protocol
- Every input that governs a run — questions, references, rubric, scoring rules, prompts — is frozen with a SHA-256 hash before that run. A changed byte means a new protocol revision, never a silent edit.
- Each solver call is made once. A failure stays in the record as a failure: no hidden retries, and no failure is scored as a zero answer.
- Stopping rules are declared before the run they govern. Every attempt, including failed and abandoned ones, is counted and will be published with the results.
- Changes of method are recorded as reversible decisions with an evidence gate and explicit reversal conditions (DEC-AUP-0065, DEC-AUP-0066, DEC-AUP-0067 in the program decision register).
Controls and what they found
Before any verdict may count toward a result, the instrument that produces it has to pass control items whose correct verdict is known in advance. Two control runs were stopped, and the method changed because of them.
-
V1
Exploratory run: figures withheld
An earlier exploratory run produced accuracy figures judged by an LLM (Luna). They are not published: on 2 of its 120 control items the judge passed a wrong answer with a false literal comparison, and its reliability was never established.
-
2026-09-29
Control run r4: stopped
In the first control batch (30 decisions) the judge restated the required value correctly and then contradicted its own restatement on 2 Noul boundary items (values 0.49 and 0.51). We tested whether the control item was at fault: the frozen rule and the judge prompt were the same rule applied to the same value, so the error was the judge’s. The batch stays recorded as failed and unchanged.
-
2026-09-30
Control run r5: stopped by a rule declared in advance
A new, explicitly declared judge instrument ran under a stopping rule fixed before the run: any contradiction ends it. In its first batch (30 decisions) the judge restated the set of allowed choices and then passed an answer outside that set. The run ended there; the remaining 23 batches (660 decisions) were never attempted, and no reworded r6 is allowed.
-
3 / 60
What the two runs measured
Across both batches the judge got 3 of 60 purely mechanical comparisons wrong — about 5%, with a 95% interval of roughly 1.7–13.7%. At that rate a 5,900-decision arm would carry on the order of 300 wrong verdicts. This is a measurement of the judge, not of either classifier.
-
DEC-AUP-0066
Successor: a deterministic scorer
Every typed verdict — Choice set membership, Score within ±0.5, Noul above or below 0.5 with exactly 0.5 as an abstention — is now computed by a frozen deterministic rule; judge output is never a verdict input. The rule is guarded by golden vectors frozen before any solver call, a mutation battery (in the reviewed battery 65 of 66 mutants were caught; the remaining one is provably equivalent) and a second, independently written implementation that must agree on every fixture. The scorer source will ship with the results.
-
next
What the LLM judge still does
Luna keeps two judgement stages only: whether a task state holds enough information to answer (adequacy), and whether a reference is supported (reference admission). Its reliability there is not measured and inherits nothing from r4 or r5; it needs new known-outcome controls, with at least 300 decisions per stage before any claim of an error rate below 1%.
Where the measurement stands
| Component | State |
|---|---|
| Request corpus (English and Russian, six questions per routing request) | selection requirements set; no request admitted or frozen yet |
| Deterministic scorer | source accepted by independent review; not yet run on solver output |
| Judge controls for adequacy and reference admission | prepared, not run — not measured |
| High-risk stratum (21 candidates per project required) | final census, attempt 6 of 6: 0 of 21 (Arcanada), 2 of 21 (Talomnia) — not measured; results will be route-only |
| JEV arm under the current protocol | not run |
| Julia arm | not run; endpoint not ready at the last check (29 September) |
| Accuracy | not measured — nothing is claimed |
High-risk stratum: census history
The high-risk stratum needs at least 21 independent candidates per project — the smallest number that gives a 90% chance of 10 qualifying at an assumed 60% qualification rate. Six attempts to find them are listed below; this is attempt 6 of 6, and there will be no seventh (DEC-AUP-0067).
| Attempt | Kind | Scope frozen before counting | Qualified (Arcanada / Talomnia) | Disposition |
|---|---|---|---|---|
| v3 | hand-built candidate list (24) | not a census | — / — | rejected by independent review at source admission |
| v4 | source-grounded hand-built candidates (20) | not a census | — / — | rejected by independent review: source and target feasibility |
| v5 | bounded source census (32 files) | not established | 0 / 0 | threshold not established |
| v6 | frozen source census (174 files) | yes | 0 / 0 | reserve gate failed |
| census_reserve_v1 | audit of 20 leads inside the v6 scope | yes (the v6 scope) | 0 / 0 | 14 of 20 never reached the risk question |
| v7 | post-hoc full-tree widening of the same three source commits | yes | 0 / 2 | paused safely; risk stratum not measured |
- v7 is a post-hoc widening adopted after the earlier scopes failed. It is not the originally planned scope, and the author of its scope had already seen the 0 of 21 result.
- Sensitivity (v3/v4 row): re-admitting the two operation families held in conflict would give 3 (Arcanada) and 5 (Talomnia) — still below 21. This row never entered the gate.
- The v7 scope was frozen by merge 1e431c04 of the program repository before the first prospect was read back. The census is final: no top-up, no re-rating, no newer commits or other repositories.
- SHA-256 commitments. Frozen v7 scope: ccff434add632bd570d3b70d9d0c7dcb92a940bf90556a339b87630dccec1955. Census outcome record (OUTCOME-V7.json, program register): b0be84cd4a9f397ef0650774fb435a0f63bfc65d0c01020b6d53ef915a306a2a.
What remains before any figure
- Admission of the request corpus with independent adequacy and reference audits, then its freeze.
- Known-outcome controls for the judge’s adequacy and reference stages, frozen and independently reviewed before any paid call.
- An isolation check of the evaluation environment before the next paid run of any kind.
- Paid runs of both solvers on identical requests, scoring by the frozen rule, and an export whose every figure can be recomputed from published files and hashes.
- In publication, accuracy will mean agreement with frozen, judge-admitted references and will be only as valid as those references. No figure will come from an LLM verdict, and numbers from before and after a change of method will never be pooled.
What this page leaves out
No private content: no host names, internal paths or credentials. We publish SHA-256 commitments of frozen protocol files; the files themselves are released with the export. Decision identifiers and the one commit identifier above refer to the program repository.