JEV task-classification benchmark: interim status

We are measuring how well two task classifiers answer the same routing requests about real engineering work. This page describes the method and where the measurement stands. It reports no accuracy: no figure has yet passed the controls we set for ourselves.

measurement in progress Status as of 30 September 2026. The page is updated as the measurement moves; nothing already recorded is rewritten.

The status table below is the whole state of the measurement: no component is complete, and none has produced a result that could support an accuracy figure.

components tracked
7
complete
0

What is measured

  • Complete routing requests derived from the source of the Arcanada and Talomnia projects, in English and Russian. Each prompt-routing request carries six named questions of three typed kinds — Choice (pick one declared option), Score (a value on an ordered scale) and Noul (the probability that a statement is true).
  • The planned selection is 1,000 requests and 5,900 decisions per solver: 490 prompt-routing and 10 pre-tool risk requests per project. These are selection requirements, not a frozen set: no request has yet passed independent admission.
  • The pre-tool risk requests are the high-risk stratum, and that stratum is not measured (see the census history below). Any figure will therefore belong to the named estimand “BENCH classification-v1 ROUTE-ONLY (risk stratum excluded)” — by our arithmetic from the contract, 980 requests and 5,880 decisions per solver — and will carry the banner “risk stratum: NOT_MEASURED”. It will never be presented as the 1,000-request result.
  • Two solvers receive byte-identical requests: JEV, a hosted classification service called through its public primitives, and Julia, a separately deployed classifier that uses the same request format. Neither sees the expected answer or its rationale. Expected answers (references) are admitted and frozen before any solver call and are never edited after solver output has been seen.
  • An earlier exploratory set (V1) was different: 1,000 English cases with one question each, built from a factor design of work objects crossed with recurring patterns, frozen on 28 September. Its figures are withheld (see the controls below).

Protocol

  • Every input that governs a run — questions, references, rubric, scoring rules, prompts — is frozen with a SHA-256 hash before that run. A changed byte means a new protocol revision, never a silent edit.
  • Each solver call is made once. A failure stays in the record as a failure: no hidden retries, and no failure is scored as a zero answer.
  • Stopping rules are declared before the run they govern. Every attempt, including failed and abandoned ones, is counted and will be published with the results.
  • Changes of method are recorded as reversible decisions with an evidence gate and explicit reversal conditions (DEC-AUP-0065, DEC-AUP-0066, DEC-AUP-0067 in the program decision register).

Controls and what they found

Before any verdict may count toward a result, the instrument that produces it has to pass control items whose correct verdict is known in advance. Two control runs were stopped, and the method changed because of them.

  1. V1

    Exploratory run: figures withheld

    An earlier exploratory run produced accuracy figures judged by an LLM (Luna). They are not published: on 2 of its 120 control items the judge passed a wrong answer with a false literal comparison, and its reliability was never established.

  2. 2026-09-29

    Control run r4: stopped

    In the first control batch (30 decisions) the judge restated the required value correctly and then contradicted its own restatement on 2 Noul boundary items (values 0.49 and 0.51). We tested whether the control item was at fault: the frozen rule and the judge prompt were the same rule applied to the same value, so the error was the judge’s. The batch stays recorded as failed and unchanged.

  3. 2026-09-30

    Control run r5: stopped by a rule declared in advance

    A new, explicitly declared judge instrument ran under a stopping rule fixed before the run: any contradiction ends it. In its first batch (30 decisions) the judge restated the set of allowed choices and then passed an answer outside that set. The run ended there; the remaining 23 batches (660 decisions) were never attempted, and no reworded r6 is allowed.

  4. 3 / 60

    What the two runs measured

    Across both batches the judge got 3 of 60 purely mechanical comparisons wrong — about 5%, with a 95% interval of roughly 1.7–13.7%. At that rate a 5,900-decision arm would carry on the order of 300 wrong verdicts. This is a measurement of the judge, not of either classifier.

  5. DEC-AUP-0066

    Successor: a deterministic scorer

    Every typed verdict — Choice set membership, Score within ±0.5, Noul above or below 0.5 with exactly 0.5 as an abstention — is now computed by a frozen deterministic rule; judge output is never a verdict input. The rule is guarded by golden vectors frozen before any solver call, a mutation battery (in the reviewed battery 65 of 66 mutants were caught; the remaining one is provably equivalent) and a second, independently written implementation that must agree on every fixture. The scorer source will ship with the results.

  6. next

    What the LLM judge still does

    Luna keeps two judgement stages only: whether a task state holds enough information to answer (adequacy), and whether a reference is supported (reference admission). Its reliability there is not measured and inherits nothing from r4 or r5; it needs new known-outcome controls, with at least 300 decisions per stage before any claim of an error rate below 1%.

Where the measurement stands

Component status, 30 September 2026
ComponentState
Request corpus (English and Russian, six questions per routing request)selection requirements set; no request admitted or frozen yet
Deterministic scorersource accepted by independent review; not yet run on solver output
Judge controls for adequacy and reference admissionprepared, not run — not measured
High-risk stratum (21 candidates per project required)final census, attempt 6 of 6: 0 of 21 (Arcanada), 2 of 21 (Talomnia) — not measured; results will be route-only
JEV arm under the current protocolnot run
Julia armnot run; endpoint not ready at the last check (29 September)
Accuracynot measured — nothing is claimed

High-risk stratum: census history

The high-risk stratum needs at least 21 independent candidates per project — the smallest number that gives a 90% chance of 10 qualifying at an assumed 60% qualification rate. Six attempts to find them are listed below; this is attempt 6 of 6, and there will be no seventh (DEC-AUP-0067).

Every attempt to build the high-risk stratum, in order
AttemptKindScope frozen before countingQualified (Arcanada / Talomnia)Disposition
v3hand-built candidate list (24)not a census— / —rejected by independent review at source admission
v4source-grounded hand-built candidates (20)not a census— / —rejected by independent review: source and target feasibility
v5bounded source census (32 files)not established0 / 0threshold not established
v6frozen source census (174 files)yes0 / 0reserve gate failed
census_reserve_v1audit of 20 leads inside the v6 scopeyes (the v6 scope)0 / 014 of 20 never reached the risk question
v7post-hoc full-tree widening of the same three source commitsyes0 / 2paused safely; risk stratum not measured
  • v7 is a post-hoc widening adopted after the earlier scopes failed. It is not the originally planned scope, and the author of its scope had already seen the 0 of 21 result.
  • Sensitivity (v3/v4 row): re-admitting the two operation families held in conflict would give 3 (Arcanada) and 5 (Talomnia) — still below 21. This row never entered the gate.
  • The v7 scope was frozen by merge 1e431c04 of the program repository before the first prospect was read back. The census is final: no top-up, no re-rating, no newer commits or other repositories.
  • SHA-256 commitments. Frozen v7 scope: ccff434add632bd570d3b70d9d0c7dcb92a940bf90556a339b87630dccec1955. Census outcome record (OUTCOME-V7.json, program register): b0be84cd4a9f397ef0650774fb435a0f63bfc65d0c01020b6d53ef915a306a2a.

What remains before any figure

  • Admission of the request corpus with independent adequacy and reference audits, then its freeze.
  • Known-outcome controls for the judge’s adequacy and reference stages, frozen and independently reviewed before any paid call.
  • An isolation check of the evaluation environment before the next paid run of any kind.
  • Paid runs of both solvers on identical requests, scoring by the frozen rule, and an export whose every figure can be recomputed from published files and hashes.
  • In publication, accuracy will mean agreement with frozen, judge-admitted references and will be only as valid as those references. No figure will come from an LLM verdict, and numbers from before and after a change of method will never be pooled.

What this page leaves out

No private content: no host names, internal paths or credentials. We publish SHA-256 commitments of frozen protocol files; the files themselves are released with the export. Decision identifiers and the one commit identifier above refer to the program repository.

All research