Skip to content
Elucidário Madeirense Home

Jev (TypeSafe System One) benchmark on Elucidário tasks

The model was jev-1.13.0, called with the official SDK typesafe-sdk 0.7.1. Code is in elucidario/bench/jev_bench.py and raw outputs are in data/bench/.

Task A: triage of out-of-lexicon tokens (OCR error vs legitimate)

Results (errors caught / false alarms / errors missed):

System Caught (of 26) False alarms Missed
Sonnet 5 (no thinking) 17 12 9
Haiku 4.5 17 23 9
Jev, Choice (2 labels), argmax 12 26 14
Jev, Choice, P(err) ≥ 0.3 17 63 9
Jev, Noul (explicit criteria) ≥ 0.5 14 31 12
Jev, Noul ≥ 0.7 8 5 18
Jev, 6-way fine Choice, argmax 6 10 20

Task B: segmentation (article / subarticle / reject), 196 reviewed candidates

System Accuracy Sub-article recall
Opus 5.5 99.5% 100%
Haiku 4.5 92.3% 96%
Sonnet 5 76.0% 45%
Jev 61.7% 11%

Caveat: the gold came from a Claude review, so the Opus score is biased upwards.

Jev's confidence does not help here either (64.6% accuracy at confidence ≥ 0.85). The decision needs indirection: using the neighbouring headwords to infer that an umbrella article is in progress. The vendor documents indirection as a weakness of jev-1.13.

Decision

Task C (Phase 6 re-test): link homonym disambiguation, English context

Recently viewed

    Pages you read will appear here.