Skip to content
Elucidário Madeirense Home

Translation strategy (result of the Phase 9 pilot)

Pilot

Quality (mean 1–5; fidelity = completeness and accuracy)

Model Fidelity en/de/hu/ru Fluency en/de/hu/ru Terminology en/de/hu/ru Chunk wins
Opus 5.5 medium 4.76 / 4.67 / 4.83 / 4.52 4.76 / 4.81 / 4.79 / 4.83 4.55 / 4.52 / 4.62 / 4.52 97
Opus 5.5 low 4.57 / 4.52 / 4.60 / 4.40 4.74 / 4.60 / 4.67 / 4.64 4.40 / 4.36 / 4.31 / 4.26 63
Sonnet 5 4.17 / 4.02 / 3.62 / 3.55 3.81 / 3.55 / 3.48 / 3.52 3.81 / 3.71 / 3.33 / 2.90 4
Haiku 4.5 3.40 / 2.60 / 2.05 / 2.74 3.52 / 3.00 / 2.05 / 2.81 3.31 / 2.81 / 2.24 / 2.50 4

Typical errors

Prompt caching

The prompt cache must be warmed before each batch. Unwarmed parallel batch requests each wrote their own copy of the cache; after a single warm-up request, 41 of 41 requests read it.

Recommendation

Metadata mini-pilot (296 English units: abstracts, chapter titles and summaries, person/place pages and notes, events → de, hu, ru)

Blind A/B judging by Opus 5.5.

Sonnet 5 (accuracy / naturalness) Opus 5.5 low (accuracy / naturalness)
de 4.84 / 4.66 4.90 / 4.85
hu 4.73 / 4.41 4.95 / 4.84
ru 4.76 / 4.50 4.87 / 4.80

Decision (owner, 2026-09-27): Sonnet 5 for metadata in all languages except Hungarian, which uses Opus 5.5 low (kb/translation_config.yaml). Its accuracy is within 0.1–0.2 of Opus. The gap is mostly naturalness, and it is largest in Hungarian, so Hungarian metadata may use Opus low. Remaining errors are minor: an idiom here and there, the work's title rendered "Elucidárium", and quotation-mark style. These are addressed in the language guides.

OpenAI Sol benchmark (same 42 chunks × en/de/hu/ru; same prompts, packages and schema; OpenAI Batch API)

The judge was Opus 5.5 at high effort, blind and shuffled, comparing four candidates: Opus low, Sonnet 5, gpt-5.6-sol (low) and gpt-6-sol (low).

Fidelity en de hu ru
Opus 5.5 low 4.55 4.50 4.50 4.17
gpt-5.6-sol 4.48 4.33 4.50 4.55
gpt-6-sol 4.10 4.07 4.45 4.19
Sonnet 5 4.05 3.88 3.21 3.33

Candidate routing for article bodies:

Languages Model
en, de, fr, it, nl Opus 5.5 low
ru, uk gpt-5.6-sol
hu gpt-6-sol

Follow-up checks and final routing (owner decision 2026-09-27)

Final routing (kb/translation_config.yaml):

Content Languages Model
Article bodies en, de, fr, it, nl Opus 5.5 low
Article bodies ru, uk gpt-5.6-sol
Article bodies hu gpt-6-sol
Metadata all gpt-6-sol

gpt-6-sol reasoning-effort test (de + ru, 42 chunks each; judged against Opus low and gpt-5.6-sol)

Setting Fidelity de / ru Reasoning tokens
gpt-6-sol low 4.07 / 4.19 (earlier run) 5k
gpt-6-sol medium 4.19 / 4.17 50k
gpt-6-sol high 4.29 / 4.43 176k
gpt-5.6-sol low 4.33 / 4.45 about 30k
Opus 5.5 low 4.31 / 4.10 –

Operational notes (en/uk/hu run, 2026-09-27)

Recently viewed

    Pages you read will appear here.