Skip to content
Elucidário Madeirense Home

Phase 3 report: OCR correction

Pipeline

  1. Rules (elu ocrfix): deterministic, context-checked fixes. 1147 edits:
    • double-period: 378
    • zero-for-O: 313
    • clitic-hyphen: 231
    • space-before-punct: 65
    • letters-in-number: 60
    • not-sign: 39
    • digit-in-word: 29
    • zero-for-o: 15
    • hyphen-space: 6
    • space-after-colon: 4
    • headword-digit: 3
    • hyphen-as-dash: 2
    • split-word: 2
  2. LLM proofreading (elu ocrproof): Claude Opus 5.5 at low effort via the Batch API. The corpus went in 644 chunks, and every paragraph was marked with the words the lexicon did not recognise. The model returns only find/replace edits. Code checks each edit before applying it: the text must be found exactly once, the change must be small, and edits that remove circumflexes or turn ph/th into f/t (modernisation) are rejected. 1464 edits applied: {'punctuation': 274, 'ocr': 903, 'split_join': 54, 'number': 94, 'hyphen': 139}. Rejected: {'circumflex-modernisation': 2, 'find-not-found': 54, 'noop': 6, 'too-different': 1, 'find-ambiguous': 1}.
  3. Manual overrides from the review go in data/03_clean/manual_edits.yaml.
  4. Structure stage: 164 paragraphs split by page breaks were merged.

Every change is logged with its source, rule and before/after text in data/03_clean/edits.*.jsonl, so all of it can be reversed.

Sampled review (Claude subagent: 400 edits, 60 random paragraphs)

Kept on purpose

The 1921/1940 orthography (pôrto, Agôsto, theatro, pharmacia, mez, …) and the authentic spellings in quoted 15th–18th century documents are kept unchanged.

Second targeted pass (added after review)

Recently viewed

    Pages you read will appear here.