[Developers]

Evidence OCR and HTR Provenance Review

Scanned records, handwritten notes, custody forms, field notebooks, and historical archives often contain the most important evidence in a case, but they are slow to review and easy to misread. Evidence OCR and HTR Prove

Category: ForensicsLast Updated: Jul 16, 2026
forensicsaiblockchain

Overview#

Scanned records, handwritten notes, custody forms, field notebooks, and historical archives often contain the most important evidence in a case, but they are slow to review and easy to misread. Evidence OCR and HTR Provenance Review turns page images into searchable text while preserving the review trail that courts and oversight teams need.

The module supports selectable recognition engines with content-based default routing, page-level review, handwritten text recognition, text revision history, and signed provenance. Reviewers can approve or reject extracted text, compare it side by side with the source page, trace every accepted span back to the page image, and accept AI-proposed entities into the case record only after human review.

Key Features#

  • Selectable Recognition Engines: Choose a fast engine for printed text or a high-accuracy engine that handles handwriting; the choice is honoured on the first automatic pass and invalid selections are rejected up front. When no engine is chosen, content-based routing sends printed text to the fast engine and handwriting to the high-accuracy engine.
  • Engine Catalogue: A capability listing shows each recognition engine's speed, availability, handwriting support, and credential requirements before a run is started.
  • Handwritten Text Recognition: Process handwritten material with the handwriting-capable engine, including external HTR service integration where an organisation uses one.
  • Page-Level Review: Reviewers work at page level, with source image, extracted text, and approval state shown together, including a side-by-side comparison view for correcting text before it enters the record.
  • Immutable Text Revisions: Each approved or corrected text version is recorded as a revision rather than overwriting the past.
  • Span-Level Provenance: Extracted words and passages can be traced back to the document, page, and region that produced them.
  • Search and Entity Enrichment: Approved text becomes available for evidence search, and enrichment runs AI entity extraction over the stored extracted text, never re-running recognition, to produce a read-only preview of proposed entities.
  • Human-in-the-Loop Entity Apply: Nothing is persisted until an analyst accepts specific proposals; accepted items create entity profiles and relationships linked to the chosen investigation, profile, entity, or case. Proposals apply individually, so one faulty item does not fail the rest of the batch.
  • Reliable Action Feedback: Success and failure notifications reflect the real outcome of upload, review, and acceptance actions, so reviewers are never told a step succeeded when it did not.

Use Cases#

  • Witness Statement Digitisation: A caseworker uploads handwritten witness statements with the high-accuracy engine, then accepts proposed people, organisations, and locations straight into the investigation.
  • Historical Abuse Inquiry: A commission processes handwritten records and typed correspondence, approving text page by page before it becomes searchable.
  • Financial Crime Disclosure: Investigators OCR scanned invoices and ledgers, then preserve the link from every extracted field back to the source page.
  • Cold Case Review: Analysts make archived paper files searchable while retaining revision history for any corrected extraction.
  • Medical and Ambulance Records: Reviewers convert scanned clinical notes into searchable text while recording every approval action.
  • Bulk Evidence Triage: A large document production is OCR processed and queued for reviewer approval before any of it enters the searchable record.

Integration#

The module connects to evidence management, document preview, review queues, entity extraction, search, redaction, disclosure export, audit logging, and provenance reporting. Extraction and entity enrichment are available from both the evidence library and the investigations case workspace, with fully localised review screens. Only approved text is treated as verified evidence text, and only reviewer-accepted proposals become linked case entities. Unreviewed or rejected recognition output remains controlled and cannot silently replace the source document.

Open Standards#

  • W3C PROV-DM: Text extraction, review, correction, and approval events are represented as provenance activities linked to source pages and reviewers.
  • RFC 8785, JSON Canonicalisation Scheme: Deterministic serialisation supports reproducible signatures over text provenance records.
  • FIPS 180-4, SHA-256: Source pages, text revisions, and provenance manifests use cryptographic hashes for tamper evidence.
  • ISO 8601: Recognition, review, correction, and approval timestamps use standard date-time formatting.

Last Reviewed: 2026-07-16 Last Updated: 2026-07-16

Ready to Build?

Get started with our APIs or contact our integration team for support.