Language

Sinhala & Tamil OCR

Optical character recognition for Sinhala and Tamil at 97% accuracy, digitising archives, court records and print collections at scale.

01 — Context

Vast amounts of Sri Lanka's written record — archives, court files, print collections — exist only on paper, in Sinhala and Tamil. Both are complex scripts, and both are largely ignored by the major OCR and AI labs, whose models are tuned for high-resource languages.

02 — The Problem

General-purpose OCR fails on Sinhala and Tamil. The scripts are dense with ligatures, modifiers and combining forms; document sources are old, degraded and inconsistently printed. Off-the-shelf accuracy was nowhere near usable for records where a single mis-read character can change a name, a date or a legal meaning.

03 — The Question

Could OCR read Sinhala and Tamil accurately enough to trust for archival and legal digitisation?

04 — Research

We treated it as a language-technology problem, not a scanning problem — studying the scripts’ structure, the failure modes of existing models, and the conditions of real source documents. Where usable training data didn’t exist, building the data was part of the work.

05 — Approach

Rather than adapt a general model, we built on models trained specifically for these scripts, with a pipeline designed around the realities of degraded, real-world documents.

06 — Engineering

A production OCR engine built to run against archives, court records and print collections at scale — engineered for throughput and for the messiness of decades-old material, not clean lab inputs.

07 — Validation

Accuracy was measured against real archival and record material, the conditions the system was actually meant to serve.

08 — Result

97% OCR accuracy on Sinhala and Tamil.

09 — Impact

Digitisation of Sinhala and Tamil archives, court records and print collections becomes feasible at scale — material that was effectively locked in paper becomes searchable and usable. [CONTENT REQUIRED: specific deployments / volumes]

10 — What We Learned

When the big labs skip your language, accuracy is an engineering problem you have to own end-to-end — from the data up.

11 — Technology
  • OCR models trained for Sinhala & Tamil
  • Document-processing pipeline
  • [CONTENT REQUIRED: precise stack]

Have a similar problem?

Start a Project →

Let's build what matters.

Start a Project →