Skip to content
Custom OCR & Multilingual NLPLanguage, Culture & Heritage2024–2026

Nawadiraat

Nawadiraat preserves and opens up South-Asian literature across nine languages. We built the platform end to end: the literature library and Lughaat dictionaries, a from-scratch multilingual spell-checker, and a custom Urdu-Nastaliq OCR model that outperforms Tesseract's baseline and is now published as an open contribution.

9
Languages provisioned (Urdu, Sindhi, Farsi…)
12,748
Image–text pairs trained (urd_naw OCR)
Open-source
urd_naw model contributed to Tesseract
Gallery
The challenge

Problem

Urdu and other Nastaliq scripts are notoriously hard for OCR: dense ligatures, diacritics and right-to-left flow break standard models. Off-the-shelf Tesseract Urdu (urd) produced garbled, unusable text, and there was no clean way to search or spell-check literature across Urdu, Sindhi, Farsi, Arabic and more.

The objective

Goal

Make a multi-script literature archive genuinely usable: accurate Urdu-Nastaliq OCR, a fast diacritic-insensitive dictionary and search, a real multilingual spell-checker, and a content platform covering nine provisioned locales including full RTL support.

The approach

AI solution

A custom Tesseract model, urd_naw, fine-tuned from the urd base on 12,748 curated Urdu image–text pairs (RTL training), delivers markedly cleaner ligature and diacritic recognition than the baseline. It runs in a dedicated Python/FastAPI OCR microservice the Laravel app calls over HTTP. Alongside it, a from-scratch Norvig-style spell-checker (Unicode-aware, diacritic/harakat-stripping) corrects Urdu and English, and dictionary search uses MySQL REGEXP normalization for diacritic-insensitive matching.

How it works

Workflow

  1. 1User uploads an Urdu page image to the OCR tool
  2. 2Laravel forwards it to the FastAPI OCR microservice
  3. 3urd_naw (Tesseract 5) extracts Nastaliq text, rendered in Noto Nastaliq
  4. 4Spell-checker normalizes and corrects Urdu/English input
  5. 5Lughaat search matches diacritic-insensitively across dictionaries
Under the hood

Model & AI components

  • urd_naw: custom Urdu-Nastaliq Tesseract model (12,748 pairs)
  • FastAPI + pytesseract OCR microservice (Dockerized)
  • Norvig-style multilingual spell-checker (Unicode/RTL aware)
  • Diacritic-insensitive dictionary search (MySQL REGEXP)
  • Per-language feature flags (dictionary / spell-check / OCR)
Capabilities

Features

  • Custom Urdu-Nastaliq OCR, open-sourced to Tesseract's tessdata_contrib
  • Multilingual spell-checker for Urdu & English
  • Lughaat: multiple dictionaries with rich lexicography
  • Literature library: poets, poetry, prose, kids' stories, recitations
  • Nine provisioned locales with full RTL support
  • Filament admin with bulk CSV/XLSX import pipelines
System design

Architecture

A Laravel 10 monolith (Blade + Livewire + Alpine, Filament 3 admin, MySQL) handles the site, content and dictionaries. OCR is isolated in a standalone Python FastAPI microservice running Tesseract 5 with the bundled urd_naw model, containerized with Docker and called over HTTP with a health check. Locale, RTL and per-language feature availability are all data-driven.

Experience

Frontend & dashboard

A fast, RTL-aware multilingual front end (literature browsing, Lughaat, spell-checker and OCR tools) with Nastaliq typography, backed by a Filament admin where the team bulk-imports words, poems and poets via CSV/XLSX.

Connected

Integrations

  • FastAPI OCR microservice (internal HTTP)
  • YouTube (recitations & video)
  • Google Analytics 4
  • Noto Nastaliq Urdu & Abdo Line fonts
  • Partners: Govt of Punjab & Sindh, PILAC, Sindhi Language Authority
In production

Deployment

The Laravel app and the Dockerized OCR service are deployed as separate containers (OCR on its own port with CORS locked to the domain and a /health check), so OCR scales independently of the website. The urd_naw model is published openly in Tesseract's tessdata_contrib repository.

The build

Tech stack

Web app
Laravel 10Livewire 3Alpine.jsBladeFilament 3MySQL
OCR microservice
PythonFastAPITesseract 5pytesseractDocker
Custom AI
urd_naw Nastaliq OCRNorvig-style spell-checkerUnicode/RTL normalization
Languages
UrduSindhiFarsiArabicPunjabiPashtoBalochiBaltiEnglish
Frontend
Tailwind CSSNoto Nastaliq / Abdo LineDiacritic-insensitive search
Outcomes

Results

9
Languages provisioned (Urdu, Sindhi, Farsi…)
12,748
Image–text pairs trained (urd_naw OCR)
Open-source
urd_naw model contributed to Tesseract
For years I wanted one place where our stories and poetry could live together, in every language, free for anyone to read. Rapit Labs made that real. And the first time the Urdu OCR read a page of Nastaliq back to me correctly, I genuinely got goosebumps.
Founder · Nawadiraat
Get started

Build something like Nawadiraat.

Tell us your workflow and goals. We'll map the highest-leverage AI use case and a clear path to production.