Nawadiraat
Nawadiraat preserves and opens up South-Asian literature across nine languages. We built the platform end to end: the literature library and Lughaat dictionaries, a from-scratch multilingual spell-checker, and a custom Urdu-Nastaliq OCR model that outperforms Tesseract's baseline and is now published as an open contribution.
Problem
Urdu and other Nastaliq scripts are notoriously hard for OCR: dense ligatures, diacritics and right-to-left flow break standard models. Off-the-shelf Tesseract Urdu (urd) produced garbled, unusable text, and there was no clean way to search or spell-check literature across Urdu, Sindhi, Farsi, Arabic and more.
Goal
Make a multi-script literature archive genuinely usable: accurate Urdu-Nastaliq OCR, a fast diacritic-insensitive dictionary and search, a real multilingual spell-checker, and a content platform covering nine provisioned locales including full RTL support.
AI solution
A custom Tesseract model, urd_naw, fine-tuned from the urd base on 12,748 curated Urdu image–text pairs (RTL training), delivers markedly cleaner ligature and diacritic recognition than the baseline. It runs in a dedicated Python/FastAPI OCR microservice the Laravel app calls over HTTP. Alongside it, a from-scratch Norvig-style spell-checker (Unicode-aware, diacritic/harakat-stripping) corrects Urdu and English, and dictionary search uses MySQL REGEXP normalization for diacritic-insensitive matching.
Workflow
- 1User uploads an Urdu page image to the OCR tool
- 2Laravel forwards it to the FastAPI OCR microservice
- 3urd_naw (Tesseract 5) extracts Nastaliq text, rendered in Noto Nastaliq
- 4Spell-checker normalizes and corrects Urdu/English input
- 5Lughaat search matches diacritic-insensitively across dictionaries
Model & AI components
- urd_naw: custom Urdu-Nastaliq Tesseract model (12,748 pairs)
- FastAPI + pytesseract OCR microservice (Dockerized)
- Norvig-style multilingual spell-checker (Unicode/RTL aware)
- Diacritic-insensitive dictionary search (MySQL REGEXP)
- Per-language feature flags (dictionary / spell-check / OCR)
Features
- Custom Urdu-Nastaliq OCR, open-sourced to Tesseract's tessdata_contrib
- Multilingual spell-checker for Urdu & English
- Lughaat: multiple dictionaries with rich lexicography
- Literature library: poets, poetry, prose, kids' stories, recitations
- Nine provisioned locales with full RTL support
- Filament admin with bulk CSV/XLSX import pipelines
Architecture
A Laravel 10 monolith (Blade + Livewire + Alpine, Filament 3 admin, MySQL) handles the site, content and dictionaries. OCR is isolated in a standalone Python FastAPI microservice running Tesseract 5 with the bundled urd_naw model, containerized with Docker and called over HTTP with a health check. Locale, RTL and per-language feature availability are all data-driven.
Frontend & dashboard
A fast, RTL-aware multilingual front end (literature browsing, Lughaat, spell-checker and OCR tools) with Nastaliq typography, backed by a Filament admin where the team bulk-imports words, poems and poets via CSV/XLSX.
Integrations
- FastAPI OCR microservice (internal HTTP)
- YouTube (recitations & video)
- Google Analytics 4
- Noto Nastaliq Urdu & Abdo Line fonts
- Partners: Govt of Punjab & Sindh, PILAC, Sindhi Language Authority
Deployment
The Laravel app and the Dockerized OCR service are deployed as separate containers (OCR on its own port with CORS locked to the domain and a /health check), so OCR scales independently of the website. The urd_naw model is published openly in Tesseract's tessdata_contrib repository.
Tech stack
- Web app
- Laravel 10Livewire 3Alpine.jsBladeFilament 3MySQL
- OCR microservice
- PythonFastAPITesseract 5pytesseractDocker
- Custom AI
- urd_naw Nastaliq OCRNorvig-style spell-checkerUnicode/RTL normalization
- Languages
- UrduSindhiFarsiArabicPunjabiPashtoBalochiBaltiEnglish
- Frontend
- Tailwind CSSNoto Nastaliq / Abdo LineDiacritic-insensitive search
Results
For years I wanted one place where our stories and poetry could live together, in every language, free for anyone to read. Rapit Labs made that real. And the first time the Urdu OCR read a page of Nastaliq back to me correctly, I genuinely got goosebumps.
Build something like Nawadiraat.
Tell us your workflow and goals. We'll map the highest-leverage AI use case and a clear path to production.