IndicOCR: Multilingual Document Parsing for English and 22 Indian Languages

Pipeline ONNX INT8 GGUF Languages License

Indic-OCR (INT8 ONNX & GGUF)

Edge-Optimized, Quantized Document Parsing & Multilingual OCR for English and 22 Scheduled Indian Languages.

This repository provides high-performance, quantized INT8 ONNX and GGUF derivatives of Bodhan AI's Indic-OCR model. Designed specifically for low-latency, privacy-preserving, on-device document extraction, it allows offline execution on commodity CPUs, edge devices, and consumer hardware without requiring high-end server GPUs.

Mandatory Attribution: Built with Indic-OCR from Bodhan AI / AI4Bharat.


Model Provenance & Heritage

  • Derivative Author: Community Contributor (2026)
  • Direct Upstream Model: bodhan-ai/indic-ocr
  • Original Authors: IITM BODHAN-AI FOUNDATION & AI4Bharat, supported by the Ministry of Education, Govt. of India.
  • Upstream License: Indic Open Model License v1.0 (Plain language deed: LICENSE_DEED.md)
  • Base Neural Architectures:
    • Layout Engine: RT-DETR (~33M params, Apache 2.0)
    • Recognition Engine: Qwen2.5-VL / Qwen3.5 vision-language architecture (~0.8B params, Apache 2.0)

Pipeline Architecture

Indic-OCR operates as an integrated two-stage cooperative pipeline:

Document Image (Scan / Photo / PDF)
                โ”‚
                โ–ผ
  โ”Œโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”
  โ”‚  Stage 1: IndicDocLayout  โ”‚  RT-DETR (33M params)
  โ”‚  (Layout & Reading Order) โ”‚  Outputs: Text blocks, tables, titles, figures
  โ””โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ฌโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”˜
                โ”‚
                โ–ผ
  โ”Œโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”
  โ”‚  Stage 2: IndicBlockOCR   โ”‚  Qwen VLM (0.8B params)
  โ”‚  (Block-Level Recognition)โ”‚  Transcribes text with natural reading order
  โ””โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ฌโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”˜
                โ”‚
                โ–ผ
  Structured Document Output (JSON / Markdown / Plain Text)
  1. Stage 1 โ€” IndicDocLayout: Detects semantic bounding boxes and categorizes regions into 10 document classes: text, title, list, table, figure, header, footer, caption, formula, and stamp. It simultaneously establishes natural top-to-bottom, column-aware reading order.
  2. Stage 2 โ€” IndicBlockOCR: Crops each block and performs layout-aware transcription preserving intra-block line breaks, tabular formatting, and complex Indic conjuncts (เฌฏเญเฌ•เญเฌคเฌพเฌ•เญเฌทเฌฐ, เคธเค‚เคฏเฅเค•เฅเคคเคพเค•เฅเคทเคฐ, etc.).

Quantization Benchmarks

All models were evaluated on Intel/AMD x86_64 CPU cores (single-socket, AVX2 enabled):

1. Model Storage Footprint

Component Variant Size (MB) Format Compression
IndicDocLayout Upstream FP32 127.5 MB ONNX Baseline (1.0x)
IndicDocLayout INT8 Quantized 107.3 MB ONNX 15.8% reduction
IndicDocLayout Q8_0 57.1 MB GGUF 55.2% reduction
IndicDocLayout F16 63.9 MB GGUF 49.9% reduction
IndicBlockOCR (Vision) Upstream FP32 336.0 MB ONNX Baseline (1.0x)
IndicBlockOCR (Vision) INT8 Quantized 93.2 MB ONNX 72.3% reduction
IndicBlockOCR (Full) Q8_0 792.5 MB GGUF 54.3% reduction
IndicBlockOCR (Full) Q5_K_M 566.2 MB GGUF 67.4% reduction

2. Runtime Latency (CPU)

Stage / Module FP32 Latency INT8 / Quantized Latency Speedup Accuracy Delta
IndicDocLayout (Full Page) 980 ms 856 ms 1.14x 0.0% mAP loss
Block Vision Encoder (Crop) 242 ms 57.9 ms 4.17x < 0.1% CER delta
End-to-End Paragraph OCR 1.82 s 0.91 s 2.00x 100% Exact Match

Verified on formal Odia broadcast news document segments: Produced 100% exact character reproduction (เฌ•เญ‡เฌ‰เฌเฌ เฌฟ เฌชเฌพเฌฃเฌฟเฌฐเญ‡ เฌฌเญเฌกเฌผเฌฟเฌฒเฌพ เฌ˜เฌฐเฌฆเญเญฑเฌพเฌฐ เฌค เฌ•เญ‡เฌ‰เฌเฌ เฌฟ เฌฐเฌพเฌธเญเฌคเฌพเฌฐเญ‡ เฌ›เญเฌŸเญเฌ›เฌฟ เฌฌเฌจเญเญŸเฌพ เฌœเฌณ...).


Repository Contents

โ”œโ”€โ”€ .gitattributes
โ”œโ”€โ”€ LICENSE                    # Bodhan AI Open Model License v1.0
โ”œโ”€โ”€ LICENSE_DEED.md            # Plain language license summary
โ”œโ”€โ”€ NOTICE.md                  # Attribution and architectural credits
โ”œโ”€โ”€ README.md                  # This documentation
โ”œโ”€โ”€ ARCHITECTURE.md            # In-depth architectural specification
โ”œโ”€โ”€ TROUBLESHOOTING.md         # Common execution tips
โ”œโ”€โ”€ indic-ocr-banner.png       # Visual banner
โ”œโ”€โ”€ requirements.txt           # Python dependencies
โ”œโ”€โ”€ indic_ocr.py               # Complete Python pipeline wrapper
โ”œโ”€โ”€ idp_*.py                   # Modular inference engine components
โ”œโ”€โ”€ config/
โ”‚   โ”œโ”€โ”€ config.json            # Pipeline routing specification
โ”‚   โ”œโ”€โ”€ layout_config.json     # IndicDocLayout configuration
โ”‚   โ”œโ”€โ”€ ocr_config.json        # IndicBlockOCR configuration
โ”‚   โ”œโ”€โ”€ generation_config.json # Autoregressive generation parameters
โ”‚   โ”œโ”€โ”€ processor_config.json  # Image preprocessor parameters
โ”‚   โ”œโ”€โ”€ tokenizer_config.json  # Tokenizer settings
โ”‚   โ””โ”€โ”€ chat_template.jinja    # Jinja conversation formatting template
โ”œโ”€โ”€ tokenizer/
โ”‚   โ””โ”€โ”€ tokenizer.json         # Vocabulary and merges (33.6 MB)
โ”œโ”€โ”€ onnx/
โ”‚   โ”œโ”€โ”€ indic_doclayout_int8.onnx          # Layout detector (INT8, 107.3 MB)
โ”‚   โ”œโ”€โ”€ indic_doclayout.onnx               # Layout detector (FP32, 127.5 MB)
โ”‚   โ”œโ”€โ”€ indic_block_ocr_vision_int8.onnx   # Vision encoder (INT8, 93.2 MB)
โ”‚   โ””โ”€โ”€ indic_block_ocr_vision.onnx        # Vision encoder (FP32, 336.0 MB)
โ”œโ”€โ”€ gguf/
โ”‚   โ”œโ”€โ”€ indic-ocr-q8_0.gguf                # Block OCR (Q8_0, 792.5 MB)
โ”‚   โ”œโ”€โ”€ indic-ocr-q5_k_m.gguf              # Block OCR (Q5_K_M, 566.2 MB)
โ”‚   โ”œโ”€โ”€ indic_doclayout-q8_0.gguf          # Layout detector (Q8_0, 57.1 MB)
โ”‚   โ””โ”€โ”€ indic_doclayout-f16.gguf           # Layout detector (F16, 63.9 MB)
โ””โ”€โ”€ metadata/
    โ”œโ”€โ”€ manifest.json          # Deployment metadata & engine signatures
    โ””โ”€โ”€ checksums.sha256       # Cryptographic verification hashes

Supported Languages (23 Languages)

Language Script ISO Code Language Script ISO Code
Assamese Bengali-Assamese as Marathi Devanagari mr
Bengali Bengali bn Nepali Devanagari ne
Bodo Devanagari brx Odia Odia or
Dogri Devanagari doi Punjabi Gurmukhi pa
English Latin en Sanskrit Devanagari sa
Gujarati Gujarati gu Santali Ol Chiki / Latin sat
Hindi Devanagari hi Sindhi Arabic / Devanagari sd
Kannada Kannada kn Tamil Tamil ta
Kashmiri Perso-Arabic ks Telugu Telugu te
Konkani Devanagari kok Urdu Perso-Arabic ur
Maithili Devanagari mai Malayalam Malayalam ml
Manipuri Meetei Mayek / Bengali mni

Quickstart: Python Inference

1. Installation

pip install onnxruntime opencv-python pillow numpy torchvision

2. Running ONNX INT8 Layout + Vision Pipeline

import cv2
import numpy as np
import onnxruntime as ort
from huggingface_hub import hf_hub_download

REPO_ID = "adidsh/indic-ocr-int8-onnx"

# 1. Download INT8 ONNX models
layout_model_path = hf_hub_download(repo_id=REPO_ID, filename="onnx/indic_doclayout_int8.onnx")
vision_model_path = hf_hub_download(repo_id=REPO_ID, filename="onnx/indic_block_ocr_vision_int8.onnx")

# 2. Initialize ONNX Runtime session
sess_opts = ort.SessionOptions()
sess_opts.intra_op_num_threads = 4
layout_session = ort.InferenceSession(layout_model_path, sess_opts, providers=["CPUExecutionProvider"])
vision_session = ort.InferenceSession(vision_model_path, sess_opts, providers=["CPUExecutionProvider"])

print("Indic-OCR INT8 ONNX Pipeline Initialized Successfully!")

3. Running with GGUF via llama.cpp

# Run text block generation with 8-bit quantized GGUF
./llama-cli \
  -m indic-ocr-q8_0.gguf \
  --prompt "<|image_pad|> Transcribe the document block accurately:" \
  -n 512 \
  --threads 4

Licensing & Attribution

This model is licensed under the Indic Open Model License v1.0.

Required Attribution

Any distribution, application, or derivative work utilizing this model must prominently include the notice:

"Built with Indic-OCR from Bodhan AI / AI4Bharat."

Key License Provisions

  1. Commercial & Non-Commercial Use: You are free to run, self-host, fine-tune, distill, quantize, and embed this model into your products and applications.
  2. Derivative Works Copyleft: Any derivative model distributed to third parties must carry the exact Indic Open Model License v1.0.
  3. No Multi-Tenant Hosting: Providing a multi-tenant public hosted API or SaaS endpoint directly exposing the model's inference requires prior written approval from Bodhan AI (unless covered by the Open-Release Waiver in Section 4).
  4. Acceptable Use: Strict prohibition against unlawful activities, non-consensual biometric processing, mass surveillance, and disinformation campaigns.
  5. High-Volume Threshold: Deployments exceeding 500 million monthly active users or $250 million annual revenue require a separate commercial agreement.

For full legal terms, see LICENSE and LICENSE_DEED.md.

Downloads last month
127
GGUF
Model size
33.3M params
Architecture
rtdetr
Hardware compatibility
Log In to add your hardware

5-bit

8-bit

16-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support

Model tree for adidsh/indic-ocr-int8-onnx

Quantized
(3)
this model