--- license: cc-by-nc-4.0 library_name: peft base_model: Qwen/Qwen3-4B base_model_relation: adapter pipeline_tag: text-classification language: - en - hi - bn - ta - te - kn - ml datasets: - AmazonScience/massive - bitext/Bitext-customer-support-llm-chatbot-training-dataset - LingoIITGN/COMI-LINGUA - PolyAI/banking77 - DeepPavlov/clinc150 tags: - decision-making - system-one - calibrated-probabilities - lora - qwen3 - indic - hinglish - hindi - non-generative --- # Nirṇaya — India-first System-One Decision Model (v1.0) **Nirṇaya** ("निर्णय", *decision*) is an open, India-first **System-One decision model** in the style of TypeSafe AI's Jev. Given a `state` (arbitrary text — often Indic-language, code-mixed, or Hinglish) plus one or more **typed questions**, Nirṇaya returns **calibrated probabilities in a single forward pass**, with no text generation. - **Backbone:** [Qwen/Qwen3-4B](https://huggingface.co/Qwen/Qwen3-4B) + LoRA (r=16, α=32, 7 target modules) - **Decision heads:** PointerHead (choice / score) + NoulHead (binary) - **Precision:** bf16 backbone, fp32 heads - **Authors:** Basab Ghosh, Sagar Sarkale, Abhijeet Katte - **Published by:** [cmul8.com](https://cmul8.com) Nirṇaya is designed for structured decision-making in Indian enterprise contexts: intent classification, script/language identification, urgency scoring, and yes/no routing on text that mixes Devanagari, Latin, and multiple Indic scripts. --- ## Question types Every question in Nirṇaya is one of three typed shapes: | Type | Description | Output | |---|---|---| | `choice` | Pick one of 2–64 options | One probability per option (sums to 1) | | `score` | Pick a level on an ordered scale (e.g. 1–5) | Full distribution over K levels + expected value | | `noul` | Yes/no | P(true) ∈ [0, 1] | Each forward pass returns probabilities and a scalar confidence per question — no generated tokens, no chain-of-thought. --- ## Intended use - **Direct use:** Structured decision APIs where you already know the option space at query time (intent routing, ticket triage, moderation flags, LID, escalation scoring). - **Downstream use:** Feature source for larger decision pipelines that want calibrated probabilities rather than free-form text. - **Non-goals:** open-ended generation, dialogue, code generation, summarisation. --- ## Evaluation All numbers are on the **first 1000 items per dataset** where dataset size permits; smaller datasets are run in full. Comparison baselines were measured under the same protocol. ### IndicJevBench v1 (cmul8-hf/IndicJevBench) | Dataset | n | Accuracy | ECE | IndicJevScore | p50 (ms) | |---|---|---|---|---|---| | fintech_banking77 | 1000 | **0.989** | 0.011 | **97.4** | 76.8 | | intent_massive | 1000 | **0.935** | 0.027 | **94.9** | 76.3 | | synthetic_enterprise | 444 | **0.831** | 0.077 | **89.7** | 71.5 | | hinglish_lid | 1000 | 0.155 | 0.553 | 40.9 | 79.9 | **Head-to-head vs. published baselines** (best prior baseline per dataset): | Dataset | Nirṇaya v1.0 | Best prior | Δ accuracy | Δ IndicJevScore | |---|---|---|---|---| | fintech_banking77 | 0.989 / 97.4 | GPT-4o-mini 0.775 / 56.1 | **+21.4pp** | **+41.3** | | intent_massive | 0.935 / 94.9 | GPT-4o-mini 0.580 / 51.3 | **+35.5pp** | **+43.6** | | synthetic_enterprise | 0.831 / 89.7 | GPT-4o-mini 0.670 / 46.6 | **+16.1pp** | **+43.1** | Weighted across the three comparable datasets (n=2444): **Nirṇaya 0.938** vs. GPT-4o-mini (best prior) **0.676** — a **+26.2pp gain**. ### JevBench v1 (English, public tiers) | Tier | n | Accuracy | Brier | ECE (15 bins) | |---|---|---|---|---| | easy | 48 | 0.896 | 0.160 | **0.095** | | original | 72 | 0.569 | 0.498 | 0.164 | | hard | 111 | 0.315 | 0.961 | 0.388 | | **weighted** | **231** | **0.515** | — | — | Nirṇaya trades some English JevBench performance for large Indic-side gains. This is a deliberate, India-first design choice — see *Limitations* below. ### Calibration ECE ≤ 0.077 on the three strong IndicJevBench datasets — well below the 0.10 practical-utility threshold. Nirṇaya's decision heads output normalised probabilities directly, so calibration is a first-class property, not a post-hoc adjustment. --- ## Limitations, biases and known failure modes - **Hinglish script detection (v1.0):** accuracy is 0.155 on the IndicJevBench `hinglish_lid` task — well below the 0.635 majority baseline. v1.0's training mix contains **no explicit script-detection examples**, so the model has no learned prior for that task shape. A continuation-training run targeting this gap is in progress; expect a follow-up release. - **English JevBench regression:** the *original* tier drops from 0.681 (mid-training) to 0.569 (final) — a −11pp regression. This is consistent with continued Indic-heavy fine-tuning drifting the decision heads away from English task phrasing. This is expected: Nirṇaya is India-first by design. - **Score-type calibration on hard tasks:** ECE on JevBench *hard* tier is 0.388 (catastrophically overconfident on the hardest questions). Practical deployment should apply per-(type, K) temperature scaling on your own calibration split. A calibration recipe will ship alongside the inference package. - **Not a generator:** Nirṇaya does not produce free-form text. Use a generative model for open-ended tasks. - **Language coverage:** trained on 7 Indic languages (`hi bn ta te kn ml`) + English + Hinglish. Marathi, Gujarati and other Indic languages will be added in future releases. - **Prompt-injection surface:** because Nirṇaya reads `state` as raw text, adversarial content in the state can still influence probabilities. Do not use Nirṇaya as the sole safety gate for adversarial inputs. --- ## Bias and fairness considerations Training data is dominated by Indic customer-support and enterprise dialog, plus translated English datasets. It over-represents formal service-desk and payments language, and under-represents casual social-media Hinglish. Users evaluating Nirṇaya on out-of-distribution informal text should expect degradation and are encouraged to run a small labelled probe before deployment. Language coverage is uneven: Hindi (both scripts) and English are the most-represented; Malayalam and Kannada have the smallest translated slices. --- ## How to use **Install** (Python ≥ 3.10, CUDA GPU with bf16 support — Ampere or newer): ```bash pip install torch --index-url https://download.pytorch.org/whl/cu124 # match your CUDA pip install git+https://github.com/cmul8/nirnaya.git@v0.1.0 ``` **Run:** ```python from nirnaya import NirnayaModel model = NirnayaModel.from_pretrained("cmul8-hf/nirnaya") result = model.decide( state="बिजली बिल का पेमेंट अभी तक नहीं हुआ, urgent hai please", questions=[ {"type": "choice", "options": ["billing", "technical", "sales"]}, {"type": "score", "k": 5, "instructions": "How urgent is this?"}, {"type": "noul", "instructions": "Should this be escalated?"}, ], ) # result[0].probs -> [0.87, 0.09, 0.04] # result[1].expected_level -> 4.2 # result[2].p_true -> 0.71 ``` All three questions above share **one** backbone forward pass. - **Inference package (source, install, tests):** [github.com/cmul8/nirnaya](https://github.com/cmul8/nirnaya) --- ## License - **Nirṇaya v1.0 weights (this repository):** [Creative Commons Attribution-NonCommercial 4.0 International (CC BY-NC 4.0)](https://creativecommons.org/licenses/by-nc/4.0/). - **Base model:** [Qwen/Qwen3-4B](https://huggingface.co/Qwen/Qwen3-4B) is released under Apache-2.0; its terms continue to apply to the base weights. **In short:** research, evaluation, personal use, and academic publication are free. Commercial use — including internal use inside a for-profit company on production workloads — is **not** permitted without a separate commercial licence from cmul8.com. --- ## Citation ```bibtex @misc{nirnaya2026v1, title = {Nirṇaya: An India-first System-One Decision Model}, author = {Ghosh, Basab and Sarkale, Sagar and Katte, Abhijeet}, year = {2026}, url = {https://huggingface.co/cmul8-hf/nirnaya}, note = {v1.0} } ``` --- ## Acknowledgements - **Backbone:** Alibaba Cloud's Qwen team ([Qwen3-4B](https://huggingface.co/Qwen/Qwen3-4B)) - **Translation:** Meta AI's [NLLB-200](https://huggingface.co/facebook/nllb-200-distilled-1.3B) - **Training frameworks:** [transformers](https://github.com/huggingface/transformers), [PEFT](https://github.com/huggingface/peft), PyTorch - **Benchmark inspiration:** TypeSafe AI's Jev / JevBench - **Companion benchmark:** [IndicJevBench](https://huggingface.co/datasets/cmul8-hf/IndicJevBench) — released alongside Nirṇaya