Gemma 4 12B (instruction-tuned) — ONNX for Transformers.js

A browser-ready ONNX export of google/gemma-4-12B-it, text only, for Transformers.js on WebGPU. Built for Oh my AI!, a chatbot that runs entirely in the browser.

As of September 2026, no other export of the 12B loads in Transformers.js: onnx-community publishes E2B and E4B only, and the existing 12B ONNX export targets onnxruntime-genai.

Files

File Content Size
onnx/embed_tokens_fp16.onnx input_ids → inputs_embeds, √hidden scale included 2.0 GB
onnx/decoder_model_merged_q4f16.onnx inputs_embeds, attention_mask, position_ids, past_key_values.* → last-position logits, present.*. MatMul weights in 4-bit blocks of 32, fp16 activations 6.7 GB, 4 shards

Total download: about 8.2 GB. It stays under the ~15.75 GiB of ArrayBuffers a Chrome tab can hold, which Transformers.js needs to keep every file in memory while loading.

Usage

import { pipeline } from '@huggingface/transformers';

const generator = await pipeline('text-generation', 'Dramatik999/gemma-4-12B-it-ONNX', {
  device: 'webgpu',
  dtype: { embed_tokens: 'fp16', decoder_model_merged: 'q4f16' },
});
const out = await generator([{ role: 'user', content: 'Why is the sky blue?' }], { max_new_tokens: 128 });

How it was built

Exported from Google's safetensors with torch.onnx.export (dynamo, opset 18, transformers 5.17, torch 2.14), no custom operators. The script is public: export_transformersjs.py.

  • The model declares gemma4_unified, unknown to Transformers.js 4.3. config.json presents it as gemma4 / Gemma4ForConditionalGeneration, which selects Transformers.js's text-only loading with two sessions. All computation lives in the graphs. The original type is kept under _source.
  • The decoder only computes last-position logits, the only ones generation reads.
  • The chat template is embedded in tokenizer_config.json, where Transformers.js's tokenizer reads it.

Validation

Against the PyTorch reference, greedy decoding:

Check Result
Embedding split, PyTorch zero logit difference
fp16 decoder, Transformers.js on CPU, 24-token prompt 32 / 32 tokens identical
q4f16 decoder, Transformers.js on CPU, 24-token prompt 8 / 32 identical, then a coherent quantization-induced divergence
q4f16, 3,679-token prompt, answer 3,000 tokens back 14 / 14 identical, correct answer
q4f16 in Chrome on WebGPU, same long prompt correct answer, about 26 s
Chat template, Transformers.js vs Python identical tokens

Limitations

  • Text only. The vision and audio encoders are not exported.
  • Key/value cache. Sliding-window layers keep their whole past instead of cropping it to 1,024 tokens. Results are correct, but the cache grows by about 340 KB per token. Oh my AI! caps the context at 8,192 tokens.
  • Speed. Attention runs as decomposed operators, without fused GroupQueryAttention.

License

Apache 2.0, like the original model.

Downloads last month
242
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for RASMUS/gemma-4-12B-it-ONNX

Quantized
(341)
this model