Instructions to use RASMUS/gemma-4-12B-it-ONNX with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers.js
How to use RASMUS/gemma-4-12B-it-ONNX with Transformers.js:
// npm i @huggingface/transformers import { pipeline } from '@huggingface/transformers'; // Allocate pipeline const pipe = await pipeline('text-generation', 'RASMUS/gemma-4-12B-it-ONNX');
Gemma 4 12B (instruction-tuned) — ONNX for Transformers.js
A browser-ready ONNX export of google/gemma-4-12B-it, text only, for Transformers.js on WebGPU. Built for Oh my AI!, a chatbot that runs entirely in the browser.
As of September 2026, no other export of the 12B loads in Transformers.js: onnx-community publishes E2B and E4B only, and the existing 12B ONNX export targets onnxruntime-genai.
Files
| File | Content | Size |
|---|---|---|
onnx/embed_tokens_fp16.onnx |
input_ids → inputs_embeds, √hidden scale included |
2.0 GB |
onnx/decoder_model_merged_q4f16.onnx |
inputs_embeds, attention_mask, position_ids, past_key_values.* → last-position logits, present.*. MatMul weights in 4-bit blocks of 32, fp16 activations |
6.7 GB, 4 shards |
Total download: about 8.2 GB. It stays under the ~15.75 GiB of ArrayBuffers a Chrome tab can hold, which Transformers.js needs to keep every file in memory while loading.
Usage
import { pipeline } from '@huggingface/transformers';
const generator = await pipeline('text-generation', 'Dramatik999/gemma-4-12B-it-ONNX', {
device: 'webgpu',
dtype: { embed_tokens: 'fp16', decoder_model_merged: 'q4f16' },
});
const out = await generator([{ role: 'user', content: 'Why is the sky blue?' }], { max_new_tokens: 128 });
How it was built
Exported from Google's safetensors with torch.onnx.export (dynamo, opset 18, transformers 5.17, torch 2.14), no custom operators. The script is public: export_transformersjs.py.
- The model declares
gemma4_unified, unknown to Transformers.js 4.3.config.jsonpresents it asgemma4/Gemma4ForConditionalGeneration, which selects Transformers.js's text-only loading with two sessions. All computation lives in the graphs. The original type is kept under_source. - The decoder only computes last-position logits, the only ones generation reads.
- The chat template is embedded in
tokenizer_config.json, where Transformers.js's tokenizer reads it.
Validation
Against the PyTorch reference, greedy decoding:
| Check | Result |
|---|---|
| Embedding split, PyTorch | zero logit difference |
| fp16 decoder, Transformers.js on CPU, 24-token prompt | 32 / 32 tokens identical |
| q4f16 decoder, Transformers.js on CPU, 24-token prompt | 8 / 32 identical, then a coherent quantization-induced divergence |
| q4f16, 3,679-token prompt, answer 3,000 tokens back | 14 / 14 identical, correct answer |
| q4f16 in Chrome on WebGPU, same long prompt | correct answer, about 26 s |
| Chat template, Transformers.js vs Python | identical tokens |
Limitations
- Text only. The vision and audio encoders are not exported.
- Key/value cache. Sliding-window layers keep their whole past instead of cropping it to 1,024 tokens. Results are correct, but the cache grows by about 340 KB per token. Oh my AI! caps the context at 8,192 tokens.
- Speed. Attention runs as decomposed operators, without fused
GroupQueryAttention.
License
Apache 2.0, like the original model.
- Downloads last month
- 242