Instructions to use AMAImedia/MiMo-V2.6-Flash-MOPD-BPW2.5-GGUF with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use AMAImedia/MiMo-V2.6-Flash-MOPD-BPW2.5-GGUF with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf AMAImedia/MiMo-V2.6-Flash-MOPD-BPW2.5-GGUF:Q8_0 # Run inference directly in the terminal: llama cli -hf AMAImedia/MiMo-V2.6-Flash-MOPD-BPW2.5-GGUF:Q8_0
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf AMAImedia/MiMo-V2.6-Flash-MOPD-BPW2.5-GGUF:Q8_0 # Run inference directly in the terminal: llama cli -hf AMAImedia/MiMo-V2.6-Flash-MOPD-BPW2.5-GGUF:Q8_0
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf AMAImedia/MiMo-V2.6-Flash-MOPD-BPW2.5-GGUF:Q8_0 # Run inference directly in the terminal: ./llama-cli -hf AMAImedia/MiMo-V2.6-Flash-MOPD-BPW2.5-GGUF:Q8_0
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf AMAImedia/MiMo-V2.6-Flash-MOPD-BPW2.5-GGUF:Q8_0 # Run inference directly in the terminal: ./build/bin/llama-cli -hf AMAImedia/MiMo-V2.6-Flash-MOPD-BPW2.5-GGUF:Q8_0
Use Docker
docker model run hf.co/AMAImedia/MiMo-V2.6-Flash-MOPD-BPW2.5-GGUF:Q8_0
- LM Studio
- Jan
- vLLM
How to use AMAImedia/MiMo-V2.6-Flash-MOPD-BPW2.5-GGUF with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "AMAImedia/MiMo-V2.6-Flash-MOPD-BPW2.5-GGUF" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "AMAImedia/MiMo-V2.6-Flash-MOPD-BPW2.5-GGUF", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }'Use Docker
docker model run hf.co/AMAImedia/MiMo-V2.6-Flash-MOPD-BPW2.5-GGUF:Q8_0
- Ollama
How to use AMAImedia/MiMo-V2.6-Flash-MOPD-BPW2.5-GGUF with Ollama:
ollama run hf.co/AMAImedia/MiMo-V2.6-Flash-MOPD-BPW2.5-GGUF:Q8_0
- Unsloth Desktop
- Pi
How to use AMAImedia/MiMo-V2.6-Flash-MOPD-BPW2.5-GGUF with Pi:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf AMAImedia/MiMo-V2.6-Flash-MOPD-BPW2.5-GGUF:Q8_0
Configure the model in Pi
# Install Pi: npm install -g @earendil-works/pi-coding-agent # Add to ~/.pi/agent/models.json: { "providers": { "llama-cpp": { "baseUrl": "http://localhost:8080/v1", "api": "openai-completions", "apiKey": "none", "models": [ { "id": "AMAImedia/MiMo-V2.6-Flash-MOPD-BPW2.5-GGUF:Q8_0" } ] } } }Run Pi
# Start Pi in your project directory: pi
- Docker Model Runner
How to use AMAImedia/MiMo-V2.6-Flash-MOPD-BPW2.5-GGUF with Docker Model Runner:
docker model run hf.co/AMAImedia/MiMo-V2.6-Flash-MOPD-BPW2.5-GGUF:Q8_0
- Lemonade
How to use AMAImedia/MiMo-V2.6-Flash-MOPD-BPW2.5-GGUF with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull AMAImedia/MiMo-V2.6-Flash-MOPD-BPW2.5-GGUF:Q8_0
Run and chat with the model
lemonade run user.MiMo-V2.6-Flash-MOPD-BPW2.5-GGUF-Q8_0
List all available models
lemonade list
- Hermes Agent
How to use AMAImedia/MiMo-V2.6-Flash-MOPD-BPW2.5-GGUF with Hermes Agent:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf AMAImedia/MiMo-V2.6-Flash-MOPD-BPW2.5-GGUF:Q8_0
Configure Hermes
# Install Hermes: curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash hermes setup # Point Hermes at the local server: hermes config set model.provider custom hermes config set model.base_url http://127.0.0.1:8080/v1 hermes config set model.default AMAImedia/MiMo-V2.6-Flash-MOPD-BPW2.5-GGUF:Q8_0
Run Hermes
hermes
- Atomic Chat
- OpenClaw
How to use AMAImedia/MiMo-V2.6-Flash-MOPD-BPW2.5-GGUF with OpenClaw:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf AMAImedia/MiMo-V2.6-Flash-MOPD-BPW2.5-GGUF:Q8_0
Configure OpenClaw
# Install OpenClaw: npm install -g openclaw@latest # Register the local server and set it as the default model: openclaw onboard --non-interactive --mode local \ --auth-choice custom-api-key \ --custom-base-url http://127.0.0.1:8080/v1 \ --custom-model-id "AMAImedia/MiMo-V2.6-Flash-MOPD-BPW2.5-GGUF:Q8_0" \ --custom-provider-id llama-cpp \ --custom-compatibility openai \ --custom-text-input \ --accept-risk \ --skip-health
Run OpenClaw
openclaw agent --local --agent main --message "Hello from Hugging Face"
⚡ Each donation funds the next large quant.
I host free GGUF or MoE quants as independent research.
Local hardware: Mechrevo Kuangshi GM7AG0M — RTX 3060 Laptop 6GB GDDR6, 64GB DDR5, i7-12700H (14C/20T, 4.7GHz), Windows 11, Samsung 990 Pro.
Good for imatrix and 0.6–35B-class work in RAM. 9B+ and searches need rented H200/Blackwell, typically $100 per quant.
🎉 Boosty🦄 | ☕ Buy Me a Coffee🦄 | ⭐ DonationAlerts🦄
💚 Thanks to Hugging Face for extra storage.🦄
NOESIS / AMAImedia
Released as part of the NOESIS Professional Multilingual Dubbing Automation Platform (framework: DHCF-FNO — Deterministic Hybrid Control Framework for Frozen Neural Operators)
- Founder: Ilia Bolotnikov
- Organization: AMAImedia.com
- X (Twitter): @AMAImediacom
- LinkedIn: Ilia Bolotnikov
- Telegram: @djbionicl
- Release date: 2026-09-27
AMAImedia
Original model: XiaomiMiMo/MiMo-V2.6-Flash-MOPD
GGUF (source of this mirror): AesSedai/MiMo-V2.6-Flash-MOPD-GGUF
Quantizer: ed's bpw-size PR (llama.cpp #15550)
Upstream model card. The block below is reproduced from the official XiaomiMiMo/MiMo-V2.6-Flash-RL card. It describes the MiMo-V2.6 series base model — architecture, training recipe and benchmark numbers for the BF16/FP8 checkpoints. It is not a measurement of this repository's BPW2.5 GGUF. For the quant recipe, verified metrics and file list of this mirror, see below.
MiMo-V2.6-Flash-RL
Scaling Reinforcement Learning Toward Self-Improvement
1. Introduction
MiMo-V2.6-Flash-RL is the efficiency-balanced checkpoint of the MiMo-V2.6 series. The series is built to scale reinforcement learning toward self-improvement — scaling RL compute, environment diversity, and grader compute together, so the model keeps expanding its capability frontier through exploration and feedback. Key features include:
- Native Omnimodal + Long Horizon: Text, image, video, and audio in one model; 1M tokens for long repositories, tool traces, and multi-session agent runs.
- You Only RL Once: One mixed RL run across coding, general agents, visual, and cybersecurity — not separate per-domain runs. Tasks and multiple harnesses are mixed in the same batch so capabilities reinforce each other and strategies transfer to harnesses never seen in training.
- Scaling RL Compute: Fully asynchronous Group Relative Policy Optimization (GRPO) on very large batches — 1,568 prompts × 16 rollouts per step, billions of tokens per update.
- Groupwise Agentic Grading (Self-Improvement Loop): Binary pass/fail cannot rank passing solutions, so the reward signal itself is scaled. An agentic grader compares rollouts within each group: Groupwise Reward Synthesis (GRS) builds task-specific rubrics offline from contrasting rollouts and fuses rubric quality with test outcomes; Groupwise Advantage Redistribution (GAR) ranks passing trajectories online and moves advantage toward higher-quality solutions. Judged against the policy’s own samples, this closes a self-improvement loop and steers toward shorter paths and fewer tokens per task.
- Aligned RL: Cold start from self-correction — the model reflects on and rewrites its own misaligned turns into grounded next steps. Throughout RL, environment hardening, adversarial screening, and verifier cross-checks keep the loop honest against reward hacking.
- Multi-Prefix Multi-Teacher On-Policy Distillation (MOPD2): After mixed RL, MOPD2 combines autonomous student rollouts with prefix-conditioned single-turn rollouts (Teacher-Prefix and SFT-Prefix), reusing histories from teacher trajectories and SFT demonstrations so decision points train without regenerating preceding turns — extending capabilities to hard-to-verify tasks.
Model Summary
- Architecture: Sparse MoE (Mixture of Experts), 309B total / 15B activated parameters
- Context Length: 1M tokens
- Modalities: Text, Image, Video, Audio
- Vision Encoder: 681M-param MiMo ViT (28 layers: 24 SWA + 4 Full)
- Audio Encoder: 308M AudioTokenizer + 127M audio patch encoder
- Multi-Token Prediction (MTP): 5-layer speculative decoder
Figure 1. MiMo-V2.6 architecture.
3. Evaluation Results
| Benchmark | MiMo-V2.6 Pro | MiMo-V2.6 Flash | MiMo-V2.5 Pro | Claude Opus 5 | GPT-5.6 Sol | Claude Fable 5 |
|---|---|---|---|---|---|---|
| Code Agent | ||||||
| DeepSWE v1.1 | 71.9 | 67.9 | 19.0 | 74.0 | 73.0 | 70.0 |
| ProgramBench | 26.5 | 26.0 | 12.5 | 37.0 | 25.0 | 33.0 |
| MiMo Code Bench | 63.2 | 61.2 | 40.4 | 68.6 | 59.3 | - |
| General Agent | ||||||
| AutomationBench v1.0.6 | 53.1 | 52.3 | 16.0 | 50.3 | 45.8 | 46.2 |
| Toolathlon-Verified | 76.9 | 73.6 | 49.1 | 80.6 | 74.9 | 77.9 |
| GDPval-AA 2.1 | 1673 | - | 1107 | 1708 | 1588 | 1595 |
| Agents’ Last Exam | 31.6 | 27.6 | 13.2 | 31.6 | 30.8 | 25.7 |
| Terminal Bench 4.0 | 34.9 | 28.8 | 1.5 | 49.0 | 39.9 | 42.4 |
| Terminal Bench 2.1 | 89.9 | 87.6 | 65.2 | 89.1 | 88.8 | 84.3 |
| OSWorld-Verified | 82.0 | 80.8 | - | 83.4 | 83.0 | 86.0 |
| JobBench | 62.0 | 61.2 | 25.0 | 65.7 | 45.4 | 57.4 |
| Cybersecurity | ||||||
| CyberGym | 94.0 | 95.1 | 40.0 | - | - | - |
| MiMo Cyber Bench | 80.2 | 77.2 | 0.0 | - | - | - |
| ExploitGym | 17.8 | 6.0 | 0.2 | 22.1 | 30.3 | 28.4 |
| ExploitBench | 47.9 | 25.3 | 16.6 | 70.0 | 78.5 | 78.0 |
| SEC Bench Pro | 66.3 | 47.5 | 17.7 | - | 79.1 | - |
| Visual Agent | ||||||
| MiMo VisualCoding | 72.3 | 71.5 | - | 70.0 | 73.4 | 69.1 |
4. Model Architecture
LLM Backbone
| Component | MiMo-V2.6-Flash-RL |
|---|---|
| Layers (Total / SWA / GA) | 48 / 39 / 9 |
| Hidden Size | 4096 |
| SWA Heads (Q/KV) | 64 / 8 |
| GA Heads (Q/KV) | 64 / 4 |
| Head Dimensions (QK / V) | 192 / 128 |
| Sliding Window Size | 128 |
| Routed Experts (Total / Activated) | 256 / 8 |
| Max Context Length | 1M |
| MTP / Speculative Decoder | 5 SWA layers, window 1024 |
The first Transformer block uses global attention with a dense FFN. Remaining blocks interleave local SWA and GA; both use sparse MoE FFNs without shared experts.
Vision Encoder (MiMo ViT)
| Configuration | Value |
|---|---|
| Layers (Total / SWA / GA) | 28 / 24 / 4 |
| Hidden Size | 1280 |
| Attention Heads (Q / KV) | 32 / 8 |
| Head Dimension | 64 |
| Patch Size (T × H × W) | 2 × 16 × 16 |
| Sliding Window (Left / Right) | 64 / 64 |
| Spatial Merge Size | 2 × 2 |
| Parameters | 681M |
Audio Encoders
AudioTokenizer encoder: 24 layers (12 SWA / 12 GA), hidden 1024, 20 RVQ codebooks, 308M parameters. Audio patch encoder: 6 layers, 127M parameters; four frames per patch (25 Hz → 6.25 Hz).
Speculative Decoder
5-layer SWA MTP drafter (DFlash-style). Predicts 7 subsequent tokens per forward pass for parallel verification.
Citation
@misc{mimo2026v26flash,
title={MiMo-V2.6-Flash-RL},
author={{Xiaomi MiMo Team}},
year={2026},
howpublished={\url{https://huggingface.co/XiaomiMiMo/MiMo-V2.6-Flash-RL}},
}
Contact
For questions or feedback, reach us at mimo@xiaomi.com or join our community:
MiMo-V2.6-Flash-MOPD — BPW2.5 GGUF
This is a mirror of the upstream BPW2.5 quant by.
The quantization was not re-run by AMAImedia. The three shards were copied
server-side and their LFS OIDs and byte sizes were verified identical to the
source repository. No re-quantization, no re-calibration, no re-packaging of the
tensor payload was performed here.
Upstream license and terms apply. Weights derive from (MIT).
What was mirrored
Only the BPW2.5 tier plus the multimodal projector and the KLD/PPL analysis
data. Other tiers (BPW2.0, BPW3.0, BPW3.5, MXFP4) remain upstream-only.
| File | Bytes | GB |
|---|---|---|
| BPW2.5/MiMo-V2.6-Flash-MOPD-BPW2.5-00001-of-00003.gguf | 6,342,912 | 0.006 |
| BPW2.5/MiMo-V2.6-Flash-MOPD-BPW2.5-00002-of-00003.gguf | 49,505,447,552 | 49.505 |
| BPW2.5/MiMo-V2.6-Flash-MOPD-BPW2.5-00003-of-00003.gguf | 47,229,347,840 | 47.229 |
| mmproj-MiMo-V2.6-Flash-MOPD-Q8_0.gguf | 1,564,252,928 | 1.564 |
Total main payload 96.74 GB. Download all three shards into the same directory and keep the projector alongside them.
Upstream quant recipe
The idea: the FFN tensors dominate total size, so quality is preserved by keeping the non-FFN tensors at high precision and quantizing FFN GATE + FFN UP + FFN DOWN down to the BPW target. The rest of the model stays at Q8_0 / Q6_K. BPW quants were produced with ed's bpw-size PR.
For BPW2.5 the non-FFN mixture is Q6_K, and the FFN tensors are quantized
down to reach 2.50 effective BPW.
Upstream measured quality
Measured by the upstream quantizer; reproduced here unchanged.
| Quant | Size | Mixture | PPL | 1-(Mean PPL(Q)/PPL(base)) | KLD |
|---|---|---|---|---|---|
| MXFP4 | 162.89 GiB (4.52 BPW) | BF16 / MXFP4 | 5.147382 ± 0.030566 | +0.0753% | -0.000000 ± 0.000000 |
| BPW3.5 | 126.18 GiB (3.50 BPW) | Q8_0 / varies | 5.248009 ± 0.031054 | +2.0317% | 0.134408 ± 0.000692 |
| BPW3.0 | 108.16 GiB (3.00 BPW) | Q8_0 / varies | 5.425520 ± 0.032234 | +5.4829% | 0.173319 ± 0.000870 |
| BPW2.5 | 90.09 GiB (2.50 BPW) | Q6_K / varies | 5.684465 ± 0.034102 | +10.5173% | 0.226887 ± 0.001093 |
| BPW2.0 | 72.07 GiB (2.00 BPW) | Q6_K / varies | 6.544785 ± 0.040620 | +27.2436% | 0.365580 ± 0.001645 |
Reference base PPL ≈ 5.143516 (derived from the MXFP4 row, which is the full-quality tier).
Per-quant detail: BPW2.5 notes, full dataset.
Usage
Requires a llama.cpp build with BPW quant type support — upstream states the BPW quants require llama.cpp PR #15550. The projector also requires MOPD multimodal support in the same runtime.
Note the model is very large. A 2.50 BPW quant of this MoE is roughly 90 GiB of weights, so plan host RAM/VRAM accordingly.
License and attribution
Upstream did not publish a separate license tag for the GGUF folder, so the upstream model card terms apply. Please cite the upstream quantizer rather than this mirror.
- Downloads last month
- 195
We're not able to determine the quantization variants.
Model tree for AMAImedia/MiMo-V2.6-Flash-MOPD-BPW2.5-GGUF
Base model
XiaomiMiMo/MiMo-V2.6-Flash-MOPD

