Gemma-4 Ecosystem
Collection
Curated collection of Gemma-4 MLX optimizations, on-policy alignment corrections, and agentic fine-tunes. β’ 11 items β’ Updated
How to use True2456/Gemma-4-12B-Agentic-LoRA with MLX:
# Download the model from the Hub pip install huggingface_hub[hf_xet] huggingface-cli download --local-dir Gemma-4-12B-Agentic-LoRA True2456/Gemma-4-12B-Agentic-LoRA
Specialist LoRA adapter fine-tuned on top of mlx-community/gemma-4-12b-it-bf16 for multi-step autonomous tool calling, structured JSON execution, and multi-turn software engineering workflows.
Designed as Expert 1 within multi-specialist MoE fusion architectures or for standalone high-precision agent execution on Apple Silicon via Apple MLX (mlx_lm).
| Parameter | Specification |
|---|---|
| Base Model | mlx-community/gemma-4-12b-it-bf16 |
| Adapter Architecture | LoRA (Low-Rank Adaptation) |
| Target Layers | 48 Transformer Layers (q_proj, k_proj, v_proj, o_proj, gate_proj, up_proj, down_proj) |
LoRA Rank (r) |
16 |
LoRA Alpha (Ξ±) |
32 |
| LoRA Scale | 10.0 |
| Dropout | 0.05 |
| Max Sequence Length | 8192 tokens |
| Training Framework | mlx-lm on Apple Silicon Metal |
Trained on the curated gemma12b_agentic_specialist pack containing 107,761 training steps organized into three quality tiers:
mlx-lm)
pip install mlx mlx-lm
from mlx_lm import load, generate
model_path = "mlx-community/gemma-4-12b-it-bf16"
adapter_path = "True2456/Gemma-4-12B-Agentic-LoRA"
model, tokenizer = load(
model_path,
adapter_path=adapter_path
)
prompt = tokenizer.apply_chat_template([
{"role": "system", "content": "You are an autonomous engineering agent with tool access."},
{"role": "user", "content": "Inspect the repository structure and list files in src/."}
], tokenize=False, add_generation_prompt=True)
output = generate(
model,
tokenizer,
prompt=prompt,
max_tokens=512,
verbose=True
)
print(output)
grad_checkpoint: true) at --max-seq-length 8192 to prevent activation spilling and memory swap.Quantized