Instructions to use HiTZ/Marmoka-es-Llama-3.1-8B-Instruct with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use HiTZ/Marmoka-es-Llama-3.1-8B-Instruct with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="HiTZ/Marmoka-es-Llama-3.1-8B-Instruct") messages = [ {"role": "user", "content": "Who are you?"}, ] pipe(messages)# pip install -U transformers accelerate # Load model directly from transformers import AutoTokenizer, AutoModelForCausalLM tokenizer = AutoTokenizer.from_pretrained("HiTZ/Marmoka-es-Llama-3.1-8B-Instruct") model = AutoModelForCausalLM.from_pretrained("HiTZ/Marmoka-es-Llama-3.1-8B-Instruct", device_map="auto") messages = [ {"role": "user", "content": "Who are you?"}, ] inputs = tokenizer.apply_chat_template( messages, add_generation_prompt=True, tokenize=True, return_dict=True, return_tensors="pt", ).to(model.device) outputs = model.generate(**inputs, max_new_tokens=256) print(tokenizer.decode(outputs[0][inputs["input_ids"].shape[-1]:])) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use HiTZ/Marmoka-es-Llama-3.1-8B-Instruct with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "HiTZ/Marmoka-es-Llama-3.1-8B-Instruct" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "HiTZ/Marmoka-es-Llama-3.1-8B-Instruct", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/HiTZ/Marmoka-es-Llama-3.1-8B-Instruct
- SGLang
How to use HiTZ/Marmoka-es-Llama-3.1-8B-Instruct with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "HiTZ/Marmoka-es-Llama-3.1-8B-Instruct" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "HiTZ/Marmoka-es-Llama-3.1-8B-Instruct", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "HiTZ/Marmoka-es-Llama-3.1-8B-Instruct" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "HiTZ/Marmoka-es-Llama-3.1-8B-Instruct", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }' - Docker Model Runner
How to use HiTZ/Marmoka-es-Llama-3.1-8B-Instruct with Docker Model Runner:
docker model run hf.co/HiTZ/Marmoka-es-Llama-3.1-8B-Instruct
Marmoka-es
A Lightweight Spanish Clinical Large Language Model
Built on Llama 3.1
Marmoka is a family of lightweight 8B-parameter clinical LLMs for English and Spanish. Marmoka-es is a Spanish member of the family, developed to address the scarcity of high-quality medical data and instructions in Spanish. It was developed through continual domain-adaptive pretraining (DAPT) of Llama-3.1-8B-Instruct on Spanish medical corpora, mixed with general-domain instructions to preserve its instruction-following ability. Marmoka-es consistently outperforms both Llama-3.1-8B-Instruct and Llama3.1-Aloe-Beta-8B on Spanish clinical multiple-choice QA, showing that robust medical LLMs can be developed for languages that remain low-resource in the medical domain.
- 📖 Paper: To Adapt or not to Adapt, Rethinking the Value of Medical Knowledge-Aware Large Language Models
How to Get Started with the Model
You can use the model with the transformers pipeline (it uses the Llama 3.1 chat template):
import torch
from transformers import pipeline
pipe = pipeline(
"text-generation",
model="HiTZ/Marmoka-es-Llama-3.1-8B-Instruct",
torch_dtype=torch.bfloat16,
device_map="auto",
)
messages = [
{"role": "system", "content": "Eres un experto médico. Responde a la pregunta de opción múltiple únicamente con la letra de la opción correcta."},
{"role": "user", "content": "¿Qué déficit vitamínico causa el escorbuto?\nA) Vitamina A\nB) Vitamina B12\nC) Vitamina C\nD) Vitamina D"},
]
outputs = pipe(messages, max_new_tokens=256)
print(outputs[0]["generated_text"][-1]["content"])
Model Description
- Developed by: Ane G. Domingo-Aldama, Iker De La Iglesia, Maitane Urruela, Aitziber Atutxa, and Ander Barrena
- Contact: Ane G. Domingo-Aldama and Iker De La Iglesia
- Language(s) (NLP): Spanish
- License: apache-2.0
- Funding:
- The Spanish Ministry of Science, Innovation and Universities, MCIN/AEI/10.13039/501100011033 projects:
- EDHIA PID2022-136522OB-C22 (also supported by FEDER, UE).
- TRUST-MED PID2025-174880OB-I00 (also supported by FEDER, UE).
- DeepR3 TED2021-130295B-C31 (also supported by the European Union NextGenerationEU/PRTR).
- Basque Government:
- Research group funding IT1570-22.
- HiTZ Center.
- Predoctoral grants: FPU23/03347 (I. De La Iglesia, Spanish Ministry of Science, Innovation and Universities), PRE_2024_1_0224 (A. G. Domingo-Aldama) and PRE_2025_1_0177 (M. Urruela), Basque Government.
- The Spanish Ministry of Science, Innovation and Universities, MCIN/AEI/10.13039/501100011033 projects:
Model Details
| Param. no. | ~8B |
| Architecture | Llama 3.1 (decoder-only) |
| Base model | Llama-3.1-8B-Instruct |
| Adaptation | DAPT on medical corpora + general-domain instructions |
| Medical corpus | ~908M (≈1B) words (Spanish) |
| General instructions | 310K (Spanish, Magpie) |
| Marmoka-es training | 3 models trained with different hyperparameters on the same corpus, merged |
| Training framework | Axolotl |
| Merging tool | MergeKit |
| Model selection | Casimedicos validation set (EN & ES) |
Training Data
Marmoka-es was trained with the hybrid strategy of Sainz et al. (2025), which combines continual domain-adaptive pretraining on medical corpora with general-domain instructions to avoid catastrophic forgetting. Unlike traditional DAPT setups, an instruction-tuned model (Llama-3.1-8B-Instruct), rather than a base model, was used as the starting point. Note that this setup includes medical corpora but no dedicated medical instructions; the amount of medical content in the Magpie instructions is minimal. The medical corpora are an updated version of those gathered in De la Iglesia et al. (2025), and were processed and deduplicated using standard approaches (deduplication code).
| Type | Source | Size |
|---|---|---|
| Medical corpora (Spanish) | WikiMed | 16.3M words |
| PubMed | 5.5M words | |
| Medical Crawler | 850.7M words | |
| MeSpEn | 6.9M words | |
| SciELO | 29M words | |
| General-domain instructions | Magpie (Llama-3.1-70B-Instruct) | 310K instructions |
Limitation and Bias
Marmoka-es is intended for research purposes. Its evaluation is based on short-form multiple-choice question answering with automated accuracy metrics and without expert clinical adjudication, so the validity of the model's underlying reasoning has not been verified, and these benchmarks may not fully capture genuine medical expertise. It is recommended to validate and test the model for specific applications before using it.
Instruction following and strict output formatting remain a limitation, especially in multi-step tasks such as summarization, where formatting failures were observed. Outputs should be checked whenever a strict format is required.
Biases
- Data Collection Bias: The training data for Marmoka-es was sourced from the web and open datasets, and may not adequately capture the full range of linguistic, demographic, or clinical variability found in real-world medical settings. This may introduce selection bias, and the model might reflect and propagate these biases in its outputs.
- Demographic and Linguistic Bias: Since the training data may not equally represent all demographic groups or linguistic nuances, the model may perform unequally across populations, potentially reinforcing existing health disparities and societal biases.
- Unexamined Ethical Considerations: No further assessments have been conducted to determine whether the training data contains biases or personally identifiable information, nor have the anonymization measures been independently verified. All data sources were previously processed, publicly released, and made available with stated limitations, including steps taken to ensure data quality and privacy compliance.
Disclaimer
Marmoka-es is intended for research purposes and is not approved for clinical use. It must not be used as a medical device or relied upon for diagnostic or therapeutic decision-making, and its outputs should not be interpreted as medical advice. Using the model without rigorous validation poses significant risks, including misinformation, misclassification, and potential harm.
When employed in actual clinical scenarios, all outputs must be independently reviewed and validated by qualified healthcare professionals. We do not take any liability for the use of this model.
Citing information
@article{domingo2026adapt,
title={To Adapt or not to Adapt, Rethinking the Value of Medical Knowledge-Aware Large Language Models},
author={Domingo-Aldama, Ane G and De La Iglesia, Iker and Urruela, Maitane and Atutxa, Aitziber and Barrena, Ander},
journal={arXiv preprint arXiv:2604.06854},
year={2026}
}
- Downloads last month
- 495