Marmoka-es
A Lightweight Spanish Clinical Large Language Model
Built on Llama 3.1


Marmoka is a family of lightweight 8B-parameter clinical LLMs for English and Spanish. Marmoka-es is a Spanish member of the family, developed to address the scarcity of high-quality medical data and instructions in Spanish. It was developed through continual domain-adaptive pretraining (DAPT) of Llama-3.1-8B-Instruct on Spanish medical corpora, mixed with general-domain instructions to preserve its instruction-following ability. Marmoka-es consistently outperforms both Llama-3.1-8B-Instruct and Llama3.1-Aloe-Beta-8B on Spanish clinical multiple-choice QA, showing that robust medical LLMs can be developed for languages that remain low-resource in the medical domain.

How to Get Started with the Model

You can use the model with the transformers pipeline (it uses the Llama 3.1 chat template):

import torch
from transformers import pipeline

pipe = pipeline(
    "text-generation",
    model="HiTZ/Marmoka-es-Llama-3.1-8B-Instruct",
    torch_dtype=torch.bfloat16,
    device_map="auto",
)

messages = [
    {"role": "system", "content": "Eres un experto médico. Responde a la pregunta de opción múltiple únicamente con la letra de la opción correcta."},
    {"role": "user", "content": "¿Qué déficit vitamínico causa el escorbuto?\nA) Vitamina A\nB) Vitamina B12\nC) Vitamina C\nD) Vitamina D"},
]

outputs = pipe(messages, max_new_tokens=256)
print(outputs[0]["generated_text"][-1]["content"])

Model Description

  • Developed by: Ane G. Domingo-Aldama, Iker De La Iglesia, Maitane Urruela, Aitziber Atutxa, and Ander Barrena
  • Contact: Ane G. Domingo-Aldama and Iker De La Iglesia
  • Language(s) (NLP): Spanish
  • License: apache-2.0
  • Funding:
    • The Spanish Ministry of Science, Innovation and Universities, MCIN/AEI/10.13039/501100011033 projects:
      • EDHIA PID2022-136522OB-C22 (also supported by FEDER, UE).
      • TRUST-MED PID2025-174880OB-I00 (also supported by FEDER, UE).
      • DeepR3 TED2021-130295B-C31 (also supported by the European Union NextGenerationEU/PRTR).
    • Basque Government:
      • Research group funding IT1570-22.
    • HiTZ Center.
    • Predoctoral grants: FPU23/03347 (I. De La Iglesia, Spanish Ministry of Science, Innovation and Universities), PRE_2024_1_0224 (A. G. Domingo-Aldama) and PRE_2025_1_0177 (M. Urruela), Basque Government.

Model Details

Training settings for Marmoka-es.
Param. no. ~8B
Architecture Llama 3.1 (decoder-only)
Base model Llama-3.1-8B-Instruct
Adaptation DAPT on medical corpora + general-domain instructions
Medical corpus ~908M (≈1B) words (Spanish)
General instructions 310K (Spanish, Magpie)
Marmoka-es training 3 models trained with different hyperparameters on the same corpus, merged
Training framework Axolotl
Merging tool MergeKit
Model selection Casimedicos validation set (EN & ES)

Training Data

Marmoka-es was trained with the hybrid strategy of Sainz et al. (2025), which combines continual domain-adaptive pretraining on medical corpora with general-domain instructions to avoid catastrophic forgetting. Unlike traditional DAPT setups, an instruction-tuned model (Llama-3.1-8B-Instruct), rather than a base model, was used as the starting point. Note that this setup includes medical corpora but no dedicated medical instructions; the amount of medical content in the Magpie instructions is minimal. The medical corpora are an updated version of those gathered in De la Iglesia et al. (2025), and were processed and deduplicated using standard approaches (deduplication code).

Data sources and sizes used to train Marmoka-es.
Type Source Size
Medical corpora (Spanish) WikiMed 16.3M words
PubMed 5.5M words
Medical Crawler 850.7M words
MeSpEn 6.9M words
SciELO 29M words
General-domain instructions Magpie (Llama-3.1-70B-Instruct) 310K instructions

Limitation and Bias

Marmoka-es is intended for research purposes. Its evaluation is based on short-form multiple-choice question answering with automated accuracy metrics and without expert clinical adjudication, so the validity of the model's underlying reasoning has not been verified, and these benchmarks may not fully capture genuine medical expertise. It is recommended to validate and test the model for specific applications before using it.

Instruction following and strict output formatting remain a limitation, especially in multi-step tasks such as summarization, where formatting failures were observed. Outputs should be checked whenever a strict format is required.

Biases

  • Data Collection Bias: The training data for Marmoka-es was sourced from the web and open datasets, and may not adequately capture the full range of linguistic, demographic, or clinical variability found in real-world medical settings. This may introduce selection bias, and the model might reflect and propagate these biases in its outputs.
  • Demographic and Linguistic Bias: Since the training data may not equally represent all demographic groups or linguistic nuances, the model may perform unequally across populations, potentially reinforcing existing health disparities and societal biases.
  • Unexamined Ethical Considerations: No further assessments have been conducted to determine whether the training data contains biases or personally identifiable information, nor have the anonymization measures been independently verified. All data sources were previously processed, publicly released, and made available with stated limitations, including steps taken to ensure data quality and privacy compliance.

Disclaimer

Marmoka-es is intended for research purposes and is not approved for clinical use. It must not be used as a medical device or relied upon for diagnostic or therapeutic decision-making, and its outputs should not be interpreted as medical advice. Using the model without rigorous validation poses significant risks, including misinformation, misclassification, and potential harm.

When employed in actual clinical scenarios, all outputs must be independently reviewed and validated by qualified healthcare professionals. We do not take any liability for the use of this model.

Citing information

@article{domingo2026adapt,
  title={To Adapt or not to Adapt, Rethinking the Value of Medical Knowledge-Aware Large Language Models},
  author={Domingo-Aldama, Ane G and De La Iglesia, Iker and Urruela, Maitane and Atutxa, Aitziber and Barrena, Ander},
  journal={arXiv preprint arXiv:2604.06854},
  year={2026}
}
Downloads last month
495
Safetensors
Model size
8B params
Tensor type
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for HiTZ/Marmoka-es-Llama-3.1-8B-Instruct

Finetuned
(3264)
this model
Quantizations
1 model

Dataset used to train HiTZ/Marmoka-es-Llama-3.1-8B-Instruct

Collection including HiTZ/Marmoka-es-Llama-3.1-8B-Instruct

Paper for HiTZ/Marmoka-es-Llama-3.1-8B-Instruct