Instructions to use BeaverAI/Fallen-Mistral-Small-3.1-24B-v1e-GGUF with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use BeaverAI/Fallen-Mistral-Small-3.1-24B-v1e-GGUF with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf BeaverAI/Fallen-Mistral-Small-3.1-24B-v1e-GGUF:Q4_K_M # Run inference directly in the terminal: llama cli -hf BeaverAI/Fallen-Mistral-Small-3.1-24B-v1e-GGUF:Q4_K_M
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf BeaverAI/Fallen-Mistral-Small-3.1-24B-v1e-GGUF:Q4_K_M # Run inference directly in the terminal: llama cli -hf BeaverAI/Fallen-Mistral-Small-3.1-24B-v1e-GGUF:Q4_K_M
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf BeaverAI/Fallen-Mistral-Small-3.1-24B-v1e-GGUF:Q4_K_M # Run inference directly in the terminal: ./llama-cli -hf BeaverAI/Fallen-Mistral-Small-3.1-24B-v1e-GGUF:Q4_K_M
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf BeaverAI/Fallen-Mistral-Small-3.1-24B-v1e-GGUF:Q4_K_M # Run inference directly in the terminal: ./build/bin/llama-cli -hf BeaverAI/Fallen-Mistral-Small-3.1-24B-v1e-GGUF:Q4_K_M
Use Docker
docker model run hf.co/BeaverAI/Fallen-Mistral-Small-3.1-24B-v1e-GGUF:Q4_K_M
- LM Studio
- Jan
- Ollama
How to use BeaverAI/Fallen-Mistral-Small-3.1-24B-v1e-GGUF with Ollama:
ollama run hf.co/BeaverAI/Fallen-Mistral-Small-3.1-24B-v1e-GGUF:Q4_K_M
- Unsloth Desktop
- Docker Model Runner
How to use BeaverAI/Fallen-Mistral-Small-3.1-24B-v1e-GGUF with Docker Model Runner:
docker model run hf.co/BeaverAI/Fallen-Mistral-Small-3.1-24B-v1e-GGUF:Q4_K_M
- Lemonade
How to use BeaverAI/Fallen-Mistral-Small-3.1-24B-v1e-GGUF with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull BeaverAI/Fallen-Mistral-Small-3.1-24B-v1e-GGUF:Q4_K_M
Run and chat with the model
lemonade run user.Fallen-Mistral-Small-3.1-24B-v1e-GGUF-Q4_K_M
List all available models
lemonade list
- Atomic Chat
curly apostrophes/feedback
What I personally dislike is its use of slanted/curly apostrophes which I cannot seem to stop.
I use sillytavern so using regex I replace them for output and even then the model insists on using them. I.e. can’t instead of can't. I tried multiple instructions (I'm really bad with that) and, as I mentioned above, automatically edit the output so that all slanted/curly apostrophes are replaced with straight ones '. For me it's annoying as I have "Trim Incomplete Sentences" turned on, trimming to the slanted/curly apostrophes like "BLA BLA BLA... Oh, I didn’ Not saying that it's your fault and up to you to fix this issue (clearly this option in sillytavern should let the users choose what and what not to trim to, I always hated they didn't give you the option and just hard coded it in) . I'm just saying that this is an issue like I had with Gemma 3 overly using curly quotes which was also uncontrollable and of course a nightmare for markdown in sillytavern.
That's all. I really like this model.