Instructions to use VladHong/Qwen3.5-9B-Ornimo-SLERP-GGUF with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use VladHong/Qwen3.5-9B-Ornimo-SLERP-GGUF with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf VladHong/Qwen3.5-9B-Ornimo-SLERP-GGUF:Q4_K_M # Run inference directly in the terminal: llama cli -hf VladHong/Qwen3.5-9B-Ornimo-SLERP-GGUF:Q4_K_M
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf VladHong/Qwen3.5-9B-Ornimo-SLERP-GGUF:Q4_K_M # Run inference directly in the terminal: llama cli -hf VladHong/Qwen3.5-9B-Ornimo-SLERP-GGUF:Q4_K_M
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf VladHong/Qwen3.5-9B-Ornimo-SLERP-GGUF:Q4_K_M # Run inference directly in the terminal: ./llama-cli -hf VladHong/Qwen3.5-9B-Ornimo-SLERP-GGUF:Q4_K_M
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf VladHong/Qwen3.5-9B-Ornimo-SLERP-GGUF:Q4_K_M # Run inference directly in the terminal: ./build/bin/llama-cli -hf VladHong/Qwen3.5-9B-Ornimo-SLERP-GGUF:Q4_K_M
Use Docker
docker model run hf.co/VladHong/Qwen3.5-9B-Ornimo-SLERP-GGUF:Q4_K_M
- LM Studio
- Jan
- vLLM
How to use VladHong/Qwen3.5-9B-Ornimo-SLERP-GGUF with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "VladHong/Qwen3.5-9B-Ornimo-SLERP-GGUF" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "VladHong/Qwen3.5-9B-Ornimo-SLERP-GGUF", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/VladHong/Qwen3.5-9B-Ornimo-SLERP-GGUF:Q4_K_M
- Ollama
How to use VladHong/Qwen3.5-9B-Ornimo-SLERP-GGUF with Ollama:
ollama run hf.co/VladHong/Qwen3.5-9B-Ornimo-SLERP-GGUF:Q4_K_M
- Unsloth Desktop
- Pi
How to use VladHong/Qwen3.5-9B-Ornimo-SLERP-GGUF with Pi:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf VladHong/Qwen3.5-9B-Ornimo-SLERP-GGUF:Q4_K_M
Configure the model in Pi
# Install Pi: npm install -g @earendil-works/pi-coding-agent # Add to ~/.pi/agent/models.json: { "providers": { "llama-cpp": { "baseUrl": "http://localhost:8080/v1", "api": "openai-completions", "apiKey": "none", "models": [ { "id": "VladHong/Qwen3.5-9B-Ornimo-SLERP-GGUF:Q4_K_M" } ] } } }Run Pi
# Start Pi in your project directory: pi
- Docker Model Runner
How to use VladHong/Qwen3.5-9B-Ornimo-SLERP-GGUF with Docker Model Runner:
docker model run hf.co/VladHong/Qwen3.5-9B-Ornimo-SLERP-GGUF:Q4_K_M
- Lemonade
How to use VladHong/Qwen3.5-9B-Ornimo-SLERP-GGUF with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull VladHong/Qwen3.5-9B-Ornimo-SLERP-GGUF:Q4_K_M
Run and chat with the model
lemonade run user.Qwen3.5-9B-Ornimo-SLERP-GGUF-Q4_K_M
List all available models
lemonade list
- Hermes Agent
How to use VladHong/Qwen3.5-9B-Ornimo-SLERP-GGUF with Hermes Agent:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf VladHong/Qwen3.5-9B-Ornimo-SLERP-GGUF:Q4_K_M
Configure Hermes
# Install Hermes: curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash hermes setup # Point Hermes at the local server: hermes config set model.provider custom hermes config set model.base_url http://127.0.0.1:8080/v1 hermes config set model.default VladHong/Qwen3.5-9B-Ornimo-SLERP-GGUF:Q4_K_M
Run Hermes
hermes
- Atomic Chat
- OpenClaw
How to use VladHong/Qwen3.5-9B-Ornimo-SLERP-GGUF with OpenClaw:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf VladHong/Qwen3.5-9B-Ornimo-SLERP-GGUF:Q4_K_M
Configure OpenClaw
# Install OpenClaw: npm install -g openclaw@latest # Register the local server and set it as the default model: openclaw onboard --non-interactive --mode local \ --auth-choice custom-api-key \ --custom-base-url http://127.0.0.1:8080/v1 \ --custom-model-id "VladHong/Qwen3.5-9B-Ornimo-SLERP-GGUF:Q4_K_M" \ --custom-provider-id llama-cpp \ --custom-compatibility openai \ --custom-text-input \ --accept-risk \ --skip-health
Run OpenClaw
openclaw agent --local --agent main --message "Hello from Hugging Face"
Qwen3.5-9B-Ornimo-SLERP (GGUF)
GGUF quantizations of a Qwen3.5-9B hybrid-architecture SLERP merge: the self-improvement RL tower of Ornith-1.5-9B blended into the agentic-SFT shell of MiMo-V2.6-Distill-Qwen-9B. An ornith (a bird) and a MiMo walk into a merge โ Ornimo.
Files
| file | size | notes |
|---|---|---|
Qwen3.5-9B-Ornimo-SLERP-Q4_K_M.gguf |
5,629,105,024 B (5.63 GB) | 4-bit K-quant; recommended for everyday use |
Qwen3.5-9B-Ornimo-SLERP-F16.gguf |
17,920,693,120 B (17.92 GB) | unquantized F16 text weights (converted from the bf16 checkpoint) |
Qwen3.5-9B-Ornimo-SLERP-mmproj-f16.gguf |
921,705,024 B (922 MB) | vision tower + projector (MiMo's, byte-exact), F16 |
That is the complete list. This repo contains GGUF weights only โ no
safetensors (those live in the sibling checkpoint repo) and no imatrix
(importance matrix): the Q4_K_M was produced by a plain, calibration-free
llama-quantize โฆ Q4_K_M run with no calibration dataset. If you prefer an
imatrix-calibrated quant, generate an imatrix from the F16 GGUF and
requantize โ the F16 is provided exactly for that.
What is this, really?
Nobody trained anything here. Both parents are fine-tunes of the same
Qwen3.5-9B backbone (32 hybrid layers โ 24 linear-attention + 8
full-attention, hidden 4096, vocab 248320, 262K context) with the
identical tokenizer (248,044 tokens, same ordering), so the text towers
align 1:1. The merge is a whole-tensor SLERP (t=0.5) of all 427 text
tensors, with MiMo's vision tower, projector, tokenizer, chat template and
configs kept byte-exact. Ornith's optional multi-token-prediction (mtp.*)
block is not part of the merge โ it is trained against Ornith's own
hidden space and would be miscalibrated on a blended one.
Usage
Both GGUFs use the Qwen3.5 hybrid arch (qwen35 in llama.cpp naming) and
run on llama.cpp b393+.
Text only:
llama-cli -m Qwen3.5-9B-Ornimo-SLERP-Q4_K_M.gguf \
-c 131072 -t 16 -cnv --repeat-penalty 1.1
Vision (image understanding) with the projector:
llama-mtmd-cli -m Qwen3.5-9B-Ornimo-SLERP-Q4_K_M.gguf \
--mmproj Qwen3.5-9B-Ornimo-SLERP-mmproj-f16.gguf \
--image <image.png> -p "What is in this image?" -c 131072 -t 16
-ccontext: the model supports up to 262144 natively; only the 8 full-attention layers cache KV, the 24 linear-attention layers keep fixed-size state.- Thinking mode follows MiMo's chat template
(
chat_template_kwargs: {"enable_thinking": true}in OpenAI-compatible servers;/no_thinkin the prompt for a direct answer). --repeat-penalty 1.1is recommended for very long reasoning chains.
Why this merge was safe (deviation analysis)
Before merging, both parents were compared tensor-by-tensor (760 shared tensors):
| check | result |
|---|---|
| tokenizer vocab ordering | identical (248,044 tokens, same order) |
| median tensor cosine (text) | 0.993โ1.000 per category |
| min tensor cosine | 0.986 (layers.22.linear_attn.in_proj_b) |
| tensors below 0.98 cosine | 0 |
| gating tally (lo=0.90, hi=0.995) | 293 full-SLERP / 134 gated / 0 anchored |
| lm_head row cosine | min 0.930, no row below 0.8 โ no clashing token families |
| vision towers | cos 1.000 median โ near-identical, copied byte-exact |
| NaNs | 0 |
Because tokenizers are identical and every output row agrees, the
unembedding (lm_head) is merged along with everything else โ no
anchor-side lm_head exception.
Verified behavior (CPU smoke checks, not a benchmark suite)
- structural verify: all 4 output shards recomputed against an independent SLERP implementation (1-ulp tolerance) โ passed
- Q4_K_M via llama.cpp: 17 ร 24 = 408 plus a clean one-line Python
function, no repetition loops,
13 tok/s generation on 24 CPU threads (100 t/s prompt processing)
Everything beyond these checks is uncharted; expect surprises and hallucinations in unknown proportions.
Credits
- XiaomiMiMo/MiMo-V2.6-Distill-Qwen-9B โ the anchor: shell, vision tower, projector, tokenizer, chat template (MIT).
- ornith-ai/Ornith-1.5-9B โ the RL-self-improvement parent whose text tower is blended in (MIT).
- Qwen3.5-9B โ the shared backbone both parents fine-tune (Apache-2.0, Copyright 2026 Alibaba Cloud).
- Downloads last month
- 641
4-bit