opencoti-llamafile

Self-contained single-file inference engine from the opencoti project β€” a Mozilla-Ocho llamafile base carrying the opencoti patch series: advanced KV residency/quantization (PolyKV, KVarN β€” see docs/features/kvarn.md), the KV rolling window (a KV cache larger than VRAM), elastic multi-session serving (elastic slots, an enforced KV admission gate with fast 429 + Retry-After, PolyKV shared-prefix pools with a REST control plane), DCA long-context extension, MTP speculative decode, sparse attention, RYS layer duplication, CUDA and Vulkan GPU backends, and a runtime introspection/control API.

Start with USAGE.md β€” it explains exactly how this engine diverges from upstream llamafile, every added feature, its flags, defaults, limitations, and which features compose.

This repo hosts the packaged release artifacts (they exceed GitHub's 2 GiB release-asset cap) plus the full patch series under patches/ and the project documentation under docs/. Release notes and per-release SHA256SUMS live under releases/ and are mirrored on the corresponding GitHub release at mann1x/opencoti.

The repo root always holds the LATEST cut's single-file executables only. Superseded cuts are archived whole β€” binaries, manifests, their dso/ set (c7 and older), release notes and SHA256SUMS β€” under releases/archived/ in one self-contained directory per release tag (e.g. releases/archived/llamafile-v0.10.3+opencoti.c4/).

Supported / target model families

This engine is not model-agnostic in what it optimizes. Upstream llama.cpp GGUF support is unchanged (any GGUF that loads upstream loads here), but the opencoti feature set is built, tuned, and validated on two families:

Family Role Models What's tuned for them
Gemma-4 PRIMARY target 26B-A4B-128e MoE ("A4B"), dense 12B / 31B, elastic E-series (E2B/E4B) head_dim 256/512 kernels, iSWA dual-cache (rolling-KV, SharedKVPool, DCA wiring), MTP via gemma4-assistant drafters, fused MoE
Qwen SECONDARY / verification Qwen3.5 / 3.6 dense + MoE + hybrid (gated-delta-net), Qwen2.5-14B-1M head_dim 128 kernels, NextN self-spec MTP (fused multi-step draft), DCA long-context on 1M-class models

Other architectures run with upstream behavior; opencoti features either fall back safely or are unvalidated on them β€” see the model-families section at the top of USAGE.md before relying on an opencoti feature elsewhere.

Artifacts

The main release is ONE file. Download the single executable for your platform; every GPU library it needs is inside.

Artifact (repository root) Runs on GPU inside
opencoti-llamafile-<ver>-<tag>-universal.llamafile.exe (~1.7 GB) Windows x86_64 and Linux x86_64 β€” one file for both CUDA 13 .dll + .so (sm_75/80/86/89/90/120) and Vulkan .dll + .so (AMD / Intel / NVIDIA); each OS takes its own
opencoti-llamafile-<ver>-<tag>-win-x86_64-gpu.llamafile.exe (~0.9 GB) Windows x86_64 CUDA 13 .dll + Vulkan .dll
opencoti-llamafile-<ver>-<tag>-x86_64.llamafile (~1.2 GB) Linux x86_64 CUDA 13 .so, the CUDA 12 legacy .so (sm_52/61/70: Maxwell, Pascal, Volta β€” chosen when every GPU is older than sm_75) and Vulkan .so
opencoti-llamafile-<ver>-<tag>-aarch64.llamafile (~0.47 GB) Linux aarch64 CUDA 13 sbsa .so (sm_110 DGX Spark GB10 / Jetson Thor, sm_121)

Every embedded library is byte-identical to the separate file of the same release (below); the publisher refuses a release whose single files embed anything else.

The split form (for integrations that pin each part by sha256 β€” the pin files under pin/, format in PIN_FORMAT.md): the bare engine components/engine/<build>/opencoti-llamafile-<ver>-<tag>-bare.llamafile (~135 MB, the same bytes as …-win-x86_64.llamafile.exe there; CPU only by itself) and one directory per library under components/ (cuda, cuda12, sbsa, vulkan, macos (Metal + the signed loader), media (codec / speech sidecars)). pin/index.txt names the current file of each.

The APE runs natively on x86_64 and aarch64, Linux, Windows, macOS and BSD; the variants differ only in which GPU libraries are embedded. Windows refuses executables over 4 GiB β€” the libraries are built compressed (--compress-mode=size) so the .exe files stay well under it. With no usable GPU library the engine runs on the CPU.

Usage

chmod +x opencoti-llamafile-*.llamafile
# server mode
sh ./opencoti-llamafile-*.llamafile --server --port 8080 \
    -m <model>.gguf -ngl 99 --flash-attn on
# without --server it starts the interactive chat CLI

On Windows run the .exe directly (opencoti-llamafile-<ver>-<tag>-universal.llamafile.exe --server …); only the NVIDIA (or AMD / Intel) driver is needed β€” no CUDA Toolkit, no MSVC.

Where the GPU library goes (c10 r2+). A GPU library cannot be loaded from inside the file, so on the first GPU run it is extracted next to the executable: <exe folder>/.llamafile/v/opencoti-<ver>-<tag>/p/<key>/ (e.g. 0.7 GB for CUDA on Windows). Only when that folder is not writable (e.g. C:\Program Files) does it go to the home folder (~/.llamafile/… / %USERPROFILE%\.llamafile\…, unless HOME is set), and the engine prints one line saying so. On Windows that profile fallback needs c10 r3 or later: r2 put the library in the current directory there. The key is the library's own checksum, so a later run of the same release reuses it and two releases never overwrite each other. The boot log names the file it loaded (cuda: loaded … (bundled, beside the executable, … bytes)). Delete the .llamafile folder to reclaim the space. (c7 – c10 r1 always extracted to the home folder.) --version prints the opencoti version string; GET /props exposes an opencoti introspection section at runtime (see USAGE.md Β§introspection).

The bare engine + a separate library

Put the library file, under its published name, in the same folder as the engine β€” the engine looks there first and loads it as is (no rename, no copy into your profile):

opencoti-llamafile-<ver>-<tag>-win-x86_64.llamafile.exe
ggml-cuda-win-x86_64.dll        # components/cuda/<build>/  (or ggml-vulkan-win-x86_64.dll for --gpu vulkan)

On Linux: ggml-cuda-x86_64.so / ggml-cuda-cu12-x86_64.so / ggml-vulkan-x86_64.so / ggml-cuda-sbsa-aarch64.so beside the APE. The boot line then reads (executable directory, … bytes). Check every file against the sha256 in its pin/<component>.txt.

Verification

Each single-file executable ships with a MANIFEST.json recording its sha256, the version string, the git commit, the full patch list and the sha256 of every embedded GPU library; releases/<full-tag>/SHA256SUMS lists every file of the release. An embedded library is the same file as the one in components/ β€” compare it with the sha256 in pin/<component>.txt without running anything:

unzip -p opencoti-llamafile-*-universal.llamafile.exe ggml-cuda.dll | sha256sum   # == pin/cuda.txt, win-x86_64 row

License

Apache-2.0 (llamafile) + MIT (llama.cpp and bundled projects). opencoti patches and packaging are MIT.

PolyKV multi-session stress benchmark (package_courier)

An agentic multi-session benchmark: LangGraph courier agents (tool-calling, 10-step delivery chains, 100 packages) run against ONE engine at --parallel 32, ctx 163840, q8_0 KV, with a tps-floor scheduler that admits new sessions only while every live session keeps β‰₯ floor tok/s (PolyKV /polykv/tps SSE telemetry drives admission). Every run below was measured on a single NVIDIA RTX 6000 Blackwell (96 GB) β€” one GPU, one engine process, no tensor/pipeline split; the whole model and all 32 slots' KV are resident, so the numbers are pure scheduler/serving behaviour rather than an offload artifact. Each worker episode is a compiled langgraph.StateGraph (capacity_gate β†’ agent ⇄ tools, with the capacity-gate node reading the engine's /polykv/pools/{id}/capacity), so the harness doubles as the reference LangGraph integration for this engine. The full harness is published in this repo at bench/package_courier/ (work server + LangGraph agent app + launcher + HTML report generator β€” see its README and the usage guide at docs/features/package_courier_harness.md); reports under bench/reports/ (per-run HTML + raw stats).

The opencoti_langgraph Python package the harness builds on β€” OpencotiChatModel (LangChain BaseChatModel with PolyKV pool/session attach), PolykvClient/AsyncPolykvClient, pool tools, and the capacity-gate / compaction LangGraph nodes β€” is published at integrations/langgraph/ (pip install ./integrations/langgraph, needs langchain-core>=1.4.9

  • langgraph>=1.2.9; see its README).

fan-out is the time-weighted median of concurrently running workers over the run β€” the number of agents actually in flight, not the scheduler's target and not the engine's attached-session count (which is cumulative and saturates at slots). Two token rates, and the distinction matters:

  • gen tok/s β€” generated tokens Γ· wall. This is what the tps-floor scheduler gates on, so it is what sets fan-out. It is not a model-speed comparison: at 36:1–74:1 prompt-to-completion it largely reflects how terse a model is.
  • work tok/s β€” (prompt tokens actually processed, i.e. net of prompt-cache hits) + generated, Γ· wall. This is the real work rate. Note the cache-hit column: each turn resends the whole conversation and the engine serves 25–76 % of it from KV cache, so counting raw prompt length would inflate this several-fold.
model (MTP) floor score wall delivery mean fan-out (med / peak) gen tok/s work tok/s cache hit
Qwen3.6-27B-Omnimerge-v4 Q6_K (NextN self-spec, dense) 15 99/100 2171 s 146 s 7 / 8 79.9 813 74 %
β€³ 5 99/100 2509 s 467 s 20 / 22 70.4 800 70 %
β€³ 1 98/100 2803 s 818 s 32 / 32 (cap) 62.9 794 67 %
Gemma-4 26B-A4B-it Q4_K_M (assistant-MTP, MoE) 15 100/100 864 s 51 s 6 / 8 64.5 2912 40 %
β€³ 5 99/100 1090 s 145 s 14 / 16 50.8 2682 30 %
β€³ 1 95/100 1340 s 346 s 32 / 32 (cap) 40.4 2296 25 %
β€³ (cap 64) 25 100/100 773 s 31 s 4 / 6 72.0 3010 45 %
β€³ (cap 64) 20 100/100 823 s 38 s 5 / 6 67.6 2862 44 %
Gemma-4 26B-A4B-it F16 (assistant-MTP, cap 64) 25 100/100 1181 s 35 s 3 / 5 47.1 1502 58 %
β€³ 20 100/100 1268 s 44 s 4 / 5 43.9 1548 54 %
β€³ 15 100/100 1369 s 59 s 5 / 6 40.7 1571 49 %
β€³ 5 100/100 1712 s 174 s 10 / 14 32.5 1567 36 %
β€³ 1 100/100 2542 s 1282 s 64 / 64 (cap) 21.9 735 56 %
Qwen3.6-35B-A3B UD-Q6_K (NextN self-spec, MoE) 25 100/100 912 s 55 s 6 / 7 110.1 1609 75 %
β€³ 20 100/100 890 s 64 s 7 / 9 112.5 1600 76 %
β€³ 15 100/100 914 s 84 s 10 / 12 109.9 1609 75 %
β€³ 5 100/100 1041 s 225 s 24 / 27 96.5 1567 72 %
β€³ 1 100/100 1574 s 786 s 64 / 64 (cap) 63.7 1264 65 %

Reading these numbers

Work rate is set by active parameters. The A4B pair runs at 2912 work tok/s against 813 for the dense 27B at floor 15 β€” 3.6Γ—, which is what a 4B-active MoE against a 27B-dense should look like. Weight precision costs on top of that: the same A4B at F16 does 1571, i.e. 1.85Γ— behind Q4_K_M, a bandwidth-bound MoE reading 4Γ— the bytes per active expert.

Terseness is a separate axis from speed, and gen tok/s conflates them. The dense 27B posts the higher generation rate at floor 15 (79.9 vs 64.5) while doing 3.6Γ— less work per second β€” because it is a reasoning model that emits <think> blocks and produces 173 k tokens to the A4B's 55.6 k for the same 100 deliveries. Wall time follows the product of both axes: A4B finishes 2.5Γ— sooner (864 s vs 2171 s) by being both faster and terser. If you take one thing from this table: do not rank engines by generation tok/s on an agentic workload β€” it rewards verbosity.

The cache-hit column is why rows are not freely comparable. Every turn resends the full conversation; the engine matches the common prefix in KV and only prefills the delta. A model that needs more turns per package re-sends a longer already-cached prefix, so its hit rate climbs and its measured work rate falls β€” not because the engine got slower, but because there was less left to do. That is why the 35B-A3B's 1609 cannot be read as "slower than the A4B's 2912": it ran at 75 % cache hit against 40 %. Compare rows freely only where the cache column is close (e.g. A4B Q4_K_M vs F16, 40 % vs 49 %).

None of this is a clean engine benchmark β€” it is a scheduler-and-serving measurement over a real agentic workload. For isolated decode/prefill throughput use docs/evaluations/mtp.md.

How fan-out is set. workers β‰ˆ gen tok/s Γ· floor β€” the scheduler admits until the generation rate the engine can deliver is exactly divided into per-session shares of floor tok/s. That rate is not a constant: it rises as the floor rises (A4B-Q4: 40.4 β†’ 72.0 tok/s from floor 1 to 25) at identical total work, so the engine does the same job faster with fewer concurrent sequences β€” consistent with per-token batch interference and KV pressure, though we have not isolated which dominates. Practically: a higher floor buys both a faster per-session experience and a shorter wall time. You only pay in parallelism.

The A4B + assistant-MTP pair finishes ~2.5Γ— faster end-to-end than the dense 27B at floor 15 (864 s vs 2171 s) while generating 3.1Γ— fewer tokens for the same 100 deliveries (55.6 k vs 173 k) β€” a terser agent on a faster engine.

All losses across all runs are agent-side (turn-budget wander on hard packages) β€” zero tool errors, zero LLM errors, engines clean throughout. The ab-pooled vs ab-naive pair isolates the SharedKVPool win at identical workload.

The A4B F16 ladder answers "what does weight precision cost under agentic load": same engine, same workload, F16 weights are ~1.55Γ— slower end-to-end than Q4_K_M at every floor (1.9Γ— at floor 1, where contention compounds it) β€” a bandwidth-bound MoE. At comfortable floors quality is saturated at Q4 (both 100/100), but at max fan-out precision shows: F16 floor 1 finished zero-loss at the full 64-worker cap grinding at ~1.1 tok/s median, while Q4_K_M at floor 1 loses 3-5 first-wave packages to turn-budget wander (95/100 original; a controlled re-run on the current StateGraph agent, same seed/cap, scored 97/100 with the identical loss signature β€” confirming the loss mode is model-precision-driven under cold-start contention, not a harness artifact).

The 35B-A3B ladder (worker cap raised to 64, floors extended to 20/25) is the cleanest full ladder: 100/100 at every floor including max fan-out, with the scheduler holding median per-turn tps on the floor (5β†’5.6, 15β†’17, 20β†’23, 25β†’28) as the fleet shrinks 64β†’6. It also locates the fan-out knee: wall time is flat (~890-1040 s) from floor 5 up, but floor 1 costs 1574 s β€” 64 workers oversubscribing 32 slots just queues (per-session tps 2.8, deliveries waiting 786 s mean). Sweet spot for this engine: floor 20-25 β€” same wall as floor 15 with the best per-package latency, at 6-7 concurrent workers each getting 23-28 tok/s. Note how little concurrency it takes to saturate: throughput is already at its maximum by floor 15, and everything above ~10 workers on this engine buys queueing, not work.


Vodafone

Computing resources kindly provided by Vodafone

Downloads last month
591
Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support