Image-Text-to-Text
Transformers
Safetensors
glm5_next
glm
exl3
tr3
vllm
sm120
nvfp4
dflash2
multimodal
shapleymcg
conversational
Eval Results (legacy)
4-bit precision
Instructions to use brandonmusic/GLM-5.3-Flash-tr3-4bpw with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use brandonmusic/GLM-5.3-Flash-tr3-4bpw with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("image-text-to-text", model="brandonmusic/GLM-5.3-Flash-tr3-4bpw") messages = [ { "role": "user", "content": [ {"type": "image", "url": "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/p-blog/candy.JPG"}, {"type": "text", "text": "What animal is on the candy?"} ] }, ] pipe(text=messages)# pip install -U transformers accelerate # Load model directly from transformers import AutoProcessor, AutoModelForMultimodalLM processor = AutoProcessor.from_pretrained("brandonmusic/GLM-5.3-Flash-tr3-4bpw") model = AutoModelForMultimodalLM.from_pretrained("brandonmusic/GLM-5.3-Flash-tr3-4bpw", device_map="auto") messages = [ { "role": "user", "content": [ {"type": "image", "url": "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/p-blog/candy.JPG"}, {"type": "text", "text": "What animal is on the candy?"} ] }, ] inputs = processor.apply_chat_template( messages, add_generation_prompt=True, tokenize=True, return_dict=True, return_tensors="pt", ).to(model.device) outputs = model.generate(**inputs, max_new_tokens=256) print(processor.decode(outputs[0][inputs["input_ids"].shape[-1]:])) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use brandonmusic/GLM-5.3-Flash-tr3-4bpw with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "brandonmusic/GLM-5.3-Flash-tr3-4bpw" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "brandonmusic/GLM-5.3-Flash-tr3-4bpw", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }'Use Docker
docker model run hf.co/brandonmusic/GLM-5.3-Flash-tr3-4bpw
- SGLang
How to use brandonmusic/GLM-5.3-Flash-tr3-4bpw with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "brandonmusic/GLM-5.3-Flash-tr3-4bpw" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "brandonmusic/GLM-5.3-Flash-tr3-4bpw", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "brandonmusic/GLM-5.3-Flash-tr3-4bpw" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "brandonmusic/GLM-5.3-Flash-tr3-4bpw", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }' - Docker Model Runner
How to use brandonmusic/GLM-5.3-Flash-tr3-4bpw with Docker Model Runner:
docker model run hf.co/brandonmusic/GLM-5.3-Flash-tr3-4bpw
Download runtime-results/v84/benchmarks/llm-decode-c1-prefill32k64k-600w.tui.log from brandonmusic/GLM-5.3-Flash-tr3-4bpw: direct link, hf CLI and curl.
- Browser
- Download file 12.3 kB
-
https://huggingface.co/brandonmusic/GLM-5.3-Flash-tr3-4bpw/resolve/main/runtime-results/v84/benchmarks/llm-decode-c1-prefill32k64k-600w.tui.log
- Command line
-
hf download hf://brandonmusic/GLM-5.3-Flash-tr3-4bpw/runtime-results/v84/benchmarks/llm-decode-c1-prefill32k64k-600w.tui.log
-
curl -L -o llm-decode-c1-prefill32k64k-600w.tui.log https://huggingface.co/brandonmusic/GLM-5.3-Flash-tr3-4bpw/resolve/main/runtime-results/v84/benchmarks/llm-decode-c1-prefill32k64k-600w.tui.log
12.3 kB
| Script started on 2026-08-28 06:40:58-04:00 [COMMAND="python3 llm_decode_bench.py --host 127.0.0.1 --port 5001 --model GLM-5.3-Flash-EXL3-4bpw --concurrency 1 --contexts 0,32k,64k --duration 10 --max-tokens 1024 --decode-warmup-seconds 3 --standalone-prefill --prefill-contexts 32k,64k --prefill-duration 10 --prefill-metric auto --token-targeting estimate --kv-budget 129473 --display-mode plain --no-resume --output /home/brandonmusic/KLC_SANDBOXES/glm53-exl3-k4-sm120/results/dflash2-v84-600w-quick/llm-decode-c1-prefill32k64k.json" <not executed on terminal>] | |
| ╭──────────────────────────── NVIDIA P2P Override ─────────────────────────────╮ | |
| │ Effective: yes │ | |
| │ Configured file: yes (/etc/modprobe.d/nvidia-p2p-override.conf) │ | |
| │ Runtime: ForceP2P=0x11; RMForceP2PType=1; RMPcieP2PType=2; │ | |
| │ GrdmaPciTopoCheckOverride=1; EnableResizableBar=1; DmaRemapPeerMmio=1 │ | |
| ╰──────────────────────────────────────────────────────────────────────────────╯ | |
| ╭─────────────────────────────── Configuration ────────────────────────────────╮ | |
| │ LLM Inference Benchmark │ | |
| │ Model: GLM-5.3-Flash-EXL3-4bpw @ 127.0.0.1:5001 │ | |
| │ Decode concurrency: [1] │ | |
| │ Decode contexts: ['0', '32k', '64k'] │ | |
| │ Duration: 10.0s per decode test | Max tokens: 1024 │ | |
| │ Pre-decode warmup: C=1 max-runnable context for 3s │ | |
| │ Prefill: standalone cold profile (auto) | Sustained decode: 3 cells │ | |
| ╰──────────────────────────────────────────────────────────────────────────────╯ | |
| KV cache budget (manual): 129,473 tokens | |
| Engine: vLLM 0.1.dev20111+g7f1e92bec.d20260827 Models: | |
| ['GLM-5.3-Flash-EXL3-4bpw'] | |
| Model context length: 98,304 tokens | |
| Prefill tests: standalone cold profile ['32k', '64k'] | |
| Calibrating padding text (run=xxeawouskdhk, up to 64k)... | |
| Token targeting: single-point estimate from 8k (use --token-targeting exact | |
| for /tokenize binary search) | |
| Calibrated: 6.18 chars/token (cached, source=8k) | |
| 32k: 202,404 chars (~32,767 tokens) | |
| 64k: 404,809 chars (~65,535 tokens) | |
| Done. | |
| llm-decode-bench v0.4.29 | |
| Prefill Speed (C=1, client ISL / TTFT) | |
| PCIe rx/tx | |
| Context Tokens TTFT (s) Client tok/s Server tok/s avg N | |
| ────────────────────────────────────────────────────────────────────────────── | |
| 32k 32,323 5.19 6,225 6,277 (2) 19267/19610 2 | |
| 64k 64,515 10.61 6,083 6,130 (1) 25574/23986 1 | |
| Client tok/s = prompt_tokens / TTFT. Integrated scout rows come from the | |
| prefix-cache scout request that decode needs anyway. Server tok/s is optional | |
| Prometheus validation when the engine exports prefill counters and the exact | |
| counter delta is uncontaminated; for vLLM this uses newly computed KV tokens, | |
| not request prompt tokens. | |
| ╭────────────────────────────────── Phase 2 ───────────────────────────────────╮ | |
| │ Sustained Decode │ | |
| │ Steady-state decode throughput after the engine has admitted the requested │ | |
| │ concurrency and passed warmup. Use this as the main tuning/regression signal │ | |
| │ for kernels, NCCL, DCP, MTP, and scheduler changes. │ | |
| ╰──────────────────────────────────────────────────────────────────────────────╯ | |
| Aggregate tok/s + TTFT/ITL | |
| ╭────────────┬─────────────╮ | |
| │ ctx \ conc │ 1 │ | |
| ├────────────┼─────────────┤ | |
| │ 0 │ 151.0 137/6 │ | |
| │ 32k │ 60.1 5k/6 │ | |
| │ 64k │ 143.3 11k/7 │ | |
| ╰────────────┴─────────────╯ | |
| Sustained Decode: aggregate tok/s uses OpenAI stream usage by default | |
| (continuous completion_tokens when the server supports it). Prometheus is kept | |
| as validation/scheduler data. | |
| Aggregate source(s): openai_continuous_usage | |
| Per-Request tok/s | |
| ╭────────────┬───────╮ | |
| │ ctx \ conc │ 1 │ | |
| ├────────────┼───────┤ | |
| │ 0 │ 151.0 │ | |
| │ 32k │ 60.1 │ | |
| │ 64k │ 143.3 │ | |
| ╰────────────┴───────╯ | |
| Client request latency: p50 | |
| / p90 ms | |
| ╭────────────┬─────────────╮ | |
| │ ctx \ conc │ 1 │ | |
| ├────────────┼─────────────┤ | |
| │ 0 │ 6.8k/6.9k │ | |
| │ 32k │ 13.0k/13.0k │ | |
| │ 64k │ 17.4k/17.4k │ | |
| ╰────────────┴─────────────╯ | |
| Aggregate cells show dim detail as TTFT ms / ITL ms for the same ctx/conc | |
| coordinate. ITL is computed from observed generated tokens, including streams | |
| stopped at the measurement boundary; a missing ITL means no stream produced at | |
| least two measured output tokens. Per-request tok/s and request latency are | |
| shown in separate per-cell matrices. Completion/sample counts and full | |
| request-level distributions remain in JSON under request_samples. | |
| Sustained mode: client latency metrics explain request UX variance; aggregate | |
| tok/s remains the primary throughput signal. | |
| ITL=(last_token_time-first_token_time)/(output_tokens-1), user tok/s=1/ITL. | |
| Hardware Summary | |
| ╭───┬─┬───────┬───────────┬───────┬─────────┬─────┬──────┬─────┬───────────────╮ | |
| │ … │ │ mode │ GPU avg/… │ Mem … │ W avg/… │ T … │ CPU… │ VR… │ PCIe rx/tx a… │ | |
| ├───┼─┼───────┼───────────┼───────┼─────────┼─────┼──────┼─────┼───────────────┤ | |
| │ 0 │ │ sust… │ 49/99% │ 28% │ 880/890 │ 69C │ 67C │ 48… │ 3018/2770 │ | |
| │ … │ │ sust… │ 50/100% │ 21% │ 938/10… │ 79C │ 67C │ 48… │ 13663/13205 │ | |
| │ … │ │ sust… │ 50/100% │ 19% │ 980/10… │ 87C │ 67C │ 48… │ 14877/17000 │ | |
| ╰───┴─┴───────┴───────────┴───────┴─────────┴─────┴──────┴─────┴───────────────╯ | |
| ╭─────────────────────── Whole-run GPU Power ────────────────────────╮ | |
| │ avg 830 W | max 1,035 W | limit 1,800 W | over 2m 35s | 66 samples │ | |
| ╰────────────────────────────────────────────────────────────────────╯ | |
| Hardware summary is sampled from nvidia-smi during the measured part of each | |
| cell. Whole-run GPU power is the sampled sum of GPU power draw across the | |
| complete benchmark run, not wall-outlet system power. PCIe rx/tx is MB/s and is | |
| a coarse live diagnostic, not a per-kernel NCCL profiler. | |
| ╭────────────────────────────────── Phase 3 ───────────────────────────────────╮ | |
| │ Burst / E2E Decode │ | |
| │ Not run. Re-run with --run-burst to append a finite client-facing request │ | |
| │ burst after Sustained Decode. This is intentionally disabled by default │ | |
| │ because it adds another full decode matrix. │ | |
| ╰──────────────────────────────────────────────────────────────────────────────╯ | |
| ╭────────────────────────────── Primary Summary ───────────────────────────────╮ | |
| │ Primary matrices repeated last so the important numbers are visible without │ | |
| │ scrolling back through diagnostics. │ | |
| ╰──────────────────────────────────────────────────────────────────────────────╯ | |
| Prefill tok/s Aggregate decode tok/s | |
| ╭─────┬────────┬────────┬───────┬───╮╭────────────┬───────╮ | |
| │ ctx │ tokens │ TTFT s │ tok/s │ N ││ ctx \ conc │ 1 │ | |
| ├─────┼────────┼────────┼───────┼───┤├────────────┼───────┤ | |
| │ 32k │ 32,323 │ 5.19 │ 6,225 │ 2 ││ 0 │ 151.0 │ | |
| │ 64k │ 64,515 │ 10.61 │ 6,083 │ 1 ││ 32k │ 60.1 │ | |
| ╰─────┴────────┴────────┴───────┴───╯│ 64k │ 143.3 │ | |
| ╰────────────┴───────╯ | |
| MTP-normalized decode | |
| steps/s (accept len) | |
| ╭────────────┬─────────────╮ | |
| │ ctx \ conc │ 1 │ | |
| ├────────────┼─────────────┤ | |
| │ 0 │ 51.6 (2.92) │ | |
| │ 32k │ 24.3 (2.47) │ | |
| │ 64k │ 51.5 (2.78) │ | |
| ╰────────────┴─────────────╯ | |
| steps/s = tok/s ÷ accept_len: engine forward passes per second, independent of | |
| MTP acceptance, so runs with different acceptance are directly comparable. | |
| (accept len) = tokens emitted per engine step. | |
| Results saved to | |
| /home/brandonmusic/KLC_SANDBOXES/glm53-exl3-k4-sm120/results/dflash2-v84-600w-qu | |
| ick/llm-decode-c1-prefill32k64k.json | |
| Script done on 2026-08-28 06:43:35-04:00 [COMMAND_EXIT_CODE="0"] | |