GLM-5.3-Flash-tr3-4bpw / runtime-results /v84 /benchmarks /llm-decode-c1-prefill32k64k-600w.tui.log
brandonmusic's picture
Publish v84 language-only runtime profile
3224669 verified
Raw History Blame Contribute Delete
12.3 kB
Script started on 2026-08-28 06:40:58-04:00 [COMMAND="python3 llm_decode_bench.py --host 127.0.0.1 --port 5001 --model GLM-5.3-Flash-EXL3-4bpw --concurrency 1 --contexts 0,32k,64k --duration 10 --max-tokens 1024 --decode-warmup-seconds 3 --standalone-prefill --prefill-contexts 32k,64k --prefill-duration 10 --prefill-metric auto --token-targeting estimate --kv-budget 129473 --display-mode plain --no-resume --output /home/brandonmusic/KLC_SANDBOXES/glm53-exl3-k4-sm120/results/dflash2-v84-600w-quick/llm-decode-c1-prefill32k64k.json" <not executed on terminal>]
╭──────────────────────────── NVIDIA P2P Override ─────────────────────────────╮
│ Effective: yes │
│ Configured file: yes (/etc/modprobe.d/nvidia-p2p-override.conf) │
│ Runtime: ForceP2P=0x11; RMForceP2PType=1; RMPcieP2PType=2; │
│ GrdmaPciTopoCheckOverride=1; EnableResizableBar=1; DmaRemapPeerMmio=1 │
╰──────────────────────────────────────────────────────────────────────────────╯
╭─────────────────────────────── Configuration ────────────────────────────────╮
│ LLM Inference Benchmark │
│ Model: GLM-5.3-Flash-EXL3-4bpw @ 127.0.0.1:5001 │
│ Decode concurrency: [1] │
│ Decode contexts: ['0', '32k', '64k'] │
│ Duration: 10.0s per decode test | Max tokens: 1024 │
│ Pre-decode warmup: C=1 max-runnable context for 3s │
│ Prefill: standalone cold profile (auto) | Sustained decode: 3 cells │
╰──────────────────────────────────────────────────────────────────────────────╯
KV cache budget (manual): 129,473 tokens
Engine: vLLM 0.1.dev20111+g7f1e92bec.d20260827 Models:
['GLM-5.3-Flash-EXL3-4bpw']
Model context length: 98,304 tokens
Prefill tests: standalone cold profile ['32k', '64k']
Calibrating padding text (run=xxeawouskdhk, up to 64k)...
Token targeting: single-point estimate from 8k (use --token-targeting exact
for /tokenize binary search)
Calibrated: 6.18 chars/token (cached, source=8k)
32k: 202,404 chars (~32,767 tokens)
64k: 404,809 chars (~65,535 tokens)
Done.
llm-decode-bench v0.4.29
Prefill Speed (C=1, client ISL / TTFT)
PCIe rx/tx
Context Tokens TTFT (s) Client tok/s Server tok/s avg N
──────────────────────────────────────────────────────────────────────────────
32k 32,323 5.19 6,225 6,277 (2) 19267/19610 2
64k 64,515 10.61 6,083 6,130 (1) 25574/23986 1
Client tok/s = prompt_tokens / TTFT. Integrated scout rows come from the
prefix-cache scout request that decode needs anyway. Server tok/s is optional
Prometheus validation when the engine exports prefill counters and the exact
counter delta is uncontaminated; for vLLM this uses newly computed KV tokens,
not request prompt tokens.
╭────────────────────────────────── Phase 2 ───────────────────────────────────╮
│ Sustained Decode │
│ Steady-state decode throughput after the engine has admitted the requested │
│ concurrency and passed warmup. Use this as the main tuning/regression signal │
│ for kernels, NCCL, DCP, MTP, and scheduler changes. │
╰──────────────────────────────────────────────────────────────────────────────╯
Aggregate tok/s + TTFT/ITL
╭────────────┬─────────────╮
│ ctx \ conc │ 1 │
├────────────┼─────────────┤
│ 0 │ 151.0 137/6 │
│ 32k │ 60.1 5k/6 │
│ 64k │ 143.3 11k/7 │
╰────────────┴─────────────╯
Sustained Decode: aggregate tok/s uses OpenAI stream usage by default
(continuous completion_tokens when the server supports it). Prometheus is kept
as validation/scheduler data.
Aggregate source(s): openai_continuous_usage
Per-Request tok/s
╭────────────┬───────╮
│ ctx \ conc │ 1 │
├────────────┼───────┤
│ 0 │ 151.0 │
│ 32k │ 60.1 │
│ 64k │ 143.3 │
╰────────────┴───────╯
Client request latency: p50
/ p90 ms
╭────────────┬─────────────╮
│ ctx \ conc │ 1 │
├────────────┼─────────────┤
│ 0 │ 6.8k/6.9k │
│ 32k │ 13.0k/13.0k │
│ 64k │ 17.4k/17.4k │
╰────────────┴─────────────╯
Aggregate cells show dim detail as TTFT ms / ITL ms for the same ctx/conc
coordinate. ITL is computed from observed generated tokens, including streams
stopped at the measurement boundary; a missing ITL means no stream produced at
least two measured output tokens. Per-request tok/s and request latency are
shown in separate per-cell matrices. Completion/sample counts and full
request-level distributions remain in JSON under request_samples.
Sustained mode: client latency metrics explain request UX variance; aggregate
tok/s remains the primary throughput signal.
ITL=(last_token_time-first_token_time)/(output_tokens-1), user tok/s=1/ITL.
Hardware Summary
╭───┬─┬───────┬───────────┬───────┬─────────┬─────┬──────┬─────┬───────────────╮
│ … │ │ mode │ GPU avg/… │ Mem … │ W avg/… │ T … │ CPU… │ VR… │ PCIe rx/tx a… │
├───┼─┼───────┼───────────┼───────┼─────────┼─────┼──────┼─────┼───────────────┤
│ 0 │ │ sust… │ 49/99% │ 28% │ 880/890 │ 69C │ 67C │ 48… │ 3018/2770 │
│ … │ │ sust… │ 50/100% │ 21% │ 938/10… │ 79C │ 67C │ 48… │ 13663/13205 │
│ … │ │ sust… │ 50/100% │ 19% │ 980/10… │ 87C │ 67C │ 48… │ 14877/17000 │
╰───┴─┴───────┴───────────┴───────┴─────────┴─────┴──────┴─────┴───────────────╯
╭─────────────────────── Whole-run GPU Power ────────────────────────╮
│ avg 830 W | max 1,035 W | limit 1,800 W | over 2m 35s | 66 samples │
╰────────────────────────────────────────────────────────────────────╯
Hardware summary is sampled from nvidia-smi during the measured part of each
cell. Whole-run GPU power is the sampled sum of GPU power draw across the
complete benchmark run, not wall-outlet system power. PCIe rx/tx is MB/s and is
a coarse live diagnostic, not a per-kernel NCCL profiler.
╭────────────────────────────────── Phase 3 ───────────────────────────────────╮
│ Burst / E2E Decode │
│ Not run. Re-run with --run-burst to append a finite client-facing request │
│ burst after Sustained Decode. This is intentionally disabled by default │
│ because it adds another full decode matrix. │
╰──────────────────────────────────────────────────────────────────────────────╯
╭────────────────────────────── Primary Summary ───────────────────────────────╮
│ Primary matrices repeated last so the important numbers are visible without │
│ scrolling back through diagnostics. │
╰──────────────────────────────────────────────────────────────────────────────╯
Prefill tok/s Aggregate decode tok/s
╭─────┬────────┬────────┬───────┬───╮╭────────────┬───────╮
│ ctx │ tokens │ TTFT s │ tok/s │ N ││ ctx \ conc │ 1 │
├─────┼────────┼────────┼───────┼───┤├────────────┼───────┤
│ 32k │ 32,323 │ 5.19 │ 6,225 │ 2 ││ 0 │ 151.0 │
│ 64k │ 64,515 │ 10.61 │ 6,083 │ 1 ││ 32k │ 60.1 │
╰─────┴────────┴────────┴───────┴───╯│ 64k │ 143.3 │
╰────────────┴───────╯
MTP-normalized decode
steps/s (accept len)
╭────────────┬─────────────╮
│ ctx \ conc │ 1 │
├────────────┼─────────────┤
│ 0 │ 51.6 (2.92) │
│ 32k │ 24.3 (2.47) │
│ 64k │ 51.5 (2.78) │
╰────────────┴─────────────╯
steps/s = tok/s ÷ accept_len: engine forward passes per second, independent of
MTP acceptance, so runs with different acceptance are directly comparable.
(accept len) = tokens emitted per engine step.
Results saved to
/home/brandonmusic/KLC_SANDBOXES/glm53-exl3-k4-sm120/results/dflash2-v84-600w-qu
ick/llm-decode-c1-prefill32k64k.json
Script done on 2026-08-28 06:43:35-04:00 [COMMAND_EXIT_CODE="0"]