Script started on 2026-08-28 06:40:58-04:00 [COMMAND="python3 llm_decode_bench.py --host 127.0.0.1 --port 5001 --model GLM-5.3-Flash-EXL3-4bpw --concurrency 1 --contexts 0,32k,64k --duration 10 --max-tokens 1024 --decode-warmup-seconds 3 --standalone-prefill --prefill-contexts 32k,64k --prefill-duration 10 --prefill-metric auto --token-targeting estimate --kv-budget 129473 --display-mode plain --no-resume --output /home/brandonmusic/KLC_SANDBOXES/glm53-exl3-k4-sm120/results/dflash2-v84-600w-quick/llm-decode-c1-prefill32k64k.json" ] ╭──────────────────────────── NVIDIA P2P Override ─────────────────────────────╮ │ Effective: yes │ │ Configured file: yes (/etc/modprobe.d/nvidia-p2p-override.conf) │ │ Runtime: ForceP2P=0x11; RMForceP2PType=1; RMPcieP2PType=2; │ │ GrdmaPciTopoCheckOverride=1; EnableResizableBar=1; DmaRemapPeerMmio=1 │ ╰──────────────────────────────────────────────────────────────────────────────╯ ╭─────────────────────────────── Configuration ────────────────────────────────╮ │ LLM Inference Benchmark │ │ Model: GLM-5.3-Flash-EXL3-4bpw @ 127.0.0.1:5001 │ │ Decode concurrency: [1] │ │ Decode contexts: ['0', '32k', '64k'] │ │ Duration: 10.0s per decode test | Max tokens: 1024 │ │ Pre-decode warmup: C=1 max-runnable context for 3s │ │ Prefill: standalone cold profile (auto) | Sustained decode: 3 cells │ ╰──────────────────────────────────────────────────────────────────────────────╯ KV cache budget (manual): 129,473 tokens Engine: vLLM 0.1.dev20111+g7f1e92bec.d20260827 Models: ['GLM-5.3-Flash-EXL3-4bpw'] Model context length: 98,304 tokens Prefill tests: standalone cold profile ['32k', '64k'] Calibrating padding text (run=xxeawouskdhk, up to 64k)... Token targeting: single-point estimate from 8k (use --token-targeting exact for /tokenize binary search) Calibrated: 6.18 chars/token (cached, source=8k) 32k: 202,404 chars (~32,767 tokens) 64k: 404,809 chars (~65,535 tokens) Done. llm-decode-bench v0.4.29 Prefill Speed (C=1, client ISL / TTFT) PCIe rx/tx Context Tokens TTFT (s) Client tok/s Server tok/s avg N ────────────────────────────────────────────────────────────────────────────── 32k 32,323 5.19 6,225 6,277 (2) 19267/19610 2 64k 64,515 10.61 6,083 6,130 (1) 25574/23986 1 Client tok/s = prompt_tokens / TTFT. Integrated scout rows come from the prefix-cache scout request that decode needs anyway. Server tok/s is optional Prometheus validation when the engine exports prefill counters and the exact counter delta is uncontaminated; for vLLM this uses newly computed KV tokens, not request prompt tokens. ╭────────────────────────────────── Phase 2 ───────────────────────────────────╮ │ Sustained Decode │ │ Steady-state decode throughput after the engine has admitted the requested │ │ concurrency and passed warmup. Use this as the main tuning/regression signal │ │ for kernels, NCCL, DCP, MTP, and scheduler changes. │ ╰──────────────────────────────────────────────────────────────────────────────╯ Aggregate tok/s + TTFT/ITL ╭────────────┬─────────────╮ │ ctx \ conc │ 1 │ ├────────────┼─────────────┤ │ 0 │ 151.0 137/6 │ │ 32k │ 60.1 5k/6 │ │ 64k │ 143.3 11k/7 │ ╰────────────┴─────────────╯ Sustained Decode: aggregate tok/s uses OpenAI stream usage by default (continuous completion_tokens when the server supports it). Prometheus is kept as validation/scheduler data. Aggregate source(s): openai_continuous_usage Per-Request tok/s ╭────────────┬───────╮ │ ctx \ conc │ 1 │ ├────────────┼───────┤ │ 0 │ 151.0 │ │ 32k │ 60.1 │ │ 64k │ 143.3 │ ╰────────────┴───────╯ Client request latency: p50 / p90 ms ╭────────────┬─────────────╮ │ ctx \ conc │ 1 │ ├────────────┼─────────────┤ │ 0 │ 6.8k/6.9k │ │ 32k │ 13.0k/13.0k │ │ 64k │ 17.4k/17.4k │ ╰────────────┴─────────────╯ Aggregate cells show dim detail as TTFT ms / ITL ms for the same ctx/conc coordinate. ITL is computed from observed generated tokens, including streams stopped at the measurement boundary; a missing ITL means no stream produced at least two measured output tokens. Per-request tok/s and request latency are shown in separate per-cell matrices. Completion/sample counts and full request-level distributions remain in JSON under request_samples. Sustained mode: client latency metrics explain request UX variance; aggregate tok/s remains the primary throughput signal. ITL=(last_token_time-first_token_time)/(output_tokens-1), user tok/s=1/ITL. Hardware Summary ╭───┬─┬───────┬───────────┬───────┬─────────┬─────┬──────┬─────┬───────────────╮ │ … │ │ mode │ GPU avg/… │ Mem … │ W avg/… │ T … │ CPU… │ VR… │ PCIe rx/tx a… │ ├───┼─┼───────┼───────────┼───────┼─────────┼─────┼──────┼─────┼───────────────┤ │ 0 │ │ sust… │ 49/99% │ 28% │ 880/890 │ 69C │ 67C │ 48… │ 3018/2770 │ │ … │ │ sust… │ 50/100% │ 21% │ 938/10… │ 79C │ 67C │ 48… │ 13663/13205 │ │ … │ │ sust… │ 50/100% │ 19% │ 980/10… │ 87C │ 67C │ 48… │ 14877/17000 │ ╰───┴─┴───────┴───────────┴───────┴─────────┴─────┴──────┴─────┴───────────────╯ ╭─────────────────────── Whole-run GPU Power ────────────────────────╮ │ avg 830 W | max 1,035 W | limit 1,800 W | over 2m 35s | 66 samples │ ╰────────────────────────────────────────────────────────────────────╯ Hardware summary is sampled from nvidia-smi during the measured part of each cell. Whole-run GPU power is the sampled sum of GPU power draw across the complete benchmark run, not wall-outlet system power. PCIe rx/tx is MB/s and is a coarse live diagnostic, not a per-kernel NCCL profiler. ╭────────────────────────────────── Phase 3 ───────────────────────────────────╮ │ Burst / E2E Decode │ │ Not run. Re-run with --run-burst to append a finite client-facing request │ │ burst after Sustained Decode. This is intentionally disabled by default │ │ because it adds another full decode matrix. │ ╰──────────────────────────────────────────────────────────────────────────────╯ ╭────────────────────────────── Primary Summary ───────────────────────────────╮ │ Primary matrices repeated last so the important numbers are visible without │ │ scrolling back through diagnostics. │ ╰──────────────────────────────────────────────────────────────────────────────╯ Prefill tok/s Aggregate decode tok/s ╭─────┬────────┬────────┬───────┬───╮╭────────────┬───────╮ │ ctx │ tokens │ TTFT s │ tok/s │ N ││ ctx \ conc │ 1 │ ├─────┼────────┼────────┼───────┼───┤├────────────┼───────┤ │ 32k │ 32,323 │ 5.19 │ 6,225 │ 2 ││ 0 │ 151.0 │ │ 64k │ 64,515 │ 10.61 │ 6,083 │ 1 ││ 32k │ 60.1 │ ╰─────┴────────┴────────┴───────┴───╯│ 64k │ 143.3 │ ╰────────────┴───────╯ MTP-normalized decode steps/s (accept len) ╭────────────┬─────────────╮ │ ctx \ conc │ 1 │ ├────────────┼─────────────┤ │ 0 │ 51.6 (2.92) │ │ 32k │ 24.3 (2.47) │ │ 64k │ 51.5 (2.78) │ ╰────────────┴─────────────╯ steps/s = tok/s ÷ accept_len: engine forward passes per second, independent of MTP acceptance, so runs with different acceptance are directly comparable. (accept len) = tokens emitted per engine step. Results saved to /home/brandonmusic/KLC_SANDBOXES/glm53-exl3-k4-sm120/results/dflash2-v84-600w-qu ick/llm-decode-c1-prefill32k64k.json Script done on 2026-08-28 06:43:35-04:00 [COMMAND_EXIT_CODE="0"]