Image-Text-to-Text
Transformers
Safetensors
glm5_next
glm
exl3
tr3
vllm
sm120
nvfp4
dflash2
multimodal
shapleymcg
conversational
Eval Results (legacy)
4-bit precision
Instructions to use brandonmusic/GLM-5.3-Flash-tr3-4bpw with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use brandonmusic/GLM-5.3-Flash-tr3-4bpw with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("image-text-to-text", model="brandonmusic/GLM-5.3-Flash-tr3-4bpw") messages = [ { "role": "user", "content": [ {"type": "image", "url": "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/p-blog/candy.JPG"}, {"type": "text", "text": "What animal is on the candy?"} ] }, ] pipe(text=messages)# pip install -U transformers accelerate # Load model directly from transformers import AutoProcessor, AutoModelForMultimodalLM processor = AutoProcessor.from_pretrained("brandonmusic/GLM-5.3-Flash-tr3-4bpw") model = AutoModelForMultimodalLM.from_pretrained("brandonmusic/GLM-5.3-Flash-tr3-4bpw", device_map="auto") messages = [ { "role": "user", "content": [ {"type": "image", "url": "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/p-blog/candy.JPG"}, {"type": "text", "text": "What animal is on the candy?"} ] }, ] inputs = processor.apply_chat_template( messages, add_generation_prompt=True, tokenize=True, return_dict=True, return_tensors="pt", ).to(model.device) outputs = model.generate(**inputs, max_new_tokens=256) print(processor.decode(outputs[0][inputs["input_ids"].shape[-1]:])) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use brandonmusic/GLM-5.3-Flash-tr3-4bpw with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "brandonmusic/GLM-5.3-Flash-tr3-4bpw" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "brandonmusic/GLM-5.3-Flash-tr3-4bpw", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }'Use Docker
docker model run hf.co/brandonmusic/GLM-5.3-Flash-tr3-4bpw
- SGLang
How to use brandonmusic/GLM-5.3-Flash-tr3-4bpw with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "brandonmusic/GLM-5.3-Flash-tr3-4bpw" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "brandonmusic/GLM-5.3-Flash-tr3-4bpw", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "brandonmusic/GLM-5.3-Flash-tr3-4bpw" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "brandonmusic/GLM-5.3-Flash-tr3-4bpw", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }' - Docker Model Runner
How to use brandonmusic/GLM-5.3-Flash-tr3-4bpw with Docker Model Runner:
docker model run hf.co/brandonmusic/GLM-5.3-Flash-tr3-4bpw
Download runtime-results/v84/benchmarks/llm-decode-c1-c4-64k.json from brandonmusic/GLM-5.3-Flash-tr3-4bpw: direct link, hf CLI and curl.
- Browser
- Download file 52.9 kB
-
https://huggingface.co/brandonmusic/GLM-5.3-Flash-tr3-4bpw/resolve/main/runtime-results/v84/benchmarks/llm-decode-c1-c4-64k.json
- Command line
-
hf download hf://brandonmusic/GLM-5.3-Flash-tr3-4bpw/runtime-results/v84/benchmarks/llm-decode-c1-c4-64k.json
-
curl -L -o llm-decode-c1-c4-64k.json https://huggingface.co/brandonmusic/GLM-5.3-Flash-tr3-4bpw/resolve/main/runtime-results/v84/benchmarks/llm-decode-c1-c4-64k.json
52.9 kB
| { | |
| "metadata": { | |
| "version": "0.4.29", | |
| "engine": "vllm", | |
| "model": "GLM-5.3-Flash-EXL3-4bpw", | |
| "server": "127.0.0.1:5001", | |
| "timestamp": "2026-08-28T01:14:26.824486", | |
| "decode_mode": "duration", | |
| "primary_decode_layer": "sustained_decode", | |
| "duration_per_test": 10.0, | |
| "request_count": 0, | |
| "warmup_request_count": 0, | |
| "run_burst": false, | |
| "prefill_mode": "integrated_decode_scout", | |
| "standalone_prefill": false, | |
| "prefill_only": false, | |
| "skip_prefill": false, | |
| "burst_e2e_status": "not_run_use_--run-burst", | |
| "burst_request_count": 0, | |
| "burst_warmup_request_count": 0, | |
| "burst_requests_per_concurrency": 5, | |
| "decode_warmup_seconds": 3.0, | |
| "decode_warmup_context": 65536, | |
| "decode_warmup_concurrency": 1, | |
| "cell_warmup_timeout_seconds": 0.0, | |
| "cell_warmup_timeout_policy": "<=32k:60s,64k:120s,>=128k:180s when override is 0", | |
| "show_capacity_limited_values": false, | |
| "max_tokens": 1024, | |
| "temperature": null, | |
| "ignore_eos": true, | |
| "max_total_tokens": 884736, | |
| "dcp_size": 0, | |
| "metrics_available": true, | |
| "metrics_warning": "", | |
| "concurrency_levels": [ | |
| 1, | |
| 2, | |
| 4 | |
| ], | |
| "context_lengths": [ | |
| 0, | |
| 65536 | |
| ], | |
| "startup_diagnostics_available": true, | |
| "nvidia_p2p_override_effective": true, | |
| "p2pmark_status": "not_run", | |
| "amd_fabric_status": "not_run" | |
| }, | |
| "startup_diagnostics": { | |
| "version": "0.4.29", | |
| "server_url": "http://127.0.0.1:5001", | |
| "hostname": "pop-os", | |
| "uname": "Linux pop-os 6.18.7-76061807-generic #202601231045~1769703228~24.04~cb87b5b SMP PREEMPT_DYNAMIC Thu J x86_64 x86_64 x86_64 GNU/Linux", | |
| "env": {}, | |
| "args": { | |
| "concurrency": "1,2,4", | |
| "contexts": "0,64k", | |
| "max_tokens": 1024, | |
| "duration": 10.0, | |
| "request_count": 0, | |
| "run_burst": false, | |
| "standalone_prefill": false, | |
| "prefill_only": false, | |
| "skip_prefill": false, | |
| "prefill_contexts": "8k,64k,128k", | |
| "prefill_metric": "client", | |
| "dcp_size": 0, | |
| "kv_budget": 0 | |
| }, | |
| "nvidia_p2p_override": { | |
| "effective": true, | |
| "configured": true, | |
| "params_path": "/proc/driver/nvidia/params", | |
| "params_available": true, | |
| "modprobe_path": "/etc/modprobe.d/nvidia-p2p-override.conf", | |
| "modprobe_available": true, | |
| "runtime": { | |
| "ForceP2P": "0x11", | |
| "RMForceP2PType": "1", | |
| "RMPcieP2PType": "2", | |
| "GrdmaPciTopoCheckOverride": "1", | |
| "EnableResizableBar": "1", | |
| "DmaRemapPeerMmio": "1" | |
| }, | |
| "expected": { | |
| "ForceP2P": "0x11", | |
| "RMForceP2PType": "1", | |
| "RMPcieP2PType": "2", | |
| "GrdmaPciTopoCheckOverride": "1", | |
| "EnableResizableBar": "1" | |
| }, | |
| "missing": [], | |
| "mismatched": {}, | |
| "registry_dwords": "ForceP2P=0x11;RMForceP2PType=1;RMPcieP2PType=2;GrdmaPciTopoCheckOverride=1;EnableResizableBar=1", | |
| "suggested_modprobe_line": "options nvidia NVreg_RegistryDwords=\"ForceP2P=0x11;RMForceP2PType=1;RMPcieP2PType=2;GrdmaPciTopoCheckOverride=1;EnableResizableBar=1\"", | |
| "suggested_reload": "stop GPU workloads, then reload NVIDIA modules or reboot; the modprobe file alone is not enough until the nvidia module is reloaded" | |
| }, | |
| "p2pmark": { | |
| "status": "not_run" | |
| }, | |
| "amd_fabric": { | |
| "status": "not_run" | |
| }, | |
| "nvidia_smi_query": { | |
| "cmd": [ | |
| "nvidia-smi", | |
| "--query-gpu=index,name,driver_version,pci.bus_id,pcie.link.gen.current,pcie.link.width.current,power.limit", | |
| "--format=csv,noheader,nounits" | |
| ], | |
| "returncode": 0, | |
| "stdout": "0, NVIDIA RTX PRO 6000 Blackwell Max-Q Workstation Edition, 610.57.04, 00000000:01:00.0, 1, 16, 300.00\n1, NVIDIA RTX PRO 6000 Blackwell Workstation Edition, 610.57.04, 00000000:21:00.0, 1, 16, 300.00\n2, NVIDIA RTX PRO 6000 Blackwell Max-Q Workstation Edition, 610.57.04, 00000000:81:00.0, 1, 16, 300.00\n3, NVIDIA RTX PRO 6000 Blackwell Workstation Edition, 610.57.04, 00000000:C1:00.0, 1, 16, 300.00", | |
| "stderr": "" | |
| }, | |
| "nvidia_smi_topo": { | |
| "cmd": [ | |
| "nvidia-smi", | |
| "topo", | |
| "-m" | |
| ], | |
| "returncode": 0, | |
| "stdout": "\u001b[4mGPU0\tGPU1\tGPU2\tGPU3\tCPU Affinity\tNUMA Affinity\tGPU NUMA ID\u001b[0m\nGPU0\t X \tNODE\tNODE\tNODE\t0-47\t0\t\tN/A\nGPU1\tNODE\t X \tNODE\tNODE\t0-47\t0\t\tN/A\nGPU2\tNODE\tNODE\t X \tNODE\t0-47\t0\t\tN/A\nGPU3\tNODE\tNODE\tNODE\t X \t0-47\t0\t\tN/A\n\nLegend:\n\n X = Self\n SYS = Connection traversing PCIe as well as the SMP interconnect between NUMA nodes (e.g., QPI/UPI)\n NODE = Connection traversing PCIe as well as the interconnect between PCIe Host Bridges within a NUMA node\n PHB = Connection traversing PCIe as well as a PCIe Host Bridge (typically the CPU)\n PXB = Connection traversing multiple PCIe bridges (without traversing the PCIe Host Bridge)\n PIX = Connection traversing at most a single PCIe bridge\n NV# = Connection traversing a bonded set of # NVLinks", | |
| "stderr": "" | |
| } | |
| }, | |
| "nvidia_p2p_override": { | |
| "effective": true, | |
| "configured": true, | |
| "params_path": "/proc/driver/nvidia/params", | |
| "params_available": true, | |
| "modprobe_path": "/etc/modprobe.d/nvidia-p2p-override.conf", | |
| "modprobe_available": true, | |
| "runtime": { | |
| "ForceP2P": "0x11", | |
| "RMForceP2PType": "1", | |
| "RMPcieP2PType": "2", | |
| "GrdmaPciTopoCheckOverride": "1", | |
| "EnableResizableBar": "1", | |
| "DmaRemapPeerMmio": "1" | |
| }, | |
| "expected": { | |
| "ForceP2P": "0x11", | |
| "RMForceP2PType": "1", | |
| "RMPcieP2PType": "2", | |
| "GrdmaPciTopoCheckOverride": "1", | |
| "EnableResizableBar": "1" | |
| }, | |
| "missing": [], | |
| "mismatched": {}, | |
| "registry_dwords": "ForceP2P=0x11;RMForceP2PType=1;RMPcieP2PType=2;GrdmaPciTopoCheckOverride=1;EnableResizableBar=1", | |
| "suggested_modprobe_line": "options nvidia NVreg_RegistryDwords=\"ForceP2P=0x11;RMForceP2PType=1;RMPcieP2PType=2;GrdmaPciTopoCheckOverride=1;EnableResizableBar=1\"", | |
| "suggested_reload": "stop GPU workloads, then reload NVIDIA modules or reboot; the modprobe file alone is not enough until the nvidia module is reloaded" | |
| }, | |
| "p2pmark": { | |
| "status": "not_run" | |
| }, | |
| "amd_fabric": { | |
| "status": "not_run" | |
| }, | |
| "hardware_run_summary": { | |
| "samples": 243, | |
| "duration_seconds": 577.387, | |
| "gpu_count": 4, | |
| "cpu_util_avg_pct": 6.97, | |
| "cpu_temp_max_c": 67.5, | |
| "gpu_util_avg_pct": 48.04, | |
| "gpu_util_max_pct": 100.0, | |
| "mem_util_avg_pct": 16.63, | |
| "mem_util_max_pct": 62.0, | |
| "temp_avg_c": 50.75, | |
| "temp_max_c": 75.0, | |
| "power_total_avg_w": 604.36, | |
| "power_total_max_w": 624.86, | |
| "power_limit_total_w": 1200.0, | |
| "vram_used_avg_mb": 189600.28, | |
| "vram_used_max_mb": 189606.0, | |
| "vram_total_mb": 391548.0, | |
| "vram_used_avg_pct": 48.42, | |
| "vram_used_max_pct": 48.42, | |
| "pcie_rx_avg_mb_s": 9462.3, | |
| "pcie_rx_max_mb_s": 24512.0, | |
| "pcie_tx_avg_mb_s": 9641.98, | |
| "pcie_tx_max_mb_s": 25466.0 | |
| }, | |
| "event_log": [ | |
| "01:04:47 benchmark start engine=vllm", | |
| "01:04:47 startup server=http://127.0.0.1:5001 model=GLM-5.3-Flash-EXL3-4bpw", | |
| "01:04:47 startup decode concurrency=1,2,4 contexts=0,64k", | |
| "01:04:47 startup NVIDIA P2P override: enabled: runtime NVIDIA P2P override matches expected RegistryDwords", | |
| "01:04:47 startup engine vLLM 0.1.dev20111+g7f1e92bec.d20260827 models=['GLM-5.3-Flash-EXL3-4bpw']", | |
| "01:04:47 startup KV cache budget from vLLM metrics: 884,736 tokens (54 blocks x 8192; local 442,368 \u00d7 CP 2; CP source: local process)", | |
| "01:04:47 startup model context length: 98,304 tokens", | |
| "01:04:47 startup prefill tests: integrated from decode scout requests ['64k']; scout-only extras ['8k']", | |
| "01:04:47 startup calibrating padding text run=erjqltqqoilw up_to=64k", | |
| "01:04:47 startup token targeting: estimate from 8k", | |
| "01:04:47 startup calibrated: 6.18 chars/token (cached, source=8k)", | |
| "01:04:47 startup context 8k: 50,601 chars (~8,191 tokens)", | |
| "01:04:47 startup context 64k: 404,809 chars (~65,535 tokens)", | |
| "01:04:47 startup startup preparation done", | |
| "01:04:47 hardware monitor interval=2s", | |
| "01:04:47 decode warmup start", | |
| "01:04:48 prefill warmup start ctx=8k", | |
| "01:04:51 prefill warmup done ctx=8k", | |
| "01:04:51 prefill scout-only start ctx=8k", | |
| "01:04:54 prefill scout-only done ctx=8k 3,897 tok/s", | |
| "01:04:54 decode warmup start C=1 ctx=64k 3s", | |
| "01:04:54 cell start C=1 ctx=64k", | |
| "01:05:28 ready C=1 ctx=64k running_reqs=1/1, queue_reqs=0, active_streams=1/1, stable=3.0s", | |
| "01:05:31 cell done C=1 ctx=64k 107.8 tok/s | norm 43.2 step/s len=2.50", | |
| "01:05:31 decode warmup done C=1 ctx=64k", | |
| "01:05:33 cell start C=1 ctx=0", | |
| "01:05:39 ready C=1 ctx=0 running_reqs=1/1, queue_reqs=0, active_streams=1/1, stable=3.0s", | |
| "01:05:49 cell done C=1 ctx=0 129.5 tok/s | norm 44.1 step/s len=2.94", | |
| "01:05:51 cell start C=1 ctx=64k", | |
| "01:05:51 integrated prefill start ctx=64k", | |
| "01:06:08 integrated prefill done ctx=64k 4,297 tok/s", | |
| "01:06:27 ready C=1 ctx=64k running_reqs=1/1, queue_reqs=0, active_streams=1/1, stable=3.0s", | |
| "01:06:46 cell done C=1 ctx=64k 122.2 tok/s | norm 41.9 step/s len=2.92", | |
| "01:06:48 cell start C=2 ctx=0", | |
| "01:07:49 warmup timeout C=2 ctx=0 running_reqs=1/2", | |
| "01:07:59 cell done C=2 ctx=0 115.9 tok/s | norm 42.2 step/s len=2.75", | |
| "01:08:01 cell start C=4 ctx=0", | |
| "01:09:02 warmup timeout C=4 ctx=0 running_reqs=1/4", | |
| "01:09:13 cell done C=4 ctx=0 119.0 tok/s | norm 42.7 step/s len=2.78", | |
| "01:09:15 cell start C=2 ctx=64k", | |
| "01:11:16 warmup timeout C=2 ctx=64k running_reqs=1/2", | |
| "01:11:42 cell done C=2 ctx=64k 12.1 tok/s | norm 3.9 step/s len=3.10", | |
| "01:11:44 cell start C=4 ctx=64k", | |
| "01:13:44 warmup timeout C=4 ctx=64k running_reqs=1/4", | |
| "01:14:24 cell done C=4 ctx=64k 0.1 tok/s" | |
| ], | |
| "prefill": { | |
| "8192": { | |
| "ttft_seconds": 2.104, | |
| "prefill_seconds": 2.104, | |
| "tok_per_sec": 3897.0, | |
| "client_ttft_seconds": 2.104, | |
| "client_tok_per_sec": 3897.0, | |
| "prompt_tokens": 8199, | |
| "samples": 1, | |
| "method": "scout_only", | |
| "server_validation": { | |
| "method": "", | |
| "tok_per_sec": 0.0, | |
| "prefill_seconds": 0.0, | |
| "prompt_tokens": 0, | |
| "request_prompt_tokens": 0, | |
| "cached_tokens": 0, | |
| "token_source": "", | |
| "samples": 0, | |
| "invalid_reason": "" | |
| }, | |
| "hardware_summary": {} | |
| }, | |
| "65536": { | |
| "ttft_seconds": 15.015, | |
| "prefill_seconds": 15.015, | |
| "tok_per_sec": 4297.0, | |
| "client_ttft_seconds": 15.015, | |
| "client_tok_per_sec": 4297.0, | |
| "prompt_tokens": 64513, | |
| "samples": 1, | |
| "method": "integrated_scout", | |
| "server_validation": { | |
| "method": "prometheus:kv_computed", | |
| "tok_per_sec": 4320.0, | |
| "prefill_seconds": 14.934, | |
| "prompt_tokens": 64513, | |
| "request_prompt_tokens": 0, | |
| "cached_tokens": 0, | |
| "token_source": "", | |
| "samples": 1, | |
| "invalid_reason": "" | |
| }, | |
| "hardware_summary": {} | |
| } | |
| }, | |
| "results": [ | |
| { | |
| "concurrency": 1, | |
| "context_tokens": 0, | |
| "benchmark_mode": "duration", | |
| "request_count_target": 0, | |
| "warmup_request_count": 0, | |
| "measurement_seconds": 9.988261, | |
| "measurement_wall_seconds": 10.000404, | |
| "client_output_tokens": 1293, | |
| "server_output_tokens": 1293, | |
| "aggregate_source": "openai_continuous_usage", | |
| "aggregate_tps": 129.45196578500713, | |
| "per_request_avg_tps": 129.45196578500713, | |
| "ttft_avg": 0.13393967098090798, | |
| "ttft_p50": 0.13393967098090798, | |
| "ttft_p90": 0.1397153286030516, | |
| "ttft_p99": 0.14101485156803392, | |
| "time_to_second_token_avg": 0.020580344484187663, | |
| "time_to_second_token_p50": 0.020580344484187663, | |
| "time_to_second_token_p90": 0.021969041670672596, | |
| "time_to_second_token_p99": 0.022281498537631707, | |
| "request_latency_avg": 8.68479720497271, | |
| "request_latency_p50": 8.68479720497271, | |
| "request_latency_p90": 8.68479720497271, | |
| "request_latency_p99": 8.68479720497271, | |
| "inter_token_latency_avg": 0.0078036521248696505, | |
| "inter_token_latency_p50": 0.0078036521248696505, | |
| "inter_token_latency_p90": 0.00825326384121595, | |
| "inter_token_latency_p99": 0.008354426477393866, | |
| "output_tps_per_user_avg": 128.81325648749555, | |
| "output_tps_per_user_p50": 128.81325648749555, | |
| "output_tps_per_user_p90": 136.2349032255562, | |
| "output_tps_per_user_p99": 137.90477374161983, | |
| "e2e_output_tps_per_user_avg": 117.9071860668988, | |
| "e2e_output_tps_per_user_p50": 117.9071860668988, | |
| "e2e_output_tps_per_user_p90": 117.9071860668988, | |
| "e2e_output_tps_per_user_p99": 117.9071860668988, | |
| "chunk_inter_token_latency_avg": 0.022398087895907616, | |
| "chunk_inter_token_latency_p50": 0.022398087895907616, | |
| "chunk_inter_token_latency_p90": 0.02262544274799395, | |
| "chunk_inter_token_latency_p99": 0.022676597589713375, | |
| "input_seq_len_avg": 78.0, | |
| "output_seq_len_avg": 1024.0, | |
| "output_seq_len_p50": 1024.0, | |
| "output_seq_len_p90": 1024.0, | |
| "output_seq_len_p99": 1024.0, | |
| "request_count": 2, | |
| "completed_request_count": 1, | |
| "request_samples": [ | |
| { | |
| "ttft": 0.12672009895322844, | |
| "time_to_second_token": 0.018844473001081496, | |
| "latency": 8.68479720497271, | |
| "inter_token_latency_avg": 0.008365666770302524, | |
| "chunk_inter_token_latency_avg": 0.022113894330799695, | |
| "input_tokens": 78, | |
| "output_tokens": 1024, | |
| "output_tps_per_user": 119.53619806491974, | |
| "e2e_output_tps_per_user": 117.9071860668988, | |
| "completed": true | |
| }, | |
| { | |
| "ttft": 0.1411592430085875, | |
| "time_to_second_token": 0.02231621596729383, | |
| "latency": 0.0, | |
| "inter_token_latency_avg": 0.007241637479436777, | |
| "chunk_inter_token_latency_avg": 0.022682281461015533, | |
| "input_tokens": 78, | |
| "output_tokens": 925, | |
| "output_tps_per_user": 138.09031491007136, | |
| "e2e_output_tps_per_user": 0.0, | |
| "completed": false | |
| } | |
| ], | |
| "total_tokens": 1293, | |
| "wall_time": 15.541175424004905, | |
| "num_completed": 1, | |
| "num_errors": 0, | |
| "server_gen_throughput": 129.23415974403488, | |
| "server_utilization": 0.6603773584905661, | |
| "server_spec_accept_rate": 0.2775789131142206, | |
| "server_spec_accept_length": 2.9430523917995446, | |
| "server_spec_drafts": 439, | |
| "server_spec_draft_tokens": 3073, | |
| "server_spec_accepted_tokens": 853, | |
| "server_spec_pos_accept": [ | |
| 0.7699, | |
| 0.5011, | |
| 0.2825, | |
| 0.18, | |
| 0.1071, | |
| 0.0638, | |
| 0.0387 | |
| ], | |
| "server_engine_steps": 440.0, | |
| "server_steps_per_s": 44.05171302815401, | |
| "server_accept_len_effective": 2.9386363636363635, | |
| "accept_norm_tps": 0.0, | |
| "accept_norm_ref_len": 0.0, | |
| "avg_running_reqs": 1, | |
| "max_running_reqs": 1, | |
| "effective_concurrency": 1, | |
| "avg_queue_reqs": 0, | |
| "max_queue_reqs": 0, | |
| "queue_fraction": 0.0, | |
| "underfilled": false, | |
| "warmup_timed_out": false, | |
| "warmup_duration": 5.529, | |
| "ready_reason": "running_reqs=1/1, queue_reqs=0, active_streams=1/1, stable=3.0s", | |
| "timeout_reason": "", | |
| "capacity_limited": false, | |
| "hardware_summary": { | |
| "samples": 4, | |
| "duration_seconds": 7.182, | |
| "gpu_count": 4, | |
| "cpu_util_avg_pct": 7.22, | |
| "cpu_temp_max_c": 66.62, | |
| "gpu_util_avg_pct": 49.5, | |
| "gpu_util_max_pct": 99.0, | |
| "mem_util_avg_pct": 26.75, | |
| "mem_util_max_pct": 62.0, | |
| "temp_avg_c": 44.25, | |
| "temp_max_c": 57.0, | |
| "power_total_avg_w": 619.0, | |
| "power_total_max_w": 619.23, | |
| "power_limit_total_w": 1200.0, | |
| "vram_used_avg_mb": 189604.0, | |
| "vram_used_max_mb": 189604.0, | |
| "vram_total_mb": 391548.0, | |
| "vram_used_avg_pct": 48.42, | |
| "vram_used_max_pct": 48.42, | |
| "pcie_rx_avg_mb_s": 2535.0, | |
| "pcie_rx_max_mb_s": 2601.0, | |
| "pcie_tx_avg_mb_s": 2322.5, | |
| "pcie_tx_max_mb_s": 2407.0 | |
| } | |
| }, | |
| { | |
| "concurrency": 1, | |
| "context_tokens": 65536, | |
| "benchmark_mode": "duration", | |
| "request_count_target": 0, | |
| "warmup_request_count": 0, | |
| "measurement_seconds": 4.417692, | |
| "measurement_wall_seconds": 10.0022, | |
| "client_output_tokens": 540, | |
| "server_output_tokens": 540, | |
| "aggregate_source": "openai_continuous_usage", | |
| "aggregate_tps": 122.23577506588123, | |
| "per_request_avg_tps": 122.23577506588123, | |
| "ttft_avg": 15.145516593009233, | |
| "ttft_p50": 15.145516593009233, | |
| "ttft_p90": 15.145516593009233, | |
| "ttft_p99": 15.145516593009233, | |
| "time_to_second_token_avg": 0.004005067981779575, | |
| "time_to_second_token_p50": 0.004005067981779575, | |
| "time_to_second_token_p90": 0.004005067981779575, | |
| "time_to_second_token_p99": 0.004005067981779575, | |
| "request_latency_avg": 23.09197867201874, | |
| "request_latency_p50": 23.09197867201874, | |
| "request_latency_p90": 23.09197867201874, | |
| "request_latency_p99": 23.09197867201874, | |
| "inter_token_latency_avg": 0.007767802618777622, | |
| "inter_token_latency_p50": 0.007767802618777622, | |
| "inter_token_latency_p90": 0.007767802618777622, | |
| "inter_token_latency_p99": 0.007767802618777622, | |
| "output_tps_per_user_avg": 128.73653581035055, | |
| "output_tps_per_user_p50": 128.73653581035055, | |
| "output_tps_per_user_p90": 128.73653581035055, | |
| "output_tps_per_user_p99": 128.73653581035055, | |
| "e2e_output_tps_per_user_avg": 44.34440264059365, | |
| "e2e_output_tps_per_user_p50": 44.34440264059365, | |
| "e2e_output_tps_per_user_p90": 44.34440264059365, | |
| "e2e_output_tps_per_user_p99": 44.34440264059365, | |
| "chunk_inter_token_latency_avg": 0.024080188118210628, | |
| "chunk_inter_token_latency_p50": 0.024080188118210628, | |
| "chunk_inter_token_latency_p90": 0.024080188118210628, | |
| "chunk_inter_token_latency_p99": 0.024080188118210628, | |
| "input_seq_len_avg": 64513.0, | |
| "output_seq_len_avg": 1024.0, | |
| "output_seq_len_p50": 1024.0, | |
| "output_seq_len_p90": 1024.0, | |
| "output_seq_len_p99": 1024.0, | |
| "request_count": 1, | |
| "completed_request_count": 1, | |
| "request_samples": [ | |
| { | |
| "ttft": 15.145516593009233, | |
| "time_to_second_token": 0.004005067981779575, | |
| "latency": 23.09197867201874, | |
| "inter_token_latency_avg": 0.007767802618777622, | |
| "chunk_inter_token_latency_avg": 0.024080188118210628, | |
| "input_tokens": 64513, | |
| "output_tokens": 1024, | |
| "output_tps_per_user": 128.73653581035055, | |
| "e2e_output_tps_per_user": 44.34440264059365, | |
| "completed": true | |
| } | |
| ], | |
| "total_tokens": 540, | |
| "wall_time": 55.68818276398815, | |
| "num_completed": 1, | |
| "num_errors": 0, | |
| "server_gen_throughput": 53.96348511679134, | |
| "server_utilization": 0.7169811320754718, | |
| "server_spec_accept_rate": 0.277992277992278, | |
| "server_spec_accept_length": 2.945945945945946, | |
| "server_spec_drafts": 185, | |
| "server_spec_draft_tokens": 1295, | |
| "server_spec_accepted_tokens": 360, | |
| "server_spec_pos_accept": [ | |
| 0.7351, | |
| 0.4919, | |
| 0.3351, | |
| 0.2, | |
| 0.0865, | |
| 0.0595, | |
| 0.0378 | |
| ], | |
| "server_engine_steps": 185.0, | |
| "server_steps_per_s": 41.87707108738524, | |
| "server_accept_len_effective": 2.918918918918919, | |
| "accept_norm_tps": 0.0, | |
| "accept_norm_ref_len": 0.0, | |
| "avg_running_reqs": 1, | |
| "max_running_reqs": 1, | |
| "effective_concurrency": 1, | |
| "avg_queue_reqs": 0, | |
| "max_queue_reqs": 0, | |
| "queue_fraction": 0.0, | |
| "underfilled": false, | |
| "warmup_timed_out": false, | |
| "warmup_duration": 35.944, | |
| "ready_reason": "running_reqs=1/1, queue_reqs=0, active_streams=1/1, stable=3.0s", | |
| "timeout_reason": "", | |
| "capacity_limited": false, | |
| "hardware_summary": { | |
| "samples": 8, | |
| "duration_seconds": 16.689, | |
| "gpu_count": 4, | |
| "cpu_util_avg_pct": 7.11, | |
| "cpu_temp_max_c": 66.88, | |
| "gpu_util_avg_pct": 49.81, | |
| "gpu_util_max_pct": 100.0, | |
| "mem_util_avg_pct": 14.59, | |
| "mem_util_max_pct": 59.0, | |
| "temp_avg_c": 47.62, | |
| "temp_max_c": 65.0, | |
| "power_total_avg_w": 617.6, | |
| "power_total_max_w": 622.46, | |
| "power_limit_total_w": 1200.0, | |
| "vram_used_avg_mb": 189604.0, | |
| "vram_used_max_mb": 189604.0, | |
| "vram_total_mb": 391548.0, | |
| "vram_used_avg_pct": 48.42, | |
| "vram_used_max_pct": 48.42, | |
| "pcie_rx_avg_mb_s": 10950.75, | |
| "pcie_rx_max_mb_s": 18611.0, | |
| "pcie_tx_avg_mb_s": 15013.5, | |
| "pcie_tx_max_mb_s": 23519.0 | |
| } | |
| }, | |
| { | |
| "concurrency": 2, | |
| "context_tokens": 0, | |
| "benchmark_mode": "duration", | |
| "request_count_target": 0, | |
| "warmup_request_count": 0, | |
| "measurement_seconds": 9.995853, | |
| "measurement_wall_seconds": 10.000935, | |
| "client_output_tokens": 1159, | |
| "server_output_tokens": 1159, | |
| "aggregate_source": "openai_continuous_usage", | |
| "aggregate_tps": 115.94808445857444, | |
| "per_request_avg_tps": 57.97404222928722, | |
| "ttft_avg": 7.834794788103965, | |
| "ttft_p50": 8.624857554968912, | |
| "ttft_p90": 9.220339812606108, | |
| "ttft_p99": 9.429953900745604, | |
| "time_to_second_token_avg": 0.022789318120986637, | |
| "time_to_second_token_p50": 0.023204773024190217, | |
| "time_to_second_token_p90": 0.024304261978249996, | |
| "time_to_second_token_p99": 0.024516121998894958, | |
| "request_latency_avg": 16.199721001743455, | |
| "request_latency_p50": 17.21378530547372, | |
| "request_latency_p90": 17.579868975607678, | |
| "request_latency_p99": 17.613597658531507, | |
| "inter_token_latency_avg": 0.008167093486157814, | |
| "inter_token_latency_p50": 0.00816347410365027, | |
| "inter_token_latency_p90": 0.008771674752686963, | |
| "inter_token_latency_p99": 0.008959866448893655, | |
| "output_tps_per_user_avg": 123.1466847171223, | |
| "output_tps_per_user_p50": 122.49686681223787, | |
| "output_tps_per_user_p90": 130.61500965849436, | |
| "output_tps_per_user_p99": 145.5801546760414, | |
| "e2e_output_tps_per_user_avg": 66.49566993826983, | |
| "e2e_output_tps_per_user_p50": 59.48782372259443, | |
| "e2e_output_tps_per_user_p90": 77.83971971955552, | |
| "e2e_output_tps_per_user_p99": 112.64835775550863, | |
| "chunk_inter_token_latency_avg": 0.023320615398816368, | |
| "chunk_inter_token_latency_p50": 0.023428110182258776, | |
| "chunk_inter_token_latency_p90": 0.023711158726643578, | |
| "chunk_inter_token_latency_p99": 0.024126927350817735, | |
| "input_seq_len_avg": 78.0, | |
| "output_seq_len_avg": 1024.0, | |
| "output_seq_len_p50": 1024.0, | |
| "output_seq_len_p90": 1024.0, | |
| "output_seq_len_p99": 1024.0, | |
| "request_count": 9, | |
| "completed_request_count": 8, | |
| "request_samples": [ | |
| { | |
| "ttft": 0.1269236400257796, | |
| "time_to_second_token": 0.01867156900698319, | |
| "latency": 8.788493758998811, | |
| "inter_token_latency_avg": 0.008466832960872953, | |
| "chunk_inter_token_latency_avg": 0.022853747015759977, | |
| "input_tokens": 78, | |
| "output_tokens": 1024, | |
| "output_tps_per_user": 118.10791645721771, | |
| "e2e_output_tps_per_user": 116.51598420394788, | |
| "completed": true | |
| }, | |
| { | |
| "ttft": 8.624857554968912, | |
| "time_to_second_token": 0.02298608800629154, | |
| "latency": 16.714498371002264, | |
| "inter_token_latency_avg": 0.007907762283512563, | |
| "chunk_inter_token_latency_avg": 0.023448234249372035, | |
| "input_tokens": 78, | |
| "output_tokens": 1024, | |
| "output_tps_per_user": 126.45802493139794, | |
| "e2e_output_tps_per_user": 61.26417779767309, | |
| "completed": true | |
| }, | |
| { | |
| "ttft": 9.453244354983326, | |
| "time_to_second_token": 0.023422114027198404, | |
| "latency": 17.61734528996749, | |
| "inter_token_latency_avg": 0.007980548323542681, | |
| "chunk_inter_token_latency_avg": 0.023595667442150758, | |
| "input_tokens": 78, | |
| "output_tokens": 1024, | |
| "output_tps_per_user": 125.30467324531975, | |
| "e2e_output_tps_per_user": 58.12453483460616, | |
| "completed": true | |
| }, | |
| { | |
| "ttft": 8.935068062972277, | |
| "time_to_second_token": 0.023124370025470853, | |
| "latency": 17.158334736945108, | |
| "inter_token_latency_avg": 0.008038383845525738, | |
| "chunk_inter_token_latency_avg": 0.023428110182258776, | |
| "input_tokens": 78, | |
| "output_tokens": 1024, | |
| "output_tps_per_user": 124.4031162503657, | |
| "e2e_output_tps_per_user": 59.679451164636404, | |
| "completed": true | |
| }, | |
| { | |
| "ttft": 9.162113677011803, | |
| "time_to_second_token": 0.024539662001188844, | |
| "latency": 0.0, | |
| "inter_token_latency_avg": 0.006791496704820366, | |
| "chunk_inter_token_latency_avg": 0.024173123864614864, | |
| "input_tokens": 78, | |
| "output_tokens": 211, | |
| "output_tps_per_user": 147.24294856687996, | |
| "e2e_output_tps_per_user": 0.0, | |
| "completed": false | |
| }, | |
| { | |
| "ttft": 8.918001865968108, | |
| "time_to_second_token": 0.021153510024305433, | |
| "latency": 17.269235874002334, | |
| "inter_token_latency_avg": 0.00816347410365027, | |
| "chunk_inter_token_latency_avg": 0.02275540601644203, | |
| "input_tokens": 78, | |
| "output_tokens": 1024, | |
| "output_tps_per_user": 122.49686681223787, | |
| "e2e_output_tps_per_user": 59.296196280552444, | |
| "completed": true | |
| }, | |
| { | |
| "ttft": 8.376473198004533, | |
| "time_to_second_token": 0.023204773024190217, | |
| "latency": 17.563807698024902, | |
| "inter_token_latency_avg": 0.008980776637361066, | |
| "chunk_inter_token_latency_avg": 0.022968336250050923, | |
| "input_tokens": 78, | |
| "output_tokens": 1024, | |
| "output_tps_per_user": 111.34894457121723, | |
| "e2e_output_tps_per_user": 58.30170869584, | |
| "completed": true | |
| }, | |
| { | |
| "ttft": 8.430108908971306, | |
| "time_to_second_token": 0.02375636500073597, | |
| "latency": 17.079744989983737, | |
| "inter_token_latency_avg": 0.008455167234616258, | |
| "chunk_inter_token_latency_avg": 0.023189372871347, | |
| "input_tokens": 78, | |
| "output_tokens": 1024, | |
| "output_tps_per_user": 118.2708717937482, | |
| "e2e_output_tps_per_user": 59.95405672628693, | |
| "completed": true | |
| }, | |
| { | |
| "ttft": 8.486361830029637, | |
| "time_to_second_token": 0.024245411972515285, | |
| "latency": 17.406307295022998, | |
| "inter_token_latency_avg": 0.008719399281518438, | |
| "chunk_inter_token_latency_avg": 0.023473540697350952, | |
| "input_tokens": 78, | |
| "output_tokens": 1024, | |
| "output_tps_per_user": 114.68679982571635, | |
| "e2e_output_tps_per_user": 58.829249802615706, | |
| "completed": true | |
| } | |
| ], | |
| "total_tokens": 1159, | |
| "wall_time": 71.05695495504187, | |
| "num_completed": 2, | |
| "num_errors": 0, | |
| "server_gen_throughput": 115.83716967539169, | |
| "server_utilization": 0.6603773584905661, | |
| "server_spec_accept_rate": 0.24949221394719026, | |
| "server_spec_accept_length": 2.746445497630332, | |
| "server_spec_drafts": 422, | |
| "server_spec_draft_tokens": 2954, | |
| "server_spec_accepted_tokens": 737, | |
| "server_spec_pos_accept": [ | |
| 0.7322, | |
| 0.4692, | |
| 0.263, | |
| 0.154, | |
| 0.0782, | |
| 0.0332, | |
| 0.0166 | |
| ], | |
| "server_engine_steps": 422.0, | |
| "server_steps_per_s": 42.217507887418826, | |
| "server_accept_len_effective": 2.7464454976303316, | |
| "accept_norm_tps": 0.0, | |
| "accept_norm_ref_len": 0.0, | |
| "avg_running_reqs": 1, | |
| "max_running_reqs": 1, | |
| "effective_concurrency": 1, | |
| "avg_queue_reqs": 1, | |
| "max_queue_reqs": 1, | |
| "queue_fraction": 1.0, | |
| "underfilled": true, | |
| "warmup_timed_out": true, | |
| "warmup_duration": 60.872, | |
| "ready_reason": "warmup_timeout", | |
| "timeout_reason": "running_reqs=1/2", | |
| "capacity_limited": true, | |
| "hardware_summary": { | |
| "samples": 5, | |
| "duration_seconds": 9.555, | |
| "gpu_count": 4, | |
| "cpu_util_avg_pct": 7.12, | |
| "cpu_temp_max_c": 66.38, | |
| "gpu_util_avg_pct": 49.4, | |
| "gpu_util_max_pct": 99.0, | |
| "mem_util_avg_pct": 25.45, | |
| "mem_util_max_pct": 61.0, | |
| "temp_avg_c": 49.65, | |
| "temp_max_c": 67.0, | |
| "power_total_avg_w": 619.02, | |
| "power_total_max_w": 619.47, | |
| "power_limit_total_w": 1200.0, | |
| "vram_used_avg_mb": 189606.0, | |
| "vram_used_max_mb": 189606.0, | |
| "vram_total_mb": 391548.0, | |
| "vram_used_avg_pct": 48.42, | |
| "vram_used_max_pct": 48.42, | |
| "pcie_rx_avg_mb_s": 2422.0, | |
| "pcie_rx_max_mb_s": 2640.0, | |
| "pcie_tx_avg_mb_s": 2273.8, | |
| "pcie_tx_max_mb_s": 2384.0 | |
| } | |
| }, | |
| { | |
| "concurrency": 4, | |
| "context_tokens": 0, | |
| "benchmark_mode": "duration", | |
| "request_count_target": 0, | |
| "warmup_request_count": 0, | |
| "measurement_seconds": 9.993663, | |
| "measurement_wall_seconds": 10.000769, | |
| "client_output_tokens": 1189, | |
| "server_output_tokens": 1189, | |
| "aggregate_source": "openai_continuous_usage", | |
| "aggregate_tps": 118.97538984991404, | |
| "per_request_avg_tps": 29.74384746247851, | |
| "ttft_avg": 19.78742812310035, | |
| "ttft_p50": 24.75903701898642, | |
| "ttft_p90": 26.1186341956025, | |
| "ttft_p99": 26.15733930614777, | |
| "time_to_second_token_avg": 0.02317402323630328, | |
| "time_to_second_token_p50": 0.023638906015548855, | |
| "time_to_second_token_p90": 0.02446553821209818, | |
| "time_to_second_token_p99": 0.024968253751285373, | |
| "request_latency_avg": 27.448205057356972, | |
| "request_latency_p50": 33.088808445987524, | |
| "request_latency_p90": 33.99806104596937, | |
| "request_latency_p99": 34.4808259572566, | |
| "inter_token_latency_avg": 0.008139797620579165, | |
| "inter_token_latency_p50": 0.008184581259989297, | |
| "inter_token_latency_p90": 0.008731430765029264, | |
| "inter_token_latency_p99": 0.008955405133080832, | |
| "output_tps_per_user_avg": 123.39942798393562, | |
| "output_tps_per_user_p50": 122.18096054449921, | |
| "output_tps_per_user_p90": 134.3267922406224, | |
| "output_tps_per_user_p99": 138.45432623611825, | |
| "e2e_output_tps_per_user_avg": 46.01047960697323, | |
| "e2e_output_tps_per_user_p50": 30.947136506158543, | |
| "e2e_output_tps_per_user_p90": 77.39302160517252, | |
| "e2e_output_tps_per_user_p99": 111.22101751525767, | |
| "chunk_inter_token_latency_avg": 0.02353125415264528, | |
| "chunk_inter_token_latency_p50": 0.02348133677266808, | |
| "chunk_inter_token_latency_p90": 0.02384579067365582, | |
| "chunk_inter_token_latency_p99": 0.02408855324857224, | |
| "input_seq_len_avg": 78.0, | |
| "output_seq_len_avg": 1024.0, | |
| "output_seq_len_p50": 1024.0, | |
| "output_seq_len_p90": 1024.0, | |
| "output_seq_len_p99": 1024.0, | |
| "request_count": 9, | |
| "completed_request_count": 8, | |
| "request_samples": [ | |
| { | |
| "ttft": 0.1278693950152956, | |
| "time_to_second_token": 0.01804825698491186, | |
| "latency": 8.90592117496999, | |
| "inter_token_latency_avg": 0.008580695777081813, | |
| "chunk_inter_token_latency_avg": 0.0241155268680074, | |
| "input_tokens": 78, | |
| "output_tokens": 1024, | |
| "output_tps_per_user": 116.54066592955093, | |
| "e2e_output_tps_per_user": 114.97968372748937, | |
| "completed": true | |
| }, | |
| { | |
| "ttft": 25.77864516695263, | |
| "time_to_second_token": 0.02290963102132082, | |
| "latency": 33.76817299297545, | |
| "inter_token_latency_avg": 0.0078099001231894645, | |
| "chunk_inter_token_latency_avg": 0.023778356625067925, | |
| "input_tokens": 78, | |
| "output_tokens": 1024, | |
| "output_tps_per_user": 128.0426105617869, | |
| "e2e_output_tps_per_user": 30.32441228647506, | |
| "completed": true | |
| }, | |
| { | |
| "ttft": 25.00813842198113, | |
| "time_to_second_token": 0.023656431993003935, | |
| "latency": 0.0, | |
| "inter_token_latency_avg": 0.008246520938160926, | |
| "chunk_inter_token_latency_avg": 0.023019065640334093, | |
| "input_tokens": 78, | |
| "output_tokens": 389, | |
| "output_tps_per_user": 121.26325846969985, | |
| "e2e_output_tps_per_user": 0.0, | |
| "completed": false | |
| }, | |
| { | |
| "ttft": 9.027650000003632, | |
| "time_to_second_token": 0.023537531029433012, | |
| "latency": 16.70896882499801, | |
| "inter_token_latency_avg": 0.00750862055229167, | |
| "chunk_inter_token_latency_avg": 0.023347473632201757, | |
| "input_tokens": 78, | |
| "output_tokens": 1024, | |
| "output_tps_per_user": 133.18025501965136, | |
| "e2e_output_tps_per_user": 61.2844521241796, | |
| "completed": true | |
| }, | |
| { | |
| "ttft": 26.10788277600659, | |
| "time_to_second_token": 0.023854755039792508, | |
| "latency": 33.4722074510064, | |
| "inter_token_latency_avg": 0.007198753347995902, | |
| "chunk_inter_token_latency_avg": 0.023603604727563485, | |
| "input_tokens": 78, | |
| "output_tokens": 1024, | |
| "output_tps_per_user": 138.9129411245067, | |
| "e2e_output_tps_per_user": 30.592544620752573, | |
| "completed": true | |
| }, | |
| { | |
| "ttft": 16.83144874899881, | |
| "time_to_second_token": 0.02357069100253284, | |
| "latency": 26.018286619975697, | |
| "inter_token_latency_avg": 0.008980291173975452, | |
| "chunk_inter_token_latency_avg": 0.023316847388266213, | |
| "input_tokens": 78, | |
| "output_tokens": 1024, | |
| "output_tps_per_user": 111.35496395684392, | |
| "e2e_output_tps_per_user": 39.35693441142345, | |
| "completed": true | |
| }, | |
| { | |
| "ttft": 24.284541705972515, | |
| "time_to_second_token": 0.025024111033417284, | |
| "latency": 33.153149329009466, | |
| "inter_token_latency_avg": 0.008669215662792717, | |
| "chunk_inter_token_latency_avg": 0.02340002011355396, | |
| "input_tokens": 78, | |
| "output_tokens": 1024, | |
| "output_tps_per_user": 115.35068902390854, | |
| "e2e_output_tps_per_user": 30.886960084482403, | |
| "completed": true | |
| }, | |
| { | |
| "ttft": 26.161639873986132, | |
| "time_to_second_token": 0.023638906015548855, | |
| "latency": 34.53446650295518, | |
| "inter_token_latency_avg": 0.008184581259989297, | |
| "chunk_inter_token_latency_avg": 0.02371905560614462, | |
| "input_tokens": 78, | |
| "output_tokens": 1024, | |
| "output_tps_per_user": 122.18096054449921, | |
| "e2e_output_tps_per_user": 29.651536673148673, | |
| "completed": true | |
| }, | |
| { | |
| "ttft": 24.75903701898642, | |
| "time_to_second_token": 0.024325895006768405, | |
| "latency": 33.02446756296558, | |
| "inter_token_latency_avg": 0.008079599749735253, | |
| "chunk_inter_token_latency_avg": 0.02348133677266808, | |
| "input_tokens": 78, | |
| "output_tokens": 1024, | |
| "output_tps_per_user": 123.76850722497329, | |
| "e2e_output_tps_per_user": 31.007312927834686, | |
| "completed": true | |
| } | |
| ], | |
| "total_tokens": 1189, | |
| "wall_time": 71.4120287669939, | |
| "num_completed": 4, | |
| "num_errors": 0, | |
| "server_gen_throughput": 118.84137161719322, | |
| "server_utilization": 0.6603773584905661, | |
| "server_spec_accept_rate": 0.25553319919517103, | |
| "server_spec_accept_length": 2.788732394366197, | |
| "server_spec_drafts": 426, | |
| "server_spec_draft_tokens": 2982, | |
| "server_spec_accepted_tokens": 762, | |
| "server_spec_pos_accept": [ | |
| 0.73, | |
| 0.4695, | |
| 0.277, | |
| 0.1714, | |
| 0.0751, | |
| 0.0399, | |
| 0.0258 | |
| ], | |
| "server_engine_steps": 427.0, | |
| "server_steps_per_s": 42.72707440362766, | |
| "server_accept_len_effective": 2.7845433255269323, | |
| "accept_norm_tps": 0.0, | |
| "accept_norm_ref_len": 0.0, | |
| "avg_running_reqs": 1, | |
| "max_running_reqs": 1, | |
| "effective_concurrency": 1, | |
| "avg_queue_reqs": 3, | |
| "max_queue_reqs": 3, | |
| "queue_fraction": 1.0, | |
| "underfilled": true, | |
| "warmup_timed_out": true, | |
| "warmup_duration": 60.889, | |
| "ready_reason": "warmup_timeout", | |
| "timeout_reason": "running_reqs=1/4", | |
| "capacity_limited": true, | |
| "hardware_summary": { | |
| "samples": 4, | |
| "duration_seconds": 7.185, | |
| "gpu_count": 4, | |
| "cpu_util_avg_pct": 7.05, | |
| "cpu_temp_max_c": 67.12, | |
| "gpu_util_avg_pct": 49.5, | |
| "gpu_util_max_pct": 99.0, | |
| "mem_util_avg_pct": 26.0, | |
| "mem_util_max_pct": 59.0, | |
| "temp_avg_c": 51.38, | |
| "temp_max_c": 69.0, | |
| "power_total_avg_w": 619.54, | |
| "power_total_max_w": 619.93, | |
| "power_limit_total_w": 1200.0, | |
| "vram_used_avg_mb": 189606.0, | |
| "vram_used_max_mb": 189606.0, | |
| "vram_total_mb": 391548.0, | |
| "vram_used_avg_pct": 48.42, | |
| "vram_used_max_pct": 48.42, | |
| "pcie_rx_avg_mb_s": 2439.75, | |
| "pcie_rx_max_mb_s": 2649.0, | |
| "pcie_tx_avg_mb_s": 2278.25, | |
| "pcie_tx_max_mb_s": 2448.0 | |
| } | |
| }, | |
| { | |
| "concurrency": 2, | |
| "context_tokens": 65536, | |
| "benchmark_mode": "duration", | |
| "request_count_target": 0, | |
| "warmup_request_count": 0, | |
| "measurement_seconds": 9.984335, | |
| "measurement_wall_seconds": 10.000465, | |
| "client_output_tokens": 121, | |
| "server_output_tokens": 121, | |
| "aggregate_source": "openai_continuous_usage", | |
| "aggregate_tps": 12.118983865880503, | |
| "per_request_avg_tps": 6.059491932940252, | |
| "ttft_avg": 35.443730441189835, | |
| "ttft_p50": 40.205350022006314, | |
| "ttft_p90": 40.74598405077122, | |
| "ttft_p99": 41.04840777766425, | |
| "time_to_second_token_avg": 0.008084286202210933, | |
| "time_to_second_token_p50": 0.008268840028904378, | |
| "time_to_second_token_p90": 0.009591756807640194, | |
| "time_to_second_token_p99": 0.009597312689293177, | |
| "request_latency_avg": 42.43303290549375, | |
| "request_latency_p50": 48.06976508401567, | |
| "request_latency_p90": 48.95471742947702, | |
| "request_latency_p99": 49.295935525210226, | |
| "inter_token_latency_avg": 0.00817623119573374, | |
| "inter_token_latency_p50": 0.008107823757570827, | |
| "inter_token_latency_p90": 0.008595517486809978, | |
| "inter_token_latency_p99": 0.008858294394729843, | |
| "output_tps_per_user_avg": 122.5777308761553, | |
| "output_tps_per_user_p50": 123.33765877264314, | |
| "output_tps_per_user_p90": 127.79598606592414, | |
| "output_tps_per_user_p99": 129.85589972455693, | |
| "e2e_output_tps_per_user_avg": 26.393212866280486, | |
| "e2e_output_tps_per_user_p50": 21.302371631183085, | |
| "e2e_output_tps_per_user_p90": 35.93885086168265, | |
| "e2e_output_tps_per_user_p99": 41.58429651962935, | |
| "chunk_inter_token_latency_avg": 0.0248280156995747, | |
| "chunk_inter_token_latency_p50": 0.024617049442487427, | |
| "chunk_inter_token_latency_p90": 0.025319819759822115, | |
| "chunk_inter_token_latency_p99": 0.02538532188198598, | |
| "input_seq_len_avg": 64513.0, | |
| "output_seq_len_avg": 1024.0, | |
| "output_seq_len_p50": 1024.0, | |
| "output_seq_len_p90": 1024.0, | |
| "output_seq_len_p99": 1024.0, | |
| "request_count": 5, | |
| "completed_request_count": 4, | |
| "request_samples": [ | |
| { | |
| "ttft": 15.913573045982048, | |
| "time_to_second_token": 0.008091808995231986, | |
| "latency": 24.258752806985285, | |
| "inter_token_latency_avg": 0.008157555973610203, | |
| "chunk_inter_token_latency_avg": 0.024617049442487427, | |
| "input_tokens": 64513, | |
| "output_tokens": 1024, | |
| "output_tps_per_user": 122.58573563393406, | |
| "e2e_output_tps_per_user": 42.211568259401204, | |
| "completed": true | |
| }, | |
| { | |
| "ttft": 39.77577421802562, | |
| "time_to_second_token": 0.004880354972556233, | |
| "latency": 48.07007792202057, | |
| "inter_token_latency_avg": 0.008107823757570827, | |
| "chunk_inter_token_latency_avg": 0.025210649556215672, | |
| "input_tokens": 64513, | |
| "output_tokens": 1024, | |
| "output_tps_per_user": 123.33765877264314, | |
| "e2e_output_tps_per_user": 21.302232995360146, | |
| "completed": true | |
| }, | |
| { | |
| "ttft": 41.0820104139857, | |
| "time_to_second_token": 0.008268840028904378, | |
| "latency": 0.0, | |
| "inter_token_latency_avg": 0.008040989966927252, | |
| "chunk_inter_token_latency_avg": 0.025392599895559743, | |
| "input_tokens": 64513, | |
| "output_tokens": 121, | |
| "output_tps_per_user": 124.36279663486951, | |
| "e2e_output_tps_per_user": 0.0, | |
| "completed": false | |
| }, | |
| { | |
| "ttft": 40.205350022006314, | |
| "time_to_second_token": 0.009582497004885226, | |
| "latency": 48.06945224601077, | |
| "inter_token_latency_avg": 0.007687294451617258, | |
| "chunk_inter_token_latency_avg": 0.02434706570899212, | |
| "input_tokens": 64513, | |
| "output_tokens": 1024, | |
| "output_tps_per_user": 130.08477901996056, | |
| "e2e_output_tps_per_user": 21.30251026700602, | |
| "completed": true | |
| }, | |
| { | |
| "ttft": 40.241944505949505, | |
| "time_to_second_token": 0.00959793000947684, | |
| "latency": 49.33384864695836, | |
| "inter_token_latency_avg": 0.008887491828943161, | |
| "chunk_inter_token_latency_avg": 0.024572713894618525, | |
| "input_tokens": 64513, | |
| "output_tokens": 1024, | |
| "output_tps_per_user": 112.51768431936922, | |
| "e2e_output_tps_per_user": 20.75653994335457, | |
| "completed": true | |
| } | |
| ], | |
| "total_tokens": 121, | |
| "wall_time": 146.70948266400956, | |
| "num_completed": 2, | |
| "num_errors": 0, | |
| "server_gen_throughput": 12.093841253636528, | |
| "server_utilization": 0.7169811320754718, | |
| "server_spec_accept_rate": 0.3082706766917293, | |
| "server_spec_accept_length": 3.1578947368421053, | |
| "server_spec_drafts": 38, | |
| "server_spec_draft_tokens": 266, | |
| "server_spec_accepted_tokens": 82, | |
| "server_spec_pos_accept": [ | |
| 0.7368, | |
| 0.5526, | |
| 0.3684, | |
| 0.1842, | |
| 0.1316, | |
| 0.1053, | |
| 0.0789 | |
| ], | |
| "server_engine_steps": 39.0, | |
| "server_steps_per_s": 3.906118766688757, | |
| "server_accept_len_effective": 3.1025641025641026, | |
| "accept_norm_tps": 0.0, | |
| "accept_norm_ref_len": 0.0, | |
| "avg_running_reqs": 1, | |
| "max_running_reqs": 1, | |
| "effective_concurrency": 1, | |
| "avg_queue_reqs": 1, | |
| "max_queue_reqs": 1, | |
| "queue_fraction": 1.0, | |
| "underfilled": true, | |
| "warmup_timed_out": true, | |
| "warmup_duration": 120.609, | |
| "ready_reason": "warmup_timeout", | |
| "timeout_reason": "running_reqs=1/2", | |
| "capacity_limited": true, | |
| "hardware_summary": { | |
| "samples": 11, | |
| "duration_seconds": 23.837, | |
| "gpu_count": 4, | |
| "cpu_util_avg_pct": 7.09, | |
| "cpu_temp_max_c": 67.25, | |
| "gpu_util_avg_pct": 49.75, | |
| "gpu_util_max_pct": 100.0, | |
| "mem_util_avg_pct": 10.07, | |
| "mem_util_max_pct": 29.0, | |
| "temp_avg_c": 54.2, | |
| "temp_max_c": 74.0, | |
| "power_total_avg_w": 617.96, | |
| "power_total_max_w": 618.92, | |
| "power_limit_total_w": 1200.0, | |
| "vram_used_avg_mb": 189606.0, | |
| "vram_used_max_mb": 189606.0, | |
| "vram_total_mb": 391548.0, | |
| "vram_used_avg_pct": 48.42, | |
| "vram_used_max_pct": 48.42, | |
| "pcie_rx_avg_mb_s": 14405.09, | |
| "pcie_rx_max_mb_s": 22097.0, | |
| "pcie_tx_avg_mb_s": 16793.09, | |
| "pcie_tx_max_mb_s": 22792.0 | |
| } | |
| }, | |
| { | |
| "concurrency": 4, | |
| "context_tokens": 65536, | |
| "benchmark_mode": "duration", | |
| "request_count_target": 0, | |
| "warmup_request_count": 0, | |
| "measurement_seconds": 10.002486, | |
| "measurement_wall_seconds": 10.002486, | |
| "client_output_tokens": 0, | |
| "server_output_tokens": 0, | |
| "aggregate_source": "prometheus_fallback", | |
| "aggregate_tps": 0.06453558471444981, | |
| "per_request_avg_tps": 0.016133896178612453, | |
| "ttft_avg": 53.79123655226431, | |
| "ttft_p50": 53.92967229100759, | |
| "ttft_p90": 83.8944338750327, | |
| "ttft_p99": 90.48217335203546, | |
| "time_to_second_token_avg": 0.006969800742808729, | |
| "time_to_second_token_p50": 0.008464488462777808, | |
| "time_to_second_token_p90": 0.00998570021474734, | |
| "time_to_second_token_p99": 0.010018377252854406, | |
| "request_latency_avg": 62.69665365225228, | |
| "request_latency_p50": 62.90479161249823, | |
| "request_latency_p90": 92.55194512308226, | |
| "request_latency_p99": 99.28581670898885, | |
| "inter_token_latency_avg": 0.008705197556195476, | |
| "inter_token_latency_p50": 0.008637062442312125, | |
| "inter_token_latency_p90": 0.009213662504711036, | |
| "inter_token_latency_p99": 0.009430095063974691, | |
| "output_tps_per_user_avg": 115.22638877102585, | |
| "output_tps_per_user_p50": 115.78047725913596, | |
| "output_tps_per_user_p90": 121.29599886317772, | |
| "output_tps_per_user_p99": 123.34338172040091, | |
| "e2e_output_tps_per_user_avg": 21.279332857760934, | |
| "e2e_output_tps_per_user_p50": 16.913611610276735, | |
| "e2e_output_tps_per_user_p90": 34.79479489896622, | |
| "e2e_output_tps_per_user_p99": 40.42771152572509, | |
| "chunk_inter_token_latency_avg": 0.02482740656725064, | |
| "chunk_inter_token_latency_p50": 0.02478472501408651, | |
| "chunk_inter_token_latency_p90": 0.025181052141669647, | |
| "chunk_inter_token_latency_p99": 0.02530338106269595, | |
| "input_seq_len_avg": 64513.0, | |
| "output_seq_len_avg": 1024.0, | |
| "output_seq_len_p50": 1024.0, | |
| "output_seq_len_p90": 1024.0, | |
| "output_seq_len_p99": 1024.0, | |
| "request_count": 4, | |
| "completed_request_count": 4, | |
| "request_samples": [ | |
| { | |
| "ttft": 16.091457222006284, | |
| "time_to_second_token": 0.010022008034866303, | |
| "latency": 24.94300672103418, | |
| "inter_token_latency_avg": 0.008652541054768228, | |
| "chunk_inter_token_latency_avg": 0.024863903087157014, | |
| "input_tokens": 64513, | |
| "output_tokens": 1024, | |
| "output_tps_per_user": 115.57298528493219, | |
| "e2e_output_tps_per_user": 41.053591150920525, | |
| "completed": true | |
| }, | |
| { | |
| "ttft": 41.04423527698964, | |
| "time_to_second_token": 0.0009282180108129978, | |
| "latency": 50.71582369500538, | |
| "inter_token_latency_avg": 0.009454143126115097, | |
| "chunk_inter_token_latency_avg": 0.024423203075797335, | |
| "input_tokens": 64513, | |
| "output_tokens": 1024, | |
| "output_tps_per_user": 105.77373186129462, | |
| "e2e_output_tps_per_user": 20.190936977739472, | |
| "completed": true | |
| }, | |
| { | |
| "ttft": 66.81510930502554, | |
| "time_to_second_token": 0.00990098196780309, | |
| "latency": 75.09375952999108, | |
| "inter_token_latency_avg": 0.008092522214042552, | |
| "chunk_inter_token_latency_avg": 0.025316973165032206, | |
| "input_tokens": 64513, | |
| "output_tokens": 1024, | |
| "output_tps_per_user": 123.57086870453684, | |
| "e2e_output_tps_per_user": 13.636286242814, | |
| "completed": true | |
| }, | |
| { | |
| "ttft": 91.21414440503577, | |
| "time_to_second_token": 0.007027994957752526, | |
| "latency": 100.03402466297848, | |
| "inter_token_latency_avg": 0.008621583829856025, | |
| "chunk_inter_token_latency_avg": 0.024705546941016007, | |
| "input_tokens": 64513, | |
| "output_tokens": 1024, | |
| "output_tps_per_user": 115.98796923333974, | |
| "e2e_output_tps_per_user": 10.236517059569747, | |
| "completed": true | |
| } | |
| ], | |
| "total_tokens": 4096, | |
| "wall_time": 160.65691649296787, | |
| "num_completed": 4, | |
| "num_errors": 0, | |
| "server_gen_throughput": 0.06453558471444981, | |
| "server_utilization": 0.6981132075471699, | |
| "server_spec_accept_rate": 0.12315270935960591, | |
| "server_spec_accept_length": 1.8620689655172413, | |
| "server_spec_drafts": 0, | |
| "server_spec_draft_tokens": 0, | |
| "server_spec_accepted_tokens": 0, | |
| "server_spec_pos_accept": [], | |
| "server_engine_steps": 0.0, | |
| "server_steps_per_s": 0.0, | |
| "server_accept_len_effective": 0.0, | |
| "accept_norm_tps": 0.0, | |
| "accept_norm_ref_len": 0.0, | |
| "avg_running_reqs": 1, | |
| "max_running_reqs": 1, | |
| "effective_concurrency": 1, | |
| "avg_queue_reqs": 3, | |
| "max_queue_reqs": 3, | |
| "queue_fraction": 1.0, | |
| "underfilled": true, | |
| "warmup_timed_out": true, | |
| "warmup_duration": 120.589, | |
| "ready_reason": "warmup_timeout", | |
| "timeout_reason": "running_reqs=1/4", | |
| "capacity_limited": true, | |
| "hardware_summary": { | |
| "samples": 17, | |
| "duration_seconds": 38.153, | |
| "gpu_count": 4, | |
| "cpu_util_avg_pct": 7.11, | |
| "cpu_temp_max_c": 67.12, | |
| "gpu_util_avg_pct": 49.72, | |
| "gpu_util_max_pct": 100.0, | |
| "mem_util_avg_pct": 10.35, | |
| "mem_util_max_pct": 30.0, | |
| "temp_avg_c": 54.96, | |
| "temp_max_c": 75.0, | |
| "power_total_avg_w": 618.56, | |
| "power_total_max_w": 620.17, | |
| "power_limit_total_w": 1200.0, | |
| "vram_used_avg_mb": 189606.0, | |
| "vram_used_max_mb": 189606.0, | |
| "vram_total_mb": 391548.0, | |
| "vram_used_avg_pct": 48.42, | |
| "vram_used_max_pct": 48.42, | |
| "pcie_rx_avg_mb_s": 14517.06, | |
| "pcie_rx_max_mb_s": 23537.0, | |
| "pcie_tx_avg_mb_s": 17394.76, | |
| "pcie_tx_max_mb_s": 25466.0 | |
| } | |
| } | |
| ], | |
| "summary_table": { | |
| "0": { | |
| "1": 129.45196578500713, | |
| "2": 115.94808445857444, | |
| "4": 118.97538984991404 | |
| }, | |
| "65536": { | |
| "1": 122.23577506588123, | |
| "2": 12.118983865880503, | |
| "4": 0.06453558471444981 | |
| } | |
| }, | |
| "burst_results": [], | |
| "burst_summary_table": {}, | |
| "methodology": { | |
| "prefill": { | |
| "name": "Prefill", | |
| "present": true, | |
| "mode": "integrated_decode_scout", | |
| "formula": "prompt_tokens / TTFT", | |
| "notes": "Default mode records the required decode scout request for each non-zero decode context, so normal runs do not pay for a separate prefill phase. Standalone mode repeats cold-prefill samples. Prometheus prefill counters, when available and uncontaminated, are stored as validation." | |
| }, | |
| "sustained_decode": { | |
| "name": "Sustained Decode", | |
| "present": true, | |
| "formula": "OpenAI stream usage completion_tokens per measured window; client chunk fallback only when continuous usage is unavailable", | |
| "notes": "Duration-based steady-state cell after warmup. This is the main tuning/regression signal for kernels, NCCL, DCP, MTP, and scheduling. Prometheus metrics are stored as validation and scheduler state, not the default headline." | |
| }, | |
| "burst_e2e_decode": { | |
| "name": "Burst / E2E Decode", | |
| "present": false, | |
| "status": "not run; use --run-burst", | |
| "formula": "sum(completion_tokens) / profiling_wall_time", | |
| "notes": "Finite client-facing request burst using OpenAI stream usage. It includes request admission, scheduling, prefill/cache behavior, and completion." | |
| }, | |
| "acceptance_normalization": { | |
| "name": "Acceptance-normalized decode (MTP / speculative)", | |
| "present": true, | |
| "formula": "engine_steps = spec_drafts + max(0, output_tokens - (accepted_tokens + spec_drafts)); accept_len_effective = output_tokens / engine_steps; steps_per_s = aggregate_tps / accept_len_effective", | |
| "notes": "With speculative decoding tok/s = steps_per_s * accept_len, so raw tok/s mixes engine speed with data-dependent acceptance. steps_per_s (target-model forward passes per second) is the acceptance-independent speed used to compare runs; server_spec_pos_accept holds per-draft-position acceptance probabilities. Counters are vLLM window deltas; SGLang falls back to its lifetime accept-length gauge." | |
| } | |
| } | |
| } |