--- license: apache-2.0 base_model: incoai/Qwen3.8-27B-DFlash2 base_model_relation: quantized library_name: exllamav3 pipeline_tag: text-generation tags: - exl3 - exllamav3 - quantization - 4-bit - speculative-decoding - dflash2 --- # Qwen3.8-27B-DFlash2 — EXL3 4.00 bpw (draft) EXL3 (ExLlamaV3) quantization of the [incoai/Qwen3.8-27B-DFlash2](https://huggingface.co/incoai/Qwen3.8-27B-DFlash2) block-diffusion draft for use with [r0b0tlab/Qwen3.8-27B-EXL3-4.00bpw](https://huggingface.co/r0b0tlab/Qwen3.8-27B-EXL3-4.00bpw). Backbone Linears quantized at 4.00 bpw; the dynamic-convolution `base_kernel` tensors and the candidate-selector codebooks stay fp16 (uncalibrated). Quantization is acceptance-neutral on this stack (overlay measurement: mean acceptance 5.474 EXL3 vs 5.463 BF16 draft). Native engine GSM8K: **5.657** AL / **162.9 tok/s**. Requires [`r0b0tlab/exllamav3`](https://github.com/r0b0tlab/exllamav3) branch `community` @ `355c6ee` — native `DFlash2DraftModel`, selected automatically via `-dm