Instructions to use allenai/Bwen-8B-Stage1 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use allenai/Bwen-8B-Stage1 with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="allenai/Bwen-8B-Stage1", trust_remote_code=True)# pip install -U transformers accelerate # Load model directly from transformers import AutoModelForCausalLM model = AutoModelForCausalLM.from_pretrained("allenai/Bwen-8B-Stage1", trust_remote_code=True, device_map="auto") - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use allenai/Bwen-8B-Stage1 with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "allenai/Bwen-8B-Stage1" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "allenai/Bwen-8B-Stage1", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }'Use Docker
docker model run hf.co/allenai/Bwen-8B-Stage1
- SGLang
How to use allenai/Bwen-8B-Stage1 with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "allenai/Bwen-8B-Stage1" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "allenai/Bwen-8B-Stage1", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "allenai/Bwen-8B-Stage1" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "allenai/Bwen-8B-Stage1", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }' - Docker Model Runner
How to use allenai/Bwen-8B-Stage1 with Docker Model Runner:
docker model run hf.co/allenai/Bwen-8B-Stage1
Bwen 8B (Stage 1 training only)
Qwen3 8B Base retrofitted to operate over bytes instead of tokens via byteification through a short additional training procedure. This checkpoint contains Stage 1 training only - inner model parameters are unchanged.
See our technical report for details: https://allenai.org/papers/bolmo.
| Name | Model | Starting Point |
|---|---|---|
| Bolmo 1B | Bolmo-1B | Bolmo-1B-Stage1 |
| Bolmo 7B | Bolmo-7B | Bolmo-7B-Stage1 |
| Bwen 8B | Bwen-8B | Bwen-8B-Stage1 |
| Llama-B 8B | Llama-B-8B | Llama-B-8B-Stage1 |
| Bolmo 1B (Stage 1) | Bolmo-1B-Stage1 | OLMo-2-1B |
| Bolmo 7B (Stage 1) | Bolmo-7B-Stage1 | Olmo-3-7B |
| Bwen 8B (Stage 1) (you are here) | Bwen-8B-Stage1 | Qwen3-8B-Base |
| Llama-B 8B (Stage 1) | Llama-B-8B-Stage1 | Meta-Llama-3-8B |
Installation
This model was tested with transformers 4.57.3 and Python 3.11:
pip install transformers>=4.57.3
It additionally requires the xlstm package (which needs Python>=3.11):
pip install xlstm==2.0.4
Inference
You can use this model with the standard HuggingFace transformers library:
from transformers import AutoModelForCausalLM, AutoTokenizer
device = "cuda"
model = AutoModelForCausalLM.from_pretrained("allenai/Bwen-8B-Stage1", trust_remote_code=True).to(device)
tokenizer = AutoTokenizer.from_pretrained("allenai/Bwen-8B-Stage1", trust_remote_code=True)
message = ["Language modeling is "]
input_ids = tokenizer(message, return_tensors="pt")["input_ids"].to(device)
# `max_new_tokens` is the amount of bytes to generate
response = model.generate(input_ids, max_new_tokens=256, do_sample=True, temperature=0.1)
print(tokenizer.decode(response[0], skip_special_tokens=True))
Model Description
- Retrofitted from model: Qwen/Qwen3-8B-Base
- Developed by: Allen Institute for AI (Ai2)
- Model type: a byte-level autoregressive language model.
- Language(s) (NLP): English
- License: This model is licensed under Apache 2.0. It is intended for research and educational use in accordance with Ai2's Responsible Use Guidelines.
- Contact: Press:
press@allenai.org
Model Sources
- Data: https://huggingface.co/datasets/allenai/bolmo_mix
- Code: https://github.com/allenai/bolmo-core
- Paper: https://allenai.org/papers/bolmo
Bias, Risks, and Limitations
Like any base language model or fine-tuned model without safety filtering, these models can easily be prompted by users to generate harmful and sensitive content. Such content may also be produced unintentionally, especially in cases involving bias, so we recommend that users consider the risks when applying this technology. Additionally, many statements from Bolmo or any LLM are often inaccurate, so facts should be verified.
Citation
@misc{bolmo,
title={Bolmo: Byteifying the Next Generation of Language Models},
author={Benjamin Minixhofer and Tyler Murray and Tomasz Limisiewicz and Anna Korhonen and Luke Zettlemoyer and Noah A. Smith and Edoardo M. Ponti and Luca Soldaini and Valentin Hofmann},
year={2025},
eprint={2512.15586},
archivePrefix={arXiv},
primaryClass={cs.CL},
url={https://arxiv.org/abs/2512.15586},
}
- Downloads last month
- 42