Temper-1-0.5B

Temper-1-0.5B is a small model fine-tuned from Qwen2.5-Coder-0.5B-Instruct for writing Rust. On the first attempt it solves more Rust tasks than models up to three times its size.

The model was trained on 10 000 Rust examples written by DeepSeek-V4-Flash and checked by the compiler and tests, then trained further with reinforcement learning on its own answers.

Website: temper-ai.pages.dev

Results

Each cell shows pass@1 / pass@10, %. pass@1 is the share of tasks solved on the first attempt, pass@10 the share solved in at least one of 10 attempts.

Temper-1-0.5B ThoxEdge-RustCoder-0.5B rust-mentor-0.6b Qwen2.5-Coder-1.5B Qwen3.5-0.8B Qwen2.5-Coder-0.5B Qwen2.5-Coder-3B (reference)
trivial¹ 81.8 / 94.0 55.8 / 88.0 35.4 / 68.0 70.0 / 94.0 45.2 / 84.0 45.6 / 88.0 77.0 / 96.0
easy¹ 59.0 / 76.7 13.0 / 50.0 13.7 / 30.0 42.7 / 73.3 13.3 / 36.7 15.0 / 50.0 51.0 / 83.3
medium¹ 18.0 / 35.0 1.5 / 10.0 0.0 / 0.0 11.0 / 45.0 3.0 / 5.0 1.0 / 5.0 25.0 / 50.0
HumanEval-RS 34.7 / 57.1 13.3 / 41.7 4.6 / 16.7 31.0 / 69.2 6.9 / 20.5 16.9 / 42.3 50.6 / 87.2
MBPP-RS 38.8 / 54.8 12.8 / 44.4 9.7 / 27.1 34.4 / 64.7 11.6 / 33.3 23.4 / 50.8 41.3 / 72.9

Bold marks the best pass@1 among models up to 1.5B. Qwen2.5-Coder-3B is shown for reference only and is not part of this comparison, it is ahead of Temper on medium, HumanEval-RS and MBPP-RS. All Qwen models are the Instruct versions. Settings are temperature 0.8, top_p 0.95 and 10 answers per task, with pass@k from the unbiased estimator in the Codex paper. Every model ran without a system prompt and without reasoning.

¹ Internal set of Rust tasks.

HumanEval-RS and MBPP-RS come from MultiPL-E at revision 28441b6. The task goes to the model as a chat message and the first closed code block of the answer is compiled with the tests. Because of this the numbers differ from leaderboards that score code completion. mbpp_67_bell_number has no correct answer and counts as failed for every model.

Usage

from transformers import AutoModelForCausalLM, AutoTokenizer

name = "temper-ai/Temper-1-0.5B"
tokenizer = AutoTokenizer.from_pretrained(name)
model = AutoModelForCausalLM.from_pretrained(name, dtype="auto", device_map="auto")

messages = [{"role": "user", "content": "Write a Rust function that returns the n-th Fibonacci number."}]
inputs = tokenizer.apply_chat_template(messages, add_generation_prompt=True, return_dict=True, return_tensors="pt").to(model.device)
out = model.generate(**inputs, max_new_tokens=512)
print(tokenizer.decode(out[0][inputs["input_ids"].shape[1]:], skip_special_tokens=True))

The default generation config samples with temperature 0.7, top_p 0.8 and top_k 20. The results above used temperature 0.8 and top_p 0.95.

A GGUF version for llama.cpp, LM Studio and Ollama is in temper-ai/Temper-1-0.5B-GGUF.

Limitations

With ten attempts Qwen2.5-Coder-1.5B solves more tasks on medium, HumanEval-RS and MBPP-RS. Qwen2.5-Coder-3B remains stronger on medium tasks and both MultiPL-E sets. The trivial, easy and medium sets are written in the same style as the Temper training data, so only HumanEval-RS and MBPP-RS give an independent comparison with other models.

The model was trained only on Rust tasks with requests in English. Requests in other languages and questions unrelated to Rust may get made-up answers or refusals. It was not trained for fill-in-the-middle and expects chat requests, so it is not a drop-in model for inline completion in an editor. Tests that the model writes itself were not checked during evaluation and may fail to compile. The training data has few tasks on async code and web frameworks.

License

Apache 2.0, same as Qwen2.5-Coder-0.5B-Instruct.

Downloads last month
229
Safetensors
Model size
0.5B params
Tensor type
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for temper-ai/Temper-1-0.5B

Finetuned
(111)
this model
Quantizations
2 models

Paper for temper-ai/Temper-1-0.5B