Qwen3.5-35B-A3B SWE-Smith LoRA

This repository contains the language-only LoRA checkpoints and reproducibility record for verl-qwen35-35b-a3b-65k-mini-swe-rloo-20260904-r10. It does not contain the Qwen3.5-35B-A3B base-model weights. Use the exact base model Qwen/Qwen3.5-35B-A3B at revision 59d61f3ce65a6d9863b86d2e96597125219dc754.

Adapter checkpoints

  • checkpoints/step-10
  • checkpoints/step-20
  • checkpoints/step-30
  • checkpoints/step-40
  • checkpoints/step-50
  • checkpoints/step-60
  • checkpoints/step-70 โ€” best SWE-Bench Verified score
  • checkpoints/step-80
  • checkpoints/step-90
  • checkpoints/step-100
  • checkpoints/step-110
  • checkpoints/step-120
  • checkpoints/step-130
  • checkpoints/step-140
  • checkpoints/step-143 โ€” final training checkpoint

Checkpoint 70 is the best evaluated adapter. Checkpoint 143 is the terminal training adapter. Raw FSDP/model/optimizer shards remain in the durable training store and are intentionally not duplicated here; their per-checkpoint manifests are included under training/checkpoint-manifests/.

Training contract

  • Algorithm: RLOO, one epoch, 143 optimizer steps
  • LoRA: rank 32, alpha 64, routed-expert projections
  • Batch: 32 prompts ร— 8 rollouts
  • Context: 65,536 tokens; maximum generated response 57,344 tokens
  • Sampling: temperature 1.0, top-p 1.0, top-k -1
  • Agent: Mini-SWE-Agent
  • Base revision pinned above; veRL commit and container digest are in training/run-identity.json.

Provenance

  • provenance/training-image/ was copied directly from the immutable, digest-pinned container that ran training. It is the authoritative source-code and dependency provenance for the historical run.
  • provenance/repository-at-export/ contains the current launch, evaluation, serving, and release tooling as of export time. It is provided for operational convenience and is not asserted to be byte-identical to the training image.
  • training/ contains the original run identity, Hydra configuration, preflight evidence, every committed checkpoint manifest, and a frozen W&B export with all 143 history rows and the remote-file manifest.

SWE-Bench Verified

All arms used the same 500-instance cohort and compatibility hash 0446d1a5f3bf182f6de1b7c7d241fc550ccfedcd624478fe18851de057c7682c. Evaluation used Mini-SWE-Agent, temperature 1.0, 65,536-token context, 150 turns, 60-second command timeout, two-hour trajectory timeout, and concurrency 128.

Checkpoint Resolved Score Agent timeouts Infra errors Truncated
0 289/500 57.8% 16 16 106
10 267/500 53.4% 75 68 67
20 289/500 57.8% 38 27 75
30 272/500 54.4% 52 45 61
40 281/500 56.2% 23 22 61
50 282/500 56.4% 17 16 71
60 290/500 58.0% 12 12 54
70 293/500 58.6% 17 17 49
140 265/500 53.0% 5 5 9
143 273/500 54.6% 0 0 10

Checkpoint 70 is the best observed arm at 58.6%. The final checkpoint 143 scored 54.6%.

Loading an adapter

Download this repository and select a checkpoint subdirectory. For example, serve checkpoint 70 with vLLM while keeping the compiled/non-eager path and prefix caching enabled:

vllm serve Qwen/Qwen3.5-35B-A3B \
  --revision 59d61f3ce65a6d9863b86d2e96597125219dc754 \
  --served-model-name Qwen3.5-35B-A3B \
  --tensor-parallel-size 4 \
  --max-model-len 65536 \
  --enable-prefix-caching \
  --enable-lora --max-lora-rank 64 \
  --lora-modules r10-step-70=/path/to/repo/checkpoints/step-70

Each checkpoint directory contains adapter_config.json, adapter_model.safetensors, and export_provenance.json. See evaluation/results.json, training/, provenance/, and RELEASE.json for the complete release record.

Downloads last month
-
Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support

Model tree for davidanugraha/Qwen3.5-35B-A3B-SWE-Smith-LoRA-Adapters

Adapter
(45)
this model