Instructions to use davidanugraha/Qwen3.5-35B-A3B-SWE-Smith-LoRA-Adapters with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- PEFT
How to use davidanugraha/Qwen3.5-35B-A3B-SWE-Smith-LoRA-Adapters with PEFT:
Task type is invalid.
- Notebooks
- Google Colab
- Kaggle
Qwen3.5-35B-A3B SWE-Smith LoRA
This repository contains the language-only LoRA checkpoints and reproducibility
record for verl-qwen35-35b-a3b-65k-mini-swe-rloo-20260904-r10. It does not contain the Qwen3.5-35B-A3B base-model
weights. Use the exact base model Qwen/Qwen3.5-35B-A3B at revision
59d61f3ce65a6d9863b86d2e96597125219dc754.
Adapter checkpoints
checkpoints/step-10checkpoints/step-20checkpoints/step-30checkpoints/step-40checkpoints/step-50checkpoints/step-60checkpoints/step-70โ best SWE-Bench Verified scorecheckpoints/step-80checkpoints/step-90checkpoints/step-100checkpoints/step-110checkpoints/step-120checkpoints/step-130checkpoints/step-140checkpoints/step-143โ final training checkpoint
Checkpoint 70 is the best evaluated adapter. Checkpoint 143 is
the terminal training adapter. Raw FSDP/model/optimizer shards remain in the
durable training store and are intentionally not duplicated here; their
per-checkpoint manifests are included under training/checkpoint-manifests/.
Training contract
- Algorithm: RLOO, one epoch, 143 optimizer steps
- LoRA: rank 32, alpha 64, routed-expert projections
- Batch: 32 prompts ร 8 rollouts
- Context: 65,536 tokens; maximum generated response 57,344 tokens
- Sampling: temperature 1.0, top-p 1.0, top-k -1
- Agent: Mini-SWE-Agent
- Base revision pinned above; veRL commit and container digest are in
training/run-identity.json.
Provenance
provenance/training-image/was copied directly from the immutable, digest-pinned container that ran training. It is the authoritative source-code and dependency provenance for the historical run.provenance/repository-at-export/contains the current launch, evaluation, serving, and release tooling as of export time. It is provided for operational convenience and is not asserted to be byte-identical to the training image.training/contains the original run identity, Hydra configuration, preflight evidence, every committed checkpoint manifest, and a frozen W&B export with all 143 history rows and the remote-file manifest.
SWE-Bench Verified
All arms used the same 500-instance cohort and compatibility hash
0446d1a5f3bf182f6de1b7c7d241fc550ccfedcd624478fe18851de057c7682c. Evaluation used Mini-SWE-Agent, temperature 1.0,
65,536-token context, 150 turns, 60-second command timeout, two-hour trajectory
timeout, and concurrency 128.
| Checkpoint | Resolved | Score | Agent timeouts | Infra errors | Truncated |
|---|---|---|---|---|---|
| 0 | 289/500 | 57.8% | 16 | 16 | 106 |
| 10 | 267/500 | 53.4% | 75 | 68 | 67 |
| 20 | 289/500 | 57.8% | 38 | 27 | 75 |
| 30 | 272/500 | 54.4% | 52 | 45 | 61 |
| 40 | 281/500 | 56.2% | 23 | 22 | 61 |
| 50 | 282/500 | 56.4% | 17 | 16 | 71 |
| 60 | 290/500 | 58.0% | 12 | 12 | 54 |
| 70 | 293/500 | 58.6% | 17 | 17 | 49 |
| 140 | 265/500 | 53.0% | 5 | 5 | 9 |
| 143 | 273/500 | 54.6% | 0 | 0 | 10 |
Checkpoint 70 is the best observed arm at 58.6%. The final checkpoint 143 scored 54.6%.
Loading an adapter
Download this repository and select a checkpoint subdirectory. For example, serve checkpoint 70 with vLLM while keeping the compiled/non-eager path and prefix caching enabled:
vllm serve Qwen/Qwen3.5-35B-A3B \
--revision 59d61f3ce65a6d9863b86d2e96597125219dc754 \
--served-model-name Qwen3.5-35B-A3B \
--tensor-parallel-size 4 \
--max-model-len 65536 \
--enable-prefix-caching \
--enable-lora --max-lora-rank 64 \
--lora-modules r10-step-70=/path/to/repo/checkpoints/step-70
Each checkpoint directory contains adapter_config.json,
adapter_model.safetensors, and export_provenance.json. See
evaluation/results.json, training/, provenance/, and RELEASE.json for
the complete release record.
- Downloads last month
- -