--- license: cc-by-nc-4.0 base_model: Qwen/Qwen3.5-9B base_model_relation: finetune pipeline_tag: image-text-to-text library_name: transformers language: - en tags: - captcha - gui-agent - computer-use - agent - reinforcement-learning - grpo datasets: - ZHEN-04/CaptchaArena - ZHEN-04/CaptchaArena-Trajectories extra_gated_heading: "Access to CaptchaAgent" extra_gated_description: "CaptchaAgent is released for non-commercial academic research only (CC-BY-NC-4.0); commercial use is prohibited. Access requests are reviewed manually by the authors." extra_gated_prompt: "By requesting access you agree to use the model solely for non-commercial academic research; any commercial use is prohibited." extra_gated_fields: Full name: text Affiliation / Institution: text Email: text Intended research use: text I confirm I will use this model for non-commercial academic research only: checkbox I confirm I will NOT use this model for any commercial purpose: checkbox extra_gated_button_content: "Request access" --- # CaptchaAgent

A Large-Scale, Fine-Grained Dataset for Training Computer-Use Agents on Interactive CAPTCHAs

Zhenhao Zhang1,*, Zhaoyu Fan2, Haohan Ying3, Jingwen Hu3, Hancen Fan1,
Junhao Zhou4, Zitian Chen1, Linchao Zhu2,†
1Columbia University  ·  2Zhejiang University
3University of Rochester  ·  4University of Illinois at Urbana-Champaign
*Project lead  ·  †Corresponding author

arXiv Code
CaptchaArena Trajectories License: CC BY-NC 4.0

![The 20 CAPTCHA types](https://raw.githubusercontent.com/X0X0X00/CaptchaArena/main/assets/overview.jpg) ## Overview CaptchaAgent is a single 9B computer-use policy for all 20 CAPTCHA types in CaptchaArena, fine-tuned from [Qwen3.5-9B](https://huggingface.co/Qwen/Qwen3.5-9B). - **Input:** screenshots of a fixed 1280×1080 viewport. **Output:** `` reasoning plus one tool call per turn (`screenshot`, `click`, `type_text`, `drag`, `hold`), at most 15 turns. - **Training:** SFT on replay-verified, reasoning-annotated trajectories, then GRPO with the environment verifier as the reward. ## Checkpoints | Subfolder | Stage | CaptchaArena Pass@1 | Pass@5 | Open CaptchaWorld | Halligan | |---|---|--:|--:|--:|--:| | `SFT/` | SFT | 70.5 | 86.0 | 47.2 | 13.6 | | `RL/` | SFT + GRPO | **71.7** | **86.8** | **51.0** | **20.0** | Each subfolder holds the full bf16 weights with tokenizer, processor and chat template. CaptchaArena scores are on the 4,000-puzzle test split (5 rollouts per puzzle, unbiased Pass@k). For reference: untrained Qwen3.5-9B 11.4, best open-weight GUI agent 35.2, best closed-source model 69.2, human 94.1. ## Usage ```bash hf auth login hf download ZHEN-04/CaptchaAgent --include "RL/*" --local-dir CaptchaAgent # or "SFT/*" ``` ```python from transformers import AutoModelForImageTextToText, AutoProcessor model = AutoModelForImageTextToText.from_pretrained("ZHEN-04/CaptchaAgent", subfolder="RL", dtype="auto", device_map="auto") processor = AutoProcessor.from_pretrained("ZHEN-04/CaptchaAgent", subfolder="RL") ``` To evaluate, serve the checkpoint with an OpenAI-compatible server (SGLang or vLLM) and run the client in the [GitHub repo](https://github.com/X0X0X00/CaptchaArena) with `SFT_EVAL_NATIVE=1`, which parses the Qwen3 XML tool calls the model emits. ## Training - **SFT:** rank-64 LoRA on the language model (vision encoder frozen), 12,000 puzzles / 37,621 per-turn samples, 3 epochs. The LoRA is merged into the released weights. - **RL:** GRPO from the SFT checkpoint on 2,643 mined puzzles, full language-model weights, KL 0.005 to the SFT policy. Full configurations are in Appendices K and L of the paper. ## Intended use CaptchaAgent is released for research on computer-use agents and CAPTCHA robustness, not for bypassing protections on services you do not own. ## License Released under CC-BY-NC-4.0 (Creative Commons Attribution–NonCommercial 4.0). Non-commercial academic research use only — commercial use is prohibited. Access is gated: you must request access and agree to these terms before downloading. Please attribute when using this model. ## Citation ```bibtex @article{zhang2026captchaarena, title = {CaptchaArena: A Large-Scale, Fine-Grained Dataset for Training Computer-Use Agents on Interactive CAPTCHAs}, author = {Zhang, Zhenhao and Fan, Zhaoyu and Ying, Haohan and Hu, Jingwen and Fan, Hancen and Zhou, Junhao and Chen, Zitian and Zhu, Linchao}, journal = {arXiv preprint arXiv:2609.31957}, year = {2026} } ```