Datasets:
task_id string | family string | domain string | level int64 | stage_a string | stage_b string | stability string |
|---|---|---|---|---|---|---|
fin1-L1 | fin1 | finance-economics | 1 | 4/8 | 15/20 | 13/20 |
fin1-L2 | fin1 | finance-economics | 2 | 6/8 | 11/20 | 14/20 |
fin1-L3 | fin1 | finance-economics | 3 | 5/8 | 11/20 | 7/20 |
fin2-L1 | fin2 | finance-economics | 1 | 7/8 | 13/20 | 14/20 |
fin2-L3 | fin2 | finance-economics | 3 | 5/8 | 14/20 | 15/20 |
ind1-L2 | ind1 | industrial-physical-systems | 2 | 2/8 | 5/20 | 4/20 |
ind2-L1 | ind2 | industrial-physical-systems | 1 | 3/8 | 4/20 | 2/20 |
ind3-L1 | ind3 | industrial-physical-systems | 1 | 3/8 | 5/20 | 9/20 |
ind3-L2 | ind3 | industrial-physical-systems | 2 | 2/8 | 11/20 | 13/20 |
ind3-L3 | ind3 | industrial-physical-systems | 3 | 5/8 | 11/20 | 7/20 |
math2-L2 | math2 | mathematics-or-formal-reasoning | 2 | 3/8 | 9/20 | 8/20 |
math2-L3 | math2 | mathematics-or-formal-reasoning | 3 | 5/8 | 9/20 | 8/20 |
math3-L2 | math3 | mathematics-or-formal-reasoning | 2 | 1/8 | 4/20 | 2/20 |
med1-L1 | med1 | media-content-production | 1 | 3/8 | 11/20 | 14/20 |
med1-L2 | med1 | media-content-production | 2 | 6/8 | 9/20 | 16/20 |
med1-L3 | med1 | media-content-production | 3 | 6/8 | 11/20 | 14/20 |
med3-L1 | med3 | media-content-production | 1 | 3/8 | 5/20 | 8/20 |
off1-L1 | off1 | office-white-collar | 1 | 5/8 | 10/20 | 11/20 |
off1-L2 | off1 | office-white-collar | 2 | 3/8 | 14/20 | 12/20 |
off1-L3 | off1 | office-white-collar | 3 | 3/8 | 13/20 | 11/20 |
off2-L1 | off2 | office-white-collar | 1 | 4/8 | 8/20 | 9/20 |
off2-L2 | off2 | office-white-collar | 2 | 2/8 | 6/20 | 6/20 |
off2-L3 | off2 | office-white-collar | 3 | 3/8 | 6/20 | 11/20 |
sci1-L2 | sci1 | natural-science | 2 | 1/8 | 5/20 | 5/20 |
sci1-L3 | sci1 | natural-science | 3 | 1/8 | 6/20 | 4/20 |
sci3-L2 | sci3 | natural-science | 2 | 4/8 | 16/20 | 11/20 |
sci3-L3 | sci3 | natural-science | 3 | 5/8 | 9/20 | 14/20 |
sec1-L2 | sec1 | cybersecurity | 2 | 4/8 | 14/20 | 11/20 |
sec2-L2 | sec2 | cybersecurity | 2 | 6/8 | 11/20 | 12/20 |
sec2-L3 | sec2 | cybersecurity | 3 | 5/8 | 16/20 | 15/20 |
sec3-L1 | sec3 | cybersecurity | 1 | 5/8 | 8/20 | 9/20 |
sec3-L2 | sec3 | cybersecurity | 2 | 1/8 | 5/20 | 8/20 |
sec3-L3 | sec3 | cybersecurity | 3 | 3/8 | 7/20 | 12/20 |
sw2-L3 | sw2 | software-engineering | 3 | 7/8 | 14/20 | 18/20 |
ind1-L1 | ind1 | industrial-physical-systems | 1 | 1/8 | 3/19 | 5/20 |
math1-L3 | math1 | mathematics-or-formal-reasoning | 3 | 2/8 | 3/18 | 8/18 |
Workbench Gym tasks
36 train tasks for the Workbench Gym OpenEnv environment. Each task id is a family and a level, for example sw2-L3. Each reset draws a fresh instance.
Each episode gives a small workspace of text files. The agent reads them with ls, read, grep and py (stdlib Python on a copy of the files), then submits a JSON answer. Reward is 1.0 when every answer field is correct, and below 0.5 for some correct fields.
Domains
- cybersecurity
- finance-economics
- industrial-physical-systems
- mathematics-or-formal-reasoning
- media-content-production
- natural-science
- office-white-collar
- software-engineering
Calibration
Each task was played with Qwen/Qwen3.8-27B (thinking off) through the Hugging Face router. stage_a and stage_b count solves out of episodes. Tasks were kept in the mixed-outcome band.
Files
tasks.jsonl: task ids and calibration counts. No answers are included.
Source and reference solvers: https://github.com/NoeFlandre/openenv-arena/tree/workbench-gym-v1/workbench-gym
- Downloads last month
- -