Dataset Viewer
Auto-converted to Parquet Duplicate
task_id
string
family
string
domain
string
level
int64
stage_a
string
stage_b
string
stability
string
fin1-L1
fin1
finance-economics
1
4/8
15/20
13/20
fin1-L2
fin1
finance-economics
2
6/8
11/20
14/20
fin1-L3
fin1
finance-economics
3
5/8
11/20
7/20
fin2-L1
fin2
finance-economics
1
7/8
13/20
14/20
fin2-L3
fin2
finance-economics
3
5/8
14/20
15/20
ind1-L2
ind1
industrial-physical-systems
2
2/8
5/20
4/20
ind2-L1
ind2
industrial-physical-systems
1
3/8
4/20
2/20
ind3-L1
ind3
industrial-physical-systems
1
3/8
5/20
9/20
ind3-L2
ind3
industrial-physical-systems
2
2/8
11/20
13/20
ind3-L3
ind3
industrial-physical-systems
3
5/8
11/20
7/20
math2-L2
math2
mathematics-or-formal-reasoning
2
3/8
9/20
8/20
math2-L3
math2
mathematics-or-formal-reasoning
3
5/8
9/20
8/20
math3-L2
math3
mathematics-or-formal-reasoning
2
1/8
4/20
2/20
med1-L1
med1
media-content-production
1
3/8
11/20
14/20
med1-L2
med1
media-content-production
2
6/8
9/20
16/20
med1-L3
med1
media-content-production
3
6/8
11/20
14/20
med3-L1
med3
media-content-production
1
3/8
5/20
8/20
off1-L1
off1
office-white-collar
1
5/8
10/20
11/20
off1-L2
off1
office-white-collar
2
3/8
14/20
12/20
off1-L3
off1
office-white-collar
3
3/8
13/20
11/20
off2-L1
off2
office-white-collar
1
4/8
8/20
9/20
off2-L2
off2
office-white-collar
2
2/8
6/20
6/20
off2-L3
off2
office-white-collar
3
3/8
6/20
11/20
sci1-L2
sci1
natural-science
2
1/8
5/20
5/20
sci1-L3
sci1
natural-science
3
1/8
6/20
4/20
sci3-L2
sci3
natural-science
2
4/8
16/20
11/20
sci3-L3
sci3
natural-science
3
5/8
9/20
14/20
sec1-L2
sec1
cybersecurity
2
4/8
14/20
11/20
sec2-L2
sec2
cybersecurity
2
6/8
11/20
12/20
sec2-L3
sec2
cybersecurity
3
5/8
16/20
15/20
sec3-L1
sec3
cybersecurity
1
5/8
8/20
9/20
sec3-L2
sec3
cybersecurity
2
1/8
5/20
8/20
sec3-L3
sec3
cybersecurity
3
3/8
7/20
12/20
sw2-L3
sw2
software-engineering
3
7/8
14/20
18/20
ind1-L1
ind1
industrial-physical-systems
1
1/8
3/19
5/20
math1-L3
math1
mathematics-or-formal-reasoning
3
2/8
3/18
8/18

Workbench Gym tasks

36 train tasks for the Workbench Gym OpenEnv environment. Each task id is a family and a level, for example sw2-L3. Each reset draws a fresh instance.

Each episode gives a small workspace of text files. The agent reads them with ls, read, grep and py (stdlib Python on a copy of the files), then submits a JSON answer. Reward is 1.0 when every answer field is correct, and below 0.5 for some correct fields.

Domains

  • cybersecurity
  • finance-economics
  • industrial-physical-systems
  • mathematics-or-formal-reasoning
  • media-content-production
  • natural-science
  • office-white-collar
  • software-engineering

Calibration

Each task was played with Qwen/Qwen3.8-27B (thinking off) through the Hugging Face router. stage_a and stage_b count solves out of episodes. Tasks were kept in the mixed-outcome band.

Files

  • tasks.jsonl: task ids and calibration counts. No answers are included.

Source and reference solvers: https://github.com/NoeFlandre/openenv-arena/tree/workbench-gym-v1/workbench-gym

Downloads last month
-