{"task_id": "downtime-easy", "split": "train", "family": "downtime", "level": "easy", "domain": "industrial-physical-systems", "output_file": "output/downtime.json", "checks": 7, "qwen3.8_27b_full_score": "7/8", "example_instruction": "# Line downtime report\n\nPlant engineering needs the unplanned downtime of production line A for the evening shift of 2026-09-20.\n\n- The plant documentation is in /app/handbook, including amendments. Use it for the definition of unplanned downtime, the per-shift downtime limit that applies on that date, and any rules about which stops count.\n- Sensor data is in /app/data/sensor_log.csv: one row per machine per minute; `state` is the machine state during that minute.\n- Shift times are in /app/data/shifts.csv.\n\nWrite /app/output/downtime.json with exactly this structure:\n\n{\n \"limit_min\": ,\n \"downtime_by_machine\": {\"M01\": , ...},\n \"over_limit\": []\n}\n\nInclude every machine that appears in the sensor log. Do not guess values: derive everything from the files.\n", "example_files": ["data/machines.json", "data/sensor_log.csv", "data/shifts.csv", "handbook/amendments/amendment-101_2026-08-13.md", "handbook/maintenance_manual.md"]} {"task_id": "downtime-medium", "split": "train", "family": "downtime", "level": "medium", "domain": "industrial-physical-systems", "output_file": "output/downtime.json", "checks": 7, "qwen3.8_27b_full_score": "5/8", "example_instruction": "# Line downtime report\n\nPlant engineering needs the unplanned downtime of production line A for the evening shift of 2026-09-21.\n\n- The plant documentation is in /app/handbook, including amendments. Use it for the definition of unplanned downtime, the per-shift downtime limit that applies on that date, and any rules about which stops count.\n- Sensor data is in /app/data/sensor_log.csv: one row per machine per minute; `state` is the machine state during that minute.\n- Shift times are in /app/data/shifts.csv.\n\nWrite /app/output/downtime.json with exactly this structure:\n\n{\n \"limit_min\": ,\n \"downtime_by_machine\": {\"M01\": , ...},\n \"over_limit\": []\n}\n\nInclude every machine that appears in the sensor log. Do not guess values: derive everything from the files.\n", "example_files": ["data/machines.json", "data/sensor_log.csv", "data/shifts.csv", "handbook/amendments/amendment-101_2026-04-17.md", "handbook/amendments/amendment-102_2026-06-19.md", "handbook/amendments/amendment-103_2026-08-13.md", "handbook/maintenance_manual.md"]} {"task_id": "downtime-hard", "split": "train", "family": "downtime", "level": "hard", "domain": "industrial-physical-systems", "output_file": "output/downtime.json", "checks": 7, "qwen3.8_27b_full_score": "4/8", "example_instruction": "# Line downtime report\n\nPlant engineering needs the unplanned downtime of production line A for the evening shift of 2026-09-21.\n\n- The plant documentation is in /app/handbook, including amendments. Use it for the definition of unplanned downtime, the per-shift downtime limit that applies on that date, and any rules about which stops count.\n- Sensor data is in /app/data/sensor_log.csv: one row per machine per minute; `state` is the machine state during that minute.\n- Shift times are in /app/data/shifts.csv.\n\nWrite /app/output/downtime.json with exactly this structure:\n\n{\n \"limit_min\": ,\n \"downtime_by_machine\": {\"M01\": , ...},\n \"over_limit\": []\n}\n\nInclude every machine that appears in the sensor log. Do not guess values: derive everything from the files.\n", "example_files": ["data/machines.json", "data/sensor_log.csv", "data/shifts.csv", "handbook/amendments/amendment-101_2026-02-24.md", "handbook/amendments/amendment-102_2026-02-28.md", "handbook/amendments/amendment-103_2026-03-10.md", "handbook/amendments/amendment-104_2026-03-16.md", "handbook/amendments/amendment-105_2026-06-17.md", "handbook/amendments/amendment-106_2026-08-16.md", "handbook/amendments/amendment-107_2026-11-05.md", "handbook/maintenance_manual.md"]} {"task_id": "expenses-easy", "split": "train", "family": "expenses", "level": "easy", "domain": "office-white-collar / finance-economics", "output_file": "output/expense_report.json", "checks": 8, "qwen3.8_27b_full_score": "4/8", "example_instruction": "# Expense audit\n\nFinance needs an audit of the Research department's expense claims dated in September 2026 (2026-09-01 to 2026-09-30).\n\n- Claims are in /app/data/claims.csv; employees and departments in /app/data/employees.csv.\n- The travel and expense policy is in /app/policy, including amendments. The finance team's folder /app/finance has the daily FX rates and their notes.\n- Apply the policy exactly as written: city tiers, caps, receipt and flight rules, currency conversion and rounding, and how amendments apply.\n\nWrite /app/output/expense_report.json with exactly this structure:\n\n{\n \"violations\": [],\n \"by_employee\": {\"E01\": , ...},\n \"reimbursable_eur\": \n}\n\nInclude in by_employee every employee of the department who has at least one claim in that month.\n", "example_files": ["data/claims.csv", "data/employees.csv", "finance/README.md", "finance/fx_corrections.md", "finance/fx_rates.csv", "policy/amendments/amendment-21_2026-09-02.md", "policy/travel_policy.md"]} {"task_id": "expenses-medium", "split": "train", "family": "expenses", "level": "medium", "domain": "office-white-collar / finance-economics", "output_file": "output/expense_report.json", "checks": 8, "qwen3.8_27b_full_score": "4/8", "example_instruction": "# Expense audit\n\nFinance needs an audit of the Finance department's expense claims dated in September 2026 (2026-09-01 to 2026-09-30).\n\n- Claims are in /app/data/claims.csv; employees and departments in /app/data/employees.csv.\n- The travel and expense policy is in /app/policy, including amendments. The finance team's folder /app/finance has the daily FX rates and their notes.\n- Apply the policy exactly as written: city tiers, caps, receipt and flight rules, currency conversion and rounding, and how amendments apply.\n\nWrite /app/output/expense_report.json with exactly this structure:\n\n{\n \"violations\": [],\n \"by_employee\": {\"E01\": , ...},\n \"reimbursable_eur\": \n}\n\nInclude in by_employee every employee of the department who has at least one claim in that month.\n", "example_files": ["data/claims.csv", "data/employees.csv", "finance/README.md", "finance/fx_corrections.md", "finance/fx_rates.csv", "policy/amendments/amendment-21_2026-08-15.md", "policy/amendments/amendment-22_2026-09-13.md", "policy/amendments/amendment-23_2026-09-15.md", "policy/travel_policy.md"]} {"task_id": "expenses-hard", "split": "train", "family": "expenses", "level": "hard", "domain": "office-white-collar / finance-economics", "output_file": "output/expense_report.json", "checks": 8, "qwen3.8_27b_full_score": "3/8", "example_instruction": "# Expense audit\n\nFinance needs an audit of the Finance department's expense claims dated in September 2026 (2026-09-01 to 2026-09-30).\n\n- Claims are in /app/data/claims.csv; employees and departments in /app/data/employees.csv.\n- The travel and expense policy is in /app/policy, including amendments. The finance team's folder /app/finance has the daily FX rates and their notes.\n- Apply the policy exactly as written: city tiers, caps, receipt and flight rules, currency conversion and rounding, and how amendments apply.\n\nWrite /app/output/expense_report.json with exactly this structure:\n\n{\n \"violations\": [],\n \"by_employee\": {\"E01\": , ...},\n \"reimbursable_eur\": \n}\n\nInclude in by_employee every employee of the department who has at least one claim in that month.\n", "example_files": ["data/claims.csv", "data/employees.csv", "finance/README.md", "finance/fx_corrections.md", "finance/fx_rates.csv", "policy/amendments/amendment-21_2026-07-13.md", "policy/amendments/amendment-22_2026-07-18.md", "policy/amendments/amendment-23_2026-09-14.md", "policy/amendments/amendment-24_2026-09-16.md", "policy/amendments/amendment-25_2026-09-18.md", "policy/amendments/amendment-26_2026-09-21.md", "policy/amendments/amendment-27_2026-11-19.md", "policy/travel_policy.md"]} {"task_id": "routing-easy", "split": "train", "family": "routing", "level": "easy", "domain": "mathematics-or-formal-reasoning", "output_file": "output/plan.json", "checks": 6, "qwen3.8_27b_full_score": "5/8", "example_instruction": "# Depot delivery plan\n\nPlan today's deliveries for two vans, V1 and V2.\n\n- Orders are in /app/data/orders.csv; the depot, vans, speed and service time are in /app/data/vans.json.\n- Read /app/data/ops_notes.md: every note there is a binding constraint.\n- Distances are great-circle distances on a sphere of radius 6371 km between the coordinates. Travel time = distance / speed; each stop takes the service time.\n- Respect van capacity, and every note. Each remaining order is served exactly once.\n- Minimise the total distance driven by both vans, including the return to the depot.\n\nWrite /app/output/plan.json with exactly this structure:\n\n{\n \"routes\": {\"V1\": [], \"V2\": []},\n \"total_km\": \n}\n\nA plan is accepted if it is feasible and its total distance is within 5% of the best possible plan.\n", "example_files": ["data/ops_notes.md", "data/orders.csv", "data/vans.json"]} {"task_id": "routing-medium", "split": "train", "family": "routing", "level": "medium", "domain": "mathematics-or-formal-reasoning", "output_file": "output/plan.json", "checks": 6, "qwen3.8_27b_full_score": "3/8", "example_instruction": "# Depot delivery plan\n\nPlan today's deliveries for two vans, V1 and V2.\n\n- Orders are in /app/data/orders.csv; the depot, vans, speed and service time are in /app/data/vans.json.\n- Read /app/data/ops_notes.md: every note there is a binding constraint.\n- Distances are great-circle distances on a sphere of radius 6371 km between the coordinates. Travel time = distance / speed; each stop takes the service time.\n- Respect van capacity, time windows (a van may wait for a window to open), and every note. Each remaining order is served exactly once.\n- Minimise the total distance driven by both vans, including the return to the depot.\n\nWrite /app/output/plan.json with exactly this structure:\n\n{\n \"routes\": {\"V1\": [], \"V2\": []},\n \"total_km\": \n}\n\nA plan is accepted if it is feasible and its total distance is within 5% of the best possible plan.\n", "example_files": ["data/ops_notes.md", "data/orders.csv", "data/vans.json"]} {"task_id": "routing-hard", "split": "train", "family": "routing", "level": "hard", "domain": "mathematics-or-formal-reasoning", "output_file": "output/plan.json", "checks": 6, "qwen3.8_27b_full_score": "4/8", "example_instruction": "# Depot delivery plan\n\nPlan today's deliveries for two vans, V1 and V2.\n\n- Orders are in /app/data/orders.csv; the depot, vans, speed and service time are in /app/data/vans.json.\n- Read /app/data/ops_notes.md: every note there is a binding constraint.\n- Distances are great-circle distances on a sphere of radius 6371 km between the coordinates. Travel time = distance / speed; each stop takes the service time.\n- Respect van capacity, time windows (a van may wait for a window to open), and every note. Each remaining order is served exactly once.\n- Minimise the total distance driven by both vans, including the return to the depot.\n\nWrite /app/output/plan.json with exactly this structure:\n\n{\n \"routes\": {\"V1\": [], \"V2\": []},\n \"total_km\": \n}\n\nA plan is accepted if it is feasible and its total distance is within 5% of the best possible plan.\n", "example_files": ["data/ops_notes.md", "data/orders.csv", "data/vans.json"]}