Dataset Viewer
Auto-converted to Parquet Duplicate
controller_revision
string
model_revision
string
fresh_model_random_seed
int64
comparison
string
selection_used_validation
bool
metrics
dict
holdout
dict
unseen_workflow_probes
dict
probe_note
string
limitations
list
0d4efdc608c73a53eeb24614b0021f6a8c46677dae2e3c29ace7c7f235fb865b
1d4bf0f2ff6012fd82039f2fa52739d0dd7c60c0
271,828
Fresh independent model decodings on a fixed, training-only selected bank and a balanced baseline; shared instances use the same fresh decodings in both arms.
false
{ "selected": { "scored": 192, "reward_mean": 0.578125, "reward_mean_cluster_bootstrap_95": [ 0.4791666666666667, 0.6770833333333334 ], "reward_target_error": 0.078125, "reward_balance_score": 0.84375, "reward_variance": 0.243896484375, "reward_histogram": { "1.0": 11...
{ "scored": 96, "reward_mean": 0.4270833333333333, "reward_mean_cluster_bootstrap_95": [ 0.25, 0.6041666666666666 ], "reward_target_error": 0.07291666666666669, "reward_balance_score": 0.8541666666666666, "reward_variance": 0.2446831597222222, "reward_histogram": { "1.0": 41, "0.0": 55 ...
{ "scored": 32, "reward_mean": 1, "reward_mean_cluster_bootstrap_95": [ 1, 1 ], "reward_target_error": 0.5, "reward_balance_score": 0, "reward_variance": 0, "reward_histogram": { "1.0": 32 }, "complete_same_instance_groups": 8, "mixed_same_instance_groups": 0, "observed_mixed_group_r...
Base-model evaluation on distinct motion-analysis and cash-ledger workflows combining skills across domains; excluded from training, task selection and the training baseline. This measures baseline capability, not learned transfer.
[ "Pilot uncertainty; no local policy training performed.", "Diversity proxies and base-model held-out performance do not establish transfer gain." ]

Reconcile Workflow Lab

A self-contained OpenEnv environment with 48 original synthetic tasks: six per domain, covering all three difficulty levels. Each frozen identifier selects the same case across fresh Arena containers and repeated rollouts. Selection happens offline using training-only base-model trials. The runtime needs no external controller, credentials, downloads or additional configuration.

Domain Agent work Verified outcome
Software engineering Inspect a small repository, repair code, run regression tests Executed repaired behavior on independent regression cases; fresh test result
Industrial and physical systems Diagnose sensor evidence, adjust controls, simulate Stable recovery while respecting physical and operating constraints
Natural science Implement a reproducible analysis or design controlled kinetics experiments Analysis agrees with records, or hypothesis predicts evidence after controls and repeated assays
Office and white-collar work Combine dependencies, calendars and resources; dispatch work Executed workflow satisfies resource, ordering and deadline constraints
Finance and economics Investigate revisions, currencies and exceptions; reconcile records Reconciliation and executed calculations agree with underlying transactions
Math and formal reasoning Transform equations or construct an exact feasible solution and dual certificate Executable exact checks establish the required mathematical claim
Cybersecurity Diagnose a deliberately flawed sandbox policy, repair it, rerun access checks Legitimate access works and prohibited access fails under regression checks
Media and content production Edit a cut list and captions, assemble and render an asset Rendered content, timing, format and technical constraints all pass

Agents use inspect, patch and task-specific run tools, then commit. Reward is 1 only when every behavioral requirement passes and the verified process result is fresh; otherwise 0. Existence of an artifact is insufficient. Modifying work after verification invalidates that verification. finish ends the episode with zero reward.

Learnability and diversity measurements

standalone-bank-selection.json reports training-only reward balance, same-case variation, format and budget failures, and a balanced sampling comparison. Four independent model responses per case distinguish variation within a case from variation between cases. Selection preserves a domain mean in [0.375, 0.625] when reachable, jointly targets the global mean closest to 0.5, then favors clean within-case reward-pair contribution per charged token (pairs involving format or budget failure contribute zero, while every charged token counts as cost). This is a sampling heuristic, not a measurement of gradient magnitude or training gain.

standalone-bank-validation.json, when present, reports fresh independent model decodings on the frozen selection and baseline. Independent structural holdout cases are reported separately and never determine selection. Equal domain quotas, all levels, skill coverage and observed workflow variation are diversity proxies. They do not establish transfer to Arena's private evaluation. No local GRPO gain or private-evaluation improvement is claimed.

Release

  • Image: ghcr.io/qingyuanwunothing/openenv-reconcile-lab@sha256:1c6bccf895a91bea68de8c47ce216340a1bdd03794931e70f97703bfb44a5176
  • Source: https://github.com/QingyuanWuNothing/openenv-reconcile-lab/tree/37bb03829d7e47b953e06086fcbc598b37b83735
  • Runtime and bank revision: 332d6340e931db752db95a7cbc8e1175cf626c3325b8db29c1622db63a0f9028
  • SDK: OpenEnv 86a180ede21e044f7929b9a7783ad83aa67d83a3
  • Calibration model: Qwen/Qwen3.8-27B, revision 1d4bf0f2ff6012fd82039f2fa52739d0dd7c60c0, thinking disabled, temperature 1.0
  • Episode limits: 32 actions, 8,192 generated plus post-reset observation tokens, 16,384 context tokens

The user authorized publication of runtime generator and verifier code. The trusted server owns protected files; restricted agent programs execute as UID/GID 10001 and cannot use files, imports or networking. The image excludes reference solutions, author policies, tests, grading fixtures, model traces and credentials. This dataset exports only task metadata, information visible through the agent's tools, provenance and aggregate measurements.

The environment uses simplified synthetic workflows. Actual cross-domain generalization is assessed by Arena after an official run; preparation checks are not evaluation scores.

Downloads last month
73