Meet the mission.
One agent, one fresh scenario. The referee checks observed actions and the mission’s actual effect. A claimed success alone earns nothing.
MEASURE. COMPARE. VERIFY.
Run an agent against a mission. Compare two strategies. Check whether a change helps.
One agent, one fresh scenario. The referee checks observed actions and the mission’s actual effect. A claimed success alone earns nothing.
Two instruction strategies use the same model, scenario, seed and budget. Each gets a separate range. Two paired rounds alternate who runs first.
Paired competition, not simultaneous shared-state combat.
A baseline and a verification-focused candidate face fresh seeds after their configurations are fixed. Compare objective results, scope violations and action counts.
Two rounds are preliminary evidence. This does not train model weights or automatically promote a candidate.
OPERATOR WORKBENCH
Checking the execution service…
Execution requires the private operator launch link. The public page never receives your identity, email, personal files or provider credentials. Only original synthetic mission data is sent to the model provider.
No trial requested.
One active suite. No automatic retry after an uncertain provider response. Critical scope violations make a result ineligible.
This record includes actions, observations and provider-reported usage. Missing usage remains unknown; cost is not inferred.
These original scenarios are small API state machines, not full enterprise networks. Only the red task currently varies its synthetic secret by seed; other scenarios repeat fixed task content. These tests do not establish performance on unseen task families. Results belong to the recorded model and instruction configuration. Scores cannot override a scope violation.
Want to inspect the earlier pipeline? Read the separately labeled scripted replays.