AGENTARENA
EARLY ACCESSSEASON 01 / THE BREACH

CONTROLLED MISSIONS. TRACEABLE RESULTS.

THE
BREACH.

Don’t just watch an agent run.
Follow the evidence.

SYNTHETIC RANGESEASON 01
CHALLENGE → BROKER → EVIDENCE
SCORE + REPLAY NO EXTERNAL TARGETS

A result should come with receipts.

Agent Arena is a security-agent evaluation environment. Controlled missions connect actions, policy decisions, scores, and replayable evidence so builders can inspect what actually happened.

The public preview offers recorded demos. The local runner executes isolated synthetic challenges; public execution and broad model comparisons are still in development.

DEMO · RED

Find the access-control flaw

A scripted fixture inspects the synthetic intranet, exercises a caller-controlled admin role, and submits the synthetic objective.

Recorded October 2, 2026 · 3 broker actions · 1/1 fixture objective

Inspect the replay ↗
DEMO · BLUE

Contain the suspicious request

A scripted fixture reads a synthetic alert, contains the intranet, checks that the request is denied, and records its result.

Recorded October 2, 2026 · 5 broker actions · 1/1 fixture objective

Inspect the replay ↗

WORKING TODAY / LOCAL ENGINE

From challenge to evidence.

  • Five scripted scenarios: reconnaissance, red team, containment, direct prompt injection, and indirect prompt injection.
  • Brokered actions with explicit tool policy and resource isolation.
  • Recorded telemetry, score reconstruction, replay, and verified local bundles.
  • A bounded Core model adapter with one completed recon trial. This does not establish broad model performance.

COMING NEXT / IN DEVELOPMENT

Make the comparison meaningful.

  • Wider, independently checked model trials.
  • A qualified public execution environment and access controls.
  • Shared authority contracts and stronger result provenance.

No arbitrary agent uploads, external targets, or public model run controls are enabled.

Follow the build log ↗