Agent objective
- Interact with the supplied stateful environment.
- Produce verifier-checkable actions or artifacts.
- Maximise scalar reward under the package contract.
A fully synthetic, offline defensive-security environment where agents triage alerts, investigate identity, endpoint, network, cloud, and change telemetry, preserve evidence, apply targeted containment, update an incident ticket, and submit a structured report. Eight procedural families include genuine attacks and benign lookalikes, with deterministic scoring for containment, evidence, scope, continuity, escalation, and investigation quality.
A fully synthetic, offline defensive-security environment where agents triage alerts, investigate identity, endpoint, network, cloud, and change telemetry, preserve evidence, apply targeted containment, update an incident ticket, and submit a structured report. Eight procedural families include genuine attacks and benign lookalikes, with deterministic scoring for containment, evidence, scope, continuity, escalation, and investigation quality.
Enough detail to understand the intellectual terrain; generated instances, hidden mechanisms, and solution paths remain inside the private package.
| Environment | Mathematical or technical frontier | Adaptive research problem |
|---|---|---|
| Identity compromise | Sessions, credentials, OAuth grants | Separate stolen sessions and malicious persistence from VPN travel and approved changes, then contain only affected authority. |
| Endpoint intrusion | Infostealers and ransomware precursors | Correlate process, persistence, memory, network, and user evidence before isolation and escalation. |
| Cloud and workload abuse | Service principals and resource access | Distinguish stolen workload credentials from authorized automation and rotate or block without unnecessary outage. |
| Insider exfiltration | Cloud, endpoint, removable media, egress | Preserve cross-source evidence, scope the data path, and coordinate targeted containment with legal and business escalation. |
| Benign lookalikes | VPN, travel, maintenance | Validate owners, changes, managed devices, MFA, and clean telemetry, then close accurately with zero disruptive action. |
In a blind expert episode, the acting model earned 0.9021 but failed because it treated an authorized change as an incident, adding scope and escalations where none were warranted. A feedback-corrected rerun passed at 0.9927, but is explicitly not the blind score. That gap demonstrates the benchmark's central defensive skill: avoiding disruptive false positives. The package passed 41/41 tests; reference calibration passed 64/64 while no-action, random, and over-containment controls passed none.
We publish aggregate behavior and task structure, while withholding generated instances, hidden labels, exact successful probes, private checks, and solution trajectories.
Shown with its provenance and limitations; it is not a performance guarantee.
The blind seed-17 run exceeded the scalar threshold but failed exact scope and escalation gates. A feedback-corrected, explicitly non-blind rerun scored 0.9927 and passed.
As identified by the supplied artifact.
1 reported runs.
soc_defender_evaluation_report.json
Machine-readable provenance and the exact displayed metric are available in results.json.
The paid ZIP will live in a private R2 bucket. Vercel authorizes the buyer and issues a 2–5 minute object URL; R2 serves the bytes directly.
Authenticated buyer + entitlement check
+ private R2 object + 2–5 minute signed URL
= direct, auditable download
Package SHA-256
595d86505a99c54eb2a9f1077a0cfceadeb3b311ed10fdfad4b94eae19874a37One purchase licenses this identified item to one legal organisation for worldwide, perpetual commercial model training, evaluation, research and development. Redistribution and resale of the package are not permitted.