Agent objective
- Interact with the supplied stateful environment.
- Produce verifier-checkable actions or artifacts.
- Maximise scalar reward under the package contract.
An ML-systems benchmark for repairing stateful training infrastructure without corrupting continuation, optimizer lineage, distributed state, schema history, or durable checkpoints. Eight private stages contain multi-module incidents with semantic defects across resume, shards, surgery, schemas, compression, recovery, transactions, and incident response under attestation, corruption, concurrency, crash, determinism, and resource checks.
An ML-systems benchmark for repairing stateful training infrastructure without corrupting continuation, optimizer lineage, distributed state, schema history, or durable checkpoints. Eight private stages contain multi-module incidents with semantic defects across resume, shards, surgery, schemas, compression, recovery, transactions, and incident response under attestation, corruption, concurrency, crash, determinism, and resource checks.
Enough detail to understand the intellectual terrain; generated instances, hidden mechanisms, and solution paths remain inside the private package.
| Environment | Mathematical or technical frontier | Adaptive research problem |
|---|---|---|
| Levels 8–9 | Full-state resume and flattened shards | Preserve microbatch, optimizer, scheduler, sampler, RNG, AMP, padding, alias, topology, and repartition semantics. |
| Levels 10–11 | Model surgery and schema lineage | Rename, split, merge, add, and delete parameters while migrating optimizer state and executing idempotent migration DAGs. |
| Levels 12–13 | Compressed state and shard recovery | Normalize dense, sparse, quantized, and low-precision representations; detect and recover stale, corrupt, duplicated, or parity-recoverable shards. |
| Level 14 | Transactional checkpoint publication | Handle concurrent writers, compare-and-swap, locks, crash boundaries, and durable visibility without split-brain state. |
| Level 15 | Lineage incident response | Select the latest valid lineage, reconcile conflicting evidence, recover state, and prove exact training continuation. |
A GPT-5.6 Pro clean-agent proxy passed all eight visible suites but earned only three perfect hidden stages, with a 0.5779 official aggregate. Development-family mean was 0.7683 versus 0.5597 on held-out stages, strong evidence that the verifier tests system semantics beyond public fixtures. All receipts verified; owner references passed, starters failed, and 96/96 semantic mutants were killed. The proxy used unsafe-local execution and is not an independently calibrated model run.
We publish aggregate behavior and task structure, while withholding generated instances, hidden labels, exact successful probes, private checks, and solution trajectories.
Shown with its provenance and limitations; it is not a performance guarantee.
Reported result from the evaluation artifact supplied with this package.
As identified by the supplied artifact.
See the source methodology.
gpt56pro_v5_score_report.json
Machine-readable provenance and the exact displayed metric are available in results.json.
The paid ZIP will live in a private R2 bucket. Vercel authorizes the buyer and issues a 2–5 minute object URL; R2 serves the bytes directly.
Authenticated buyer + entitlement check
+ private R2 object + 2–5 minute signed URL
= direct, auditable download
Package SHA-256
41b126727d7f2fd0dc187b37e3b043edfb4d7fdf9c805908e3fefbc6afa14c35One purchase licenses this identified item to one legal organisation for worldwide, perpetual commercial model training, evaluation, research and development. Redistribution and resale of the package are not permitted.