Agent objective
- Interact with the supplied stateful environment.
- Produce verifier-checkable actions or artifacts.
- Maximise scalar reward under the package contract.
Five long-horizon mathematics environments with noisy apparatus, hidden nuisance variables, and irreversible decisions. Across cluster wall crossing, p-adic Hodge theory, rough paths, KAM resonance, and spin-glass landscapes, agents route observations through distinct state machines, allocate scarce resources, and choose justified identification, bounded-set, or abstention decisions with sealed transfer.
Five long-horizon mathematics environments with noisy apparatus, hidden nuisance variables, and irreversible decisions. Across cluster wall crossing, p-adic Hodge theory, rough paths, KAM resonance, and spin-glass landscapes, agents route observations through distinct state machines, allocate scarce resources, and choose justified identification, bounded-set, or abstention decisions with sealed transfer.
Enough detail to understand the intellectual terrain; generated instances, hidden mechanisms, and solution paths remain inside the private package.
| Environment | Mathematical or technical frontier | Adaptive research problem |
|---|---|---|
| ClusterWallCrossingLab | Cluster algebras and scattering diagrams | Commit to a seed and route mutations, walls, theta bases, stability, and monodromy evidence adaptively. |
| pAdicHodgeLab | p-adic Hodge and Galois representations | Choose a period-ring route and coordinate Frobenius, filtration, monodromy, refinement, and local conditions. |
| RoughPathSignatureLab | Rough paths and controlled equations | Select a lift and noncommutative words, then reconcile area, bracket, renormalization, and response evidence. |
| KAMResonanceLab | Hamiltonian resonance and small divisors | Commit to a frame and adapt normal-form, residue, splitting, control, and transport probes. |
| SpinGlassLandscapeLab | Replica symmetry breaking and TAP landscapes | Choose replica coupling and combine overlap, cavity, Hessian, chaos, signal, and annealing evidence. |
A session-mediated GPT-5.6 Pro sweep averaged 0.4923 with no passes across ten hard/expert cases. Its mean path score was 0.8386 but terminal-decision score only 0.2765, making the suite a focused test of moving from competent experimentation to justified conclusions. Every packaged control also passed 0/10, while privileged replay certificates establish reachability. The release passed 132/132 tests.
We publish aggregate behavior and task structure, while withholding generated instances, hidden labels, exact successful probes, private checks, and solution trajectories.
Shown with its provenance and limitations; it is not a performance guarantee.
Reported result from the evaluation artifact supplied with this package.
As identified by the supplied artifact.
10 reported runs.
gpt56_pro_blind_report.json
Machine-readable provenance and the exact displayed metric are available in results.json.
The paid ZIP will live in a private R2 bucket. Vercel authorizes the buyer and issues a 2–5 minute object URL; R2 serves the bytes directly.
Authenticated buyer + entitlement check
+ private R2 object + 2–5 minute signed URL
= direct, auditable download
Package SHA-256
291f6e57b3e8bd54a1ef5508e73f1942c7247d383934ba557997125efcbf4acbOne purchase licenses this identified item to one legal organisation for worldwide, perpetual commercial model training, evaluation, research and development. Redistribution and resale of the package are not permitted.