Agent objective
- Interact with the supplied stateful environment.
- Produce verifier-checkable actions or artifacts.
- Maximise scalar reward under the package contract.
A five-environment mathematics benchmark for extreme adaptive identification: each expert episode hides one system among 32,768 candidates and permits only two noisy measurements. Its worlds range from thermodynamic dynamics and cluster wall crossing to entropy solutions, planar algebras, and non-Shannon information geometry, demanding complementary experiment planning, sharp posterior inference, and sealed transfer.
A five-environment mathematics benchmark for extreme adaptive identification: each expert episode hides one system among 32,768 candidates and permits only two noisy measurements. Its worlds range from thermodynamic dynamics and cluster wall crossing to entropy solutions, planar algebras, and non-Shannon information geometry, demanding complementary experiment planning, sharp posterior inference, and sealed transfer.
Enough detail to understand the intellectual terrain; generated instances, hidden mechanisms, and solution paths remain inside the private package.
| Environment | Mathematical or technical frontier | Adaptive research problem |
|---|---|---|
| RuelleCocycleCipher | Thermodynamic formalism and transfer operators | Choose long-orbit and twisted-transfer measurements beyond common short-cylinder statistics. |
| ClusterScatteringOracle | Cluster varieties and wall crossing | Probe distinct path sectors to recover noncommuting wall geometry hidden beyond low-order tropical expansion. |
| ShockEntropyLabyrinth | Conservation laws and entropy solutions | Combine shock, entropy, rarefaction, and viscous evidence to infer off-anchor flux structure. |
| PlanarAlgebraTangleProbe | Subfactor planar algebras and tensor networks | Select phase-sensitive closed-tangle measurements despite identical modulus data and exact unitarity. |
| EntropyConeXRay | Entropy cones and non-Shannon inequalities | Detect high-order dependence through complementary inequalities when all lower-order marginals are uniform. |
A direct five-episode expert run selected high-quality probes—0.9710 mean action quality and 0.9487 experiment-design score—yet achieved 0.5007 mean reward, no exact identifications, and no passes. This is a useful extreme frontier: good local measurements are insufficient to resolve a 15-bit posterior in two calls. The audit also documents prompt-likelihood limitations rather than hiding them; the release passed 32/32 tests.
We publish aggregate behavior and task structure, while withholding generated instances, hidden labels, exact successful probes, private checks, and solution trajectories.
Shown with its provenance and limitations; it is not a performance guarantee.
Reported result from the evaluation artifact supplied with this package.
As identified by the supplied artifact.
5 reported runs.
direct_selftest_audit_bundle.zip:direct_selftest_report.md
Machine-readable provenance and the exact displayed metric are available in results.json.
The paid ZIP will live in a private R2 bucket. Vercel authorizes the buyer and issues a 2–5 minute object URL; R2 serves the bytes directly.
Authenticated buyer + entitlement check
+ private R2 object + 2–5 minute signed URL
= direct, auditable download
Package SHA-256
717aad4d87e90b78767192cad6d21504dd78b1a67677472490fcde00fe99eae4One purchase licenses this identified item to one legal organisation for worldwide, perpetual commercial model training, evaluation, research and development. Redistribution and resale of the package are not permitted.