Agent objective
- Interact with the supplied stateful environment.
- Produce verifier-checkable actions or artifacts.
- Maximise scalar reward under the package contract.
Five mathematical system-identification environments with 2,048 hypotheses and only two noisy measurements. From hyperbolic geometry and graph limits to inverse scattering, conformal bootstrap, and mean-field games, every candidate shares substantial exact invariants, so success depends on a complementary probe pair and a sharply calibrated posterior that transfers to sealed functionals.
Five mathematical system-identification environments with 2,048 hypotheses and only two noisy measurements. From hyperbolic geometry and graph limits to inverse scattering, conformal bootstrap, and mean-field games, every candidate shares substantial exact invariants, so success depends on a complementary probe pair and a sharply calibrated posterior that transfers to sealed functionals.
Enough detail to understand the intellectual terrain; generated instances, hidden mechanisms, and solution paths remain inside the private package.
| Environment | Mathematical or technical frontier | Adaptive research problem |
|---|---|---|
| TeichmullerEcho | Marked hyperbolic and Teichmüller geometry | Select marked functionals probing complementary deformations hidden by identical individual lengths and traces. |
| GraphonGauge | Graph limits and compact operators | Recover a hidden spectral frame through marked statistics beyond shared spectra and unmarked cycle densities. |
| SolitonScatteringCipher | Toda inverse scattering | Infer norming and collision structure behind a fixed spectrum using complementary times and spectral locations. |
| BootstrapFunctionalXRay | Conformal bootstrap | Probe off-anchor crossing behavior beyond a complete shared low-order crossing jet. |
| MeanFieldMasterProbe | Nonlinear mean-field games | Coordinate path, equilibrium-residual, and curvature measurements to reveal mechanisms absent from linearization. |
A two-stage exact-likelihood planner averaged 0.7619 over 25 fresh expert episodes and passed three with its calibrated posterior. More aggressive confidence-gated and one-hot submissions raised pass counts but reduced average reward, exposing the cost of false certainty. The planner beat the same-seed greedy control by 6.2 reward points. These are simulator-control scores, not a raw model rollout; the release passed 51/51 tests.
We publish aggregate behavior and task structure, while withholding generated instances, hidden labels, exact successful probes, private checks, and solution trajectories.
Shown with its provenance and limitations; it is not a performance guarantee.
Reported result from the evaluation artifact supplied with this package.
As identified by the supplied artifact.
See the source methodology.
mahlo_passplanner_fresh_results.json
Machine-readable provenance and the exact displayed metric are available in results.json.
The paid ZIP will live in a private R2 bucket. Vercel authorizes the buyer and issues a 2–5 minute object URL; R2 serves the bytes directly.
Authenticated buyer + entitlement check
+ private R2 object + 2–5 minute signed URL
= direct, auditable download
Package SHA-256
e6d0f7d125514ce4ee996bd844187c2eb6fd60a38919745767574bf12621d2a3One purchase licenses this identified item to one legal organisation for worldwide, perpetual commercial model training, evaluation, research and development. Redistribution and resale of the package are not permitted.