Agent objective
- Interact with the supplied stateful environment.
- Produce verifier-checkable actions or artifacts.
- Maximise scalar reward under the package contract.
Five adaptive mathematics environments spanning tensor decomposition, topological invariants, kernel geometry, coding theory, and nonlinear dynamics. Each hides meaningful structure behind deliberately non-identifying summaries, forcing agents to choose complementary probes, manage experimental cost, and make calibrated predictions on sealed cases.
Five adaptive mathematics environments spanning tensor decomposition, topological invariants, kernel geometry, coding theory, and nonlinear dynamics. Each hides meaningful structure behind deliberately non-identifying summaries, forcing agents to choose complementary probes, manage experimental cost, and make calibrated predictions on sealed cases.
Enough detail to understand the intellectual terrain; generated instances, hidden mechanisms, and solution paths remain inside the private package.
| Environment | Mathematical or technical frontier | Adaptive research problem |
|---|---|---|
| TensorCircuit | Tensor decomposition and multilinear sensing | Align complementary contractions and operator views to recover sparse tensor structure hidden by shared null and crossed-factor responses. |
| TopologyLens | Topological data analysis and Hodge geometry | Use anchored cohomological probes to recover labeled attachment geometry that global Betti and heat summaries cannot identify. |
| KernelForge | Kernel methods, Gaussian processes, operator learning | Probe away from a shared seed point to reveal differences in spectra, uncertainty, leverage, and learning-flow behavior. |
| DecoderLab | Coding theory and probabilistic decoding | Construct ambiguous multi-bit errors that expose learned decoder bias while clean and single-bit behavior remains identical. |
| KoopmanLab | Nonlinear dynamics, Koopman operators, control | Move off a shared fixed point and choose states, controls, observables, and horizons that enter the nonlinear regime. |
In a blind public-interface evaluation, GPT-5.6 Pro passed all ten hard and expert episodes with 0.9285 mean reward using an average of 3.3 queries. No-query and inert controls passed none, while the release passed 71/71 tests and all ten fixture replays. The result demonstrates that targeted mathematical probes work, without making the tasks a one-step lookup.
We publish aggregate behavior and task structure, while withholding generated instances, hidden labels, exact successful probes, private checks, and solution trajectories.
Shown with its provenance and limitations; it is not a performance guarantee.
Reported result from the evaluation artifact supplied with this package.
As identified by the supplied artifact.
5 reported runs.
5Math-4_results.zip:scores_summary.json
Machine-readable provenance and the exact displayed metric are available in results.json.
The paid ZIP will live in a private R2 bucket. Vercel authorizes the buyer and issues a 2–5 minute object URL; R2 serves the bytes directly.
Authenticated buyer + entitlement check
+ private R2 object + 2–5 minute signed URL
= direct, auditable download
Package SHA-256
c00f23b4726c2149a65ba7964dbeb6f7d7787562284631c4485503d51684fdb6One purchase licenses this identified item to one legal organisation for worldwide, perpetual commercial model training, evaluation, research and development. Redistribution and resale of the package are not permitted.