Agent objective
- Interact with the supplied stateful environment.
- Produce verifier-checkable actions or artifacts.
- Maximise scalar reward under the package contract.
Five advanced mathematics environments built around adaptive inference from deliberately non-identifying evidence. Across arithmetic representations, optimal stopping, free probability, oriented matroids, and filtered homological algebra, agents face 128 shuffled hypotheses and at most two noisy measurements, rewarding complementary probes, calibrated uncertainty, and transfer to unseen conditions.
Five advanced mathematics environments built around adaptive inference from deliberately non-identifying evidence. Across arithmetic representations, optimal stopping, free probability, oriented matroids, and filtered homological algebra, agents face 128 shuffled hypotheses and at most two noisy measurements, rewarding complementary probes, calibrated uncertainty, and transfer to unseen conditions.
Enough detail to understand the intellectual terrain; generated instances, hidden mechanisms, and solution paths remain inside the private package.
| Environment | Mathematical or technical frontier | Adaptive research problem |
|---|---|---|
| FrobeniusCipher | Finite representations and Frobenius actions | Select primes and noncommutative readouts whose activation patterns jointly reveal hidden arithmetic twists. |
| SnellEnvelope | Optimal stopping and Bellman free boundaries | Combine value and policy-boundary evidence when similar values conceal different stop/continue mechanisms. |
| FreeMomentProbe | Free probability and mixed operator moments | Recover relative eigenbases through mixed words although every marginal spectrum and pure moment agrees. |
| ChirotopeCipher | Oriented matroids and Plücker geometry | Choose sign-sensitive geometric measurements when Gram magnitudes and the underlying unoriented matroid reveal nothing. |
| SpectralSequenceXRay | Filtered complexes and extension data | Pair filtration-sensitive probes with analytic operators while all associated-graded and homological invariants agree. |
A task-adaptive exact-Bayesian policy averaged 0.7881 on 60 fresh expert episodes, identifying 33 systems and passing 19. A one-hot submission raised formal passes to 29 but reduced mean reward when confidently wrong, exposing a useful calibration-versus-commitment tradeoff. FrobeniusCipher passed 10/12 while two families passed none under calibrated submission. The package passed 47/47 tests.
We publish aggregate behavior and task structure, while withholding generated instances, hidden labels, exact successful probes, private checks, and solution trajectories.
Shown with its provenance and limitations; it is not a performance guarantee.
The report identifies this as a source-informed or executable-policy result, not a clean black-box benchmark.
As identified by the supplied artifact.
60 reported runs.
ulam_rlvr_scores.json
Machine-readable provenance and the exact displayed metric are available in results.json.
The paid ZIP will live in a private R2 bucket. Vercel authorizes the buyer and issues a 2–5 minute object URL; R2 serves the bytes directly.
Authenticated buyer + entitlement check
+ private R2 object + 2–5 minute signed URL
= direct, auditable download
Package SHA-256
f4a0845fddb6827ca1063ad29a58c61aa241e9beeee2b3edadebce8c26fe4d37One purchase licenses this identified item to one legal organisation for worldwide, perpetual commercial model training, evaluation, research and development. Redistribution and resale of the package are not permitted.