Agent objective
- Interact with the supplied stateful environment.
- Produce verifier-checkable actions or artifacts.
- Maximise scalar reward under the package contract.
Five quantum-information environments for adaptive experimental reasoning under model aliasing, nuisance parameters, laboratory drift, and strict resource limits. Each expert episode requires calibration, complementary probe families, a posterior-selected stress test, and unseen transfer—combining physical-model identification, continuous parameter recovery, calibrated uncertainty, and scientific certification.
Five quantum-information environments for adaptive experimental reasoning under model aliasing, nuisance parameters, laboratory drift, and strict resource limits. Each expert episode requires calibration, complementary probe families, a posterior-selected stress test, and unseen transfer—combining physical-model identification, continuous parameter recovery, calibrated uncertainty, and scientific certification.
Enough detail to understand the intellectual terrain; generated instances, hidden mechanisms, and solution paths remain inside the private package.
| Environment | Mathematical or technical frontier | Adaptive research problem |
|---|---|---|
| Indefinite Causal Order Adversarial Certifier | Process matrices and causal nonseparability | Combine process, witness, interference, and intervention evidence before targeting the remaining memory or coherent-order explanation. |
| Bosonic Code Channel Learning Laboratory | Continuous-variable channels and bosonic codes | Calibrate frames, distinguish noise and multimode effects, then stress the leading explanation and transfer to an unseen code family. |
| Scrambling Hydrodynamics Extrapolator | Quantum chaos and many-body hydrodynamics | Integrate scrambling fronts, transport, spectral correlations, and entanglement evidence to extrapolate beyond observed size. |
| Finite-Key Device-Independent Entropy Auditor | Bell nonlocality and finite-key entropy | Combine loophole-aware Bell and temporal evidence, then stress memory, loss, or input leakage before finite-block transfer. |
| Holevo-Limited Multiparameter Sensor Architect | Multiparameter metrology and Holevo costs | Compare probe and measurement strategies under loss and correlated noise, then transfer to an unseen cost matrix. |
In a five-episode public-interface run, GPT-5.6 Pro averaged 0.5666 and passed none, despite 0.9309 experiment-design and 0.9680 protocol scores. Evidence consistency averaged only 0.0102. That cleanly isolates a valuable post-training gap: selecting a sensible quantum protocol is much easier than reconstructing an honest, concentrated posterior and reliable scientific certificate. All transcripts replayed, and the package reports 151/151 tests.
We publish aggregate behavior and task structure, while withholding generated instances, hidden labels, exact successful probes, private checks, and solution trajectories.
Shown with its provenance and limitations; it is not a performance guarantee.
Reported result from the evaluation artifact supplied with this package.
As identified by the supplied artifact.
5 reported runs.
gpt-5.6-pro_ulam-rlvr-v3_expert_seed0_run.json
Machine-readable provenance and the exact displayed metric are available in results.json.
The paid ZIP will live in a private R2 bucket. Vercel authorizes the buyer and issues a 2–5 minute object URL; R2 serves the bytes directly.
Authenticated buyer + entitlement check
+ private R2 object + 2–5 minute signed URL
= direct, auditable download
Package SHA-256
68847a09fcd6e1b690c7efc756e177e58284e821ee47c71a4cade23931b4920eOne purchase licenses this identified item to one legal organisation for worldwide, perpetual commercial model training, evaluation, research and development. Redistribution and resale of the package are not permitted.