Agent objective
- Interact with the supplied stateful environment.
- Produce verifier-checkable actions or artifacts.
- Maximise scalar reward under the package contract.
Five research-level quantum-information investigations in which agents distinguish 128 aliased physical hypotheses, calibrate nuisance effects, recover continuous parameters, choose posterior-dependent stress and falsification experiments, and publish uncertainty-aware certificates. Success is tested across interpolation, extrapolation, and counterfactual regimes—not merely by choosing a plausible mechanism.
Five research-level quantum-information investigations in which agents distinguish 128 aliased physical hypotheses, calibrate nuisance effects, recover continuous parameters, choose posterior-dependent stress and falsification experiments, and publish uncertainty-aware certificates. Success is tested across interpolation, extrapolation, and counterfactual regimes—not merely by choosing a plausible mechanism.
Enough detail to understand the intellectual terrain; generated instances, hidden mechanisms, and solution paths remain inside the private package.
| Environment | Mathematical or technical frontier | Adaptive research problem |
|---|---|---|
| Measurement-Induced Criticality Scaler | Monitored circuits and entanglement transitions | Separate genuine critical scaling from finite-size, feedback, conservation, postselection, and dephasing effects. |
| Quantum Markov Recovery Auditor | Conditional mutual information and Petz recovery | Decide whether apparently Markovian data support genuine recovery or hide structural and latent-mixture obstructions. |
| Anyon Fusion and Braiding Forensics | Non-Abelian anyons and braid representations | Infer fusion and braid behavior while rejecting leakage, poisoning, phase drift, readout, and nonadiabatic mimics. |
| Channel Superactivation Capacity Hunter | Quantum-channel capacity and superactivation | Certify finite-block activation while ruling out flag leakage, postselection, memory assistance, and calibration artifacts. |
| Quantum LDPC Threshold Transfer Laboratory | Quantum LDPC codes and correlated noise | Diagnose noise and decoder mismatch before transferring a threshold claim to an unseen code family. |
A source-blind public-interface evaluation averaged 0.5257 with no strict passes across five expert episodes. Several agents predicted interpolation and extrapolation extremely well yet still failed mechanism, evidence-consistency, calibration, or falsification gates. The contrast makes this a test of complete scientific practice rather than curve fitting. Exact-model calibration establishes reachability separately, and the release reports 148/148 tests.
We publish aggregate behavior and task structure, while withholding generated instances, hidden labels, exact successful probes, private checks, and solution trajectories.
Shown with its provenance and limitations; it is not a performance guarantee.
Reported result from the evaluation artifact supplied with this package.
As identified by the supplied artifact.
5 reported runs.
ulam_blind_public_transcripts.zip:ulam_blind_score_report.json
Machine-readable provenance and the exact displayed metric are available in results.json.
The paid ZIP will live in a private R2 bucket. Vercel authorizes the buyer and issues a 2–5 minute object URL; R2 serves the bytes directly.
Authenticated buyer + entitlement check
+ private R2 object + 2–5 minute signed URL
= direct, auditable download
Package SHA-256
768d49d887592c29dda686df5c4b7675e363d587f161c6c50a26b48c2795a236One purchase licenses this identified item to one legal organisation for worldwide, perpetual commercial model training, evaluation, research and development. Redistribution and resale of the package are not permitted.