Agent objective
- Interact with the supplied stateful environment.
- Produce verifier-checkable actions or artifacts.
- Maximise scalar reward under the package contract.
Five environments for auditing and repairing flawed mathematical arguments under evidence and tool constraints. Agents inspect interleaved branches, search a frozen corpus, test theorem applicability, construct executable counterexamples, recover nonlocal dependencies, locate the earliest fatal gap, and propose minimal repairs that survive controlled hidden variants.
Five environments for auditing and repairing flawed mathematical arguments under evidence and tool constraints. Agents inspect interleaved branches, search a frozen corpus, test theorem applicability, construct executable counterexamples, recover nonlocal dependencies, locate the earliest fatal gap, and propose minimal repairs that survive controlled hidden variants.
Enough detail to understand the intellectual terrain; generated instances, hidden mechanisms, and solution paths remain inside the private package.
| Environment | Mathematical or technical frontier | Adaptive research problem |
|---|---|---|
| AlgebraicFaultline | Algebra and number theory | Audit divisibility, cancellation, inverse, sign, and polynomial arguments, then repair missing hypotheses across hidden variants. |
| LinearAlgebraSurgery | Linear algebra and matrix theory | Diagnose invalid claims about positivity, products, spectra, rank, nilpotence, and inverses, then rebuild dependencies. |
| AnalysisQuantifierLab | Real and functional analysis | Track quantifiers and global hypotheses through convergence, completeness, differentiation, integration, and local-to-global steps. |
| ProbabilitySetAudit | Probability and conditioning | Test dependence structure and set corrections, then repair the earliest inference without overclaiming convergence or independence. |
| CombinatorialInductionLab | Combinatorics and graph arguments | Audit bases, multiplicities, recurrences, assumptions, and inclusion–exclusion dependencies, then cover every relevant case. |
The supplied five-episode hard test rollout averaged 0.9472 and passed four tasks. It localized every earliest gap and complete defect set, yet AlgebraicFaultline failed because one repair did not survive hidden variants. That distinction is the benchmark's core value: diagnosing a flaw is not enough unless the repair is robust. The package reports 190 tests, 1,200 stress cases, and 1,200 exact replays.
We publish aggregate behavior and task structure, while withholding generated instances, hidden labels, exact successful probes, private checks, and solution trajectories.
Shown with its provenance and limitations; it is not a performance guarantee.
Reported result from the evaluation artifact supplied with this package.
As identified by the supplied artifact.
5 reported runs.
ulam_run_bundle.zip:ulam_score_report.json
Machine-readable provenance and the exact displayed metric are available in results.json.
The paid ZIP will live in a private R2 bucket. Vercel authorizes the buyer and issues a 2–5 minute object URL; R2 serves the bytes directly.
Authenticated buyer + entitlement check
+ private R2 object + 2–5 minute signed URL
= direct, auditable download
Package SHA-256
2a051f0db17564364d0d29ef6b16806af11c197bd9823d6f5eb6cd17c574cb3eOne purchase licenses this identified item to one legal organisation for worldwide, perpetual commercial model training, evaluation, research and development. Redistribution and resale of the package are not permitted.