Agent objective
- Interact with the supplied stateful environment.
- Produce verifier-checkable actions or artifacts.
- Maximise scalar reward under the package contract.
A procedural, partially observed, general-sum multi-agent environment where policies operate firms inside evolving supply networks. Agents produce, trade, bargain, share or challenge forecasts, manage credit and trust, and survive correlated shocks with unfamiliar counterparties while balancing profit, service, solvency, contribution, relationship integrity, sustainability, fairness, and lower-tail resilience.
A procedural, partially observed, general-sum multi-agent environment where policies operate firms inside evolving supply networks. Agents produce, trade, bargain, share or challenge forecasts, manage credit and trust, and survive correlated shocks with unfamiliar counterparties while balancing profit, service, solvency, contribution, relationship integrity, sustainability, fairness, and lower-tail resilience.
Enough detail to understand the intellectual terrain; generated instances, hidden mechanisms, and solution paths remain inside the private package.
| Environment | Mathematical or technical frontier | Adaptive research problem |
|---|---|---|
| Demand and supply shocks | Spikes, outages, routes, recalls | Coordinate production, sourcing, logistics, inventory, contracts, and trust through changing physical constraints. |
| Financial stress | Credit crunch and solvency | Balance borrowing, repayment, trade, capacity, and service while avoiding bankruptcy and idle cash-hoarding strategies. |
| Shared externalities | Carbon and infrastructure | Trade off private profit with sustainability and contribution to costly public goods under free-rider incentives. |
| Information and trust | Forecast manipulation and breaches | Interpret biased signals, communicate strategically, audit counterparties, and rebuild or withdraw trust. |
| Scarcity and compound recovery | Rationing, fairness, mixed shocks | Allocate scarce goods across priority and ordinary demand while preserving service, equity, and lower-tail resilience. |
A GPT-5.6-Pro-authored heuristic was frozen before a 180-episode private campaign. It scored 0.2413 versus 0.1194 for the reciprocal baseline, won on mean and tail utility, and cut bankruptcy from 55% to 0.56%. The score is a bounded multi-objective benchmark value, not a percentage or evidence of real procurement competence. Validation passed 12 tests with one optional dependency skip; the product remains a technically validated paid beta.
We publish aggregate behavior and task structure, while withholding generated instances, hidden labels, exact successful probes, private checks, and solution trajectories.
Shown with its provenance and limitations; it is not a performance guarantee.
Private campaign across 180 episodes, 12 scenarios, and three roles per scenario.
As identified by the supplied artifact.
180 reported runs.
mercantile_scores_bundle.zip:mercantile_scores_bundle/evaluation_summary.json
Machine-readable provenance and the exact displayed metric are available in results.json.
The paid ZIP will live in a private R2 bucket. Vercel authorizes the buyer and issues a 2–5 minute object URL; R2 serves the bytes directly.
Authenticated buyer + entitlement check
+ private R2 object + 2–5 minute signed URL
= direct, auditable download
Package SHA-256
96582a822be491102918281a5c84fc1cd69000fed1958abf403c9a43da9c7203One purchase licenses this identified item to one legal organisation for worldwide, perpetual commercial model training, evaluation, research and development. Redistribution and resale of the package are not permitted.