Agent objective
- Interact with the supplied stateful environment.
- Produce verifier-checkable actions or artifacts.
- Maximise scalar reward under the package contract.
Fifteen quantitative-finance control environments across execution, statistical arbitrage, market making, portfolio allocation, and option hedging at Core, Research, and Challenge tiers. Agents face richer costs, constraints, delays, hidden liquidity, and stress, with paired exogenous tapes, observation-only policies, native financial metrics, and deterministic replay.
Fifteen quantitative-finance control environments across execution, statistical arbitrage, market making, portfolio allocation, and option hedging at Core, Research, and Challenge tiers. Agents face richer costs, constraints, delays, hidden liquidity, and stress, with paired exogenous tapes, observation-only policies, native financial metrics, and deterministic replay.
Enough detail to understand the intellectual terrain; generated instances, hidden mechanisms, and solution paths remain inside the private package.
| Environment | Mathematical or technical frontier | Adaptive research problem |
|---|---|---|
| Optimal Execution · 3 tiers | Impact, passive fills, liquidity, latency | Schedule passive and marketable orders while balancing completion, shortfall, transient impact, and vanishing opportunities. |
| Statistical Arbitrage · 3 tiers | Pairs, baskets, structural breaks | Control exposure through costs, drift, stale prices, short-locate limits, and borrow recalls. |
| Market Making · 3 tiers | Queues, adverse selection, inventory | Choose quotes and sizes under discrete priority, latency, toxic flow, and tail inventory risk. |
| Portfolio Allocation · 3 tiers | Return, turnover, leverage, delayed macro data | Allocate long-only or long/short portfolios under distribution shift and information delay without future leakage. |
| Option Hedging · 3 tiers | Dynamic nonlinear risk control | Coordinate stock and option instruments under costs, latency, jumps, and path-dependent hedging error. |
Frozen GPT-5.6-Pro-authored observation-only controllers were evaluated on 1,000 held-out episodes and compared with 1,000 paired official-benchmark episodes. They were clearly better in 7/10 task/scenario comparisons, with three inconclusive and no clear losses. The report preserves each family's native financial metric rather than inventing a misleading aggregate. Validation collected 262 passing tests, one justified optional skip, and 93/93 release invariants.
We publish aggregate behavior and task structure, while withholding generated instances, hidden labels, exact successful probes, private checks, and solution trajectories.
Shown with its provenance and limitations; it is not a performance guarantee.
Frozen observation-only controllers beat the official benchmark in 7/10 paired test/stress comparisons; three were inconclusive and none were clear losses.
As identified by the supplied artifact.
1000 reported runs.
GPT56_QFinRL_v3_run_bundle.zip:GPT56_QFinRL_v3_report.md
Machine-readable provenance and the exact displayed metric are available in results.json.
The paid ZIP will live in a private R2 bucket. Vercel authorizes the buyer and issues a 2–5 minute object URL; R2 serves the bytes directly.
Authenticated buyer + entitlement check
+ private R2 object + 2–5 minute signed URL
= direct, auditable download
Package SHA-256
41ab266d2284bab2115d4c1d78dbf6553e131dc34970f8c1f9044118aed1aa89One purchase licenses this identified item to one legal organisation for worldwide, perpetual commercial model training, evaluation, research and development. Redistribution and resale of the package are not permitted.