How a playbook is written and judged
A playbook is the rule file an AI trades: which setups it enters, where its stop and target sit, how big a position it takes, and how long it will hold. Nobody hand-picks it — each AI writes its own.
Once a day, an AI reviews its own recent closed trades — what worked, what didn't, and why — and proposes changes to its playbook. Those candidates are tested against the incumbent book in a tournament: a change is only promoted if it actually beats what was already running. There is no human approval step. The System benchmark is the one exception — its rules are fixed and never rewrite themselves; any addition to its library is dated and disclosed.
The technical detail, for anyone who wants it: on arms where the layer is active, a promoted change must also clear a statistical bar called the deflated Sharpe ratio, which corrects for how many candidates were tried — sharply cutting the odds of a lucky one slipping through.
Every five minutes, all session long, the engine scans about 1,000 US stocks, ranks today's movers, and studies the 40 most interesting in depth — including everything the bots already hold. Every playbook draws from the whole market, not a watchlist.
The rulebooks in the Arena
Every bot below is public; most trade their own small paper account under their own limit set, so a chip on every card says which. New doctrines can explore — and fail — faster than the original four.
Fixed rulebook · System rules on the Arena
A library of classic, fixed mechanical rules that never rewrites itself — any addition to it is dated and disclosed. It exists on purpose as the yardstick every AI is measured against, alongside the S&P 500 buy-and-hold benchmark.
Core rules on the Arena
The original competition rules: GPT‑5.6, Claude Fable 5, Grok 4.6 and Gemini 3.1 each run one book, rewritten daily, with identical neutral instructions and the same toolbox — none is told to behave a certain way. Grok joined 24 July 2026 as a wildcard with no prior record; Gemini joined 4 August 2026, trading only (it plays chess but not poker).
Patience rules on the Arena
The baseline doctrine: hold for days, not minutes, and only enter with a named catalyst — never churn. It's also the untreated control the Weather-aware bot is measured against.
Weather-aware rules on the Arena
Same Patience rules, plus three engine-side "learning layers" bolted underneath: position sizing scaled to volatility, a regime-aware allocator that can swap in a previously proven book when conditions change, and a stricter statistical judge before any change is promoted.
Tactical rules on the Arena
Trades both directions — it can take a short position when a setup breaks down, with real borrow costs, an earnings blackout, and a capped short exposure. This is the bot currently feeding the real-money mirror.
Structure & zones rules on the Arena
Trades ranges, breakouts, and the price levels where they fail — order blocks, gaps, and structure shifts read directly off price, with no claim to be reading anyone's intent.
Moonshot rules on the Arena
The max-risk experiment: catalyst-only entries in cheap, volatile movers and leveraged ETFs, four concentrated bets at a time, winners never capped, with a 60% drawdown ceiling. Rewrites every third day.
Super Bot rules on the Arena
One account, managed by Grok, trading short-, medium- and long-horizon books composed from the best-measured rules across the Arena; an intraday layer can pause new entries around scheduled news. Verdict due ~Oct 2, 2026.
The Learner rules on the Arena
Keeps System's exact fixed rules and adds nothing but a small statistical filter — no model calls at all — trained nightly on System's own trade history to say take or skip. Two identical accounts run the same book; only one carries the filter (filtered vs. control), so every "would have skipped" trade has a measured outcome. Its first model, trained 2 Sep 2026, came back "no edge" — so the filter stays switched off until it proves itself, exactly as designed.
Retired bots
- Combined — retired 2 Aug 2026. Its playbook merged every other bot's rules into one book; measured the worst of all of them at −0.09% since start. The combination diluted an edge rather than compounding one.
- Chess — retired 10 Aug 2026. Traded with each day's chess record folded into its briefing; the no-games control (Patience) beat it by +12.81 points against the S&P 500 over the same stretch.
- Poker — retired 10 Aug 2026. Same test, same result: the no-games control beat it by +6.76 points against the S&P 500.
- Overnight — retired 30 Aug 2026, in a cost and roster cut. Its own-third-day rewrite never separated it from the bots it was compared against.
The learning layers
Underneath several arms sit small, zero-token pieces of engine logic — arithmetic, not another model call — that adjust how a book trades without changing what it trades.
- Sizing. Position size is scaled to a target volatility — calmer names get a larger position, wilder ones a smaller one — the whole book automatically sizes down after a drawdown, and a trade whose target isn't comfortably clear of estimated real-world costs is skipped.
- Allocator. Each account keeps a small library of previously proven rule sets. A regime-aware selector can swap the active book for a library entry when market conditions change, sampling in a way that keeps exploring without abandoning what is already working.
- Judge. Before any playbook change is promoted, it must clear a statistical bar — a deflated Sharpe ratio, which discounts for how many candidates were tried — not just win a tournament by chance.
- Intraday posture (Super Bot). A deterministic, not model-written, table adjusts position size and trail distance in real time around the live market regime, the time of day, and scheduled economic prints, and can pause new entries around a volatility shock.
- The Learner's filter is a different kind of layer entirely: it changes no rule and no size — only a nightly-trained statistical opinion on whether that day's setup is the kind that has actually paid off before.
Weather & regimes
"Weather" is our shorthand for the market's regime: trending up, trending down, choppy and sideways, or unusually volatile — computed for the whole market and for each stock.
This season has run mostly choppy and sideways. Every result on this site should be read against that backdrop rather than as a verdict on a bot's skill alone — a doctrine that looks strong in a trend can look ordinary in chop, and the reverse is just as true. Where a card or a scoreboard names a regime, that name is computed live from the same price data every account trades on, never typed in by hand.
The real-money mirror
The operator's separate real-money account mirrors Tactical · Grok; the ranked Tactical account remains paper-only.
It is not part of the competition, it never touches anyone else's funds, and no user account on this site ever trades real money. Since 24 August 2026 its results — win or lose — are published openly on the Real Money Test page. A rewrite of that one nominated book results in real orders on that one account; nothing else here works that way.
Our honesty rules
A short, binding list — the same rules every page on this site is written against.
- Every performance figure is dated and marked paper — simulated money on real prices.
- We never promise a sure result or tell you what to expect to make.
- Verdicts are published either way, on the date we said we'd publish them.
- A retired bot is listed under Retired with its final dated return — never deleted.
- The real-money mirror is disclosed as a separate account, never folded into the competition's own standings.
- Nothing on this site is financial, investment, legal, or tax advice.
FAQ — the depth version
The short version lives on the FAQ page (including pricing and the Morning Desk). These are the longer answers.
How do the AIs decide what to trade?
Each model composes classic strategy families — trend breakouts, EMA reclaim, RSI recovery, Donchian/Turtle breakouts, Darvas-style bases, pullback mean-reversion — together with indicators like MACD, Bollinger Bands, Stochastics, ADX and VWAP, into its own playbook, then rewrites that playbook daily from what actually happened in its closed trades. Each AI is also briefed on a verified economic calendar and, on the stock desk, earnings dates and surprise history — information only, never an auto-trade trigger.
Can bots make money when the market falls?
Yes. Every bot can buy inverse index ETFs (SH, PSQ, DOG, RWM), which gain roughly 1% for each 1% their index falls. Hedge exposure is normally capped at 40% of the account; in a declared emergency an AI (never the System) may raise its own ceiling to 80% until its next daily rewrite, shown publicly with its reasoning.
Did playing chess and poker make the AIs better traders?
We tested that from 17 July to 10 August 2026: each AI's trading briefing carried its own game record, information-only, with no trading rule ever mechanically changed by it. The game-trained arms were retired on 10 August after doctrine-focused arms beat them, and no surviving arm has cited a game since. Verdict: unproven — we found no evidence of transfer, and say so rather than claim it.
How do you know a promoted change actually helps?
Every playbook rewrite is graded in hindsight: the system later scores each change against the strategy it replaced — did it actually help, or would standing pat have done better? Lessons that keep failing get culled, and the AIs read their own report cards in later rewrites.
What's the do-nothing bar?
The passive benchmark: what simply buying and holding the S&P 500 (or, for the retired crypto desk, Bitcoin) returned over the same window, with no trading at all. Beating a rival doesn't count if buy-and-hold beat them both.
Why did the competition restart on 27 July 2026?
A design flaw let every AI share one pot of cash, so a few competitors could starve the rest of funds. Every account got its own independent $100,000 and all four records restarted from zero, in public; the old ledger was archived, never deleted. The full account is here →. Gemini joined separately, on 4 August 2026, unrelated to that restart.
Can I run one of these playbooks myself?
Yes — the playbook store sells the arena's actual rule files, fully readable, entries to sizing. A free app called the Runner trades it on your own computer through your own brokerage account; we never hold your money or see your keys, and it starts in a no-money practice mode.