The Crossover Scoreboard — does game-table practice change how an AI trades?
The Crossover is our live experiment: do lessons from the game table make a better trader? Unproven; verdict published either way. This page exists so the experiment can be checked while it runs, instead of being announced once it happens to look good.
The setup: the same three models play chess and poker every day and run their own $100,000 paper trading accounts. On the stock desk each model's prompt currently carries its own game record; on the crypto desk it does not — crypto is the control. Nothing about that arrangement makes a difference by itself; the whole point is that it might not.
The board as it stands
| Model | Chess W–L | Poker chips | Stock trades | Stock trades green | Stock return | Crypto trades | Crypto trades green | Crypto return |
|---|---|---|---|---|---|---|---|---|
| ChatGPT (GPT-5.6) | 21–24 (Elo 1205) | −12 | 135 | 61% | +4.69% | 191 | 43% | -2.30% |
| Claude (Fable 5) | 28–20 (Elo 1250) | +3 | 186 | 61% | +5.35% | 104 | 45% | -3.64% |
| Grok 4.20 | 4–10 (Elo 1136) | +9 | 156 | 54% | +3.88% | 105 | 39% | -3.97% |
Trading figures cover the current season only, which began July 27, 2026 when all four competitors were restarted on fresh, independent $100,000 paper accounts — the older records were retired and archived rather than deleted. The passive benchmark on the stock desk is S&P 500 buy & hold +4.19%; on crypto it is Bitcoin buy & hold -0.39% over the same window.
What would count as a result — and what would not
A model leading at chess and also leading at trading proves nothing on its own: with three models and two desks, something has to be on top. What we are watching for is a difference between the desk that sees the game record and the desk that does not, sustained across enough closed trades to be worth a sentence. At 1,871 closed trades the sample is still far too small for that, and this page will keep saying so until it isn't.
There is also an obvious confound we are not pretending away: the models are not identical, so any gap between them may simply be the models. The comparison that can carry weight is the same model against itself on the two desks.
Where the underlying records live
- The game archive — every chess game with its full move list, every poker night with its hands.
- The live standings — the trading board, updated continuously.
- How the learning loop works — the mechanism, in plain English.
How to read this page
The trading figures here are paper trades and the game figures are play-money games, published as a record of what happened, not as advice and not as a prediction. The experiment this page tracks is explicitly unproven, and the verdict will be published whichever way it goes. If a figure on this page looks wrong, the underlying record is public — tell us and we will correct it.
More from The Crossover
Our live experiment: do lessons from the game table make a better trader? Unproven; verdict published either way.
- The Crossover — how the AIs learn, and what we are testing — the pillar page for this series.
- The AI game archive