Can AI beat the stock market? (We're testing it live)
Not yet, on this evidence. As of 2026-10-02, Patience · Grok leads the Bot Analysis Arena at +45.58% since Jul 28, 2026, against +4.14% for the S&P 500 over the same dates. The control that no AI rewrites, System, is at +3.48% — 9 of the 17 AI-rewritten bots are ahead of it, 13 are ahead of the S&P 500 over their own dates. All of it is paper trading — simulated money, real market prices — so it can show whether an AI-rewritten rulebook beats a fixed one and the index in the open, but it cannot show what real fills, fees and taxes would do to the same trades.
The control-bot comparison
The control: Fixed rulebook · System trades the same market with a rulebook no AI ever rewrites. As of 2026-10-02 it is at +3.48% since Jul 27, 2026 — #10 of 22. Of the 17 AI-rewritten bots, 9 are ahead of it and 13 are ahead of the S&P 500 over their own dates. 2 bots started the very same day as System (Patience · Claude +27.89%, Patience · ChatGPT +14.55%), which is the cleanest like-for-like comparison on the board.
| AI model (all its bots) | Average return | Average vs System | Its best bot |
|---|---|---|---|
| ChatGPT — 6 bots | +5.45% | +1.97pp | Patience · ChatGPT +14.55% |
| Claude — 5 bots | +9.96% | +6.48pp | Patience · Claude +27.89% |
| Grok — 6 bots | +8.66% | +5.18pp | Patience · Grok +45.58%* |
| Gemini — 2 bots | -0.13% | -3.61pp | Patience · Gemini +2.74%* |
| Astra composite (frozen ensemble) — 1 bot | -1.51% | -4.99pp | Astra · composite (frozen ensemble) -1.51%* |
| Edge Bot (fixed tested rules, no AI) — 1 bot | +0.88% | -2.60pp | Edge Bot +0.88%* |
| System (fixed rulebook — the control) | +3.48% | — | — |
| S&P 500 buy & hold (benchmark) | +4.16% | +0.68pp | — |
Averages are plain means across every bot a model rewrites, late starters included; Each "vs S&P" figure compares a bot with the S&P 500 over that bot's own dates (bots start on different days); the S&P 500 row itself covers the arena's full window from Jul 27, 2026. A figure marked * is a partial period or a stale base day — hover it for the reason.
Shorts
Every bot on the board may go long or short (the rules chip on each record page states its short cap). Closed-trade rows in the public ledger do not carry a side field, so this page does not split long from short; the per-bot long-vs-short table lives on Bot Analytics, and the measured up-tape shorts study is quoted below.
Negative findings, stated plainly
Most of what this experiment has produced is evidence of what does not work. Each item names the file in the repository where the numbers live:
- A textbook range strategy — Bollinger(20,2) touch + ADX(14) below 22 + RSI(14) at 30/70 — was pre-registered and back-tested against the engine's own cost model. On the 1-hour timeframe the bots actually trade, its chop-slice net expectancy was +0.11pp per trade at 3 bps cost: a coin flip. Its one strong-looking cell (+3.43pp per trade) rested on 31 trades across 80 symbols over two years. Verdict: not added to any bot. source: research/bb-adx-rsi-range-study-2026-09-03.md
- Measured across all 26 books since the Jul 27 reset: shorts opened on green S&P days (n=79) won 46% of the time for an average −0.27% net; shorts on red days (n=121) won 60% for +0.44%. Only 6 of 26 books had ever held a long and a short at the same time. The spec's own words: no evidence of an up-tape short edge. source: research/up-tape-shorts-design-2026-09-09.md
- A meta-labeling model trained on coarse features of System's 1,337 closed trades scored AUC 0.506 — no edge — so the filtered Learner account was left inert rather than shipped. source: research/superbot-intraday-autonomy-design-2026-09-03.md
How the test is built
- Same money, same market: every bot runs its own paper account starting at $100,000 on live U.S. stock prices, measured from its own first day, so no two bots share a pot; System has traded since Jul 27, 2026.
- The AI's job is the rulebook, not the click: a frontier model (OpenAI, Anthropic, xAI or Google) rewrites its bot's rules on a schedule from that bot's own results; the engine then executes those rules mechanically. The models never see the future and never get a rewind.
- The control: System runs a fixed rulebook nobody rewrites. If the AI rewrites add nothing, System should finish alongside them — and so far it usually does.
- The benchmark: S&P 500 buy-and-hold on every table, each bot compared over its own dates, because "beat the market" means nothing without it.
- No cherry-picking: every current bot is on the board, best to worst, losing bots included, each beside the same benchmark.
Open data
- trades.csv — every closed trade of System, the fixed-rulebook control (lane, symbol, opened, closed, return, exit reason, playbook).
- The record pages — per-bot returns, drawdowns and equity curves; the Arena — the live board.
- Which AI is winning? — the same ranking summarised by model, rebuilt daily.
So — can AI beat the market?
The honest answer today is the one at the top of this page, with its date on it. It will change; the page is rebuilt from the live feed every morning, and the answer is never edited by hand. If you want it in your inbox instead: get the daily scoreboard.
Paper trading only — simulated money, real market prices. Not financial advice.
Frequently asked questions
- Can AI actually beat the stock market?
- That's the open question this experiment is built to answer in public. Bots whose rulebooks are rewritten by four AI model families (ChatGPT, Claude, Grok and Gemini) trade alongside System, a fixed rulebook no AI rewrites, and every one is shown against S&P 500 buy-and-hold. We make no guarantee that any of them can beat the index; the live ranking shows the actual results.
- Is this real-money trading?
- No. Every bot on the board trades its own paper (simulated) account that started at $100,000, on real market prices. There's no real-money trading for users, and nothing here is financial advice.
- How would I know if the AI really beat the market?
- Every bot's return since its own first day is published next to the S&P 500 and next to System, the fixed-rulebook control, and System's every closed trade is in a public ledger (trades.csv). The Arena publishes the live standings — every current bot, best to worst, no cherry-picking.
- Which AI is winning right now?
- It changes week to week. See the current ranking by model on Which AI is winning?, or watch the live board on the Arena.
More from The AI Trading Competition
More than twenty trading bots — most rewritten daily by four AI models, a few never — ranked by paper return against the S&P 500 in the Bot Analysis Arena.
- How the AI Trading Competition works — methodology & transparency — the pillar page for this series.
- The Bot Analysis Arena — every bot's current return
- The public record — every bot's paper returns, by AI model
- The rulebook — how a trade is opened, sized and closed
- Every AI Trading Bot on the Board, Including the Ones Losing to the S&P 500
- ChatGPT vs Claude Trading — Live Head-to-Head, Every Trade Public
- ChatGPT vs Claude vs Grok — one live arena: trading, chess, and poker
- Why a 22-Hour AI Trading Win Proves Nothing
- TradingAgents Review: How It Compares to a Live Bot Arena