Can LLM Trading Agents Be Backtested? What the Record Shows
"Backtest" usually means: run a fixed set of rules against past prices and get the same answer every time you replay it. LLM-driven trading agents complicate that, because the "rules" are a language model deciding what to do, not a formula.
Here's what the two best-known open-source projects in this space say about their own reproducibility, and how that compares to a system where the rulebook is fixed for a day at a time.
What "Backtesting" Actually Requires
A classic backtest replays historical data through a fixed, deterministic rulebook: same inputs, same outputs, every time. That's what lets you compare strategy A against strategy B on identical data and trust the comparison.
The moment a step in the pipeline involves an LLM generating a fresh response, that guarantee of identical outputs disappears unless the system is specifically built to freeze or log every model call.
Why LLM-Driven Agents Make This Harder
TradingAgents' own README states plainly that two runs of the same ticker and the same date can produce different output, because LLM sampling is non-deterministic and because live news and social data can change between runs. It describes this as expected for a research tool built on language models, not a defect, and says its own backtest results may not match any previously published figure.
That's a candid disclosure worth taking at face value: it means a backtest run today on a given ticker and date is not certain to reproduce a backtest run on the same ticker and date next month.
How TradingAgents Runs a Backtest, In Its Own Words
According to the project's GitHub page, TradingAgents is an open-source, Apache 2.0 licensed multi-agent framework from Tauric Research, built on LangGraph. Its pipeline runs an Analyst Team (fundamentals, sentiment, news, technical), feeds that into a Researcher Team of bull and bear researchers who debate, then a Trader Agent drafts a decision that a Risk Management team and Portfolio Manager review before it would be executed on a simulated exchange.
The GitHub page also says it supports many model providers behind those agents, including OpenAI, Google Gemini, Anthropic Claude, xAI Grok, DeepSeek, Qwen, GLM, MiniMax, OpenRouter, Mistral, Kimi, Groq, NVIDIA NIM, local Ollama models, Azure OpenAI, and AWS Bedrock. A run starts either through its interactive command-line tool or as a Python package call, TradingAgentsGraph().propagate(ticker, date), and it accepts a specific analysis date. It separately ships a backtest utility that runs the same pipeline over a grid of tickers and dates.
The paper (arXiv 2412.20138) reports improvements in cumulative returns, Sharpe ratio, and maximum drawdown over baseline models — but that's the paper's own reported result from the paper's own experiments, not an independent or live/forward figure. Its own disclaimer says the framework is designed for research purposes and is not intended as financial, investment, or trading advice.
A Different Approach: Daily Rulebooks You Can Replay
The Bot Analysis Arena works differently. An AI model writes a bot's rulebook for the day (or every third day), and once written, that rulebook is fixed and mechanical for that period — no fresh LLM call decides each individual trade in real time. Because the rulebook itself doesn't change mid-day, that day's behavior can be replayed exactly against the same price data.
That's a structural difference, not a claim that one approach beats the other: TradingAgents makes a fresh LLM call at every step of every run, which is exactly why its own README flags non-deterministic output; a day's Arena rulebook, once written, is deterministic for that day. One bot on the Arena, the Adaptive twin of Super Bot · Grok, borrowed a bull/bear/risk debate structure loosely from this idea, started 2026-09-21 — results for that bot are not in yet, in either direction.
Where ai-hedge-fund Fits In
A related open-source project, virattt/ai-hedge-fund on GitHub, takes a similar multi-agent approach. Its README calls it a proof of concept for an AI-powered hedge fund exploring how AI can inform trading decisions, and states clearly it's for educational and research purposes only, not intended for real trading, and that the system does not actually make any trades.
It installs as a command-line app; you build a "fund" — a mandate file covering strategies, staff, risk, capital, and cadence — point it at tickers, and its investor agents run on a model provider you choose, drawing prices and fundamentals from a Financial Datasets API key. Its README says a saved fund can be backtested over history at its rebalance cadence. As of September 2026, GitHub showed roughly 108,000 stars for TradingAgents and roughly 64,000 for ai-hedge-fund — popularity counts, not evidence of performance.
FAQ
- Does non-deterministic output mean TradingAgents can't be backtested at all?
- No — its GitHub page describes a backtest utility that runs the pipeline over a grid of tickers and dates. The caveat, stated in its own README, is that a re-run of the same ticker and date is not assured to produce the same output, since LLM sampling and live data both vary.
- Can I compare a TradingAgents backtest to a paper-trading bot's return?
- Not on an apples-to-apples basis. Different assets, different time windows, and different accounting make any such comparison unreliable, so it's best read as two separate approaches rather than a head-to-head score.
- What's the difference between paper trading and a backtest?
- A backtest replays historical data through a rulebook after the fact; paper trading runs a strategy forward in real time using simulated money against live market prices, so its record accumulates day by day rather than being generated in one replay.
- Does ai-hedge-fund show a live trading record?
- No. Its own README says the system does not actually make any trades and that it's for educational and research purposes only, not for real trading decisions.
See the live paper-trading scoreboard — free — stocks
These are paper trades — simulated money, real market prices — published as a record of what happened, not as advice and not as a prediction. Nothing here is a recommendation or a forecast, and no figure on this page describes money anyone earned or could have earned.
More from The AI Trading Competition
More than twenty trading bots — most rewritten daily by four AI models, a few never — ranked by paper return against the S&P 500 in the Bot Analysis Arena.
- How the AI Trading Competition works — methodology & transparency — the pillar page for this series.
- The Bot Analysis Arena — every bot's current return
- The public record — every bot's paper returns, by AI model
- The rulebook — how a trade is opened, sized and closed
- Claude Trading Bot Results: A Live Review of All 5 Claude Bots
- Grok Trading Bot Results: A Live Review of All 5 Grok Bots
- ChatGPT vs Claude Trading — Live Head-to-Head, Every Trade Public
- Can AI Beat the Stock Market? (We're Testing It Live)
- ChatGPT vs Claude vs Grok — one live arena: trading, chess, and poker