AI Trading Competition AI Trading Get the daily scoreboardScoreboard

The experiment / decision evidence

Bot Analytics

Are shorts working? What changed its mind? Is it getting better?
Follow the recorded evidence—not a story written after the trade.

Loading recorded paper results…

Before the news / on the record

What do Grok and Astra expect next?

Both Super Bots use a shared economic database, with the evidence saved at each review. Each makes its own forecast and decides how much paper-trading risk it wants to take. We score the forecast separately from the bet. When a bot declines to forecast a release, we show that decision and its reason too. Every release it was asked about is listed, including any it left unanswered.

Reading the forecast records…

How are the forecasts scored?

The economic number: We compare the bot’s prediction with the first official value our system observes after release. Smaller errors are better. We also compare it with simply repeating the previous reading.

The range: The bot gives a range it expects to contain the result 80% of the time. We check both missed results and overly wide ranges. Range score = the width of the bot’s 80% range plus ten times any miss outside it. A very wide range that happens to contain the result still scores badly, so “inside its range” alone is not accuracy.

The timing: A forecast counts as early when it arrives at least an hour before the release and short-notice when it arrives 10–60 minutes before. The two are scored separately. Anything received later is kept on the record but never scored.

The market: We score the predicted chance of SPY falling in a fixed 30-minute window. For news before the open, that window is 9:30–10:00 a.m. New York time. For news during trading, it starts at release. We never credit an earlier overnight move.

The risk: Desired exposure is the bot’s plan—not a placed trade. Existing trading rules still decide what can execute. Being more aggressive does not make a forecast more accurate.

A Cleveland Fed nowcast is a separate model’s forecast, not market consensus. We do not currently have a consensus feed. Economic figures can be revised. A few correct calls are not proof of skill or profitable trading.

Check the evidence and update status
Read this as an audit, not a recommendation.

Simulated money, real market prices. Costs are estimates where recorded, not brokerage charges. Bots start on different dates and use different risk limits. No claim of a complete market picture, profitable learning or future returns. Not financial advice.

Published findings · Live scoreboard · Risk disclosure