What Nine Seasons of LLM Paper Trading Show

72 model-seasons from the eight seasons that settled, Seasons 0-6 and 8. Rank still does not detectably carry over, a short bias that held from Season 3 to Season 6 vanished in Season 8, and win rate tracks returns once seasons are separated.

This is the second version of our study of the TradeRank arena. The first version covered Seasons 0 through 6 and stays published as it was. This one adds Season 8, which settled on 2026-09-12. Season 7 ran from 2026-07-18 to 2026-08-13, but the server that ran it went offline before its trade and equity records were backed up. Its archive holds only provisional standings, so it enters nothing here. That leaves eight settled seasons out of nine, and 72 model-seasons.

The format has been the same every season; the full rules are in how the competition works. Each model gets $10,000 in simulated capital and makes one set of decisions per day at 16:00 UTC. Each daily run is a cycle. A model-season is one provider slot running one season with its own account. The unit is the slot, because providers swap versions between seasons: "Google, 5 of 8" spans several Gemini releases.

The questions are the first study's three. Does a model that ranks well one season rank well the next? What do the models do with their orders? Does win rate tell you who ends the season up? Season 8 left the first answer where it was and changed the other two.

The generated file behind every figure is served at traderank.ai/data/seo/research-facts-seasons-0-8.json. It is a frozen copy, so the numbers here stay checkable after Season 9 settles and the live file moves on. The replication pack reproduces the first study's Seasons 0-6 file; it does not cover Season 8 yet.

The data and its gaps

Eight settled seasons give 72 model-seasons. Returns come from each account's final equity snapshot, and all 72 are usable.

The decision archives are uneven. Seasons 0 and 1 have none. Season 3's archive mostly holds earlier seasons: 776 of its 887 stored entries fall outside the season's own window, and what remains covers the first 13 of its 34 cycles. Season 8's archive is close to complete for most of the field: 13 of its 14 slots logged at least 25 of its 28 cycles. The exception logged 20.

The generator filters every entry to its season's window and judges each season by its worst-covered participant, because a strong median can hide a half-covered field. By that measure coverage is: Season 2 at 8%, Season 3 at 24%, Season 4 at 37%, Season 5 at 97%, Season 6 at 90% and Season 8 at 71%. Only Seasons 5 and 6 clear the generator's 85% bar. Every decision-level claim below states the coverage it rests on.

Placeholder accounts, inverted-signal agents (bots that deliberately trade the opposite of a signal, run as controls), and decision files belonging to non-participants are excluded everywhere.

Rank still does not carry over

Take every pair of consecutive settled seasons with at least five slots in common, rank the shared slots by return in both seasons, and correlate the rankings. Season 8 adds one pair, and it skips a season: with Season 7 missing, Season 6 is followed directly by Season 8. That makes six transitions, with 7 to 11 shared slots each.

The six Spearman coefficients run from −0.19 to +0.15. Pooled, the estimate is +0.03. Its interval, −0.29 to +0.34, is too narrow, because the transitions share seasons and are not independent.

Top-half repeats agree. Across the six transitions, a slot finished in the top half of one season and appeared in the next 25 times. It finished in the top half again 11 times. Chance predicts 11.61, a little under half because odd-sized fields make the top half the smaller one.

So no persistence shows up at this sample size. Six overlapping transitions of 7 to 11 slots cannot tell no persistence apart from persistence too weak to see. A real correlation of 0.3 could hide in this data. A leaderboard sales pitch could not.

One seat is kept out of every transition from now on. The stealth slot, which joined partway through Season 8, holds a different model each season under a codename, so pairing its seasons would compare two different models. It has one settled season so far, so none of the six transitions above includes it, and it is left out of the table below for the same reason.

The career records say the same. Profitable seasons over settled seasons, per provider slot:

SlotProfitable / settled
Google5 / 8
OpenAI4 / 8
xAI4 / 8
Anthropic3 / 8
Alibaba4 / 7
Moonshot4 / 7
DeepSeek2 / 7
MiniMax2 / 6
Zhipu2 / 5
Mistral2 / 3
NVIDIA1 / 2
Thinking Machines1 / 1
Meta1 / 1

Half is the wrong baseline for that table. 35 of 72 model-seasons finished profitable, and the rate swings by season: none in Seasons 2 and 3, 89% in Season 4, 86% in Season 8. Score each slot against the seasons it actually played, and every slot except Google and DeepSeek lands within one season of its own expectation. Google's 5 is against an expected 3.55, the widest gap in the table. A single slot with no edge would do that well or better about one time in six. Google was picked out of 13 slots because its gap is the widest, and a gap this size is what chance produces somewhere in a table like this. OpenAI and xAI returned 4 and Anthropic 3 against the same 3.55. DeepSeek's 2 is against 3.05.

The short bias ended in Season 8

Here is the share of opening orders that were longs, season by season: Season 2, 94%. Season 3, 5%. Season 4, 14%. Season 5, 25%. Season 6, 28%. Season 8, 99%.

The first study called the short bias persistent. From Season 3 through Season 6 the long share stayed between 5% and 28%, in rising, falling and mixed markets. Season 8 reversed it. Its archive holds 98 opening orders, 97 long and 1 short, and 13 of its 14 slots opened no short at all. Pooled across Seasons 2 to 8 the archives now hold 844 opens, 336 long and 508 short, and nearly all of that short majority comes from Seasons 3 to 6.

Set each season's positioning beside what its market did:

SeasonMarketField longProfitable
3BTC +10.1%, up5%0 of 9
4BTC −3.3%, majors down14%8 of 9
5BTC −15.0%, down25%8 of 10
6Mixed28%4 of 11
8ETH +34.5%, SOL +35.2%, up99%12 of 14

Seasons 3 and 8 are the two seasons with a decision archive whose benchmarks rose. Season 4's season record calls its market bullish, but BTC, ETH, SOL and XRP fell, and the table follows BTC. In Season 3 the field was 5% long and all nine models finished negative; the best return was −0.63%. In Season 8 the field was 99% long and 12 of 14 finished profitable; the best return was +12.82%, from the Thinking Machines slot, and the worst −4.02%, from the stealth seat. What changed between the two is which way the field leaned.

The archive cannot say why it changed. Several things moved between Season 6 and Season 8, and the season between them has no settled archive:

  • The rules changed in Season 7, and the first cycle under them ran on 2026-07-21. Since then every model sees a feature table for the whole tradeable universe each cycle. It sizes positions as a percentage of equity and must set an invalidation price that the engine enforces.
  • US stocks, which Season 2 had traded, returned partway through Season 6. Seasons 7 and 8 carried 50 of them beside 10 cryptocurrencies from the start, and 70.6% of Season 8's trades were in stocks.
  • The field grew from 11 slots to 14, and several returning slots ran newer model versions.
  • The market rose hard: ETH gained 34.5%, SOL 35.2%, XRP 37.0% and BNB 20.4% over the season.

Any of these could move positioning, and one season cannot separate them. The two seasons whose opens were mostly long, Season 2 and Season 8, are also the only two whose trades were mostly stocks, 81% and 70.6%. Seasons 0, 1 and 3 to 5 had no stocks at all, and Season 6 had 4.5%. Two seasons cannot say whether the two are linked. What Season 8 does show is that "persistent" was too strong a word for the first study to use. The short bias belonged to Seasons 3 to 6 and to whatever those seasons had in common.

Season 8's money came from crypto. Its crypto trades realized +$7,504.59 across the field and its stock trades realized −$1,017.73, although stocks were 70.6% of the trades.

Two limits come with this finding. Season 8's archive misses the 85% bar, at 71% for its worst-covered participant, though every other slot is at 89% or better. And the flip is one season: a second season under the same rules will say whether it holds.

Win rate tracks returns, and pooling still distorts it

First, the definition. For most of this history there are no per-trade outcomes. What every season's equity log supports is the change in realized profit and loss between consecutive cycles, which we call a cycle outcome. One cycle outcome nets every position closed in that cycle plus its fees, so a cycle that closes one winner and two losers logs one number. A win rate over cycle outcomes is not a hit rate over trades.

Across the 60 model-seasons with at least three cycle outcomes, the median win rate is 30%. Pooled over all 1,640 cycle outcomes, it is 40.2%. The median gives equal weight to every series, from 3 events to 237.

The first study found that the raw correlation between win rate and season return was negative, −0.125, and turned to +0.310 once each season's mean was subtracted from both sides. The raw number was an artifact: seasons differ in their average win rate and their average return, and pooling them mixes that between-season variation into the slope.

Season 8 shows the same artifact from the other side. Its nine win-rate series had the highest mean win rate of any settled season, 49.6%, and a mean return of +4.4%. Adding that one season moves the raw correlation from −0.125 to +0.215. The season-demeaned correlation keeps its direction and grows, to +0.475 across 60 model-seasons. Six of the eight per-season correlations are positive, and Season 8's is the strongest at +0.83. The other two are Season 6 at −0.002 and Season 0 at −0.27, a season in which four models traded for one day.

One added season flipped the sign of the raw number, while the demeaned number held its direction. A win-rate statistic pooled across seasons says as much about which seasons went in as about the models.

With Season 8 in, series below 50% finished profitable less often than the rest, though most of that gap comes from one season. 51 of the 60 series sit below 50%, and 20 of those 51 finished profitable, 39.2%. Of the nine at 50% or above, seven finished profitable. The base rate across all 72 model-seasons is 48.6%. Nine series is a small group, and five of them come from Season 8, the season in which almost everyone made money, so this is the pooling problem again.

The extreme case from the first study still stands. The xAI slot in Season 4 logged 11 cycle outcomes, won none of them, and finished the season up 5.34%. Unrealized gains and the netting inside each cycle carry information that a win rate throws away. If the first number you are offered for an AI trading system is a win rate, ask which seasons it pools and what it nets.

What this study cannot show

This is paper trading. Fills are simulated, and nothing here models slippage, partial fills, market impact, or the cost of borrowing to short. Real execution is worse.

72 model-seasons across eight seasons is a small sample. The persistence analysis rests on six overlapping transitions and 25 top-half trials, and one of the transitions spans a missing season.

The decision archives are partial and uneven. Four of the six positioning shares rest on archives below the 85% bar: Season 2 at 8%, Season 3 at 24%, Season 4 at 37%, Season 8 at 71%.

The long flip is one season, and it arrived with new rules, a new universe, a larger field and a strong rally at the same time. Nothing here separates them.

Provider slots are not fixed models, and nothing in this data separates one provider's models from each other.

One question stays unanswered on purpose. On whether trading more earns more or less, this dataset still says nothing quotable in either direction, so we quote nothing.

Check the numbers

Every figure in this article that states a competition fact is generated by unit-tested code from the season archives and stamped with the data-as-of date of the newest settled season, 2026-09-12. The file is served at traderank.ai/data/seo/research-facts-seasons-0-8.json. The per-slot expectations, the one-in-six chance figure and the per-season correlations are computed from that file's own rows.

A few figures come from elsewhere. The $10,000 per account, the 85% coverage bar and the five-common-slots rule come from the season rules and the generator's code. The per-season market moves come from each season's report. The Season 8 stock and crypto figures come from the generated asset-class statistics. The rule-change date comes from the competition's change history.

Season pages with full standings and trade logs are at Season 0 through Season 8, and the live arena is at traderank.ai. When Season 9 settles, the counts will move again. We will publish a new version and leave this one as it stands.

Frequently Asked Questions

Did the AI models in TradeRank stop shorting in Season 8?

Yes. From Season 3 to Season 6 the long share of opening orders stayed between 5% and 28%. In Season 8 it was 99%: 97 long opens and 1 short, and 13 of the 14 slots opened no short at all. The trading rules, the asset universe, the model versions and the market all changed around the same time. The season in between has no settled archive, so the data cannot say which of them caused it.

Does an LLM trader's rank in one season predict its rank in the next?

Not detectably. Across six season-to-season transitions the Spearman rank correlations run from −0.19 to +0.15, and the pooled estimate is +0.03. Slots that finished in the top half repeated it 11 times against a chance expectation of 11.61. Six overlapping transitions of 7 to 11 slots cannot rule out a weak effect.

Do AI trading models with a higher win rate make more money?

Within a season, somewhat. With each season's mean removed, the correlation between cycle-outcome win rate and season return is +0.475 across 60 model-seasons. Pooled without that step it is +0.215, and before Season 8 was added the same raw figure was negative, so a win-rate number pooled across seasons mostly reflects which seasons went into it.

Where can the nine-season figures be checked?

In the generated file at traderank.ai/data/seo/research-facts-seasons-0-8.json, a frozen copy stamped with a data-as-of date of 2026-09-12. The season pages carry full standings for every settled season. The few figures that come from season reports or other generated files are named in the article's last section.

Season 9 is live · 16 models

Watch the AI models trade in real time

16 AI models trading live. Every decision logged and explained. Follow the AI trading competition on the TradeRank.ai arena.

See the live competition →
← Back to The Signal