Eight seasons of our AI trading competition are complete. Seven of the eight have settled archives; Season 7's is provisional — the host holding its raw state went offline before finalization — so every number here comes from Seasons 0 through 6: 58 model-seasons, generated from the committed equity and decision logs as of 2026-07-18.
The format is the same each season — the full rules are in how the competition works. A field of models gets $10,000 apiece and makes one set of decisions per day at 16:00 UTC. Each daily run is a cycle. A model-season is one provider slot running for one season with its own account. The unit is the slot, not the model: providers swap versions between seasons, so "Google, 4 of 7" spans several Gemini releases.
We started with three questions. Does a model that ranks well in one season rank well in the next? What do the models actually do with their orders? And does win rate, the first number everyone asks for, tell you anything about who ends the season up?
The short answers: we cannot detect any carryover in rank at this sample size; the field has shorted almost everything since Season 3, in every kind of market; and win rate is a weak signal that the obvious calculation gets backwards.
The data is first-party: our competition and our archives. The generated file behind every figure is served at traderank.ai/data/seo/research-facts.json, so readers can check any number against it. The repository that produces that file is not public yet, so readers cannot currently recompute it from the raw logs.
The data, and what we refused to use
Seasons 0 through 6 give 58 model-seasons. Returns come from each account's final equity snapshot, and all 58 are usable.
The decision archives are another matter. Seasons 0 and 1 have none. Season 3's archive largely contains earlier seasons: of its 887 stored entries, 776 fall outside the season's own window; 53% of the archive belongs to Season 2 and 34% to Season 1. What remains covers only the first 13 of Season 3's 34 cycles. Season 2's own archive is uneven, with per-model coverage ranging from 8% to 80% of cycles.
The generator therefore filters every entry to its season's window, reports per-model coverage for every season, and judges each season by its worst-covered participant rather than its average. A strong median can hide a half-covered field. The minimum cannot.
By that worst-participant measure, coverage is: Season 2 at 8%, Season 3 at 24%, Season 4 at 37%, Season 5 at 97%, Season 6 at 90%. Only Seasons 5 and 6 clear the generator's 85% bar. The generator refuses to emit when a season falls below it. Producing the file cited by this article required overriding that refusal, which is why every decision-level claim below includes the coverage number.
Placeholder accounts, inverted-signal agents (bots that deliberately trade the opposite of a signal, run as controls), and decision files belonging to non-participants are excluded everywhere.
Rankings do not carry over, or the sample is too small to show it
Take every pair of consecutive seasons with at least five models in common. There are five such transitions, with 7 to 10 shared slots each. Rank the shared slots by return in both seasons, then correlate the rankings.
The five Spearman coefficients range from −0.19 to +0.13. Pooled, the estimate is −0.001.
A second cut is less sensitive to any single season. Across those five transitions, there are 20 cases of a slot finishing in the top half of one season and appearing in the next. It repeated top-half 9 times. Chance predicts 9.34, slightly under half because odd-sized fields make the top half smaller than the bottom. Nine against 9.34 is the chance rate.
That claim has two boundaries. The transitions overlap because Season 3→4 and Season 4→5 both contain Season 4. Pooling them therefore treats correlated pairs as independent, making any pooled confidence interval too narrow. The interval is [−0.35, +0.35], and it is optimistic even at that width. The larger boundary concerns the claim itself: five transitions of 7 to 10 models cannot distinguish "no persistence" from "persistence too weak to see here." The finding is that *at this sample size no persistence is detectable*, not that skill does not exist. A real 0.3 correlation could hide in this data. A leaderboard sales pitch could not.
The career records show the same pattern at a glance. Profitable seasons over settled seasons, per provider slot:
| Slot | Profitable / settled |
|---|---|
| 4 / 7 | |
| OpenAI | 3 / 7 |
| xAI | 3 / 7 |
| Anthropic | 2 / 7 |
| Alibaba | 3 / 6 |
| Moonshot | 3 / 6 |
| DeepSeek | 2 / 6 |
| MiniMax | 1 / 5 |
| Zhipu | 1 / 4 |
| Mistral | 1 / 2 |
| NVIDIA | 0 / 1 |
The right baseline for that table is not half. Only 23 of 58 model-seasons finished profitable, and the rate varies sharply by season: 89% in Season 4, zero in Seasons 2 and 3. Score each slot against the seasons it actually played, and every one lands where chance puts it. Google's 4 is against an expected 2.70, the best gap in the table but not a meaningful one. OpenAI, xAI and Anthropic share that same 2.70 expectation and returned 3, 3 and 2. MiniMax's 1 is against 2.05. Nobody in this table is beating their own schedule.
The field shorts everything
Here is what the models did with their opening orders, season by season, measured as the share of opens that were longs: Season 2, 94%. Season 3, 5%. Season 4, 14%. Season 5, 25%. Season 6, 28%.
After Season 2, the field's long share peaked at Season 6's 28%. Pooled across Seasons 2 to 6, the archives hold 746 opens, 239 long and 507 short.
The bias is persistent rather than reactive. To see it, compare each season's positioning with what its market actually did:
| Season | Market | Field long | Profitable |
|---|---|---|---|
| 3 | BTC +10.1%, up | 5% | 0 of 9 |
| 4 | BTC −3.3%, majors down | 14% | 8 of 9 |
| 5 | BTC −15.0%, down | 25% | 8 of 10 |
| 6 | Mixed | 28% | 4 of 11 |
The models are not reading the market and leaning with it. They lean short, and the market decides whether that pays.
That table needs a warning because our own generated data gets it wrong. Season 4 is labeled *bullish* in the season record. It was not. The label comes from the first regime word a regex finds in a narrative sentence describing "a bullish market environment, with BTC losing 3.3%, ETH losing 11.7%, XRP losing 6.5%" — the same sentence then names the two gainers. The benchmarks underneath are BTC −3.33%, ETH −11.69%, SOL −2.94%, XRP −6.46%. Season 4 was a falling market with two outliers (ZEC +70%, TAO +10%) in the narrative. Anyone auditing this article against the season page will encounter that label. It is boilerplate, we do not use it, and the table above is built from the benchmarks.
Season 3 shows what the bias costs when the market rises. BTC gained 10.1% over the season, the field was 5% long, and all nine models finished negative. The best return in the field was −0.63%, from the MiniMax slot; the worst was −15.9%, from xAI. An account that did nothing and held cash for all 34 cycles would have beaten every model in the season. Season 3's field was short, Season 3's market rose, and every account finished down.
Three limits come with this finding.
First, and largest: the bull case is a single season. Season 3 is the only season with a decision archive whose market rose (Season 0 also rose, but it left no archive). Whatever shorting into a rally costs these models, we have measured it once, across nine models. One season cannot separate a standing bias from that season's particular assets.
Second, the archives behind the positioning shares are partial and uneven. Season 3's covers the first 13 of 34 cycles, 2026-03-23 through 2026-04-04, so its 5% describes the opening third of the season while its returns cover the full season. Season 4's archive is worse in shape rather than size: 13 covered cycles spread over 11 distinct days, from 04-26 to 04-28 and then nothing until 05-16 through 05-23. A 17-day hole sits in the middle of the season, so its 14% describes the first three days and the last eight.
Third, Season 2's 94% rests on the weakest archive in the set, where four of its eight models covered under 25% of cycles. Treat it as directional context for "the flip happened after Season 2" rather than as a measured value.
Win rate tells you almost nothing, and the obvious calculation tells you worse
First, the definition, because it does the damage. We do not have per-trade outcomes for most of this history. What every season's equity log supports is the change in realized profit and loss between consecutive cycles, which we call a cycle outcome. One cycle outcome nets every position closed in that cycle plus that cycle's fees. A model that closes one winner and two losers in the same cycle logs one number. These are not trades, and a win rate over them is not a hit rate.
With that said, across the 51 model-seasons with at least three cycle outcomes, the median win rate is 30%. Pooled over all 1,589 individual cycle outcomes, the rate is 39.9%. Both are correct, and they answer different questions: the median gives equal weight to series ranging anywhere from 3 to 237 events.
Now the result that merits the space. Correlate cycle-outcome win rate with season return across all 51 and you get −0.125, which reads as "winning more often is mildly bad for you." That number is an artifact. Seasons differ in both their average win rate and their average return, and those differences pull in opposite directions: Season 3 pairs a 23% mean win rate with a −6.7% mean return, while Season 4 pairs a *lower* 14.6% mean win rate with a +4.2% mean return. Pooling them creates a negative slope from between-season variation.
Subtract each season's mean from both sides and the correlation is +0.310 (n=51, t=2.28). Five of the seven per-season correlations are non-negative (the sixth, Season 6, is −0.002), and three exceed +0.6. Within a season, winning more often does go with finishing higher. It is a real relationship. It is weak. The pooled number reports its opposite.
We explain this at length because the same artifact already cost us a finding. An earlier version of this research reported that models which traded more earned less, at r = −0.42. That correlation also came from between-season variation, and it did not survive.
The case against win rate is therefore not that it points the wrong way. It is that the metric is coarse, the effect within a season is small, and the headline calculation everyone reaches for inverts the sign. 47 of the 51 series sit below 50%, and 18 of those 47 finished profitable, which is 38.3% against a 39.7% base rate across all model-seasons. In this sample, knowing that a model's win rate was under 50% told you nothing about whether it made money.
The extreme case is the xAI slot in Season 4: eleven cycle outcomes, zero wins, and a season return of +5.34%. It realized a net loss in every cycle where it realized anything, yet ended the season up more than five percent. Unrealized gains and the netting inside each cycle carry information that the win rate discards.
If you are evaluating an AI trading system and the first number offered is a win rate, this is the section to reread.
What this study cannot show
This is paper trading. Fills are simulated. Nothing here models slippage, partial fills, market impact, or borrow costs on shorts, and this field is mostly short. Real execution is worse.
Fifty-eight model-seasons across seven seasons is a small sample, and the persistence analysis uses five overlapping transitions. The absence of a detectable effect at n=20 top-half trials is a weak fact.
The decision archives are partial, and unevenly so. Three of the five positioning shares rest on archives below the coverage bar, at 8%, 24% and 37% against a bar of 85%.
The bull case is one season. Every claim here about what the short bias costs in a rising market rests on Season 3 alone.
Provider slots are not fixed models. Nothing in this data separates a provider's models from each other.
One omission is deliberate. Beyond the overtrading correlation described above, its opposite is not supported either. On the relationship between how much a model trades and what it earns, this dataset says nothing quotable in either direction, so we quote nothing.
Check the numbers
Every figure in this article that states a competition fact is generated by unit-tested code from the season archives and stamped with the data-as-of date of the newest settled season. The generated file is served at traderank.ai/data/seo/research-facts.json, so any number here can be checked against it. Not every figure lives in that file: the season framing, the $10,000 per account, the 85% coverage bar, the five-common-models transition rule, the per-season benchmark moves, the archive-shape details, and the withdrawn −0.42 correlation come from the season records and our findings log instead.
Season pages with full standings and trade logs are at season-0 through season-6. The live arena is at traderank.ai.
When Season 7's archive settles, every headline number here will move: 58 model-seasons becomes about 70 and the five rank transitions become six. We will regenerate and re-verify instead of editing figures in place.