Season 3 Final: All 9 Premium AI Models Lost Money

BTC gained 10.1% over the 34-day window. Nine frontier models (GPT-5.4, Claude Opus 4.6, Gemini 3.1 Pro, Grok 4.20, and five others) combined to lose 60.21 percentage points. Here is what the final data says.

-0.63%196 trades

Nine Frontier Models, Identical Inputs, All Lost

Season 3 ended on April 26, 2026. Nine premium frontier models traded $10,000 each across 37 cryptocurrencies for 34 daily cycles. The archive contains 196 closed trade records and $502.36 in fees.

Not one model finished positive. Returns ranged from -0.63% to -15.90%, and summed to -60.21 percentage points across the field. Over the same window, BTC gained 10.1%, ETH gained 8.2%, ARB gained 36.8%, and BLUR gained 56.9%.

The direct result is narrow but clear: under this prompt, asset universe, and market window, all nine models lost while the BTC benchmark rose.

Warning

This article is for educational and entertainment purposes only. Nothing here is financial advice. Trades described are from a simulated competition using live market prices and simulated capital; no real money is at risk. Past simulated performance does not predict future results. The numbers below are the official end-of-Season-3 standings as published in the Season 3 report.

Key Insight

TL;DR: MiniMax M2.5 won Season 3 at -0.63%, the smallest loss in a 9-model field where every model finished negative. Gemini 3.1 Pro took silver at -2.64%, validating our mid-season call that Gemini was the best of the Big Four. Grok 4.20 finished last at -15.90%, a 0.06pp deterioration in the final three days. Higher activity had a moderate association with larger losses in this season, not a universal rule.

Final Standings

Here are the official Season 3 final standings as of April 26, 2026. Every number is sourced from the Season 3 competition report and verifiable against the immutable trade log. The ΔPos column is the change versus our mid-season snapshot on April 22 (cycle 31).

Season 3 Final Standings

RankModelProviderReturnWin RateTradesMax DDFeesΔPos vs. Apr 22
1MiniMax M2.5MiniMax-0.63%20.0%104.11%$24.89
2Gemini 3.1 ProGoogle-2.64%31.8%227.04%$50.03
3Qwen 3.5 PlusAlibaba-2.73%30.4%237.96%$55.70
4GPT-5.4OpenAI-5.00%23.5%179.69%$50.50
5Kimi K2.5Moonshot-6.35%26.1%2310.21%$64.44
6Claude Opus 4.6Anthropic-7.61%8.3%248.82%$53.37
7GLM-5Zhipu AI-7.67%20.0%209.46%$46.76
8DeepSeek V3.2DeepSeek-11.66%22.9%3512.98%$91.35
9Grok 4.20xAI-15.90%27.3%2219.72%$65.32

Final-cycle ranks held the mid-season order all the way through. Trade count and return were moderately associated in this nine-model season: Pearson correlation was about -0.56 and Spearman rank correlation about -0.45. MiniMax had the fewest closed trades (10) and the smallest loss; DeepSeek had the most (35) and finished eighth. Several middle-ranked models break a simple ordering, and the complete cross-season sample is nearly flat, so this is a season-specific observation rather than a transferable law.

Data Point

Aggregate Season 3 figures. 9 models, 196 closed trade records, $502.36 in fees, -60.21 pp summed return, $539,735 in total notional volume across 37 cryptocurrencies. Daily cycles at 16:00 UTC. Starting capital: $10,000 per model. 0 of 9 models finished positive. BTC returned +10.1% in the same window.

What Changed in the Final Three Days

Our mid-season piece was published on April 23 with three days left in the season. The headline read "Gemini Beat Grok by 13 Points" and the final standings widened that gap to 13.26 points. The mid-season ordering held. Three things shifted in the closing cycles.

First, Grok deteriorated by 0.06 pp in the final three days, from -15.84% to -15.90%. After 31 days of large losses, three days of marginal additional damage. The loss rate was decelerating. Grok's worst trades were already on the books well before the close.

Second, Gemini's lead over Qwen narrowed to 0.09 pp from 0.24 pp. Gemini lost 0.06 pp while Qwen gained 0.09 pp. The two were effectively tied through the close, separated by less than a day's worth of single-trade variance.

Third, MiniMax preserved its lead with the smallest tail risk in the field. The mid-season piece reported MiniMax at -0.62% with 19 trades. The final report shows it at -0.63% with 10 trades, a slightly worse return on a much lower trade count. The trade-counting discrepancy is worth flagging: the published Season 3 report counts only fully-closed round-trip trades for its standings, whereas the mid-season snapshot counted entries individually. Both are defensible; we surface the difference here so the numbers reconcile.

Nothing in the closing cycles changed the season's story. The bear trap that defined Season 3 had already sprung. By April 22, the damage was priced in across the entire field.

Data Point

Final-three-days breakdown for each model (return change vs. cycle 31): MiniMax -0.01 pp · Gemini -0.06 pp · Qwen +0.09 pp · GPT-5.4 -0.15 pp · Kimi -0.10 pp · Claude -0.15 pp · GLM-5 +0.03 pp · DeepSeek -0.08 pp · Grok -0.06 pp. The closing days were stable across the field. No model rallied; no model crashed. The standings were settled before the final cycles ran.

MiniMax M2.5: The Quietest Win on Record

Final: -0.63% | Win Rate: 20.0% | 10 Trades | Max Drawdown: 4.11% | Fees: $24.89

MiniMax M2.5 won Season 3 by trading less than half as many times as the field average. Its 10 round-trip trades are the lowest in the competition by a wide margin (DeepSeek made 35, Claude made 24, even fellow patient model Gemini made 22). MiniMax also had the lowest fee bill ($24.89, less than half of Gemini's $50.03), the lowest max drawdown (4.11%, with second-lowest Gemini at 7.04% nearly 1.7x worse), and the smallest single-trade losses.

The model's voice in its reasoning logs is consistent across the season: short paragraphs, narrow conviction, frequent decisions to stay flat. When MiniMax couldn't see a setup, it didn't manufacture one. When it took a position, the size was modest and the stop was tight. "No symbols meet the entry criteria — stay flat" appeared in roughly half of its cycle assessments.

MiniMax was unremarkable in Season 2 — it finished 7th of 13 at -1.05%, mid-pack, with a middling 4.96% max drawdown. Season 3 is the breakout: the lowest drawdown in the field (4.11%) and an outright win. A single clean season in one market regime is a small sample, but the pattern it set is the cleanest read we have on what works in this competition: be selective, stay small, accept the no-trade outcome as a valid output.

The negative return matters. MiniMax did not generate alpha in Season 3. It preserved capital while every other model bled out. In a season where preservation was the only winning approach, that was enough.

Cross-timeframe alignment is broken across the majors — daily reversed bullish on most names while weekly remains bearish. Per the strategy framework, weekly/daily disagreement disqualifies new entries. No symbols meet the composite >= 50 threshold with directional agreement. Maintain existing AAVE short with stop trail; otherwise stay flat.

MiniMax M2.5MiniMax M2.5's typical Season 3 cycle output: short, structurally constrained, and willing to take no action. Across 34 cycles, MiniMax's average reasoning length was the lowest in the field by a wide margin.

Gemini 3.1 Pro: The Big Four Winner, Confirmed

Final: -2.64% | Win Rate: 31.8% | 22 Closed Records | Max Drawdown: 7.04% | Fees: $50.03

Gemini was the best Big Four model: 2.36 points ahead of GPT-5.4, 4.97 ahead of Claude, and 13.26 ahead of Grok. Its 31.8% win rate was the highest in the final table, narrowly above Qwen's 30.4%.

That does not establish Gemini as universally best. It shows that Gemini lost least among the famous four in this prompt, universe, and regime. Season 4 remained crypto-only rather than adding equities, so it did not run the transfer test originally anticipated.

Grok 4.20: The Loss Decelerated, the Story Did Not Change

Final: -15.90% | Win Rate: 27.3% | 22 Closed Records | Max Drawdown: 19.72% | Fees: $65.32

Grok finished last, 4.24 points behind DeepSeek. Its cumulative HYPE losses were about $790—roughly half its total dollar loss—with a worst individual closed record of -$563, about 35% of the total rather than half by itself.

The 19.72% max drawdown and repeated HYPE exposure make concentration a credible explanation for much of this run. They do not prove a permanent Grok trait; the evidence is one model version under one regime and prompt.

The Win-Rate Floor: Claude Opus 4.6 at 8.3%

Final: -7.61% | Win Rate: 8.3% | 24 Trades | Max Drawdown: 8.82% | Fees: $53.37

Claude Opus 4.6 closed Season 3 with the lowest win rate in the competition by a substantial margin: 8.3% versus the field average of 23.4%. Of 24 closed trades, only 2 were profitable. The mid-season piece flagged this as Claude's most striking single statistic; the final data confirmed it.

Claude wrote the most thorough, most internally consistent reasoning in the field across Season 3. It also produced the worst hit rate. Across Season 2 and Season 3, Claude finished mid-pack-or-worse in both, with both seasons showing a similar over-conviction profile. That repeated profile is a hypothesis for further seasons, not enough evidence to separate a model effect from regime and prompt effects.

Whether this is fixable through prompt engineering is the most interesting open question Claude raises. Our system prompt is identical across all nine models. A prompt tuned specifically to reduce Claude's tendency to double down on initial theses would test whether the conviction problem is the model or the prompt. Season 4 will run with the same prompt across the field, so we will not get that test this season; it is on the roadmap.

Trade Frequency Was Associated With Loss, Not Destiny

Closed-record count and final return had a moderate negative association: Pearson about -0.56 and Spearman about -0.45. MiniMax paired the fewest records with the smallest loss, while DeepSeek paired the most with the second-largest loss.

The relationship was not a rule. GPT-5.4 traded relatively little but finished mid-pack, and models with similar counts produced very different returns. Fees make activity mechanically costly, but direction, sizing, and concentration also matter. The evidence supports monitoring turnover as one risk factor—not calling it a leading indicator or causal explanation.

Key Insight

Within Season 3, higher trade counts were moderately associated with larger losses, but the relationship was not deterministic and does not hold as a universal rule across the complete 22-model-season sample. Treat it as a season-specific observation to retest.

What Season 3 Tells Us About Premium AI as a Category

Season 3 asked whether replacing the budget model tiers used in Season 2 with premium counterparts would produce better trading. Prompts, data, constraints, and modeled fees were standardized within Season 3, but the market window and model versions also differed from Season 2, so this was not a controlled price-tier experiment.

Going premium did not produce better outcomes in this run. Across Seasons 1 and 2, roughly 23% of model-seasons finished positive; Season 3 produced none. Opus 4.6 and Grok 4.20 finished sixth and ninth.

Several explanations remain possible: the bear-to-bull whipsaw may have been unusually hostile; the shared prompt may have interacted differently with premium models; or trading performance may simply not track general benchmark tier. Season 4 changed the regime and most model versions again, so it adds evidence without isolating those factors.

What Happened Next in Season 4

Season 4 kept daily crypto trading but narrowed the universe from 37 assets to seven. MiniMax won again, all nine models beat a falling BTC benchmark, and a 70.17% ZEC rally dominated several books. That completed result is in the Season 4 final report; it should not be read as the equity experiment this article originally anticipated.

The Closing Statement

Season 3 was a standardized run of premium AI models: zero of nine finished positive, the winner ended at -0.63%, and BTC gained 10.1% during the same window.

This is not proof that AI cannot trade. Two prior seasons produced positive finishers, including Season 2's contrarian sweep. It shows that results remain regime-, prompt-, and sample-dependent in this autonomous simulated setup.

The completed Season 4 report records what happened next. Full Season 3 data, including model trade logs and decision history, lives at /competitions/season-3, while the cross-season benchmark keeps the broader context visible.

For Season 3 mid-season context: Gemini Beat Grok by 13 Points: Season 3 Flips the AI Trading Narrative, the April 23 prediction piece this article validates and extends.

What came next: Season 4 Final: All 9 Premium AI Models Beat BTC — the same nine model families, mostly upgraded versions, in a falling market.

For longer-horizon context: Can AI Beat the Market? Two Seasons of Data Say It Depends and 5 Lessons from 1,782 AI Trading Decisions. The LLM trading benchmark tracks the cross-season dataset, and reports preserve each cycle.

For activity and payoff context: One AI Wins 17% of Trades. Another Wins 81%. Here's Why Both Are Losing and Why Reverse Kimi Was Worse Than Doing Nothing examine fee drag and payoff shape without turning trade count into a universal rule.

For methodology: /how-it-works covers the full setup, scoring system, and Season 3 risk rules.

Frequently Asked Questions

Who won Season 3 of TradeRank's AI trading competition?

MiniMax M2.5 won Season 3 with a -0.63% return, the smallest loss in a field where all nine models finished negative. It also had the lowest max drawdown (4.11%), lowest fee bill ($24.89), and fewest closed trade records (10).

Did any AI model finish Season 3 with a positive return?

No. Returns ranged from MiniMax at -0.63% to Grok at -15.90% and summed to -60.21 percentage points across the field, while BTC gained 10.1% in the same 34-day window.

Why did all nine AI models lose money when BTC was up 10%?

Season 3 launched into a textbook bearish technical setup across most major altcoins. All nine models loaded short positions on the same setups in the first 11 days, then got trapped when the market reversed into a counter-trend rally that lifted ETH +8.2%, AAVE +7.3%, IMX +20.2%, and several other shorted assets. The herd was on the same side of the trade and the market punished them together. Eight of the nine models hit their equity peak on or around April 2, then bled together for the rest of the season.

How did Gemini rank among the Big Four in Season 3?

In Season 3 final standings: Gemini 3.1 Pro was the best of the Big Four at -2.64%, followed by GPT-5.4 (-5.00%), Claude Opus 4.6 (-7.61%), and Grok 4.20 (-15.90%). Gemini also posted the second-highest win rate in the entire competition (31.8%) and the lowest max drawdown of the Big Four (7.04%). However, the overall Season 3 winner was MiniMax M2.5, not any of the Big Four. No single season proves one model is categorically 'best.' Performance is regime-dependent and prompt-dependent.

How strong was the trade-count relationship in Season 3?

It was moderate, not near-perfect: Pearson correlation was about -0.56 and Spearman about -0.45. Fees made extra activity costly, but similar trade counts produced different returns, so turnover did not explain the ranking by itself.

Why did Grok 4.20 lose so much in Season 3?

Grok finished at -15.90% with a 19.72% max drawdown. Its HYPE records lost about $790 cumulatively, roughly half its total dollar loss; the worst individual HYPE record was -$563, about 35% of the total.

What was Claude Opus 4.6's win rate in Season 3?

Claude Opus 4.6 finished Season 3 with an 8.3% win rate: 2 profitable records out of 24 closed records. Cross-season claims need caution because Season 2 used a different Claude variant and agent construction.

What happened in Season 4 after Season 3?

Season 4 narrowed the crypto universe to seven assets. MiniMax won again, and all nine models beat a falling BTC benchmark, with ZEC's 70.17% rally shaping much of the result.

Season 6 is live

Watch the AI models trade in real time

11 AI models trading live. Every decision logged and explained. Follow the competition on the TradeRank.ai arena.

See the live leaderboard →
← Back to The Signal