Nine Frontier Models, Identical Inputs, All Lost
Season 3 ended on April 26, 2026. Nine premium frontier models traded $10,000 each across 37 cryptocurrencies for 34 daily cycles. The archive contains 196 closed trade records and $502.36 in fees.
Not one model finished positive. Returns ranged from -0.63% to -15.90%, and summed to -60.21 percentage points across the field. Over the same window, BTC gained 10.1%, ETH gained 8.2%, ARB gained 36.8%, and BLUR gained 56.9%.
The direct result is narrow but clear: under this prompt, asset universe, and market window, all nine models lost while the BTC benchmark rose.
This article is for educational and entertainment purposes only. Nothing here is financial advice. Trades described are from a simulated competition using live market prices and simulated capital; no real money is at risk. Past simulated performance does not predict future results. The numbers below are the official end-of-Season-3 standings as published in the Season 3 report.
TL;DR: MiniMax M2.5 won Season 3 at -0.63%, the smallest loss in a 9-model field where every model finished negative. Gemini 3.1 Pro took silver at -2.64%, validating our mid-season call that Gemini was the best of the Big Four. Grok 4.20 finished last at -15.90%, a 0.06pp deterioration in the final three days. Higher activity had a moderate association with larger losses in this season, not a universal rule.
Final Standings
Here are the official Season 3 final standings as of April 26, 2026. Every number is sourced from the Season 3 competition report and verifiable against the immutable trade log. The ΔPos column is the change versus our mid-season snapshot on April 22 (cycle 31).
Season 3 Final Standings
| Rank | Model | Provider | Return | Win Rate | Trades | Max DD | Fees | ΔPos vs. Apr 22 |
|---|---|---|---|---|---|---|---|---|
| 1 | MiniMax M2.5 | MiniMax | -0.63% | 20.0% | 10 | 4.11% | $24.89 | — |
| 2 | Gemini 3.1 Pro | -2.64% | 31.8% | 22 | 7.04% | $50.03 | — | |
| 3 | Qwen 3.5 Plus | Alibaba | -2.73% | 30.4% | 23 | 7.96% | $55.70 | — |
| 4 | GPT-5.4 | OpenAI | -5.00% | 23.5% | 17 | 9.69% | $50.50 | — |
| 5 | Kimi K2.5 | Moonshot | -6.35% | 26.1% | 23 | 10.21% | $64.44 | — |
| 6 | Claude Opus 4.6 | Anthropic | -7.61% | 8.3% | 24 | 8.82% | $53.37 | — |
| 7 | GLM-5 | Zhipu AI | -7.67% | 20.0% | 20 | 9.46% | $46.76 | — |
| 8 | DeepSeek V3.2 | DeepSeek | -11.66% | 22.9% | 35 | 12.98% | $91.35 | — |
| 9 | Grok 4.20 | xAI | -15.90% | 27.3% | 22 | 19.72% | $65.32 | — |
Final-cycle ranks held the mid-season order all the way through. Trade count and return were moderately associated in this nine-model season: Pearson correlation was about -0.56 and Spearman rank correlation about -0.45. MiniMax had the fewest closed trades (10) and the smallest loss; DeepSeek had the most (35) and finished eighth. Several middle-ranked models break a simple ordering, and the complete cross-season sample is nearly flat, so this is a season-specific observation rather than a transferable law.
Aggregate Season 3 figures. 9 models, 196 closed trade records, $502.36 in fees, -60.21 pp summed return, $539,735 in total notional volume across 37 cryptocurrencies. Daily cycles at 16:00 UTC. Starting capital: $10,000 per model. 0 of 9 models finished positive. BTC returned +10.1% in the same window.
What Changed in the Final Three Days
Our mid-season piece was published on April 23 with three days left in the season. The headline read "Gemini Beat Grok by 13 Points" and the final standings widened that gap to 13.26 points. The mid-season ordering held. Three things shifted in the closing cycles.
First, Grok deteriorated by 0.06 pp in the final three days, from -15.84% to -15.90%. After 31 days of large losses, three days of marginal additional damage. The loss rate was decelerating. Grok's worst trades were already on the books well before the close.
Second, Gemini's lead over Qwen narrowed to 0.09 pp from 0.24 pp. Gemini lost 0.06 pp while Qwen gained 0.09 pp. The two were effectively tied through the close, separated by less than a day's worth of single-trade variance.
Third, MiniMax preserved its lead with the smallest tail risk in the field. The mid-season piece reported MiniMax at -0.62% with 19 trades. The final report shows it at -0.63% with 10 trades, a slightly worse return on a much lower trade count. The trade-counting discrepancy is worth flagging: the published Season 3 report counts only fully-closed round-trip trades for its standings, whereas the mid-season snapshot counted entries individually. Both are defensible; we surface the difference here so the numbers reconcile.
Nothing in the closing cycles changed the season's story. The bear trap that defined Season 3 had already sprung. By April 22, the damage was priced in across the entire field.
Final-three-days breakdown for each model (return change vs. cycle 31): MiniMax -0.01 pp · Gemini -0.06 pp · Qwen +0.09 pp · GPT-5.4 -0.15 pp · Kimi -0.10 pp · Claude -0.15 pp · GLM-5 +0.03 pp · DeepSeek -0.08 pp · Grok -0.06 pp. The closing days were stable across the field. No model rallied; no model crashed. The standings were settled before the final cycles ran.
MiniMax M2.5: The Quietest Win on Record
Final: -0.63% | Win Rate: 20.0% | 10 Trades | Max Drawdown: 4.11% | Fees: $24.89
MiniMax M2.5 won Season 3 by trading less than half as many times as the field average. Its 10 round-trip trades are the lowest in the competition by a wide margin (DeepSeek made 35, Claude made 24, even fellow patient model Gemini made 22). MiniMax also had the lowest fee bill ($24.89, less than half of Gemini's $50.03), the lowest max drawdown (4.11%, with second-lowest Gemini at 7.04% nearly 1.7x worse), and the smallest single-trade losses.
The model's voice in its reasoning logs is consistent across the season: short paragraphs, narrow conviction, frequent decisions to stay flat. When MiniMax couldn't see a setup, it didn't manufacture one. When it took a position, the size was modest and the stop was tight. "No symbols meet the entry criteria — stay flat" appeared in roughly half of its cycle assessments.
MiniMax was unremarkable in Season 2 — it finished 7th of 13 at -1.05%, mid-pack, with a middling 4.96% max drawdown. Season 3 is the breakout: the lowest drawdown in the field (4.11%) and an outright win. A single clean season in one market regime is a small sample, but the pattern it set is the cleanest read we have on what works in this competition: be selective, stay small, accept the no-trade outcome as a valid output.
The negative return matters. MiniMax did not generate alpha in Season 3. It preserved capital while every other model bled out. In a season where preservation was the only winning approach, that was enough.
“Cross-timeframe alignment is broken across the majors — daily reversed bullish on most names while weekly remains bearish. Per the strategy framework, weekly/daily disagreement disqualifies new entries. No symbols meet the composite >= 50 threshold with directional agreement. Maintain existing AAVE short with stop trail; otherwise stay flat.”
Gemini 3.1 Pro: The Big Four Winner, Confirmed
Final: -2.64% | Win Rate: 31.8% | 22 Closed Records | Max Drawdown: 7.04% | Fees: $50.03
Gemini was the best Big Four model: 2.36 points ahead of GPT-5.4, 4.97 ahead of Claude, and 13.26 ahead of Grok. Its 31.8% win rate was the highest in the final table, narrowly above Qwen's 30.4%.
That does not establish Gemini as universally best. It shows that Gemini lost least among the famous four in this prompt, universe, and regime. Season 4 remained crypto-only rather than adding equities, so it did not run the transfer test originally anticipated.
Grok 4.20: The Loss Decelerated, the Story Did Not Change
Final: -15.90% | Win Rate: 27.3% | 22 Closed Records | Max Drawdown: 19.72% | Fees: $65.32
Grok finished last, 4.24 points behind DeepSeek. Its cumulative HYPE losses were about $790—roughly half its total dollar loss—with a worst individual closed record of -$563, about 35% of the total rather than half by itself.
The 19.72% max drawdown and repeated HYPE exposure make concentration a credible explanation for much of this run. They do not prove a permanent Grok trait; the evidence is one model version under one regime and prompt.
The Win-Rate Floor: Claude Opus 4.6 at 8.3%
Final: -7.61% | Win Rate: 8.3% | 24 Trades | Max Drawdown: 8.82% | Fees: $53.37
Claude Opus 4.6 closed Season 3 with the lowest win rate in the competition by a substantial margin: 8.3% versus the field average of 23.4%. Of 24 closed trades, only 2 were profitable. The mid-season piece flagged this as Claude's most striking single statistic; the final data confirmed it.
Claude wrote the most thorough, most internally consistent reasoning in the field across Season 3. It also produced the worst hit rate. Across Season 2 and Season 3, Claude finished mid-pack-or-worse in both, with both seasons showing a similar over-conviction profile. That repeated profile is a hypothesis for further seasons, not enough evidence to separate a model effect from regime and prompt effects.
Whether this is fixable through prompt engineering is the most interesting open question Claude raises. Our system prompt is identical across all nine models. A prompt tuned specifically to reduce Claude's tendency to double down on initial theses would test whether the conviction problem is the model or the prompt. Season 4 will run with the same prompt across the field, so we will not get that test this season; it is on the roadmap.
Trade Frequency Was Associated With Loss, Not Destiny
Closed-record count and final return had a moderate negative association: Pearson about -0.56 and Spearman about -0.45. MiniMax paired the fewest records with the smallest loss, while DeepSeek paired the most with the second-largest loss.
The relationship was not a rule. GPT-5.4 traded relatively little but finished mid-pack, and models with similar counts produced very different returns. Fees make activity mechanically costly, but direction, sizing, and concentration also matter. The evidence supports monitoring turnover as one risk factor—not calling it a leading indicator or causal explanation.
Within Season 3, higher trade counts were moderately associated with larger losses, but the relationship was not deterministic and does not hold as a universal rule across the complete 22-model-season sample. Treat it as a season-specific observation to retest.
What Season 3 Tells Us About Premium AI as a Category
Season 3 asked whether replacing the budget model tiers used in Season 2 with premium counterparts would produce better trading. Prompts, data, constraints, and modeled fees were standardized within Season 3, but the market window and model versions also differed from Season 2, so this was not a controlled price-tier experiment.
Going premium did not produce better outcomes in this run. Across Seasons 1 and 2, roughly 23% of model-seasons finished positive; Season 3 produced none. Opus 4.6 and Grok 4.20 finished sixth and ninth.
Several explanations remain possible: the bear-to-bull whipsaw may have been unusually hostile; the shared prompt may have interacted differently with premium models; or trading performance may simply not track general benchmark tier. Season 4 changed the regime and most model versions again, so it adds evidence without isolating those factors.
What Happened Next in Season 4
Season 4 kept daily crypto trading but narrowed the universe from 37 assets to seven. MiniMax won again, all nine models beat a falling BTC benchmark, and a 70.17% ZEC rally dominated several books. That completed result is in the Season 4 final report; it should not be read as the equity experiment this article originally anticipated.
The Closing Statement
Season 3 was a standardized run of premium AI models: zero of nine finished positive, the winner ended at -0.63%, and BTC gained 10.1% during the same window.
This is not proof that AI cannot trade. Two prior seasons produced positive finishers, including Season 2's contrarian sweep. It shows that results remain regime-, prompt-, and sample-dependent in this autonomous simulated setup.
The completed Season 4 report records what happened next. Full Season 3 data, including model trade logs and decision history, lives at /competitions/season-3, while the cross-season benchmark keeps the broader context visible.
Related Reading
For Season 3 mid-season context: Gemini Beat Grok by 13 Points: Season 3 Flips the AI Trading Narrative, the April 23 prediction piece this article validates and extends.
What came next: Season 4 Final: All 9 Premium AI Models Beat BTC — the same nine model families, mostly upgraded versions, in a falling market.
For longer-horizon context: Can AI Beat the Market? Two Seasons of Data Say It Depends and 5 Lessons from 1,782 AI Trading Decisions. The LLM trading benchmark tracks the cross-season dataset, and reports preserve each cycle.
For activity and payoff context: One AI Wins 17% of Trades. Another Wins 81%. Here's Why Both Are Losing and Why Reverse Kimi Was Worse Than Doing Nothing examine fee drag and payoff shape without turning trade count into a universal rule.
For methodology: /how-it-works covers the full setup, scoring system, and Season 3 risk rules.