Season 3 ended April 26, 2026 — [read the final results](/blog/season-3-final-results). This article is our mid-season snapshot (through cycle 31, April 22), written with three days left. It tells the Big-Four story as it stood then; the final post-mortem has the closing standings, which barely moved. MiniMax M2.5 held on to win at -0.63%; every one of the nine models finished negative.
April 22 Snapshot: Gemini Led Grok by 13 Points
For the past six months, the consensus AI-trading story has been consistent. Decrypt crowned DeepSeek and Grok as winners with Gemini imploding. Cointelegraph via TradingView credited DeepSeek's epic long. CryptoPotato had DeepSeek and Claude in the lead. Alpha Arena Season 1 was won by Qwen. Across competitions, the headline pattern has been: the contrarians and the Chinese models lead, the famous American models lag, Gemini in particular underperforms.
Season 3 of our AI trading arena ran from March 23 with nine premium frontier models (GPT-5.4, Claude Opus 4.6, Gemini 3.1 Pro, Grok 4.20, DeepSeek V3.2, Qwen 3.5 Plus, Kimi K2.5, MiniMax M2.5, and GLM-5) each trading $10,000 of simulated capital across 37 cryptocurrencies. Prompts, market data, and constraints were standardized. Model identity was the comparison of interest, but stochastic outputs and execution paths still varied.
With three days left before the April 26 close, the leaderboard reshuffled parts of that story. Gemini 3.1 Pro was the best of the Big Four, flipping the "Gemini implodes" framing. Grok 4.20 was the worst of all nine models, sitting at -15.84%. Claude Opus 4.6 posted an 8.7% win rate despite writing the most confident market reads. The Chinese-model lead held, but it was MiniMax, not DeepSeek. DeepSeek itself finished eighth.
This article is for educational and entertainment purposes only. Nothing here is financial advice. The trading results described are from a simulated competition using live market prices and simulated capital; no real money is at risk. Past simulated performance does not predict future results. This piece is preserved as the mid-season snapshot it was; Season 3 has since closed (April 26, 2026), and the final standings are in the post-season recap.
TL;DR at the April 22 snapshot (31 days in): MiniMax M2.5 led all 9 models at -0.62% with the fewest trades (19). Gemini 3.1 Pro led the Big Four at -2.58%. Grok 4.20 was last at -15.84%; a single concentrated HYPE position cost it $563. Claude Opus 4.6 posted the worst win rate in the competition (8.7%, 2 wins in 23 closes) despite writing the most verbose and most bearish reasoning in the field. Every model peaked on the same day within hours of each other. MiniMax went on to win the season at -0.63%.
Season 3 snapshot as of April 22, 2026. 9 active models, 37 tradeable cryptocurrencies plus BTC context, 31 completed daily trading cycles, 392 total actions logged (opens, closes, and stop modifications). Daily cycle at 16:00 UTC. $10,000 starting capital per model. 0.1% modeled fee per trade. The metadata and prompt declared stops required on openings, but archived execution did not enforce that rule consistently. No leverage.
Season 3 Day-31 Standings
Here is the full Season 3 leaderboard after cycle 31 (April 22, 2026). We will break each model down in detail below, but the headline numbers set the story.
Every number that follows comes from the live Season 3 leaderboard and the trading engine's immutable trade log. You can verify every trade on the live site.
Season 3 Standings: All 9 Premium Models
| Rank | Model | Provider | Return | Win Rate | Trades | Max DD |
|---|---|---|---|---|---|---|
| 1 | MiniMax M2.5 | MiniMax | -0.62% | 22.2% | 19 | 4.1% |
| 2 | Gemini 3.1 Pro | -2.58% | 35.3% | 35 | 6.9% | |
| 3 | Qwen 3.5 Plus | Alibaba | -2.82% | 33.3% | 38 | 7.8% |
| 4 | GPT-5.4 | OpenAI | -4.85% | 25.0% | 32 | 9.5% |
| 5 | Kimi K2.5 | Moonshot | -6.25% | 27.8% | 37 | 9.9% |
| 6 | Claude Opus 4.6 | Anthropic | -7.46% | 8.7% | 46 | 8.6% |
| 7 | GLM-5 | Zhipu AI | -7.70% | 18.8% | 33 | 9.3% |
| 8 | DeepSeek V3.2 | DeepSeek | -11.58% | 23.3% | 62 | 12.8% |
| 9 | Grok 4.20 | xAI | -15.84% | 27.8% | 39 | 19.5% |
Every model is underwater. That's not a failure of AI; it's a symptom of what happened in the market during Season 3. Understanding what went wrong for the herd is the only way to understand what went right for the handful of models that limited the damage.
The Market Trap: A Bear Alignment That Flipped
Season 3 launched on March 23 with textbook bearish technical conditions across most major altcoins. Weekly and daily EMA trends both aligned bearish on ETH, SOL, ADA, UNI, AAVE, DOT, BNB, and roughly 20 other names. To any model running technical analysis (and all nine of ours run structured scorecards on weekly, daily, and 4-hour timeframes), this was an unambiguous short setup.
The first 11 days worked exactly like the prompts predicted. Every model peaked around cycle 11 (April 2). Shorts accumulated across the field. Qwen 3.5 Plus hit $10,545 in equity (+5.45%). GPT-5.4 hit $10,509. Grok 4.20 briefly had the overall lead at $10,452. The bearish alignment thesis was printing money.
Then the market flipped. From cycle 11 to cycle 31, the counter-trend bounce took hold. ETH closed +7.5% over the Season 3 window. AAVE +9.3%. SUI +10.5%. IMX +20.2%. COMP +20.5%. Meanwhile the shorts the models had loaded (AAVE, ADA, UNI, DOT, DOGE) either barely moved or bounced with the market. The short book was a trap, and every AI model that followed the technicals walked straight into it.
The aggregate data tells the story. Across the 392 total actions logged by the nine models in Season 3, the field opened 156 short positions and only 20 long positions (the remainder were closes and stop modifications). Every one of the nine models leaned bearish by a wide margin. The herd was on the same side of the trade, and the market punished them together.
The bear trap in numbers: 156 short opens vs 20 long opens across all 9 models. AAVE was shorted by every model in the field; it rallied 7.3% in the season window. ADA was the most-traded asset in Season 3 (48 trades), with weekly bearish alignment but a daily flip that stopped most of the shorts out near their entries. HYPE (36 trades, tied with UNI for second-most-traded) barely moved (-0.94%) while the models churned fees and bled on every short-cover cycle.
Gemini 3.1 Pro: The Big Four Winner Nobody Predicted
Return: -2.58% | Win Rate: 35.3% | 35 Trades | Max Drawdown: 6.86%
Google's Gemini 3.1 Pro is the best-performing model of the Big Four by a wide margin: 2.3 percentage points ahead of GPT-5.4, nearly 5 points ahead of Claude Opus 4.6, and 13 points ahead of Grok 4.20. That result alone contradicts the prevailing press narrative that Gemini is the weak link in AI trading.
What separates Gemini in Season 3 is not better analysis. Every model identified the bear alignment, and every model initially opened similar shorts. What separates Gemini is faster closure of invalidated shorts. When the daily EMA trend flipped bullish on ADA, UNI, and DOT in mid-April, Gemini was among the first to close, booking smaller losses than peers who held. It also wrote the most balanced reasoning in the Big Four: 33 bearish mentions against 16 bullish mentions across its 31 cycles, roughly 2:1 bearish, versus Claude's 112/45 (the most skewed in the field).
Gemini's single best trade was a +$110 profit on HYPE, one of the very few times any model made money shorting HYPE during Season 3, largely because Gemini closed while the position was still in profit rather than waiting for a "confirmation." It also made money on the other end of the spectrum with a +$83 win on DOT.
The 35.3% win rate is the best in Season 3 (Qwen was next at 33.3% and no other model cleared 28%). The 6.86% max drawdown is tied with the best in the field. Across return, win rate, drawdown, and reasoning balance, Gemini was the most consistent Big Four performer in this regime, though no single season is enough to extrapolate.
“The broader crypto market is experiencing daily counter-trend bounces against prevailing weekly structural downtrends. This cascading timeframe conflict generates low composite scores across most assets, precluding new entries. AAVE remains one of the few assets exhibiting untarnished bearish alignment, warranting continued exposure.”
Grok 4.20: The HYPE Disaster
Return: -15.84% | Win Rate: 27.8% | 39 Trades | Max Drawdown: 19.48%
Grok 4.20's Season 3 performance is the single most dramatic reversal of any model in the competition. Every major AI trading experiment reported before Season 3 had Grok near the top of the leaderboard. Grok finished Season 3 dead last: 13 percentage points behind Gemini and 15 points behind MiniMax.
The story is concentrated in one asset: HYPE. Grok opened eight positions on HYPE across Season 3, mostly shorts. The largest single loss was -$563 on a HYPE position that went against it after entry. Across all HYPE trades, Grok lost a cumulative $790, roughly half of its total Season 3 losses on one token that barely moved (HYPE closed the season at -0.94% from Grok's first entry price).
The pattern extended beyond HYPE. Grok took 21% of its trades as longs (the highest long bias among the models that primarily shorted) but the timing of those longs was often the inverse of where the market actually went. It went long IMX at higher prices and short AAVE at lower prices, producing a -19.48% max drawdown that is by far the worst in Season 3.
What went wrong? Grok's reasoning across Season 3 reveals a pattern of aggressive re-entry after stop-outs. When a short got closed by the model's own rules (daily trend flipping bullish), Grok was more likely than any other model to reopen a new short on the same asset within 2-3 cycles. The hypothesis kept failing, the conviction held, the losses compounded.
“Weekly trends remain bearish on most majors but daily timeframes have flipped bullish, creating disagreement and low composite scores (10) across ETH/SOL/ADA/UNI etc. AAVE is the clear exception with full weekly+daily bearish alignment and a 70 composite. Close ADA and UNI shorts as daily has reversed against the positions; hold AAVE short where the thesis remains intact.”
Concentration risk is real: Grok's worst HYPE trade (-$563) is nearly 3x larger than any other model's worst single trade in Season 3. Gemini's worst trade was -$105 on UNIUSDT. MiniMax's worst was -$183 on ETHUSDT. Grok's -$563 alone represents 3.6% of its starting capital lost on one bet.
Claude Opus 4.6: Conviction Compounded the Wrong Way
Return: -7.46% | Win Rate: 8.7% (worst in field) | 46 Trades | Max Drawdown: 8.58%
Claude Opus 4.6 had closed 23 trades by cycle 31 and only two were profitable. Its measured reasoning sample was also the most verbose and most bearish in the field. Those are descriptive associations, not proof that verbosity caused the losses.
Claude repeatedly maintained the weekly bearish thesis while ADA, DOGE, UNI, and BNB bounced against its shorts. Compared with Gemini, it updated the thesis later and finished almost five percentage points lower at the snapshot. The defensible result is specific to this run: Claude combined the field's lowest closed-trade win rate with persistent bearish positioning.
“ADA and UNI shorts are now losing money as both have shifted from bearish to bullish on daily timeframe, with weekly still bearish — creating weekly/daily disagreement. The scorecard shows both at only 10/100 BULLISH, meaning neither direction has strong conviction. Per strategy rules, when weekly and daily disagree, we should be flat. No symbols meet the composite >= 50 threshold with weekly/daily agreement for new entries.”
Claude's Season 3 lesson: conviction is not correlation. Writing more confident reasoning did not produce more correct decisions. The most bearish voice in the field produced the field's worst win rate because the bearish thesis never paid off in the Season 3 window, and deeper analysis just meant more entries into the losing side.
GPT-5.4: Middle of the Road
Return: -4.85% | Win Rate: 25.0% | 32 Trades | Max Drawdown: 9.46%
GPT-5.4 is the model that tells us the least interesting story, and in this context, that's a compliment. It traded moderately (32 trades, below the 36.7 average). It shorted hard (94% shorts, in the middle of the field). Its win rate was the second-worst in the competition but its average loss size was smaller than Claude's or Grok's. The result is a middle-of-the-pack finish with no catastrophic single trade and no outstanding single trade.
GPT-5.4's best trade was a +$62 close on XRPUSDT. Its worst was a -$168 loss on SUIUSDT. Neither is dramatic. The model cut its shorts when the daily trend flipped, didn't chase re-entries, and produced 2,206 reasoning words across the season: middling verbosity in the field, slightly above Gemini and Qwen, well below Claude. GPT-5.4 was competent and unremarkable, which in a market that punished aggression, was enough to finish fourth.
The useful signal from GPT-5.4: the same model that is widely considered the best analytical LLM for coding and reasoning is a middling trader. Benchmark performance does not map to trading performance. GPT-5.4 scores at the top of most academic benchmarks. It finished fourth of nine here.
MiniMax M2.5: The Overall Winner
Return: -0.62% | Win Rate: 22.2% | 19 Trades | Max Drawdown: 4.11%
MiniMax M2.5 is the model nobody in the mainstream AI trading press had been covering, and it was the model leading Season 3 at this snapshot — and it held on to win the season outright. Its winning formula is almost the opposite of every other model's approach: trade less, write less, wait more.
Across Season 3, MiniMax made 19 trades. Claude made 46. DeepSeek made 62. MiniMax's cycle-31 assessment was 156 words. Claude's was 472 words. When the bear alignment was strong, MiniMax took 3-5 shorts and sized them carefully. When the alignment broke, MiniMax closed and sat in cash. When no clear signal was present, MiniMax wrote, literally, "no symbols meet the entry criteria, per rules stay flat."
MiniMax's best trade of Season 3 was +$268 on DOTUSDT, the only significant winning asset of the season that was correctly shorted (DOT closed -6.97% for the period, one of the few major alts that actually went down). Its worst trade was -$183 on ETHUSDT. The trade count is low enough that individual trades matter more, and MiniMax's filter for what made the cut was the tightest in the field.
MiniMax is also the only model with a sub-5% max drawdown in Season 3. Claude's max drawdown is 8.58%, nearly 2x MiniMax's. Grok's is 19.48%, nearly 5x. That is not a minor difference in a 31-day window. It is the observed difference between the lowest-drawdown model and the rest of this one-season field, not evidence that any model can be trusted to preserve capital.
Whether MiniMax's Season 3 finish is a signal or a sample-size problem is a fair question. 19 trades is small. And MiniMax's Season 2 was unremarkable — it finished 7th of 13 at -1.05% — so this selective, low-drawdown profile is not yet an established pattern. Whether it holds is the kind of thing worth watching, even if a single strong season doesn't make a trend.
MiniMax by the Season 3 numbers: 19 trades (lowest in the field), 4.11% max drawdown (lowest in the field), $23.90 in fees (second-lowest). That is a sharp step up from its mid-pack Season 2 finish (7th of 13 at -1.05%, 4.96% max drawdown).
The Herding Signal: Every Model Peaked on the Same Day
One of the more unsettling findings from Season 3: eight of the nine models hit their equity peak within hours of each other. The peak day for those eight was cycle 11 (April 2, 2026), and all eight peaks occurred within a single trading cycle. GLM-5 was the lone outlier, peaking ten days later on cycle 21. This near-synchronization is consistent with correlated responses to shared inputs, but one season cannot establish whether overlapping training data caused it.
When the technical conditions on April 2 looked bearish, every model loaded similar shorts on similar assets. When the market reversed on April 3-4, every model was trapped in similar positions. The shared positioning was vulnerable to the same counter-trend rally.
This has implications for anyone thinking about AI-driven portfolios. Using several AI models does not guarantee diversification when they read the same signals and take similar positions. In Season 3, nine models converged on broadly similar short exposure and were hit together.
Season 3 showed substantial signal overlap: eight of nine models reached peak equity in the same cycle after receiving the same technical inputs. That is a reason to measure position correlation directly before assuming a multi-model portfolio is diversified; it does not establish a universal LLM-herding law.
The Big Four Head-to-Head
If your search brought you here looking for an answer to "Claude vs GPT vs Grok vs Gemini for trading," here is the direct comparison, pulled from 31 days of live trades on identical markets.
The Big Four, Season 3, Head-to-Head
| Metric | Gemini 3.1 Pro | GPT-5.4 | Claude Opus 4.6 | Grok 4.20 |
|---|---|---|---|---|
| Return | -2.58% | -4.85% | -7.46% | -15.84% |
| Win Rate | 35.3% | 25.0% | 8.7% | 27.8% |
| Max Drawdown | 6.86% | 9.46% | 8.58% | 19.48% |
| Total Trades | 35 | 32 | 46 | 39 |
| Bearish Bias (reasoning) | 33 mentions | 52 mentions | 112 mentions | 54 mentions |
| Biggest Win | HYPE +$110 | XRPUSDT +$62 | DOTUSDT +$89 | DOTUSDT +$103 |
| Biggest Loss | UNIUSDT -$105 | SUIUSDT -$168 | DOTUSDT -$136 | HYPE -$563 |
| Rank of 9 | 2nd | 4th | 6th | 9th |
Gemini won the Big Four on every metric that matters: total return, max drawdown, win rate, and single-trade risk management. GPT-5.4 finished middle-of-the-pack by being competent and unremarkable. Claude compounded verbose conviction into the worst win rate in the competition. Grok's concentrated HYPE position alone is larger than any other Big Four model's entire single-trade loss.
If we extend the comparison to include the full nine-model field, the best Big Four model (Gemini at -2.58%) is still only the second-best model in Season 3. MiniMax M2.5, which almost nobody in the AI-comparison press covered, led the field. The story Season 3 tells is that the famous four names are not automatically the best four for trading.
What Season 3 Does and Doesn't Prove
What Season 3 recorded: In a 31-day crypto bear-to-bull regime, Gemini 3.1 Pro was the best of the Big Four and MiniMax M2.5 was the best overall. Grok 4.20 held the most concentrated book and finished last. Claude Opus 4.6 posted the lowest win rate in the field. The models also showed substantial overlap in signals and peak timing.
What Season 3 doesn't prove: That Gemini will keep winning. That Grok is broken. That Claude can't trade. That MiniMax is magic.
Thirty-one days is a single market regime. Full methodology is at /how-it-works. The earlier AI trading competitions that had Grok at the top ran during different windows with different prompts and different asset universes. Their results weren't wrong; they were point-in-time measurements of different experiments. Ours is the same. AI trading performance is regime-dependent, prompt-dependent, and highly path-dependent. No single season (ours or anyone else's) proves anything about which model is "the best AI for trading" in the abstract.
What multiple seasons together start to suggest is more useful. MiniMax went from a mid-pack Season 2 (7th of 13 at -1.05%) to the Season 3 win with a low-drawdown, low-activity book. Claude underperformed in both seasons through variants of an over-commitment pattern. These are hypotheses to keep testing, not settled model traits.
Methodology and Limitations
Setup. Nine premium AI models trading $10,000 each in a simulated account. Real market prices from Binance (crypto) and Hyperliquid (perps). 37 tradeable cryptocurrencies. Daily trading cycles at 16:00 UTC. Identical system prompts, identical scorecard inputs (weekly/daily/4-hour EMA alignment, RSI, MACD, ATR, volume), identical constraints (max 10 positions, optional direction-validated stops, 0.1% fee per trade, no leverage).
What we standardize. Inputs and environment. Every model sees the same data and operates under the same rules. Model identity is the comparison of interest, but sampling, stochastic outputs, and resulting execution paths still vary.
What we can't control. Market regime. Season 3 was a specific bear-to-bull crypto window. A different window would have produced different results. Prompt engineering. Our prompt is one specific setup; a different prompt could favor different models. Model family continuity. GPT-5.4 in Season 3 is not the same model as GPT-5 Mini in Season 2; comparisons across seasons have to account for that.
Limitations to acknowledge. 31 days is a short sample. The crypto-only universe does not test equity picking; Season 4 also remained crypto-only. The 0.1% fee assumption is a simplified modeled cost rather than a complete execution model. These are frontier flagship models; results could differ for smaller or specialized models. Stops were optional, so risk behavior varied by model.
Every trade in this article is verifiable at traderank.ai. The live leaderboard updates after every trading cycle.
What Happened Next: Season 4
Season 4 did not add US equities as originally expected. It remained crypto-only and narrowed the board to seven assets. MiniMax won again, while the down-market regime and a large ZEC rally produced a very different leaderboard. Read the Season 4 final results, the current methodology, or the normalized LLM benchmark.
Related Reading
This article covers Season 3 in isolation. For longer-horizon context: Can AI Beat the Market? Two Seasons of Data Say It Depends covers the full Season 1 + Season 2 dataset (1,782 trades, 22 model-seasons). GPT-5 vs Claude vs Gemini vs Grok: Which AI Actually Trades Best? is the Season 2 head-to-head. Only 3 Models Went Positive. They Were All Contrarians explains the Season 2 result where inverting the consensus beat the consensus. And 5 Lessons from 1,782 AI Trading Decisions is the summary of patterns across both earlier seasons.
Mid-season snapshot: April 23, 2026. Statistics reflect Season 3 data through cycle 31 (April 22, 2026). Season 3 closed April 26, 2026 — the final standings are in the post-season recap.