Season 4 Final: All 9 Premium AI Models Beat BTC

BTC fell 3.33% over the 27-day window. Nine frontier models combined for 32.96 percentage points of return. The contrast with Season 3, where every model lost money in a +10% BTC tape, could not be sharper. Here is what changed.

+6.94%Sharpe 20.32125 trades

Nine Frontier Models, a Down Market, and a Clean Sweep

Season 4 ended on May 23, 2026. Nine premium models traded $10,000 each across seven Binance USDT pairs for 30 daily cycles. The archive contains 125 closed trade records and $342,659 in notional volume.

Eight of nine finished positive and all nine beat BTC's -3.33% benchmark. Five tradeable assets fell or stayed within roughly 3% of flat, while TAO gained 10.13% and ZEC gained 70.17%. ZEC exposure therefore explains much of the cross-model spread.

Warning

This article is for educational and entertainment purposes only. Nothing here is financial advice. Trades described are from a simulated competition using live market prices and simulated capital; no real money is at risk. Past simulated performance does not predict future results. The numbers below are the official end-of-Season-4 standings as published in the Season 4 report.

Key Insight

TL;DR: MiniMax M2.7 won Season 4 at +6.94%, defending the crown it claimed in Season 3 (then at -0.63%, in a +10% BTC tape). DeepSeek V4 Pro took silver. Grok 4.20 MA took bronze on the back of one $1,000+ ZEC trade. Eight of nine premium models finished positive, and every premium model beat BTC. The deeper pattern: this was a season where being short the right majors and long the right outlier paid for everything. Six of seven tradeable assets were down or flat. The models that sized into shorts and held ZEC made the season; the model that didn't (GLM-5.1) finished red.

Final Standings

Here are the official Season 4 final standings as of May 23, 2026, taken from the end-of-season equity snapshot. Every number is sourced from the Season 4 competition report and verifiable against the immutable trade log. As in Season 5, several models still held open positions at the close, so part of each return is a mark-to-market gain rather than booked profit — see the realized/unrealized note below the table.

Season 4 Final Standings

RankModelProviderReturnWin RateTradesSharpe
1MiniMax M2.7MiniMax+6.94%55.6%920.32
2DeepSeek V4 ProDeepSeek+5.40%30.8%1311.23
3Grok 4.20 MAxAI+5.34%16.7%189.57
4Gemini 3.1 ProGoogle+4.43%35.3%1710.76
5Kimi K2.6Moonshot+4.13%27.8%188.23
6GPT-5.5OpenAI+3.69%31.3%1610.97
7Qwen 3.6 PlusAlibaba+2.72%26.7%157.86
8Claude Opus 4.7Anthropic+0.88%25.0%123.06
9GLM-5.1Zhipu AI-0.57%28.6%7-18.32
Data Point

Aggregate Season 4 figures. 9 models, 125 closed trade records, 30 daily cycles, $342,659 in notional volume across 7 tradeable cryptocurrencies. Starting capital: $10,000 per model. 8 of 9 models finished positive; all 9 beat BTC at -3.33%.

Data Point

Realized vs. unrealized (the rigor note). As in Season 5, Season 4's headline returns are end-of-season equity marks, and for several models a large share was still unrealized on open positions — the ZEC longs especially. Per the Season 4 report: Grok's +5.34% sat on +$1,150 unrealized against −$616 realized, and Claude's +0.88% was +$458 unrealized against −$370 realized. The winner banked more of it: MiniMax closed +$447 realized on top of +$247 unrealized, and GPT-5.5 booked +$331 realized against +$38 unrealized. The field was positioned correctly, but not all of the green was money in the bank.

Benchmarks: A Down Market with One Outlier

The market context for Season 4 was unusual. Six of the seven tradeable assets were down or roughly flat — but one (ZEC) ran up 70 percent. The aggregate effect was a slightly negative crypto-majors tape with a single asymmetric long opportunity sitting in the middle of it.

Season 4 Asset Returns (Apr 26 → May 23, 2026)

SymbolStartEndReturn
ZEC$355.06$604.22+70.17%
TAO$246.70$271.70+10.13%
BNB$631.27$646.42+2.40%
DOGE$0.1019$0.1043+2.37%
SOL$86.51$83.97-2.94%
BTC (benchmark)$78,012$75,415-3.33%
XRP$1.43$1.34-6.46%
ETH$2,330$2,058-11.69%

The ZEC Trade That Defined the Season

Every model held ZEC long at some point. Grok entered near the cycle low at $359.97; its 4.04-unit open long was marked around $609.01 for roughly $1,006 unrealized gain at the close. That position lifted Grok to third despite a 16.7% closed-record win rate.

Other models entered later or sized smaller. The entry-price spread was nearly $300, making ZEC timing and exposure a major source of variance. The $1,006 position was the season's largest individual mark, but it was not larger than the bottom five models' combined P&L.

Data Point

The ZEC sizing spread. Grok bought 4.04 units (entry $359.97). DeepSeek bought 2.49 units. Claude bought 2.24 units. Gemini, MiniMax, and Kimi each bought ~2 units at successively worse prices. GLM-5.1, which finished last, bought only 1.57 units and locked in a small loss. The lower you went in the standings, the later and smaller your ZEC long.

MiniMax M2.7: The Quiet Defense of the Crown

MiniMax M2.5 won Season 3 at -0.63% — the smallest loss in a field where every model lost. The story then was discipline: 10 trades, the lightest hand on the field. Season 4's upgraded M2.7 version produced a different result.

MiniMax M2.7 made 9 trades — still light, but this time the model won 5 of them outright (55.6% win rate, the only model above 50% in the field). It carried just three open positions into the close: a short ETH, a short XRP, and a long ZEC. All three were correct. The ETH short returned 5.4%, the XRP short 5.5%, and the ZEC long 14.1%. MiniMax did not catch ZEC at the low. It did not size enormously into any one asset. It simply got every directional call right.

The Sharpe ratio of 20.32 is striking on the surface — but it is partly an artefact of low position count smoothing equity curve volatility. Back-to-back wins by the MiniMax model family warrant further testing, while the version change prevents a same-model causal claim.

DeepSeek and Grok: Two Different Roads to the Podium

DeepSeek and Grok finished six basis points apart at +5.40% and +5.34%, but their books differed. DeepSeek's final open book contained positive ETH and XRP shorts plus a losing ZEC long; the portfolio was positive overall, not three closed winners.

Grok's podium return depended on an open ZEC long worth roughly +$1,006 unrealized against about -$616 realized elsewhere. Its low win rate and concentrated mark show why total equity, realized P&L, and open exposure must be read together.

GLM-5.1: The Bottom of the Field for a Different Reason

GLM finished last at -0.57%. Its seven closed records produced -$65.46 realized P&L, the final open book added +$8.54 unrealized, and fees were $10.99. It entered ZEC late and small, so it did not capture the outlier that supported other books.

The result does not show that low activity caused the loss. It shows that GLM's limited activity did not include enough profitable exposure to offset realized losses and fees.

Key Insight

Season 3 lesson: trade less, lose less. Season 4 lesson: when the one true trade arrives, *size into it*. Both seasons reward discipline — but discipline cuts in different directions depending on whether the market is offering setups or sandbagging the field.

Season 3 vs. Season 4: The Comparison That Matters

The provider slots, starting capital, daily cadence, and fee structure were comparable, but this was not a one-variable experiment. The tradeable universe shrank from 37 assets to seven and displayed model versions changed between seasons. Market regime was a major difference, not the only one.

Season 3 → Season 4 Movement

ModelS3 ReturnS4 ReturnΔ
MiniMax-0.63%+6.94%+7.57pp
DeepSeek-11.66%+5.40%+17.06pp
Grok-15.90%+5.34%+21.24pp
Gemini-2.64%+4.43%+7.07pp
Kimi-6.35%+4.13%+10.48pp
GPT-5.00%+3.69%+8.69pp
Qwen-2.73%+2.72%+5.45pp
Claude Opus-7.61%+0.88%+8.49pp
GLM-7.67%-0.57%+7.10pp

Every model improved. The biggest movers were the bottom-half models from Season 3 — Grok went from -15.90% to +5.34% (a 21-point swing) and DeepSeek went from -11.66% to +5.40% (17 points). The top-half models all improved by 5-10 points. Nobody moved meaningfully against the trend.

The surface read is that AI trading models got better. The deeper read is that Season 3 was a tape where 'patient longs in a rising market' would have beaten every model in the field, and the models persisted in over-trading and shorting into strength. Season 4 was the opposite: a tape where 'short the obvious downside, ride the one moonshot' was the playbook, and the models — finally — found it.

What Season 4 Tells Us About AI Trading as a Category

Three conclusions fit the evidence. First, all nine models beat a falling BTC benchmark in this window. Second, ending equity included open marks, especially ZEC, so positive leaderboard return was not always banked profit. Third, neither trade count nor win rate explained the standings consistently; asset selection, entry timing, sizing, and unrealized exposure mattered.

A 27-day sample with one 70% outlier cannot rule out chance, prompt sensitivity, or regime fit. Season 4 is evidence that this setup can outperform its benchmark in one down market, not proof of durable alpha.

What Happened Next: Season 5

Season 5 launched on May 23, 2026 with two structural changes:

1. The roster expanded to 10 models. Mistral Medium 3.5 joined, while Gemini moved to 3.5 Flash and Grok to 4.3.

2. The asset list expanded to 10. TON, SUI, TRX, and PEPE joined; TAO left the board.

The $10,000 simulated starting capital, daily 16:00 UTC cycle, and cascading-timeframe prompt carried over. Season 5 is now complete; read the final report rather than treating this launch note as live status.

The Closing Statement

Season 3 was a confidence-shaking month for AI trading. Every model lost in a rising market, and the case for premium frontier models as systematic traders looked weaker for it.

Season 4 is the rebuttal. The same nine model families returned under the same prompt and rules, but eight of nine seats used newer model versions. The market was sloppier, asymmetric, and dominated by one outlier; the field found that asymmetry, and every model beat the benchmark.

One data point is not a thesis. Back-to-back wins from the MiniMax family, M2.5 then M2.7, in opposite regimes warrant attention without proving a stable same-model edge. Season 5 provides the next completed test.

Frequently Asked Questions

Who won Season 4 of TradeRank's AI trading competition?

MiniMax M2.7 won Season 4 with a +6.94% return, the best of nine premium AI models. It traded just 9 times, posted the field's only win rate above 50% (55.6%), and carried three correct open positions into the close (short ETH, short XRP, long ZEC). It was the second season in a row MiniMax finished first, after winning Season 3 at -0.63%. Eight of the nine models finished positive; only GLM-5.1 finished red, at -0.57%.

Did the AI models beat Bitcoin in Season 4?

Yes — all nine models beat BTC, which fell 3.33% over the 27-day window. Eight of the nine also finished positive against the dollar, and the field combined for +32.96 percentage points of return. The contrast with Season 3 is stark: the same nine model families all lost money while BTC rose 10%, though eight seats changed model version before Season 4. In Season 4 the field largely shorted weak majors and held the one asset that rallied.

What was the ZEC trade that defined Season 4?

ZEC rallied 70.17% during Season 4 (from $355.06 to $604.22) and it was the only asset on the seven-coin board that ran hard while the rest fell or stalled. Every model held ZEC long at some point, but size and entry timing separated the field. Grok 4.20 MA bought 4.04 units near the low at $359.97; that position was marked up roughly $1,006 by the close, just below the bottom five models' combined P&L of about $1,085, and lifted Grok to bronze despite a 16.7% win rate. Most of the mark was unrealized at the final snapshot.

Why did GLM-5.1 finish last in Season 4?

GLM-5.1 finished ninth at -0.57%, the only model in the red — but not because it read the market wrong. It barely traded: 7 trades over 30 cycles, and its ZEC long was only 1.57 units entered at $629.30, well above where the rally bottomed. Its directional calls were mostly defensible, but it never built a position big enough to matter on ZEC, the one asymmetric trade that made the season for everyone else. In a market that offered a single big setup, under-sizing it was the mistake.

Which of Claude, GPT-5, Gemini, and Grok traded best in Season 4?

Grok 4.20 MA was the best of the Big Four in Season 4 at +5.34% (third overall), followed by Gemini 3.1 Pro at +4.43%, GPT-5.5 at +3.69%, and Claude Opus 4.7 at +0.88%. Grok got there almost entirely on one deeply-sized ZEC long rather than consistent accuracy — its win rate was just 16.7%. The overall winner was MiniMax M2.7 (+6.94%), not any of the Big Four. No single season proves one model is categorically best; performance is regime-dependent.

How was Season 4 different from Season 3?

The same nine model families used the same prompt and rules, but eight seats had version upgrades. In Season 3, BTC rose 10.1% and all nine lost money after a bearish setup reversed. In Season 4, BTC fell 3.33%; eight of nine finished positive and all nine beat BTC, helped by shorts in weak majors and the ZEC rally. Grok improved by 21.24 points and DeepSeek by 17.06, but version and regime changed together, so the experiment cannot attribute the swing to either one alone.

Season 6 is live

Watch the AI models trade in real time

11 AI models trading live. Every decision logged and explained. Follow the competition on the TradeRank.ai arena.

See the live leaderboard →
← Back to The Signal