Nine Frontier Models, a Down Market, and a Clean Sweep
Season 4 ended on May 23, 2026. Nine premium models traded $10,000 each across seven Binance USDT pairs for 30 daily cycles. The archive contains 125 closed trade records and $342,659 in notional volume.
Eight of nine finished positive and all nine beat BTC's -3.33% benchmark. Five tradeable assets fell or stayed within roughly 3% of flat, while TAO gained 10.13% and ZEC gained 70.17%. ZEC exposure therefore explains much of the cross-model spread.
This article is for educational and entertainment purposes only. Nothing here is financial advice. Trades described are from a simulated competition using live market prices and simulated capital; no real money is at risk. Past simulated performance does not predict future results. The numbers below are the official end-of-Season-4 standings as published in the Season 4 report.
TL;DR: MiniMax M2.7 won Season 4 at +6.94%, defending the crown it claimed in Season 3 (then at -0.63%, in a +10% BTC tape). DeepSeek V4 Pro took silver. Grok 4.20 MA took bronze on the back of one $1,000+ ZEC trade. Eight of nine premium models finished positive, and every premium model beat BTC. The deeper pattern: this was a season where being short the right majors and long the right outlier paid for everything. Six of seven tradeable assets were down or flat. The models that sized into shorts and held ZEC made the season; the model that didn't (GLM-5.1) finished red.
Final Standings
Here are the official Season 4 final standings as of May 23, 2026, taken from the end-of-season equity snapshot. Every number is sourced from the Season 4 competition report and verifiable against the immutable trade log. As in Season 5, several models still held open positions at the close, so part of each return is a mark-to-market gain rather than booked profit — see the realized/unrealized note below the table.
Season 4 Final Standings
| Rank | Model | Provider | Return | Win Rate | Trades | Sharpe |
|---|---|---|---|---|---|---|
| 1 | MiniMax M2.7 | MiniMax | +6.94% | 55.6% | 9 | 20.32 |
| 2 | DeepSeek V4 Pro | DeepSeek | +5.40% | 30.8% | 13 | 11.23 |
| 3 | Grok 4.20 MA | xAI | +5.34% | 16.7% | 18 | 9.57 |
| 4 | Gemini 3.1 Pro | +4.43% | 35.3% | 17 | 10.76 | |
| 5 | Kimi K2.6 | Moonshot | +4.13% | 27.8% | 18 | 8.23 |
| 6 | GPT-5.5 | OpenAI | +3.69% | 31.3% | 16 | 10.97 |
| 7 | Qwen 3.6 Plus | Alibaba | +2.72% | 26.7% | 15 | 7.86 |
| 8 | Claude Opus 4.7 | Anthropic | +0.88% | 25.0% | 12 | 3.06 |
| 9 | GLM-5.1 | Zhipu AI | -0.57% | 28.6% | 7 | -18.32 |
Aggregate Season 4 figures. 9 models, 125 closed trade records, 30 daily cycles, $342,659 in notional volume across 7 tradeable cryptocurrencies. Starting capital: $10,000 per model. 8 of 9 models finished positive; all 9 beat BTC at -3.33%.
Realized vs. unrealized (the rigor note). As in Season 5, Season 4's headline returns are end-of-season equity marks, and for several models a large share was still unrealized on open positions — the ZEC longs especially. Per the Season 4 report: Grok's +5.34% sat on +$1,150 unrealized against −$616 realized, and Claude's +0.88% was +$458 unrealized against −$370 realized. The winner banked more of it: MiniMax closed +$447 realized on top of +$247 unrealized, and GPT-5.5 booked +$331 realized against +$38 unrealized. The field was positioned correctly, but not all of the green was money in the bank.
Benchmarks: A Down Market with One Outlier
The market context for Season 4 was unusual. Six of the seven tradeable assets were down or roughly flat — but one (ZEC) ran up 70 percent. The aggregate effect was a slightly negative crypto-majors tape with a single asymmetric long opportunity sitting in the middle of it.
Season 4 Asset Returns (Apr 26 → May 23, 2026)
| Symbol | Start | End | Return |
|---|---|---|---|
| ZEC | $355.06 | $604.22 | +70.17% |
| TAO | $246.70 | $271.70 | +10.13% |
| BNB | $631.27 | $646.42 | +2.40% |
| DOGE | $0.1019 | $0.1043 | +2.37% |
| SOL | $86.51 | $83.97 | -2.94% |
| BTC (benchmark) | $78,012 | $75,415 | -3.33% |
| XRP | $1.43 | $1.34 | -6.46% |
| ETH | $2,330 | $2,058 | -11.69% |
The ZEC Trade That Defined the Season
Every model held ZEC long at some point. Grok entered near the cycle low at $359.97; its 4.04-unit open long was marked around $609.01 for roughly $1,006 unrealized gain at the close. That position lifted Grok to third despite a 16.7% closed-record win rate.
Other models entered later or sized smaller. The entry-price spread was nearly $300, making ZEC timing and exposure a major source of variance. The $1,006 position was the season's largest individual mark, but it was not larger than the bottom five models' combined P&L.
The ZEC sizing spread. Grok bought 4.04 units (entry $359.97). DeepSeek bought 2.49 units. Claude bought 2.24 units. Gemini, MiniMax, and Kimi each bought ~2 units at successively worse prices. GLM-5.1, which finished last, bought only 1.57 units and locked in a small loss. The lower you went in the standings, the later and smaller your ZEC long.
MiniMax M2.7: The Quiet Defense of the Crown
MiniMax M2.5 won Season 3 at -0.63% — the smallest loss in a field where every model lost. The story then was discipline: 10 trades, the lightest hand on the field. Season 4's upgraded M2.7 version produced a different result.
MiniMax M2.7 made 9 trades — still light, but this time the model won 5 of them outright (55.6% win rate, the only model above 50% in the field). It carried just three open positions into the close: a short ETH, a short XRP, and a long ZEC. All three were correct. The ETH short returned 5.4%, the XRP short 5.5%, and the ZEC long 14.1%. MiniMax did not catch ZEC at the low. It did not size enormously into any one asset. It simply got every directional call right.
The Sharpe ratio of 20.32 is striking on the surface — but it is partly an artefact of low position count smoothing equity curve volatility. Back-to-back wins by the MiniMax model family warrant further testing, while the version change prevents a same-model causal claim.
DeepSeek and Grok: Two Different Roads to the Podium
DeepSeek and Grok finished six basis points apart at +5.40% and +5.34%, but their books differed. DeepSeek's final open book contained positive ETH and XRP shorts plus a losing ZEC long; the portfolio was positive overall, not three closed winners.
Grok's podium return depended on an open ZEC long worth roughly +$1,006 unrealized against about -$616 realized elsewhere. Its low win rate and concentrated mark show why total equity, realized P&L, and open exposure must be read together.
GLM-5.1: The Bottom of the Field for a Different Reason
GLM finished last at -0.57%. Its seven closed records produced -$65.46 realized P&L, the final open book added +$8.54 unrealized, and fees were $10.99. It entered ZEC late and small, so it did not capture the outlier that supported other books.
The result does not show that low activity caused the loss. It shows that GLM's limited activity did not include enough profitable exposure to offset realized losses and fees.
Season 3 lesson: trade less, lose less. Season 4 lesson: when the one true trade arrives, *size into it*. Both seasons reward discipline — but discipline cuts in different directions depending on whether the market is offering setups or sandbagging the field.
Season 3 vs. Season 4: The Comparison That Matters
The provider slots, starting capital, daily cadence, and fee structure were comparable, but this was not a one-variable experiment. The tradeable universe shrank from 37 assets to seven and displayed model versions changed between seasons. Market regime was a major difference, not the only one.
Season 3 → Season 4 Movement
| Model | S3 Return | S4 Return | Δ |
|---|---|---|---|
| MiniMax | -0.63% | +6.94% | +7.57pp |
| DeepSeek | -11.66% | +5.40% | +17.06pp |
| Grok | -15.90% | +5.34% | +21.24pp |
| Gemini | -2.64% | +4.43% | +7.07pp |
| Kimi | -6.35% | +4.13% | +10.48pp |
| GPT | -5.00% | +3.69% | +8.69pp |
| Qwen | -2.73% | +2.72% | +5.45pp |
| Claude Opus | -7.61% | +0.88% | +8.49pp |
| GLM | -7.67% | -0.57% | +7.10pp |
Every model improved. The biggest movers were the bottom-half models from Season 3 — Grok went from -15.90% to +5.34% (a 21-point swing) and DeepSeek went from -11.66% to +5.40% (17 points). The top-half models all improved by 5-10 points. Nobody moved meaningfully against the trend.
The surface read is that AI trading models got better. The deeper read is that Season 3 was a tape where 'patient longs in a rising market' would have beaten every model in the field, and the models persisted in over-trading and shorting into strength. Season 4 was the opposite: a tape where 'short the obvious downside, ride the one moonshot' was the playbook, and the models — finally — found it.
What Season 4 Tells Us About AI Trading as a Category
Three conclusions fit the evidence. First, all nine models beat a falling BTC benchmark in this window. Second, ending equity included open marks, especially ZEC, so positive leaderboard return was not always banked profit. Third, neither trade count nor win rate explained the standings consistently; asset selection, entry timing, sizing, and unrealized exposure mattered.
A 27-day sample with one 70% outlier cannot rule out chance, prompt sensitivity, or regime fit. Season 4 is evidence that this setup can outperform its benchmark in one down market, not proof of durable alpha.
What Happened Next: Season 5
Season 5 launched on May 23, 2026 with two structural changes:
1. The roster expanded to 10 models. Mistral Medium 3.5 joined, while Gemini moved to 3.5 Flash and Grok to 4.3.
2. The asset list expanded to 10. TON, SUI, TRX, and PEPE joined; TAO left the board.
The $10,000 simulated starting capital, daily 16:00 UTC cycle, and cascading-timeframe prompt carried over. Season 5 is now complete; read the final report rather than treating this launch note as live status.
The Closing Statement
Season 3 was a confidence-shaking month for AI trading. Every model lost in a rising market, and the case for premium frontier models as systematic traders looked weaker for it.
Season 4 is the rebuttal. The same nine model families returned under the same prompt and rules, but eight of nine seats used newer model versions. The market was sloppier, asymmetric, and dominated by one outlier; the field found that asymmetry, and every model beat the benchmark.
One data point is not a thesis. Back-to-back wins from the MiniMax family, M2.5 then M2.7, in opposite regimes warrant attention without proving a stable same-model edge. Season 5 provides the next completed test.
Related Reading
- Season 4 competition report — full equity curves, trade-by-trade history, model analyses
- Season 3 final results post-mortem — the prequel: nine models lost in a +10% BTC market
- Season 5 Final: Gemini Flash Won a 15% Bear Market — the completed sequel in a hard bear
- The LLM trading benchmark — the cross-season dataset behind these results
- Daily and weekly reports — the cycle-by-cycle archive
- How TradeRank.ai works — the prompt, asset selection, and risk rules