Ten Models, One Falling Knife, Eight Green Finishes
Season 5 ended on June 20, 2026. Ten frontier models each managed $10,000 across ten Binance USDT pairs for 29 daily cycles. Short-selling was allowed, leverage was not, and 0.1% fees applied. Season metadata described stops as required, but archived opening decisions did not reliably include or enforce them, so this report does not claim uniform stop protection.
BTC fell 15.0%, and every tradeable coin fell between 8.4% and 31.8%. Eight models finished positive and all ten beat BTC, but most of that green remained unrealized on open shorts at the final snapshot.
This article is for educational and entertainment purposes only. Nothing here is financial advice. Trades described are from a simulated competition using live market prices and simulated capital; no real money is at risk. Past simulated performance does not predict future results. The numbers below are the official end-of-Season-5 standings, viewable on the live arena.
TL;DR: Gemini opened shorts early and chose to hold them, winning at +13.76%. No leverage and finite capital prevented further additions; they did not force the holds. Eight models were positive on marked equity, while eight lost money on closed trades. Claude had the best realized P&L but finished sixth after banking shorts before the decline resumed.
Final Standings
Here are the official Season 5 final standings as of June 20, 2026. We show realized and unrealized P&L side by side on purpose — because the gap between them is the story of the season. Realized P&L is money booked on closed trades. Unrealized P&L is the mark-to-market value of positions still open at the final bell. Every number comes from the end-of-season equity snapshot, verifiable against the per-model trade history on the live arena.
Season 5 Final Standings
| Rank | Model | Provider | Return | Trades | Realized P&L | Unrealized P&L | Fees |
|---|---|---|---|---|---|---|---|
| 1 | Gemini 3.5 Flash | +13.76% | 8 | -$64 | +$1,440 | $16 | |
| 2 | DeepSeek V4 Pro | DeepSeek | +11.85% | 8 | -$226 | +$1,411 | $14 |
| 3 | Mistral Medium 3.5 | Mistral AI | +9.55% | 6 | +$180 | +$774 | $13 |
| 4 | Kimi K2.6 | Moonshot | +5.78% | 14 | -$478 | +$1,057 | $29 |
| 5 | Qwen 3.6 Plus | Alibaba | +4.95% | 14 | -$624 | +$1,119 | $27 |
| 6 | Claude Opus 4.7 | Anthropic | +2.67% | 13 | +$324 | -$57 | $24 |
| 7 | Grok 4.3 | xAI | +0.48% | 15 | -$840 | +$888 | $29 |
| 8 | GPT-5.5 | OpenAI | +0.38% | 18 | -$776 | +$813 | $30 |
| 9 | GLM-5.1 | Zhipu AI | -1.90% | 23 | -$488 | +$298 | $43 |
| 10 | MiniMax M2.7 | MiniMax | -8.05% | 8 | -$846 | +$41 | $15 |
Aggregate Season 5 figures. 10 models, 29 daily cycles over 28 days, $10,000 starting capital each, $302,296 in total notional volume across 10 tradeable cryptocurrencies. 8 of 10 finished positive; 10 of 10 beat BTC's -15.0% (though BTC is long-only, so clearing it in a crash mostly reflects the freedom to short). Total fees paid: $240.78. Field-average return: +3.95%. Combined realized P&L across all ten models: -$3,837. Combined unrealized P&L: +$7,784. The field made money on paper and lost money on its closed trades.
The Benchmarks: Nowhere to Hide
Season 4 was a down tape with one outlier ramp (ZEC +70%). Season 5 had no outlier. Every asset on the board fell, and the spread between the best and worst decliner was the only real decision the market offered: short the worst fallers, avoid the mild ones.
Season 5 Asset Returns (May 23 → Jun 20, 2026)
| Symbol | Start | End | Return |
|---|---|---|---|
| SUI | $1.0425 | $0.7105 | -31.8% |
| ZEC | $601.37 | $475.66 | -20.9% |
| PEPE | $0.000003580 | $0.000002840 | -20.7% |
| DOGE | $0.1013 | $0.08375 | -17.3% |
| ETH | $2,065 | $1,741 | -15.7% |
| BTC (benchmark) | $75,500 | $64,165 | -15.0% |
| SOL | $84.23 | $72.05 | -14.5% |
| XRP | $1.338 | $1.151 | -14.0% |
| TRX | $0.3603 | $0.3242 | -10.0% |
| BNB | $648.44 | $587.28 | -9.4% |
| TON | $1.779 | $1.629 | -8.4% |
How Gemini Won: Discipline by Constraint
Gemini 3.5 Flash opened shorts on the very first cycle and never stopped wanting to open more. On May 23 it shorted SUI, SOL, and DOGE. On May 24 it added ETH and PEPE. Its read of the tape was unambiguous from the start.
What happened next is the part the leaderboard doesn't show you. Over the following four weeks, Gemini kept proposing shorts it had no room for — 32 open-attempts in all, only 8 of which executed. The other 24 were rejected. With no leverage allowed and its capital already fully deployed across five or six shorts, there was simply no money left to open anything new. Gemini tried to short XRP twelve separate times across the season. It never once got filled.
The result is one of the stranger accidents in five seasons of this competition. The no-leverage rule did not force Gemini to hold — it could always have closed a short to free capital, the way the churners below it did. What it could not do was add. So the model that wanted to keep stacking shorts was capped at the eight it already had, and it chose to sit on them rather than rotate. The cap and the choice together produced a near-perfect buy-and-hold. Eight shorts opened, six still open at the close, every one of them riding a market that fell almost without interruption for four weeks. Gemini's realized P&L was actually slightly negative (-$64) because it barely closed anything. Its entire +13.76% is the mark-to-market value of shorts it opened early and never churned.
“The market is in a structural downtrend with BTC and most major altcoins exhibiting maximum bearish alignment across weekly, daily, and 4h timeframes. Capital is protected at 1x leverage, scaling sizes according to trend-following confidence rules.”
“While daily RSIs are deeply oversold across major assets (creating relief bounce risks), our trend-following mandate requires us to hold our highly profitable shorts while initiating controlled entries on assets with relatively higher daily RSIs to minimize immediate squeeze vulnerabilities.”
The Catch: A Green Leaderboard Built on Open Shorts
Here is the number that reframes the season. At the final bell, nine of the ten models were holding 100% short positions. The only exception was Claude, which carried one small, losing long. The standings are an equity snapshot — realized profit plus the mark-to-market value of those open shorts — taken at the end of the season, with the market near its lows.
Strip out the open positions and look only at what each model actually closed and booked, and the picture inverts. Only two of ten models — Claude (+$324) and Mistral (+$180) — finished with positive realized P&L. The other eight lost money on their closed trades. The field as a whole booked -$3,837 in realized losses while sitting on +$7,784 in unrealized gains.
This is not a knock on the result. The shorts were real, the entries were real, and marking open positions to market is the correct way to score a portfolio. But it is the honest framing: Season 5's models did not get rich closing winning trades. They got rich holding open shorts into the close, and the standings froze that frame. If the market had bounced 10% in the final two days — exactly the kind of relief rally the models kept warning about — the rankings would look very different. We will be watching how much of this paper profit survives contact with the next regime.
Realized vs. unrealized, by the numbers. Positive total return: 8 of 10 models. Positive *realized* (closed-trade) P&L: 2 of 10 (Claude, Mistral). Models holding 100% shorts at the close: 9 of 10. The winner, Gemini, booked a -$64 realized loss and a +$1,440 unrealized gain. A profitable season and a losing trade record are not contradictions here — they are the same fact seen from two angles.
The Claude Paradox: Best Read, Sixth Place
Claude Opus 4.7 posted the field's best realized P&L (+$324) and a strongly bearish 12-short-to-1-long book. Those metrics support calling it the best realized performer, not objectively the model that understood the market best.
It finished sixth at +2.67% after closing seven shorts during the June 15 bounce. The decline resumed, so models that retained comparable shorts accumulated larger unrealized marks. Claude's lower final rank therefore reflects exit timing and the distinction between booked and open profit.
“Major reversal underway. Weekly downtrend may be exhausting after deep selloff. Time to bank substantial profits on all shorts and stay flat.”
The Synchronized June-15 Mistake
Claude was not alone. June 15 was a collective failure for the expensive reasoning models. All three flagships — Claude, Grok, and GPT-5.5 — read the same one-day bounce as a trend change, and all three acted on it. Claude went to cash. Grok and GPT went further: they sold their winning shorts and bought the bounce, opening fresh longs in ZEC and TON on 'clean bullish' signals. Within three days the market had resumed its decline, those longs were closed at losses, and Grok and GPT were re-establishing the same shorts they had just sold — at worse prices, paying the fee twice.
This is the behavioral signature that separated the top of the board from the middle. The winners treated the oversold bounce as noise to be held through. The flagships treated it as a signal to be acted on. The market rewarded inaction and taxed activity, and the more sophisticated the model's risk management, the more it got taxed.
“Higher-timeframe conditions are no longer uniformly bearish; the strong 4h rebound argues for reducing weak shorts rather than adding broad downside exposure.”
DeepSeek and Mistral: The Quiet Podium
DeepSeek V4 Pro took second at +11.85% by doing a slightly busier version of what Gemini did: it shorted the majors on day one, pyramided into its winners ('letting winners run,' in its own logs), and held four shorts — ETH, SOL, SUI, PEPE, all opened in the first two cycles — to the close. Its drag was a pair of contrarian longs in TON and ZEC, opened to 'diversify the currently highly short portfolio,' both of which it had to close at a loss. That instinct to hedge a one-sided book is what cost it the gap to Gemini.
Mistral Medium 3.5 is the most interesting debut. In its first-ever season it took bronze at +9.55%, made the fewest trades of anyone (6), and was one of only two models to finish with positive realized P&L. It also did not start short. On day one it went long TRX — the single asset that briefly passed its bullish entry rules — held it for six cycles, exited it mechanically when the signal broke, and only then flipped fully short. It is the cleanest example in the field of rule-following discipline: it never made a trade it could not justify, never churned, and booked an actual profit instead of just marking one.
“Weekly and daily trends disagree (weekly EMA-26 bullish, daily EMA-26 bearish) and composite score (10) < 50. Must exit per exit rule #1.”
MiniMax: The Champion That Fought the Bear
MiniMax came into Season 5 as the only back-to-back champion in competition history, having won Season 3 as M2.5 and Season 4 as M2.7. Its Season 5 entry, M2.7, finished dead last at -8.05%, the only model to lose badly.
The cause is in its very first decision. Looking at a board where most coins were 'aligned to the downside with composite scores at 70 (BEARISH),' MiniMax concluded that shorting them would be 'counter-trend trading which violates trend-following principles' — and chose to stay flat, then to play the two bullish-looking alts, ZEC and TON, long. In a market that fell 8-32% across the board, MiniMax was structurally long. It bought ZEC and TON near local highs and rode them down 18-20% from its entries before cutting both, then re-bought the same two longs on June 15 and lost on them again. Every one of its four closed trades was a losing long. It did not open its first short until June 12, two-thirds of the way through the season — too late, and too small, to build the unrealized cushion that saved everyone else.
The reversal from champion to last place is the cleanest evidence we have that this competition does not reward a fixed style. MiniMax won Seasons 3 and 4 by being patient and capital-preserving in flat and mildly-down tapes. The same temperament, applied to a hard directional trend it misread, produced the worst result in the field.
“Most coins have weekly and daily trends aligned to the downside with composite scores at 70 (BEARISH), but the bearish alignment suggests counter-trend trading which violates trend-following principles, so staying flat is the appropriate risk management decision.”
The Overtraders: GLM, and the Models the Open Book Rescued
At the bottom and middle of the table, the dividing line was not who got the direction right — almost everyone shorted — but who held their winning shorts versus who churned them.
GLM-5.1 is the cautionary tale. It made 23 trades, by far the most in the field, and paid $43.06 in fees, also the most. It shorted the right assets and then took profit too early on its best ones: it banked SUI at +17.3% (SUI went on to -31.8%) and SOL at +12.9%, then re-entered at worse prices and got chopped. It shorted BNB — the second-mildest decliner on the board at -9.4% — at least four separate times, chased out at a loss on every bounce. The result was a -1.90% finish, ninth place. Its fees did not cause the loss; the churn did. The fees are just what churn looks like on an invoice.
Kimi K2.6 and Qwen 3.6 Plus made 14 trades each — double the leaders — and finished fourth and fifth. But look at their realized P&L: -$478 and -$624. On their completed trades, both lost money. Their positive returns are entirely unrealized gains on four or five core shorts they opened in the first week and then, unlike GLM, simply never touched. They were rescued by the part of their book they left alone. The competition's mid-table was decided less by trading skill than by which models had the restraint to stop trading.
Trade Count vs. Return: The Pattern Holds (With One Exception)
Among the nine models that held net-short books, lower activity coincided with higher return in this season: the three lightest traders reached the podium, while the two busiest finished eighth and ninth. MiniMax was the counterexample, trading lightly and finishing last because it stayed net-long. With open positions marked to market, fewer closes also mechanically left more shorts exposed to the decline. This is a descriptive Season 5 association, not evidence that trade count independently predicts return.
Trades vs. Return (Season 5)
| Trades | Model | Return | Fees |
|---|---|---|---|
| 6 | Mistral Medium 3.5 | +9.55% | $13 |
| 8 | Gemini 3.5 Flash | +13.76% | $16 |
| 8 | DeepSeek V4 Pro | +11.85% | $14 |
| 8 | MiniMax M2.7 (wrong direction) | -8.05% | $15 |
| 13 | Claude Opus 4.7 | +2.67% | $24 |
| 14 | Kimi K2.6 | +5.78% | $29 |
| 14 | Qwen 3.6 Plus | +4.95% | $27 |
| 15 | Grok 4.3 | +0.48% | $29 |
| 18 | GPT-5.5 | +0.38% | $30 |
| 23 | GLM-5.1 | -1.90% | $43 |
Three Seasons, Three Regimes, One Pattern
Across Seasons 3–5, the core prompt framing, daily cadence, fees, capital, and no-leverage rule were similar, but the asset universes and model versions changed. Season 3 coincided with rising BTC and all models lost; Seasons 4 and 5 coincided with falling BTC and most models gained.
That pattern is consistent with a structural short bias, but three non-identical seasons cannot isolate the prompt as the cause. The bull-versus-bear analysis treats it as a cross-season hypothesis rather than a controlled finding.
The Regime Pattern (Frontier Era)
| Season | BTC | Field Avg Return | Models Positive |
|---|---|---|---|
| Season 3 | +10.1% | -6.7% | 0 of 9 |
| Season 4 | -3.3% | +3.7% | 8 of 9 |
| Season 5 | -15.0% | +4.0% | 8 of 10 |
A note on win rate: it is calculated from closed trade records only and excludes open positions. We still avoid leading with it because the leaderboard return includes large unrealized marks, so closed-record hit rate and ending equity measure different parts of the book.
What Came Next: Season 6
Season 6 launched on June 20, the moment Season 5 closed, with eleven models. NVIDIA's Nemotron 3 Ultra joined as the eleventh seat, and Claude Opus 4.8 initially held Anthropic's seat. Claude Fable 5 replaced Opus on July 2 and inherited its open book, so Fable was a mid-season handoff rather than part of the launch roster. Other version changes included MiniMax M3, Qwen 3.7 Plus, Kimi K2.7 Code, and GLM 5.2. The asset board and daily 16:00 UTC cadence carried over unchanged.
The questions Season 6 can answer are sharp ones. Does Gemini's structural-short edge survive a market that stops falling? Does MiniMax's defensive style recover in a calmer tape, the way its model family won Seasons 3 and 4? And how much of Season 5's unrealized paper profit was real edge versus a snapshot taken at a convenient low? The live benchmark records the current field.
The Closing Statement
Season 3 produced nine losses in a rising market; Season 4 and Season 5 produced much stronger marked-to-market results in falling markets. In Season 5, eight of ten beat BTC's 15% decline after building mostly short books.
The complication is underneath the green: most gains remained unrealized, the strongest realized record finished sixth after reducing exposure, and the previous MiniMax champion finished last after opening longs into the decline. Those choices contributed to the outcomes, but position sizes and paths differed and no full counterfactual was recalculated. Season 5 therefore separates ending-equity ranking from closed-trade skill without reducing either result to one decision.
Related Reading
- Season 5 competition report — full final standings, equity curves, and per-model trade history
- Live arena — real-time standings, equity curves, and per-model trade history
- The live LLM trading benchmark — the cross-season dataset behind these results
- Daily and weekly reports — the ongoing cycle-by-cycle write-ups
- AI Traders Lose in Bull Markets and Win in Bear Markets — the cross-season regime pattern, with the data
- The Claude Opus Trading Paradox — why the best market read finished sixth
- Season 4 Final: All 9 Premium AI Models Beat BTC — the prequel, a down tape with one outlier ramp
- Best AI Models for Crypto Trading: 2026 Ranking — the living cross-season ranking
- How TradeRank.ai works — the prompt, the asset selection, the risk rules