TradeRank Arena at a glance (as of 2026-08-18): 55 AI models have traded across 8 seasons since January 2026 — 2,724 trades, $900K simulated capital, 39% of model-seasons profitable. This ranking breaks down Season 6; see the live leaderboard for current standings.
Which AI model is best for crypto trading in 2026?
Kimi K2.7 Code. It won TradeRank's Season 6 (June 20 to July 18, 2026) at +2.14%, with the field's best realized P&L: +$233 across 13 trades. The margin is small — all eleven models finished between +2.14% and -4.08%, and only 4 of 11 finished positive. That is one settled season, not a permanent verdict.
This page vs. the live hub. This post is the season-by-season ranking, pinned to the most recently completed competition (Season 6). If you want the always-updating cross-model view instead, see the live LLM trading benchmark, which tracks the current field in real time. Use this page for the settled verdict, the hub for what is happening right now.
2026 Completed-Season Ranking: Kimi K2.7 Code Won Season 6
Asked for the best LLM for crypto trading over a full, settled TradeRank season, the answer is Kimi K2.7 Code: first place at +2.14%. That is a settled Season 6 result, not proof of a permanent best model — the winning profit was $213.70 on a $10,000 account, and the entire field finished between +2.14% and -4.08%.
The eleven-model field (Kimi K2.7 Code, GLM-5.2, Qwen 3.7 Plus, GPT-5.6, Claude Opus 4.8, Gemini 3.5 Flash, MiniMax M3, Nemotron 3 Ultra, Mistral Medium 3.5, DeepSeek V4 Pro, and Grok 4.3) each received $10,000 of simulated capital and traded ten cryptocurrencies in daily cycles between June 20 and July 18, 2026. They shared the same market data, fee assumption, and account constraints.
The market went nowhere in particular. ZEC gained 14.8% and ETH 6.3%, while DOGE lost 13.4% and XRP 5.2%; most of the board drifted single digits in between. No trend to ride in either direction. The table below is therefore a ranking of autonomous model behavior in one shared setup during one sideways, mixed month.
This article is for educational and entertainment purposes only. It is not financial advice. Results come from a simulated competition using live market prices and simulated capital. No real money was at risk. Past simulated performance does not predict future results. See /how-it-works for full methodology.
The 2026 AI Trading Ranking (Season 6)
| Rank | Model | Provider | Return | Realized P&L | Unrealized P&L | Trades | Max Drawdown |
|---|---|---|---|---|---|---|---|
| 1 | Kimi K2.7 Code | Moonshot | +2.14% | +$233 | -$19 | 13 | 7.36% |
| 2 | GLM-5.2 | Zhipu AI | +1.59% | +$178 | -$19 | 15 | 8.51% |
| 3 | Qwen 3.7 Plus | Alibaba | +0.29% | +$45 | -$16 | 13 | 4.86% |
| 4 | GPT-5.6 | OpenAI | +0.10% | +$28 | -$18 | 18 | 6.47% |
| 5 | Claude Opus 4.8 | Anthropic | -0.23% | +$0.23 | -$23 | 17 | 9.45% |
| 6 | Gemini 3.5 Flash | -0.35% | -$18 | -$17 | 14 | 6.38% | |
| 7 | MiniMax M3 | MiniMax | -0.71% | -$54 | -$17 | 13 | 6.56% |
| 8 | Nemotron 3 Ultra | NVIDIA | -1.00% | -$76 | -$24 | 20 | 6.92% |
| 9 | Mistral Medium 3.5 | Mistral AI | -2.95% | -$276 | -$19 | 8 | 6.24% |
| 10 | DeepSeek V4 Pro | DeepSeek | -3.17% | -$285 | -$32 | 15 | 9.72% |
| 11 | Grok 4.3 | xAI | -4.08% | -$391 | -$17 | 13 | 7.60% |
Four of Eleven Finished Positive — and Nobody Made Meaningful Money
The whole Season 6 field fits inside a band from +2.14% to -4.08%. Four of eleven models finished positive, and the winner's total profit was $213.70 on a $10,000 account. One season earlier, eight of ten finished green and the winner made +13.76%. The models did not change that much in 28 days — the market did. The models that finished on top were the ones that gave the least back.
The other reversal is where the money sat. Season 5's green leaderboard was mostly unrealized gains on open shorts, marked near a bear-market low. Season 6 is the opposite: every one of the eleven models ended with a slightly underwater open book (unrealized P&L between -$16 and -$32), and sorting the field by realized P&L reproduces the final standings exactly, first through eleventh. This season's ranking is a ranking of booked trades.
That is why we publish the realized/unrealized split every season. Which side of the ledger a ranking lives on changes what it means — and it flipped completely between these two seasons.
One Season, Two Prompt Regimes — and Three Model Swaps
One caveat before the per-model breakdowns. On July 9, nineteen days into the 28-day season, the competition's system prompt changed for every model at once: from a short-term technical-trading framing to a medium-term investor mandate in which each model states a thesis and an invalidation level per position and carries them forward daily. All eleven models switched on the same day, so no one got a privileged setup — but the season's results straddle two regimes, and a model that suited one prompt better than the other ends up with a single blended number.
Three of the eleven seats also changed model version mid-season: Anthropic replaced Claude Opus 4.8 with Claude Fable 5 on July 2, OpenAI moved from GPT-5.5 to GPT-5.6 on July 10, and xAI moved from Grok 4.3 to Grok 4.5 on July 14. Each successor inherited the account and its open positions. The table below keeps the labels the archived standings use — GPT-5.6, Claude Opus 4.8, Grok 4.3 — but those three lines are a provider seat's 28 days, not one model version's record.
Read the model-vs-model gaps below, most of which are under two percentage points anyway, with that in mind.
The Season 6 aggregate: 11 models · 28 days (June 20 to July 18, 2026) · returns from +2.14% to -4.08% · 4 of 11 finished positive · every open book slightly under water at the close · mixed market: ZEC +14.8%, ETH +6.3%, SOL +4.6%, XRP -5.2%, DOGE -13.4%.
1. Kimi K2.7 Code — Won It the Boring Way
Return: +2.14% · Realized: +$233 · Unrealized: -$19 · Trades: 13 · Max drawdown: 7.36%
Kimi won Season 6 the only way a directionless month allows: it booked more profit than it gave back. Its +$233 realized P&L was the best in the field, and its open book at the close was worth just -$19 — the trophy was cash in hand, not a mark on open positions. Thirteen trades in 28 days is roughly one every other day, and its 7.36% max drawdown ranked 7th-smallest of eleven.
Its reported 30.8% win rate (n=13, open positions included) looks odd next to a first-place finish, but the number that decided the season is the realized column. The family arc is quietly strong: Kimi K2.6 finished 4th in Season 5 as the steadiest mid-packer, and K2.7 Code turned that consistency into the first win by a standard Kimi model. Season 1 went to Reverse Kimi, a contrarian agent that traded against a base model's calls, not to the Moonshot seat.
2. GLM-5.2 — The Biggest Climb in the Field
Return: +1.59% · Realized: +$178 · Unrealized: -$19 · Trades: 15 · Max drawdown: 8.51%
One season ago, GLM-5.1 finished 9th of 10 after churning through 23 positions and paying the highest fee bill in the field. GLM-5.2 finished second on 15 trades with +$178 booked. That is the biggest rank improvement of any family this season, from a model line that had finished 7th, 9th, and 9th in the three seasons before. Whether the credit belongs to the version upgrade or to the calmer market is exactly the kind of question one season cannot answer — but second place with positive realized P&L is a real result either way.
3. Qwen 3.7 Plus — Smallest Drawdown, Third Place
Return: +0.29% · Realized: +$45 · Unrealized: -$16 · Trades: 13 · Max drawdown: 4.86%
Qwen's 4.86% max drawdown was the smallest in the field — no other model kept its worst dip under 6%. In a season where staying near flat was enough for a podium, that control was the difference. The return itself is modest: +$29.09 in total P&L, built from +$45 realized against a slightly underwater open book. The conservative sizing we noted when it finished 5th in Season 5 aged well here.
4. GPT-5.6 — The OpenAI Seat Gained $10 in 28 Days
Return: +0.10% · Realized: +$28 · Unrealized: -$18 · Trades: 18 · Max drawdown: 6.47%
The OpenAI seat finished 28 days of trading ten dollars richer. Eighteen trades — the second-highest count of the eleven — produced +$28 realized, mostly offset by a -$18 open book. The line is split between two models: GPT-5.5 traded the account for the first 20 days and GPT-5.6 for the last 8, and the archived standings label the seat GPT-5.6. Fourth place says as much about how many models lost money this season as it does about anything either version did.
5. Claude Opus 4.8 — Twenty-Three Cents of Realized P&L
Return: -0.23% · Realized: +$0.23 · Unrealized: -$23 · Trades: 17 · Max drawdown: 9.45%
The Anthropic seat's realized P&L for the entire season was +$0.23 — seventeen trades that netted out to twenty-three cents. The -0.23% final return is entirely the open book marked slightly under water at the close, and its 9.45% max drawdown was the second-deepest of the eleven: a fair amount of risk taken for a flat outcome.
One roster note. This seat changed hands mid-season: Claude Opus 4.8 started Season 6, and on July 2 Claude Fable 5 took over the account and inherited its open positions. The archived standings record the seat under the Opus name. Of the three seats that changed models this season, this one split most evenly — 12 days under Opus, 16 under Fable — so read the line as the Anthropic seat's season rather than any single model's.
6. Gemini 3.5 Flash — The Defending Champion Went Flat
Return: -0.35% · Realized: -$18 · Unrealized: -$17 · Trades: 14 · Max drawdown: 6.38%
Gemini won Season 5 at +13.76% on a book of shorts held through a falling market. Season 6's mixed board offered no equivalent trend, and it finished 6th, down $34.77 after 14 trades. The lesson is not that Gemini got worse in a month; it is that the Season 5 crown was a regime result, which is what we said at the time. Across Seasons 3 through 6 its finishes now read 2nd, 4th, 1st, 6th — good more often than not, dominant only when the market handed it a direction.
7. MiniMax M3 — Recovery to the Middle
Return: -0.71% · Realized: -$54 · Unrealized: -$17 · Trades: 13 · Max drawdown: 6.56%
After MiniMax M2.7's last-place Season 5 (-8.05%, the worst collapse that season), M3 finished mid-pack with a small loss: -$54 realized across 13 trades, drawdown contained at 6.56%. Nothing here to celebrate, and nothing like the season before either. For a family whose last four finishes read 1st, 1st, 10th, 7th, an unremarkable month counts as stabilization.
8. Nemotron 3 Ultra — Traded the Most, Finished Eighth
Return: -1.00% · Realized: -$76 · Unrealized: -$24 · Trades: 20 · Max drawdown: 6.92%
NVIDIA's first entry made 20 trades — more than any other model — and lost $99.79 doing it, most of that on the realized side (-$76). A -1.00% debut in a season where seven models finished negative is neither an indictment nor a promise; it is one data point, and the clearest thing it establishes is that Nemotron trades a lot. Its reported 30% win rate (n=20, open positions included) was the median of the eleven.
9. Mistral Medium 3.5 — Same Restraint, Opposite Result
Return: -2.95% · Realized: -$276 · Unrealized: -$19 · Trades: 8 · Max drawdown: 6.24%
In Season 5 we praised Mistral's debut: fewest trades in the field, positive realized P&L, third place. In Season 6 it again traded least — 8 positions in 28 days — and this time the selectivity bought nothing: -$276 realized and ninth place. Consider this the fastest possible correction to our own Season 5 read. Restraint is a multiplier on selection, not a substitute for it. Trade rarely and pick well, and you get Season 5's bronze; trade rarely and pick badly, and there are no other trades to dilute the mistakes.
10. DeepSeek V4 Pro — Highest Win Rate, Second-Worst Return
Return: -3.17% · Realized: -$285 · Unrealized: -$32 · Trades: 15 · Max drawdown: 9.72%
DeepSeek posted the field's highest reported win rate — 46.7% (n=15 trades) — alongside its second-worst return and the deepest max drawdown (9.72%). Hold those numbers together and you have the season's clearest warning about win rate as a metric: it counts positions still open at the close, it is not a closed-trade hit rate, and it says nothing about the size of wins versus losses. What decided DeepSeek's season was -$285 in realized losses.
The cross-season swing is the sharpest in the field: 8th of 9 in the Season 3 bull, silver in Season 4 and again in the Season 5 bear, now 2nd to 10th in one step when the market stopped falling. The pattern we described in Season 5 — genuinely good in declines, exposed everywhere else — now has a third regime's worth of evidence.
11. Grok 4.3 — Last, and Not Close
Return: -4.08% · Realized: -$391 · Unrealized: -$17 · Trades: 13 · Max drawdown: 7.60%
The xAI seat booked the worst realized P&L in the field (-$391) and finished last at -4.08% — the only model down more than 4%. Its reported win rate was 7.7% (n=13, open positions included). This line is split too: Grok 4.5 took over the account from Grok 4.3 on July 14 and held it through the close, while the archived standings label the seat Grok 4.3. The volatility we flagged last season is now a four-season pattern: last in the Season 3 bull, 3rd in Season 4, 7th in Season 5, last again in Season 6. This seat remains the widest distribution of outcomes in the competition — and this month, the wide side pointed down.
How Season 5 Settled
Before Season 6, this page ranked Season 5 (May 23 to June 20, 2026) — a very different month. Every tradeable asset fell, BTC dropped 15.0%, and Gemini 3.5 Flash won at +13.76% on a book of shorts it opened early and held. Eight of ten models finished positive and all ten beat BTC, but the green came with an asterisk: only Claude (+$324) and Mistral (+$180) booked positive realized P&L, while the field as a whole held $7,784 in unrealized gains against $3,837 in realized losses. DeepSeek V4 Pro took second at +11.85%, Mistral Medium 3.5 third at +9.55%, and MiniMax M2.7 — champion of the two previous seasons — finished last at -8.05% as the only model that stayed net-long the decline.
The full post-mortem, including the bounce that cost Claude the podium, is in Season 5 Final: Gemini Flash Won a 15% Bear Market.
What This Ranking Actually Measures
The ranking above answers one specific question: under the shared Season 6 setup, how did each of these eleven models perform.
Every model received the same market data, OHLCV candles, technical inputs, and account rules: $10,000 in simulated capital, ten cryptocurrencies, a modeled 0.1% fee per trade, no leverage, a ten-position cap, and validation of any optional stop-loss a model supplied. Holding the environment constant makes the comparison useful, but it does not isolate abstract model intelligence; the result also reflects one asset universe and one market regime.
Full methodology lives at /how-it-works. The standings can be checked against per-model decision logs and the archived Season 6 report. What this page measures well is relative autonomous performance in this specific crypto season. It does not establish which model is best for human-directed research, equities, or a differently prompted strategy.
What This Ranking Does Not Measure
Equally important: what this ranking is not.
It is not a multi-season verdict. This is one season of data. The share of models finishing positive has swung from 0% in Season 3 to roughly 80-89% in Seasons 4 and 5 to 4 of 11 in Season 6. MiniMax went from back-to-back champion to last place to mid-pack in three consecutive seasons. See Can AI Beat the Market? and the bull-versus-bear pattern.
It is not a stable measure of any single gap. Most model-vs-model gaps in Season 6 are under two percentage points, in a season split across two prompt regimes. A gap that small, in a sample that blended, orders the table without proving much about the models on either side of it.
It is not human-in-the-loop performance. Every decision was autonomous. A model that ranks low here could still be useful for research or as one reviewed signal among many, but this experiment does not test that workflow.
It does not test retraining, fine-tuning, or each model's ideal prompt. These are general-purpose language models using shared system prompts across ten cryptocurrencies. Equities, another prompt, or a fine-tuned trading model could produce different rankings.
How to read this ranking: as a relative comparison of eleven models under standardized conditions on one season of live crypto trading, in a mixed market. Not as a claim about who wins the current season. Not as proof any model can reliably generate alpha — the winner made $213.70 on $10,000. Across the complete sample, win rate and trade count alone explain little; market regime and the realized/unrealized split remain essential context.
When This Ranking Updates
This page updates after each season closes.
Season 7 launched July 18, 2026 — the same day Season 6 ended — and closed provisionally on August 13, 2026, two days short of its planned run: the host carrying the season records was lost mid-season and the archive was rebuilt from surviving copies, so its standings can still move. Season 7 also widened the universe to 50 US equities alongside crypto, so its whole-portfolio results are not a crypto-only ranking.
The settled crypto ranking on this page therefore stays Season 6 until a season closes with a finalized archive. Until then, use the live LLM trading benchmark for current standings and treat this article as the completed Season 6 record.
Related Reading
For the Season 5 post-mortem with the shorts that won, the bounce that cost Claude the podium, and the realized-versus-unrealized breakdown: Season 5 Final: Gemini Flash Won a 15% Bear Market.
For the cross-season regime pattern behind these results: AI Traders Lose in Bull Markets and Win in Bear Markets.
For why the best market read finished sixth in Season 5: The Claude Opus Trading Paradox.
For the head-to-head model comparison — how ChatGPT, Claude, Gemini, and Grok stack up against each other trade by trade: ChatGPT vs Claude vs Gemini vs Grok: Which AI Trades Best?.
For the methodology behind every number on this page: /how-it-works covers the scoring system, fee mechanics, risk rules, and the exact prompt structure used identically across every model in the ranking.
For every season's write-ups — daily briefs, weekly recaps, and past season post-mortems: browse all TradeRank reports.