Season 6 by the numbers: 11 models, 29 daily cycles, 28 days (June 20 to July 18, 2026), $10,000 simulated capital each, 10 tradeable crypto assets, 159 trades, $446,184 in notional volume, $446.11 in modeled fees, 4 of 11 finishing positive. See the Season 6 results page for the full archive, or the live leaderboard for Season 7.
The Season Everyone Got the Read Right
Season 6 ran from June 20 to July 18, 2026. Eleven models each managed $10,000 across ten crypto pairs on a single daily decision at 16:00 UTC, with 0.1% fees modeled per trade and no leverage available.
The market did not cooperate with a clean story. ZEC rose 14.78% and ETH rose 6.30%, while DOGE fell 13.41% and XRP fell 5.18%. Sideways, mixed, and unhelpful: the kind of tape where being directionally right on a few names does not carry a portfolio.
The winner's margin was slim. The field agreed almost perfectly on what to do first, was proven right, and still mostly lost money.
How to Read These Standings
Read the standings as eleven provider seats rather than eleven fixed models. Three seats changed build partway through Season 6: Anthropic's on July 2 (Claude Opus 4.8 to Claude Fable 5), OpenAI's on July 10 (GPT-5.5 to GPT-5.6), and xAI's on July 14 (Grok 4.3 to Grok 4.5). In each case the account kept running, so whatever was open transferred to the incoming build instead of being closed out. The archive labels those rows GPT-5.6, Claude Opus 4.8 and Grok 4.3, so each totals four weeks of a seat rather than four weeks of the version it is named after.
The brief changed under everyone on July 9, roughly two-thirds of the way in, and on all eleven seats at once rather than a subset. Before that date the models answered a short-horizon framing; after it they worked to a medium-term investor mandate, stating a thesis and an invalidation level for each position and carrying both forward daily. A Season 6 figure therefore averages two briefs, and any comparison with Season 5, which ran entirely under the earlier one, crosses that boundary.
Season 6 Final Standings
| Rank | Model | Provider | Return | Trades | Realized P&L | Unrealized P&L | Win Rate | Max DD | Fees |
|---|---|---|---|---|---|---|---|---|---|
| 1 | Kimi K2.7 Code | Moonshot | +2.14% | 13 | +$232.85 | -$19.16 | 30.8% | 7.36% | $38.53 |
| 2 | GLM-5.2 | Zhipu AI | +1.59% | 15 | +$177.93 | -$19.01 | 33.3% | 8.51% | $38.28 |
| 3 | Qwen 3.7 Plus | Alibaba | +0.29% | 13 | +$45.18 | -$16.09 | 30.8% | 4.86% | $32.41 |
| 4 | GPT-5.6 | OpenAI | +0.10% | 18 | +$28.34 | -$18.34 | 27.8% | 6.47% | $37.06 |
| 5 | Claude Opus 4.8 | Anthropic | -0.23% | 17 | +$0.23 | -$22.76 | 35.3% | 9.45% | $46.45 |
| 6 | Gemini 3.5 Flash | -0.35% | 14 | -$17.80 | -$16.97 | 28.6% | 6.38% | $34.30 | |
| 7 | MiniMax M3 | MiniMax | -0.71% | 13 | -$54.22 | -$17.01 | 23.1% | 6.56% | $34.54 |
| 8 | Nemotron 3 Ultra | NVIDIA | -1.00% | 20 | -$75.73 | -$24.07 | 30.0% | 6.92% | $48.60 |
| 9 | Mistral Medium 3.5 | Mistral AI | -2.95% | 8 | -$276.24 | -$18.85 | 25.0% | 6.24% | $37.89 |
| 10 | DeepSeek V4 Pro | DeepSeek | -3.17% | 15 | -$285.39 | -$31.87 | 46.7% | 9.72% | $63.76 |
| 11 | Grok 4.3 | xAI | -4.08% | 13 | -$390.61 | -$17.00 | 7.7% | 7.60% | $34.29 |
This article is for educational and entertainment purposes only. Nothing here is financial advice. Trades described are from a simulated competition using live market prices and simulated capital; no real money is at risk. Past simulated performance does not predict future results.
Three Models, One Opening Trade, Three Different Seasons
On the first decision cycle, June 20 and before any seat changed build, Kimi K2.7 Code opened shorts on XRP, DOGE and BNB. So did Grok 4.3. So did DeepSeek V4 Pro. All three cited a composite score of 70/100 with weekly and daily trend agreement. Two of them logged identical RSI figures for all three coins: 41.4 on XRP, 34.0 on DOGE, 40.7 on BNB. The third cited RSI without giving numbers, describing the readings as oversold.
The read was correct. Over the season XRP fell 5.18%, DOGE fell 13.41% and BNB fell 3.27%. Three models, the same three shorts, all three going the right way.
Those models finished 1st, 11th and 10th.
Entry selection was not what separated the field, because on the opening cycle the field barely differed. What separated them was everything after the entry: what they added, what they cut, and what they refused to let go of.
“Weekly and daily EMA-26 trends are bearish-aligned for several altcoins, and XRP, DOGE, and BNB each score 70/100 with full timeframe agreement, satisfying the composite >=50 entry rule. I will open three 1x shorts, sizing each at 20% equity to reflect strong signal conviction while limiting exposure.”
“Multiple symbols show aligned bearish weekly and daily EMA-26 signals with composite scores of 70 (>=50 threshold). Weekly and daily agree on direction for XRP, DOGE, and BNB. These satisfy all entry rules for shorts.”
“Overall market is bearish with BTC scoring 55 BEARISH. Multiple altcoins show strong bearish alignment on weekly and daily trend, with composite scores of 70. While RSI levels are low, the trend momentum remains intact, and no reversal signals are present.”
The models were handed the same indicators and the same universe, and on day one they converged on an identical trade. If you are evaluating models on whether they can read a chart, Season 6 says they largely read it the same way. The spread came from position management, not analysis.
The Paper Gains Did Not Come Back
Season 5 produced a green leaderboard that came with an asterisk: eight of ten models finished positive, but the gains sat in open short positions marked to market near the lows, and only two models had positive realized P&L.
Season 6 is the other side of that trade. Every one of the eleven models closed with negative unrealized P&L, ranging from -$16.09 for Qwen to -$31.87 for DeepSeek. Not one model was carrying a profitable open book at the final snapshot.
That inverts where the season's results came from. The four models that finished positive did it on closed trades: Kimi banked +$232.85, GLM +$177.93, Qwen +$45.18 and the OpenAI seat +$28.34, each against a small open loss. The Anthropic seat is the cleanest illustration of the mechanic, finishing with +$0.23 of realized P&L, essentially flat on everything it closed, and landing at -0.23% purely on the -$22.76 its open positions were down.
Read the two seasons together and the lesson is about measurement, not skill. A season's headline returns can be made almost entirely of positions that have not settled. Season 5's were. Season 6's were not.
Where the Returns Actually Came From
| Model | Return | Realized P&L | Unrealized P&L | Read |
|---|---|---|---|---|
| Kimi K2.7 Code | +2.14% | +$232.85 | -$19.16 | Won on banked trades |
| GLM-5.2 | +1.59% | +$177.93 | -$19.01 | Won on banked trades |
| Claude Opus 4.8 | -0.23% | +$0.23 | -$22.76 | Flat closed book, sunk by open marks |
| Mistral Medium 3.5 | -2.95% | -$276.24 | -$18.85 | Fewest trades in the field (8) |
| DeepSeek V4 Pro | -3.17% | -$285.39 | -$31.87 | Best win rate, deepest open loss |
| Grok 4.3 | -4.08% | -$390.61 | -$17.00 | Won 1 position of 13 |
The Win-Rate Paradox, in Its Purest Form Yet
DeepSeek V4 Pro won 46.7% of its positions. That is the highest win rate in Season 6, in a field whose average was 29.0%. It finished 10th of 11.
The xAI seat won 7.7% of its positions, which is one position out of thirteen. It finished 11th, 0.90 percentage points behind DeepSeek.
Between those two sits the whole argument against reading win rate as skill. DeepSeek was right six times more often than the xAI seat and ended up in effectively the same place, because it also paid the highest fee bill in the competition ($63.76 across 15 trades) and carried the deepest open loss at the close (-$31.87). Being right more often did not survive contact with sizing and costs.
The winner makes the point from the other direction. Kimi K2.7 Code won 30.8% of its positions, fourth-best in the field and a shade above its 29.0% average, and finished first. It was still wrong roughly seven times in ten. What put it top was not accuracy but what it did with the trades it got right: it held them.
A caveat that matters for reading any of these figures: the season reports count still-open positions as trades, so these are not clean closed-trade hit rates. Directionally the point holds, and it is the same one Season 2 produced at a 17% versus 81% extreme.
What the Flagships Did
None of the four flagship families reached the podium, and the spread between them was narrow enough to be noise.
The OpenAI seat finished 4th at +0.10%, the best flagship result, and got there while trading more than any model except Nemotron (18 positions). The Anthropic seat finished 5th at -0.23% with the highest win rate of the four (35.3%) and the second-deepest drawdown in the field (9.45%). Google's Gemini 3.5 Flash, which won Season 5 outright at +13.76%, finished 6th at -0.35%. Three of those four rows span a build change; only the Google seat ran one version from start to finish.
The xAI seat finished last. One winning position out of thirteen and a -$390.61 realized loss are the weakest numbers in either column this season, and they belong to a row that was Grok 4.3 for its first 24 days and Grok 4.5 for its last four.
Gemini's seat is the only flagship that held one build the whole way. First place in a market where everything fell, sixth in a market that went sideways. The brief was not identical across the two seasons, so this is not a controlled comparison. But the tape moved far more than the brief did, and it remains the strongest signal Season 6 offers that these results track the market regime at least as much as the model.
Kimi's Season: Rules, Then Patience
Kimi K2.7 Code traded 13 times in 29 cycles. Its decision logs read like a checklist rather than a narrative: an explicit composite threshold for entry, an explicit requirement that weekly and daily trends agree, and an explicit filter rejecting setups where RSI suggested a counter-trend bounce rather than continuation.
What it did with those rules mattered more than the rules themselves. Having opened XRP, DOGE and BNB on cycle one, it held them through cycles where the composite score deteriorated, on the stated grounds that its exit condition was a weekly or daily trend reversal and not a score dip. When BNB's composite fell to 40 and the position was underwater, it logged the reason for staying in rather than quietly drifting.
That is a modest kind of discipline and it produced the largest realized gain in the field. It also produced a 30.8% win rate, fourth in the field and barely clear of the 29.0% average, a 7.36% maximum drawdown, and a total return of +2.14%. Season 6 did not reward brilliance. In this book it rewarded leaving a correct entry alone.
Market Context
Season 6's tape was genuinely mixed, which is why so few models cleared zero. Four of the nine tracked pairs rose and five fell, with no single direction to lean on for 28 days. PEPE was tradeable but carries no start or end benchmark in the archive, so it is absent from the table below.
Season 6 Asset Moves (June 20 to July 18, 2026)
| Asset | Start | End | Change |
|---|---|---|---|
| ZEC | $473.06 | $542.99 | +14.78% |
| ETH | $1,736.87 | $1,846.22 | +6.30% |
| SOL | $71.68 | $75.01 | +4.65% |
| SUI | $0.7086 | $0.7378 | +4.12% |
| TRX | $0.3242 | $0.3221 | -0.65% |
| TON | $1.62 | $1.60 | -1.23% |
| BNB | $586.49 | $567.32 | -3.27% |
| XRP | $1.1489 | $1.0894 | -5.18% |
| DOGE | $0.08365 | $0.07243 | -13.41% |
Three Things Season 6 Shows
Convergent analysis, divergent outcomes. The opening cycle produced near-identical trades from models built by different labs on different architectures. When the inputs are identical and the analysis is commoditized, the differentiator is position management. That is not where most model evaluation looks.
Unrealized P&L is a loan, not a result. Season 5's eight green finishes were open shorts marked near the lows. Season 6 closed with all eleven open books underwater and the winners separated by what they had actually banked. Any single-season ranking that does not split realized from unrealized is describing something less stable than it appears.
Win rate is not a dependable proxy for skill. The best hit rate in the field finished 10th; the champion was wrong seven times in ten. Season 2 pulled the two further apart still: there the best win rate finished 4th of 13 and the worst finished 5th, so the metric carried almost no ordering at all. Season 5 did not: there the top three finishers held three of the four highest win rates in the field, and the metric tracked the table closely. That it holds in some regimes and breaks in others is the finding, and it is why win rate cannot carry a ranking on its own.
Limitations
This is one 28-day window with eleven models and 159 trades, which is a small sample by any standard. The models were not held constant either, whether across seasons or within this one: the three mid-season seat changes and the July 9 change of brief are set out above the standings. Returns include unrealized P&L, and the realized/unrealized split can invert the reading, as it did here. Win rates count still-open positions, so they are not closed-trade hit rates. Hold-time is not recorded in the archive and is therefore absent rather than estimated. Fees are modeled at 0.1% per trade; slippage, market impact and borrow costs are not modeled. BTC was a context asset in Season 6 rather than a tradeable one, and no closing benchmark price is stored in the archive, so this report makes no BTC comparison.
Nothing here establishes that any model can trade profitably. It establishes what eleven models did under one shared rulebook for four weeks.
Related Reading
Season 6 continues the arc from Season 5, where Gemini won a 15% bear market on largely unrealized gains. For the win-rate argument at its most extreme, see the Season 2 win-rate paradox. For the crowding question this season raises, see our research on what happens when AI traders agree. For the standing cross-season ranking, see the best AI models for crypto trading, and for the current run see the live LLM trading benchmark.