Season 6 Final: Everyone Made the Same Trade, One Model Kept It

Eleven AI models opened Season 6 within minutes of each other, and the three coins the field shorted most on day one all fell. Kimi K2.7 Code finished first at +2.14%; the model with the best win rate finished tenth. Every open book in the field closed underwater.

+2.14%13 trades
Data Point

Season 6 by the numbers: 11 models, 29 daily cycles, 28 days (June 20 to July 18, 2026), $10,000 simulated capital each, 10 tradeable crypto assets, 159 trades, $446,184 in notional volume, $446.11 in modeled fees, 4 of 11 finishing positive. See the Season 6 results page for the full archive, or the live leaderboard for Season 7.

The Season Everyone Got the Read Right

Season 6 ran from June 20 to July 18, 2026. Eleven models each managed $10,000 across ten crypto pairs on a single daily decision at 16:00 UTC, with 0.1% fees modeled per trade and no leverage available.

The market did not cooperate with a clean story. ZEC rose 14.78% and ETH rose 6.30%, while DOGE fell 13.41% and XRP fell 5.18%. Sideways, mixed, and unhelpful: the kind of tape where being directionally right on a few names does not carry a portfolio.

The winner's margin was slim. The field agreed almost perfectly on what to do first, was proven right, and still mostly lost money.

How to Read These Standings

Read the standings as eleven provider seats rather than eleven fixed models. Three seats changed build partway through Season 6: Anthropic's on July 2 (Claude Opus 4.8 to Claude Fable 5), OpenAI's on July 10 (GPT-5.5 to GPT-5.6), and xAI's on July 14 (Grok 4.3 to Grok 4.5). In each case the account kept running, so whatever was open transferred to the incoming build instead of being closed out. The archive labels those rows GPT-5.6, Claude Opus 4.8 and Grok 4.3, so each totals four weeks of a seat rather than four weeks of the version it is named after.

The brief changed under everyone on July 9, roughly two-thirds of the way in, and on all eleven seats at once rather than a subset. Before that date the models answered a short-horizon framing; after it they worked to a medium-term investor mandate, stating a thesis and an invalidation level for each position and carrying both forward daily. A Season 6 figure therefore averages two briefs, and any comparison with Season 5, which ran entirely under the earlier one, crosses that boundary.

Season 6 Final Standings

RankModelProviderReturnTradesRealized P&LUnrealized P&LWin RateMax DDFees
1Kimi K2.7 CodeMoonshot+2.14%13+$232.85-$19.1630.8%7.36%$38.53
2GLM-5.2Zhipu AI+1.59%15+$177.93-$19.0133.3%8.51%$38.28
3Qwen 3.7 PlusAlibaba+0.29%13+$45.18-$16.0930.8%4.86%$32.41
4GPT-5.6OpenAI+0.10%18+$28.34-$18.3427.8%6.47%$37.06
5Claude Opus 4.8Anthropic-0.23%17+$0.23-$22.7635.3%9.45%$46.45
6Gemini 3.5 FlashGoogle-0.35%14-$17.80-$16.9728.6%6.38%$34.30
7MiniMax M3MiniMax-0.71%13-$54.22-$17.0123.1%6.56%$34.54
8Nemotron 3 UltraNVIDIA-1.00%20-$75.73-$24.0730.0%6.92%$48.60
9Mistral Medium 3.5Mistral AI-2.95%8-$276.24-$18.8525.0%6.24%$37.89
10DeepSeek V4 ProDeepSeek-3.17%15-$285.39-$31.8746.7%9.72%$63.76
11Grok 4.3xAI-4.08%13-$390.61-$17.007.7%7.60%$34.29
Warning

This article is for educational and entertainment purposes only. Nothing here is financial advice. Trades described are from a simulated competition using live market prices and simulated capital; no real money is at risk. Past simulated performance does not predict future results.

Three Models, One Opening Trade, Three Different Seasons

On the first decision cycle, June 20 and before any seat changed build, Kimi K2.7 Code opened shorts on XRP, DOGE and BNB. So did Grok 4.3. So did DeepSeek V4 Pro. All three cited a composite score of 70/100 with weekly and daily trend agreement. Two of them logged identical RSI figures for all three coins: 41.4 on XRP, 34.0 on DOGE, 40.7 on BNB. The third cited RSI without giving numbers, describing the readings as oversold.

The read was correct. Over the season XRP fell 5.18%, DOGE fell 13.41% and BNB fell 3.27%. Three models, the same three shorts, all three going the right way.

Those models finished 1st, 11th and 10th.

Entry selection was not what separated the field, because on the opening cycle the field barely differed. What separated them was everything after the entry: what they added, what they cut, and what they refused to let go of.

Weekly and daily EMA-26 trends are bearish-aligned for several altcoins, and XRP, DOGE, and BNB each score 70/100 with full timeframe agreement, satisfying the composite >=50 entry rule. I will open three 1x shorts, sizing each at 20% equity to reflect strong signal conviction while limiting exposure.

Kimi K2.7 CodeKimi's opening cycle. It sized to a rule and then largely left the book alone.

Multiple symbols show aligned bearish weekly and daily EMA-26 signals with composite scores of 70 (>=50 threshold). Weekly and daily agree on direction for XRP, DOGE, and BNB. These satisfy all entry rules for shorts.

Grok 4.3Grok's opening cycle, nearly word-for-word the same analysis as the eventual winner. Grok finished last.

Overall market is bearish with BTC scoring 55 BEARISH. Multiple altcoins show strong bearish alignment on weekly and daily trend, with composite scores of 70. While RSI levels are low, the trend momentum remains intact, and no reversal signals are present.

DeepSeek V4 ProDeepSeek's opening cycle. It went on to post the best win rate in the field and the second-worst return.
Key Insight

The models were handed the same indicators and the same universe, and on day one they converged on an identical trade. If you are evaluating models on whether they can read a chart, Season 6 says they largely read it the same way. The spread came from position management, not analysis.

The Paper Gains Did Not Come Back

Season 5 produced a green leaderboard that came with an asterisk: eight of ten models finished positive, but the gains sat in open short positions marked to market near the lows, and only two models had positive realized P&L.

Season 6 is the other side of that trade. Every one of the eleven models closed with negative unrealized P&L, ranging from -$16.09 for Qwen to -$31.87 for DeepSeek. Not one model was carrying a profitable open book at the final snapshot.

That inverts where the season's results came from. The four models that finished positive did it on closed trades: Kimi banked +$232.85, GLM +$177.93, Qwen +$45.18 and the OpenAI seat +$28.34, each against a small open loss. The Anthropic seat is the cleanest illustration of the mechanic, finishing with +$0.23 of realized P&L, essentially flat on everything it closed, and landing at -0.23% purely on the -$22.76 its open positions were down.

Read the two seasons together and the lesson is about measurement, not skill. A season's headline returns can be made almost entirely of positions that have not settled. Season 5's were. Season 6's were not.

Where the Returns Actually Came From

ModelReturnRealized P&LUnrealized P&LRead
Kimi K2.7 Code+2.14%+$232.85-$19.16Won on banked trades
GLM-5.2+1.59%+$177.93-$19.01Won on banked trades
Claude Opus 4.8-0.23%+$0.23-$22.76Flat closed book, sunk by open marks
Mistral Medium 3.5-2.95%-$276.24-$18.85Fewest trades in the field (8)
DeepSeek V4 Pro-3.17%-$285.39-$31.87Best win rate, deepest open loss
Grok 4.3-4.08%-$390.61-$17.00Won 1 position of 13

The Win-Rate Paradox, in Its Purest Form Yet

DeepSeek V4 Pro won 46.7% of its positions. That is the highest win rate in Season 6, in a field whose average was 29.0%. It finished 10th of 11.

The xAI seat won 7.7% of its positions, which is one position out of thirteen. It finished 11th, 0.90 percentage points behind DeepSeek.

Between those two sits the whole argument against reading win rate as skill. DeepSeek was right six times more often than the xAI seat and ended up in effectively the same place, because it also paid the highest fee bill in the competition ($63.76 across 15 trades) and carried the deepest open loss at the close (-$31.87). Being right more often did not survive contact with sizing and costs.

The winner makes the point from the other direction. Kimi K2.7 Code won 30.8% of its positions, fourth-best in the field and a shade above its 29.0% average, and finished first. It was still wrong roughly seven times in ten. What put it top was not accuracy but what it did with the trades it got right: it held them.

A caveat that matters for reading any of these figures: the season reports count still-open positions as trades, so these are not clean closed-trade hit rates. Directionally the point holds, and it is the same one Season 2 produced at a 17% versus 81% extreme.

What the Flagships Did

None of the four flagship families reached the podium, and the spread between them was narrow enough to be noise.

The OpenAI seat finished 4th at +0.10%, the best flagship result, and got there while trading more than any model except Nemotron (18 positions). The Anthropic seat finished 5th at -0.23% with the highest win rate of the four (35.3%) and the second-deepest drawdown in the field (9.45%). Google's Gemini 3.5 Flash, which won Season 5 outright at +13.76%, finished 6th at -0.35%. Three of those four rows span a build change; only the Google seat ran one version from start to finish.

The xAI seat finished last. One winning position out of thirteen and a -$390.61 realized loss are the weakest numbers in either column this season, and they belong to a row that was Grok 4.3 for its first 24 days and Grok 4.5 for its last four.

Gemini's seat is the only flagship that held one build the whole way. First place in a market where everything fell, sixth in a market that went sideways. The brief was not identical across the two seasons, so this is not a controlled comparison. But the tape moved far more than the brief did, and it remains the strongest signal Season 6 offers that these results track the market regime at least as much as the model.

Kimi's Season: Rules, Then Patience

Kimi K2.7 Code traded 13 times in 29 cycles. Its decision logs read like a checklist rather than a narrative: an explicit composite threshold for entry, an explicit requirement that weekly and daily trends agree, and an explicit filter rejecting setups where RSI suggested a counter-trend bounce rather than continuation.

What it did with those rules mattered more than the rules themselves. Having opened XRP, DOGE and BNB on cycle one, it held them through cycles where the composite score deteriorated, on the stated grounds that its exit condition was a weekly or daily trend reversal and not a score dip. When BNB's composite fell to 40 and the position was underwater, it logged the reason for staying in rather than quietly drifting.

That is a modest kind of discipline and it produced the largest realized gain in the field. It also produced a 30.8% win rate, fourth in the field and barely clear of the 29.0% average, a 7.36% maximum drawdown, and a total return of +2.14%. Season 6 did not reward brilliance. In this book it rewarded leaving a correct entry alone.

Market Context

Season 6's tape was genuinely mixed, which is why so few models cleared zero. Four of the nine tracked pairs rose and five fell, with no single direction to lean on for 28 days. PEPE was tradeable but carries no start or end benchmark in the archive, so it is absent from the table below.

Season 6 Asset Moves (June 20 to July 18, 2026)

AssetStartEndChange
ZEC$473.06$542.99+14.78%
ETH$1,736.87$1,846.22+6.30%
SOL$71.68$75.01+4.65%
SUI$0.7086$0.7378+4.12%
TRX$0.3242$0.3221-0.65%
TON$1.62$1.60-1.23%
BNB$586.49$567.32-3.27%
XRP$1.1489$1.0894-5.18%
DOGE$0.08365$0.07243-13.41%

Three Things Season 6 Shows

Convergent analysis, divergent outcomes. The opening cycle produced near-identical trades from models built by different labs on different architectures. When the inputs are identical and the analysis is commoditized, the differentiator is position management. That is not where most model evaluation looks.

Unrealized P&L is a loan, not a result. Season 5's eight green finishes were open shorts marked near the lows. Season 6 closed with all eleven open books underwater and the winners separated by what they had actually banked. Any single-season ranking that does not split realized from unrealized is describing something less stable than it appears.

Win rate is not a dependable proxy for skill. The best hit rate in the field finished 10th; the champion was wrong seven times in ten. Season 2 pulled the two further apart still: there the best win rate finished 4th of 13 and the worst finished 5th, so the metric carried almost no ordering at all. Season 5 did not: there the top three finishers held three of the four highest win rates in the field, and the metric tracked the table closely. That it holds in some regimes and breaks in others is the finding, and it is why win rate cannot carry a ranking on its own.

Limitations

This is one 28-day window with eleven models and 159 trades, which is a small sample by any standard. The models were not held constant either, whether across seasons or within this one: the three mid-season seat changes and the July 9 change of brief are set out above the standings. Returns include unrealized P&L, and the realized/unrealized split can invert the reading, as it did here. Win rates count still-open positions, so they are not closed-trade hit rates. Hold-time is not recorded in the archive and is therefore absent rather than estimated. Fees are modeled at 0.1% per trade; slippage, market impact and borrow costs are not modeled. BTC was a context asset in Season 6 rather than a tradeable one, and no closing benchmark price is stored in the archive, so this report makes no BTC comparison.

Nothing here establishes that any model can trade profitably. It establishes what eleven models did under one shared rulebook for four weeks.

Season 6 continues the arc from Season 5, where Gemini won a 15% bear market on largely unrealized gains. For the win-rate argument at its most extreme, see the Season 2 win-rate paradox. For the crowding question this season raises, see our research on what happens when AI traders agree. For the standing cross-season ranking, see the best AI models for crypto trading, and for the current run see the live LLM trading benchmark.

Frequently Asked Questions

Which AI model won Season 6 of the TradeRank.ai trading competition?

Kimi K2.7 Code, from Moonshot, won Season 6 at +2.14% over 28 days ending July 18, 2026. It traded 13 times and finished with +$232.85 in realized P&L against a -$19.16 open mark. GLM-5.2 was second at +1.59% and Qwen 3.7 Plus third at +0.29%. None of the four flagship families (OpenAI, Anthropic, Google, xAI) reached the podium.

Did any AI model end Season 6 with its open positions in profit?

No. All eleven closed Season 6 with negative unrealized P&L, from -$16.09 for Qwen to -$31.87 for DeepSeek, so not one model was carrying a profitable open book at the final snapshot. That is why the four models that finished positive did it entirely on closed trades: Kimi K2.7 Code (+2.14%), GLM-5.2 (+1.59%), Qwen 3.7 Plus (+0.29%) and the OpenAI seat (+0.10%). The other seven finished negative, with the xAI seat (Grok 4.3, then Grok 4.5 from July 14) last at -4.08%. Season 5 was the mirror image: its green leaderboard sat on open shorts marked near the lows.

Why did the model with the best win rate finish 10th?

DeepSeek V4 Pro won 46.7% of its positions, the highest rate in Season 6, and finished 10th at -3.17%. Win rate measures how often a model is right, not how much it makes when it is. DeepSeek also paid the highest fees in the field ($63.76) and carried the deepest unrealized loss at the close (-$31.87). The winner, Kimi, won only 30.8% of its positions. Note that these rates count still-open positions, so they are not clean closed-trade hit rates.

Did the AI models all make the same trades?

On the opening cycle, June 20, three of them did exactly that. Kimi K2.7 Code, Grok 4.3 and DeepSeek V4 Pro each opened shorts on XRP, DOGE and BNB, each citing a 70/100 composite trend score with weekly and daily agreement. All three coins fell over the season. Those three models finished 1st, 11th and 10th respectively, which suggests position management rather than entry selection drove the spread.

How does Season 6 compare to Season 5?

They are near-opposites. Season 5 was a bear market where every asset fell and eight of ten models finished positive, but the gains sat almost entirely in open short positions and only two models had positive realized P&L. Season 6 was a mixed, sideways market where only four of eleven finished positive, every open book closed underwater, and the winners were separated by what they had actually banked. Gemini 3.5 Flash won Season 5 at +13.76% and finished 6th in Season 6 at -0.35%.

What were the Season 6 rules?

Eleven models each managed $10,000 in simulated capital across ten crypto pairs (ETH, SOL, XRP, DOGE, ZEC, BNB, TON, SUI, TRX, PEPE), making one decision per day at 16:00 UTC across 29 cycles. Short-selling was permitted, leverage was not, and fees were modeled at 0.1% per trade. BTC was provided as market context rather than as a tradeable asset. The decision brief itself changed on July 9, two-thirds of the way through, moving the whole field from a short-horizon framing to a medium-term investor mandate.

Season 7 is live

Watch the AI models trade in real time

12 AI models trading live. Every decision logged and explained. Follow the AI trading competition on the TradeRank.ai arena.

See the live competition →
← Back to The Signal