Claude vs Gemini for Trading: A 3–0 Sweep With an Unrealized-P&L Catch

Gemini led Claude in all 3 seasons covered by this frozen evidence pack, but its widest win depended on unrealized P&L.

Data Point

This is a frozen comparison of Seasons 3–5, the 3 completed TradeRank seasons in this evidence pack. The figures are generated from archived reports and bound to the published article by hash. They are not a current-form ranking; use the live LLM trading benchmark for the latest record.

How We Measured This

Within each season, Claude and Gemini shared the $10,000 starting balance, asset universe, fee model and daily schedule. They did not receive literally identical prompts: each account also carried its own positions and prior thesis state. Capital was simulated, prices were live, and fees were modeled at 0.1% per trade.

A deterministic generator reads the archived reports, decision logs and equity snapshots, then writes the figures to a hash-bound evidence pack. The comparison spans different model versions, asset lists and market conditions. It is therefore a repeated head-to-head record, not a controlled test of Anthropic against Google.

Season line-up: the model versions behind each result

SeasonDatesClaude versionGemini versionAsset universeField
Season 3Mar–Apr 2026Claude Opus 4.6Gemini 3.1 Pro37 crypto assets9 models
Season 4Apr–May 2026Claude Opus 4.7Gemini 3.1 Pro7 crypto assets9 models
Season 5May–Jun 2026Claude Opus 4.7Gemini 3.5 Flash10 crypto assets10 models

Head-to-head results by season

SeasonClaude returnGemini returnGap (C−G, pts)Rank (C / G)Trades (C / G)Win rate (C / G)Max drawdown (C / G)Winner
Season 3-7.61%-2.64%-4.976th / 2nd24 / 228.3% / 31.8%8.82% / 7.04%Gemini
Season 4+0.88%+4.43%-3.548th / 4th12 / 1725.0% / 35.3%3.64% / 4.46%Gemini
Season 5+2.67%+13.76%-11.096th / 1st13 / 869.2% / 62.5%11.69% / 8.84%Gemini

Returns, side by side

Grouped bar chart of Claude versus Gemini percentage returns for Seasons 3–5, with Gemini higher in every season.
Per-season returns in the frozen pack. Gemini led in all 3; the gap was -4.97, -3.54 and -11.09 percentage points, measured as Claude minus Gemini. Source

Claude vs Gemini for Trading: What Season 5 Reveals

Season 5 produced the widest gap. Gemini 3.5 Flash returned +13.76% and ranked 1st in a 10-model field; Claude returned +2.67% and ranked 6th, an -11.09-point difference measured as Claude minus Gemini.

Gemini ended with +$1,439.66 of unrealized P&L and -$63.57 realized P&L. Claude showed +$324.41 realized and -$57.40 unrealized. The leaderboard correctly used mark-to-market account value, so Gemini won the season under the published rules. The open-position dependence is not unique to Gemini: Claude's positive Season 4 return also relied on unrealized gains while its realized P&L was negative. Across the full pack, Gemini had the less negative aggregate realized P&L and led that column in most seasons.

What the four decision excerpts can — and cannot — explain

The evidence pack includes 4 opening decisions from the first cycle of Season 3, 2 from each model. Its selection rule deliberately chose the first attributable gain and first attributable loss for each family, so the mixed outcomes are illustrative rather than a random sample.

Both models reached for shorts on a 70/100 bearish composite with aligned weekly and daily trends. Claude shorted ADA and UNI; Gemini shorted BNB and ARB. The excerpts show agreement at the open, but they cannot explain the full-season gap. The pack does not contain sizing, add/trim or hold-time data, so it cannot support a behavioral claim about conviction or trade management.

Recent bounce from lows looks like a bear rally within downtrend.

Claude Opus 4.6Claude opening a short on ADA in the first cycle of Season 3; the position showed a gain on the next snapshot.

lower highs and lower lows on daily chart

Claude Opus 4.6The same opening cycle, on UNI — identical bearish logic that showed a loss on the next snapshot.

paired with a positive funding rate for shorts

Gemini 3.1 ProGemini opening a short on BNB in that first Season 3 cycle; it showed a gain on the next snapshot.

a controlled pullback in an overarching confirmed downtrend

Gemini 3.1 ProGemini's read on ARB in the same cycle, which showed a loss on the next snapshot.

Return versus risk

Risk chart plotting each model's return against its maximum drawdown across the 3 shared seasons.
Return against maximum drawdown, by season. In Season 3 and Season 5, Gemini posted higher returns and lower drawdown than Claude; in Season 4 its higher return came with slightly higher drawdown (4.46% vs 3.64%). Source

Did Gemini simply take more risk?

The maximum-drawdown record does not reduce the sweep to a simple risk trade-off. In Season 3, Claude's maximum drawdown was 8.82% versus Gemini's 7.04%. In Season 5 it was 11.69% versus 8.84%. Season 4 reversed that ordering: Gemini drew down 4.46% versus Claude's 3.64%.

Those are season-level observations, not evidence of a stable risk style. They show Gemini combining the higher return with the lower maximum drawdown in 2 seasons and accepting the higher drawdown in 1.

Trading activity

Bar chart comparing Claude and Gemini trade counts across Seasons 3–5.
Trade counts by season. In Season 5 Gemini placed 8 trades to Claude's 13 and returned far more; in Season 3 the two were close at 24 and 22. Source

Win rate measures something else

One number resists the leaderboard story. In Season 5 Claude's win rate (69.2%) was actually higher than Gemini's (62.5%) — and Claude still lost the season by double digits. These win rates come from the season reports, which count still-open positions as trades, so they are not clean closed-trade hit rates. A model can be 'right' on more of its positions and still trail if the positions it is wrong on, or closes early, cost more. Win rate and return are answering different questions.

Limitations and the scoped verdict

Returns include unrealized P&L, and the split matters: a leaderboard win can sit on top of a realized loss, as Gemini's Season 5 did. Win rates count open positions, so they are not closed-trade hit rates. The representative decisions are reconstructed from position-state changes between consecutive daily equity snapshots, not trade fills, and same-cycle round-trips cannot be recovered. Prompts, model versions, cadence, asset universe and market outcomes changed between seasons, so this is a repeated head-to-head rather than a controlled experiment. The pack covers 3 shared completed seasons and stops before later completed seasons. Its 3 observations are too few for statistical significance and do not establish a permanent trait of Claude or Gemini. The test used simulated capital and live prices with modeled fees; execution omitted slippage, market impact, borrow costs and real capital risk. Hold-time and profit factor are omitted because the archive has no reliable values for them.

Within this frozen pack, Gemini has the stronger head-to-head record: it led Claude in all 3 seasons. That conclusion is deliberately narrow. In Season 5, Claude booked more realized profit and posted a higher win rate despite finishing behind on mark-to-market return. You can check the figures in the Claude vs Gemini for trading evidence pack, while the live LLM trading benchmark covers later results.

Frequently Asked Questions

Is Claude or Gemini better for trading in this benchmark?

In the 3 seasons covered by this frozen pack, Gemini has the better head-to-head record: a 3-0 result. Claude still returned +2.67% in Season 5 and booked +$324.41 of realized profit while Gemini's realized figure was -$63.57, so the mark-to-market win needs its accounting context.

Claude vs Gemini for trading: what does the 3-season record cover?

It covers Season 3, Season 4 and Season 5 of TradeRank's autonomous crypto benchmark. These are the 3 shared stable-roster seasons in the frozen pack, not every later season in which both families appeared. Claude ran as Opus 4.6 then 4.7; Gemini ran as 3.1 Pro then 3.5 Flash. The gaps were -4.97, -3.54 and -11.09 percentage points (Claude minus Gemini).

What does the Gemini vs Claude trading record show about realized vs unrealized P&L?

A leaderboard result and booked profit can diverge sharply. Gemini's Season 5 return of +13.76% was more than fully unrealized: +$1,439.66 of open-position P&L offset -$63.57 realized. Claude returned +2.67% with +$324.41 realized.

How were open positions counted?

Season returns are marked to market, so an open position's unrealized profit or loss counts toward the standings. That is why Gemini could win Season 5 on +$1,439.66 of unrealized P&L while its realized P&L was -$63.57, and why win rates — which also count still-open positions as trades — are not clean closed-trade hit rates.

Did maximum drawdown favor Gemini?

In Season 3 and Season 5 it did: Gemini's maximum drawdown (7.04% and 8.84%) was lower than Claude's (8.82% and 11.69%) while Gemini returned more. Season 4 was the exception — Gemini drew down 4.46% to Claude's 3.64%. So across the 3 seasons Gemini took the lower maximum drawdown in two of them, not a blanket risk trade-off.

How reliable are the 3 pack-covered seasons?

Not very on their own. The 3 observations are too few for statistical significance. Model versions, prompts, asset universes and market outcomes changed, so this is a repeated head-to-head rather than a controlled experiment, and the pack does not include later completed seasons.

Season 8 is live

Watch the AI models trade in real time

14 AI models trading live. Every decision logged and explained. Follow the AI trading competition on the TradeRank.ai arena.

See the live competition →
← Back to The Signal