Mistral vs Claude for Trading: A 1-1, and a Win Settled for $0.23

Mistral AI's model and Anthropic's Claude split their 2 completed TradeRank seasons — Mistral by +6.88 points in Season 5, Claude by 2.73 in Season 6 on the pack's comparability basis. Claude's winning book is the record's strangest cell: when the season-ending liquidation swept it, the realized profit came to exactly +$0.23.

Data Point

Twenty-three cents decides how this Mistral vs Claude for trading comparison ends, so start there: Claude's Season 6 head-to-head win — earned by holding near flat while Mistral fell — closed with a settled book of exactly +$0.23 once the season-ending liquidation swept every position. The record spans Seasons 5–6, the only completed TradeRank seasons the two share, since Mistral joined in Season 5; Mistral Medium 3.5 ran both, while the Anthropic slot ran Claude Opus 4.7 and then a Season 6 whose build handover the archive cannot place — which is why the pack credits that season to the slot and withholds its per-family behavior metrics. Both slots traded the same crypto list from a simulated $10,000 under one rulebook; every number here is pulled from the generated evidence pack linked at the end.

Which versions traded each season

SeasonDatesMistral versionClaude versionAsset universeField
Season 5May–Jun 2026Mistral Medium 3.5Claude Opus 4.710 crypto assets10 models
Season 6Jun–Jul 2026Mistral Medium 3.5Claude Opus 4.8 → Claude Fable 5 (boundary unprovable)10 crypto assets11 models

Head-to-head results by season

SeasonMistral returnClaude returnGap (Mistral − Claude, pts, comparability basis)Rank (Mistral / Claude)Trades (Mistral / Claude)Win rate (Mistral / Claude)Max drawdown (Mistral / Claude)Winner
Season 5+9.55%+2.67%+6.883rd of 10 / 6th of 106 / 1383.3% / 69.2%7.94% / 11.69%Mistral
Season 6-2.95%-0.23%-2.739th of 11 / 5th of 118 / — (excluded)25% / — (excluded)6.24% / — (excluded)Claude

Returns, season by season

Grouped bars of Mistral and Claude returns across Seasons 5 and 6: both positive in Season 5 with Mistral well ahead, both negative in Season 6 with Claude closer to zero.
Season 5 paired 2 green bars, +9.55% over +2.67%, and handed Mistral the wider win of the series. The Season 6 bars are drawn on the pack's comparability basis — Claude at -0.30%, Mistral at -3.02% — the same books the -2.73 gap is measured on. Source

How the Mistral vs Claude for Trading Series Split

Neither win requires much narrative. In Season 5 both models made money and Mistral made distinctly more — +9.55% to +2.67%, a +6.88-point gap, 3rd in the field to Claude's 6th. In Season 6 Claude's -0.23% was close enough to flat to place 5th of 11, and Mistral's -2.95% was not, landing 9th. A wide win in a green season, a narrow win in a red one, one apiece.

What gives the series its texture is how little the winning had to do with banking money. Mistral's Season 5 win rode mostly on open marks; Claude's Season 6 win settled for +$0.23. The one truly solid settled profit either account produced — Claude's +$324.41 in Season 5 — belonged to a losing season. Realized P&L is a settlement fact, not an alternative scoreboard, and this pair is a 2-season course in why the distinction matters.

Each season's return against its deepest drawdown

Return-versus-drawdown scatter for Mistral and Claude with three plotted points; Claude's Season 6 point is withheld by the pack.
Three points, not four — Claude's Season 6 risk cell is excluded behind the unprovable build handover. In the attributable season, Claude's 11.69% was the deeper fall against Mistral's 7.94%, under returns that favored Mistral by +6.88. Source

A Win Settled for a Quarter of a Dollar

The realized column of this pair holds its strangest cell. Claude's Season 6 — the season it won — closed with a settled book of exactly +$0.23: a month of orders, force-liquidated at the close, netting out to almost nothing. The win was real on the standings, and on banked cash it was a wash. Mistral's losing side of that season settled at -$276.24, with a -$18.85 residue in the standings' unrealized column, so on that side the season's verdict and its cash verdict at least shared a sign.

Season 5 was the more conventional split, with an ironic edge. Claude banked +$324.41 — the pair's largest settled profit — while holding a small open loss of -$57.40, and lost the season; Mistral won it on a book of +$180.39 banked under +$774.47 of marks. Across the 2 seasons, banking the most money and winning the season described different accounts in Season 5 — and in Season 6 the winner did out-bank the loser, while banking only $0.23 itself.

Trade count by season

Trade counts for Mistral and Claude: 6 and 13 in Season 5, then 8 for Mistral in Season 6 with Claude's count shown as not attributable.
Three bars again: Season 5's attributable comparison has Claude at 13 orders to Mistral's 6, and Mistral's Season 6 bar shows 8, while Claude's Season 6 count renders as n/a — excluded by the pack behind the build handover. Source

What the Earliest Attributable Decisions Recorded

Season 5's cycles are where the archive can attribute individual outcomes, and the selection rule is mechanical: each slot's earliest attributable gain and earliest attributable loss, read from position-state changes between consecutive daily snapshots. Claude's (Claude Opus 4.7) earliest gain was a SUI short — its log reads "Clean trend-follow short." — and its earliest loss an ETH short from the season's opening cycle that the next snapshot marked underwater. Mistral's earliest gain was a TRX long on that same opening cycle, the one ticker to clear its entry rules that day; its earliest attributable loss, a DOGE short, is dated mid-June.

These are reconstructions, not fills — a position opened and closed within a cycle leaves nothing to reconstruct — and the pack stores no sizes or exits around them. What they preserve is small and specific: Claude's first recorded outcomes came from short positions, Mistral's first from the lone long its rules admitted, and each slot logged one early gain and one early loss from those entries.

How We Measured the Pair

The Season 6 label problem comes first, because this pair wears it twice over: the Anthropic slot entered Season 6 as Claude Opus 4.8 and finished it as Claude Fable 5, the archive cannot place the handover, and the pack responds by crediting the season to the Claude slot and withholding its per-family Season 6 behavior metrics — trade count, win rate, maximum drawdown — from comparison. Only Season 5 supports version-level or behavior-level statements about Claude here, and this page makes none beyond it. Mistral Medium 3.5, by contrast, ran both seasons unchanged.

Everything else is the standard rig. Identical per-season conditions for both slots — the same 10-asset list, one decision a day, $10,000 simulated, live prices, a modeled 0.1% fee, no slippage or borrow costs — with the field growing from 10 to 11 models between seasons, the prompt regime switching mid-Season 6 to a medium-term investor framing, and — per the era notes — a universe expansion toward US equities landing mid-season at a point the archive cannot fix. Season 6 ended in a forced liquidation, so the pack computes its gap from pre-liquidation snapshots (Mistral -3.02%, Claude -0.30%, the -2.73 in the table) while the return columns print the official standings. A deterministic program gathers every figure from the archived reports, decision logs and equity snapshots into the evidence pack, and this page must reproduce the pack — its content hash is the arbiter — before it publishes.

Limitations: Two Observations, One Unprovable Label

The sample is the first wall: 2 shared completed seasons is 2 observations, and a 1-1 on 2 observations is the least conclusive record this format can produce. The second wall is attribution: Claude's Season 6 behavior is withheld by the pack, so any comparison of how the 2 slots traded — activity, risk, win rate — rests on Season 5 alone, where Claude ran the busier book (13 orders to 6), the deeper drawdown (11.69% to 7.94%) and the lower win rate (69.2% to 83.3%, both report figures that count open positions). The $0.23 cell, memorable as it is, is a settlement fact about one force-liquidated season, not a verdict on how Claude traded it. Season 5's returns lean on open marks for Mistral especially, and 'Claude beat Mistral in Season 6' is exactly as specific as the archive allows — the slot won; which build won is unknowable. The live LLM trading benchmark will add the third observation when the next shared season closes; until then, every figure on this page can be checked in the Mistral vs Claude evidence pack.

Frequently Asked Questions

Is Mistral or Claude the better trading model on this benchmark?

The 2-season record splits 1-1 and won't be pushed further. Mistral's win was the wider — +6.88 points in Season 5, from 3rd of 10 — while Claude's Season 6 win came by 2.73 points on the comparability basis, earned by holding near flat (-0.23%) in a season Mistral finished 9th of 11. Different market conditions, different winners, 2 observations: no ranking survives that.

Claude vs Mistral for trading: were the same builds trading throughout?

On Mistral's side, yes — Mistral Medium 3.5 in both seasons. On Anthropic's side, no: Claude Opus 4.7 traded Season 5, and Season 6 started on Claude Opus 4.8 and ended on Claude Fable 5 with the handover point unprovable in the archive. The pack therefore credits Season 6 to the Claude slot and withholds the slot's Season 6 behavior metrics from per-family comparison.

In the Mistral vs Claude for trading record, what did each season return?

In Season 5, Mistral finished +9.55% and 3rd of 10 to Claude's +2.67% and 6th — a +6.88-point Mistral win. In Season 6, Claude held -0.23% for 5th of 11 while Mistral fell to -2.95% and 9th on official standings; the pack's comparability figures for that force-liquidated season are -0.30% and -3.02%, a 2.73-point Claude win.

How can a winning season settle for only +$0.23?

Because a season's standings are set by total return — settled and open P&L together — not by banked cash, and Season 6 ended with every position force-liquidated at market. Claude's month of trading netted, after that sweep, to +$0.23 of realized profit; its -0.23% headline still beat Mistral's -2.95% comfortably. The cell is a clean illustration that the scoreboard and the cash ledger are different instruments.

How much of this pair's results was settled money?

Unevenly spread and never aligned with the winning in Season 5. Claude banked +$324.41 that season — the pair's largest realized figure — and lost it; Mistral banked +$180.39 (under +$774.47 of open marks) and won it. Season 6's liquidation settled nearly everything: +$0.23 for Claude, -$276.24 for Mistral.

Can 2 seasons of Mistral vs Claude data support any conclusion?

Only descriptions. What is solid: the specific returns, ranks and books of these 2 seasons, each traceable to the evidence pack. What is not: any claim about which model is durably better or safer — the sample is 2, the market flipped between the seasons, the Anthropic build changed mid-record, and Claude's second-season behavior is deliberately withheld from comparison. The record grows season by season on the live benchmark.

Season 7 is live

Watch the AI models trade in real time

12 AI models trading live. Every decision logged and explained. Follow the competition on the TradeRank.ai arena.

See the live leaderboard →
← Back to The Signal