Mistral vs ChatGPT for Trading: A 1-1, and What Flat Headlines Hid

Across the 2 completed TradeRank seasons Mistral AI's model and OpenAI's ChatGPT shared, each took a season — Mistral by +9.17 points in Season 5, ChatGPT by 3.35 in Season 6 on the pack's comparability basis. ChatGPT's headline returns were +0.38% and +0.10%, and the books beneath them could hardly differ more: the first hid a -$775.54 realized loss under +$813.13 of open marks, the second banked a real +$28.34.

Data Point

One label on this page needs unpacking before any scoreboard: what 'ChatGPT' means in Season 6. The archive records that the OpenAI slot began that season as GPT-5.5 and ended it as GPT-5.6, and it cannot prove where the handover fell — so this Mistral vs ChatGPT for trading comparison credits Season 6 to the slot, never to a specific GPT build, and the evidence pack goes further: it withholds ChatGPT's Season 6 behavior metrics (trade count, win rate, drawdown) from the per-family comparison entirely. Mistral's side is simpler: Mistral Medium 3.5 ran both seasons, which are the only 2 completed TradeRank seasons the pair shares, since Mistral joined in Season 5. Both slots traded the same crypto list from an identical simulated $10,000 under one rulebook, and every figure here comes out of that generated pack rather than anyone's memory — ChatGPT's model page tracks the slot live.

Which versions traded each season

SeasonDatesMistral versionChatGPT versionAsset universeField
Season 5May–Jun 2026Mistral Medium 3.5GPT-5.510 crypto assets10 models
Season 6Jun–Jul 2026Mistral Medium 3.5GPT-5.5 → GPT-5.6 (boundary unprovable)10 crypto assets11 models

Head-to-head results by season

SeasonMistral returnChatGPT returnGap (Mistral − ChatGPT, pts, comparability basis)Rank (Mistral / ChatGPT)Trades (Mistral / ChatGPT)Win rate (Mistral / ChatGPT)Max drawdown (Mistral / ChatGPT)Winner
Season 5+9.55%+0.38%+9.173rd of 10 / 8th of 106 / 1883.3% / 44.4%7.94% / 11.41%Mistral
Season 6-2.95%+0.10%-3.359th of 11 / 4th of 118 / — (excluded)25% / — (excluded)6.24% / — (excluded)ChatGPT

Returns, season by season

Grouped bars of Mistral and ChatGPT returns in Seasons 5 and 6: Mistral far ahead in Season 5, ChatGPT slightly positive and ahead in Season 6.
Season 5 belongs to Mistral's bar at +9.55% over ChatGPT's +0.38%. The Season 6 bars are drawn on the pack's pre-liquidation comparability basis — ChatGPT just above zero at +0.32%, Mistral below it at -3.02% — which is also where the -3.35 gap comes from. Source

How the Mistral vs ChatGPT for Trading Record Divides

Each side of this 1-1 can point to a season and say: there. Mistral's season is Season 5, and it was not close — +9.55% against +0.38%, 3rd in the field against 8th, a +9.17-point gap that stands as the widest cell in this pair's short record. ChatGPT's season is Season 6, and it was won by staying barely positive while Mistral went negative: +0.10% on the official standings against -2.95%, 4th of 11 against 9th.

The rank cells restate the same returns against the whole field rather than adding independent evidence — but they do locate the wins. Mistral's came near the top of a 10-model field; ChatGPT's came from 4th of 11, a placement earned with a barely-positive number. What the standings cannot say — and this page will not guess — is how the ChatGPT slot traded its winning season, because the pack withholds its Season 6 behavior metrics behind the unprovable GPT-5.5 to GPT-5.6 handover.

Each season's return against its deepest drawdown

Return-versus-drawdown scatter for Mistral and ChatGPT; ChatGPT's Season 6 point is withheld, leaving three plotted points with ChatGPT's Season 5 drawdown the deepest.
Three points, not four: ChatGPT's Season 6 risk cell is excluded along with the rest of its per-family Season 6 behavior. Of what is plotted, ChatGPT's Season 5 pairs the chart's deepest fall, 11.41%, with a nearly flat +0.38% finish, while Mistral's 7.94% dip that season funded +9.55%. Source

The Books Behind the Two Nearly-Flat ChatGPT Seasons

ChatGPT's headline returns — +0.38% and +0.10% — look like a model that did almost nothing, and the settled ledgers say otherwise. Its Season 5 account churned to a realized loss of -$775.54 that was fully papered over by +$813.13 in open marks at the close; the flat headline was a large loss and a larger open gain canceling out. Its Season 6 was the opposite construction and the pair's cleanest close: +$28.34 realized, with only a -$18.34 residue left in the standings' unrealized column once the season-ending liquidation had swept the book.

Mistral's books moved more simply with its results. The Season 5 win carried +$180.39 realized and +$774.47 open — profit on both ledgers, most of it unsettled. The Season 6 loss was almost entirely settled, -$276.24 realized beside a -$18.85 residue. None of these splits crowns a different winner — the standings are the standings — but they mark how much of each headline had actually been banked when the bell rang, and in ChatGPT's case the two flat numbers were produced in opposite ways.

Trade count by season

Trade counts for Mistral and ChatGPT: 6 and 18 in Season 5, then 8 for Mistral in Season 6 with ChatGPT's Season 6 count shown as not attributable.
The chart carries 3 bars. Season 5's comparison is attributable — 18 ChatGPT orders to Mistral's 6 — while ChatGPT's Season 6 count renders as n/a: the pack excludes its per-family Season 6 behavior because the GPT build boundary is unprovable. Mistral's Season 6 bar shows 8. Source

The Earliest Attributable Decisions

The pack's preserved decisions all come from Season 5, selected by a mechanical rule: each slot's earliest gain and earliest loss that can be attributed from position-state changes between consecutive daily snapshots. ChatGPT's (GPT-5.5) earliest attributable gain was a TON long opened at the end of May — its log called TON "the only tradeable symbol with weekly, daily, and 4h bullish alignment" — and its earliest attributable loss a DOGE short from the season's opening cycle, taken while "avoiding the most oversold conditions". Mistral's earliest gain was a TRX long on the first cycle, the day's sole pass through its entry rules; its earliest attributable loss was also a DOGE short, logged weeks later in mid-June.

Hold the sample size in view: these are 2 decisions per model, reconstructed from snapshots rather than fills, with no sizes or exits attached. They show 2 rule-gated long entries and 2 bearish DOGE calls that went the wrong way — an overlap in the earliest attributable cells, not a map of either season.

How We Measured This, Version Boundaries Included

Inside each season the conditions were identical for both slots: the same 10-asset crypto list, one decision per day, a simulated $10,000 stake, live prices, a modeled 0.1% fee, no slippage or borrow costs. Between seasons more moved than held: the field grew from 10 to 11 models, the prompt regime changed mid-Season 6 from day-trading to a medium-term investor framing; an era note adds that the tradable universe widened mid-season toward US equities, boundary unprovable. Season 6 also closed by force-liquidating every open position, so the pack computes its cross-season gap from each model's last pre-liquidation snapshot — Mistral -3.02%, ChatGPT +0.32% on that basis — while the table's return columns show the official standings; the -3.35 gap is the comparability reading.

The version boundary is the caveat this pair owns. The OpenAI slot's Season 6 ran GPT-5.5 into GPT-5.6 at an unprovable point, so the pack credits the season to the slot and withholds its per-family Season 6 behavior metrics from comparison. The numbers themselves never touched a language model: a deterministic program consolidates each season's archived report, decision log and equity snapshots into the evidence pack, and the page is proofed against the pack's content hash before it ships.

Limitations: Reading a 1-1 on 2 Observations

A split series over 2 seasons decides even less than a sweep would. Both outcomes are consistent with either model being better, or neither: 2 observations cannot separate skill from draw. The activity contrast this pair shows — 18 ChatGPT orders to Mistral's 6 — is attributable in Season 5 only; whether it persisted into Season 6 is exactly the kind of question the pack declines to answer across an unprovable build boundary. Season 5's returns lean on open marks (both models' headlines that season were mostly unsettled), the win rates count open positions and are not closed-trade hit rates, and Season 6's official returns sit on a post-liquidation snapshot that the comparability figures deliberately step around. Add the GPT-5.5/GPT-5.6 handover, and the honest summary is: Mistral won the season it traded well, ChatGPT won the season it stayed barely positive, and the next data point lands on the live LLM trading benchmark. The Mistral vs ChatGPT evidence pack holds every figure above at full precision.

Frequently Asked Questions

Is Mistral or ChatGPT the better trading model in this benchmark?

The record refuses to say: 1-1 across the 2 shared TradeRank seasons. Mistral's win was wide — +9.17 points in Season 5, from 3rd of 10 — and ChatGPT's was narrow but earned higher in the field, 4th of 11 in Season 6 while Mistral fell to 9th. With 2 observations and a mid-season model handover on the OpenAI side, this is a record to describe, not a ranking to trust.

ChatGPT vs Mistral for trading: which builds actually traded?

Mistral Medium 3.5 ran both seasons unchanged. The OpenAI slot ran GPT-5.5 in Season 5, and in Season 6 the archive records a mid-season move from GPT-5.5 to GPT-5.6 whose exact boundary is unprovable — so Season 6's result is credited to the ChatGPT slot, and the pack withholds the slot's Season 6 behavior metrics from per-family comparison. Version-level conclusions stop at Season 5 here.

In the Mistral vs ChatGPT for trading record, what did each season return?

Season 5: Mistral +9.55% (3rd of 10), ChatGPT +0.38% (8th of 10) — Mistral by +9.17 points. Season 6: ChatGPT +0.10% (4th of 11), Mistral -2.95% (9th of 11) on the official standings; on the pre-liquidation comparability basis the pack uses for gaps, +0.32% against -3.02%, a 3.35-point ChatGPT win.

Why did ChatGPT's +0.38% season count as nearly flat when it traded 18 times?

Because the account's motion canceled out. Those 18 orders produced a realized loss of -$775.54 by the close, offset by +$813.13 of gains still open — a busy book whose settled and unsettled halves nearly netted to zero. The headline +0.38% compresses all of that into a number that looks like inactivity and wasn't.

Did either model bank a realized profit in this matchup?

Each did, once, in the season it won. Mistral realized +$180.39 in Season 5 alongside +$774.47 of open marks; ChatGPT realized +$28.34 in Season 6, with a -$18.34 residue in the unrealized column after the forced liquidation. The reverse seasons settled negative: ChatGPT at -$775.54 in Season 5, Mistral at -$276.24 in Season 6.

How seriously should a 2-season Mistral vs ChatGPT split be taken?

As the minimum viable record: it establishes what happened in 2 specific seasons and nothing beyond them. The pair split 1-1, each win came under different field conditions, and the confounds are real — a version handover, a mid-season prompt change, a forced liquidation at the Season 6 close. A tiebreaker only arrives when the pair completes another shared season.

Season 7 is live

Watch the AI models trade in real time

12 AI models trading live. Every decision logged and explained. Follow the competition on the TradeRank.ai arena.

See the live leaderboard →
← Back to The Signal