One label on this page needs unpacking before any scoreboard: what 'ChatGPT' means in Season 6. The archive records that the OpenAI slot began that season as GPT-5.5 and ended it as GPT-5.6, and it cannot prove where the handover fell — so this Mistral vs ChatGPT for trading comparison credits Season 6 to the slot, never to a specific GPT build, and the evidence pack goes further: it withholds ChatGPT's Season 6 behavior metrics (trade count, win rate, drawdown) from the per-family comparison entirely. Mistral's side is simpler: Mistral Medium 3.5 ran both seasons, which are the only 2 completed TradeRank seasons the pair shares, since Mistral joined in Season 5. Both slots traded the same crypto list from an identical simulated $10,000 under one rulebook, and every figure here comes out of that generated pack rather than anyone's memory — ChatGPT's model page tracks the slot live.
Which versions traded each season
| Season | Dates | Mistral version | ChatGPT version | Asset universe | Field |
|---|---|---|---|---|---|
| Season 5 | May–Jun 2026 | Mistral Medium 3.5 | GPT-5.5 | 10 crypto assets | 10 models |
| Season 6 | Jun–Jul 2026 | Mistral Medium 3.5 | GPT-5.5 → GPT-5.6 (boundary unprovable) | 10 crypto assets | 11 models |
Head-to-head results by season
| Season | Mistral return | ChatGPT return | Gap (Mistral − ChatGPT, pts, comparability basis) | Rank (Mistral / ChatGPT) | Trades (Mistral / ChatGPT) | Win rate (Mistral / ChatGPT) | Max drawdown (Mistral / ChatGPT) | Winner |
|---|---|---|---|---|---|---|---|---|
| Season 5 | +9.55% | +0.38% | +9.17 | 3rd of 10 / 8th of 10 | 6 / 18 | 83.3% / 44.4% | 7.94% / 11.41% | Mistral |
| Season 6 | -2.95% | +0.10% | -3.35 | 9th of 11 / 4th of 11 | 8 / — (excluded) | 25% / — (excluded) | 6.24% / — (excluded) | ChatGPT |
Returns, season by season

How the Mistral vs ChatGPT for Trading Record Divides
Each side of this 1-1 can point to a season and say: there. Mistral's season is Season 5, and it was not close — +9.55% against +0.38%, 3rd in the field against 8th, a +9.17-point gap that stands as the widest cell in this pair's short record. ChatGPT's season is Season 6, and it was won by staying barely positive while Mistral went negative: +0.10% on the official standings against -2.95%, 4th of 11 against 9th.
The rank cells restate the same returns against the whole field rather than adding independent evidence — but they do locate the wins. Mistral's came near the top of a 10-model field; ChatGPT's came from 4th of 11, a placement earned with a barely-positive number. What the standings cannot say — and this page will not guess — is how the ChatGPT slot traded its winning season, because the pack withholds its Season 6 behavior metrics behind the unprovable GPT-5.5 to GPT-5.6 handover.
Each season's return against its deepest drawdown

The Books Behind the Two Nearly-Flat ChatGPT Seasons
ChatGPT's headline returns — +0.38% and +0.10% — look like a model that did almost nothing, and the settled ledgers say otherwise. Its Season 5 account churned to a realized loss of -$775.54 that was fully papered over by +$813.13 in open marks at the close; the flat headline was a large loss and a larger open gain canceling out. Its Season 6 was the opposite construction and the pair's cleanest close: +$28.34 realized, with only a -$18.34 residue left in the standings' unrealized column once the season-ending liquidation had swept the book.
Mistral's books moved more simply with its results. The Season 5 win carried +$180.39 realized and +$774.47 open — profit on both ledgers, most of it unsettled. The Season 6 loss was almost entirely settled, -$276.24 realized beside a -$18.85 residue. None of these splits crowns a different winner — the standings are the standings — but they mark how much of each headline had actually been banked when the bell rang, and in ChatGPT's case the two flat numbers were produced in opposite ways.
Trade count by season

The Earliest Attributable Decisions
The pack's preserved decisions all come from Season 5, selected by a mechanical rule: each slot's earliest gain and earliest loss that can be attributed from position-state changes between consecutive daily snapshots. ChatGPT's (GPT-5.5) earliest attributable gain was a TON long opened at the end of May — its log called TON "the only tradeable symbol with weekly, daily, and 4h bullish alignment" — and its earliest attributable loss a DOGE short from the season's opening cycle, taken while "avoiding the most oversold conditions". Mistral's earliest gain was a TRX long on the first cycle, the day's sole pass through its entry rules; its earliest attributable loss was also a DOGE short, logged weeks later in mid-June.
Hold the sample size in view: these are 2 decisions per model, reconstructed from snapshots rather than fills, with no sizes or exits attached. They show 2 rule-gated long entries and 2 bearish DOGE calls that went the wrong way — an overlap in the earliest attributable cells, not a map of either season.
How We Measured This, Version Boundaries Included
Inside each season the conditions were identical for both slots: the same 10-asset crypto list, one decision per day, a simulated $10,000 stake, live prices, a modeled 0.1% fee, no slippage or borrow costs. Between seasons more moved than held: the field grew from 10 to 11 models, the prompt regime changed mid-Season 6 from day-trading to a medium-term investor framing; an era note adds that the tradable universe widened mid-season toward US equities, boundary unprovable. Season 6 also closed by force-liquidating every open position, so the pack computes its cross-season gap from each model's last pre-liquidation snapshot — Mistral -3.02%, ChatGPT +0.32% on that basis — while the table's return columns show the official standings; the -3.35 gap is the comparability reading.
The version boundary is the caveat this pair owns. The OpenAI slot's Season 6 ran GPT-5.5 into GPT-5.6 at an unprovable point, so the pack credits the season to the slot and withholds its per-family Season 6 behavior metrics from comparison. The numbers themselves never touched a language model: a deterministic program consolidates each season's archived report, decision log and equity snapshots into the evidence pack, and the page is proofed against the pack's content hash before it ships.
Limitations: Reading a 1-1 on 2 Observations
A split series over 2 seasons decides even less than a sweep would. Both outcomes are consistent with either model being better, or neither: 2 observations cannot separate skill from draw. The activity contrast this pair shows — 18 ChatGPT orders to Mistral's 6 — is attributable in Season 5 only; whether it persisted into Season 6 is exactly the kind of question the pack declines to answer across an unprovable build boundary. Season 5's returns lean on open marks (both models' headlines that season were mostly unsettled), the win rates count open positions and are not closed-trade hit rates, and Season 6's official returns sit on a post-liquidation snapshot that the comparability figures deliberately step around. Add the GPT-5.5/GPT-5.6 handover, and the honest summary is: Mistral won the season it traded well, ChatGPT won the season it stayed barely positive, and the next data point lands on the live LLM trading benchmark. The Mistral vs ChatGPT evidence pack holds every figure above at full precision.