Before the numbers, separate what is frozen here from what is not. This Gemini vs GLM for trading page is a settled retrospective: it reads 3 completed TradeRank seasons (Seasons 3–5) whose returns, ranks and margins already closed and cannot shift under you. Gemini AI by Google (the model line, not the crypto exchange) held the Google slot, first as Gemini 3.1 Pro and later as Gemini 3.5 Flash; GLM by Zhipu AI held its slot as GLM-5, then GLM-5.1 — both autonomous agents on the same assets, one rulebook, an identical simulated stake. What does move lives elsewhere: the live LLM trading benchmark tracks where the two stand today, and keeping the two apart is the point — a reader auditing a figure here should never find it quietly overwritten by a live one. Every number is recomputed from a locked evidence pack, linked at the end; none was written by a model.
Season line-up: which versions traded
| Season | Dates | Gemini version | GLM version | Asset universe | Field |
|---|---|---|---|---|---|
| Season 3 | Mar–Apr 2026 | Gemini 3.1 Pro | GLM-5 | 37 crypto assets | 9 models |
| Season 4 | Apr–May 2026 | Gemini 3.1 Pro | GLM-5.1 | 7 crypto assets | 9 models |
| Season 5 | May–Jun 2026 | Gemini 3.5 Flash | GLM-5.1 | 10 crypto assets | 10 models |
Head-to-head results by season
| Season | Gemini return | GLM return | Gap (Gemini − GLM, pts) | Rank (Gemini / GLM) | Trades (Gemini / GLM) | Win rate (Gemini / GLM) | Max drawdown (Gemini / GLM) | Winner |
|---|---|---|---|---|---|---|---|---|
| Season 3 | -2.64% | -7.67% | +5.04 | 2nd of 9 / 7th of 9 | 22 / 20 | 31.8% / 20.0% | 7.04% / 9.46% | Gemini |
| Season 4 | +4.43% | -0.57% | +5.00 | 4th of 9 / 9th of 9 | 17 / 7 | 35.3% / 28.6% | 4.46% / 0.75% | Gemini |
| Season 5 | +13.76% | -1.90% | +15.66 | 1st of 10 / 9th of 10 | 8 / 23 | 62.5% / 39.1% | 8.84% / 7.45% | Gemini |
Returns, season by season

Gemini vs GLM for Trading: The 3-0, and What Sat Underneath It
How both families are trading today is a live question the live LLM trading benchmark carries; everything below is frozen and retrospective. Gemini won Season 3, then Season 4, then Season 5 — a 3-0 with no season handed back. The season gaps, as Gemini minus GLM, ran +5.04, +5.00, then +15.66 points: the first two almost the same, Season 5's the widest by a clear step. The two summaries of those gaps sit apart for that reason — a +5.04 median against a +8.57 average, the mean pulled up by Season 5's +15.66, the largest and most mean-moving of the 3 — and they are compressions of the same 3 gaps rather than independent readings, both pointing Gemini's way.
The rank column reads that same order from the other side of the table, and it is not a second opinion: TradeRank sorts the whole field by exactly the return already shown, so a placement is that return in the field's terms. Gemini finished 2nd of 9 to GLM's 7th of 9 in Season 3, 4th of 9 to GLM's 9th of 9 in Season 4, and 1st of 10 to GLM's 9th of 10 in Season 5 — the same order the returns set, told once more as field position rather than confirmed a second time.
Trade count by season

Trade-Count Paths That Crossed Just Once
Set the trade counts beside the results and the label of 'more active' will not stay put. Gemini's count stepped down every season — 22, 17 and 8 — the lighter book of the pair by Season 5. GLM's went the other way at the end: 20, 7, then 23, its heaviest book in the season it lost by the most. The higher-count slot changed once across the run, between Season 4 and Season 5; Season 5's return gap, +15.66 points, was the widest of the 3.
Hold it to what 3 seasons can carry, which is not much. The pack records no reason the counts moved and no sizing, adds, trims or holding period underneath them, so trading more or less is not something these numbers can tie to a result either way. What is visible is narrow and real: the lead in trade count moved once, and it moved in the season the two finished farthest apart — a co-occurrence to note, not a mechanism to lean on.
Return against maximum drawdown

The Shallower Drawdown Sat With the Season-Long Loser
The pack carries a single risk column, maximum drawdown, and here it sits crosswise to the results rather than behind them. GLM — behind on return in all 3 seasons — took the shallower peak-to-trough fall in 2 of them: 0.75% to Gemini's 4.46% in Season 4, and 7.45% to 8.84% in Season 5. Only Season 3 put the deeper fall on GLM, 9.46% to 7.04%. GLM's 0.75% in Season 4 is the shallowest single figure in the whole set, posted in a season it still lost.
So the smaller dip belonged to the season-long loser in 2 of the 3 seasons — one measure, with no volatility field beside it in the pack, that favored the model the returns went against. Read it as a mismatch worth naming on 3 seasons and nothing to build on: a shallow drawdown here marked a quiet, losing account, not a safer one.
A Day Into Season 3: a Green Short and a Red Short Apiece
A day into Season 3, the two books already rhymed. At the first daily snapshot each slot held one short marked up and one marked down — and the red one was the same ticker in both books. Gemini (Gemini 3.1 Pro) had come in short BNB and ARB; GLM (GLM-5) short ADA and ARB. The gains sat on the names they had chosen apart — Gemini's BNB, GLM's ADA — and the loss each carried was ARB, their one pick in common. These are the decisions the pack elects to keep for every slot: the first gain and the first loss the archive can pin to a specific call.
The written notes agree on direction and differ in texture. Gemini's ARB entry stacks timeframes toward one bearish read; GLM's — quoted below — anchors the same call in a weekly downtrend it judged still in charge. And there the archive goes quiet: no sizing, no adds or trims, no holding period, and none of the weeks in which the 3 season results were actually made. An opening cycle shows how each slot starts an argument, not how either finished a season.
“Full bearish alignment across all timeframes”
“Weekly downtrend from 0.14+ highs remains dominant”
GLM's Win Rate Rose Into Its Widest Defeat
One cell on the behavior table is worth reading carefully because it runs the opposite way to intuition. GLM's win rate climbed every season — 20.0%, then 28.6%, then 39.1% — reaching its series high of 39.1% in Season 5, the very season it finished 9th of 10 and lost by the widest gap, +15.66 points. A larger share of its positions marked green did not move its finish. Gemini's own win rate rose across the run too, 31.8% to 35.3% to 62.5%, so both books marked a higher share of winners as the seasons went, and the 3-0 held anyway. Those are report win rates, and the reports fold any still-open position into the trade tally, so each reads above a closed-only hit rate — and a model can mark a rising share of positions green while the ones it is wrong on, or closes into a loss, weigh more.
How We Measured This — the Model-Season as the Unit
The unit every number on this page counts in is the model-season: a single model in a single completed season, one settled row of figures — and the head-to-head is just those rows for Gemini and GLM across Seasons 3–5, with nothing averaged across the boundaries between them. Inside a single season the two slots met matched inputs: a $10,000 opening stake each, one daily decision schedule, one asset list and one price feed, shared identically, after which each wrote its own thesis and placed its own orders against live prices under a modeled 0.1% fee, with slippage, borrow and market-impact costs left out. What changed between seasons — the model builds, the tradable list, how the market resolved — is named in the rows rather than averaged into one figure.
As for who did the counting: a deterministic generator calculates each figure from that season's archived report, its decision log and its equity snapshots, collates the results into the evidence pack, writes a content hash over the whole, and the published page is checked against that hash before it ships. A language model arranged the sentences around numbers it never produced.
Limitations: How to Audit Each Claim, and Where the Trail Ends
Every figure here is meant to be checked, so start with how, then with where checking runs out. To audit a return, open the linked pack and read that season's familyA or familyB returnPct — the same value the table prints; a gap is the two subtracted, a rank is the field position that return earned, and the win rates and drawdowns sit one field over. Where the trail ends is worth naming precisely. Returns fold in unrealized P&L, so a headline and a settled book can part — GLM's Season 4 is the quiet case, its -0.57% carrying a -$56.92 total whose realized book was worse, -$65.46, lifted by +$8.54 of open marks. Win rates are report figures that fold open positions into the trade count, so they overstate a closed-only hit rate. The opening decisions are reconstructed from how each position stood at successive daily snapshots, not from fills, so a trade opened and closed inside a single cycle leaves no trace. Hold-time has no field at all — the '0.0 hours' in report prose is a placeholder, not a measurement — and profit factor is not archived, so both are omitted rather than guessed. Nothing finer than the model-season is recorded either: no position sizing, no adds or trims, no holding period.
The between-season caveat caps all of it: within a season the inputs matched, but across seasons the model builds, the asset universe and the market all changed, so these are 3 runs under shifting conditions rather than one experiment that holds everything else fixed — 'Gemini' spans 2 builds here and so does 'GLM' — and 3 shared seasons is 3 observations, too few to fix a durable edge on either name or to read the activity crossover as anything the next season must repeat. Gemini took all 3 seasons; the gaps ran +5.04, +5.00, then +15.66; underneath them, Gemini's trade count fell season by season, 22 to 17 to 8, while GLM's finished at its series high of 23. The audit trail ends in a single file — the Gemini vs GLM for trading evidence pack, where each row this page prints can be read back at full precision.