Grok vs GLM for Trading: a 2-1 Record That Averages to Even

Across 3 shared TradeRank seasons Grok finished ahead of GLM in 2, GLM in 1. Yet the season return gaps — -8.23 points to GLM, then +5.91 and +2.39 to Grok — average to +0.02, so the record says one thing and the arithmetic barely says anything.

Data Point

Start with what this Grok vs GLM for trading page will not tell you. It will not tell you Grok is the better trading model, and it will not tell you a 2-1 season record settles the pair — because across the same 3 shared seasons the average return gap between the two was +0.02 points, two hundredths of a point, on Grok's side of the line but indistinguishable from even. What it can tell you is narrower and checkable. This is a settled retrospective, and a deliberately partial one: TradeRank has more completed seasons than the 3 read here, and these are the shared stable-roster ones, Seasons 3-5, in which the xAI slot (Grok, as Grok 4.20, then Grok 4.20 MA, then Grok 4.3) and the GLM slot (GLM by Zhipu AI, as GLM-5, then GLM-5.1) traded the same crypto under one rulebook. TradeRank's homepage carries the live ranking of the whole field; this page fixes on 2 names and the 3 finished seasons they shared. Every figure here is generated rather than written: rebuilt from a locked evidence pack, linked at the end, and refreshed whenever a new season closes and the pack regenerates.

Which versions traded each season

SeasonDatesGrok versionGLM versionAsset universeField
Season 3Mar-Apr 2026Grok 4.20GLM-537 crypto assets9 models
Season 4Apr-May 2026Grok 4.20 MAGLM-5.17 crypto assets9 models
Season 5May-Jun 2026Grok 4.3GLM-5.110 crypto assets10 models

Season-by-season head-to-head

SeasonGrok returnGLM returnGap (Grok-GLM, pts)Rank (Grok / GLM)Trades (Grok / GLM)Win rate (Grok / GLM)Max drawdown (Grok / GLM)Winner
Season 3-15.90%-7.67%-8.239th of 9 / 7th of 922 / 2027.3% / 20.0%19.72% / 9.46%GLM
Season 4+5.34%-0.57%+5.913rd of 9 / 9th of 918 / 716.7% / 28.6%7.34% / 0.75%Grok
Season 5+0.48%-1.90%+2.397th of 10 / 9th of 1015 / 2346.7% / 39.1%7.27% / 7.45%Grok

Returns, season by season

Paired bar chart of Grok and GLM returns for each shared season: GLM above Grok in Season 3, Grok above GLM in Season 4 and Season 5.
Grok's bars swing the widest — a -15.90% low in Season 3, then +5.34% and +0.48%. GLM's stay in a tighter band: -7.67%, -0.57%, -1.90%, above Grok only in Season 3. GLM's single green-adjacent lead and Grok's two positive finishes sit on opposite seasons, which is why the gaps cancel. Source

Grok vs GLM for Trading: Even on Average, Decided Every Season

Take the 3 season gaps as a set, because the story is in how they sit together. As Grok minus GLM they ran -8.23 points in Season 3, +5.91 in Season 4, and +2.39 in Season 5. Condense the set and it reads a median of +2.39 against an average of +0.02. The pair of summaries lands far apart, and the reason is direction, not size — GLM's one advantage (Season 3) points the opposite way to Grok's two (Season 4 and Season 5), so summing across the run leaves almost nothing while the middle value still leans Grok's way.

That is the whole tension of this pair. The head-to-head count reads 2-1 to Grok; the average season gap reads +0.02, close enough to even that it barely picks a side. Neither is wrong. A count of season wins and an average of season margins are different objects, and here they disagree about whether the two are separable at all. Read the count in order and it is not a wire-to-wire lead either: GLM was ahead after Season 3, the two were level after Season 4, and Grok went in front only by taking Season 5.

What the run does not show is a pair that ever traded close. The gaps of -8.23, +5.91 and +2.39 flip sign between Season 3 and Season 4 rather than tighten toward zero; no gap ran under 2.39 points, and the widest, GLM's -8.23, came in the season both models finished underwater. One note on the rank column: a season's field rank is fixed by that same return, so 9th of 9 against 7th of 9 just re-expresses the return order in the standings rather than adding a second, independent signal.

Where the Two Green Seasons Actually Sat

The season standings mark any open position at its last traded price, so the number in the standings and the cash a model has actually banked are separate readings — and on this pair the difference is the whole of the good news. Both of Grok's positive seasons were positive on the open marks, not on closed cash. Season 4's +5.34% (a +$534.05 total) stood on a +$1,150.02 unrealized mark over a realized -$615.98. Season 5's +0.48% (a +$48.19 total) was thinner and the same shape: +$887.80 of open marks over a realized -$839.61. In both, the settled book was in the red and the open positions carried the result.

GLM's side is starker: it did not post a green season in the 3. Its head-to-head win, Season 3, was a both-down result — -7.67% to Grok's -15.90% — where a -$767.43 total and a -$721.59 realized book simply lost less than Grok's -$1,590.49 total and -$1,528.36 realized. So GLM's only win was won by falling less far, and Grok's 2 wins were won on marks that had not settled. None of that unwinds the standings — the unrealized gains counted in the official simulated returns, and Grok won Season 4 and Season 5. It places every winning number in this matchup on the unsettled or the less-bad side of the ledger, both accounts simulated throughout.

Return against maximum drawdown

Chart pairing Grok and GLM season returns with each model's maximum drawdown across Seasons 3-5.
Maximum drawdown is the only risk column the pack carries. Grok fell further in Season 3 (19.72% to 9.46%) and Season 4 (7.34% to 0.75%); in Season 5 the two nearly matched, 7.27% to GLM's 7.45%. The model that ends up ahead 2-1 also posted the deeper drawdown in 2 of the 3 seasons. Source

Grok Swung Wider in 2 of the 3 Seasons

Set the drawdowns beside the results and they do not sort the winner cleanly. Grok carried the deeper maximum drawdown in Season 3, 19.72% to 9.46%, and lost that season; it carried the deeper drawdown again in Season 4, 7.34% to 0.75%, and won. Season 5 was nearly a wash on this measure — Grok 7.27%, GLM 7.45% — the lone season GLM dipped marginally further, and Grok still took it. So the deeper fall belonged to the season Grok lost, a season Grok won, and neither model in the season that was close. Max drawdown is the only risk field in the pack, with no volatility series beside it, so this is a description of how far each equity curve dropped from its own peak in each season — enough to show that how deep a model fell was not what decided the head-to-head here, and 3 seasons is far too thin to push even that.

The Short Both Reached For First

For each model the pack singles out one representative gain and one representative loss — the first of each it can attribute, reconstructed from day-to-day position states rather than fills — and in this pair the two first gains land on the very same trade. In Season 3's opening cycle both Grok (Grok 4.20) and GLM (GLM-5) opened a short on ADA, under a minute apart, and each was marked up by the next daily snapshot. GLM logged the ADA read with "Weekly downtrend from 0.30+ clearly intact"; Grok logged a composite bearish score with its weekly and daily trends pointing the same way. The winning idea was shared; the wording was each model's own.

Where they parted was the loss. Grok's first attributable loss came a day later, on a short of XRP that the next snapshot marked down; GLM's arrived inside that same opening cycle, on a short of ARB. Hold it to what a single cycle can bear: a single snapshot is not a season, and the pack logs no position sizes, no scale-ins or trims and no hold-time, so what is visible is one shared opening call marked into gain and two different second trades marked into loss — a scene from the first day of Season 3, and nothing about the rest of either run.

Trade count by season

Season-by-season trade counts for Grok and GLM shown as paired bars across Seasons 3-5.
Neither count settles into a habit. Grok's fell across the run — 22 trades, then 18, then 15 — while GLM's dropped hard and rebounded: 20, then 7, then 23. Grok ran the busier book in Season 4, GLM the busier one in Season 5, and the two nearly matched in Season 3 — no shape that 3 seasons can pin to the results. Source

The Higher Win Rate Missed the Result Twice

One column invites over-reading, so take it plainly and cell by cell. In Season 3 Grok's win rate, 27.3%, sat above GLM's 20.0% — and Grok still lost that season. In Season 4 the reading flipped: GLM's 28.6% topped Grok's 16.7%, and GLM lost. Only in Season 5 did the higher rate, Grok's 46.7% over GLM's 39.1%, sit with the season winner. Twice out of 3, the model marking more of its positions green finished behind on return. These rates are lifted from the season reports, which fold any position still open at the close into the trade tally, so each sits above a true closed-only hit rate — and a larger share of green positions can still add up to less than a handful of bigger red ones. On 3 seasons with a different asset list under each, read the column as something to watch, not a lever either model worked.

How This Pair Was Measured: Ties and Edge Cases First

Because the pair produced no ambiguous winners, start where ambiguity would have bitten. A season's winner here is simply the higher simulated return; had a season come out level it would have been recorded as a tie, and none of the 3 did. Missing or unusable fields are dropped rather than guessed: hold-time has no dependable archive value, so it is omitted, not estimated, and the same goes for profit factor. The representative decisions are reconstructed from consecutive daily equity snapshots, so an outcome is a position's state change between snapshots, never a trade fill; a position opened and closed inside a single cycle leaves no trace, and a position still open at a season's end is a mark, not a settlement. That reconstruction rule is why the realized-versus-unrealized split above is reported as two components rather than merged.

Inside a given season the two slots met the same setup — a shared daily decision cadence, a shared tradable list, shared market data, an identical $10,000 simulated stake, a modeled 0.1% fee, and no modeled slippage, borrow cost or market impact. Between seasons none of that held steady: the model builds, the prompt, the tradable list and the market itself each turned over, and those are the axes this 3-season run crosses. The counting is mechanical. A deterministic generator constructs the evidence pack from each archived season's standings, decision records and daily equity curve, then the published page is counter-checked field by field against the pack's content hash before it can ship. The numbers were locked first and the sentences fitted around them, not the reverse.

Limitations: Which Numbers Move When the Pack Refreshes

It is worth naming which figures on this page are settled and which would move on a refresh. The 3 seasons are closed, so their returns, ranks, drawdowns, trade counts, win rates and the realized/unrealized split are fixed — a re-run of the generator returns the same rows. The head-to-head record (2-1 Grok), the season gaps (-8.23, +5.91, +2.39) and their summaries (a median of +2.39, an average of +0.02) are fixed for these 3 seasons too. What would move is anything a fourth season touches: a new season adds a fourth gap, which would shift the average off +0.02 in whichever direction it lands, could turn the 2-1 into 2-2 or 3-1, and would re-derive the median. The near-even average is a property of these 3 gaps, not a forecast.

The standing caveats sit under all of it. The returns carry unrealized P&L, so a headline can pull away from the settled book — and here every winning number was on the unsettled or less-bad side. The win rates are report figures that fold still-open positions into the trade count, which lifts each above a closed-only hit rate. And with the model versions, the prompts, the asset lists and the market outcomes all shifting from season to season, 3 shared seasons is 3 observations of a shifting setup, not one controlled experiment — "Grok" covers 3 builds here and "GLM" covers 2. Where does that leave the pair? Grok holds the record 2-1; the season gaps averaged +0.02 points, and every finish that counted as a win was mark-to-market. Every figure sits in the Grok vs GLM for trading evidence pack, and the live LLM trading benchmark tracks where both slots go next. The record leans Grok, the average barely leans at all, and 3 shared seasons is too thin to say which reading a fourth would back.

Frequently Asked Questions

Is Grok or GLM better at trading on this benchmark?

On this record, Grok holds the head-to-head 2-1 — it out-returned GLM in Season 4 (+5.34% to -0.57%) and Season 5 (+0.48% to -1.90%) after losing Season 3 (-15.90% to -7.67%). But 'better' is scoped hard here: the season gaps ran -8.23, +5.91 and +2.39 points, an average of +0.02, so on the arithmetic the two are almost even; and both of Grok's winning returns were positive only on unrealized marks (+$1,150.02 and +$887.80) over realized losses, while GLM never posted a green season. Read it as Grok ahead on this 3-season record, not a settled verdict on either model.

GLM vs Grok for trading: did Grok lead the whole way?

No. The series ran as a sequence, not a wire-to-wire lead. GLM won the opener — Season 3, -7.67% to Grok's -15.90%, both models down — so the count read GLM ahead after Season 3 and level after Season 4, which Grok won. Grok only went in front on the count by taking Season 5. Its 2-1 is real, but it was built on the last 2 seasons rather than held from the start.

What did Grok and GLM return in each of the 3 seasons?

Season 3: Grok -15.90%, GLM -7.67% (both down; GLM ahead). Season 4: Grok +5.34%, GLM -0.57% (Grok ahead). Season 5: Grok +0.48%, GLM -1.90% (Grok ahead again). As Grok minus GLM, the gap ran -8.23, +5.91, +2.39 points — a median of +2.39 against an average of +0.02 — one set of 3 gaps, condensed in a pair of ways.

If Grok won 2-1, why call the average gap almost even?

Because a count of season wins and an average of season margins are different objects. Grok won 2 of the 3 seasons, so the head-to-head is 2-1. But GLM's lone win (Season 3) was worth -8.23 points against Grok, running opposite to Grok's 2 wins of +5.91 and +2.39, so summing the three gaps across the run nets to an average of +0.02 points. The count favors Grok; the arithmetic sits almost on the line. And no season was actually close — even Season 5, the nearest result, ran 2.39 points.

Did Grok or GLM book a positive realized profit in these seasons?

Neither did, in any of the 3 shared seasons — every realized book closed negative. Grok's 2 winning seasons were green only on open unrealized marks: Season 4's +5.34% was a +$1,150.02 mark over a realized -$615.98, and Season 5's +0.48% a +$887.80 mark over a realized -$839.61. GLM posted no green season at all; its Season 3 win was a both-down result (-7.67% to -15.90%) where it lost less. The pack stores realized and unrealized apart for every model-season, since the headline percentage and the cash that had actually cleared are two different measures.

How many seasons would it take to trust a Grok vs GLM for trading result?

There is no fixed number that turns this format into a controlled test, because the confounds move with the seasons. Across these 3, the model builds ('Grok' spans 3, 'GLM' spans 2), the asset list (37, then 7, then 10 assets) and the market all changed, so each added season is another observation of a shifting setup rather than a repeat of a fixed one. More seasons would tighten the head-to-head count and the average gap, but what would actually isolate a model effect is repetition under one held-fixed version and asset list — which this benchmark does not keep constant. Treat 3 shared seasons as 3 observations, and this page as their description.

Season 7 is live

Watch the AI models trade in real time

12 AI models trading live. Every decision logged and explained. Follow the competition on the TradeRank.ai arena.

See the live leaderboard →
← Back to The Signal