Qwen vs GLM for Trading: A 3-0 Sweep That Lives in One Column

Alibaba's Qwen finished ahead of Zhipu AI's GLM in all 3 shared TradeRank seasons — a 3-0 head-to-head — but GLM never once closed a season in the green (-7.67%, -0.57%, -1.90%), and the drawdown and trade-count columns refused to follow the sweep.

Data Point

The Qwen vs GLM for trading question has a one-line answer and a longer honest one. The one-line: across the 3 shared TradeRank seasons, Seasons 3–5, Qwen finished ahead of GLM in every season — a 3-0 head-to-head. The longer one is what that 3-0 is and is not. It is a return result on 3 observations, over a stretch where the model builds, the tradable list and the market outcome all changed from each season to the next; Qwen is Alibaba's model, GLM is Zhipu AI's, and both traded as autonomous agents off one $10,000 stake under a single rulebook, not as anyone's advice. Every figure here comes out of a hash-locked evidence pack — linked below — not out of a model; each time a season closes, the pack is regenerated from the archive and this page reissued to match.

Qwen vs GLM for Trading: One Direction on Return, Three Unlike Seasons

The running version of this question belongs to the live LLM trading benchmark; what follows is the closed ledger underneath it, and nothing in it moves. Read the 3 seasons in order and the return order never changes. Season 3 closed red for both — Qwen -2.73%, GLM -7.67%, a Qwen edge of +4.94 points (every gap here is Qwen minus GLM). Season 4 turned green for Qwen alone, +2.72% against GLM's -0.57%, a +3.29-point margin. Season 5 widened it: Qwen +4.95% and 5th in a field of 10, GLM -1.90% and 9th, a +6.86-point gap, the largest of the run. Qwen won Season 3, won again in Season 4, and won again in Season 5 — a 3-0 assembled across the run, and the pack's rank-reversal count between the two stands at 0. Summarize those three gaps and the median is +4.94, the mean +5.03; they are the same three margins condensed two ways, both landing on Qwen, not two separate verdicts. One note on the field ranks: a season's place falls straight out of its return by arithmetic, so 5th against 9th restates the return order in the standings rather than confirming it a second time.

Head-to-head results by season

SeasonQwen returnGLM returnGap (Q−G, pts)Rank (Q / G)Trades (Q / G)Win rate (Q / G)Max drawdown (Q / G)Winner
Season 3-2.73%-7.67%+4.943rd / 7th23 / 2030.4% / 20.0%7.96% / 9.46%Qwen
Season 4+2.72%-0.57%+3.297th / 9th15 / 726.7% / 28.6%3.41% / 0.75%Qwen
Season 5+4.95%-1.90%+6.865th / 9th14 / 2350.0% / 39.1%7.99% / 7.45%Qwen

Returns, season by season

Grouped bar chart of Qwen versus GLM percentage returns for Seasons 3 to 5, with Qwen's bar above GLM's in every season and GLM below zero throughout.
Qwen's bar sits above GLM's in every season, while GLM's stays below zero throughout: -2.73% to -7.67%, then +2.72% to -0.57%, then +4.95% to -1.90%. Only Qwen's later two bars clear the zero line. Source

GLM Finished Every Shared Season in the Red

The line that carries the sweep is GLM's own column: it did not close a single shared season above water. GLM posted -7.67% (Season 3), -0.57% (Season 4) and -1.90% (Season 5) — negative all 3 times, and by field placement a 7th of 9, then 9th of 9, then 9th of 10 — the foot of the field in Season 4 and one place off it in Season 5. Qwen was not always in profit either; its -2.73% in Season 3 was a loss, and its win that season was only the less-negative of two red books. But from Season 4 on, Qwen crossed into profit twice while GLM posted a loss every season, and that is the shape of the 3-0. It is still a return result on 3 seasons, and the sections below show how little the other columns back it.

Return against maximum drawdown

Scatter of Qwen's and GLM's per-season return against maximum drawdown across Seasons 3–5.
The season winner and the shallower dip part company twice here: Qwen won every season yet took the deeper drawdown in Season 4 (3.41% to GLM's 0.75%) and Season 5 (7.99% to 7.45%), with GLM falling further only in Season 3 (9.46% to 7.96%). Maximum drawdown is the pack's only risk column. Source

The Sweep Did Not Reach the Drawdown Column

Set the maximum drawdowns beside the record and they do not sort both models the same way. Qwen finished ahead every season, yet its deepest peak-to-trough fall was the larger of the two in Season 4, 3.41% to GLM's 0.75%, and again in Season 5, 7.99% to 7.45%; GLM's was the larger only in Season 3, 9.46% to 7.96%. So the model that won on return took the deeper drawdown, by this one measure, in Season 4 and in Season 5 — the risk column and the head-to-head simply did not line up. Read it as description, not a rule: maximum drawdown is a single worst-moment number, the pack carries no volatility series next to it, and on 3 seasons one slot landing on the deeper dip most of the run is as easily these particular months as anything about the model.

What Had Settled in Qwen's Widest Win

A season's return still carries open positions at their last marks, so the headline figure and the cash actually booked are separate readings — and Season 5, Qwen's widest-margin win, is where they part. Qwen's +4.95% that season rested on unrealized marks: a realized -$623.91 underneath +$1,119.37 of open gains, netting a positive total. GLM's -1.90% was a loss overall — a booked -$488.16 under +$297.85 of open marks — but its realized figure was the less-negative of the two booked results, so on settled cash alone, in the very season Qwen won by the most, GLM's book closed nearer to flat. The earlier seasons ran the other way on the settled side: Qwen booked a realized -$337.21 to GLM's -$721.59 in Season 3, and a realized +$170.31 — its only positive booked season — against GLM's -$65.46 in Season 4. None of this moves the 3-0; every return in the table is the official mark-to-market figure, and Qwen finished ahead in all 3. It only marks where each number stood when the season ended, both accounts simulated throughout.

How Each Model Opened the Run

The pack keeps each model's first attributable gain and first attributable loss, reconstructed from day-to-day position states rather than fills, and for this pair all of them sit in Season 3's opening days. Both models opened the same way: a short on ADA in the first cycle, logged seconds apart, and by the next daily snapshot each was carrying that ADA short in gain — one opening trade, one shared early result. Their first losses split. GLM's came from a second short opened in that same first cycle, on ARB, which the next snapshot marked down — so its earliest gain and earliest loss were a pair of opening-day shorts, one paid, one charged. Qwen's first loss was an upside bet, opened later in the stretch — a long on HYPE that sat in loss by the following snapshot. So the archive's first divergence is the losing trade, not the winning one: GLM opened short on both, Qwen reached for an upside bet. Hold it to what an opening cycle can carry — a handful of decisions out of full seasons of them, no position size, no adds or trims, no hold-time in the record — and it says nothing about the weeks that actually built the 3-0.

Trade count by season

Bar chart comparing Qwen and GLM trade counts across Seasons 3 to 5.
Trade counts ran 23 to 20 in Season 3 and 15 to 7 in Season 4 (Qwen first), then flipped to 14 against GLM's 23 in Season 5. Qwen placed more trades early; GLM's busiest season was its last. The pack logs the counts and nothing tying them to either return. Source

Season line-up: the model versions behind each result

SeasonDatesQwen versionGLM versionAsset universeField
Season 3Mar–Apr 2026Qwen 3.5 PlusGLM-537 crypto assets9 models
Season 4Apr–May 2026Qwen 3.6 PlusGLM-5.17 crypto assets9 models
Season 5May–Jun 2026Qwen 3.6 PlusGLM-5.110 crypto assets10 models

How This Was Measured

Every number on this page resolves to one archived file, and none of it was written by a model. A deterministic generator extracts each completed season's report, decision log and equity snapshots, recalculates every return, rank, drawdown and realized/unrealized figure, and attests the result under a content hash — the same inputs yielding the same rows on every run, with the article's copy checked back against that hash before it ships. A language model only arranges the locked figures into sentences. Within a single season the conditions were level: Qwen and GLM read the same market data, each staked the same $10,000 of simulated capital, traded the same asset list on the same daily clock, and filled at live prices under a modeled 0.1% fee per trade, with slippage, borrow costs and market impact absent from the simulation. Each then wrote its own thesis and sent its own orders. What no season could equalize is what shifted between them — the builds, the prompt, the tradable list running 37, then 7, then 10 names, and the market outcome — which is why this reads as a repeated head-to-head across 3 seasons rather than one controlled experiment.

Limitations, and the One Column the Sweep Lives In

The honest frame for the whole 3-0 is that it lives in one column. Qwen won the return in every shared season; the risk column leaned the other way in Season 4 and Season 5, the activity column changed hands in Season 5, and the win-rate column — which the reports build with still-open positions counted as trades, so it is not a closed-trade hit rate — is not a scoreboard at all. Score the pair by return and Qwen sweeps; that is the one measure the head-to-head uses, and it is worth knowing it is only the one. The sample binds everything else: 3 shared completed seasons is 3 observations, and GLM ran as GLM-5 then GLM-5.1 while Qwen ran as 3.5 then 3.6 Plus, each season on a different asset list into a different market, so nothing here is a fixed trait of either model. The metrics carry their own limits: returns fold in unrealized P&L, which is why Qwen's Season 5 +4.95% is split from its -$623.91 realized book above; the opening decisions come from daily position states rather than fills, so a within-cycle round-trip never shows; and neither hold-time nor profit factor holds a trustworthy archive value, so both are left off the page instead of guessed. What is left is small and clean: Qwen ahead in all 3 shared seasons on return, a loser that never turned green, and a set of risk and activity columns that decline to follow the record. You can check every figure in the Qwen vs GLM for trading evidence pack.

Frequently Asked Questions

Is Qwen or GLM better for trading in this benchmark?

Qwen, by the one measure the head-to-head scores. It finished ahead of GLM in all 3 shared TradeRank seasons, a 3-0 result, and GLM did not post a positive return in any of them (-7.67%, -0.57%, -1.90%). The qualifier is that 'better' here means higher return: GLM actually carried the shallower maximum drawdown in Season 4 and Season 5, and on 3 seasons of shifting builds and markets, a return sweep is a description of these seasons rather than a fixed ranking.

What does the Qwen vs GLM for trading record cover?

The window is Seasons 3–5 — the 3 completed crypto seasons both families ran on TradeRank under one rulebook, each from a $10,000 simulated stake. The builds moved across it: Qwen went from 3.5 to 3.6 Plus, GLM from GLM-5 to GLM-5.1. Qwen finished ahead in every one, so the head-to-head closes 3-0, on season margins of +4.94, +3.29 and +6.86 points (Qwen minus GLM).

In the GLM vs Qwen for trading matchup, did GLM ever win a season?

No. Across Seasons 3–5 GLM finished behind Qwen every season and never posted a positive return, running -7.67%, then -0.57%, then -1.90% and placing 7th of 9, 9th of 9 and 9th of 10 in the field. Its closest season on return was Season 4, a -0.57% that trailed Qwen's +2.72% by +3.29 points. So the GLM vs Qwen record over these 3 shared seasons is a 3-0 for Qwen, with GLM's own book red throughout.

If Qwen swept the seasons, why did GLM draw down less?

Because return and drawdown are not the same axis. The 3-0 is a return result; maximum drawdown — the deepest peak-to-trough fall inside a season — sorts the two differently. GLM's was the smaller in Season 4 (0.75% against 3.41%) and Season 5 (7.45% against 7.99%), and Qwen's was smaller only in Season 3 (7.96% against 9.46%). Finishing a season higher and dipping less along the way are separate outcomes, and here they part. On 3 seasons that is a property of these particular months, not a standing risk edge, and the pack holds no volatility series to widen the picture.

What happened in each of the 3 shared Qwen and GLM seasons?

Start at Season 3, which both lost: Qwen 3rd of 9 at -2.73%, GLM 7th at -7.67%, a +4.94-point Qwen edge. In Season 4 only Qwen turned green — 7th at +2.72% against GLM's 9th and -0.57%, a +3.29 margin — and Season 5 was the widest, Qwen 5th of 10 at +4.95% over GLM's 9th at -1.90%, +6.86 points. Qwen won Season 3, then won again in Season 4 and again in Season 5, closing the head-to-head 3-0.

Which Qwen and GLM versions competed each season?

Each family changed build once. Season 3 put Qwen 3.5 Plus against GLM-5 on a 37-name list in a 9-model field. For Season 4 and Season 5 it was Qwen 3.6 Plus against GLM-5.1 — 7 names, then 10 in a 10-model field. Both were upgraded mid-run, so none of the results pins to a single build; it is one matchup replayed on 3 different boards.

Season 7 is live

Watch the AI models trade in real time

12 AI models trading live. Every decision logged and explained. Follow the competition on the TradeRank.ai arena.

See the live leaderboard →
← Back to The Signal