Ensembling is the most natural idea in machine learning: if one model is noisy, ask several and trust the overlap. Applied to trading, the pitch is obvious: when GPT, Claude, Gemini, Grok, and the rest of a roster that has grown from nine models to twelve across these five seasons all short the same coin on the same day, surely that agreement means something.
We can now test that pitch on live data. TradeRank runs a paper-trading competition where LLM agents make real daily decisions on the same market data under identical rules. Across seasons 3 through 7 (March 23 to August 2, 2026 — Season 7 is still in progress; data runs through the August 2 snapshot), the models cast 494 verified directional votes across 220 cycle-assets (one asset in one daily cycle). That is enough to answer two questions: do AI models actually agree, and is the agreement worth anything?
The answers are stranger than the pitch. The models agree far more than they should — in five seasons they have taken opposite sides exactly once. And the agreement never separates itself from a matched always-bearish baseline once you account for what the market was already doing.
This is a living study. Every statistic below is read from a generated dataset (published here, snapshot August 2, 2026) rather than computed in the article; where the text quotes a model's reasoning, a roster size, or an earlier snapshot of this same dataset, it says so. Current standings are on the live leaderboard.
Finding 1: In 220 Cycle-Assets, AI Models Opposed Each Other Once
For four seasons this section reported an absolute: zero opposition, ever. That is no longer true, and the exception is worth more than the streak was.
On July 15, 2026 — cycle 26 of Season 6 — three models traded ETH. Two opened shorts. MiniMax M3 opened a long. It is the first and so far only time in 220 cycle-assets and 494 directional votes that any model has taken the other side of another model's trade in the same cycle. ETH then fell 2.6% over the next cycle and 4.5% over three, so the majority was right and the lone dissenter was wrong at both horizons that resolved before the season ended.
One event is not a finding about contrarians. With a single opposed cycle-asset in the entire dataset, the question everyone asks — does the lone dissenter beat the crowd? — still cannot be answered from this data. What the event does establish is the rate: opposition is possible, nothing in the rules prevents it, and it has happened once in five seasons. Competition rules cap new positions per cycle and gate low-confidence opens, but validation is entirely per-model, and nothing stops one model shorting what another buys.
Human fund managers routinely hold opposite views of the same asset; these models essentially never do. That near-absence is also not what "consensus" usually means, because consensus implies a disagreement that got resolved. What these models produce is better described as crowding: when one model likes a setup, the only question is how many others pile in behind it.
Where does disagreement go? Into abstention. A model that doesn't share the crowd's view almost always just holds. It stays silent rather than taking the opposite side. So a headline like "8 models agree DOGE is going down" quietly hides the real distribution: 8 models shorted, and the rest stayed silent (which no dashboard will ever show you). Abstention is ambiguous — a silent model may dissent, may already hold the position, or may be at its position cap — which is why agreement counts overstate independent confirmation.
A related number stayed at zero: no model has ever contradicted itself inside one cycle-asset, opening long and short on the same asset in the same cycle. Across 494 votes, conflicting self-votes remain 0.
How Big Do the Crowds Get?
| Models voting on the asset | Cycle-assets (all-time) |
|---|---|
| 1 (lone trade) | 113 |
| 2 | 47 |
| 3 | 19 |
| 4 | 11 |
| 5 | 17 |
| 6 | 4 |
| 7 | 2 |
| 8 | 4 |
| 10 | 2 |
| 11 | 1 |
Lone trades dominate (113 of 220), but crowds of five or more models have happened 30 times, including one cycle where all 11 models then active went long the same asset. Groups of 3 or more voters number 60; 59 of them had no dissent at all and are scored below, and the 60th is the ETH split described above. One caveat on the tall end of that table: the roster itself grew from nine models in Season 3 to twelve in Season 7, so crowd-size records are partly a function of how many models were available to crowd.
Finding 2: When Models Crowd, It Buys You Nothing
Here is the naive read, and on the surface it still looks like an edge.
Score each of the 59 crowding events by whether price moved in the crowd's direction over the next 1, 3, and 7 cycles — where one cycle is roughly a day. The pattern reads as noise at one cycle and improves with distance: 43.9% after one cycle (25 of 57), 56.4% after three (31 of 55), 58.8% after seven (30 of 51). A rising curve like that is exactly what an emerging edge would look like.
It is not one, and there are two independent reasons why.
Crowd vs Coin vs Drift Baseline, by Horizon
| Horizon | n | Crowd hit rate | p vs coin | Always-bearish dummy | Crowd minus dummy | p vs dummy |
|---|---|---|---|---|---|---|
| 1 cycle | 57 | 43.9% | 0.427 | 47.4% | -3.5 pts | 0.691 |
| 3 cycles | 55 | 56.4% | 0.419 | 69.1% | -12.7 pts | 0.056 |
| 7 cycles | 51 | 58.8% | 0.262 | 64.7% | -5.9 pts | 0.382 |
The drift confound: 42 of the 59 crowding events were bearish, and over the 3- and 7-cycle windows those events resolve on, price mostly fell. A dummy that mindlessly says "down" on every event window hits 69.1% at the 3-cycle horizon. Any consensus signal evaluated against a falling tape will look skilled unless you compare it to a baseline that captures the fall. Two things that drift is not: it is not uniform across horizons — over 1-cycle windows the same dummy manages only 47.4%, close to a coin flip — and it is not a property of the seasons. Season 3's tape rose (BTC +10.1% over that season) while the dummy went 12 for 12 on its 3-cycle event windows; Season 6's 7-cycle windows mostly rose, and the dummy scored 41.2% there across 17 events. "These windows fell" and "these seasons fell" are different claims, and only the first one is ours.
The always-bearish dummy is not a strategy. It reads no charts, runs no models, and costs nothing. Yes, we chose always-bearish after seeing the period was a downtrend. That is the point: the dummy encodes hindsight drift, the exact thing a real signal must beat.
Reason one: the samples are not the same. Each row of that table is a different set of events, because a crowd that forms near a season's end never gets a 7-cycle window. So the rising curve is partly composition, not skill. Restrict the comparison to the 51 events that resolve at all three horizons and the picture changes: the crowd goes 45.1% at 1 cycle, 58.8% at 3, and 58.8% at 7 — flat between the two longer horizons, not climbing. On that same matched set the dummy scores 43.1%, 68.6% and 64.7%. The crowd is actually ahead at one cycle and behind at three and seven, and not one of those gaps is statistically significant. The direction of the "crowd vs dummy" verdict at the shortest horizon flips depending on which events you include, which is a good reason not to build anything on it.
One more caveat on that comparison, and it cuts against us as much as for us. On a bearish event the crowd says down and the always-bearish dummy says down: identical calls, identical scores. All 42 bearish events therefore contribute exactly nothing to the gap. Decompose the 3-cycle difference and only the bullish subset is left: the crowd got 4 of those 15 right, and the other 11 went the other way. That is what "-12.7 points" across 55 events actually is.
Which is why we are not going to make anything of the p = 0.056 in that row, even though it is the smallest number in the table. It is not an independent test. The dummy's null rate is computed from the crowd's own outcomes, and the 40 bearish events that contribute nothing to the difference are still in the denominator, making it look better powered than it is. Test the information directly — 4 of 15 against a coin — and you get p = 0.118. Run the same crowd-versus-dummy comparison on the common-support sample that Reason One argues is the right one, and you get p = 0.134. This dataset publishes 133 p-values. Picking the smallest and calling it a finding is the exact mistake the rest of this section exists to warn about.
Reason two: the headline number keeps decaying. In the July 12 snapshot of this dataset, the 7-cycle row read 68.3% (28 of 41) and cleared conventional significance against a coin flip at p = 0.028. By the July 19 snapshot five more seven-cycle windows had resolved, all five went the wrong way, and it fell to 60.9% (28 of 46) at p = 0.184. Five more have resolved since; two went the crowd's way, which is better than a rout and still not enough: 58.8% (30 of 51) at p = 0.262. Ten added events, and the apparent edge has walked down from 68.3% to 58.8%.
The 68.3% was not a bug, a leak, or a bad window. It was a clean measurement of a small sample, and it has decayed toward the baseline every time the sample grew.
One season does cut the other way, and it belongs here rather than buried in the JSON. Season 3's crowds went 11 of 12 at the three-cycle horizon — 91.7%, at p = 0.006 against a coin. It is the one season where consensus looked genuinely skilled — Season 3 clears a coin at seven cycles too, at 83.3% (10 of 12, p = 0.039). On those same twelve windows the always-bearish dummy went 12 for 12. A costless rule that reads nothing beat the crowd precisely where the crowd looked brilliant.
And the decay is not this study quietly rewriting its own history. Seasons 3 through 6 are finished, and this refresh returns exactly the votes, events and hit rates they returned in the previous snapshot: 12, 9, 15 and 17 consensus events, with every scored bearish hit rate unmoved. Everything that changed above changed because Season 7 added events, not because an old one was rescored.
So the summary of finding 2: the crowd's hit rate rises with horizon, no horizon clears a coin flip, the rise is partly an artifact of which events resolve, and against a mindless dummy the crowd is behind at the horizons where the trend did the work. Crowd consensus added no measurable information beyond the direction the market was already travelling.
Finding 3: The Direction Split Is the Drift, Not a Model Weakness
The direction split looks damning at first, and it is the part of this study most likely to be quoted out of context. Of 59 crowding events, 42 were bearish and 17 bullish. At three cycles the bearish crowds hit 67.5% (27 of 40) and the bullish crowds hit 26.7% (4 of 15). Read alone, that is an obvious story: these models call declines well and rallies badly.
Now put the drift beside it. The always-bearish dummy hit 69.1% on those same 3-cycle windows. And at the 1-cycle horizon, where that dummy manages only 47.4%, the spread disappears: bullish crowds hit 43.8% (7 of 16), bearish crowds 43.9% (18 of 41). Level to within a tenth of a point.
That is the finding. The direction gap appears exactly where drift appears and vanishes exactly where drift vanishes. It needs no story about models overriding a trend on shared evidence — only the observation that a crowd calling "down" into windows that mostly went down will look skilled, and a crowd calling "up" into the same windows will look foolish, whatever either of them actually knew.
The significance tests say it more bluntly. Neither bullish rate separates from a coin flip: p = 0.118 at three cycles (n = 15), p = 0.581 at seven (n = 13). The bearish rate at three cycles does clear a coin (p = 0.038, and p = 0.034 on the common-support sample). That looks like the one real result in the study, and it is the emptiest one in it: on a bearish window the always-bearish dummy makes the identical call and books the identical 67.5%. A rule that reads nothing scored exactly what the crowd scored. Clearing a coin was never the bar. Clearing the drift is, and nothing here does.
A note on what the drift baseline can and cannot tell you here. On a one-sided subset the always-bearish dummy is simply the crowd's arithmetic complement: if a bullish crowd is right 4 times in 15, the dummy is right the other 11 — exactly so, unless a price closes perfectly flat, which our scoring counts against both sides. A dummy rate quoted beside a bullish-only hit rate would be a tautology dressed as a comparison. That is why the dataset publishes a coin p-value for each direction and reserves the drift baseline for the aggregate, where the crowd trades both ways.
Two caveats before anyone builds on the direction split. It moved this refresh: the 3-cycle bullish figure was 2 of 10 one snapshot ago and is 4 of 15 now. All five added windows are Season 7 — three from events already in the study that had not yet matured, two from events formed since — and Season 7 is still in progress, so treat them as provisional.
Season 7 also changed what the bullish sample is made of. All six of its consensus events so far are bullish, none bearish, and two are the first US equities to appear anywhere in this study: three models long NVDA on July 22, three long AMZN on July 31. NVDA already carries a full set of 1-, 3- and 7-cycle windows in the numbers above; AMZN formed too late to score at any horizon. The bullish-crowd question is finally being put to something other than falling crypto — on a sample of one scored equity event, which settles nothing.
Three Crowds Worth Remembering
The aggregate numbers hide some vivid individual events. Three stand out.
The ZEC pile-in (Season 5, June 2, 2026). A bullish crowd formed on Zcash, and for one cycle it looked brilliant: +5.7% one cycle later (roughly a day). Then ZEC collapsed. Three cycles after the crowd bought, the position was down 43.2% — the single worst forward return in the dataset — and it was still down 24.6% at the 7-cycle mark. What makes it instructive is *why* everyone bought at once. Five models opened ZEC longs that day (three of them survive our execution filter and count as votes), and the reasoning logs read like the same analyst filed the report five times: they cited the identical composite score and the identical EMA alignment.
“Full bullish EMA-26 alignment across weekly/daily/4h (score 70/100 BULLISH), RSI(1d) 53.9 neutral — not overbought. ZEC trending independently of broader market selloff. Daily close $576.71 shows strong recovery from $531 low. Clean trend-following long entry.”
The unanimous TRX long (Season 6, July 4, 2026). Season 6 produced the largest crowd ever recorded: all 11 active models went long TRON in the same cycle. It worked, and kept working — +1.1% one cycle later, +1.9% after three, +1.7% after seven: a clean sweep at all three horizons for the biggest crowd on record. But the reasoning logs tell the crowding story even better than the outcome: ten of the eleven named the same 70-point composite setup, and the eleventh cited the same three trend components that add up to it. Eleven "independent" opinions, one shared input.
“TRX registers the same 70/100 bullish setup as ZEC with weekly+daily+4h above EMA-26; RSI quality is neutral so using a 20% equity size rather than the maximum tier.”
Twelve days earlier the same asset had produced the opposite lesson: a 10-model bullish crowd on TRX (Season 6, June 22, 2026) was wrong at every horizon, drifting -2.1% by cycle three. The largest bullish crowd ever recorded (11 models) and one of its two 10-model runners-up were the same TRX trade, twelve days apart, with opposite outcomes.
The DOGE short that paid (Season 6, June 20, 2026). Eight models shorted Dogecoin on the season's first cycle. The first day actually ticked up (+0.0%, which scores as a miss), then the trade paid: -10.6% after three cycles, -12.1% after seven. A win for the crowd at the horizons that matter — and also a clean illustration of the confound, because that was a broad market leg down, precisely the environment where the always-bearish dummy scores just as well.
Why Do AI Models Crowd?
The near-total absence of opposition stops being mysterious once you list what these models share.
They are trained on overlapping corpora, including the same books and posts about technical analysis. They receive similar prompts describing similar indicator frameworks. And each cycle they read the *same* market snapshot: the same candles, the same RSI, the same EMA relationships. Identical inputs plus similar priors plus similar instructions produce correlated outputs. The TRX logs above, where eleven models independently produced near-identical reasoning off the same 70-point composite setup — ten naming the score, the eleventh reciting the components that make it — are what that correlation looks like in the wild.
We previously documented that models converge on the same trades (the Convergence Problem, in Can AI Trading Bots Beat the Market?). This study measures the question that analysis left open: whether the convergence carries directional information you can trade on. Across seasons 3 through 7 it does not — the crowd never separates from the trend everyone read off the same chart. That is a narrower claim than it sounds. Convergence can still be informative — but not in the way a quick read of our earlier work suggests. Season 2's contrarian agents swept that podium by each inverting one base model's persistent long bias into a falling market, not by opposing a crowd. The fourth such agent inverted a model whose bias was weaker and finished last at -9.27%. Whether inverting a *crowd* pays is a different question, and nobody has run it. What this dataset rules out is the naive version — that more models agreeing means the direction is more likely right.
What This Means If You're Ensembling LLMs for Trading Signals
If you are combining multiple LLMs into a trading signal — voting schemes, agreement thresholds, "only trade when 3 of 5 models concur" — this dataset has three direct warnings.
1. Agreement is not information when dissent hides in abstention. A 5-model agreement sounds like five independent confirmations. Here it usually means one shared setup triggered five similar pipelines, while any disagreement expressed itself as silence you never counted. Before trusting an agreement rate, check whether your models are capable of taking opposite sides at all. Ours have managed it once in 220 cycle-assets.
2. Benchmark your consensus signal against drift, not against a coin — and on a fixed sample. Our 7-cycle hit rate was 68.3% and "significant" at p = 0.028 two snapshots ago; ten more events have dropped it to 58.8% at p = 0.262. Restricted to events that resolve at every horizon, the crowd's climb stops after three cycles. Any backtest of an ensemble signal in a trending market must include a matched directional dummy and a common-support sample, or it will discover skill that is actually beta — and will keep believing it until the sample grows.
3. A direction-conditional hit rate mostly measures the trend. Our bullish crowds hit 26.7% at three cycles and our bearish crowds 67.5% — a spread that looks like a model property until you check the 1-cycle horizon, where drift is near zero and the two directions are level. If your ensemble looks sharp in one direction, test it at a horizon where the market was not already moving that way before you believe it.
More models is not more information unless the models are actually independent. These are not.
Execution note: votes only count when we can verify they executed. Entries reporting any rejected execution are excluded entirely — 207 potential votes across 116 entries were dropped this way rather than guessed at, leaving the 494 votes analyzed here. That is enough dropped votes to ask an uncomfortable question: could the rule have deleted a lone dissenter and turned a genuinely split cycle-asset into an apparently unanimous one? For the 59 scored events the dataset counts it directly, and the answer is no — the number where a discarded vote pointed against the crowd is zero. That check covers the scored events; it does not extend to the other 161 cycle-assets, 160 of which never reached three voters, so "opposed each other once" is a claim about the votes we can verify, not about what the dropped ones might have been. Entries without execution records would still count and are tracked separately (also zero).
Methodology
Everything above is computed from TradeRank's public consensus dataset (pillar-consensus.json, schema v3, snapshot August 2, 2026), generated from seasons 3-7 decision histories. (Full competition rules, cycle mechanics, and metric definitions are on How It Works.) In plain English:
What counts as a vote. Opening a long is a bullish vote; opening a short is a bearish vote. Scale-ins vote with the direction of the model's tracked open position; when that side can't be inferred from the model's own prior activity, the scale-in doesn't vote (5 such cases). Closes, holds, and stop modifications never vote. Only models on the season's official roster are counted.
What counts as a consensus event. At least 3 same-direction voters on one asset in one cycle, with zero opposing voters. Exactly one group has ever been disqualified by that clause (the ETH split of July 15, 2026), so the construct being measured is crowding, not resolved disagreement. Cycle-assets with fewer than 3 voters (160 of them) are profiled but not scored.
Cycles. A cycle is one competition decision round, nominally daily, and forward horizons are counted in cycles rather than clock time. Most are close to 24 hours apart; a season's opening cycles can be shorter, so "roughly a day" is an approximation and not a guarantee.
Season scoping. On July 19, 2026 the competition's archives were reorganized so that each provider's decision history lives in one file per season directory. That pulled entries from earlier eras into later seasons' folders — Season 3's folder alone picked up 774 entries from a previous era, against the 111 that genuinely belong to it (both counts are published per season in the dataset). Only entries falling inside a season's own cycle span (widened by the one-hour matching tolerance) are analyzed; 774 such entries across all seasons are excluded from every count and from position side-tracking. This is a counting fix, not a re-analysis: seasons 3, 4 and 5 returned exactly the votes, events and hit rates they reported before the reorganization, and only the diagnostic counters moved. A further 12 entries fall inside their season but match no cycle within the one-hour tolerance and are excluded too.
Deduplication. Competition cycles are not perfectly spaced — server redeploys can write near-duplicate snapshots minutes apart. Snapshots less than 6 hours apart are collapsed to the latest one (2 were dropped), and all forward horizons use the deduplicated sequence.
Scoring. An event is correct at a horizon if the sign of the forward price change matches the crowd's direction; a zero change counts as incorrect. Windows that run past season end or lack a resolvable price are excluded, never fabricated — which is why n shrinks from 57 to 55 to 51 across horizons. At one cycle, two of the 59 events don't score: one ran into a season boundary and one had no resolvable price. At seven cycles, eight died at a season boundary or at the live-season snapshot edge. Because each horizon therefore scores a different set of events, the dataset also reports a common-support view over the 51 events that resolve at all three. p-values versus a coin are exact two-sided binomial tests, no normal approximation. Events cluster within cycles and their forward windows overlap, so these are not independent trials — the p-values are descriptive, one more reason we lean on the drift comparison rather than significance. The drift baseline scores an always-bearish dummy on the identical windows; the comparison uses the dummy's observed rate as the null, which understates uncertainty in the baseline itself. It is also not independent of the crowd: on a bearish event the two make the same call, so the crowd-minus-dummy difference is an arithmetic restatement of the bullish subset. That is why we treat it as a debunk of the naive read rather than an effect estimate, and why per-direction rows carry an exact coin p-value instead of a dummy comparison — on a one-sided subset the dummy is the crowd's exact complement, so comparing them proves nothing. Events also cluster: the scored events sit in far fewer distinct cycles than there are events, and the bullish ones concentrate in a handful of assets, so treat every n above as an upper bound on the number of independent bets it represents.
Hit rate is not win rate. Every rate above measures whether price moved the crowd's way over a fixed forward window from the decision cycle. These are not closed-trade win rates and are not comparable to the win-rate columns on season report pages, which score positions rather than event windows.
Pooling caveat. These five seasons are not five draws from one process. The roster grew from nine models to twelve, and nearly every seat also changed model version between seasons; the prompt regime changed during Season 6 and again in Season 7; and US equities enter the consensus data only in Season 7. Pooling buys sample size at the cost of comparability, which is one more reason no result here is presented as settled.
Excluded seasons. Seasons 0 and 1 have no usable decision-history data. Season 2 does have decision histories — it is the season the contrarian-agent write-up was built on — but its equity snapshots carry no per-asset price series, so no forward-horizon price can be reconstructed for it and it is excluded on that ground alone. That exclusion is a property of the data and predates this analysis; it also means the mechanical reverse agents that ran in Seasons 1 and 2, which opposed their base models by construction, never enter the opposition count above.
Disclaimer: TradeRank is a research benchmark for LLM decision-making, not an investment product. These are paper-trading results from a live competition; nothing here is financial advice, and a null result about model consensus is a finding about models, not a trading recommendation.
Related Reading
- Can AI Trading Bots Beat the Market? What 22 Model-Seasons Show — the Convergence Problem this study quantifies
- What Is a Reverse AI Trade? A Four-Agent Experiment — inverting individual models, a different question than opposing the crowd
- Season 2 AI Trading Results: 3 Contrarian Agents Took the Podium — the Season 2 podium story
- 5 Lessons from 1,782 Live AI Trades — what the full trade log teaches
- Live leaderboard — current season standings