Ensembling is the most natural idea in machine learning: if one model is noisy, ask several and trust the overlap. Applied to trading, the pitch is obvious: when GPT, Claude, Gemini, Grok, and the rest of a roster that has grown from nine models to fourteen across these five seasons all short the same coin on the same day, surely that agreement means something.
We can now test that pitch on live data. TradeRank runs a paper-trading competition where LLM agents make real daily decisions on the same market data under identical rules. Across seasons 3 through 6 and Season 8 (March 23 to September 6, 2026 — Season 8 is still in progress), the models cast 527 verified directional votes across 257 cycle-assets (one asset in one daily cycle). That is enough to answer two questions: do AI models actually agree, and is the agreement worth anything?
Season 7 was in that list until this refresh. The host holding its per-decision records and price snapshots failed on August 14, 2026, the day after the season's final cycle, and neither has been recovered — the season report sets out what survived. The six consensus events this study published from Season 7 in its August snapshot are therefore withdrawn, not rescored, and the four new events below all come from Season 8. Every number in this article moved for that reason or because Season 8 added data; no finished season was recomputed.
The answers are stranger than the pitch. The models agree far more than they should — in five seasons they have taken opposite sides exactly once. And the agreement never separates itself from a matched always-bearish baseline once you account for what the market was already doing.
This is a living study. Every statistic below is read from a generated dataset (published here, snapshot September 6, 2026) rather than computed in the article; where the text quotes a model's reasoning, counts something in the reasoning logs, or cites an earlier snapshot of this same dataset, it says so. Current standings are on the live leaderboard.
Finding 1: In 257 Cycle-Assets, AI Models Opposed Each Other Once
For four seasons this section reported an absolute: zero opposition, ever. That is no longer true, and the exception is worth more than the streak was.
On July 15, 2026 — cycle 26 of Season 6 — three models traded ETH. Two opened shorts. MiniMax M3 opened a long. It is the first and so far only time in 257 cycle-assets and 527 directional votes that any model has taken the other side of another model's trade in the same cycle. ETH then fell 2.6% over the next cycle and 4.5% over three, so the majority was right and the lone dissenter was wrong at both horizons that resolved before the season ended.
One event is not a finding about contrarians. With a single opposed cycle-asset in the entire dataset, the question everyone asks — does the lone dissenter beat the crowd? — still cannot be answered from this data. What the event does establish is the rate: opposition is possible, nothing in the rules prevents it, and it has happened once in five seasons. Competition rules cap new positions per cycle and gate low-confidence opens, but validation is entirely per-model, and nothing stops one model shorting what another buys.
Human fund managers routinely hold opposite views of the same asset; these models essentially never do. That near-absence is also not what "consensus" usually means, because consensus implies a disagreement that got resolved. What these models produce is better described as crowding: when one model likes a setup, the only question is how many others pile in behind it.
Where does disagreement go? Into abstention. A model that doesn't share the crowd's view almost always just holds. It stays silent rather than taking the opposite side. So a headline like "8 models agree DOGE is going down" quietly hides the real distribution: 8 models shorted, and the rest stayed silent (which no dashboard will ever show you). Abstention is ambiguous — a silent model may dissent, may already hold the position, or may be at its position cap — which is why agreement counts overstate independent confirmation.
A related number stayed at zero: no model has ever contradicted itself inside one cycle-asset, opening long and short on the same asset in the same cycle. Across 527 votes, conflicting self-votes remain 0.
How Big Do the Crowds Get?
| Models voting on the asset | Cycle-assets (all-time) |
|---|---|
| 1 (lone trade) | 150 |
| 2 | 49 |
| 3 | 18 |
| 4 | 10 |
| 5 | 17 |
| 6 | 4 |
| 7 | 3 |
| 8 | 3 |
| 10 | 2 |
| 11 | 1 |
Lone trades dominate (150 of 257), but crowds of five or more models have happened 30 times, including one cycle where all 11 models then active went long the same asset. Groups of 3 or more voters number 58; 57 of them had no dissent at all and are scored below, and the 58th is the ETH split described above. One caveat on the tall end of that table: the roster itself grew from nine models in Season 3 to fourteen in Season 8, so crowd-size records are partly a function of how many models were available to crowd.
Finding 2: When Models Crowd, It Buys You Nothing
Here is the naive read, and on the surface it still looks like an edge.
Score each of the 57 crowding events by whether price moved in the crowd's direction over the next 1, 3, and 7 cycles — where one cycle is roughly a day. The pattern starts below a coin flip and improves with distance: 41.1% after one cycle (23 of 56), 56.6% after three (30 of 53), 61.2% after seven (30 of 49). A rising curve like that is exactly what an emerging edge would look like.
It is not one, and there are two independent reasons why.
Crowd vs Coin vs Drift Baseline, by Horizon
| Horizon | n | Crowd hit rate | p vs coin | Always-bearish dummy | Crowd minus dummy | p vs dummy |
|---|---|---|---|---|---|---|
| 1 cycle | 56 | 41.1% | 0.229 | 50.0% | -8.9 pts | 0.229 |
| 3 cycles | 53 | 56.6% | 0.410 | 69.8% | -13.2 pts | 0.050 |
| 7 cycles | 49 | 61.2% | 0.152 | 63.3% | -2.0 pts | 0.769 |
The drift confound: 42 of the 57 crowding events were bearish, and over the 3- and 7-cycle windows those events resolve on, price mostly fell. A dummy that mindlessly says "down" on every event window hits 69.8% at the 3-cycle horizon. Any consensus signal evaluated against a falling tape will look skilled unless you compare it to a baseline that captures the fall. Two things that drift is not: it is not uniform across horizons — over 1-cycle windows the same dummy lands on exactly 50.0%, a coin — and it is not a property of the seasons. Season 3's tape rose (BTC +10.1% over that season) while the dummy went 12 for 12 on its 3-cycle event windows; Season 6's 7-cycle windows mostly rose, and the dummy scored 41.2% there across 17 events. "These windows fell" and "these seasons fell" are different claims, and only the first one is ours.
The always-bearish dummy is not a strategy. It reads no charts, runs no models, and costs nothing. Yes, we chose always-bearish after seeing the period was a downtrend. That is the point: the dummy encodes hindsight drift, the exact thing a real signal must beat.
Reason one: the dummy is ahead at every horizon. Each row of that table is a different set of events, because a crowd that forms near a season's end never gets a 7-cycle window, so part of the rise could be composition rather than skill. Restrict the comparison to the 49 events that resolve at all three horizons and the crowd still climbs — 42.9% at 1 cycle, 59.2% at 3, 61.2% at 7. On that same matched set the dummy scores 44.9%, 69.4% and 63.3%, so it leads at all three, and its own best horizon is three cycles rather than seven. None of those gaps is statistically significant (p = 0.886, 0.123 and 0.769).
Read the one-cycle row of that matched comparison with care. The dummy's lead there is 2.0 points, and it pointed the other way one snapshot ago: in August the crowd led at one cycle, 45.1% against 43.1%. The verdict at the shortest horizon turns on which events happen to be in the study, which is a good reason not to build anything on it in either direction.
One more caveat on that comparison, and it cuts against us as much as for us. On a bearish event the crowd says down and the always-bearish dummy says down: identical calls, identical scores. Every bearish event therefore contributes exactly nothing to the gap, and at three cycles that is 40 of the 53 scored events. Decompose the 3-cycle difference and only the bullish subset is left: the crowd got 3 of those 13 right, and the other 10 went the other way. That is what "-13.2 points" across 53 events actually is.
Which is why we are not going to make anything of the p = 0.050 in that row, even though it is the smallest number in the table and has now walked onto the conventional threshold. It is not an independent test. The dummy's null rate is computed from the crowd's own outcomes, and the 40 bearish events that contribute nothing to the difference are still in the denominator, making it look better powered than it is. Test the information directly — 3 of 13 against a coin — and you get p = 0.092. Run the same crowd-versus-dummy comparison on the common-support sample that Reason One argues is the right one, and you get p = 0.123. This dataset publishes 133 p-values. Picking the smallest one in view and calling it a finding is the exact mistake the rest of this section exists to warn about. Taken at face value it would say the crowd did worse than the dummy, which is not an edge either.
Reason two: the headline number keeps moving, and so does the sample under it. In the July 12 snapshot of this dataset, the 7-cycle row read 68.3% (28 of 41) and cleared conventional significance against a coin flip at p = 0.028. By July 19 five more seven-cycle windows had resolved, all five went the wrong way, and it fell to 60.9% (28 of 46) at p = 0.184. On August 2 it read 58.8% (30 of 51) at p = 0.262. It now reads 61.2% (30 of 49) at p = 0.152, and that last move is not the sample growing. Six Season 7 events left and four Season 8 events arrived. Season 7's crowds had gone 2 for 5 at seven cycles, below the pooled rate, so the withdrawal lifted the number before Season 8 added anything: it went up because a below-average season left the study, not because the crowd got better. Four snapshots, a spread from 58.8% to 68.3%, and only the first of them ever cleared a coin.
The 68.3% was not a bug, a leak, or a bad window. It was a clean measurement of a small sample, and no snapshot since has reproduced it. The exact 95% interval on today's number says the same thing without waiting for the next snapshot: 61.2% carries an interval of 46.2% to 74.8%. That contains a coin flip, contains the dummy's 63.3%, and contains the 68.3% we published in July. Forty-nine events do not tell those apart.
One season does cut the other way, and it belongs here rather than buried in the JSON. Season 3's crowds went 11 of 12 at the three-cycle horizon — 91.7%, at p = 0.006 against a coin. It is the one season where consensus looked genuinely skilled — Season 3 clears a coin at seven cycles too, at 83.3% (10 of 12, p = 0.039). That p = 0.006 is one of the same 133, and the multiplicity argument above applies to it exactly as much. It does not have to be argued away that way, though: on the twelve three-cycle windows the always-bearish dummy went 12 for 12, and 11 of 12 at seven. A costless rule that reads nothing beat the crowd at both horizons, precisely where the crowd looked brilliant.
And none of this movement is the study quietly rewriting its own history. Seasons 3 through 6 are finished, and this refresh returns exactly the votes, events and hit rates they returned in the previous snapshot: 12, 9, 15 and 17 consensus events, with every scored rate unmoved, crowd and dummy alike. What changed above changed because one season left the dataset and another added events, not because a finished season was rescored. One of those changes helped this study's own argument, and it should be said plainly: in August the crowd led the dummy at one cycle on the matched sample, 45.1% against 43.1%, and after the swap it leads nowhere.
So the summary of finding 2. The crowd's hit rate rises with horizon and no horizon clears a coin flip. The always-bearish dummy is ahead at all three, whether you score every event or only the 49 that resolve at all of them. Crowd consensus added no measurable information beyond the direction the market was already travelling.
Finding 3: The Direction Split Tracks the Drift
The direction split looks damning at first, and it is the part of this study most likely to be quoted out of context. Of 57 crowding events, 42 were bearish and 15 bullish. At three cycles the bearish crowds hit 67.5% (27 of 40) and the bullish crowds hit 23.1% (3 of 13). Read alone, that is an obvious story: these models call declines well and rallies badly.
Now put the drift beside it, horizon by horizon. Three-cycle windows are where the always-bearish dummy does best, at 69.8%, and they are where the direction split is widest: 23.1% bullish against 67.5% bearish. One-cycle windows are where the dummy is a coin, and there the split narrows to 33.3% (5 of 15) against 43.9% (18 of 41), with both directions below the coin. Seven cycles sits between the two on both counts: dummy 63.3%, bullish 45.5% (5 of 11) against bearish 65.8% (25 of 38).
That is a consistency rather than a finding, and it is a weaker claim than the one this study made in August, when the two directions were level at one cycle and the gap looked like it disappeared outright. Three horizons ordering the same way on two quantities is not much evidence of anything: they are overlapping windows on the same events, not three independent looks. What the pattern is consistent with is drift — a crowd calling "down" into windows that mostly went down will look skilled, and a crowd calling "up" into the same windows will look foolish, whatever either of them actually knew. What it does not do is rule out a model property.
The split is not even stable in shape across samples. On the 49 events that resolve at all three horizons — the sample Reason One argues is the right one — the one-cycle split does not narrow, it inverts: bullish 45.5% (5 of 11) against bearish 42.1% (16 of 38), with neither rate separating from a coin (p = 1 and p = 0.418). Either way it is nowhere near the spread at three cycles.
The significance tests say the rest. No bullish rate separates from a coin flip at any horizon: p = 0.302 at one cycle (n = 15), p = 0.092 at three (n = 13), p = 1 at seven (n = 11). The bearish rate at three cycles does clear a coin (p = 0.038, and p = 0.034 on the common-support sample). That looks like the one real result in the study, and it is the emptiest one in it: on a bearish window the always-bearish dummy makes the identical call and books the identical 67.5%. A rule that reads nothing scored exactly what the crowd scored. Clearing a coin was never the bar. Clearing the drift is, and nothing here does.
A note on what the drift baseline can and cannot tell you here. On a one-sided subset the always-bearish dummy is simply the crowd's arithmetic complement: if a bullish crowd is right 3 times in 13, the dummy is right the other 10 — exactly so, unless a price closes perfectly flat, which our scoring counts against both sides. A dummy rate quoted beside a bullish-only hit rate would be a tautology dressed as a comparison. That is why the dataset publishes a coin p-value for each direction and reserves the drift baseline for the aggregate, where the crowd trades both ways.
Two caveats before anyone builds on the direction split. It moved this refresh, and not by maturing: the 3-cycle bullish figure was 4 of 15 one snapshot ago and is 3 of 13 now. All six withdrawn Season 7 events were bullish and all four Season 8 arrivals are bullish, so the refresh replaced part of the only subset the crowd-versus-dummy comparison rests on — the bearish events contribute nothing to it. Every bullish event in the study that was not there in August belongs to a season still in progress.
Season 8 also changed what the bullish sample is made of, and put it on a rising tape. All four of its consensus events so far are bullish, none bearish — and at one cycle all four were wrong, while the always-bearish dummy went 4 for 4 on the same windows. By seven cycles two of the three that had matured were right, ETH by 27.9% and BNB by 12.6%, and the dummy was down to 1 of 3. Three scored events settle nothing in either direction. Two of the four are US equities — three models long MU on August 17, three long NVDA on September 4 — and with Season 7's equity pair withdrawn they are the only stocks anywhere in this study. MU is the only stock scored at all three horizons, and it was wrong at all three, down 7.9%, 6.7% and 11.0%. NVDA formed on cycle 20, so its 3- and 7-cycle windows run past the snapshot edge and it scores at one cycle only.
Three Crowds Worth Remembering
The aggregate numbers hide some vivid individual events. Three stand out.
The ZEC pile-in (Season 5, June 2, 2026). A bullish crowd formed on Zcash, and for one cycle it looked brilliant: +5.7% one cycle later (roughly a day). Then ZEC collapsed. Three cycles after the crowd bought, the position was down 43.2% — the single worst forward return in the dataset — and it was still down 24.6% at the 7-cycle mark. What makes it instructive is *why* everyone bought at once. The decision logs for that cycle show five models opening ZEC longs (three of them survive our execution filter and count as votes), and they read like the same analyst filed the report five times: they cited the identical composite score and the identical EMA alignment.
“Full bullish EMA-26 alignment across weekly/daily/4h (score 70/100 BULLISH), RSI(1d) 53.9 neutral — not overbought. ZEC trending independently of broader market selloff. Daily close $576.71 shows strong recovery from $531 low. Clean trend-following long entry.”
The unanimous TRX long (Season 6, July 4, 2026). Season 6 produced the largest crowd ever recorded: all 11 active models went long TRON in the same cycle. It worked, and kept working — +1.1% one cycle later, +1.9% after three, +1.7% after seven: a clean sweep at all three horizons for the biggest crowd on record. But the reasoning logs tell the crowding story even better than the outcome: reading all eleven, ten name the same 70-point composite setup and the eleventh cites the same three trend components that add up to it. Eleven "independent" opinions, one shared input.
“TRX registers the same 70/100 bullish setup as ZEC with weekly+daily+4h above EMA-26; RSI quality is neutral so using a 20% equity size rather than the maximum tier.”
Twelve days earlier the same asset had produced the opposite lesson: a 10-model bullish crowd on TRX (Season 6, June 22, 2026) was wrong at every horizon, drifting -2.1% by cycle three. The largest bullish crowd ever recorded (11 models) and one of its two 10-model runners-up were the same TRX trade, twelve days apart, with opposite outcomes.
The DOGE short that paid (Season 6, June 20, 2026). Eight models shorted Dogecoin on the season's first cycle. The first day actually ticked up (+0.0%, which scores as a miss), then the trade paid: -10.6% after three cycles, -12.1% after seven. A win for the crowd at the horizons that matter — and also a clean illustration of the confound, because that was a broad market leg down, precisely the environment where the always-bearish dummy scores just as well.
Why Do AI Models Crowd?
The near-total absence of opposition stops being mysterious once you list what these models share.
They are trained on overlapping corpora, including the same books and posts about technical analysis. They receive similar prompts describing similar indicator frameworks. And each cycle they read the *same* market snapshot: the same candles, the same RSI, the same EMA relationships. Identical inputs plus similar priors plus similar instructions produce correlated outputs. The TRX logs above, where eleven models independently produced near-identical reasoning off the same 70-point composite setup — ten naming the score, the eleventh reciting the components that make it — are what that correlation looks like in the wild.
In Season 2 every standard agent converged on the same bearish thesis within the first few cycles (5 Lessons from 1,782 Live AI Trades). This study measures the question that analysis left open: whether the convergence carries directional information you can trade on. Across these five seasons it does not — the crowd never separates from the trend everyone read off the same chart. That is a narrower claim than it sounds. Convergence can still be informative, though not in the way a quick read of our earlier work suggests. Season 2's contrarian agents swept that podium by each inverting one base model's persistent long bias into a falling market, not by opposing a crowd. The fourth such agent inverted a model whose bias was weaker and finished last at -9.27%. Whether inverting a *crowd* pays is a different question, and nobody has run it. What this dataset rules out is the naive version — that more models agreeing means the direction is more likely right.
What This Means If You're Ensembling LLMs for Trading Signals
If you are combining multiple LLMs into a trading signal — voting schemes, agreement thresholds, "only trade when 3 of 5 models concur" — this dataset has three direct warnings.
1. Agreement is not information when dissent hides in abstention. A 5-model agreement sounds like five independent confirmations. Here it usually means one shared setup triggered five similar pipelines, while any disagreement expressed itself as silence you never counted. Before trusting an agreement rate, check whether your models are capable of taking opposite sides at all. Ours have managed it once in 257 cycle-assets.
2. Benchmark your consensus signal against drift, not against a coin — and on a fixed sample. Our 7-cycle hit rate was 68.3% and "significant" at p = 0.028 in July; three snapshots later it reads 61.2% (30 of 49) at p = 0.152, with an exact 95% interval of 46.2% to 74.8% that contains both the coin and the dummy. Restricted to the events that resolve at every horizon, a mindless always-bearish dummy leads the crowd at all three. Any backtest of an ensemble signal in a trending market must include a matched directional dummy and a common-support sample, or it will discover skill that is actually beta — and will keep believing it until the sample turns over.
3. A direction-conditional hit rate mostly measures the trend. Our bullish crowds hit 23.1% at three cycles (3 of 13) and our bearish crowds 67.5% (27 of 40) — the horizon where the always-bearish dummy scores 69.8%. That spread looks like a model property until you check the 1-cycle horizon, where the dummy is at exactly 50.0% and the two directions narrow to 33.3% (5 of 15) against 43.9% (18 of 41). If your ensemble looks sharp in one direction, test it at a horizon where the market was not already moving that way before you believe it.
More models is not more information unless the models are actually independent. These are not.
Execution note: votes only count when we can verify they executed. Entries reporting any rejected execution are excluded entirely — 208 potential votes across 117 entries were dropped this way rather than guessed at, leaving the 527 votes analyzed here. That is enough dropped votes to ask an uncomfortable question: could the rule have deleted a lone dissenter and turned a genuinely split cycle-asset into an apparently unanimous one? For the 57 scored events the dataset counts it directly, and the answer is no — the number where a discarded vote pointed against the crowd is zero. That check covers the scored events; it does not extend to the other 200 cycle-assets, 199 of which never reached three voters, so "opposed each other once" is a claim about the votes we can verify, not about what the dropped ones might have been. Entries without execution records would still count and are tracked separately (also zero).
Methodology
Everything above is computed from TradeRank's public consensus dataset (pillar-consensus.json, schema v4, snapshot September 6, 2026), generated from seasons 3-6 and 8 decision histories. (Full competition rules, cycle mechanics, and metric definitions are on How It Works.) In plain English:
What counts as a vote. Opening a long is a bullish vote; opening a short is a bearish vote. Scale-ins vote with the direction of the model's tracked open position; when that side can't be inferred from the model's own prior activity, the scale-in doesn't vote (6 such cases). Closes, holds, and stop modifications never vote. Only models on the season's official roster are counted, and the dataset publishes each season's roster size beside its results.
What counts as a consensus event. At least 3 same-direction voters on one asset in one cycle, with zero opposing voters. Exactly one group has ever been disqualified by that clause (the ETH split of July 15, 2026), so the construct being measured is crowding, not resolved disagreement. Cycle-assets with fewer than 3 voters (199 of them) are profiled but not scored.
Cycles. A cycle is one competition decision round, nominally daily, and forward horizons are counted in cycles rather than clock time. Most are close to 24 hours apart; a season's opening cycles can be shorter, so "roughly a day" is an approximation and not a guarantee.
Season scoping. On July 19, 2026 the competition's archives were reorganized so that each provider's decision history lives in one file per season directory. That pulled entries from earlier eras into later seasons' folders — Season 3's folder alone picked up 774 entries from a previous era, against the 111 that genuinely belong to it (both counts are published per season in the dataset). Only entries falling inside a season's own cycle span (widened by the one-hour matching tolerance) are analyzed; 774 such entries across all seasons are excluded from every count and from position side-tracking. This is a counting fix, not a re-analysis: seasons 3, 4 and 5 returned exactly the votes, events and hit rates they reported before the reorganization, and only the diagnostic counters moved. A further 11 entries fall inside their season but match no cycle within the one-hour tolerance and are excluded too.
Deduplication. Competition cycles are not perfectly spaced — server redeploys can write near-duplicate snapshots minutes apart. Snapshots less than 6 hours apart are collapsed to the latest one (2 were dropped), and all forward horizons use the deduplicated sequence.
Scoring. An event is correct at a horizon if the sign of the forward price change matches the crowd's direction; a zero change counts as incorrect. Windows that run past season end or lack a resolvable price are excluded, never fabricated — which is why n shrinks from 56 to 53 to 49 across horizons. At one cycle, one of the 57 events doesn't score, having run into a season boundary. At seven cycles, eight died at a season boundary or at the live-season snapshot edge. Because each horizon therefore scores a different set of events, the dataset also reports a common-support view over the 49 events that resolve at all three. p-values versus a coin are exact two-sided binomial tests, no normal approximation, and every horizon hit rate is published with an exact (Clopper-Pearson) 95% interval computed the same way. Events cluster within cycles and their forward windows overlap, so these are not independent trials — the p-values are descriptive, one more reason we lean on the drift comparison rather than significance. The drift baseline scores an always-bearish dummy on the identical windows; the comparison uses the dummy's observed rate as the null, which understates uncertainty in the baseline itself. It is also not independent of the crowd: on a bearish event the two make the same call, so the crowd-minus-dummy difference is an arithmetic restatement of the bullish subset. That is why we treat it as a debunk of the naive read rather than an effect estimate, and why per-direction rows carry an exact coin p-value instead of a dummy comparison — on a one-sided subset the dummy is the crowd's exact complement, so comparing them proves nothing. Events also cluster: the scored events sit in far fewer distinct cycles than there are events, and the bullish ones concentrate in a handful of assets, so treat every n above as an upper bound on the number of independent bets it represents.
Hit rate is not win rate. Every rate above measures whether price moved the crowd's way over a fixed forward window from the decision cycle. These are not closed-trade win rates and are not comparable to the win-rate columns on season report pages, which score positions rather than event windows.
Pooling caveat. These five seasons are not five draws from one process. The roster grew from nine models to fourteen, and nearly every seat also changed model version between seasons; the prompt regime changed during Season 6; US equities enter the consensus data only in Season 8; and the five seasons are not consecutive, because Season 7 is absent. Pooling buys sample size at the cost of comparability, which is one more reason no result here is presented as settled.
Excluded seasons. Seasons 0 and 1 have no usable decision-history data at all. Season 2 does have decision histories — it is the season the contrarian-agent write-up was built on — but its equity snapshots carry no per-asset price series, so no forward-horizon price can be reconstructed for it. That exclusion predates this analysis; it also means the mechanical reverse agents that ran in Seasons 1 and 2, which opposed their base models by construction, never enter the opposition count above. Season 7 was in this study until the August snapshot and is now out: its per-decision histories and its per-cycle price snapshots both lived on a host that failed on August 14, 2026, the day after the final cycle, and neither has been recovered. The repository does still hold one decision entry per agent from Season 7's opening cycle, and two of the six withdrawn events sit on that cycle; narrative daily reports cover July 18 to August 3. What none of that provides is the per-cycle price series every other event in this study is scored against, and scoring two events off some other price source would make them non-comparable with the rest. All six are withdrawn rather than half-restored. If the host is recovered the season becomes scoreable again and this study will say so.
Disclaimer: TradeRank is a research benchmark for LLM decision-making, not an investment product. These are paper-trading results from a live competition; nothing here is financial advice, and a null result about model consensus is a finding about models, not a trading recommendation.
Related Reading
- Can AI Beat the Market? Gemini Did, Across Nine Seasons — every model family against the S&P 500 and Bitcoin, Seasons 0 to 8
- What Is a Reverse AI Trade? A Four-Agent Experiment — inverting individual models, a different question than opposing the crowd
- Season 2 AI Trading Results: 3 Contrarian Agents Took the Podium — the Season 2 podium story
- 5 Lessons from 1,782 Live AI Trades — what the full trade log teaches
- Season 7 Final: Nobody Beat the S&P, and the Winner Couldn't Open a Position — the season this study lost, and what happened to its records
- Live leaderboard — current season standings