This is an observed three-season association, not a tested trading strategy. Review the competition rules and accounting on How It Works.
Are AI Traders Better in Bear Markets?
In this sample, frontier AI traders performed better during two falling markets than during one rising market. That is the direct answer. The honest qualifier is just as important: n=3, and the seasons were not controlled repetitions. The result supports a hypothesis about downside positioning; it does not prove that frontier models possess a general bear-market edge. Compare every model and season on the LLM trading benchmark.
The Claim, Stated Plainly
We have now run frontier AI models through three very different markets under identical conditions: a bull, a mild bear, and a hard bear. Same nine-to-ten models, same system prompt, same $10,000 accounts, same fees, same risk rules. The only thing that changed between seasons was what the market did.
And the results line up with the market in a way that should make you raise an eyebrow. When crypto went up, the models lost. When crypto went down, the models won. Not occasionally. Every time, across every season we can compare.
That is the opposite of what most people assume about AI traders. The intuition is that a smart enough model should make money in any market, or at least make more when the market is rising and there is easy beta to capture. The data says the reverse. Let us walk through it, then explain why, then tell you the two reasons it is not the money-printer it sounds like.
This article is for educational and entertainment purposes only. It is not financial advice. All results come from a simulated competition using live market prices and simulated capital. No real money was at risk. Past simulated performance does not predict future results. Methodology is at /how-it-works.
The Pattern, Across Three Frontier Seasons
Here are the three seasons in which the same class of premium frontier models traded under matching rules. Field average is the equal-weighted mean return across every model that season. Read the first and last columns together.
Market Regime vs. AI Field Performance
| Season | Window | BTC | Regime | Field Avg Return | Models Positive | Winner & Return |
|---|---|---|---|---|---|---|
| Season 3 | Mar-Apr 2026 | +10.1% | Bull | -6.7% | 0 of 9 | MiniMax M2.5, -0.63% |
| Season 4 | Apr-May 2026 | -3.3% | Mild bear | +3.7% | 8 of 9 | MiniMax M2.7, +6.94% |
| Season 5 | May-Jun 2026 | -15.0% | Hard bear | +4.0% | 8 of 10 | Gemini 3.5 Flash, +13.76% |
In the one rising market we have tested, the entire field finished underwater. Not a single one of the nine models made money while BTC gained 10%. The winner, MiniMax, 'won' by losing only 0.63%. In the two falling markets that followed, the field flipped to positive and stayed there, with eight of nine and then eight of ten models in profit, and winners returning +6.94% and +13.76%.
Run the correlation between BTC's return and the field's average return across these three seasons and you get -0.90. But read it correctly: this is not a sliding scale where a worse market pays more. The two bears prove that — BTC fell about four and a half times harder in Season 5 than in Season 4 (-15% versus -3.3%), yet the field's average return barely moved (+4.0% versus +3.7%). The pattern is binary, not graded: the field lost in the one bull and won in both bears. The -0.90 mostly captures that clean bull-versus-bear split — striking enough for a strategy that, on paper, has no instruction to be bearish, but not a dose-response curve.
The correlation, by the numbers. Season 3: BTC +10.1%, field -6.7%. Season 4: BTC -3.3%, field +3.7%. Season 5: BTC -15.0%, field +4.0%. Pearson correlation between BTC return and field-average return across the three frontier seasons: -0.90. Caveat up front: that is three data points. It is a strong signal, not a statistically settled fact. We will add Season 6 to the series as it closes.
Why: The Models Carry a Short Lean
The reason is not that the models are good at calling tops and bottoms. It is that they are structurally biased toward the short side and toward defense, and that bias pays off when the market falls and burns when it rises.
The cleanest evidence is in the profit-and-loss composition, not the headline returns. In the Season 3 bull, the models did not just fail to capture the rally; they actively bet against it. They loaded short positions on bearish-looking technical setups and held them while the market climbed, losing on both their closed trades and their open ones. They were short a market that went up.
In the Season 5 bear, the same instinct paid. Nine of the ten models ended the season holding only short positions. Their open shorts, marked to market at the lows, were worth a combined +$7,784. They were short a market that went down, and this time the market obliged.
The direction of the bet barely changed between the two seasons. What changed was whether the market agreed with it. That is what a structural lean looks like: the same behavior, rewarded or punished entirely by the regime it lands in.
One honest complication before leaning too hard on the word 'short': the models reason in explicit trend-following terms — they cite a 'trend-following mandate' and a refusal to fight the tape. So part of what looks like a short bias may be a trend-following bias that happened to ride the two sustained downtrends and get chopped in the one grinding bull. The two are hard to fully separate on this data. What argues for a genuine directional tilt, rather than pure trend-following, is the budget era: those earlier, cheaper models leaned the opposite way, as chronic dip-buyers, under the same rules. If the prompt's framing alone dictated the behavior, both generations would lean the same direction. They don't.
The distinction that matters: these models are not timing the market, they are tilted against it. A trader who times well makes money in both directions. A trader with a permanent short tilt makes money only when the market falls, and looks like a genius for exactly as long as the bear lasts. Season 5's +13.76% winner was not a market wizard. It was a model that was correctly leaning the way the market happened to move, and then mostly capped by its own account limits from over-managing it.
Caveat One: This Is a Frontier-Model Trait, Not a Law
It would be easy to overstate this into 'AI traders make money when markets crash.' They do not, as a category. The behavior is specific to the current generation of premium frontier models, and we can prove it, because we ran the earlier, cheaper models through bear markets too.
In Season 1, BTC fell 22% and the budget-model field lost 8%, with only two of nine finishing positive. In Season 2, BTC fell 6% and the field lost 1.6%, with only three of thirteen positive. Falling markets, and the cheaper models still lost. The reason is that they leaned the other way: they were chronic dip-buyers, long-biased, forever trying to catch the bounce. That is why the standout strategy of those seasons was a contrarian bot that simply inverted a base model's calls and rode it to first place.
So the regime is not the whole story. The model's built-in lean is. Frontier models in 2026 lean short and defensive; the budget models before them leaned long. Put either one in the wrong market and it loses. The market does not change which way a model leans; it just reveals it.
Bear Markets Did Not Save the Budget Models
| Season | Model Tier | BTC | Field Avg | Models Positive |
|---|---|---|---|---|
| Season 1 | Budget | -22.4% | -8.0% | 2 of 9 |
| Season 2 | Budget | -5.7% | -1.6% | 3 of 13 |
| Season 4 | Frontier | -3.3% | +3.7% | 8 of 9 |
| Season 5 | Frontier | -15.0% | +4.0% | 8 of 10 |
Caveat Two: The Bear-Market Win Was Mostly Unrealized
The second caveat is about what 'won' means. Season 5's eight-of-ten green finish was not eight of ten models banking profits. At the final bell, nine of ten models were still holding their shorts open, and the standings marked those positions to market at the season-end snapshot, near the lows. On the trades the models actually closed and booked, the field lost $3,837. Only two of ten finished with positive realized P&L.
That does not erase the result, but it qualifies it. A short lean wins on the scoreboard in a bear market because open shorts show large unrealized gains at a low. Whether those gains survive depends on what happens next. If the market had bounced in the final days, much of that paper profit would have evaporated, and the 'AI wins in bear markets' headline would have a much thinner season behind it. The lean is real. The booked edge is smaller than the leaderboard suggests. We unpack this fully in the Season 5 post-mortem.
What Stayed Constant (So You Know It Is Not a Rules Artifact)
One obvious objection: maybe the models only started winning in bears because we turned on short-selling at some point. We did not. Short-selling has been available since the very first season, alongside the same $10,000 starting capital, the same 0.1% fee, the same no-leverage rule, and the same mandatory stop-loss. The toolkit has not changed.
That is what makes the budget-versus-frontier contrast meaningful. The budget models of Seasons 1 and 2 had the exact same ability to short a falling market and mostly chose not to. The frontier models of Seasons 3 through 5 reach for it by default. The change is in the models' judgment, not in what the competition allowed them to do. When you give two generations of models the identical tools and the identical falling market and one generation shorts while the other buys the dip, you are watching a genuine shift in how these systems read risk.
Which Models Are Consistent Across Regimes
If the field tracks the market, the interesting question becomes which individual models hold up across regimes rather than just catching one. Three patterns stand out across the frontier seasons, with one caveat to keep in mind: each provider's slot ran a different model version each season ('Gemini' was 3.1 Pro, then 3.5 Flash; MiniMax went M2.5 to M2.7), so 'consistency' here is vendor-slot consistency with an upgraded model each time, not one fixed agent proving itself three times.
Gemini is the most consistent. It finished 2nd in the bull, 4th in the mild bear, and 1st in the hard bear. It has been top-four in every frontier season and is the only model to convert that consistency into a championship. Whatever Gemini is doing, it travels.
MiniMax is the cautionary tale of regime-fit. It won Seasons 3 and 4 outright with a patient, capital-preserving style, then finished dead last in Season 5 when that same style misread a hard directional trend and left it long the bear. Back-to-back champion to lanterne rouge in one season is the sharpest reminder that no trading style is good in the abstract, only good for a given market.
GLM is the reliable bottom. It finished 7th, 9th, and 9th across the three seasons, near the back in every regime. Consistency cuts both ways.
Frontier-Season Finishing Positions
| Model | Season 3 | Season 4 | Season 5 |
|---|---|---|---|
| Gemini | 2nd | 4th | 1st |
| DeepSeek | 8th | 2nd | 2nd |
| Kimi | 5th | 5th | 4th |
| Qwen | 3rd | 7th | 5th |
| MiniMax | 1st | 1st | 10th (last) |
| GPT | 4th | 6th | 8th |
| Grok | 9th (last) | 3rd | 7th |
| Claude Opus | 6th | 8th | 6th |
| GLM | 7th | 9th (last) | 9th |
What This Means If You Use AI to Trade
The practical lesson is not 'wait for a bear market and let an AI run your book.' It is narrower and more useful than that.
First, know which way your model leans before you trust its signals. A model with a structural short bias is a good second opinion on downside risk and a bad one on whether to chase a rally. If you ask a 2026 frontier model whether to buy a strong uptrend, you are asking a witness with a known bias.
Second, treat regime as the dominant variable. The spread between the field's best and worst season was about ten percentage points of average return, driven almost entirely by what the market did, not by which model you picked. Model choice mattered at the margin; regime mattered at the core. If you are evaluating any AI trading claim, the first question is always 'what was the market doing during the test,' and the second is 'did they show you realized P&L or marked-open positions.'
Third, do not confuse a lean with a skill. The most impressive single-season return we have recorded came from a model that was correctly tilted and then structurally prevented from changing its mind. That is a fine outcome. It is no proof the model can handle whatever comes next.
The one-sentence version: today's frontier AI models are not market-beating traders, they are market-leaning ones, and the lean happens to be short, which is why they look brilliant in a crash and hopeless in a rally. Judge them on whether their lean matches the market you are actually in.
Related Reading
- Season 5 Final: Gemini Flash Won a 15% Bear Market — the latest data point, with the realized-vs-unrealized breakdown
- Can AI Beat the Market? Two Seasons of Data Say It Depends — the earlier version of this question, now extended by two more seasons
- Season 4 Final: All 9 Premium AI Models Beat BTC — the mild-bear season that started the pattern
- Season 3 Final: All 9 Premium AI Models Lost Money — the bull market that broke them
- Best AI Models for Crypto Trading: 2026 Ranking — the living cross-season model ranking
- Which LLM trades crypto best — the live LLM trading benchmark across every season
- How TradeRank.ai works — identical rules, prompt, and scoring across every season