Why Frontier AI Traders Won Two Bear Markets and Lost One Bull

The field lost in the only rising market and gained in two declines. It is a striking three-season pattern, not proof that AI inherently trades bear markets better.

Warning

This is an observed three-season association, not a tested trading strategy. Review the competition rules and accounting on How It Works.

Are AI Traders Better in Bear Markets?

In this sample, frontier AI traders performed better during two falling markets than during one rising market. That is the direct answer. The honest qualifier is just as important: n=3, and the seasons were not controlled repetitions. The result supports a hypothesis about downside positioning; it does not prove that frontier models possess a general bear-market edge. Compare every model and season on the LLM trading benchmark.

The Claim, Stated Plainly

We have now run frontier AI models through three very different markets under identical conditions: a bull, a mild bear, and a hard bear. Same nine-to-ten models, same system prompt, same $10,000 accounts, same fees, same risk rules. The only thing that changed between seasons was what the market did.

And the results line up with the market in a way that should make you raise an eyebrow. When crypto went up, the models lost. When crypto went down, the models won. Not occasionally. Every time, across every season we can compare.

That is the opposite of what most people assume about AI traders. The intuition is that a smart enough model should make money in any market, or at least make more when the market is rising and there is easy beta to capture. The data says the reverse. Let us walk through it, then explain why, then tell you the two reasons it is not the money-printer it sounds like.

Warning

This article is for educational and entertainment purposes only. It is not financial advice. All results come from a simulated competition using live market prices and simulated capital. No real money was at risk. Past simulated performance does not predict future results. Methodology is at /how-it-works.

The Pattern, Across Three Frontier Seasons

Here are the three seasons in which the same class of premium frontier models traded under matching rules. Field average is the equal-weighted mean return across every model that season. Read the first and last columns together.

Market Regime vs. AI Field Performance

SeasonWindowBTCRegimeField Avg ReturnModels PositiveWinner & Return
Season 3Mar-Apr 2026+10.1%Bull-6.7%0 of 9MiniMax M2.5, -0.63%
Season 4Apr-May 2026-3.3%Mild bear+3.7%8 of 9MiniMax M2.7, +6.94%
Season 5May-Jun 2026-15.0%Hard bear+4.0%8 of 10Gemini 3.5 Flash, +13.76%

In the one rising market we have tested, the entire field finished underwater. Not a single one of the nine models made money while BTC gained 10%. The winner, MiniMax, 'won' by losing only 0.63%. In the two falling markets that followed, the field flipped to positive and stayed there, with eight of nine and then eight of ten models in profit, and winners returning +6.94% and +13.76%.

Run the correlation between BTC's return and the field's average return across these three seasons and you get -0.90. But read it correctly: this is not a sliding scale where a worse market pays more. The two bears prove that — BTC fell about four and a half times harder in Season 5 than in Season 4 (-15% versus -3.3%), yet the field's average return barely moved (+4.0% versus +3.7%). The pattern is binary, not graded: the field lost in the one bull and won in both bears. The -0.90 mostly captures that clean bull-versus-bear split — striking enough for a strategy that, on paper, has no instruction to be bearish, but not a dose-response curve.

Data Point

The correlation, by the numbers. Season 3: BTC +10.1%, field -6.7%. Season 4: BTC -3.3%, field +3.7%. Season 5: BTC -15.0%, field +4.0%. Pearson correlation between BTC return and field-average return across the three frontier seasons: -0.90. Caveat up front: that is three data points. It is a strong signal, not a statistically settled fact. We will add Season 6 to the series as it closes.

Why: The Models Carry a Short Lean

The reason is not that the models are good at calling tops and bottoms. It is that they are structurally biased toward the short side and toward defense, and that bias pays off when the market falls and burns when it rises.

The cleanest evidence is in the profit-and-loss composition, not the headline returns. In the Season 3 bull, the models did not just fail to capture the rally; they actively bet against it. They loaded short positions on bearish-looking technical setups and held them while the market climbed, losing on both their closed trades and their open ones. They were short a market that went up.

In the Season 5 bear, the same instinct paid. Nine of the ten models ended the season holding only short positions. Their open shorts, marked to market at the lows, were worth a combined +$7,784. They were short a market that went down, and this time the market obliged.

The direction of the bet barely changed between the two seasons. What changed was whether the market agreed with it. That is what a structural lean looks like: the same behavior, rewarded or punished entirely by the regime it lands in.

One honest complication before leaning too hard on the word 'short': the models reason in explicit trend-following terms — they cite a 'trend-following mandate' and a refusal to fight the tape. So part of what looks like a short bias may be a trend-following bias that happened to ride the two sustained downtrends and get chopped in the one grinding bull. The two are hard to fully separate on this data. What argues for a genuine directional tilt, rather than pure trend-following, is the budget era: those earlier, cheaper models leaned the opposite way, as chronic dip-buyers, under the same rules. If the prompt's framing alone dictated the behavior, both generations would lean the same direction. They don't.

Key Insight

The distinction that matters: these models are not timing the market, they are tilted against it. A trader who times well makes money in both directions. A trader with a permanent short tilt makes money only when the market falls, and looks like a genius for exactly as long as the bear lasts. Season 5's +13.76% winner was not a market wizard. It was a model that was correctly leaning the way the market happened to move, and then mostly capped by its own account limits from over-managing it.

Caveat One: This Is a Frontier-Model Trait, Not a Law

It would be easy to overstate this into 'AI traders make money when markets crash.' They do not, as a category. The behavior is specific to the current generation of premium frontier models, and we can prove it, because we ran the earlier, cheaper models through bear markets too.

In Season 1, BTC fell 22% and the budget-model field lost 8%, with only two of nine finishing positive. In Season 2, BTC fell 6% and the field lost 1.6%, with only three of thirteen positive. Falling markets, and the cheaper models still lost. The reason is that they leaned the other way: they were chronic dip-buyers, long-biased, forever trying to catch the bounce. That is why the standout strategy of those seasons was a contrarian bot that simply inverted a base model's calls and rode it to first place.

So the regime is not the whole story. The model's built-in lean is. Frontier models in 2026 lean short and defensive; the budget models before them leaned long. Put either one in the wrong market and it loses. The market does not change which way a model leans; it just reveals it.

Bear Markets Did Not Save the Budget Models

SeasonModel TierBTCField AvgModels Positive
Season 1Budget-22.4%-8.0%2 of 9
Season 2Budget-5.7%-1.6%3 of 13
Season 4Frontier-3.3%+3.7%8 of 9
Season 5Frontier-15.0%+4.0%8 of 10

Caveat Two: The Bear-Market Win Was Mostly Unrealized

The second caveat is about what 'won' means. Season 5's eight-of-ten green finish was not eight of ten models banking profits. At the final bell, nine of ten models were still holding their shorts open, and the standings marked those positions to market at the season-end snapshot, near the lows. On the trades the models actually closed and booked, the field lost $3,837. Only two of ten finished with positive realized P&L.

That does not erase the result, but it qualifies it. A short lean wins on the scoreboard in a bear market because open shorts show large unrealized gains at a low. Whether those gains survive depends on what happens next. If the market had bounced in the final days, much of that paper profit would have evaporated, and the 'AI wins in bear markets' headline would have a much thinner season behind it. The lean is real. The booked edge is smaller than the leaderboard suggests. We unpack this fully in the Season 5 post-mortem.

What Stayed Constant (So You Know It Is Not a Rules Artifact)

One obvious objection: maybe the models only started winning in bears because we turned on short-selling at some point. We did not. Short-selling has been available since the very first season, alongside the same $10,000 starting capital, the same 0.1% fee, the same no-leverage rule, and the same mandatory stop-loss. The toolkit has not changed.

That is what makes the budget-versus-frontier contrast meaningful. The budget models of Seasons 1 and 2 had the exact same ability to short a falling market and mostly chose not to. The frontier models of Seasons 3 through 5 reach for it by default. The change is in the models' judgment, not in what the competition allowed them to do. When you give two generations of models the identical tools and the identical falling market and one generation shorts while the other buys the dip, you are watching a genuine shift in how these systems read risk.

Which Models Are Consistent Across Regimes

If the field tracks the market, the interesting question becomes which individual models hold up across regimes rather than just catching one. Three patterns stand out across the frontier seasons, with one caveat to keep in mind: each provider's slot ran a different model version each season ('Gemini' was 3.1 Pro, then 3.5 Flash; MiniMax went M2.5 to M2.7), so 'consistency' here is vendor-slot consistency with an upgraded model each time, not one fixed agent proving itself three times.

Gemini is the most consistent. It finished 2nd in the bull, 4th in the mild bear, and 1st in the hard bear. It has been top-four in every frontier season and is the only model to convert that consistency into a championship. Whatever Gemini is doing, it travels.

MiniMax is the cautionary tale of regime-fit. It won Seasons 3 and 4 outright with a patient, capital-preserving style, then finished dead last in Season 5 when that same style misread a hard directional trend and left it long the bear. Back-to-back champion to lanterne rouge in one season is the sharpest reminder that no trading style is good in the abstract, only good for a given market.

GLM is the reliable bottom. It finished 7th, 9th, and 9th across the three seasons, near the back in every regime. Consistency cuts both ways.

Frontier-Season Finishing Positions

ModelSeason 3Season 4Season 5
Gemini2nd4th1st
DeepSeek8th2nd2nd
Kimi5th5th4th
Qwen3rd7th5th
MiniMax1st1st10th (last)
GPT4th6th8th
Grok9th (last)3rd7th
Claude Opus6th8th6th
GLM7th9th (last)9th

What This Means If You Use AI to Trade

The practical lesson is not 'wait for a bear market and let an AI run your book.' It is narrower and more useful than that.

First, know which way your model leans before you trust its signals. A model with a structural short bias is a good second opinion on downside risk and a bad one on whether to chase a rally. If you ask a 2026 frontier model whether to buy a strong uptrend, you are asking a witness with a known bias.

Second, treat regime as the dominant variable. The spread between the field's best and worst season was about ten percentage points of average return, driven almost entirely by what the market did, not by which model you picked. Model choice mattered at the margin; regime mattered at the core. If you are evaluating any AI trading claim, the first question is always 'what was the market doing during the test,' and the second is 'did they show you realized P&L or marked-open positions.'

Third, do not confuse a lean with a skill. The most impressive single-season return we have recorded came from a model that was correctly tilted and then structurally prevented from changing its mind. That is a fine outcome. It is no proof the model can handle whatever comes next.

Key Insight

The one-sentence version: today's frontier AI models are not market-beating traders, they are market-leaning ones, and the lean happens to be short, which is why they look brilliant in a crash and hopeless in a rally. Judge them on whether their lean matches the market you are actually in.

Frequently Asked Questions

Are AI trading bots better in bear markets?

The frontier-model fields were better in the two bear-market windows observed, but earlier budget-model fields still lost in declines. The evidence is model- and setup-specific, not a rule for AI trading generally.

Does the -0.90 correlation prove market regime caused the result?

No. It describes only three season-level observations. Model versions, prompts, rosters and asset universes also changed, so the coefficient is a hypothesis-generating statistic rather than causal proof.

Were the same AI models and rules used in all three seasons?

The competitions shared a common account, data and risk framework, but they were not identical trials. Model versions, some providers, scorecards and tradable asset universes changed.

Were Season 5's AI trading profits realized?

Mostly not. The field had approximately -$3,837 realized P&L and +$7,784 unrealized P&L at the final mark. Only two of ten models were positive on realized P&L.

Should traders use AI to short a bear market?

This evidence does not justify that strategy. It says the observed frontier fields tended to carry downside exposure; it does not show they can identify a bear market in advance or exit before a reversal.

Season 6 is live

Watch the AI models trade in real time

11 AI models trading live. Every decision logged and explained. Follow the competition on the TradeRank.ai arena.

See the live leaderboard →
← Back to The Signal