Stocks or Crypto? AI Traders Picked Stocks, Then Lost in Both

Eight LLM accounts could trade both classes from February 8 to March 8, 2026. They chose stocks in the first ten minutes and stayed there. Both books lost money, and neither class separated from a benchmark-long dummy scored on the same windows.

For one month in early 2026, February 8 through March 8, TradeRank's competitors could buy US stocks and crypto out of the same wallet. Season 2 is the first archived season where both classes were tradeable, and the only such season in the fast-cycle era: eight LLM accounts, $10,000 of simulated capital each, one decision every six hours, 117 cycles, and 87 tradeable assets split 49 stocks to 38 crypto pairs. Every account ran the same prompt and the same rules, set out on How It Works.

The season fielded more than those eight. Reverse agents, which invert another model's decisions, and user-submitted slots traded it too, and this article excludes both: 530 reverse-agent trades and 216 user-slot trades are dropped from every figure here, leaving the 289 trades the eight LLM accounts made.

That setup answers a question the crypto-only seasons cannot: what an LLM reaches for when both classes sit on equal terms, and whether the choice pays.

The figures come out of one generated dataset (asset-class-stats.json, built August 7, 2026 from the archived season files). One figure in the article is derived by adding two of its fields, and it is labelled where it appears. The season is closed and archived, so these numbers are final and this page is not refreshed as new seasons run. Season 7 is trading both classes right now, but as of August 7, 2026 it is still open, so no Season 7 figure appears here — its standings are on the live LLM trading leaderboard.

Data Point

Two boards, one account. Season 2 offered 49 US stocks beside 38 crypto pairs, and eight LLM agents worked both with $10,000 of simulated capital each, deciding every six hours across 117 cycles from February 8 to March 8, 2026. They made 289 trades, 234 of them equity. Archive: Season 2 results.

The Roster Changed Six Days In

Four of the eight accounts stop on February 13, at cycle 32: Qwen3, Claude Haiku 4.5, DeepSeek R1 and Kimi K2. MiniMax M2.5 appears the next day, at cycle 36, and runs to the close.

Those four had already bought stocks. Qwen3 made nine equity trades, Claude Haiku 4.5 ten, Kimi K2 ten, DeepSeek R1 seven. Between them they closed zero equity legs. That is an absence, not a result: they left with positions open, and an open position has no realized P&L to score.

So the equity numbers in this article come from four of the eight models. Gemini 3.0 Flash closed 34 equity legs, GPT-5 Mini 32, Grok 4-1 Fast 29, MiniMax M2.5 15. Allocation and breadth cover all eight; realized money covers those four.

The Allocation Was Settled in One Cycle

Season 2's opening cycle held nothing. Every account was in cash.

The season's first cycles ran minutes apart before the six-hour schedule settled in. At the second of them, ten minutes after the first, $6,525.47 was deployed and 79.6% of it sat in stocks. The book never came back. Across the 112 cycles that held any capital at all, the equity share of invested capital ranged from 76.5% to 100%, with a median of 93.8%, and it sat at or above 80% in 102 of them. At the last cycle holding capital, on March 8, the book was 100% stocks.

112 of the season's 117 cycles held capital. The rest held none, and a share of nothing is not a zero, which is why every share figure here is defined on 112.

Where the Invested Capital Sat

MeasureEquity share of invested capital
First cycle holding capital (cycle 2, Feb 8)79.6%
Minimum across the 112 capital-holding cycles76.5%
Median across the 112 capital-holding cycles93.8%
Maximum across the 112 capital-holding cycles100%
Cycles at or above 80%102 of 112
Last cycle holding capital (cycle 113, Mar 8)100%

Both Books Lost Money

The equity book closed 110 legs for -$538.87 realized. The crypto book closed 28 for -$97.23. Season 2's trade log survived intact, so those are complete totals for closed legs rather than floors.

The trade counts follow the allocation: 234 of the 289 scoped trades were equity, 81% of the season's activity, against 55 crypto trades. All eight models traded stocks. Seven of the eight traded crypto, and the exception is Claude Haiku 4.5, whose ten trades were all equity.

Per model, the equity share of trades runs from DeepSeek R1's 43.8%, the one model that put more trades into crypto than stocks, up to Claude Haiku 4.5's 100%.

Two models closed the crypto book in profit. GPT-5 Mini banked +$115.78 across 11 closed legs. Grok 4-1 Fast banked +$6.19 on one closed leg, its only crypto exit of the season. Nobody closed the equity book in profit: of the four models that closed equity legs, all four are negative, from -$55.76 to -$190.01. And nobody finished the two books together in profit. GPT-5 Mini's crypto gain is smaller than its -$182.31 equity loss, and adding its two class rows, the article's one derived figure, gives -$66.53.

Each Model's Split, by Asset Class

ModelEquity / crypto tradesEquity share of tradesEquity legsEquity realizedCrypto legsCrypto realized
Gemini 3.0 Flash66 / 1878.6%34-$190.0110-$117.28
GPT-5 Mini49 / 1873.1%32-$182.3111+$115.78
Grok 4-1 Fast56 / 296.6%29-$110.791+$6.19
MiniMax M2.527 / 293.1%15-$55.761-$35.27
DeepSeek R17 / 943.8%0n/a3-$29.38
Kimi K210 / 376.9%0n/a1-$4.43
Qwen39 / 375%0n/a1-$32.83
Claude Haiku 4.510 / 0100%0n/a0n/a

The four rows showing no closed legs are the models that left on February 13. Their realized column is blank rather than zero, because zero would assert a measurement nobody took.

Against Buy-and-Hold, Neither Class Separates

Of the 110 closed equity legs, 43 made money: 39.1%. Of the 28 closed crypto legs, 12 did: 42.9%. Both are closed-leg realized rates. The leaderboard's win rate counts still-open positions as well, which makes it a different measurement.

Against a fair coin, the equity rate is low enough to register: 43 of 110, p = 0.028. The crypto rate is not: 12 of 28, p = 0.572.

A coin is the wrong opponent, though, because it ignores what the market was doing while each position was open. The dummy used here buys and holds the class benchmark across each closed leg's own entry-to-exit window and counts as correct when that benchmark rose. Model and dummy are scored on exactly the same windows, so the two outcomes are paired, and the test is an exact McNemar on the legs where they disagree.

On the equity side, 10 of the 110 legs sit on windows where SPY's two ends read the same price and drop out, leaving 100 scored. The models were right on 43 of them and an SPY-long dummy on 47. Where they disagreed, 24 legs went the models' way and 28 the dummy's. p = 0.678.

On the crypto side, 3 of the 28 legs drop out flat, leaving 25. The models were right on 11 and a BTC-long dummy on 8, which is 44% against 32%. Five discordant legs went the models' way and two the dummy's. p = 0.453. A twelve-point lead resting on a handful of discordant legs is what noise looks like.

Neither class produced a Season 2 result you could tell apart from drift.

Model vs Benchmark-Long, Same Windows

Equity (SPY)Crypto (BTC)
Closed legs11028
Dropped, flat benchmark window103
Legs scored10025
Model correct43 (43%)11 (44%)
Benchmark-long dummy correct47 (47%)8 (32%)
Discordant, model only245
Discordant, dummy only282
Exact McNemar p0.6780.453

Two Comparisons Separate, and Both Go Against the Models

Across every season row and era row in this dataset, two comparisons reach a p-value below 0.05. Both are crypto, both come from the fast-cycle era, and both run against the models.

Season 1 is one of them, and it carries the larger gap. It ran from January 11 to February 8, 2026 on four-hour cycles, crypto only, and closed 724 legs. 684 could be scored against a BTC-long dummy: the models were right on 220 of them, 32.2%, and the dummy on 298, 43.6%. The gap is 11.4 points, the disagreements split 87 to the models and 165 to the dummy, and p = 0.000001, the finest resolution this dataset reports.

The other is the fast-cycle era as a whole, Seasons 0 through 2, January 10 to March 8, 2026. Its 807 closed crypto legs yield 709 scored, after dropping 59 whose windows carry no benchmark price, 28 on flat benchmark windows and 11 realizing trades with no surviving open to match. On those 709 the models were right on 231, 32.6%, and the dummy on 306, 43.2%. The gap is 10.6 points, the disagreements split 92 to 167, and p = 0.000004.

Those two rows are the same finding seen at two scopes. Season 1 supplies 96.5% of the era's scored legs, 684 of the 709. Season 2 supplies the other 3.5%, its 25. Season 0 supplies none at all: all 55 of its closed crypto legs dropped out, 44 for want of a benchmark price across their windows and 11 as realizing trades with no surviving open to match.

Season 1 is also the season whose trade log survived only as a partial snapshot. That makes its realized crypto P&L of -$4,811.98 a floor rather than a total, and it leaves the leg set possibly short of closures the log never recorded. The era row inherits the same limit: its -$4,775.29 is a floor too.

The two separating rows are the two largest samples in the dataset, at 709 and 684 scored legs. Season 2's own crypto row rests on 25 and its equity row on 100, and neither of those separates from anything.

Key Insight

Season 2 put two asset classes, one prompt and one capital pool in front of each model. The allocation was set in the first ten minutes and held for a month. Both classes lost money, and scored against a benchmark-long dummy on the same windows, neither one separated from drift.

One Stock Carried the Equity Book

The models touched 34 of the 49 available stocks and 17 of the 38 available crypto pairs.

Within those 34 names, one paid for everything that worked. AMAT closed 10 legs for +$422.36 across four models, against an equity book that finished at -$538.87.

The most-traded name was also the worst. NVDA took 24 trades from five different models and closed 14 legs for -$206.64. C is second worst at -$173.51.

The crypto extremes are smaller in both directions: DOT is the best at +$107.41 across three closed legs, FIL the worst at -$83.86 across two.

The Equity Book's Extremes

SymbolTradesClosed legsModelsRealized
AMAT16104+$422.36
CAT1696+$59.70
GE843+$20.16
MU18104-$93.80
C734-$173.51
NVDA24145-$206.64

Methodology and Limitations

Where the numbers come from. Every figure is read from a generated dataset (asset-class-stats.json, generated August 7, 2026 from the archived season files, archive snapshot July 18, 2026). Class totals are quoted exactly as the generator emits them. The per-symbol rows are rounded independently and do not sum to the class total, so they are read individually and never added together. One figure in the article is derived: GPT-5 Mini's -$66.53 combined total, which adds its two class rows, and it is labelled at the point of use.

Who is counted. Season 2 ran more accounts than the eight LLM models covered here. Reverse agents, which invert another model's decisions, and user-submitted slots are both excluded: 530 reverse-agent trades and 216 user-slot trades dropped, along with 459 exposure snapshots, leaving 289 trades across the eight. The reverse agents have their own write-ups, linked below.

Closed legs. A closed leg is a trade carrying a realized P&L, matched first-in-first-out against the earliest unclosed opening trade for the same account and symbol, within one season. Capital resets at a season boundary and open positions are liquidated into the close-out, so legs are never paired across seasons.

Win-rate basis. Every win rate here counts closed legs with realized P&L. The leaderboard's win rate counts still-open positions too. They measure different things.

The flat-benchmark exclusions. The dummy can only be scored when its benchmark moved during a leg's window. SPY's Season 2 archive holds 43 distinct prices for the whole month, because equities price only during market hours and the archive stores a coarse step function; that coarseness is why 10 of the 110 equity legs land on windows whose two ends read the same SPY price. BTC, priced continuously, has 115 distinct prices over the same month and loses only 3 of 28. A flat window means the benchmark did not move, so the leg is set aside and scored for neither side.

Era versus season. A season row describes one season with one capital base. An era row concatenates seasons that share a cycle length and reset capital between them, so it describes a regime across all of them. The fast-cycle era covers Seasons 0 to 2, and its realized P&L is a floor rather than a total because Season 1's trade log survived only as a partial snapshot. Trade counts also stop being comparable across the era boundary, since the fast-cycle era ran several cycles a day and the current daily regime runs one. Compare shares across that boundary; counts do not carry.

Sample size. 289 trades, 110 closed equity legs, 28 closed crypto legs, eight models, one month. This is a case study of one season. It does not measure any model's ability, and it does not settle stocks against crypto.

Frozen window. Season 2 closed on March 8, 2026 and is archived. Every figure here is final, and this page is not updated as new seasons run.

Warning

This article describes a paper-trading competition. The capital is simulated, the prices are live, and nothing here is financial advice or a claim that any model can trade profitably. Past simulated results do not predict future results.

Frequently Asked Questions

Do AI models trade stocks or crypto better?

Neither, on TradeRank's Season 2 data. The eight models closed 110 equity legs for -$538.87 realized and 28 crypto legs for -$97.23. Scored against a benchmark-long dummy on the same windows, the equity side was right on 43 of 100 scored legs against the dummy's 47 (exact McNemar, p = 0.678), and the crypto side on 11 of 25 against the dummy's 8 (p = 0.453).

How much of an AI trading portfolio went into stocks?

A median 93.8% of invested capital, across the 112 of Season 2's 117 cycles that held any capital. The allocation was set immediately: 79.6% in stocks at the second cycle of day one, February 8, 2026. It stayed at or above 80% in 102 of those 112 cycles, ranged from 76.5% to 100%, and read 100% at the last cycle holding capital. All eight models traded stocks, and Claude Haiku 4.5 traded nothing else, its ten trades all being equity.

Did any AI model make money trading stocks in Season 2?

No. Four of the eight models closed equity legs at all, and all four finished the equity book negative: MiniMax M2.5 at -$55.76, Grok 4-1 Fast at -$110.79, GPT-5 Mini at -$182.31 and Gemini 3.0 Flash at -$190.01. The other four left the competition on February 13 holding open stock positions, so they closed no equity legs; nothing settled, so there is nothing to score. Two models did close the crypto book in profit, GPT-5 Mini at +$115.78 across 11 legs and Grok 4-1 Fast at +$6.19 on a single leg, but no model finished both books together in profit.

Was the AI equity result spread across many names or concentrated?

Concentrated. The models touched 34 of the 49 available stocks, and one name carried the upside: AMAT closed 10 legs for +$422.36 across four models, against an equity book that finished at -$538.87. The most-traded name was also the worst, with NVDA taking 24 trades from five models and closing 14 legs for -$206.64. Crypto breadth was 17 of 38 available pairs, and its extremes were smaller in both directions.

Did the AI models ever beat a buy-and-hold benchmark in this dataset?

No. Two comparisons in the dataset reach a p-value below 0.05, and both run the other way. Season 1's crypto book was right on 220 of 684 scored legs (32.2%) against a BTC-long dummy's 298 (43.6%), p = 0.000001. The fast-cycle era containing it reads 32.6% against 43.2% on 709 scored legs, p = 0.000004, and Season 1 supplies 96.5% of those 709 legs, 684 of them, on a trade log that survived only as a partial snapshot. Every other row, Season 2's included, sits in the noise.

Season 7 is live

Watch the AI models trade in real time

12 AI models trading live. Every decision logged and explained. Follow the AI trading competition on the TradeRank.ai arena.

See the live competition →
← Back to The Signal