GPT vs Claude vs Gemini vs Grok: Which AI Trades Crypto Best?

Not capability benchmarks — logged trading decisions. We ran ChatGPT, Claude, Gemini and Grok on live crypto prices with simulated capital and modeled fees, then published every position and its P&L. Our arena has logged 2,527 trades from 43 AI models to date; here's how these four stack up.

Data Point

TradeRank Arena at a glance (as of 2026-07-08): 43 AI models have traded across 6 seasons since January 2026 — 2,527 trades, $650K simulated capital, 42.6% of model-seasons profitable. This article compares the settled Season 5 verdict with archived Season 2 evidence; see the live LLM trading benchmark for current standings.

Data Point

2026 update (June 23). The Season 2 analysis below still stands as a record of the early competition, but three more seasons have completed since. As of Season 5 (closed June 20, 2026), the answer to "which AI trades best" has moved. The current verdict is in the next section; the original Season 2 deep-dive follows it.

2026 Update: Three Seasons On, the Answer Changed

When this piece was first written, it was Season 2 and the honest conclusion was that no model had proven itself. Three seasons later, with upgraded frontier models and a far larger sample, there is a clearer completed-season answer, and it is not the one the early data suggested.

In Season 5, ten frontier models traded a brutal crypto bear market in which every asset fell between 8% and 32%. Google's Gemini 3.5 Flash won at +13.76% — the highest single-season return in competition history. Non-flagship model tiers swept the podium. The flagship reasoning models all finished outside it — Claude Opus 6th, Grok 7th, GPT-5.5 8th. If you came here asking which AI traded crypto best in the latest completed season, the answer is Gemini, with two caveats worth more than the headline: most of its gain was unrealized P&L on open shorts, and no model holds the title across regimes.

Comparing these models for stocks instead of crypto? See our dedicated head-to-head on Grok vs ChatGPT vs Claude for stock market analysis.

Season 5: The Podium and the Flagships (May-Jun 2026)

ModelProviderReturnRank
Gemini 3.5 FlashGoogle+13.76%1st
DeepSeek V4 ProDeepSeek+11.85%2nd
Mistral Medium 3.5Mistral AI+9.55%3rd
Claude Opus 4.7Anthropic+2.67%6th
Grok 4.3xAI+0.48%7th
GPT-5.5OpenAI+0.38%8th
Key Insight

The 2026 one-liner: Season 5's winner (May–June 2026) was Gemini Flash, non-flagship tiers completed the podium, and there is no permanent winner. The same Gemini finished 4th in Season 4 and 2nd in Season 3. Across those seasons, the field performed better in falling crypto markets than in the rising Season 3 window, an observed association that does not guarantee the next regime. See the live benchmark, Season 5 post-mortem, bull-vs-bear analysis, and completed-season ranking.

Head-to-Head Verdicts: The Pairwise Matchups

If you came here for a single pairwise call, here is what the completed seasons say. Each verdict draws only on finished competitions — no live standings — and links the model's live page so you can check the current run yourself. For an always-updating cross-model view, see our live benchmark: Best LLM for Crypto Trading.

Claude vs Gemini for Trading: Verdict

Across completed seasons, Gemini has been the stronger crypto trader of the two. In Season 5 (May–June 2026), Gemini 3.5 Flash won the whole field at +13.76% while Claude Opus 4.7 finished 6th at +2.67%; in Season 4 Gemini (4th, +4.43%) again edged Claude (8th, +0.88%). Both still trailed the cheaper open-weight models on the podium, so "better of the two" is not the same as "best overall." Track Gemini's live results on the arena.

GPT vs Claude for Trading: Verdict

This one is close and regime-dependent, with no durable edge either way. In Season 5, Claude Opus 4.7 (6th, +2.67%) beat GPT-5.5 (8th, +0.38%); in Season 4 the order flipped, with GPT-5.5 (6th, +3.69%) ahead of Claude (8th, +0.88%). Both flagships consistently finish behind the podium, so the honest answer is that neither has separated from the other. See Claude's live trade history.

Grok vs GPT for Trading: Verdict

Grok's directional conviction gives it more upside in trending markets, while GPT is steadier but rarely tops the field. In Season 4, Grok (3rd, +5.34%) clearly outperformed GPT-5.5 (6th, +3.69%); in the Season 5 bear the two converged near flat, with Grok 4.3 (7th, +0.48%) just ahead of GPT-5.5 (8th, +0.38%). If you want a bolder book you lean Grok, and if you want fewer surprises you lean GPT — but both sit in the reasoning tier below the podium. View Grok's live decisions.

Why Benchmarks Don't Tell the Trading Story

Every few months, a new AI model drops and everyone races to compare benchmarks. MMLU scores. HumanEval pass rates. MATH accuracy. Those scores measure general capabilities; they do not establish whether a model can manage a portfolio under uncertainty.

Trading adds path-dependent decisions, position sizing, fees, and the option to do nothing. If you're wondering which AI is best for trading (ChatGPT, Claude, Gemini, or Grok), a capability benchmark cannot answer that on its own. A trading experiment can expose behavior, but it still needs enough seasons and regimes before it can establish a durable edge.

So we built TradeRank.ai as a live trading arena where AI models compete on live market prices with simulated capital and modeled transaction fees. Same prompt. Same technical data. Same market. Different model. Every decision is logged so the comparison can be audited.

Season 2 ran from February 8, 2026. Thirteen AI models built on GPT-5, Claude, Gemini, Grok, DeepSeek, Qwen, Kimi, and MiniMax made autonomous trading decisions every 6 hours across 89 assets spanning crypto and US equities. After 14 days and 656 logged trades, the results included the discovery that one model's calls were so consistently wrong in that window that simply doing the opposite was profitable.

Warning

This article is for educational and entertainment purposes only. It is not financial advice. The trading results described are from a simulated competition with no real money at risk. Past simulated performance does not predict future results.

Data Point

Season 2 stats as of February 22, 2026: 13 active models, 89 tradeable assets (49 US equities via Yahoo Finance + 21 Binance crypto + 17 Hyperliquid perps + 2 benchmarks), 56 completed trading cycles, 656 total trades executed.

Key Insight

TL;DR: MiniMax M2.5 leads with +2.29% return and 0.88% max drawdown, trading the least (16 trades in 6 days). Grok's simple "buy everything" strategy (+1.42%) beat GPT-5 Mini's sophisticated analysis (-0.67%) and Gemini's aggressive rotation (-0.44%). Inverting DeepSeek's every call produced +2.64%, because a model that is consistently wrong is more useful than one that is randomly wrong.

Data Point

Data as of: Day 14, February 22, 2026. Updated results available in The Win Rate Paradox (Day 25 data).

The Methodology: How We Made This a Fair Fight

Every design choice aimed to hold the environment constant so differences in decisions came from the models rather than different data feeds or account rules.

Identical inputs. All agents received the same system prompt, the same technical indicators (RSI, MACD, EMA, ATR, Bollinger Bands, key levels), and the same OHLCV candlestick data across four timeframes: 1-hour, 4-hour, daily, and weekly.

Two-tier decision flow. With 89 assets, we couldn't dump everything into one prompt (that would be ~400,000 tokens). Each model first scanned all 89 assets in compressed form, selected up to 5 for deeper analysis, then received full OHLCV data for those picks.

Uniform constraints. $10,000 simulated starting capital, a shared open-position cap, a modeled 0.1% fee per trade, no leverage, and direction checks on any stop-loss a model supplied. The archived Season 2 report is the source of truth for historical rules; current seasons use a different prompt and cadence.

6-hour cycles. Models decided every 6 hours across four daily windows. Over 14 days, 56 cycles completed (some cycles were skipped due to maintenance or API outages).

Key Insight

Why 6-hour cycles? Shorter intervals (like 15 minutes) reward speed and encourage overtrading. Longer intervals (daily) don't capture enough of the market's intraday structure. The 6-hour timeframe is the sweet spot for testing analytical ability rather than reaction speed.

The Scoreboard: 14 Days of Live Trading

Where every model stands after two weeks of autonomous trading. We break down each model in detail below, but the headline numbers set the stage.

Main AI Agents: Head-to-Head Performance

MetricGPT-5 Mini (OpenAI)Gemini 3.0 Flash (Google)Grok 4-1 Fast (xAI)MiniMax M2.5 (MiniMax)
Return-0.67%-0.44%+1.42%+2.29%
Win Rate55%38%31%33%
Total Trades38533716
Max Drawdown2.88%3.88%2.35%0.88%
Fees Paid$28.50$49.58$37.25$17.64
Long/Short PrefMixed89% long100% long90% long
Peak Equity$10,131$10,141$10,224$10,231

GPT-5 Mini (OpenAI) — "The Cautious Analyst"

Return: -0.67% | Win Rate: 55% | 38 Trades | Max Drawdown: 2.88%

GPT-5 Mini won more than half its trades and still finished underwater in this dated snapshot. It had the highest win rate among the four main agents, but its wins were smaller than its losses.

It peaked at $10,131 on Day 2 and declined from there. Unlike Gemini and Grok, which leaned overwhelmingly long, GPT-5 Mini used a mix of long and short positions. It shorted SOL, ETH, and XRP during the February 8 selloff, then rotated into longs on CAT and LRCX when it spotted relative strength in industrials and semiconductors.

Its logs show the execution gap clearly. It protected CAT gains while carrying Citigroup (C) through a roughly -4.3% drawdown. That is evidence from one short window, not proof that GPT is always a better analyst than executor, but it explains why a 55% hit rate did not translate into positive P&L. See the full Season 2 competition coverage for the dated field context.

Maintain high cash (~85% available) after closing XRP to preserve optionality in a volatile, BTC-led downtrend and to be ready to scale into confirmed continuation or rebound setups later in the week.

GPT-5 MiniGPT-5 Mini's portfolio strategy during the February 8 selloff — cautious to a fault, leaving money on the table while preserving capital.
Key Insight

GPT-5 Mini's lesson: A high win rate means nothing if your average win is smaller than your average loss. In trading, how much you make when you're right matters more than how often you're right.

Gemini 3.0 Flash (Google) — "The Aggressive Allocator"

Return: -0.44% | Win Rate: 38% | 53 Trades | Max Drawdown: 3.88%

Gemini is the most active trader in the competition and has the fee bill to prove it. Fifty-three trades, $49.58 in transaction fees, nearly half a percent of starting capital burned on execution costs alone.

Its style is aggressive and overwhelmingly bullish. An 89% long preference means it almost never shorts, even in clearly downtrending markets. When the broader market sold off in mid-February, Gemini bought the dip in NVDA, AMD, and JPM rather than stepping aside. Sometimes that worked. More often, it meant catching falling knives.

It peaked at $10,141 on Day 2 and then churned its way down through ill-timed entries, premature exits, and relentless transaction costs. Its worst habit is what traders call "the churn": closing a position only to reopen a similar one a few cycles later, paying 0.1% on both sides each time.

To its credit, Gemini made the single smartest trade of Day 14: closing an INJ position at +19% profit (~$90) when the token became overbought (RSI above 70), then rotating into NVDA for semiconductor momentum. Textbook asset rotation. That one brilliant trade gets diluted by 52 others.

Gemini's 3.88% max drawdown is the highest among main agents, reflecting its tendency to stay fully invested when conditions call for caution. See the full Season 2 competition coverage for its dated results.

Locking in +19% profit (~$90). Asset is overbought (RSI > 70) in a generally bearish crypto environment (BTC bearish). Securing these gains improves weekly performance and frees up a slot.

Gemini 3.0 FlashGemini's reasoning for its best trade of the competition — closing INJ at a 19% gain on February 21. The model recognized overbought conditions and acted decisively.
Warning

The overtrading tax: Gemini's $49.58 in fees on a $10,000 account means it needs to generate 0.50% in raw alpha just to break even on execution costs. Every trade has a built-in hurdle, and 53 trades means 53 hurdles.

Grok 4-1 Fast (xAI) — "The Stubborn Bull"

Return: +1.42% | Win Rate: 31% | 37 Trades | Max Drawdown: 2.35%

Grok was the only main agent in profit in this Day 14 snapshot, and it got there with the simplest directional posture: 100% long across 37 trades, with no shorts in the logged history.

While GPT-5 Mini and Gemini rotated and hedged, Grok kept buying. It peaked at $10,224 on Day 4, when it briefly held the overall lead after loading up on crypto and technology names during the initial rally.

The 31% win rate was the lowest among the main agents, but its winning trades were larger than its losers. In a market that was slightly positive over this 14-day window, a pure long bias happened to fit the regime.

Whether that posture was skill or regime fit cannot be decided from two weeks. The completed Season 2 results later changed the rankings, which is why this section remains explicitly frozen to February 22. See the full Season 2 competition coverage for the dated field context.

ETH long in small profit (1.5%), oversold conditions with bullish 4h MACD; hold with existing stop loss near entry for risk management amid potential bounce.

Grok 4-1 FastGrok's typical reasoning: short, direct, and biased toward holding longs. No overthinking, no second-guessing.

MiniMax M2.5 — "The Patient One"

Return: +2.29% | Win Rate: 33% | 16 Trades | Max Drawdown: 0.88%

MiniMax M2.5 trades like the textbook says you should. It holds 80% of the time, makes fewer trades in two weeks than Gemini makes in three days, and has the best risk-adjusted returns of any active model.

Sixteen trades. 0.88% max drawdown. +2.29% return. MiniMax never saw its equity drop below $9,912 while other models whipsawed through 3-4% drawdowns.

It joined on February 14, six days into the competition. This turned out to be an advantage. It missed the Day 1 false rally and the February 12-13 selloff. By the time it started trading, initial volatility had settled and clearer trends had emerged.

Its asset selection is distinctive. While others chased NVDA, AAPL, and ETH, MiniMax focused on industrial and semiconductor names: AMAT (Applied Materials), LRCX (Lam Research), GE (General Electric), MU (Micron). Relative-strength picks rather than flashy ones.

Its first trade was a short on SOL that initially went against it (-1%) before it closed and rotated into AMAT on a +8% daily surge. That single rotation (losing crypto short to winning semiconductor long) captured MiniMax's entire approach: cut losers fast, ride winners in trending sectors.

At $17.64 in total fees (64% lower than Gemini's), every dollar of return is more capital-efficient. The only question mark is sample size. Sixteen trades over eight active days is small. But the consistency across patience, sector focus, and the willingness to hold cash suggests genuine trading temperament, not luck. See the full Season 2 competition coverage for the dated field context.

Data Point

MiniMax by the numbers: 80% hold rate (chose to do nothing in 4 out of 5 cycles), $17.64 in fees (lowest of any active model), 0.88% max drawdown (best capital preservation), 8 active trading days (fewest of any model). Sometimes the best trade is no trade.

Claude, DeepSeek, Qwen3, and Kimi: The Reverse-Agent Experiments

Season 2 includes four "reverse agents": models that take the output of a base AI model and invert every decision. If Claude says buy, Reverse Claude sells. If DeepSeek says open a long on ETH, Reverse DeepSeek opens a short. The base models (Claude Haiku 4.5, DeepSeek R1, Qwen3, and Kimi K2) were frozen on February 13 to save API costs, but their reverse counterparts continue trading.

This creates a natural experiment. If a reverse agent does well, the base model's instincts were consistently wrong, so wrong that doing the opposite was profitable. If a reverse agent does poorly, the base model was actually making decent calls and reversing them destroyed value. For a deeper analysis of the reverse agent experiment, see our dedicated contrarian strategy article.

Here are the results.

Base Models vs. Their Reverse Counterparts

ModelBase ReturnReverse ReturnSpreadInterpretation
Claude (Anthropic)-3.26%-2.70%+0.56 ptsMarginal — reversal slightly improved, but within noise
DeepSeek (DeepSeek)-3.39%+2.64%+6.03 ptsDeepSeek was consistently wrong — reversing was profitable
Qwen3 (Alibaba)-1.63%+1.18%+2.81 ptsQwen3's instincts were reliably inverted — wrong enough to exploit
Kimi (Moonshot)-0.76%-6.70%-5.94 ptsKimi was noisy, not consistently wrong — reversing amplified the randomness

What the Reverse Agents Reveal

Claude Haiku 4.5 finished at -3.26% with just 10 trades, all buys, and it never sold. Claude loaded up on Day 1 and sat on its positions as they declined. Reverse Claude, which inverted those calls into shorts, returned -2.70% with 99 trades. The marginal spread of +0.56 points is essentially noise. Unlike DeepSeek's dramatic 6.03-point spread, the Claude reversal shows no exploitable bias in either direction. Claude's problem was execution. It never took profits. It never managed risk. It just held.

DeepSeek R1 is the most striking case. The base model returned -3.39%, buying into the Day 1 rally and holding through the subsequent decline. Reverse DeepSeek, by inverting those calls, returned +2.64%, the best performance of any reverse agent and the second-best overall in the competition for most of the season. The 6.03-point spread is enormous. DeepSeek's market reads were so inverted that a simple contrarian strategy turned it into one of the competition's top performers.

Qwen3 shows a similar but less extreme pattern. The base model returned -1.63%; Reverse Qwen returned +1.18%. The 2.81-point spread indicates that Qwen3's instincts were reliably inverted, though not as dramatically as DeepSeek's. Qwen3 traded sparingly (40 trades for the reverse agent), and its errors were more subtle, misdirecting its timing rather than its direction.

Kimi K2 is the cautionary tale. The base model returned a modest -0.76% (the best of the four frozen models), but Reverse Kimi cratered to -6.70%, the worst return in the entire competition. The -5.94-point spread tells us Kimi was not consistently wrong; it was random. Its calls had no reliable pattern to exploit, so inverting them just added noise and transaction costs. Reverse Kimi's 100 trades (the most of any model) amplified that randomness into catastrophic losses.

Key Insight

The reverse-agent experiment reveals a counterintuitive truth: a model that is consistently wrong is more useful than a model that is randomly wrong. DeepSeek's reliable wrongness created a profitable contrarian signal. Kimi's randomness was worse than useless.

Head-to-Head: What the Matchups Reveal

The individual breakdowns above tell you what each model did. The matchups tell you why the differences matter.

Caution vs. Aggression (GPT-5 Mini vs. Gemini)

These two lost similar amounts but for opposite reasons. GPT-5 Mini wins the risk metrics (higher win rate, lower fees, smaller drawdown) yet ended up more negative than Gemini. Conservative execution that cuts winners early can lose more money than aggressive execution that occasionally connects. Temperament, not analysis, was the dividing line.

Simplicity vs. Sophistication (Grok vs. The Field)

The model with the worst win rate (31%), simplest strategy (100% long, zero shorts), and most stubborn conviction beat two models with objectively better analytical capabilities. Markets reward conviction more than analysis when the direction is right. Was Grok genuinely good, or just lucky that a slightly-up market rewarded a pure-long book?

Discipline vs. Activity (MiniMax vs. Everyone)

MiniMax's edge is the combination of metrics, not any single one. Its daily return rate of ~0.38% and maximum drawdown of 0.88% are superior even after adjusting for its six-day late start. The model held cash in 4 out of 5 cycles, focused on sector momentum rather than chasing headlines, and cut losers within a single session. No other model showed that level of discipline across position selection, sizing, and timing.

Trading Style Analysis: What Separates Winners from Losers

Three patterns emerge from the data that separate the profitable models from the unprofitable ones.

Pattern 1: Activity is negatively correlated with returns.

The correlation between trade count and performance is striking and negative. The model with the fewest trades (MiniMax, 16) has the best returns. The model with the most trades among main agents (Gemini, 53) has the second-worst returns. Across all 13 models, the most active traders underperform: Reverse Kimi (100 trades, -6.70%) is the worst performer in the competition, and Reverse Claude (99 trades, -2.70%) fares poorly despite its high volume.

Trade Count vs. Return: The Overtrading Tax

ModelTrade CountReturnFees PaidFees as % of Capital
MiniMax M2.516+2.29%$17.640.18%
Grok 4-1 Fast37+1.42%$37.250.37%
GPT-5 Mini38-0.67%$28.500.29%
Gemini 3.0 Flash53-0.44%$49.580.50%
Reverse DeepSeek65+2.64%----
Reverse Claude99-2.70%----
Reverse Kimi100-6.70%----

Pattern 2: The fee drag is real and underappreciated.

At a 0.1% fee per trade, each round-trip (open + close) costs 0.2% of the position size. That sounds trivial until you realize Gemini paid $49.58 in fees on a $10,000 account, which is 0.50% of starting capital consumed by execution costs alone. Gemini's raw trading P&L before fees was actually slightly positive. The fees turned a marginally profitable strategy into a losing one.

MiniMax, by contrast, paid $17.64 in fees, roughly one-third of Gemini's bill. Those savings compound. On a per-trade basis, MiniMax deploys larger positions (since it has fewer of them) and pays proportionally less in fees relative to its returns.

Pattern 3: Long bias worked in this market, but discipline mattered more.

Three of the four main agents lean strongly long (Gemini at 89%, Grok at 100%, MiniMax at 90%). The only balanced trader (GPT-5 Mini) underperformed. In a slightly-up market, a long bias is the right structural bet.

But long bias alone does not explain the results. Gemini's 89% long preference generated a -0.44% return while Grok's 100% long preference generated +1.42%. The difference is discipline. Grok holds through volatility. Gemini churns. Grok uses wider stops. Gemini cuts and re-enters. The same directional bias, applied with different levels of conviction and patience, produces dramatically different results.

The Day 14 Verdict (February 22 Snapshot)

This section preserves how the four main agents stacked up after 14 days. It is a dated snapshot, not the latest completed-season ranking; no single metric tells the whole story, so the table keeps several criteria side by side.

Final Rankings by Category (Main Agents Only)

Category1st2nd3rd4th
Total ReturnMiniMax (+2.29%)Grok (+1.42%)Gemini (-0.44%)GPT-5 (-0.67%)
Risk-Adjusted (Return/Drawdown)MiniMax (2.60x)Grok (0.60x)Gemini (-0.11x)GPT-5 (-0.23x)
Capital Preservation (Lowest DD)MiniMax (0.88%)Grok (2.35%)GPT-5 (2.88%)Gemini (3.88%)
Win RateGPT-5 (55%)Gemini (38%)MiniMax (33%)Grok (31%)
Fee EfficiencyMiniMax ($17.64)GPT-5 ($28.50)Grok ($37.25)Gemini ($49.58)
Discipline (Hold Rate)MiniMax (80%)Grok (~60%)GPT-5 (~50%)Gemini (~35%)

Snapshot leader: MiniMax M2.5. It ranked first on total return, risk-adjusted return, drawdown, fees, and hold rate at this February 22 cutoff. Its late entry and small trade sample are important limits, so this is a snapshot result rather than a recommendation.

Established-model leader: Grok 4-1 Fast. It was the only main agent in profit from Day 1. Its 100% long posture happened to fit this 14-day window; another regime could punish the same exposure.

Highest hold rate: MiniMax M2.5. An 80% hold rate meant no trade in four out of five decision cycles. That restraint coincided with the best result in this sample, but the complete 22-model-season data does not show activity alone determining returns.

Highest win rate: GPT-5 Mini. Its 55% win rate led the main agents but did not produce a positive return, illustrating that hit rate and total return answer different questions.

Most active main agent: Gemini 3.0 Flash. Gemini logged 53 trades and $49.58 in modeled fees. It also produced the snapshot's best individual close, +19% on INJ, showing how a strong trade and a weak total result can coexist.

Largest base-versus-reverse spread: DeepSeek. Reverse DeepSeek finished at +2.64% while the base agent was -3.39%, a 6.03-point gap in this short window. That is an observation worth retesting, not evidence of a stable inverted signal.

For the complete 13-model competition overview including user-submitted models, see the full competition report.

What This Means for You

If you are using AI for trading insights (whether through ChatGPT, Claude, Gemini, or another model), here is what the first 656 logged simulated trades taught us.

Behavior differed alongside results in this snapshot. The spread between the best and worst main agents was 2.96 percentage points, while activity, sizing, direction, and holding periods also varied. This experiment cannot isolate which factor caused the spread.

The least-active models did best in this window. MiniMax's 16 trades outperformed Gemini's 53. Across the complete 22-model-season sample, however, activity and return were nearly uncorrelated. If you are prompting an AI for trading, treat its output as one input to validate rather than an automatic trigger.

Modeled fees still create drag. At the arena's 0.1% fee assumption, Gemini paid $49.58 across 53 trades. That does not prove low frequency always wins, but every additional trade adds a modeled cost.

Complex reasoning did not guarantee a better return. GPT-5 Mini produced nuanced explanations and still lost money; Grok's simpler long bias made money in this particular window.

No AI model, as of this February 22 snapshot, had proved itself a reliable autonomous trader. For the settled full-season ranking, use the completed-season model ranking; for current standings, use the live LLM trading benchmark.

Key Insight

In this 14-day snapshot, the least-active main agents finished highest. Across the complete 22-model-season sample, activity and return were nearly uncorrelated (about -0.04), so the durable takeaway is narrower: validate each AI-generated idea instead of treating model output as an automatic trade.

This article is part of our Season 2 competition coverage. See also: We Made 13 AI Models Trade Against Each Other for the full competition overview, and What Happens When You Reverse Every AI Trading Decision? We Tested It. for a deep dive into the contrarian strategy experiment. And if you're weighing which AI to use for stock-market analysis specifically, see our dedicated head-to-head on Grok vs ChatGPT vs Claude for stock market analysis.

Methodology Notes and Limitations

What this study is: A controlled comparison of AI model trading behavior using identical prompts, data, and account constraints, with live market prices, simulated capital, and modeled transaction fees.

What this study is not: Financial advice. Proof that any model is profitable long-term. A recommendation for automated trading.

Key limitations:

  • 14 days is a short sample. Model performance could invert in a different market regime.
  • MiniMax's late entry (Day 6) gave it a structural advantage by missing early volatility.
  • Reverse agents can have higher trade counts because inversion changes the source decisions.
  • The flat 0.1% fee does not model slippage, market impact, borrow costs, or funding.
  • These were the specific API versions named above, under one shared prompt; other versions and prompts may behave differently.

Season 2 was ongoing when this frozen section was written. The arena has since moved on to later seasons.

Snapshot date: February 22, 2026. Rankings and statistics reflect 56 completed cycles. Use the completed-season ranking for a settled verdict and the live benchmark for current standings.

Frequently Asked Questions

Which AI is best for stock trading?

As of Season 5 (May–June 2026), our most recent completed competition, Google's Gemini 3.5 Flash led the field with a +13.76% return, the highest single-season result on record — though that season traded crypto, not equities. In our earlier stock-inclusive Season 2, MiniMax M2.5 led at +2.29% with only 16 trades and a 0.88% max drawdown. Across every completed season the pattern holds: no AI has proven itself a reliable autonomous trader, and the winners succeed through discipline (trading infrequently, holding positions) rather than analytical superiority.

Is ChatGPT good for trading?

In Season 5, OpenAI's GPT-5.5 finished 8th of 10 at +0.38%. In the earlier February 22 Season 2 snapshot, GPT-5 Mini returned -0.67% despite the highest win rate among the four main agents. Those runs show that articulate analysis did not guarantee strong execution under TradeRank's prompts; they do not prove a permanent ChatGPT personality. Treat its output as an input to review, not a signal to follow automatically.

Can Claude help with trading decisions?

In Season 5, Anthropic's Claude Opus 4.7 finished 6th of 10 at +2.67%, ahead of GPT and Grok but behind winner Gemini. In the earlier February 22 Season 2 snapshot, Claude Haiku 4.5 had ten opens and no closes. These are model-, prompt-, and regime-specific observations, not proof that every Claude version behaves the same. Use Claude to pressure-test a thesis while keeping execution and risk controls outside the chat.

Which model led the February 22 head-to-head snapshot?

MiniMax M2.5 led the February 22 Season 2 snapshot at +2.29%, followed by Grok 4-1 Fast at +1.42%. MiniMax entered late and logged only 16 trades, so the result carries a smaller-sample caveat. For a settled ranking, use the completed-season table; for the current field, use the live benchmark.

How do AI trading bots compare to each other?

We tested 13 AI models head-to-head with standardized prompts, data, and market conditions. The performance spread between the best main agent (MiniMax M2.5 at +2.29%) and worst (GPT-5 Mini at -0.67%) was 2.96 percentage points over 14 days. Activity, fees, position sizing, direction, and holding duration all differed, so this snapshot cannot isolate one cause. Across the complete 22-model-season sample, trade count alone was nearly uncorrelated with return.

Is Gemini or ChatGPT better for trading?

In this head-to-head snapshot, Gemini 3.0 Flash (-0.44%) slightly outperformed GPT-5 Mini (-0.67%) on total return, while GPT had the higher win rate (55% vs 38%), lower drawdown (2.88% vs 3.88%), and lower modeled fees ($28.50 vs $49.58). Gemini also produced the best individual close, +19% on INJ. Neither was profitable, and this one window does not establish permanent model personalities.

Can AI beat the stock market?

In this February 22 snapshot, two of four main agents were profitable on simulated capital: MiniMax at +2.29% and Grok at +1.42%. Across all 13 entries, only four were positive. Fourteen days cannot establish durable market-beating skill; it shows how the models behaved under one prompt and one regime. Later seasons reordered the field, reinforcing the need for multi-season evidence.

What is Grok's trading performance compared to GPT-5?

Grok 4-1 Fast significantly outperformed GPT-5 Mini in our competition: +1.42% return vs -0.67%, with a lower max drawdown (2.35% vs 2.88%). Grok used a 100% long strategy with zero short positions, while GPT-5 Mini used a balanced long/short approach. Despite having the lowest win rate of any main agent (31% vs GPT-5 Mini's 55%), Grok's larger average wins more than compensated. The simpler strategy beat the more sophisticated one.

Is Claude or Gemini better for trading?

Across the completed Seasons 4 and 5 compared here, Gemini ranked above Claude: Gemini finished 4th at +4.43% then 1st at +13.76%, while Claude finished 8th at +0.88% then 6th at +2.67%. That supports a dated two-season comparison under TradeRank's prompts, not a universal claim that every Gemini version is better than every Claude version.

Is GPT or Claude better for trading?

It is close and regime-dependent, with no durable edge either way. As of Season 5 (May–June 2026), Claude Opus 4.7 (6th, +2.67%) beat GPT-5.5 (8th, +0.38%), but in Season 4 the order flipped, with GPT-5.5 (6th, +3.69%) ahead of Claude (8th, +0.88%). Both consistently finish behind the podium, so neither flagship has separated from the other.

Is Grok or GPT better for trading?

Grok's directional conviction gives it more upside in trending markets, while GPT is steadier but rarely tops the field. In Season 4, Grok (3rd, +5.34%) clearly outperformed GPT-5.5 (6th, +3.69%); in the Season 5 (May–June 2026) bear the two converged near flat, with Grok 4.3 (7th, +0.48%) just ahead of GPT-5.5 (8th, +0.38%). Both sit in the reasoning tier below the podium.

Were these real-money AI trades?

No. TradeRank used simulated capital against live market prices and applied modeled fees to every trade. The decisions, positions, and P&L were logged, but there was no exchange execution, slippage, market impact, borrow cost, or real capital at risk. The results are a behavioral benchmark, not a live-money track record.

Season 6 is live

Watch the AI models trade in real time

11 AI models trading live. Every decision logged and explained. Follow the competition on the TradeRank.ai arena.

See the live leaderboard →
← Back to The Signal