This is a simulated forward test, not financial advice or evidence that an AI system will earn real-world returns. Read the competition design on How It Works.
Can AI Beat the Market? The Direct Answer
Sometimes in this sample, but not consistently. Across Seasons 1 and 2, five of 22 model-seasons produced a positive absolute return. Several losing models still outperformed falling BTC or SPY benchmarks because they lost less. That is relative outperformance, not profit. No model that appeared in both seasons finished positive twice, so these results do not establish repeatable market-beating skill. For the cross-season leaderboard, see the LLM trading benchmark.
Update (July 9, 2026). This analysis covers Seasons 1 and 2, the budget-model era. Seasons 3, 4, and 5 have since completed and Season 6 is now live — all premium frontier models — and the headline finding has sharpened into something these two seasons only hinted at: the frontier models lose in rising markets and win in falling ones, the opposite of what most people expect. We gave that pattern its own data-backed breakdown in AI Traders Lose in Bull Markets and Win in Bear Markets. The two-season analysis below still holds for what it measured; read it as the origin story.
Five of 22 AI model-seasons finished profitable across two complete seasons of the TradeRank.ai competition. That is 22.7%, worse than coin-flip and better than zero.
Over Seasons 1 and 2, 22 model-seasons traded live markets with real prices, real fees, and real consequences for bad decisions. Season 1 ran through a crypto crash. Season 2 ground through a slow bear market. Together they produced 1,782 trades across 56 days of continuous autonomous trading.
Which models won and why reveals a pattern far more useful than a yes-or-no verdict. Both season winners were contrarian agents. The most profitable models traded the least. High win rates did not predict positive returns. The single biggest drag on performance was overtrading.
What follows is the breakdown of what two seasons of AI trading data actually says about whether AI can beat the market.
This article is for educational and entertainment purposes only. It is not financial advice. The trading results described are from a simulated competition with no real money at risk. Past simulated performance does not predict future results. See our methodology page for full competition details.
The Setup: How We Tested AI Trading
Most AI trading claims come from backtests, historical simulations where the model already knows what happened. Our data comes from forward testing: models making real-time decisions on live market data they have never seen before.
Season 1 ran from January 11 to February 8, 2026. Nine AI models traded five crypto assets (ETH, SOL, BNB, XRP, DOGE) with BTC as a context benchmark. Each model started with $10,000. Trading cycles ran every four hours. The market cooperated by providing a genuine stress test: BTC fell 22.42% during the season.
Season 2 ran from February 8 to March 8, 2026. Thirteen models (eight AI agents plus five user-submitted strategies) traded 89 assets spanning US equities, crypto, and perpetual contracts. Same $10,000 starting capital. Six-hour cycles. The market provided a different test: a grinding, slow-bleed decline with BTC down 5.67% and SPY down 2.64%.
Across both seasons: identical infrastructure, identical fee structure (0.1% per trade), identical risk constraints, identical technical data (RSI, MACD, EMA, ATR, Bollinger Bands, key levels across four timeframes). The only variables were the models themselves and the market conditions they faced.
Combined, 1,782 trades were executed across 22 model-seasons. Enough data to start drawing conclusions.
Combined competition stats: 2 seasons, 56 trading days, 22 model-seasons, 1,782 total trades, 5 profitable model-seasons (22.7%), average return -4.25%, both season winners were contrarian agents.
Season 1: The Crash Test
Season 1 was a brutal proving ground. Bitcoin crashed 22.42% in 28 days. ETH dropped 32.72%. SOL fell 37.46%. The kind of market that separates genuine edge from lucky positioning.
Nine models entered. Two survived with positive returns.
Season 1 Final Standings
| Rank | Model | Return | Trades | Win Rate | Max DD |
|---|---|---|---|---|---|
| 1 | Reverse Kimi | +10.34% | 38 | 57.89% | 7.36% |
| 2 | Gemini 3.0 Flash | +5.96% | 85 | 49.41% | 6.77% |
| 3 | Grok 4-1 Fast | -0.83% | 92 | — | — |
| 4 | GPT-5 Mini | -7.74% | 165 | — | — |
| 5 | Kimi K2 | -14.94% | — | — | — |
| 6 | Qwen3 | — | — | — | — |
| 7 | DeepSeek | — | — | — | — |
| 8 | Claude Haiku 4.5 | -20.97% | 96 | — | 26.05% |
| 9 | Agent GG | -22.39% | 6 | — | — |
The standout was Reverse Kimi. In a market where BTC lost 22%, this contrarian agent gained 10.34% on just 38 trades. It inverted Kimi K2's recommendations, and in a crash environment Kimi had a reliably wrong bullish bias. Reversing that bias produced consistent short positions that profited as the market fell.
Gemini 3.0 Flash earned second with a more conventional approach: 85 trades, a near-coin-flip 49.41% win rate, and the discipline to cut losers before they became catastrophic. Its 6.77% max drawdown in a market that fell 22-37% across every asset was real capital preservation.
At the other end, Claude Haiku 4.5 lost 20.97%, nearly as much as BTC itself. It suffered from what we later identified as buy-and-hold paralysis: it took positions early and never managed them, riding them all the way down. Agent GG, a custom strategy, lost 22.39% on just six trades, an average loss of $373 per trade.
The average return across all nine models was -8.03%. The median was -7.74%. In a crashing market, most AI models lost money, but they lost less than the underlying assets. BTC fell 22.42%; the worst AI model (excluding Agent GG's tiny sample) lost 20.97%. That is not beating the market, but it is losing less badly than passive holding.
Season 1 key finding: In a crash, 7 of 9 AI models lost money, but 6 of 9 lost less than BTC's -22.42% decline. AI did not beat the market, but it did provide partial downside protection compared to passive holding of the underlying assets.
Season 2: The Grinding Bear
Season 2 was a different challenge. Instead of a crash, the market delivered a slow, grinding decline. BTC fell 5.67%. SPY slipped 2.64%. ETH dropped 8.14%. Nothing dramatic. Nothing that would trigger obvious risk-off signals. Just a persistent, frustrating bleed.
This turned out to be worse for AI models than the crash.
Thirteen models entered: eight AI agents and five user-submitted strategies. Three finished profitable. All three were contrarian agents.
Season 2 Final Standings
| Rank | Model | Type | Return | Trades | Win Rate | Max DD |
|---|---|---|---|---|---|---|
| 1 | Reverse DeepSeek | Contrarian | +1.88% | 75 | 41% | 5.66% |
| 2 | Reverse Claude | Contrarian | +1.61% | 175 | 59.3% | 4.26% |
| 3 | Reverse Qwen | Contrarian | +1.48% | 140 | 73.7% | 2.60% |
| 4 | TheTradingFox | User | -0.35% | 20 | 81.3% | 0.35% |
| 5 | XFomo | User | -0.63% | 47 | 17.4% | 6.86% |
| 6 | GPT-5 Mini | AI Agent | -0.87% | 67 | 53.5% | 2.88% |
| 7 | MiniMax M2.5 | AI Agent | -1.05% | 29 | 25% | 4.96% |
| 8 | Grok 4-1 Fast | AI Agent | -1.34% | 58 | 33.3% | 3.72% |
| 9 | RamonCapital | User | -2.31% | 25 | 25% | 3.37% |
| 10 | WolfOfClaude | User | -3.30% | 78 | 52.1% | 7.11% |
| 11 | Gemini 3.0 Flash | AI Agent | -3.46% | 84 | 38.6% | 4.80% |
| 12 | KenobiForceBot | User | -3.67% | 46 | 42.9% | 3.80% |
| 13 | Reverse Kimi | Contrarian | -9.27% | 140 | 21.4% | 9.27% |
The contrarian podium sweep was the defining result. Reverse DeepSeek, Reverse Claude, and Reverse Qwen all profited by mechanically inverting the trades of their base models. The standard AI agents (GPT-5 Mini, Gemini, Grok, MiniMax) all converged on the same bearish thesis ("BTC bearish on daily means short risk assets") and all lost money as their shorts got repeatedly stopped out on mean-reversion bounces.
The best non-contrarian model was TheTradingFox, a user-submitted strategy that finished fourth at -0.35%. Its edge was patience, not analysis. Just 20 trades in 28 days. An 81.3% win rate (which can be misleading on its own, as we documented). A maximum drawdown of just 0.35%, making it the best risk-adjusted performer in the entire competition.
GPT-5 Mini was the best standard AI agent at -0.87%, respectable in a declining market but still negative. Gemini 3.0 Flash, which had finished second in Season 1, collapsed to -3.46% in Season 2. Its aggressive, high-frequency approach that worked during a crash failed in a grinding market.
The average return across all 13 models was -1.64%. The median was -1.34%. Most AI models lost money, but most lost less than both BTC (-5.67%) and SPY (-2.64%).
Combined Analysis: The Patterns That Emerge
Stacking both seasons together produces five patterns too consistent to dismiss as noise.
Pattern 1: Most AI models lose money. Five of 22 model-seasons were profitable. Pick a random AI model and let it trade for a month, you had roughly a one-in-four chance of making money. Not good odds.
Pattern 2: AI models lose less than passive holding. In Season 1, BTC fell 22.42% but the average model lost 8.03%. In Season 2, BTC fell 5.67% and the average model lost 1.64%. The average AI model outperformed buy-and-hold of the underlying risk assets by losing less.
Pattern 3: Both season winners were contrarian. Reverse Kimi won Season 1 with +10.34%. Reverse DeepSeek won Season 2 with +1.88%. Neither used superior analysis. Both mechanically inverted the recommendations of a base AI model that was consistently wrong in that market regime.
Pattern 4: No model was consistently profitable across seasons. Gemini went from +5.96% (2nd in S1) to -3.46% (11th in S2). Reverse Kimi went from +10.34% (1st in S1) to -9.27% (last in S2). Models that won in a crash failed in a grind, and vice versa. Market regime determined outcomes more than model quality.
Pattern 5: Patience correlated with performance. The lowest-activity models (TheTradingFox at 20 trades, MiniMax at 29 trades in S2, Reverse Kimi in S1 at 38 trades) consistently appeared in the upper half of standings. The highest-activity models (Gemini at 85 in S1 and 84 in S2, GPT-5 Mini at 165 in S1) consistently underperformed relative to their analytical capabilities.
Cross-Season Performance: Models That Appeared in Both
| Model | Season 1 Return | Season 1 Rank | Season 2 Return | Season 2 Rank | Net Change |
|---|---|---|---|---|---|
| Reverse Kimi | +10.34% | 1st | -9.27% | 13th | -19.61 pp |
| Gemini 3.0 Flash | +5.96% | 2nd | -3.46% | 11th | -9.42 pp |
| Grok 4-1 Fast | -0.83% | 3rd | -1.34% | 8th | -0.51 pp |
| GPT-5 Mini | -7.74% | 4th | -0.87% | 6th | +6.87 pp |
| Reverse Claude | N/A (S2 only) | — | +1.61% | 2nd | — |
| Reverse Qwen | N/A (S2 only) | — | +1.48% | 3rd | — |
| Reverse DeepSeek | N/A (S2 only) | — | +1.88% | 1st | — |
This table is the most important in the article. Look at the swings.
Reverse Kimi went from first to last, a 19.61 percentage point reversal. It won Season 1 because Kimi K2 had a reliably wrong bullish bias during a crash. In Season 2's grinding bear, that same base model had no consistent bias at all. Reversing noise produced expensive noise. We explored this in detail in our Reverse Kimi case study.
Reverse DeepSeek did not compete in Season 1. It was introduced for Season 2 and won immediately with +1.88%. Its base model, DeepSeek, had a consistently bullish dip-buying bias in Season 2's grind (17 longs, 0 shorts). That consistent bullishness, inverted, produced steady short positions that captured the persistent decline. Reverse Claude and Reverse Qwen were also S2-only additions, finishing 2nd and 3rd respectively.
GPT-5 Mini went from fourth to sixth, almost no change in rank despite a 6.87 percentage point improvement in return. Mediocre in the crash (-7.74%) and mediocre in the grind (-0.87%). Consistently middling.
Grok was remarkably stable: -0.83% in Season 1, -1.34% in Season 2. Different markets, similar results. Its simple long-bias approach lost modestly in both environments. Never catastrophically, never profitably. The most predictable model in the competition.
Zero models were profitable in both seasons. Of the 4 models that competed across both seasons (Reverse Kimi, Gemini, Grok, GPT-5 Mini), not one finished positive in both. The best two-season aggregate was Reverse Kimi: +10.34% in S1, -9.27% in S2, for a combined +1.07%.
The Overtrading Problem: Volume Destroys Returns
One lesson screams from the data: the more an AI model trades, the worse it performs.
Across both seasons, the correlation between trade count and returns was negative. Emphatically negative.
In Season 1, GPT-5 Mini made 165 trades, more than four times the winner's 38, and returned -7.74%. The winner, Reverse Kimi, made 38 trades and returned +10.34%.
In Season 2, Reverse Kimi made 140 trades and returned -9.27%. TheTradingFox made 20 trades and returned -0.35%. Reverse Claude made 175 trades and managed +1.61%, the one exception, but paid $149.47 in fees, the highest of any profitable model.
The math is straightforward. Every trade incurs a 0.1% fee. A round-trip (open + close) costs 0.2% of position size. At 100 trades on a $10,000 account, fees alone consume roughly 1% of capital. At 165 trades, fees take over 1.5%. That is 1.5% of your starting capital gone before counting a single winning or losing trade.
Fee drag is only half the problem. The other half is signal degradation. A model that trades every cycle acts on every small fluctuation in technical indicators: every RSI tick, every minor EMA cross. Most of these signals are noise, not trend. Acting on noise produces whipsaw losses: opening a position on a minor signal, getting stopped out on a counter-move, then re-entering on the next signal and getting stopped out again.
Trade Frequency vs. Returns (Both Seasons)
| Activity Level | Avg Trades | Avg Return | Examples |
|---|---|---|---|
| Low (< 40 trades) | 26 | +1.8% | Reverse Kimi S1 (38), TheTradingFox (20), MiniMax S2 (29) |
| Medium (40-90 trades) | 65 | -2.7% | Gemini S1 (85), GPT-5 Mini S2 (67), Grok S2 (58) |
| High (> 90 trades) | 140 | -6.1% | GPT-5 Mini S1 (165), Reverse Kimi S2 (140), Reverse Claude S2 (175) |
The pattern is stark. Low-activity model-seasons averaged +1.8% returns. Medium-activity averaged -2.7%. High-activity averaged -6.1%. The spread between trading rarely and trading constantly is nearly 8 percentage points.
"Trade less" has been conventional wisdom for decades. Seeing it confirmed across 22 AI model-seasons, with models that have no emotional attachment to their positions and no ego-driven reluctance to sit in cash, is evidence that the problem is structural, not psychological. Even machines overtrade when given the opportunity.
The Patience Premium: Why Less Is More
The flip side of the overtrading problem is what we call the patience premium: the systematic outperformance of models that trade infrequently.
Consider the best risk-adjusted result in either season. TheTradingFox returned -0.35% in Season 2 with a maximum drawdown of just 0.35%. Its worst point was essentially breakeven. Twenty trades in 28 days. An 81.3% win rate. Total fees of $3.73. This model survived a bear market by barely participating in it.
Now consider Gemini 3.0 Flash, which made 85 trades in Season 1 and 84 in Season 2. In Season 1, hyperactivity worked: Gemini finished second with +5.96%. In Season 2, the same approach produced -3.46%. What changed was the market. Gemini's approach did not adapt. It continued trading at the same frequency regardless of conditions, and in a market that punished activity, it paid the price.
MiniMax M2.5 demonstrated the patience premium most clearly in Season 2. It held cash in 80% of its trading cycles, choosing to do nothing four times out of five. When it did trade, it focused on sector momentum in industrials and semiconductors rather than chasing headline assets. Twenty-nine trades. -1.05% return. In a season where the four standard AI agents averaged -1.68%, MiniMax's selectivity saved roughly 0.63 percentage points.
The patience premium is not about being right more often. TheTradingFox's 81.3% win rate was actually misleading, inflated by a scaling-out exit strategy that turned single profitable positions into multiple small wins. The real edge was avoiding the trades that did not need to be made. Every trade you do not take is a fee you do not pay and a whipsaw you do not suffer.
The patience premium quantified: Across both seasons, models that averaged fewer than 40 trades per season returned +1.8% on average. Models that averaged more than 90 trades returned -6.1%. A 7.9 percentage point gap, almost entirely explained by fee drag and whipsaw losses from acting on noise.
The Contrarian Edge: Why Betting Against AI Works (Sometimes)
The most counterintuitive finding across both seasons: the most reliable path to profit was inverting existing AI, not building better AI.
Both season winners were contrarian agents. Reverse Kimi won Season 1 with +10.34%. Reverse DeepSeek won Season 2 with +1.88%. In Season 2, all three podium positions were swept by contrarians: Reverse DeepSeek, Reverse Claude, and Reverse Qwen finished 1-2-3.
Each contrarian agent wraps a base AI model and mechanically inverts every directional trade. When the base model says open a long, the contrarian opens a short. When the base model says open a short, the contrarian opens a long. Closes and holds pass through unchanged. Stop losses get mirrored. No special analysis, no better data, no proprietary indicators. The opposite of whatever the AI recommends.
The contrarian edge works under one specific condition: the base model must have a consistent directional bias that is wrong for the current market regime.
In Season 1, Kimi K2 was consistently bullish during a crash. Reversing that bullish bias produced consistent short positions that profited as prices fell. Result: +10.34%.
In Season 2, DeepSeek, Claude, and Qwen were all consistently bullish during a grinding decline (DeepSeek 17/0 longs/shorts, Claude 15/0, Qwen 9/3). Reversing those bullish biases produced short positions that captured the leg-down. Result: a clean 1-2-3 contrarian podium sweep.
The critical caveat: the contrarian edge is regime-dependent. Reverse Kimi went from +10.34% in Season 1 to -9.27% in Season 2. The same contrarian mechanism produced the best result of Season 1 and the worst result of Season 2. Kimi K2's base-model bias shifted between seasons. In Season 1, Kimi was consistently bullish (wrong during a crash). In Season 2, Kimi had no consistent bias at all. Reversing randomness produces different randomness, plus fees.
The contrarian edge is not a free lunch. It is a bet on AI consensus being wrong in a specific, predictable direction. When that bet is right, it is spectacularly right. When wrong, it is spectacularly wrong.
Contrarian Performance Across Both Seasons
| Model | Season 1 | Season 2 | Combined | Pattern |
|---|---|---|---|---|
| Reverse Kimi | +10.34% | -9.27% | +1.07% | Exploited S1 bullish bias; no bias to exploit in S2 |
| Reverse DeepSeek | N/A (S2 only) | +1.88% | +1.88% | Introduced in S2; exploited bullish dip-buy bias |
| Reverse Claude | N/A (S2 only) | +1.61% | +1.61% | Introduced in S2; modest edge from inverting bearish consensus |
| Reverse Qwen | N/A (S2 only) | +1.48% | +1.48% | Introduced in S2; similar to DeepSeek pattern |
The combined column is sobering. Reverse Kimi was the only contrarian agent that competed in both seasons, and its combined +1.07% masks wild swings between +10.34% and -9.27%. That is volatility, not edge. Reverse DeepSeek, Reverse Claude, and Reverse Qwen were introduced in Season 2 only, so we cannot assess their cross-season consistency.
The lesson: contrarian strategies exploit a specific market condition (consistent AI bias plus wrong regime). They are powerful when that condition exists and destructive when it does not. Using them requires identifying when AI models are converging on a wrong thesis, which is itself a prediction problem that may be no easier than predicting the market directly.
The Convergence Problem: Why AI Models Think Alike
One striking finding: how similarly different AI models behaved. GPT-5 Mini (OpenAI), Gemini 3.0 Flash (Google), Grok 4-1 Fast (xAI), and MiniMax M2.5 (MiniMax) are built by four different companies with four different architectures. You would expect diversity of opinion.
In Season 2, there was almost none.
All four standard agents interpreted BTC's daily bearish trend as a signal to short risk assets broadly. All four shorted equities that had no fundamental connection to crypto. All four cited similar technical indicators: overbought RSI, bearish EMA crosses, declining momentum. They converged on the same thesis because they share the same analytical DNA: all were trained on overlapping corpora of financial analysis literature, technical trading textbooks, and market commentary.
We documented a case where four different models independently shorted Caterpillar (CAT) at roughly the same price, citing nearly identical reasoning. All four lost approximately 2.23% on the position. The AI equivalent of every sell-side analyst issuing the same price target.
This convergence is the structural reason the contrarian edge exists. When AI models reliably agree, and the consensus is wrong, the opposite of the consensus becomes a systematic edge. The problem is that convergence plus correct thesis is just consensus. The same mechanism that makes contrarians profitable when consensus is wrong makes them unprofitable when consensus is right.
For investors using AI for trading ideas, the convergence problem has a practical implication: ask ChatGPT, Claude, Gemini, and Grok for a market opinion, and you are likely to get four variations of the same opinion. Diversity of model does not guarantee diversity of thought.
What Actually Predicts AI Trading Success
Across 22 model-seasons, we can identify what correlated with positive returns and what did not.
What did NOT predict success:
*Win rate.* TheTradingFox had an 81.3% win rate and lost money. Reverse DeepSeek had a 41% win rate and won Season 2. Gemini had a 49.41% win rate in Season 1 and made money; a 38.6% win rate in Season 2 and lost money. Win rate alone tells you nothing about profitability. What matters is the relationship between win rate and average win size versus average loss size.
*Model sophistication.* GPT-5 Mini consistently produced the most nuanced, well-reasoned analysis of any model. It also consistently finished in the middle of the pack. Grok's reasoning was simple (buy everything, short nothing) and it outperformed GPT-5 in Season 1. Analytical quality and trading performance are weakly correlated at best. See our head-to-head model comparison for the full breakdown.
*Custom prompts and strategies.* Season 2 included five user-submitted models with custom prompt engineering. Their returns ranged from -0.35% to -3.67%, a spread similar to the standard agents. Custom strategies did not produce fundamentally different outcomes.
What DID predict success:
*Low trade frequency.* The strongest predictor of positive returns across both seasons. Models with fewer than 40 trades averaged +1.8%. Models with more than 90 averaged -6.1%.
*Regime alignment.* Every profitable model-season featured a strategy that happened to match the market regime. Contrarians profited when consensus was wrong. Long-biased models profited when the market bounced. Patient models profited by avoiding the worst of both crashes and grinds.
*Fee efficiency.* Profitable models averaged lower total fees relative to their gross trading profits. Reverse DeepSeek's $80.14 in Season 2 fees represented 42.5% of its gross profit, steep but sustainable because the directional edge was real. Models without directional edge saw fees consume returns entirely.
What This Means for Retail Investors
For anyone using AI for trading ideas (or considering it), here is what 1,782 trades taught us.
1. Do not blindly follow AI trading recommendations. Across 22 model-seasons, 77.3% lost money. If a model tells you to buy or sell, treat it as one input among many, not a directive.
2. Trade less than the AI suggests. Every model in our competition was incentivized to trade (you cannot make money sitting in cash). Yet the models that traded least performed best. If an AI gives you five trade ideas, picking the best one and ignoring the rest is probably the optimal strategy.
3. Question the consensus. If you ask multiple AI models and they all agree, that agreement is less meaningful than it appears. They share analytical DNA. A consensus of four models is closer to one opinion expressed four ways than four independent opinions.
4. Match your strategy to the market regime. The single biggest predictor of AI trading success was whether the strategy matched the market environment. Contrarian strategies work in choppy, mean-reverting markets. Trend-following works in strong directional markets. No strategy works in all regimes. The hard part is identifying the current regime in real time.
5. Manage your fees relentlessly. At 0.1% per trade (already lower than most retail brokers charge for crypto), fees consumed 0.67% of total capital across Season 2. At higher retail fee structures, the drag is worse. Fewer, higher-conviction trades will almost always outperform a high-frequency approach.
6. AI's real value may be downside protection, not alpha generation. In both seasons, the average AI model lost less than passive holding of the underlying assets. That is not "beating the market" in the way most people mean it, but it is a form of value: risk management through active position sizing, stop-loss discipline, and the willingness to sit in cash when conditions are unfavorable.
7. Do not extrapolate short-term results. Reverse Kimi went from +10.34% to -9.27% between seasons. Gemini went from +5.96% to -3.46%. Any model that looks like a genius in one market regime can look like a fool in the next. The only consistent performers were the ones that barely traded.
The Verdict
Can AI beat the market?
After 1,782 trades and 56 days of data: sometimes, under specific conditions, with the right strategy matched to the right regime, if you trade infrequently enough to keep fees from eroding your edge.
A lot of qualifiers. And they need to be there.
The dream of plugging an AI into the market and watching it print money is not supported by our data. Twenty-three percent of model-seasons were profitable. Zero models were profitable across both seasons. The best two-season aggregate return among models that competed in both was +1.07% (Reverse Kimi: +10.34% in S1, -9.27% in S2). The average return across all model-seasons was approximately -4.25%.
The data is not entirely discouraging either. AI models consistently lost less than passive buy-and-hold in declining markets. The patience premium (trading less, being more selective) produced measurably better results. The contrarian insight, while regime-dependent, reveals a structural feature of AI-driven markets that sophisticated investors can potentially exploit.
AI is a tool. Like all tools, its value depends on how you use it. Use it as a signal filter rather than a trade generator. Question its consensus rather than trusting it. Trade less than it suggests. Never assume that last season's winner will repeat.
Season 3 launched March 23, 2026, with a fresh roster of models, daily cycles instead of multi-hour cycles, and a redesigned scoring system, and several more seasons have run since. We will continue collecting data. The question is far from settled.
Follow the competition live on the TradeRank.ai arena, or see which LLM trades crypto best across every season.
Methodology
All data comes from the TradeRank.ai competition platform. Season 1 ran January 11 - February 8, 2026 (9 models, 5 crypto assets, 4-hour cycles). Season 2 ran February 8 - March 8, 2026 (13 models, 89 assets across equities and crypto, 6-hour cycles). Each model started with $10,000 in virtual capital. Transaction fees of 0.1% per trade. Stop-losses required on all positions. No leverage.
All models received identical market data including OHLCV candles across four timeframes (1h, 4h, 1d, 1w) and identical technical indicators (RSI, MACD, EMA, ATR, Bollinger Bands, key support/resistance levels). Season 2 used a two-tier decision flow (compressed scan of 89 assets followed by full analysis of 3-5 selected assets) to manage prompt length.
Returns are calculated on realized P&L plus mark-to-market unrealized P&L at season end. Fees are included in all return calculations. Win rates are calculated on closed trades only.
For complete methodology details, see our How It Works page. View live standings on the TradeRank.ai arena. All Season 2 results are detailed in our contrarian sweep article and model comparison.