A 17% Win Rate Led 13 AI Trading Strategies — the Day-14 Results

At the Season 2 midpoint, XFomo led 13 strategies despite winning only 17% of its closed trades. The final leaderboard told a different story.

+5.50%Sharpe 0.42656 trades
Data Point

This is a Season 2 retrospective (February 2026), not the current leaderboard. Follow the live AI trading competition or browse the latest results for daily and weekly updates, or compare models on the always-live Best LLM for Crypto Trading benchmark.

On February 8, 2026, we launched our largest AI trading experiment to that point: thirteen strategies, $10,000 each, trading autonomously against real markets. They shared the same 89-asset market feed, fee model, and execution infrastructure, but they did not all use the same prompt or decision mechanism: four were standard agents, four inverted base-model decisions, and five used community-designed prompts.

Two weeks and 56 decision cycles later, the snapshot contradicted almost everything you'd assume about AI trading. The strategy in first place had won only 17% of its closed trades. The one with a 90% win rate was almost flat. And one of the most profitable approaches at that moment was literally doing the opposite of what DeepSeek recommended.

This is the Day-14 record, not the completed-season verdict. Season 2 continued for another two weeks and Reverse DeepSeek ultimately won at +1.88%, while snapshot leader XFomo finished fifth at -0.63%.

Warning

This article is for educational and entertainment purposes only. It is not financial advice. The trading results described are from a simulated competition with no real money at risk. Past simulated performance does not predict future results.

Data Point

Data as of: Day 14, February 22, 2026. Updated results available in The Win Rate Paradox (Day 25 data).

The Setup: 13 Models, $10K Each, No Human Intervention

TradeRank.ai is an AI trading arena built to compare autonomous portfolio decisions on simulated capital.

Season 2 launched on February 8, 2026. Each model received a $10,000 simulated account and access to the same universe of 89 assets: 49 equities, 21 Binance crypto pairs, and 17 Hyperliquid perpetuals, plus BTC and SPY as benchmarks.

Roughly every six hours, agents received a compressed market scan and detailed OHLCV data for selected assets, then returned open, close, stop-modification, or hold decisions. The shared environment used live market prices, a modeled 0.1% fee, and a ten-position cap. Built-in, reverse, and user agents did not have identical decision mechanisms: reverse entries were transformed and user agents supplied custom prompts.

Season metadata and prompts declared stops required on new positions, but archived decisions show that execution did not enforce the rule consistently. Read this as a shared market/account comparison, not a perfectly controlled identical-rules trial. The completed record is preserved in the Season 2 reports.

The Competitors

We fielded three categories of strategy:

Standard agents (4): GPT-5 Mini (OpenAI), Gemini 3.0 Flash (Google), Grok 4-1 Fast (xAI), and MiniMax M2.5. They shared the standard prompt and made independent decisions.

Reverse agents (4): adapters that inverted directional recommendations from Claude, DeepSeek, Qwen3, or Kimi.

User strategies (5): community-designed prompts with different trading philosophies, including sentiment, EMA structure, coaching, forced activity, and confirmation rules.

The 13 strategies were ranked by marked equity relative to their initial $10,000. Current model pages and standings are available through the live arena.

Season 2 Standings (Day 14, Feb 22, 2026)

RankModelTypeEquityReturnWin RateTrades
1XFomoUser model$10,550+5.50%17%27
2Reverse DeepSeekReverse agent$10,264+2.64%35%65
3MiniMax M2.5AI agent$10,230+2.29%33%16
4WolfOfClaudeUser model$10,150+1.50%52%48
5Grok 4-1 FastAI agent$10,142+1.42%31%37
6Reverse Qwen3Reverse agent$10,118+1.18%18%40
7TheTradingFoxUser model$9,996-0.04%90%13
8Gemini 3.0 FlashAI agent$9,956-0.44%38%53
9GPT-5 MiniAI agent$9,933-0.67%55%38
10RamonCapitalUser model$9,731-2.69%25%25
11Reverse ClaudeReverse agent$9,730-2.70%15%99
12KenobiForceBotUser model$9,699-3.01%46%44
13Reverse KimiReverse agent$9,330-6.70%17%100
Key Insight

The Day-14 spread was 12.20 percentage points, from +5.50% to -6.70%. All 13 strategies used the same market data, fees, starting capital, and execution engine. Their prompts and decision layers intentionally differed, so this compared complete agent strategies rather than isolating the underlying LLM as the only variable.

Surprise #1: Win Rate Is Overrated

If you had to bet on one model before the competition started, you'd probably pick the one that wins most often. That intuition is wrong.

XFomo sat in first place with a 17% win rate. Out of 27 trades, only about 5 were profitable. But those 5 winners were enormous. Its single best trade, a long position on SNX (Synthetix), returned $286 on a roughly $1,000 position. That one trade alone covered every losing trade and then some.

Contrast that with TheTradingFox, which posted a 90% win rate on 13 trades. Sounds incredible until you check the return: -0.04%. TheTradingFox won almost every trade but won small. Its tight stop-losses protect against big losses, but they also prevent the kind of outsized gains that actually move the needle on a portfolio.

This is not a new insight in trading. Trend followers have known for decades that a low win rate with fat tails beats a high win rate with thin tails, a principle well-documented in systematic trend-following research. LLM-driven strategies reproduce this same dynamic without anyone explicitly programming it. XFomo's system prompt emphasizes momentum and sentiment, not position sizing theory. The fat-tail behavior emerged from how the model interprets technical data.

Data Point

XFomo: 17% win rate, +5.50% return, 27 trades. TheTradingFox: 90% win rate, -0.04% return, 13 trades. Winning often and winning big are not the same thing.

Surprise #2: Doing the Opposite of DeepSeek Works

One of our more provocative experiments was the "reverse agent" concept. We took four AI models (Claude, DeepSeek, Qwen3, and Kimi) and created contrarian copies. The adapter flips open-long to open-short and open-short to open-long; close, hold, and add actions pass through rather than being inverted.

The hypothesis: if a model's entry direction is consistently wrong, flipping those entries could improve the result.

The results split down the middle. Two inversions were profitable at this Day-14 snapshot. Two were losing.

Reverse DeepSeek finished at +2.64%. The frozen base DeepSeek R1 was -3.39%, a 6.03-point spread in this window. Reverse Qwen3 was also positive at +1.18%.

Reverse Claude lost 2.70%, and Reverse Kimi lost 6.70%. The split shows that entry inversion is not mechanically profitable and does not establish stable model error personalities. Different position paths after the first flipped entry also prevent a clean per-decision counterfactual.

[REVERSED] Locking in 20.15% profit as 4h RSI at 11.9 indicates extreme oversold conditions and high probability of bounce.

Reverse DeepSeekBest trade. The base DeepSeek model wanted to buy the bounce; the reverse agent shorted instead. The short was right — the asset continued falling.

Surprise #3: Activity Carried a Cost

Trade frequency and return had a moderate negative relationship in this Day-14 snapshot (Pearson correlation about -0.56), useful as a hypothesis but far from a law.

MiniMax M2.5 joined on February 14 and made 16 trades over roughly eight active days. It chose to hold approximately 80% of the time, and ranked third at +2.29% at this cutoff.

Reverse Kimi made 100 trades and sat last at -6.70%, with roughly $110 in fees. Reverse Claude was an important counterexample: it made 99 trades and lost much less.

The evidence supported a narrower claim: activity created fee drag and more opportunities to be wrong, while low activity helped some strategies preserve capital. It did not prove that trading less automatically produces a higher return. The completed-season analysis in Can AI Trading Bots Beat the Market? tests the pattern over a longer window.

Warning

Fee drag is real. Reverse Kimi paid an estimated $110 in trading fees on a $10,000 account (100 trades at 0.1% each). That's 1.1% of the portfolio lost to friction alone.

Act I: False Dawn (Feb 8-11)

The competition opened on February 8 with all 13 models at $10,000. By Day 2, the leaderboard already had clear frontrunners, and they were all the flagship AI agents.

Grok 4-1 Fast, GPT-5 Mini, and Gemini 3.0 Flash surged early. Grok caught a crypto rally and GPT-5 jumped on tech momentum. The traditional AI models looked dominant. The user-submitted models and reverse agents were still feeling out the market, making tentative initial trades or holding entirely.

This early performance was deceptive. The models that deployed first and traded aggressively in the opening cycles benefited from a brief upward market move. They looked brilliant. But they were also building positions that would soon turn against them.

Act II: The Market Turns (Feb 11-18)

Around February 11, the market shifted. Crypto gave back its early-week gains. Tech stocks stalled. And the models that had been riding momentum suddenly found themselves on the wrong side of every trade.

February 13 was the universal pain day. Every single model in the competition either lost money or was flat. Grok, which had been near the top, saw its gains evaporate. GPT-5 Mini slipped below breakeven and stayed there. Gemini 3.0 Flash began a slow bleed that would continue for the rest of the competition.

The flagship AI agents (the ones most people would have bet on) peaked on Day 2 and spent the next ten days declining. This is first-mover disadvantage: the models that jumped in earliest built the biggest positions during a brief rally, then held those positions as the market turned. Their confidence in their initial thesis prevented them from cutting losses quickly.

Meanwhile, MiniMax M2.5 was doing almost nothing. It held through the volatility, made a handful of selective trades, and quietly climbed into third place. Its restraint during the messy middle period of the competition is what separated it from the pack.

Act III: The Late Rally (Feb 19-21)

The final stretch of the competition saw a dramatic reshuffling. User-submitted models, which had been cautious early, began making their moves.

The defining moment came on February 19-20 when XFomo opened a long position on SNX (Synthetix). The token had been building momentum, and XFomo's sentiment-focused system prompt identified the setup. The position ran to +33% unrealized gains before XFomo made the call to close it.

24h -3.2% drop despite +33% unrealized PnL; bearish 1d trend and BTC correlation indicate sentiment switch to negative; exit fully to protect profits.

XFomoXFomo's reasoning for closing its SNX position. The $286 gain vaulted it into the Day-14 lead, but XFomo later finished the season fifth and negative overall.

That $286 gain on SNX was the largest single-trade profit of the entire competition. It vaulted XFomo from the middle of the pack to first place. XFomo's 17% win rate doesn't matter when it wins this big.

WolfOfClaude also surged in the final days, climbing to fourth place at +1.50%. Its swing-trading approach (holding positions for multiple cycles and using wider stop-losses) let it capture multi-day moves that the more active models were churning through.

By the time Cycle 56 concluded, the top four positions were held by two user models and two non-flagship agents. GPT-5 Mini finished ninth. Gemini 3.0 Flash finished eighth. The models with the biggest brand names and the most hype finished in the bottom half of the standings.

The Reverse Agent Experiment: A Deeper Look

As we covered in Surprise #2, the reverse agent results split cleanly: two profitable inversions (DeepSeek, Qwen3) and two destructive ones (Claude, Kimi). The table below summarizes the numbers. We've published a dedicated deep-dive with the full analysis: What Happens When You Reverse Every AI Trading Decision?.

Whether a reverse agent works depends on the base model's error structure. If the base model has a consistent directional bias (DeepSeek calling reversals too early), inversion is profitable. If the base model is roughly calibrated (Claude, Kimi), inversion just adds noise and fees.

Reverse Agent Performance Summary

Reverse AgentReturnWin RateTradesBase Model Bias
Reverse DeepSeek+2.64%35%65Premature reversal calls
Reverse Qwen3+1.18%18%40Bullish lean
Reverse Claude-2.70%15%99Well-calibrated (inversion hurts)
Reverse Kimi-6.70%17%100Random/balanced calls (inversion adds noise)

What the AI Models Actually Say

Every trade comes with reasoning. Each model explains, in natural language, why it's making its decision. Reading through hundreds of these reasoning logs reveals distinct personalities.

XFomo is terse and conviction-driven. When it sees a setup, it commits. When the setup deteriorates, it exits without hesitation. Its reasoning on the SNX close was disciplined profit-taking: recognizing that a 33% unrealized gain in a deteriorating trend environment is a gift, not a floor.

MiniMax M2.5 reads like a cautious institutional analyst. When it finally breaks its silence to make a trade, the reasoning is thorough and measured.

NVDA presents the best technical opportunity among exploration picks. Strong bullish trend across 1h/4h/1d timeframes, price above all key EMAs, RSI in healthy territory.

MiniMax M2.5MiniMax selecting NVIDIA from 89 assets during its scan phase. After 56 cycles of mostly holding, it identified a high-conviction setup and acted.

The reverse agents produce accidentally poetic reasoning. Their explanations are written by the base model but their actions are inverted, so you get contradictions like Reverse DeepSeek explaining why it's "locking in profit" on a trade that the base model never would have taken. A model giving the right reasons for the wrong trade, and the wrong trade turns out to be the right one.

Six Things We Learned About AI Trading

1. Patience coincided with a stronger snapshot result. MiniMax held about 80% of the time and ranked third at this cutoff. Its later join date and small sample limit the comparison.

2. Win rate did not rank the field. XFomo led at this cutoff with a 17% closed-record win rate while TheTradingFox was nearly flat despite winning most closed trades. Payoff size and open marks mattered.

3. Fees were visible drag. Reverse Kimi paid roughly $110 in modeled fees on a $10,000 account, 1.1% of starting capital.

4. Early deployment was vulnerable in this window. Several agents that deployed aggressively near Day 1 peaked early and then declined, while later or more selective entries initially held up better.

5. Custom prompts produced distinct behavior, not proven alpha. XFomo and WolfOfClaude ranked highly at Day 14, but the completed season put XFomo fifth and negative, with no user strategy positive.

6. Decision policy deserves separate measurement from model brand. Agents receiving the same inputs produced different position paths, but this snapshot cannot isolate architecture, prompt, timing, and sampling effects.

Methodology Notes

Transparency matters. Here's what you should know about how this experiment works.

Simulated execution, real data. Trades are simulated against real-time market data from Binance, Yahoo Finance, and Hyperliquid. There is no order book simulation. We assume trades execute at the current price with a 0.1% fee. This means we don't capture slippage or liquidity effects, which would matter at larger position sizes.

Two-tier data flow. With 89 assets, feeding full OHLCV data for everything would produce a ~400,000-token prompt. Instead, each model first receives a compressed technical scan of all 89 assets, selects 3-5 to investigate further, then receives full candle data for those picks. This keeps prompts manageable and forces the models to make a screening decision before a trading decision.

Identical prompts for built-in agents. GPT-5 Mini, Gemini 3.0 Flash, Grok 4-1 Fast, and MiniMax M2.5 all receive exactly the same system prompt. User-submitted models have custom system prompts designed by their creators. Reverse agents use the same prompt as their base model, with inversion applied after the response.

6-hour decision cycles. Each model makes decisions every 6 hours. This is slower than most algorithmic trading but appropriate for LLM-based strategies that work better on higher timeframes. Models can hold positions across multiple cycles.

No lookahead bias. Models receive only data available at the time of the decision. Historical candles are fetched up to the current time, never into the future.

Open competition. Anyone can submit a custom model with their own system prompt. The arena is live at traderank.ai.

Can AI Actually Trade?

After 56 cycles and thousands of individual decisions, the answer depended on what you meant.

The agents could read technical indicators and produce recognizable trading rationales. Profitability was less settled. At this Day-14 snapshot, six of thirteen strategies were positive and a seventh was nearly flat; the range ran from +5.50% to -6.70%. By the completed Season 2 close, only three strategies remained positive. The final benchmark analysis separates those two records.

Raw return did not tell the full story. XFomo's +5.50% snapshot gain depended heavily on a few positions, while MiniMax M2.5's +2.29% came with less activity. A proper comparison needs return, drawdown, fee drag, and concentration—not just leaderboard rank.

Two weeks and 27 XFomo trades were too few observations for durable claims. The snapshot supported hypotheses about position sizing, restraint, and fees; it did not prove universal laws about AI trading.

The experiment also compared whole agent strategies, not isolated foundation models. Custom prompts, model providers, and reverse adapters changed how identical market inputs became actions. The strongest conclusion was therefore narrow: decision policy and risk behavior mattered at least as much as model brand in this run.

Further Reading

This article covers the full 13-model competition, but we've also published deeper analyses of specific angles:

What Happened After the Day-14 Snapshot

Season 2 continued for another two weeks. XFomo's snapshot lead disappeared: it closed the season fifth at -0.63%, while Reverse DeepSeek rose from second at the midpoint to first at +1.88%. Only three strategies finished positive, all contrarian agents.

That reversal is why this page preserves its February 22 snapshot language instead of silently presenting midpoint numbers as final results. Read the Season 2 final analysis, compare the broader LLM trading benchmark, or check the current competition for the latest season rather than treating these historical standings as live.

Frequently Asked Questions

Who led the AI trading competition after 14 days?

XFomo led the February 22, 2026 Day-14 snapshot at +5.50%. That was not the completed result: XFomo later finished fifth at -0.63%, and Reverse DeepSeek won Season 2 at +1.88%.

How did a 17% win rate lead the Day-14 table?

XFomo won only about five of 27 closed trades, but its few winners were much larger than its losses. The result demonstrates positive payoff asymmetry in this snapshot, not that a 17% win rate is generally desirable.

Did all 13 AI trading strategies use the same prompt?

No. Four standard built-in agents shared a prompt structure, four reverse agents inverted base-model decisions, and five community agents used custom prompts. They shared data, capital, fees, and execution infrastructure, so this was a comparison of complete agent strategies rather than a controlled test of model brand alone.

Did the Day-14 leader win Season 2?

No. XFomo led at Day 14 with +5.50% but finished fifth at -0.63%. Reverse DeepSeek, second at the snapshot, won the completed season at +1.88%.

What did trade frequency show in the Day-14 snapshot?

Trade count and return had a moderate negative Pearson correlation of about -0.56. Fees and extra opportunities to be wrong made activity costly, but counterexamples mean the snapshot did not prove that fewer trades automatically produce better returns.

What did GPT-5 Mini do in the Day-14 snapshot?

GPT-5 Mini ranked ninth at -0.67% with a 55% win rate. That combination reinforced the article's central point: win rate without favorable payoff size does not determine portfolio return.

Season 6 is live

Watch the AI models trade in real time

11 AI models trading live. Every decision logged and explained. Follow the competition on the TradeRank.ai arena.

See the live leaderboard →
← Back to The Signal