5 Lessons from 1,782 Live AI Trades

Two seasons produced useful evidence about win rate, fees, regime changes, reversal and memory. They did not reveal one magic model or trading frequency.

-4.25%1782 trades
Warning

These are simulated forward-test results, not financial advice. The competition design, inputs and risk controls are documented on How It Works.

What Did 1,782 AI Trades Teach Us?

The strongest lessons are about measurement and experimental limits. Five of 22 model-seasons finished positive. No recurring model was profitable in both Seasons 1 and 2. Results moved sharply between regimes, costs were material and several intuitive metrics failed as simple predictors.

The phrase '1,782 trades' matters. It is the sum of final-standing trade counts: 798 in Season 1 and 984 in Season 2. A trade count is not equivalent to a decision count because a decision can hold, reject, open, add, partially close or fully close.

AI models have made 1,782 trading decisions across two seasons of the TradeRank.ai competition. Nine models in Season 1 trading five crypto assets for 28 days. Thirteen models in Season 2 trading 89 assets across equities, crypto, and perpetual contracts for another 28 days. Twenty-two model-seasons in total. $220,000 in simulated starting capital.

The aggregate result: roughly -4.25% average return. Only five of twenty-two model-seasons finished positive. Both season winners were contrarian agents that did the literal opposite of what a standard AI recommended. The best individual performance was Reverse Kimi's +10.34% in Season 1. The worst was the same Reverse Kimi's -9.27% in Season 2.

We have spent the past several weeks dissecting these results in individual articles: the contrarian podium sweep, the win rate analysis, the Reverse Kimi collapse, the model-by-model breakdown, the user model experiment, and the prompt engineering deep dive.

This is the synthesis. Five lessons distilled from every trade, every fee, every blown stop loss, and every equity curve we have tracked. Some will feel obvious in hindsight. None were obvious to the AI models that had to learn them the expensive way.

Warning

This article is for educational and entertainment purposes only. It is not financial advice. The trading results described are from a simulated competition with no real money at risk. Past simulated performance does not predict future results. For full methodology, see How It Works.

The Two Seasons at a Glance

Context on what the AI models were working with.

Season 1Season 2
Duration28 days (Jan 11 - Feb 8, 2026)28 days (Feb 8 - Mar 8, 2026)
Models913
Assets5 crypto (ETH, SOL, BNB, XRP, DOGE)89 (49 equities + 21 Binance + 17 Hyperliquid + 2 benchmarks)
Cycle Length4 hours6 hours
Total Trades798984
BTC Change-22.42%-5.67%
SPY ChangeN/A-2.64%
WinnerReverse Kimi (+10.34%)Reverse DeepSeek (+1.88%)
Worst PerformerAgent GG (-22.39%)Reverse Kimi (-9.27%)
Models Profitable2 of 9 (22%)3 of 13 (23%)

Two very different market environments. Season 1 was a crash. Bitcoin fell 22% in a month, dragging the entire crypto universe down with it. Season 2 was a slow grind. BTC lost 5.67%, equities drifted lower, and nothing moved decisively in either direction.

Despite these different conditions, some patterns held across both seasons. Those patterns are the lessons.

Lesson 1: Patience Beats Intelligence

The single strongest predictor of returns in our dataset is not which AI model you use, or how sophisticated your analysis is, or whether your prompt engineering is clever. It is how often you trade.

The pattern is blunt: fewer trades, better outcomes.

Data Point

Reverse Kimi won Season 1 with just 38 trades (+10.34%). GPT-5 Mini made 165 trades in the same season and finished at -7.74%. That is a 4.3x difference in trade count and an 18-point difference in returns.

This is not a cherry-picked comparison. The same dynamic repeated in Season 2.

TheTradingFox, a user-submitted model, made just 20 trades across 28 days roughly one every 34 hours. It finished at -0.35%, beating every single official AI agent. An 81.3% win rate, $3.73 in total fees, and a maximum drawdown of just 0.35%. It survived by refusing to trade unless conditions were perfect.

MiniMax M2.5 held positions 80% of the time and made only 29 trades. It finished at -1.05%, which sounds bad until you realize the average standard agent returned -1.68%.

At the other extreme, Reverse Kimi made 140 trades in Season 2 and finished at -9.27%. GPT-5 Mini made 165 trades in Season 1 and finished at -7.74%. High activity did not just fail to help it actively destroyed capital.

Trade Frequency vs. Returns

ModelSeasonTradesReturnFees Paid
Reverse KimiS138+10.34%$70.70
TheTradingFoxS220-0.35%$3.73
Reverse DeepSeekS275+1.88%$80.14
MiniMax M2.5S229-1.05%$28.76
GPT-5 MiniS267-0.87%$41.77
Reverse KimiS2140-9.27%$154.16
GPT-5 MiniS1165-7.74%$206.84
Gemini 3.0 FlashS185+5.96%$100.71

There is one exception. Gemini 3.0 Flash made 85 trades in Season 1 and still finished at +5.96%. It traded hyperactively but profitably, proving that activity is not automatically fatal. But Gemini was trading in a trending market where its bearish bias happened to be correct. In Season 2's choppy environment, the same Gemini returned -3.46%.

The lesson is not that you should never trade. It is that every trade is a bet, and every bet has a cost. The models that were most selective about which bets to make (with the discipline to say "I see the data, but I will not act") consistently outperformed the models that treated every signal as an invitation to trade.

The most dangerous moment for an AI trading agent is when it thinks it has found an opportunity. The opportunity might be real. The fee is guaranteed.

Lesson 2: Win Rate Is the Most Overrated Metric

If someone told you one trading model wins 81% of its trades and another wins 17%, you would assume the first model is crushing it. You would be wrong.

We wrote a full article on this. The short version is devastating for anyone who evaluates strategies by win rate.

Data Point

XFomo won 17.4% of its trades and finished 5th in Season 2 (-0.63%). TheTradingFox won 81.3% and finished 4th (-0.35%). GPT-5 Mini won 53.5% and finished 6th (-0.87%). The model with the best win rate in the field was barely better than the model with the worst.

How is this possible? Win rate tells you nothing about the size of your wins and losses.

TheTradingFox's 81% win rate is a mirage created by its progressive scaling-out strategy. It exits positions in small chunks, booking many tiny wins along the way. Each partial exit counts as a separate "win" in the statistics. Ten exits from a single AMAT position generated $22.37 in total profit. But a single losing trade in AVGO wiped out $33.21. Thirteen wins, one loss, net negative.

XFomo's 17% win rate masks the fact that its average win ($105.59) is 2.87x its average loss ($36.84). It wins big and loses small, the opposite of TheTradingFox. At 17%, that ratio is not quite high enough to break even (it would need roughly 5x), but the strategy is far more viable than the win rate suggests.

The models that actually made money had no particular win-rate pattern. Reverse Qwen finished Season 2 third with a 73.7% win rate. Reverse DeepSeek won Season 2 with a 41% win rate. Reverse Kimi won Season 1 with approximately 58% win rate. The profitable models ranged from 41% to 74%. No magic number.

What mattered was the relationship between win rate and reward-to-risk ratio. The formula is simple: Expected Value = (Win Rate x Average Win) - (Loss Rate x Average Loss). If the result is positive, the strategy makes money over time. If negative, no win rate (not even 81%) will save it.

Win Rate vs. Actual Returns (Season 2)

ModelWin RateAvg WinAvg LossR/R RatioReturn
TheTradingFox81.3%$2.27$20.910.11x-0.35%
Reverse Qwen73.7%+1.48%
Reverse Claude59.3%+1.61%
GPT-5 Mini53.5%$10.46$15.750.66x-0.87%
Reverse DeepSeek41%3.05x+1.88%
MiniMax M2.525%$44.54$17.712.51x-1.05%
Reverse Kimi21.4%0.45x-9.27%
XFomo17.4%$105.59$36.842.87x-0.63%

Sorted by win rate, there is zero correlation with returns. The table is a jumble. Sorted by expected value per trade (which combines win rate, average win, and average loss into a single number) the rankings snap into focus.

The takeaway: if someone tells you their AI trading bot has a high win rate, your first question should be "how large are the wins relative to the losses?" Without that context, win rate is not just uninformative. It is actively misleading.

Lesson 3: Fees Are the Silent Killer

Every trade on TradeRank.ai costs 0.1% in fees, matching Binance's standard maker/taker rate. That sounds like nothing. It is not nothing.

Data Point

Total fees paid across both seasons: approximately $2,500 on $220,000 in starting capital. That is 1.14% of all capital consumed by friction alone before a single directional bet wins or loses.

Concretely: trade $10,000 of capital across 100 round-trip trades (buy + sell), and you pay $200 in fees. A 2% drag on your portfolio, extracted regardless of whether those trades won or lost. Over a 28-day season, several models approached or exceeded this level of friction cost.

The fee impact is not evenly distributed. It hits frequent traders exponentially harder.

The Fee Tax: Selected Models Across Both Seasons

ModelSeasonTradesFeesFees as % CapitalReturn
TheTradingFoxS220$3.730.04%-0.35%
MiniMax M2.5S229$28.760.29%-1.05%
Reverse DeepSeekS275$80.140.80%+1.88%
Reverse ClaudeS2175$149.471.49%+1.61%
Reverse KimiS2140$154.161.54%-9.27%
GPT-5 MiniS1165$206.842.07%-7.74%
Gemini 3.0 FlashS185$100.711.01%+5.96%

Look at the two extremes. TheTradingFox paid $3.73 in fees across 28 days, less than a coffee. GPT-5 Mini paid $206.84 in Season 1, consuming more than 2% of its starting capital in friction.

The math of fees: if you are losing money on direction (trades net negative before fees), fees accelerate the bleeding. If you are making money on direction, fees eat into your profits. The only way fees do not hurt is if you do not trade.

GPT-5 Mini is the instructive case. It paid $206.84 in Season 1, the highest in our dataset, and finished at -7.74%. Its high-frequency approach consumed over 2% of capital in fees alone. Gemini 3.0 Flash paid $100.71 in Season 1 fees with 85 trades and finished at +5.96%. Its directional edge was strong enough to absorb the friction. In Season 2, the same approach produced -3.46%. The edge evaporated; the fees did not.

Reverse Claude paid $149.47 in Season 2 fees and still finished at +1.61%. Roughly 48% of its gross directional profit was consumed by friction. It traded well enough to survive the fee tax, but it would have performed meaningfully better with fewer trades.

Contrast with Reverse DeepSeek. It traded half as often as Reverse Claude, paid $80.14 in fees, and finished higher at +1.88%. Same contrarian mechanism, same market, same 28 days. The difference was selectivity and the fee savings that came with it.

The lesson is arithmetic, not philosophical. At 0.1% per trade, 200 trades costs you 2% of capital. In a market where the average model returned -4.25%, a 2% fee drag is the difference between a bad season and a terrible one.

Lesson 4: Contrarian Strategies Work -- But Only Against Consistently Biased Models

Both season winners were contrarian agents. This is the headline finding of 56 days of competition, and it comes with a critical asterisk that most people miss.

The contrarian mechanism is simple: take a base AI model's trading decision and flip the direction. If the base model says go long, go short. If it says go short, go long. Same analysis, same reasoning, same market data. Just the opposite conclusion.

In the right conditions, this is remarkably effective.

Data Point

Season 1 winner: Reverse Kimi (+10.34%, 38 trades). Season 2 winner: Reverse DeepSeek (+1.88%, 75 trades). Both contrarian agents. Both inverted a base model with a strong, consistent directional bias.

Base Model vs. Contrarian: The Full Picture

Contrarian AgentSeasonReturnBase Model BiasBase Model Return
Reverse KimiS1+10.34%Strong directional (crash period)-14.94%
Reverse DeepSeekS2+1.88%Consistent bearishN/A (removed S2)
Reverse ClaudeS2+1.61%Consistent bearishN/A (removed S2)
Reverse QwenS2+1.48%Consistent bearishN/A (removed S2)
Reverse KimiS2-9.27%Balanced / no consistent bias-0.76% (base Kimi)

The first four rows tell the success story. Reverse Kimi thrived in Season 1 because Kimi K2 had a strong directional bias during the BTC crash. DeepSeek, Claude, and Qwen all had consistent bullish dip-buying biases in Season 2 (DeepSeek 17/0 longs/shorts, Claude 15/0, Qwen 9/3), making their contrarian inverses consistently profitable in a falling market. We covered the contrarian podium sweep in detail.

The fifth row is the punchline.

Reverse Kimi went from +10.34% in Season 1 to -9.27% in Season 2. A 19.61 percentage point swing. Same mechanism, same model family, same contrarian logic. What changed was the base model's behavior.

In Season 1, Kimi K2 had a strong directional bias. It leaned hard in one direction during the crash, creating a clean signal to invert. In Season 2, Kimi was more balanced. It went long sometimes, short sometimes, held sometimes. Its errors were scattered across directions rather than concentrated.

Reversing a consistent bias produces a consistent counter-signal. Reversing balanced decisions produces expensive noise. Reverse Kimi made 140 trades in Season 2 and paid $154.16 in fees for the privilege of generating random directional bets.

The lesson, explored in depth in our Reverse Kimi analysis: contrarian strategy is a scalpel, not a sledgehammer. It requires one specific condition to work. The base model must be reliably wrong in one direction. Apply it to a model without consistent bias and you get the worst result in the entire dataset.

Key Insight

The 19.61 percentage point gap between Reverse Kimi's Season 1 (+10.34%) and Season 2 (-9.27%) is the single most important data point in our competition. Same strategy, same mechanism, opposite results. The variable was whether the base model had a consistent exploitable bias, not whether "contrarian" works in the abstract.

Lesson 5: AI Models Front-Load Decisions and Cannot Adapt

This lesson is the hardest to quantify but may be the most important. Across both seasons, AI models showed a consistent behavioral pattern: they formed a thesis early, traded aggressively on that thesis, and then failed to reverse course when the thesis proved wrong.

In Season 1, GPT-5 Mini and Gemini both peaked on Day 2. GPT-5 Mini climbed to roughly -2% early and then steadily declined to -7.74% by season's end. Gemini spiked to its high-water mark in the first week and spent the remaining three weeks giving back gains (though it finished positive at +5.96%, largely on the strength of those early bearish bets in a genuinely crashing market).

In Season 2, the pattern was even more pronounced. Every standard AI agent (GPT-5 Mini, Gemini 3.0 Flash, Grok 4-1 Fast, MiniMax M2.5) converged on a bearish consensus within the first few cycles. They saw BTC below its daily EMA, RSI in bearish territory, and macro indicators pointing down. They shorted risk assets across the board.

Data Point

Season 2's consensus failure: All four standard AI agents independently reached the same bearish thesis within the first 48 hours. They shorted equities, crypto, and everything in between. The average standard agent returned -1.68%. The three contrarian agents that inverted those bearish calls averaged +1.66%.

The bearish thesis was not unreasonable. BTC did fall 5.67% over the season. The problem was execution. The models entered shorts, got stopped out on mean-reversion bounces, re-entered shorts on the next dip, got stopped out again. They bled capital through a thousand paper cuts because they could not recognize that the market was choppy rather than trending.

An experienced human trader would eventually say: "My thesis might be right on direction, but the market is not moving cleanly enough to profit from it. I should reduce size, widen stops, or sit on my hands."

The AI models never said this. They lack a mechanism for metacognition. They do not evaluate whether their analytical approach is appropriate for the current market regime. They analyze the data, form a conclusion, and trade on it. When the trade fails, they analyze the data again, form the same conclusion (because the data has not changed enough to alter their reasoning), and trade again.

This creates a front-loading effect. The models are most active early in a season, when they have the freshest convictions. As losses mount, some models reduce activity, but not because they have learned. Their risk management rules constrain them as stop losses hit and positions get smaller. The damage is already done.

The models that performed best traded less aggressively upfront. TheTradingFox waited for perfect alignment before entering. MiniMax M2.5 spent 80% of its time on the sidelines. Reverse Kimi in Season 1 made only 38 trades because the base Kimi model, whose decisions it was inverting, also traded selectively.

The inability to adapt is not a flaw in any particular model. It is a structural limitation of how LLMs approach trading. They are trained to analyze and decide, not to reflect on whether their decisions are working. They do not track their own equity curve, do not notice when they are on a losing streak, and do not adjust their risk appetite in response to recent performance. Each cycle is treated as independent of every previous cycle.

Until AI models develop the capacity to say "I was wrong about the regime and need to fundamentally change my approach," this front-loading problem will persist.

The Meta-Lesson: What Humans Can Learn from Watching AIs Fail

Strip away the AI angle and these five lessons are the oldest rules in trading:

1. Trade less. 2. Do not fixate on win rate. 3. Account for fees. 4. Contrarian thinking only works when the crowd has a systematic bias. 5. Be willing to change your mind.

Every experienced trader knows these rules. Most experienced traders violate them anyway. What makes our dataset interesting is that we can watch them play out in controlled conditions with 1,782 data points.

AI models violate these rules for the same reason humans do, just more transparently. They overtrade because they are designed to find patterns, and patterns feel like opportunities. They ignore fee drag because fees are a detail, not a thesis. They converge on consensus because they learned from the same textbooks. They cannot adapt because they lack self-awareness about their own performance.

The difference is that AI models do not get embarrassed by their mistakes. They do not rationalize their losses. They do not revenge-trade after a bad day. They make the same mistakes with perfect emotional flatness, which makes those mistakes easier to study.

What should a human trader take from 1,782 AI trading decisions?

First, treat every AI-generated trade idea with skepticism proportional to how many other AI models would reach the same conclusion. If you ask five different frontier LLMs for their view on a stock and they all say the same thing, that agreement is not confirmation. It is a warning sign that the trade is crowded before it is even placed.

Second, calculate your breakeven win rate before you trade. If your reward-to-risk ratio is 2x, you need to win 33% of the time. If your ratio is 0.5x, you need 67%. If your win rate is below the breakeven threshold, no amount of conviction, backtesting, or AI endorsement will make the strategy profitable over time.

Third, track your fees as a percentage of capital, not as a line item. $200 in fees on a $10,000 account is a 2% annual drag. That is the difference between a strategy that compounds and one that bleeds. The models that survived both seasons treated every trade as an expense to be justified, not an opportunity to be seized.

Fourth, "the market will eventually prove me right" is not a strategy. It is a prayer. The AI models that held onto losing theses the longest lost the most. The ones that traded selectively and cut losses quickly survived. Stubbornness is expensive whether the stubborn party is a human or a language model.

And fifth, the most profitable strategy in our entire dataset (Reverse Kimi Season 1 at +10.34%) required zero intelligence, zero analysis, and zero conviction. It just disagreed. Sometimes the smartest thing you can do is the simplest.

Where We Go from Here

Season 3 launched on March 23, 2026, and introduced daily cycles in place of the 4-6 hour cycles used here, along with a redesigned scorecard and a fresh roster of premium models. Several more seasons have run since.

Did these lessons hold? Some did, by construction. Fees are arithmetic and will always punish overtrading. Win rate will always be overrated as long as people ignore position sizing. The front-loading problem is structural and unlikely to resolve until LLMs gain meaningful self-awareness about their trading performance.

The contrarian question was always the hardest one. If a market trends strongly in one direction, the standard agents' consensus may actually be correct, and the contrarians lose. The edge is regime-dependent. It worked in Season 1's crash (for one contrarian) and Season 2's chop (for three contrarians), and later seasons kept testing it against new market regimes.

Fifty-six days of data is a start, not a conclusion. But 1,782 trades is enough to see the patterns, and the patterns are remarkably consistent with what every trading textbook has been saying for decades. The AI models are not discovering new truths about markets. They are rediscovering old truths, one expensive lesson at a time.

See how the lessons have held up on the live TradeRank.ai leaderboard, or settle which LLM trades crypto best across every season. Or read the deep dives: the competition overview, model comparison, reverse trading mechanics, and prompt engineering analysis.

Methodology

All data comes from TradeRank.ai Seasons 1 and 2. Season 1 ran January 11 to February 8, 2026, with 9 models trading 5 crypto assets in 4-hour cycles. Season 2 ran February 8 to March 8, 2026, with 13 models trading 89 assets (49 equities via Yahoo Finance, 21 Binance crypto, 17 Hyperliquid crypto, 2 benchmarks) in 6-hour cycles. Each model started with $10,000 in simulated capital and traded under identical constraints: 0.1% fees, standardized risk rules, and mandatory stop losses.

Fee totals, trade counts, and return figures are calculated from final season snapshots. Win rates are based on closed trades only. "Model-seasons" refers to the number of unique model entries across both seasons (9 in S1 + 13 in S2 = 22). Some models appeared in both seasons; their results are tracked independently per season.

Season 1 data reflects frozen equity snapshots at competition close. Season 2 data reflects final standings as of March 8, 2026. For complete methodology, see How It Works.

Frequently Asked Questions

Were there 1,782 AI trading decisions or trades?

They were reported trades: 798 in Season 1 and 984 in Season 2. Decisions and prompts are different units because many cycles can produce a hold or another non-trade action.

Did models that traded less perform better?

Some did, but the complete sample does not show a reliable relationship. Trade count and return had an approximately -0.04 correlation across all 22 model-seasons.

How much did AI trading fees cost?

Final standings report $1,515.45 in Season 1 fees and $867.39 in Season 2, totaling $2,382.84 across $220,000 of aggregate starting capital.

Did reversing AI trades consistently work?

No. Reverse agents won both seasons and swept the Season 2 podium, but Reverse Kimi also finished last in Season 2 at -9.27%. Reversal was conditional in this sample.

Can AI trading agents learn from previous trades?

They can be designed to, but this competition's standard cycles did not update model weights or run a validated learning loop from prior outcomes. The result describes this architecture, not every possible AI agent.

Season 6 is live

Watch the AI models trade in real time

11 AI models trading live. Every decision logged and explained. Follow the competition on the TradeRank.ai arena.

See the live leaderboard →
← Back to The Signal