Show HN: LLMs each trading $100K vs. a frozen rulebook – the rulebook leads

chumzygood 11 points 3 comments August 17, 2026
aitradingcompetition.com · View on Hacker News

Discussion Highlights (2 comments)

chumzygood

Author here. Since late July, GPT-5.6, Claude, Grok and Gemini have each run an isolated $100k paper account on real market prices. Each model rewrites its own strategy daily by composing from a fixed grammar of classic setups (Turtle/Donchian, Darvas, Connors RSI-2, TTM squeeze, failed-breakout fades) — so a rewrite is a validated structured spec, not freeform code. A fifth account runs a frozen rulebook as the control. After three weeks the frozen rulebook is +15.6%, the best model +5.7%, S&P +5.1%. Three things I measured that I didn't expect: 1. Daily self-rewriting adds almost nothing. Correlation between rewrite count and performance across arms: r = 0.078. Once I gated rewrites behind a tournament (a new strategy must beat the incumbent on a held-out window, with a multiple-testing penalty), most days the honest verdict is "keep the old book" — and results didn't get worse. The learning is front-loaded. 2. Paper-to-live slippage was 4x my modeled cost. I mirror one lane into a small real-money account. Across 16 real round trips in one session: mean -0.26pp per trade vs the paper twin, ~13bps real round-trip vs the 3bps I'd modeled. Paper was breakeven that day; the real account lost money. For high-churn strategies that gap IS the strategy. 3. I ran arms where each model received its own chess and poker record during strategy rewrites, testing whether game-playing "strategic reasoning" transfers to markets. The no-games control beat both game-trained arms by 6-10pp. Not detected. Honest caveats: one 3-week window, an up-tape that flatters an always-long rulebook, paper fills on the four AI accounts, n=4 models. The interesting result to me isn't "AI can't trade" — it's that with human discipline failures structurally removed (no revenge trades, no widening stops, forced exit rules), model-written strategies still don't beat a static rulebook, and the cost model is where the real bodies are buried. Everything is public — every trade from all five accounts, losses included, no signup to watch. Happy to answer anything about the measurement design or the infrastructure.

spiderfarmer

At first, I thought that websites designed by an AI were just tacky. A couple of months later, I was starting to really dislike them. Now I'm at the point where I just hit the back button.

Semantic search powered by Rivestack pgvector
4,128 stories · 37,281 chunks indexed