Notes on gotchas while migrating 35kb preprompts from Opus to self-hosted Ollama
0o_MrPatrick_o0
129 points
69 comments
September 14, 2026
Related Discussions
Found 5 related stories in 77.7ms across 6,607 title embeddings via pgvector HNSW
- Claude: System Prompts tosh · 607 pts · August 16, 2026 · 52% similar
- Ask HN: Advice on Migrating from 1Password? 0xbadcafebee · 102 pts · September 03, 2026 · 48% similar
- Analysis of Opus 5's 'Contrition' slowmovintarget · 17 pts · July 28, 2026 · 47% similar
- I obtained Claude Opus 5 system prompt juleiie · 22 pts · July 30, 2026 · 45% similar
- Lotus Notes and the dangers of starting from scratch maguay · 161 pts · September 09, 2026 · 45% similar
Discussion Highlights (12 comments)
SyneRyder
TLDR: Local models have a smaller context window, so your 35kB prompts that worked fine against a hosted 1 Million token window, crash out when you only have a 65K (!) token window locally. I dislike being negative, but I was really hoping for more substance when reading this. It would have been an interesting topic.
dell2024
I had hoped to get some new information out of this topic, but unfortunately found the same local "dead-ends" that I explored myself. It unfortunately feels like we will be stuck waiting for a burst bubble before local hardware can be reasonably acquired for personal LLM usage.
DiabloD3
The article doesn't really describe the problem: if your prompt is 35kb, your prompt is confusing, unfocused, and doesn't work right on any LLM, and is needlessly bloating your context. At this point in time, due to how most people and companies run their inference engine, regardless of the model (yes, this includes the newest from OpenAI and Anthropic and the Chinese Tigers and Dragons), you run out of useful context that the model can accurately attend to around the 250k mark no matter how much they advertise their context size is . You need to cut your prompt up. If you believe LLMs work, have the LLM help you shape the overall plan, and then have multiple sessions run each step in the plan without being bloated with the context of previous successful steps. I don't see LLMs being production-ready until the context rot and sampling problem is fixed forever. This has not occurred, and the big inference providers aren't even bothering to integrate any of the research on that subject. If anything, many of the bigger companies are actively making inference quality worse just to extend their runway a tiny bit farther before they go bankrupt. The only thing the article gets right is this: if you're serious about LLMs, abandon Big AI and infer locally only. This is the only way you have control over the quality of the output.
andai
> Everyone who begins learning exploitation hits a phase of exploitability grief about 3 month into dedicated, practiced study. They hack something they didn’t think they had the skill to break into and it terrifies them. They’re smart enough to know that, relatively speaking, they are an idiot, and if an idiot can do this then nothing is safe. That feeling is correct.
stackedinserter
The main gotcha for local models is insane hardware requirements. Even for $10K you get mediocre performance.
cube00
Friends Don't Let Friends Use Ollama https://news.ycombinator.com/item?id=47788385
robotswantdata
Why are you using Ollama? Just use llama.cpp
roschdal
Self-hosted Ollama is the best.
fghorow
I've been using Claude Code Extension in VSCode (no phone-home configured), backed by DwarfStar on a LAN local MBPro 128GB M5. The context bloat is horrendous, leading to 5-10 minute prefills. I've recently been exploring tools like headroom to help manage context, with some limited "success" (for some definition of success). What do others with similar setups do? (I kind of hate to abandon Claude Code, as it seems to be the most capable coding assistant of the limited set of tools I've tried. But that horrendous context bloat is really painful!)
krttherealest
attention is the key
Havoc
Can a 27b model even do meaningful security tasks? I thought the interesting cyber stuff is really at the edge of frontier
onesandofgrain
What model are you using?