Show HN: Raven – The harness of harnesses, built for RSI

cyfyifanchen 54 points 50 comments September 29, 2026
github.com · View on Hacker News

Discussion Highlights (17 comments)

cyfyifanchen

We're working on Raven: the harness of harnesses, built for RSI (recursive self-improvement). Raven brings Claude Code, Codex, and its own Research, Code, Design, and Oncall agents into a shared task graph. The idea is to let different agents handle the parts of a project they are suited to, with shared memory across subagents and context carried across sessions. The RSI work extends to the harness itself: prompts, policies, strategy code, and playbooks. Raven's specialist harnesses and orchestration layer can be improved independently. Candidate changes are evaluated before adoption. This concerns Raven's own components; it doesn't rewrite Claude Code or Codex internals. We've used Raven for long-running research and experimentation workflows and for building a Godot game. The repository includes examples and outputs, along with installation instructions. We're also exploring how to develop and refine specialist agents for particular domains. Raven is pre-alpha and Apache-2.0 licensed. The self-improvement work is experimental; the Curator currently ships in the repository rather than the installed package. Code and examples: https://github.com/EverMind-AI/Raven Where do you find coordination between agents breaks down today? We'd also be interested in what evidence you'd want before trusting an agent-generated change to its own harness.

ssddanbrown

RSI here is for "recursive self-improvement", instead of a harness being built to help users with repetitive strain injury like I first thought when reading the post title

PcChip

I wish they had also compared it to omp and dsh, I’m curious how it stacks up

aatd86

Interesting. Only thing, from someone who has built something similar, is that it moght tend to duplicate certain capabilities that those harnesses handle on their own. Some overlap is bound to happen.

arminluschin

This looks similar to https://paseo.sh/ , if I understand correctly. I’ve recently tried it and liked it a lot. Would be nice to see a comparison. When’s the harness of harnesses of harnesses coming?

mpalmer

Agents editing their prompts and writing memory to markdown files is not and will never be RSI.

revexos

Recursion has started. Here we go!

jeffnash

Reminds me a lot of omnigent (which I am a huge fan of) with a persistent memory layer. Unlike omnigent's subagent threads, the DAG it uses to coordinate other harnesses doesn't look to be durable; I am curious as to whether this is by design or is a forthcoming feature, as this essentially makes or breaks my use case of long-running project-sized implementation sessions. In any event, it's great to see competition in this meta-harness space, which is likely one that none of the frontier labs will touch since it, by definition, would utilize their competitors' products.

mawadev

Why do none of the projects it built work on my machine? Even the fps game looks cool but can't be downloaded? :0

Kuyawa

That's one of the most beautiful readmes I've seen in my whole life. Threshold is stunning. I am so impressed even RSI got overshadowed

julesrms

It's all very glitzy, but I'm failing to understand how "One prompt in. One result out." is of any importance. This and other recent AI hype-fests all seem to be obsessed with making agents do more work unattended. But surely in the real world, anyone who's got a real product to make is going to want to steer what's happening. It's ridiculous to think that anyone with a deadline would write a prompt so perfect that they walk away for 4 days and come back to find the finished product ready to ship. If you really can write a prompt so complete and perfect that it needs nothing further, then any regular harness could probably also do the job. But if like normal people you need to try something, think about it, iterate, and repeat.. then you also just need a regular harness.

calebhwin

Dumb but honest question - do repos like this buy stars? How do they have thousands of stars with very little presence across HN/Reddit/X?

fraywing

It feels like the only thing left people are building are harnesses? Harnesses of harnesses?

riskable

I always figured that the free/open weights models like qwen3.8:27b would perform just as well if not better than Claude's latest if you just fed it back into itself enough times. This project seems to prove that this is indeed the case. What I'd like to see now is how good it can get when you feed the micro models like qwen3.5:0.8b into itself to solve problems. Will it be like toddlers discussing neighborhood politics at a pretend tea party or will it actually get some decent results? Another game-changer (if this style works out): Just get a model like qwen3.8:27b onto one of those model-on-a-chip cards that makes it 1000x faster and see how fast it can go using the same method.

joshstrange

I continue to yearn for a harness of harnesses but each one I try (or build) takes me uncomfortably far from the work being done. I don't want to be a prompt shuttle, though I feel that way sometimes. Performing the same dance for each ticket I work on. My issue is that, to bastardize a common joke/phrase, 50% of the things the agent stops for are things it (or another agent) could answer for me, but it's a different 50% task to task. With HoH's I constantly feel like I'm getting peppered with unimportant questions or being kept out of the loop of things that really need my eyes on it. Threading that needle has been particularly difficult.

redhale

I don't want to be too negative, but ... all this for a 0.8% improvement in SWE-bench Verified (90.2 for OpenCode vs 91)? And why is this (saturated) benchmark the one coding benchmark chosen to showcase on the homepage? Without trying it, this seems like its probably just a massive waste of tokens.

topheroo

“Built for RSI” sounds like the kind of meaningless thing an AI would throw into a tagline to garner clicks.

Semantic search powered by Rivestack pgvector
8,041 stories · 75,100 chunks indexed