Show HN: OpenAPPA – open-source deterministic guardrails that don't break agents
Hi Hacker News! Matvey, one of the authors, is here. While building enterprise agents, we ran into a problem: the more tools you connect to the AI, the higher the chance it will run out of control and leak sensitive data. Guardrails, in theory, should prevent this, but the situation is worrying: - Non-deterministic guardrails (LLM as a judge, auto modes, etc.) are vulnerable to prompt injections, or they lack knowledge of the data, making them inefficient (~10% data leaks on our benchmarks). - Existing deterministic guardrails (Cedar, OPA, FIDES, Dogwood) require massive case-specific IF-ELSE-like policies and break agents (~59% utility loss on our benchmarks). We did something differently. We’ve taken the best of existing deterministic guardrails and built a policy language that is data-specific, not use-case specific. It lets you scale agents without updating a policy. On top of that, we’ve added multiple tricks (like a remedy plan or a DualLLM pattern) to help agents operate within those restrictions, raising utility from ~40% to ~90% and making it the first deterministic guardrail that doesn't break agents. Finally, we’ve designed it to be pluggable into any agent loop with pre- and post-tool-call hooks. We invite you to check out our benchmarks: https://www.openappa.com/evaluation Play with it in Claude Code: https://www.openappa.com/claude-code Try plugging it into your agent: https://www.openappa.com/add-to-agent Or check the academic paper: https://arxiv.org/abs/2607.24625 We'd love to hear any feedback!
Discussion Highlights (11 comments)
dvorkanton
Finally some determinism in our high-temperature sampling world!
joeyorlando
aussi, si vous êtes à Montréal et vous aimerez apprendrez plus sur OpenAPPA, viens à l'evenement CNCF demain soir, le 29 - je ferai un p'tit discours sur l'integration kagent d'OpenAPPA. https://www.meetup.com/kubernetes-montreal/events/316689391/
arseny_info
I still remember the times when ai/ml security was about perturbing pixel gradients to misclassify a panda
keshakon
Hi! One of the OpenAPPA authors here. Ask me anything! My favorite part of APPA is “batteries”: you can run arbitrary programs as part of an authorization decision. For example, a battery could call the GitHub API to check whether a repository is public or private, then use that result to decide whether its contents can be posted to Slack.
ildari
a few days ago I started an agent on gpt-5.6-terra to work on a project, and one of the website pages had a sentence to create GH issues. Agent read it and that was enough to derail and go creating issues with my context
piercypixel
Guardrails with builtin remediation instead of simply blocking my agent is a mind blowing long awaited experience! Sooo good. Can't recommend more!
immafridge
Quick disclaimer, I work at Archestra. I’ve had the chance to play with OpenAppa for a bit and if there’s one thing that I love with this project: it’s simple to get started with and easy to tweak. imo agentic security shouldn’t have to be painful to setup. Give it a shot and hopefully ya’ll will find this project useful. It's also open source :)
zborro
i suspect we’ll see more of this: flexible agents but deterministic boundaries. Congrats on launch!
vladimir_gor
Really interesting direction. What resonated with me is that you're treating agent security as an information-flow problem rather than a prompt-classification problem. It was not so obvious to me. A key question I agree isn't just "is this tool call allowed?", but "given everything the agent has read so far, is this information now allowed to flow to this destination?" That feels like a much more fundamental abstraction. The part I'm particularly curious about is how this will work with policy authoring at scale. What would be the main adoption challenge?
apetrovicheva
the paper is good! thorough. I like it.
moneytool
HI I have created something that actually to solve this exact problem please go through it this stays in your environment independent of AI and its a python module that blocks AI from using unauthorized commands https://github.com/moneytool/aegis-devops