Show HN: Open-source playground to red-team AI agents against public prompts
zachdotai
13 points
4 comments
August 09, 2026
Related Discussions
Found 5 related stories in 152.9ms across 8,795 title embeddings via pgvector HNSW
- Show HN: Benchmark your eng team's AI agent maturity in 5 minutes adamgold7 · 13 pts · July 14, 2026 · 71% similar
- Show HN: Clawfight.ai MCP-driven agentic game play wesleyhales · 13 pts · September 11, 2026 · 69% similar
- Show HN: TinyAIArena watch AI agents battle it out hp6 · 105 pts · September 27, 2026 · 68% similar
- Show HN: I built a web tool to see and edit what an AI thinks before it answers ada1981 · 25 pts · July 09, 2026 · 66% similar
- Show HN: Open-source simulation testing infra for voice agents Nischalj10 · 14 pts · September 10, 2026 · 65% similar
Discussion Highlights (3 comments)
yousefh409
Why would any company leave the enforcement of rules to an agent? If something is truly a rule, there should be code that deterministically enforces it.
mmvaid
Interesting. Is there any way companies or individuals can submit an agent before they release it and use this as a way to pentest their agent?
nourzahzah
Does Nyx solve these before you publish them? Curious whether humans are still finding breaks your own agent misses, or just the same ones slower.