FLAWED's Flaws and What This Means for Industry Research
tob_scott_a
19 points
4 comments
September 24, 2026
Related Discussions
Found 5 related stories in 84.0ms across 7,510 title embeddings via pgvector HNSW
- Why Software Factories Fail (or: harness engineering is not enough) dhorthy · 247 pts · July 23, 2026 · 50% similar
- AI Is Breaking This Thing We Call Trust matheusml · 76 pts · September 10, 2026 · 47% similar
- Investigating three real-world incidents in our cybersecurity evaluations surprisetalk · 154 pts · July 30, 2026 · 47% similar
- When Genius Fails: The Intellectual Arrogance of the AI Labs gmays · 171 pts · August 14, 2026 · 46% similar
- AI-found bugs aren't proving any easier to exploit despite the hype Tomte · 14 pts · July 29, 2026 · 46% similar
Discussion Highlights (4 comments)
Qwuke
The "peer review" comment in the original seems particularly bizarre given the errors https://en.wikipedia.org/wiki/Scholarly_peer_review
tolugenius
> If you believe these issues should not be raised because FLAWED was critical of OpenAI, please know that research is not sports. OpenAI is not the Spurs, 1Password is not the Knicks, and Off-by-1 Labs is not Jalen Brunson. Got a hearty laugh out of this.
unprovable
> A paper is not a blog post True words. This post is actually a very good enumeration of the status quo for this debacle.
darkamaul
A timeline to help understand what happened : - Aug 6: Off By 1, the security lab of 1Password publishes their blog post (current title: Why AI-generated vulnerability patches still require expert human review) This article gets some specialized and mainstream press coverage - September 15 : Trail of Bits publishes a rebuttal on their blog 1Password's AI patching benchmark is misleading - September 22: Suha publishes the current blog post The main idea in the rebuttal is that the headline number (26 percent only of the AI generated patches are flawless) is misleading because the experiments were not well designed.