FLAWED's Flaws and What This Means for Industry Research

tob_scott_a 19 points 4 comments September 24, 2026
suhacker.ai · View on Hacker News

Discussion Highlights (4 comments)

Qwuke

The "peer review" comment in the original seems particularly bizarre given the errors https://en.wikipedia.org/wiki/Scholarly_peer_review

tolugenius

> If you believe these issues should not be raised because FLAWED was critical of OpenAI, please know that research is not sports. OpenAI is not the Spurs, 1Password is not the Knicks, and Off-by-1 Labs is not Jalen Brunson. Got a hearty laugh out of this.

unprovable

> A paper is not a blog post True words. This post is actually a very good enumeration of the status quo for this debacle.

darkamaul

A timeline to help understand what happened : - Aug 6: Off By 1, the security lab of 1Password publishes their blog post (current title: Why AI-generated vulnerability patches still require expert human review) This article gets some specialized and mainstream press coverage - September 15 : Trail of Bits publishes a rebuttal on their blog 1Password's AI patching benchmark is misleading - September 22: Suha publishes the current blog post The main idea in the rebuttal is that the headline number (26 percent only of the AI generated patches are flawless) is misleading because the experiments were not well designed.

Semantic search powered by Rivestack pgvector
7,510 stories · 69,432 chunks indexed