LLM Judges Verify Presence, Not Absence: Omission Blindness in AI Clinical Notes

sbulaev 20 points 11 comments September 02, 2026
arxiv.org · View on Hacker News

Discussion Highlights (4 comments)

Planktonne

The fact that even the abstract is very clearly AI-generated does not fill me with confidence in the rigour of the research.

udik2

May be AI written. Used my AI to get the gist of it. One thing that I believe is LLMs need to be considered a pure play tool and can be guided to point out misses as well. The reason is simple. To find whats missing, one needs to ask the right questions on how to judge and when that is there.

velrim

Not sure about the paper but the results make sense, we see it in PDF extraction too. Fields that aren't in the document are being made up 11 to 40% of the time depending on the API. The whole thing is nasty partly because it isn't just AI problem. When checking the human labels that we used in our evals 40 out of 142 answer keys claiming absence were wrong. Tricky one.

tarboreus

"the floored note before its clean twin" People using AI like this should be run out of their jobs.

Semantic search powered by Rivestack pgvector
5,346 stories · 48,358 chunks indexed