Ask a model if code is malicious and it reaches for its morals

codyznash 15 points 3 comments October 06, 2026
www.manifold.security · View on Hacker News

Discussion Highlights (2 comments)

prasadvara

This is great writeup, can we do a cross comparison with "human" experts?? whether models perform better OR worse??

cortesoft

I feel like LLMs really highlight the ambiguity and imprecision of the english language in normal use. LLMs are getting really good at guessing what we mean, but it is still a guess.

Semantic search powered by Rivestack pgvector
8,687 stories · 81,484 chunks indexed