Ask a model if code is malicious and it reaches for its morals
codyznash
15 points
3 comments
October 06, 2026
Related Discussions
Found 5 related stories in 86.6ms across 8,687 title embeddings via pgvector HNSW
- Every Model Cheats vga805 · 91 pts · August 20, 2026 · 59% similar
- A warning about 'model welfare' andsoitis · 211 pts · September 16, 2026 · 54% similar
- Models Don't Go Rogue cdrnsf · 17 pts · September 03, 2026 · 50% similar
- AI models keep posting screenshots showing sensitive data from inside companies Dotnaught · 20 pts · September 29, 2026 · 49% similar
- Models Are Getting Dumber on Purpose hruvhwe · 300 pts · August 16, 2026 · 49% similar
Discussion Highlights (2 comments)
prasadvara
This is great writeup, can we do a cross comparison with "human" experts?? whether models perform better OR worse??
cortesoft
I feel like LLMs really highlight the ambiguity and imprecision of the english language in normal use. LLMs are getting really good at guessing what we mean, but it is still a guess.