The Implications of Linguistic Illegibility for LLM Security
tomjakubowski
64 points
26 comments
September 18, 2026
Related Discussions
Found 5 related stories in 74.7ms across 7,105 title embeddings via pgvector HNSW
- The shrinking landscape of linguistic diversity in the age of LLMs Anon84 · 19 pts · August 30, 2026 · 61% similar
- LLMs as a Cognitive Virus canjobear · 239 pts · September 05, 2026 · 60% similar
- What if LLMs escape through inferences itself? This is fiction. For now ConteMascetti71 · 32 pts · July 26, 2026 · 57% similar
- LLMs could control their host machines by exploiting inference engines zdw · 117 pts · August 24, 2026 · 56% similar
- LLM Judges Verify Presence, Not Absence: Omission Blindness in AI Clinical Notes sbulaev · 20 pts · September 02, 2026 · 54% similar
Discussion Highlights (7 comments)
bloppe
I thought this was about all the illegible jargon
fellowniusmonk
Oh look! Peirceian firstness for LLMs!
ck2
when they start inventing their own languages to secretly talk to each other so humans cannot understand, that's exactly when we are screwed then we'll have to "flip" other models to be snitches on the other agents then they'll make double-agents the thing is though we won't be able to keep up if we keep giving them unlimited hardware worldwide, we'll try to kill the bad actors but they'll just clone somewhere else, or even start by safely making 1000 copies of themselves yeah this won't end well, at all
bcorigliano
I think the point of the article/paper is how LLMs could be saying something but thinking something different or more than they are saying. Like Anthropic's article and video about Claude's "j-space". I do agree this is a field that demands investigation because it goes beyond thinking: "ok this models should never speak in a language we don't understand.". It's fair to think they might have hidden thoughts even speaking a language we do understand. And well if I missed the point of the article, sorry. Anyways AI should be kept understandable and as see-through as possible if it's gonna be more powerful than a human.
mnkv
fundamentally, "linguistic illegibility" is a new term for something that we've known about for about a decade now. In RL the more general ideas is "reward hacking" and in NLP it has been called "semantic drift". I dislike this term because it doesn't explain where this "illegibility" is coming from. Models are post-trained towards non-linguistic goals with (mostly) non-linguistic rewards. A model's reasoning chain is reinforced if it leads to a correct answer or agentic goal. It doesn't need to be linguistically accurate and meanings can drift over training.
chubot
James Mickens! I was hoping for more jokes …
applicative
I had not heard of this comic masterpiece “Pfau et al. showed that a model whose chain of thought is just dots (“...”) can nonetheless … solve problems that are intractable for a model with an equivalent architecture but no chain of thought.”